class: title-slide <img class="title-australia-map" src="assets/australia_pixel_outline_transparent.png" alt="Pixel outline map of Australia" style="position:absolute;right:5%;top:25%;width:29%;height:auto;z-index:2;filter:brightness(0) invert(1);opacity:0.85;pointer-events:none;"> <div class="title-content"> <h1>Towards an Australian<br>Parliamentary Speech Corpus</h1> <p>Steven Coats<br>University of Oulu, Finland<br> <a href="mailto:steven.coats@oulu.fi">steven.coats@oulu.fi</a><br> CLARIN Annual Conference, Brighton<br> September 30th, 2026</p> </div> --- layout: true <div class="my-header"><img border="0" alt="Oulu logo" src="https://cc.oulu.fi/~scoats/oululogonewEng.png" width="80" height="80"></div> <div class="my-footer"><span>Steven Coats · Australian Parliamentary Speech | CLARIN 2026</span></div> --- ### Overview 1. **Motivation**: Spoken reality vs. the Hansard record 2. **Background and Data**: Parliamentary corpora and Australian sources 3. **Proof of Concept**: Speech alignment and metadata enrichment 4. **Preliminary extensions:** Acoustic speaker matching and Hansard-supported name correction 5. **Outlook:** Verbatim speech and AU ParlaSpeech .footnote[Slides: [cc.oulu.fi/~scoats](https://cc.oulu.fi/~scoats)] --- ### The official record: Hansard <div style="display: flex; gap: 36px; align-items: center;"> <div style="flex: 0 0 190px; text-align: center;"> <a href="https://commons.wikimedia.org/wiki/File:Hansard-1832.jpg" target="_blank" title="View on Wikimedia Commons"> <img src="assets/Hansard-1832.jpg" style="max-width: 100%; height: 345px; object-fit: contain; border-radius: 3px; box-shadow: 2px 2px 8px rgba(0,0,0,0.15);" alt="Hansard 1832 title page"> </a> <p class="caption" style="font-size: 12px; color: #777; margin-top: 6px; line-height: 1.3;">Hansard (1832)<br><a href="https://commons.wikimedia.org/wiki/File:Hansard-1832.jpg" target="_blank" style="color: #888; font-size: 11px; text-decoration: underline;">Wikimedia Commons / PD</a></p> </div> <div style="flex: 1;"> <ul> <li><strong>Commonwealth tradition:</strong> Named after English publisher Thomas Curson Hansard; standard in the UK, Australia, New Zealand, Canada, Singapore, and elsewhere</li> <br> <li><strong>"Substantially verbatim":</strong> Official transcripts omit repetitions, false starts, and redundancies, while regularizing grammar</li> <br> <li><strong>Linguistic limitation:</strong> Not ideal for some linguistic research questions; can directly alter the <strong>interpretation of discourse</strong> (illustrated in the next examples)</li> </ul> </div> </div> --- class: example-slide exclude: true ### From speech to the official record: Self-repair <div class="example-grid"> <div> <video controls playsinline preload="metadata" src="data:image/png;base64,#../presentation_examples/01_self_repair_trimmed.mp4"></video> <p class="caption">Sam Birrell · House of Representatives<br>30 March 2026 · 7 seconds</p> </div> <div> <h4>ASR</h4> <p>“… Dad ordered a load of diesel <strong>last week, sorry, three weeks ago</strong>, and it hasn’t arrived yet …”</p> <h4>Hansard</h4> <p>“… Dad ordered a load of diesel <strong>three weeks ago</strong> and it hasn’t arrived yet …”</p> </div> </div> <p style="font-size: 20px; margin-top: 18px; color: #333;"><em>The corrected information survives in Hansard; the process of self-correction does not.</em></p> --- class: example-slide ### From speech to the official record: Grammar <div class="example-grid"> <div> <video controls playsinline preload="metadata" src="data:image/png;base64,#../presentation_examples/03_grammatical_regularization.mp4"></video> <p class="caption">Llew O’Brien · House of Representatives<br>17 August 2026 · 12 seconds</p> </div> <div> <h4>ASR</h4> <p>“Indeed, if every international obligation Australia had ever undertaken <strong>was</strong> swept away and only one could remain, for me it would be the one that protects the fundamental human rights of our children.”</p> <h4>Hansard</h4> <p>“Indeed, if every international obligation Australia had ever undertaken <strong>were</strong> swept away …”</p> </div> </div> - Evidence consistent with the ongoing shift away from subjunctive “were” in Australian English (cf. Filppula, 2022; Vaughan & Mulder, 2014) --- class: example-slide ### From speech to the official record: Managing participation <div class="example-grid"> <div> <video controls playsinline preload="metadata" src="data:image/png;base64,#../presentation_examples/02_turn_management.mp4"></video> <p class="caption">The Speaker · House of Representatives<br>17 August 2026 · 18 seconds</p> </div> <div> <h4>ASR</h4> <p>“She’s done that. So we’re going to… <strong>No, resume your seat. Resume your seat.</strong> We’re going to deal with this properly and efficiently and effectively.”</p> <h4>Hansard</h4> <p>“She’s done that. We’re going to deal with this properly, efficiently and effectively.”</p> </div> </div> - Hansard omits the directive, the repetition, and the interruption of the Speaker’s own formulation --- ### Why preserve these features? **Linguistic authenticity:** - Parliamentary corpora are records of actual human speech, not only repositories of political decisions - Spoken grammar (e.g. *was* vs. subjunctive *were*) captures ongoing variation and language change which editorial regularization may erase **Turn interaction and decision-making:** - Interactional details (directives, interruptions, floor management) reveal how political differences and agreements emerge in real time and are essential for understanding how legislative debate actually unfolds --- ### Background: Parliamentary corpora and AU Hansard .pull-left[ #### CLARIN infrastructure - **ParlaMint:** European parliamentary transcripts with metadata <span class="small">(Erjavec et al., 2023)</span> and **ParlaSpeech:** Aligned audio for some Slavic languages <span class="small">(Ljubešić et al., 2025)</span> #### Australian Hansard research - ***Australian Diachronic Hansard Corpus:*** Aggregated Hansard transcripts 1940s–2010s, used to study (e.g.) colloquialization and style shifts <span class="small">(Collins & Smith, 2021; Kruger & Smith, 2018)</span> - Comparison of a small corpus of manually transcribed parliament recordings with the official Hansard records <span class="small">(Kotze et al., 2023)</span> - [RAPID](https://research.qut.edu.au/centre-for-justice/projects/rapid-reusable-and-accessible-public-interest-documents/) (Reusable and Accessible Public Interest Documents) project; Metapragmatics of "unparliamentary language" over a century <span class="small">(Hames et al., 2025)</span> ] .pull-right[ #### Goals of AU ParlaSpeech pilot - **Verbatim speech:** Retain non-standard usages and turn structures/content, also self-repairs, interjections, hesitations - **Synchronized audio & alignment:** Word-level timing from WhisperX - **Speaker identities & metadata:** Speaker attribution from Hansard; MP information and electoral demographics - cf. Asonitis et al. (2026): pairs AU/NZ audio to Hansard for ASR training ] --- ### Data sources .small[ .pull-left[ **Parliament of Australia — ParlView (https://aph.gov.au)**: - House sitting audio, delivered through HLS media streams - Sitting dates and ParlView recording IDs - Hansard XML files: Speaker names and roles, party and electorate metadata, and speech-turn structure - Archived WebVTT live captions **OpenAustralia**: - Mirror of Parliament's Hansard XML ] .pull-right[ **Wikidata**: - Supplementary speaker metadata (date of birth, gender, education) **Digital Atlas of Australia**: - Constituency latitude-longitude coordinates **Australian Bureau of Statistics**: - Constituency median weekly household income ]] --- class: figure-slide ### The proof-of-concept pipeline <img class="pipeline" src="assets/paper_pipeline.svg" alt="ParlView download to ASR to word alignment; Hansard XML feeds lexical alignment; contextual rules and limited LLM review precede metadata enrichment and dataset export."> --- ### ASR and Hansard transcript matching 1. Download the recording, transcribe with faster-whisper large-v3 <span class="small">(Radford et al., 2023)</span> 2. WhisperX <span class="small">(Bain et al., 2023)</span> supplies word-level timing 3. Convert the ASR and Hansard texts into normalized word streams 4. Use <strong><code>difflib</code></strong> to identify matching stretches 5. Transfer Hansard speaker labels to aligned ASR segments using matched content words 6. Use fuzzy turn matching to assign ASR segments not resolved by the global alignment 7. Refine the assignments with contextual and parliamentary procedural rules <p style="font-size: 17px; margin-top: 10px; color: #666;">Proof of concept: a final Gemini 2.5 Flash pass reviewed assignments in local Hansard context and could split mixed-speaker segments.</p> --- ### Metadata enrichment #### Linked information - **Hansard:** speaker, party, electorate, parliamentary role - **Wikidata API:** date of birth, gender, institutions attended - **Digital Atlas of Australia:** constituency boundaries and calculated centroids - **Wikidata SPARQL → ABS:** electorate identifiers and census income <p style="clear: both; font-size: 20px; padding-top: 10px;"><strong>Possible extensions:</strong> age at recording, birthplace, qualifications, previous occupation, parliamentary experience; richer electorate demographics.</p> --- exclude: true ### Why not simply use ParlView captions? - An additional, valuable transcript source—but not automatically a gold standard. - Rolling caption displays need reconstruction to avoid counting repeated display text as repeated speech. - Captions still require checking for lexical errors and integration with speaker metadata and word timing. - Agreement with Hansard alone cannot establish acoustic accuracy: Hansard edits and omits spoken material. - **Guiding principle:** Keep the recording as the reference point; compare transcript layers according to the research question. --- class: section-slide ## Preliminary extensions - Better matching: combine lexical and acoustic evidence - Hansard-supported onomastic correction of plausible name errors --- ### Acoustic speaker matching: method and preliminary comparison - Raw faster-whisper segments sometimes overlap speakers .pull-left[ #### Method 1. Use **pyannote** to group speech into acoustic speaker clusters 2. Use ASR–Hansard matches to attach names to parts of each cluster 3. Name a cluster when one speaker consistently dominates the matched evidence 4. Extend that name to other speech in the cluster ] .pull-right[.small[ #### Results on three additional sittings: Percent of ASR segments assigned to a speaker | Sitting | Hansard text only | Acoustic + Hansard | Agreement where both named | |:---|---:|---:|---:| | 17 Aug. | 99.1% | 94.6% | 94.9% | | 18 Aug. | 98.6% | 94.2% | 95.5% | | 20 Aug. | 97.8% | 92.7% | 95.3% | ]] - Adding heuristics to the acoustic matching method (speech of the Clerk, some Speaker utterances) should improve this --- ### Hansard-supported name correction - Extract names, places, and institutions from Hansard <span class="small">(spaCy <em>en_core_web_trf</em>)</span> - Locate the corresponding ASR span using matching words on both sides - Require a fuzzy orthographical match - Save a separate corrected layer and a record of each substitution --- ### First name-correction test: 17 August | Original ASR | Hansard-supported replacement | |:---|:---| | Greg Kavanagh | Greg Cavanagh | | Hillsville | Healesville | | Unadera | Unanderra | | Sam Ray | Sam Rae | | CSRIO’s | CSIRO’s | | TEXA | TEQSA | **49 substitutions, 45 distinct pairs**, across approximately **58,000 ASR words**. --- exclude: true ### Discussion: speech enriched by Hansard - **Linguistic value:** Retains repairs, repetitions, and procedural interventions alongside the official record. - **Complementary evidence:** ASR captures spoken wording; Hansard provides speaker identities and metadata; the recording remains the ground truth. - **Key trade-off:** ASR preserves interactional nuance but introduces acoustic errors; Hansard guarantees clarity but sanitizes the interaction. - **Core objective:** A speech corpus enriched by Hansard—not an audio-linked copy of edited Hansard. --- ### Verbatim speech: Keeping disfluencies **Disfluencies as cognitive and interactional signals:** - Filled pauses (*um*, *uh*) and self-repairs can - Signal planning delay and upcoming complexity <span class="small">(Clark & Fox Tree, 2002; Arnold et al., 2003; Levelt, 1983; Bortfeld et al., 2001)</span> - Mark disagreement or resistance <span class="small">(Kendrick & Torreira, 2015)</span> - Be analyzed as sociolinguistic variables subject to regional variation and diachronic change <span class="small">(Fruehwald, 2016; Tottie, 2011; Wieling et al., 2016)</span> **Verbatim ASR possibilities:** - ParlaSpeech-Pause: use Wav2Vec2-BERT to add filled pauses - Verbatim models such as *CrisperWhisper2* <span class="small">(Wagner et al., 2026)</span> - Fine-tuned Whisper on Conversation Analysis (CA) data to "keep the ums" <span class="small">(Coats, forthcoming)</span> --- ### Towards an Australian ParlaSpeech: what remains to be done #### 1. Evaluation (primary priority) - **Verbatim gold standard:** Human-verified reference transcripts to quantify actual ASR Word Error Rate (WER) and disfluency capture - **Attribution audit:** Ground-truth evaluation of acoustic clustering, diarization boundaries, and speaker assignment precision #### 2. Corpus scaling - Scale beyond the 1-sitting proof of concept to full multi-year legislative coverage #### 3. ParlaSpeech layers & verbatim ASR - **ParlaSpeech enrichments:** UD syntax, sentiment, and topic classification - **Interoperable formats:** Parla-CLARIN TEI-XML, Praat TextGrids, BlackLab/NoSketchEngine concordancers --- class: center, middle ## Thank you! Proof-of-concept dataset: https://huggingface.co/datasets/stcoats/au-parliament-poc-turns-2026-03-30 --- class: references-all ### References .small[ .hangingindent[ Arnold, J. E., Fagnano, M., & Tanenhaus, M. K. (2003). [Disfluencies signal theee, um, new information](https://doi.org/10.1023/A:1021980931292). *Journal of Psycholinguistic Research, 32*(1), 25–36. Asonitis, A., Lanzendörfer, L. A., Berdoz, F., & Wattenhofer, R. (2026). [WorldSpeech: A multilingual speech corpus from around the world](https://arxiv.org/abs/2605.09167). *arXiv preprint arXiv:2605.09167*. Bain, M., Huh, J., Han, T., & Zisserman, A. (2023). [WhisperX: Time-accurate speech transcription of long-form audio](https://doi.org/10.21437/Interspeech.2023-78). *Proceedings of Interspeech 2023*, 4489–4493. Bortfeld, H., Leon, S. D., Bloom, J. E., Schober, M. F., & Brennan, S. E. (2001). [Disfluency rates in conversation: Effects of age, relationship, topic, role, and gender](https://doi.org/10.1177/00238309010440020101). *Language and Speech, 44*(2), 123–147. Clark, H. H., & Fox Tree, J. E. (2002). [Using uh and um in spontaneous speaking](https://doi.org/10.1016/S0010-0277%2802%2900017-3). *Cognition, 84*(1), 73–111. Coats, S. (forthcoming). Keeping the ums: Conversation-analysis-informed verbatim ASR for web speech. *Proceedings of the 13th Web as Corpus Workshop (WAC-XIII)*. Collins, P., & Smith, A. (2021). [Diachronic register change in Australian English parliamentary debates](https://doi.org/10.1007/978-981-15-8938-6_7). In P. Collins & A. Smith (Eds.), *Varieties of English in the Indo-Pacific* (pp. 147–172). Springer. Erjavec, T., et al. (2023). [The ParlaMint corpora of parliamentary proceedings](https://doi.org/10.1007/s10579-021-09574-0). *Language Resources and Evaluation, 57*, 415–448. Filppula, M. (2022). [The variable fortunes of the were-subjunctive in varieties of English](https://doi.org/10.4324/9781003025078-3). In S. Lucek & C. P. Amador-Moreno (Eds.), *Expanding the landscapes of Irish English research* (pp. 54–64). Routledge. Fruehwald, J. (2016). [Filled pause choice as a sociolinguistic variable](https://repository.upenn.edu/pwpl/vol22/iss2/7/). *University of Pennsylvania Working Papers in Linguistics, 22*(2), 41–49. Hames, S., Haugh, M., & Musgrave, S. (2025). [“How is that unparliamentary?”: The metapragmatics of ‘unparliamentary’ language in the Australian Federal Parliament](https://doi.org/10.1016/j.lingua.2025.103932). *Lingua, 320*, 103932. Kendrick, K. H., & Torreira, F. (2015). [The timing and construction of preference: A quantitative study](https://doi.org/10.1080/0163853X.2014.955997). *Discourse Processes, 52*(4), 255–289. Kotze, H., Korhonen, M., Smith, A., & van Rooy, B. (2023). [Salient differences between Australian oral parliamentary discourse and its official written records: A comparison of 'close' and 'distant' analysis methods](https://doi.org/10.1075/scl.111.02kot). In M. Korhonen, H. Kotze, & J. Tyrkkö (Eds.), *Exploring language and society with big data: Parliamentary discourse across time and space* (pp. 54–88). John Benjamins. Kruger, H., & Smith, A. (2018). [Colloquialization versus densification in Australian English: A multidimensional analysis of the Australian Diachronic Hansard Corpus (ADHC)](https://doi.org/10.1080/07268602.2018.1470452). *Australian Journal of Linguistics, 38*(3), 293–328. Levelt, W. J. M. (1983). [Monitoring and self-repair in speech](https://doi.org/10.1016/0010-0277%2883%2990026-4). *Cognition, 14*(1), 41–104. Ljubešić, N., Rupnik, P., & Koržinek, D. (2025). [The ParlaSpeech collection of automatically generated speech and text datasets from parliamentary proceedings](https://doi.org/10.1007/978-3-031-77961-9_10). In *Speech and Computer* (pp. 137–150). Springer. OpenAustralia Foundation. (n.d.). [OpenAustralia: Australian parliamentary Hansard and XML archive](https://www.openaustralia.org.au/). Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). [Robust speech recognition via large-scale weak supervision](https://proceedings.mlr.press/v202/radford23a.html). *Proceedings of ICML 2023*, 28492–28518. Tottie, G. (2011). [Uh and um as sociolinguistic markers in British English](https://doi.org/10.1075/ijcl.16.2.02tot). *International Journal of Corpus Linguistics, 16*(2), 173–197. Vaughan, J., & Mulder, J. (2014). [The survival of the subjunctive in Australian English: Ossification, indexicality and stance](https://doi.org/10.1080/07268602.2014.929086). *Australian Journal of Linguistics, 34*(4), 486–505. Wagner, L., Zusag, M., & Thallinger, B. (2026). [Transcription policy as a latent variable: Activating controllable verbatim ASR with word-level timing](https://arxiv.org/abs/2607.18934). *arXiv preprint arXiv:2607.18934*. Wieling, M., Grieve, J., Bouma, G., Fruehwald, J., Coleman, J., & Liberman, M. (2016). [Variation and change in the use of hesitation markers in Germanic languages](https://doi.org/10.1163/22105832-00602001). *Language Dynamics and Change, 6*(2), 199–234. ] ]