Zihao Tao
2026
Disentangling Annotator Skill from Verifier Strictness in Cross-Verified Dialogue Annotation
Zihao Tao | John A. Prado | Ignazio Steven LaManna | Ryan Puterbaugh | Mim Datta | Julia Hirschberg
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Zihao Tao | John A. Prado | Ignazio Steven LaManna | Ryan Puterbaugh | Mim Datta | Julia Hirschberg
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Some dialogue corpus projects use a verify-after-annotation workflow: a second team member reviews a submitted file and records corrections. The resulting correction count mixes two signals, annotator accuracy and verifier strictness. We separate these signals for RASwDA, an audio-anchored re-alignment of 1,045 Switchboard file sides (105,005 corrections across 977 change logs) produced by five team members in 2024-2025. Verification is crossed: each of three identified annotators was checked by four different verifiers, with overlap in both directions. A cross-classified mixed-effects model on per-file corrections-per-interval assigns 10.8% of the variance to annotator identity, while the verifier random effect collapses to zero (singular fit, stable across seven leave-one-out and response-choice refits). Thus, for this boundary-realignment task, we find no detectable verifier identity effect once annotator identity, batch, and file length are controlled. Boundary placement, not label selection, accounts for 58.5% of corrections corpus-wide. This helps explain why verifier-specific strictness has little room to appear: timestamp adjustments are anchored in the audio, while DA-label changes account for only 3.0% of corrections. We release the action-typed change logs so other projects can run the same annotator-verifier decomposition on their own verification data.
Completing and Validating the Re-Aligned Switchboard Dialog Act Corpus
Run Chen | Zihao Tao | John Prado | Ignazio LaManna | Ryan Puterbaugh | Mim Datta | Julia Hirschberg
Proceedings of the 20th Linguistic Annotation Workshop (LAW XX)
Run Chen | Zihao Tao | John Prado | Ignazio LaManna | Ryan Puterbaugh | Mim Datta | Julia Hirschberg
Proceedings of the 20th Linguistic Annotation Workshop (LAW XX)
Although widely used in dialog act prediction and generation, the Switchboard Dialog Act (SwDA) corpus has performed poorly in models incorporating prosodic information because of misalignment between speech and text data. In this paper, we report our completion of the work begun in Chen et al. (2024) in addressing these misalignment issues with an improved SwDA corpus called RASwDA (Re-Aligned Switchboard Dialog Act Corpus). Now fully re-aligned and validated, RASwDA finally meets standards of accuracy allowing for classification models trained on it to exceed classification benchmarks set by models trained on other Switchboard subcorpora.