PAPER / ARXIV:2609.13168
Maria Teleki , Kimi Wenzel , Anna Seo Gyeong Choi , Tobias Weinberg , Shree Harsha Bokkahalli Satish , Stephanny Sanchez , Belu Ticona , Ariadna Sanchez , Yash Sonkar , Aarti Mathur , Christoph Minixhofer , Abraham Glasser , Raja Kushalnagar , James Caverlee , Minha Lee , Shaomei Wu , Alyssa Hillary Zisk , Éva Székely , Dylan Gaines , Angelika Seeschaaf Veres , Seray Ibrahim , Nicholas Cummins , Allison Koenecke
RESUMO
Speech AI, any AI system that recognizes, transforms, or generates speech, is built and evaluated across two communities with only a small overlap: technical natural language processing (NLP) venues (e.g., ACL, ICASSP, Interspeech), and sociotechnical HCI venues (e.g., ASSETS, CHI, FAccT). In this position paper, we work toward a cross-community synthesis, organizing our critique around three problems: speech AI operates with an incomplete model of communication; it operates with an incomplete model of identity; and its metrics measure the wrong constructs. We draw on AAC as a setting where these failures are most visible and their stakes highest, alongside other underserved speakers - people who stutter, multilingual speakers, and non-binary and transgender users. For each problem we offer solution sketches oriented toward designing for human variability, nearly all of which require quantitative and qualitative methods in combination. We close on the venue structures that hold these methods apart, and on what program committees and individual authors can do to bring them together.
NO MESMO MAPA