PAPER / ARXIV:2609.07549
Chuanmeng Bian , Daren Chen , Peixin Chen , Zhigao Chen , Zhiyun Fan , Zhifu Gao , Bo Gong , Qing Gu , Jiajun He , Yawei Hu , Yunjie Ji , Jingbei Li , Xiangang Li , Xu Li , Zengxi Li , Zheng Li , Chengdong Liang , Baiji Liu , Ying Liu , Bin Ma , Yiping Peng , Yuezhang Peng , Zhendong Peng , Yu Pu , Yang Shi , Xin Shu , Jian Tang , Biao Tian , Peiyao Wang , Tianzi Wang , Wen Wang , Wupeng Wang , Cheng Wen , Yuzhong Wu , Zijian Xia , Yunchong Xiao , Nan Yang , Jianwei Yu , Jixing Yu , Binbin Zhang , Lei Zhang , Sitong Zhao , Guangdong Zhou , Yuan Zhou , Jianheng Zhuo
RESUMO
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.
NO MESMO MAPA