PAPER / ARXIV:2609.17639
Haoke Xiao , Yueyang Liu , Yuhui Zhang , Xiang Chen , Yufei Liu , Jia Xu , Yalong Guan , Xiaolan Zhu , Xiaoyu Zhang , Shijun Wang , Shuang Yang , Zijie Meng , Zejian Zhang , Ruochen Yang , Xiangyu Wu , Tingting Gao , Han Li , Lantao Hu , Cheng Luo , Kun Gai
RESUMO
We presented SARA, an industrial framework that transforms sparse articulated user rationales into scalable recommendation signals. Its data engine curates questionnaire responses into SARA-HQ, providing explicit preference supervision for aligning SARA-7B through SFT and Quality-Refining DPO. This alignment extends rationale generation from $86{,}564$ questionnaire-covered authors to the full $10$M-author space. SARA-Ranker translates the generated positive and negative rationales into features for user--author interaction modeling and negative-feedback history modeling, connecting articulated reasons to production ranking. Evaluation on unseen authors demonstrates that SARA-7B generates more specific, relevant, and grounded rationales than the evaluated general-purpose MLLMs. On top of a strong industrial ranking baseline with multimodal features, separate online A/B tests show that positive-rationale integration increases watch time by $0.99\%$, while negative-rationale integration reduces Hate feedback by $8.16\%$. Daily refresh and more than $30$ days of production deployment further demonstrate the operational feasibility of the approach. These findings establish articulated rationales as a useful complement to behavioral and content signals, and demonstrate a practical role for MLLMs in scaling sparse human explanations into preference information that improves industrial recommendation.
NO MESMO MAPA