PAPER / ARXIV:2609.09338
Fengxiang Bie, Yuqing Jian, Yifan Yu, Zhongzhu Zhou, Zelei Shao, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu, Tianyi Zhang
RESUMO
Speculative decoding drafters are typically trained for a single target model, and their acceptance rate collapses under workload shifts, since pretraining is hard to reuse across targets. Osprey bootstraps drafters from off-the-shelf pretrained small language models, treating pretraining as a reusable, target-agnostic asset via pruning to a shallow backbone and target-agnostic next-token pretraining, then adapting per target through vocabulary alignment and distillation. A single pretrained Osprey backbone improves mean acceptance length by 16-23% across Qwen3-8B, Llama-3.3-70B, and MiniMax-M2.5.
NO MESMO MAPA