PAPER / ARXIV:2609.20751
Anton Xue, Litu Rout, Aditya Akella et al.
RESUMO
Adapting a pretrained autoregressive model is a cost-efficient route to diffusion language models. While nearly all such adaptations start from a full-attention transformer, autoregressive modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such hybrid backbones can become effective diffusion language models by adapting Qwen3.5 at 0.8B, 2B, 4B, 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation against full-attention control, dQwen3.5 reaches a given training loss in about half the tokens. Across dQwen3.5 scales, it resembles full-attention diffusion language models, has any-order decoding behavior and performs strongly under parallel decoding.
NO MESMO MAPA