PAPER / ARXIV:2609.19868
Leonid Sinev, Ilya Koziev, Vladislav Leshchuk
RESUMO
Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and incoherent generation arising from learning dependencies over an intractable space of token combinations. We introduce Zarya, a family of hybrid language models that jointly optimizes an autoregressive (AR) objective and masked-diffusion objective within a single architecture. Zarya structures training data into variable-size slots and employs a curriculum that gradually increases slot granularity, enabling smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At Zarya inference, provides two distinct decoding paradigms through a unified interface: (i) MDM sampling with first-hitting denoising, (ii) slotted speculative decoding that interleaves inter-slot diffusion-based selection with intra-slot AR infilling, achieving full KV cache reuse. The inference regimes are fully decoupled, allowing a model trained with any configuration to be deployed in either mode. Extensive configurability including grouped noise patterns (Prefix Completion, Fill-In-the-Prefix, Fill-In-the-Middle), ordered sampling schedules, and noise-level permutation strategies enables flexible research exploration. We release Zarya publicly in model sizes 0.6B, 1.7B, and 4B, demonstrating performance on standard benchmarks while offering principled integration of AR and diffusion paradigms.
NO MESMO MAPA