PAPER / ARXIV:2609.20539
Chen, S.; Wang, L.
RESUMO
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among many competing paradigms for dLLMs, from masked uniform Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of three leading approaches and prove the following: Uniform and Gaussian can sample in a number of passes that scales with the dual total correlation underlying distribution, a measure of intrinsic complexity that can be much smaller than context length. Previously, it was only known how to achieve this using masked diffusion. For a certain family of random empirical measures, we show that Θ̃(√d) passes are necessary and sufficient for uniform or Gaussian diffusion, yet there exist approximate score oracles for which Ω̃(d) passes are needed for masked diffusion. This establishes the first provable separation in parallelism between three prevailing dLLM paradigms. Contrary to popular intuition that diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in sampling diffusion are asymptotically narrower than those in sampling.
NO MESMO MAPA