PAPER / ARXIV:2609.06498
Yaojie Zhang, Linfeng Zhang, Bin Cui, Xupeng Miao
RESUMO
Block diffusion speculative decoding improves LLM inference efficiency by proposing a block of future tokens and verifying them in one forward pass, but existing methods discard verifier representations at rejected positions, forcing the drafter to reconstruct them from scratch. DFlow reuses hidden states from rejected positions to guide subsequent drafting rounds without extra target computation, using a self-condition training strategy. On Qwen3 models, DFlow consistently improves draft quality and acceptance length over DFlash.
NO MESMO MAPA