Accelerating Masked Image Generation by Learning Controlled Latent Dynamics — a tiny
learned module that skips the heavy 8B backbone on most MaskGIT decoding steps of
Lumina-DiMOO, giving >4x faster
text-to-image generation with almost no quality loss.
The shortcut steps slider below sets how many of the decoding steps still run the full
backbone ("budget"); every other step is computed by the lightweight shortcut module instead.