Friday, August 7, 2026

Scaling Categorical Circulate Maps – Apple Machine Studying Analysis


Steady diffusion and move matching fashions might characterize a strong various to autoregressive approaches for language modelling (LM), as they unlock a bunch of benefits presently reserved for steady modalities, together with accelerated sampling and tilting. Just lately, a number of works have demonstrated the potential for producing discrete information repeatedly by a easy move matching course of between a Gaussian and the one-hot encoded information distribution. They’ve additional proven the feasibility of accelerated sampling by way of Categorical Circulate Maps (CFMs), leading to aggressive pattern high quality within the few-step regime. Nevertheless, this methodology had solely been evaluated at comparatively modest scales (< 1B), leaving the query of its scalability utterly open. On this article, we prepare a 1.7B-parameter base move mannequin on 2.1T tokens and self-distill it right into a CFM that generates various, high-quality textual content in as few as 4 inference steps whereas sustaining near-data-level token entropy. Moreover, we introduce a probability sure for CFMs within the semi-discrete setting, and present that they can be utilized to attain the mannequin on commonplace LM benchmarks, reaching leads to the identical vary as discrete diffusion strategies. Lastly, we uncover among the challenges that come up from coaching these fashions at scale, and we offer prescriptive insights on loss weighting and time scheduling.

Related Articles

Latest Articles