Siri Expressive Voices synthesize wealthy, configurable speech in actual time and fully on system, powered by AFM 3 Core Superior, Appleās strongest on-device basis mannequin. This work presents the memory-efficient audio synthesis structure behind that functionality: a detokenizer that converts the semantic audio tokens emitted by the muse mannequin into high-fidelity audio inside the tight compute and reminiscence price range of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) illustration with a three-component designāa streaming encoder, a temporal decoder, and a depth decoderāthat systematically decouples temporal and depth processing. A single reusable depth decoder with Diffusion Transformer (DiT)-style stage conditioning generates all RVQ ranges autoregressively, changing the devoted per-level decoders of prior multi-decoder architectures, whereas causal sliding window consideration with fixed-window key-value caching yields fixed reminiscence complexity unbiased of sequence size. Deployed on the AMX, the detokenizer sustains roughly 10ms per era stepāabout 16x sooner than actual timeāwith a peak runtime reminiscence of solely ā¼21MB and 329MB of on-device property, enabling steady streaming synthesis of 20ā320 seconds of audio alongside the on-device basis mannequin. This fixed, small footprint replaces the linear and quadratic reminiscence scaling of typical transformer- and GAN-based approaches. Complete ablation research validate the effectiveness of key architectural parts, together with DiT conditioning mechanisms, temporal lookahead processing, and unified depth decoding methods. Audio high quality evaluation by phonetic discriminability evaluation, perceptual high quality metrics, and neural high quality estimation confirms that the proposed structure maintains synthesis constancy whereas reaching computational effectivity good points over current methodologies. The proposed structure is deployed in manufacturing as a part of Siri Expressive Voices, powering a voice overhaul with Tempo and Expressivity customizations sliders in Apple Units and assist for customized assistant voices. Working at a 1-billion-parameter activation dimension inside AFM 3 Core Superior, it improves Imply Opinion Rating (MOS) by +0.28 general (4.15 vs. 3.87) and by +0.42 on conversational speech (4.24 vs. 3.82) over the prior on-device text-to-speech system.
