Wednesday, July 29, 2026

Reminiscence Environment friendly Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers


Siri Expressive Voices synthesize wealthy, configurable speech in actual time and fully on system, powered by AFM 3 Core Superior, Apple’s strongest on-device basis mannequin. This work presents the memory-efficient audio synthesis structure behind that functionality: a detokenizer that converts the semantic audio tokens emitted by the muse mannequin into high-fidelity audio inside the tight compute and reminiscence price range of the Apple Matrix Coprocessor (AMX). We convert semantic audio tokens to a residual vector quantization (RVQ) illustration with a three-component design—a streaming encoder, a temporal decoder, and a depth decoder—that systematically decouples temporal and depth processing. A single reusable depth decoder with Diffusion Transformer (DiT)-style stage conditioning generates all RVQ ranges autoregressively, changing the devoted per-level decoders of prior multi-decoder architectures, whereas causal sliding window consideration with fixed-window key-value caching yields fixed reminiscence complexity unbiased of sequence size. Deployed on the AMX, the detokenizer sustains roughly 10ms per era step—about 16x sooner than actual time—with a peak runtime reminiscence of solely ∼21MB and 329MB of on-device property, enabling steady streaming synthesis of 20–320 seconds of audio alongside the on-device basis mannequin. This fixed, small footprint replaces the linear and quadratic reminiscence scaling of typical transformer- and GAN-based approaches. Complete ablation research validate the effectiveness of key architectural parts, together with DiT conditioning mechanisms, temporal lookahead processing, and unified depth decoding methods. Audio high quality evaluation by phonetic discriminability evaluation, perceptual high quality metrics, and neural high quality estimation confirms that the proposed structure maintains synthesis constancy whereas reaching computational effectivity good points over current methodologies. The proposed structure is deployed in manufacturing as a part of Siri Expressive Voices, powering a voice overhaul with Tempo and Expressivity customizations sliders in Apple Units and assist for customized assistant voices. Working at a 1-billion-parameter activation dimension inside AFM 3 Core Superior, it improves Imply Opinion Rating (MOS) by +0.28 general (4.15 vs. 3.87) and by +0.42 on conversational speech (4.24 vs. 3.82) over the prior on-device text-to-speech system.

Related Articles

Latest Articles