Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

πŸ“… 2026-09-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
ζœ¬ζ–‡η ”η©Άι€šθΏ‡ζ½œεœ¨η©Ίι—΄θ’Έι¦εŽ‹ηΌ©ζ΅εΌη₯žη»ιŸ³ι’‘ηΌ–η ε™¨οΌŒδ»₯ε‡ε°‘ε‚ζ•°ζ•°ι‡οΌŒι™δ½ŽεŠŸθ€—ε’Œε»ΆθΏŸοΌŒεŒζ—ΆδΏζŒζ€§θƒ½γ€‚
πŸ“ Abstract
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
Problem

Research questions and friction points this paper is trying to address.

Streaming Neural Audio Encoders
Latent-Space Distillation
Tokenizer Compression
On-Device Dictation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent-Space Distillation
Pre-Quantizer Latent
Squared-Error Objective
Affine Layer
πŸ”Ž Similar Papers
No similar papers found.