The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?

📅 2026-02-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether speech large language models (Speech LLMs) are behaviorally and mechanistically equivalent to cascaded systems comprising an automatic speech recognition (ASR) model followed by a large language model (LLM). By aligning the LLM backbone architectures, the authors systematically compare four Speech LLMs against a Whisper→LLM cascade across six tasks and propose the “cascade equivalence hypothesis.” Leveraging logit lens analysis, LEACE-based concept erasure, noise robustness evaluations, and causal interventions, they provide the first evidence for the necessity of textual representations in such models. Results show that Ultravox closely mirrors its cascaded counterpart (κ = 0.93), with performance collapsing upon textual representation erasure, while Qwen2-Audio exhibits significant deviation. Moreover, most Speech LLMs outperform their cascaded equivalents by up to 7.6% under noisy conditions. The findings demonstrate that cascade equivalence is architecture-dependent and not universally valid.

Technology Category

Application Category

📝 Abstract
Current speech LLMs largely perform implicit ASR: on tasks solvable from a transcript, they are behaviorally and mechanistically equivalent to simple Whisper$\to$LLM cascades. We show this through matched-backbone testing across four speech LLMs and six tasks, controlling for the LLM backbone for the first time. Ultravox is statistically indistinguishable from its matched cascade ($κ{=}0.93$); logit lens reveals literal text emerging in hidden states; LEACE concept erasure confirms text representations are causally necessary in both architectures tested, collapsing accuracy to near-zero. Qwen2-Audio genuinely diverges, revealing cascade equivalence is architecture-dependent, not universal. For most deployed use cases, current speech LLMs are expensive cascades, and under noise, they are worse ones, with clean-condition advantages reversing by up to 7.6% at 0 dB.
Problem

Research questions and friction points this paper is trying to address.

speech LLMs
ASR
cascade equivalence
behavioral equivalence
architecture dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cascade Equivalence Hypothesis
speech LLM
implicit ASR
matched-backbone testing
concept erasure
💼 Related Jobs
No related jobs found.
J
Jayadev Billa
Unaffiliated researcher; previously at ISI@USC, Yahoo, Nuance, and BBN.