Archive/Beyond Transcript Alignment: Diagnosing Paralinguistic Information Flow in Frozen Speech-to-LLM Adapters
Beyond Transcript Alignment: Diagnosing Paralinguistic Information Flow in Frozen Speech-to-LLM Adapters
Nurgali Kadyrbek, Madina Mansurova
30 juillet 2026
en

Abstract

Frozen speech-to-LLM systems train an adapter between a frozen audio encoder and a text LLM. We test whether adapters preserve sentence stress beyond transcripts and whether the LLM uses it. In a five-seed Qwen3-8B/WavLM baseline, transcript alignment lowers linear adapter Probe-K below the text-only KT baseline (0.211 vs. 0.290). The R1.8 configuration improves MLP-2 Probe-K from 0.245 to 0.306 (paired p=0.012) and passes three controls, but because warmup and augmentation also change, the gain is not attributed to Lcf alone; its crossing of the linear KT floor is neither statistically established nor capacity-matched. Response Probe-G remains 0.512 in both cohorts; single-seed LoRA and styled-teacher pilots do not improve it. A post hoc trace through all 36 Qwen blocks finds persistent absolute stress decodability in speech-slot states, but no reliable R1.8-over-R0 advantage at any state and no speech-conditioned answer margin; an explicit text tag instead yields a +6.40-nat margin. Attention differences depend on slot-length normalization, and late left/system concentration is shared across modalities rather than speech-specific. Thus, the tested system remains an end-to-end negative: adapter decodability does not imply causal response use. Scope is limited to sentence stress, mostly synthetic voices, one encoder, and one LLM.

IPC Classification

G06A61

Keywords

beyondtranscriptalignmentdiagnosingparalinguisticinformationflowfrozenspeech-to-llmadaptersdatacognitivecomputingsystemstrainadapteraudioencodertexttestwhetherpreservesentencestress
Citer cette publication

€ 4.00