Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition
Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in a grounded way when the raw audio is already available. We study this question in speech emotion recognition (SER) by deriving six interpretable acoustic concept tokens from the standardised eGeMAPS paralinguistic feature set. These tokens summarise energy, pitch, dynamics, brightness, formants, and voice quality, and are appended to the textual pr
Record details
Published: 5 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Misaligned AI as a New Insider Risk
arXiv · 4 June 2026
STELLAR: Spatio-Temporal Environmental Learning with Latent Alignment and Refinement for Long-Tailed Species Distribution Modeling
arXiv · 7 June 2026
ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs
arXiv · 9 June 2026
Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
arXiv · 1 June 2026
MIMO: Multilingual Information Retrieval via Monolingual Objectives
arXiv · 29 May 2026
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
arXiv · 29 May 2026
How to cite this record
ethics.ai (5 June 2026), “Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition,” evidence record 1364, https://ethics.ai/record/1364 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.