Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unc
Record details
Published: 16 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
arXiv · 13 June 2026
ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs
arXiv · 9 June 2026
GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning
arXiv · 23 June 2026
IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control
arXiv · 25 June 2026
STELLAR: Spatio-Temporal Environmental Learning with Latent Alignment and Refinement for Long-Tailed Species Distribution Modeling
arXiv · 7 June 2026
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
arXiv · 26 June 2026
How to cite this record
ethics.ai (16 June 2026), “Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement,” evidence record 910, https://ethics.ai/record/910 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.