EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames within the high-redundancy, uncurated visual streams. In
Record details
Published: 5 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare
Retrieved: 6 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
arXiv cs.LG · 5 August 2026
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv cs.CY · 5 August 2026
Towards cross-center head and neck cancer detection: a multi-level domain alignment exploration
Frontiers in Artificial Intelligence · 6 August 2026
Functional Misalignment in Human-AI Interactions on Digital Platforms
arXiv cs.CY · 6 August 2026
DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation
arXiv · 6 August 2026
Towards Interpretable Foundation Models for Retinal Fundus Images
HuggingFace Daily Papers · 3 August 2026
How to cite this record
ethics.ai (5 August 2026), “EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment,” evidence record 16667, https://ethics.ai/record/16667 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.