AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models
Pre-trained vision-language models (VLMs) exhibit strong zero-shot generalization but remain vulnerable to adversarial perturbations. Existing classification-guided adversarial fine-tuning methods often disrupt pre-trained cross-modal alignment, weakening visual-textual correspondence and degrading zero-shot performance. In this paper, we propose an Alignment-Guided Fine-Tuning (AGFT) framework that enhances zero-shot adversarial robustness while preserving the cross-modal semantic structure. Un
Record details
Published: 31 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
arXiv · 31 March 2026
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
arXiv · 31 March 2026
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
arXiv · 31 March 2026
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
arXiv · 31 March 2026
Physiological and Semantic Patterns in Medical Teams Using an Intelligent Tutoring System
arXiv · 31 March 2026
Structured Intent as a Protocol-Like Communication Layer: Cross-Model Robustness, Framework Comparison, and the Weak-Model Compensation Effect
arXiv · 31 March 2026
How to cite this record
ethics.ai (31 March 2026), “AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models,” evidence record 6559, https://ethics.ai/record/6559 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.