RAGShield: Detecting Numerical Claim Manipulation in Government RAG Systems
Retrieval-Augmented Generation (RAG) systems are deployed across federal agencies for citizen-facing tax guidance, benefits eligibility, and legal information, where a single incorrect number causes direct financial harm. This paper proves that all embedding-based RAG defenses share a fundamental blind spot: changing a tax deduction by $50,000 produces cosine similarity 0.9998, invisible to every known detection threshold. Across 174 manipulation pairs and two embedding models, the mean sensitiv
Record details
Published: 1 April 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SentinelAgent: Intent-Verified Delegation Chains for Securing Federal Multi-Agent AI Systems
arXiv · 3 April 2026
CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks
arXiv · 30 March 2026
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
arXiv · 5 April 2026
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
arXiv · 10 April 2026
Empathic and agentic artificial intelligence in nursing: perspectives on a human-centered framework for cancer care navigation in the United States
arXiv · 11 April 2026
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
arXiv · 12 April 2026
How to cite this record
ethics.ai (1 April 2026), “RAGShield: Detecting Numerical Claim Manipulation in Government RAG Systems,” evidence record 6512, https://ethics.ai/record/6512 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.