Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications
LLM applications are AI systems whose nondeterministic outputs and evolving model behavior make traditional testing insufficient for release governance. We present an automated self-testing framework that introduces quality gates with evidence-based release decisions (PROMOTE/HOLD/ROLLBACK) across five empirically grounded dimensions: task success rate, research context preservation, P95 latency, safety pass rate, and evidence coverage. We evaluate the framework through a longitudinal case study
Record details
Published: 13 March 2026
Source: arXiv
Category: Research
Topics: Regulation
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
arXiv · 29 March 2026
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing
arXiv · 1 June 2026
LLM Constitutional Multi-Agent Governance
arXiv · 13 March 2026
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
arXiv · 13 March 2026
Governing the Agentic Enterprise: A Governance Maturity Model for Managing AI Agent Sprawl in Business Operations
arXiv · 13 March 2026
HR-Agents: Using Multiple LLM-based Agents to Improve Q&A about Brazilian Labor Legislation
arXiv · 13 March 2026
How to cite this record
ethics.ai (13 March 2026), “Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications,” evidence record 7263, https://ethics.ai/record/7263 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.