Evidence record 7263 · automatically gathered

Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications

LLM applications are AI systems whose nondeterministic outputs and evolving model behavior make traditional testing insufficient for release governance. We present an automated self-testing framework that introduces quality gates with evidence-based release decisions (PROMOTE/HOLD/ROLLBACK) across five empirically grounded dimensions: task success rate, research context preservation, P95 latency, safety pass rate, and evidence coverage. We evaluate the framework through a longitudinal case study

Record details

Published: 13 March 2026
Source: arXiv
Category: Research
Topics: Regulation
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (13 March 2026), “Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications,” evidence record 7263, https://ethics.ai/record/7263 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.