Evidence record 3140 · automatically gathered

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates

Record details

Published: 6 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Agents & autonomy · Biotech
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (6 July 2026), “Evaluating calibrated refusal and safe usefulness in dual-use biology settings,” evidence record 3140, https://ethics.ai/record/3140 (originally published by arXiv red teaming query).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.