Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, Geoffrey Irving. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Record details
Published: 1 January 2022
Source: OpenAlex
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Interpretable machine learning: Fundamental principles and 10 grand challenges
OpenAlex · 1 January 2022
A Systematic Review of Explainable Artificial Intelligence in Terms of Different Application Domains and Tasks
OpenAlex · 27 January 2022
The Alignment Problem: Machine Learning and Human Values
OpenAlex · 1 December 2021
Economic effects of the COVID-19 pandemic on entrepreneurship and small businesses
OpenAlex · 12 September 2021
Environment of Peace: Security in a New Era of Risk
OpenAlex · 29 April 2022
Transparency of AI in Healthcare as a Multilayered System of Accountabilities: Between Legal Requirements and Technical Limitations
OpenAlex · 30 May 2022
How to cite this record
ethics.ai (1 January 2022), “Red Teaming Language Models with Language Models,” evidence record 9190, https://ethics.ai/record/9190 (originally published by OpenAlex).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.