MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit a
Record details
Published: 13 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Reference-Based Face Super-Resolution Using the Spatial Transformer
arXiv cs.LG · 13 July 2026
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
arXiv · 13 July 2026
Trustworthy synthetic data for campaign decision support: strategy simulation fidelity and the PolicySynth framework
arXiv cs.LG · 13 July 2026
Towards Predictive, Aligned, and Scalable Robot Learning
arXiv · 13 July 2026
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
HuggingFace Daily Papers · 13 July 2026
Longitudinal Multi-View Breast Cancer Risk Prediction
arXiv · 13 July 2026
How to cite this record
ethics.ai (13 July 2026), “MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment,” evidence record 3124, https://ethics.ai/record/3124 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.