RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to "jailbreak" LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstru
Record details
Published: 29 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 31 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
arXiv cs.LG · 29 July 2026
Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
arXiv cs.LG · 30 July 2026
APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
arXiv · 30 July 2026
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
arXiv · 28 July 2026
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
arXiv · 31 July 2026
Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI
arXiv cs.CY · 28 July 2026
How to cite this record
ethics.ai (29 July 2026), “RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation,” evidence record 15242, https://ethics.ai/record/15242 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.