BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The dataset combines 596 exam-style free-text legal case tasks across multiple levels of legal education and 531 short doctrinal reasoning tasks. It includes a controlled validation subset of timed human-written solutions under both unaided and human-AI co-creation conditions. We evaluate 12 contemporary LLM systems - closed flagship, efficiency-o
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Children & education
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
arXiv · 27 May 2026
Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
arXiv · 26 May 2026
Less is More: Early Stopping Rollout for On-Policy Distillation
arXiv · 26 May 2026
AI Sovereignty as National Learning Capacity: A Human-Centered Learning Mechanics Viewpoint on France, the United States, and China
arXiv · 30 May 2026
Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment
arXiv · 24 May 2026
Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach
arXiv · 2 June 2026
How to cite this record
ethics.ai (27 May 2026), “BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law,” evidence record 3609, https://ethics.ai/record/3609 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.