Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-Model Consensus Study on Japanese Listed Firms
Automated scoring of ESG narrative disclosures with large language models (LLMs) is gaining traction, yet whether reasoning-heavy frontier models add value commensurate with their cost remains empirically unsettled. We evaluate this question on a corpus of ten Japanese listed firms across three rubric axes -- quantitative targets, progress-tracking infrastructure, and external-standard alignment -- using a four-model consensus design that combines a reasoning-on frontier model with three reasoni
Record details
Published: 22 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Privacy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How Far Will They Go? Red-Teaming Online Influence with Large Language Models
arXiv · 20 May 2026
Measuring the Depth of LLM Unlearning via Activation Patching
arXiv · 23 May 2026
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
arXiv · 24 May 2026
Differentially Private Preference Data Synthesis for Large Language Model Alignment
arXiv · 29 May 2026
Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences
arXiv · 29 May 2026
ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows
arXiv · 13 May 2026
How to cite this record
ethics.ai (22 May 2026), “Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-Model Consensus Study on Japanese Listed Firms,” evidence record 3890, https://ethics.ai/record/3890 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.