Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. This paradigm leverages LLMs broad world knowledge and ease of deployment, but limited task-specific data may reduce alignment on complex scoring tasks. In particular, its impact on scoring partially correct responses that require nuanced interpretation remains underexplored. We investigate the relationship between the degree of task-specific adaptat
Record details
Published: 8 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Representation Alignment Rests on Linear Structure
arXiv · 22 May 2026
Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
arXiv · 26 May 2026
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
arXiv · 14 April 2026
The Geometric Inductive Bias of Grokking: Bypassing Phase Transitions via Architectural Topology
arXiv · 5 March 2026
CAMAL: Improving Attention Alignment and Faithfulness with Segmentation Masks
arXiv · 8 May 2026
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
arXiv · 8 May 2026
How to cite this record
ethics.ai (8 May 2026), “Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation,” evidence record 4737, https://ethics.ai/record/4737 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.