Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, \textit{GPT-5 mini}, \textit{Gemini 3 Flash}, and \
Record details
Published: 29 June 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Impact of Security and Privacy Controls on Users' Emotional Engagement with Generative AI Chatbots
arXiv cs.HC · 7 July 2026
Platform Choice, Trust, and Privacy in the Consumer AI Assistant Market
arXiv cs.CY · 17 July 2026
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
arXiv cs.CY · 21 July 2026
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
arXiv cs.AI · 22 July 2026
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
arXiv cs.CY · 23 July 2026
Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code
arXiv cs.CY · 28 July 2026
How to cite this record
ethics.ai (29 June 2026), “Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns,” evidence record 428, https://ethics.ai/record/428 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.