Evidence record 16527 · automatically gathered

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education

LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital

Record details

Published: 4 August 2026
Source: arXiv cs.HC
Category: Research
Topics: Safety & alignment · Children & education
Retrieved: 5 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (4 August 2026), “Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education,” evidence record 16527, https://ethics.ai/record/16527 (originally published by arXiv cs.HC).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.