Evaluating LLM Alignment With Human Trust Models
Trust plays a pivotal role in enabling effective cooperation, reducing uncertainty, and guiding decision-making in both human interactions and multi-agent systems. Although it is significant, there is limited understanding of how large language models (LLMs) internally conceptualize and reason about trust. This work presents a white-box analysis of trust representation in EleutherAI/gpt-j-6B, using contrastive prompting to generate embedding vectors within the activation space of the LLM for dia
Record details
Published: 6 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A Hazard-Informed Data Pipeline for Robotics Physical Safety
arXiv · 6 March 2026
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
arXiv · 6 March 2026
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
arXiv · 5 March 2026
Boosting deep Reinforcement Learning using pretraining with Logical Options
arXiv · 6 March 2026
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
arXiv · 6 March 2026
Alignment Backfire: Language-Dependent Reversal of Safety Interventions Across 16 Languages in LLM Multi-Agent Systems
arXiv · 5 March 2026
How to cite this record
ethics.ai (6 March 2026), “Evaluating LLM Alignment With Human Trust Models,” evidence record 7625, https://ethics.ai/record/7625 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.