When AI Says It Feels
Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes. This policy is designed using a top-down approach and may conflict with the goal of training models to exhibit human-like intelligence using human-generated texts. Here, we performed an experiment called Human-like Model eXpressions of Feeling (HMX-feel), in which LLMs were encouraged to express feelings, intentions, and self-awareness through self-reward
Record details
Published: 4 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents
arXiv · 4 June 2026
Misaligned AI as a New Insider Risk
arXiv · 4 June 2026
CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model
arXiv · 4 June 2026
Unsupervised Pattern Analysis in Japanese Veterinary Toxicology: A Regulatory-Compliant Framework for Cross-Species Risk Assessment
arXiv · 4 June 2026
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
arXiv · 3 June 2026
AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task
arXiv · 2 June 2026
How to cite this record
ethics.ai (4 June 2026), “When AI Says It Feels,” evidence record 1450, https://ethics.ai/record/1450 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.