Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains
Foundation models are increasingly deployed in socially sensitive domains such as education, mental health, and caregiving, where failures are often cumulative and context-dependent. Existing guardrail approaches -- ranging from training-time alignment to prompting, decoding constraints, and post-hoc moderation -- primarily provide empirical risk reduction rather than enforceable behavioral guarantees, and largely treat safety as a property of individual outputs rather than interaction trajector
Record details
Published: 19 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare · Children & education · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning
arXiv cs.CY · 30 July 2026
Evaluating Cognitive Age Alignment in Interactive AI Agents
arXiv · 18 May 2026
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
arXiv · 24 May 2026
ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows
arXiv · 13 May 2026
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
arXiv · 28 May 2026
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
arXiv · 7 May 2026
How to cite this record
ethics.ai (19 May 2026), “Robotics-Inspired Guardrails for Foundation Models in Socially Sensitive Domains,” evidence record 4016, https://ethics.ai/record/4016 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.