{
  "count": 50,
  "items": [
    {
      "id": 20183,
      "url": "https://thenextweb.com/news/anthropic-risk-report-bio-classifiers-human-feedback-gap",
      "title": "Anthropic ran 133 million contractor chats with its bioweapon filters off",
      "summary": "Anthropic published the Risk Report on 14 August, covering the period to 15 July. Axios got the company on the record and led on the misalignment rating, as did most of the coverage. Anthropic raised its estimate of catastrophic harm from misalignment in high-stakes settings. It now calls that risk low, up from very low [&hellip;] This story continues at The Next Web",
      "authors": "Ana Maria Constantin",
      "category": "news",
      "topics": "safety-alignment,biotech",
      "published_at": "2026-08-15T14:14:18.000Z",
      "source": "The Next Web AI",
      "ethics_ai_record_url": "https://ethics.ai/record/20183"
    },
    {
      "id": 19112,
      "url": "https://arxiv.org/abs/2608.12356",
      "title": "Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio",
      "summary": "arXiv:2608.12356v1 Announce Type: new Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is difficult, because the instruments available to curriculum committees, namely advisory boards, tracer studies, and employer surveys, are slow, narrow, and hard to reproduce. We apply one uniform, taxonomy-",
      "authors": "Sherzod Turaev, Saja Aldabet, Mary John, Namya Musthafa, Mamoun Awad, Nazar Zaki, Khaled Shuaib",
      "category": "research",
      "topics": "safety-alignment,jobs-economy",
      "published_at": "2026-08-14T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/19112"
    },
    {
      "id": 19122,
      "url": "https://arxiv.org/abs/2608.12323",
      "title": "Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance",
      "summary": "arXiv:2608.12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show th",
      "authors": "Mika Okamoto, Ansel Kaplan Erol, Kutluhan Erol",
      "category": "research",
      "topics": "regulation,safety-alignment,healthcare,agents-autonomy",
      "published_at": "2026-08-14T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/19122"
    },
    {
      "id": 19123,
      "url": "https://arxiv.org/abs/2608.12346",
      "title": "Position: The Alignment Community is Unintentionally Building a Censor's Toolkit",
      "summary": "arXiv:2608.12346v1 Announce Type: cross Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a \"perfectly aligned\" model inadvertently also provides malicious actors with an ever-improving tool for informational d",
      "authors": "Sarah Ball, Phil Hackemann",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-14T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/19123"
    },
    {
      "id": 19127,
      "url": "https://arxiv.org/abs/2608.12372",
      "title": "Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning",
      "summary": "arXiv:2608.12372v1 Announce Type: cross Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data sh",
      "authors": "Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-14T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/19127"
    },
    {
      "id": 19129,
      "url": "https://arxiv.org/abs/2603.29545",
      "title": "Stand-Alone Complex or Vibercrime? Exploring the adoption and innovation of GenAI tools, coding assistants, and agents within cybercrime ecosystems",
      "summary": "arXiv:2603.29545v2 Announce Type: replace Abstract: Existential risk scenarios relating to Generative Artificial Intelligence often involve advanced systems or agentic models breaking loose and using hacking tools to gain control over critical infrastructure. In this paper, we argue that the real threats posed by generative AI for cybercrime are rather different. We apply innovation theory and evolutionary economics - treating cybercrime as an ecosystem of small- and medium-scale tech start-ups,",
      "authors": "Jack Hughes, Ben Collier, Daniel R. Thomas",
      "category": "research",
      "topics": "safety-alignment,agents-autonomy",
      "published_at": "2026-08-14T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/19129"
    },
    {
      "id": 19546,
      "url": "https://link.springer.com/article/10.1007/s10462-026-11648-w",
      "title": "Data augmentation in multimodal frameworks: a survey",
      "summary": "Training machine learning models with more than one data modality has enhanced predictive performance in most contexts. Thus, many recent applications of machine learning use data from different sources and forms. Multimodal data augmentation (MMDA) addresses critical challenges in multimodal learning, such as data scarcity, modality imbalance, and cross-modal alignment. This survey systematically reviews 68 state-of-the-art MMDA approaches, and, as result, proposes a taxonomy for the area. For",
      "authors": null,
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-14T00:00:00.000Z",
      "source": "Artificial Intelligence Review",
      "ethics_ai_record_url": "https://ethics.ai/record/19546"
    },
    {
      "id": 19131,
      "url": "https://www.wired.com/story/openai-safety-security-ai-agents-culture",
      "title": "The Safety Reckoning Inside OpenAI",
      "summary": "OpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led to it.",
      "authors": "Maxwell Zeff",
      "category": "news",
      "topics": "safety-alignment,agents-autonomy",
      "published_at": "2026-08-13T22:37:19.000Z",
      "source": "Wired",
      "ethics_ai_record_url": "https://ethics.ai/record/19131"
    },
    {
      "id": 19282,
      "url": "https://fortune.com/2026/08/13/buried-in-openais-latest-research-no-correlation-between-ai-use-and-revenue-per-employee",
      "title": "Buried in OpenAI’s latest research: No correlation between AI use and revenue per employee",
      "summary": "The AI lab grapples with a question that's existential for its future: Do ChatGPT corporate customers get a clear ROI?",
      "authors": "Emily Forlini",
      "category": "news",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T20:28:57.000Z",
      "source": "Fortune AI",
      "ethics_ai_record_url": "https://ethics.ai/record/19282"
    },
    {
      "id": 19335,
      "url": "https://www.nextgov.com/policy/2026/08/gao-official-suggests-addressing-ai-risks-through-existing-legislation/415410",
      "title": "GAO official suggests addressing AI risks through existing legislation",
      "summary": "Nick Marinos, the managing director of IT and cybersecurity issues at GAO, said leveraging existing avenues for cybersecurity information sharing can expedite getting AI safety policies into law.",
      "authors": "Alexandra Kelley",
      "category": "news",
      "topics": "regulation,safety-alignment",
      "published_at": "2026-08-13T19:09:00.000Z",
      "source": "NextGov/FCW",
      "ethics_ai_record_url": "https://ethics.ai/record/19335"
    },
    {
      "id": 19160,
      "url": "https://arxiv.org/abs/2608.13482v1",
      "title": "Synthetic Persona Pretraining: Alignment from Token Zero",
      "summary": "As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the d",
      "authors": "Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui et al.",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T17:12:04.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19160"
    },
    {
      "id": 19167,
      "url": "https://arxiv.org/abs/2608.13345v1",
      "title": "Rules or Character? Scaling Laws for AI Safety Design",
      "summary": "Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes saf",
      "authors": "Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T15:15:09.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19167"
    },
    {
      "id": 19170,
      "url": "https://arxiv.org/abs/2608.13317v1",
      "title": "StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems",
      "summary": "Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that token identities alone cannot capture. Recent work proposes latent communication as an alternative, where agents transmit hidden representations directly without converting them to text. However, existing latent methods either inject working memory la",
      "authors": "Yanwen Peng, Delvin Ce Zhang, Xi Wang, Nikolaos Aletras",
      "category": "research",
      "topics": "safety-alignment,agents-autonomy",
      "published_at": "2026-08-13T14:40:59.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19170"
    },
    {
      "id": 19474,
      "url": "https://arxiv.org/abs/2608.13190v1",
      "title": "ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning",
      "summary": "Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable. However, existing methods often infer environments from a separate reference model and select representations before fitting the classifier used at deployment, leaving both decisions misaligned with the deployed predictor. In this work, we formulate group robustness without training-group labels as the endogenous environments with repair-aware selection (ERAS) problem, and",
      "authors": "Qianqian Wang, Yunshan Li, Dawei Huang, Wenwu Gong, Lili Yang",
      "category": "research",
      "topics": "safety-alignment,environment",
      "published_at": "2026-08-13T12:57:41.000Z",
      "source": "arXiv cs.LG",
      "ethics_ai_record_url": "https://ethics.ai/record/19474"
    },
    {
      "id": 19183,
      "url": "https://arxiv.org/abs/2608.13069v1",
      "title": "Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds",
      "summary": "Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively paralleliz",
      "authors": "Lucia Malíčková",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T10:33:00.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19183"
    },
    {
      "id": 19184,
      "url": "https://arxiv.org/abs/2608.13043v1",
      "title": "From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion",
      "summary": "Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache).",
      "authors": "Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang",
      "category": "research",
      "topics": "regulation,safety-alignment",
      "published_at": "2026-08-13T10:08:47.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19184"
    },
    {
      "id": 20127,
      "url": "https://www.contractsfinder.service.gov.uk/Notice/5d296690-8730-485d-9ebc-b3700dc46366",
      "title": "Project Libra AI security support",
      "summary": "Finding a low latency, contextual guardrail solution that understands AI-specific threats to ensure a secured, thorough review of all applications the FCA reviews Classification: Research and Development services on security and defence materials.",
      "authors": "FCA",
      "category": "policy",
      "topics": "regulation,safety-alignment,military-security",
      "published_at": "2026-08-13T06:18:03.000Z",
      "source": "UK Contracts Finder — AI procurement",
      "ethics_ai_record_url": "https://ethics.ai/record/20127"
    },
    {
      "id": 19472,
      "url": "https://arxiv.org/abs/2608.12821v1",
      "title": "HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models",
      "summary": "Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchi",
      "authors": "Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T04:49:51.000Z",
      "source": "arXiv red teaming query",
      "ethics_ai_record_url": "https://ethics.ai/record/19472"
    },
    {
      "id": 18719,
      "url": "https://arxiv.org/abs/2608.11955",
      "title": "Philosophical vertigo with artificial intelligence",
      "summary": "arXiv:2608.11955v1 Announce Type: new Abstract: Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabi",
      "authors": "Thomas A. Pollak (King's College London), Hamilton Morrin (King's College London), Murray Shanahan (Imperial College London)",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/18719"
    },
    {
      "id": 19191,
      "url": "https://arxiv.org/abs/2608.12788v1",
      "title": "ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs",
      "summary": "The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic compon",
      "authors": "Jiale Cui, Yueyao Yuan, Kaixi Zhong, Xiaogang Xu, Jiafei Wu, Zhe Liu",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-13T03:48:07.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19191"
    },
    {
      "id": 19194,
      "url": "https://arxiv.org/abs/2608.12689v1",
      "title": "Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging",
      "summary": "Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tas",
      "authors": "Zhi Qiao, Xintong Wu, Yichu He, Feng Shi",
      "category": "research",
      "topics": "safety-alignment,healthcare",
      "published_at": "2026-08-13T01:12:34.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/19194"
    },
    {
      "id": 18765,
      "url": "https://techcrunch.com/2026/08/12/as-ai-safety-concerns-mount-three-pioneers-make-the-case-for-staying-open",
      "title": "As AI safety concerns mount, three pioneers make the case for staying open",
      "summary": "At Ai4, three of the world's most respected AI experts — Geoffrey Hinton, Fei-Fei Li, and Andrew Ng — debated regulation, open source access, and how America can compete as China advances in Asia.",
      "authors": "Kate Park",
      "category": "news",
      "topics": "regulation,safety-alignment",
      "published_at": "2026-08-12T17:51:00.000Z",
      "source": "TechCrunch",
      "ethics_ai_record_url": "https://ethics.ai/record/18765"
    },
    {
      "id": 19067,
      "url": "https://arxiv.org/abs/2608.12084v1",
      "title": "NAE: Normalizing AutoEncoder",
      "summary": "We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d<D$) settings, and group these models under the term flow autoencoders. We present a theoretical investigation into their training dynamics and prove that the proposed loss used by existing approaches is suboptimal; specifically, both encoder and decoder surrogates must be optimized in alignment with reconstruction loss. Guided by these insights",
      "authors": "Muhammad Abdur Rafae, Niels Landwehr",
      "category": "research",
      "topics": "safety-alignment,finance-investment",
      "published_at": "2026-08-12T14:07:09.000Z",
      "source": "arXiv cs.LG",
      "ethics_ai_record_url": "https://ethics.ai/record/19067"
    },
    {
      "id": 19057,
      "url": "https://arxiv.org/abs/2608.12077v1",
      "title": "A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models",
      "summary": "Recent studies have shown that binary-to-image representations can enable effective machine learning-based results for malware detection and classification. However, performance can vary significantly, depending on the technique used to convert binaries to images. Furthermore, the explainability and interpretability of image-based models is largely unexplored within the malware domain. In this research, we employ Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool,",
      "authors": "Vibha Bhavikatti, Mark Stamp",
      "category": "research",
      "topics": "safety-alignment,transparency",
      "published_at": "2026-08-12T14:00:44.000Z",
      "source": "arXiv cs.CR (AI security)",
      "ethics_ai_record_url": "https://ethics.ai/record/19057"
    },
    {
      "id": 18783,
      "url": "https://arxiv.org/abs/2608.11955v1",
      "title": "Philosophical vertigo with artificial intelligence",
      "summary": "Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabilisation in which inherited criteria become newl",
      "authors": "Thomas A. Pollak, Hamilton Morrin, Murray Shanahan",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T11:40:50.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18783"
    },
    {
      "id": 18786,
      "url": "https://arxiv.org/abs/2608.11816v1",
      "title": "How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment",
      "summary": "State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21",
      "authors": "Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T08:58:37.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18786"
    },
    {
      "id": 18789,
      "url": "https://arxiv.org/abs/2608.11766v1",
      "title": "Instruction Alignment for Binary Code Representation Learning",
      "summary": "Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they largely ignore fine-grained instruction-level correspondences. This limitation misses valuable supervision signals available from compiler debug information, which can support the learning of more accurate and interpretable binary code representations",
      "authors": "Huaijin Wang, Shuai Wang",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T08:07:26.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18789"
    },
    {
      "id": 18472,
      "url": "https://www.sciencedirect.com/science/article/pii/S2666920X26001268?dgcid=rss_sd_all",
      "title": "Neuro-symbolic pedagogical alignment (NSPA) for long-horizon classroom discourse analysis: Mitigating dialect bias via counterfactual preference optimization",
      "summary": "Publication date: Available online 10 August 2026 Source: Computers and Education: Artificial Intelligence Author(s): Qianyi Fang, Wenhe Liu",
      "authors": null,
      "category": "research",
      "topics": "bias-fairness,safety-alignment,children-education",
      "published_at": "2026-08-12T05:10:43.828Z",
      "source": "Computers and Education: Artificial Intelligence",
      "ethics_ai_record_url": "https://ethics.ai/record/18472"
    },
    {
      "id": 19068,
      "url": "https://arxiv.org/abs/2608.11656v1",
      "title": "Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models",
      "summary": "Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and datasets. However, dominant pretraining paradigms face key challenges: masked autoencoding tends to prioritize low-level signal reconstruction over task-relevant semantics, while autoregressive modeling creates a mismatch between continuous neural dynamics and discrete token spaces. To address these challenges, ne",
      "authors": "Myeong-Ju Cho, Hye-Bin Shin, Seo-Hyun Lee, Seong-Whan Lee",
      "category": "research",
      "topics": "safety-alignment,environment",
      "published_at": "2026-08-12T04:54:43.000Z",
      "source": "arXiv cs.LG",
      "ethics_ai_record_url": "https://ethics.ai/record/19068"
    },
    {
      "id": 18797,
      "url": "https://arxiv.org/abs/2608.11631v1",
      "title": "CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement",
      "summary": "In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when cl",
      "authors": "Kuangzhao Yang, Ziliang Zhao, Zhicheng Dou",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T04:27:45.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18797"
    },
    {
      "id": 18798,
      "url": "https://arxiv.org/abs/2608.11623v1",
      "title": "FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting",
      "summary": "Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational overhead and failing to leverage the rich spectral dynamics inherent in time-series data. To enable prompt-free, frequency-aware adaptation of frozen LLMs, we propose FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series Forecasting), an autoreg",
      "authors": "Rentao Gu, Yihang Ding, Junjie Li, Yi Ding, Weijing Sang, Xiaoli Huo et al.",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T04:09:52.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18798"
    },
    {
      "id": 18356,
      "url": "https://arxiv.org/abs/2608.09937",
      "title": "Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory",
      "summary": "arXiv:2608.09937v1 Announce Type: cross Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries an",
      "authors": "Krishna Pothugunta, John P. Lalor",
      "category": "research",
      "topics": "safety-alignment,environment",
      "published_at": "2026-08-12T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/18356"
    },
    {
      "id": 18369,
      "url": "https://arxiv.org/abs/2608.11171",
      "title": "From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop",
      "summary": "arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in e",
      "authors": "Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T04:00:00.000Z",
      "source": "arXiv cs.CY",
      "ethics_ai_record_url": "https://ethics.ai/record/18369"
    },
    {
      "id": 18800,
      "url": "https://arxiv.org/abs/2608.11583v1",
      "title": "Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models",
      "summary": "Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the ef",
      "authors": "Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T02:44:30.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18800"
    },
    {
      "id": 18802,
      "url": "https://arxiv.org/abs/2608.11537v1",
      "title": "Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment",
      "summary": "Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel dis",
      "authors": "Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-12T01:00:04.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18802"
    },
    {
      "id": 18474,
      "url": "https://www.frontiersin.org/articles/10.3389/frobt.2026.1898805",
      "title": "Effect of a robotic insole-type active assist device on horizontal ground reaction force and center-of-pressure stability during stepping in patients with medial knee osteoarthritis",
      "summary": "An insole-type active assist device has been developed as a robotic system to dynamically correct ankle alignment at heel contact in patients with medial knee osteoarthritis. Although our previous feasibility study demonstrated that the device could be safely used during an on-the-spot stepping task, its effects on loading behavior remain unclear. This study aimed to investigate whether dynamic ankle alignment correction using the device alters horizontal ground reaction force variability and ce",
      "authors": "Taku Itami",
      "category": "research",
      "topics": "safety-alignment,agents-autonomy,finance-investment",
      "published_at": "2026-08-12T00:00:00.000Z",
      "source": "Frontiers in Robotics and AI",
      "ethics_ai_record_url": "https://ethics.ai/record/18474"
    },
    {
      "id": 18805,
      "url": "https://arxiv.org/abs/2608.11493v1",
      "title": "From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation",
      "summary": "Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact",
      "authors": "Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T23:06:39.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18805"
    },
    {
      "id": 18483,
      "url": "https://simonwillison.net/2026/Aug/11/stealing-reasoning-traces",
      "title": "Stealing Reasoning Traces from Proprietary LLM APIs",
      "summary": "Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext You can see an example of these encrypted blocks by running: curl https://a",
      "authors": null,
      "category": "org",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T22:40:45.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/18483"
    },
    {
      "id": 18753,
      "url": "https://arxiv.org/abs/2608.11878",
      "title": "ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents",
      "summary": "Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering an",
      "authors": "Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu",
      "category": "research",
      "topics": "safety-alignment,agents-autonomy,environment",
      "published_at": "2026-08-11T20:00:00.000Z",
      "source": "HuggingFace Daily Papers",
      "ethics_ai_record_url": "https://ethics.ai/record/18753"
    },
    {
      "id": 18807,
      "url": "https://arxiv.org/abs/2608.11392v1",
      "title": "AI Guardrail Survival under Single-Cycle Agentic Self-Summarization",
      "summary": "Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations acrossmany models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safetyrule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not asafety check: when compaction does not",
      "authors": "Ted Kwartler, Alan Aqrawi, Arian Abbasi",
      "category": "research",
      "topics": "regulation,safety-alignment,agents-autonomy",
      "published_at": "2026-08-11T19:57:55.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18807"
    },
    {
      "id": 18812,
      "url": "https://arxiv.org/abs/2608.11335v1",
      "title": "Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation",
      "summary": "Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates feat",
      "authors": "Md Maklachur Rahman, Tracy Hammond",
      "category": "research",
      "topics": "safety-alignment,healthcare",
      "published_at": "2026-08-11T18:40:52.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18812"
    },
    {
      "id": 18499,
      "url": "https://thehill.com/homenews/media/6022910-hunter-biden-donald-trump-existential-threat",
      "title": "Hunter Biden: Trump 'the one existential threat' to U.S.",
      "summary": "Hunter Biden, the son of former President Joe Biden, took a swipe at President Trump Monday, describing him as an “existential threat” to the U.S. During an interview with conservative media personality Tucker Carlson, the former president's son blamed wealthy tech leaders for the country’s current divide. He said Trump's cordial relationships with people like...",
      "authors": "Tara Suter",
      "category": "news",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T18:24:37.000Z",
      "source": "The Hill Technology",
      "ethics_ai_record_url": "https://ethics.ai/record/18499"
    },
    {
      "id": 18405,
      "url": "https://arxiv.org/abs/2608.11181v1",
      "title": "How to Verify Consistency of Probabilistic Claims",
      "summary": "When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. We construct an interactive PCP as follows. Let a predictive model be specified by a probability circuit P and a circuit Q which outputs confidence in predictions. Together, P",
      "authors": "Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T17:41:39.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18405"
    },
    {
      "id": 18406,
      "url": "https://arxiv.org/abs/2608.11171v1",
      "title": "From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop",
      "summary": "The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). W",
      "authors": "Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma et al.",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T17:30:16.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18406"
    },
    {
      "id": 18693,
      "url": "https://arxiv.org/abs/2608.11167v1",
      "title": "MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment",
      "summary": "Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Cod",
      "authors": "Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T17:28:52.000Z",
      "source": "arXiv cs.LG",
      "ethics_ai_record_url": "https://ethics.ai/record/18693"
    },
    {
      "id": 18458,
      "url": "https://80000hours.org/podcast/episodes/geoffrey-irving-superintelligence-alignment-theory",
      "title": "Geoffrey Irving on how to solve alignment before superintelligence arrives",
      "summary": "The post Geoffrey Irving on how to solve alignment before superintelligence arrives appeared first on 80,000 Hours .",
      "authors": "Tom Reed",
      "category": "org",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T15:59:39.000Z",
      "source": "80,000 Hours",
      "ethics_ai_record_url": "https://ethics.ai/record/18458"
    },
    {
      "id": 18685,
      "url": "https://arxiv.org/abs/2608.11025v1",
      "title": "Data Attribution of Emergent Misalignment with Persona Features",
      "summary": "Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diff",
      "authors": "Clemens Vetter, David Kaczér, Lucie Flek, Florian Mai",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T15:05:24.000Z",
      "source": "arXiv red teaming query",
      "ethics_ai_record_url": "https://ethics.ai/record/18685"
    },
    {
      "id": 18686,
      "url": "https://arxiv.org/abs/2608.10933v1",
      "title": "SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense",
      "summary": "Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstruction, which not only incur high computational latency but also tend to distort semantics. To address these issues, we experimentally and systematically analyze the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the fi",
      "authors": "Siyuan Liang, Yupeng Qiu, Junfeng Fang, Rong-Cheng Tu, Jiaxing Huang, Dacheng Tao",
      "category": "research",
      "topics": "regulation,safety-alignment,military-security",
      "published_at": "2026-08-11T14:01:24.000Z",
      "source": "arXiv red teaming query",
      "ethics_ai_record_url": "https://ethics.ai/record/18686"
    },
    {
      "id": 18410,
      "url": "https://arxiv.org/abs/2608.10929v1",
      "title": "FedCGR: Federated Cross-Domain Generative Recommendation",
      "summary": "Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interaction signals, are often sparse, unavailable, or privacy-sensitive across clients. To address this tension, we revisit federated CDR as generation over a stable semantic item language. By representing items as discrete semantic ID (SID) sequences de",
      "authors": "Zhuodong Liu, Hugen Lv, Xiangyu Li, Bohan Guo, Peiyu Hu",
      "category": "research",
      "topics": "safety-alignment,privacy-surveillance",
      "published_at": "2026-08-11T13:58:55.000Z",
      "source": "arXiv",
      "ethics_ai_record_url": "https://ethics.ai/record/18410"
    },
    {
      "id": 18665,
      "url": "https://arxiv.org/abs/2608.10839v1",
      "title": "The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset",
      "summary": "This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the inter",
      "authors": "Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko, Youngwoo Yoon, Gustav Eje Henter",
      "category": "research",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T12:06:16.000Z",
      "source": "arXiv cs.HC",
      "ethics_ai_record_url": "https://ethics.ai/record/18665"
    }
  ],
  "attribution": "via ethics.ai"
}