AI safety report: research, incidents and governance
A living AI safety report showing coverage volume, sources and the latest source-linked evidence, updated daily. Coverage counts are signals of attention—not measures of importance, harm or consensus.
Prepared by the ethics.ai evidence desk · automatically refreshed · editorial scope reviewed against the methodology and corrections policy
Tracks technical safety research, evaluations, red-teaming, alignment work and the governance of advanced systems. It does not combine unlike risks into a single severity score.
Questions to take into the evidence
What safety evidence is being published and by whom?
Which evaluation and assurance methods are gaining attention?
How are technical and governance approaches interacting?
Anthropic published the Risk Report on 14 August, covering the period to 15 July. Axios got the company on the record and led on the misalignment rating, as did most of the coverage. Anthropic raised its estimate of catastrophic harm from misalignment in high-stakes settings. It now calls that risk low, up from very low […] This story continues at The Next Web
arXiv:2608.12356v1 Announce Type: new Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is difficult, because the instruments available to curriculum committees, namely advisory boards, tracer studies, and employer surveys, are slow, narrow, and hard to reproduce. We apply one uniform, taxonomy-
arXiv:2608.12323v1 Announce Type: cross Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show th
arXiv:2608.12346v1 Announce Type: cross Abstract: This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational d
arXiv:2608.12372v1 Announce Type: cross Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data sh
arXiv:2603.29545v2 Announce Type: replace Abstract: Existential risk scenarios relating to Generative Artificial Intelligence often involve advanced systems or agentic models breaking loose and using hacking tools to gain control over critical infrastructure. In this paper, we argue that the real threats posed by generative AI for cybercrime are rather different. We apply innovation theory and evolutionary economics - treating cybercrime as an ecosystem of small- and medium-scale tech start-ups,
Training machine learning models with more than one data modality has enhanced predictive performance in most contexts. Thus, many recent applications of machine learning use data from different sources and forms. Multimodal data augmentation (MMDA) addresses critical challenges in multimodal learning, such as data scarcity, modality imbalance, and cross-modal alignment. This survey systematically reviews 68 state-of-the-art MMDA approaches, and, as result, proposes a taxonomy for the area. For
This report is assembled automatically from source metadata and keyword classifications. It summarizes what the tracked source fleet published; it does not independently validate every linked claim. Source-fleet growth can inflate historical comparisons. Cite the individual evidence record and original publisher for substantive claims.
We'd like to use Google Analytics to see which pages get read — no ads, no selling data. See the privacy policy.