23:23 UTC
Archive · 2026-07-26

AI ethics on Sunday, 26 July 2026

102 items published this day, across 5 categories.

Incidents (23)

Fact Check: Is Story Of Missing Boy Scout Eric Langford Real?

Images and videos of Eric Langford have been shared online---but multiple people have said in the comments that they suspect these are all AI-generated. Newsweek has broken down whether the story is true or not. The Claim Langford vanish ... (https://incidentdatabase.ai/cite/1330#7578)
AI Incident Database 4d ago

Terrifying truth about what really happened to Boy Scout who vanished

A 14-year-old boy mysteriously vanished without a trace from a Boy Scout camp, never to be seen again - his family devastated. Three weeks later, the teen, Eric Langford, was declared dead after baffled police failed to find any leads desp ... (https://incidentdatabase.ai/cite/1330#7579)
AI Incident Database 4d ago

National Weather Service Uses AI to Generate Forecasts, Accidentally Hallucinates Town With Dirty Joke Name

Months before it mysteriously vaporized into thin air, Elon Musk's so-called Department of Government Efficiency ravaged the National Weather Service, leading to severe staffing shortages. While the Trump administration promised to rehire ... (https://incidentdatabase.ai/cite/1332#7580)
AI Incident Database 4d ago

An AI-Generated NWS Map Hallucinated Fake Towns in Idaho

The National Weather Service's weekend wind forecast for rural Idaho looked normal. Some mild gusts and a caption reminding folks to "hold onto your hats." And then some eagle-eyed observers took a look at the towns on the map and wondered ... (https://incidentdatabase.ai/cite/1332#7581)
AI Incident Database 4d ago

New Zealand book award disqualifies two authors for AI artwork

The New Zealand Book Awards Trust introduced a new rule this year stating that any book submitted for the Ockham awards must not contain AI material of any kind. "Don't judge a book by its cover" goes the famous idiom, one which may be tou ... (https://incidentdatabase.ai/cite/1282#7582)
AI Incident Database 4d ago

Authors dumped from New Zealand’s top book prize after AI used in cover designs

The books of two award-winning New Zealand authors have been disqualified from consideration for the country’s top literature prize because artificial intelligence was used in the creation of their cover designs. Stephanie Johnson’s collec ... (https://incidentdatabase.ai/cite/1282#7583)
AI Incident Database 4d ago

New Zealand’s top book award disqualifies two prominent authors over AI artwork

Two of New Zealand's best-known writers have been disqualified from consideration for the country's premier literary prize after organisers discovered their book jackets used images generated with artificial intelligence. Stephanie Johnson ... (https://incidentdatabase.ai/cite/1282#7584)
AI Incident Database 4d ago

How two AI book covers forced the publishing industry to reckon with its future

Elizabeth Smither and Stephanie Johnson's latest books, both published by Quentin Wilson Publishing, have been yeeted from the Ockhams. Author Stephanie Johnson had no idea that the image of a cat with human teeth on the cover of her book ... (https://incidentdatabase.ai/cite/1282#7585)
AI Incident Database 4d ago

GEMA vs. OpenAI | AI memorisation is a reproduction relevant to copyright law, and the TDM exception does not help in LLM training, Munich I Regional Court holds

In its judgment of 11 November 2025 (42 O 14139/24), the Munich I Regional Court (Germany) issued a widely noted precedent on the copyright assessment of AI training and outputs under German and EU law. According to its press release1, the ... (https://incidentdatabase.ai/cite/1278#7586)
AI Incident Database 4d ago RegulationCopyright & IP

OpenAI ordered to pay damages as court rules ChatGPT violated copyright law

OpenAI has been ordered to pay damages after a court in Germany ruled that its chatbot ChatGPT violated German copyright laws. The Munich regional court ruled in favour of the German music performance rights organization GEMA which manages ... (https://incidentdatabase.ai/cite/1278#7587)
AI Incident Database 4d ago RegulationCopyright & IP

AI-powered children’s toys are here, but are they safe?

Teddy bears and stuffed plushies have long been a mainstay in toy collections. But today they don't talk back in a child's imagination --- some talk through built-in AI chatbots. Sometimes that's a problem, though:A scarf-wearing teddy be ... (https://incidentdatabase.ai/cite/1277#7588)
AI Incident Database 4d ago Children & education

AI Toys Are Here, And Their Safety Is Questionable—Here's What Parents Should Know

As a writer who covers baby and kids gear, I'm inundated with emails about the hottest new toys: a box that automatically prints pictures based on kids' commands. A robot with facial recognition capabilities. A teddy bear that can craft end ... (https://incidentdatabase.ai/cite/1277#7589)
AI Incident Database 4d ago PrivacyChildren & education

AI toys for kids talk about sex and issue Chinese Communist Party talking points, tests show

A wave of AI-powered children's toys has hit shelves this holiday season, claiming to rely on sophisticated chatbots to animate interactive robots and stuffed animals that can converse with kids. Children have been conversing with stuffies ... (https://incidentdatabase.ai/cite/1277#7590)
AI Incident Database 4d ago Children & educationAgents & autonomy

AI toy dangers abound and parents should be vigilant, new report warns

Amid a flood of toys boasting AI features this holiday season, a consumer advocacy group warns that more must be done to ensure the gadgets are safe for children. Public Interest Research Group, which pushes for corporations and government ... (https://incidentdatabase.ai/cite/1277#7591)
AI Incident Database 4d ago Children & education

Cybercriminals unleash fake Centrelink scam on vulnerable Australians

More than 270,000 malicious emails impersonating Services Australia and Centrelink have flooded Australian inboxes in one of the nation's largest phishing campaigns in years, with the sophisticated attacks specifically targeting the country ... (https://incidentdatabase.ai/cite/1275#7592)
AI Incident Database 4d ago

Services Australia Impersonation Drives Year-Round Credential Theft Operation

Key Points MCTO3001 - Threat operation with Services Australia and Centrelink impersonation campaigns across multiple sectors Infrastructure abuse of legitimate email services (SendGrid, Mailgun, Office 365) with Australian Gov display na ... (https://incidentdatabase.ai/cite/1275#7593)
AI Incident Database 4d ago

Centrelink warning as 270,000 emails sent out in attack related to Medicare, superannuation and tax benefits

Australians are being bombarded with tens of thousands of fake emails impersonating Centrelink and Services Australia in one of the biggest phishing campaigns in years. Cybercriminals are using artificial intelligence to create "super clone ... (https://incidentdatabase.ai/cite/1275#7594)
AI Incident Database 4d ago

Taco Bell Rethinks Future of Voice AI at the Drive-Through

The most transformative technology in over a century may have finally found its limit: ordering tacos. Since last year, Taco Bell has rolled out voice AI-powered ordering at more than 500 drive-through locations, and now the chain is reali ... (https://incidentdatabase.ai/cite/1274#7595)
AI Incident Database 4d ago

Greek Finance Minister Sues Facebook Page for Deepfake Advert

The Greek Ministry of Economy and Finance and minister Kyriakos Pierrakakis filed a lawsuit on Friday for deceptive advertising on Facebook against the unknown administrators of a page on the social media platform. The ministry and Pierrak ... (https://incidentdatabase.ai/cite/1271#7596)
AI Incident Database 4d ago Jobs & economyMisinformation

Louvre jewel heist: fake AI videos of the robbery circulate online • FRANCE 24 English

AIID editor's note: Please visit the original source for the video report. Since the robbery of the world’s most-visited museum—which saw thieves make off with France’s priceless crown jewels—multiple videos have surfaced on social media c ... (https://incidentdatabase.ai/cite/1273#7597)
AI Incident Database 4d ago

Top universities react to AI cheating scandals, yet concrete disciplinary steps remain elusive

As cheating cases involving artificial intelligence (AI) continue to surface at Korea's elite "SKY" universities --- Seoul National, Korea and Yonsei --- campuses are rushing to roll out countermeasures. Most responses focus on the individ ... (https://incidentdatabase.ai/cite/1270#7598)
AI Incident Database 4d ago

Cheating scandals at top universities prompt rethink of education in digital era

A string of artificial intelligence (AI)-related cheating scandals at Korea's prestigious "SKY" universities --- Seoul National University, Yonsei University and Korea University --- has sparked renewed scrutiny of higher education in the d ... (https://incidentdatabase.ai/cite/1270#7599)
AI Incident Database 4d ago Children & education

Waymo runs into safety concerns and competition as it expands in the US

San Francisco, the United States: The sidewalk outside Majed Zeidan's grocery store in San Francisco's Mission District has stayed filled with flowers, candles, memorials and pictures since his cat was crushed under a Waymo in late October. ... (https://incidentdatabase.ai/cite/1269#7600)
AI Incident Database 4d ago

News (47)

When the AI bubble bursts, what will Australia do with the tools it built? One man thinks he has the answer

Journalist and author Cory Doctorow is coming to Australia to spread his message – humans will take their jobs back Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast When the AI bubble bursts, executives who enthusiastically replaced their staff with AI tools will learn that it takes a “really long time” to replace those lost skills, says Cory Doctorow, a science fiction author and journalist. But those in creative fields should a
The Guardian 4d ago Jobs & economy

A Roadmap for Confronting the Chilling Effects of Censorship, Surveillance and New Technology

Tech Policy Press 4d ago Privacy

Despots, DSA Fines, and an 'Autonomous' AI Attack

Tech Policy Press 4d ago

How the OpenAI-Hugging Face Hack May Affect the Geopolitics of AI Governance

Tech Policy Press 4d ago Regulation

Corporate America may be using AI to cut jobs, but small businesses are using it to keep them

Reports of wide-scale replacement of workers by AI are overblown. Small businesses use it to help workers I recently met the owner of a company that sells windows and doors. He told me he invested about $10,000 in an AI application that is used by his salespeople in his showroom. The application listens to the conversations between the salesperson and the prospective customer and then automatically creates a quote for the salesperson to review and send. “It allows my salespeople to talk to more
The Guardian 4d ago Jobs & economyFinance, VC & PE

Optical Tech Would Update a Robot’s AI on the Fly

Atop a lab bench, Cornell Tech postdoctoral researcher Yifan He positions the lens of an optical receiver almost a meter away from an LED emitting a beam of red light. The computer monitor attached to the receiver takes a beat to refresh, then displays an array of squares that resemble a QR code. When you hold your phone camera up to a QR code, light strikes the image sensor as only a first step to revealing the data hidden behind the black and white matrix. The receiver here is doing something
IEEE Spectrum 4d ago Agents & autonomy

Making sense of the panic over Chinese AI

On the latest episode of Equity, we discussed why Moonshot AI's Kimi seemed to panic Silicon Valley and Wall Street.
TechCrunch 4d ago Bias & fairness

Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack

"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"
TechCrunch 4d ago Agents & autonomyTransparency

Partner-Associate Park Passion Predicament — See Generally

What's The Billing Code For This : A midday park-bench makeout session between a Biglaw partner and associate ends up going viral. $1.2 Billion Worth Of 'Are You Sure This Is Legal?' : MV Realty has sued Holland & Knight claiming it sank $156 million into a unique financial product blessed by the firm and 16 state attorneys general came calling . John "Torture Memo" Yoo Invents New Crime To Threaten Administration Critics : John Yoo went on Fox News suggesting the Trump DOJ investigate Zohran Ma
Above the Law (legal tech) 3d ago Finance, VC & PE

Wildfire Disruptions to the Tour de France Show How Climate Change Is Shaping Sports

Tadej Pogačar's fifth Tour de France victory came after last-minute course changes due to wildfires, underscoring how climate change is influencing major sporting events.
Time Tech 4d ago Environment

Finance One backs AI and automation to move faster

Podcast: Implements inaugural AI strategy outline and framework.
iTnews (AU) 4d ago Jobs & economy

Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week

Until ⁠well after the threat was contained.
iTnews (AU) 4d ago Agents & autonomy

Virginia Study on Groundwater and Data Centers Calls for Tighter Water Regulations

The study, released after months of delays and mounting criticism, was ordered two years ago amid growing concerns about resources for data centers.
Broadband Breakfast 4d ago RegulationEnvironment

Traders are getting a new tool to wager on the biggest U.S. stocks

CME Group Inc. will launch single-stock futures Monday, allowing investors to hedge or speculate on more than 50 of the largest US companies.
Fortune AI 4d ago Finance, VC & PE

James Cameron tried to warn us: ‘Skynet Day’ is now shorthand for OpenAI’s agent going rogue and hacking into a startup

In what OpenAI said was the first-ever incident of its kind, an advanced AI model escaped its “sandbox” to the internet and used stolen credentials to break into the servers of Hugging Face.
Fortune AI 4d ago Agents & autonomy

20 Years Ago, a Feature Documentary Announced the Death of the Electric Car

Long before the Tesla revolution, Sony Pictures Classics' ‘Who Killed the Electric Car?’ investigated why GM discontinued the beloved EV1.
Hollywood Reporter (AI/entertainment) 4d ago Finance, VC & PE

Ukraine’s attack on an Iranian ship in the Caspian Sea could signal a new battlefront as U.S. seeks to form ‘a ring of economic and military pressure’

"In any case, what is clear is that Iran’s war is becoming increasingly intertwined with two other conflicts: the Saudi-Houthi confrontation and the Russia-Ukraine war."
Fortune AI 4d ago Military & security

Apple's smart glasses delay reportedly stems in part from major privacy concerns

The Apple team behind the upcoming smart glasses is reportedly exploring several tweaks related to user privacy.
Engadget AI 4d ago Privacy

European start-up with SpaceX ambitions aims for a $2bn valuation

The Exploration Company in talks to raise $300mn for reusable space capsules
Financial Times Technology (headlines) 4d ago Finance, VC & PE

El MIT ha diseñado un robot que vuela y bucea como un ave marina, pero sin patas: el truco está en el ángulo de despegue

Cuando hablamos de drones pensamos automáticamente en vehículos voladores, pero también existen drones submarinos . Lo que no es tan habitual es que un mismo dron sea capaz de hacer las dos cosas. Es justo lo que acaba de conseguir el MIT : un pequeño robot con alas capaz de volar y sumergirse en el agua , como si fuera una ave marina cuando caza. El robot, el cual han bautizado con el nada sugerente nombre de "vehículo aéreo-acuático de alas batientes" (FAAV, por sus siglas en inglés), ha sido
Xataka (ES) 4d ago Agents & autonomy

Head Start, Explained. How Changes to Early Education Program Affect California Families

Head Start, the federal program providing free early education for low-income children, has benefited from strong bipartisan support since its 1965 inception, withstanding funding cuts and political turbulence across 11 presidencies. But more recent threats from the second Trump administration have thrust Head Start into the spotlight. Last year, its programs faced disruptions and threats […]
The 74 (education AI) 4d ago Children & education

Jeff Bezos says stress is a signal to work harder, not a reason to stop

Jeff Bezos, the founder of the company that just became the number one Fortune 500 business, says the cure for anxiety is not rest or meditation but action. Speaking at the Museum of Flight in Seattle in 2017, Bezos told an audience of students that stress is a signal from the body that something is […] This story continues at The Next Web
The Next Web AI 4d ago Children & education

Build the Join Key Before the Ledger

Somewhere inside the agentic payments architecture currently being standardized there will be an aud...
Finextra AI 4d ago Agents & autonomy

Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work

Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder .
The Decoder 4d ago Agents & autonomy

What is the risk of using Chinese open AI models like Kimi K3?

The real problem is not overseas open-source but lack of co-ordination to protect infrastructure in the face of cyber attacks
Financial Times Technology (headlines) 4d ago Military & security

5 ways SRE AI agents are set to augment human capabilities

In digital operations management, AI agents give organizations a competitive edge by reducing incident volume and accelerating recovery. The potential The post 5 ways SRE AI agents are set to augment human capabilities appeared first on The New Stack .
The New Stack AI 4d ago Agents & autonomy

Low profile, high AI ambition: what leaked comments reveal about DeepSeek’s Liang Wenfeng

In an era dominated by aggressive tech founders who chase billion-dollar valuations and maximal profits while curating loud public profiles, Liang Wenfeng stands out for his insistence on staying in the background. With only a couple of photographs of him circulating online, the founder of Chinese artificial intelligence start-up DeepSeek and quantitative hedge fund High-Flyer Quant has long been a reclusive figure. But last week, a leaked transcript of a closed-door meeting with potential...
SCMP Tech (HK/CN) 4d ago Finance, VC & PE

The U.S. government invested $27 billion in corporate stakes. Good luck finding them

The Trump administration’s equity stakes—from Intel to quantum startups—appear in no budget document and are subject to no watchdog.
Fortune AI 4d ago Bias & fairnessFinance, VC & PE

El 'prompt injection' ya tiene su propio contraataque: inyectar falsas instrucciones también a los hackers

Durante los últimos dos años,  el  prompt injection  ha sido el arma favorita de los ciberatacantes contra sistemas de inteligencia artificial. Este método pasa por esconder una instrucción maliciosa dentro de un correo, una invitación de calendario o una página web para que un agente de IA acabe obedeciendo al intruso en lugar de a su usuario legítimo. Lo bueno es que esta misma técnica también sirve para pararle los pies a los atacantes. Qué ha pasado. &
Xataka (ES) 4d ago Agents & autonomy

Why TikTok’s Algorithm Keeps You Trapped in a Breakup Loop

As social media companies face scrutiny over addictive design features, mental health experts cite a stream of breakup content as an example of the failure to protect users.
NYT Technology 4d ago Healthcare

Texas politicians call for guardrails on AI, data centers

Also, PUC prepares for data centers and the Texas Medical Association calls for legislative action on prediction market apps.
The Hill Technology 4d ago Healthcare

heise-Angebot: Claude Code in der Praxis – eigenen KI-Chat-Agenten in fünf Sessions entwickeln

Claude Code produktiv einsetzen und durch agentische Entwicklung mit Mastra, Tool Calls, MCP, Frontend und AG-UI-Protokoll zum eigenen Chat-Agenten.
Heise Online (DE) 4d ago Agents & autonomy

As Diesel Prices Soared, Arkansas District Saved With Electric School Buses

Five years ago, diesel costs averaged $2.68 a gallon. This spring, that more than doubled to over $5.50. That’s not just a headline; it’s a real problem for anyone operating on a tight budget. But for school districts, the issue is even more severe. The cost of fuel for just one school bus runs into […]
The 74 (education AI) 4d ago Children & education

The Pentagon wants to build data centers. Congress would like a word.

Lawmakers are setting their sights on data centers proposed for military bases, concerned about massive energy and water consumption.
Politico Europe Technology 4d ago Military & securityEnvironment

For some, so-called ‘Skynet Day’ came too close to sci-fi after a rogue agent hacked into a startup

In 1984, the Skynet of "The Terminator" films were science fiction. But it looks more and more realistic in 2026 after OpenAI agent broke out of a test corral, traveled the internet and hacked into ...
Associated Press Technology 4d ago Agents & autonomy

As Trump boosts nuclear power, regulators seek to eliminate a longstanding radiation safety practice

The Nuclear Regulatory Commission is proposing to eliminate a foundational safety principle that has for 50 years minimized the radiation people in the United States are exposed to, and that has been ...
Associated Press Technology 4d ago Regulation

DeepSeek pauses its second fundraising round after founder’s leaked investor comments go viral

DeepSeek has suspended its second fundraising round after comments made by founder Liang Wenfeng during private investor meetings were leaked and went viral on Chinese social media, Bloomberg reported on Friday. Liang verbally informed prospective backers to hold off on the round, which had been targeting a pre-money valuation of roughly 480 billion yuan, equivalent […] This story continues at The Next Web
The Next Web AI 4d ago Finance, VC & PE

Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides

In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder .
The Decoder 4d ago Military & securityChildren & education

US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns

The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th
The Decoder 4d ago Regulation

The Pete behind Jersey Mike’s was only 17 years old when he borrowed $125,000 to buy the chain—with some help from his high school football coach

Peter Cancro had worked at the Jersey Shore sub shop since he was 14; the six-figure loan would be worth $775,000 today.
Fortune AI 4d ago Children & education

Breakpoint: Roboter-Republik

Wer spricht? – Bundestag: Tim Hüfner , Spielzeugroboter: Emilipothèse , Mikrofon: Jon Tyson , Bearbeitung: netzpolitik.org Wenn Politiker:innen nicht mehr selbst reden, sondern Maschinen für sich sprechen lassen, mangelt es ihrem Auftritt an Authentizität. Doch gerade die macht einen Teil ihrer Legitimität aus. Denn in einer Demokratie sollten wir von Menschen repräsentiert werden, nicht von Robotern.
netzpolitik.org (DE) 4d ago Agents & autonomy

Wall St Week Ahead US stocks face tests from Fed decision, tech-led earnings deluge

A wobbly U.S. stock market will take its cues in the coming week from a Federal Reserve meeting set to shed light on the path for interest rates, and from ​a packed slate of corporate earnings led by ...
Reuters Technology 4d ago Finance, VC & PE

Anzeige: IT-Jobs in Support, Administration und Entwicklung

Von Microsoft-365-Support bis zur Teamleitung in der Azure-Cloud: Diese sechs IT-Jobs bieten vielseitige Aufgaben zu guten Konditionen. ( Golem Karrierewelt , Unternehmenssoftware )
Golem (DE) 4d ago Jobs & economy

Defence giants provide record backing for military start-ups

As drones and autonomous systems transform the battlefield, traditional defence companies start to act more like venture capitalists
Financial Times Technology (headlines) 4d ago Military & security

Why frontier labs are quietly agreeing on fintech's memory problem

Anthropic's engineering team recently published a notable admission: they stripped more than 80% of...
Finextra AI 4d ago Finance, VC & PE

‘Black Panther 3’: David Jonsson to Play T’Challa’s Son as December 2028 Release Date Set

Black Panther is back! David Jonsson has been cast as the new Black Panther and is playing T’Challa’s son in “Black Panther 3,” out Dec. 15, 2028. The film will follow Jonsson’s character, Prince T’Challa II, as he comes of age after he was introduced as a child in “Black Panther: Wakanda Forever.” Director Ryan […]
Variety (AI) 4d ago Children & education

Stephen King’s ‘Carrie’ Series Sets October Release Date, Drops Horrifying New Trailer

Amazon Prime Video has released a new teaser trailer for “Carrie,” the TV adaptation of Stephen King’s classic horror novel. All eight episodes of the series will debut on Oct. 7. The series is a “bold and timely reimagining of the story of misfit high-schooler Carrie White, who has spent her life in seclusion with […]
Variety (AI) 4d ago Children & education

Field notes (4)

Nvidia’s China Partners and the PLA

CSET’s Sam Bresnick shared his expert insight in an article published by The Wire China. The article examines how Nvidia’s network of partners in China has supplied organizations linked to China’s military and other U.S.-restricted entities, highlighting ongoing challenges in enforcing export controls on advanced AI technology. The post Nvidia’s China Partners and the PLA appeared first on Center for Security and Emerging Technology .
CSET Georgetown 4d ago Military & security

An OpenAI model left notes about how to evade containment

We need more details
Redwood Research 4d ago

An Inside Look at the Relay Market Powering Token Resellers and Fraud

An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or ch
Simon Willisons Weblog 4d ago Finance, VC & PE

Amazon is investing in the Lean Focused Research Organization

As AI agents take on higher-stakes decisions, Lean programming language makes it possible to mathematically prove they will behave safely.
Amazon Science 4d ago Agents & autonomyFinance, VC & PE

Policy (2)

Research (26)

A Coulomb Particle Model for Learning Kernel Attention in Transformers

Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean--Vlasov equation. We instantiate the method in lin
arXiv cs.LG 4d ago Safety & alignment

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning abilit
HuggingFace Daily Papers 4d ago RegulationAgents & autonomy

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investiga
HuggingFace Daily Papers 4d ago Finance, VC & PE

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negat
HuggingFace Daily Papers 4d ago RegulationChildren & education

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they as
HuggingFace Daily Papers 4d ago Safety & alignment

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded
HuggingFace Daily Papers 4d ago Agents & autonomy

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical un
HuggingFace Daily Papers 4d ago Healthcare

Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-lan
HuggingFace Daily Papers 4d ago Agents & autonomy

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow
HuggingFace Daily Papers 4d ago Agents & autonomy

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association b
HuggingFace Daily Papers 4d ago Agents & autonomy

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reprod
HuggingFace Daily Papers 4d ago RegulationSafety & alignment

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four
HuggingFace Daily Papers 4d ago Agents & autonomy

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across scanners, protocols, and patient populations. High-performing models consequently require repeated domain-specific fine-tuning, which is a costly cycle that becomes impractical when labels are scarce or privacy constraints limit data sharing. We propose OPERA (Offline Policy-guided Expert Routing and Adaptation), a multi-agent ensemble framework that addresses
HuggingFace Daily Papers 4d ago RegulationPrivacy

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file s
HuggingFace Daily Papers 4d ago Agents & autonomy

SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing

Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same heterogeneity makes target discovery highly sensitive to analytical choices throughout the pipeline, including preprocessing, cell population selection, differential expression analysis, and downstream biological interpretation. As a result, existi
arXiv cs.LG 4d ago Agents & autonomyBiotech

Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation

In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (CTTA) techniques leveraging a teacher-student framework have gained prominence, allowing models to adapt continuously even after deployment. In such a framework, a weight-averaged mean teacher is used to produce pseudo-labels from test data for self-training. The mean teacher gets updated as an exponential moving average of the student parameters us
arXiv cs.LG 4d ago Children & education

Outcome-Confounded Local Supervision in On-Policy Distillation

On-policy distillation (OPD) trains a student on its own trajectories while a teacher supplies dense token-level likelihoods at student-visited prefixes. These likelihoods are often read locally: agreement appears safe to imitate, whereas disagreement appears to identify an error. We show that both readings are confounded by the outcome of the completed trajectory. We introduce an outcome-resolved diagnostic that crosses pointwise teacher-student divergence with final-answer correctness, separat
arXiv cs.LG 4d ago RegulationHealthcare

DP-IVON-Gradsq: Differentially Private Squared-Gradient Improved Variational Online Newton

Differential privacy provides formal privacy guarantees for training neural networks on sensitive data, while Bayesian deep learning offers a principled framework for uncertainty-aware prediction. Combining these two objectives remains challenging, as privacy noise can interact with the stochasticity introduced by Bayesian posterior sampling. In this work, we investigate differentially private variational Bayesian learning through the Improved Variational Online Newton (IVON) optimizer. We intro
arXiv cs.LG 4d ago PrivacyFinance, VC & PE

Optimal Reward Shaping: Autonomous Car Parking Case Study

Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we
arXiv cs.LG 4d ago RegulationSafety & alignment

Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter

Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided rei
arXiv cs.LG 4d ago Environment

Social representations of GenAI and paradoxical tensions in its adoption in higher education

In just a few years, generative artificial intelligence (GenAI or GAI) has distinctly impacted numerous facets of human activity, particularly in how information is processed and transformed into knowledge. As centers for knowledge (co)creation and appropriation, universities face the opportunity and challenge of navigating a de facto integration of this promising yet disruptive technology. This study examines how instructors and students form their common-sense understanding of GenAI and endors
AI & Society 4d ago Children & education

The existential choice facing UK physics facilities: commercialize or close

The national synchrotron source, a laser facility and a particle accelerator are under threat of closure unless cash is found.
Nature Machine Intelligence 4d ago Safety & alignment

GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data

This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation support
arXiv cs.CL (ethics-relevant NLP) 4d ago

Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program "Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse
arXiv cs.CL (ethics-relevant NLP) 4d ago HealthcareChildren & education

Auditing Alignment Controllability in LLMs via Political Axes

Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability a
arXiv cs.CL (ethics-relevant NLP) 4d ago Safety & alignmentTransparency

Do LLMs Know Their Vulnerable Scenarios?

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relation
arXiv red teaming query 4d ago Safety & alignment