Archive · 2026-07-26
AI ethics on Sunday, 26 July 2026
102 items published this day, across 5 categories.
Incidents (23)
Fact Check: Is Story Of Missing Boy Scout Eric Langford Real?
Images and videos of Eric Langford have been shared online---but multiple people have said in the comments that they suspect these are all AI-generated. Newsweek has broken down whether the story is true or not. The Claim Langford vanish ... (https://incidentdatabase.ai/cite/1330#7578)
Terrifying truth about what really happened to Boy Scout who vanished
A 14-year-old boy mysteriously vanished without a trace from a Boy Scout camp, never to be seen again - his family devastated. Three weeks later, the teen, Eric Langford, was declared dead after baffled police failed to find any leads desp ... (https://incidentdatabase.ai/cite/1330#7579)
National Weather Service Uses AI to Generate Forecasts, Accidentally Hallucinates Town With Dirty Joke Name
Months before it mysteriously vaporized into thin air, Elon Musk's so-called Department of Government Efficiency ravaged the National Weather Service, leading to severe staffing shortages. While the Trump administration promised to rehire ... (https://incidentdatabase.ai/cite/1332#7580)
An AI-Generated NWS Map Hallucinated Fake Towns in Idaho
The National Weather Service's weekend wind forecast for rural Idaho looked normal. Some mild gusts and a caption reminding folks to "hold onto your hats." And then some eagle-eyed observers took a look at the towns on the map and wondered ... (https://incidentdatabase.ai/cite/1332#7581)
New Zealand book award disqualifies two authors for AI artwork
The New Zealand Book Awards Trust introduced a new rule this year stating that any book submitted for the Ockham awards must not contain AI material of any kind. "Don't judge a book by its cover" goes the famous idiom, one which may be tou ... (https://incidentdatabase.ai/cite/1282#7582)
Authors dumped from New Zealand’s top book prize after AI used in cover designs
The books of two award-winning New Zealand authors have been disqualified from consideration for the country’s top literature prize because artificial intelligence was used in the creation of their cover designs. Stephanie Johnson’s collec ... (https://incidentdatabase.ai/cite/1282#7583)
New Zealand’s top book award disqualifies two prominent authors over AI artwork
Two of New Zealand's best-known writers have been disqualified from consideration for the country's premier literary prize after organisers discovered their book jackets used images generated with artificial intelligence. Stephanie Johnson ... (https://incidentdatabase.ai/cite/1282#7584)
How two AI book covers forced the publishing industry to reckon with its future
Elizabeth Smither and Stephanie Johnson's latest books, both published by Quentin Wilson Publishing, have been yeeted from the Ockhams. Author Stephanie Johnson had no idea that the image of a cat with human teeth on the cover of her book ... (https://incidentdatabase.ai/cite/1282#7585)
GEMA vs. OpenAI | AI memorisation is a reproduction relevant to copyright law, and the TDM exception does not help in LLM training, Munich I Regional Court holds
In its judgment of 11 November 2025 (42 O 14139/24), the Munich I Regional Court (Germany) issued a widely noted precedent on the copyright assessment of AI training and outputs under German and EU law. According to its press release1, the ... (https://incidentdatabase.ai/cite/1278#7586)
OpenAI ordered to pay damages as court rules ChatGPT violated copyright law
OpenAI has been ordered to pay damages after a court in Germany ruled that its chatbot ChatGPT violated German copyright laws. The Munich regional court ruled in favour of the German music performance rights organization GEMA which manages ... (https://incidentdatabase.ai/cite/1278#7587)
AI-powered children’s toys are here, but are they safe?
Teddy bears and stuffed plushies have long been a mainstay in toy collections. But today they don't talk back in a child's imagination --- some talk through built-in AI chatbots. Sometimes that's a problem, though:A scarf-wearing teddy be ... (https://incidentdatabase.ai/cite/1277#7588)
AI Toys Are Here, And Their Safety Is Questionable—Here's What Parents Should Know
As a writer who covers baby and kids gear, I'm inundated with emails about the hottest new toys: a box that automatically prints pictures based on kids' commands. A robot with facial recognition capabilities. A teddy bear that can craft end ... (https://incidentdatabase.ai/cite/1277#7589)
AI toys for kids talk about sex and issue Chinese Communist Party talking points, tests show
A wave of AI-powered children's toys has hit shelves this holiday season, claiming to rely on sophisticated chatbots to animate interactive robots and stuffed animals that can converse with kids. Children have been conversing with stuffies ... (https://incidentdatabase.ai/cite/1277#7590)
AI toy dangers abound and parents should be vigilant, new report warns
Amid a flood of toys boasting AI features this holiday season, a consumer advocacy group warns that more must be done to ensure the gadgets are safe for children. Public Interest Research Group, which pushes for corporations and government ... (https://incidentdatabase.ai/cite/1277#7591)
Cybercriminals unleash fake Centrelink scam on vulnerable Australians
More than 270,000 malicious emails impersonating Services Australia and Centrelink have flooded Australian inboxes in one of the nation's largest phishing campaigns in years, with the sophisticated attacks specifically targeting the country ... (https://incidentdatabase.ai/cite/1275#7592)
Services Australia Impersonation Drives Year-Round Credential Theft Operation
Key Points MCTO3001 - Threat operation with Services Australia and Centrelink impersonation campaigns across multiple sectors Infrastructure abuse of legitimate email services (SendGrid, Mailgun, Office 365) with Australian Gov display na ... (https://incidentdatabase.ai/cite/1275#7593)
Centrelink warning as 270,000 emails sent out in attack related to Medicare, superannuation and tax benefits
Australians are being bombarded with tens of thousands of fake emails impersonating Centrelink and Services Australia in one of the biggest phishing campaigns in years. Cybercriminals are using artificial intelligence to create "super clone ... (https://incidentdatabase.ai/cite/1275#7594)
Taco Bell Rethinks Future of Voice AI at the Drive-Through
The most transformative technology in over a century may have finally found its limit: ordering tacos. Since last year, Taco Bell has rolled out voice AI-powered ordering at more than 500 drive-through locations, and now the chain is reali ... (https://incidentdatabase.ai/cite/1274#7595)
Greek Finance Minister Sues Facebook Page for Deepfake Advert
The Greek Ministry of Economy and Finance and minister Kyriakos Pierrakakis filed a lawsuit on Friday for deceptive advertising on Facebook against the unknown administrators of a page on the social media platform. The ministry and Pierrak ... (https://incidentdatabase.ai/cite/1271#7596)
Louvre jewel heist: fake AI videos of the robbery circulate online • FRANCE 24 English
AIID editor's note: Please visit the original source for the video report. Since the robbery of the world’s most-visited museum—which saw thieves make off with France’s priceless crown jewels—multiple videos have surfaced on social media c ... (https://incidentdatabase.ai/cite/1273#7597)
Top universities react to AI cheating scandals, yet concrete disciplinary steps remain elusive
As cheating cases involving artificial intelligence (AI) continue to surface at Korea's elite "SKY" universities --- Seoul National, Korea and Yonsei --- campuses are rushing to roll out countermeasures. Most responses focus on the individ ... (https://incidentdatabase.ai/cite/1270#7598)
Cheating scandals at top universities prompt rethink of education in digital era
A string of artificial intelligence (AI)-related cheating scandals at Korea's prestigious "SKY" universities --- Seoul National University, Yonsei University and Korea University --- has sparked renewed scrutiny of higher education in the d ... (https://incidentdatabase.ai/cite/1270#7599)
Waymo runs into safety concerns and competition as it expands in the US
San Francisco, the United States: The sidewalk outside Majed Zeidan's grocery store in San Francisco's Mission District has stayed filled with flowers, candles, memorials and pictures since his cat was crushed under a Waymo in late October. ... (https://incidentdatabase.ai/cite/1269#7600)
News (47)
When the AI bubble bursts, what will Australia do with the tools it built? One man thinks he has the answer
Journalist and author Cory Doctorow is coming to Australia to spread his message – humans will take their jobs back Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast When the AI bubble bursts, executives who enthusiastically replaced their staff with AI tools will learn that it takes a “really long time” to replace those lost skills, says Cory Doctorow, a science fiction author and journalist. But those in creative fields should a
A Roadmap for Confronting the Chilling Effects of Censorship, Surveillance and New Technology
Despots, DSA Fines, and an 'Autonomous' AI Attack
How the OpenAI-Hugging Face Hack May Affect the Geopolitics of AI Governance
Corporate America may be using AI to cut jobs, but small businesses are using it to keep them
Reports of wide-scale replacement of workers by AI are overblown. Small businesses use it to help workers I recently met the owner of a company that sells windows and doors. He told me he invested about $10,000 in an AI application that is used by his salespeople in his showroom. The application listens to the conversations between the salesperson and the prospective customer and then automatically creates a quote for the salesperson to review and send. “It allows my salespeople to talk to more
Optical Tech Would Update a Robot’s AI on the Fly
Atop a lab bench, Cornell Tech postdoctoral researcher Yifan He positions the lens of an optical receiver almost a meter away from an LED emitting a beam of red light. The computer monitor attached to the receiver takes a beat to refresh, then displays an array of squares that resemble a QR code. When you hold your phone camera up to a QR code, light strikes the image sensor as only a first step to revealing the data hidden behind the black and white matrix. The receiver here is doing something
Making sense of the panic over Chinese AI
On the latest episode of Equity, we discussed why Moonshot AI's Kimi seemed to panic Silicon Valley and Wall Street.
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"
Partner-Associate Park Passion Predicament — See Generally
What's The Billing Code For This : A midday park-bench makeout session between a Biglaw partner and associate ends up going viral. $1.2 Billion Worth Of 'Are You Sure This Is Legal?' : MV Realty has sued Holland & Knight claiming it sank $156 million into a unique financial product blessed by the firm and 16 state attorneys general came calling . John "Torture Memo" Yoo Invents New Crime To Threaten Administration Critics : John Yoo went on Fox News suggesting the Trump DOJ investigate Zohran Ma
Wildfire Disruptions to the Tour de France Show How Climate Change Is Shaping Sports
Tadej Pogačar's fifth Tour de France victory came after last-minute course changes due to wildfires, underscoring how climate change is influencing major sporting events.
Finance One backs AI and automation to move faster
Podcast: Implements inaugural AI strategy outline and framework.
Its AI agent spent days hacking a company, but sources say OpenAI did not notice for a week
Until well after the threat was contained.
Virginia Study on Groundwater and Data Centers Calls for Tighter Water Regulations
The study, released after months of delays and mounting criticism, was ordered two years ago amid growing concerns about resources for data centers.
Traders are getting a new tool to wager on the biggest U.S. stocks
CME Group Inc. will launch single-stock futures Monday, allowing investors to hedge or speculate on more than 50 of the largest US companies.
James Cameron tried to warn us: ‘Skynet Day’ is now shorthand for OpenAI’s agent going rogue and hacking into a startup
In what OpenAI said was the first-ever incident of its kind, an advanced AI model escaped its “sandbox” to the internet and used stolen credentials to break into the servers of Hugging Face.
20 Years Ago, a Feature Documentary Announced the Death of the Electric Car
Long before the Tesla revolution, Sony Pictures Classics' ‘Who Killed the Electric Car?’ investigated why GM discontinued the beloved EV1.
Ukraine’s attack on an Iranian ship in the Caspian Sea could signal a new battlefront as U.S. seeks to form ‘a ring of economic and military pressure’
"In any case, what is clear is that Iran’s war is becoming increasingly intertwined with two other conflicts: the Saudi-Houthi confrontation and the Russia-Ukraine war."
Apple's smart glasses delay reportedly stems in part from major privacy concerns
The Apple team behind the upcoming smart glasses is reportedly exploring several tweaks related to user privacy.
European start-up with SpaceX ambitions aims for a $2bn valuation
The Exploration Company in talks to raise $300mn for reusable space capsules
El MIT ha diseñado un robot que vuela y bucea como un ave marina, pero sin patas: el truco está en el ángulo de despegue
Cuando hablamos de drones pensamos automáticamente en vehículos voladores, pero también existen drones submarinos . Lo que no es tan habitual es que un mismo dron sea capaz de hacer las dos cosas. Es justo lo que acaba de conseguir el MIT : un pequeño robot con alas capaz de volar y sumergirse en el agua , como si fuera una ave marina cuando caza. El robot, el cual han bautizado con el nada sugerente nombre de "vehículo aéreo-acuático de alas batientes" (FAAV, por sus siglas en inglés), ha sido
Head Start, Explained. How Changes to Early Education Program Affect California Families
Head Start, the federal program providing free early education for low-income children, has benefited from strong bipartisan support since its 1965 inception, withstanding funding cuts and political turbulence across 11 presidencies. But more recent threats from the second Trump administration have thrust Head Start into the spotlight. Last year, its programs faced disruptions and threats […]
Jeff Bezos says stress is a signal to work harder, not a reason to stop
Jeff Bezos, the founder of the company that just became the number one Fortune 500 business, says the cure for anxiety is not rest or meditation but action. Speaking at the Museum of Flight in Seattle in 2017, Bezos told an audience of students that stress is a signal from the body that something is […] This story continues at The Next Web
Build the Join Key Before the Ledger
Somewhere inside the agentic payments architecture currently being standardized there will be an aud...
Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work
Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the new system, which separates planners from workers, eventually scored 100 percent on the test suite. The old swarm choked on merge conflicts of its own making. The article Cursor's agent swarm suggests cheaper models can handle most coding when frontier models plan the work appeared first on The Decoder .
What is the risk of using Chinese open AI models like Kimi K3?
The real problem is not overseas open-source but lack of co-ordination to protect infrastructure in the face of cyber attacks
5 ways SRE AI agents are set to augment human capabilities
In digital operations management, AI agents give organizations a competitive edge by reducing incident volume and accelerating recovery. The potential The post 5 ways SRE AI agents are set to augment human capabilities appeared first on The New Stack .
Low profile, high AI ambition: what leaked comments reveal about DeepSeek’s Liang Wenfeng
In an era dominated by aggressive tech founders who chase billion-dollar valuations and maximal profits while curating loud public profiles, Liang Wenfeng stands out for his insistence on staying in the background. With only a couple of photographs of him circulating online, the founder of Chinese artificial intelligence start-up DeepSeek and quantitative hedge fund High-Flyer Quant has long been a reclusive figure. But last week, a leaked transcript of a closed-door meeting with potential...
The U.S. government invested $27 billion in corporate stakes. Good luck finding them
The Trump administration’s equity stakes—from Intel to quantum startups—appear in no budget document and are subject to no watchdog.
El 'prompt injection' ya tiene su propio contraataque: inyectar falsas instrucciones también a los hackers
Durante los últimos dos años, el prompt injection ha sido el arma favorita de los ciberatacantes contra sistemas de inteligencia artificial. Este método pasa por esconder una instrucción maliciosa dentro de un correo, una invitación de calendario o una página web para que un agente de IA acabe obedeciendo al intruso en lugar de a su usuario legítimo. Lo bueno es que esta misma técnica también sirve para pararle los pies a los atacantes. Qué ha pasado. &
Why TikTok’s Algorithm Keeps You Trapped in a Breakup Loop
As social media companies face scrutiny over addictive design features, mental health experts cite a stream of breakup content as an example of the failure to protect users.
Texas politicians call for guardrails on AI, data centers
Also, PUC prepares for data centers and the Texas Medical Association calls for legislative action on prediction market apps.
heise-Angebot: Claude Code in der Praxis – eigenen KI-Chat-Agenten in fünf Sessions entwickeln
Claude Code produktiv einsetzen und durch agentische Entwicklung mit Mastra, Tool Calls, MCP, Frontend und AG-UI-Protokoll zum eigenen Chat-Agenten.
As Diesel Prices Soared, Arkansas District Saved With Electric School Buses
Five years ago, diesel costs averaged $2.68 a gallon. This spring, that more than doubled to over $5.50. That’s not just a headline; it’s a real problem for anyone operating on a tight budget. But for school districts, the issue is even more severe. The cost of fuel for just one school bus runs into […]
The Pentagon wants to build data centers. Congress would like a word.
Lawmakers are setting their sights on data centers proposed for military bases, concerned about massive energy and water consumption.
For some, so-called ‘Skynet Day’ came too close to sci-fi after a rogue agent hacked into a startup
In 1984, the Skynet of "The Terminator" films were science fiction. But it looks more and more realistic in 2026 after OpenAI agent broke out of a test corral, traveled the internet and hacked into ...
As Trump boosts nuclear power, regulators seek to eliminate a longstanding radiation safety practice
The Nuclear Regulatory Commission is proposing to eliminate a foundational safety principle that has for 50 years minimized the radiation people in the United States are exposed to, and that has been ...
DeepSeek pauses its second fundraising round after founder’s leaked investor comments go viral
DeepSeek has suspended its second fundraising round after comments made by founder Liang Wenfeng during private investor meetings were leaked and went viral on Chinese social media, Bloomberg reported on Friday. Liang verbally informed prospective backers to hold off on the round, which had been targeting a pre-money valuation of roughly 480 billion yuan, equivalent […] This story continues at The Next Web
Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides
In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder .
US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns
The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th
The Pete behind Jersey Mike’s was only 17 years old when he borrowed $125,000 to buy the chain—with some help from his high school football coach
Peter Cancro had worked at the Jersey Shore sub shop since he was 14; the six-figure loan would be worth $775,000 today.
Breakpoint: Roboter-Republik
Wer spricht? – Bundestag: Tim Hüfner , Spielzeugroboter: Emilipothèse , Mikrofon: Jon Tyson , Bearbeitung: netzpolitik.org Wenn Politiker:innen nicht mehr selbst reden, sondern Maschinen für sich sprechen lassen, mangelt es ihrem Auftritt an Authentizität. Doch gerade die macht einen Teil ihrer Legitimität aus. Denn in einer Demokratie sollten wir von Menschen repräsentiert werden, nicht von Robotern.
Wall St Week Ahead US stocks face tests from Fed decision, tech-led earnings deluge
A wobbly U.S. stock market will take its cues in the coming week from a Federal Reserve meeting set to shed light on the path for interest rates, and from a packed slate of corporate earnings led by ...
Anzeige: IT-Jobs in Support, Administration und Entwicklung
Von Microsoft-365-Support bis zur Teamleitung in der Azure-Cloud: Diese sechs IT-Jobs bieten vielseitige Aufgaben zu guten Konditionen. ( Golem Karrierewelt , Unternehmenssoftware )
Defence giants provide record backing for military start-ups
As drones and autonomous systems transform the battlefield, traditional defence companies start to act more like venture capitalists
Why frontier labs are quietly agreeing on fintech's memory problem
Anthropic's engineering team recently published a notable admission: they stripped more than 80% of...
‘Black Panther 3’: David Jonsson to Play T’Challa’s Son as December 2028 Release Date Set
Black Panther is back! David Jonsson has been cast as the new Black Panther and is playing T’Challa’s son in “Black Panther 3,” out Dec. 15, 2028. The film will follow Jonsson’s character, Prince T’Challa II, as he comes of age after he was introduced as a child in “Black Panther: Wakanda Forever.” Director Ryan […]
Stephen King’s ‘Carrie’ Series Sets October Release Date, Drops Horrifying New Trailer
Amazon Prime Video has released a new teaser trailer for “Carrie,” the TV adaptation of Stephen King’s classic horror novel. All eight episodes of the series will debut on Oct. 7. The series is a “bold and timely reimagining of the story of misfit high-schooler Carrie White, who has spent her life in seclusion with […]
Field notes (4)
Nvidia’s China Partners and the PLA
CSET’s Sam Bresnick shared his expert insight in an article published by The Wire China. The article examines how Nvidia’s network of partners in China has supplied organizations linked to China’s military and other U.S.-restricted entities, highlighting ongoing challenges in enforcing export controls on advanced AI technology. The post Nvidia’s China Partners and the PLA appeared first on Center for Security and Emerging Technology .
An OpenAI model left notes about how to evade containment
We need more details
An Inside Look at the Relay Market Powering Token Resellers and Fraud
An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or ch
Amazon is investing in the Lean Focused Research Organization
As AI agents take on higher-stakes decisions, Lean programming language makes it possible to mathematically prove they will behave safely.
Policy (2)
IMF Executive Board Concludes 2026 Article IV Consultation with Kingdom of The Netherlands—The Netherlands
Entering 2026 from a position of relative strength, the energy price shock from the war in the Middle East is projected to slow growth to about 1 percent and lift inflation to around 3 percent, amid ...
All UNESCO news on education
We use cookies on this site to enhance your user experience. For more information on how we use cookies, read our privacy notice.
Research (26)
A Coulomb Particle Model for Learning Kernel Attention in Transformers
Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature distribution. We propose a particle-based method that learns this distribution by optimizing kernel-target alignment while regularizing particles with a Riesz/Coulomb repulsive potential. The resulting Hamiltonian yields diverse, task-adaptive random features and admits a mean-field description through a McKean--Vlasov equation. We instantiate the method in lin
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning abilit
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investiga
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negat
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they as
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical un
Data Pyramid for Embodied Manipulation
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-lan
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association b
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reprod
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four
OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across scanners, protocols, and patient populations. High-performing models consequently require repeated domain-specific fine-tuning, which is a costly cycle that becomes impractical when labels are scarce or privacy constraints limit data sharing. We propose OPERA (Offline Policy-guided Expert Routing and Adaptation), a multi-agent ensemble framework that addresses
Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent Retrieval Bench, a file-level benchmark for this upstream retrieval problem. Samples are built from real coding-workflow signals and evaluated against frozen base-commit repositories, with relevance defined by what an agent needs next rather than direct query-file s
SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing
Identifying therapeutic target genes from single-cell RNA sequencing (scRNA-seq) data remains a fundamental challenge in translational biology. Unlike bulk assays, scRNA-seq captures heterogeneous cellular states and rare subpopulations, but this same heterogeneity makes target discovery highly sensitive to analytical choices throughout the pipeline, including preprocessing, cell population selection, differential expression analysis, and downstream biological interpretation. As a result, existi
Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation
In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (CTTA) techniques leveraging a teacher-student framework have gained prominence, allowing models to adapt continuously even after deployment. In such a framework, a weight-averaged mean teacher is used to produce pseudo-labels from test data for self-training. The mean teacher gets updated as an exponential moving average of the student parameters us
Outcome-Confounded Local Supervision in On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories while a teacher supplies dense token-level likelihoods at student-visited prefixes. These likelihoods are often read locally: agreement appears safe to imitate, whereas disagreement appears to identify an error. We show that both readings are confounded by the outcome of the completed trajectory. We introduce an outcome-resolved diagnostic that crosses pointwise teacher-student divergence with final-answer correctness, separat
DP-IVON-Gradsq: Differentially Private Squared-Gradient Improved Variational Online Newton
Differential privacy provides formal privacy guarantees for training neural networks on sensitive data, while Bayesian deep learning offers a principled framework for uncertainty-aware prediction. Combining these two objectives remains challenging, as privacy noise can interact with the stochasticity introduced by Bayesian posterior sampling. In this work, we investigate differentially private variational Bayesian learning through the Improved Variational Online Newton (IVON) optimizer. We intro
Optimal Reward Shaping: Autonomous Car Parking Case Study
Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we
Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision. To address this, we propose an anticipatory risk-guided rei
Social representations of GenAI and paradoxical tensions in its adoption in higher education
In just a few years, generative artificial intelligence (GenAI or GAI) has distinctly impacted numerous facets of human activity, particularly in how information is processed and transformed into knowledge. As centers for knowledge (co)creation and appropriation, universities face the opportunity and challenge of navigating a de facto integration of this promising yet disruptive technology. This study examines how instructors and students form their common-sense understanding of GenAI and endors
The existential choice facing UK physics facilities: commercialize or close
The national synchrotron source, a laser facility and a particle accelerator are under threat of closure unless cash is found.
GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data
This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other in counsellor strategies and client emotions. The gap is detectable but small. A generative validation support
Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration
Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program "Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse
Auditing Alignment Controllability in LLMs via Political Axes
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability a
Do LLMs Know Their Vulnerable Scenarios?
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear. Meanwhile, mechanistic interpretability studies have characterized both refusal directions and jailbreak-associated features, without explaining the relation