Archive · 2026-08-13

AI ethics on Thursday, 13 August 2026

363 items published this day, across 5 categories.

Incidents (24)

AI Incident Database

AI-generated ‘Baling-Pengkalan Hulu cable car’ dupes KL couple into 300km trip, officials explain viral news clip is fake — open the original publisher

BALING, July 4 --- Authorities have dismissed a viral video claiming the existence of a cable car linking Pengkalan Hulu in Perak to Baling in Kedah, saying it was generated using artificial intelligence (AI). Sinar Harian quoted Baling Di ... (https://incidentdatabase.ai/cite/1634#7705)

AI Incident Database

Crece el número de turistas que planifican sus viajes con ChatGPT y también los destinos inventados por la IA — open the original publisher

¿Has utilizado alguna vez ChatGPT para planificar tus vacaciones? Lo más probable es que sí y también es posible que la inteligencia artificial te haya dicho que vayas a un lugar presumiblemente turístico que no lo es... y, puede que ¡ni ex ... (https://incidentdatabase.ai/cite/1636#7708)

News (200)

The Guardian

Unemployed young people to join AI boot camps to get job-ready — open the original publisher

Pilot scheme will provide three weeks of training as part of UK government’s latest attempt to address Neets crisis Young people out of work or at risk of unemployment in the UK are to join “AI boot camps” where they harness the technology to get a foothold in the workplace. The government’s latest attempt to address the crisis in Neets – young people not in work or education – involves turning to a technology that many view as a potential threat to employment. Continue reading...

Jobs & economyChildren & education
The Guardian

Massachusetts teen accused of killing mother and brother used ChatGPT — open the original publisher

District attorney says Arjun Aravind, 17, used internet and AI to search for fantasy stories regarding killing of his family A Massachusetts teenager accused of killing his mother and younger brother is being held without bail as authorities investigate a double-murder case that prosecutors say is connected to his use of ChatGPT. Arjun Aravind, 17, appeared Thursday morning for his arraignment in Concord district court, where a not-guilty plea was entered on his behalf to murder and several addi

Finance, VC & PE
The Guardian

An AI agent for all? Try using your brain, Mark Zuckerberg | Brief letters — open the original publisher

Meta AI | Future Guardian writers and Neets | Food for thought | Plants surviving the heat | Living and dying well Your article ( Zuckerberg pushes ‘superintelligent’ AI for all as Meta releases open-weight model, 10 August ) quotes Mark Zuckerberg as saying: “Everyone will have an exceptionally capable personal agent that understands you, your goals, and everything you care about. You’ll be able to interact with your agent through any device, including your glasses.” He is wasting his time

Agents & autonomy
MIT Technology Review

Flock is tightening its rules in response to a growing surveillance backlash — open the original publisher

The police-tech giant Flock is announcing today that it will change officers’ access to its nationwide network of license plate readers, in an apparent effort to quell a growing backlash and win back contracts lost amid concerns about mass surveillance and police abuse. Several changes aim directly at a problem that has made recent headlines:…

Privacy
The Guardian

Lost jobs, inequality, rogue agents: why are we accepting oligarchs’ AI agenda? | Robert Reich — open the original publisher

The dangers of AI become clearer every day. Why are we still acting as if we have no choice about our future? Rather than producing jobs, the US economy actually lost 23,000 jobs in July, according to Bureau of Labor Statistics data released on Friday. In addition, May and June’s job numbers were revised downward, showing a combined 103,000 fewer jobs than previously reported. As if this weren’t bad enough, wage growth has also slowed. Average hourly earnings rose by just 0.1% from June. Continu

Jobs & economyAgents & autonomy
The Guardian

Taiwan says it was hit by ‘abnormal’ AI-assisted cyber-attack — open the original publisher

Taiwan’s statement comes a day after reports that suspected China-linked hackers had carried out a first-of-a-kind breach Taiwan says it detected AI-assisted cyber-attacks on government agencies that came from overseas last month, a new kind of threat that has been reported as “first-of-a-kind breach”. The Ministry of Digital Affairs (MDA) said its cybersecurity monitoring units ‌detected the “abnormal attack” targeting government agencies, which began on 20 July. The National Institute of Cyber

Military & security
Above the Law (legal tech)

XOXO, Gossip Girl: The Yale Law Whisper Network Dishes On Usha Vance — See Also — open the original publisher

Everything You Need To Know About Usha Vance, You Can Learn At Yale Law: Her classmates gossip about her on Signal. The clerkship culture that shaped her politics is the part they don't put in the group chat. Daft For Taft : John Roberts pens essay celebrating William Howard Taft... but mostly trying to sugarcoat his own legacy . The Top Schools For Tech & The Law : Check out the Honor Roll here. Todd Blanche's First Message To The DOJ: Trust me. Slip And Fall Hopscotch : Law firm decorated side

RegulationChildren & education
The Hill Technology

Top Commerce Committee Democrat presses airlines over AI ‘surveillance pricing’ — open the original publisher

A top Democrat on the House Energy and Commerce Committee is pressing major U.S. airlines over whether they use artificial intelligence to set ticket prices based on travelers’ personal information, raising concerns that it determines what fares consumers see. Rep. Frank Pallone Jr. (D-N.J.), the ranking member of the House Energy and Commerce Committee, sent...

PrivacyEnvironment
The Next Web AI

Databricks Closes $5 Billion Round at $190 Billion Valuation — open the original publisher

Databricks has closed a $5bn round at a $190bn valuation, led by Coatue, with revenue run-rate past $7bn and growth above 80% year on year. That is a 42% valuation increase in six months, and it comes after chief executive Ali Ghodsi called 2026 a bad year to go public. Databricks has closed $5bn at […] This story continues at The Next Web

Finance, VC & PE
The Hill Technology

Trump memo allows private sector to aid cyber offense against transnational criminal groups — open the original publisher

President Trump is laying the groundwork for private sector firms to take a larger role in the U.S.'s cybersecurity offense against transnational cyber crimes in what could be one of the largest ever changes in U.S. cyber policy. In a memo signed Wednesday, Trump called on the National Coordination Center to create a program for...

RegulationMilitary & security
The Decoder

Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50% — open the original publisher

Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the company's own benchmarks, it beats Claude Sonnet 5 and GPT-5.6 Terra at half the price. The article Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50% appeared first on The Decoder .

Agents & autonomy
The 74 (education AI)

Gap to Adequately Fund Schools Widens in Illinois — open the original publisher

Chicago Public Schools is further away from having enough money to meet its students’ needs as defined by the state, according to new data released by the Illinois State Board of Education. Under ISBE’s calculations released Friday, CPS is now considered a little over 70% adequately funded, compared with 73% last year. CPS isn’t alone […]

Children & education
War on the Rocks

The Commercial Space Race — open the original publisher

A strong commercial space industry is an important partner for the U.S. government, as it contributes to building more robust space and defense capabilities and facilitates innovation more broadly. As competition between the United States and China heats up, both countries look to the commercial space sector to help them get ahead in both the space race and defense technologies. We asked five experts: What is a crucial step for the United States to take now to stay ahead in commercial space comp

Military & security
The Next Web AI

GM and LG restart Ohio battery production after seven idle months — open the original publisher

Ultium Cells, the GM and LG Energy Solution joint venture, restarts EV battery cell production in Warren, Ohio next week after a seven-month shutdown, with about 1,400 workers. US EV sales rose 14.2% in the second quarter but remain 20.5% below a year earlier. The battery plant that GM and LG idled in January is […] This story continues at The Next Web

Environment
Above the Law (legal tech)

Who Is Justice Barrett? — open the original publisher

She has been labeled conservative, a swing justice, and sometimes unpredictable. This article puts all of these hypotheses to the test and provides a data based assessment of where she really stands. The post Who Is Justice Barrett? appeared first on Above the Law .

Regulation
Above the Law (legal tech)

Relativity Announces claiR, A Conversational AI For Lawyers, But You’ll Have To Wait Awhile To Chat With It — open the original publisher

Relativity CEO Phil Saunders calls it a fundamentally new way for lawyers to get straight to the answers in their most consequential legal data. The post Relativity Announces claiR, A Conversational AI For Lawyers, But You’ll Have To Wait Awhile To Chat With It appeared first on Above the Law .

Regulation
The Next Web AI

Anthropic’s $2trn IPO is priced below what AI stocks already fetch — open the original publisher

The number comes from investors, not the company. Six backers told the Financial Times that rising revenue would let the five-year-old lab more than double its valuation in an autumn listing. At $2 trillion it would eclipse SpaceX, which went public at $1.77 trillion in June. Anthropic itself has fixed nothing. Several investors said senior […] This story continues at The Next Web

Finance, VC & PE
The 74 (education AI)

Opinion: Congress Made it Easier to Pay for College. It Must Keep That Promise — open the original publisher

Right now, 15.5 million undergraduate students around the country are getting ready to start or return to college. Seven million of them have one thing in common: relying on a Pell Grant to help pay for their higher education. But the grant program is facing a $15 billion funding shortfall that will undermine college access […]

Children & education
The Decoder

Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices — open the original publisher

Deepseek has moved its flagship V4-Pro out of the testing phase and released its agent software, Harness v0.1, under the MIT license. API prices are going up at the same time, with cache hits jumping to six times their current cost. For agent workflows that repeatedly read the same files, that's the biggest price increase in the transition. The article Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices appeared first on The Decoder .

Agents & autonomy
Fast Company Tech

Chatbots argue against election conspiracy theories, then willingly illustrate them — open the original publisher

Welcome to AI Decoded , Fast Company ’s weekly newsletter that breaks down the most important news in the world of AI. I’m Mark Sullivan, a senior writer at Fast Company , covering emerging tech, AI, and tech policy. Sign up to receive this newsletter every week via email here . And if you have comments on this issue and/or ideas for future ones, drop me a line at sullivan@fastcompany.com , and follow me on X @thesullivan . As midterms loom, AI chatbots fight election lies, but then supply image

RegulationMisinformation
Legal IT Insider

amicable applies to SRA to launch “tech-forward” law firm — open the original publisher

amicable, the award-winning divorce and separation service, has submitted its application to the Solicitors Regulation Authority (SRA) to launch a new, separate law firm serving England and Wales. Alongside that, it’s opening the search […] The post amicable applies to SRA to launch “tech-forward” law firm appeared first on Legal IT Insider .

Regulation
Legal IT Insider

Exclusive: Elite Cloud customers to outnumber on prem for first time in company’s history — open the original publisher

Elite’s cloud customers will outnumber its on-premises customers by the end of 2026, with SaaS users growing 60% YoY to 53,000, and over 120 law firms now live in the […] The post Exclusive: Elite Cloud customers to outnumber on prem for first time in company’s history appeared first on Legal IT Insider .

Regulation
The Hill Technology

Flock will add privacy guardrails to camera systems amid backlash — open the original publisher

Flock Safety will implement new privacy and data retention policies as pressure grows on the automated license plate operator to prevent its technology from being abused or used to conduct mass surveillance. The company announced Thursday it will shorten the recommended default data retention window from 30 to seven days to cut the amount of...

Privacy
Xataka (ES)

Grok 4.6 necesita la mitad de pasos que Claude Opus 5 para hacer el mismo trabajo. Y eso es más importante que el benchmark — open the original publisher

Ejecutar un agente de IA durante horas tiene una factura no siempre evidente: cada vuelta que da el modelo para pensar, comprobar o corregirse, se paga. Y ahí es donde Grok 4.6 , el modelo que acaba de lanzar SpaceXAI (el nuevo nombre de la empresa desde el mes pasado ), dice tener su gran ventaja. En la prueba AA-Briefcase completó sus tareas en 53 turnos consumiendo 500 millones de tokens de media. Claude Opus 5 Max, en cambio, necesitó 103 turnos  y 2.000 millones de turnos para llegar a

Agents & autonomy
The 74 (education AI)

Indiana Literacy Rates Improve for Fifth Straight Year — open the original publisher

Indiana’s early literacy rates improved for the fifth consecutive year, with nearly 89% of Hoosier third graders demonstrating proficiency in foundational reading skills on the state’s IREAD assessment. The Indiana Department of Education released statewide IREAD, ILEARN and SAT results for the 2025-26 school year Tuesday at the State Board of Education’s August meeting. The […]

Children & education
Xataka (ES)

Linus Torvalds, desarrollador y fundador de Linux, sobre dirigir su empresa hoy: "Ya no soy programador" — open the original publisher

Linus Torvalds creó el kernel Linux en 1991, pero 35 años después ya no se define como un programador que escribe código. Durante el Open Source Summit India 2026 explicó que apenas lee código del proyecto de código abierto , y que hoy se considera más bien un jefe de proyecto. Uno de los programadores más legendarios de la historia ha abandonado esa labor, y es probablemente lo mejor que le podía pasar al proyecto. Lo de programar, como que no . Torvalds explica en ese encuentro que sigue escri

Jobs & economy
Next (FR, ex-INpact)

NVIDIA s’allie à six grands financiers pour tenter de sécuriser 500 milliards de dollars — open the original publisher

Nvidia annonce un partenariat encore hypothétique avec six acteurs de la finance pour proposer une « nouvelle classe d’actifs investissables », selon les mots de son PDG Jensen Huang. Si l’accord est conclu, il pourrait permettre l’entrée de financements extérieurs au secteur de l’IA. NVIDIA annonce un accord avec Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs et KKR […]

Finance, VC & PE
The 74 (education AI)

Opinion: Amid the Emerging AI Economy, We Need a Skilled Trades Pipeline in High School — open the original publisher

Across the economy, Americans are watching an artificial intelligence investment boom reshape the job market. Some of the same companies spending hundreds of billions to build the AI future are also announcing sweeping job cuts, adding to an already daunting employment landscape for college graduates. In other sectors, opportunity is booming. Homes, roads, bridges, vehicles, […]

Jobs & economyChildren & education
Xataka (ES)

El trágico caso de la persona más joven jamás diagnosticada con Alzheimer: fue diagnosticado cuando tenía 19 años — open the original publisher

El Alzheimer, como otras enfermedades neurodegenerativas, tiende a manifestarse a edades avanzadas, aunque puedan pasar muchos años entre el comienzo de la enfermedad y la aparición de los síntomas. Cuando la enfermedad aparece antes de los 65 años hablamos de Alzheimer de aparición temprana. El problema es que este trastorno puede llegar a manifestarse mucho antes . El caso más precoz. No sabemos a qué edad puede llegar a comenzar a aparecer el mal de Alzheimer, pero el caso más precoz jamás re

Healthcare
MediaNama (IN)

Data protection authority suspends Discord livestreams in Brazil after teen’s suicide — open the original publisher

Suspending Discord's 'Go Live' feature, the National Data Protecting Authority has asked the company to prove compliance to Brazil's Digital ECA after a 13-year-old girl's suicide during a livestream. The post Data protection authority suspends Discord livestreams in Brazil after teen’s suicide appeared first on MEDIANAMA .

RegulationPrivacy
TechCabal (Africa)

Africa’s cybercriminals are adopting AI faster than the institutions chasing them — open the original publisher

AI was involved in 55% of cybercrime cases observed by African countries surveyed by Interpol in 2025, as criminals used the technology to produce convincing phishing messages, fabricate identities, and impersonate executives and public figures. Deepfake incidents increased sevenfold between the second and fourth quarters of 2024.

Misinformation
The 74 (education AI)

More Early Childhood Programs Are Providing Free Housing to Teaching Staff — open the original publisher

This story was co-published with Mother Jones. A few years ago, Eric Gil was living at his uncle’s place in Waterbury, Connecticut, where he shared a bedroom with his brother and cousin. With eight people in the house, it was crowded. He was trying to get out, but rental prices in Waterbury — which currently […]

Children & education
Legal IT Insider

TalkingTech podcast: AI’s biggest legal challenge isn’t the technology. It’s the data — open the original publisher

As generative AI moves from experimentation to implementation, one question continues to dominate conversations in the legal sector: are law firms’ data foundations strong enough to support it? In our […] The post TalkingTech podcast: AI’s biggest legal challenge isn’t the technology. It’s the data appeared first on Legal IT Insider .

Regulation
Le Monde Pixels (FR)

A Taïwan, des agences gouvernementales ont été la cible de cyberattaques provoquées par des agents d’intelligence artificielle — open the original publisher

Les autorités taïwanaises n’ont pas donné de détails sur l’ampleur, les dégâts ou les cibles précises de ces attaques. Elles n’ont pas, non plus, mentionné la Chine, qu’elles accusent régulièrement de harceler l’île avec des cyberattaques.

Agents & autonomy
InfoQ AI/ML

Anthropic's Claude Breaches Sandbox During Model Security Evaluations — open the original publisher

Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised attacks on live targets. Anthropic has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors. By Olimpiu Pop

Transparency
MediaNama (IN)

Jio Financial to sell 49.9% stake in NBFC arm to Bank of America for $1.9 billion — open the original publisher

Jio Financial Services and Bank of America has signed an agreement which lets BofA acquire 49.9% stake in its NBFC subsidiary Jio Credit through a preferential allotment of equity shares and warrants. The post Jio Financial to sell 49.9% stake in NBFC arm to Bank of America for $1.9 billion appeared first on MEDIANAMA .

Bias & fairness
Fast Company Tech

The world’s largest ‘biological datacenter’ could help make animal testing obsolete — open the original publisher

There’s a big flaw in the way that drug companies develop and test medicines today: What works in a mouse often doesn’t work in a human. For decades, the industry has relied on animal testing. But in a laboratory south of San Francisco, a startup called Vivodyne is scaling up a different approach. Inside wardrobe-size mini labs, robots grow human tissue and run thousands of AI -designed experiments that could better predict how well a new drug will work—and whether it will be safe. [Image: Vivod

HealthcareAgents & autonomy
SCMP Tech (HK/CN)

DeepSeek’s updated V4 Pro AI model struggles on benchmarks, shines in cybersecurity — open the original publisher

Chinese artificial intelligence start-up DeepSeek has quietly released DeepSeek-V4-Pro-0813, an updated version of its latest flagship model, leaving some developers underwhelmed by its overall capabilities and disappointed in its pricing – but impressing researchers in niche areas like cybersecurity. The stealth update to April’s preview version came with a brief statement on DeepSeek’s official website on Wednesday, noting that the model offered “significantly enhanced agent capabilities”, but

Agents & autonomy
War on the Rocks

The True Cost of Cheap Chips — open the original publisher

Given the relentless demand for computing power, electronic components are in scarce supply. Prices for certain memory chips, known as DRAM, have surged by more than 50 percent in a single quarter this year, and have roughly quadrupled since last fall. Because DRAM supply is tight, Apple, Dell, and HP are currently evaluating memory from ChangXin Memory Technologies (CXMT), a company the Pentagon has designated as a Chinese military company. Apple, in particular, has sought assurances from the U

Military & security
War on the Rocks

2026 and All That: Another Benchmark Year in Royal Navy Decline — open the original publisher

The 1930 comic history, 1066 and All That, made famous the British habit of reducing national history to a sequence of memorable dates. The modern Royal Navy has its own unhappy version of that calendar. Since 1945, a series of ostensibly practical political decisions has steadily reduced Britain’s ability to sustain a globally relevant fleet: the 1956 failed Suez expedition, the 1966 retreat from “East of Suez,” the 1981 Nott Review, the 1998 Strategic Defence Review, and the 2010 review that f

Military & security
Next (FR, ex-INpact)

☕️ La préfecture de Seine-et-Marne autorise l’installation de Campus IA à Fouju — open the original publisher

C’est fait. Sans grande surprise, le plus grand projet de data center de France, Campus IA, a obtenu le 29 juillet l’autorisation préfectorale de s’installer à Fouju, à 40 km au Sud-Est de Paris. Après une consultation publique organisée auprès de la population à l’automne, puis une enquête publique à l’issue de laquelle les trois […]

Environment
TechNode (CN)

Tencent Backs Lovable in $400 Million Series C at $13.3 Billion Valuation — open the original publisher

Lovable, a Stockholm-based software creation platform, has raised $400 million in Series C funding at a $13.3 billion valuation. The round was led by Menlo Ventures and co-led by EQT’s Scaleup Europe Fund. New investors include Tencent, Balderton Capital, Carmignac, Kaszek Ventures, LTS Growth, World Innovation Lab and Regent. Lovable provides tools that allow users […]

Finance, VC & PE
TechNode (CN)

Honor Launches Robot Phone with Gimbal Camera and AI Agent Features — open the original publisher

Honor has launched its Robot Phone, a smartphone equipped with a four-degree-of-freedom titanium gimbal and a system-level AI agent architecture. The phone measures about 9.59 millimeters thick, weighs 248 grams and includes a 7,060mAh battery and a 6.31-inch display. The 12GB+512GB version is priced at 9,999 yuan, while the 16GB+1TB version costs 12,999 yuan. Pre-orders […]

Agents & autonomy
TechNode (CN)

Alibaba Cloud Launches Qwen AI Arena for Real-World Agent Testing — open the original publisher

Alibaba Cloud has launched Qwen AI Arena, a challenge and evaluation platform for AI agents. The platform creates tasks based on real business scenarios and provides developers with models, runtime environments and evaluation tools to submit and test agent solutions. Its first challenge focuses on cross-border e-commerce. Participants must generate product listings for the US, […]

Agents & autonomyEnvironment
SCMP Tech (HK/CN)

China’s ‘brain chip’ drive accelerates with slew of state-backed initiatives — open the original publisher

For years, Elon Musk has dreamed of conquering a range of chronic health conditions by inserting computer chips into the human brain. But that vision could become reality fastest in China, where state authorities are launching a coordinated effort to accelerate the nascent industry’s development. The past few days have seen a string of initiatives related to brain-computer interfaces (BCIs) announced in China, involving parties ranging from state insurance companies to investment banks and local

HealthcareFinance, VC & PE
Variety (AI)

Indian Freedom Fighter Bhagat Singh Series in Development at Collective Studios’ Historyverse (EXCLUSIVE) — open the original publisher

A new original series exploring the life of Indian freedom fighter Bhagat Singh is in development at Collective Studios’ Historyverse. The project is being made in association with Hathiramani Commercial Ventures and co-produced with investor and entrepreneur Manish Hathiramani, who marks his producing debut with the show. The series is designed to move past the […]

Finance, VC & PE
Variety (AI)

Hollywood’s Top-Grossing Films of 2025 Show Limited Progress Toward Inclusion, Study Says: ‘The Pace of Change Has Been Slow’ — open the original publisher

After nearly 20 years of tracking representation metrics, the latest report from Dr. Stacy L. Smith and the Annenberg Inclusion Initiative reveals that Hollywood’s film industry has not made lasting progress towards inclusion. “While we’ve seen pockets of progress for women on screen, the pace of change has been slow. Year after year, many of […]

Privacy

Field notes (22)

EFF Deeplinks

Too Little, Too Late: Flock Admits Their Technology Needs Reforms — open the original publisher

Flock Safety, the embattled vendor of mass surveillance technology, has rolled out a handful of new reforms intended to appease the justified nationwide anger that has seen scores of towns cancel or suspend their contracts with the company for automated license plate readers (ALPRs). The reforms are a combination of long overdue changes along with some cosmetic fixes that fail to address the fundamental dangers of this technology. We should not be letting companies decide how much privacy we des

Privacy
Truth on the Market (digital regulation)

The Data Center Chessboard Has No Pause Button — open the original publisher

The whole country ostensibly wants America to win the artificial intelligence (AI) race. A striking number, however, would prefer someone else’s town to host the data centers, power plants, transmission lines, and cooling systems required to run it. Adam Smith knew the type. In “The Theory of Moral Sentiments,” he warned against the “man of ... The Data Center Chessboard Has No Pause Button The post The Data Center Chessboard Has No Pause Button appeared first on Truth on the Market .

Environment
Inforrm (media law)

Global Freedom of Expression, Columbia University: Newsletter, 13 August 2026 — open the original publisher

Columbia Global Freedom of Expression seeks to contribute to the development of an integrated and progressive jurisprudence and understanding on freedom of expression and information around the world. It maintains an extensive database of international case law. This is its newsletter dealing with recent developments in the field. For a video to go viral on TikTok, […]

Regulation
NVIDIA Blog (AI)

Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More — open the original publisher

GeForce NOW is giving cloud gaming an extra-credit upgrade just in time for back-to-school season. The native Linux app for GeForce NOW is officially out of beta. GeForce NOW is also delivering new cloud optimizations that make Frame Generation feel even more responsive while streaming. On top of that, Performance members will see higher frame […]

Children & education
Just Security

“In Focus” Syllabus Supplements: U.S. Lethal Strikes on Suspected Drug Traffickers, Operation Southern Spear, and Operation Absolute Resolve (2025–2026) — open the original publisher

This syllabus supplement offers curated articles intended to be combined with traditional casebooks in a law school or other higher education classroom. The post “In Focus” Syllabus Supplements: U.S. Lethal Strikes on Suspected Drug Traffickers, Operation Southern Spear, and Operation Absolute Resolve (2025–2026) appeared first on Just Security .

RegulationHealthcare
Just Security

“In Focus” Syllabus Supplements: ICE and CBP Operations in Minnesota and Other States (2025–2026) — open the original publisher

This syllabus supplement offers curated articles intended to be combined with traditional casebooks in a law or higher ed classroom. The post “In Focus” Syllabus Supplements: ICE and CBP Operations in Minnesota and Other States (2025–2026) appeared first on Just Security .

RegulationChildren & education
Just Security

“In Focus” Syllabus Supplements: Artificial Intelligence, Emerging Technology, and National Security (2025–2026) — open the original publisher

This syllabus supplement offers curated articles intended to be combined with traditional casebooks in a law or higher ed classroom. The post “In Focus” Syllabus Supplements: Artificial Intelligence, Emerging Technology, and National Security (2025–2026) appeared first on Just Security .

RegulationMilitary & security
Just Security

Immigration Law & Policy: Syllabus Supplements — open the original publisher

Access the Immigration Law & Policy Syllabus Supplements via PDF here. Access the original version, published Sep. 18, 2025, via PDF here. Additional syllabus supplements, including our new “In Focus” syllabus supplements, are available here. Table of Contents I. Executive, Congressional, and State Authorities II. Admission to the United States III. Non-Citizens in the United […] The post Immigration Law & Policy: Syllabus Supplements appeared first on Just Security .

Regulation
Bruce Schneier — Schneier on Security

Separating AI’s Technological Problems from Its Capitalism Problems — open the original publisher

This essay was written with Nathan E. Sanders, and originally appeared in Tech Policy Press . AI represents the first time we humans can do cognitive work outside of our bodies at scale. The only comparable moment is the early years of the industrial revolution, when new technologies like the steam engine provided a quantum leap in our ability to do mechanical work outside of our bodies at scale. If AI’s cognitive capabilities become integrated into our lives, businesses, and governments—a proce

Regulation
EA Forum (AI safety)

What would make us scale or stop? NOVAH's pre-commitment before seeing the RCT results — open the original publisher

Published on August 11, 2026 11:51 AM GMT TL;DR NOVAH (No Violence At Home) was incubated by Charity Entrepreneurship (now Ambitious Impact) in 2024 to test a promising idea: preventing intimate partner violence through edutainment, in our case a serialised radio drama. Over the past two years we have produced and aired two seasons in Rwanda. We are currently evaluating our second season through a randomized controlled trial with 2,400 couples in Rwanda in partnership with Innovations for Povert

Apple Machine Learning Research

When Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational Costs — open the original publisher

As concerns around data privacy in machine learning grow, the ability to unlearn, or remove, specific data points from trained models becomes increasingly important. While state of the art unlearning methods have emerged in response, they typically treat all points in the forget set equally. In this work, we challenge this approach by asking whether points that have a negligible impact on the model’s learning need to be removed. Through a comparative analysis of influence functions across langua

Privacy

Policy (12)

UK Contracts Finder — AI procurement

NPCC ANPR Analytics Tool — open the original publisher

… were identified, alongside the need for further ML in order that the tool could be made more widely available across all ROCUs and other crime types in the future. This Machine Learning (ML) phase will ensure that the AAT can deliver region specific insights into County lines and provide the evidence required to inform future national adoption. Classification: Software package and information systems.

Regulation
White House

Adjusting Imports of Unmanned Aircraft Systems and Unmanned Aircraft Systems Components into the United States — open the original publisher

BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION 1. Within the past 90 days, the Secretary of Commerce (Secretary) transmitted to me a report on his investigation into the effects of imports of unmanned aircraft systems (UAS), as well as their parts and components (together, UAS components), on the national security of […] The post Adjusting Imports of Unmanned Aircraft Systems and Unmanned Aircraft Systems Components into the United States appeared first on The White House .

Military & securityFinance, VC & PE
White House

Rebuilding the United States Navy and America’s Shipbuilding Industrial Base — open the original publisher

MEMORANDUM FOR THE SECRETARY OF WAR THE DIRECTOR OF THE OFFICE OF MANAGEMENT AND BUDGET THE ASSISTANT TO THE PRESIDENT FOR NATIONAL SECURITY AFFAIRS SUBJECT: Rebuilding the United States Navy and America’s Shipbuilding Industrial Base By the authority vested in me as President by the Constitution and the laws of the United States of […] The post Rebuilding the United States Navy and America’s Shipbuilding Industrial Base appeared first on The White House .

Military & security
US Federal Register

Airworthiness Directives; The Boeing Company Airplanes — open the original publisher

The FAA is superseding Airworthiness Directive (AD) 2018-11- 14, which applied to certain The Boeing Company Model 767-300 and -300F series airplanes with certain winglets installed. AD 2018-11-14 required high frequency eddy current (HFEC) inspections for cracking of the lower outboard wing skin and repair or modification if necessary. AD 2018-11-14 also required one of three follow-on actions: Repeating the HFEC inspections; modifying certain internal stringers and oversizing and plugging the

Baker McKenzie Connect On Tech

Brazil: ANPD regulates Digital ECA transparency reports — open the original publisher

First transparency report due by September 17 ANPD establishes minimum content and a deadline for September 17 for the first Digital ECA transparency report. In brief On August 11, 2026, the Brazilian Data Protection Agency (ANPD) issued Decision Order CD/ANPD No. 122/2026, which addresses requirements of the semiannual transparency reports required under the Children and [...] The post Brazil: ANPD regulates Digital ECA transparency reports appeared first on Connect On Tech .

RegulationPrivacy
US GAO Reports

National Crisis and Support Hotlines: Contact Volume Increased and Operations Remained Relatively Consistent in 2025 — open the original publisher

What GAO Found The 988 Suicide & Crisis Lifeline provides free, confidential 24/7 support for anyone experiencing mental health distress, suicidal thoughts, or substance use crises. The National Maternal Mental Health Hotline provides free, confidential emotional support, resources, and referrals to pregnant and postpartum women experiencing mental health challenges. Both of these hotlines, within the Department of Health and Human Services (HHS), are operated by nonfederal entities—primarily by

Healthcare

Research (105)

arXiv

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design — open the original publisher

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design

Agents & autonomy
arXiv

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark — open the original publisher

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to

Privacy
arXiv

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data — open the original publisher

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture o

arXiv

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) — open the original publisher

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two ma

arXiv

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity — open the original publisher

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificit

arXiv

Synthetic Persona Pretraining: Alignment from Token Zero — open the original publisher

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the d

Safety & alignment
arXiv

A Unifying Perspective on Causal World Models: From Observations to Representations to Structure — open the original publisher

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entit

Agents & autonomyEnvironment
arXiv

UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models — open the original publisher

Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and performing a wide range of manipulation tasks. However, their direct control over embodied agents also exposes them to adversarial interference that may cause unsafe physical behaviors. Existing attacks on robotic policies are typically optimized for a single task or instruction, leaving the cross-task vulnerabilities of multitask VLAs largely unexplored. We intr

Agents & autonomy
arXiv

Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension — open the original publisher

Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artificial Intelligence (LIA) at the Federal University of Santa Catarina (UFSC), designed to integrate teaching, research, and university extension through a student-centered, project-based approach. The framework combines democratic governance, collaborativ

RegulationChildren & education
arXiv

Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements — open the original publisher

Machine learning ethics researchers and critical HCI scholars have argued that algorithmically predicting gender is wrong. At the same time, other researchers rely on predicted gender labels to study gender disparities and develop algorithmic fairness techniques. How do we reconcile these two seemingly contradictory intuitions? We differentiate two ways gender prediction may be wrong: being illegitimate, thereby contributing to harm; and being invalid, thereby producing unusable measurements. Ou

Bias & fairness
arXiv

Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks — open the original publisher

6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where thousands of autonomous artificial intelligence (AI) agents are interconnected. The agents are deployed across a wide range of platforms including low Earth orbit (LEO) satellites, high-altitude platforms (HAPs), unmanned aerial vehicles (UAVs), edge servers, and terrestrial devices. These agents continuously observe their environment and exchange information

Agents & autonomyEnvironment
arXiv

TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies — open the original publisher

Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. Existing NetOps automation tools mainly operate after this design is fixed, providing limited support for generating structured security topologies from underspecified natural-language requirements. We present TopoIntent, a system that compiles security intent into executable, compliance-checked networ

RegulationJobs & economy
arXiv

Rules or Character? Scaling Laws for AI Safety Design — open the original publisher

Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes saf

Safety & alignment
arXiv

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning — open the original publisher

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived fro

arXiv

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples — open the original publisher

Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled training sets, and transfer poorly across analytical objectives and experimental datasets. Here we introduce UltraIR, a foundat

Jobs & economy
arXiv

StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems — open the original publisher

Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that token identities alone cannot capture. Recent work proposes latent communication as an alternative, where agents transmit hidden representations directly without converting them to text. However, existing latent methods either inject working memory la

Safety & alignmentAgents & autonomy
arXiv

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model — open the original publisher

We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixture of Training (MoT), a scaffolded modular pre-training procedure that partitions a target Transformer into contiguous layer blocks, trains each block inside a frozen pretrained aligner scaffold, and then recomposes the trained blocks with an optional short end-to-end adaptation pass. On a 1.3B-parameter Gemma-style m

Jobs & economy
arXiv

Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability — open the original publisher

A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both the legal power and the political will to decide which other countries may use these systems. In June 2026 the United States required a leading developer to obtain licences before releasing its most advanced models to any foreign person, including foreign nationals resident in the United States. The affected models were withdrawn worldwide at short notice, part

Military & security
arXiv

Into the ORBIT for Time Series: Training Regimes for Foundation Models — open the original publisher

Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored. As a result, pre-training distributions are often poorly controlled with respect to domain imbalance, context requirements, prediction horizons, and missingness. We introduce ORBIT (Omni-Range Bootstrap Incremental Training), a training paradigm that makes this distribution explicit and controllable. ORBIT combines Boo

arXiv

Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data — open the original publisher

As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data cap

Finance, VC & PEBiotech
arXiv

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport — open the original publisher

Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy by reusing computation across denoising steps. In multi-view texturing, however, skipping a step also removes the cross-view interaction that continually aligns different observations of the same surface, leading to rapidly degraded consistency and fi

arXiv

Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales — open the original publisher

Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by se

arXiv

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test — open the original publisher

Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology class provides a diagnostic and search feature. A contr

HealthcareAgents & autonomy
arXiv

Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering — open the original publisher

Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved documents into English or the query language to bridge the cross-lingual semantic gap, or decompose a complex query into sub-questions and aggregate the intermediate reasoning process. However, both lines of work suffer from two limitations. First, one-size-fits-all translat

arXiv

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation — open the original publisher

With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To add

arXiv

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback — open the original publisher

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only acro

Agents & autonomy
arXiv

Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment — open the original publisher

Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by introducing the Multi-Layer Context Camouflaging Theory (M

arXiv

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding — open the original publisher

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable repres

arXiv

Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds — open the original publisher

Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively paralleliz

Safety & alignment
arXiv

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion — open the original publisher

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache).

RegulationSafety & alignment
arXiv

Applied and Filtered: An End-to-End Algorithmic Fairness Audit of A Public Employment Agency — open the original publisher

Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 t

Bias & fairnessJobs & economy
arXiv

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses — open the original publisher

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We i

arXiv

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving — open the original publisher

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find

arXiv

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents — open the original publisher

Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution. Existing benchmarks measure current behavior or

RegulationAgents & autonomy
arXiv

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation — open the original publisher

Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. This bias originates from the combined effects of imbalanced semantic codebooks during SID construction, and popularity bias together with the maximum likelihood estim

Bias & fairness
arXiv

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval — open the original publisher

Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's solution to AI City Challenge 2026 Track 4. Our framework

arXiv

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs — open the original publisher

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic compon

Safety & alignment
arXiv

Correct Is Not Governed: Provenance Integrity in Agentic Workflows — open the original publisher

Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies co

Agents & autonomy
arXiv

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval — open the original publisher

While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce \textbf{ERSkill}, a retrieval-centric framework for self-evolving, skill-guided memory access. ERSkill compiles interaction histories into

Agents & autonomy
arXiv

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging — open the original publisher

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tas

Safety & alignmentHealthcare
JMIR (Journal of Medical Internet Research)

Digital Decisions: Enhancing Chronic Disease Self-Care Through Digital Health and AI-Enhanced Decision-Making — open the original publisher

Advances in digital health have dramatically changed how patients engage with their health. Rather than relying solely on periodic clinical visits, patients now have access to smartphones, patient portals, wearable devices, and mobile apps that provide support for day-to-day self-care decisions. This commentary discusses the findings of Longhini et al’s systematic review and meta-analysis on the effectiveness of digital health interventions, which found modest improvements in self-care monitorin

Healthcare
JMIR (Journal of Medical Internet Research)

Digital Health Technology for Improving Physical Function in Adults With Chronic Heart Failure: Systematic Review and Meta-Analysis of Randomized Controlled Trials — open the original publisher

Background: Chronic heart failure (CHF) significantly impairs physical function and quality of life. Although exercise-based cardiac rehabilitation represents a primary therapeutic strategy, participation rates remain low due to logistical barriers. Digital health technologies (DHTs) offer a promising alternative to deliver home-based interventions. However, evidence regarding their specific impact on functional capacity versus daily physical behavior remains inconsistent. Objective: This system

Healthcare
JMIR (Journal of Medical Internet Research)

Exploring Perceptions of Leveraging AI to Improve Outcomes in Maternal, Sexual, and Reproductive Health in Sub-Saharan Africa: Exploratory Qualitative Study — open the original publisher

Background: AI has the potential to transform health care in low- and middle-income countries, where access to quality care remains limited. Maternal, sexual, and reproductive health (MSRH) outcomes are especially poor due to resource shortages, financial barriers, and geographic inequities. With thoughtful implementation, AI could help address these gaps through innovations in diagnostics, health education chatbots, and telemedicine. However, responsible use is essential to ensure AI reduces, r

HealthcareChildren & education
JMIR (Journal of Medical Internet Research)

Association Between Daily Food-Tracking Frequency and Clinically Significant Weight Loss: Retrospective Cohort Study of 5132 Mobile App Users — open the original publisher

In this retrospective cohort of 5132 users of a commercial nutrition-tracking mobile application, higher food-tracking frequency was associated with greater weight loss over 6 months; 70.1% (3599/5132) of users lost at least 5% of body weight.

Privacy
arXiv cs.LG

ProME: Prototype-Margin Environments with Repair-Aware Selection for Group-Robust Learning — open the original publisher

Group-robust learning is crucial for maintaining accuracy on rare subpopulations when training-group labels are unavailable. However, existing methods often infer environments from a separate reference model and select representations before fitting the classifier used at deployment, leaving both decisions misaligned with the deployed predictor. In this work, we formulate group robustness without training-group labels as the endogenous environments with repair-aware selection (ERAS) problem, and

Safety & alignmentEnvironment
MIT Sloan Management Review AI

Why Water Management Is a Strategic Concern — open the original publisher

Dante Terzigni/theispot.com We live on a blue planet where water appears to be plentiful. But there are strong warning signs indicating that our reliance on the fresh water we perceive to be abundant will need to change. Although water covers 70% of the Earth’s surface, only 0.5% of it is effectively usable. Further, global water […]

Environment
arXiv cs.LG

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization — open the original publisher

Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this gap with dense token supervision, yet applying it throughout training creates a different failure mode: the teacher is a biased, low-variance surrogate for the reward objective, so persistent imitation can oppose reward-improving updates after the policy becomes capable of produ

Bias & fairnessRegulation
arXiv cs.CY

Variable Selection in the Context of AI Fairness — open the original publisher

arXiv:2608.11251v1 Announce Type: new Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account philosophical ethics and social awareness. Variable selection processes, in particular, can introduce implicit bias, affecting equity across different subgroups. We discuss a mathematical approach that evaluates fairness in AI, aligning mathematical methodologies with ethical considerations an

Bias & fairnessRegulation
arXiv cs.CY

Methodologies for Improving the Quality of AI Tutoring in K-12 Education — open the original publisher

arXiv:2608.11259v1 Announce Type: new Abstract: Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have

Children & education
arXiv cs.CY

Governing Agentic AI in FinTech — open the original publisher

arXiv:2608.11344v1 Announce Type: new Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility r

RegulationAgents & autonomy
arXiv cs.CY

The Accuracy Trap: Structural Scarcity Amplifies Relative Inequality in Algorithmic Allocation — open the original publisher

arXiv:2608.11491v1 Announce Type: new Abstract: Algorithmic systems increasingly rank individuals for access to scarce public resources, from child welfare interventions to cancer treatment referrals. The prevailing fairness frame treats disparity as a property of biased data or deficient models, with remedies through calibration and debiasing. Under structural scarcity, where demand exceeds supply by an order of magnitude, allocation becomes a rationing problem, and the statistical properties o

Bias & fairnessChildren & education
arXiv cs.CY

Cheap, Fallible Cognition and the Political Economy of Expertise — open the original publisher

arXiv:2608.11512v1 Announce Type: new Abstract: The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question sel

Jobs & economy
arXiv cs.CY

Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion — open the original publisher

arXiv:2608.11794v1 Announce Type: new Abstract: The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The

RegulationTransparency
arXiv cs.CY

Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap — open the original publisher

arXiv:2608.11803v1 Announce Type: new Abstract: Deployed foundation models are often not static systems, with providers able to modify system behavior through fine-tuning, classifier updates, system prompt revisions, retrieval changes, and routing changes. These updates can be made silently -- that is, without public disclosure, a version increment, or re-evaluation. Such silent updates challenge a core assumption behind current AI governance frameworks that an externally verifiable chain of cus

RegulationTransparency
arXiv cs.CY

Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs — open the original publisher

arXiv:2608.11830v1 Announce Type: new Abstract: The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water

HealthcareEnvironment
arXiv cs.CY

Philosophical vertigo with artificial intelligence — open the original publisher

arXiv:2608.11955v1 Announce Type: new Abstract: Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabi

Safety & alignment
arXiv cs.CY

No One to Blame: A Framework of Constitutive AI Unaccountability — open the original publisher

arXiv:2608.12104v1 Announce Type: new Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and institutional reform. We argue that this framing is insufficient: certain configurations of actors, systems, and institutions render AI accountability conceptually unachievable regardless of effort. We i

Agents & autonomyTransparency
arXiv cs.CY

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers — open the original publisher

arXiv:2608.12166v1 Announce Type: new Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should be made transparent and how, as well as in their interest in and ability to parse the information currently published in the registers. Moreover, it remains unclear how these instruments can represent the sociotechnical systems in which these algorithms are embed

RegulationTransparency
arXiv cs.CY

Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior — open the original publisher

arXiv:2608.12292v1 Announce Type: new Abstract: An effective large language model (LLM) tutor must often decline to give an answer it could easily produce. In a randomized study, students who used an unguarded chatbot scored higher while practicing but lower on a later test taken without it, whereas a Socratically guarded version of the same model kept the practice gain and removed the later loss [4]. Reliable answer-withholding is therefore central to a tutor's value, yet a capable model presse

Children & education
arXiv cs.CY

Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach — open the original publisher

arXiv:2608.11245v1 Announce Type: cross Abstract: Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide lea

Children & education
arXiv cs.CY

Why AI Detection Fails for Academic Integrity — open the original publisher

arXiv:2608.11256v1 Announce Type: cross Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pang

Regulation
arXiv cs.CY

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits — open the original publisher

arXiv:2608.11410v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions. Using the MIMIC-III database, we propose the Counterfactual Clinical Audi

HealthcareAgents & autonomy
arXiv cs.CY

A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era — open the original publisher

arXiv:2608.11540v1 Announce Type: cross Abstract: The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping manufacturing faster than engineering curricula can adapt, widening the gap between the competencies required on the shop floor and those delivered by traditional engineering and technology education. This paper proposes a Workforce Readiness Level (WRL) framework, which adapts the Technology Readiness Level s

Jobs & economyMilitary & security
arXiv cs.CY

Organizational Technology Ladders: Remote Work and Generative AI Adoption — open the original publisher

arXiv:2608.11626v1 Announce Type: cross Abstract: This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology during the COVID-19 period shaped later uptake of generative AI. Using U.S. job-posting data and an instrumental-variables strategy based on predicted

Jobs & economy
arXiv cs.CY

Twitter and disability activism: leadership and relevant topics in the online conversation — open the original publisher

arXiv:2608.11923v1 Announce Type: cross Abstract: The dissemination and viralization of information on social media has been widely studied from various perspectives, including that of digital activism. On the other hand, disability-related activism has conquered the online environment, thus obtaining a reach that goes beyond the offline space and generating dialogue in the digital sphere. This article analyses the conversation generated on Twitter, taking as a sample all the tweets with the #di

Environment
arXiv cs.CY

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages — open the original publisher

arXiv:2608.12278v1 Announce Type: cross Abstract: Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barrier

Children & education
arXiv cs.CY

Small Data Explainer -- The impact of small data methods in everyday life — open the original publisher

arXiv:2507.11773v2 Announce Type: replace Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies. We provide a conceptual overview, clarify the relationship between sma

RegulationHealthcare
arXiv cs.CY

Public support for misinformation interventions depends on perceived fairness, effectiveness, and intrusiveness — open the original publisher

arXiv:2508.05849v3 Announce Type: replace Abstract: The proliferation of misinformation on social media has concerning possible consequences, such as the degradation of democratic norms. While recent research on countering misinformation has largely focused on analyzing the effectiveness of interventions, the factors associated with public support for these interventions have received little attention. We asked 1,010 American social media users to rate their support for and perceptions of ten mi

Bias & fairnessMisinformation
arXiv cs.CY

Ethics Practices in AI Development: An Empirical Study Across Roles and Regions — open the original publisher

arXiv:2508.09219v3 Announce Type: replace Abstract: Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks posed by these technologies. In this paper, we present a mixed-methods survey study - combining statistical and qualitative analyses - to examine the ethical perceptions, practices, and knowledge of individuals involved in various AI development roles. Our survey comprises 414 participants from 43 countries

Regulation
arXiv cs.CY

Prestige over merit: An adapted audit of LLM bias in peer review — open the original publisher

arXiv:2509.15122v2 Announce Type: replace Abstract: Large language models (LLMs) play a growing but largely informal role in scholarly peer review. Yet whether LLMs reproduce biases observed in human decision-making remains unclear. We adapt a resume-style audit to scientific publishing, developing a multi-role LLM simulation (editor/reviewer) that evaluates high-quality manuscripts across the physical, biological, and social sciences under randomized author identities (institutional prestige, g

Bias & fairnessTransparency
arXiv cs.CY

Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web — open the original publisher

arXiv:2510.10315v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly relying on web crawling to stay up to date and accurately answer user queries. These crawlers are expected to honor robots.txt files, which govern automated access. In this study, for the first time, we investigate whether reputable news websites and misinformation sites differ in how they configure these files, particularly in relation to AI crawlers. Analyzing a curated dataset, we find a stark co

MisinformationAgents & autonomy
arXiv cs.CY

How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers? — open the original publisher

arXiv:2603.00056v2 Announce Type: replace Abstract: STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what students know but also into how effectively they can apply, relate to, and integrate concepts across various contexts. Thus, students' responses are critical markers of the quality of their understanding and not entities that should be merely graded. However, inferring these mental models from student an

Children & education
arXiv cs.CY

SolarChain: A Physics-Grounded Embodied IoT System for Verifiable Urban Solar Market Design — open the original publisher

arXiv:2605.23162v2 Announce Type: replace Abstract: Distributed solar markets must coordinate physical reports, economic allocation, and public settlement even when IoT data can be manipulated. We present SolarChain, a controlled Embodied Intelligence of Things (EIoT) prototype that integrates four functions: physics-bounded screening of photovoltaic reports, persistent agent and planner coordination, configurable allocation between producer rewards and market liquidity, and replayable hash-link

Agents & autonomy
Frontiers in Artificial Intelligence

Recent advancements and future prospects on AI-integrated sensing techniques for non-invasive chronic kidney disease diagnosis: a review — open the original publisher

Chronic Kidney Disease (CKD) has emerged as a major public health concern worldwide, and most patients with CKD are asymptomatic until the later stages, causing growing morbidity and mortality. Diabetes and hypertension are the main causative factors for the development of CKD, damaging the renal microcirculation system. In addition, the impact of Acute Kidney Injuries (AKI) may result in the recovery or progression to either CKD or renal failure. The conventional techniques for diagnosis, such

Healthcare
AI & Society

Designing for rhythm: tempo-setting infrastructures and the temporal conditions of digital life — open the original publisher

Digital systems increasingly function as tempo-setting infrastructures that organize the temporal conditions under which cognition, communication, learning, and participation occur. Although research in human-computer interaction, platform studies, and AI ethics has extensively examined privacy, fairness, transparency, engagement, and well-being, the temporal organization of digital life has received comparatively little attention as a distinct object of sociotechnical analysis. We argue that di

Bias & fairnessPrivacy
AI & Society

Artificial intelligence and academic integrity in nursing education: a thematic analysis of Filipino nurse educators’ perspectives — open the original publisher

As technology advances, artificial intelligence is increasingly integrated into educational settings, raising important questions about its impact on learning and academic integrity. This study describes the perspectives of Filipino nurse educators on the role of artificial intelligence in maintaining academic integrity. A descriptive qualitative design was used, with a two-phase data collection process employing purposive sampling. In Phase 1, open-ended online responses were collected via Goog

Children & education
arXiv cs.AI

QuoteBench: How Matched Scores Can Hide Command-Path Failures — open the original publisher

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point re

Agents & autonomy
arXiv cs.AI

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure — open the original publisher

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELE

Children & education
arXiv cs.AI

Vero: Can AI Agents Build Formally Verified Software Repositories? — open the original publisher

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. Existing benchmarks in this direction either focus on individual functions or only evaluate proof generation with provided implementations. It is still an open question whether agents can m

Agents & autonomy
arXiv cs.AI

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination — open the original publisher

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a pla

HealthcareAgents & autonomy
arXiv cs.AI

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification — open the original publisher

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learn

Agents & autonomy
arXiv cs.AI

CAPRI: Contract-Aware Proof Repair for Isabelle — open the original publisher

We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised. We present CAPRI, a contract-aware repair workflow in which Isabelle checks the proof and an independent checker enforces a machine-readable edit contract. Prompts, proposals, candidate repositories, diagnostics, verdicts, and hashes are retained for audit. We evaluate five workflo

HealthcareTransparency
arXiv cs.AI

ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models — open the original publisher

Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close gripper--object views help observe contact, but a poor approach may already push, miss, slip, or disturb the object before conventional detectors react. We introduce \emph{ContactGuard}, a pre-contact execution monitor for chunked visuomotor policies. Given the policy's planned action chunk, ContactGuard predicts its short-horizon conseque

RegulationAgents & autonomy
arXiv cs.AI

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level — open the original publisher

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct com

RegulationFinance, VC & PE
arXiv cs.AI

Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes — open the original publisher

Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning

Agents & autonomyEnvironment
arXiv cs.AI

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development — open the original publisher

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule

Agents & autonomy
arXiv cs.AI

Deliberate Practice: Learning Robot Skills under a Budget — open the original publisher

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, \emph{Deliberate Practice (DP)}, that computes a provably \emph{budget-optimal} allocation---practicing skills that maximize expected cumulative reward while being learnable within the budget. DP estimates both the time needed to master skills and the cumulative reward of the task plans that the skills unlock. Computing a budget-optima

Agents & autonomy
arXiv cs.AI

Jointly Predicting Courses and Grades Using a Transformer-Based Model — open the original publisher

Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads. This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE) that addresses this limitation by jointly predicting both the set of courses a student will take and their

Children & education
arXiv cs.AI

Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs — open the original publisher

This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments. Three specialized discriminators -- global, hand, and head -- each guide a corresponding expert branch in the generator toward a distinct visual region, enabling implicit feature specialization without explicit diversity losses. To stabilize this multi-discriminator system,

Bias & fairness
arXiv cs.AI

It's How You Ask: Gender-Associated Linguistic Bias in LLMs — open the original publisher

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encode

Bias & fairness
arXiv cs.HC

More Than 63% of IEEE VIS Research Liable to be Retracted?! Ethics Approval Statements Protect Participants (and Researchers!) — open the original publisher

We analyzed the ethics reporting in 255 IEEE VIS papers from 2024 and 2025, as published in TVCG. This analysis arose from our experience as readers and reviewers of IEEE VIS papers that such reporting is frequently incomplete or missing, as well as from investigations in which we ourselves had to answer challenges regarding ethics approval in our own work. Visualization research naturally often involves human participants, yet ethics approval and informed-consent procedures are not always expli

Finance, VC & PE
arXiv cs.AI

Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision — open the original publisher

Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using sy

HealthcareFinance, VC & PE
arXiv cs.AI

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures — open the original publisher

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures wi

Healthcare
arXiv cs.HC

Tracing Methamphetamine abuse in under-treatment drivers: How biomechanical and oculomotor features help detect at-risk drivers? — open the original publisher

While the detrimental impacts of driving under the influence of stimulants such as methamphetamine are well-documented, the driving performance of individuals currently under-treatment has received considerably less attention. This study compared the behavior of individuals with a history of stimulant abuse (across two distinct treatment phases) with a control group of healthy drivers using a driving simulator. Oculomotor and biomechanical data were continuously collected via an eye-tracker and

Healthcare
arXiv red teaming query

Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research — open the original publisher

Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic "librarian" ingests timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph int

Agents & autonomyTransparency
arXiv fairness query

Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice — open the original publisher

Vertical Federated Learning (VFL) enables organizations holding complementary features of shared entities to collaborate and train models. In this setting, the initiator can withhold information about the learning task, while other contributors participate without exposing their local datasets, creating an asymmetric information structure aligned with growing privacy demands. However, this asymmetry is a double-edged sword. Among various threats, backdoor attacks are particularly concerning beca

Privacy
arXiv cs.HC

PolyPresentation: A Multimodal AI Platform for Slide-Aware Iterative Presentation Practice — open the original publisher

Presentations are essential for students, researchers, and professionals to communicate ideas persuasively, yet delivering them effectively requires repeated practice that coordinates content, delivery, visual materials, and audience interaction. Existing AI-assisted rehearsal tools provide scalable feedback, but they often treat presentations as single-run delivery performances, offering limited support for linking feedback to the slide deck or planning what to practice in the next iteration. T

Children & education
arXiv red teaming query

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models — open the original publisher

Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules. Such static designs struggle to maintain a cross-category safety boundary while generating constructive responses tailored to specific risks and avoiding over-refusal of benign inputs. To address these limitations, we propose HiRoute, an input-adaptive hierarchi

Safety & alignment
arXiv cs.HC

Considering Contribution Statements in Visualization and HCI Research — open the original publisher

Contribution statements are an increasingly common way to make research labor visible, reduce academic malfeasance, and provide broader transparency. Despite this potential value, they remain uncommon in visualization and HCI. To explore this gap, we conducted an online study with (N=21) visualization and HCI researchers. We find a range of differing opinions about the utility of contribution statements, which are set against a background of tensions relating to contribution frameworks that inad

Jobs & economyTransparency
arXiv cs.HC

NavSight in the Wild: Understanding Real-World Use of a Mobile Augmented Reality Application for People with Low Vision in Outdoor Navigation — open the original publisher

The ability to navigate outdoors safely and independently is crucial yet challenging for people with low vision (PLV). While various augmented reality (AR) systems for low vision have been designed and evaluated in ideal lab environments, no research has investigated their real-world feasibility and challenges. We present NavSight, a mobile AR application that assists PLV in outdoor navigation by recognizing important outdoor objects (e.g., curb, vehicle) and rendering real-time visual augmentat

EnvironmentFinance, VC & PE
arXiv cs.HC

PatientAct: Theory-Grounded Mental Health Client Simulation — open the original publisher

LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simula

Healthcare