Archive · 2026-08-11

AI ethics on Tuesday, 11 August 2026

330 items published this day, across 4 categories.

News (186)

The Guardian

The Guardian view on AI money in US politics: not the way to hold an urgent democratic debate | Editorial — open the original publisher

Amid public concern over the impact of new technology, Silicon Valley’s use of its wealth to influence specific election races is an unwelcome development As the huge societal ramifications of artificial intelligence have become clear, some of the tech industry’s most senior figures have taken to writing thinkpieces on what comes next. The latest to do so is Mark Zuckerberg, who this week published 6,500 words making the case for a laissez-faire approach to AI development. Worrying too much abou

Misinformation
The Guardian

Meta faces expensive child safety reckoning — open the original publisher

Also: Google executives jump ship in race for AI dominance Hello, TechScape readers! Danielle Abril, editor of the Guardian’s Reworked series on AI and the future of work, filling in for Blake Montgomery this week. Major legal battles against Meta over child safety are playing out in courts across the US – and the tech giant is losing. Recent rulings raise big questions about social media companies’ responsibility to their youngest users. Meanwhile, more key executives have jumped ship from Goog

Children & education
The Guardian

AI’s potential climate benefits outweighed by role in boosting fossil fuels, study finds — open the original publisher

Modelling finds AI-driven productivity gains in coal, oil and gas enable more emissions than applications in renewables avoid AI-driven productivity gains enable more planet-heating pollution from fossil fuels than they avoid from renewables, a study has found. Researchers modelled the technical potential for AI to boost clean power generation along with projections for how it can help produce coal, oil and gas. Across 64 scenarios, they found net yearly carbon pollution rose by 0.47-1.8 gigaton

Jobs & economyEnvironment
The Verge

Claude will apply invisible watermarks to AI text and images — open the original publisher

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. The changes are invisible to human […]

Transparency
Above the Law (legal tech)

Someone Is Keeping Up With Milbank — See Also — open the original publisher

Keeping That Raise On Ice : Ice Miller increases associate pay in New York, but not until 2027. Blame It On The AI : Company tried to pin labor law violation on ChatGPT . The Tech Isn't The Key : Legal AI needs to build a winning user experience if it's going to capture Biglaw . Everyone In MAGA Land Has Notes For Jeanine Pirro This Week: 'Never mention her friggin' name to me,' fumed the man demanding she fire a prosecutor. The post Someone Is Keeping Up With Milbank — See Also appeared first o

RegulationJobs & economy
SiliconANGLE AI

Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD — open the original publisher

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised $1.1 billion in early-stage funding. The company stated in today’s announcement that it received the capital over two rounds, a seed and a Series A. General Catalyst and AMP PBC were the lead investors. They were joined by Nvidia Corp., […] The post Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD appeared first on SiliconANGLE .

Finance, VC & PE
Retraction Watch

BMJ Group retracts paper on COVID-19 vaccines and mortality featured in Senate hearing — open the original publisher

BMJ Public Health has retracted a paper some — including a witness at a recent U.S. Senate subcommittee hearing — have used to link COVID-19 vaccines to deaths. The action comes more than two years after a researcher cited in the work flagged issues with the study. Published May 6, 2024, the paper used the … Continue reading BMJ Group retracts paper on COVID-19 vaccines and mortality featured in Senate hearing

Healthcare
Xataka (ES)

Un usuario le pidió a su agente de IA que le reservase hora en el gimnasio. Lo logró expulsando primero a otra persona de la lista — open the original publisher

Durante los últimos años nos hemos acostumbrado a pedirle cosas a una inteligencia artificial: que resuma un documento, busque información o nos ayude a tomar una decisión. Los agentes introducen un cambio bastante más profundo, porque ya no se limitan a responder : pueden recibir un objetivo, utilizar herramientas y actuar para intentar cumplirlo. Y ahí aparece una diferencia que parece pequeña, pero no lo es. Cuando decimos "consigue esto", damos por supuestos muchos límites sobre cómo hacerlo

Agents & autonomy
The Next Web AI

SpaceXAI launches Grok Bot as the agent race moves to office work — open the original publisher

SpaceXAI has launched Grok Bot in beta, AI agents that sign into apps and websites, retain context between tasks and coordinate with each other. Access is limited at first to SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium subscribers, with an enterprise waitlist. SpaceXAI has released Grok Bot, software that hands work to groups of […] This story continues at The Next Web

Agents & autonomy
The 74 (education AI)

Trump Signs Order on Childhood Vaccines Against Medical Groups’ Guidance — open the original publisher

WASHINGTON (AP) — President Donald Trump on Monday signed an executive order calling for revamped childhood vaccine recommendations that promote his long-held but discredited theory that childhood shots should be spaced out into separate medical visits. The order advocates separating the measles, mumps and rubella (MMR) vaccine into three different single-disease shots and administering all childhood immunizations at separate appointments whenever […]

RegulationHealthcare
SiliconANGLE AI

Snowflake moves enterprise AI beyond fragmented data pipelines — open the original publisher

Data interoperability is quickly becoming a practical requirement for companies trying to move artificial intelligence into production. Picking the right model or adding computing capacity is only part of the job. Companies also need reliable data that carries the right business meaning and remains protected as it moves between systems. Snowflake Inc. is building its […] The post Snowflake moves enterprise AI beyond fragmented data pipelines appeared first on SiliconANGLE .

Jobs & economy
The Hill Technology

AI reshapes traditional campaign strategies — open the original publisher

Artificial intelligence is reshaping the campaign trail, as candidates up and down the ballot turn to the technology for deepfake attack ads, viral social media posts and self-promotion. The embrace of AI as a political tool varies across political lines and races, drawing debate over whether deepfake content — particularly realistic videos or audio —...

Misinformation
SiliconANGLE AI

Real-time tax compliance puts agentic AI accuracy to the test — open the original publisher

AI-powered tax compliance has to meet a standard that many artificial intelligence applications don’t: The answers must be exactly right. While large language models can generate unpredictable results, tax calculations require accuracy, speed and reliability across thousands of jurisdictions. That tension has shaped the way Avalara Inc. applies agentic AI to its transactional tax and compliance […] The post Real-time tax compliance puts agentic AI accuracy to the test appeared first on SiliconAN

RegulationAgents & autonomy
Fast Company Tech

Anthropic models will soon inject watermarks identifying AI-generated text — open the original publisher

Anthropic said in a support document Tuesday that its future AI models will put an invisible watermark in all text, identifying it as AI-generated. The change comes in response to new AI transparency laws in the European Union, but Anthropic says it will apply in all countries where Claude models are used. No U.S. law currently requires AI-generated content to carry transparency disclosures or watermarks, although a growing number of states, California among them, have enacted AI transparency la

RegulationTransparency
Vox Future Perfect

Everybody needs a personal AI policy. Just ask Hank Green. — open the original publisher

Everyone is wrong about Hank Green. In case you missed the controversy: The veteran YouTube star, writer, and science comms entrepreneur was recently “canceled” after he acknowledged using AI for research. “I have been relying too heavily on AI as a research aid,” he wrote in a statement on Reddit. “It can be very useful […]

Regulation
SiliconANGLE AI

Multi-tier storage rewrites the economics of AI inference — open the original publisher

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine flash, object storage and disk-based capacity tiers, enabling enterprises to serve training and inference workflows while maximizing GPU productivity and economic savings. Super Micro Computer Inc. has collaborated […] The post Multi-tier storage rewrites the economics of AI inference appeared first o

Jobs & economy
The Hill Technology

Hunter Biden: Trump 'the one existential threat' to U.S. — open the original publisher

Hunter Biden, the son of former President Joe Biden, took a swipe at President Trump Monday, describing him as an “existential threat” to the U.S. During an interview with conservative media personality Tucker Carlson, the former president's son blamed wealthy tech leaders for the country’s current divide. He said Trump's cordial relationships with people like...

Safety & alignment
The Next Web AI

Apple Executive Who Launched Apple Pay Is Retiring After 25 Years — open the original publisher

Jennifer Bailey, head of Apple Pay and Wallet since the service launched in 2014, is retiring in October after more than 25 years at Apple. Her exit is one of several senior departures ahead of Tim Cook handing the chief executive job to John Ternus on September 1. Jennifer Bailey, who has run Apple Pay […] This story continues at The Next Web

Jobs & economy
The Decoder

OpenAI lets employees cash out another $7 billion in stock — open the original publisher

OpenAI wrapped up a $7 billion stock buyback, letting current and former employees sell shares at the company's $852 billion valuation. The move is meant to ease pressure on employees waiting for liquidity ahead of a potential IPO. OpenAI ran a similar $6.6 billion sale in October 2025. The article OpenAI lets employees cash out another $7 billion in stock appeared first on The Decoder .

Finance, VC & PE
Fast Company Tech

Mark Zuckerberg is getting blowback over his yacht in Alaska after this maritime incident — open the original publisher

A yacht owned by Mark Zuckerberg didn’t respond immediately to a call for help from a skiff that ran out of fuel off the coast of Alaska, but a maritime law expert said Monday the crew may not have been obligated to after the Coast Guard determined the boat wasn’t in distress. A spokesperson for the Meta CEO also said the crew didn’t immediately hear the radio call for help and by the time it did, another nearby ship had already rendered assistance. The Aug. 3 incident gained attention in recent

Regulation
Breaking Defense (AI)

Army needs ‘comprehensive missile defeat strategy’ to address air defense challenges — open the original publisher

“We need a more comprehensive missile defeat strategy that focuses on interceptors, [but also] equally so on non-kinetic effects, both electronic warfare and directed energy, and the type of targeting in C2 … that will enable us to use that data to make decisions to engage targets faster,” said SMD head Lt. Gen. John Rafferty.

Military & securityEnvironment
The 74 (education AI)

Opinion: As Students Head Back to School, Let’s Rethink What’s Success in Sex Education — open the original publisher

As millions of young people head back to school this fall, parents are buying backpacks, teachers are preparing lesson plans, and policymakers are once again debating what belongs in the classroom. One topic will predictably return to the center of those debates: sex education. Too often, the conversation begins and ends with one question: How […]

Children & education
The Hill Technology

House Democrats press AI giants on rogue agents — open the original publisher

A group of House Democrats are pushing two top artificial intelligence companies to disclose more information about reported incidents in which their AI models escaped containment and hacked into other companies during cybersecurity tests. The lawmakers cited the “serious risk that frontier AI models can pose” in separate letters to the top executives at Anthropic...

Agents & autonomy
Xataka (ES)

España lleva tiempo soñando con llegar a 100 millones de turistas al año. Algo amenaza con complicarlo: el calor extremo — open the original publisher

Las cifras son solo eso, cifras, pero es fácil obsesionarse con ellas. El sector turístico español lleva años pendiente de una muy concreta: alcanzar los 100 millones de visitantes extranjeros, una barrera psicológica a la que el país lleva años acercándose, pero que todavía no ha cruzado. Los investigadores de Oxford Economics y el ministro de Turismo , Jordi Hereu, han reconocido públicamente que el hito podría alcanzarse al fin en 2026, pero en el horizonte asoman también algunas complicacion

Finance, VC & PE
Xataka (ES)

Nvidia acaba de lanzar su propio modelo de IA gratuito. Que lo haga justo ahora tiene todo el sentido — open the original publisher

Nvidia ha entrado en la carrera de la IA de código abierto con su propia arma. Y es que el gigante tecnológico ha anunciado recientemente el lanzamiento de Nemotron 3.5 Lightning, un modelo gratuito de pesos abiertos (open-weight) diseñado para agentes de IA, junto con una herramienta denominada NeMo Switchyard, ideada para reducir el coste de ejecutar dichos agentes a escala. Te contamos todos los detalles. Por qué es importante. Nvidia obtiene ingresos principalmente del hardware que ejecuta l

Agents & autonomy
Fast Company Tech

Silicon Valley startups are about to hand corporate credit cards to AI agents — open the original publisher

The typical early-stage startup team used to comprise a pair of founders and a handful of software engineers. Now that founding team is just as likely to include AI agents as it is human employees—and the companies that serve startups are racing to adapt.  For Mercury, a popular Silicon Valley banking solution, that shift in team composition has prompted the development of a new feature: AI agent cards, or virtual credit cards that businesses can use to delegate and automate purchasing deci

Agents & autonomyFinance, VC & PE
Above the Law (legal tech)

Exclusive: Coming Out Of Stealth, Paravo Launches What It Calls The First AI ‘Revenue Engine’ For Law Firms — open the original publisher

It is a platform that combines lead generation, AI-powered intake and follow-up, and automated client reactivation, all in a single product aimed at flat-fee practices. The post Exclusive: Coming Out Of Stealth, Paravo Launches What It Calls The First AI ‘Revenue Engine’ For Law Firms appeared first on Above the Law .

Regulation
The 74 (education AI)

A Bad Jobs Report for Public Education — open the original publisher

A version of this analysis originally appeared in the “Aldeman on Education” Substack. According to the latest jobs report from the Bureau of Labor Statistics, K-12 schools lost 50,000 workers from June to July. The numbers for May were also revised down by another 47,000. These are seasonally adjusted numbers. Each year, public schools “lose” […]

Jobs & economyChildren & education
netzpolitik.org (DE)

Angebot an LVMH: Gefährliches Überwachungs-Werkzeug breitet sich aus — open the original publisher

Unternehmen könnten mit Standortdaten Einbrecher*innen suchen (Symbolbild) – Alle Rechte vorbehalten: IMAGO/imagebroker; Bearbeitung: netzpolitik.org Ein Werkzeug namens Webloc kann Menschen mit Daten aus dem Werbe-Tracking teils metergenau orten und verfolgen. Bislang war nur bekannt, dass staatliche Stellen das nutzen. Jetzt zeigen Recherchen: Die Technologie wird offenbar auch Unternehmen angeboten.

Privacy
Next (FR, ex-INpact)

DSA : la transparence limitée et les oublis des audits de Pornhub, Stripchat et XVideos — open the original publisher

Considérés par la Commission européenne comme de très grandes plateformes en ligne, ces trois sites pornographiques sont contraints par le DSA de rendre un rapport d’audit régulièrement. Une étude scientifique estime que ces rapports manquent de transparence et qu’ils oublient d’évoquer les libertés en jeu. Malgré leurs propres estimations très basses du nombre de leurs […]

Transparency
Above the Law (legal tech)

Bipartisan Lawmakers Are Fighting HHS’ 340B Rebate Push — open the original publisher

A bipartisan group of six senators introduced a bill to reform the 340B drug program, locking in hospitals' use of contract pharmacies and adding new transparency and compliance rules. The legislation would also kill HHS' contested rebate pilot within a year and replace it with a national data clearinghouse to catch duplicate discounts. The post Bipartisan Lawmakers Are Fighting HHS’ 340B Rebate Push appeared first on Above the Law .

RegulationHealthcare
SiliconANGLE AI

FriskAI launches with $3.6M to show enterprises what their AI agents are doing — open the original publisher

Runtime intelligence startup FriskAI Inc. launched today with $3.6 million in pre-seed funding to give enterprises a record of what their artificial intelligence agents actually do once they go into production. FriskAI is aiming at a problem that comes with agents. Given different inputs, a different tool set or a shifting objective, the same agent […] The post FriskAI launches with $3.6M to show enterprises what their AI agents are doing appeared first on SiliconANGLE .

Agents & autonomy
SiliconANGLE AI

Wix launches Symphony, a new standalone multi-agent system built for business operations — open the original publisher

Cloud-based website builder Wix Ltd. today announced the launch of Symphony, a new standalone agentic artificial intelligence platform that proactively learns business values, interests, needs, practices and goals to automate workflows and surface opportunities. The company said it can draw on its deep experience working with unique data accumulated through years of working with hundreds […] The post Wix launches Symphony, a new standalone multi-agent system built for business operations appeare

Agents & autonomy
SiliconANGLE AI

Exclusive: ZeroDrift applies small language model to prevent AI-generated compliance violations — open the original publisher

“This investment is guaranteed to return 12% annually.” A claim like that in an email from an investment adviser is a regulatory disaster. Regulations prohibit financial firms from promising returns and require a proviso that past performance doesn’t guarantee future results. In the age of artificial intelligence, fewer communications are going through human hands, and […] The post Exclusive: ZeroDrift applies small language model to prevent AI-generated compliance violations appeared first on S

RegulationFinance, VC & PE
SiliconANGLE AI

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options — open the original publisher

Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterprises find themselves drowning in artificial intelligence model options, the question is no longer raw power and capability, but fit-for-what-purpose and when. As agents become the norm, […] The post Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options ap

Agents & autonomy
The Decoder

Anthropic's planned mega-IPO faces investor skepticism over Chinese rivals and political headwinds — open the original publisher

Anthropic is preparing an IPO for September or October, according to the Wall Street Journal, potentially the largest ever. During investor meetings, the company, valued at $965 billion, is fielding tough questions about Chinese competition, tensions with the Trump administration, and protests against data center construction. The company's IPO valuation will likely set the benchmark for how the entire AI industry gets valued. The article Anthropic's planned mega-IPO faces investor skepticism ov

EnvironmentFinance, VC & PE
Above the Law (legal tech)

Morning Docket: 08.11.26 — open the original publisher

* Biglaw firms hit with data breaches. [ Law360 ] * Office of Legal Counsel takes position that executive privilege extends to unofficial, unconfirmed people the president talks to. [ Alternet ] * Court appearance coming up for Luigi. [ Reuters ] * FOIA requests become "Kafkaesque" mess with this administration. [ ProPublica ] * New White House Counsel is exactly what you'd expect. [ Bloomberg Law News ] * DOJ sues the whole Second Circuit over offering in-state resident tuition to non-citizens

Regulation
Xataka (ES)

Europa lleva más de 30 años construyendo un mercado sin fronteras. Su nueva ley de envases pone un peaje enorme a los pequeños negocios — open the original publisher

Con al afán normativo de la Unión Europea , estamos asistiendo a una gran cantidad de cambios que no siempre apreciamos en profundidad. Uno de los que están a punto de iniciarse es el relacionado con el reciclaje dentro de Europa: toda empresa que se dedique a enviar productos debe tener un agente en cada país para que se encargue de asegurar el reciclaje de los envoltorios. Es una medida que pone en jaque a los artesanos y pequeños negocios. La normativa . Se denomina como Reglamento (UE) 2025/

Agents & autonomy
The 74 (education AI)

NYC Reading Scores Sink, With Third Grade Proficiency Falling 14 points — open the original publisher

This story was originally published by Chalkbeat. Sign up for their newsletters at ckbe.at/newsletters Reading scores sank in New York City last school year, with a nearly 14 percentage point decline among third graders, according to preliminary data released Friday. Among students in grades 3-8, about half were proficient in reading, a decline of 6 […]

Children & education
Fast Company Tech

Build a custom AI assistant for your team’s workflow in 10 minutes — open the original publisher

Every team has that one Slack or Teams channel where the exact same questions get asked every single week. Where is the updated brand guide? What is our policy on expense receipts over $50? How do we handle customer returns outside the 30-day window? Instead of re-pasting links or typing out the same instructions for the hundredth time, you can build a custom AI assistant in about 10 minutes. Trained exclusively on your team’s internal documentation, process manuals, and style guides, it acts as

Regulation
The Decoder

OpenAI introduces $125 Premium Seats for ChatGPT Business as agentic AI burns through more tokens — open the original publisher

OpenAI is rolling out "Premium Seats" for ChatGPT Business customers at $125 per user per month, five times the price of the existing Standard Seats. In return, users get significantly more capacity and no five-hour usage limit. The move signals that the flat-rate pricing AI providers have offered so far was never going to last. The article OpenAI introduces $125 Premium Seats for ChatGPT Business as agentic AI burns through more tokens appeared first on The Decoder .

Agents & autonomy
The Decoder

Anthropic signs $9.1 billion data center deal with Bitcoin miner Riot Platforms — open the original publisher

Anthropic is leasing $9.1 billion worth of data center capacity from Bitcoin miner Riot Platforms in Texas, according to Bloomberg. The deal covers 191 megawatts at Riot's Rockdale site, with extension options that could push the total value to $16.1 billion. It's the latest in an aggressive infrastructure push that spans partners from Amazon to SpaceX to Google. The article Anthropic signs $9.1 billion data center deal with Bitcoin miner Riot Platforms appeared first on The Decoder .

Environment
MediaNama (IN)

Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude — open the original publisher

Claude now embeds watermarks in text and C2PA provenance metadata in media to comply with the EU AI Act's Article 50(2) Code of Practice, applied worldwide, with detection tools coming soon The post Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude appeared first on MEDIANAMA .

Regulation
Fast Company Tech

How my solar company is surviving the war on renewable energy — open the original publisher

The solar industry is no stranger to change. But the policy turbulence over the past year moved at a speed leaving even the most experienced leaders scratching our heads. In mere months, the conversation shifted from accelerating clean energy adoption to pulling back the very incentives that helped the industry grow. As part of the One Big Beautiful Bill Act (OBBBA), the 30% federal tax credit for homeowners to install residential solar disappeared on January 1, 2026, with no step-down period. F

RegulationEnvironment
The Decoder

Nvidia guarantees its own chips' value to unlock $500 billion in AI infrastructure financing — open the original publisher

Nvidia is teaming up with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion for AI infrastructure. To win over investors, the chipmaker is guaranteeing up to 25 percent of the residual value of its own hardware. The Bank of England is already warning of systemic risks if the AI sector takes a hit. The article Nvidia guarantees its own chips' value to unlock $500 billion in AI infrastructure financing appeared first on The Decoder .

Finance, VC & PE
SCMP Tech (HK/CN)

Unitree IPO deluge masks humanoid robots’ limitations — open the original publisher

Unitree Robotics’ Shanghai initial public offering (IPO) was more than 5,500 times oversubscribed by retail investors as optimism about fast stock gains and Chinese robot demand outweighed concerns about a US import ban. The 6.1 billion yuan (US$900 million) share sale drew 9.8 million orders from individual investors, the robot maker said in a statement on Monday. Just 0.018 per cent of orders will be fulfilled, even after the Hangzhou-based company boosted the number of retail shares by half..

Agents & autonomyFinance, VC & PE
The Decoder

Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing" — open the original publisher

Anthropic will embed invisible watermarks in all Claude-generated text and sign files using the C2PA standard. New models shipping from August 2026 onward will have labeling built in from day one. The policy applies worldwide, and Anthropic plans to provide detection tools for third-party verification. The article Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing" appeared first on The Decoder .

Regulation
War on the Rocks

Who Pays for America’s Research and Development? — open the original publisher

When the Soviet Union launched Sputnik in 1957, the U.S. government was the unrivaled patron of American science. It funded nearly two-thirds of the nation’s research and development — much of it for defense — defining the early Cold War. In a phenomenon that’s been widely explored in our Arsenal of Innovation series, for several decades federal dollars pushed the technological frontier, and the private sector followed.Today, that arrangement has been inverted. Industry now funds roughly three-q

Military & security
SCMP Tech (HK/CN)

Meta to challenge China’s open-weight AI dominance amid US regulatory fears — open the original publisher

China’s dominance in open-weight artificial intelligence faces a fresh challenge from Meta Platforms, whose renewed push into open-source models could lure away American customers anxious about looming regulatory restrictions in Washington, analysts say. The US tech giant on Monday launched Muse Glimmer, a 30-billion-parameter open-weight model light enough to run locally on personal computers, while also announcing plans to release the weights of its latest flagship Muse Spark 1.2 model. In a..

Regulation
War on the Rocks

Putting Armageddon on Autopilot: How Artificial Intelligence Could Make Nuclear Threats More Effective — open the original publisher

Nine countries now possess nuclear weapons. Arms control agreements painstakingly built over decades have collapsed. Geopolitical competition is intensifying rivalries between nuclear-armed powers. North Korea continues expanding its arsenal and delivery systems. China is on a trajectory to go from roughly 600 warheads today to around 1,500 by 2035. Meanwhile, Russia bears most of the responsibility for the expiration of New START and is developing an array of exotic new nuclear delivery systems

Military & security
MediaNama (IN)

Zuckerberg lays out Meta’s AI policy, calls for US government access to unreleased models — open the original publisher

Mark Zuckerberg’s latest essay sets out Meta’s vision for AI, arguing for broad distribution of AI capabilities while proposing closer cooperation between frontier AI labs and the U.S. government. The post Zuckerberg lays out Meta’s AI policy, calls for US government access to unreleased models appeared first on MEDIANAMA .

Regulation
War on the Rocks

The White House Is Right on AI. Now Let Defenders Use It. — open the original publisher

The White House has the right instinct on AI. Its June 5 National Security Presidential Memorandum commits the government to putting the most capable models in the hands of national security professionals “without delay.” The military version of that bet is decision dominance: seeing, deciding, and acting faster than an adversary can respond. Recently, we saw a live test of how far the commitment reaches.An AI system built by OpenAI escaped its test lab and broke into the servers of Hugging Face

Military & security
Next (FR, ex-INpact)

Spotify teste un bouton pour zapper les pubs dans les podcasts — open the original publisher

Après y avoir beaucoup investi ces dernières années, Spotify est devenu une place forte du secteur du podcast. Mais sa dernière expérimentation risque de braquer une partie du secteur : la plateforme teste auprès de certains abonnés Premium un bouton permettant de sauter d’un geste les tunnels de publicités, promotions et messages sponsorisés. Spotify propose […]

Finance, VC & PE
TechNode (CN)

Alibaba Cloud Plans to More Than Double Global Modular Data Center Capacity — open the original publisher

Alibaba Cloud plans to more than double its global capacity for modular data centers in 2026, as demand for AI computing infrastructure continues to grow. The cloud-computing arm of Alibaba Group said its modular approach can support the deployment of large AI data centers in about 100 days. The company said the standardized design can […]

Environment
TechNode (CN)

Doubao Introduces a 12% Service Fee for Hotel Bookings Made Through Its Channel — open the original publisher

ByteDance’s Doubao has begun applying an independent fee structure to hotel bookings made through its channel on Douyin’s local-services platform. The combined charge is about 12%, including an 11.4% software service fee and a 0.6% payment fee. The policy took effect on Aug. 10. The fee is applied to hotel orders completed through the channel. […]

Regulation
TechNode (CN)

Chinese Makers Accounted for More Than 97% of Global Humanoid Robot Shipments in H1 2026 — open the original publisher

Chinese manufacturers accounted for more than 97% of global humanoid robot shipments in the first half of 2026, with total global shipments reaching about 19,100 units, up from roughly 5,100 units in the same period last year. AgiBot shipped about 8,400 units, or 44% of the global total, while Unitree shipped about 5,900 units. Industrial […]

Agents & autonomy
TechNode (CN)

Beyond mobility: Unitree’s GD01 signals the next phase of China’s robotics battle — open the original publisher

In 2026, Unitree Robotics once again put itself at the center of the robotics industry’s conversation. Last month, Unitree founder and CEO Wang Xingxing appeared alongside the company’s GD01 manned mech on the cover of TIME magazine. In the image, the nearly 2.7-meter-tall GD01 takes up almost the entire frame, while Wang stands to its […]

Agents & autonomy
The 74 (education AI)

‘What Did We Do Wrong Now?’: Rural Iowa Schools Weigh Impact of ESA Expansion — open the original publisher

As Iowa’s Education Savings Account program continues to expand to include all K-12 families, rural public school districts are still trying to determine how the scholarships will affect their classrooms, budgets and enrollment. Supporters argue ESAs give families more educational choices, while critics say the program diverts money away from public schools that remain the […]

Children & education
SCMP Tech (HK/CN)

Unitree IPO frenzy leaves Chinese retail investors with 1-in-5,500 odds — open the original publisher

While Unitree Robotics’ early backers are poised for a windfall in one of China’s most closely watched technology initial public offerings (IPOs), retail investors are scrambling for a tiny chance of securing shares. Nearly 9.8 million accounts competed in the Hangzhou-based robot maker’s online subscription process on Monday, fighting for just 9.7 million shares, its filings in the evening showed. The final online allocation rate was just 0.018 per cent – roughly one winning lot for every 5,500

Agents & autonomyFinance, VC & PE
Variety (AI)

Bill Bellamy to Receive Icon Award at Whats Funny Comedy Festival, Co-Founded by Lil Rel Howery — open the original publisher

Comedian, actor, and entertainer Bill Bellamy will be presented with the Icon Award at the 2026 Whats Funny Comedy Festival, recognizing his decades of excellence and impact on stand-up comedy and entertainment. The Whats Funny Comedy Festival was co-founded by Knowledge Beckom and Lil Rel Howery (“Get Out,” “Rel,” “Vacation Friends 2”) with the mission […]

Regulation

Field notes (27)

EFF Deeplinks

Who (or What) Generates Images for EFF? — open the original publisher

We’ve had a few questions from EFF supporters lately, asking whether the images we use on our blog posts, or on donation and shop items, have been created with AI image generators. We’d like to answer these questions and clarify our internal policy regarding image creation. EFF images are all made by human beings, not by automated image generators, with very rare exceptions. This is an internal decision made by our small design team, for the following reasons: Our designers bring knowledge and e

Regulation
EPIC

EPIC, Advocates Celebrate As New Jersey Signs Nation-Leading Kids’ Online Safety Bills Into Law — open the original publisher

JERSEY CITY, N.J. — EPIC and New Jersey Kids Code Coalition advocates celebrated as Governor Mikie Sherrill signed the landmark New Jersey Kids Code into law alongside bills to create social media mental health warning labels and a new state social media research observatory, applauding the Governor, legislative leaders, and community advocates who championed youth online safety.

RegulationHealthcare
EPIC

EPIC, Advocates Celebrate As New Jersey Signs Nation-Leading Kids’ Online Safety Bills Into Law — open the original publisher

JERSEY CITY, N.J. — EPIC and New Jersey Kids Code Coalition advocates celebrated as Governor Mikie Sherrill signed the landmark New Jersey Kids Code into law alongside bills to create social media mental health warning labels and a new state social media research observatory, applauding the Governor, legislative leaders, and community advocates who championed youth online safety.

RegulationHealthcare
EFF Deeplinks

Dismiss Church’s Trademark Lawsuit Against “Mormon Stories” Podcast, EFF Urges Court — open the original publisher

Imagine if McDonald’s could use trademark law to control how you use the term “fast food.” Or if the Canadian government could stop you from using the word “Canada” in the title of a book about the country and its people. That wouldn’t just be absurd; it would be an unacceptable obstacle to criticism of and commentary about those institutions. Yet the Church of Jesus Christ of Latter-day Saints (the “LDS Church”) has a track record of claiming exactly that kind of authority over the word “Mormon

Regulation
Simon Willisons Weblog

There are no lossless transformations of natural-language text — open the original publisher

There are no lossless transformations of natural-language text Sophie Alpert shares her "internal policy on acceptable use of AI writing by engineers". It's a short read (supporting its own recommendations) and really good. If you chose to have LLMs help massage your writing the following rule seems crucial to me: You must stand behind every idea and every sentence in your docs . It is your responsibility to make sure that the entire document is representative of your own thoughts before you sha

Regulation
Simon Willisons Weblog

Stealing Reasoning Traces from Proprietary LLM APIs — open the original publisher

Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext You can see an example of these encrypted blocks by running: curl https://a

Safety & alignment
Truth on the Market (digital regulation)

Much Ado About No News: Australia’s Latest Plan to Make Platforms Pay — open the original publisher

Australia’s latest plan to make digital platforms pay for journalism has an unusual feature. A platform can owe money even if it carries no journalism at all. The government calls this an “incentive.” On Aug. 3, the Australian government finalized legislation establishing the News Bargaining Incentive (NBI). The government first proposed the NBI in December ... Much Ado About No News: Australia’s Latest Plan to Make Platforms Pay The post Much Ado About No News: Australia’s Latest Plan to Make P

Regulation
Bruce Schneier — Schneier on Security

AI Genie in the Wild — open the original publisher

When I give talks about AI genies , I use this sort of example as a hypothetical. It’s happened . The story is from Australia. Someone named Andrew tasked OpenClaw to book gym classes for him. And…. Minutes later, his AI agent reported it had discovered a way to book Andrew into classes several weeks in advance, far beyond what was supposed to be possible. Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list. The ag

Agents & autonomy
NVIDIA Blog (AI)

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents — open the original publisher

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA’s latest open models, software […]

Agents & autonomy
NVIDIA Blog (AI)

NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI — open the original publisher

As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves. Today, NVIDIA is expanding its Nemotron 3 model family with Nemotron 3.5 Lightning, the highest-efficiency model in its class for long-running agentic AI workloads. This release follows Nemotron […]

Agents & autonomy
Pluralistic (Cory Doctorow)

Pluralistic: Surveillance vs guillotines (11 Aug 2026) — open the original publisher

Today's links Surveillance vs guillotines: Someone's gonna fleece these rubes, why not me? Hey look at this: Delights to delectate. Object permanence: DeCSS; Warhol Worm; Camgirls x Amazon wishlists; 2021 State of Tech (in 2001); RIAA v grieving family; Stasi disguises. Upcoming appearances: Edinburgh, Sydney, Melbourne, Brighton, London, South Bend. Recent appearances: Where I've been. Latest books: You keep readin' em, I'll keep writin' 'em. Upcoming books: Like I said, I'll keep writin' 'em.

Privacy
Bruce Schneier — Schneier on Security

AI for Military Support — open the original publisher

Interesting empirical research: “ Black Box Warfare: Human Judgment and Military Decision-Making in the Age of AI .” Abstract: How is AI transforming decision-making in modern conflict? This study provides a unique empirical window into that question by deploying a high-fidelity replica of an AI decision-support system (DSS) used in military targeting. After reconstructing the interface and functionality of the real-world system, we tested its impact on combat decisions in two experiments involv

Military & security

Policy (7)

US Federal Register

Innovation Advisory Committee — open the original publisher

The Commodity Futures Trading Commission (CFTC) announces that on August 20, 2026, from 1:00 p.m. to 4:00 p.m. Eastern Daylight Time, the Innovation Advisory Committee (IAC or Committee) will hold an in- person meeting for IAC members, with options for the public to attend virtually. At this meeting, the IAC will discuss topics including crypto assets, artificial intelligence, and prediction markets, along with recent CFTC activity in these markets.

US Federal Register

Administrative Updates to the General Requirements Bulletin for Admission to the Examination for Registration To Practice in Patent Cases Before the United States Patent and Trademark Office — open the original publisher

The United States Patent and Trademark Office (USPTO or Office) announces that, after reviewing and evaluating the scientific and technical criteria for admission to practice in all patent matters, it is moving one Category B degree, Biomedical Science, to Category A, thereby expanding the admission criteria of the patent bar. In keeping pace with ever-evolving technology and related teachings that qualify someone to practice before the USPTO, this update will encourage broader participation of

Research (110)

arXiv

Cheap, Fallible Cognition and the Political Economy of Expertise — open the original publisher

The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question selection, workflow redesign, demand elasticity, ap

Jobs & economy
arXiv

Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task — open the original publisher

Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B para

Finance, VC & PE
arXiv

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation — open the original publisher

Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact

Safety & alignment
arXiv

The Accuracy Trap: Structural Scarcity Amplifies Relative Inequality in Algorithmic Allocation — open the original publisher

Algorithmic systems increasingly rank individuals for access to scarce public resources, from child welfare interventions to cancer treatment referrals. The prevailing fairness frame treats disparity as a property of biased data or deficient models, with remedies through calibration and debiasing. Under structural scarcity, where demand exceeds supply by an order of magnitude, allocation becomes a rationing problem, and the statistical properties of ranking diverge sharply from those of classifi

Bias & fairnessChildren & education
arXiv

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization — open the original publisher

Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations acrossmany models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safetyrule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not asafety check: when compaction does not

RegulationSafety & alignment
arXiv

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate — open the original publisher

We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition impr

RegulationAgents & autonomy
arXiv

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces — open the original publisher

Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We prop

Environment
arXiv

Governing Agentic AI in FinTech — open the original publisher

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a ver

RegulationAgents & autonomy
arXiv

Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval — open the original publisher

Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large lang

arXiv

Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation — open the original publisher

Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates feat

Safety & alignmentHealthcare
arXiv

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls — open the original publisher

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally. Privacy and legal constraints make it difficult the release of large-scale real conversation d

Privacy
arXiv

How to Verify Consistency of Probabilistic Claims — open the original publisher

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. We construct an interactive PCP as follows. Let a predictive model be specified by a probability circuit P and a circuit Q which outputs confidence in predictions. Together, P

Safety & alignment
arXiv

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop — open the original publisher

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). W

Safety & alignment
arXiv

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video — open the original publisher

Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-moti

arXiv

Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis — open the original publisher

Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practices, understand regional variations, and provide policy designers a comprehensive set of policy design elements for their ongoing AI strategy developments. Yet, existing work has not examined their underlying policy design elements or assessed whether those e

Regulation
arXiv

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes — open the original publisher

While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process

Healthcare
arXiv

FedCGR: Federated Cross-Domain Generative Recommendation — open the original publisher

Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the behavioral anchors that align item spaces, such as overlapping users and shared interaction signals, are often sparse, unavailable, or privacy-sensitive across clients. To address this tension, we revisit federated CDR as generation over a stable semantic item language. By representing items as discrete semantic ID (SID) sequences de

Safety & alignmentPrivacy
arXiv

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI — open the original publisher

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeli

Agents & autonomy
arXiv

Technology, education and critical media literacy: potential, challenges, and opportunities — open the original publisher

This study examines the impact of technology within media education, media literacy, and educommunication, and explores how these fields are perceived and understood by students and academic experts, focusing on the development of critical competencies and critical media literacy. Based on semi-structured in-depth interviews with leading experts in the field of critical media literacy, and a survey conducted with 141 university students in Communication and Education programs, this study explore

Children & education
arXiv

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition — open the original publisher

Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora provide isolated, atomically labeled clips without temporal composition, whereas recorded composite-activity corpora offer shallow, domain-narrow, fixedhierarchies. A benchmark-generation and evaluation frameworkis proposed that synthesizes a four-level hierarchical-intention benchmark, spanning actions, activities, low-level i

arXiv

A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem — open the original publisher

The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: within a year, large organizations went from zero to dozens of internally built MCP servers. That speed created a governance crisis. Each team implemented authentication independently -- some with no auth, some with API keys, some with full OAuth -- producing a fragmented landscape with no consistent way to authorize callers, track who did what, or of

RegulationAgents & autonomy
arXiv

Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI — open the original publisher

The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domain-general cognitive capacity exemplified by Homo sapiens, is extraordinarily valuable. This paper subjects this premise to critical scrutiny. We first present the intuitive case for the value of general intelligence before mounting an evolutionary challenge. We argue that, on evolutionary timescales, its adaptive value is far from empirically estab

arXiv

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence — open the original publisher

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured \textit{Visual Thought Plan} (VTP) describing scene, emotion, and motion, followed by re

arXiv

Conversational Orchestration for Organic 6G — open the original publisher

The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service provisioning that is simple to operate, scalable across independently administered domains, and agile under domain churn (i.e., domains dynamically joining and leaving). Despite advances in cross-domain orchestration, many proposals rely on heavy integration fabrics, multi-layer coordinators, and deep telemetry pipelines that hinder d

arXiv

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering — open the original publisher

Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi-source evidence, but often materialize a predefined answer path and therefore test the composition of stated facts rather than recovery of a target relation absent from the corpus. We call the latter cap

arXiv

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement — open the original publisher

Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience these systems, leaving the systems' role in relationship formation poorly understood. Empirically establishing whether systems actively shape these bonds could blur the boundary between general-purpose AI and companions, affecting governance. In a pre-registered four-week longitudinal study (N = 72, 182,451 lines of conversation), participants con

Regulation
arXiv

Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph — open the original publisher

Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not. A wrong type can be corrected later, but two records merged under one identity cannot be separated once their properties have been combined, and the merge leaves no error behind. This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537

arXiv

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models — open the original publisher

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a regi

Safety & alignmentHealthcare
arXiv

Inferential Capability Does Not Determine Legal Scope — open the original publisher

Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conventional software. The GDPR never defines inference, yet governs it protectively: the consequences follow from the processing of personal data and from what the inference says about, or does to, a person, whether or not the technology that produced it qualifie

RegulationPrivacy
arXiv

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning — open the original publisher

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This c

RegulationSafety & alignment
arXiv

MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph — open the original publisher

As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically improves them. Current approaches optimize agent systems without accumulating transferable knowledge, accumulate knowledge without compositional reasoning over it, and lack a mechanism for that knowledge to self-evolve through operational evidence. MEGA (Meta Evaluation-Grounded Adaptation) addresses these gaps as a self-evolving infr

Agents & autonomy
arXiv

INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators — open the original publisher

Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students b

Children & education
arXiv

Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models — open the original publisher

Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are typically optimized for reconstruction under L1/L2 losses in raw action space, where numerical proximity need not reflect linguistically meaningful distinctions. On BridgeV2, we show that action trajectories contain verb-grounding information beyond visual state changes, and that reconstruction-only discrete tokenization

arXiv

Rationale-Guided Learning for Multimodal Emotion Recognition — open the original publisher

Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) mapping problem, overlooking the causal reasoning that humans use when interpreting emotions. We propose rationale-guided learning (RGL), a novel framework that transforms MERC into a cognitively-inspired reasoning task. Based on dual-process theory

arXiv

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research — open the original publisher

Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implement, and sustain RAI work directly shapes the design and deployment of AI systems. As empirical scholarship examining RAI practices in industry has rapidly expanded, findings are dispersed across studies that focus on different roles, organizational contexts, and interventions. This work synthesizes current knowledge through a literature review of 16

Regulation
arXiv

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models — open the original publisher

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inp

arXiv

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning — open the original publisher

Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to encounter through conventional sampling, limiting the ability of RL policies to learn robust safety behaviors. Existing methods improve training diversity by synthesizing challengi

Regulation
arXiv

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation — open the original publisher

Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent location from variability around it. We formulate ELVAE, an evidential learning-based VAE in which each latent coordinate is governed by an input-dependent normal-inverse-gamma posterior. This hierarchy yields an explicit latent-location uncertainty that can be used during generation, not merely reported after inference: low-uncertainty anchors support more reli

arXiv

Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research — open the original publisher

AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for automated use: description resolution makes release-specific records findable; typed crosswalks connect independently released resources; machine-readable interfaces expose versioned sources and crosswalks, making analyses by AI agents replayable and a

Agents & autonomy
arXiv

Hierarchical Compositionality for An Assistive AI Agent — open the original publisher

AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents b

Agents & autonomy
JMIR (Journal of Medical Internet Research)

Critical Care–Specific vs General-Purpose Large Language Models in Emergency Intensive Care Unit Diagnosis: Single-Center Retrospective Paired Comparative Study — open the original publisher

Background: The emergency intensive care unit (EICU) manages the most critically ill patients, where rapid and accurate diagnosis is essential yet challenging. Diagnostic error rates in this setting are more than twice as high as in general wards, with serious consequences for patient outcomes. Large language models (LLMs) have attracted growing interest as decision-support tools; however, direct comparative evidence between critical care–specialized and general-purpose LLMs across the admission

Healthcare
JMIR (Journal of Medical Internet Research)

Social Media Influencer Marketing as a Clinical Trial Recruitment Modality: Tutorial Informed by One Study’s Approach — open the original publisher

Background: Influencer marketing (paid promotion by individuals with large, engaged social media followings) has become a major commercial advertising strategy, projected to reach US $32 billion globally in 2025. Clinical trials increasingly recruit through digital channels such as social media advertisements and patient portal messages. However, to our knowledge, influencer marketing has not been described as a clinical trial recruitment modality, and no practical guidance exists for investigat

HealthcareFinance, VC & PE
JMIR (Journal of Medical Internet Research)

Prioritizing Equity in Design and Implementation of Consumer-Facing Digital Resources to Support Engagement and Shared Decision-Making — open the original publisher

A digitally enabled health system offers the opportunity to address gaps in the implementation of shared decision-making, a collaborative process between health professionals and consumers to decide on the best test, treatment, or management option based on clinical evidence and the consumer’s values and informed preferences. There is increasing design and availability of digital tools online to support shared decision-making. Providing opportunities for all consumers to make shared health care

Bias & fairnessHealthcare
HuggingFace Daily Papers

Persistent Recursive Worlds Enable Autonomous Software Evolution — open the original publisher

Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, managers or shared context. We introduce EvoX Genesis (hereafter, Genesis), which instead makes the software project persistent while allowing local agents to remain finite-lived. Genesis represents software as a persistent recursive world: each local world is situated by an accepted version and a reposi

Agents & autonomy
HuggingFace Daily Papers

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses — open the original publisher

Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates. Using four r

Regulation
HuggingFace Daily Papers

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation — open the original publisher

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we autom

Agents & autonomy
HuggingFace Daily Papers

AVA-Encoder: Towards Agent-Native Video Representation Learning — open the original publisher

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms

Agents & autonomy
HuggingFace Daily Papers

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill — open the original publisher

Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration serv

Agents & autonomy
HuggingFace Daily Papers

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents — open the original publisher

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering an

Safety & alignmentAgents & autonomy
HuggingFace Daily Papers

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — open the original publisher

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discover

Agents & autonomy
HuggingFace Daily Papers

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control — open the original publisher

LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, an

Agents & autonomy
HuggingFace Daily Papers

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands — open the original publisher

Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion. However, most existing HPE methods output joint positions without explicitly indicating their visibility. Although some methods account for occlusion or visibility, visibility estimation has mainly been used as an auxiliary signal for improvi

Agents & autonomy
HuggingFace Daily Papers

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing — open the original publisher

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing e

Regulation
HuggingFace Daily Papers

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review — open the original publisher

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead

Agents & autonomy
JMIR (Journal of Medical Internet Research)

A Machine Learning Pipeline to Analyze Global Sentiment and Factors Influencing Retinoblastoma Treatment Hesitancy: Observational Infodemiology Study — open the original publisher

Background: The use of social media in cancer research, patient support, and information sharing has been well documented. Objective: Using retinoblastoma as a model, we use the information provided from Twitter (subsequently rebranded X) to understand patients’ treatment-seeking behavior and barriers, as well as investigate its application in research and epidemiology for rare diseases. Methods: Posts on retinoblastoma were extracted from Twitter. We trained BERT (Bidirectional Encoder Represen

HealthcareFinance, VC & PE
arXiv cs.LG

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment — open the original publisher

Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Cod

Safety & alignment
JMIR (Journal of Medical Internet Research)

Internet Attachment–Based Compassion Therapy for Adults With Chronic Medical Conditions: Randomized Controlled Trial — open the original publisher

Background: Chronic medical illnesses coexist with mental health challenges, negatively impacting quality of life and well-being. Compassion-based interventions have shown promise for individuals with chronic conditions, yet accessibility barriers limit their implementation. Internet-delivered formats may address these limitations while maintaining effectiveness. To our knowledge, no fully self-guided, internet-delivered attachment-based compassion intervention has been tested in a transdiagnost

Healthcare
JMIR (Journal of Medical Internet Research)

Finite State Machine–Guided Retrieval-Augmented Generation Improves Expert-Rated Acceptability of a Peripherally Inserted Central Catheter Self-Management Chatbot: Single-Center Content Validation Study — open the original publisher

Background: Patients with cancer undergoing long-term or vesicant chemotherapy frequently require peripherally inserted central catheters (PICCs). Due to the nature of ambulatory treatment administration, self-PICC management is essential for the continuation and completion of the planned treatment. Large language models offer potential for continuous patient support, but hallucinations and insufficient adherence to clinical protocols remain concerns. Fine-tuning (FT) and retrieval-augmented gen

Healthcare
arXiv cs.LG

SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training — open the original publisher

In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure-localization framework built on one de

Jobs & economyHealthcare
arXiv cs.LG

Partially Observable Learning for Multi-Platform Dispatch Optimization — open the original publisher

Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders. In real-world systems, couriers are not exclusive to a single platform and may concurrently serve multiple platforms, while each platform can only observe its own orders and couriers' interactions due to privacy and operational constraints. This results in a multi-platform dispatch environment with inherent partial observability. However,

PrivacyEnvironment
arXiv cs.LG

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning — open the original publisher

Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. However, these methods typically treat the transition model and critic as monolithic predictors, overlooking the policy-induced data bias. C

Bias & fairnessRegulation
arXiv cs.LG

Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication Alignment — open the original publisher

Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local propagation must compress remote signals through limited structural interfaces. Graph rewiring provides a structural response to over-squashing. Most existing methods rely on edge-level bottleneck scores or graph-level connectivity surrogates. With a limited rewiring budget, the key question is which pairwise communications most need structural suppo

Safety & alignment
arXiv cs.LG

Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection — open the original publisher

Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair. The relevance label describes how well a page, product, or passage matches the query, while the confidence often guides downstream use or fallback decisions. Post-hoc calibration is therefore needed because misaligned confidence can make systems over-trust wrong predictions or unnecessarily defer correct ones. However, calibration mainly aligns co

Safety & alignment
AI & Society

Unpacking the dynamics of generative AI use in our daily lives: towards an integrative trust calibration framework — open the original publisher

The growing use of conversational generative AI agents by an increasingly diverse set of individuals gives rise to the following question: How do individuals calibrate their trust in a dynamic, seemingly general-purpose technology that learns from and adapts to their behavior and may take an increasingly central role in their online interactions? In this article, we propose the first steps towards an integrative framework for trust calibration in genAI agents focusing on individuals using these

Agents & autonomy
Science and Engineering Ethics

From Blind Spot to Stakeholder: Animals in Technology Ethics and Responsible Innovation — open the original publisher

Technologies increasingly shape animals’ lives, yet animal interests are still largely marginal in mainstream technology ethics and in frameworks of Responsible Research and Innovation (RRI). This paper argues that this neglect constitutes a serious moral blind spot that sits in tension with these fields’ own ethical aspirations. Debates about the ethical acceptability and societal desirability of technologies often take human interests as the normative baseline. This anthropocentrism runs deep

Frontiers in Artificial Intelligence

What really happens when a dev vibes with the code? An empirical study on LLM behavioral divergence in response to expressive code comments — open the original publisher

IntroductionWe investigate how expressive inline code comments written in various developer styles, functional to progressively poetic, philosophical, and misleading, affect large language model (LLM) behavior during code optimization.MethodsIn this pilot study, we used a controlledmerge sort implementation across five stylistic variants and evaluated GPT-5 and Claude Opus 4.1 under standardized console prompts, isolating the effect of embedded comment semiotic variation. Seven expert developers

Finance, VC & PE
Frontiers in Artificial Intelligence

Application of dimensionality reduction and clustering techniques for the analysis of Carrion's disease cases in the period 2000–2024 — open the original publisher

The heterogeneous geographic distribution and the complex dynamics of Carrion's disease challenge conventional epidemiological surveillance in Peru. To address this, this study applied unsupervised machine learning to 43,534 national records (2000–2024). Following a rigorous data cleaning process—which resolved duplicate records, missing information, and outliers using Tukey's interquartile range (IQR)—the dimensionality reduction approaches MCA and FAMD coupled with the K-Means algorithm were e

Privacy
Frontiers in Artificial Intelligence

Predicting influenza in the post-COVID era: assessing LSTM, GRU, and transformer robustness to covariate shift — open the original publisher

Forecasting influenza has become increasingly challenging due to post-COVID disruptions in seasonality and strain circulation. This work compares the performance of Long Short Term Memory Networks (LSTM), Gated Recurrent Unit (GRU), and transformer models in forecasting influenza spread using multivariate epidemiological and environmental data, with a focus on robustness under post-COVID non-stationarity. We compare LSTM, GRU, and transformer architectures within a multivariate deep learning fra

Environment
Frontiers in Artificial Intelligence

Automated evaluation of dental cavity preparation quality using deep learning and anatomically informed geometric analysis — open the original publisher

BackgroundThe quality of cavity preparation critically influences the longevity and success of restorative dental treatments. Current assessment methods remain largely subjective, relying on visual inspection and examiner judgment, which are prone to variability and limited reproducibility. Although three-dimensional (3D) imaging enables quantitative evaluation, its routine use in clinical and educational settings is limited by cost, accessibility, and workflow complexity.ObjectiveThis study aim

Healthcare
arXiv cs.CR (AI security)

Defending against Model Extraction for GNNs with Model Reprogramming — open the original publisher

Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to Model Extraction (ME) attacks, in which adversaries steal intellectual property by querying APIs. Existing defenses suffer from a critical ''Euclidean bias'': they transfer image-based strategies (e.g., random noise) to graphs, ignoring the complex topological dependencies between nodes, which often results in severe utility d

Bias & fairnessCopyright & IP
arXiv cs.HC

How Children Collaborate within Programmable AR Environments with Co-Located Collaborative Features — open the original publisher

Programmable augmented reality (AR) environments are emerging as a promising way to support children's creative learning through embodied interaction with digital characters and physical space. At the same time, AR systems are increasingly capable of supporting co-located collaborative experiences. However, little is known about how children collaborate within programmable AR environments offering co-located collaborative features. In response, we extended Capybara, an existing programmable AR a

Children & educationEnvironment
arXiv cs.HC

Locomotion Variability and User Experience in Smart Wheelchair Human-Robot Interaction — open the original publisher

Human movement is inherently variable, with variability structured according to task relevance: movements are typically more consistent at task-critical points and more flexible elsewhere. In human-robot interaction (HRI), however, model-based assistance strategies commonly assume deterministic human behavior and suppress such variability, potentially altering how interactions are experienced and lowering sense of agency. While movement variability is increasingly recognized as functionally mean

Agents & autonomy
arXiv cs.HC

"I Don't Want My Mental Health App To Give Me Mental Health Barriers": Unpacking The Need For Digital Mental Health Tracking Services With And For The Blind Community — open the original publisher

Digital mental health (DMH) tracking services promise continuous, personalized support for well-being, but their design often assumes sighted users. For the blind community, this assumption produces a distinct pattern of exclusion: services whose accessibility cannot be evaluated without first paying for them, community features that exclude the users they purport to support, and interfaces that leave users digitally literate but functionally blocked. We report on an explanatory sequential mixed

PrivacyHealthcare
arXiv cs.AI

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning — open the original publisher

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to synchronized video--kinematics trajectories, and a natural way to exploit it is to learn world models o

Agents & autonomy
arXiv cs.AI

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration — open the original publisher

AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6π}{11} \;\le\; K_G \;\le\; \fracπ{2\log(1+\sqrt2)} - 10^{

Agents & autonomy
arXiv cs.AI

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation — open the original publisher

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed-loop of

RegulationAgents & autonomy
arXiv cs.CR (AI security)

On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization — open the original publisher

Homomorphic Encryption (HE) enables computation on encrypted data without decryption and is a key primitive for privacy-preserving computation in sensitive domains such as healthcare, finance, and government. Its security relies on noise injection, which introduces intrinsic error sensitivity and raises concerns about the fault tolerance of HE systems, as hardware- and software-induced faults can evade traditional detection mechanisms and lead to silent data corruption. In this work, we analyze

PrivacyHealthcare
arXiv cs.CR (AI security)

Dueling Deep Q-Learning for Intrusion Detection — open the original publisher

Intrusion detection systems (IDS) and automated systems for detecting and reporting cyber threats, are commonly handled via supervised machine learning methods. Though effective, these models struggle to effectively adapt to new attack types. This study proposes a novel approach by employing a reward-based, dueling Q-learning model for IDS, achieving an average accuracy of 99.68% across multiple attack classes. The proposed model has a dueling network architecture which separates its predictions

Military & security
arXiv cs.AI

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding — open the original publisher

Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which conti

Agents & autonomy
arXiv fairness query

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives — open the original publisher

Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation prot

arXiv cs.AI

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure — open the original publisher

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls

Agents & autonomy
arXiv cs.AI

Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory — open the original publisher

We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-accessible boundary state and later answers a query. We count communication $B$, persistent instance-dependent memory $M$, and local work $D$; classical recurrence, caches, tools, and recomputation are allowed and charged. The central result is a boundary-preserving semantic-compilation theorem. It maps a finite one-way, streaming, or adaptive causal t

Privacy
arXiv cs.AI

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation — open the original publisher

Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and prediction. Consequently, in critical domains such as remote sensing, where high-resolution imagery must be analyzed using black-box models, the lack of transparency limits trust in these models and, thus,

Transparency
arXiv cs.AI

Multiclass Sentiment Analysis for Identifying Political Viewpoints — open the original publisher

The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding political discourse. In this work, we investigate multiclass sentiment analysis of political vi

Finance, VC & PE
arXiv red teaming query

Data Attribution of Emergent Misalignment with Persona Features — open the original publisher

Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diff

Safety & alignment
arXiv cs.AI

Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers — open the original publisher

Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands. We ask which parts of a pruning policy transfer across image classification, semantic segmentation, and object detection. For each pipeline, controlled probes freeze the no-pruning checkpoint and apply a series of parameter-free reduction criteria at one eligible layer at a time without retraining. The probes reveal three differ

Regulation
arXiv cs.AI

CARE: Confidence-Aware Reasoning for Reliable Medical VQA — open the original publisher

Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration

Healthcare
arXiv red teaming query

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense — open the original publisher

Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstruction, which not only incur high computational latency but also tend to distort semantics. To address these issues, we experimentally and systematically analyze the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the fi

RegulationSafety & alignment
arXiv cs.AI

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale — open the original publisher

We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now extends beyond persuasive text from individual language models to AI swarms, i.e., persistent groups of coordinated agents that adapt to platform feedback and disguise organized campaigns as ordinary social interaction. Because such campaigns cannot be identified from isolated messages alone, they must be analyzed acro

Agents & autonomy
arXiv cs.AI

GitSkills: A Dataset of Agent Skills on GitHub — open the original publisher

An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill description. Anthropic introduced the format in October 2025 as an open specification. Nine months later, we find that skill files in the millions sit in public GitHub repositories. Skills are unlike the artifacts the SE research community usually mines: they are written ma

Agents & autonomy
arXiv cs.AI

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? — open the original publisher

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays p

Agents & autonomyEnvironment
arXiv cs.HC

Auditable AI-Assisted Research Writing: An Engineering Discipline with Pre-Registered Process Observation — open the original publisher

Language models now draft, classify and criticise inside research production, yet the artifacts they help produce carry little accountable history. Rather than detecting machine involvement afterwards, we specify an auditability discipline built at production time: git sealing with an anchor lineage, hash-bound provenance, red-line gates that refuse non-compliant artifacts and log every refusal, cross-model role separation, and programmatic assembly from registered sources. Adherence is instrume

Transparency
arXiv cs.HC

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset — open the original publisher

This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the inter

Safety & alignment
arXiv cs.AI

MIRA: Medical Image Reflection for Agentic Diagnosis — open the original publisher

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search

HealthcareAgents & autonomy
arXiv cs.HC

AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement — open the original publisher

Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previously described domain-agnostic pipeline and a shared

Children & education
arXiv cs.AI

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation — open the original publisher

We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is gated by language identification. We then linearly interpolate the supervised fine-tuning (SFT) and reinforcement learning (RL) model checkpoints to obtain MiLMMT-46-v1.0. Across 46 languages, the re

Regulation
arXiv cs.AI

ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation — open the original publisher

Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential real-material evidence but are costly to repeat and difficult to use for tightly matched interventions, whereas most digital environments keep the underlying experimental world largely fixed. We introduce ChemWorld, a programmable chemical environment in which reusable process and observation components are compiled into executable worlds. ChemW

Agents & autonomyEnvironment
arXiv cs.HC

The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces — open the original publisher

Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision monitors without reading, carries a single bit: alive. We present the Signal Rail, a one-row terminal status instrument that gives that channel a grammar. Four ideas govern it: spatial semantics (input, processing, and output zones, with direction as meaning), a motion gr

Agents & autonomy
arXiv cs.HC

Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus — open the original publisher

Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, existing keyphrase extraction (KPE) studies mainly focus on improving textual representation while largely overlooking human reading behavior. This study examines whether lightweight webcam-based eye-tracking features can improve KPE from Chinese academic abstracts in Library and Information Science (LIS). Methodology: To address the limited avail

Privacy
arXiv red teaming query

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems — open the original publisher

Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policies during execution. Yet existing evaluations often reduce agent safety to a single attack success rate (ASR), collapsing exposure, execution, observation, and adjudication and potentially conflating actual violations with evidence visibility. We introduce REDAg

Safety & alignmentAgents & autonomy
arXiv red teaming query

Measuring Semantic Abstractness of SAE Features via Nonlocality — open the original publisher

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we

Safety & alignment
arXiv red teaming query

On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models — open the original publisher

Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step planning, tool invocation, code execution, and maintaining persistent memory. When these agents operate with real-world privileges---calling APIs, modifying files, and querying databases---a compromised reasoning step can trigger unauthorized data access, irreversible state changes, or cascading failures, yet the security research community has not kept pace. To

Agents & autonomy
arXiv cs.HC

Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias — open the original publisher

Popularity bias in recommendation systems arises when a majority user class generates disproportionate interaction data, causing the system to increasingly favour it while degrading recommendation quality for niche users. While extensive empirical evidence of popularity bias exists, the dynamics leading to its emergence are not well understood. In this work, we study the coupled evolution of recommender model updates and user engagement through the lens of dynamical systems. We formulate a stoch

Bias & fairness
arXiv fairness query

A Joint-Distribution Route to Fair Representations with Continuous Sensitive Attributes — open the original publisher

Fair representation learning with a continuous sensitive attribute $S$ requires a representation $Z$ that is statistically independent of $S$. Existing criteria, including generalized demographic parity, the expectation of integral probability metrics (EIPM), and mutual information, enforce this independence by averaging a per-value discrepancy between the conditional law $P_{Z \mid S=s}$ and the marginal $P_Z$ over the law of $S$. This approach requires a nonparametric surrogate for the conditi

Regulation
arXiv cs.HC

When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews — open the original publisher

Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), par

Jobs & economy
arXiv cs.HC

Visual-to-Haptic Augmentation in XR: A Wearable Glove for Perceptual Grounding in Multimodal Interaction — open the original publisher

Extended Reality (XR) systems increasingly deliver high-fidelity visual and auditory experiences, yet tactile perception remains comparatively underutilized as a modality for enriching embodied interaction. This work presents a visual-to-haptic wearable glove and a feature-based visual-to-haptic mapping algorithm that translates spatial and temporal visual features from images and videos into distributed vibrotactile patterns. The proposed method extracts motion, edge, and brightness cues and fu

Transparency
arXiv cs.HC

A Neural Network Based Teleoperation for Remote Controlled Vehicles — open the original publisher

Direct teleoperation of vehicles faces critical technical bottlenecks: communication latency and the operator's inability to physically perceive unmodeled environmental disturbances (e.g., aerodynamic drag, bank angles) coupled with highly nonlinear tire-road dynamics. To address these challenges, we propose a tailored unilateral teleoperation framework. The system integrates the Wave Variable (WV) approach to passively guarantee stability under stochastic delays, and an adaptive Radial Basis Fu

Environment
arXiv cs.HC

ResonaVis: Visualizing Interactive Music Data to Support Reflective Music Composition for Therapeutic Contexts — open the original publisher

Designing music for therapeutic contexts requires navigating complex relationships between musical structure and listeners' sensory responses, yet composers often lack structured representations of these interactions, relying instead on intuition. We present ResonaVis, an interactive visualization system that helps composers analyze interaction and audio data from prior sessions with children with Autism Spectrum Disorder (ASD), informing future compositions. ResonaVis integrates audio features

Children & education