Archive · 2026-07-13
AI ethics on Monday, 13 July 2026
310 items published this day, across 5 categories.
Incidents (1)
News (143)
What Anthropic’s latest AI discovery does—and doesn’t—show
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. It’s looking into whether AI models can feel pain, for example,…
Can Labor save us from the risks of AI? – podcast
The AI revolution is here, and with it a fear that soon it will replace many of us in the workplace. The Australian government is grappling with how to deal with the multi-layered disruption but so far reform has been slow as it weighs up regulation against the claims of investment opportunities an AI boom presents. Could that change on Wednesday when the prime minister delivers a landmark speech addressing the government’s approach to the technology? The chief political correspondent, Dan Jervi
Albanese to compare pivotal moment in AI to renewable energy transition as he outlines approach
Labor sources say the PM will discuss safety concerns in speech this week but will not provide an update on copyright reforms to protect artists Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast Anthony Albanese will describe the progress of AI as an inflection point for society on par with the renewable energy transition, but is not expected to detail progress on copyright reforms to protect creative industries. The prime ministe
Assigning Responsibility When AI Discriminates Against Job Applicants
The Public Rejects OMB's Federal Financial Assistance Rule
Wall Street’s Big Week for Earnings and Economic Data
Earnings season kicks off in earnest on Tuesday with big banks up first. Corporate America faces a high bar to beat investor expectations.
Workers in Asia Are Fighting for Protections as AI Threatens Jobs
The New York nurses replaced by AI: ‘It should concern every patient who cares about quality of care’
The union for 12 nurses laid off by Montefiore hospital say company broke contract they recently won through a strike Marilyn Shuler has worked as a utilization review nurse for 39 years at Montefiore hospital in the Bronx in New York City, helping to read patient charts and communicate with insurance companies over coverage. After nearly four decades in her job, Shuler is one of 12 nurses who were laid off Sunday after being replaced with AI-powered software, according to the New York State Nur
Senator Warner Makes a First Foray into Agentic AI Regulation
Europe Wants Platforms to Prove They Are Safe for Children
China’s massive AI rollout - podcast
Senior China correspondent Amy Hawkins on China’s embrace of AI, from medical avatars to food delivery drones and state surveillance While the spread of AI has been met perhaps with a lot of scepticism in the west, China has fully embraced the technology, explains Amy Hawkins , from millions of users talking to AI doctors, to the use of intelligent robots in factories, and drones delivering food on the Great Wall of China. AI has also been eagerly taken up by the state, not least in the opportun
More than 50% of Australian university assignments used AI. How should unis respond?
A big challenge for universities is distinguishing whether students are using AI to help or as a substitute for learning.
The 6 wildest claims in Apple’s lawsuit against OpenAI
When Apple employees interviewed for jobs at OpenAI, the AI startup's hardware head allegedly asked them to show up with something unusual: components they were working on and unreleased product samples. That's according to a blockbuster lawsuit filed by Apple, which accuses OpenAI of stealing confidential documents, spying on hardware prototypes, and tricking one of […]
The AI Arms Race in Technical Interviews Is Escalating
Software engineering jobs are under threat from AI . Some applicants are fighting back by using AI in the interview process, employing AI assistants that suggest responses on the fly during remote technical interviews. Meanwhile, some employers are countering with—you guessed it—AI. They’re applying AI-powered tools to detect telltale signs of AI use during interviews. This two-sided dynamic is turning hiring into an AI arms race with no clear winners. Yet as interviewers and interviewees naviga
Now, defenders are embracing the prompt injection, too
"Context bombing" tricks hacking agents into shutting down before they can do harm.
Despite the growth of some AI schools like Alpha, research doesn’t show that AI tutors are better than human teachers
Instead of schools trying to replace teachers with AI, teachers could use AI to become better educators.
Building a Foundation Stack for General-Purpose Robots
This article is brought to you by X Square Robot . Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the fi
Hermes agent maker Nous Research in talks for new funding at $1.5B valuation
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Here’s how America builds again
We need new blueprints to scale production, workforce capability, and industrial competitiveness simultaneously.
Here’s how America builds again
We need new blueprints to scale production, workforce capability, and industrial competitiveness simultaneously.
Here’s how America builds again
We need new blueprints to scale production, workforce capability, and industrial competitiveness simultaneously.
Here’s how America builds again
We need new blueprints to scale production, workforce capability, and industrial competitiveness simultaneously.
Here’s how America builds again
We need new blueprints to scale production, workforce capability, and industrial competitiveness simultaneously.
How MIT students are helping to prevent cyberattacks
Students from the MIT Cybersecurity Clinic help local governments and other vulnerable organizations defend against digital threats.
AI agents create virtual playgrounds to help robots get crucial training data
“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.
The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
Apple’s trade secrets lawsuit against OpenAI contains allegations that range from employees joking about unauthorized access to Apple’s systems to claims that job candidates were asked to bring Apple hardware to interviews. Here are the complaint’s most eye-catching claims.
Sam Altman’s space data center trash talk is what most experts already believe
Responding to Musk accusing him of being a scammer, Altman said, "homeboy you're the one sellling [sic] public market investors on short-term space datacenters."
Massive AI spending is driving up prices on laptops and electricity, as the Fed watches closely
Investment in AI data centers could exceed $700 billion in 2026, pushing up inflation.
Massive AI spending is driving up prices on laptops and electricity, as the Fed watches closely
Investment in AI data centers could exceed $700 billion in 2026, pushing up inflation.
Massive AI spending is driving up prices on laptops and electricity, as the Fed watches closely
Investment in AI data centers could exceed $700 billion in 2026, pushing up inflation.
New method aims to keep kids safe from illegal AI-generated content
Researchers developed an auditing technique to test generative AI models for malicious capabilities, without prompting them for illegal outputs.
AI agents are becoming the enterprise's most privileged users. And most organisations don't know it yet
AI agents are outpacing identity governance.
Spatial AI startup Augmodo raises $21M to expand beyond retail stores
Spatial artificial intelligence startup Augmodo Inc. said today it has raised $21 million in an impromptu funding round that lifts its valuation to $350 million. TQ Ventures served as the lead investors, with participation from Lerer Hippeau, Jefferson River Capital, Arena Holdings, Chemist Warehouse, New Fare, Interlace and Webb Investment Network The Seattle-based startup is […] The post Spatial AI startup Augmodo raises $21M to expand beyond retail stores appeared first on SiliconANGLE .
‘Trump v. Trump’ Did Not Impress The Judge — See Also
Benchslapped By His Own Petard : Judge nixes Trump v. IRS settlement citing administration's unitary executive nonsense . Todd Blanche Loses Lindsey Graham : Contentious confirmation hearing set to begin this week will have to go ahead without Lindsey Graham after the senator died over the weekend (perhaps he'll be dealing with Graham's sister, who has been tapped to fill out his term ). Yale Law Lost The Top Spot In The 2026 U.S. News Rankings : And the school's placement is even worse in this
The Explainability Trade-Off Was Never Real
By Swagatam Sen, Founder & CEO, ControlOne Financial crime teams have been offered a choice for ...
Pennsylvania Senate Advances Bill Allowing 18-Month Data Center Moratoriums
Lawmakers must weigh three competing state bills proposing bans ranging from 180 days to three years.
Biglaw Firm Is Growing Its Sports Law Practice
And hiring an insider to hard launch the practice. The post Biglaw Firm Is Growing Its Sports Law Practice appeared first on Above the Law .
Yale Law Students Advise Their Own University: Knock Off This ‘Obey In Advance’ Nonsense
Over 150 students and 26 organizations told Yale to stop negotiating away the rule of law. The post Yale Law Students Advise Their Own University: Knock Off This ‘Obey In Advance’ Nonsense appeared first on Above the Law .
States Tell FCC to Leave Pole Attachment Authority Intact
Utility regulators in Washington, Ohio, and Connecticut call FCC’s proposed recertification unnecessary.
Carolyn Herzog On Why Legal Leaders Need Optimism, Not Certainty
The old legal instinct to pause, assess, and draw hard boundaries no longer works. The post Carolyn Herzog On Why Legal Leaders Need Optimism, Not Certainty appeared first on Above the Law .
Weak Security Continues to Fuel Russian Cyberattacks
In a first, the UK and the EU jointly imposed sanctions on Russian individuals and entities for cyberattacks and disinformation campaigns in the region.
States are building their own election defense networks as federal support evaporates
Election officials are facing an impossible choice: follow federal directives they don’t trust, or risk becoming targets of a criminal investigation. The post States are building their own election defense networks as federal support evaporates appeared first on CyberScoop .
DOD suspends CMMC Phase 2, launches 60-day ‘reform’ review
Citing prohibitive costs for small and mid-size contractors, the Defense Department will keep Phase I self-assessments in place while a new task force studies the cyber and supply chain security program's future.
EU leaders eye social media ban for children under age 13
“While ultimately it is up to parents to decide when children get their first smartphones, what we already have is a consensus that there needs to be a start date for the age children can join social media,” says European Commission President Ursula van der Leyen.
Prominent economists, tech executives call for new regulatory approach to AI
A group of economists and tech industry figures has called on policymakers to address the risks posed by artificial intelligence more directly. The signatories outlined their concerns in a public letter published today. The initiative was organized by economics professors Erik Brynjolfsson, Ajay Agrawal, Anton Korinek and Tom Cunningham. More than half of the 200-plus […] The post Prominent economists, tech executives call for new regulatory approach to AI appeared first on SiliconANGLE .
Arguments That Embarrass Only The Advocate
Thinking about issues is a good idea. Making silly arguments that reflect poorly on your intelligence is a less good idea. The post Arguments That Embarrass Only The Advocate appeared first on Above the Law .
Microsoft carbon emissions rose 25 percent in 2025 as AI data centers grow
Microsoft reported a 25 percent jump in carbon emissions in 2025 as it sets out to grow its use of artificial intelligence (AI) data centers. The emissions increase was "driven primarily by the expansion of our datacenter infrastructure and pausing our use of non-additional, unbundled renewable energy certificates as we prioritize investments that bring net...
SpaceX stock sinks for a second-straight day, nearing $135 IPO price
SpaceX went public a month ago in a record IPO. Elon Musk's space and AI company was added to the Nasdaq-100 last week.
Reditus readies first launch of its re-entry vehicle/hypersonic target
The Missile Defense Agency is evaluating the company’s ENOS spacecraft as a hypersonic target/testbed under the SHIELD contract, said Reditus CEO Stef Crum.
Los Angeles law enforcement will stop using Flock cameras
The Los Angeles police department did not renew its contract with Flock due to data privacy concerns.
Rebecca Slaughter Has A Message For Yale: Grow A Spine
The former FTC commissioner isn't buying the argument that a $44 billion university has no choice but to settle. The post Rebecca Slaughter Has A Message For Yale: Grow A Spine appeared first on Above the Law .
U.S. military uses Corsair maritime drones to attack Iran
The robotic platforms, which are built by Saronic, were recently employed in an attack against Iran for the first time. The post U.S. military uses Corsair maritime drones to attack Iran appeared first on DefenseScoop .
Pentagon announces ‘immediate suspension’ of CMMC Phase II mandates
Top Pentagon officials said as currently executed, CMMC is too prohibitively burdensome on the Defense Industrial Base.
Judge Cites Supreme Court’s Newfound Unitary Executive Theory To Blow Up Trump’s IRS Settlement
If the President really is the entire executive branch, then he can't be on both sides of the 'v.' The post Judge Cites Supreme Court’s Newfound Unitary Executive Theory To Blow Up Trump’s IRS Settlement appeared first on Above the Law .
Summer camps remain a battleground over what it means to be American
Defense Secretary Pete Hegseth's demand that Scouting America drop diversity initiatives revives a fight settled in 2019, when it began admitting girls.
Meta's Louisiana data center investment to reach $50 billion, aided by generous tax incentives
Meta said the planned Hyperion data center supercluster in Richland Parish, Louisiana, will be a 5GW facility and cost more than $50 billion.
As Gas Plants Rise to Power AI, Renewable Energy Allies are Fighting For Cleaner Alternatives
wealthy companies putting billions of dollars into data centers can afford to build renewable energy sources to power them.
Well, This Doesn’t Help Todd Blanche
Lindsey Graham's death ahead of this week's confirmation hearing leaves a very conspicuous empty chair. The post Well, This Doesn’t Help Todd Blanche appeared first on Above the Law .
How Medicaid agencies can prepare for community engagement requirements
Explore practical strategies for improving Medicaid beneficiary engagement, streamlining exemption verification and strengthening audit readiness.
Maine, 14 Other States Sue Trump Administration to Block School Mental Health Funding Cuts
Maine joined 15 states on Friday in suing the Trump administration to prevent millions of dollars in cuts to school-based mental health funding. The new lawsuit is part of an ongoing legal battle between Democratic-led states and the U.S. Department of Education over a mental health grant program that Congress established following the 2018 school […]
'Yellow Teams' Are Defining the Future of AI Security
In some companies, engineers are building defense and attack tools to test the potential of artificial intelligence for cybersecurity — and its threat.
In first, US uses sea drones in combat in Iran strikes: CENTCOM
The US military posted a video of the unmanned surface vessels appearing to approach docks and explode.
Shein executive chairman to step down as IPO nears completion, sources say
Shein Executive Chairman Donald Tang will step down as his mission of taking the company public nears completion, three sources with direct knowledge of the matter said on Monday, retreating to an advisory role after three years as the public face of the global fast-fashion retailer. A Chinese-American billionaire who began his career in banking, Tang has acted as the Western proxy of secretive Shein founder Sky Xu, liaising with politicians and regulators around the world while also...
A ‘Projection’ Of The 2027 U.S. News Law School Rankings
Which law schools are excluded from the T14 in this version of the rankings? And what happened to Yale? The post A ‘Projection’ Of The 2027 U.S. News Law School Rankings appeared first on Above the Law .
Turing Award winner Rich Sutton founds Oak Lab to build AI agents that learn on their own
Richard Sutton, 2024 Turing Award winner and co-founder of modern reinforcement learning, has launched a new startup called Oak Lab in Toronto. He calls current deep learning methods "weak and inefficient" and wants to build AI agents that learn continuously from their environment. The article Turing Award winner Rich Sutton founds Oak Lab to build AI agents that learn on their own appeared first on The Decoder .
Legal Ethics Roundup: Lawyer ‘Negligence’ For Not Using AI, Cameras At SCOTUS, Law School Laptop Ban & More
Your tour of all things related to lawyer and judicial ethics, with University of Houston law professor Renee Knake Jefferson. The post Legal Ethics Roundup: Lawyer ‘Negligence’ For Not Using AI, Cameras At SCOTUS, Law School Laptop Ban & More appeared first on Above the Law .
Meta expanding plans for its largest data center
Meta will expand its largest data center to 5 gigawatts of compute capacity as investment in the project hits more than $50 billion, the company announced Monday. The Hyperion data center in Richland Parish, La., was announced in October and was originally projected to cost more than $27 billion as part of Meta's joint venture...
Trump calls for Congress to pass Clarity Act crypto bill to honor Lindsey Graham
The Senate Banking Committee approved the bill 15-9 in May, with two Democrats joining Republicans to advance the legislation.
Narmi releases AI to streamline account opening for communitty banks and credit unions
Narmi, a leading digital banking platform provider for banks and credit unions, today announced the upcoming launch of AI Decision Assist, a new agentic AI capability designed to help financial institutions automate and accelerate account opening reviews while still maintaining control, transparency, and compliance.
Massive AI spending is driving up prices on laptops and electricity, as the Fed watches closely
American consumers — and the Federal Reserve — are being hit with another high-cost headache. The gusher of investment in data centers — likely topping $700 billion this year — to power artificial intelligence has made memory chips, computer processors and other equipment, as well as electricity, more expensive. Economists expect it will continue to push up inflation at least through the end of this year. While it won’t be as large a spike as occurred in 2021-2023, when inflation peaked at 9.1%,
States Try New Measures To Get Chronically Absent Students Back to Class
This year, at least six states enacted laws trying to reduce the number of students chronically absent from school. The measures include requiring monitoring of absences and publicly releasing data, developing new guidance on the best ways to address the problem and increasing punishments for parents and guardians of chronically absent students. Chronic absenteeism is […]
How a Teacher Revived Backyard Baseball
Backyard Baseball, a favorite game for 1990s kids, had been off the market for years. But Lindsay Barnett was determined.
Europe strikes out against Russia’s Turla over espionage, ‘destructive attacks’
The EU, its members and the U.K. took action against Russian government officials and others while attributing the winter cyberattacks against Poland’s energy grid to the FSB. The post Europe strikes out against Russia’s Turla over espionage, ‘destructive attacks’ appeared first on CyberScoop .
Larry David Casts His Biglaw Lawyer In New Show
Latham partner Andrew Clubok appears as Senator Potter in Life, Larry, and the Pursuit of Unhappiness. The post Larry David Casts His Biglaw Lawyer In New Show appeared first on Above the Law .
S&P stuft Oracle auf BBB- herab – nur noch eine Stufe über Ramschniveau
S&P Global hat Oracles Rating wegen massiver KI-Investitionen herabgestuft. Der Konzern ist nur noch eine Stufe vom spekulativen Bereich entfernt.
Nobel laureates and AI leaders warn the window to prepare for AI's economic impact is closing fast
More than 200 economists and AI researchers, including 16 Nobel laureates and representatives from Google, OpenAI, and Anthropic, are calling for immediate action in a coordinated statement. The AI transformation could surpass the Industrial Revolution but unfold in a fraction of the time. The paper doesn't propose concrete measures, and studies so far have found no significant AI-driven effects on the labor market. The article Nobel laureates and AI leaders warn the window to prepare for AI's e
Trump urges Senate to pass crypto bill in honor of Graham
President Trump called on the Senate to pass a cryptocurrency regulation bill Monday in honor of the late Sen. Lindsey Graham (R-S.C.), who unexpectedly died over the weekend. “In honor of Senator Lindsey Graham, a big supporter, the U.S. Senate should pass the Clarity Act,” Trump wrote in a post on Truth Social. “China, and...
This tech stock just got another Wall Street boost. Why we're on the same page
The Investing Club holds its "Morning Meeting" every weekday at 10:20 a.m. ET.
Overstock Loon Patrick Byrne Loses To Himself In Court, Must Pay Hunter Biden
RIP to a real one. The post Overstock Loon Patrick Byrne Loses To Himself In Court, Must Pay Hunter Biden appeared first on Above the Law .
The Alters and Frostpunk developer 11 Bit Studios is laying off 20 employees
The studio said it has managed to mitigate the number of job cuts by transferring employees to other internal teams.
Warren rips Senate leaders over crypto bill's lack of ethics restrictions
Sen. Elizabeth Warren (Mass.), the top Democrat on the Senate Banking Committee, slammed Senate leaders Monday over the lack of ethics restrictions in a cryptocurrency regulation bill that is poised to hit the Senate floor in the coming weeks. Warren, a longtime crypto skeptic, raised concerns about President Trump’s recent financial disclosures, which showed he...
Pentagon disburses Havana Syndrome compensation, rebrands team focused on ‘Directed Energy Bio-Effects’
The Anomalous Health Incidents cross-functional team has been renamed as officials focus on "non-kinetic threats." The post Pentagon disburses Havana Syndrome compensation, rebrands team focused on ‘Directed Energy Bio-Effects’ appeared first on DefenseScoop .
Officials once again warn defenders that Russian hackers are targeting network devices
State-sponsored attackers are targeting critical infrastructure networks in defense, communications, energy, finance, government and health care. The post Officials once again warn defenders that Russian hackers are targeting network devices appeared first on CyberScoop .
Judge dismisses Epidemic Sound’s second copyright lawsuit against Meta – but leaves the door open to amend
US District Judge Jacqueline Scott Corley granted Meta's motion to dismiss on Friday, July 10 Source
Hundreds of economists say ‘we must act now’ on AI’s economic impact and job displacement risks
Hundreds of economists urge immediate action to address AI's potential impact on the economy. In an open letter released Monday, they warn that AI could transform the economy and displace many jobs.
WhatsApp Usernames in India: Privacy Upgrade or New Fraud Surface for Digital Finance?
WhatsApp usernames may reduce phone number exposure, but in India, they also raise a deeper question...
Surprise, Surprise: More Evidence That What You Say To Your Chatbot Isn’t Always Private
Those of us in the legal profession have a responsibility to sound the alarm about what our clients and the public as a whole put in chatbots The post Surprise, Surprise: More Evidence That What You Say To Your Chatbot Isn’t Always Private appeared first on Above the Law .
Gegenwind für Bundesregierung: Mehr als eine halbe Million Menschen wollen Informationsfreiheit retten
Sollte die Bundesregierung die Pläne umsetzen, wird es für Bürger:innen und Presse noch schwieriger an staatliche Dokumente zu kommen. (Symbolbild) – Gemeinfrei-ähnlich freigegeben durch unsplash.com: Anastassia Anufrieva Damit hat Schwarz-Rot offenbar nicht gerechnet: Heftige Kritik am Angriff auf die staatliche Transparenz kommt nicht nur von der Opposition, sondern aus der Koalition selbst. Dazu erreicht eine Petition gegen das Vorhaben bemerkenswerten Zulauf.
US quantum computing needs a national buyer
Washington’s demand signal could do for the technology what early defence contracts did for silicon
Opinion: Small Changes, Big Relief: How States Can Support School Districts
State education agencies are being asked to do something they have rarely been asked to do before: lead. As the federal government’s influence over education recedes, leaving confusion in its wake, calls for guidance, clarity and strategic direction are shifting to states. And they are shifting fast, to agencies that are often understaffed, under-resourced and […]
Nadella calls out AI labs like OpenAI and Anthropic for banning distillation while training on everyone else's data
Microsoft CEO Satya Nadella is calling out OpenAI and Anthropic for what he calls a "reverse information paradox." They train on public data under fair use but ban distillation of their own models, all while learning from customer interactions. Nadella wants companies to control their own learning infrastructure. Microsoft, of course, sells exactly that. The article Nadella calls out AI labs like OpenAI and Anthropic for banning distillation while training on everyone else's data appeared first
The European Union’s chief is considering social media restrictions for kids under 13
A top European Union official on Monday called for limits to be placed on children using social media as a special EU panel looking into the challenge recommended forbidding access for those under 13 until tech companies can prove their platforms are safe. Growing awareness of the dangers social media poses for young, developing brains has shown up in a wave of new restrictions globally. Australia , the U.K. , Turkey, Indonesia and others have passed bans on kids under 16 or 15 from using platfo
The Path to Sovereign Data: Challenges and Priorities in Local-First Computing
A panel on data ownership challenged the definition of "ownership," arguing it must extend beyond simple account control to include structural independence, interoperability, and community governance. Speakers like Zenna Fiscella, Paul Frazee, Boris Mann, and Robin Berjon emphasised the need for shared standards, unbundled platforms, and better tools to support user sovereignty. By Olimpiu Pop
How DoorDash Built an AI Shopping Assistant That Doesn’t Rely on the LLM Alone
DoorDash details the architecture behind Ask DoorDash, its AI-powered conversational shopping assistant, combining LLMs, specialized AI agents, MCP-based tooling, and an intelligence layer with persistent consumer memory and live backend data. Early results show up to 24% higher checkout conversion, 17% larger baskets, and improved intent accuracy using memory-backed sessions. By Leela Kumili
Marine Serre Enters Receivership, Seeks Investor
The Paris Commercial Tribunal granted its request for court protection. The post Marine Serre Enters Receivership, Seeks Investor appeared first on Above the Law .
Europe’s military advantage depends on sovereign command of the ground truth
[Sponsored] As NATO and European governments accelerate investment in space-based intelligence, decision advantage will depend on how quickly sovereign and commercial capabilities can be fused into a trusted operational picture.
What an ex-NSA red teamer wants every SOC to stop doing
Security teams have spent years trying to see more. More endpoints, more cloud services, more identities, more telemetry. For the The post What an ex-NSA red teamer wants every SOC to stop doing appeared first on The New Stack .
Orange Rag Legal Tech Clinic: Day one thinking, for law firms that aren’t on day one
“Should we be looking at Harvey? Legora? Maybe Claude? Should we buy or build? Could we run something open source, and what’s the overhead if we did?” And almost every time, in some form or other: […] The post Orange Rag Legal Tech Clinic: Day one thinking, for law firms that aren’t on day one appeared first on Legal IT Insider .
China works on AI safety benchmark as regulators target large model risks
China’s Ministry of Industry and Information Technology (MIIT) has started building a safety benchmark to evaluate artificial intelligence models, as regulators in the United States and Europe strengthen oversight of AI security. The MIIT-led National Industrial Information Security Development Research Centre is now recruiting companies and experts to co-build the benchmark, with applications due on Tuesday, according to a notice published on Monday. The institute said that current frameworks..
Morning Docket: 07.13.26
* Lindsey Graham died over the weekend, removing a key Trump ally from the Judiciary Committee in advance of Todd Blanche's already controversial nomination hearings. [ PBS ] * Civil rights coalition calls for Senate to reject Blanche. [ Ms ] * Audit reveals the broken California alternative bar exam process. [ ABA Journal ] * Judges embark on whistlestop tour to explain the increasing threats against the judiciary. [ Washington Post ] * DOJ opens investigation into UAW president. If only a work
UK regulates Microsoft, Google, Amazon in finance sector; India sticks to indirect oversight
Cloud computing companies Google, Microsoft, AWS & Oracle have been brought under finance sector regulation as "critical third parties" as UK banks face outage or cyberattack risks due to their reliance on the big four. The post UK regulates Microsoft, Google, Amazon in finance sector; India sticks to indirect oversight appeared first on MEDIANAMA .
Europe Takes Step Toward Possible Social Media Ban for Children
After the release of a new report, the European Commission is considering changing the rules across the 27-nation bloc.
Exclusive: BUILT Bags $2 Mn To Bring ‘Natural Movement’ Footwear To India
D2C footwear startup BUILT has raised $2 Mn (about ₹17 Cr) in a pre-seed funding round from Singapore-based VC firm…
Auterion, Ukrainian drone-maker Skyfall to supply 50,000 FPVs
Skyfall said that the Shrike has destroyed a wide range of Russian military assets, with hits on targets including an Mi-8 helicopter, armored vehicles, electronic warfare and artillery systems and a TOS-1A SoIntsepyok heavy flamethrower.
Turning the Tables on Email Scammers With 'ScamBuster'
An open source, AI-driven system adopts victim personas to engage with phishing attackers, allowing organizations and law enforcement to gather relevant data on cybercriminal operations.
I loved ChatGPT Desktop until OpenAI gutted it to make room for Codex and Work
OpenAI just merged the ChatGPT desktop app with Codex - and removed my favorite productivity features. What were they thinking?
Opinion: The Science of Reading Goes to High School
For anyone who cares about student literacy, the past few years have given us reason to cheer. While trends in national test scores, young people’s reading habits, and talk of a “learning recession” are clear causes for concern, there’s another side to the story. More than 40 states now mandate evidence-based reading instruction in public […]
Iran strikes, Lindsey Graham, Apple takes OpenAI to court and more in Morning Squawk
Here are five key things investors need to know to start the trading day.
AI is changing older workers' careers, research finds — here's how
AI may either prompt some older workers to leave their jobs or help make their roles more efficient, research finds. Here's which careers may be most affected.
Goldman Sachs picks two stocks that could benefit from a chip designer shortage
A structural shortage in labor supply across the semiconductor industry means the firms could be uniquely placed to grow EDA revenues.
Centre’s Spacetech Fund Makes Maiden Investment, Backs Dhruva Space With ₹60 Cr
Spacetech startup Dhruva Space has netted ₹60 Cr ($6.3 Mn) in funding from IN-SPACe’s recently commissioned VC fund Antariksh Venture…
Can Ozempic prevent cancer? A doctor explains why the headlines are easy to misread
Several studies suggest GLP-1 drugs may lower cancer risk. But that benefit may be due to the patients themselves: those who are healthier, wealthier and with better access to care.
Massive AI Buildout Poses Latest Inflation Threat as Consumers Pay More for Laptops and Electricity
Economists expect the $700 billion investment in data centers will continue to push up inflation at least through the end of this year.
Phoebe Gates: Tochter von Bill Gates wird Trickserei bei ihrer Shopping-App Phia vorgeworfen
Das Unternehmen von Phoebe Gates steht aktuell heftig in der Kritik. Wie Recherchen von »Bloomberg« ergaben, soll Phia Provisionen für Verkäufe eingestrichen haben, die eigentlich anderen Unternehmen zustanden.
Trotz roter Zahlen: Intel investiert Milliarden in irische Halbleiterfabrik
Intel will in Irland mehr Chips für Prozessoren herstellen, um von der hohen Servernachfrage zu profitieren.
Ismail Eleburuike is building the operating system for African schools with SchoolTry
Within a year of deploying SchoolTry for higher education, the startup generated more revenue than it made in three years deploying SchoolTry for K-12.
Exclusive: 34 CEOs on what thrills and terrifies them about agentic AI
When businesses leaders think about AI right now, they’re thinking about how agentic tools will change the very nature of their work. “We’ve crossed a line,” says Varun Krishna, CEO of the fintech giant Rocket Companies . “AI is no longer just creating. It is thinking, deciding and acting. That changes everything, from client interaction to security.” As agentic tools get more sophisticated, they “expose how many organizations a
China’s drug industry pivots to AI-powered candidates to drive next wave of deals
After China’s cross-border deals for innovative drugs hit a record US$110 billion in the first half of 2026, the sector is now pivoting towards artificial intelligence-powered candidates to drive the next wave of transactions. China accounted for about 30 per cent of all new drugs currently under development worldwide, ranking second globally, according to Lan Gongtao, deputy director general of the Department of Drug Registration at the National Medical Products Administration. China’s...
New Mexico AG Calls for Reform After Report Finds “Substantial Racial Disparities” in One School District
The post New Mexico AG Calls for Reform After Report Finds “Substantial Racial Disparities” in One School District appeared first on ProPublica .
Africa’s crypto payment experiment is finding its first believers at local stores
Bitcoin communities and fintech startups are testing two competing models for making crypto payments work in everyday commerce across Africa.
iPhone 18 Pro Max: Komponentenpreise geschätzt um 300 US-Dollar erhöht
Eine Bill-of-Materials-Berechnung besagt, dass Apple auch beim iPhone 18 Pro Max die Preise erhöhen dürfte. Die Komponentenkosten steigen deutlich.
EU moves towards social media ban for children
Brussels will propose gradual access for different ages in response to concerns over child safety online
Chinese internet firms sign AI agent data protection pact
The China Internet Association released a self-regulatory pact on personal information protection for AI agents at a forum in Beijing, with Baidu, Tencent, Alibaba, Volcengine, and 27 other internet companies among the first signatories. The pact is aimed at standardizing how AI agents collect, process, and use personal data as agent-based services spread across internet […]
China sets 2030 target for next-generation internet infrastructure
China’s Ministry of Industry and Information Technology and three other agencies issued guidelines on July 13 to upgrade the country’s internet basic resources, targeting “systematic breakthroughs” by 2030 and a more advanced national internet infrastructure by 2035. The document calls for research into agent-to-agent networks, satellite internet, digital identity infrastructure, IPv6 upgrades, and the integration […]
EU sanctions Russian cyber spies for years-long hacking
Moscow's Federal Security Service is behind the cyber espionage and sabotage campaigns, EU says.
Ant Group unveils AI safety models for agents and multimodal systems
Ant Group’s AI Safety Lab has open-sourced SingGuard-NSFA, a safety guardrail model for autonomous agents, and disclosed details of SingGuard, a multimodal safety model. SingGuard-NSFA is designed to detect risks such as prompt injection, sensitive data theft, malicious code execution, resource abuse, and permission misuse before agents take action. The model covers seven major risk […]
Google’s SensorFM turns messy wearable sensor data into a general-purpose health intelligence layer
Google Research's SensorFM is a foundation model trained on more than a trillion minutes of wearable data from five million Fitbit and Pixel Watch users. It beats existing benchmarks on 34 of 35 health and behavioral tasks. SensorFM could eventually power Google's AI health coach, but the company hasn't announced any integration plans yet. The article Google’s SensorFM turns messy wearable sensor data into a general-purpose health intelligence layer appeared first on The Decoder .
AI-generated code has made security debt a governance problem
Moving from tool approval to true governance is the only way for CISOs to keep pace with the accelerating velocity of software risk. The post AI-generated code has made security debt a governance problem appeared first on CyberScoop .
A U.S.-Mexico Impasse Will Test How Far the Trump Administration Will Go to Fight Drug Trade
The post A U.S.-Mexico Impasse Will Test How Far the Trump Administration Will Go to Fight Drug Trade appeared first on ProPublica .
Europe's Anduril rival Helsing raises $1.8 billion at $18 billion valuation
Helsing said "investor demand significantly exceeded the available allocation" for its $1.8 billion funding round.
An Unlearned Lesson: The Sorry Record of Regime Change Operations in the Middle East
On Feb. 28, 2026, President Donald Trump announced the commencement of Operation Epic Fury, a joint U.S.-Israeli military operation against Iran. Among the mission’s goals was the overthrow of the Islamic Republic. “When we are finished,” Trump told the Iranian people, “take over your government. It will be yours to take.” Israeli Prime Minister Benjamin Netanyahu echoed Trump’s message, saying: “Our joint action will create the conditions for the brave Iranian people to take their destiny into
Fast-tracking AI: Linking data, governance, easy access
Artificial intelligence agents can both read data and write actions, meaning security and governance are critically important.
Digital-Health-Podcast: Wie ein mehrfach Betroffener Datenlecks verhindern will
Cyberangriffe und Datenlecks im Gesundheitswesen sorgen für Unsicherheit. Wie ein mehrfach Betroffener selbst handelt und wo es hakt, zeigt diese Podcast-Folge.
Guide: Complexity overload: Why SME cyber security breaks, and how to fix it
A practical framework for IT generalists and small cyber security teams to cut complexity and strengthen security posture.
Ban social media for under 13s, says von der Leyen's child safety panel
The European Union should restrict access to social media for children aged under 13, a special online child safety advisory panel told Commission President Ursula von der Leyen on Monday. As she ...
FY26 Financial Tracker: Tracking The Financial Performance Of Indian Startups
The Indian startup ecosystem continued to mature in FY26, with 22 new-age tech companies making their public market debut as…
WTF is SPUR’s publisher-run Content Telemetry Framework?
SPUR is publisher‑run and fixated on one thing: turning AI’s use of their content from opaque scraping into a transparent, usage‑based licensing system they control.
Prediction market users spend nearly $200 million on midterm election bets: Report
Prediction market users have wagered in excess of $197 million on midterm election results, according to NBC News. The outlet analyzed 1,408 open markets on Kalshi and Polymarket for its report, published on Friday. On both platforms, users can bet on a variety of topics, including sports, global events and political elections. As of Sunday...
Social Media Wars, GoKwik Axes 120 Jobs & More
Uniform Norms For Social Media The IT ministry is crafting uniform standards for all messaging apps. But the timing is…
Former Mayo Clinic Leader Sues System Over Alleged AI Cover-Up: 6 Things to Know
A former Mayo Clinic research director claims she was silenced, demoted and ultimately fired for sounding the alarm on AI safety and patient privacy lapses at the health system. Traci Tamiko Eto is now suing Mayo for retaliation. The post Former Mayo Clinic Leader Sues System Over Alleged AI Cover-Up: 6 Things to Know appeared first on MedCity News .
Inside The Numbers Lifting IPO-Bound Cult.fit
After years of prioritising expansion over profitability, Cult.fit’s pre-IPO papers and draft red herring prospectus (DRHP) suggest that the fitness…
Early AI adopters set to dominate European banking market - Visa study
Artificial intelligence will fundamentally reshape retail banking by 2030, with early adopters set to significantly outperform laggards, according to a Visa survey of European industry players.
Study: The National AI Policy Landscape in K–12 Education
A snapshot of where districts stand on AI and what it reveals.
Field notes (23)
Washington Is Looking to Keep China From Training Its AI on US Models
CSET’s Colin Shea-Blymyer shared his expert insight in an article published by Bloomberg. The article examines growing concerns in Washington over the use of "distillation" by Chinese AI companies to train models using outputs from leading US systems, and the resulting debate over intellectual property, competition, and national security in the global AI race. The post Washington Is Looking to Keep China From Training Its AI on US Models appeared first on Center for Security and Emerging Technol
What will be left for us to work on?
My keynote at ICML 2026
Empowering India’s next generation of innovators with ATL Saathi
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
EPIC Urges D.C. Council to Strengthen Proposed Government Data Privacy and Protection Act
EPIC submitted testimony on Monday to the D.C. Council Committee on Public Works & Operations urging them to further strengthen B26-0670, the DC Government Data Privacy and Protection Act of 2026.
FIRST Global and Experiential Bring Agentic AI Learning Experience to 190+ Countries, Advancing Robotics Education
FIRST Global Joins the UN ITU AI Skills Coalition; FIRST Global and XRP Kits Offer New Agentic AI module powered by FYI.AI For Student Robotics Teams Worldwide GENEVA, Switzerland — July 8, 2026 — At the AI for Good Global Summit hosted... The post FIRST Global and Experiential Bring Agentic AI Learning Experience to 190+ Countries, Advancing Robotics Education appeared first on AI for Good .
EPIC, CFA, Fairplay Submit Recommendations Ahead of Potential Rulemaking for Colorado Automated Decision-Making and Chatbot Laws
EPIC, the Consumer Federation of America, and Fairplay submitted comments on Monday in response to the Colorado Department of Law’s request for stakeholder input ahead of potential rulemaking for the state’s recently passed bills on chatbot safety and automated decision-making technology (ADMT) in consequential decisions.
ChinAI #366: Most Companion Robots Die by Day 30
Greetings from a world where…
Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition
According to information from the STEM Education Promotion Alliance (SEPA) on July 11th, the two Vietnamese teams excellently won second place in the Junior and Senior categories of the Robotics for Good Youth Challenge 2026 Global Finals. The post Vietnam won two second prizes at the global finals of the Robotics for Good 2026 competition appeared first on AI for Good .
What VAR tells us about AI
FIFA's semi-automated refereeing system promises neutrality but brims with bias. It's widely despised yet embraced by elites. It is, in other words, a lot like AI.
datasette code-frequency chart on GitHub
datasette code-frequency chart on GitHub Out of curiosity I decided to see if I could find a useful illustration of the impact of coding agents and Opus 4.5 class models on my own output. The best I've found so far is this GitHub chart of frequency of code changes to my Datasette open source project: The big spike in activity at the end aligns with Opus 4.8, GPT-5.5, Fable 5 and GPT-5.6 Sol. Tags: github , ai , datasette , generative-ai , llms , ai-assisted-programming , coding-agents
Private Credit, Public Panic: Why Life Insurers Are Stronger Than the Headlines Suggest
Private credit has become the financial system’s latest designated villain: opaque, fast-growing, and—depending on the headline—one bad quarter away from dragging insurers, banks, and retirees down with it. For the past two years, warnings about life insurers’ private-credit investments have become a staple of financial commentary. In 2024, the International Monetary Fund cautioned that private ... Private Credit, Public Panic: Why Life Insurers Are Stronger Than the Headlines Suggest The post P
The Republics and Their Alliance, If We Can Keep It: The United States and Republic of Korea’s Time-Forged and Battle-Tested Alliance
The U.S.–South Korea relationship is well institutionalized, rooted in shared values, mutual aims of national security, and cutting-edge business cooperation. Despite external and internal ...
Heatwaves and Legal Remedies
Europe suffered an unprecedented heatwave this June, with debilitating effects felt across various walks of life: thousands of deaths, particularly among the elderly, individuals and families suffering in “heat-trap” apartments, hospitals full and caught unprepared, school closures, and productivity losses. Adaptation measures are indispensable for coping with these soaring temperatures, which have cost lives and severely affected people’s well-being. However, rights-based litigation involving a
Civilian Protection in the Age of Military AI: What Congress’s New Legislative Proposals Reveal About Emerging Safeguards
Members of the Senate are taking steps to regulate and restrict how the Department of Defense develops and uses AI in its operations. The post Civilian Protection in the Age of Military AI: What Congress’s New Legislative Proposals Reveal About Emerging Safeguards appeared first on Just Security .
Litigation Tracker: Legal Challenges to Trump Administration Actions
A public resource tracking all the legal challenges to the Trump administration's executive orders and actions. The post Litigation Tracker: Legal Challenges to Trump Administration Actions appeared first on Just Security .
Dignity Without Autonomy
In Prajwala v. Union of India, the Supreme Court held that victims of trafficking for commercial sexual exploitation have a right to rehabilitation under Article 23 read with the right to dignity under Article 21. While the judgment has been celebrated for its three-dimensional dignity framework, it is a missed opportunity to articulate a constitutional basis for protecting the rights of sex workers. The Court's dignity framework – calibrated against objectification in trafficking – is insuffici
📈 Data to start your week
A new scaling law; AI job fears fade; Dementia in retreat++
Law and Media Round Up – 13 July 2026
On 7 July 2026, Mr Justice Nicklin handed down judgment following the lengthy trial of the misuse of private information and breach of confidence claims brought by seven Claimants against Associated Newspapers Limited (“Associated”), the publisher of the Daily Mail, Mail on Sunday and MailOnline; Baroness Lawrence of Clarendon OBE & Ors v Associated Newspapers […]
Patricia Evangelista on journalism
The ‘worst’ and ‘best possible’ job at once: what it’s like to dedicate your life to documenting crimes against humanity - by Aeon Video Watch on Aeon
The state of enforcement: Part II — Kids and teens' privacy
State regulators are ramping up children's privacy enforcement, focusing on parental consent, company knowledge of minors, and protective safeguards.
The EU AI Act Newsletter #106: Calls to Enforce General-Purpose AI Rules
Experts call for robust enforcement of the rules on general-purpose AI with systemic risk, as the Commission unveils a new plan on advanced AI and cybersecurity.
What the UK government should do on AI and tech policy
Published 13 July 2026 — 3 minute READ Image — Kanishka Narayan, the UK’s Minister for Artificial Intelligence and Online Safety, speaks at an event celebrating the AI Impact Summit 2026, in New Delhi ...
The Agentic Age Needs A Cognitive Operating Model
Last October, I published a blog proposing a different mental model for AI agents: Treat them as cognitive skills and products, not as digital employees. That framing has since resonated strongly with Forrester clients, particularly technology leaders building agentic capabilities inside the enterprise. But the concept of a cognitive skill in that blog was deliberately loose. […]
Policy (20)
Special panel report: Child safety online protecting and empowering minors in a digital world
Special panel report: Child safety online protecting and empowering minors in a digital world Anonymous (not verified) Mon, 07/13/2026 - 11:15 As announced in the 2025 State of the Union address, President von der Leyen set up a special panel of experts to develop a strong and practical European approach to keep children safe online. The panel's co-chairs presented their final report in July 2026. The report highlights the critical challenges children face online and provides recommendations and
Regulatory Relief for Certain Stationary Sources to Promote American Chemical Manufacturing Security
BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION 1. The United States relies on a strong chemical manufacturing sector to support industries like energy, national defense, agriculture, and health care. These facilities produce essential inputs for critical infrastructure, advanced manufacturing, medical sterilization, semiconductors, and national defense systems. Maintaining a robust domestic chemical […] The post Regulatory Relief for Certain Stationary Sources to Promote America
Nominations Sent to the Senate
NOMINATIONS SENT TO THE SENATE: Keith Sonderling, of Florida, to be Secretary of Labor. Andrew A. De Mello, of Virginia, to be a Judge of the United States Tax Court for a term of fifteen years. The post Nominations Sent to the Senate appeared first on The White House .
Statistical Policy Directive No. 8: North American Industry Classification System (NAICS)-Request for Comments on Proposed Updates for 2027
The Office of Management and Budget (OMB) seeks public comment on the advisability of adopting the proposed North American Industry Classification System (NAICS) updates for 2027 recommended by its Economic Classification Policy Committee (ECPC). The ECPC recommends an update of NAICS to clarify existing industry definitions and content, recognize new and emerging industries, and combine industries. There are two parts in the SUPPLEMENTARY INFORMATION section below. Part I summarizes the backgro
FY 2026 Job Placement and Training-Native American Technology and Manufacturing Grant Pilot Program (IGNITE: Indigenous Growth in New & Innovative Trade Employment); Solicitation of Proposals
Through this notice, the Bureau of Indian Affairs (BIA), Office of Indian Services (OIS), Division of Workforce Development (DWD), Job Placement and Training announces a forthcoming FY 2026 Job Placement and Training discretionary grant pilot program titled, "Native American Technology and Manufacturing Grant Pilot Program-- (IGNITE: Indigenous Growth in New & Innovative Trade Employment)" Notice of Funding Opportunity (NOFO) for Tribal workforce development strategies preparing Tribal participa
Financial Market Reforms Could Lift Europe's Growth
French startup Mistral AI’s recent financing round was led by ASML, a Dutch maker of semiconductor manufacturing equipment. Such cross-border investments—even when small relative to US deals—are not ...
News from UK Parliament
Welsh Affairs Committee: cross-border healthcare patients are “falling through the gaps” The Welsh Affairs Committee is concerned by the lack of urgency in addressing long-standing issues affecting ...
OECD Forum on Tax Administration Plenary to be held in Cape Town on 18-20 November
How to apply effective governance to harness the benefits of A.I. and mitigate its risks ...
Thought for the week: Web scraping for generative AI is subject to the GDPR
This article was originally published by IAPP linked here. Organizations using scraped data for AI training should prepare for heightened expectations around data minimization, transparency and accountability. On 7 July, the European Data Protection Board approved “Guidelines on web scraping in the context of generative AI.” Perhaps not surprisingly, the EDPB considers that web scraping [...] The post Thought for the week: Web scraping for generative AI is subject to the GDPR appeared first on C
FTC Secures $12 Million in Penalties for Pre-Merger Reporting Act Violations
FTC alleges Edwards Lifesciences and Genesis structured JC Medical deal to avoid federal antitrust review The Federal Trade Commission secured $12 million in penalties to settle charges alleging that Edwards Lifesciences Corp. acquired medical device maker JC Medical from Genesis MedTech Group Limited without complying with the notification and waiting period requirements of the Hart-Scott-Rodino Act (HSR). View Press Release
Priority Open Recommendations: U.S. Department of Agriculture
What GAO Found In May 2025, GAO identified 5 priority recommendations for the U.S. Department of Agriculture (USDA). Since then, USDA has implemented two of those recommendations. In June 2026, GAO identified an additional 3 priority recommendations, bringing the total to 6. GAO is highlighting the following three areas that warrant timely and focused attention: Improving IT modernization, Deterring SNAP retailer fraud, and Improving data sharing on foreign investment in U.S. agricultural land.
Military Health Care: Clinical Quality Management in Operational Settings Like Field Hospitals
What GAO Found In 2023, the Department of Defense (DOD) directed the military departments—Army, Navy, and Air Force—to update their policies on clinical quality management to align with Defense Health Agency (DHA) procedures to ensure high-quality care in operational settings. In December 2024, GAO reported that the military departments had not yet issued policies, specifically on provider credentialing and privileging, and recommended that they do so. As of March 2026, GAO found that Army and A
Human Capital: A Guide for Developing and Assessing Strategic Training and Development Efforts in the Federal Government
What GAO Found Training and development programs help federal agencies achieve their mission and goals by improving individual and, ultimately, organizational performance. This report is a guide that federal agencies can use to ensure their training and development investments are targeted strategically. In recent years, training and development have shifted from primarily classroom-based instruction to more integrated, blended learning approaches that reflect changes in the workplace and advanc
Southern Border Security: DOD Used Multiple Strategies to Fund Operations
What GAO Found The Department of Defense (DOD) used multiple strategies to fund support for southern border operations since the start of fiscal year 2025 and into fiscal year 2026. Specifically, DOD realigned $1.74 billion in funding from amounts appropriated for fiscal year 2025 from various funding categories; transferred $608 million from or through DOD’s Drug Interdiction and Counter-Drug Activities, Defense account; relied on military construction authorities to fund border barrier project
Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers
What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
Health Insurance Marketplaces: CMS Needs Stronger Controls to Prevent Unauthorized Actions by Agents and Brokers
What GAO Found Millions of consumers rely on the assistance of health insurance agents and brokers to purchase health insurance plans through federal and state Marketplaces established by the Patient Protection and Affordable Care Act. The federal Marketplace is maintained by the Centers for Medicare & Medicaid Services (CMS). To assist consumers in the federal Marketplace, agents and brokers must be licensed to sell health plans and be registered with the Marketplace, among other things. CMS co
Iran and Nuclear Weapons Production
Forward-Funded Federal Education Programs: Frequently Asked Questions
Veterans Law: Unaccredited Representatives
Young firms, job quality and inclusiveness
How to apply effective governance to harness the benefits of A.I. and mitigate its risks ...
Research (123)
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Cri
Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing
Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans and three large language models (GPT-4o, Gemini-2.5-Pro, Claude-Sonnet-4.5) using verbal fluency data. By applying trajectory-based NLP metrics to the items generated by 82 human participants and LLM output across eight temperature settings, we quantified three complementary dimensions: entropy (step size predictability), distance to next (successive semanti
Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems
Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants. We present Cost-Governed RAG, an architecture that integrates a codebook-oblivious vector index (TurboVec) with a multi-tenant LLM governance gateway, creating a unified observability st
TRAIL: A Platform for Configurable Human--AI Teaming Experiments
An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions. Studying this rigorously demands infrastructure no existing tool provides: reproducible configuration of an AI teammate embedded in instrumented, real-time collaboration sustained over time. We present the Team Research and AI Integration Lab (TRAIL), a web platform that makes the AI teammate a configurable, reproducible design object, pairing a Big Five pe
It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance
Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradi
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability ceiling, and the Genesis Mission has placed autonomous e
CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform
Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validated against empirical mobility patterns. We present CityBehavEx, an interactive LLM-assisted urban simulation platform that scales to city-size populations, exposes agent behavior for inspection, supports empirical validation, and generates mobility patterns that better match real-world spatial, temporal, and semantic distributions. Instead of inv
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation
Representation and Reference Selection in Training-Free Synthetic Image Attribution
Synthetic image attribution aims at identifying the generator responsible for a given AI-generated image. Training-free reference-based attribution methods are easily scalable, since newly emerging generators can be incorporated by adding source-specific references rather than retraining a task-specific classifier. Their performance depends on two coupled factors: the representation space used for comparison and the way source-specific references are constructed. However, the interaction between
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning. In this class, we theoretically prove that the training dynamics of attention models can be confined to a hig
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning. In this class, we theoretically prove that the training dynamics of attention models can be confined to a hig
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manip
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
From maritime trade to commercial nuclear power, insurance has been the enabler of major economic and technological developments by pricing risk, limiting downside, and spreading best practices. The emerging AI agent economy, projected to handle trillions of dollars in transactions by 2030, looks to be the next such development. Yet insurers' exposure to AI agent risk currently sits largely unpriced across existing insurance lines; between this silent coverage and growing exclusions, coverage is
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTelle
Uncovering Students' Mental Models of Generative Artificial Intelligence
In this paper we present a study of students' mental models of generative AI (GenAI). A student's mental model of GenAI influences not only how they perceive the technology's capabilities and limitations but also how they choose to integrate it into their academic work. Whether they view it as a collaborative partner, a shortcut to complete tasks, or something in between, depends on how they conceptualize its use. This study addresses the following questions: (I) What mental models do undergradu
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences. However, progress remains fragmented: models use incompatible action spaces and prediction targets, datasets and tasks follow different conventions, and runtime systems e
Interaction Scaling: Grounding the Third Axis of Test-Time Compute
There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal. Every extra token comes from the same frozen weights and the same prompt, so neither can tell the model anything it does not already know. We study a third way, interaction: the model proposes an artifact, an external instrument observes how it actually behaves, and the model revises. Each cycle imports a real observation,
Structure-Feature Aligned Graph Learning via Alternating Constrained Optimization
We introduce a constrained two-view framework for node prediction that aligns structure-conditioned GNN embeddings with a structure-free feature prior learned by an anchor model. Conventional Graph Neural Networks (GNNs) couple feature transformation and neighborhood aggregation, which renders them vulnerable to topology noise and heterophilous connections. To decouple this dependency, our framework utilizes an independent anchor network to capture intrinsic attribute features via a self-supervi
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-tr
LightMem-Ego: Your AI Memory for Everyday Life
Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously accumulate, organize, and retrieve long-term experiences, which remains challenging. To address this challenge, we present LightMem-Ego, a lightweight streaming multimodal memory system for everyday-life assistance. The system continuously captures egocentric
A Multimodal Dataset for Large Language Model Applications in the Energy Domain
This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observation
BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation
Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI. While current Text-to-Audio (TTA) frameworks successfully synthesize isolated sound effects, they struggle with narrative cohesion, temporal alignment, and cinematic emotional depth. We present BackgroundMellow, a framework that treats story-to-audio generation as a precise orchestration and signal processing problem. This framework is enabled without ground-
Longitudinal Multi-View Breast Cancer Risk Prediction
Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection. Recent deep learning methods have shown the value of longitudinal data and explicit temporal alignment. However, existing approaches either perform explicit alignment using a single mammographic view or model multiple views without explicit longitudinal alignment, limiting their ability to exploit the complementary spatial-temporal information used in c
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science education. In most systems, a general-track subject -- Digital Literacy, ICT, TIC, or SNT -- bears the weight of universal AI literacy, while a specialist Informatics course serves STEM pathways separately. Yet the content and depth of the general track are shaped by governance decisions made largely with reference to the specialist one. This paper presents a compar
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. This study presents a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses regarding the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. We uncover four interconnected patterns of \emph{epistemic paternalism}: (1)~\textbf{Differential Refusal}, where safety-aligned models block
Towards Predictive, Aligned, and Scalable Robot Learning
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training e
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precision. We propose Pix2Act, an imitation learning method that addresses these challenges by generating continuous image-space keypoint trajectories in each camera plane and losslessly recovering end-effector poses via triangulation. This reformulates high
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form p
NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management
Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints. Current assessment practice, however, remains poorly aligned with this setting: many studies rely on static examinations or report only terminal portfolio returns, while the intermediate evidence, analyst judgments, and execution steps that produced those returns stay largely invisible. We introduc
A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery
The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are given access to a flat, monolithic registry of tools, the model must evaluate hundreds or thousands of options simultaneously. This leads to decision-space explosion, context window saturation, and degraded routing accuracy. To address these limitations, this paper presents a hierarchical, skill-based architecture for agentic orchestration. Capabilities are or
Adoption-Ready Project-Based Learning for Computing Education: The FORAP Framework and a Multi-Scale Project Portfolio
This innovative practice full paper presents FORAP (Framework for Organizing Reusable and Adaptable PjBL Projects) and a portfolio of 14 adoption-ready project-based learning (PjBL) project packages built with the framework. PjBL in computing education offers strong educational benefits, yet its adoption remains limited by high instructor workload and recurring student technical challenges. FORAP addresses these barriers by organizing each package around a project designed with aligned learning
NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study
Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generation, or statistical analysis. However, deployment in institutional biomedical environments requires governed mechanisms for research planning, data access, workflow orchestration, evidence tracking, reproducibility, and human oversight. We present NVAITC AI Scientist (NAIS), a governed end-to-end agentic research system designed to support domain-general scie
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-a
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Dat
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynC
LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI
Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby anatomy, complicating noninvasive prediction. These fine perineural cues are easily attenuated by routine downsampling or overly global feature aggregation, reducing the effectiveness of conventional volumetric models. We present LoSA-Net, a localized
Prism: Automating Science-of-Evals Research
From Solution Traps to Solution Patchwork: Easing Tensions in Designing Digital Health in the Global Context
Two recently published viewpoint articles in JMIR highlighted the two faces of medical informatics research. On the one hand, they spotlight the significant advances digital technologies bring to health management and delivery. Technological advances of the last decade have transformed healthcare worldwide. Consumers have access to digital health tools, gathering an unprecedented amount of data available for gaining personalized insights. Digital technologies support medical professionals in cli
From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness
Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms. We show this conflates two distinct claims: decoder-geometry alignment and encoder-activation behavior. We reproduce the superposition phase diagram of Elhage et al. (2022), identifying a convergence artifact at high sparsity and an un
Digital Outpatient Care for Patients With Type 1 Diabetes (DigiDiaS): Pragmatic Observational Pre-Post Study
Background: Patient-reported outcomes in digital health solutions can offer patients with type 1 diabetes an opportunity to voice their needs in outpatient care, enabling clinicians to tailor support. Evidence on long-term health impact and routine integration of such digital solutions outside controlled settings is limited. Objective: This study aimed to compare a flexible digital supplement to outpatient care for type 1 diabetes (DigiDiaS) with usual care over 1 year, with self-management as t
Promoting Problem-Solving Among Low-Income Adults With Type 2 Diabetes: Cluster-Randomized Controlled Trial of a Mobile Health Intervention With SMS Text Messaging (Mobile Diabetes Detective)
Background: Problem-solving is essential for the self-management of type 2 diabetes but remains challenging for underserved individuals. Although mobile health (mHealth) interventions can improve diabetes self-management, few focus on problem-solving. Objective: This study evaluates the efficacy of Mobile Diabetes Detective (MoDD), a fully automated web-based intervention with SMS text messaging that provides problem-solving support tailored to self-monitoring data, for improving glycemic contro
Tracing Agentic Failure from the Flow of Success
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and tr
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy pass@1 nearly vanishes after compression, yet pass@k recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repeti
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements evolve, the harness must be continually modified. Before such a change can be made, a developer or coding agent must identify all code locations that implement the target behavior. This is difficult because production harnesses are large, tightly coupled, and behavio
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception via a "skeptical" reasoning paradigm. Unlike holistic models, ChartCynics decouples perception from verification: a Diagnostic Vision Path captures structural anomalies (e.g., inverted axes) through strategic ROI cropping, while
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We explo
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning pr
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we int
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for predefined goals such as capture-the-flag, remote code execution, exploit reproduction, or trajectory similarity, in simplified or narrow settings. These tools are valuable for measuring bounded capabilities, yet they do not adequately capture the complexity, open
Self-Improvements in Modern Agentic Systems: A Survey
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains. We offer a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple tas
From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review effic
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to
Edge-Aware Thermal Infrared UAV Swarm Tracking
Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging due to limited appearance cues, frequent occlusions, and rapid maneuvers. Despite significant progress driven by benchmarks such as the Anti-UAV challenge, existing methods primarily prioritize accuracy while overlooking the computational constraints of real-time edge deployment. The standard Kalman Filter (KF) offers the efficiency required for
Paradoxes of Game Theoretic Equilibria and Price of Anarchy
For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing multi-agent learning to static equilibrium and black-box regret analysis obscures underlying dynamic disequilibrium and game theoretic bounds. First, interior Nash equilibria lack $C^1$ vector field information, meaning
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated
Latent-Identity Tuning in Text-to-Image Personalization Models
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the l
Metacognition in LLMs: Foundations, Progress, and Opportunities
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to adv
Evidence-Backed Video Question Answering
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a b
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post hoc, model-specific,
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment
Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable representations across diverse graph domains. Recent advancements in GFMs have been largely dominated by two paradigms: Graph Neural Network and Large Language Model (LLM) based methods. However, these methods often face a fundamental dilemma between training with limited data and a heavy reliance on textual attributes. Tabular foundation models (TFMs) offer a potential alternative, as node features and
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-tr
Trustworthy synthetic data for campaign decision support: strategy simulation fidelity and the PolicySynth framework
Decision support systems (DSS) increasingly run retention what-if analysis on synthetic customer populations, because privacy constraints preclude unrestricted use of real data. Such a system is trustworthy only if the synthetic data lead managers to the same decisions as the real data would; yet prevailing criteria certify distributional similarity, not decision alignment, so a synthetic population can match every marginal distribution while still steering a marketing team toward the wrong camp
LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models
Pathology Foundation Models (PFMs) offer powerful Whole Slide Image (WSI) representations but suffer from massive computational costs. While Knowledge Distillation (KD) can create efficient student models, existing multi-teacher methods often use suboptimal uniform weighting that ignores tissue heterogeneity. We propose LaGuadia (Language-Guided Adaptive DistillAtion), a framework that develops a compact pathology image encoder by dynamically integrating expertise from multiple PFMs under clinic
Multi-Agent LLMs Fail to Explore Each Other
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do so, often exhibiting myopic and polarized interaction patterns that lead to suboptimal coordination and increased regret. We formalize this challenge as the Multi-Agent Exploration problem, modeling it as a partially observable stochastic game (POSG) problem in w
Reference-Based Face Super-Resolution Using the Spatial Transformer
Face super-resolution is the task of increasing the resolution of an image containing a face thereby adding finer detail. It is a ubiquitous task in many computer vision applications and quite often the user isn't even aware that it is being performed. However, doing it with high fidelity is challenging as it is an ill-posed problem. In this paper we present a reference-based solution for face super-resolution that uses higher resolution reference images to aid in the task. We show an alignment
Teacher-regulated generative AI support, student agency, and perceived learning gains in higher education: the moderating role of perceived fairness
IntroductionGenerative artificial intelligence is increasingly used in higher education, yet its educational value depends not only on technological access but also on how its use is pedagogically regulated. This study examined how teacher-regulated generative AI support is associated with university students' perceived learning gains in higher education. It further tested whether student agency mediates this association and whether perceived fairness conditions the strength of the association b
Lions and tigers and AI, oh my: an ethical framework for human-AI interaction based on the five freedoms of animal welfare
If an AI entity is conscious, it deserves moral status and to have its welfare protected. Just as society has granted certain animals moral status with legislation to protect their welfare because they have been deemed to be sentient, this paper shows that for any AI that can be confidently determined to be conscious, that AI deserves the same status and protection as sentient entities. This paper’s focal point is developing a framework for how an AI entity’s welfare can be protected, based on t
Cognitive load gating system in motor imagery BCIs: a dual-task EEG study with differential entropy-based reliability estimation
Brain–computer interface (BCI) systems based on motor imagery hold significant clinical value for individuals who have lost voluntary movement, but most studies test BCI assistive devices like powered wheelchairs under ideal conditions where motor imagery signals are not interfered by simultaneous cognitive load. This work records electroencephalography (EEG) from 13 participants across four tasks: baseline rest, mental arithmetic (Easy, Medium, Hard), pure left/right motor imagery and both task
Vertics Oy was born from the vision of two upper secondary school students
Atte Pohjanmaa, who founded a company when he was still an upper secondary school student, succeeded in growing his company even during the coronavirus years. Atte Pohjanmaa, Vertics founder Atte ...
From Chaos to Clarity: A Framework for Program-Level AI Learning Outcomes
Industry is leaning into generative artificial intelligence (GenAI), and higher education is under pressure to prepare graduates for a GenAI-augmented workforce. Yet, there is still no clear structure for defining AI readiness across disciplines, programs, courses, and assignments. Current approaches often rely on broad institutional policies or individual course-level decisions, which can also create mixed messages for students, fragmented expectations across programs, and limited visibility fo
Analysis of Mutual and Referential Human and Robot Gazes in a Collaborative Word Association Game
Robot gaze is a major component of human-robot dialogue coordination. Most studies of gaze in human-robot dialogue focus on face-to-face social conversations, but little is known about gaze in demanding task-focused interactions. In this paper, we investigate how the gaze of a robot game partner affects human visual attention and if humans tend to direct confirmation-seeking gazes towards the robot. In our study, we let participants play a collaborative word association game with a NAO robot act
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. In this work, we adapt and enhance preconditioned gradient methods to overcome the practical challenges of large-scale LLM pretraining. We first identify instabilities in SOAP at large batch sizes and propose algorithmic modifications including per-step QR orthogonalization and improved preconditioning strategies that e
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagnostic reasoning, while the length of tasks such systems can complete at 50% reliability doubled roughly every seven months. These crossings are rapid and broad, but the frontier is jagged: humans retain decisive advantage
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents. Existing web design, SEO, and ge
Metacognition in LLMs: Foundations, Progress, and Opportunities
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to adv
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven judges, seven bias types, and nine benchmarks. Geometry: baseline judging inputs occupy a tight acti
Evidence-Backed Video Question Answering
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise
LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components. Rather than assuming a fixed fusion structure, the framework first integrates more closely related modalities and then in
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations). An automated scenario generation pipeline produces
Supporting Reflection in LLM-based Exploratory Search
Large Language Models (LLMs) can make exploratory search more efficient but may undermine the reflection and iterative sensemaking needed in unfamiliar domains. Existing LLM tools often prioritize rapid answers over supporting users in tracking how their understanding evolves and how well their strategies align with their goals. We present TrailLM, a system that helps users reconstruct and revisit their exploration paths to support reflection and metacognitive engagement during information seeki
Natural hazards articles from across Nature Portfolio
Soil plasticity and water content drive elevated soil cracking risk across ~60% of China’s land during summer peaks, with increasing risk under future warming, suggests an analysis of 4,000 laboratory ...
A Guide to the Convergence of Electronic Warfare and Cyber Operations
McIlvenny, J. (2026, July 14). A Guide to the Convergence of Electronic Warfare and Cyber Operations. Retrieved July 16, 2026, from https://doi.org/10.58012/ssy8-4t48 ...
NTU College of Computing and Data Science
At the College of Computing and Data Science (CCDS), education is designed for a world where artificial intelligence shapes how problems are understood, solved, and scaled. Our programmes combine ...
HandPad: A Bimanual Hand Interface for Fluid Window Interactions in VR
Virtual Reality (VR) offers potential for productivity work by creating expansive displays anywhere, yet current systems often rely on external input devices that limit the on-the-go use of mobile VR. We introduce HandPad, a suite of bare-hand interaction techniques that leverage the benefits of asymmetric bimanual coordination and self-haptic support. HandPad assigns the non-dominant hand (NDH) to establish spatial frames and interaction contexts, while the dominant hand (DH) performs fine-grai
"We are all in big trouble! *Shock Emoji": Personal Narratives in Expressing Emotions, Opinions, and Data Regarding Climate Change in TikTok Short Videos
Climate change is a source of anxiety about the future. Understanding how people express themselves about climate change enables us to address such concerns. To study climate change expression on social media, we analyzed 200 TikTok videos tagged with #climatechange, identifying four categories of content: expression-feelings, views-appeals, news-information, and trend-hijacking. We found that creators use humor to package sharp critiques, avoiding direct confrontation. They replace complex disc
Securing LLMs in the Wild: Privacy and Security Challenges at the Edge
Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edge platforms. While cloud deployments offer scalable compute, concerns over data sovereignty, compliance, latency, and third-party dependence are driving organizations toward edge and on-premise LLMs. This shift introduces new security and privacy challenges: limited compute and memory force aggressive optimizations, including quantization, pruning,
Deborah Hughes Hallett
Deborah Hughes Hallett is Adjunct Professor of Public Policy and Professor of Mathematics at the University of Arizona. She graduated from Cambridge University in England and has taught at Middle East ...
Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game
Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination game, models of conceptual semantics cannot explain how shared meaning emerges and changes in groups of people; however, existing games assume that players share payoffs in a partnership setting. We model conceptual alignment as a non-partnership game and illustrate differences in actual and percei
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal
Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although prior research has proposed frameworks, patterns, and user-centered approaches to support explainability, there is limited empirical understanding of how existing Requirements Engineering (RE) practices support explainability requirements across the RE lifecycle, especially in an industrial context. This paper reports early findings from an ongoing industry-b
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronge
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories
Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop. A major complication is decisive: long post-operation time-lags caused by concrete curing, watertightness ponding tests, and paint drying, during which a module is blocked while its workstation stays free. On benchmark instances grounded in an official national prefabrication guidebook, these lags inflate even the optimal reference makespan by
Active Offline-to-Online Reinforcement Learning
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the po
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams often span over multiple weeks or months, gradually build trust and request for money or sensitive information. Existing scam-detection systems mainly focus on isolated messages, which renders them inadequate against this evolving threat. This paper extends single-message phishing detection and presents an explainable agentic system for detecting sophisticated
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evalu
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety failures directly actionable. Red-teaming must therefore keep pace with evolving models and tools. Existing approaches mainly optimize attack success and preserve artifacts such as benchmarks, payloads, or attack programs, which record where attacks succeed but not the enabling conditions behind unsafe agent behavior. We study automated red-teaming for productio
Requirement-Driven Design of Whole-Body Social Tactile Sensing via Virtual Human-Robot Interaction
Tactile sensing for social-physical human-robot interaction (spHRI) is designed in a hardware-driven manner, where predefined sensor configurations constrain coverage, spatial resolution, and the range of recognizable gestures. We propose a requirement-driven framework that derives sensing requirements, specifically spatial resolution and placement, directly from interaction data. Using a VR-based platform with haptic feedback, we collected high-resolution whole-body contact distributions across
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely
Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using
Extending LLM Context via Associative Recurrent Memory
Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recurrent Memory Transformer (ARMT) as a practical approach for enabling long-context processing in LLMs, constant memory scaling, and better efficiency. We make three main contributions. First, we construct two domain-specific long-context datasets desig
Auditing the Risk Claims of Distributional Reinforcement Learning
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominan
Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection
The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced languages like Bangla, due to cultural context, implicit expressions, and informal linguistic patterns. This study aimed to expose the crisis of Bangla HS detection systems by diagnosing how and why benchmark-trained models fail to identify implicit, context-dependent HS. S
DiffEEG: A Self-Supervised Denoising Diffusion Model for Learning EEG Generic Representations
Deep learning for EEG-based seizure detection faces critical challenges: severe annotation scarcity and extreme class imbalance, where ictal events comprise less than 10\% of clinical recordings. We present DiffEEG, a 9.6M-parameter self-supervised foundation model that addresses both limitations through denoising diffusion pre-training and reinforcement learning (RL)-based fine-tuning. Pre-trained on 1.3M unlabeled segments from the Temple University Hospital Seizure Corpus (TUHSZ), DiffEEG lea
ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions
As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI
Heuristic Learning for Active Flow Control Using Coding Agents
Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design particularly challenging. Deep reinforcement learning (DRL) has emerged as a powerful framework for such problems, but its success typically relies on large numbers of simulator interactions and produces neural-network policies whose decision process often remains difficult to interpret. In this work, we investigate a different paradigm: instead of optimizing
PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing
Researchers organize the papers they collect into personal folder hierarchies in reference managers, and route each new paper into the folder where it belongs. This task differs from standard hierarchical text classification. A user's folder hierarchy is not a fixed, shared taxonomy but a private and evolving folksonomy whose folder meanings may be topical, shorthand, venue-based, or process-oriented, and are often defined by the papers already stored inside them. We formalize this setting as pe
Technical Report on the CVPR 2026@AdvML Workshop Challenge
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images an
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet most VLAs observe the scene in the camera frame, creating a frame mismatch between where the scene is observed and where actions are defined. The mismatch is benign under a fixed viewpoint, where the policy can memorize a single observation-to-action mapping, but grows harder as large-scale datasets aggregate demonstra
Agentic Skill Optimization over Lie Algebroids
Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and traces. Skill edits are not independent coordinates in a vector space: they are local repairs to structured artifacts whose effects are observed only after rollout, validation, and critique. Distinct edits can have the same immediate visible effect while differing in routing context, template state, guardrail scope, or future composability. The order of edits can
Enhancing Query Efficiency for d-DNNF Representations Through Preprocessing
In this paper, we investigate preprocessing techniques aimed at improving the efficiency of accessing models of propositional formulas represented in conjunctive normal form (CNF). We focus on three fundamental tasks: uniform sampling, direct model access, and model enumeration. Our analysis reveals that most state-of-the-art preprocessors, when they do not preserve formula equivalence, are generally unsuitable for these tasks. In contrast, we demonstrate that preprocessors which preserve model
ManiScope: LLM-Assisted Visual Analytics of Cryptocurrency Manipulation Risk
Cryptocurrency markets are vulnerable to trade-based manipulation, such as wash trading, which can distort price signals and mislead investors. Prior research has mainly focused on detecting manipulation using fixed rules or labeled examples, offering limited flexibility and interpretability for assessing potential risks. Existing visual analytics tools can reveal basic manipulation-related signals, such as token distribution, but still require substantial manual effort to integrate holder relat
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool trajectories, or free-form histories, making it difficult to track what has been grounded, what remains missing, and when the evidence is sufficient to answer. We propose Omni-Decision, a training-free evidence-state system that turns omni-modal QA i
DiffLens: A Visualization System to Explore Local Differences in Graph Sampling
Graph sampling techniques have been widely used to simplify network computation and visualization, which also results in inevitable differences between the sampled networks and the original networks in terms of nodes, edges and structures. Investigating such differences can inform graph sampling technique users of the pros and cons of different techniques and select the appropriate one, and can also help graph sampling developers evaluate their own technique. However, there are still no systemat
Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation
Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applications in urban planning, forest monitoring, and climate policy, where both accuracy and reliability are critical. Yet most deep learning models yield only deterministic predictions, providing no indication of per-pixel reliability. These regression tasks are inherently challenging due to heterogeneous land surfaces, skewed target distributions, sensor noise, and
FAD-SA-GRU: Enhancing Hate Speech Detection in Algerian Dialect Through Feature-Augmented Self-Attention GRU Networks
The widespread adoption of social media platforms has transformed online communication by enabling users to exchange information and opinions instantly. However, these platforms have also facilitated the dissemination of abusive and hateful content, posing major social, psychological, and ethical challenges. Hate speech can incite discrimination, harassment, and violence against individuals or communities based on attributes such as ethnicity, religion, gender, nationality, or political affiliat
When cheap gradients fail: the measurement cost of attacking quantum classifiers
Adversarial perturbations threaten machine learning classifiers, including variational quantum classifiers. We show that finite quantum measurement statistics (shot noise) act as a built-in defense against gradient-based test-time attacks whose cost scales unfavorably for the attacker. Because every gradient component must be inferred from repeated circuit executions under any unbiased gradient-estimation rule, white-box extraction consumes a dimension-dependent measurement budget that measureme
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit a
Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration
Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities. However, passive feed browsing creates a paradox: the authentic peer content that provides emotional grounding also triggers potentially detrimental upward social comparison and cognitive overload. Previous work has either structured online user-generated content to reduce noise without changing the passive browsing modality, or built AI-powered career exploration systems th