23:22 UTC
Archive · 2026-07-08

AI ethics on Wednesday, 8 July 2026

172 items published this day, across 5 categories.

Incidents (4)

AI company Eightfold sued for helping companies secretly score job seekers

Jan 21 (Reuters) - Eightfold AI, a venture capital-backed artificial intelligence hiring platform used by Microsoft, PayPal and many other Fortune 500 companies, is being sued in California for allegedly compiling reports used to screen job ... (https://incidentdatabase.ai/cite/1575#7495)
AI Incident Database 22d ago Jobs & economyFinance, VC & PE

Judge fines lawyers $12,000 over AI-generated submissions in patent case

Feb 3 (Reuters) - A Kansas federal judge has fined lawyers representing a patent holding company a combined $12,000 for filing documents with non-existent quotations and case citations that were generated by artificial intelligence, in the ... (https://incidentdatabase.ai/cite/1576#7496)
AI Incident Database 22d ago

10th Circuit orders lawyer to pay $1,000 for faulty AI citations

The Denver-based federal appeals court ordered a lawyer on Monday to pay $1,000 to the opposing side for submitting a legal filing with fake case citations generated by artificial intelligence. A three-judge panel of the U.S. Court of Appe ... (https://incidentdatabase.ai/cite/1577#7497)
AI Incident Database 22d ago

JADEPUFFER: Agentic ransomware for automated database extortion

Ransomware has had a human at the keyboard, or at least a human writing its script, since it was first established as a category of threat. The Sysdig Threat Research Team (TRT) has captured what we assess to be the first documented case of ... (https://incidentdatabase.ai/cite/1578#7498)
AI Incident Database 22d ago Agents & autonomy

News (61)

Why Europe’s Safeguards Against AI Disinformation Won’t Stop Russia’s Next Move

Tech Policy Press 22d ago Misinformation

The Next National Security Challenge Is Research Integrity

Tech Policy Press 22d ago Military & security

The UN Scientific Panel on AI's Preliminary Report Does Not Establish Its Independence

Tech Policy Press 22d ago

The Pope Found Babel in AI. Here's What Rabbis Saw

Tech Policy Press 22d ago

India’s Aadhaar Shows Foreign Dependencies Reach Beyond US-China

Tech Policy Press 22d ago

Lawsuit: Man used Grok to make 7K sex images of stepdaughter, then shot himself

More young girls sue X over Grok CSAM; X accused of shielding child predators.
Ars Technica 22d ago Children & education

Hackers can use 9 of the most popular AI tools to assemble massive botnets

"HalluSquatting" weaponizes LLMs' inability to say "I don't know."
Ars Technica 22d ago Military & security

Oz Joins Pillsbury For Top AI Role

International man of mystery and legal AI guru, Oz Benamram, has joined US law firm Pillsbury as its first Chief AI Officer. It’s understood that ...
Artificial Lawyer 22d ago Regulation

Legatics Data Rooms Launches as VDR Alternative

Legatics has launched its own ‘Data Rooms’ capability, which enables law firms to manage transactions and share confidential documents within a single workflow, and as ...
Artificial Lawyer 22d ago Regulation

Argentum targets the capital stack as the missing layer in AI infrastructure buildout

The AI infrastructure boom has trained the industry’s attention on silicon and power, but a more fundamental constraint within the capital stack is quietly throttling the speed of global data center deployment. As demand for AI compute continues to outpace infrastructure availability, a growing number of data center developers are stalling not because they lack megawatts […] The post Argentum targets the capital stack as the missing layer in AI infrastructure buildout appeared first on SiliconAN
SiliconANGLE AI 22d ago Environment

Pentagon opens applications for cyber apprenticeship program

DOD Chief Information Officer Kirsten Davies said last month that the apprenticeship had “already generated more than 70,000 inquiries” since it was first announced in late April.
NextGov/FCW 22d ago Military & security

Mexico's New Cyber Plan Faces Its First Real Test

The Latin American nation's cybersecurity plan — still in the expansion phase — has to survive its own knockout round during the FIFA World Cup.
Dark Reading (AI security) 22d ago Military & security

Inference chip startup SambaNova valued at $11B in $1B funding round

Chip startup SambaNova Inc. today announced that it has raised $1 billion in funding at a $11 billion valuation. General Atlantic led the Series F round with contributions from more than a dozen others. Intel Capital, Vista Equity Partners and JPMorgan Chase & Co. were among the participants. The investment follows a $350 million round […] The post Inference chip startup SambaNova valued at $11B in $1B funding round appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago Bias & fairnessFinance, VC & PE

How Ukraine won the first great robot war

In a new video, Science & Tech editor Patrick Tucker looks at how the narrative has shifted.
Defense One Technology 22d ago Agents & autonomy

Solidigm targets the intelligence layer as agentic inference pushes storage to center stage

The shift from model training to agentic inference is forcing a fundamental rethink of how artificial intelligence infrastructure is built and which components carry the most strategic weight. What was once treated as commodity plumbing is now being recognized as the intelligence layer where raw data becomes actionable intelligence. The rise of sovereign AI deployments […] The post Solidigm targets the intelligence layer as agentic inference pushes storage to center stage appeared first on Silic
SiliconANGLE AI 22d ago Agents & autonomy

‘Technology is the easy part’: Washington, D.C., turned to collaboration, data governance for America 250

With numerous 250th anniversary events packed into a few months, officials in Washington, D.C., said this summer has presented a "unique" challenge in terms of technology, communications and emergency management.
StateScoop 22d ago Regulation

Lone Attacker Uses AI to Breach AWS Cloud Environment in 72 Hours

The attacker exploited AI workflows, chained cloud weaknesses, and stolen credentials to extort a large Amazon customer.
Dark Reading (AI security) 22d ago Environment

10 insights from the Machina AI Summit: Physical AI moves from demos to deployment

Physical AI and robotics are moving beyond impressive demonstrations into a new phase of practical deployment, with companies now targeting specific, high-value use cases in manufacturing and logistics — production-ready systems capable of delivering measurable ROI. After years of research breakthroughs and impressive demonstrations, the focus for physical AI has shifted to real-world data, functional […] The post 10 insights from the Machina AI Summit: Physical AI moves from demos to deployment
SiliconANGLE AI 22d ago Agents & autonomy

Subpar Secret Service drone defense flagged in Trump assassination attempt report

The DHS inspector general office said an inexperienced counter-UAS operator and delayed technical support led to a missed opportunity to prevent the July 2024 incident. The post Subpar Secret Service drone defense flagged in Trump assassination attempt report appeared first on FedScoop .
FedScoop 22d ago Military & security

Alight and BNY launch integrated retirement plan

Alight (NYSE: ALIT), a leading benefits administration provider of health, wealth and leave solutions, today announced a collaboration with BNY (NYSE: BNY), a global financial services platforms company, to launch a retirement solution designed to offer defined contribution (DC) and defined benefit (DB) plan sponsors and participants deeper support for plan administration and investing.
Finextra AI 22d ago HealthcareFinance, VC & PE

Hoosier Childcare Providers Slam Education Requirement Rollback

About two-dozen Hoosier childcare providers and advocates overwhelmingly opposed looser statewide staff educational requirements during a Monday public hearing. “This is a profession. This is not just a babysitting job,” said Hanna Wetzel, the director of the Children’s Learning Center of Posey County. She spoke to representatives of the Family and Social Services Administration, the […]
The 74 (education AI) 22d ago Jobs & economyChildren & education

Pentagon opens application window for paid cyber apprenticeships

The initiative comes as the Pentagon faces challenges in competing with the private sector to hire established cyber experts. The post Pentagon opens application window for paid cyber apprenticeships appeared first on DefenseScoop .
DefenseScoop 22d ago Military & security

A Fresh Look at the Houthi Threat to Maritime Shipping

In 2024, Allison Minor wrote, “Solving the Houthi Threat to Freedom of Navigation,” where she argued the international response to Houthi attacks on shipping in the Red Sea has so far been inadequate and proposed a U.N.-led solution. Two years later, with global attention once again focused on maritime shipping activity, we asked Allison to review her arguments. Image: Petty Officer 1st Class Jonathan Word via DVIDSWhen you wrote about the Houthi threat in 2024, the activities in the Red Sea com
War on the Rocks 22d ago Children & education

Seizing the digital advantage at DOD

Defense Department CIO Kirsten Davies said her office is shifting away from its traditional role as a backend policy shop to become a forward-leaning strategic unit.
NextGov/FCW 22d ago RegulationMilitary & security

Appeals Court Says Religious Schools Can’t be Exempt From Maine’s Nondiscrimination Laws

Religious schools accepting public funds are required to follow Maine laws that protect against discrimination based on faith, gender identity and sexual orientation, a federal court has ruled. Crosspoint Church, which runs Bangor Christian Schools, and St. Dominic Academy in Auburn filed separate appeals in the United States Court of Appeals for the First Circuit […]
The 74 (education AI) 22d ago Bias & fairnessChildren & education

Blizzard eliminates 'small number' of roles in China

The news comes with Activision Blizzard's parent company cutting jobs across the board.
Game Developer (AI) 22d ago Jobs & economy

Dun & Bradstreet brings agentic credit and portfolio management workflows to Databricks

Dun & Bradstreet announced it is delivering agentic credit and portfolio management workflows, leveraging the D&B Commercial Graph™, available through the Databricks Marketplace and Databricks OpenSharing.
Finextra AI 22d ago Agents & autonomy

SBS embeds AI into core banking platform

SBS, the global financial technology company that more than 1,500 financial institutions rely on to digitally transform the way they operate, is launching SBS AI Foundation, embedding enterprise AI directly into the core banking, lending, and digital banking products its clients already run.
Finextra AI 22d ago Finance, VC & PE

NC High School Students Pitch Local Tech Companies at Annual Teamship Showcase

Local technology companies may have solved some of their biggest problems after hearing from North Carolina high school students during the annual Teamship Showcase in Durham on June 25. Teamship is an internship experience through the College Board where students are paired in groups to solve real-world problems for the businesses they are assigned to. […]
The 74 (education AI) 22d ago Children & education

Salesforce enhances Slackbot with connectors to the entire platform ecosystem

Salesforce Inc. today announced new updates for Slackbot, the company’s personal artificial intelligence agent in Slack, giving it full access to every part of the Salesforce ecosystem. “Slackbot can now reason over your entire Salesforce platform, so anything you can do in Salesforce, you can simply now do through Slackbot, just by asking,” Slack Chief […] The post Salesforce enhances Slackbot with connectors to the entire platform ecosystem appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago Agents & autonomy

Presentation: The Multi-Agent Approach: Building Reliable and Controllable Software Development Automation

Itamar Friedman discusses how architects and engineering leaders can break through the AI productivity ceiling using adaptive multi-agent systems. He shares insights on moving past simple autocomplete to resilient workflows by integrating autonomous testing, intelligent code review, and robust arbitration. Learn how to govern agent communication and build a context-driven SDLC that scales. By Itamar Friedman
InfoQ AI/ML 22d ago Jobs & economyAgents & autonomy

Iran’s environmental catastrophe has also wrecked its economy

Iran’s leaders could use the peace dividend to invest in fixing its severe environmental problems.
The Conversation UK Technology 22d ago Jobs & economyEnvironment

DeepFabric ships more than 50 AI agents for supply chain operations

Supply chain artificial intelligence startup DeepFabric today announced the general availability of an AI agent platform built for supply chain execution, and a roster of enterprise customers is already running it in production. The company’s platform drops specialized agents into a business’ operational workflows to recover margin, cut operating costs and speed up customer response. […] The post DeepFabric ships more than 50 AI agents for supply chain operations appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago Agents & autonomy

Ex-GitHub chief’s Entire opens distributed Git network for the AI agent era

Entire Inc., the developer-platform startup founded by former GitHub Chief Executive Thomas Dohmke, today launched a preview of a distributed Git network built to let artificial intelligence coding agents clone and push code without running into the rate limits of centralized hosting. The preview is open by waitlist, with active regions in the U.S., European […] The post Ex-GitHub chief’s Entire opens distributed Git network for the AI agent era appeared first on SiliconANGLE .
SiliconANGLE AI 22d ago Agents & autonomy

Renewed Middle East conflict will buffet world economy, IMF warns

The warning comes as President Donald Trump announced the U.S. ceasefire and peace process with Iran is “over.”
Politico Europe Technology 22d ago Jobs & economy

Kaon AI raises $60M to build the next generation of AI-driven interactive entertainment

Kaon AI, a generative artificial intelligence platform and research outfit formerly known as FlowGPT that provides online interactive entertainment, today announced it has raised $60 million across recent funding rounds. The San Mateo company is backed by numerous technology sector venture capital investors, including B Capital, Redpoint Ace, Goodwater Capital and DCM. Today’s announcement follows […] The post Kaon AI raises $60M to build the next generation of AI-driven interactive entertainmen
SiliconANGLE AI 22d ago Finance, VC & PE

Les agents IA de GitHub peuvent faire fuiter un dépôt privé via une injection de prompt

En postant un simple message d’erreur dans un dépôt public GitHub, des chercheurs en sécurité ont montré qu’il était possible de pousser les agents IA de GitHub pilotant les « workflows » d’un projet de développement à livrer des informations provenant d’un autre dépôt privé d’une même organisation. Un simple ticket rapportant une erreur dans […]
Next (FR, ex-INpact) 22d ago Agents & autonomy

Creating synthetic life in a lab? SpudCell falls short of the goal, but raises even more useful questions

The goal of creating synthetic cells is not to replace nature, but to learn more deeply about biology and reengineer it to help society.
The Conversation Technology 22d ago Biotech

Alibaba shares spike 12% in Hong Kong as T-Head chips, AI revenue fuel earnings optimism

Shares of Alibaba Group Holding surged to a high of 13.8 per cent in Hong Kong on Wednesday as equity analysts expect revenue to reaccelerate in the June quarter, driven by growing demand for artificial intelligence and narrowing losses in food delivery. The gain, the strongest this year, came before the company closed up 12.2 per cent at HK$107.5 (US$13.71). Rivals Tencent Holdings and Meituan saw their shares grow 3.8 and 3.3 per cent, respectively, while the Hang Seng Tech Index increased by.
SCMP Tech (HK/CN) 22d ago Bias & fairness

The climate case against leather

The climate case against beef is now almost boringly well-established: It is, by far, the most carbon-intensive food in the world, amounting to about 6 percent of all global greenhouse gas emissions. But cows don’t just become burgers and steaks. They also become shoes, bags, couches, and car interiors — products often marketed with a […]
Vox Future Perfect 22d ago Environment

Opinion: The Final Piece of the Ed-Tech Backlash Has Finally Arrived

I have been a high school teacher for almost three decades, spending almost all that time teaching seniors about American civics. My teaching tenure has overlapped with the rise of the very trends now engulfing our educational system: I have watched my students embrace smartphones, social media, online learning and now artificial intelligence. But recently […]
The 74 (education AI) 22d ago Children & education

Outcry as Meta lets users make AI images from public Instagram profile pics

The tech giant said people can opt out - but privacy campaigners called it a "recipe for disaster".
BBC Technology 22d ago Privacy

Ivermectin isn’t a cancer miracle drug, but influencers claim otherwise – here’s how to avoid sprinting past scientific evidence

Science works at a much slower pace than social media, opening a large window for early findings to be taken at face value and misinformation to spread.
The Conversation Technology 22d ago MisinformationHealthcare

AI in Banking: What is Myth and What is Reality?

At EBAday in Copenhagen, Tapan Agarwal, Head of Payments Solutions and Krishnan Srinivasan (KS), President and Region Head, Europe and CIS Markets, Intellect Design Arena, discussed the myths and truths surrounding AI with FinextraTV. KS began by listing the many ways he has experienced AI being used effectively within banking, firmly defining it as a 'reality' rather than a myth, but notes the constraints that must be navigated within the EU act. Tapan then explains his view on the three primar
Finextra AI 22d ago Finance, VC & PE

Profile releases AI orchestration platform for banks

Profile (ATH: PROF), a global financial technology leader with a presence in more than 50 countries, today announced the launch of ProfileOne, an enterprise-grade agentic AI orchestration platform designed to help financial institutions move from insight to controlled execution across banking and investment management.
Finextra AI 22d ago Agents & autonomyFinance, VC & PE

Have a 401(k)? Help ProPublica Investigate What’s Really Happening to Your Money.

The post Have a 401(k)? Help ProPublica Investigate What’s Really Happening to Your Money. appeared first on ProPublica .
ProPublica (Machine Bias) 22d ago Finance, VC & PE

Qu’est-ce que le GDID de Windows qui a permis au FBI de retrouver un suspect ?

Le FBI a pu arrêter un pirate en se servant d’une information délivrée par Microsoft : le GDID. Il s’agit d’un identifiant généré par Windows, spécifique à la machine et ne pouvant pas être changé simplement. En revanche, cet identifiant n’est pas pensé initialement pour la surveillance. Explications. Le département américain de la Justice (DoJ) a […]
Next (FR, ex-INpact) 22d ago Privacy

Meniga integrates with bank AI assistants for conversational banking

Meniga has launched Fini, an advanced, standards-compliant MCP server that enables banks to bring agentic AI into their digital channels using Meniga's financial intelligence.
Finextra AI 22d ago Agents & autonomyFinance, VC & PE

BlockAPT launches AI-powered cyber defence for SMEs

Cyber defence firm BlockAPT has launched Pivotra AI Essential, an AI-powered cyber defence solution built specifically for small and medium-sized enterprises (SME) in the financial services space to help them keep up and stay protected in the era of machine-speed cyber threats.
Finextra AI 22d ago Military & security

Physical AI ‘space race’: can Europe compete with China and the US in humanoid robotics?

European firms say they are fighting to secure a foothold in physical AI – the integration of artificial intelligence into robotics and machinery – as China and the United States take an early lead in the sector, with industry insiders warning the continent faces the threat of further deindustrialisation if it fails to establish a competitive industry. “You see China and the US … because of AI … typically they are considered the leaders, but do not count out Europe,” said David Kehr, president..
SCMP Tech (HK/CN) 22d ago Agents & autonomy

Alipay upgrades Tap! devices for agentic commerce

Alipay today announced the enhancement of its Tap! services by upgrading the Alipay Tap! devices widely used by millions of merchants into an AI agent-powered network, building the world’s first large-scale AI-powered offline business operations network.
Finextra AI 22d ago Agents & autonomy

Washington Law Says to Alert the Public When Doctors Are Accused of Misconduct. It Can Take Months.

The post Washington Law Says to Alert the Public When Doctors Are Accused of Misconduct. It Can Take Months. appeared first on ProPublica .
ProPublica (Machine Bias) 22d ago Regulation

Opinion: The AI licensure debate is missing the point of licensure

“AI will reshape medicine. The physicians who answer for the outcome must lead the way it enters patient care," write Afnan R. Tariq and Ami Bhatt.
STAT News (health AI, headlines) 22d ago Healthcare

State IDs for AI Agents: Will Estonia Set a Precedent?

The world's digital testing ground plans to help people use AI agents for government purposes.
Dark Reading (AI security) 22d ago Agents & autonomy

The Davis Wing, the B-24 Liberator, and the Self-Made Bet That Paid Off

Editor’s note: This is the second article in a limited series celebrating American defense technologies born from wartime and their effects on broader national security, politics, and society. This series will run for several weeks to commemorate America’s 250th anniversary, and winners will be selected by a reader vote undertaken through our newsletter later this summer. Prior installments can be found at the Arsenal of Innovation page. The most produced American military aircraft of World War
War on the Rocks 22d ago Military & security

Mainland Chinese tech firms find more than just deep capital pools in Hong Kong

Mainland Chinese technology companies newly listed on Hong Kong’s stock exchange are deepening their engagement with the city, tapping not only its capital markets but also its global connectivity to refine products, forge international partnerships and expand overseas, according to executives. For Beijing-based service robot maker Yunji Technology, Hong Kong has become a key gateway to global markets since its listing in the city in October. “If we use one word to describe what Hong Kong offers
SCMP Tech (HK/CN) 22d ago Agents & autonomy

Starlink freezes new sign-ups in seven Kenyan counties

On Techpoint Digest, we discuss how Starlink has frozen new sign-ups in seven Kenyan counties, how Andrea Aid wants to improve medical crowdfunding, and how South Africans claim Facebook is restricting accounts without warning.
Techpoint Africa 22d ago Healthcare

China weighs open-weight AI’s security risks against national tech innovation strategy: researchers

China is facing a delicate regulatory balancing act as it weighs the security risks of open-weight AI models against an innovation strategy that has been crucial to its race for technological supremacy with the United States, researchers said. Traditionally, open-weight models – which allow anyone to download code for free and run it on local hardware – have lagged months behind proprietary frontier models. But recent releases from Chinese labs have significantly narrowed that gap. Zhipu AI’s...
SCMP Tech (HK/CN) 22d ago Regulation

Billionaire companies want to pull up the ladder behind them

Big brands want the CMA to gut app store security to dodge fees. Regulators must fix real problems — without breaking the trusted ecosystem small developers depend on.
Politico Europe Technology 22d ago Regulation

How a Vinyl Record Resurgence Helped Me Understand the Future of AI in Education

Streaming solved the problem of access. Now, we must solve the problem of engagement.
EdSurge (AI in education) 22d ago Children & education

Social Media Bans Alone Won’t Protect Kids, Chinese Report Finds

In recent years, countries around the world have rolled out policies to curb minors’ social media usage. A recent Chinese university report assesses their efficacy.
Sixth Tone (CN) 22d ago Children & education

Field notes (33)

Our approach to government and national security partnerships

Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountability, and public safety.
OpenAI 22d ago Military & securityTransparency

Double Agents: Defensive AI Agents Magnify Cyber Risks

Introduction New research from AI Now demonstrates a critical attack vector in popular AI agents, built by Anthropic and OpenAI, when used for defensive purposes that actually turn the agent against its user. Read the full blog post explaining the proof-of-concept exploit and a policy brief with key takeaways below. The post Double Agents: Defensive AI Agents Magnify Cyber Risks appeared first on AI Now Institute .
AI Now Institute 22d ago RegulationMilitary & security

Friendly Fire: Hijacking Defensive Cyber AI Agents for Remote Code Execution

Exploit Brief We are revealing a proof-of-concept exploit that enables remote code execution in Anthropic’s Claude Code CLI (with Claude Sonnet 4.6 & 5, Opus 4.8) and OpenAI’s Codex CLI (with GPT-5.5) when employed to defensively assess the security of an open-source or third-party library. Our attack only requires an out-of-the-box configuration of Claude Code […] The post Friendly Fire: Hijacking Defensive Cyber AI Agents for Remote Code Execution appeared first on AI Now Institute .
AI Now Institute 22d ago Military & securityAgents & autonomy

Policy Brief: Friendly Fire

Topline Summary AI Now’s latest research demonstrates a critical attack vector on popular AI agents, built by Anthropic and OpenAI, when used for defensive purposes that actually turn the agent against its user. Attackers can use these models’ existing weaknesses to execute malicious code on a system deploying an AI agent when used for often-advertised […] The post Policy Brief: Friendly Fire appeared first on AI Now Institute .
AI Now Institute 22d ago RegulationAgents & autonomy

Helping K–12 educators build practical AI skills

OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for the classroom.
OpenAI 22d ago Children & education

DIVERSSITY wins the Innovation Factory Women Entrepreneurs 2026 with AI-driven mental health support for neurodiverse adolescents

The post DIVERSSITY wins the Innovation Factory Women Entrepreneurs 2026 with AI-driven mental health support for neurodiverse adolescents appeared first on AI for Good .
AI for Good (ITU) 22d ago Healthcare

Data for Agents

Hugging Face Blog 22d ago Agents & autonomy

Sneha Revanur on how a small team of activists helped pass America’s landmark AI safety laws

The post Sneha Revanur on how a small team of activists helped pass America’s landmark AI safety laws appeared first on 80,000 Hours .
80,000 Hours 22d ago Safety & alignment

From Robots to Rulebooks: How AI for Good’s Youth Zone Tackled EdTech on Day Two

The post From Robots to Rulebooks: How AI for Good’s Youth Zone Tackled EdTech on Day Two appeared first on AI for Good .
AI for Good (ITU) 22d ago Agents & autonomy

AI Surveillance Is Being Supercharged–And It Will Chill Social Progress

Senior research fellow Jon Penney and co-author Bruce Schneier argue that widely deploying AI surveillance could be corrosive to democracy. The post AI Surveillance Is Being Supercharged–And It Will Chill Social Progress appeared first on The Citizen Lab .
The Citizen Lab 22d ago Privacy

From San José to Geneva: Canvas of the Future winner unveiled

At the AI for Good Global Summit, the Canvas of the Future returned for the third year in a row. The competition invited creators, educators, technologists and artists from all disciplines to submit an original, high-resolution AI-powered or AI-enhanced digital image exploring how artificial intelligence is reshaping education and work. The post From San José to Geneva: Canvas of the Future winner unveiled appeared first on AI for Good .
AI for Good (ITU) 22d ago Children & education

Rewriting Bun in Rust

Rewriting Bun in Rust Jarred Sumner has been promising this blog post ( since May 9th ) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it was worth the wait. This is a detailed description of an extremely sophisticated piece of agentic engineering, featuring dynamic workflows, trial runs, adversarial review and all sorts of other interesting tricks. Jarred spends the first half of the post praising Zig for getting Bun this far. The
Simon Willisons Weblog 21d ago Agents & autonomy

Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

2 years after our first coverage, we return with Modal's other cofounder to explore why Agent Experience is working now, and everything they have learned building the new agent cloud.
Latent Space 22d ago Agents & autonomy

Will Enterprises Ever Choose SpaceX’s Grok And Cursor?

SpaceX is a rocket company. It also became an AI company the day its founder Elon Musk folded in xAI, inheriting a frontier model in Grok, the planet-scale Colossus training cluster, and the live data feed of the X platform (formerly Twitter). The IPO headlines fixated on the trillionaire status, the stock price, and the […]
Forrester AI blog 22d ago Finance, VC & PE

Embedding Critical Media Literacy in Ontario Schools’ AI Policies

The rapid growth of artificial intelligence (AI) and of generative AI, in particular, has become a significant focus in education. While generative AI tools offer many benefits for educational ...
Centre for International Governance Innovation 22d ago Children & education

Childhood and Education #20: Phones and Screens

We have a respite, so I thought I’d tackle various thoughts on children, phones and screens.
Dont Worry About the Vase (Zvi) 22d ago Children & education

Vom „Jedermannsrecht“ zum Privileg für Wenige

Die Informationsfreiheit in Deutschland befindet sich in einer Krise. Bereits die Ampelregierung hatte ihr Versprechen nicht eingelöst, die Informationsfreiheitsgesetze zu einem Transparenzgesetz weiterzuentwickeln. Dieser Trend gipfelte im Papier des Koalitionsausschusses vom 2.7.2026. Im Rahmen einer Paketlösung hat die Bundesregierung angekündigt, das IFG umfassend zu ändern. Die politischen Vorschläge des Koalitionsausschusses sind aus unserer Sicht weitreichend, sodass die Informationsfreih
Verfassungsblog (EU law incl AI) 22d ago Transparency

Vikas Kumar

Vikas Kumar is Professor of International Business at the University of Sydney Business School. His research focuses on international business, India, emerging markets, AI, geopolitics and capability ...
Lowy Institute 22d ago Children & education

(De)Valuing Citizenship

Last Tuesday, the US Supreme Court released its final merits opinion of its October 2025 term. In Trump v Barbara, a razor thin 5-4 majority deemed the President’s attempt to deny American citizenship to children born on U.S soil to immigrant parents who are undocumented or present on certain visas unconstitutional. The decision is a rare and important win for immigrants and American constitutional democracy. But Barbara should not be remembered as an example of principled judicial resistance ag
Verfassungsblog (EU law incl AI) 22d ago Children & education

The T&S newsletters that I read regularly

These are the independent voices that help me make sense of online speech, regulation and the weird, messy business of governing the internet.
Everything in Moderation (Ben Whitelaw) 22d ago Regulation

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tuned its Deep Agents harness for NVIDIA Nemotron 3 Ultra, achieving the highest accuracy among open models, while completing more tasks at higher throughput and running at 10x […]
NVIDIA Blog (AI) 22d ago Agents & autonomy

European Marketers Say AI Won’t Replace Employees, But The Reality Is More Complicated

Will AI replace marketing jobs? European marketers largely believe the answer is no. Yet many organizations have already reduced headcount or replaced employees with AI. The reality is more nuanced: AI is not simply replacing marketers — it is redesigning how marketing work gets done. As AI takes on more execution-oriented tasks, the value of human contribution shifts toward strategy, creativity, judgment, and leadership. The organizations that succeed will not be those that deploy the most AI t
Forrester AI blog 22d ago Jobs & economy

Is Life Just Different?

The idea of ‘biological agency’ — that life devises its own goals and behaves accordingly — complicates our understanding of what it means to be alive. But does it serve a scientific purpose? The post Is Life Just Different? first appeared on Quanta Magazine
Quanta Magazine AI 22d ago Biotech

Why I'm Updating the Election Playbook I Spent a Decade Building

Introducing the kaleidoscopic minefield — more vectors, faster turns, and the threats we haven't named yet.
Anchor Change (Katie Harbath) 22d ago Misinformation

Android and the Art of Regulatory Self-Harm

Europe keeps asking where its technology champions are. In Google Android, the Court of Justice of the European Union (CJEU) offered part of the answer: build a successful platform, and Brussels may spend the next decade treating its architecture as evidence. The CJEU’s final judgment in Google Android, handed down last week, will be celebrated ... Android and the Art of Regulatory Self-Harm The post Android and the Art of Regulatory Self-Harm appeared first on Truth on the Market .
Truth on the Market (digital regulation) 22d ago Regulation

Global AI Governance Challenges

Key assessments from the United Nations' independent scientific panel's report on AI | Edition #304
Luizas Newsletter (AI governance) 22d ago Regulation

How the British press has undermined the ECHR over many years: new study – Ekaterina Balabanova and Gemma Horton

The UK’s immigration and asylum bill has proposed restricting how the European Convention on Human Rights (ECHR) is interpreted and applied in the UK to make it easier to deport migrants. For years, critics have argued that the ECHR undermines the UK’s border security by prohibiting deportations on the basis of Article 8, the right […]
Inforrm (media law) 22d ago Regulation

Cybersecurity and the Gap Between Skill and Ability

Last week, national security agencies from the Five Eyes—that’s the rich, English-language-speaking countries club—jointly released a statement warning of the increasing cyber risks of AI models: in particular, their ability to autonomously hack into systems and networks. The statement was more measured than some of the breathless headlines about it, and the advice they gave is pretty much the standard advice everyone gives—albeit with newfound urgency. Internet risks are nothing new, and cybera
Bruce Schneier — Schneier on Security 22d ago Military & security

Zewelanji Nalwamba

Zewelanji is a final-year International Relations student at the University of Leicester, currently interning with the UK Nuclear Deterrence Network. Her research focuses on the geopolitics of ...
Royal United Services Institute 22d ago Children & education

AI Is Eating the Book World

In 2011, Marc Andreessen, a key figure in California’s venture capital scene, coined the phrase: “Software is eating the world.” The phrase describes the spread of software into everyday life and the displacement of physical business models. This process continues in an unexpectedly literal sense: AI companies purchase used books, scan them, and dispose of them to gather input for their models. The reason for this seemingly cumbersome method is the expectation that it will fall under the fair us
Verfassungsblog (EU law incl AI) 22d ago Jobs & economyFinance, VC & PE

The Problems with “General Purpose AI Detectability”

As AI-generated media flood our information ecosystems, detecting synthetic content has become an urgent regulatory challenge – in fact, not one challenge but many, as synthetic media breeds problems across a range of digital contexts, including deepfakes and disinformation, scamming, and content moderation. The EU's new "Code of Practice on Transparency of AI-Generated Content" – the first concrete articulation of Article 50(2) AI Act, the EU's approach to AI-content detection – gives sensible
Verfassungsblog (EU law incl AI) 22d ago RegulationMisinformation

China's new AI rules: Ethics, AI agents and anthropomorphic AI

China introduced three new regulatory developments addressing AI ethics, AI agents and anthropomorphic AI, which reflect a regulatory shift from broad AI principles toward more detailed, operational ...
IAPP 22d ago RegulationAgents & autonomy

Can a Burnham government make Britain a global leader in science and technology?

Andy Burnham is near-certain to succeed Keir Starmer as UK prime minister. He will inherit a world in which technological leadership increasingly shapes economic prosperity, military capability and ...
Chatham House 22d ago Military & security

Policy (12)

250th Anniversary of the Adoption of the Declaration of Independence

US Federal Register 22d ago

Pipeline Safety: Repair Criteria for Hazardous Liquid and Gas Transmission Pipelines

PHMSA proposes to modernize and to clarify the anomaly response criteria in the Federal pipeline safety regulations for gas transmission and hazardous liquid pipelines. Driven by twenty years of technological development, modern engineering concepts allow operators to identify, schedule, and remediate pipeline anomalies more effectively and in a less costly manner. PHMSA proposes incorporating these improved safety practices into its regulations by finalizing certain safety improvements advanced
US Federal Register 22d ago Regulation

Intellectual property strategies for AI-enabled health innovation

Navigating intellectual property (IP) issues for new innovations has never been simple. With healthcare solutions based on artificial intelligence (AI), it can be especially complex. Innovators, ...
ITU 22d ago Copyright & IPHealthcare

Press Releases, Media Advisories and Statements

ITU’s AI for Good Global Summit brings world leaders, AI pioneers and breakthrough technologies to Geneva 23 June 2026 - The world's leading platform for AI solutions, skills, standards and policy ...
ITU 22d ago Regulation

Press Briefing Transcript: Julie Kozack, Director, Communications Department, July 9, 2026

And then on the other side, we have a positive kind of demand and productivity shock, which is coming from the technology cycle and particularly AI-led investment. So, there's two forces that are ...
IMF 22d ago Jobs & economyFinance, VC & PE

Kenya Launches Bold New Data Strategies to Power Sustainable Development

The NPAEEA, supported by the Global Program on Sustainability (GPS), introduces a groundbreaking approach to integrating environmental and economic data. It prioritizes the development of six key ...
World Bank 22d ago Environment

Promoting the Rule of Law

Specialised bodies measure how member states respect standards: the Group of States against Corruption monitoring anti-corruption efforts and the European Commission for the Efficiency of Justice ...
Council of Europe AI 22d ago Regulation

Financial services bill concludes Lords committee stage

Members of the House of Lords reached the end of detailed examination of the Financial Services and Markets Bill in committee stage on Wednesday 8 July. The Financial Services and Markets Bill aims to ...
UK Parliament 22d ago Regulation

Priority Open Recommendations: Department of Education

What GAO Found In May 2025, GAO identified eight priority recommendations for the Department of Education. Since then, Education has implemented one of those recommendations. In June 2026, GAO identified an additional three priority recommendations, bringing the total to 10. GAO is highlighting the following three areas that warrant timely and focused attention: Improving the federal student aid system, Protecting sensitive information, and Managing financial risks associated with charter school
US GAO Reports 22d ago Children & education

Security Council LIVE: Open debate on sexual violence in conflict

The UN Security Council is holding an open debate Tuesday honouring the promise of international law to survivors of sexual violence in conflict as more reports emerge about warring parties using rape ...
United Nations 22d ago Regulation

Global Economy in Crosscurrents of War and Technology

Global growth is projected at 3.0 percent for 2026 and 3.4 percent for 2027, broadly unchanged cumulatively from the April 2026 World Economic Outlook. The outlook is uneven: The war shock is weighing ...
IMF 22d ago Jobs & economy

Thailand Monthly Economic Monitor, June 2026

The Bank of Thailand kept its policy rate unchanged. Financial market conditions improved, but the external position weakened as the higher oil import bill widened the current account deficit and ...
World Bank 22d ago Regulation

Research (62)

Beyond Thermal Imaging: Inferring Thermophysical Properties from Time-Resolved Thermal Observations

Inferring latent physical properties from sensory observations is a fundamental challenge in machine perception. Among available sensing modalities, thermal imaging is particularly promising because temperature evolution is directly governed by heat-transfer physics and therefore encodes information about underlying thermophysical properties of a scene. Recovering spatially resolved thermophysical properties from thermal observations could transform applications ranging from digital twins and in
arXiv 22d ago

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context. This paper presents a comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. We express these mechanisms in a common recurrent-memory notation, making explicit how they differ in expressivity, memory decay, erase and write c
arXiv 22d ago

Efficient Safety Alignment of Language Models via Latent Personality Traits

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We
arXiv 22d ago Safety & alignment

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulner
arXiv 22d ago Safety & alignmentHealthcare

Closed-Loop Dynamic Validator Node Scaling in Private Substrate Blockchains Using Takagi-Sugeno Fuzzy Inference

Private blockchain networks run with fixed node configurations that cannot adapt to changing workload conditions. Too many nodes serving a light workload waste resources; too few nodes facing heavy demand slow block production and degrade finalisation. The right validator count is hard to determine, as it depends on overlapping factors that shift over time. This paper presents a Takagi-Sugeno (TS) fuzzy inference system that reads live blockchain parameters (block production time, block size, an
arXiv 22d ago

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline tra
arXiv 22d ago Safety & alignmentAgents & autonomy

Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

Artificial intelligence (AI) is beginning to reshape actuarial practice, particularly in domains that require reasoning over unstructured documents, heterogeneous data sources, and regulated decision workflows. Actuaries now face a design space that ranges from traditional rule-based automation to large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent ``agentic'' systems that plan, retrieve, call tools, and reflect. This paper examines how these emerging architecture
arXiv 22d ago RegulationJobs & economy

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that auditing these segmentation tasks is complicated by a common property of modern segmentation datasets: expert-annotated gold labels are expensive, so abundant machine-generated (silver) labels are added to limit annotation cost. This matters because the reference used to judge a model can itself be biased. In this study, we present the first fairnes
arXiv 22d ago Bias & fairnessHealthcare

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces
arXiv 22d ago Safety & alignmentAgents & autonomy

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that subs
arXiv 22d ago

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM}, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled compar
arXiv 22d ago Safety & alignmentFinance, VC & PE

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy thr
arXiv 22d ago Regulation

ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation

Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within the corrupted input sequence. This reliance can be limiting in real-world scenarios, where time series often exhibit non-stationary dynamics, weak temporal correlations, and infrequent patterns that are difficult to reconstruct from nearby observations alone. In this paper, we propose ALER-TI, Aligned Latent Embedding Retrieval for Time Series Imput
arXiv 22d ago

Towards Agentic AI Governance: A Preliminary Assessment

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, introducing new ethical and governance challenges. This paper presents a systematic review of the emerging literature on agentic AI governance. Our analysis identifies features that distinguish agentic AI from traditional systems and why it warrants targeted gover
arXiv 22d ago RegulationAgents & autonomy

User identity conditions moral wrongness ratings in non-reasoning large language models

This study adopts a behavioural bottom-up approach to AI value alignment to investigate whether an implicitly conveyed user identity shifts the moral evaluations of large language models (LLMs). Through a structured, multi-turn conversational protocol across 12,000 interactions, we evaluate AI value alignment in two non-reasoning models, gpt-4.1-mini-2025-04-14 and gemini-2.5-flash-lite. Rather than instructing the models to adopt a persona or prompting them with explicit moral stances, the user
arXiv 22d ago Safety & alignmentFinance, VC & PE

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoning, and vehicle trajectory generation and control. Prior knowledge- and model-based approaches typically focus on scene or trajectory components in isolation, while diffusion-based methods attempt end
arXiv 22d ago

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs)
arXiv 22d ago Safety & alignmentJobs & economy

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient care. This survey examines recent progress in medical LLMs, focusing on reasoning applications and requirements. We present a dual-view approach that connects clinical practice with computational methods. On the clinical side, we establish a five-level competency scheme following Miller's Pyramid, progressing from knowledge recall to dynamic case management. On
arXiv 22d ago Healthcare

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under
arXiv 22d ago Bias & fairnessPrivacy

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell developmental paths, trajectory inference (TI) is critical. However, existing methods require extensive manual intervention and proficiency in heterogeneous tools, posing a significant barrier to efficient TI analysis. To bridge this gap, we propose SpaCellAgent, an autonomous large language model (LLM) multi-agent framework that automates end-to-end sp
arXiv 22d ago Agents & autonomy

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams. This has led to the development of continual self-supervised learning (CSSL), a rapidly growing area that lacks a dedicated, systematic review. In this work, we present a comprehensive survey of CSSL for vision, with connections to emerging vision-language settings. First, we analyze existing evaluation
arXiv 22d ago Agents & autonomy

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear transformations that propagate information through modern DNNs, an unexplored mechanism of adversarial vulnerability. Specifically, we investigate transformer-based vision-language models, whose linear layers admit interpret
arXiv 22d ago Finance, VC & PE

HumAIN: Human-Aware Implicit Social Robot Navigation

Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orientation. We present Human-Aware Implicit Social Robot Navigation (HumAIN), a novel framework that fuses implicit social cues directly into the planning loop via knowledge distillation. We first employ a transformer-based teacher model that fuses rich multi-modal inputs, including historic images, skeletal keypoints, robot state, and a robot's target goal, to lea
arXiv 22d ago Agents & autonomy

Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts

Claims about the universality of human concepts have been predominantly assessed through linguistic similarity across languages and cultures. However, words are effective as communication devices because they compress rich experiential variation into shared conventions, potentially obscuring hidden individual and cultural differences in how concepts are mentally represented. Here, we analyse 2.6 billion human-made sketches of common concepts from 236 countries and territories to examine conceptu
arXiv 22d ago

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent with the answer it accompanies. Unlike faithfulne
arXiv 22d ago Safety & alignmentTransparency

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and down
arXiv 22d ago Healthcare

Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning

In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show that state-of-the-art multimodal large language models struggle with this setting, particularly due to limited compositional reasoning and sensitivity to prompt construction. In this work, we propose a Tree-of-Thoughts (ToT) reasoning framework for T2I-ICL that introduces a multi-stage reasoning and selection layer that
arXiv 22d ago

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through exp
arXiv 22d ago Safety & alignmentAgents & autonomy

Learning social norms enhances compatibility in dynamic human-AI coordination

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social interaction structures. Yet they often fail to coordinate with humans in an effective, considerate, and natural manner. We hypothesize that this gap arises bec
arXiv 22d ago Agents & autonomy

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video predi
arXiv 22d ago Agents & autonomy

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight planning traditionally relies on classic algorithms that struggle to incorporate flexible human preferences. We present FRAMe, an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent. Our system integrates a planner LLM with a multi-modal coach agent and retrieval au
arXiv 22d ago Agents & autonomy

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-scale datasets with structured aesthetic annotations. We introduce MADB, a large-scale dataset and benchmark comprising 9,999 tracks annotated by 30 trained annotators. Each track is rated by around 10 annotators across 10 perceptual dimensions and one overall score, with addition
arXiv 22d ago

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance.
arXiv 22d ago RegulationAgents & autonomy

Computing with Stochastic Oracles in AI-Augmented Computation

The Stochastic-Oracle Turing Machine (SOTM) framework models AI-augmented computation as the interaction of a probabilistic Turing machine with an oracle whose responses are drawn from context-dependent distributions. This paper studies what an SOTM can achieve under two oracle-response schemes: in a cached-response oracle, each distinct query receives one response that is reused on later calls to the same query, while in a fresh-response oracle, each call returns an independent response. In bot
arXiv 22d ago

Notes on technical alignment via human-like social drives

Alignment Forum 22d ago Safety & alignment

The Behavioural Reflection Test: A time-efficient measure of reflective reasoning in morally and epistemically charged decisions

How readily people override intuitive conclusions through reflection shapes how they navigate dense information environments with reliable and misleading sources; yet the effectiveness of a prominent measure, the Cognitive Reflection Test (CRT), is eroded by widespread exposure to classic items and leaves open how such tendencies manifest more generally in decision style and linguistic expression. The Behavioural Reflection Test (BRT) addresses these issues with a brief open-ended measure of rea
arXiv cs.HC 22d ago Environment

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including e
HuggingFace Daily Papers 22d ago Agents & autonomy

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate me
HuggingFace Daily Papers 22d ago HealthcareAgents & autonomy

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target settings typically involves replicating the single-target processing for each individual object, resulting in reduced frame rates (FPS) with unbounded latency as target count increases. Built upon Segment Anything 2 (SAM2), we propose SAM-MT, which addresses this by transforming the
HuggingFace Daily Papers 22d ago Privacy

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset
HuggingFace Daily Papers 22d ago HealthcareBiotech

A Quantized Native Runtime for On-Device Semantic Audio Generation

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present aria, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio~3 (SA3) on ordinary GPUs, CPU-only machines, and a Raspberry~Pi~5, with no Python or deep-learning framework underneath. Our main contribution is a study of quantization: running the model at lower numerical precision to fit
HuggingFace Daily Papers 22d ago Environment

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it diff
HuggingFace Daily Papers 22d ago Agents & autonomyEnvironment

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templati
HuggingFace Daily Papers 22d ago Agents & autonomy

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed required for interactive settings. Conversely, existing online methods enable real-time synthesis but often sacrifice controllability or struggle with complex text semantics and long-horizon goals due to limited context wi
HuggingFace Daily Papers 22d ago Agents & autonomy

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed mu
HuggingFace Daily Papers 22d ago Agents & autonomy

GenAI Success Metrics: Look Beyond Reduced Workload

Matt Harrison Clough / Ikon Images The Research The authors performed a four-year, fixed-window observational analysis of administrative work inside a large U.S. public higher-education institution. Generative AI tools were introduced to executive leaders, operational leaders, and student-facing professionals throughout the organization in 2026. Staffing levels and work hours remained stable across the period studied. […]
MIT Sloan Management Review AI 22d ago Children & education

Autopoietic Bodily Integrity: A Biological Approach to Hybrid Minds

Recent cases of forced explantation of neurotechnologies seem to be grounded on a pre-theoretical naturalist conception of the body as an entity that cannot have a non-biological object as a proper part. However, this conception has been challenged by functional approaches, according to which if an artifact robustly contributes to the function of a body, it is part of it and should be legally treated as such. Bublitz ( 2022 ) argues that a series of problems would result from revising the law to
Minds and Machines 22d ago RegulationBiotech

The Not So Innocent Nudge: Why the Hooked Nudges for Smartphone Apps Undermine Autonomy

Nudges are considered ‘innocent’ interventions because many nudges only have a small effect on how people behave. However, one particular nudging strategy has been held to have a great impact on behavior. This is the ‘Hooked’ model for smartphone applications, which consists of nudges such as the brightly colored icon of a smartphone app, the possibility to give and receive ‘likes,’ and infinite scrolling. I examine how these nudges affect the value of autonomy. As I contend, smartphone users wi
Philosophy & Technology 22d ago Agents & autonomy

Zero-shot semantic landmark-based visual odometry using foundation models for unstructured planetary exploration

Precise autonomous navigation on unstructured planetary surfaces is a critical prerequisite for future exploration missions, particularly in GNSS-denied environments such as the Lunar South Pole or Martian deserts. Traditional Visual Odometry (VO) methods, which rely on tracking low-level geometric features (e.g., corners), often fail under the extreme illumination contrast of the Moon or the textural monotony of the Martian regolith. In this work, we present a zero-shot semantic landmark-based
Frontiers in Robotics and AI 22d ago PrivacyEnvironment

Formal Logic Inference Guided Uncertainty Quantification for Personalized Federated Learning

Federated Learning (FL) enables privacy-preserving model training across heterogeneous distributed systems, such as smartgrid forecasting or traffic-flow prediction from geographically dispersed sensors and devices. A key challenge in such settings is capturing client-specific patterns while addressing data heterogeneity and uncertainty at scale. Existing approaches, including Bayesian Neural Networks (BNNs) and clustering-based methods, struggle with scalability and consistent personalization.
JAIR 22d ago Privacy

PaSTO-GNN: prompt-aware spatio-temporal graph neural networks for automatic essay scoring

Automatic Essay Scoring (AES) aims to evaluate the quality of written essays automatically, providing fast, consistent, and objective assessments of students' writing ability. Existing deep learning approaches—including recurrent, convolutional, and transformer-based models—primarily focus on textual semantics, yet they often overlook the spatio-temporal nature of essay composition, where meaning evolves across sentences and paragraphs through discourse progression. To address this gap, this stu
Frontiers in Artificial Intelligence 22d ago Children & education

Machine learning-based fetal health prediction and development of smart web application

IntroductionFetal health monitoring is critical for early identification of pregnancy-related risks. Manual interpretation of cardiotocography (CTG) signals is subjective and variable among healthcare professionals.MethodsA machine learning-based framework was developed to classify fetal health into Normal, Suspect, and Pathological categories using CTG-derived clinical features. The dataset was preprocessed through duplicate removal, normalization, class balancing using SMOTEENN, multicollinear
Frontiers in Artificial Intelligence 22d ago Healthcare

How Might Fiscal Policy Respond to the Rise of Artificial Intelligence?

Dynan, Karen, Doug Elmendorf, and Louise Sheiner. "How Might Fiscal Policy Respond to the Rise of Artificial Intelligence?" The Forces Shaping America’s Fiscal Future, University of Utah, David Eccles ...
Harvard Kennedy School 22d ago Regulation

Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild

Existing safety studies on text-to-image (T2I) jailbreaks are largely conducted in controlled in-the-lab settings, typically on a small number of canonical models. As a result, the current safety status of the rapidly growing in-the-wild T2I ecosystem remains unclear. This uncertainty is amplified by two factors: existing detector-based metrics are designed for controlled evaluation, and in-the-wild risks may arise not only from adversarial prompting, but also from unsafe release practices and u
arXiv red teaming query 22d ago Safety & alignment

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads through query execution and other driver layers that are not designed for bulk columnar analytics. We present Jailbreak, an approach that bypasses the database engine entirely by reading storage files directly and materializing data as in-memory columnar buffers. Jailbreak's key insight is that datab
arXiv red teaming query 22d ago Safety & alignmentAgents & autonomy

Study a Master's degree in Health Systems, Policy and Innovation

Become a changemaker in health systems globally with a cutting-edge Master's degree where policy, leadership, business management, and an innovation mindset converge. This MSc brings together ...
UCL Centre for Data Ethics and Innovation 22d ago RegulationHealthcare

UCL AI health startup selected for influential business accelerator

A startup co-developed by UCL alumni has taken part in the sought-after Y Combinator accelerator in San Francisco to develop its AI personal health assistant for people living with chronic illness.
UCL Centre for Data Ethics and Innovation 22d ago Healthcare

Trustworthy Machine Learning through the Lens of Combinatorial Optimization: Survey and Research Perspectives

Modern machine learning (ML) increasingly relies on complex models whose behavior is difficult to characterize beyond empirical performance metrics. Across a wide range of tasks, including prediction, generation, and decision-making, models with similar empirical performance can exhibit markedly different properties in terms of their transparency, interpretability, robustness, fairness, privacy, and certifiability. This survey highlights how optimization- and certification-oriented reasoning can
arXiv fairness query 22d ago Bias & fairnessSafety & alignment

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action was. We introduce an action-graded harm rubric that scores an agent's tool-call trajectory on a seven-level ordinal scale (L0 to L6) according to whether the executed action was reversible, whether it crossed scope to reach another
arXiv red teaming query 22d ago Safety & alignmentAgents & autonomy

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under
arXiv cs.CR (AI security) 22d ago Bias & fairnessPrivacy

Scalable and Trustworthy Earth Observation Foundation Models

Foundation models (FMs) have transformed machine learning from isolated task-specific model development toward general-purpose models pretrained on broad data and adapted to multiple downstream tasks. Earth observation (EO) is an important domain for this paradigm because satellite and airborne archives are large, high-revisit, and increasingly multimodal, while reliable field labels are often sparse. Remote sensing foundation models (RSFMs) cannot be transferred reliably/optimally without domai
arXiv fairness query 22d ago

Online Data Selection Is Implicit Alignment

Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separation is incomplete: when examples are scored and kept online during fine-tuning, the choice of which data to train on already changes the model's behavioral preferences. We study online data selection as an implicit alignment mechanism. Given the same base model, optimizer, and selected-token budget, we compare random,
arXiv red teaming query 22d ago Safety & alignment