A day after Google unveiled a new feature that allowed users to create AI-generated satellite images with the click of a button on its Google Earth platform, the company announced it was removing the tool. "We've seen geospatial profession ... (report_number: 7681)
Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment (ROI) from AI hinges on having the right foundation, with inadequate infrastructure and data…
After Australia’s first reported automated hacking accident, experts warn deployers – and possibly developers – of AI agents could be held liable for the actions of their bots Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast The law is clear, says Prof Jeannie Paterson. “If I deploy an AI agent and it causes harm to someone else, I am responsible for that harm. “Even if I didn’t intend for that to happen, it was foreseeable, and
The AI jobs apocalypse never showed up. Still, jobs are changing and economists expect more to come The prediction was stark: artificial intelligence advancements would wipe out jobs en masse. “Half” of all entry-level white collar jobs would vanish, Anthropic’s CEO, Dario Amodei, said in May 2025. A month later, OpenAI’s CEO, Sam Altman , went further, foreseeing the end of “certain job categories”. Companies began citing AI in their layoffs. Workers organized. And students reconsidered their f
Open models may soon be added to an updated AI framework, sources tell WIRED, as the White House continues to grapple with how to regulate a technology it has tried not to regulate.
The $399 Google Pixel Watch 5 isn't about the hardware. Sure, there's a new satin pyrite case finish, a few new strap colors, and a Steph Curry Special Edition. Under the hood, there's a slightly faster Qualcomm processor and an itty-bitty battery bump. There's a $50 price hike from last year, too, because the Pixel […]
SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent "AI teammates" that can do your work for you. The bots share their own cloud-based computer environment, and can sign into apps, tools, and websites you already use to complete multi-step workplace tasks, only coming back when their assigned work […]
From pessimism around dating to AI reshaping culture, cyber-ethnographer Ruby J. Thelot tells WIRED why people are putting too much stock into things that go viral.
At Ai4, three of the world's most respected AI experts — Geoffrey Hinton, Fei-Fei Li, and Andrew Ng — debated regulation, open source access, and how America can compete as China advances in Asia.
Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Altimeter Capital.
Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center of multiple legal cases.
It's shelling out more money than Chevron, Verizon, AT&T and OpenAI this quarter to beat California legislation that would limit mark ups to no more than 10 percent more than the original price tag.
Stone said on a podcast that 'The Pitt' was so similar to her own work as a nurse that she nearly had a "panic attack" reading lines before an audition.
Get ready for more action and more Alan Ritchson in “Reacher” Season 4, which premieres Wednesday, Aug. 12 on Prime Video. Our hunky, ex-military hero (Ritchson) is once again joined by a slew of new cast members as “Reacher” moves on to a fresh adventure based on Lee Child’s books. Per Amazon Prime, the logline […]
Process intelligence company Skan AI said today it raised $63 million in a Series C round to help further develop a platform that records how enterprise work actually gets done and feeds that record to artificial intelligence agents. Founded in 2019, the company offers software that sits on employee desktops, grabs screenshots and then processes […] The post Skan AI raises $63M to give AI agents a map of enterprise work appeared first on SiliconANGLE .
Before they ran American foreign policy, they billed hours together. The post Two Brother Left The Same Biglaw Firm And Took Over U.S. Foreign Policy appeared first on Above the Law .
Lovable Labs Inc. said today it has raised $400 million in Series C funding at a $13.3 billion valuation. The Swedish artificial intelligence coding startup was worth $6.6 billion in December. The company sells a vibe coding service. Users describe what they want in plain language and the platform builds it. Hosting is included. Chief […] The post Vibe coding startup Lovable doubles valuation to $13.3B with $400M raise appeared first on SiliconANGLE .
Turnabout, as they say, is fair play. The post Ken Paxton Won’t Discuss His Marriage With Thousands Of Strangers. He’ll Just Regulate Yours. appeared first on Above the Law .
The Defense Department will not decide whether to produce operational space-based interceptors until after industry demonstrates their feasibility, Gen. Michael Guetlein said.
AI code review specialist CodeRabbit announced its Agentic Change Management control layer on Wednesday. The service is intended to help The post “Issue tracking is dead”; How the pull request became the last chokepoint in the SDLC bottleneck appeared first on The New Stack .
Assistant Secretary of Defense for Science and Technology Joseph Jewell briefed DefenseScoop on the DOD’s new structure to enable innovation. The post Pentagon’s S&T boss touts early benefits of re-org, including ‘virtuous cycles of collaboration’ appeared first on DefenseScoop .
“This action is necessary to continue [electronic health record modernization] deployment activities and associated support services until all VA Medical Centers (VAMCs) and related facilities have fully transitioned to the EHR system,” the department said in a Sam.gov solicitation notice.
Former DHS officials tell TIME that the devices aren’t necessary as an enforcement method against civilians—and that their use would pose numerous risks.
Judges, lawyers, educators, and higher-ed leaders will examine how legal education is regulated, with an eye toward independence. The post Legal Education In America Is Getting A Fresh Look, Without The ABA Accreditor At The Table appeared first on Above the Law .
The economics of making a blockbuster game are starting to look as demanding as the games themselves. Development cycles stretch across years, budgets climb into extraordinary territory, and a single failed release can erase enormous amounts of investment. The industry’s latest workforce data makes the pressure difficult to ignore. 28% of respondents to the 2026 […] This story continues at The Next Web
AI data centers in space sound great, but practically speaking, they may be next to impossible. For tech bros, it The post Why space is actually a terrible place to cool a data center appeared first on The New Stack .
“We have an aging fleet, and we have modernization on the horizon, but it’s not coming fast enough,” Brig. Gen. Beth Behn, commanding general of TACOM, told Breaking Defense.
Give an authoritarian a cookie. The post Former SPLC Expert Arrested On Fraud Charges As Malicious Prosecution Expands appeared first on Above the Law .
“Leading up to [2028], we’re going to do a lot of experimentation and test with this unit to make sure that we’re ready to fight with it when it’s delivered,” said Lt. Gen. John Rafferty.
73 Strings, the AI-powered platform for private markets valuation and portfolio intelligence, today announced a series of senior leadership appointments following a period of rapid client growth and increasing adoption among the private capital community.
The Data Foundation’s new AI assistant surfaces statistical evidence and policy-supporting data underpinning the ultra-dense 2018 law. The post ‘Unlocking’ the Evidence Act: Ex-feds’ AI tool aims to bring policy docs to life appeared first on FedScoop .
A new $244M Pentagon memo skips the Federal Acquisition Regulation's process for sole-source justifications and urges agencies to find ways to work with the company.
The long-standing mantra that certain operational systems must be continuously available and can’t be brought down for a patch is becoming harder to defend as cyber threats accelerate, US officials warn.
The 50,000-square-foot facility at ProvPort is a $6 million investment that aims to leverage local partnership to provide “additional jobs” in the community, according to the firm.
Rep. Michael Baumgartner asked why VA OIG’s electronic health record modernization recommendations haven’t been implemented while the department presses on with deployment. The post VA’s EHRM oversight holes must be addressed, House Republican says appeared first on FedScoop .
The latest kids online safety trial against Meta kicked off Wednesday in a California federal court, where the technology giant is accused of designing its platforms to be addictive for young users. Jury selection began Wednesday morning in the U.S. District Court in Oakland, Calif., for the case, which consolidates lawsuits from attorneys general across the country....
xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex workflows in about 53 steps where Claude Opus 5 needs 103, at a price more than 60 percent lower. The article SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price appeared first on The Decoder .
Remember when Trump lost the birthright citizenship case? Because he doesn't. The post Biglaw Surrenders Were Worse Than We Thought appeared first on Above the Law .
Jaleel is attending his first day of classes this month at a middle school opening its doors for the first time that specializes in air traffic control, commercial drones and engineering. His parents decided not to send him to the majority-Black, F-rated campus closer to their home in Fort Bend ISD. Still, they fear their […]
The National Institute of Standards and Technology is angling to modernize how it reports and responds to cyber vulnerabilities with the help of artificial intelligence.
While infrastructure requirements may be commonly shared in enterprise IT, the path for AI deployment can vary significantly depending on the organization. Healthcare, financial services, manufacturing, telecommunications and public sector organizations bring unique data, governance and operational challenges to the table when it comes to AI implementation. This means that the future of AI is […] The post Vertical AI pushes infrastructure beyond one-size-fits-all appeared first on SiliconANGLE .
ICE agents will soon have body cameras, but footage of serious incidents may only be released if it’s in the agency’s “best interests.” Who does that protect?
The do-over election hastily thrown together by Republicans proved as chaotic as expected. The post Alabama’s Primary Drew 5 Percent Turnout Thanks To Supreme Court Destroying Voting Rights Act appeared first on Above the Law .
Elsevier’s International Journal of Biological Macromolecules has issued more than 100 retractions and counting for manipulated peer review in guest-edited special issues. The publisher has said one more retraction is coming. Since June the publisher has retracted 120 papers. According to the retraction notices, the journal identified “[s]ystematic, coordinated, and widespread manipulation of the peer-review … Continue reading Biology journal pulls 100+ articles and counting from special issues
The Pentagon’s secretive research agency says new advances are needed to counter adversary air defenses, along with innovative manufacturing techniques that can produce hypersonic weapons affordably and at scale.
Israeli cyber firm Dream said the framework adapted mid-operation, corrected its mistakes and expanded as it went along. The post Researchers observe first ‘near-autonomous’ AI attack on government target in Taiwan appeared first on CyberScoop .
Associates are already sitting on six-figure bonuses, and another payout is still ahead. The post Elite Litigation Boutique Is Already On Its Third Round Of Associate Bonuses appeared first on Above the Law .
Vaya por delante algo: soy la antítesis de una creadora de contenido. Apenas subo cosas a mis redes y cuando lo hago, lo que comparto son cosas como una croqueta que se me ha quemado, un vídeo de mis perras donde la mitad de las veces se escapan del plano o una foto de la playa de Benidorm. Soy como mi tía Marisa o tu tío Antonio, vamos. Así que cuando vi por primera vez el resultado de una noche loca entre un móvil y un gimbal pensé: eso no es para mí. Menos de medio año después, resulta
Originally designed to make software predictable for humans, Google Go is now positioning itself as a language tailored for machine The post Code that passes every test can still break the next AI agent that touches it appeared first on The New Stack .
During AI DevCon in London this summer, Lamis Mukta, member of technical staff at Anthropic, hosted a stage presentation session The post Anthropic gave agents the ability to dream. Then developers woke up. appeared first on The New Stack .
An Army official told Breaking Defense the service’s original plan to build a software backbone itself makes less sense now that there’s real industry investment in the capability.
After making AI tools widely accessible for employees, leaders are putting caps on usage and retraining staff to understand small and cheaper models can often do the job.
Blacksmith Software Inc. today announced it has raised $45 million in new funding for its continuous integration service, which combines code development with cloud-based testing instead of on the developer’s computer. Peak XV Partners led the Series B round, with existing investors Y Combinator and GV also participating. The funding brings the company to a valuation of […] The post Blacksmith raises $45M to aid AI code validation as agentic development grows appeared first on SiliconANGLE .
Countless times over the last few years, district leaders have reached out to my organization with a problem: They rolled out a new evidence-based curriculum and trained all their teachers, but many didn’t use it properly — or at all — and the school saw only pockets of success. Talking through what happened, they inevitably […]
James Uthmeier saw a white-grievance layup and took the shot. The post Florida’s AG Wants To Charge WNBA Player With Crime Of Playing Basketball appeared first on Above the Law .
In this prestige-obsessed profession, this opportunity may not come again for a long time. The post If You Really Want To Work For The Department Of Justice, Then Do It Regardless Of Who The President Is appeared first on Above the Law .
El jueves 16 de julio de 2026, en La Mierla, al norte de la provincia de Guadalajara fue " una cosechadora agrícola ". En Burgohondo, al sur de Ávila, fue "una imprudencia grave por el empleo de maquinaria pesada" . En Los Gallardos, al norte de Almería, fue un cable olvidado y las Peñas de Riglos, en la hoya de Huesca, las sospechas apuntan a una campa de unas obras en la A-132 . Faltan datos oficiales y meses de investigaciones, pero algo está claro algo no funciona en la normativa. España es
Tech leaders say AI will drive down costs. But slow corporate adoption and the data center build-out create inflation pressures that complicate the Fed’s job.
Gallup’s State of the Global Workplace report shows employee engagement fell to 20% in 2025, its lowest level since 2020, with disengagement costing the global economy an estimated $10 trillion annually. Research also shows that engagement is closely tied to how employees experience their workplace relationships. Gallup’s workplace research found that highly engaged teams achieve […] This story continues at The Next Web
Desde que ChatGPT empezó a ser lo que es y la IA comenzó a llegar a todo el mundo, ya que antes era algo reservado a la investigación en diferentes ámbitos, una palabra nos ha acompañado: entrenamiento . De ese entrenamiento sabíamos que era algo costoso por los equipos que se necesitaban para llevarlo a cabo, pero también por los recursos naturales y energéticos que consume (así como todo el conocimiento robado y destruido por el camino ), pero desde hace no tanto, parece que nos acercamo
The New York City Council announced Wednesday it is investigating four prediction markets for alleged “predatory marketing practices” targeting young people. Council Speaker Julie Menin sent letters to Polymarket, Kalshi, Coinbase and Gemini Titan on Monday, requesting more information about their advertising in New York City. In a statement shared by her office on Wednesday,...
It is not possible to bomb an entire country into submission from the air using conventional weapons. The post Lobbing Multimillion-Dollar Munitions At Targets Worth Less Than The Missiles Themselves Remains A Bad War Strategy appeared first on Above the Law .
Ha llovido mucho desde 2019, año en el que Asus aún lanzaba móviles que intentaban cosas distintas y cuando vimos el Zenfone 6 . Ese teléfono tenía una cámara trasera que podía convertirse en la frontal. Más allá de ser la solución al problema del notch en las pantallas para conseguir un frontal despejado, lo de Asus era una declaración de intenciones: tenemos músculo tecnológico para innovar también en smartphones. Han pasado siete años y acabamos de conocer el Honor Robot Phone . No nos pilla
Three years ago, a handful of Maine school districts struggling with student mental health needs applied for a federal grant to help them hire school counselors and social workers. The grant applications included a commitment to hiring staff that represented their particular community and understood its needs. The grant — part of the Bipartisan Safer […]
Small group insurers are proposing a median 14% premium increase for 2027, driven by rising healthcare costs. The post Small Businesses Could Face Double-Digit Premium Hikes In 2027 appeared first on Above the Law .
The request by the Pentagon’s No. 2 official that CEOs for defense firms bring their board members in for discussions is unprecedented, but in line with Feinberg’s industrial maneuvering.
Robert Mahari is Anthropic's first "Head of Claude for Legal," responsible for deploying and expanding the Claude AI model across the legal industry. The article Legal startup founder Robert Mahari joins Anthropic to lead Claude's push into law practices appeared first on The Decoder .
Curious Kids is a series for children of all ages. If you have a question you’d like an expert to answer, send it to CuriousKidsUS@theconversation.com . How is a text-to-speech voice made? – Sarah G٫ age 11٫ Seguin٫ Texas When you talk to computerized assistants like Siri or Alexa , they reply in voices that sound very human. But how do computers, smartphones and apps actually talk like a person? They use a technology called text-to-speech . When you speak, your lungs push air up your windpipe a
Tencent stock was down 26% so far in 2026 as the company faces intense competition in China in AI and investors grow jittery about its rising spending.
FE fundinfo today confirmed it is accelerating investment in Nexus for Financial Advisers, a unified, AI-powered workflow platform built on the breadth of FE fundinfo's data, designed to support the entire advice journey by bringing together the disconnected systems most advice firms rely on today.
As agentic AI infrastructure moves from experimentation into production, enterprises are confronting a more complex question than which model to use: how to control the cost, data exposure and infrastructure supporting production AI applications. That shift is pushing organizations to rethink how much they should rely on public cloud AI services alone, especially as agentic […] The post Agentic AI infrastructure shifts enterprise focus from model choice to platform control appeared first on Sili
CodeRabbit Inc., the creator of a popular tool that automatically reviews artificial intelligence-generated code, is becoming more ambitious after closing on its latest $143 million Series C round of funding. Alongside the round, it announced the launch of a new Agentic Change Management layer that’s meant to help companies govern, evaluate and prioritize code changes […] The post CodeRabbit bags $143M to help companies get a grip on the explosion of AI-generated code appeared first on SiliconAN
Marketing intelligence company Ahrefs Pte. Ltd. today launched Letaido, an agent-powered marketing workspace built to take over the recurring research, reporting and monitoring work that fills up a marketing team’s week. In most marketing departments, generative artificial intelligence is still something people use on their own. A writer drafts with it. An analyst pulls numbers. […] The post Ahrefs launches AI agent workspace Letaido for marketers and agencies appeared first on SiliconANGLE .
A venture capital fund backed by Yangtze Memory Technologies Corp (YMTC), China’s top producer of NAND flash memory, has taken a stake in domestic semiconductor manufacturer SOI Micro, adding a high-profile investor to the country’s push for an alternative chipmaking ecosystem. Led by a prominent figure in China’s semiconductor industry, SOI Micro is focused on developing fully depleted silicon-on-insulator (FD-SOI) technology – a low-power chipmaking technology that offers a complementary route
Rimini Street, Inc. (Nasdaq: RMNI), the Software Support and Agentic AI ERP Company™ and the leading third-party support provider for Oracle, SAP and VMware software, today announced the appointment of Keith Costello as chief operating officer (COO) and Alexander Guasch as chief delivery officer (CDO).
* Elite firms warn that AI spending could impact profits, though most continue to believe they'll be fine. Ahem... then where are the associate raises? [ ABA Journal ] * ICE wants $20 million worth of electric shock gloves, in case you were wondering where we were on the "turning law enforcement into Spider-Man villains" spectrum. [ Huffington Post ] * Paul Weiss and Gibson Dunn lead Yankees financing deal to inject $2.6B for the club to not spend during the looming lockout. [ Bloomberg Law News
Orbital data centers, space habitats, mining the Moon – despite taking place in space, the practices many companies are imagining will still affect the Earth.
Nearly 25 years ago, I walked into the U.S. Department of Education at what felt like a remarkable moment in the history of reading instruction. The National Reading Panel had just been released, and for the first time in my professional life, there was broad agreement that educational policy should be grounded in scientific evidence. […]
Santa Clara-based technology services firm Apexon Inc. today expanded AgentRise, its agentic artificial intelligence platform, with three new components. The additions are named AgentRise Polaris, AgentRise Lodestone and AgentRise Harness. Each maps to one of three disciplines Apexon has built its client work around, called Domain & Strategy, Cognitive Architecture and Harness Engineering. Polaris covers […] The post Apexon targets stalled AI pilots with three AgentRise additions appeared first
Autonomous coding agents ignore contribution rules in open source communities, finds a new study from researchers at Peking University. That’s The post Coding agents ignore open source contribution guidelines, researchers find. appeared first on The New Stack .
Nvidias neues offenes Modell soll Agenten-Workloads schneller abwickeln. In der Open Source Bibliothek NeMo Switchyard gibt es Bausteine für Modell-Router.
Honor acaba de celebrar una presentación por todo lo alto en Shenzhen para presentar algo que realmente ya vimos hace bastante tiempo y conocíamos desde hacía aún más : el 'Robophone'. Oficialmente bautizado como Honor Robot Phone , estamos ante una propuesta de lo más curiosa que mezcla en un sólo dispositivo una cámara de acción con gimbal y un móvil con prestaciones de gama Ultra. A continuación, vamos con todas las características de este peculiar nuevo móvil de Honor que no sabemos cu
Chinese smartphone maker Honor has launched a handset with a motorised robotic camera that nods back at users, betting that its interactive hardware will spark growth in a cooling global market for the devices. The new Robot Phone, unveiled on Wednesday, features a three-axis mechanical gimbal with four degrees of motion, allowing for video stabilisation and interaction via nodding and tilting. Activated by a hand gesture, the gimbal camera unfurls from the back of the device – a mechanism that.
India's banking sector must be proactive in its adoption of AI and ensure that it determines how the technology reshapes the industry rather than let the technology dictate the inevitable changes.
OpenAI a annoncé, le 11 août 2026, lancer une version preview de son application ChatGPT pour Linux. Elle réunit au sein d’une même interface ChatGPT, l’agent professionnel Work et l’assistant de programmation Codex.
Technological change—whether it was the internet, the smartphone or even the original printing press—can make people uneasy. Each wave has changed how you access and share information. But AI is different, and bigger: a technology that can decide who gets hired, who gets treated, even who gets targeted. This isn’t just a technology that allows you to access information; it’s a technology designed to autonomously make decisions. Too often, questions about regulating AI are either framed by an acc
Unless your shower turns cold, you probably don’t give much thought to your water heater. But it’s a major energy user in the home, and unlike solar panels or heat pumps, it’s gotten relatively little attention as a piece of smart home tech. A new startup is trying to change that. The company, called Reservoir , designed a heat pump water heater that predicts when you’ll need hot water and heats up in advance, when electricity is cheapest. The upfront cost is comparable to other water heaters, b
BMG is the first significant rightsholder to ink a licensing deal with Suno since Warner Music Group announced an alliance with the AI company nine months ago. Source
The E.U.’s new electronic entry system has been in full effect for only a few months, but some airports and airlines are already calling for it to be overhauled.
Zepto and Blinkit face regulatory action after food safety officials found unhygienic conditions, improper storage, labelling violations and pest-control failures at facilities in Bengaluru and Mumbai. The post Zepto in Bengaluru, Blinkit in Mumbai: Quick commerce platforms face regulatory heat over food safety violations appeared first on MEDIANAMA .
The MCP 2026-07-28 specification removes the initialize handshake and session header, and adds required method and tool-name headers so gateways can route agent traffic without parsing JSON. Reaction split between developers calling it a rediscovery of REST and those arguing the standard itself was always the point. By Steef-Jan Wiggers
Reena SenGupta speaks with chairman Mike Schmidtberger and new partner hire, John Budetti, an ex Kirkland and Paul Hastings partner. In June, Norm Ai, an AI company focused on regulatory […] The post Norm Law: The new disrupter? appeared first on Legal IT Insider .
The global identity verification and compliance firm has introduced a local bank account verification solution to strengthen security and governance in financial institutions.
Tencent Holdings nearly tripled its capital spending in the second quarter to supercharge its AI push, as accelerating investments in computing power and models helped drive better-than-expected revenue. Revenue for the quarter ending in June reached 204.8 billion yuan (US$30.4 billion), up 11 per cent year on year from 184.5 billion yuan, the Chinese tech giant reported on Wednesday. The figure beat average analyst estimates of 202.8 billion yuan compiled by Bloomberg. Adjusted net income...
A HCLTech investigation found that the employee data leak claim by a hacker group "may be limited and dated to a few years back" and there's no evidence of breach in its own or client's systems. The post HCLTech finds no evidence of system breach, says the leaked data is years old appeared first on MEDIANAMA .
MCP server for Nutanix Cloud Platform brings secure, natural-language, agentic AI automation to hybrid multicloud environments without sacrificing control.
The second space race has birthed a handful of public space companies that finally allow investors to own their own piece of the new orbital economy. Yes, there’s newly public SpaceX, but there are also plenty of others, like AST SpaceMobile, Rocket Lab, Firefly Aerospace, and Virgin Galactic. Together, these companies form an emerging category: the space stock. Today, this new class might include as many as two dozen companies. An incomplete list includes Redwire, Planet Labs, Globalstar, Viasa
Satirical candidates are a long-standing tradition in British politics, but few have made as much of an impression as Count Binface. TIME speaks to the novelty candidate ahead of the Clacton by-election.
The post “They’re Putting Kids’ Lives at Risk”: How Abuse in a Tennessee Businessman’s Juvenile Prisons Remained Under Wraps appeared first on ProPublica .
Editor’s note: This is the eighth article in a limited series celebrating American defense technologies born from wartime and their effects on broader national security, politics, and society. This series will run for several weeks to commemorate America’s 250th anniversary, and winners will be selected by a reader vote undertaken through our newsletter later this summer. Prior installments can be found at the Arsenal of Innovation page.For more than 70 years, nuclear power has propelled America
Anthropic va appliquer des watermarks sur les contenus générés par Claude, ce qui permettra de les repérer plus facilement. Cet engagement est aussi une obligation réglementaire en vertu de l’AI Act européen. Depuis le 2 août dans l’UE, les fournisseurs d’IA générative doivent intégrer un marquage technique de détection dans les images, les sons, les […]
When war changes, schools change. New technology, the politics of mass mobilization and campaign design, and novel tactics that defined Napoleonic warfare led to Gerhard von Scharnhorst’s reforms and the idea of lifelong education as a professional obligation. The pattern repeated itself in the interwar period. In 1929, James Carson Breckinridge, a Marine officer, published “Some Thoughts on Service Schools” in the Marine Corps Gazette, calling for a new educational paradigm consistent with John
Chinese AI start-up ModelBest has kicked off its pre-initial public offering (IPO) tutoring process for a listing in mainland China, capitalising on growing demand for compact artificial intelligence models that run locally on devices like smartphones and laptops, and in cars. The four-year-old company also appears to be banking on its small-model strategy to overcome China’s acute shortage of advanced American computing chips by training lightweight systems optimised for domestic...
The sandbox marks Nigeria’s latest move to regulate the virtual asset industry. The CBN will now oversee virtual assets used for payments, including stablecoins, payment, settlement, custody, wallet management, and other transaction-based infrastructure services. The Nigerian Securities and Exchange Commission (SEC) will oversee digital assets that behave like securities.
A handful of elite Bay Area schools have enlisted parent-donors from Sequoia, Lightspeed, and Battery to steer small pools of donated capital into early-stage startups. With massive IPOs back en vogue, schools are lining up for what could come next.
The Groups Pushing Back on McMahon’s ‘Call to Action’ Katherine Knott Wed, 08/12/2026 - 03:00 AM As university leaders announce plans to consult faculty on how to respond to the education secretary’s questions, two groups call for unity and resistance. Byline(s) Katherine Knott
Linda McMahon Legal Battle Grinds On Josh Moody Wed, 08/12/2026 - 03:00 AM The education secretary was sued in 2024 for allegedly ignoring sexual abuse when she was a WWE executive. Almost two years in, the largely overlooked lawsuit is playing out slowly. Byline(s) Josh Moody
A Pre-Orientation Designed for Student Parents Joshua.Bay Wed, 08/12/2026 - 03:00 AM Wichita State’s new pre-orientation program builds belonging and connects student parents with resources before the semester begins. Byline(s) Joshua Bay
The Legislation That Eats Away at Higher Education’s Financial Footing kjohnsonbowles… Wed, 08/12/2026 - 03:00 AM Legislation is increasing financial burdens on students and institutions. Here are some things to think about when trying to fix the business model. Byline(s) Kathy Johnson Bowles
The Odd Couple: We’re Not Back in School rachel.toor Wed, 08/12/2026 - 03:00 AM Another academic year starts and neither Gordon nor Rachel will be on campus. Byline(s) Rachel Toor E. Gordon Gee
Academic Forests Are Higher Ed’s Hidden Jewels Elizabeth Redden Wed, 08/12/2026 - 03:00 AM Forests are an asset for employee and student well-being and offer invaluable opportunities for teaching, research and recreation. Universities should protect them. Byline(s) Robert F. Baldwin Elizabeth Dennis Baldwin Olin Thompson Mefford Danny Weathers Robert H. Jones
The George Washington University recently sold its Virginia campus to Amazon, and The University of Michigan proposed a new $1.2 billion data center development.
In today's edition: Kenya’s cyber cafes are becoming data collectors || Uber discontinues UberX in South Africa || BRICS wants faster cross-border payments || South Africa’s central bank takes over the payment rail
Honor has scheduled the China launch of its Robot Phone for Aug. 12. The device combines a smartphone with a motorized gimbal camera and AI-assisted subject tracking. Honor has also highlighted an imaging collaboration with ARRI. Pricing, sales channels and broader availability were not disclosed in the announcement. [IT Home, in Chinese]
DeepSeek is recruiting for an IDC data-center team in Beijing, Hangzhou and Ulanqab, with roles covering data-center planning, construction, testing and operations. The listings indicate that the company is expanding its infrastructure efforts beyond model research and software development. The job descriptions seek candidates in electrical engineering, HVAC, automation, energy, communications, computer science, environmental engineering […]
For most regular tech-using mortals, thinking about AI means thinking about ChatGPT or Gemini . For those in the know, though, there’s also Anthropic’s Claude . And whether you’re among the Claude-embracing crowd or someone who’s never ventured into its embrace, this off-the-beaten-path productivity powerhouse is packed with potential that you’ve probably never noticed. Claude came online five years ago, when a group of former OpenAI employees decided to branch off and build their own generativ
The competition around large AI models is moving beyond models and chips and increasingly into data center infrastructure. Over the past few years, tech giants around the world have continued to ramp up AI computing capacity, with GPU purchases and server expansions becoming almost standard practice. But as demand for computing power continues to surge, […]
Last month, all 193 United Nations member states gathered in Geneva for the inaugural Global Dialogue on AI Governance. For two days they debated how the world should govern AI, yet paid little ...
Jenna Ortega recently told Esquire that, as a child actor, she would go entire days without asking for food or water to make sure she was not “in the way of anybody.” “I really had it going for me as a child ‘cause I can’t think of one mistake I made,” Ortega said. “Is that […]
South Korean chipmaker SK Hynix’s potential deal to offload a packaging plant in southwest China is part of a strategic pivot towards higher-margin artificial intelligence (AI) memory products, according to analysts. But a sale might not be straightforward, as the memory chip giant would need to navigate valuation hurdles and a volatile chip cycle, they cautioned. SK Hynix said on Monday it was “looking into various solutions to enhance the competitiveness of its packaging business”, after...
His résumé includes ‘A Time for Burning,’ now in the National Film Registry, ‘Super Chief: The Life and Legacy of Earl Warren’ and ‘The Rise and Fall of Jim Crow.’
Align Technology, Inc. ("Align") (Nasdaq: ALGN), a leading global medical device company that designs, manufactures, and sells the Invisalign® System of clear aligners, today announced that the Jinan Intermediate People's Court in China issued a judgment in favor of Align in a patent infringement action against Angelalign Technology's operating subsidiaries in China ("Angel") (Hong Kong Stock Exchange: 6699.HK).
American Express's corporate venture arm has invested in Fazeshift, an AI-native platform deploying autonomous agents to execute end-to-end accounts receivable workflows.
In the remote mountains of China’s Sichuan province, viral online soccer star Jifu Lama is spreading his love of the game to help ‘left behind’ Yi ethnic children find a sense of purpose.
On August 11, New Jersey became the newest state to enact a design code law aimed at minor online safety after Governor Sherrill signed A4015, the “New Jersey Age-Appropriate Design Code” (NJAADC). The new law is among the broadest in the country, most closely resembling a blend of the design codes enacted in South Carolina […]
AI and genomics are creating new possibilities for earlier diagnosis and more personalized care. Experts from WHO, Harvard University and industry discuss what it will take to make these advances accessible, reliable and equitable. The post How AI and Genomics Are Changing Disease Detection and Precision Medicine appeared first on AI for Good .
Last year, we introduced the concept of the AI-native cloud, as we observed that the cloud industry was moving beyond commodity infrastructure toward platforms purpose-built for generative AI and agentic AI. We also identified two emerging paths in this transformation: AI infrastructure cloud platforms (neoclouds) and AI-centric neoPaaS. One year later, those two paths have evolved from emerging concepts into distinct market categories, […]
Nigeria has looked south and seen a $40 million payday for the press. The trouble is that it misread both the price tag and the fine print—and its attempt to collect may leave Nigerian publishers with fewer readers and no comparable payday. On July 6, Nigeria’s Federal Competition and Consumer Protection Commission (FCCPC) announced investigations ... Copy, Paste, Compensate: Nigeria’s Misguided Bid to Make Big Tech Pay for News The post Copy, Paste, Compensate: Nigeria’s Misguided Bid to Make B
AI is creating real value, but the biggest opportunities will go to organizations that build the right foundations first. Explore five no-regrets investments that can help technology leaders govern, scale, and realize value from AI regardless of how the technology evolves.
The math, which combines chaos, quantum theory, and infinitely complex fractal structures, has been called a “foundational result.” The post Graduate Student Proves a Quantum Uncertainty Principle for Fractals first appeared on Quanta Magazine
NVIDIA founder and CEO Jensen Huang is ranked No. 1 on Glassdoor’s Best CEOs list for 2026. In the just-released ranking, recognition is earned directly from the people who know their leadership the best — employees. Huang topped the list, with 99% of employees approving of the job he does. “As AI and shifting expectations […]
AI agents aren’t coming — they’re already here. And they’re not waiting for your security architecture to catch up. Learn how Forrester's new AEGIS framework can help CISOs secure, govern, and manage AI agents and agentic infrastructure.
An immigration expert considers questions that might have been asked at DHS General Counsel Percival's UCLA Law Federalist Society event. The post Free Speech, Immigration Law, and Truth Telling at UCLA Law’s Recent Federalist Society Event appeared first on Just Security .
Anthropic v. Department of War reveals why courts must distinguish genuine national security judgments from pretextual ones and how to do it. The post Deference Should Follow Expertise, Not Pretext appeared first on Just Security .
Sign up to receive the Early Edition in your inbox here. A curated weekday guide to major news and developments over the last 24 hours. Here’s today’s news: IRAN WAR The U.S. military said yesterday that a U.S. Navy MH-60 helicopter fired two Hellfire missiles to disable the steering gear of a Panama-flagged cargo ship, adding […] The post Early Edition: August 12, 2026 appeared first on Just Security .
This seems to work : Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down. Examples are a prompt that
We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capital to support the buildout of AI infrastructure over time. This is a major milestone for NVIDIA and the AI industry. We have moved from an era in which companies […]
To build capability in software development and agile ways of working across teams within the AI delivery environment to improve operations. Classification: Project management consultancy services.
Purchase of an AI patient safety and experience platform Additional information: This is a call off contract from framework Mid and South Essex NHS Foundation Trust Framework Agreement for the Consultancy Solutions and Advisory Services (reference number: 2023/S 000-024003) Classification: IT services: consulting, software development, Internet and support.
MEMORANDUM FOR THE VICE PRESIDENT THE SECRETARY OF STATE THE SECRETARY OF THE TREASURY THE SECRETARY OF WAR THE ATTORNEY GENERAL THE SECRETARY OF COMMERCE THE SECRETARY OF ENERGY THE SECRETARY OF HOMELAND SECURITY THE ASSISTANT TO THE PRESIDENT AND CHIEF OF STAFF THE DIRECTOR OF NATIONAL INTELLIGENCE THE ASSISTANT TO THE PRESIDENT FOR SCIENCE […] The post Expanding Capabilities to Combat Transnational Cyber-Enabled Crime appeared first on The White House .
BY THE PRESIDENT OF THE UNITED STATES OF AMERICA A PROCLAMATION This National Substance Use Primary Prevention Month, we renew our pledge to safeguard our children, friends, families, and communities from the devastating reach of illicit drugs by stopping addiction before it ever begins. Substance abuse remains among the gravest threats confronting our country, poisoning […] The post National Substance Use Primary Prevention Month, 2026 appeared first on The White House .
The Commodity Futures Trading Commission (CFTC or Commission) is publishing this notice to announce the renewal of the Innovation Advisory Committee (IAC). The Commission has determined that the renewal of the IAC is necessary and in the public's interest.
The National Institute of Standards and Technology (NIST) established and operates the National Vulnerability Database (NVD), which provides the U.S. government repository of standards-based vulnerability management data. NIST seeks stakeholder input on opportunities, challenges, and priorities for modernizing the NVD in an evolving cybersecurity landscape increasingly shaped by artificial intelligence (AI) and machine-consumable security data. NIST's goal is to improve the NVD's scalability, au
As artificial intelligence (AI) rapidly transforms everything from how we learn, to the job market, children are adopting the technology more than three times faster than adults – exposing them to ...
The Network’s specific objectives and scope of work are defined through a participatory and inclusive process, allowing members to actively participate in shaping its direction and activities.
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination. To address these challenges, we propose DreamFly, a diffusion-b
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unnecessarily costly trajectory. Prior work studies selection manipulation, malicious skill instructions, and tool-chain resource amplification largely separately, leaving their end-to-end composition unc
Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics differ in their expectations of what should be made transparent and how, as well as in their interest in and ability to parse the information currently published in the registers. Moreover, it remains unclear how these instruments can represent the sociotechnical systems in which these algorithms are embedded, and how system-level transparency can facil
Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated content, table structure degradation, and lack governed workflows extending beyond extraction to validation and artifact generation. This leaves enterprises to perform this manually, consuming 2-3 days per document. To address this, we introduce GUIDE, a governed multi-agent framework built on a shared versioned rule store
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent
GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to reflect on issues that extend beyond these activities. Our findings show that GenAI expands the capab
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety w
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence. This paradigm enables a unified generation interface for item IDs, histories, and item text, but it also creates a structured optimization bottleneck during reward-based post-training: when an early semantic token enters the wrong branch of the item-token space, finite rollout groups rarely reach the ground-truth ite
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. Learning effective reflection, however, is challenging because reflection is performed locally within the current branch, whereas its utilit
Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabilisation in which inherited criteria become newl
Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behavior and parameter updates. We argue that a useful role should instead be an executable control variable: it should summarize behavior predictive of future utility, guide subsequent interaction, and identify the trainable capacity responsible for that behavior. We introduce ExRole, a trajectory-to-role framework that le
Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, a
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-
Deployed foundation models are often not static systems, with providers able to modify system behavior through fine-tuning, classifier updates, system prompt revisions, retrieval changes, and routing changes. These updates can be made silently -- that is, without public disclosure, a version increment, or re-evaluation. Such silent updates challenge a core assumption behind current AI governance frameworks that an externally verifiable chain of custody links the model referred to in evaluation r
Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they largely ignore fine-grained instruction-level correspondences. This limitation misses valuable supervision signals available from compiler debug information, which can support the learning of more accurate and interpretable binary code representations
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the schol
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation f
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $λ$ to every token. This can drive the student to fit extreme peaks in the implicit reward, causing reward hacking and unstable training, and the optimal $λ$ varies across domains, requiring costly sweeps. We prop
Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understanding and task-planning capabilities, vision-language models (VLMs) are promising candidates for JMSU. However, directly applying existing VLMs to JMSU is non-trivial due to scarce cross-scene supervision and attention dispersion ca
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, we propose a multimodal framework that leverages pre-trained vision-language models
A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided. We provide such a foundation b
Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an explicit control plane, unmediated updates by models, tools, and background workers risk stale overwrites, un-audited exposures, and self-authorizing privilege escalation. We argue that agent state governance is an infrastructural activation problem, defining continuity as an unbroken, authorized lineage of accepted branch heads. We present the Conti
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when cl
Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introducing nontrivial computational overhead and failing to leverage the rich spectral dynamics inherent in time-series data. To enable prompt-free, frequency-aware adaptation of frozen LLMs, we propose FM-LLM (Frequency-Enhanced Mixture-of-Experts for adapting LLMs to Time Series Forecasting), an autoreg
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users' natural conversational feedback. Moreover, the available online feedback is hete
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the ef
Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability, or generate only local road structures rather than complete HD maps. We present RoadWeaver, a coarse-to-fine framework for from-scratch generation of diverse, large-scale HD maps. RoadWeaver first synthesizes a global road layout, expands it into a
Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization. We present Semantic Prism, a conditional semantic-image generation-and-refinement framework with deterministic inference. A diffusion-distilled one-step generator renders a semantic RGB image; per-pixel dis
Background: Online social support, the interaction among individuals in which one helps another during difficult situations through online platforms such as online forums or social media, has proliferated as a vital tool for personal mental health care. Despite the growing usage and importance of online social support, prior studies have mainly focused on either understanding the characteristics of support seekers or merely identifying types of support, which leaves room for improvement in 2 key
Background: Internet addiction (IA) has been consistently associated with adverse mental health outcomes, but less is known about whether adolescents with IA seek mental health support, and whether associations between help-seeking and mental health problems differ across pathways. Objective: This study aimed to describe mental health help-seeking patterns across internet use and IA status, and examine the independent and interactive associations of IA and help-seeking with mental health problem
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, w
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transfor
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We stud
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design
Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to extract, summarize, or update memories. This design makes memory construction increasingly costly as conversations grow. Coarse summarization can reduce construction cost but risks discarding fine-grained contextual evidence, whereas larger retrieval contexts or multi-hop LLM reasoning shift the ov
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or w
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific do
Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few-step students using bidirectional teachers that score complete clips. The score fo
Background: Online reviews of health care services represent a growing source of unsolicited, citizen-generated data that can complement traditional instruments for monitoring public perception of health systems. However, longitudinal analyses examining how citizens’ perceptions evolved before, during, and after the COVID-19 pandemic remain scarce, and existing studies have rarely differentiated between levels of care. Objective: This study aimed to examine the longitudinal evolution of public p
Background: The current global status of breastfeeding is marked by both progress and challenges. Digital health interventions (DHIs) have emerged as a promising strategy for improving breastfeeding practices, yet evidence regarding their impact on breastfeeding outcomes remains limited. Objective: This study aimed to evaluate the impact of DHIs on breastfeeding practices and outcomes. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we s
Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly. We benchmark six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference. First, weak overlap is governe
Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable. As a case study of this problem, we consider post-cardiac-arrest neurological prognostication using a cohort of 2,497 patients, including 1,429 patients whose outcomes were rendered indeterminate by treatment decisions. These patients with indeterminate outcomes were revi
Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction.
An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. We show this is structural: any risk certificate is defined over a decision contract, the inputs a system acts on plus the semantic relation under which an output counts correct, and weakening either hides base-classifier error. We develop a decision-contract theory: an error-conservation law showing error is only reassigned among harmful aut
Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy. We address this gap for AddGraph, the foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has n
Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder, a frozen ResNet-50. We attach 17 linear classification probes to the intermediate layers of the Res
We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d<D$) settings, and group these models under the term flow autoencoders. We present a theoretical investigation into their training dynamics and prove that the proposed loss used by existing approaches is suboptimal; specifically, both encoder and decoder surrogates must be optimized in alignment with reconstruction loss. Guided by these insights
Open-source investigators and OSINV journalists have spent the last decade turning publicly available data into evidence of war crimes, environmental harms, and human rights abuses. More recently, they have trained the same lens on a different target: the algorithms, bot networks, and generative AI systems that industrialise deception and information disorder. The post Turning OSINV inside the machine: Open-source investigations and information disorder first appeared on HKS Misinformation Revie
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the teacher-reference log-likelihood ratio to move beyond direct imitation, but apply a single global coefficient $λ$ to every token. This can drive the student to fit extreme peaks in the implicit reward, causing reward hacking and unstable training, and the optimal $λ$ varies across domains, requiring costly sweeps. We prop
Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subjects, recording environments, and datasets. However, dominant pretraining paradigms face key challenges: masked autoencoding tends to prioritize low-level signal reconstruction over task-relevant semantics, while autoregressive modeling creates a mismatch between continuous neural dynamics and discrete token spaces. To address these challenges, ne
Spatially continuous quantification of forest above-ground biomass (AGB) is what makes carbon accounting credible and mitigation strategies actionable. While field inventories provide high localized accuracy, they are spatially sparse; conversely, spaceborne LiDAR from the Global Ecosystem Dynamics Investigation (GEDI) offers broad biomass samples but lacks spatial continuity and systematic underestimation of high-biomass forests. This paper presents an operational framework centered on a single
arXiv:2608.10194v1 Announce Type: new Abstract: Humans are increasingly expected to interact with AI systems that observe and make inferences about them - but do these systems actually work? A standard approach to answering this question is AI auditing. Conducting an AI audit requires identifying how a system behaves (i.e., determining what types of inputs to audit it with and then observing and documenting actual system behavior) and contrasting that with how a system should behave (i.e., deter
arXiv:2608.10329v1 Announce Type: new Abstract: Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or aggregate-corpus level, too coarse to capture the discrete regulatory obligations where commenters seek change. We introduce obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-c
arXiv:2608.10601v1 Announce Type: new Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer constitutively: it is the central feature separating the regulated category from conventional software. The GDPR never defines inference, yet governs it protectively: the consequences follow from the processing of personal data and from what the inference says about, or does to, a person, whether
arXiv:2608.10773v1 Announce Type: new Abstract: The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on journalism and its democratic function. We explore these risks through a case study of GenAI in Norwegian Newsrooms during the 2025 parliamentary election campaign. Based on interviews with managers and journalists over a ten-month period, we analyse how ambitious visions fared in the face of technological and pr
arXiv:2608.10778v1 Announce Type: new Abstract: This study examines the impact of technology within media education, media literacy, and educommunication, and explores how these fields are perceived and understood by students and academic experts, focusing on the development of critical competencies and critical media literacy. Based on semi-structured in-depth interviews with leading experts in the field of critical media literacy, and a survey conducted with 141 university students in Communic
arXiv:2608.11006v1 Announce Type: new Abstract: Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and regional AI strategies to identify their convergences and divergences can uncover their common practices, understand regional variations, and provide policy designers a comprehensive set of policy design elements for their ongoing AI strategy developments. Yet, existing work has not examined their underlying po
arXiv:2608.09937v1 Announce Type: cross Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries an
arXiv:2608.09940v1 Announce Type: cross Abstract: The Metaverse is a convergent space integrating virtual reality (VR) and augmented reality (AR) technologies, with market projections rising from \$65.5 billion in 2022 to \$1.3 trillion by 2030. Despite rapid adoption in education, the specific contributions of visual elements, environmental design, and communication features to user experience (UX) remain underexplored, limiting evidence-based design and resource allocation. This study examined
arXiv:2608.09945v1 Announce Type: cross Abstract: When representing digital circuits, 2 dimensional hand drawings free us from the linear structure of hardware description languages, enabling intuitive reasoning and making structure explicit. However these drawings are imprecise and inert: they do not enforce that the circuits are well defined and cannot be tested. We want both intuitive visual representations and well defined testable ones but students can struggle to link one to the other. To
arXiv:2608.09998v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental implications, primarily due to their high energy demands and associated carbon emissions. This concern is particularly relevant in light of the increasing deployment of large-scale models, especially Deep Learning (DL) architectur
arXiv:2608.10046v1 Announce Type: cross Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and hiring manager interviews capture what employers ask for. How candidates themselves articulate these competencies has not been studied, and existing CV-mining work is both keyword-based, so it cannot see skills conveyed thro
arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequ
arXiv:2608.10268v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of in
arXiv:2608.10276v1 Announce Type: cross Abstract: Student-generated metaphors about mathematics can reveal students' attitudes, beliefs, identities, and experiences, but human expert coding of these thematically and semantically complex open-ended responses is time-intensive and difficult to scale. This study examines whether LoRA-based supervised fine-tuning of large language models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We used a
arXiv:2608.10412v1 Announce Type: cross Abstract: Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM inte
arXiv:2608.10492v1 Announce Type: cross Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework th
arXiv:2608.10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reli
arXiv:2608.10818v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previousl
arXiv:2608.10858v1 Announce Type: cross Abstract: Language models now draft, classify and criticise inside research production, yet the artifacts they help produce carry little accountable history. Rather than detecting machine involvement afterwards, we specify an auditability discipline built at production time: git sealing with an anchor lineage, hash-bound provenance, red-line gates that refuse non-compliant artifacts and log every refusal, cross-model role separation, and programmatic assem
arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in e
arXiv:2605.15850v3 Announce Type: replace Abstract: In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrictedly. While most prior research has focused on how to pedagogically scaffold its usage, the question of when to allow off-the-shelf GenAI remains understudied and lacks pedagogically grounded empirical inve
arXiv:2606.07270v2 Announce Type: replace Abstract: Contribution: This paper presents a novel two-phase algorithmic approach that decouples preference satisfaction from fairness optimization in student team formation, achieving both objectives without compromise. The method applies simulated annealing -- a core materials science technique -- to an educational challenge, demonstrating pedagogical integration of administrative processes. Background: Forming effective teams in large engineering coh
arXiv:2507.19538v2 Announce Type: replace-cross Abstract: Long school bus rides adversely affect student performance and well-being. Rural school bus rides are particularly long, incentivizing parents to drive their children to school rather than to opt for the school bus. This in turn exacerbates the traffic congestion around schools, further compounding the problem of long bus rides, creating a vicious cycle. It also results in underutilized school buses and higher bus operating costs per ride
arXiv:2606.17441v2 Announce Type: replace-cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies. However, existing approaches often lack realism and controllability, often oversharing information unprompted, and failing to capture the wide variability of patient behavior. Here, we introduce PatientsWithPersonality (PWP), a patient simulation framework that generates realis
arXiv:2608.09548v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses ed
Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning. Teams therefore train the bandit on a fast proxy reward, and separately must judge whether a contextual bandit is worth its complexity over sending one best message. Settling both decisions with the usual offline checks - a batch off-policy estimate, a margina
Imitation learning enables robots to acquire complex manipulation skills from human demonstrations, but current methods rely solely on low-level sensorimotor data while ignoring the rich semantic knowledge humans naturally possess about tasks. We present ConceptACT, an extension of Action Chunking with Transformers that leverages episode-level semantic concept annotations during training to improve learning efficiency. Unlike language-conditioned approaches that require semantic input at deploym
An insole-type active assist device has been developed as a robotic system to dynamically correct ankle alignment at heel contact in patients with medial knee osteoarthritis. Although our previous feasibility study demonstrated that the device could be safely used during an on-the-spot stepping task, its effects on loading behavior remain unclear. This study aimed to investigate whether dynamic ankle alignment correction using the device alters horizontal ground reaction force variability and ce
Generative AI exposes the historical fragility of Romantic myths of the sovereign, self-transparent writer without inaugurating a crisis of authorship itself. Tracing a genealogy from symbolic AI to large language models, we show how creativity has always depended on distributed infrastructures, archives, and labour that the figure of the solitary author conceals. Drawing on Barthes, Foucault, Butler, Haraway, Hayles, and recent legal and bibliometric debates, we argue that “AI authorship” is a
As generative AI becomes embedded in platform-based cultural production, new forms of creative labour are emerging that challenge established understandings of what it means to be a “creator.” This article examines AI slop creators in China, a group that uses generative tools to benchmark viral content, test formats, operate matrix account systems, and formalise production, transforming virality from an unpredictable outcome into an object of systematic management. Drawing on ethnographic resear
Large language models (LLMs) increasingly generate outputs that resemble introspection, including self-reference, epistemic modulation, and claims about their internal states. This study investigates whether such behaviors reflect stable underlying patterns or merely surface-level generative artifacts. We evaluated five open-weight, stateless LLMs using a structured battery of 21 introspective prompts. The main corpus comprised 1050 completions collected under a baseline decoding condition ( tem
As artificial intelligence (AI) becomes increasingly embedded in social, economic, and public-service contexts, understanding how citizens perceive its benefits and risks is important for responsible technology governance. This study examines perceptions of AI among 280 African participants in a 4-week AI bootcamp held in 2025. Using a mixed-methods design, the study analyzed closed-ended survey responses on perceived AI benefits, collaboration comfort, community impact areas, risk perceptions,
Architecture’s relationship with materiality faces an unprecedented crisis as artificial intelligence systems transform how materials are understood, selected, and specified. This paper examines the dissolution of embodied material knowledge through three interconnected developments: the historical abstraction of matter via industrial catalogues, the emergence of circular economy databases as attempted alternatives, and the recent shift toward AI-generated architecture that operates primarily th
In this article, we argue that the benefits and harms of generative AI in the content and conditions of creative labour are not evenly distributed across the cultural industries and between cultural workers. In light of this, we propose a conceptual typology that allows for a more granular and nuanced analysis of the asymmetrical impact of GenAI in creative labour structured around three dimensions: contractual relations, forms of automation, and workflows. Regarding contractual relations, we di
As generative artificial intelligence (GenAI) becomes increasingly embedded in organizational innovation and digital transformation strategies, understanding its implications for workforce adaptation has emerged as a critical managerial and strategic issue. This study examines how employees’ perceptions of GenAI, specifically perceived competence and perceived warmth, together with information overload, relate to AI anxiety and subsequent avoidance behavior in the workplace. Drawing on technostr
This study investigates the phenomenon of digital decoupling by proposing an AI Dual-Track model designed to quantify how generative artificial intelligence (AI) simultaneously displaces and augments labor across 846 U.S. occupations. By leveraging an ensemble-based AI classification protocol, the research constructs a validated set of 63 O*NET competencies to map occupational tasks into substitution and facilitation tracks. Central to this analysis is the exploration of how educational attainme
Deploying robot fleets in complex, real-world environments requires human operators to supervise multiple robots simultaneously. Managing operator attention is a fundamental challenge of designing multi-robot supervision interfaces, encompassing both feed layout and feed content (i.e., robot behavior design). Thus far, designers lack empirical guidance on the latter-how to change a robot's behavior to capture, sustain, or relinquish operator attention during multi-robot supervision. In our visio
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCIT
Designers require different design spaces across creative stages: broad during exploration, and targeted during refinement. Yet existing agent-driven tools assume a fixed or continuously expanding space, leaving designers to manage and navigate it themselves. Informed by a formative study with five designers, we propose an axis-centered workflow that adaptively broadens and narrows the design space to support structured exploration and refinement. We implemented this workflow in Surprise2Refine,
Norway is among the most digitalized countries in the world, where access to essential services increasingly depends on digital systems. Although universal design of ICT is legally required across public and private sectors, ensuring cognitive accessibility for older adults involves more than technical compliance. We analyzed responses from 294 participants aged 55 to 90 to examine the barriers they encounter when using digital services. Our findings identify four tensions, namely navigational,
Dance imitation integrates motor planning, sensorimotor integration, and social cognition, offering a sensitive framework to characterize motor behavior in autism. In this work, we explore a computational analysis framework to identify potential biomarkers that allow the design and development of improved medical and human-machine systems. We analyzed 3D motion capture data from autistic and neurotypical adults performing dance imitation under solo and socially-framed duo conditions. Methodologi
At the studied research institute, one professorship oversees approximately 20 theses per semester, while day-to-day supervision is distributed among doctoral and postdoctoral researchers. To manage this supervision demand, the institute uses an exposé-first workflow in which students prepare a research proposal before entering the main thesis-writing phase. This paper asks how students, supervisors, and administrators experience the exposé-first workflow as a structured process for early thesis
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study strong-to-weak scaffolding: whether a stronger builder model can construct inference-time harnesses that help a weaker target model solve tasks more reliably without any parameter updates. Using four r
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dat
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically relies on expert interpretation of technical documentation, limiting scalability for complex systems. This study presents a framework for automated construction of DML models from system descriptions and their representation as Knowledge Graphs (KG-DML), using Retrieval-Augmented Generation and Large
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanatio
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperparameters to yield drastically different outputs often necessitating inefficient, brute-force trial-and-error processes. To address these limitations, we introduce the ``Agentic Self-Improvement" frame
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source rea
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theo
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this work on at production scale, and we set out to measure how far such delegation can reach. In this work, three prompt-specialized agent roles operate under a version-controlled specification that the agents themselves authored and revised, while humans
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark
We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterpri
With advancements in generative AI technology, an increasing number of researchers have begun exploring AI-native games in which gameplay rules are directly driven by generative AI. This paper presents "Pharos Night: Crown Pursuit," an AI-native deck-building and tactical arena game based on a multi-agent system. The game uses large language models to generate materials and cards, support NPC decision-making, and mediate natural-language interactions. During play, players collect materials, desc
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. Motivated by these challenges, we propose a hybrid planning architecture that combines the advantages of machine learning with the verifiability and the determini
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral analysis and vector calculus operators, can enhance segmentation beyond data-driven learning alone. Methods: We propose M-Net (Math-Augmented Network), which integrates three complementary mathematical p
This case study presents IF: CARGO, an experimental puzzle game that uses a large language model as a semantic compiler rather than an autonomous game-playing agent. Players author IF/THEN rules in natural language, which the model translates into a constrained command schema for deterministic validation and execution by the game engine. This architecture creates a playable loop of expression, execution, observation, and revision, framing AI interaction as semantic debugging. A mixed-methods pla
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Ove
With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time detection and autonomous response. In this paper, we propose a reinforcement learning-based dynamic cyber defense framework. We deploy a Deep Q-Network (DQN) to train effective defensive strategies to counteract the evolving cyberattacks. We leverage the CICIDS2017 dataset for model creation and the UNSW-NB15 dataset for external validation, involving preproc
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines,
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framewor
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, an
The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and institutional reform. We argue that this framing is insufficient: certain configurations of actors, systems, and institutions render AI accountability conceptually unachievable regardless of effort. We introduce the concept of constitutive AI unaccoun
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. This limits their adoption in high-stakes and safety-critical settings. Counterfactual explanations address this by revealing the minimal structural modifications that would change a model's prediction. On graphs, however, such a modification is hard to produce. The search space
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as an inductive bias for world models that are more sample-efficient and generalize better. Yet prior object-centric world models (OCWMs) take the slot encoder as given and evaluate only in-distribution, leaving open whether the object-centric bias actually deliv
Recent studies have shown that binary-to-image representations can enable effective machine learning-based results for malware detection and classification. However, performance can vary significantly, depending on the technique used to convert binaries to images. Furthermore, the explainability and interpretability of image-based models is largely unexplored within the malware domain. In this research, we employ Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool,
Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead. Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facil
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discover
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging d
The integration of Artificial Intelligence (AI), generally as Machine Learning (ML) algorithms, in all levels and aspects of cellular networks demonstrates the success of data-driven algorithms; for example, the Radio Intelligence Controller (RIC) of the O-RAN paradigm bestows the network with optimised radio resource allocation, load balancing or energy efficiency functions, among others. Nevertheless, this dependency on data opens new security vulnerabilities, as attackers can alter data prope
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public be
Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenarios. Existing approaches to building SLMs typically follow two paths: training compact models from scratch, or compressing larger pre-trained models using methods such as pruning, quantization, or distillation. As language models become increasingly integrated into real-world applications, ensuring their trustworthiness
Understanding how lower-limb muscle groups coordinate is important for studying movement impairment, rehabilitation, and physical performance. Reproducible analysis of this coordination requires multimodal recordings that relate local muscle-related signals with body-level kinematics. Complementing neural-level electrical activation captured by EMG, AMG provides a valuable mechanical approach to monitoring muscle activity. Here, we introduce a synchronized, multimodal dataset for healthy-adult l
Training and validation of Embodied AI for social navigation critically depends on realistic simulation environments, yet many current approaches fail to find a balance between realism and simulability. We propose D3D-GEN, a novel world generation system that combines a domain agent with a retrieval-augmented generation (RAG) pipeline grounded in that domain. Our system enables users to rapidly generate domain-grounded, fully interactive 3D worlds by automating both the collection of domain know
The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomize
Avatar interaction shapes how engaging and immersive a metaverse experience feels, and for that interaction to feel natural, avatars need to respond to users without forcing them through a controller-based interface first. This paper describes a gesture-driven interaction layer built for a browser-based metaverse onboarding environment, where users explore a set of virtual rooms as an avatar and interact with embedded video, document, and quiz content using hand, arm, and head gestures instead o
Browser-based webcam gaze trackers are increasingly used for crowd-scale data collection and in clinical settings where lab eye trackers are impractical, but the reported latency numbers may not represent real world functionality. The common practice of timestamping each gaze sample when it is emitted, rather than when its source frame was captured, makes the measured inference latency read about $0\,$ms no matter how slow the engine really is. We show how to measure it honestly, recovering a pe