Archive · 2026-07-19
AI ethics on Sunday, 19 July 2026
104 items published this day, across 4 categories.
News (46)
AI job worries grow. But some human skills can’t be replaced by machines | Gaynor Parkin and Dave Winsborough
The fear of becoming obsolete is a rising source of anxiety as artificial intelligence emerges in the workplace. Staying strongly connected with others will help The modern mind is a column where experts discuss mental health issues they are seeing in their work Earlier this year Paul* discovered via a flurry of media stories that his business group had been targeted for job cuts, alongside a number of other public-sector agencies. For him, this was the second time, after his technical managemen
The World Cup Ends Today. The Surveillance Will Continue.
ICE and the Architecture of Authoritarianism
Step Into the ‘Zone of Genius’ (Before A.I. Takes Your Job)
A decades-old self-help concept is gaining new purchase among people searching for meaningful work in the age of artificial intelligence.
Government use of automated AI decision-making to be curbed under new Australian rules
New national plan is accompanied by Labor push for digital duty of care legislation Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast The use of AI in automated decision-making by government departments and agencies will be subject to tough rules under a new national plan, expected to extend to consumer protections, workplace safety and privacy. As the Albanese government grapples with the rapid growth in the use of artificial int
The Silent World Cup Winners are the Companies Normalizing Surveillance
Victoria announces new social media ‘demasking’ powers for accounts accused of vilification
New laws would give Vcat power to force social and AI platforms to identify anonymous users in move premier says will protect children Get our breaking news email , free app or daily news podcast Social media companies could be forced to identify anonymous accounts accused of online vilification, under new laws being proposed in Victoria. The Victorian premier, Jacinta Allan, announced a suite of social media reforms on Sunday, saying families needed new ways to protect their children online. Co
New trial of AI-powered traffic lights will be a test for who gets priority on public roads
Australia’s first AI traffic-light trial could cut delays. It also raises hard questions about fairness, safety and who gets priority on public roads.
Can an Apple lawsuit derail OpenAI’s hardware plans?
On the latest episode of Equity, we debate whether Apple's lawsuit will cast over OpenAi's much-discussed plans to get into hardware and go public.
Flight Centre to embrace AI agents, ecommerce consolidation
Appoints new executives including chief AI officer.
U.S. military announces the death of another service member, blaming the ‘controlled detonation’ of a downed Iranian drone in Iraq
Since the war began, 17 U.S. service members have been killed.
The U.K. is getting a new prime minister, but bond investors are really in charge and ‘hyper-reactive,’ Wall Street veteran says
"But investors know it’s the bond market that will call the shots in the $4.2 trillion economy no matter who sits in the prime minister’s office."
Shocked AWS customers receive astronomical bill estimates
Billing estimate engine update goes wrong.
Trump Administration Scales Back Endangered Species Protections—Again
The Interior Department’s Fish and Wildlife Service revised regulations implementing the Endangered Species Act. Conservationists warn the changes could put vulnerable species at greater risk.
An Agency That Supports Health-Care Improvement Cuts Off Grant Funding
An Agency That Supports Health-Care Improvement Cuts Off Grant Funding Ryan Quinn Sun, 07/19/2026 - 02:30 PM The Agency for Healthcare Research and Quality sent researchers letters last week saying they wouldn’t receive long-awaited funding continuations. The agency cited the “best interest of the federal government.” Byline(s) Ryan Quinn
‘Shark Tank’ Star Kevin O’Leary Sued For Claiming Utah Data Center Opposition Has Ties to China
Kevin O’Leary is being sued for defamation by a group fighting data center construction in Utah after the “Shark Tank” star went on Fox News and claimed some of the individuals involved with the opposition have ties to China, according to The Hill. The suit, which was also against Fox News, was filed on Wednesday […]
USAF Chief Wilsbach on hardened shelters, T-7 flight and F-35 delays
In a one-on-one with Breaking Defense, Air Force Chief of Staff Gen. Kenneth Wilsbach laid out his big concerns for the service.
Los expertos creen que nos faltan árboles: "Una zona verde sin arbolado suficiente pierde gran parte de su capacidad para actuar"
Las olas de calor extremo son un desafío estructural cada vez más intenso y duradero y, ante unas ciudades convertidas en trampas de asfalto con una clara falta de árboles, el concepto de "refugio climático urbano" ha ganado fuerza como infraestructura de emergencia. Sin embargo, no cualquier espacio verde actúa como escudo térmico, puesto que una reciente investigación liderada por la Universidad de Granada evidencia que el diseño, la morfología y, sobre todo, el tipo de arbolado son factores c
Morning Bid: Rising oil, yields rain on AI party
The rise in 30-year Treasury yields above 5.0% carries a warning for equity valuations. Just a glance at a chart shows yields have spent little time above that barrier in the past two decades, and ...
Conflicts, aircraft orders in focus as Farnborough Airshow kicks off
Farnborough Airshow opens on Monday with Boeing and Airbus pursuing aircraft deals and defence firms vying for a share of booming military budgets fuelled by wars in Ukraine and the Middle East.
Asia shares shaky as oil climbs, earnings loom
Asian share markets slipped on Monday as the escalating conflict in the Gulf lifted oil prices and fanned fears of inflation, while a packed week of major tech earnings will further test investor ...
TSMC sees multi-year demand for AI chips, ramps up Arizona investment
TSMC is seeing strong, multi-year demand for its AI chips as it invests a further $100 billion to expand its Arizona facilities, but it needs to address several challenges, such as a shortage of ...
Could AI chip boom make ASML Europe's first trillion-dollar firm?
AMSTERDAM, July 20 (Reuters) - The global artificial intelligence boom has propelled ASML (ASML.AS), opens new tab to the top of Europe's stock market, as soaring demand for AI computer chips flows to ...
Seoul AI ETF howler morphs into zombie volatility
Wild price swings already have watchdogs rueing their May approval of leveraged funds tracking mega-cap stocks SK Hynix and Samsung. But the limits they’re now imposing don’t address the root of ...
Book Bans, Censorship and Funding Fears Challenge Ohio Public School Librarians
Public school librarians in Ohio are raising alarms about book bans and funding cuts. School librarians have been navigating challenges in their work as long as they’ve been among the stacks in their local districts. Proposed legislation to filter the reading choices students can make has brought concern, and budget reductions make some worry about […]
Trump pushes UN ‘free speech’ declaration in veiled attack on EU tech regulation
The White House has repeatedly condemned EU rules that govern online platforms as censorship. Now they want the world to put the critique into writing.
Top Pentagon official blasts OpenAI's Dean Ball
Secretary of Defense Pete Hegseth and Emil Michael, under secretary of defense, at the Pentagon last year. Photo: Win McNamee/ ...
Alibaba says newest Qwen AI model is second only to Anthropic’s Claude Fable 5
Alibaba Group Holding has previewed its next-generation artificial intelligence model Qwen3.8, which the company says is “second only” to Anthropic’s Claude Fable 5. Qwen3.8-Max-Preview, the preview version of the strongest model of the Qwen family, had been made available on Alibaba’s Token Plan subscription service, as well as its Qoder and QoderWork agentic platforms, the Chinese tech giant said on an official X post on Sunday. With 2.4 trillion parameters, it was “one of the most powerful...
Self-healing GPU nodes in Kubernetes: What we learned building the EKS node monitoring agent
When you run Kubernetes at the scale we do on Amazon EKS, nodes break constantly. GPUs fall off the PCIe The post Self-healing GPU nodes in Kubernetes: What we learned building the EKS node monitoring agent appeared first on The New Stack .
US renews strikes on Iran after two military personnel killed by Iranian attack
The U.S. said it had completed an eighth straight night of attacks against Iran after earlier announcing that two U.S. military personnel were killed in Jordan, while U.S. allies in the region ...
How Chinese tech giants from Ant to Tencent use AI agents to win over enterprise clients
Chinese tech giants are doubling down on enterprise artificial intelligence agents with new products unveiled at the country’s top AI summit, signalling heightened domestic rivalry to win over business clients as agent-based AI adoption accelerates. At the four-day World Artificial Intelligence Conference (WAIC) in Shanghai which concludes on Monday, major tech companies including Ant Group, Tencent Holdings, Alibaba Group Holding and Baidu launched or showcased offerings designed to integrate..
Europa tiene un problema con el tamaño de sus coches: el "carspreading" amenaza con devorar miles de aparcamientos
Los coches nuevos que se venden en Europa son cada año un poco más largos , más altos y más anchos. El fenómeno tiene ya nombre propio, "carspreading", y según un nuevo informe de las organizaciones ecologistas Transport & Environment (T&E) y Clean Cities, si esta tendencia no se frena, tendrá consecuencias directas tanto en la seguridad vial como en el aparcamiento disponible en nuestras ciudades. Te contamos los detalles. Qué está pasando. El informe, publicado por T&E y Clean Cities, ha anali
NYC’s Mayor Asked for Parents’ Opinions on Schools. Here’s What Some Have to Say
When Zohran Mamdani became mayor of New York City, he vowed to gather community feedback on how the public schools are working. His chancellor, Kamar Samuels, has assembled working groups “tasked with creating customized proposals to accelerate the work of building a stronger school system, rooted in academically rigorous, safe and integrated schools.” So I […]
Meta avisa: "Llevamos 20 años construyendo infraestructura para humanos, quizás tengamos 20 meses para reconstruirla para los agentes"
Las empresas han abrazado el boom de la IA agéntica, hasta el punto de que algunas se están ahogando en tantos agentes de IA . Los empleados están creando agentes sin control, disparando el consumo de tokens y provocando que muchos de esos agentes dupliquen tareas. Pero el verdadero problema de fondo es otro: la infraestructura sobre la que está corriendo todo esto no está preparada para soportarlo. La advertencia de Meta. Lo cuentan en Venture Beat . Durante la charla VB Transform 2026, el vice
Google's AlphaEvolve Reaches General Availability with Evolutionary Code Optimization as a Service
Google's AlphaEvolve reached general availability on the Gemini Enterprise Agent Platform, turning the DeepMind research project into an evolutionary code optimization service. Evaluators run client-side so code never leaves the customer's infrastructure. Klarna doubled ML training throughput; practitioners note it only works where a measurable evaluation function exists. By Steef-Jan Wiggers
Divide grows between AI employees and executives over policy battles
Tech employees in Silicon Valley are increasingly finding themselves at odds with industry executives who are spending millions to push for light-touch AI regulation. The latest disagreement is playing out in a new political spending fight between OpenAI’s rank-and-file employees and the firm’s co-founder and president Greg Brockman. A group of current and former OpenAI...
These Parents Are Asking Reddit to Photoshop the Most Painful Pictures They Own. The Reason Why Is Tragic.
After losing a child, many parents turn to one online community—and find unexpected support.
"Lanzamos nuestro primer robot hace 20 años y seguimos intentando perfeccionarlo": Will Kerr, VP de nuevos productos de Dyson
¿Cómo se gestiona una de las empresas de hardware más innovadoras del mundo? En un mercado tan feroz como el de los aspiradores por la saturación de marcas, competencia y guerra de precios, la multinacional británica Dyson ha comenzado una profunda reestructuración estratégica para tratar de esquivar los males que acechan al sector tecnológico: la dependencia de un solo producto, la pérdida de agilidad operativa y la dosificación de lanzamientos. Viajamos a su sede central de Singapur para
El salario de los junior lleva años en caída libre. Un nuevo estudio sugiere que el impacto de la IA no es tan obvio
Llevamos un par de años viendo titulares sobre el destrozo que la IA está haciendo en el empleo junior . Y la réplica habitual: que no, que es excusa, que las empresas recortan cuando suben los tipos y la IA solo es una justificación cómoda. Las dos lecturas tienen su parte de razón. Ese es justo el problema. El estudio . El Stanford Digital Economy Lab ha estado analizando nóminas reales. No encuestas ni estimaciones: nóminas. Publicó un estudio a finales de 2025 pero mantiene actualizado un da
China state investors pledge further purchases of stocks
Chinese state-owned capital operators China Reform Holdings Corp (CRHC) and China Chengtong Holdings Group on Sunday announced fresh purchases of Chinese equities, according to separate statements, ...
Changzhou says it is building China’s first city-level clean-power AI token factory
The Chinese city of Changzhou is building what it claims will be the country’s first city-level “green token factory”, as local governments race to answer Beijing’s call to power massive artificial intelligence computing demands with clean energy. The project would be powered by green energy and could produce 60 trillion tokens per year once completed, the municipal government said in a statement on Sunday during the World Artificial Intelligence Conference in Shanghai. The city in the eastern..
AI chatbots reading X-rays can be dangerously confident even when they're wrong
The RadLE 2.0 benchmark tests whether AI models in radiology can tell when they should leave a diagnosis to a human. Many models deliver wrong findings with full confidence, and human radiologists are still well ahead. Before AI can diagnose on its own, it needs to learn when it's better to say nothing. The article AI chatbots reading X-rays can be dangerously confident even when they're wrong appeared first on The Decoder .
Investors blink on the AI trade
Why it matters: Chipmakers are the picks-and-shovels of the AI gold rush, and they are now looking shaky. A recent selloff, combined with new stress over competition from Chinese open-source AI models ...
Anzeige: IT-Jobs in Sicherheit, Administration und Entwicklung
Von Informationssicherheit �ber IT-Betrieb bis Netzwerktechnik - sechs IT-Jobs mit vielseitigen Aufgaben und moderner Technik. ( Golem Karrierewelt , Unternehmenssoftware )
Defence giants to maintain grip on weapons market despite drone boom
Report finds industry’s ‘primes’ will still account for 80% of global sales well into the next decade
These Hip-Hop Artists Were Already Teaching
A new degree program credentials hip-hop artists to bring communal, real-world learning to the classroom.
Field notes (6)
AllHere Chatbot Scandal Shows How Not to Deploy AI
Want proof that policymakers and school staff need help with artificial intelligence? Look no further than the Los Angeles Unified School District. Superintendent Alberto Carvalho resigned June 21 as ...
Event Recap July 17, 2026 AISI Workshop: Shaping the Future of AI Regulation
Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.
PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.
How AI is Transforming Scientific Discovery While Keeping Humans at the Center
From designing new antibodies to simulating 1,000 years of climate in a day, AI is transforming what's possible—but humans remain the ones deciding what matters.
New Approach to Scaling Laws Could Change How AI Models Are Trained
Leveraging statistical concepts from measurement science and education, AI researchers have greatly reduced the computational demand of predicting how the largest of large language models will scale up in the future. It could save millions of dollars in training costs.
🔮 Kimi K3 surprise & AI economics; the solar paradox; AI's right to learn, cancer vaccine & junior jobs++
Plus: Palantir’s power, cancer vaccines and AI on iPhones
Policy (2)
IMF Executive Board Concludes 2026 Article IV Consultation with Singapore
In 2026Q1, annualized q/q GDP expanded by 5.3 percent, reflecting continued AI-related semiconductor demand and ongoing infrastructure projects. A sharp increase in global energy prices following the ...
Unlocking the Potential: AI in Sub-Saharan Africa
Artificial intelligence (AI) is emerging as a general-purpose technology with the potential to reshape productivity, labor markets, and growth trajectories worldwide. For sub-Saharan Africa (SSA), the ...
Research (50)
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences--effective in low-light and privacy-sensitive environment--they face two major challenges: 1) learning and effectively exploiting interaction cues from skeletal data, and 2) compensating for the lack of visual information absent in skeletons alone. To address these challenges, we propose skeletal token alignment and rearrangement
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries. This paper presents Agentic ERP, an expert-system architecture that combines role-aligned large-language-model (LLM) agents with a risk-tiered human-in-the-loop harness and a graph-base
The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent Coordination
The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collaborative systems. Resource allocation is based on optimized collective arrangements accounting for agents' needs. Such coordination should not only be computationally efficient but also account for fairness, i.e., equitable redistribution of costs incurred by all agents. Recent literature has proposed several algorithms that efficiently determine optimal plan
SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks, yet remain scarce due to privacy constraints and annotation costs. We present SAGA (Synthetic Agentic Graph Architecture), a system for generating large-scale, semantically rich temporal graphs via a four-phase pipeline. Our Skeleton-First, Semantics-Second architecture decouples structure from semantics: (S) an O(1)-per-edge skeleton generator produces pow
AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals. The emerging AI-Generated Bidding (AIGB) paradigm widely adopts generative modeling to optimize bidding strategies, yet suffers from the limited mode coverage of offline datasets and inadequate task-state understanding, hindering effective exploration of optimal strategies. Large Language Models (LLMs), with prior world knowledge and reasoning capabilities,
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to improve manipulation performance. However, these modalities often differ not only in information content but also in sensing rates and inference latencies. Existing multimodal diffusion policies typically rely on synchronous fusion or manually designed multi-frequency architectures, which either slow down high-frequency feedback or limit extensibility to new m
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate. Reinforcement learning with verifiable rewards is a natural fit for this task, but standard GRPO reward relies on sparse outcome-level feedback and gives no signal about where a reasoning trajectory goes wrong, nor
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch. However, model repositories often provide incomplete machine-readable documentation about model provenance, licenses, datasets, limitations, and external references, creating transparency and governance gaps across the AI supply chain. Artificial Intelligence Bills of Materials (AIBOMs) address these gaps by documenting AI artifacts, including models, metadata, licens
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment
Test-time scaling empowers Large Reasoning Models (LRMs) to tackle complex tasks via extensive Chain-of-Thought (CoT). However, this often induces the "overthinking" paradox, where redundant reasoning increases computational overhead without guaranteeing accuracy. Existing test-time efficiency optimization methods primarily fall into two categories: information-theoretic approaches, which are prone to "deceptive convergence" where low uncertainty masks hallucinations, and latent representation a
A Diagnostic Framework for AI Agent Behavior
AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer def
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non-contiguous expert weights, the non-deterministic scatter-gather traffic generated by input-dependent token routing, and the tail-latency dependency imposed by synchronous expert output aggregation. To address these chal
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing serving systems and auto-parallelism compilers commit to limited transformations and fixed workload assumptions, so achieving high performance on a new application requires hand-crafting an efficient implementation. We present Flas
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when reading tool output. Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Concretely, a small head turns the agent's own internal rep
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot ma
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy thro
DiFA: Inference-Time Forward-Process Alignment for Diffusion Models
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this work, we propose Forward-Process Aligned Diffusion prediction (DiFA), a training-free framework that reframes inference-time data prediction refinement as a sequential state estimation problem. Rather than reusing past outputs sole
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably decline
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their str
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predi
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this class of threats as self-state attacks. In this paper, we investigate the OS resilience to this class of attacks. Formally, we characterize a four-axis attack space (Target, Mechanism, Granularity, Temporal); investigate the structural limits of prevention, detect
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable a
Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural networks. We investigated Differentiable Logic Gate Networks (Diff-Logic) as a hardware-native alternative that compiles models into pure Boolean circuits executable via bitwise CPU operations. Through rigorous iso-parameter experiments across four EEG datasets spanning two classification tasks, binary dementia detection and 3-class emotion recognition, we compared Diff-Logic agai
SLAM in Low-Light Environments: Project Report
Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous operations in real-world scenarios. Under low illumination, reduced contrast, sensor noise, and motion blur degrade both feature extraction and feature matching, while compensating with LiDAR, depth, or thermal sensors raises cost, power draw, and integration complexity. Existing benchmarks remain dominated by well-lit indoor or daylight sequences, leaving open how far SLAM with
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-inten
Three-Body Scattering for Generative Modeling
Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct regression supervision for a one-step generator. Three-Body Scattering Modeling (TBSM) for generation turns the energy distance into a constant-size per-projectile interaction: each projectile is attracted toward one real source and repelled from one independent
An Iterative Geometric Approach to Optimizing Separating Hyperplanes
Given a binary-labeled linearly separable dataset, and the objective is to compute the maximum-margin separating hyperplane, also known as the hard-margin Support Vector Machine (SVM) classifier. This paper investigates whether, if given an initial separating hyperplane, can it be exploited to reach this unique optimum more efficiently. We present a geometric approach that gradually improves the alignment of the hyperplane, starting from an initial separating hyperplane, while preserving separat
Working remotely, unequally: Evidence from a digitally divided labor market
Publication date: September 2026 Source: Telecommunications Policy, Volume 50, Issue 8 Author(s): José Wilmar Quintero-Peña
A comparative and integrative mapping of thematic trajectories, the digital turn, and future research frontiers in telecommunications policy
Publication date: September 2026 Source: Telecommunications Policy, Volume 50, Issue 8 Author(s): Menglan Luo, Yong Jiang, Yi-Shuai Ren
How demography shapes the macroeconomic returns to automation
Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Izabella Kuncz, Petra Németh, Eszter Szabó-Bakos
“You Might Like Each Other”: Mass Customization and the Structuring of Sociality Among Chinese Generation Z on the Soul App
Social Media + Society, Volume 12, Issue 3, July-September 2026. We employed the framework of mass customization to investigate how the Soul app, a leading interest-based social platform, is reshaping sociality among Chinese Generation Z. Six months of participatory observation, from December 2024 to June 2025, in Soul’...
Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instance-level transfer while near-matching clause density. At aligned size bins, with near-matched density and matched maximum clause width, we compare proof-hard expander-Tseitin and proof-easy ladder-Tseitin formulas, pigeonhole anchors, and density-mismatched controls. Theory separates their resolution hardness; a solver-specific Glucose mean-conflict proxy diff
The value of contact in legged locomotion: a survey of sensing channels, artificial intelligence and control
Legged robots traverse unstructured terrain through brief, intermittent foot–ground contacts whose support conditions are difficult to perceive and predict in real time. In such regimes, haptic feedback provides early and trustworthy evidence of traction limits, partial support, and incipient slip. This structured survey asks two questions: first, what locomotion-relevant contact evidence can be acquired and preserved under real deployment constraints; and second, how that evidence is translated
The shared blind spot: why diverse AI governance approaches fail for the same reason
AI governance instruments are proliferating, and so are their difficulties. Across major jurisdictions and international bodies, reform efforts built on substantially different premises encounter a recognizably similar pattern of failure. I argue that anticipatory regulatory governance rests on three operational premises —categorical stability, epistemic accessibility, and manageable pace—and that AI’s emergence, opacity, and velocity violate all three in compound. These premises form a distinct
School of Biological Sciences
NTU School of Biological Sciences (SBS) explores new knowledge and discoveries in molecular biology that will have major impact on the life sciences, widely acknowledged as the next technological ...
Dr Tomislav Karačić to investigate faith and digital futures with LSE grant
Dr Tomislav Karačić, Assistant Professor of Information Systems at the Department of Management and DSI Affiliate, has been awarded LSE Research and Innovation funding for his project, “Faith and AI: ...
From Perception to Assistance: Open-Vocabulary Shared Autonomy for Robotic Manipulation
Teleoperating a robotic manipulator in industrial environments demands precision that camera-based interfaces alone struggle to deliver. The operator must align the end-effector with a target in clutter, under limited depth perception, and without colliding with the surrounding structures. This paper presents a shared-autonomy framework that assists the operator throughout this process. A single RGB-D camera captures the operator's arm motion and hand gestures without wearables, fiducials, or a
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks. Existing jailbreak methods, mostly adapted from text-to-image attacks, suffer notable drawbacks when applied to T2V systems. They fail to fully leverage temporal consistency, an inherent characteristic of video generation. Besides, these methods demand heavy video query optimization, which is infeasible in practical black-box scenarios. Their adversarial prom
Semantic Context Matters: Analysis of Color Names Across Domains
Color naming is influenced not only by physical color values but also by the semantic context in which colors are used. This paper investigates context-dependent color naming by mapping color-name datasets from Cosmetics, Crayola, and Car-color vocabularies onto the 86 fuzzy color categories of the COLIBRI color model. Contextual variation is analyzed using category coverage, Shannon entropy, and maximum lift. The results show that the three contexts occupy the COLIBRI color space differently: C
PocketPPD: Screening for Postpartum Depression Risk Using Passive Smartphone Sensing
Postpartum depression (PPD) is a serious perinatal mental health condition affecting approximately 20% of new mothers worldwide. Common screening approaches for PPD, such as self-report questionnaires and active digital logs, rely heavily on user input and thus impose a substantial burden on participants, limiting their feasibility for long-term use. Recent passive mobile sensing (PMS) approaches have enabled low-burden detection of depressive symptoms using machine learning methods with multi-m
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting
Strategic Gaze: Attention Allocation and Transition Patterns Across Functional Areas of Interest by Gameplay Outcome
Video games present players with complex, spatially distributed information across interface elements, with attention shaped by visual features and task goals. Eye tracking provides a useful method for examining player attention through gaze behaviour during gameplay. Yet empirical game research has relied on accumulated fixation measures that capture where attention is directed and how long it is maintained within regions, leaving less known about how gaze moves between regions to coordinate di
Teach it to stop, not just to click
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across three cells (crossed data-draw $\times$ seed grid, bootstrap CIs) finds evaluation variance negligible ($σ_{\mathrm{eval}} \approx 0$) and the training-seed effect small everywhere ($\leq 10\%$); ins
SAVEstate: A Method for Documenting Player Reflection in Digital Games
In recent years, interest in eudaimonic player experiences (PX) - concerning reflection, meaning-making, and personal growth - has increased. However, most games user research methods are not well-suited to study eudaimonic PX, as they have been developed to evaluate features of hedonic PX, such as flow, immersion, and playability. To more deeply explore eudaimonic PX, we require methods that can 1) investigate how moment-to-moment PX shapes player reflection and 2) explore how players reengage
A Multi-Model Hybrid Defense Approach Against White-box Adversarial Attacks in Computer Network Traffic
It is crucial to safeguard computer networks from evolving network security threats and unknown cyberattacks. An essential tool for protecting computer networks against unknown cyber threats is Network Intrusion Detection System (NIDS). However, NIDS faces a major security concern due to its susceptibility to adversarial attacks. Adversarial attacks aim to deceive NIDS by crafting and injecting adversarial examples into the system. These adversarial inputs can deceive the NIDS into misclassifyin
Federated Lightweight Intrusion Detection in Drone Swarms with Knowledge Distillation
Drone swarms are increasingly deployed in critical applications such as surveillance, disaster response, and infrastructure monitoring. However, their reliance on open communication channels and their limited computational resources make them vulnerable to a wide range of cyber-threats. There is a growing interest in intrusion detection systems (IDS) specifically designed for drone environments and operations. However, the conventional solutions including Machine Learning (ML)-based approaches r