01:52 UTC

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce
arXiv yesterday Safety & alignment

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconst
arXiv yesterday Agents & autonomyEnvironment

FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we intr
arXiv yesterday RegulationPrivacy

Contrastive ESA: Human Evaluation of Multiple Translations at Once

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to
arXiv cs.HC yesterday Children & education

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR as a training-free, retrieval-augmented task, in which statistical descriptions
arXiv cs.LG yesterday PrivacyHealthcare

WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recommendation models (FRMs). Existing approaches typically enhance recommendation with explicit Chain-of-Thought (CoT) under the Think-then-Answer paradigm. However, generating lengthy rationales introduces substantial inference overhead, while fixed CoT templates struggle to model diverse, dynamic, and context-dependent user interests. We propose WhisperRec, an ef
arXiv yesterday

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a
arXiv yesterday

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key
arXiv yesterday RegulationAgents & autonomy

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. The natural fix is a guard-agnostic recover-and-decode amplifier that transcribes image content and restates encoded text into its plain payload before the guard, so any off
arXiv red teaming query yesterday Safety & alignmentMilitary & security

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together. This monolithic design obscures workflow structure, inflates scaling overhead, forces users to manage low-level GPU coordination, and limits fine-grained fairness in multi-tenant
arXiv yesterday Bias & fairness

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with pr
arXiv red teaming query yesterday Safety & alignment

A Design Study on Voice-based Interaction for Immersive Network Visualization and Analysis

Visual network analysis leverages network visualization authoring techniques to facilitate sensemaking, serendipitous discovery, and hypothesis verification on network data. However, transferring the same paradigm to immersive environments is non-trivial due to insufficient UI affordance for authoring operations. Researchers have studied combining multiple modalities for interactions, but the high learning curve of such input systems limits their adoption by typical data analysts, let alone for
arXiv cs.HC yesterday Environment

Parameterized Fair Resource Allocation under Diversity Constraints

Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and course allocation, and is commonly formulated as an optimization problem with diversity constraints to ensure group fairness. Existing approaches typically enforce these constraints as hard conditions, which overly restrict the feasible solution space and often lead to suboptimal allocations. In this paper, we propose PRA, a parameterized framework for
arXiv fairness query yesterday Bias & fairnessAgents & autonomy

Navigating the DEIverse: A comprehensive review and research agenda on diversity, equity, and inclusion in the metaverse

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Paloma Almodóvar, Alberto Ferraris
Technology in Society yesterday Bias & fairness

Reconfiguring the global factory: The synergistic role of additive manufacturing, explainable AI, and knowledge acquisition in MNC operations

Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Femi Olan, Konstantina Spanaki, Uchitha Jayawickrama
Technology in Society yesterday Transparency

Artificial Intelligence Ethical Awareness of University Students in Ghana: A Network and Latent Profile Analyses

Publication date: Available online 27 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Bernard Yaw Sekyi Acquah, Iddrisu Salifu, Francis Arthur, Mark Inkoom, Sheriff Dagimah Suradji, Christian Inkoom, Emmanuel Quayson, Silas Afutu Quaye, Francis Obeng Gyedu, Sharon Abam Nortey
Computers and Education: Artificial Intelligence yesterday Children & education

From Simulation to Flight: Simulation-Assisted Drone Learning with Teacher-AI Co-Designed Scaffolds for Secondary Students’ STEM Knowledge and Competencies

Publication date: Available online 27 July 2026 Source: Computers and Education: Artificial Intelligence Author(s): Richard Chung Yiu Yeung, Chi Ho Yeung, Daner Sun, Therese Keane, Yuqin Yang
Computers and Education: Artificial Intelligence yesterday Children & education

How experience moderates the impact of AI suggestions on researchers' perceptions of their ideas

Publication date: October 2026 Source: Research Policy, Volume 55, Issue 8 Author(s): Matthias Tröbinger, Anil R. Doshi, Sen Chai
Research Policy yesterday Regulation

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters,
arXiv yesterday Agents & autonomy

Simulating Single Transferable Voting for the Colorado House of Representatives

arXiv:2607.25105v1 Announce Type: new Abstract: Social choice theory research demonstrates that single transferable voting (STV) results in more proportionally representative legislative bodies. We aim to understand how using multi-member districts and ranked ballots with STV would affect the representation of political parties in the Colorado House of Representatives. We investigated this objective by producing 10,000 multi-member districting plans of Colorado, generating ranked ballots for eac
arXiv cs.CY yesterday Finance, VC & PE

Passive wearable physiology tracks a state-level material-hardship gradient in resting heart rate

arXiv:2607.25301v1 Announce Type: new Abstract: Resting heart rate is an established marker of cardiovascular risk, but population-scale measurement has depended on clinical or survey instruments. We ask whether passively sensed consumer-wearable physiology recovers the socioeconomic gradient established in clinical cohorts. Using 19.1 million quality-filtered photoplethysmography readings from 18,734 opt-in users of the Welltory app, we computed cohort-adjusted mean daytime resting heart rate p
arXiv cs.CY yesterday Healthcare

Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data

arXiv:2607.25526v1 Announce Type: new Abstract: How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respon
arXiv cs.CY yesterday Transparency

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

arXiv:2607.25648v1 Announce Type: new Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - under
arXiv cs.CY yesterday Regulation

Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course

arXiv:2607.24755v1 Announce Type: cross Abstract: This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming (OOP) courses. Generative artificial intelligence (GenAI) tools are increasingly used by students in programming education, yet evidence on their educational impact remains mixed. In particular, little is known about how students integrate GenAI tools when learning OOP, and how different patterns of use relate
arXiv cs.CY yesterday Children & education

From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs

arXiv:2607.24757v1 Announce Type: cross Abstract: This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to make students' programming processes easily visible to teachers. We developed a web application that analyzes log files generated by Thonny (an IDE for Python) and produces interpretable views of students' programming processes. Teachers can upload a log, a ZIP archive, or a folder containing logs for a group
arXiv cs.CY yesterday Children & education

Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents

arXiv:2607.24759v1 Announce Type: cross Abstract: Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across s
arXiv cs.CY yesterday Agents & autonomy

PATHFinder Agent for Tailored Prenatal Care

arXiv:2607.24768v1 Announce Type: cross Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through
arXiv cs.CY yesterday HealthcareAgents & autonomy

Empathy and the Human-Moment Gaps of AI Chatbots: Insights from Empathy Displacement Theory

arXiv:2607.24775v1 Announce Type: cross Abstract: Artificial intelligence (AI) chatbots are increasingly deployed in domains where empathy is essential, including healthcare, education, and customer service. However, their capacity to sustain authentic human moments remains structurally limited. This paper introduces two interlinked conceptual models to explain and address this limitation. First, the Human-Moment Gap Framework (HMGF) identifies three structural empathy deficits in AI-mediated in
arXiv cs.CY yesterday Jobs & economyHealthcare

The AI Wave and the Reinvention of Game Discovery: Oversupply, Structural Correction, and Agentic Player-Game Matching

arXiv:2607.25010v1 Announce Type: cross Abstract: AI-assisted production has sharply reduced the cost and team size required to ship a video game, producing a supply shock on open marketplaces. Recent estimates put Steam release volume at roughly sixty new titles per day, with median per-title revenue for a large share of releases falling below the platform's own submission fee [1]. This paper asks whether the resulting oversupply constitutes an emerging market crash or a structural correction,
arXiv cs.CY yesterday Agents & autonomy

Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

arXiv:2607.25057v1 Announce Type: cross Abstract: As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Dr
arXiv cs.CY yesterday Jobs & economy

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

arXiv:2607.25425v1 Announce Type: cross Abstract: Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper report
arXiv cs.CY yesterday Bias & fairness

Detecting CSAM Text-to-Image LoRAs From Weights

arXiv:2607.25750v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's update
arXiv cs.CY yesterday Children & education

Effort Matters in Score-Based Admissions: How Retaking and Aggregation Shape Test Scores

arXiv:2607.25974v1 Announce Type: cross Abstract: Observed standardized test scores are the result of an endogenous process: students strategically allocate effort across multiple retake attempts to improve their outcomes. Because students differ in their ability to make these investments, the interaction between applicant strategy and institutional scoring rules---such as the widely used Single-Sitting and Superscoring policies---can disparately distort observed scores. We develop a strategic f
arXiv cs.CY yesterday Children & educationFinance, VC & PE

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

arXiv:2507.11548v3 Announce Type: replace Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias relative to human judgment. However, this framing leaves a prior question unresolved: whether these systems are capable of performing the evaluative task at all. This study presents a two-part audit of eight widely used AI platforms used for resume screening. Drawing on the concept of the Illusion of Neutra
arXiv cs.CY yesterday Bias & fairnessTransparency

Three Lessons from Citizen-Centric Participatory AI Design

arXiv:2602.08554v2 Announce Type: replace Abstract: This workshop paper examines challenges in designing agentic AI systems from a citizen-centric perspective. Drawing on three participatory workshops conducted in 2025 with members of the general public and cross-sector stakeholders, we explore how societal values and expectations shape visions of future AI agents. Using constructive design research methods, participants engaged in storytelling and lo-fi prototyping to reflect on potential commu
arXiv cs.CY yesterday Agents & autonomy

LLM-generated personalized nudges for improving pro-environmental behavior: Field evidence from resource conservation

arXiv:2604.03881v2 Announce Type: replace Abstract: Encouraging pro-environmental behavior remains a major challenge for sustainable cities. Conventional feedback nudges can show individuals how their current behavior compares with environmental goals but often provide limited guidance on what to do differently in daily life. This study examines whether supplementing weekly feedback on participants' behavior with LLM-generated personalized action suggestions improves pro-environmental behavior,
arXiv cs.CY yesterday Environment

Less Deliberate in Teams: Student LLM Use Across Individual and Collaborative Work

arXiv:2606.30860v2 Announce Type: replace Abstract: As large language models (LLMs) become common in computing courses, we need to understand how the social setting shapes how students use them. This paper reports findings from a semester-long study of 96 undergraduate students who completed six assignments, alternating between individual homework and team project milestones. We tracked LLM usage, prompting habits, and how students verified AI-generated output across all six assignments. LLM usa
arXiv cs.CY yesterday Children & education

Generative AI Availability, Grades, and Student Satisfaction at a Large University

arXiv:2607.21534v2 Announce Type: replace Abstract: The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as
arXiv cs.CY yesterday Children & education

"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research

arXiv:2311.18424v3 Announce Type: replace-cross Abstract: Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across dis
arXiv cs.CY yesterday Healthcare

When Algorithms Meet Artists: Semantic Compression and Stake-holder Marginalisation in Public AI-Art Discourse (2013-2025)

arXiv:2508.03037v5 Announce Type: replace-cross Abstract: Artists occupy a paradoxical position in generative AI. Their own work trains models that now compete with them, replicate their styles, and reshape the creative economy they inhabit. Yet whether artist concerns achieve proportional representation in the public discourse that shapes AI governance remains an open empirical question. We mapped the semantic landscape of public AI-art discourse from 2013 to 2025, drawing on 1,736 text chunks
arXiv cs.CY yesterday RegulationJobs & economy
← Newer Older →