Topic · updated daily · RSS feed for this topic
Healthcare
Clinical AI, diagnostic bias, patient safety and medical-device regulation — the healthcare front of AI ethics, daily.
Medical robotics beyond automation: Human-robot collaboration and the RONNA system as a socio-technical case study
Publication date: September 2026 Source: Technology in Society, Volume 88 Author(s): Marina Raguž, Domagoj Dlaka, Marko Švaco, Petar Marčinković, Dominik Romić, Filip Šuligoj, Bojan Šekoranja, Darko Chudy, Bojan Jerbić
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
arXiv:2607.26317v1 Announce Type: new Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natur
Hearsay: Vision-Language Medical Diagnoses Without an Image
arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts t
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and exam
The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning
arXiv:2507.04398v3 Announce Type: replace-cross Abstract: Generative AI is becoming part of academic writing, but its educational value depends on how control is shared between learner and system. This study examined an agency gap: performance differences that may arise when AI agent initiative is misaligned with learners' generative AI literacy. Seventy-nine medical and nursing students completed two multimodal analytical writing tasks using healthcare simulation data visualisations. They were
OpenAI CFO Sarah Friar tells employees that annualized revenue in July topped all of Q2
OpenAI is trying to reassure employees that the business is healthy as competition emerges from Anthropic as well as a host of open-source players.
Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI
Weng previously served as the VP of AI Safety Research at OpenAI.
Examining the Roles of Technology Across the Health Care Journey for Individuals With Obsessive-Compulsive Disorder: Qualitative Interview Study
Background: As digital technologies become increasingly embedded in daily life, their roles in mental health care have expanded and diversified. Digital tools are being explored as interventions for obsessive-compulsive disorder (OCD) across the care continuum, including symptom recognition, access to care, treatment, and self-management. However, there is limited empirical understanding of how individuals living with OCD use digital technologies in situ to navigate their health care journeys or
The Silicon Valley Health Trend Making Doctors Nervous
Is more data about your body actually good for you?
Experts disagree on how to fight AI disinformation, but agree that health and politics need different solutions
When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). The post Experts disagree on how to fight AI disinformation, but agree that health and politics need different solutions first appeared on HKS Misinformation Review .
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
Background: Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. Objective: This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospad
Detecting Narcissistic Personality Disorder Traits on Forums: Proof-of-Concept Study
Background: Identifying traits of narcissistic personality disorder (NPD) is clinically challenging, yet early detection can significantly improve outcomes. Online forums have become a major source of self-expression, offering new opportunities to understand mental health. However, analyzing this complex language requires new tools. Objective: This study aims to determine whether a machine learning model could be trained to reliably detect language patterns associated with NPD traits in Reddit p
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a
🏃 Fitness Tracker Privacy Fails | EFFector 38.14
Watches, bands, and rings—if you want to digitally monitor your fitness, more companies than ever are selling devices to do it. And more Americans than ever now own at least one wearable health device. But what are the companies that make fitness trackers doing to protect our sensitive data from prying eyes? A lot less than they could be, it turns out. We're explaining what companies can do to protect your health data, and more, with our EFFector newsletter . JOIN OUR NEWSLETTER For over 35 year
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20
A new benchmark for evaluating patient-facing health AI agents
PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.
How a medical database developed at MIT evolved into a global standard of data-sharing
The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated
In this edition of AI Prognosis: A conversation about benchmarking leading clinical chatbots, investor view on AI in biopharma, and more.
Hearsay: Vision-Language Medical Diagnoses Without an Image
When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply
AI in healthcare is an evolving landscape of new technologies, productivity benefits and legal uncertainties
A standardized framework for regulating the safety and efficacy of AI in healthcare has yet to be established.
FTC and States Act Against Hims & Hers for Deceptive and Unlawful Privacy Practices
Complaint alleges telehealth provider shared consumers’ sensitive health information with third-party advertising platforms despite promising patient privacy The Federal Trade Commission, joined by Utah and California, by and through Los Angeles County Counsel, today sued Hims & Hers alleging that the telehealth provider shared consumers’ sensitive health information about medical conditions with third-party advertising platforms despite claiming its services maintain consumers’ privacy and dece
Kentucky Governor to McConnell: Prove Fitness to Serve or ‘Resign’
Democrat Andy Beshear sent a letter to the Republican Senator who has been absent from his public duties and faces questions about his health status.
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework compri
ENHANCE (Tailored Intervention for Brain Health and Cognitive Enrichment for Cognitive Health), a Coach-Supported Digital Intervention for Dementia Prevention in Underserved Older Adults: Co-Design and Usability Study
Background: Digital multidomain interventions hold promise for dementia risk reduction; however, populations at higher dementia risk, including those experiencing socioeconomic and educational disadvantage, remain underrepresented in trials, and engagement with digital interventions often declines over time. Coproduction and blended models that combine digital tools with human support may improve reach, acceptability, usability, and sustained engagement. Designing interventions that are usable a
AI Is Hyper-Scaling Digital Inequality
Artificial intelligence is rapidly becoming part of everyday infrastructure–in some places. It helps write emails and software code, filters job applications, powers recommendation systems, and is increasingly being integrated into education, health care, finance, and public administration. Industry leaders talk about “AI for everyone,” while governments rush to publish national AI strategies and build sovereign compute. Yet over the past decade, working on digital inclusion and digital literacy
The Hidden Cost of Stress at Work
Policymakers, employers, and professional organizations have opportunities to design work environments that make us healthier, write Lilly Springer, David Slusky, and Anupam B. Jena.
Kaggle removes problematic stroke dataset for copyright infringement
The online data repository Kaggle has removed a problematic dataset for violating intellectual property rights months after it was flagged by sleuths for containing celebrity photos and lacking information on data provenance or confirmed medical diagnoses. Researchers had used the images to train machine learning algorithms to detect stroke, but the dataset instead contained images … Continue reading Kaggle removes problematic stroke dataset for copyright infringement
RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment
Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies. Retrieval-Augmented Generation for Human Activity Recognition (RAG-HAR) addresses this by framing HAR as a training-free, retrieval-augmented task, in which statistical descriptions
STAT+: Clinical chatbots are taking medicine by storm. Should doctors trust them?
As LLMs compete for doctors' attention, some developers say the science of benchmarking their AI tools for safety and accuracy is flawed.
Altron appoints Collin Govender as MD of Altron HealthTech
Govender succeeds Leslie Moodley, who announced his early retirement earlier this month.
Oregon Institution Lacks Funds to Move Away From Animal Testing
Oregon Institution Lacks Funds to Move Away From Animal Testing Ryan Quinn Wed, 07/29/2026 - 03:00 AM The Trump administration is pushing to end animal testing, but Oregon Health and Science University said federal dollars weren’t available to turn its National Primate Research Center into a sanctuary. Byline(s) Ryan Quinn
What We Know About the Government’s Investigations Into Medical Schools
What We Know About the Government’s Investigations Into Medical Schools Johanna Alonso Wed, 07/29/2026 - 03:00 AM Initial findings from probes into medical schools’ admissions practices have left experts questioning the government’s evidence and analysis. Byline(s) Johanna Alonso
Passive wearable physiology tracks a state-level material-hardship gradient in resting heart rate
arXiv:2607.25301v1 Announce Type: new Abstract: Resting heart rate is an established marker of cardiovascular risk, but population-scale measurement has depended on clinical or survey instruments. We ask whether passively sensed consumer-wearable physiology recovers the socioeconomic gradient established in clinical cohorts. Using 19.1 million quality-filtered photoplethysmography readings from 18,734 opt-in users of the Welltory app, we computed cohort-adjusted mean daytime resting heart rate p
PATHFinder Agent for Tailored Prenatal Care
arXiv:2607.24768v1 Announce Type: cross Abstract: Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through
Empathy and the Human-Moment Gaps of AI Chatbots: Insights from Empathy Displacement Theory
arXiv:2607.24775v1 Announce Type: cross Abstract: Artificial intelligence (AI) chatbots are increasingly deployed in domains where empathy is essential, including healthcare, education, and customer service. However, their capacity to sustain authentic human moments remains structurally limited. This paper introduces two interlinked conceptual models to explain and address this limitation. First, the Human-Moment Gap Framework (HMGF) identifies three structural empathy deficits in AI-mediated in
"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research
arXiv:2311.18424v3 Announce Type: replace-cross Abstract: Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across dis
FTC sues Hims & Hers for allegedly sharing patient information with third-party platforms
Popular telehealth provider Hims & Hers "shared consumers’ sensitive health information with third-party advertising platforms such as Meta and Snap despite promising to protect patient privacy," the federal government alleges.
Statement of Organization, Functions, and Delegations of Authority
The Food and Drug Administration's (FDA) plans to centralize and enhance key functions across the agency. These changes will reduce redundancies, improve efficiency, and advance alignment to better serve the American public.
Requiem Without an Orchestrator: A Commentary on Hollanek and Nowaczyk-Basińska (2024)
Hollanek and Nowaczyk-Basińska (2024) analyse the harms of AI-enabled re-creation services and offer four recommendations to providers. I accept their diagnosis but argue that their prescription shares a premise with the industry it seeks to reform: that a responsible provider remains present to carry it out. All four recommendations (retirement procedures, meaningful transparency, adult-only access, and mutual consent) are addressed to a continuing operator. Yet the paper’s own opening example,
DrugPred: an EdgeConv-GNN and Bio_ClinicalBERT based polypharmacy ADR prediction and specialist recommendation model
Adverse drug reactions (ADRs) are caused by medication and are considered a serious issue in healthcare when there is simultaneous use of different medications resulting in drug-drug interaction (DDI). Traditional approaches mostly focus on the effects caused by a single drug, and they fail to capture the side effects from drug combinations. In this research work, a deep learning-based DrugPred framework is proposed to predict ADR risks by integrating individual drug effects, interaction statist