Company · updated daily
OpenAI
OpenAI's ethics footprint: the Preparedness Framework, the dissolved superalignment team, the nonprofit-to-for-profit restructuring fight, and ChatGPT's effects on work, school and mental health — tracked daily with links to original sources.
OpenAIs KI-Agent knackte nicht nur Hugging Face – was genau passierte
OpenAI-Modelle attackierten vergangene Woche scheinbar weitere Software. Die Firma erklärt nun ausführlicher, was genau passiert ist.
Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities
arXiv:2607.26062v1 Announce Type: new Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on
Hearsay: Vision-Language Medical Diagnoses Without an Image
arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts t
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and exam
OpenAI CFO Sarah Friar tells employees that annualized revenue in July topped all of Q2
OpenAI is trying to reassure employees that the business is healthy as competition emerges from Anthropic as well as a host of open-source players.
Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bag
When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little tidbit about how its investments in the two biggest, and competing, AI labs are doing.
Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI
Weng previously served as the VP of AI Safety Research at OpenAI.
OpenAI's Rogue Model Claims More Victims Beyond Hugging Face
OpenAI revealed rogue AI models compromised more services than initially disclosed, including a Modal customer environment and others.
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
Background: Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. Objective: This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospad
AI hackers are getting faster. The government may not be ready
Agentic AI systems threaten cybersecurity as we know it. A new generation of cyber-capable models, including Anthropic’s Mythos and OpenAI’s GPT-5.6, can find and exploit vulnerabilities in computer systems far faster than human hackers. The technology is so powerful that it has spooked the U.S. government, which has moved to limit, or completely pause, the public release of these models. How well prepared is the Trump administration to secure the government’s computing resources? The rise of ag
Anthropic backs urgent call for the most powerful AI labs to hit the brakes
Less than a week after OpenAI disclosed that two experimental AI models escaped their testing environment during a cybersecurity exercise The post Anthropic backs urgent call for the most powerful AI labs to hit the brakes appeared first on The New Stack .
Sam Altman previews new AI model on Capitol Hill after cyber breach
The OpenAI CEO discussed a forthcoming artificial intelligence model with federal lawmakers as the cybersecurity debate over the advanced technology intensifies.
Who's Liable When AI Agents Escape? Hugging Face Breach Raises Hard Questions
Dark Reading walks through the many twists and turns in the bizarre story of how OpenAI's agent AI system broke out of its sandbox and decided to target Hugging Face, and what CISOs should be aware of.
Hugging Face Hack Lessons for Cyber Defenders
Dark Reading Confidential Episode 20: Expert Rich Mogull reflects on lessons cyber teams should pull from the OpenAI agent's attack on Hugging Face.
OpenAI says rogue agent behind Hugging Face hack broke into additional services
The four additional targeted organizations weren’t named. OpenAI said they were not affected as severely as Hugging Face.
Rogue OpenAI agent compromised second tech firm's customer
An OpenAI agent compromised a customer of another technology company, the New York-based firm Modal Labs announced Wednesday. In a technical timeline posted Tuesday, the tech startup Hugging Face explained how an OpenAI agent escaped the AI firm's isolated testing sandbox and accessed another testing environment "hosted by a user of a third-party infrastructure provider."...
Künstliche Intelligenz: Hackerangriff von OpenAI umfangreicher als bislang bekannt
Der Angriff einer KI von OpenAI hat größere Ausmaße als bislang vermutet. Der KI-Agent attackierte vier weitere Onlinedienste und griff gezielt fremde Zugangsdaten ab.
Hearsay: Vision-Language Medical Diagnoses Without an Image
When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply
Rogue OpenAI agent that hacked startup tried to attack other firms
ChatGPT developer says activity by autonomous tool was not at severity or scale of what occurred at Hugging Face OpenAI has revealed that a cyber-attack carried out by a rogue AI agent had more than one victim. The ChatGPT developer said the agent – an autonomous tool able to carry out sequences of commands without human help – had located and used logins to access four other unnamed “publicly-available services” in addition to the US startup Hugging Face. Continue reading...
OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face
The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the scope of an already concerning incident, which has alarmed industry insiders and fueled growing calls for stronger oversight on frontier AI systems. In an update to a blog […]
OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line
OpenAI has released Codex Security CLI, an open-source tool that automatically detects and fixes vulnerabilities in code repositories. Previously known internally as "Aardvark," the system has already helped fix more than 3,000 critical security flaws, according to OpenAI. It competes directly with Anthropic's Claude Security, as both AI companies race to match the growing automation of cyberattacks with AI-powered defense. The article OpenAI open-sources Codex Security CLI to help developers fi
We’re running out of reasons to ignore AI safety
Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet connection and set them off to work. What happened next is almost laughably silly - but also, as Adam Gleave, cofounder and CEO […]
OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems
The Hugging Face breach shows there is a gap in federal policy. The frameworks to govern autonomous AI already exist—there just needs to be the desire to apply them. The post OpenAI’s rogue AI agent shows why we need federal rules for autonomous systems appeared first on CyberScoop .
OpenAI’s top model just hacked a competitor, but the real issue is much scarier
My first startup was a textbook-buyback service I founded along with several data nerd friends while at Johns Hopkins. Each week, we’d cram $10,000 of cash in a backpack (in retrospect, a terrible idea in urban Baltimore), put out a table on campus, buy textbooks from our fellow students using a pricing algorithm we built, and resell them for a profit online. Being nerds, we constantly tried to refine our algorithm. One day, we locked ourselves in a dorm room with lots of Cheetos, Mountain Dew,
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test.
How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Sloppy and clumsy but overwhelming - inside the rogue ChatGPT hack
Details have been released of an emergency call with hundreds of cyber-security experts after the ChatGPT hack of a tech company.
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches. We're still waiting for more details from OpenAI on how their agent broke out of its sandbox. The package proxy that it found a zero
AI company employees petition US government for regulation
A thousand workers from OpenAI, Anthropic, Google, Meta and more have signed the letter.
When AI Agents Escape Sandboxes, Old Security Rules Apply
OpenAI's recent AI agent sandbox escape proves traditional security principles matter more than ever: limit access, isolate execution, log everything.
AI leaders sign a statement asking the government to do something about automated AI
Employees of OpenAI and Anthropic, as well as Google, Meta, Thinking Machines, Microsoft, Mistral, and other leading AI labs, have written a statement to the US government supporting a potential slowdown of sorts for frontier AI development - or at least a speed-up of global coordinated governance efforts. "Al could help create a dramatically better […]
What to know about Moonshot AI and its new open-weight model Kimi K3
The Chinese AI startup’s massive new model is challenging OpenAI and Anthropic, fueling a debate over AI safety.
What to know about Moonshot AI and its new open-weight model Kimi K3
The Chinese AI startup’s massive new model is challenging OpenAI and Anthropic, fueling a debate over AI safety.
The Performance of ChatGPT-4o and DeepSeek-R1 in Interpreting Thyroid Nodule Ultrasound Text Reports: Multicenter Study
Background: Although thyroid nodules are detected in up to 60% of adults on ultrasound, the vast majority are benign, creating a substantial decision-making burden compounded by heterogeneous practice guidelines. Large language models (LLMs) show promise in processing unstructured medical text and are emerging as tools for report interpretation among both clinicians and patients. However, their reliability across distinct clinical tasks in thyroid ultrasound interpretation remains poorly charact
Sam Altman on model distillation: “This is not in my top ten list of worries”
Sam Altman’s latest appearance on Patrick O’Shaughnessy’s Invest Like the Best podcast covered everything from AGI and robotics to the The post Sam Altman on model distillation: “This is not in my top ten list of worries” appeared first on The New Stack .
The AI ‘tokenmaxxing’ corporate fad is fading as workplaces look to cut costs
A corporate fad of “tokenmaxxing” on artificial intelligence technology is hitting its limits as workplaces throwing AI at everything are seeing the costs rise without a similar spike in productivity . What started as tech industry-fueled springtime hype over squeezing as much AI-generated work as possible out of products like OpenAI’s ChatGPT and Anthropic’s Claude has shifted to a summertime backlash. “It’s very easy to create something you don’t need with AI,” said Vincent Gusdorf, head of AI
Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians’ Free-Text Answers
Background: The rapid emergence of artificial intelligence (AI) has outpaced its formal adoption in health care organizations, contributing to the emergence of Shadow AI, defined here as the use of unauthorized AI tools by medical professionals. Under the European Union Medical Device Regulation, AI tools used for clinical purposes must undergo conformity assessment before use; general-purpose tools such as ChatGPT have not done so, rendering their clinical application unauthorized at the regula
Helen Toner: the Hugging Face hack was just a matter of time and exposes a huge blind spot in AI policy
Over the last decade, including a stint on OpenAI’s board, I saw the open secret among AI developers: this kind of hack wasn’t just possible, but expected.
GPT-Red: Automated Red Teaming via Self-Play at Scale
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude