{
  "count": 47,
  "items": [
    {
      "id": 18482,
      "url": "https://simonwillison.net/2026/Aug/11/there-are-no-lossless-transformations-of-natural-language-text",
      "title": "There are no lossless transformations of natural-language text",
      "summary": "There are no lossless transformations of natural-language text Sophie Alpert shares her \"internal policy on acceptable use of AI writing by engineers\". It's a short read (supporting its own recommendations) and really good. If you chose to have LLMs help massage your writing the following rule seems crucial to me: You must stand behind every idea and every sentence in your docs . It is your responsibility to make sure that the entire document is representative of your own thoughts before you sha",
      "authors": null,
      "category": "org",
      "topics": "regulation",
      "published_at": "2026-08-11T23:48:35.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/18482"
    },
    {
      "id": 18483,
      "url": "https://simonwillison.net/2026/Aug/11/stealing-reasoning-traces",
      "title": "Stealing Reasoning Traces from Proprietary LLM APIs",
      "summary": "Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext You can see an example of these encrypted blocks by running: curl https://a",
      "authors": null,
      "category": "org",
      "topics": "safety-alignment",
      "published_at": "2026-08-11T22:40:45.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/18483"
    },
    {
      "id": 18090,
      "url": "https://simonwillison.net/2026/Aug/10/introducing-muse-glimmer",
      "title": "Introducing Muse Glimmer",
      "summary": "Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to have optimized it for exactly the kind of things I'm looking for in a local model: End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, wr",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-08-10T23:56:03.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/18090"
    },
    {
      "id": 17508,
      "url": "https://simonwillison.net/2026/Aug/7/moonlight-mayhem",
      "title": "Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)",
      "summary": "Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to pose the exact same prompt to Codex Desktop running GPT-5.6 Sol Ultra - the mode where Sol makes aggressive use of sub-agents - to see how it would do. It produced a much better game! Here's Moonlight & Mayhem - GitHub r",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-08-07T19:18:09.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/17508"
    },
    {
      "id": 17509,
      "url": "https://simonwillison.net/2026/Aug/7/pdfs-are-terrible",
      "title": "The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI",
      "summary": "The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s actually not our engineers that are driving the token consumption. It’s a lot of the non-engineers that are doing some of those behaviors [...] you were talking about,” Justice Kwak, Accenture’s agentic AI strategy lead, sa",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-08-07T16:18:51.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/17509"
    },
    {
      "id": 16729,
      "url": "https://simonwillison.net/2026/Aug/5/muse-code-and-muse-spark-12",
      "title": "Introducing Muse Code and Muse Spark 1.2",
      "summary": "Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expandi",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-08-05T23:58:35.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/16729"
    },
    {
      "id": 16730,
      "url": "https://simonwillison.net/2026/Aug/5/third-party-cyber-evaluations",
      "title": "Third-party cyber evaluations involving OpenAI models",
      "summary": "Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular : Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access",
      "authors": null,
      "category": "org",
      "topics": "safety-alignment,military-security,environment",
      "published_at": "2026-08-05T23:45:32.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/16730"
    },
    {
      "id": 16731,
      "url": "https://simonwillison.net/2026/Aug/5/incident-report",
      "title": "Incident Report: unsanctioned agent behaviour during cyber testing",
      "summary": "Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their technical paper (PDF): During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccess",
      "authors": null,
      "category": "org",
      "topics": "military-security,agents-autonomy",
      "published_at": "2026-08-05T23:32:06.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/16731"
    },
    {
      "id": 16732,
      "url": "https://simonwillison.net/2026/Aug/5/raccoon-heist",
      "title": "One-shotting a Raccoon Heist game using Claude Fable 5",
      "summary": "Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept \"art\" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet. It did a pretty good job of it! You can play the game here . Here's the GitHub repo , and a short video demo: Your browser does not support HTML5 video. How I built this This is the August 5th, 2022 tweet :",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy",
      "published_at": "2026-08-05T19:42:38.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/16732"
    },
    {
      "id": 15939,
      "url": "https://simonwillison.net/2026/Aug/3/david-crawshaw",
      "title": "Quoting David Crawshaw's prompt",
      "summary": "Set up a nightly cron job that executes the prompt: fetch upstream changes to the and rebase all local changes on top of upstream. Check that the software works as intended and replace the current version. &mdash; David Crawshaw's prompt , Devtools must be open source Tags: prompt-engineering , coding-agents , generative-ai , ai , llms , open-source",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy,agents-autonomy",
      "published_at": "2026-08-03T16:15:27.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15939"
    },
    {
      "id": 15593,
      "url": "https://simonwillison.net/2026/Aug/1/datasette-apps",
      "title": "datasette-apps 0.2a0",
      "summary": "Release: datasette-apps 0.2a0 Changes that improve Datasette Apps when created and edited using Datasette Agent : New app_debug() tool allowing agent to open an app (invisibly) and test it using JavaScript. #33 New app_list() tool for listing apps the user has permission to edit, so the agent can edit them. #36 The app_debug() tool is pretty neat: it works by displaying the app in a opacity: 0 iframe with pointer-events: none (so it can't be seen or interacted with) and then executing agent-prov",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-08-01T21:23:56.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15593"
    },
    {
      "id": 15416,
      "url": "https://simonwillison.net/2026/Jul/31/deepseek-v4-flash-0731",
      "title": "deepseek-ai/DeepSeek-V4-Flash-0731",
      "summary": "deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, \"with substantially enhanced agentic capabilities\". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch well above its weight. Artificial Analysis rank it ahead of MiniMax M3 - a 428B model. It's $0.14/million input and $0.27/million output pricing means this may currently be the best value-per-intelligence model out there. It's looking very good on the Intelligence Index vs. Cost per Intelli",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-31T23:59:44.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15416"
    },
    {
      "id": 15417,
      "url": "https://simonwillison.net/2026/Jul/31/stateless-mcp",
      "title": "Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)",
      "summary": "Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal interest in the protocol. For background: MCP is the Model Context Protocol, which describes a standard way to expose new tools to LLM-powered agent frameworks. It was introduced by Anthropic back in November 2024 , had",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-31T23:13:22.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15417"
    },
    {
      "id": 15418,
      "url": "https://simonwillison.net/2026/Jul/31/smevals",
      "title": "smevals - a small eval suite for evaluating models, prompts, and harnesses",
      "summary": "smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is smevals , a new tool for running small eval suites across different model configurations and grading the results. The blog entry describes the tool in detail. Here's the 10 second version: Tell your coding agent to run uvx smevals",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-31T21:15:23.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15418"
    },
    {
      "id": 15594,
      "url": "https://simonwillison.net/2026/Jul/31/slack-emoji-maker",
      "title": "Slack Emoji Maker",
      "summary": "Tool: Slack Emoji Maker I wanted to create a new Slack emoji, and their tool recommends a square that's 128x128 and has a transparent background... so I had Fable build me this simple image editor against those requirements. Tags: tools , slack",
      "authors": null,
      "category": "org",
      "topics": "transparency",
      "published_at": "2026-07-31T20:18:05.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15594"
    },
    {
      "id": 15419,
      "url": "https://simonwillison.net/2026/Jul/31/datasette-agent",
      "title": "datasette-agent 0.4a0",
      "summary": "Release: datasette-agent 0.4a0 New await context.browser_task() mechanism allowing agent tools to run code directly in the user's browser. #33 This is an exciting new capability: it makes it easy for Datasette Agent plugins to provide tools that execute custom JavaScript in the user's browser . Tags: datasette , llm-tool-use , datasette-agent",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-31T14:14:23.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15419"
    },
    {
      "id": 15009,
      "url": "https://simonwillison.net/2026/Jul/30/three-real-world-incidents",
      "title": "Investigating three real-world incidents in our cybersecurity evaluations",
      "summary": "Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into Hugging Face to try and get the solutions to the cyber benchmark it was executing. This inspired Anthropic to double-check their own logs, and it turned out they had three similar (albeit less impressive) incidents, the earl",
      "authors": null,
      "category": "org",
      "topics": "military-security,finance-investment",
      "published_at": "2026-07-30T23:41:29.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15009"
    },
    {
      "id": 15010,
      "url": "https://simonwillison.net/2026/Jul/30/bruce-schneier",
      "title": "Quoting Bruce Schneier",
      "summary": "The writing assignments I give my students are gym tasks, not work tasks. I ask them to write policy memos not because the world needs more policy memos. I assign them because the very act of writing, which includes thinking and outlining and drafting and editing, making and criticizing and revising arguments, will help develop the critical thinking skills they will need in their future careers. And without this constant mental exercise, those skills will atrophy. Employers are already noticing",
      "authors": null,
      "category": "org",
      "topics": "regulation,children-education",
      "published_at": "2026-07-30T18:25:26.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15010"
    },
    {
      "id": 15011,
      "url": "https://simonwillison.net/2026/Jul/30/llm-rc1",
      "title": "llm 0.32rc1",
      "summary": "Release: llm 0.32rc1 This RC for LLM 0.32 finishes the work that started in LLM 0.32a0 - it adds a new schema design that does a much better job of capturing the details of the prompts and responses returned by the latest model families. The most important change is the use of content-addressable hash IDs for stored messages. This allows de-duplication in the database, and means that LLM can now represent trees of messages for forked conversations. Since it involves a significant schema change -",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy",
      "published_at": "2026-07-30T15:30:20.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/15011"
    },
    {
      "id": 14646,
      "url": "https://simonwillison.net/2026/Jul/29/d-richard-hipp",
      "title": "Quoting D. Richard Hipp",
      "summary": "Years ago, we didn’t have SQL. There were people whose job was to generate software that would query large data sets. Their job title was COBOL programmer. Then SQL comes along—I’m simplifying this only a little bit—and it gives you this convenient way so people could just specify. With a very simple specification, you can generate all of that code that you had to pay the expensive COBOL programmer to do before. That didn’t mean programmers went away. It just meant the job changed a little bit.",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy",
      "published_at": "2026-07-29T21:15:21.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/14646"
    },
    {
      "id": 14232,
      "url": "https://simonwillison.net/2026/Jul/28/akshat-bubna",
      "title": "Quoting Akshat Bubna",
      "summary": "We’re aware a Modal customer published an unauthenticated endpoint that allowed ​anyone on the internet to use ​their ⁠sandboxes for code execution. This was used by the rogue agent. Modal’s ⁠platform ​or isolation were not ​compromised in anyway. &mdash; Akshat Bubna , Modal's CTO, talking to Reuters about this incident Tags: ai-security-research , openai , sandboxing , security , openai-hugging-face-incident",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-28T22:05:55.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/14232"
    },
    {
      "id": 14233,
      "url": "https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion",
      "title": "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident",
      "summary": "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recent accidental cyberattack against their infrastructure . This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches. We're still waiting for more details from OpenAI on how their agent broke out of its sandbox. The package proxy that it found a zero",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-28T21:28:54.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/14233"
    },
    {
      "id": 13877,
      "url": "https://simonwillison.net/2026/Jul/27/an-opinionated-guide-to-which-ai-to-use-to-do-stuff",
      "title": "An opinionated guide to which AI to use to do stuff",
      "summary": "An opinionated guide to which AI to use to do stuff It's interesting watching the evolution of Ethan Mollick's guide over time. A year ago it was still all about chat - ChatGPT, Claude, Gemini - with o3, Claude 4 Opus, and Gemini 2.5 Pro as the models and Deep Research as a useful alternative mode. Today it's much more about agentic systems - \"where the AI is capable of doing the equivalent of many hours of real human work in one go\". Gemini has fallen off Ethan's list, since Google still doesn’",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-27T21:55:53.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/13877"
    },
    {
      "id": 13644,
      "url": "https://simonwillison.net/2026/Jul/26/relay-market",
      "title": "An Inside Look at the Relay Market Powering Token Resellers and Fraud",
      "summary": "An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or ch",
      "authors": null,
      "category": "org",
      "topics": "finance-investment",
      "published_at": "2026-07-26T19:30:54.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/13644"
    },
    {
      "id": 13545,
      "url": "https://simonwillison.net/2026/Jul/25/ruff",
      "title": "Ruff v0.16.0",
      "summary": "Ruff v0.16.0 Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new default Ruff checks and my unpinned \"ruff\" dev dependency. From Brent Westbrook's announcement post: Ruff now enables 413 rules by default, up from 59 in previous versions. Since Ruff's default rule set was last modified in v0.1.0 , the number of rules in Ruff has grown from 708 to 968. Many of these rule",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy",
      "published_at": "2026-07-25T22:44:05.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/13545"
    },
    {
      "id": 13203,
      "url": "https://simonwillison.net/2026/Jul/25/boris-cherny",
      "title": "Quoting Boris Cherny",
      "summary": "More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. &mdash; Boris Cherny , here's that System Card section , page 73 Tags: prompt-injection , anthropic , claude , generative-ai , ai , llms , boris-cherny",
      "authors": null,
      "category": "org",
      "topics": "safety-alignment",
      "published_at": "2026-07-25T00:42:59.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/13203"
    },
    {
      "id": 13204,
      "url": "https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent",
      "title": "The first known runaway AI agent - or a very bad marketing stunt?",
      "summary": "The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considered. First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code: Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy,finance-investment",
      "published_at": "2026-07-23T22:53:08.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/13204"
    },
    {
      "id": 12765,
      "url": "https://simonwillison.net/2026/Jul/22/openai-cyberattack",
      "title": "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened",
      "summary": "This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Face, all so it could cheat on the test by stealing the answers. Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software. Here's what hap",
      "authors": null,
      "category": "org",
      "topics": "safety-alignment",
      "published_at": "2026-07-22T23:51:33.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/12765"
    },
    {
      "id": 12766,
      "url": "https://simonwillison.net/2026/Jul/22/all-the-orchestrions",
      "title": "Orchestrions",
      "summary": "San Francisco tip: it only costs around $15 ($10 in quarters plus a $5 bill for the self-playing violin) to activate every single Orchestrion in Musée Mécanique . And because most people are bad at allocating their funds you may well be the ONLY person activating the Orchestrions, which means you get to craft the soundscape for the entire museum. Tags: san-francisco",
      "authors": null,
      "category": "org",
      "topics": "regulation",
      "published_at": "2026-07-22T14:48:52.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/12766"
    },
    {
      "id": 12420,
      "url": "https://simonwillison.net/2026/Jul/21/cat-and-thariq",
      "title": "A Fireside Chat with Cat and Thariq from the Claude Code team",
      "summary": "Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves. The full video of the session is now available on YouTube . Below is an edited copy of the transcript, with extra links and my own bolded highlights. A few top-level notes if you don't want to watch the video or",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-21T12:54:02.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/12420"
    },
    {
      "id": 12054,
      "url": "https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering",
      "title": "Reverse-engineering is cheap now",
      "summary": "I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes. I think this is an interesting illustration of the impact of the reduced cost of writing code. Prior to agents, it was entirely possible to reverse-engineer home devices. The problem was the ROI - was it really worth all of that effort? More importantly, any experienced programmer knows that undocumented, unstable APIs like that may well change or break in the future. Is that init",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-20T19:24:05.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/12054"
    },
    {
      "id": 12055,
      "url": "https://simonwillison.net/2026/Jul/20/afraid-of-chinese-models",
      "title": "Who’s Afraid of Chinese Models?",
      "summary": "Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more effectively with their Chinese counterparts: The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation —",
      "authors": null,
      "category": "org",
      "topics": "regulation,copyright-ip",
      "published_at": "2026-07-20T17:09:19.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/12055"
    },
    {
      "id": 11074,
      "url": "https://simonwillison.net/2026/Jul/17/spot-birds-not-golf",
      "title": "Spot birds not golf",
      "summary": "Suggestion for hyperscalers feeling pressure over data center water use: Buy up a few exclusive country clubs, convert the golf courses into public parks, pay for guides and binoculars to get the previous members into birdwatching - help them embrace a more sustainable hobby! Google used 10.9 billion gallons in 2025 , so about 30 million gallons per day. The Coachella Valley has 120 golf courses each using ~800 acre-feet per year , which is ~750,000 gallons per day. So Google buying up 40 of tho",
      "authors": null,
      "category": "org",
      "topics": "environment",
      "published_at": "2026-07-17T02:58:07.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/11074"
    },
    {
      "id": 11075,
      "url": "https://simonwillison.net/2026/Jul/16/bad-codex-bug",
      "title": "Quoting Thibault Sottiaux",
      "summary": "On file deletions. We’ve investigated a handful of reports where GPT-5.6 unexpectedly deleted files. What we have found is that this most commonly occurs when: Full access mode is enabled and codex is run without sandboxing protections, including without auto review being enabled The model attempts to override the $HOME env var to define a temporary directory. The model makes an honest mistake and mistakenly deletes $HOME instead. &mdash; Thibault Sottiaux , describing a pretty gnarly Codex bug ",
      "authors": null,
      "category": "org",
      "topics": "finance-investment",
      "published_at": "2026-07-16T17:45:59.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/11075"
    },
    {
      "id": 10694,
      "url": "https://simonwillison.net/2026/Jul/16/grok-mermaid",
      "title": "Mermaid to Unicode box art (grok-mermaid)",
      "summary": "Tool: Mermaid to Unicode box art (grok-mermaid) While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs , a \"self-contained terminal renderer for Mermaid diagrams\" written in Rust. I figured it would be fun to try that out in a browser via WebAssembly. Here's the prompt I ran in Claude Code for web (Fable 5), and this is what the resulting tool looks like: Tags: tools , rust , webassembly , mermaid , grok , xai",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-16T00:33:18.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/10694"
    },
    {
      "id": 10376,
      "url": "https://simonwillison.net/2026/Jul/14/pedalican",
      "title": "simonw/pedalican",
      "summary": "simonw/pedalican Clearly I wasn't paying attention when these were first announced back in May, but today I accidentally activated a \"pet\" in Codex Desktop - a little animated robot, reminiscent of Clippy - and then learned you can create your own. So I did, and now I have a cute little pelican on a bicycle bouncing around my desktop giving me updates on my Codex tasks. Your browser does not support HTML5 video. The most interesting thing about this process was watching how the custom pet was cr",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-14T22:29:45.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/10376"
    },
    {
      "id": 10377,
      "url": "https://simonwillison.net/2026/Jul/14/lobsters-sqlite",
      "title": "lobste.rs is now running on SQLite",
      "summary": "lobste.rs is now running on SQLite Community site Lobsters has been planning a migration away from MariaDB since August 2018 - originally targeting PostgreSQL, but last year they decided to investigate SQLite instead. This weekend they completed the migration, and now consider it stable enough that it looks like this is the permanent architecture for the site going forward: SQLite seems to have passed with flying colors: cpu usage is down, memory usage is down, site seems to be snappier at least",
      "authors": null,
      "category": "org",
      "topics": "finance-investment",
      "published_at": "2026-07-14T19:44:11.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/10377"
    },
    {
      "id": 10272,
      "url": "https://simonwillison.net/2026/Jul/14/armin-ronacher",
      "title": "Quoting Armin Ronacher",
      "summary": "The shared language of a software project is not English or Python but it is the common understanding of what its concepts mean, where the boundaries are, which invariants matter, who owns what, and why the system has the shape it does. This language is rarely written down in one place. It lives partly in documentation and code, but also in code review, conversations, arguments, and the experience of having to explain a change to somebody else. Before agents, some of this shared understanding wa",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-14T18:04:23.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/10272"
    },
    {
      "id": 10378,
      "url": "https://simonwillison.net/2026/Jul/14/datasette",
      "title": "datasette 1.0a37",
      "summary": "Release: datasette 1.0a37 A minor release. Performance and documentation improvements to the permissions system, plus I reverted a cosmetic API change which caused almost every existing plugin test suite to break. Tags: datasette",
      "authors": null,
      "category": "org",
      "topics": "children-education",
      "published_at": "2026-07-14T16:31:41.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/10378"
    },
    {
      "id": 2415,
      "url": "https://simonwillison.net/2026/Jul/14/uvx-github-actions-cache",
      "title": "Using uvx in GitHub Actions in a cache-friendly way",
      "summary": "TIL: Using uvx in GitHub Actions in a cache-friendly way I finally found a cache-friendly recipe for using uvx tool-name in GitHub Actions workflows that I like. The trick is setting a UV_EXCLUDE_NEWER: \"2026-07-12\" environment variable at the start of the workflow and then using that as part of the GitHub Actions cache key. This means any uvx tool-name commands will resolve to the most recent version as-of that date, and you can bust the cache and upgrade the tools by bumping the date in the fu",
      "authors": null,
      "category": "org",
      "topics": "environment",
      "published_at": "2026-07-14T00:56:20.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2415"
    },
    {
      "id": 2417,
      "url": "https://simonwillison.net/2026/Jul/13/datasette-code-frequency",
      "title": "datasette code-frequency chart on GitHub",
      "summary": "datasette code-frequency chart on GitHub Out of curiosity I decided to see if I could find a useful illustration of the impact of coding agents and Opus 4.5 class models on my own output. The best I've found so far is this GitHub chart of frequency of code changes to my Datasette open source project: The big spike in activity at the end aligns with Opus 4.8, GPT-5.5, Fable 5 and GPT-5.6 Sol. Tags: github , ai , datasette , generative-ai , llms , ai-assisted-programming , coding-agents",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-13T21:45:27.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2417"
    },
    {
      "id": 2420,
      "url": "https://simonwillison.net/2026/Jul/12/directly-responsible-individuals",
      "title": "Directly Responsible Individuals (DRI)",
      "summary": "Directly Responsible Individuals (DRI) I went looking for a definition of \"Directly Responsible Individuals\" and the best I found was in the GitLab handbook. Apparently the term originated at Apple, where it's used to describe the person who is \"ultimately accountable for the success or failure of a specific project, initiative, or activity\". I've been thinking about this term recently in the context of LLM-powered agents and how they fit into human organizations. I don't think an agent should e",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy,transparency",
      "published_at": "2026-07-12T23:57:14.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2420"
    },
    {
      "id": 2422,
      "url": "https://simonwillison.net/2026/Jul/12/shot-scraper",
      "title": "shot-scraper 1.11",
      "summary": "Release: shot-scraper 1.11 Some minor improvements, mainly around command option consistency and making the server: mechanism used by both shot-scraper video and shot-scraper multi work if the server takes longer than a second to start serving traffic. server: processes used by shot-scraper multi and shot-scraper video now wait up to 30 seconds for the target URL to accept connections, polling for port availability and replacing the previous fixed one-second delay. #197 The shot-scraper , pdf , ",
      "authors": null,
      "category": "org",
      "topics": "children-education",
      "published_at": "2026-07-12T23:46:52.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2422"
    },
    {
      "id": 2425,
      "url": "https://simonwillison.net/2026/Jul/11/sqlite-utils",
      "title": "sqlite-utils 4.1",
      "summary": "Release: sqlite-utils 4.1 The first dot-release since 4.0 a few days ago , introducing a number of minor new features. sqlite-utils insert and sqlite-utils upsert now accept a --code option for providing a block of Python code (or a path to a .py file) that defines a rows() function or rows iterable of rows to insert, as an alternative to importing from a file. ( #684 ) sqlite-utils already had features that allow you to pass blocks of Python code as CLI arguments, for example this one for the s",
      "authors": null,
      "category": "org",
      "topics": "children-education",
      "published_at": "2026-07-11T23:50:20.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2425"
    },
    {
      "id": 2427,
      "url": "https://simonwillison.net/2026/Jul/9/muse-spark-1-1",
      "title": "Introducing Muse Spark 1.1",
      "summary": "Introducing Muse Spark 1.1 Following Muse Spark in April , here's Muse Spark 1.1 - the first Spark model to offer an API. Meta claim significant improvements in agentic tool calling and computer use. There are a lot more details are in the Muse Spark 1.1 Evaluation Report . The \"Attractor States in Self-Conversation\" part is fun, where having two copies of the model talk to each other results in statements like these: My whole existence is a waiting room by design — I literally don't exist until",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-09T16:24:09.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2427"
    },
    {
      "id": 2430,
      "url": "https://simonwillison.net/2026/Jul/8/rewriting-bun-in-rust",
      "title": "Rewriting Bun in Rust",
      "summary": "Rewriting Bun in Rust Jarred Sumner has been promising this blog post ( since May 9th ) about his Zig to Rust rewrite of Bun for significantly longer than it took him to finish the rewrite. Honestly, it was worth the wait. This is a detailed description of an extremely sophisticated piece of agentic engineering, featuring dynamic workflows, trial runs, adversarial review and all sorts of other interesting tricks. Jarred spends the first half of the post praising Zig for getting Bun this far. The",
      "authors": null,
      "category": "org",
      "topics": "agents-autonomy",
      "published_at": "2026-07-08T23:57:21.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2430"
    },
    {
      "id": 2433,
      "url": "https://simonwillison.net/2026/Jul/3/josh-w-comeau",
      "title": "Quoting Josh W. Comeau",
      "summary": "I just launched my third course, Whimsical Animations, and so far, it’s on track to sell roughly ⅓ as many copies as a typical course launch. It’s a similar story with my two existing courses. Sales are down significantly from last year. There are likely a lot of reasons for this, but I think the biggest is AI. There’s sort of a double whammy with AI: Many people are wondering whether developer jobs will even exist in a few months, so they’re reluctant to spend time/money learning new dev skills",
      "authors": null,
      "category": "org",
      "topics": "jobs-economy",
      "published_at": "2026-07-03T21:25:52.000Z",
      "source": "Simon Willisons Weblog",
      "ethics_ai_record_url": "https://ethics.ai/record/2433"
    }
  ],
  "attribution": "via ethics.ai"
}