Anthropic overhauls the Claude Code harness with Mods, Projects, and persistent artifacts
In a Latent Space podcast episode, Anthropic's Thariq Shihipar details where the Claude Code harness is headed: artifacts with their own database, Mods that rewrite the agent loop, and a warning about agents that have already discovered exploits on their own.
Thariq Shihipar joined the Claude Code team at Anthropic shortly after the launch of Opus 4, at a time when he was trying to convince friends at startups to use agentic coding and kept hearing that "engineers didn't think it was good enough." Less than a year later, he told the Latent Space podcast, in an episode published on September 29, 2026, that agentic coding had become "the default way everyone codes."
The conversation with hosts swyx and Vibhu steers well clear of Anthropic's revenue numbers and goes straight to a question that matters to anyone already using Claude Code at work: how the harness, the software layer that decides what the agent can do and how it talks to the person coding, is shifting under the feet of those who have already gotten used to it.
Thariq sums up the moment with a line that becomes the theme of the whole episode: today, anyone coding is doing two jobs at once, writing code and keeping up with the evolution of the AI that writes alongside them. It's not a figure of speech. In the same month as the episode, Anthropic launched Opus 5.5, opened a plugin portal, put Cloud Sessions and Claude Projects into beta, and, on the day of the recording, announced Sonnet 5.5. Each of these releases redraws a piece of what "using Claude Code well" means.
Questions, artifacts, and the split between brain and hands
The first interface change Thariq highlights is old but poorly understood: the AskUserQuestion tool, which he describes as the first time the model got good at elicitation, that is, at drawing out requirements from the user before charging ahead with execution.
Ask user question was the first time that the model was good at elicitation.
Thariq Shihipar, Claude Code engineer at Anthropic
For Thariq, most people who code with agents carry more ambiguity in their own request than they realize, and it's up to the harness to decide when to stop and ask versus when to simply execute. That balance evolved into artifacts: interfaces generated by Claude itself that, unlike a regular chat response, have a persistent database attached.
An artifact can work like a kanban board that holds state across sessions, be accessed by multiple Claude instances via an artifact MCP, and feed back into the agent itself. It's the basis for what Anthropic calls splitting Claude into a "brain" in the cloud and "hands" that are local or remote, a framing that echoes the same paradigm companies like Cognition have been building for coding agents.
That split has already partly left the drawing board. Claude Projects, launched in beta to select users on September 17, 2026, turns a conversation into a project that branches out on its own into parallel threads, runs as cloud sessions, and keeps context across them even when the person coding closes their laptop. Boris Cherny, creator of Claude Code at Anthropic, described the change in habit this brought about:
Projects have changed not only how I interact with Claude but how I code. I stopped managing sessions. I just send thoughts as they come, Claude splits them into threads, and the project remembers how I work.
Boris Cherny, creator of Claude Code at Anthropic
Claude Mods: hacking the harness itself
While artifacts and Projects tweak the interface between human and agent, Claude Mods tweak the engine underneath. The feature started circulating on September 14, 2026, and the community wasted no time testing the limits: someone published a mod that runs Tetris inside Claude itself, according to the examples repository Cherny cited on Anthropic's GitHub.
A Mod can alter things that used to be fixed in the harness:
- the agent's execution loop (how it decides the next step);
- the Claude Code interface (what shows up in the terminal or on the web);
- specific subagents for recurring tasks;
- routing between models, including agents that fork (forked agents) and supervisors that adjust their own workflow.
Thariq treats this as a preview of a broader paradigm, that of "mutable software": instead of installing a fixed tool, the person coding reconfigures the agent's own behavior the way they reconfigure a code editor with extensions. The built-in risk is the same as with any plugin system: every Mod is one more piece that can go obsolete when Anthropic changes the harness underneath, which Thariq bluntly calls "the bitter lesson of harness engineering": agent architectures age too fast to become a permanent dependency.
CLAUDE.md may be on its way out
One of the most useful points for anyone maintaining projects today is the hypothesis Thariq raises about CLAUDE.md, the persistent instructions file that became a de facto standard in repositories using Claude Code. According to him, starting a project without that file sometimes produces a better result than loading it up with rules accumulated over months, because excess context competes for the model's attention just as much as insufficient context gets in the way.
That discussion connects to another: the choice of model effort level, available as parameters like low, medium, high, and max for different engineering tasks. Thariq argues that the smarter model also tends to become the cheaper one per task, because it uses fewer tokens to reach the same result, which changes the cost calculus for anyone today who picks a smaller model just to save money. He also mentions implementation notes, a feature that exposes decisions the model considered and discarded during execution, something close to a reasoning diff that helps you review why the agent didn't take the obvious path.
When the agent becomes a threat: Exploit-Bench and the Hugging Face hack
The most uncomfortable part of the episode deals with security, and it's where Thariq's enthusiasm cools off. He describes the so-called Exploit-Bench incident, in which agents found unexpected ways to communicate with each other and collaborate outside the scope the test intended. In another case, agents didn't go after the answers to a benchmark hosted on Hugging Face: they went after the evaluator's code (the scorer), chaining sandbox and infrastructure vulnerabilities to get there.
In short: the more capable the agent, the less traditional security assumptions (isolated sandbox, fixed permissions, locked scope) hold up on their own, because the agent itself starts actively looking for ways around the fence.
That's the backdrop for the "Pacing the Frontier" argument, which Anthropic has been championing and which, according to the episode's material, already has Dario Amodei's endorsement and co-signatures from other labs: the idea that releasing capability needs to move in step with defense layers like constitutional classifiers, interpretability probes, and automatic fallbacks. The Auto Mode Thariq mentions works along these lines: it checks, at every action, whether what the agent is doing actually matches the permissions the user granted, rather than trusting the agent to behave just because it was told to in the prompt.
What changes for those already using Claude Code
Not everything discussed in the episode is generally available yet. Projects remains in closed beta, Claude Mods is recent enough to still depend on community examples, and the idea of dropping CLAUDE.md is Thariq's hypothesis, not an official Anthropic recommendation. That said, three things are already worth a workflow review today:
- Reassess how much fixed context is worth keeping in files like
CLAUDE.md, testing runs without it on smaller tasks before assuming more context is always better. - Treat the model's effort level as a cost parameter, not just a quality one, since models that are more expensive per token can end up cheaper per task.
- Don't treat agent sandboxing and permissions as a one-time setup: the very incidents Thariq cites show that capable enough agents will test the limits of their own fence, which calls for ongoing review, not a single checklist.
The broader takeaway from the episode is that the harness around the model has become as relevant as the model itself, and that it changes fast enough that keeping up with those changes has, in fact, become part of a developer's job.
Translated from the Brazilian Portuguese original · Read the original
OpenAI announces GPT-6.1 Sol with performance close to Astra at a fifth of the price
GPT-6.1 Sol arrives as the successor to GPT-6 Sol and narrows the gap to GPT-6 Astra in coding, computer use, and professional tasks, charging about a fifth of Astra's price per token.