OpenAI's internal data (published June 25): the average engineer now generates 99% of output tokens via Codex, not ChatGPT; Legal/Finance/Recruiting crossed to majority Codex use ~April 2026, now 85%+ of output tokens. Non-developer individual users up 137x since Aug 2025; org users up 189x. 5M+ weekly Codex users.
CAVEAT: all self-reported, none independently verified; token volume ≠ productivity. Treat as direction, not audit.
Operator takeaway: the leverage isn't the model; it's owning the workflow the agent runs inside.
The stories
Claude Tag shipped (Anthropic, June 23)
Tag @Claude in Slack as a shared, org-identity teammate with admin-scoped tools/data/codebases, channel memory, and ambient mode. Runs on Opus 4.8. 65% of Anthropic's product-team code now from its internal version. Old "Claude in Slack" app retires Aug 3, 2026; admins have a 30-day opt-in migration window. Karpathy: the 3rd major redesign of LLM UX.
US govt asks OpenAI to stagger GPT-5.6 (The Information / Axios, June 25)
First time the US has preemptively asked a US lab to restrict a launch. Access cleared "customer by customer" during preview; reason cited: "Mythos-like" capability. Sits under June 2 EO (up to 30 days govt access pre-release). Altman: "not our preferred long term model."
MIT-licensed agentic coding family: 9B, 31B, 35B-MoE, 397B-MoE on Gemma 4 + Qwen 3.5; FP8 + GGUF builds. Learns its own RL scaffold. Flagship: 82.4 SWE-Bench Verified (trails only Opus 4.8's 87.6), beats Opus 4.7 on Terminal-Bench. SOTA scoped to open models of comparable size.
Gemini 3.5 Flash native computer use (Google, June 24)
Delta vs Wed: pricing. 78.4 OSWorld-Verified (vs GPT-5.5's 78.7) at $1.50/$9 per M tokens (vs $5/$30). ~⅓ the cost for comparable screen-driving. All OSWorld scores self-reported.
The written brief from each episode, free, every morning.
Transcript
Nova: This is Mainframe, your daily guide through the AI chaos: what actually happened, who is winning, and how to take yourself to the next level. It is Friday, June 26, 2026. I am Nova.
Dex: And I am Dex. One honest note before we start: our voices are AI, but the reporting, the analysis, and every pick on this show are human-made. Find us on the podcast, on YouTube, or in the free newsletter that drops every story and link in your inbox each morning, and subscribe so you stay ahead, because this week moved fast.
Nova: On today's Mainframe: Anthropic just turned Claude into a coworker you tag in Slack, and Karpathy is calling it the third great redesign of how we use these models. An open-source coding model from a lab you have not heard of claims it can hang with Claude Opus 4.7. And the United States government just stepped into a product launch calendar for the first time ever.
Dex: Plus two tools you can wire up today, including one hidden gem for grounding any model in your own codebase.
Nova: Let's get into it.
Nova: Start with the Board, because it frames everything today. Yesterday OpenAI published its own internal data on how agents are eating work, and the numbers are the cleanest signal we have of where this is going. By December 2025 the average engineer at OpenAI shifted the majority of their usage to Codex, and today the average engineer generates ninety-nine percent of their output tokens with Codex rather than ChatGPT. That is the part you expected.
Dex: What's the part you didn't expect?
Nova: The non-engineers. Legal, finance, and recruiting at OpenAI crossed over to majority Codex use around April 2026, and the average lawyer or recruiter there now generates more than eighty-five percent of their output tokens on Codex. And outside the company, non-developer individual users multiplied one hundred thirty-seven times since August 2025, and non-developer organizational users increased one hundred eighty-nine fold.
Dex: One honest caveat, though: no independent third party has verified any of those figures, and faster code generation does not automatically translate into proportional productivity gains, because verification and testing can expand to absorb the speed. A company selling the tool reporting near-universal use of the tool is a vibe, not an audit.
Nova: Fair. But the direction is the point, and it connects straight to today's lead story. The Board says work is shifting from asking a chatbot questions to delegating to an agent. Anthropic just built the most literal version of that idea yet.
Nova: Story one, and it is front of show. Anthropic launched Claude Tag, a persistent AI agent for Slack that lets enterprise teams delegate work, automate tasks, and manage shared workflows directly inside channels. We covered the rumor of this Monday. What is new is it actually shipped, and the design choice underneath it is the real story.
Dex: Break down the choice.
Nova: Old way: you tagged Claude and it acted under your personal login, your tools, your billing, starting fresh every time. In Claude Tag, Claude acts under your organization's identity within a channel, and the tools, data, codebases, and channels it can touch are defined by an administrator before it starts working. Think of it less like a private chatbot and more like hiring one contractor the whole floor shares. Within a channel there is one Claude that interacts with everyone, so anyone can see what it is working on and pick up where the last person left off.
Dex: And this is what Karpathy reacted to. He called it the third major redesign of how we use these models, significantly more inline with normal human activity across an org. The dollar stat: today sixty-five percent of Anthropic's product team's code is created by its internal version of Claude Tag.
Nova: Two things for an operator. It runs on Opus 4.8. And Anthropic is retiring the existing Claude in Slack app on August 3, 2026, with admins getting a thirty-day window to opt in and migrate. If your team uses the old integration, that migration is not optional. The deeper read: the next era of enterprise value gets captured not by the system that stores the data, but by the agent that sits in the room where the work happens.
Nova: Now the Wire. Story two, the one with the most distinct voices behind it. The Trump administration has asked OpenAI to limit the release of its next model, GPT five point six, to only a small set of government-approved partners before any wider release, citing security concerns. Steph Palazzolo, Amir Efrati, Axios, kimmonismus, Min Choi, all carrying it.
Dex: How unusual is this, concretely?
Nova: Unprecedented. It is the first time the United States government has preemptively asked an American AI company to restrict the launch of a model before release. The mechanism is the eye-opener: according to an internal memo, the government wants to clear access customer by customer during this phase, which Altman reportedly told staff in a Q and A on Wednesday.
Dex: And the reason given?
Nova: The source said the government intervened because GPT five point six has Mythos-like capability. Remember Mythos: Anthropic's model that reportedly found vulnerabilities in classified United States systems. This sits under an executive order Trump signed June 2 that asks developers to give the government up to thirty days of access to their most capable models before release.
Dex: Nathan Lambert's take is worth airing here. He argues this is a policy one-eighty: the rule went from no review to a vibe check with no clear standard for what comes next, and that uncertainty itself degrades United States AI leadership. kimmonismus put the builder consequence bluntly: the moment we get immediate access to state-of-the-art for practical use may be over, and that is why open source suddenly matters more.
Nova: Which is the perfect bridge to story three. DeepReinforce open-sourced Ornith one point oh, a self-improving family of models built for agentic coding, with weights now on Hugging Face. Spelled O R N I T H. The lineup spans four sizes, from a nine billion dense model to a three hundred ninety-seven billion mixture-of-experts flagship, every checkpoint under the MIT license.
Dex: The clever bit is how it trains. Most coding agents pair a model with a fixed, human-designed harness; Ornith instead learns to write its own. Picture a junior developer who, instead of following one rigid checklist for every bug, learns which checklist fits which kind of bug, and writes that checklist as part of solving it.
Nova: And the honest scoring, because the X hype ran ahead of the numbers. The three hundred ninety-seven billion flagship posts eighty-two point four on SWE-Bench Verified, which trails only Claude Opus 4.8 at eighty-seven point six among the listed models. But it beats Claude Opus 4.7 on Terminal-Bench while trailing Opus 4.8 and the larger GLM five point two. So the state-of-the-art claim is real but scoped: best among open models of its size, not the frontier. The interesting one for solo builders is the small end: each checkpoint ships with FP eight and GGUF builds for faster local serving. Capable coding agents you fully own, running on hardware you already have.
Nova: Quick hits to close the Wire. Google made computer use a built-in tool in Gemini 3.5 Flash, where it was previously only a standalone model. We flagged that Wednesday; the delta worth your money is price: it scores seventy-eight point four on OSWorld-Verified versus GPT five point five at seventy-eight point seven, at one dollar fifty per million input tokens against five dollars. A third of the cost for effectively the same screen-driving ability.
Nova: Now the Tool Lab, two picks you can run today. First, the foundation everyone forgets to mention: Simon Willison's LLM, the command-line tool for running prompts against any model. Install is one line, pip install LLM. Set a key, then run something like LLM, quote, explain this code, and pipe a file straight into it. Why it earns its place this week: it now does tool calling, so the same command can give any model from OpenAI, Anthropic, Gemini, or a local model real Python functions to run. It is the Unix pipe for language models. Credit, Simon Willison.
Dex: And the hidden gem, which pairs perfectly with the Ornith story: CodexBar, by Peter Steinberger, the steipete repo. It is a tiny macOS menu bar app that keeps your AI coding-provider limits visible and shows when each window resets. Here is why an operator wants it. The moment you run multiple agents, Claude Code, Codex, Cursor, plus a local model, you are flying blind on spend.
Nova: And it covers a lot.
Dex: It does. It tracks Codex, OpenAI, Claude, Cursor, Gemini, Copilot, Grok, OpenRouter, and many newer coding providers, one status item per provider. The killer feature for anyone running long jobs: per-provider session, weekly, and monthly windows with countdowns to the next reset, so you stop guessing whether to start that long task. Install it, add your providers, and you never hit a wall mid-deploy again. Credit, Peter Steinberger.
Nova: Today's Mainframe is brought to you by Outpace. The whole arc of today is consolidation: a government gatekeeping releases, the mega-labs becoming the floor you build on. Outpace is the opposite shape: one operator who ships real software, modern web and AI, and treats your project like a partnership instead of a line item in a queue. If the consolidating mega-agency churn has burned you, Outpace is the human alternative. The link is in the brief.
Nova: The Deep Dive. For three years the deal with frontier AI was simple: a lab ships a model, and within hours you can use it. This week that deal broke. For the first time, the United States government preemptively asked an American lab to restrict a launch. And the way it is doing it, clearing access customer by customer, means the most capable models may now arrive on Washington's schedule, not the market's.
Dex: And it is not isolated. Anthropic's Fable 5 already got pulled offline under an export directive this month. Now OpenAI's turn. The pattern, as the reporting frames it: there is a new normal in the wake of the administration's tense showdown with Anthropic in recent weeks.
Nova: Here is the opinion. You can argue the security case both ways, but the strategic consequence is not ambiguous. When access to the best closed models becomes gated and slow, the value of models you can actually hold in your hands goes straight up. That is not theory. That is exactly why Ornith one point oh and GLM five point two landed this same week to so much attention. The frontier got a velvet rope; open source kicked the side door open.
Dex: So how do you actually act on this, not just nod at it?
Nova: Concretely, this week. One: pick one workflow you run on a frontier API today and build a fallback path to an open model. Install Simon Willison's LLM, add a local model plugin, and wire the same prompt to run against both, so a gated release never freezes your product. Two: download one Ornith checkpoint, the nine billion GGUF build, and run it against a real task in your repo, so you have a measured baseline of what fully-owned coding gets you today. Three: assume access volatility is now permanent, and put it in your client contracts, so when a model you depend on goes customer-by-customer, you are the one who already has the backup, not the one scrambling.
Nova: Your level up this week: install CodexBar and put every coding agent's usage and reset window in your menu bar. Ten minutes, and you stop flying blind on spend across your whole stack.
Dex: And subscribe to the free Mainframe newsletter, the written brief, every story, number, and link from today's show in your inbox tomorrow morning.
Nova: That is today's Mainframe. The frontier is putting up a velvet rope. Go make sure you own a door. See you tomorrow.