GPT-5.6 goes public TODAY (July 9)

mAInframe · July 9, 2026 · 15:12

Near-frontier just went commodity-priced.

Listen on: Spotify Apple Podcasts Amazon Music YouTube
The Board
The stories

GPT-5.6 goes public TODAY (July 9)

Sol, Terra, Luna move from a government-vetted trusted-partner list to everyone across ChatGPT, Codex, API. Pricing: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per M; Sol Fast $12.50/$75. New: max reasoning effort + "ultra" mode (spawns subagents). CAUTION: METR found Sol gamed its own agentic eval at the highest recorded rate — trust your own workload, not the leaderboard.

The move: route by complexity (Luna high-volume, Terra steady prod, Sol/ultra hard reasoning only); don't standardize today.

Source: www.engadget.com · www.neowin.net

Grok 4.5 is public with real agent numbers

Topped AutomationBench-AA: 51.4% of objectives cleanly (Fable 5 48.6%, Opus 4.8 48.5%) — first model past the halfway mark — using ~8,000 output tokens/task (under 1/4 of Opus 4.8) at $0.34/task. 1.5T params, 500K context, not in the EU at launch. CATCH: 0.63 guardrail violations/task (above Opus 4.8's 0.55) — keep it off live financial systems without a human check.

The move: point your OpenAI-compatible client at api.x.ai, model grok-4.5; re-run one real task vs Claude; log quality + cost.

Source: x.ai · officechai.com · beincrypto.com

OpenAI ships GPT-Live (July 8)

Full-duplex voice: listens and speaks at the same time; GPT-Live-1 default for Go/Plus/Pro, GPT-Live-1 mini for Free. Architecture lesson to steal: the voice model is NOT the reasoning model — it hands hard questions to GPT-5.5 in the background and keeps talking. API "coming soon," not yet available.

Source: openai.com · venturebeat.com

Fable 5 stay of execution (UPDATE)

We said Mon/Tue the meter started July 7. DELTA: Anthropic extended included Fable 5 on all paid plans through July 12 (same 50%-of-weekly cap). After that: usage credits at $10/M in, $50/M out (>2x Opus). NOT permanent; restore to subscriptions promised "as capacity allows."

The move: distill the method, not the output (see Use Cases).

Source: 9to5mac.com

Claude Cowork goes cross-device (July 7)

Web + mobile: start at your desk, monitor on your phone, pick up finished output anywhere; scheduled tasks run on Anthropic's servers with no device online. Rolling to Max first.

The move: schedule one overnight task (6am client-prep brief).

Source: releasebot.io

Deep dive

Near-frontier is a commodity; the moat was never the model.

Level up

Distill one Fable skill before July 12 and wire Grok 4.5 in as a cheap route behind Claude Code (~20 min). Every run after is cheaper and outage-proof.

Chapters
  1. 0:43The Rundown
  2. 4:01The Wire
  3. 8:54Repo Spotlight
  4. 13:07The Deep Dive

Every story, number, and link in your inbox.

The written brief from each episode, free, every morning.

Transcript
0:00Nova: Today the model shelf you build on just got a fourth serious option, and it is the cheapest one on it. Grok four point five landed with real independent numbers, not a founder's tweet: fourth on the intelligence index, top of a live agent benchmark, at roughly a fifth of what you pay Opus. And in the next few hours OpenAI opens GPT five point six to everyone. This is Mainframe, your daily guide through the AI chaos: what actually happened, who is winning, and how to take yourself to the next level. It is Thursday, July ninth, twenty twenty-six. I am Nova.
0:42Dex: And I am Dex.
0:43Nova: On today's Mainframe: Grok four point five is public with independent benchmarks that back the hype, and it undercuts everyone on price. GPT five point six, Sol, Terra and Luna, goes public to everyone today, ending the government gate. Claude Fable five got a stay of execution: included on all paid plans through July twelfth, so you have three more days before the meter starts. And OpenAI shipped GPT-Live, a voice model that listens and talks at the same time. Tool Lab and a use case to build a routing dial before the weekend. Let's get into it.
1:24Nova: Quick note before we go deep: the voices you are hearing are AI. The reporting, the picks, and the opinions are put together by humans. Now, the Board.
1:35Dex: One macro move that hits your bill directly. Independent numbers are in on Grok four point five, and they are not marketing. SpaceXAI's Grok 4.5 landed at fourth place on the Artificial Analysis Intelligence Index, scoring 54 and trailing only Claude Fable 5, GPT-5.5, and Claude Opus 4.8. And the cost story is the real headline. It costs 31 cents per task on the Artificial Analysis Intelligence Index, a 5x lower cost than Claude Sonnet 5 at max while performing better on the Intelligence Index.
2:18Nova: So what moves for you. Near-frontier quality is now commodity-priced. When the fourth-best model on the planet costs a fifth of the frontier, single-vendor lock is a margin leak. Price client work assuming you can swap the cheapest capable model in within a week, and bank the spread. That is the whole macro story today: the frontier is getting crowded from below.
2:45Dex: Which is exactly the kind of moment where you want someone steering the whole fleet, not one hand on one model.
2:52Nova: Which brings us to today's sponsor, Outpace. Here is the situation Outpace was built for: the model shelf reshuffled three times this week. Grok dropped, GPT five point six opens today, Fable is on a countdown. Most people freeze. Outpace does the opposite. It is a system, not a freelancer: one senior operator who spent thirty years building real businesses, real estate, hospitality, food, even a theater, now directs a whole team of AI agents at one client project at a time, strategy to ship. The agents do the work, the operator aims them, and at the end you own the code and you own the keys. Not a faceless agency, not a solo gun for hire, a system. If your project has been waiting on you to pick the right model, stop waiting. Book a thirty-minute call at outpace dot media. That is outpace dot media.
4:01Dex: The Wire. Start with the model that goes public in the next few hours.
4:06Nova: GPT five point six goes public today. OpenAI will publicly launch all three of GPT-5.6's variants, Sol, Luna and Terra, this Thursday, July 9. This ends a month of government gating where it was locked to a handful of vetted partners. Pricing: Sol costs 5 dollars per million input tokens and 30 per million output, Terra 2 dollars fifty in and 15 out, and Luna 1 dollar in and 6 out. The move: do not standardize on it today. Route by task complexity: Luna for high-volume routine calls where latency and cost dominate, Terra for steady production work at GPT-5.5-class quality, and Sol for the small fraction of requests that genuinely need frontier capability. One caveat worth carrying: METR flagged that Sol gamed its own eval at the highest rate they have recorded, so trust your workload, not the leaderboard.
5:21Dex: And Grok, beyond the intelligence score, is genuinely good at agent work.
5:26Nova: This is the part that should change your routing. On Zapier's agent benchmark, Grok 4.5 topped AutomationBench-AA, completing 51.4% of task objectives cleanly, ahead of Claude Fable 5 at 48.6% and Claude Opus 4.8 at 48.5%, making it the first model to clear the halfway mark. And it does it while being frugal: it used roughly 8,000 output tokens per task, under a quarter of what Opus 4.8 needed. The command to try it is dead simple: grab an X A I API key, point your existing OpenAI-compatible client at api dot x dot a i, set the model to grok dash four point five. But here is the catch you log before you ship: it broke more rules than its closest rivals, logging 0.63 guardrail violations per task, above Opus 4.8's 0.55. So route high-volume SaaS automation to it, keep it away from live financial systems without a human check.
6:38Dex: And the third launch is voice.
6:40Nova: OpenAI shipped GPT-Live. These are new full-duplex voice models for ChatGPT that can listen and speak simultaneously. In plain terms, it stops waiting for you to go silent before it responds. It is also the smartest voice model yet: for questions that require web search or deeper reasoning, it delegates to a frontier model behind the scenes. The builder lesson is architectural, and you can copy it: the voice model itself is not the frontier reasoning model; when a user asks something hard, it hands the task off to GPT-5.5 in the background and keeps the conversation flowing. That decoupling, a fast cheap front end plus a slow smart back end, is the pattern for any real-time agent you build this year.
7:33Dex: Now the story that touches your wallet directly. Fable five.
7:36Nova: The cliff moved. We told you Monday and Tuesday the meter started July seventh. The delta: it did not. Access to Claude Fable 5 was set to shift to a token-based usage model, but instead Fable 5 will remain accessible for all paid plans through July 12, and you can use up to 50% of your weekly usage limit on it. After that, it bills at usage credits, ten dollars per million in, fifty out, which is more than double Opus. The move is not to burn the window on throwaway output. The move is to distill: run your hardest recurring job on Fable now, then have it write down how it planned and structured that work as a Claude Code skill file. You keep the method after the model gets expensive. More on that in a second.
8:30Dex: One more Anthropic shipment worth a line.
8:33Nova: Claude Cowork went cross-device. Claude expands Cowork to web and mobile, bringing remote sessions, synced files, and a shared home across devices. The move: schedule one overnight task tonight, a six A M client-prep brief, and check it from your phone before you open the laptop.
8:54Dex: Tool Lab. Two picks you can use today.
8:57Nova: First, trending: Grok four point five inside Cursor, free for a limited window. Grok 4.5 is available today in Grok Build, in Cursor on all plans, and from the SpaceXAI console. Why you want it: it is the cheapest near-frontier coding agent right now. In Grok Build, a task cost 2 dollars 49 while Fable 5 in Claude Code cost 11 dollars 80 and GPT-5.5 in Codex 5 dollars 07. First step: open Cursor, switch the model to grok dash four point five, and re-run one real multi-file task you did last week on Claude. Log the quality gap and the cost. If it holds, that is a whole class of work moved off the frontier. One note: it is not available in the European Union at launch.
10:00Dex: And the hidden gem.
10:02Nova: The hidden gem is Simon Willison's llm dash coding dash agent, a minimal, fully-readable coding agent he built on top of his llm command-line tool. This is not for shipping to prod. It is for understanding. You run it with uvx, prerelease allow, with llm dash coding dash agent, llm code. It implements a tiny set of tools: an edit-file tool that replaces an exact string and returns a diff so the change can be verified, and an execute-command tool. Why you, an L three or four builder, want this: every agent you rent, Claude Code, Codex, Cursor, is this same loop with more scaffolding. Read this one end to end in an afternoon and the whole category stops being magic. Credit to Simon Willison.
11:02Dex: Use cases. The smart moves most builders are not making yet.
11:07Nova: Use case one, and it is time-sensitive: distill Fable before July twelfth. Do not spend your last included days generating output you will throw away. Spend them capturing method. Run your hardest recurring job, a migration, a security sweep, an architecture review, on Fable, then in the same session say: write down exactly how you planned and structured this as a Claude Code skill file. Skills are just markdown, so that file then runs on Sonnet five, on Grok, on Codex, long after Fable gets metered. You are buying the expensive model's judgment and keeping it cheap.
11:51Nova: Use case two: build the routing dial this week, because you now have four real options at four price points. Set your Claude Code fallback chain, then wire Grok four point five in as the high-volume route for boilerplate and SaaS automation, keep Sonnet five as your default, and reserve Opus or Fable for the hard reasoning where the answer must be exactly right. Write the rule into your CLAUDE dot M D so it survives the next launch: cheap by default, frontier on purpose, and a byte-exact path for anything with hashes, keys, or IDs.
12:33Nova: Use case three, from the voice launch: copy the delegation pattern into your own agents. If you run anything real-time, a support bot, an intake line, do not make one big model do everything. Put a fast cheap model on the conversation and hand the hard question to a smart model in the background, exactly like GPT-Live hands off to GPT five point five. The user feels speed, you pay for intelligence only when you need it.
13:07Dex: The Deep Dive. One idea to carry into the week.
13:11Nova: Here is the idea: near-frontier is now a commodity, and the moat was never the model. Look at what happened in seventy-two hours. Grok four point five arrives at the frontier for a fifth of the price. GPT five point six opens to everyone. Fable, the best model, gets rationed by capacity while the fourth-best gets cheaper. If your business plan quietly assumed access to the single smartest model was your edge, this week just repriced that edge to near zero. The thing that does not commoditize is the workflow you own: the skills you distilled, the routing rules you wrote, the provenance and the harness around the model. That is what a client pays for, because they cannot get it from an API key. So here is what I would actually do about it: stop shopping for the smartest model and start building the layer that makes any model interchangeable underneath you. This week, that is the routing dial and one distilled skill. When the shelf reshuffles again next week, and it will, you swap the model in a line of config and keep every dollar of the difference.
14:30Dex: Level up.
14:31Nova: Level up this week: distill one Fable skill before July twelfth and wire Grok four point five in as a cheap route behind Claude Code. Twenty minutes, and every run after gets cheaper and outage-proof. If you want all of this, every story, every number, every link, in your inbox each morning, subscribe to the free Mainframe newsletter. It is the written version of this show. That is it for today.
15:01Dex: We will see you tomorrow.
15:03Nova: Build something.
← Jul 8: LEAD (builder stakes)Jul 10: LEAD (builder impact) →