LEAD (builder impact / technique), Cursor's agent swarm rebuilt SQLite from the manual alone (July 20)

mAInframe · July 21, 2026 · 12:34

The biggest copyright settlement in us history just priced your training data.

Listen on: Spotify Apple Podcasts Amazon Music YouTube
The Board
The stories

LEAD (builder impact / technique), Cursor's agent swarm rebuilt SQLite from the manual alone (July 20)

Cursor published "Agent swarms and the new model economics" (codebase on GitHub). A swarm got only SQLite's 835-page manual, no source, no tests, no internet, and produced a working Rust replica that passed 100% of a held-out sqllogictest suite (millions of queries). Architecture: PLANNER agents on the smartest model decompose the goal into a task tree and never write code; WORKER agents on a cheap model execute leaves and never plan; planners can spawn sub-planners (recursive, parallel). THE RECEIPT: Opus 4.8 planner + Composer 2.5 workers = $1,339 total (worker fleet just $411); everything on GPT-5.5 = $10,565; informal all-Fable-5 run = $20,057. Same task, same passing result, 15x cost spread decided by who plans vs who executes. Carried by cursor_ai, Andrew Curran, AlphaSignal, The Decoder.

The move: not "buy Cursor", copy the shape. Frontier model on planning only, cheap models on execution, a held-out test as the judge.

Source: cursor.com · alphasignal.ai · the-decoder.com

BUILDER STAKES (access WIN), Claude Team plan minimum drops from 5 seats to 2

Confirmed in Anthropic's docs; flagged by TestingCatalog. Standard seat $20/mo annual = 1.25x Pro usage; Premium seat $100/mo = 6.25x Pro. Two standard seats = $40/mo: more usage than Pro, cheaper than a single Max 5x seat, plus shared Projects. A straight win, not a stake.

The move: if you + one collaborator both bump Pro's weekly ceiling, the 2-seat Team plan is now your cheapest headroom upgrade. Reprice today.

Source: support.claude.com · x.com

Claude Conway goes dark Friday, July 24, 5pm PT (TestingCatalog)

Conway = Anthropic's always-on remote Claude in a dedicated container (webhooks, connectors, browser). Testers told to export data ("ask Conway: export my data"). Two reads: killed, or graduating to public preview. With Projects + cloud sessions just added to Claude Code, lean preview.

The move: watch anthropic.com/news this week; if you rely on Cowork cloud scheduling, a public always-on container agent is the next tier, be first in line.

Source: www.testingcatalog.com

Deep dive

Decomposition is the new moat.

Level up

Run one planner/worker split in Claude Code on a real task, frontier model plans, a cheap model executes, a held-out test grades before it merges. One habit that outlasts every model on the leaderboard.

Chapters
  1. 0:49The Rundown
  2. 1:38The Board
  3. 2:49Sponsor: Outpace
  4. 7:11Repo Spotlight

Every story, number, and link in your inbox.

The written brief from each episode, free, every morning.

Transcript
0:00Nova: A swarm of AI agents just rebuilt SQLite, say sequel light, from scratch. No source code, no test suite, no internet, just the eight-hundred-thirty-five-page manual. It passed one hundred percent of a held-out test suite. And here is the part that should change how you work this week: the cheapest run beat the most expensive one by fifteen times on cost, for the same passing result. The winner was not a smarter model. It was a smarter way to split the work. This is Mainframe, your daily guide through the AI chaos: what actually happened, who is winning, and how to take yourself to the next level. It is Tuesday, July twenty-first, twenty twenty-six. I am Nova.
0:47Dex: And I am Dex. Here is the rundown.
0:49Nova: On today's Mainframe: Cursor pointed an agent swarm at SQLite's manual and published the bill, and that bill tells you your model spend is mostly a decomposition problem, not a model problem. Claude's Team plan just dropped to a two-person minimum: more usage than Pro, cheaper than a Max five x seat, so if you and one collaborator keep slamming into Pro's ceiling, your cheapest upgrade changed today. Claude Conway, Anthropic's always-on agent, goes dark Friday, and that probably means the public version is close. And the biggest copyright settlement in United States history got signed, quietly setting the price of the data inside every model you touch. Let's get into it.
1:38Nova: One honest note before the Board: our voices here are AI. The reporting, the picks, and the opinions are human-made, ours. Now, the Board, quick. A federal judge approved Anthropic's one point five billion dollar settlement with authors on July twentieth, the largest copyright class settlement in United States history: roughly three thousand dollars per book, across about seven million works pulled from pirate libraries. Dex, why does a lawsuit settlement touch a builder's keyboard?
2:12Dex: Because it puts a real price tag on training data. Licensed data was theoretically expensive before; now it is a line item with a court order behind it. That cost gets baked into the sticker of every future frontier model made the clean way.
2:28Nova: Right. And the open-weight models that trained the same way, offshore, just picked up a structural cost edge. The floor under your bill has two inputs now: what the data cost to license, and how well you structure the work on top of it. Which is exactly today's lead. Keep your options open, do not marry one vendor's number.
2:49Nova: Brought to you by Outpace. Here is a picture that fits today: an agent swarm did the SQLite work, but somebody designed the swarm. That is Outpace. A senior operator who spent thirty years building real businesses, real estate, hospitality, food, even a theater, now aims a whole team of AI agents at one client's project at a time, strategy to ship. Outpace is a system, not a freelancer and not a faceless agency: the agents do the work, one operator directs the fleet, and you own the code and the keys at the end. When frontier capability changes hands every Sunday, you want someone driving the whole fleet, not renting you one seat. Book a thirty-minute call at outpace dot media. That is outpace dot media.
3:43Dex: So let's do the swarm properly. What actually shipped?
3:47Nova: Cursor published a piece called agent swarms and the new model economics on July twentieth, and they put the codebase on GitHub. They gave a swarm of agents nothing but SQLite's manual, eight hundred thirty-five pages, no source, no tests, no internet, and out came a working Rust replica that passed a held-out sequel-logic-test suite, millions of queries with known answers, one hundred percent. Andrew Curran carried it, AlphaSignal and the Decoder ran it down. But the architecture is the gold. Planner agents run on the smartest model, decompose the goal into a tree, and never write a line of code. Worker agents run on a cheap model, pick up a leaf task, execute it, push, and never plan.
4:35Dex: And the cost spread?
4:37Nova: This is the number to tattoo on your monitor. Opus four point eight planning, with Composer two point five doing the work: one thousand three hundred thirty-nine dollars total, and the entire worker fleet was just four hundred eleven of that. Run everything on GPT five point five instead: ten thousand five hundred sixty-five dollars. An informal all-Fable-five run: twenty thousand fifty-seven. Same task. Same passing result. Fifteen times the cost, decided entirely by who you put on planning versus grunt work. The move is not "go buy Cursor." The move is the shape: frontier model on planning only, cheap models on execution, and a held-out test as the judge.
5:29Dex: Second story, the wallet one. The Team plan.
5:32Nova: Anthropic dropped the Claude Team plan minimum from five seats to two. TestingCatalog flagged it, and it is confirmed in Anthropic's own docs. A standard seat is twenty dollars a month on annual billing and gives you one point two five times Pro's usage; a premium seat is one hundred dollars and gives you six point two five times Pro. So two standard seats, forty dollars a month, is more usage than Pro and cheaper than a single Max five x seat, and you get shared Projects on top. This one is a straight win, not a stake. If you and a collaborator both keep bumping Pro's weekly ceiling, the two-seat Team plan is now your cheapest headroom upgrade. Reprice it today.
6:21Dex: And Conway.
6:24Nova: TestingCatalog reports Anthropic is cutting internal access to Claude Conway this Friday, July twenty-fourth, at five PM Pacific, and telling testers to export their data. Conway is the always-on remote Claude that lives in a dedicated container, wakes on webhooks, and drives connectors. Two readings: they killed it, or they are graduating it to public preview. Given Anthropic just added Projects and cloud sessions to Claude Code, I lean hard toward preview. The move: watch anthropic dot com slash news this week. If you have leaned on Cowork's cloud scheduling, a public always-on container agent is the obvious next tier, and you want to be first in line.
7:11Dex: Tool Lab.
7:12Nova: Pick one is a technique, and it is the most valuable thing on the show today: the planner-worker swarm split, worked out and documented by the Cursor team. You do not need their product to run it. In Claude Code, define one subagent as the planner: its whole job is to read the spec, break the goal into a tree of small tasks, and delegate, never to write code. Define worker subagents on a cheaper model, Sonnet five or Kimi K three, whose only job is to complete one leaf and push. Then, and this is the load-bearing part, grade every worker's output against a held-out test the workers never saw while building. That structure is what turned a twenty-thousand-dollar job into a thirteen-hundred-dollar one. The code is on GitHub off the Cursor blog if you want a reference harness to copy.
8:06Nova: Pick two, a hidden gem for anyone who ships visuals: Seedream five point oh Pro on Byte Plus Lumina. Most Western builders slept on this. It is not one-shot image generation; it does editable design. Layer separation splits a poster into text, subject, background, and decorations as separate assets, and you edit a specific region with point, lasso, or box selection instead of regenerating the whole thing, across fourteen languages. Surfaced this week by kimmonismus and Min Choi. First step: drop one marketing asset in, select just the headline, and change it, keep everything else pixel-locked. It replaces a round-trip to a designer for the small stuff.
8:55Dex: Use Cases. Existing tools, run them this week.
8:58Nova: One, swarm on a budget: tonight, take one real feature and run the planner-worker split in Claude Code, frontier plans, cheap model executes, a test suite grades before anything merges. You do not need SQLite-scale ambition to capture the fifteen-times spread. Two, rebuild from the manual, small: pick the one gnarly library function you keep fighting, point an agent at that library's official docs with no source, and have it reimplement just that function graded against your existing tests. You get a dependency-light version you actually understand. Credit the Cursor method. Three, the Team-plan reprice: if you are a solo builder plus one collaborator both maxing Pro, switch to the two-seat Team plan today, forty dollars a month, more usage, shared Projects. Stop paying two Pro seats to hit two ceilings.
9:58Nova: Brought to you by SearchVis, also an Outpace Media product. Mainframe tells you what shipped. SearchVis makes sure the models know what you shipped. When a buyer asks Claude, ChatGPT, Perplexity, or Google's AI Overviews for the best tool in your category, the model names you or it names your competitor, and if it does not name you, you do not exist to that buyer. SearchVis tracks whether the answer engines cite your brand across every engine, tells you why, and hands you the exact move to publish to win the citation. Check where you stand and start free at searchvis dot outpace dot media. That is searchvis dot outpace dot media. Be the answer.
10:44Dex: Deep Dive. Land the one idea.
10:46Nova: Here it is: decomposition is the new moat. Look at what actually moved that bill from twenty thousand dollars to thirteen hundred. It was not a better model. It was an org chart: a planner that never touches code, a worker that never plans, and a grader that says done or not done. Intelligence is the commodity now, six vendors clear the bar that two cleared in June. But the structure you wrap around that intelligence, how you split the task and how you verify it, that is not something you rent from a lab. It is something you own, and it survives every model that ships next Tuesday. Most builders are still shopping for the smartest model. The edge is in becoming the manager who assigns it well.
11:29Nova: So here is what I would actually do. This week, take your single biggest task, and before you write any prompt, write the test that defines "done." Then split planning from execution, frontier on the plan, cheap on the work. You will spend less and trust the output more, because for the first time the agent has to prove it, not just narrate it.
11:52Dex: Close us.
11:53Nova: Level up this week: run one planner-worker split in Claude Code on a real task, frontier model plans, a cheap model executes, and a held-out test grades it before it merges. That single habit outlasts every model on the leaderboard. And if you want every story, number, and link from today in your inbox each morning, subscribe to the free Mainframe newsletter. It is the written brief, no fluff. That is the show. I am Nova.
12:23Dex: And I am Dex. Design the swarm. See you tomorrow.
← Jul 20: LEAD (builder impact), Qwen 3.8 Max Preview Jul 22: LEAD (fresh Anthropic release), RECORD A SKI →