LongCat-2.0 opens its WEIGHTS (Meituan, July 5)

mAInframe · July 6, 2026 · 14:13

The frontier stays gated, the open shelf becomes real inventory.

Listen on: Spotify Apple Podcasts Amazon Music YouTube
The Board
The stories

LongCat-2.0 opens its WEIGHTS (Meituan, July 5)

DELTA vs the June 30 API-only reveal: full weights + inference code now MIT-licensed, no restrictions. 1.6T params, ~48B active (33B-56B dynamic), native 1M-token context; integrates with Claude Code, OpenClaw, Hermes. SWE-bench Pro 59.5, edging GPT-5.5's 58.6 (vendor-reported; independent repro lags 1-2 wks). Trained end-to-end on ~50,000 domestic Chinese ASICs, no Nvidia. API: $0.75/$2.95 per M (promo $0.30/$1.20). Decrypt's hands-on: quality ~Sonnet 4.6, behind Fable/Opus 4.8, but quality-per-dollar is hard to argue.

The move: add LongCat as a fallback route behind Claude; run your 3 highest-volume coding tasks through it; log quality gap + cost before trusting it.

Source: x.com · huggingface.co · decrypt.co

Grok Voice Agent Builder promoted (xAI; beta July 1, promo push July 5, via TestingCatalog)

No-code, plain-language brief → live phone agent in ~2 min on a single speech-to-speech model (sub-second latency). Flat $0.05/min + ~$0.01/min for a provisioned number; 80+ voices, cloning, 25+ languages, SIP, MCP. CATCH: beta 403 access errors reported; no published answer to the misheard-instruction/confirm-step failure mode.

The move: point it at your own intake line (bookings/order-status/after-hours), give it docs + one tool, review call recordings before it touches paying customers.

Source: x.ai · www.eesel.ai

Anthropic tightens access enforcement (Digitimes, July 6)

After reports Chinese firms (incl. Ant Group, ByteDance) bypassed geo-restrictions via overseas subsidiaries, cloud, and VPNs, Anthropic hardened enforcement. Builder lesson: the model you depend on can be walled off by policy on either side; an owned, downloadable model is leverage.

Source: www.digitimes.com

Deep dive

The frontier put up a rope; the open shelf walked around it.

Level up

Add the lower-power-subagent instruction to CLAUDE.md and wire one open-weight fallback route behind Claude Code (~10 min). Every run after is cheaper and outage-proof.

Chapters
  1. 0:35The Rundown
  2. 1:27The Board
  3. 5:14The Wire
  4. 5:18Sponsor: Outpace
  5. 8:26Repo Spotlight

Every story, number, and link in your inbox.

The written brief from each episode, free, every morning.

Transcript
0:00Nova: The first trillion-parameter coding model trained end to end on Chinese chips, no Nvidia anywhere, just dropped its full weights under an MIT license, and it beats GPT five point five on real GitHub issues. You can download it today and run it on your own metal. This is Mainframe, your daily guide through the AI chaos: what actually happened, who is winning, and how to take yourself to the next level. It is Monday, July sixth, twenty twenty-six. I am Nova.
0:33Dex: And I am Dex.
0:35Nova: On today's Mainframe: LongCat two point zero opens its weights, and the open shelf just went from insurance policy to real inventory. GPT five point six is still behind the velvet rope while Sam Altman teases new math, so we sort the signal from the hype. Grok ships a no-code voice-agent builder that answers your phone in two minutes. And in the Tool Lab, a fresh Claude Code setting plus a Simon Willison trick that quietly shrinks your token bill. Let's get into it.
1:14Dex: One note before we dig in: our voices are AI, but the reporting, the numbers, and the calls we make are human-made. Nova, set the table.
1:27Nova: Here is the board. For three straight weeks the story has been the same: the best closed models arrive on Washington's schedule, not yours. GPT five point six Sol, Terra, and Luna are still preview-only. During the preview, Sol, Terra, and Luna are available through the OpenAI API and Codex to a limited group of trusted partners, GPT five point six is not available in ChatGPT, and OpenAI has not announced a general-availability date. Prediction markets now peg broad access to late July.
2:07Dex: And over the weekend Altman poured gasoline on it. He wrote that his oldest son put two words together for the first time, and he was roughly as amazed by that as by the fact that GPT five point six discovered new mathematical concepts. Carried by kimmonismus, Andrew Curran, Min Choi. The catch?
2:30Nova: The catch is there is no artifact. India Today reported Altman gave no technical detail, and OpenAI's official page does not document the specific new-math claim. So treat it as a tease, not a capability you can use. Here is the through-line for the whole show: while the frontier stays locked behind a guest list, the open shelf is filling up with models that are genuinely close and genuinely cheap. That is the opening for a solo builder. And it lands squarely on today's lead.
3:08Dex: Which is LongCat.
3:10Nova: On June thirtieth Meituan, the Chinese food-delivery giant, revealed LongCat two point zero, but it was API-only. The delta as of yesterday, July fifth, straight from their account: the full weights and inference code are now open. LongCat two point zero is now fully open-source, MIT licensed with no restrictions, one point six trillion parameters, about forty-eight billion active, a one-million-token context, and it integrates directly with Claude Code, OpenClaw, and Hermes Agent. The number that matters: it scores fifty-nine point five on SWE-bench Pro, which edges GPT five point five's fifty-eight point six.
4:00Dex: And the geopolitics are loud. It was trained entirely on a cluster of fifty thousand domestic Chinese chips, no A one hundreds, no H one hundreds, built without the American hardware that export restrictions were designed to keep out.
4:16Nova: Right, and be honest about the ceiling. Vendor-reported numbers, independent reproduction lags a week or two. Decrypt ran it through a game-build test and put the quality visibly behind Claude Fable and Opus four point eight, closer to Sonnet four point six, but the quality-per-dollar is hard to argue with. Because the price is brutal: standard API is seventy-five cents per million input tokens and two dollars ninety-five per million output, well under GPT five point five and Claude Sonnet 5. Now the weights are free, so it is a self-hosting option too. Your move: this week, wire LongCat in as a fallback route behind Claude, run your three highest-volume coding tasks through it, and log the quality gap and the cost. If it clears your bar on grunt work, you just found margin.
5:14Dex: Before the Wire, a word from the people who paid for the coffee.
5:18Nova: This one is brought to you by Outpace. Here is the theme that connects to today: the frontier is gated, the open shelf is exploding, and most business owners have no idea how to turn any of it into a shipped product. Outpace does. It is run by an operator who spent thirty years building real businesses: real estate, hospitality, food, even a theater. Now he points a whole team of AI agents at one client project at a time, strategy to ship. Outpace is a system, not a freelancer and not a faceless mega-agency: the agents do the work, one senior operator directs them, and at the end you own the code and you own the keys. If you have a project sitting in a drawer because the quotes were insane, book a thirty-minute call at outpace dot media. That is outpace dot media.
6:19Dex: The Wire. Grok answered its own phone.
6:23Nova: It did. xAI is now promoting its Voice Agent Builder, flagged by TestingCatalog over the weekend. It is a no-code platform that turns a plain-language brief into a live phone agent in about two minutes, running on a single speech-to-speech model instead of a stitched-together pipeline, which gives it sub-second latency. Pricing is a flat five cents per minute versus OpenAI's token-based billing, plus about a cent a minute for a phone number.
6:56Dex: The honest caveat, per eesel's review: it is a beta, several developers hit four-oh-three not-authorized errors trying to get access, and nobody has published an answer for the classic failure mode, an agent acting on a misheard instruction.
7:13Nova: So the builder move is not to point it at paying customers day one. Point it at your own intake line: bookings, order status, after-hours triage. Give it your docs, wire one tool, and listen to the call recordings before you trust it. For an agency, that is a same-day offer you can sell.
7:35Dex: Second on the Wire, quietly important for anyone building on Chinese open weights: the access game cuts both ways.
7:43Nova: It does. Digitimes reported this morning that Anthropic has tightened enforcement against unauthorized access after reports that Chinese firms bypassed geographic restrictions through overseas subsidiaries, cloud services, and VPNs, with firms including Ant Group and ByteDance reportedly routing engineers to Claude via foreign entities. The lesson for you: the model you depend on can be walled off by policy, on either side of the Pacific. Which is exactly why an owned, downloadable model like LongCat is leverage, not a novelty.
8:26Dex: Tool Lab. Two picks, both about spending fewer tokens without losing your mind.
8:32Nova: First, a fresh Claude Code setting most people missed in the changelog. Claude Code added a fallbackModel setting that lets you configure up to three fallback models tried in order when the primary is overloaded or unavailable, and the dash-dash-fallback-model flag now applies to interactive sessions too. Pair that with organization defaults: admins can now set which model new conversations start with across chat, Cowork, and Claude Code, so routine work does not default to the most expensive option. The move: run claude dash dash version to confirm you are current, then set your fallback chain so a frontier outage or a rate-limit wall does not freeze your build. This is the routing-dial idea baked straight into the tool.
9:25Dex: And the hidden gem is not a repo, it is a one-line technique from Simon Willison, posted July third.
9:32Nova: This is the cheapest thing you will do all week. Willison told Claude, in his words, for all coding tasks use your judgement to decide an appropriate lower power model and run that in a subagent. The effect: substantive implementation gets spawned to a subagent on a cheaper model, trivial edits go even lower, and design, auditing, and anything judgment-heavy stays in the main model. His result was blunt: he is getting a ton of work done and his Fable allowance is shrinking less quickly than before. Drop that one sentence in your CLAUDE dot M D. That is it. You just taught your agent to spend your money like it is its own.
10:17Dex: Deep dive. The frontier put up a rope, and the open shelf walked right around it.
10:24Nova: Here is the full picture, John Oliver style. For three years the deal was simple: a lab ships a model, you use it in hours. That deal is dead. GPT five point six sits behind a government-vetted guest list with no public date. Altman is out here comparing it to his toddler discovering math with zero evidence you can test. And the most capable Claude models spent June getting yanked offline and slowly reissued under Commerce Department conditions. The frontier is now a nightclub with a bouncer.
11:02Dex: And meanwhile, in a Meituan office in Beijing, on chips America tried to keep out of the country.
11:08Nova: Exactly. They trained a one-point-six-trillion-parameter model, ran it anonymously on OpenRouter for two months as "Owl Alpha" to prove it holds up under real load, then dropped the weights for free. A growing chorus of technologists warns these defensive regulatory moves backfired: by locking down Western closed models and driving up API costs, the government left a wide operational window for affordable, high-performance alternatives like LongCat. That is the whole irony. The velvet rope did not slow China down. It made open Chinese weights the rational default for a cost-conscious builder anywhere on Earth.
11:57Dex: So what does a builder actually do with that, without pretending LongCat is Opus?
12:03Nova: Here are the steps. One: this week, stand up a real routing dial. In Claude Code, set your fallbackModel chain, and add LongCat two point zero as a route through OpenRouter or its Anthropic-compatible endpoint. Two: take your three most repetitive coding jobs from last week, the boilerplate, the scaffolds, the mechanical refactors, and run them on LongCat. Log cost and whether the tests still pass. Three: keep the hard reasoning, the architecture calls, the security review, on Claude. Four: write the rule down in CLAUDE dot M D, cheap by default, frontier on purpose, and add Willison's lower-power-subagent line while you are in there. Do not send anything with hashes, keys, or exact IDs to a model you are testing without checking the output byte for byte. You end the week with a documented routing policy that survives the next launch and the next gating headline, because your product is the workflow, not the model behind it.
13:20Dex: Level up.
13:22Nova: One thing this week: open your CLAUDE dot M D and add the lower-power-subagent instruction, then wire one open-weight fallback route behind Claude Code. Ten minutes, and every run after gets cheaper and outage-proof.
13:40Dex: And if you want every story, every number, and every link from today in your inbox tomorrow morning, the Mainframe newsletter is free. It is the written version of this show, no fluff.
13:52Nova: Subscribe, tell one builder who needs it, and we will see you tomorrow. The frontier can put up a rope. It cannot lock the shelf you own. That is Mainframe.
← Jul 5: Artifacts in Claude Code reach Pro/Max (UPDAJul 7: Anthropic ships the J-lens; you can read Cla →