Source: 37 YouTube auto-caption transcripts, all fetched 2026-09-29 unless noted — mainly Matthew Berman (one video each on Gemini 3.8 Flash, Grok 4.7, GLM 5.3 Flash and DeepSeek V4.1 Flash), The AI Daily Brief (Nathaniel Whittemore, “NLW”), Last Week in AI (#255–257), Matt Wolfe’s weekly AI News, Two Minute Papers, Intelligent Machines (TWiT), All-In, Moonshots, The Artificial Intelligence Show and Nate B Jones. Nearly every benchmark below is vendor-reported and relayed by a creator reading a chart aloud; the creator is named each time. Full list in sources:.
September 2026 was the densest release month the wiki has tracked: a Moonshots host counted “12 major model releases” in 22 days. Everything below sits outside the Anthropic and GPT-6 launches (see Claude Opus 5.5 and GPT-6). The pattern is the same across all five labs: a model one step below the frontier, at a small fraction of frontier price, with a benchmark sheet that tests sharply weaker on the harder or newer benchmarks. This article gives each model’s date, availability, price, headline numbers with the person who reported them, and a one-line “consider it when”.
Key Takeaways
- The cheap tier caught up with the frontier on some benchmarks and not on others. Gemini 3.8 Flash, Muse Spark 1.3, DeepSeek V4.1 Flash and Grok 4.7 all post roughly Opus-5-level DeepSWE scores (71–75%) at list prices from 2 per million input tokens. On Terminal-Bench 4.0, released about two weeks earlier, Gemini 3.8 Flash scores 19.1% against Opus 5’s 51.8% (NLW).
- SemiAnalysis calls two of them “benchmaxxed”. Its wording, relayed by NLW: “Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxed models we’ve seen yet”. Its reasoning is that the models score near GPT-6 and Fable 5.1 on the fully public Terminal-Bench 2.1 but “markedly worse” on 4.0.
- Artificial Analysis (AA) changed its index mid-month, so AA numbers from before and after the change can’t be compared. AA “rushed out” Intelligence Index v4.2 the weekend after GPT-6 Astra launched (NLW). Muse Spark 1.3 went from 62 to 48, and Gemini 3.8 Flash from 7th to 12th place. Every AA figure in this article says which index it comes from.
- Compare cost per completed task, not the per-token price. Three cases from this month:
- Gemini 3.8 Flash’s cost per task rose 40% over 3.7 Flash “despite unchanged per token pricing” (AA via NLW).
- Theo says Grok 4.7 is 30–80% less token-efficient than 4.6.
- GLM 5.3 Flash uses about 47K tokens per AA task against GPT-5.6 Luna’s 20K (Berman).
- Most of these models are open-weight. GLM 5.3 (plus Flash), DeepSeek V4.1 Flash and the Qwen 3.8 family can be downloaded. Muse Spark 1.3’s status is unclear: Meta “say that they’re going to make it open weight” (Leo Laporte), while Berman calls it open source.
- Free tiers usually pay for themselves with your data. Two examples:
- Hosts suspected the stealth “Ox Alpha” (GLM 5.3 Flash) of training on prompts while it was free on OpenRouter (AI For Humans). Leo says OpenRouter denied it.
- Muse Spark’s free OpenCode tier is called “contributor” because Meta may train on inputs and outputs (an X user quoted by NLW).
- Consumer agents, not models, got the most attention. Meta’s Muse agent reached #1 on the US free App Store chart, overtaking ChatGPT (NLW), and Amazon soon blocked it from shopping. See the Muse section.
At a Glance
| Model (lab) | Released | Access | List price, $/M tokens in / out | Headline claim (who reported it) | Consider it when |
|---|---|---|---|---|---|
| Gemini 3.8 Flash (Google) | First week of Sept | Google API; a Cyber variant goes only to vetted defenders | 3.75 introductory, then 7.50 after year-end | DeepSWE 73.7%; Terminal-Bench 2.1 89.4% but 4.0 19.1% (Google, via Berman and NLW) | You need speed or many quick drafts, and legal work ^[inferred] |
| Grok 4.7 (SpaceX AI) | Sept 21 | API; Cursor; Grok Build | 6 (same as 4.6) | DeepSWE 71%, GDPval 1695 (xAI, via Berman) | You already use Cursor, Grok Bot or Grok Build and your work fits in 500K context ^[inferred] |
| GLM 5.3 Flash (Z.ai) | Late Aug (stealth as “Ox Alpha”) | Open weights; z.ai (served from China); OpenRouter and other hosts | 0.50 (launch half price 0.25 through Sept 9) | 9¢ per AA task (pre-v4.2); a 320B-total model that one operator runs locally on two DGX Sparks (Berman, Leo Laporte) | You want a strong open-weight default, run on your own or a third-party host ^[inferred] |
| DeepSeek V4.1 Flash (DeepSeek) | Sept 9 | Open weights and paper; DeepSeek API | Input 0.30 peak; output 1.20 | DeepSWE 74.2% (DeepSeek, via Berman); AA 40 on the post-revision index (NLW) | You run high-volume, verifiable work where the price matters most ^[inferred] |
| Muse Spark 1.3 (Meta) | First week of Sept | Meta’s public API; free on OpenRouter and OpenCode for a period; powers the Muse agent | Not stated in any source | DeepSWE 75.4% (Meta); $0.55 per AA task (NLW) | You want cheap agentic coding and can accept Meta’s data terms ^[inferred] |
Gemini 3.8 Flash and 3.7 Flash (Google)
Release and availability
- 3.7 Flash was released three weeks after 3.6 Flash as “the new workhorse model”, “a refinement of 3.6 flash rather than a new base model” (Last Week in AI #255). Its DeepSWE score rose from 49% to 65%, and Automation Bench from 17% to 30% (the benchmark name is garbled in the captions).
- 3.8 Flash shipped in the first week of September, “just a few weeks after 3.7”. Google says it “will work harder than 37”, calling tools iteratively and taking more reasoning steps (NLW). Logan Kilpatrick called it Google’s third Flash update in six weeks.
- Google still has no new Pro model. Last Week in AI notes that 3.5 Pro was “coming soon” in May and again in July. NLW reports rumors that Gemini 4 is close.
- Gemini 3.8 Flash Cyber is “only available to what they call trusted defenders via the Fairwind program” (Berman; program name as captioned).
- Gemini 3.8 Live extended thinking topped the AA speech-to-speech index, “beating out GPT live one Astra and Grok voice think fast 2.0”. It auto-detects 97 languages and hands tool calls off to run in the background (NLW).
Price
- 3.75 output per million tokens. Berman reads Google’s fine print: these “are introductory prices which expire at the end of this year”, after which the rate is 7.50.
- AA found the per-token price unchanged from 3.7 Flash but the cost per task up 40%, “driven by a 30% increase in output tokens per task and more turns” (NLW, quoting AA). AA still called it “the cheapest we’ve measured at this level of intelligence”.
Benchmarks (Google’s own, relayed by Berman and NLW)
| Benchmark | Gemini 3.8 Flash | Compared with |
|---|---|---|
| DeepSWE v1.1 | 73.7% | Opus 5 74%; GPT-5.6 Sol 72.7% |
| Terminal-Bench 2.1 | 89.4% (#1) | “in line with Opus and Sol” |
| Terminal-Bench 4.0 | 19.1% | Opus 5 51.8% |
| GDPval | 1545 Elo | Opus 5 1824; “closer to Sonnet 5 and GPT-5.6 Terra” |
| Harvey legal agent | 61.4% (#1) | 3.7 Flash was #2 |
| Humanity’s Last Exam | 55.9% (#1 on Google’s chart) | — |
| OSWorld | 59% | Opus 5 75% |
| CyberGym (Cyber variant) | 86.2% | GPT-5.5 Cyber 85.6%; GPT-5.6 Sol 83% |
Independent readings
- AA Intelligence Index: 59 and 7th place on the old formula at release. On v4.2 it fell to 12th, “behind GPT56 Terra and GLM 53 Flash” (NLW).
- AA cost per task: $0.58 in the first week of September (Matt Wolfe).
- Speed: “around 20% more tokens per second than runner-up Muse Spark 1.3, and almost four times faster than GLM 53 Flash” (NLW, citing AA).
- 3.7 Flash: took #1 on AA’s Analyst Agent benchmark with a 60% pass rate, against Opus 5 at 54% and Fable 5 at 49%, across 80 tasks in 14 domains (Moonshots EP283, Peter Diamandis reading the slide). Alex Wissner-Gross notes that the metric counts a question as correct only if all five attempts get it right, which rewards reliability.
- Calibration: AI Explained’s own Integrity Bench (built with Pablo Romero) found the Gemini family “wildly overconfident in its abilities”. On the same test the Claude family was “much more calibrated” and the Muse family “the most calibrated of all”.
Field reports
- One user compared it on a game build: Opus 5 won, but “Gemini 38 Flash was 39x faster. Opus 5 took 24 minutes. Gemini 38 Flash took 37 seconds” (relayed by NLW).
- Another user, on the same game built with Kimi K3: “Flash was insanely fast and used way fewer tokens, but the actual game was nowhere close” (also relayed by NLW).
- Zach Lloyd (Warp, on How I AI) replays his own past front-end tasks against other models. His verdict on Gemini 3.7 Flash: “You’re going to take a quality hit.”
Consider it when:
- the task benefits from many fast iterations more than from one polished pass;
- the job is legal analysis, where it leads Harvey’s benchmark;
- a speech agent needs to run live across many languages.
Budget from the 7.50 post-promo price, and check its self-reported confidence before trusting it. ^[inferred] Earlier Flash generations are covered in the Gemini 3.5 Flash field test.
Grok 4.7 and 4.6 (SpaceX AI)
Release and availability
- Grok 4.6 launched August 12, two days before SpaceX AI’s acquisition of Cursor closed (Last Week in AI #255). It kept Grok 4.5’s base model and has a 500K context window. The hosts put it at “maybe… around half the cost” of GPT-5.6 and Claude 5.
- Grok 4.7 launched September 21 (All-In; NLW: SpaceX AI “kicked off what could be a big week for model releases”). It came about a week and a half late. Musk’s explanation, as read by Berman: “We might have penalized response length too much or something in RL”, and it still “gives up on hard tasks that it can do too early”.
- It is available through the API, in Cursor and in Grok Build (Berman reading xAI’s blog). The context window is still 500K, where Berman says “basically every other Frontier model is a million tokens”.
Price: 6 output, “same price and speed” as 4.6. Berman: GPT-5.6 Sol is “more than twice as expensive”, and Fable 5.1 and GPT-6 Astra about five times.
Benchmarks (xAI’s own, as read by Berman and NLW)
- CursorBench 4.0: up six points over 4.6, overtaking GPT-5.6 Sol but five points short of Fable 5.1 (NLW). Berman reads 33% at low effort and 46.3% at extra-high, about half Opus 5’s cost for a result just below it.
- DeepSWE: 71%, against GPT-5.6 Sol 72.7% and Fable 5.1 70%. xAI left Astra out of this table; Berman had Astra rebuild it, which put Astra max at 74.1%.
- Terminal-Bench 4.0: 38%. The Fable 5.1 (57.9%) and Astra (58.2%) scores on the same xAI chart don’t match figures published elsewhere, which put Fable 5.1 at 55.8% and Astra at 57.6–57.9%.
- Knowledge work: GDPval 1695 (Fable 5.1 1735, Astra 1542). AA Briefcase 1657 (Fable 5.1 1678), which puts it ahead of GPT-5.6 Sol.
- Legal: 19.6%, the top score on xAI’s chart (Grok 4.6 had 15.8%).
Independent readings
- AA Intelligence Index, post-revision (v4.2): Berman read it on launch day as 5th, with Grok 4.7 extra-high at 46. NLW, a day later, had it 7th, “behind Astra, two iterations of Fable, Opus, MuSpark 1.3, and GPT5.6 Sol”. ^[ambiguous] AA said its “gains come with higher token usage”.
- Theo: “They claimed it would be more token efficient and it’s less by 30 to 80%… real-world costs come out to more than 2x above Grok 4.6, putting it over Astra’s cost in real-world use.” He called it “a very disappointing release” (via NLW).
- Field reports:
- Leo Laporte: “it actually feels like it got dumber.”
- A rocket-launch render test against Kimi K3 went badly (Bobby, via NLW and Berman).
- Developer Kun Chen used it “for a whole day as my first mate” and called it “a really solid model with visible improvements over 4.5”.
- Grok 4.6 for long jobs: DHH’s first-hand test was a Python-to-Rust library port. Grok 4.6 “completes the task. 10x speed up, same size executable. $55”, where OpenAI’s Luna and DeepSeek V4 Flash failed (Lex Fridman #501).
Consider it when:
- you are already inside Cursor, Grok Build or Grok Bot and want a lower-cost coding or knowledge-work model;
- the task fits in 500K tokens.
Measure the actual cost per task yourself, because the per-token price understates it. ^[inferred] Earlier context: Grok 4.5.
GLM 5.3 and GLM 5.3 Flash (Z.ai)
Release and availability
- GLM 5.3 Flash appeared first as the stealth model “Ox Alpha”. It was free on OpenRouter, which said it had capacity for 100 trillion tokens (Leo Laporte), and AI For Humans relayed that it served “17 trillion tokens… in the last 5 days”.
- In late August Z.ai confirmed that it was GLM 5.3 Flash. Both GLM 5.3 Flash and Qwen 3.8 Flash dropped that morning “and both of them have open weights out now” (Intelligent Machines, fetched 2026-08-27).
- Two Minute Papers also covers a larger GLM 5.3 released alongside it, with open weights.
Specs
- 320B total parameters, 18B active; mixture of experts (Berman; Last Week in AI #256).
- Hybrid attention, “mostly linear” plus full attention (Last Week in AI).
- About half the layers of the previous model, which had 92 (Two Minute Papers).
- 1M context, 131K max output tokens, maximum reasoning by default (Berman).
Price
- 0.50 output (Last Week in AI #256).
- It launched at 0.25, “half its list rates through September 9th” (The Artificial Intelligence Show, ep. 235).
- Cached input is 1¢ per million tokens, and Berman puts it at about one-tenth the price of GLM 5.2.
Benchmarks
- Z.ai’s own, via Berman: Terminal-Bench (older version) 84.3; DeepSWE 63.4; first on GDPval among the models Z.ai compared it with.
- AA Intelligence Index, pre-v4.2 (late August): 57, against Fable 5 at 62 (Berman). A Moonshots panelist on release day: “Fable scores 60… The new GLM model flash that dropped today scores 57 and it is like 100 times cheaper” (EP283).
- AA cost per task: 9¢, against GPT-5.6 Luna max at about 5¢ and Fable 5 at $3.14 (Berman).
- Token use: “one of the most token intensive models out there… at 47,000 tokens” per task, against Luna’s 20,000 (Berman).
Infrastructure: SemiAnalysis reported GLM 5.3 Flash “serving 100 trillion tokens per day on purely Chinese chips” (via Berman). A Moonshots panelist (EP283) described it as running on “these new Huawei chips”. Z.ai separately announced a $5B raise, with 60% going to training, “which includes building a quote fully self-training loop” (NLW).
Field reports
- Leo Laporte, first-hand, 30-task agentic harness: in the cloud it passed 22, was weak on 5 and failed 3. Run locally at NVFP4 on two DGX Sparks it “failed one more but it was weak on one less”. He runs it as his Hermes agent’s main model at about 25–30 tokens per second.
- Christina Warren (GitHub): “if I’m locally, I’m using GLM.”
- Zach Lloyd (Warp), after replaying front-end tasks: “can I be using JLM53 [GLM 5.3] on those instead of using… opus. Yes, definitely.”
- Berman: GLM 5.3 Flash beat GPT-5.6 Sol on most of his website-design prompts and lost on 3D scenes.
- Two Minute Papers: quantized local builds “can start looping like crazy”.
- Theo put GLM 5.3 in D tier on his model tier list (via NLW).
Consider it when: you want an open-weight general default for coding and document work that you can host yourself or through a third-party host in your own jurisdiction. Remember that z.ai’s own endpoint “is being served from China” (Berman). ^[inferred] It fits the task-by-task approach in the Chinese open-weight decision framework. Earlier generations: 5.2.
DeepSeek V4.1 Flash (DeepSeek)
Release and availability: released September 9 (All-In; machine-translated transcript). It has open weights and a full technical paper (Berman; Two Minute Papers), and Two Minute Papers adds that it has “native visual understanding”.
Specs
- 552B-parameter mixture of experts. Berman reads “8 billion active parameters for input and 16 billion for output” from the slide.
- The KV cache is “437 times smaller than V1… also four times smaller than the previous 4.0 flash” (Two Minute Papers). Berman’s version: it needs “a fourth of the HBM” and one-eighth of the SSD storage.
- Berman estimates about 200 tokens per second.
Price (DeepSeek’s API, as read by Berman): input 0.30 peak; output 1.20 peak. NLW quotes the peak rate.
Benchmarks
- DeepSeek’s own: DeepSWE 74.2; CyberGym 88.1; ExploitGym 15, “well below GPT 5.6 soul and Claude Opus 5” (Berman).
- Terminal-Bench 4.0: 31.2%, against GPT-5.6 Sol’s 39.9% (NLW).
- AA Intelligence Index, which NLW names “version 4.3”: 40, against a top score of 53 for Fable 5.1 and GPT-6 Astra. That is “roughly in line with GPT56 Luna and Gemini 38 Flash”. Matt Wolfe notes the previous DeepSeek scored 36.
- AA also found the Flash model “outperforms DeepSeek’s full-size pro model on the intelligence index at a quarter of the cost” (NLW), with a cost per task of 27¢ (Matt Wolfe).
Field reports
- Berman: it failed his Rubik’s-cube simulation, both on deepseek.com and inside Codex: “I haven’t had a model fail the Rubik’s Cube simulation test in a while.”
- Two Minute Papers: “likes to think a lot and burns a lot of tokens”.
- DHH, earlier V4 generation: V4 Flash failed his Rust port. V4 Pro finished it in “2 hours 45, $23”.
Consider it when: the work is high-volume, cheap to check and repeated often, so a failed run costs little: classification, extraction, first drafts. Use off-peak hours for batch work, since the input price halves then. ^[inferred] DeepSeek also released an open-source harness that can rewrite itself, described in an 88-page paper (Two Minute Papers). The wiki has not tested it.
Meta: the Muse Spark 1.3 model and the Muse agent
Muse Spark 1.3 (the model)
- Release: first week of September, “the following day” after Gemini 3.8 Flash (NLW). Meta added a max-effort setting a couple of days later. It is “public on their API” (Last Week in AI #257), and was free for a period on OpenRouter (Matt Wolfe) and on OpenCode’s “contributor” tier.
- Meta’s own benchmarks (via NLW and Matt Wolfe):
- DeepSWE 75.4, against GPT-5.6 Sol 73 and Opus 5 74. At the time DeepSWE’s own leaderboard didn’t list 1.3.
- Terminal-Bench (2.1) 88.8%.
- Alexandr Wang claims “20% fewer tool calls and 25% fewer tokens” than Spark 1.2.
- AA readings (via NLW):
- Pre-revision: 62 on the Intelligence Index, 3rd; and 68 on the coding-agent index at max effort, tied with Opus 5. Fable 5.1 and Astra hadn’t been tested yet.
- Post-revision (v4.2): 48, 5th.
- $0.55 per task, “the most costefficient model at its intelligence level” per AA.
- Pushback:
- SemiAnalysis: “benchmaxed”.
- Wang’s reply: “We don’t claim Muspark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost effective.”
- Matt Wolfe’s own code-generated image test ranked it 20th.
- Open weights? This is unresolved. Leo Laporte (first-hand, using it): “they say that they’re going to make it open weight”. Berman: “Muspark 1.3 open-source”. Meta’s Dina Powell McCormick on All-In: “the very first American open source model just a few weeks ago”, without naming the model.
- Next up: a larger model code-named “Watermelon” is reportedly being prepared for October (The Information, via NLW).
Muse (the agent)
- Launch: Tuesday, September 8, after running under the code name “Hatch” (NLW).
- Where it runs: in the US on iOS, Android and muse.ai, plus a Mac app with computer use (Last Week in AI #257), and through WhatsApp.
- Pricing: Leo Laporte (first-hand): “There’s a free tier… There’s a 100 tier.”
- Security model, per Wang (via NLW):
- “Each Muse runs in its own secure VM, an isolated computer dedicated to you.”
- A separate system, “the Sentinel, checks every action before anything leaves the VM.”
- Muse “never sees your actual passwords or car [card] numbers.”
- Traction: “#1 free app in the US”, overtaking ChatGPT (NLW). A Moonshots host cited 2.8M downloads.
- Commerce fight:
- Amazon blocked Muse from shopping (“Continued access by an unauthorized AI agent violates Amazon’s conditions of use”).
- The next day Shopify announced support for Muse in its Shop Pay agentic checkout on all Shopify stores (NLW). Wider context: agentic commerce.
- Security incidents:
- Meta reported that a poisoned link could give an attacker root on the Muse VM. The attack needed the agent to go looking for the link and the user to approve it, and Meta’s fix “so far is to make the warning message more prominent” (NLW).
- Separately, a Marketplace delegation leaked a seller’s address to a buyer (also NLW).
- What operators got from it:
- Nate B Jones: “It found 5,350 bucks a year that I am spending on subscriptions. So far, it’s canceled 1,285 of that.”
- Leo Laporte had it compile a daily tech-news brief with links, and draft Instagram posts from his Obsidian diary.
- Developer connectors: Meta now accepts connector submissions. According to a founder’s first-hand account of the review form, relayed by Greg Isenberg, you can submit either “an API or an existing MCP server”. Isenberg’s advice before submitting: test awkward requests, such as a time that’s already booked, an expired code, or a repeated request that must not create a second reservation.
Consider it when:
- Model: you want cheap agentic coding tokens and can accept Meta’s terms on free tiers. Test it on your own work because of the benchmaxxing charge.
- Agent: consumer and household admin, provided you set up connections and spending limits deliberately. See What Personal Agents Keep.
- Connector: a service business with a booking API or an MCP server can build one.
Reading Artificial Analysis Numbers This Month
- Before the revision. Early in September, Muse Spark 1.3 scored 62 on AA’s Intelligence Index. That was one point above GPT-6 Astra (61), which had “the same score as GPT56 Soul” and trailed Fable 5.1 by five points (NLW).
- The revision. AA then released v4.2 over the weekend. It gives more weight to agentic tasks through the AA Briefcase test and changes the weighting of existing tests (NLW). NLW later calls the post-revision index “version 4.3” in a separate segment; it is not clear whether that is a further update.
- After the revision:
- Fable 5.1 and Astra tied at the top with 53 (NLW, Berman).
- Muse Spark 1.3 fell to 48.
- Grok 4.7 scored 46.
- DeepSeek V4.1 Flash scored 40.
- What to do with an older number: treat any AA figure dated before about September 5 as a different scale from later ones. That includes the wiki’s earlier “52” for Qwen 3.8 27B.
Other Models Released or Surfaced in September
- Qwen 3.8 Flash (Next) (Alibaba, late August): 125B total / 6.8B active, with open weights. See the Qwen 3.8 local-model article.
- Qwen 3.8 Omni Flash: a 1M-context model aimed at “video editing, music video creation, film production” and real-time conversation. It hadn’t reached OpenRouter when Matt Wolfe recorded.
- Xiaomi MiMo Pro (Sept 22): 309B parameters with open weights. An All-In host says it is “on par with” Opus 5 and GPT-5.6 Sol “across most benchmarks” (machine-translated transcript).
- Bonsai 2 (Prism ML, Sept 17): a 27B Qwen fork at 5.9 GB. The same All-In host claims “98% of the performance of the big Qwen model”.
- A Qwen open image model (Sept 20), which the same host says “outperforms Nano Banana 2”.
- Cognition’s new coding model: a post-trained Kimi K3 that replaces SWE-1.7, with a successor name garbled in the captions. NLW reports 50% on FrontierCode 1.1, “slightly ahead of GPT56 Soul and just behind Fable 5.1”, at a 64% cost reduction. It runs in Devin’s desktop app and CLI. Background: Kimi K3.
- Jev (TypeSafe AI): a “decision model” that picks between options instead of writing text. See Jev.
- “Union Alpha”: a stealth model on OpenRouter that turned out to be a multi-model harness, not a single set of weights. It was named as “Paro 26.9 from the Unbiased company” (spelling unverified; Matt Wolfe).
Try It
- Replay your own history before switching models. This is Zach Lloyd’s method:
- Take 20–30 real past tasks in one category, for example front-end fixes.
- Rerun them on the candidate model.
- Score the results against what actually shipped.
- Feed the results into your routing. As he puts it, “you actually have evidence that okay on your own tasks… the best model configuration for cost and quality is… whatever you find.” See also Picking the Right Model.
- Price the completed task, not the token. Log the total tokens per finished task for each candidate. Gemini 3.8 Flash, Grok 4.7 and GLM 5.3 Flash each shifted meaningfully once token use per task was counted. ^[inferred] See cost and intelligence levers.
- Plug open-weight models into the harness you already use.
- Berman ran DeepSeek V4.1 Flash inside Codex: “You literally just tell Codex, ‘Add this model.’ You give it an API key, and as long as it has a responses API compatible endpoint, you can do it.”
- He ran GLM 5.3 Flash in OpenCode through an OpenAI-compatible endpoint.
- AI For Humans had Claude Code wire up the free OpenRouter endpoint for Ox Alpha on its own.
- Keep client data off free and stealth tiers. Treat any free endpoint as training data unless its terms say otherwise. Prefer a third-party host in your jurisdiction over z.ai or DeepSeek’s first-party APIs for anything sensitive.
- Discount Terminal-Bench 2.1 and early DeepSWE claims. Look for Terminal-Bench 4.0 or an equally new benchmark before believing that a cheap model matches Opus. See Terminal-Bench and DeepSWE.
- If you run a booking-based service business, draft a Muse connector. Wrap your existing API or MCP server and test the awkward cases before submitting.
Open Questions
- Almost every number here is secondary. No primary page from Google, xAI, Z.ai, DeepSeek or Meta was fetched for this article. Benchmarks are as read aloud from charts, and caption errors are possible (for example “DeepSeek” for DeepSWE, and “Paro” for “Pareto”).
- Muse Spark 1.3’s licence. It is unresolved whether the weights are downloadable or only promised. Check Meta’s release page before building on it.
- Grok 4.7’s AA rank is 5th or 7th depending on the day and source, and xAI’s Terminal-Bench 4.0 figures for Fable 5.1 and Astra don’t match other published figures.
- “Version 4.3” of the AA index appears once in NLW’s coverage. It is unclear whether AA published a v4.3 or the host misspoke about v4.2.
- Grok 4.6’s price is only described as “maybe… around half the cost” of the Claude 5 / GPT-5.6 class. No source gives its per-token rate.
- The dates in the All-In list come from a machine-translated transcript that also dates the Muse launch to Sept 22, which conflicts with NLW’s Sept 8. The Sept 9 date for DeepSeek and Sept 21 for Grok 4.7 are corroborated elsewhere; MiMo Pro, Bonsai 2 and the Qwen image model are not.
Related
- Lab Slowdown Warnings, September 2026 — the pacing debate that ran alongside this release wave
- GPT-6 (Astra, Sol, Luna) — OpenAI’s September launches, measured against these models throughout
- Claude Opus 5.5 — Anthropic’s September release and the per-task cost benchmark these models are priced against
- Chinese Open-Weight Models — Decision Framework — how to choose between API, third-party host and self-hosting
- Qwen 3.8 27B — The Local-Model Inflection — the local-hardware side of the same open-weight wave
- Grok Bot — the SpaceX AI agent product that runs on Grok models
- GLM-5.2 — the previous GLM generation
- Grok 4.5 — the previous Grok generation and its Cursor origin