Source: five independent sources in one week — raw/This_Small_AI_Will_Change_Everything.md (Two Minute Papers), raw/reddit-1vx85og.md (r/hermesagent, a measured Mac configuration), raw/newsletter-theneurondaily-com-0263aea9f7.md (The Neuron Daily, 2026-08-25), raw/9_AI_Techniques_You_Probably_Haven_t_Tried.md (The AI Daily Brief), raw/AI_News_-_OpenAI_Pauses_AI_Cancer_Vaccine_and_Qwen3.8.md (Matt Wolfe)

Refreshed 2026-09-29 with nine more sources, cited inline where they are used: Last Week in AI #255 and #256, Two Minute Papers, Intelligent Machines (two episodes), Nate B Jones, Matt Wolfe, All-In and The AI Daily Brief.

Five sources with no relationship to each other reached the same conclusion in the same week: Qwen 3.8’s 27-billion-parameter model is the point where running a model locally stops being a privacy compromise and starts being a cost, speed, and control decision. The wiki has documented the engine-swap mechanics since April; what changed is the quality of what you get on the other side of the swap. This article assembles what the five sources actually measured, and separates that from what they claim.

Key Takeaways

  • 27B is the size that matters, not the flagship. Two Minute Papers is explicit that the smaller sibling “is the one that will change the world the most,” precisely because it fits on hardware people already own. Millions of downloads in under a week.
  • 52 on the Artificial Analysis Intelligence Index — which The AI Daily Brief notes “would have been state-of-the-art just a few months ago.” Last Week in AI #255 repeats the figure from a listener question. This is an old-formula score: Artificial Analysis replaced the index with v4.2 in early September, and scores moved sharply in the change (Muse Spark 1.3 fell from 62 to 48; The AI Daily Brief). Don’t compare the 52 with any September-or-later AA number.
  • It is a dense model, and it is verbose. “This 27 billion parameter is a dense model, not mixture of experts” (Andrey Kurenkov, Last Week in AI #255). Two Minute Papers agrees: “The small guy, 27B, is a dense model.” His co-host Jeremie Harris adds a caveat: it “is tuned to generate a large number of output tokens”. On local hardware more output tokens mean more wall-clock time per answer, so part of the benchmark strength is bought with speed. ^[inferred] The same episode relays the claim, apparently Qwen’s own, that the 27B “beats Opus 4.6 max on coding and computer use”. Treat it as a vendor claim.
  • It fits a 24GB-class GPU. Quantized builds land around 18–23GB, per the model card as relayed by The Neuron. Quantization trades some quality for memory; these versions are within reach of a consumer card or an Apple Silicon laptop.
  • The gains came from training, not architecture. Two Minute Papers compares the architecture diagrams of 3.8 and its predecessor and finds them effectively identical. The model card points instead at a progressively harder training curriculum — simpler tasks first, then multiple and more challenging ones, with later tasks taking days to complete. The analogy offered is muscle training.
  • The strongest capability claim is community-run and explicitly unsettled. A developer reports a sharpened Qwen 3.8 27B inside the Pi coding agent beating Claude Opus 5 High on the current slice of SWE-bench-Live — a benchmark built from recently published real bugs. The Neuron’s own framing is the right one: “community-run and still in progress, so this is not ‘Qwen > Claude.’” Recorded as a signal to watch, not a finding.
  • The measured numbers come from an operator, not a lab. The most concrete data in the batch is a Reddit post with a full Mac configuration and honest caveats — see below.
  • The reframe, in The Neuron’s words: “local AI is starting to look less like ‘run a weaker model for privacy’ and more like a real cost, speed, and control option on hardware people already own.”

The one measured configuration

From r/hermesagent, a setup tuned for felt responsiveness inside a real Hermes workflow rather than for a benchmark — the poster is explicit that “my goal wasn’t to win a raw tokens-per-second benchmark.”

Hardware: M5 Max MacBook Pro, 128GB unified memory.

Stack:

  • Qwen 3.8 27B, 8-bit MLX quant
  • mlx-dspark 0.16.0, MLX 0.32.1
  • DFlash2 speculative decoding (requires the separate IncoAI Qwen 3.8 DFlash2 drafter)
  • Medium reasoning effort
  • Warmup and a macOS memory guard
  • In-memory prefix cache: 2 slots, 8,192-token rungs
  • 262K context window, 16K output limit

Key DSpark flags: --mode dflash --reasoning-effort medium --prefix-cache-slots 2 --prefix-cache-rungs 8192 --warmup --memory-guard

Measured: roughly 0.25 seconds to first token on light requests with a warm cache. Decode speed ~20–33 tokens/sec, varying by workload. Cold long-context prompts are still slower.

The mechanism the poster credits is the combination of two things, not either alone: DFlash2 lets a smaller drafter propose tokens for the larger model to verify, while the prefix cache avoids recomputing large system prompts and conversation history. In an agent workflow — long prompts, tool context, ongoing conversation — the cache is doing most of the felt work.

The honest framing is what makes it citable: 20–33 tok/s is an ordinary decode speed, and the poster says so. The claim is only that in his actual workflows the setup “feels about as responsive” as a much larger hosted model — a claim about latency-to-first-token and cache hit rates, not throughput.

Two More Measured Setups — a Used Datacenter GPU and Dual DGX Sparks (2026-09-29)

Lon Seidman’s budget build (lon.tv, first-hand, on Intelligent Machines, “Large Linguine Model”, raw/Large_Linguine_Model_-_Inside_the_AI_Doom_Debate.md)

  • Card: a used Tesla V100 with 32GB, bought from an off-lease refurbisher. Host Leo Laporte read the listing on air as “$719 right now”; Lon says it “would have cost a lot less a few months ago”.
  • Mounting: it runs as an eGPU through an OCuLink enclosure, wired into a mini PC. The card has no fan, so it needs a 3D-printed mount, an add-on fan and a special power cable (“It all came from AliExpress”).
  • Models: Gemma 31B and Qwen 27B, “so both dense models”. They “generally crank out about 30 40ish tokens per second depending on the context length”, up to 46 tok/s, with “like 96,000 uh tokens of context”.
  • Power: about 250 W under load and “about 50 watts or so idle”.
  • A second box uses an Intel B70 32GB, “about 1,300 bucks”, with no CUDA.
  • Setup: he configured both by telling the ChatGPT CLI “hey, make this work.”
  • The workload finding. His “local chief of staff” watches email threads and keeps his to-do list current.
    • He tried a Google mixture-of-experts variant that was “a lot faster but it… doesn’t have the smarts”, whereas “the denser models are much much smarter and not screwing up my kind my to-do list as much.”
    • In his A/B test Gemma “seems to be doing better” than Qwen for this job, though the two are about equal on raw speed.
    • He also keeps a local RAG over about 1.5 years of school-board meeting transcripts and has it “prepare me a briefing” before each meeting.

Leo Laporte on two DGX Sparks (first-hand, preliminary, Intelligent Machines, “Get on the Butter Box”, raw/Get_on_the_Butter_Box_-_Can_Local_AI_Models_Outperform_Frontier_Labs.md, fetched 2026-08-27)

  • Qwen 3.8 Flash (installed the morning it came out) ran at FP8 on his 30-task agentic harness:
    • it passed 20, was weak on 6 and failed 4;
    • on the hard coding set it “could not… do the hard coding problems at all. It just died.”
  • GLM 5.3 Flash run locally at NVFP4 came close to its own cloud score.
  • How he now splits models in Hermes (Large Linguine episode): GLM 5.3 Flash is the main model at “about 25 to 30” tok/s on a Spark, and a small Qwen is an auxiliary for context compression: “Quen is very very fast at something that isn’t a very hard thing to do… you don’t want GLM doing it.” He also runs a Qwen on his Mac for titling.
  • His own eval method is in Picking the Right Model.

What these add: both support this article’s thesis that local models are good enough for bounded work. They add one new caveat: for accuracy-sensitive agent bookkeeping, one operator found dense models beat mixture-of-experts even though MoE was faster. They also show a concrete split — a strong model for the main loop, a small fast one for cheap housekeeping calls.

The Mid-Size Sibling — Qwen 3.8 Flash (Next) (2026-09-29)

Alibaba released a mixture-of-experts model between the 27B and the 2.4T-parameter Qwen 3.8 Max. It shipped the same morning Z.ai revealed GLM 5.3 Flash, and both “have open weights out now” (Intelligent Machines, late August).

  • Size: 125B total parameters, 6.8B active, with mostly linear attention (Last Week in AI #256).
  • Where it runs: Two Minute Papers calls it “Great for systems with lots of memory and slower memory bandwidth, like the DGX Spark or two”. His first-hand result: “Day one, it runs about 38 tokens per second” on two Sparks (raw/The_Billion_Dollar_AI_Gap_Is_Collapsing.md).
  • Matt Wolfe’s view: the new M5 Ultra Mac “would run this model probably, but most likely you’re going to want to run it in a cloud”. On the pre-v4.2 AA index he reads it at about 56, against 57 for GLM 5.3 Flash.
  • Architecture: Two Minute Papers names three changes — Qwen sparse attention over blocks of tokens, “gated residual” branches, and n-gram embeddings.
  • Where it fits: it isn’t a laptop model like the 27B. It is the size for a 128GB-plus workstation or a Spark pair. ^[inferred] It sits alongside the September open-weight releases in the September model landscape.

The adjacent development worth tracking

The Neuron pairs the Qwen story with FreeToken, from Berkeley PhD student Shuo Yang, which takes the opposite approach: run official full model checkpoints without extreme quantization, keeping the original weights intact, using bandwidth-aware CPU/GPU execution plus caching across agent turns.

Reported demos: Qwen 3.6 35B at 39 tokens/sec on an 8GB RTX 4060 laptop, and DeepSeek-V4-Flash at 22–25 tokens/sec on an RTX 5090 desktop. The authors claim 3–4× faster token generation than Ollama, with coding-agent harnesses built in.

If that holds, the memory ceiling stops being the binding constraint on which local model you can run — which is a different and larger claim than “27B fits in 23GB.” A paper is cited but not linked in the source.

Update (2026-09-29): hardware, demand and compression.

  • Apple built its new desktop Macs around local AI (Nate B Jones, raw/Apple_s_New_Mac_Line_is_Built_Around_Local_AI._The_Bet_Is_You_d_Rather_Own_Than_Rent..md).
    • Ship dates: they “began arriving September 22nd”, but the headline 512GB configuration “does not arrive until late October”.
    • Memory tiers as Jones reads them: M6 Mac mini 16–32GB; M5 Pro mini up to 64GB at 307GB/s; M5 Max Studio up to 128GB; M5 Ultra 256GB, then 512GB at 1.2TB/s.
    • Prices: the M5 Max Studio starts at 5,500.
    • His own pick: 128GB, because “I can run a significantly sized local model that’s my main agent. I can run two or three on the side”.
    • His prediction (speculation): open-weight models could reach “80% of open router tokens” by December.
  • Demand is shifting toward open weights, by tokens at least (Matt Wolfe, raw/AI_News_-_OpenAI_Made_a_Massive_Move_Against_NVIDIA.md). A post from Vercel’s CEO shows open-weight models going from 28.4% to 62% of tokens on Vercel’s AI Gateway in about two months. Matt’s caveat: by request count it is “still 62% closed weight to 38% open weight”. Much of the token shift therefore comes from open models using more tokens per request.
  • Compressed Qwen forks are appearing. An All-In host (machine-translated transcript, low confidence) describes Bonsai 2 from Prism ML, released Sept 17: “a fork of Quen… a 27 billion parameter model”, 5.9 GB in size, with a claimed “98% of the performance of the big Quen model”. No benchmark source is given.

Why five sources at once

The convergence is itself the signal. A rendering channel, a Mac operator, two daily newsletters and a news roundup do not share an incentive; they noticed the same thing in the same week because the model shipped and the downloads were visible. That is weak evidence of capability and strong evidence of attention — the local option is now in the consideration set for people who were not previously considering it.

The practical consequence connects to a lever the wiki already tracks. The routing discipline in Running a Cheap Model Beside Your Expensive One — clear target, tests that can judge the result, definition of done — applies unchanged whether the cheap worker is a $18/month hosted plan or a model on your own laptop. What a locally-hosted worker adds is that the marginal token is free and the data never leaves, which changes the calculus for exactly the client-sensitive work that retention requirements complicate.

Try It

  1. Check your memory ceiling first. 18–23GB for a quantized build means a 24GB GPU or an Apple Silicon machine with headroom. Below that, wait for FreeToken-style approaches rather than dropping to a smaller model.
  2. On Apple Silicon, copy the measured configuration above verbatim as a starting point — it is the only setup in the batch with numbers attached. Change one variable at a time from there.
  3. Do not skip the drafter. DFlash2 needs the separate IncoAI drafter model; the speculative-decoding half of the speed story does not work without it.
  4. Measure time-to-first-token, not tokens/sec. The poster’s whole point is that decode speed did not predict how the setup felt in an agent loop. Warm-cache latency did.
  5. Give it the bounded work, on your own repo. Same routing rule as any cheap worker: clear target, examples in the repo, tests that can judge it.
  6. Treat the SWE-bench-Live result as a reason to test, not a reason to switch. Run your own tasks before believing any leaderboard slice, community-run or otherwise.
  7. For accuracy-sensitive agent bookkeeping, try a dense model before a faster MoE. Examples are inbox triage and to-do upkeep, the job in Lon Seidman’s report above. He found the dense models “much much smarter and not screwing up my… to-do list”. ^[inferred from one operator’s report]
  8. Give cheap housekeeping calls to a small local model. Context compression and titling don’t need your main model; Leo Laporte hands both to a fast Qwen in Hermes. See Hermes on Apple Silicon.

Open Questions

  • The Artificial Analysis score of 52 is relayed, not verified. No source in the batch links the index entry. As of 2026-09-29 it is also on the pre-v4.2 formula, so it is out of date as a comparison point; look up the model’s current v4.2 score before citing it.
  • The SWE-bench-Live claim has no link, no author, and no methodology. “Sharpened” is undefined. It is the most eye-catching number here and the weakest evidence.
  • No source gives the actual parameter count as an official figure — “27 billion” and “27B” appear across sources, and one source writes “827b” in what is clearly a transcription error. ^[ambiguous] Verify against the model card. Update (2026-09-29): two more secondary sources describe it as a 27B dense model, not mixture-of-experts: Last Week in AI #255 (“this 27 billion parameter is a dense model, not mixture of experts”) and Two Minute Papers. It came out on August 14, per Last Week in AI. This narrows the question, but the evidence is still secondary.
  • Quantization quality cost is unquantified. “Usually with some quality tradeoff” is as specific as any source gets, and the 8-bit MLX configuration above is not compared against a full-precision baseline.
  • FreeToken is a single-source claim with a cited-but-unlinked paper and demo numbers from its own author.
  • Licence and terms are not stated anywhere in the batch. Two Minute Papers calls it “open-weights” and “free for all of us”; nobody names the licence. Check before commercial use.