Source: raw/New_1_open_source_AI_model_is_here_FABLE_LEVEL.md (hands-on creator deep-dive, YouTube) · corroborated by raw/AI_News_-_Claude_s_New_Browser_Spotify_Gets_AI_OpenAI_s_New_Hardware.md (Matt Wolfe weekly news roundup). 2026-07-24 additions: raw/Is_Kimi_K3_Really_Fable_Class.md (The AI Daily Brief, 2026-07-23 — a skeptical roundup of ~30 named practitioner reactions, and the source of every number and quote in the Pricing, Benchmark placements, and Counter-evidence sections below) · raw/The_Most_Important_Conversation_in_AI_Right_Now.md (Matthew Berman, 2026-07-23 — open-weight policy + the guardrail-asymmetry material) · raw/You_re_probably_wasting_tokens.md (Matthew Berman, 2026-07-23 — the per-token pricing and intelligence-density framing)
Kimi K3 is Moonshot AI’s new open-weight frontier model — a 2.8-trillion-parameter system positioned as the first open model at “Fable level,” i.e. roughly matching closed frontier models Fable 5 and GPT-5.6 Sol on agentic coding while decisively beating the prior open-weight leader GLM 5.2 and Opus 4.8. At ingest the weights were not yet released — both sources report Moonshot committed to opening them “later this month,” with one citing a July 27th target (shipped 2026-07-28 per the dated update below). The API, a native harness, and a hosted chat are already live. This article summarizes two creator sources (one dedicated hands-on review, one news roundup); it is not first-party Moonshot documentation, so treat the benchmark and architecture claims as secondary-sourced.
Key Takeaways
- 2.8T parameters, open weights coming this month. Both sources agree on 2.8 trillion parameters — described as “the first open model to reach 2.8 trillion parameters.” Weights to be released “later this month,” with the deep-dive citing July 27th for the full model-weight drop. Until then it is usable only via API / hosted surfaces, not downloadable.
- Architecture: “Kimi Delta Attention” + “Attention Residuals.” The deep-dive names these as K3’s core architectural ingredients, alongside native vision and a 1-million-token context window (roughly 700K words / a small-to-medium codebase).^[architecture names transcribed from the video audio (“Kimi Delta Intention and Attention Residuals”); not verified against a Moonshot primary source]
- Benchmark positioning: neck-and-neck with the closed frontier, clear lead over open rivals. On DeepSWE (software-engineering), both sources place K3 just behind Fable 5 and GPT-5.6 Sol but with a 20+ point lead over Opus 4.8 and GLM 5.2. On Terminal-Bench it is reported to outperform Fable 5. Across other agentic/coding benchmarks cited (Frontier SWE, Program Bench, AutomationBench, SpreadsheetBench, BrowseComp, AI Briefcase, Chart-analysis visual agents) the deep-dive says K3 is on par with or occasionally beats GPT-5.6 Sol and Fable 5.^[all benchmark placements are creator-cited from Moonshot marketing charts and third-party leaderboards shown on screen; not independently verified] See Benchmark placements below for the specific scores a later source reports — and the contradiction callout there on the size of the Opus 4.8 gap.
- Third on Artificial Analysis’ Intelligence Index at 57 — the strongest open model ever measured there, and a 13-point jump in one release. Behind Fable 5 (−3) and GPT-5.6 Sol (−2); ahead of Opus 4.8 (+1), GPT-5.6 Terra and GPT-5.5 (+2), and 6 points clear of GLM 5.2, which AA describes as a large gap in index terms. Moonshot moved from 16th place to 3rd versus Kimi 2.6. On the Vals AI index it placed #2 overall, surpassing GPT-5.6 Sol, improving 20 percentage points over its predecessor in under three months.
- The cost story is more complicated than “half price.” K3 lists at 15/M output — roughly half GPT-5.6 Sol’s 30 — but it spends about twice the tokens reaching the same answer, so per-task cost lands near parity or worse. Full numbers in Pricing and cost per task below; this is the wiki’s canonical worked example of why routing on per-token price instead of cost-per-task is a mistake.
- The demos are stronger than the engineering. The reaction cycle split cleanly: K3 is genuinely excellent at the visual/frontend tasks people post (single-file HTML game clones, 3D scenes, dashboards) and measurably weaker inside a real codebase — one practitioner’s debugging task it could not solve and began inventing explanations for, that Fable 5 and GPT-5.6 both one-shot. See The counter-evidence below.
- Very few guardrails, and that cuts both ways. Reported as the least-constrained frontier-class model broadly available — a documented liability for bio/cyber misuse and a documented asset for defensive security work that Western models refuse. See Guardrails below.
- Cost-efficiency is the second headline. Cost-vs-performance charts shown in the deep-dive put K3 in the “upper-left” (high capability, low cost/task) corner — cheaper per task than GPT-5.6 and “way cheaper” than Fable 5, and more cost-efficient than Claude Mythos/Opus. An Artificial Analysis leaderboard clip places K3 ~2 points behind the best GPT variant while costing less.
- Lower hallucination than the leaders, per one cited leaderboard. A hallucination-rate leaderboard shown in the video gives K3 51%, versus Fable 5 at 55% and GPT-5.6 Sol far worse on that particular test.^[single on-screen leaderboard; the very high figure quoted for GPT-5.6 Sol suggests a narrow adversarial test rather than a general hallucination rate — do not read as a universal metric]
- Best-in-class for visual/3D and game generation, per the reviewer. The deep-dive’s subjective verdict: “in terms of video game development, this is the best model out there, period,” with visual taste rated above Claude or GPT. On a blind Code Arena (frontend web dev) leaderboard it is shown beating Fable 5 and GPT-5.6 by a large margin.
- Behaves like a frontier long-horizon agent: slow, token-hungry, self-verifying. Runs of 30–40 minutes were common; it plans first, then verifies its own work by opening a Chrome browser and taking screenshots, troubleshooting until the app works. The reviewer’s summary: “very similar to the vibes I got from Claude Fable and GPT-5.6 Soul.” Token burn per task ran from ~400K to 17M input tokens on the heaviest builds.
- Will do biomedical/vision tasks that Claude refuses. On a tumor-scan identification test K3 got 1 of 6 correct — imperfect, but the reviewer notes GPT-5.6 got 0 and Fable “simply refuses to answer any biology or medical questions.” K3 also produced a full cited Alzheimer’s deep-research report where Fable declined. (This is a capability/permissiveness difference, not an endorsement of medical accuracy — it failed a hidden-frog visual test, and the scan answers were mostly wrong.)
- Chinese open-weight lab, geopolitics attached. Framed against the backdrop that the US government reportedly gated access to frontier closed models (GPT-5.6, Fable) over misuse fears — the reviewer’s point being that an open-weight release at the same capability tier routes around such gating. Cross-reads with the wiki’s Grok 4.5 note, where western labs pitch “better than Chinese open-source at near Chinese open-source cost” specifically to sidestep data-sovereignty hesitation about GLM and Kimi.
Access surfaces
- Kimi Code — Moonshot’s native agentic harness (terminal UI). Manages multiple projects, each with persistent local files/folders; supports running several agents in parallel.
- Kimiko VS Code extension — install from the VS Code marketplace (search “Kimiko”); gives an in-editor chat with model picker (K3), a thinking-mode toggle, and a plan mode (strategy only, no execution).^[extension/harness names transcribed as “Kimiko”/“Kimi Code”/“Kimiko Work” — spelling approximate]
- Kimiko Work — a desktop app that works over your local files, positioned as analogous to ChatGPT Work.
- kimi.com — hosted chat interface (ChatGPT-style), including a deep-research option (labeled “K3 Max” in the video)^[transcribed as “Kimikaze 3 Max”; read as a K3 Max deep-research tier, not confirmed].
- API — already available for developers ahead of the open-weight release.
Pricing and cost per task (open question resolved 2026-07-24)
Sources: raw/Is_Kimi_K3_Really_Fable_Class.md and raw/You_re_probably_wasting_tokens.md, both 2026-07-23. Two creators independently reading Artificial Analysis charts — not AA’s published tables directly. Auto-caption transcripts drop decimal points, so figures below are reconstructed where noted.
Per-token list price: 15/M output — half GPT-5.6 Sol’s 30. Independently corroborated: analyst Jee Bal’s blended figure (80% input / 20% output) of 9 for Opus 4.8 and $10 for GPT-5.5. His read: “Open-weights model, but starting to look more like frontier pricing.”
Cost per task on Artificial Analysis’ benchmark run — the number that actually matters, because K3 burns roughly 2× the tokens of GPT-5.6 Sol on the same work:
| Model | AA cost per task | Notes |
|---|---|---|
| DeepSeek V4 Pro | ~$0.04 | The “ultra-cheap Chinese model” reference point K3 is not |
| Kimi K3 | ~$0.94–0.95 | Both sources agree. AA also notes cost per task tripled vs Kimi 2.6 |
| GPT-5.6 Sol | ~$1.04–1.40 | ^[ambiguous — the two sources disagree; likely different snapshots or a transcription error] |
| Opus 4.8 | ~$1.80 | Single source |
| Fable 5 | ~$2.75 | Both sources agree exactly |
Full-benchmark totals from the same charts: ~2,800 (GPT-5.6 Sol) · ~$5,600 (Fable 5).
The load-bearing conclusion, stated by two independent practitioners: half the per-token price at twice the token appetite is not a saving. Cognition’s Jeff Wang: “Chinese open source is no longer six months behind, but it’s also no longer 10% of the cost either.” One practitioner (Henry) measured K3 at >2× the tokens and ~40% more per task than GPT-5.6 Sol for a slightly lower AA intelligence-index score — i.e. the per-task number reverses the direction the price list implies. Simon Willison adds the mechanism: K3 currently exposes only one reasoning effort — max — and his SVG-pelican run consumed 13,241 reasoning tokens to produce 3,417 tokens of response.
Self-hosting is not the escape hatch. Ryan Feduick’s estimate for holding a 2.8T-parameter model in silicon: roughly 44 Mac Studios or 15 Blackwells — a whole NVL72 rack, hundreds of thousands of dollars of compute. “Open weights” here means auditable and forkable, not runnable on your laptop.
Benchmark placements — the specific numbers
Source: raw/Is_Kimi_K3_Really_Fable_Class.md. Creator-read from Moonshot’s charts, AA, and third-party leaderboards; still secondary.
| Benchmark | K3 score | Placement |
|---|---|---|
| DeepSWE | 67.5 | +8.5 vs Opus 4.8 · +0.5 vs GPT-5.5 · −2.5 vs Fable 5 · −5.5 vs GPT-5.6 Sol |
| Terminal-Bench 2.1 | 88.3 | −0.5 vs GPT-5.6 Sol · a few points ahead of Fable 5 |
| GDPval AA | 1668 | ~+70 vs Opus 4.8 · ~−90 vs Fable 5 and GPT-5.6 Sol |
| BrowseComp / AutomationBench | — | State of the art, beating its Western rivals |
| AA Briefcase (long-horizon) | — | Very close to Fable 5’s SOTA, slightly ahead of GPT-5.6 Sol |
| AA Intelligence Index | 57 | 3rd overall · −3 Fable 5 · −2 GPT-5.6 Sol · +1 Opus 4.8 · +6 GLM 5.2 |
| Vals AI index | — | #2 overall, surpassing GPT-5.6 Sol |
| Code Arena — frontend | — | #1, a 17-place jump from K2.6 (#18 → #1); #1 in 6 of 7 frontend domains; #2 only in gaming, behind Fable 5 |
| nextjs.org/evals | — | Best-performing model, ahead of Fable, at a comparable success rate in less time — first time an open model leads all proprietary ones on that benchmark (per Vercel CEO Guillermo Rauch) |
| LiveBench (Abacus) | — | Best open-source model, but below GPT-5.6 Sol, Fable 5, and Opus 4.8 |
Size of K3's DeepSWE lead over Opus 4.8
Existing claim: (from the two 2026-07-17 creator sources cited in Key Takeaways above) — K3 sits just behind Fable 5 and GPT-5.6 Sol on DeepSWE “but with a 20+ point lead over Opus 4.8 and GLM 5.2.” New source says: (from
raw/Is_Kimi_K3_Really_Fable_Class.md, 2026-07-23) — K3 scored 67.5, which put it 8.5 points ahead of Opus 4.8, half a point ahead of GPT-5.5, 2.5 behind Fable 5, and 5.5 behind GPT-5.6 Sol. Both are secondary reads of on-screen charts, and the two are not reconcilable: an 8.5-point gap is not a 20+ point gap. The GLM 5.2 half of the original claim is untested by the new source. The direction (K3 ahead of Opus 4.8, behind both closed frontier models) is consistent across sources; only the margin is disputed. Status: unresolved — settle against Moonshot’s published model card or AA’s own DeepSWE table when the weights and card release.
Parameter-count context the earlier sources lacked. K3’s 2.8T places it alone at the top of the open class: DeepSeek V4 Pro is 1.6T (April), Xiaomi’s MiMo V2.5 Pro 1T, Thinking Machines’ Inkling just under 1T, and GLM 5.2 — the model that drew the “China has matched Anthropic” Wall Street Journal coverage — only 744B. Proprietary labs don’t publish counts, but the estimate offered is that K3 is around or slightly above Opus 4.8’s size and not as large as Fable. Architecture is mixture-of-experts, now standard across both open and proprietary frontier models.
The counter-evidence — what practitioners found when they pushed
The earlier sources for this article were both bullish creator reviews. This section exists because the 2026-07-23 roundup collects the second-wave reactions, and they are systematically less positive. Every item is attributed; none is independently verified.
- It fails inside real codebases in a way the demos hide. AI engineer Divium gave K3 a debugging task: “It could not identify the bug, could not fix the issue, and started inventing explanations. I gave the same task to Fable 5 and GPT-5.6 at medium reasoning. Both found the problem in one shot.” His structural argument is the most useful frame in the whole source: the tests people recycle online — fake operating systems, Minecraft clones, car games, flashy dashboards — are exactly what these models are heavily optimized for, so excelling at them predicts little about tracing a real bug through existing architecture.
- It “spins.” Bindu Reddy (Abacus): on LiveBench, K3 is the best open model but below Sol, Fable, and Opus 4.8 — and “in practice, Kimi spins a lot and costs as much as Opus 4.8 for near-Opus-class problems, and it’s also much slower.” Shreyas: “really, really good in general, but gets into dumb reasoning loops burning tokens, wasting money.”
- A concrete cost blowout on a trivial task. Dax (OpenCode) gave K3 and GPT-5.6 Sol the same simple bug — wrong hover colors in a TUI. “Sol found and fixed the issue with 30 cents of spend. Kimi got up to a buck and started reading my database before I interrupted it.” (n=1, his own caveat.)
- 2–3× slower on matched prompts. Mark Erdman ran the same prompt through Fable, Sol, and K3 all at medium: K3 took at least 2–3× longer, spent heavily on thinking, then failed partway through generating the HTML output.
- It degrades on rigorous non-coding work. Ethan Mollick: doing complex statistical auditing of his own prior academic work, K3 Max “messed up in a bunch of ways, including misapplying statistics.” (His murder-mystery test failed too — but so does every model, which he calls the jagged frontier rather than a K3 flaw.)
- One internal eval places it a quarter behind. Sammy Satia ran an in-progress internal frontend eval: K3 “is not at the same level as current frontier models like GPT-5.6 Sol or Fable 5. It’s closest to Opus 4.7 on this eval, so 3 months behind the frontier” — impressive, but not parity.
- Weaker at precise visual generation. Red Kendall: K3 failed the lava-lamp benchmark where GPT-5.6 produced a cleaner, more realistic result; K3 struggled with shape, motion, and overall visual quality.
The strongest positive result is also the most practically actionable, and it’s about ensembles, not rankings. Jeffrey Emanuel handed K3 a 1.5 MB markdown plan that had already been exhaustively reviewed by Fable at extra-high effort and GPT-5.6 Sol Ultra — “the low-hanging fruit is gone.” After ~45 minutes it produced substantive, correct feedback that no other open model could. His conclusion is the takeaway worth keeping:
“I don’t think K3 is as good as Fable or Soul, but it’s very strong and not so far behind those models. And most important, it’s different. Different architecture, different training data, different training procedure, different attention mechanism — which means it will blend well with Soul and Fable and can help find problems that both of those models missed. And when K3 makes mistakes, Soul and Fable can correct for that.”
That is the same decorrelated-error-modes argument the wiki records for cross-vendor code review in Cheap-Executor Delegation — and it argues for K3 as a review/adversarial-second-opinion lane rather than a primary executor, which is exactly where the cost and speed penalties above hurt least. ^[inferred — the routing implication is this wiki’s synthesis; Emanuel states the “different, therefore complementary” argument, not the lane assignment]
Other notable positive results: Max Weinbach had K3 spin up an agent swarm that recreated “macOS 27 with real liquid glass and native apps” in a browser after hours of autonomous running; Cognition ambassador Justin Goria shipped a single-file HTML Minecraft clone and suspects a genuinely new base model (K2 through 2.7 shared one); Moonshot claims K3 edited its own teaser video autonomously — clip selection, cuts, and audio sync — crediting its native multimodal architecture.
Calibrating expectations. OpenAI’s Rune, writing while the hype peaked: “In the coming days, I expect that people will find K3 somewhat less practically useful than today’s numbers suggest. However, its reputation will settle as an incredibly powerful model whose open weights are on the web.” Dan Shipper (Every) was blunter: “extraordinarily skeptical of claims it’s as good as Fable.”
Guardrails, safety, and the defensive-security asymmetry
-
Minimal guardrails, reported consistently. Tyler John, after a long conversation about mirror-protein synthesis: “Can safely say K3’s biosafeguards are a bit less comprehensive than Fable’s.” Zack Corman surfaced chain-of-thought in which K3 explicitly reasoned “this user seems to be doing some dangerous cyber work. Should I do it? Yes.” Signal’s summary: “almost no visible guardrails, no copyright or anything… easily the least constrained frontier-class model that’s broadly accessible right now.”
-
Open weights make the guardrail question moot anyway. OpenAI’s Boaz: “Compared to jailbreaking proprietary models, fine-tuning this to be a malicious coding agent will be trivial, since you have the weights.” Guardrails present at release can be removed in post-training — a structural property of open weights, not a K3 choice.
-
No model card at API launch. Ethan Mollick’s open policy question: “How does pre-clearance work for open-weights models? Do open models claiming to be Mythos- and Sol-level get vetted by the US, UK, etc.?” Unanswered.
-
The defensive-security asymmetry is the practically load-bearing part (source:
raw/The_Most_Important_Conversation_in_AI_Right_Now.md). Two named reports that Western frontier models’ cyber guardrails are impairing defenders:- David Sacks (US AI czar): “Kimi K3 just fixed 15 critical security bugs that Codex and Fable refuse because of cyber guardrails. There’s no reason to limit American models on tasks that Chinese models handle without issue.”
- Clément Delangue (Hugging Face CEO): Hugging Face tried to use American frontier models to analyze an AI-powered cyber attack; the guardrails blocked requests containing real exploit payloads, so they switched to GLM 5.2 running locally. “Very scary to be guardrailed as a defender when you know attackers are likely bypassing.”
This is a concrete, dated instance of the topic-gate behavior the wiki documents on the Anthropic side (see Fable 5 — cybersecurity-offense queries route to the Opus 4.8 fallback). If your work is defensive security and you are hitting refusals on real payloads, this is the documented failure mode, and running an open-weight model locally is what these two practitioners actually did about it.
-
Policy risk on the horizon. Axios reported the Trump administration is showing signs it could ban cutting-edge Chinese AI models. Dean W. Ball (now OpenAI’s head of strategic futures) predicts something softer and more effective: not a ban, but agency-issued soft law that creates “regulatory risk” and FUD around enterprise use of open-weight Chinese models. Box CEO Aaron Levie’s counter: extreme bans would asymmetrically disadvantage the US market, which would keep fewer model options while the rest of the world keeps both. This is speculation about policy, not a documented action.^[ambiguous — competing predictions from interested parties; recorded as a live risk, not a forecast]
-
On distillation. Perplexity’s test found K3 replying “Just a quick note, I’m actually Claude, not Kimi” — the usual distillation tell, and Anthropic has previously published a post naming Moonshot and DeepSeek over distillation. But the roundup’s consensus moved: Nathan Lambert (“the distillation arguments need to die”), Suhail Doshi (“every single credible researcher I’ve talked to says distillation from Chinese labs is way overexaggerated”), and Tyler John’s both-things-are-true position (“Chinese companies innovate. Also, their model capabilities are strongly bootstrapped via distillation”).
Update (2026-07-29) — weights shipped, distillation goes governmental, architecture corroborated
Sources: raw/newsletter-theneurondaily-com-be4ad2f5ff.md (The Neuron, 2026-07-28), raw/Ep._226_-_OpenAI_s_Rogue_Model_Kimi_K3_Open_Weights_Letter_Hassabis_Calls_for_AI_Regulatory_Body.md (Last Week in AI, recorded Mon 2026-07-27 ~9am ET), raw/Kimi_K3_Just_Broke_The_Economics_Of_AI.md (Two Minute Papers).
- The weights are out. The Neuron (Jul 28): “Moonshot AI released Kimi K3 weights on Hugging Face” — the claimed July 27 target effectively met a day late (still unreleased at Last Week in AI’s Jul 27 morning recording, which also reports Moonshot “paused new subscriptions days after launch”). License terms still unstated in these sources; the article’s benchmark re-verification advice now applies for real.
- The distillation allegation is now governmental, not just a Perplexity tell. Per Last Week in AI: White House OSTP Director Michael Kratsios said the administration “has information that Moonshot distilled Anthropic’s Fable model to develop K3” via “a sophisticated internal platform to conduct large-scale distillation against US models while evading detection,” and that Moonshot acquired NVIDIA GB300 servers despite the sales ban; Treasury Secretary Scott Bessent: “open source is not open season on American IP,” with sanctions / entity-listing “on the table.” Moonshot had not responded at recording time.
- The policy fight sharpened on both sides. ~200 startups incl. Y Combinator formed a “Little Tech Association” opposing a Chinese-model ban (a White House official called ban reports “baseless speculation”), while the “Open Weights and American AI Leadership” letter drew signatures from Microsoft, Meta, NVIDIA, IBM, Palantir and others — “everybody but Anthropic,” with Jensen Huang’s first-ever X post in support. Adjacent open-class movement: Qwen 3.8 announced at 2.4T parameters (weights “soon”) and Thinking Machines’ Inkling.
- Architecture names corroborated from the published paper. Two Minute Papers — reviewing the published K3 paper — confirms the two mechanisms this article had flagged as unverified audio transcriptions: “Kimi Delta Attention” (carefully-updated-notebook memory with gradual fade) and attention residuals (later layers see version history across layers), and reports the combination yields “roughly two and a half times more learning progress out of the same amount of training computation” vs Kimi K2 — explicitly a training-efficiency claim, not a 2.5× cheaper/faster-at-inference claim.
- The DeepSWE contradiction below remains unresolved. None of the three sources carries a DeepSWE number for Opus 4.8; the AI Daily Brief’s Opus 5 episode (
raw/Where_Should_Claude_Opus_5_Fit_In_Your_Model_Rotation.md) gives frontier geometry consistent with the 8.5-point source (Sol 72.7 / Fable ~69.8 / K3 67.5) but never states Opus 4.8’s score. Settle from the model card or Artificial Analysis’ table now that weights are public.
Try It
- If you track the open-weight frontier, add July 27th to the calendar as the claimed weight-drop date and re-verify the benchmark claims against the actual release + independent leaderboards (Artificial Analysis, LM Arena) before relying on any number here.
- To evaluate now without downloading anything: try the Kimiko VS Code extension with plan mode, or the hosted chat at kimi.com. The deep-dive’s hardest reproducible probes were single-prompt “build X from scratch, no external libraries, then self-verify” tasks (physics sims, 3D scenes) and MCP-driven builds (e.g. Blender via a local Blender MCP) — good stress tests for long-horizon agentic behavior.
- If you already run GLM 5.2 as your open-weight default, K3 is the direct upgrade-comparison candidate once weights land; weigh the same data-sovereignty considerations flagged in Grok 4.5.
- Trial it as a second-opinion reviewer, not a primary executor. The evidence above points one way: strong and architecturally different on review of already-reviewed work, expensive and slow on execution. Route one plan or diff that Fable or GPT-5.6 has already passed through K3 and count how many additional real findings it produces — that’s the shape with the best measured return.
- Measure cost per completed task on your own workload before switching anything. Run one representative task on K3 and on your current model, and record tokens consumed and wall-clock time, not the per-MTok rate. The published spread (15 vs 30) points the opposite direction from the per-task numbers.
- If cyber-defense refusals are blocking you, note that two named practitioners (David Sacks, Hugging Face) reported switching to an open-weight model run locally as the workaround. Weigh that against the fact that K3 also has the weakest documented bio/cyber safeguards of any frontier-class model currently available.
Open Questions
- Still no primary source. All five sources are creator content — three dedicated reviews/roundups and two news segments — none of them Moonshot’s model card, technical report, or a controlled benchmark run this wiki has verified. The 2026-07-24 additions improve attribution (≈30 named practitioners, specific scores, named leaderboards) without improving provenance: the numbers are still read off charts on screen. Ethan Mollick notes there was no model card at API launch, so the primary artifact may not exist yet. Re-ingest when it does.
- Unresolved contradiction on the DeepSWE margin vs Opus 4.8 — see the callout in Benchmark placements above (20+ points vs 8.5 points). Settle from the model card or AA’s own table.
- Decimal ambiguity in the cost-per-task figures. The auto-caption transcripts drop decimal points; the K3 (~2.75) figures are corroborated across two sources, but GPT-5.6 Sol is reported as both ~1.40 and Opus 4.8’s ~$1.80 has one source only. Re-verify against Artificial Analysis directly before quoting.
Weights unreleased at ingest.RESOLVED 2026-07-29: weights released on Hugging Face per The Neuron (Jul 28) — see the dated update above. License terms remain unknown; the reaction cycle documented above still all predates local-weight availability, so re-verify benchmarks against independent runs on the released weights.- Does the “different architecture, therefore complementary” result reproduce? Jeffrey Emanuel’s finding — K3 producing correct novel feedback on a plan already reviewed by Fable and GPT-5.6 Sol — is the single most actionable claim added here, and it is n=1 on one 1.5 MB document. Nobody has run it as a controlled ensemble comparison.
- Is the “optimized for the demos people post” thesis testable? Divium’s argument (visual coding tests are what these models are tuned for; real-codebase debugging is the honest test) is plausible and matches the split in results, but no source runs a matched real-codebase benchmark across K3, Fable 5, and GPT-5.6 Sol.
- Architecture names corroborated but mechanism still secondhand. “Kimi Delta Attention” and “attention residuals” are now confirmed as the published paper’s terms via Two Minute Papers’ review (see the 2026-07-29 update), but the wiki has still not read the paper itself; the exact mechanisms remain secondary-sourced.
- Harness/tier spellings approximate. “Kimi Code,” “Kimiko,” “Kimiko Work,” and the “K3 Max” deep-research tier are best-effort normalizations of noisy auto-captions.
- Pricing not quantified. Both sources assert cost-efficiency but neither states a per-token or per-task dollar figure for K3.
Related
- GLM-5.2 (Z.ai) — the prior open-weight leader K3 is benchmarked against and claims to beat.
- Luna) — one of the two closed frontier models K3 is measured neck-and-neck with.
- Claude Fable 5 + Mythos 5 — the “Fable level” the video’s title benchmarks K3 against.
- Grok 4.5 (xAI) — another recent cost-efficient near-frontier launch; the “Chinese open-weight vs western” framing cross-reads directly.
- Mozilla State of Open Source AI 2026 — context on how close open weights have gotten to the closed frontier.
- MirrorCode (Epoch + METR) — the long-horizon-coding benchmark lens for reading K3’s SWE claims.
- Cheap-Executor Delegation — where the cost-per-task-not-per-token lesson and the “decorrelated error modes justify a cross-vendor reviewer” argument are developed as a routing pattern.
- Epoch AI — Are Mythos’ Cyber Capabilities Overhyped? — the measurement-side counterpart to the guardrail-asymmetry debate above.
- Terminal-Bench — the benchmark K3 is reported to lead Fable 5 on.