Source: First-party: TypeSafe AI, “Introducing System One Models and Jev” (Diogo Almeida, typesafe.ai/blog, dated Sep 14, 2026), saved at ai-research/jev-typesafe-system-one-launch-2026-09.md. Commentary and hands-on reports (YouTube auto-captions, all fetched 2026-09-29):

  • Two Minute Papers, “Yes Jev Is Insane But There’s A Catch” (youtube.com/watch?v=qBBRRsH0rQc) — raw/Yes_Jev_Is_Insane_But_There_s_A_Catch.md
  • MattVidPro AI, “New Jev Model Acts in Real Time… Minecraft Broke It” (youtube.com/watch?v=DI3iRGUlczQ) — raw/New_Jev_Model_Acts_in_Real_Time..._Minecraft_Broke_It.md
  • Claire Vo, How I AI, “I’m using Jev more than Opus 5.5 or GPT-6” (youtube.com/watch?v=-KIBgpGA_XI) — raw/I_m_using_Jev_more_than_Opus_5.5_or_GPT-6._Here_s_why..md
  • NLW, The AI Daily Brief, “How People Are Actually Using Jev” (youtube.com/watch?v=uV6h3Uo4Nh8) — raw/How_People_Are_Actually_Using_Jev.md
  • NLW, The AI Daily Brief, “Why a New Class of AI Judgement Models Could Have Big Business Implications” (youtube.com/watch?v=sdR7FZ0SeqA) — raw/Why_a_New_Class_of_AI_Judgement_Models_Could_Have_Big_Business_Implications.md
  • Nate B Jones, “Why Developers Are Losing Their Minds Over AI That Can’t Write” (youtube.com/watch?v=tYugqJ9YytQ) — raw/Why_Developers_Are_Losing_Their_Minds_Over_AI_That_Can_t_Write.md
  • Greg Isenberg with Ryan Vogel (OpenCode founding team), “Jev is HERE. How to use it” (youtube.com/watch?v=4mTLpuQpB80) — raw/Jev_is_HERE._How_to_use_it.md
  • Matt Wolfe, “AI News - Opus 5.5, GPT-6 Sol, Jev, Muse and More” (youtube.com/watch?v=aDpIra7NFuE) — raw/AI_News_-_Opus_5.5_GPT-6_Sol_Jev_Muse_and_More.md
  • AI For Humans, “Instinct, Muse: AI Agents Can Run Your Life” (youtube.com/watch?v=-VL4VKHrnpU) — raw/Instinct_Muse_AI_Agents_Can_Run_Your_Life._We_Let_One_Try..md
  • Matthew Berman, “We need to talk about Jev…” (youtube.com/watch?v=2z-7pIj57f8) and “8 Jev Use Cases That Feel Like Cheating” (youtube.com/watch?v=jGD_UR4wMJc), both Zapier-sponsored — raw/We_need_to_talk_about_Jev....md, raw/8_Jev_Use_Cases_That_Feel_Like_Cheating.md
  • The Neuron Daily newsletter, 2026-09-25 and 2026-09-29 — raw/newsletter-theneurondaily-com-dceb623961.md, raw/newsletter-theneurondaily-com-094aff157e.md

Jev is the first “System One Model” from TypeSafe AI: it reads text and returns a typed decision (a choice, a score, or a yes-probability) with calibrated probabilities, and it cannot write prose. That trade makes it far cheaper and faster than an LLM for the many small “which bucket / how urgent / is this true” calls inside software and agent workflows. Eleven creator videos, two newsletters and TypeSafe’s own launch post agree on what it is and what it costs; they differ on how big the speed and accuracy gains are in practice. The practical use is as a cheap judgment layer beside an LLM: Jev tags, routes, scores and filters, and the LLM handles only the items that need writing or reasoning.

Key Takeaways

  • What it is. “Unstructured state in, typed probabilistic decisions out” (TypeSafe post). Text only; no images (Claire Vo; TypeSafe’s own Doom demo runs on structured text state, “not on images (yet…)”).
  • Three output types. A choice from up to 255 options you list; a score on a 2–10 level scale you describe in words; and a yes/no probability from 0 to 1 (NLW, How People…; the 255 cap is also in the TypeSafe post).
  • Price. 42 per billion”); output tokens free, “too cheap to meter” (TypeSafe post). Every creator who quotes a price gives the same figure (4 to 4.2 cents per million). TypeSafe says it “can’t prove it isn’t subsidized.”
  • Speed. 70–500 ms end to end per call (TypeSafe post). Ryan Vogel measured about 200 ms per query “no matter … the input output structure.”
  • Vendor multipliers vary by document ^[ambiguous]: the founder’s launch post on X said 20–200× faster and 40–400× cheaper (read aloud by Berman, NLW and MattVidPro); the blog table says 40–200× faster; the home-page figures of 193.6× faster and 444.6× cheaper come from TypeSafe’s own workflow evals, which it says are “on the higher end of real world gains.”
  • The recurring pattern across sources ^[inferred]: “LLM proposes options, Jev decides, code executes” (Nathan Flurry, quoted by NLW). Claire Vo: “Jev alone is okay. Jev with an LLM buddy is super powerful.”
  • Where it fails. Long-horizon goals (it loops in Minecraft on its own, per MattVidPro), multi-hop questions, maths and dates (NLW relaying TypeSafe), and anything needing reasoning or prose. It lost a chess game to Fable on material but won on the clock (Berman).
  • “Can’t hallucinate” is a claim about output shape. TypeSafe’s 0% figure “is not empirical. Schema matching is guaranteed” (TypeSafe post). In a test by Every it caught six of seven planted mistakes where Fable 5.1 caught all seven (NLW). ^[inferred] Treat it as always well-formed, not always right.
  • Adoption signal. Nate B Jones reports it was the fastest-adopted model in Vercel AI Gateway history within 24 hours. TypeSafe raised a 200M valuation and was reportedly in talks to raise up to 10B or more (NLW citing The Information; The Neuron, 2026-09-25).

What Jev Is

  • Maker and founder. TypeSafe AI; founder Diogo Almeida. The captions spell him “Diego” and “Dooo”; the post’s byline settles it. He writes that at OpenAI he “helped build the methods that made language models useful at following instructions … That work ended up as the research behind ChatGPT.” In his X post he calls himself a ChatGPT co-inventor (read aloud by Berman and NLW). That is his own claim.
  • Release. Early access opened after “two years in stealth.” The post is dated Sep 14, 2026; Nate B Jones dates the launch to September 15 ^[ambiguous] (possibly a time-zone difference
  • Name. “System One” comes from Kahneman’s fast/slow thinking. “Jev” is named after William Stanley Jevons: cheaper intelligence should raise total demand for it, as cheaper steam power did for coal (TypeSafe post; Nate B Jones).
  • How it differs from an LLM, per TypeSafe:
    • a new model architecture;
    • a “parallel sampler” that produces every output in one pass rather than one token at a time;
    • a training method called RLCD (Reinforcement Learning for Calibrated Decisions), which optimises for “answers with epistemically honest probabilities” where RLHF optimises for human preference.
  • Why calibration matters (Two Minute Papers). If it says 80% confident, it should be right about 80 times in 100. You can then “look at the confidence value and decide when to trust it or fall back to a heavier, smarter model.”
  • Output is directly usable by code. Numbers come back as number objects, “not like text that is a number” (Ryan Vogel). Several questions about the same item run in parallel. Per NLW relaying TypeSafe’s tests, 13 questions in one call were 12.2× cheaper and 10× faster than 13 separate calls, with identical answers.
  • The yes/no type’s name is unclear ^[ambiguous]. The captions render it as “new”, “null” and “nule”; NLW says the name is short for “Bernoulli.”
  • Skeptic’s view (Two Minute Papers). “There is no official research paper,” only a high-level description. The core idea resembles a classifier “that dates back around 90 years,” and “similar ideas have been explored for years.” Is it hyped? “Yes. Is it new? Partly.” What is new is the combination of the three ingredients above.
  • Another framing (Matt Stockton, quoted by NLW). Jev is an LLM-style interface to classic machine-learning classification. You get the classifier without labelling data, training a model or hosting it.

Pricing, Speed and Access

ItemValueSource
Input price$0.042 / M tokensTypeSafe post; Berman (“4.2 cents”); Nate B Jones; Claire Vo (“4 cents”)
Output priceFreeTypeSafe post
Worked cost1,000 calls × 1,000 tokens = 4 cents; 10,000 calls = 42 cents; 1M calls = $42Nate B Jones
Latency70–500 ms per call (TypeSafe evals run from laptops on the US West Coast)TypeSafe post
Measured latency~200 ms per query; 173 ms for one 12-question ad callRyan Vogel; NLW (Berman’s ad test)
Starting credit$5. Ryan’s team ran heavy demos for two days without using it upBerman, 8 Use Cases; Ryan Vogel
Evidence sizeUnder 32,000 tokens of textNLW, How People…
Choice cardinalityUp to 255; above that TypeSafe scores the options first, then choosesTypeSafe post; NLW
  • Where to get it:
    • TypeSafe’s own early access and waitlist (TypeSafe post; “invite only” per Greg Isenberg).
    • Vercel AI Gateway for instant access. MattVidPro added credits, got a key and let Codex call it: “Hours of testing and I spent less than a dollar.” Claire Vo says Jev was “currently free on the AI gateway by Vercel” when she recorded.
    • OpenRouter: MattVidPro says “it looks like” it is there too ^[ambiguous] (unconfirmed).
  • Access was throttled at one point. Matt Wolfe reported new sign-ups paused “because of the overwhelming demand.”
  • No-code route. Zapier added a “TypeSafe Jev” action (Berman; Zapier sponsored the video).
  • Agent route. Paste TypeSafe’s “setup prompt for agents” from its site into Claude Code, Codex or ChatGPT. The agent installs the TypeSafe skill; you create an account and an API key, and the agent connects it to your project (Nate B Jones). Claire Vo relies most on the docs’ “primitives” and “cookbooks” pages.

How People Use It: Four Patterns

These come from Nate B Jones. The mapping of NLW’s six categories onto them is this article’s

  1. A shim between messy input and routing code. Jev reads a ticket or email and picks a category, urgency and risk flag; ordinary code does the routing. NLW’s “triaging what comes in” and “instant response” categories fit here.
  2. Attention triage over a big corpus. Ask every item in a pile the same few questions, then count or rank the answers. This matches NLW’s “analyzing what you already have” and “searching by meaning.”
  3. The outer-loop chooser. Jev decides whether the next step is a tool, a cheap LLM, a frontier model or a human (Nate credits James Ward with “putting classifiers into the outer loop of a harness”). NLW’s “speeding up AI agents” belongs here.
  4. Live UI composition. Jev picks components from a fixed design system as the user types (Nate; Berman’s UI-assembly demo).
  • A fifth use: checking work against rules (NLW). Turn “please review this” into specific yes/no questions and run them on every paragraph or agent trace. Harrison Chase (LangChain) calls it “great for evals, especially online evals” (per NLW). Mike Taylor (Every) calls it a “linter for knowledge work” (per NLW).
  • Safety gate. Nate suggests having Jev review each proposed agent command (e.g. deleting a build folder, force-pushing) and return proceed / stop / ask a human.
  • Planner plus reflexes. MattVidPro’s conclusion from Minecraft: an LLM should direct “the instantaneous Jev body.” Paired with GPT-6 Astra as coordinator, Jev “plays Minecraft super well” (Woo Young Zoo’s demo, as MattVidPro describes it).

Use cases people reported

Numbers are as the named users reported them, relayed by the creators listed. None was independently reproduced here. Several names are caption spellings.

Bulk analysis of what you already have

  • Engineering-effort split (Claire Vo).
    • Method: pairwise “are these two PRs related?” calls cluster the PRs, then Gemini Flash Lite names the clusters.
    • Marketing site: 112 PRs for 1.1 cents.
    • ChatPRD: 1,700 PRs and 17,000 pairs for 9 cents in about 2 minutes. “Almost 30%” of PRs were platform, security and infrastructure.
  • Her local Claude Code and Codex session logs (Claire Vo). In January almost all her tasks were engineering; by September it was under 40%.
  • Product-insights graph (Claire Vo). About 1,100 signals (PRs, support tickets, Granola notes, Linear tickets) went through “over like 200,000 classification and pair-wise groupings” for about $4 of Jev. Astra did the analysis and Sol/Luna the generation.
  • Ad archive (Matthew Berman, per NLW). 724 live ads from 37 brands × 12 questions (hook, format, offer, CTA…) took 40 seconds and 9 cents.
    • He then ran 723 ads past 30 buyer archetypes: 21,690 stop-or-scroll decisions for 22 cents.
    • NLW’s caution: “this is not data” — treat it as a hypothesis generator before paying for real tests.
  • Post archive (Ian Nuttall, per NLW). Eight questions (topic, hook, tone…) about roughly 3,300 past X posts, to compare with engagement. The cost is garbled in the captions.
  • Inbox and messages.
    • Ryan Vogel: 1,700 emails scored on category, priority, spam and reply-worthiness. 4.2M input and 500K output tokens cost 18 cents in total, which matches free output
    • “Zach” (per Nate B Jones): 20,000 emails, Slack messages and transcripts in 7 minutes for about $1.
  • Tax documents (Nakshatra Saxena, per Nate B Jones). Swapping an LLM pipeline to Jev cost 34× less and ran 6× faster.
  • Research triage (Derya Unutmaz, per Nate B Jones). Picked the top 100 of 10,000 literature-grounded immunology questions.

Search and filter by meaning

  • SEO internal-link map (Boura, per NLW).
    • Jev read all 586 site pages (8,790 yes/no calls) in 45.1 seconds; its cost is garbled in the captions.
    • Claude Opus 5 got through 21 pages for $1.43.
    • Boura’s framing: internal linking “is a classification problem and we’ve been paying frontier prices to do it one page at a time.”
  • Brand news filter (Elvis, per NLW). Checked 384 morning news stories against 15 brands in 24.9 seconds; Opus 5 got through four stories in the same time.
  • Property search (Justine Moore, a16z, per NLW). Classified Zillow listings by things you can’t normally filter on: architectural style, renovation state, closeness to freeways.
  • Video clipping.
    • Clips from a 90-plus-minute video in under 2 seconds (Clipfast, per Berman; Burhan, per NLW, “under two cents”).
    • Ryan Vogel’s clip finder scored 17 moments in about 3 seconds.
  • Audience comments (Claire Vo).
    • About 4,500 YouTube comments via YouTube Data API v3, scored for sentiment and “contains an episode idea” (58 did), with a dashboard and live search.
    • She notes the speed needed architecture: “you have to like batch the results and then score them and then push the high scores up.”
  • Slop and ad blocking.
    • Kitsy’s Unclutter ad-and-slop blocker; bring your own key; open source per Berman.
    • A slop detector on madewithjev.com rated anthropic.com “26% AI slop” (Berman).
    • Robin Billgill’s real-time X slop filter (NLW).
  • Fuzzy find-in-page (Berman). A Google product manager’s open-source tool finds what you mean rather than the exact keyword.

Triage and routing

  • Email priority (Jonathan Unikowski, per NLW and Matt Wolfe). 100 emails rated in 453 ms for about a tenth of a cent; he reports Jev’s ratings matched his own on every one.
  • Downloads folder (Marcel Pio, CTO of Beyond Code, per NLW). Detects invoices, renames them and moves them, “no other LLM calls involved.”
  • Malicious links (Steven Tey, Dub, per NLW). Fed Jev 10,000 known-bad domains to flag links on the free shortener. “With Jev, we solved it in 2 hours.”
  • Moderation and routing.
    • Dev Ed moderates live chat (NLW; Matt Wolfe).
    • Box has tried incident triage for customer impact, severity and escalation route (NLW).
    • A CRM vendor proposed scoring every new lead on priority, buying readiness and spam (NLW).
  • Gmail cleanup (Claire Vo). Scored each email’s deletability from subject line and snippet, then handed the kept clusters to an agent.
  • Inbound leads for a small agency. Ryan Vogel’s girlfriend scores contact-form enquiries “is good lead” from 0 to 1.
  • No-code (Berman). A Zapier Zap where Jev accepts or declines calendar invites against your criteria.

Checking work against rules

  • Every’s planted-mistake test (per NLW).
    • Mistakes were planted in 12 passages. Jev caught six of seven in 0.35 s; Fable 5.1 caught all seven in 8.83 s.
    • Jev was “about 580 times cheaper,” and roughly 25× faster by those timings
  • Mike Taylor’s AI-tell check (Every, per NLW). 37 documents × 21 questions returned 777 judgments in under 0.7 s for “an estimated quarter of a cent.”

Speeding up agents and cutting tokens

  • Model and effort routing.
    • Riley Brown’s prompt-to-model router (Berman).
    • Vchen (per NLW) had Jev change GPT-6’s reasoning effort inside Codex mid-task, reporting “50% lower Astra costs” and faster runs.
  • Skill selection (Daniel Son, per NLW). Jev picks the one Claude Code skill that matches each request and injects only that skill: an “88% decrease in tokens and cost.”
  • A harness that learns (AJ Aspher, per NLW). Repetitive steps move from LLM calls to code as the harness runs. Cost per compliance alert fell from about $2.95 at the start (the end figure is garbled in the captions), and they claim the harness “cuts the cost of repetitive work by 90%.”
  • A cheaper tagger (AI For Humans). Gavin swapped Claude Haiku for Jev to tag link-blog posts on the AI For Humans site; his Claude did the swap. Result: “faster, better, and cheaper.”

Real-time and interface

  • Games (TypeSafe post; Berman; MattVidPro).
    • Doom at 10 queries a second, which TypeSafe prices at about $7 an hour.
    • Wikiracing: five hops in half a second, where the other models took 4–5 s.
    • Also “Melee” in real time (Alex from Berman’s team) and a simulated self-driving car (Justin Schroeder).
  • Crowd simulation (Berman). 50 town NPCs each chose a reaction in 0.6 s.
  • Browser use (Ryan Vogel; Nate B Jones). In Browser Use’s agent, Jev picks an operation and an element. It picked a Zurich→London flight in 7.1 s.
  • Voice computer use (Jack Chang, per MattVidPro). “Put that there” shape editing at conversational speed.
  • Voice mood app (Claire Vo). OpenAI Realtime voice feeds Jev, which scores a list of hex colours for the speaker’s mood and filters a quotes API.
  • Spreadsheet column (per Nate B Jones). Type “urgency” as a column header and Jev classifies every row.
  • Smart forms and paste. Copy a résumé and the application fields fill themselves (Marcus Low and Norman, per NLW). Match a service request to the best local provider for an instant quote (Ryan Vogel’s startup idea).

Where It Falls Down

  • Long-horizon goals. In Minecraft on its own it “would consistently predict the same actions over and over again, getting itself in loops.” It is “not so good at long-term goal keeping” (MattVidPro).
  • Open-ended judgment. A Bitcoin buy/hold/sell signal “does not seem to be doing well”: “I would not put this model in front of like your stock portfolio” (Ryan Vogel).
  • Deep reasoning. Fable outplayed it at chess (+16 material by move 29) but ran out of time at 6–15 s per move against Jev’s 2.6 s (Berman).
  • TypeSafe’s own list (per NLW). Accuracy drops with each hop in multi-step questions. It is weak at counting, maths and dates, so extract the facts with Jev and calculate elsewhere. It has consistency and intent-reading issues.
  • High-stakes calls. NLW advises caution on hiring, money and security: Jev “is going to return with a number that doesn’t have any reasoning attached.”
  • No reasoning trace. Nothing comes back but the typed answer (Ryan Vogel).
  • Input limits. Text only, and under 32K tokens of evidence (NLW). Matt Wolfe describes game-playing Jev as looking “at the screen” ^[ambiguous]. That conflicts with TypeSafe’s note that its game demo uses structured text state, and with Claire Vo and MattVidPro saying it can’t take images.
  • The headline speed needs engineering. Real-time search over thousands of items needed batching and ranking (Claire Vo). The TypeSafe post concedes its side-by-side demo used a short, dense input that “paints our model in an advantageous light.”

Vendor Claims vs Evidence

  • Benchmark.
    • TypeSafe’s “workflow evals” score each model against the average answer of GPT-6 Astra and Fable 5.1. The four workflows were built by its own capability team, and the LLMs ran through TypeSafe’s System One LLM wrapper (TypeSafe post).
    • Berman, reading a vendor chart, puts Jev about level with Luna, Terra and Sonnet 5 and above Opus 5 and Sol. That is vendor-run.
  • Hallucination. Berman frames Jev as suitable where “any hallucination is catastrophic” (healthcare, military targeting). TypeSafe’s actual guarantee is about schema matching. The Every test shows wrong answers still happen.
  • Price sustainability. TypeSafe: “We can’t prove it isn’t subsidized,” though it expects prices to fall. Claire Vo’s heavy use ran partly on Vercel’s free period.
  • Accuracy against labelled data. No source publishes it. The closest are Claire Vo (“it is accurate”), Unikowski’s 100 emails matching his own ratings, and Every’s six of seven.

Try It

  1. Find a Jev-shaped step. The Neuron’s test: “if you can write every valid answer on a sticky note, you probably don’t need a giant model composing prose to pick one.” NLW’s four fit criteria:
    • the answers can be written in advance;
    • there is volume (a pile or a stream);
    • a wrong answer is cheap or easy to catch;
    • the evidence fits in under 32K tokens of text.
  2. Get a key. Add credit on Vercel AI Gateway and create an API key (MattVidPro spent under $1 in hours of testing), or join TypeSafe’s early access. Without code, use the Zapier “TypeSafe Jev” action.
  3. Let your agent wire it up. Paste TypeSafe’s setup prompt for agents into Claude Code or Codex, then give it Nate B Jones’s prompt: “use the type safe skill to find a place in the project we’re working on together where asking an LLM to choose among defined outcomes is suboptimal. Build a Jev version instead.” Have it compare results, speed and cost and report back.
  4. Write the questions well (NLW).
    • Ask one judgment per question. “Is this a good lead?” is several: industry fit, company size, intent.
    • Describe every score level in words.
    • Ask many questions per item in one call; parallel questions add tokens, not time.
  5. Test before you trust it. Label 50 items yourself and compare (NLW). Then set a confidence threshold: act automatically above it, and send lower-confidence items to a frontier model or a person (Two Minute Papers; TypeSafe post on calibration).
  6. Marketing starters:
    • Turn each style-guide rule or banned phrase into a yes/no question and run it on every paragraph of a draft (NLW).
    • Score an ad or post archive on hook, format and offer, then join the results to performance data (Berman and Nuttall, per NLW).
    • Cluster your last few hundred PRs or support tickets pairwise and have a cheap generator name the clusters (Claire Vo).

Implementation

Tool/Service: Jev, TypeSafe AI’s first System One Model. Setup:

  • API key via Vercel AI Gateway, or TypeSafe early access plus its agent setup prompt, which installs a TypeSafe skill.
  • Docs: “primitives” and “cookbooks” pages. There is also a playground (Matt Wolfe).
  • Zapier action for no-code flows.

Cost: 5 of credit (Berman). Integration notes:

  • Define the question and every allowed answer up front; the response is typed (choice key, score, probability) and needs no parsing (Ryan Vogel).
  • Text in only.
  • More than 255 options needs two stages: score, then choose (TypeSafe post).
  • Batch items and questions for throughput (Claire Vo; NLW).
  • Pair it with a generative model for anything that must be written or reasoned.

Open Questions

  • No paper and no independent accuracy study. Are RLCD’s calibration claims reproducible on your own labelled data?
  • Data retention and terms. None of the sources saved here covers them; check before sending customer data.
  • Price after the free and subsidised period. Rate limits, and whether the sign-up pause Matt Wolfe reported has ended.
  • Launch date. Sep 14 (post dateline) or Sep 15 (Nate B Jones).
  • The yes/no type’s exact name. Captions give “new” and “null”; NLW ties it to “Bernoulli.” Check TypeSafe’s primitives docs.
  • OpenRouter availability. MattVidPro’s “it looks like” is unconfirmed.
  • The reported 10B or more. Relayed from The Information via NLW and The Neuron; not confirmed by TypeSafe.
  • The founder’s co-invention claim. It rests on his own statements.