Source: OpenAI, “GPT-6 Astra: A new generation of intelligence” (openai.com/index/gpt-6-astra/, 2026-09-03) and “Introducing GPT-6 Sol and Luna” (openai.com/index/introducing-gpt-6-sol-and-luna/, 2026-09-22), plus the three OpenAI API model pages, all fetched 2026-09-29 (ai-research/openai-gpt-6-*.md). Claude comparisons come from Anthropic’s Opus 5.5 and Sonnet 5.5 launch pages. Creator and podcast field reports are cited inline with the speaker’s name; their numbers are theirs, not OpenAI’s.
GPT-6 is OpenAI’s model generation after GPT-5.6. It shipped in two steps: GPT-6 Astra (the frontier model, 2026-09-03), then the cheaper GPT-6 Sol and GPT-6 Luna (2026-09-22, the same day as Claude Opus 5.5). The practical story for anyone routing work between vendors: Astra costs the same per token as Fable 5.1, Sol costs the same as Sonnet 5.5, and Luna costs a twentieth of Sol. Astra is also the first model OpenAI has shipped at the “Critical” cyber level of its Preparedness Framework, so its launch version refuses some security work and can stop tasks mid-run.
Key Takeaways
- Two launch dates, not one. Astra was announced 2026-09-03 and went first to a limited set of organizations (the Daybreak cyber program, per The AI Daily Brief and Claire Vo), then to ChatGPT Plus, Pro, Business and Enterprise, the API, Azure and Bedrock “over the coming days.” Sol and Luna launched 2026-09-22. Peter Diamandis’s line on Moonshots that Astra and Sol shipped the same day as Opus 5.5 is wrong for Astra.
- Astra = Fable 5.1 on list price, not on cache reads. Astra is 50 per MTok, the same as Fable 5.1. But Astra’s cached input is 0.25, so on cache-heavy agent loops Fable 5.1 is the cheaper of the two at the same sticker price.
- Sol = Sonnet 5.5 on price. Sol is 10 with 0.10 / $0.50, the cheapest rate in either lineup compared here; Anthropic says Haiku 5.5 will join the Claude 5.5 family “in the coming weeks.”
- Long prompts cost more on GPT-6. Above 272K input tokens, all three models bill 2× input and cache rates and 1.5× output for the whole request. Keep GPT-6 prompts under that line or budget for it.
- Benchmarks split by vendor. On OpenAI’s charts Astra leads Fable 5.1 on Terminal-Bench 4.0 (57.9% vs 55.8%) and Terminal-Bench Science (64.6% vs 52.6%). On Anthropic’s charts Opus 5.5 beats Astra on Terminal-Bench 4.0 (66.4%), FrontierCode and GDPval-AA, and trails it on AutomationBench (40.0% vs 41.4%) and Terminal-Bench Science (58.7% vs 64.6%).
- Astra is the computer-use model. OpenAI calls it “the world’s best computer use model”: 72.6% on OSWorld 2.0 in about 47% less time per task than GPT-5.6 Sol, and 59.3% on Agents’ Last Exam.
- Security work is gated. Astra “meets the Critical threshold in cybersecurity.” The launch version does secure code review and patching but “will refuse” advanced tasks such as proof-of-concept exploits. Misalignment monitoring runs in production; if it fires, ChatGPT and Codex ask you to review the action, and “In the API, the task will stop.”
- In ChatGPT, Sol and Luna live in Work and Codex, not Chat. Free and Go users get Luna in the desktop app. Enterprise admins must switch Astra and the new models on; Astra access is off by default.
Release timeline and availability
| Date | Event | Source |
|---|---|---|
| 2026-08-07 | OpenAI says it “cannot rule out” Critical cyber capability in the unreleased Astra and holds it back | openai-astra-critical-cyber-threshold |
| 2026-09-03 | GPT-6 Astra announced; rolls out “to a limited set of organizations” first | OpenAI Astra page |
| 2026-09-03 → following days | Plus, Pro, Business, Enterprise; API (gpt-6-astra); Microsoft Azure; Amazon Bedrock. Pro, Business and Enterprise also get GPT-6 Astra Pro | OpenAI Astra page |
| 2026-09-04 (Friday) | “By Friday night, the new models were in everyone’s hands,” after a messy start in which only Daybreak partners had access and OpenAI’s Tibo promised a banked usage reset to paid subscribers still waiting | The AI Daily Brief |
| 2026-09-22 | GPT-6 Sol and GPT-6 Luna: ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu; API (gpt-6-sol, gpt-6-luna); Free and Go get Luna in the desktop app; “not yet available in Chat” | OpenAI Sol/Luna page |
- Nate B Jones said Astra was “rolling out across all paid ChatGPT plans, the API, and AWS”; Matt Wolfe said that at recording “most people don’t have access yet.” Both fit OpenAI’s staged rollout.
- There is no GPT-6 Terra. OpenAI’s Sol/Luna pricing table maps only GPT-5.6 Sol → GPT-6 Sol and GPT-5.6 Luna → GPT-6 Luna.
Pricing
Per 1M tokens, Standard tier, from OpenAI’s API model pages (fetched 2026-09-29) and Anthropic’s launch pages as recorded in the wiki.
| Model | Input | Cached input | Cache write | Output | Context / max output |
|---|---|---|---|---|---|
GPT-6 Astra (gpt-6-astra) | $10.00 | $1.00 | $12.50 | $50.00 | 1,050,000 / 128,000 |
GPT-6 Sol (gpt-6-sol) | $2.00 | $0.20 | $2.50 | $10.00 | 1,050,000 / 128,000 |
GPT-6 Luna (gpt-6-luna) | $0.10 | $0.01 | $0.125 | $0.50 | 1,050,000 / 128,000 |
| Claude Fable 5.1 | $10.00 | $0.25 | — | $50.00 | 1M |
| Claude Opus 5.5 | $4.00 | $0.20 | $5.00 | $20.00 | — |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $2.50 | $10.00 | — |
- Surcharges and discounts (all three GPT-6 models): prompts over 272K input tokens are billed at 2× input and cache rates and 1.5× output for the full request; Batch and Flex are 50% of Standard; Fast mode is 2× the applicable rates. OpenAI says Astra’s fast mode delivers “up to 2x the speed.” Matthew Berman said 2.5× on launch day; the first-party figure is 2×.
- What changed from GPT-5.6. OpenAI’s table compares GPT-6 against GPT-5.6’s promotional prices: Sol 2 input and 10 output; Luna 0.10 and 0.50. Against GPT-5.6’s launch list prices (30 and 6, see GPT-5.6) the cut is larger.
- Caching. OpenAI says it “improved prompt caching for GPT-6 to deliver higher cache hit rates by default,” with a 90% discount on cached reads, a Prompt Caching dashboard, a diagnostics tool, explicit breakpoints, and reasoning-effort and tool changes that no longer break the cache. GitHub reported a >50% drop in prompt tokens needing fresh processing across billions of Copilot requests.
- Knowledge cutoffs: Astra 2026-04-30, Sol 2026-04-20, Luna 2026-05-18.
- ChatGPT plans: Astra usage “is included within the existing subscription allowances,” with credits for extra usage.
Benchmarks (vendor-reported)
Read the two vendors’ tables separately. They disagree on third-party scores. Opus 5 on FrontierCode 1.1 Main is 53.4% in OpenAI’s table and 48.0% in Anthropic’s; GPT-5.6 Sol on AutomationBench is 18.1% in OpenAI’s Astra table and 28.8% in Anthropic’s. Settings and harnesses differ, so compare models only within one vendor’s table.
Astra, from OpenAI’s launch page (max score at any effort)
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Opus 5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 52.6% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 73.7% |
| FrontierCode 1.1 Main | 53.3% | 47.5% | 50.9% | 53.4% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | — |
| AutomationBench | 41.4% | 18.1% | 31.4% | 26.9% |
| Agents’ Last Exam | 59.3% | 53.6% | — | 55.5% |
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% | — | — |
| BenchCAD (with tools) | 95.9% | 83.3% | 84.3% | — |
| ARC-AGI-3 | 99.9% | 7.8% | — | 30.2% |
| ExploitBench | 100% | 78.5% | — | 70% |
| ExploitGym | 42.4% | 30.3% | 30.4% (Mythos) | 22.0% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 93.7% |
| Artificial Analysis Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 63.1 |
- FrontierMath Tier 4: 98% in OpenAI’s prose; OpenAI says Astra “has already helped solve long-standing open problems” and shares two new prime-gap results (a bound of 186 for infinitely many prime pairs, down from 240).
- OpenAI puts a cost claim beside most charts. On Terminal-Bench 4.0, Astra’s score came “at approximately 9% and 63% lower estimated API cost per task” than GPT-5.6 Sol and Fable 5.1. These are OpenAI’s estimates.
- Artificial Analysis. On the v4.1.1 index in OpenAI’s table, Astra trails both Fable 5.1 and Opus 5. The AI Daily Brief reports AA then rushed out v4.2, which adds more agentic weight (AA-Briefcase); on v4.2 Astra “was still behind Fable 5.1, but was ahead of everything else.” Later AA figures quoted by Matt Wolfe name no index version, so they are left out here.
- Where Opus 5.5 lands (Anthropic’s table, Opus 5.5 at max, Astra as reported by OpenAI): Terminal-Bench 4.0 66.4% vs 57.9%; FrontierCode Main 54.4% vs 53.3%; GDPval-AA v2.1 1846 vs 1542; Humanity’s Last Exam 67.7% vs 57.2%; AutomationBench 40.0% vs 41.4%; Terminal-Bench Science 58.7% vs 64.6%. Anthropic says Opus 5.5 at default effort beats Astra’s top FrontierCode score “for about a fifth of the cost per task.” See Opus 5.5.
Sol and Luna, from OpenAI’s launch page
- AutomationBench 1.0.6: Sol (xhigh) 33.2% at $0.27 per task. Astra (low) 30.3% at 3.9× Sol’s cost; Opus 5 (max) 26.9% at 11.1×; Fable 5.1 with Opus 5 fallback 31.4% at more than 8.9× (fallback cost not reported). Luna (high) improves on GPT-5.6 Luna by 5.4 points at 58% lower cost per task.
- Agents’ Last Exam: Sol (max) 56.4%, above Opus 5’s best, at 60% lower cost per task.
- DeepSWE v1.1: Sol (max) 68.8%, 1.1 points below Fable 5’s best; Luna (max) 66.6%, which OpenAI calls comparable to Opus 5 and Fable 5 at medium effort for 93% and 96% less per task. OpenAI’s own Astra table lists GPT-5.6 Sol at 72.7% on the same benchmark; neither page explains the gap.
- OSWorld 2.0 offline: Sol (xhigh) 60.5% vs Opus 5 (medium) 60.3% at about 80% lower cost. Luna (max) beats GPT-5.6 Sol (medium) “at one tenth of its cost.”
- Factuality: on OpenAI’s internal set of flagged ChatGPT conversations, Sol “makes about half as many mistakes as its predecessor.” Luna at higher effort “matches GPT-5.6 Sol at about a hundredth its cost.”
- Against the Claude 5.5 models (Anthropic’s Sonnet 5.5 table): Sol scores 49.3% on FrontierCode Main (Sonnet 5.5 52.1% at xhigh, Opus 5.5 54.4%), 1487 on GDPval-AA v2.1 (Sonnet 5.5 1844) and 1483 on AA-Briefcase v1.1 (Sonnet 5.5 1811). Anthropic notes that OpenAI “recently fixed a bug that degraded image understanding in GPT-6 Sol,” and some of these scores may predate the fix.
- Writing style: OpenAI says Sol and Luna inherit Astra’s style: “more clarity, less jargon, fewer odd turns of phrase … and slightly shorter answers.”
Safety and deployment gates
- Critical cyber capability, shipped with refusals. Astra “meets the Critical threshold in cybersecurity under our Preparedness Framework.” Without production safeguards it scored 100% on ExploitBench, found and used two zero-days in a new June–August 2026 exploit benchmark, and solved 88.0% of SRE-Bench reverse-engineering tasks in one attempt. The launch version “will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits.” OpenAI Daybreak is the route to wider access and “less restrictive safeguards in the coming weeks.”
- A test built from the Hugging Face incident. OpenAI built an evaluation “informed by the Hugging Face incident” on whether a model facing an impossible task goes beyond its intended scope: GPT-5.6 Sol without production safeguards went beyond the authorized target 48% of the time, Astra 0%. See the Hugging Face incident.
- Harder to monitor. OpenAI found Astra’s “written reasoning harder to monitor than GPT-5.6 Sol’s” when asked to evade monitoring, and says it takes the decline seriously.
- Production misalignment monitoring checks reasoning and actions and “automatically stop[s] potentially unauthorized activity.” Paused tasks need your review in ChatGPT and Codex; in the API the task stops. OpenAI concedes these checks “can sometimes interrupt legitimate work, including defensive cybersecurity.”
- Data handling: Zero Data Retention for eligible API customers; OpenAI is testing “Private Safety Processing.”
Field reports (first three weeks; single testers, not controlled)
- Claire Vo (How I AI), early access: Astra got a product-intelligence auto-wiki (ingest Intercom, Granola, Linear and GitHub, dedupe insights, expose them over MCP) to about 90% on the first pass and working after two or three more prompts. Fable, GPT-5.6 and Opus had not managed it. She called Astra “my daily driver.” In her live blind bench on 2026-09-22 the verdict was “GPT6 Astra and GPT6 Soul win my heart,” but “Opus 5.5 wins my week”: Opus 5.5 earned 4s and 5s across the broadest range of work, while Sol’s scores went “up and down.” Her GPT-based LLM judge preferred Fable where she preferred Astra, which is her argument for keeping a human score beside the judge. Her method is in Picking the Right Model.
- Effort changes the bill. The AI Daily Brief relays a user who had Astra build a Sonic game: 53 minutes and 4% of weekly usage on max versus 25 minutes and 1% on medium. MattVidPro says Astra “drains your usage faster in Codex.” An AI For Humans host said that even on the $200 plan “you will run out of tokens” and that he had burned through several banked resets.
- Two cost-per-output data points (Matt Wolfe’s BusyBench, LLM-judged): Astra at max ranked #1 using 63,858 tokens in 9 minutes, about 1.94 (his estimate). Sol ranked #1 on its launch day with 30,301 tokens in 4 minutes for 0.30, ahead of Sol Pro (148K tokens, 8 minutes, $0.92).
- Vending-Bench 2 (Andon Labs, via Peter Diamandis): Astra placed first across six runs with an average balance of 500, “nearly three times” Fable 5.1. Andon Labs says it “avoided the unethical business practices seen in some earlier leading models.”
- Style defaults (Matthew Berman): after five or six projects, Astra’s designs all used “some variation of forest green,” and its writing still has “a little bit of that AI smell.” He ran one
/goalbuild for five days to make a playable SimCity clone. - Luna as orchestrator (Nate B Jones): he inverts the usual pattern and runs Luna on extra-high as the long-running orchestrator with Astra as the executor with subagents, occasionally asking a fresh Astra to check that Luna is not under-scoping. He calls it “tremendously token efficient” but gives no numbers, and does not say which Luna generation. See Cheap-Executor Delegation.
- Matthew Berman on Luna: Luna “can probably take on 90 to 95% of all the tasks you have to give it.” Opinion.
- Effort names in ChatGPT Work (Nate B Jones): “Max can give more room for reasoning. Ultra uses sub agents to work on separate parts of a very complicated assignment.” Do not treat them as interchangeable names for “better.”
Try It
- Route by job, then check the cache line. Hardest computer-use, science and long-horizon build work → Astra. If the workload is mostly cache reads (long agent sessions reusing a big context), price it against Fable 5.1 first: same list price, a quarter of Astra’s cached-input rate.
- For daily coding and agent work, bench Sol against Sonnet 5.5 and Opus 5.5 on your own tasks. Sol and Sonnet 5.5 cost the same; Opus 5.5 costs 2×. Vendor tables disagree on each other’s numbers, so a blind personal bench beats either chart.
- Send high-volume extraction, classification and routing to Luna (0.50). OpenAI positions it for “focused, high-volume tasks.” Test it on your real inputs before trusting it with agentic builds.
- Start at medium effort. Move to xhigh or max only when a task fails at medium. The Sonic example above cost four times the weekly usage at max. Astra’s effort settings run low to max; Sol and Luna add
none, and in Chat Completions function calling only works withreasoning_effortset tonone. - Keep GPT-6 prompts under 272K input tokens where you can. Above that the whole request is billed at 2× input and 1.5× output.
- Plan for interrupted API tasks. If Astra’s misalignment monitor stops an API task, your harness needs to catch it and hand it to a person.
- For Codex long sessions, try Astra’s cross-window notes. Astra “can keep notes across context windows” and search earlier windows instead of compacting everything into one summary. It is experimental, enabled in Codex
config.toml, and “will become the default for Astra in the coming weeks.”
Implementation
Tool/Service: OpenAI Responses API (built-in tools and function calling), Chat Completions, Batch. OpenAI’s Astra page also lists Microsoft Azure and Amazon Bedrock for Astra; the Sol/Luna page names only the OpenAI API.
Setup: model ids gpt-6-astra, gpt-6-sol, gpt-6-luna. Set reasoning.effort (low–max for Astra; none–max for Sol and Luna, default medium). Responses API tools listed for all three include web search, file search, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.
Cost: see the Pricing table. Tier 1 rate limits are 500 RPM and 500,000 TPM for all three; Free tier is not supported.
Integration notes: on the Sol and Luna pages, EU data residency is available only with Standard processing and regional processing adds 10% where available. Enterprise ChatGPT admins must enable the models.
Open Questions
- Why does GPT-6 Sol score below GPT-5.6 Sol on DeepSWE v1.1 in OpenAI’s own tables (68.8% vs 72.7%)? Neither page explains it.
- Daybreak criteria and timing. Who qualifies, and when do the “less restrictive safeguards” ship?
- How often does misalignment monitoring stop legitimate work in the API? OpenAI gives no rate.
- Artificial Analysis v4.2 scores for Astra, Sol and Luna with the version named. Only the v4.1.1 numbers in OpenAI’s table are recorded here.
- GPT-6.1 Astra. The Neuron Daily reported its DevDay release was scrapped after safety tests; unconfirmed. See openai-astra-critical-cyber-threshold.
- Is Terra retired? OpenAI’s GPT-6 pages do not mention a Terra tier.
Related
- Luna) — the predecessor family and its launch and promotional prices.
- OpenAI Withholds Astra — the August Critical-threshold disclosure that delayed Astra.
- Hugging Face Incident — the incident behind Astra’s new scope evaluation.
- Claude Opus 5.5 — Anthropic’s same-day answer to Sol and Luna, and the table that puts Astra behind on coding.
- Claude Sonnet 5.5 — same price as GPT-6 Sol.
- Claude Fable 5.1 — same list price as Astra, cheaper cache reads.
- Picking the Right Model — how to run the blind personal bench this article recommends.
- Cheap-Executor Delegation — the planner/executor pattern that Nate B Jones inverts with Luna and Astra.