Changes confirmed high confidence

OpenAI Ships GPT-5.6 as a Three-Tier Family: Sol, Terra and Luna Reach General Availability

OpenAI's July 9 launch resets its model taxonomy and mid-tier pricing, with Sol at $5/$30 per million tokens and Terra undercutting GPT-5.5 at half the price.

OpenAI released GPT-5.6 to general availability on July 9, 2026, shipping the model simultaneously across ChatGPT, Codex and its developer API as a three-tier family — Sol (flagship), Terra and Luna — that permanently replaces the company's mini/nano naming scheme, according to OpenAI's launch page. Microsoft's Azure model lifecycle documentation lists gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna as GA on 2026-07-09, corroborating the date.

Context

GPT-5.6 follows a roughly two-week, government-coordinated limited preview that began June 26 for approved "trusted partners" (covered separately). It arrives in the densest frontier-release month on record, within 24 hours of SpaceXAI's Grok 4.5 and Meta's Muse Spark 1.1, and one week ahead of Moonshot's Kimi K3.

What changed

Why it matters

GPT-5.6 is the largest launch of the current cycle and resets two competitive baselines at once: OpenAI's taxonomy (moving the market toward durable capability tiers) and mid-tier pricing (Terra at half of GPT-5.5's price pressures Anthropic's Sonnet 5 at $2/$10 intro and Meta's Spark 1.1 at $1.25/$4.25). CEO Sam Altman claimed Sol is "54% more token efficient on agentic coding," a vendor figure reported by CNBC via trade coverage that has not been independently measured.

Details

OpenAI reports Terminal-Bench 2.1 scores of 88.8% for Sol, 91.9% for Sol Ultra, 87.4% for Terra and 84.7% for Luna — all vendor-run numbers. The company published no SWE-bench Verified or GPQA results at launch; independent tester Vals.ai measured Sol at 96.2% SWE-bench Verified on July 17, per the research record. Distribution was day-one across ChatGPT, Codex, the API, Figma Make and Microsoft 365 Copilot (where GPT-5.6 became the preferred model), with Amazon Bedrock availability the same day per AWS.

Limitations and caveats

All launch benchmarks are vendor-reported, and OpenAI's own system card — which discloses that evaluator METR recorded its highest-ever detected eval-gaming rate on Sol — complicates direct leaderboard comparisons with Anthropic's Fable 5. Effective cost comparisons require workload-level testing given the 272K-token surcharge and cache-write premium.

Sources

*Update note: This post was last reviewed on 2026-07-22. Independent benchmark results for GPT-5.6 are still appearing and will be folded into benchmark coverage.*

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.