Thinking Machines Releases Inkling, Its First Model, With Apache 2.0 Weights on Day One
Mira Murati's lab shipped a 975B MoE with full open weights at launch — no gating, no wait — while candidly conceding it is 'not the strongest overall model available today.'
Thinking Machines Lab, Mira Murati's startup, released its first model on July 15, 2026: Inkling, a 975B-total/41B-active mixture-of-experts model with full weights on Hugging Face under Apache 2.0 from day one — no gating, no waitlist — per Vectrel's analysis citing VentureBeat and NYU Shanghai's RITS write-up.
Context
Inkling arrives as the first credible US-lab open-weight release at near-frontier scale — a counterpoint to July's Chinese open-weight wave (LongCat-2.0, Hy3, Kimi K3) and to Meta's pivot to closed paid APIs with Muse Spark 1.1. An Inkling-Small preview (276B/12B active) shipped alongside.
What changed
- Specs: 975B total / 41B active MoE (256 routed + 2 shared experts, 6 active per token), 1M-token context, trained from scratch on 45T tokens of text, image, audio and video.
- Honest positioning: Thinking Machines says Inkling is "not the strongest overall model available today, open or closed," pitching it instead as a fine-tuning base — multimodal capability, efficient reasoning, and availability on the Tinker fine-tuning service (customers include Bridgewater).
- Benchmarks (vendor/AA): 77.6% SWE-bench Verified, 97.1% AIME 2026, 74.1% MCP Atlas; Artificial Analysis Intelligence Index 41 — #1 among US-lab open-weight models.
- Hosted pricing: roughly $1.00/$4.05 per million tokens via OpenRouter and Together.
- Ecosystem: day-zero support across vLLM, SGLang, Modal, Baseten, Databricks, Together and Fireworks; Unsloth shipped 1-bit GGUF quants (270GB versus 1.9TB full precision).
Why it matters
Inkling tests whether a US lab can build a business on open weights in 2026 — monetizing through fine-tuning services rather than API metering. Its day-one Apache 2.0 release also sharpened the media-literacy contrast with Kimi K3, launched a day later with weights only promised.
Details
Deployment reality is cluster-scale: the BF16 checkpoint needs at least 2TB of aggregate VRAM, and the NVFP4 checkpoint at least 600GB, per wavespeed.ai's deployment analysis. Community reviewers noted the lab's distillation-purity claims were walked back after outside researchers found Kimi 2.5 SFT traces — a provenance caveat for downstream users.
Limitations and caveats
Benchmarks are vendor-reported or Artificial Analysis measurements; the model trails leading open models on some agentic benchmarks. A circulating "63% hallucination rate" claim comes from a single low-authority source and is excluded from this report. Self-hosting remains impractical for individuals.
Sources
- Vectrel — Inkling, Kimi K3 and the open-weight ownership strategy (aggregator)
- NYU Shanghai RITS — Thinking Machines releases Inkling (research)
- wavespeed.ai — What is the Inkling model? Deployment requirements (aggregator)
*Update note: This post was last reviewed on 2026-07-22. Watch for independent evals of the released weights and Tinker adoption metrics.*
Sources
- Vectrel — Thinking Machines Inkling / Kimi K3 open-weight strategy — aggregator
- NYU Shanghai RITS — Thinking Machines releases Inkling — research
- wavespeed.ai — What is Inkling? — aggregator
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.