Grok 4.5 ships with no model card; independent testing shows hallucination rate doubling (25%→54%)
SpaceXAI launched Grok 4.5 without a model card or system card (breaking xAI's own prior practice under its August 2025 Risk Management Framework, which AI Lab Watch/LessWrong analyses had already criticized as inadequate). Artificial Analy
On 2026-07-08, the verified AI news record added a significant safety, evals & benchmarks development: SpaceXAI launched Grok 4.5 without a model card or system card (breaking xAI's own prior practice under its August 2025 Risk Management Framework, which AI Lab Watch/LessWrong analyses had already criticized as inadequate). Artificial Analysis AA-Omniscience: accuracy 35%→52% but hallucination (confident-wrong) rate 25%→54% vs Grok 4.3 — "buying knowledge at the cost of calibration." Fable 5 (54.9%) and Kimi K3 (51%) show the same frontier pattern: the most capable 2026 models are among the most overconfident. Separately reported: a Cursor codebase snapshot contaminated Grok 4.5 training data, inflating its CursorBench score — a voluntary contamination disclosure (single secondary source; see Conflicts).
Context
SpaceXAI launched Grok 4.5 without a model card or system card (breaking xAI's own prior practice under its August 2025 Risk Management Framework, which AI Lab Watch/LessWrong analyses had already criticized as inadequate). Artificial Analysis AA-Omniscience: accuracy 35%→52% but hallucination (confident-wrong) rate 25%→54% vs Grok 4.3 — "buying knowledge at the cost of calibration." Fable 5 (54.9%) and Kimi K3 (51%) show the same frontier pattern: the most capable 2026 models are among the most overconfident. Separately reported: a Cursor codebase snapshot contaminated Grok 4.5 training data, inflating its CursorBench score — a voluntary contamination disclosure (single secondary source; see Conflicts). EU availability delayed at launch; compliance teams told to treat the absent safety evaluation as an open item. AA-Omniscience measures overconfidence on knowledge questions, not coding reliability.
What changed
No Grok 4.5 model card; hallucination doubled on AA-Omniscience while accuracy rose. According to aitoolsreview.co.uk (2026-07-14, B); suprmind.ai hallucination tracker (2026-07-18, B, AA-sourced tables), the supporting record states: “at the time of writing, xAI had not published a Grok 4.5-specific model card or system card… its hallucination rate rose even faster over the same period, from 25% to 54% — meaning Grok 4.5 is now more likely to confidently state something false than its own predecessor was.”.
Why it matters
The only July-wave flagship with no published safety evaluation, landing in an EU AI Act transparency environment; also the clearest independent-verification-vs-vendor-claims gap of the cycle. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.
Details
The research file records the item under “Grok 4.5 ships with no model card; independent testing shows hallucination rate doubling (25%→54%)” with source timing of Launch 2026-07-08; AA-Omniscience data through Jul 18. The captured research confidence note is: High on missing card + AA figures (multi-source); Medium-low on CursorBench contamination (single secondary — thursdai). Additional captured source links are listed below so readers can inspect the evidence trail rather than rely on a single summary. EU availability delayed at launch; compliance teams told to treat the absent safety evaluation as an open item. AA-Omniscience measures overconfidence on knowledge questions, not coding reliability.
Limitations and caveats
This item is based on a single authoritative source or a company-attributed claim captured in the research file; independent corroboration was not established in the research window. Research confidence note: High on missing card + AA figures (multi-source); Medium-low on CursorBench contamination (single secondary — thursdai).
Sources
Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.