GPT-Live system card: disclosed (non-significant) regressions vs Advanced Voice Mode
OpenAI's GPT-Live system card introduces voice-native safety evals built from real (consented, de-identified) user audio and compares against AVM predecessors. Disclosures: GPT-Live-1 shows a slight regression on emotional reliance (0.88→0.
On 2026-07-08, the verified AI news record added a significant safety, evals & benchmarks development: OpenAI's GPT-Live system card introduces voice-native safety evals built from real (consented, de-identified) user audio and compares against AVM predecessors. Disclosures: GPT-Live-1 shows a slight regression on emotional reliance (0.88→0.82) and GPT-Live-1 mini on sexual content (0.97→0.95); OpenAI notes neither is statistically significant; evals are deliberately hard and not prevalence-weighted.
Context
OpenAI's GPT-Live system card introduces voice-native safety evals built from real (consented, de-identified) user audio and compares against AVM predecessors. Disclosures: GPT-Live-1 shows a slight regression on emotional reliance (0.88→0.82) and GPT-Live-1 mini on sexual content (0.97→0.95); OpenAI notes neither is statistically significant; evals are deliberately hard and not prevalence-weighted. Card notes comparison values for prior models are from latest snapshots, so launch-time values may differ — a recurring comparability caveat across OpenAI cards.
What changed
GPT-Live shows small, non-significant regressions on emotional reliance and sexual content vs AVM. According to OpenAI — GPT-Live System Card, Deployment Safety Hub (PRIMARY, fetched), the supporting record states: “GPT-Live-1 shows a slight regression on emotional reliance from 0.88 to 0.82, and GPT-Live-1 mini shows a slight regression on sexual content from 0.97 to 0.95. Note that neither of these are statistically significant.”.
Why it matters
Rare explicit regression disclosure in a launch system card; establishes audio-native eval categories (self-harm, psychosis/mania, emotional reliance) as a new safety-eval surface for voice agents. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.
Details
The research file records the item under “GPT-Live system card: disclosed (non-significant) regressions vs Advanced Voice Mode” with source timing of 2026-07-08. The captured research confidence note is: High (primary). Card notes comparison values for prior models are from latest snapshots, so launch-time values may differ — a recurring comparability caveat across OpenAI cards.
Limitations and caveats
The research file did not identify a blocking caveat, but vendor-supplied claims should still be read as company statements unless independently confirmed.
Sources
Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.
Sources
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.