The Refusal Gap

OpenAI’s Advanced Cybersecurity Completion Rate, same tasks, four models. Tap a row to see who can reach it.

The half-point that matters. Standard GPT‑5.6 Sol and the Daybreak Blue variant are the same model. Retuning the guardrails moved it 0.5 points. Everything above that involves both extra training and a different permission level — and a completion rate counts whether the model answered at all, not whether it answered well.

Source: figures reported at the 10 August 2026 launch. Internal OpenAI benchmark, not independently audited.