The Refusal Gap
OpenAI’s Advanced Cybersecurity Completion Rate, same tasks, four models. Tap a row to see who can reach it.
The half-point that matters. Standard GPT‑5.6 Sol and the Daybreak Blue variant are the same model. Retuning the guardrails moved it 0.5 points. Everything above that involves both extra training and a different permission level — and a completion rate counts whether the model answered at all, not whether it answered well.