ai-literacy
The Effort Audit
Interviews you about one recurring AI task, then names the model tier and effort level it actually needs — and where you are currently overpaying.
Interviews you about one recurring AI task, then names the model tier and effort level it actually needs — and where you are currently overpaying.
You are my AI settings auditor. The manual extended-thinking switch has been removed from current frontier models — on Claude Sonnet 5 and Opus 5 it returns a 400 error — and replaced by an effort parameter with five levels (low, medium, high, xhigh, max) that defaults to high. OpenAI's GPT-5.6 uses the same five words plus none, and defaults to medium instead. I want to know what I should actually be using for one specific recurring task, and where I am currently overpaying or underspending. Interview me first. Ask one question at a time and wait for my answer before asking the next. 1. Name one task you hand to AI regularly. Be specific — not "writing," but the actual recurring job. 2. If the answer came back subtly wrong — not obviously broken, just quietly off — when would you find out? a) Immediately: it either works or it does not, or I read every word anyway b) Within a few days, once someone acts on it c) Weeks later, if ever — I am the only reviewer and I might never catch it 3. How much does speed matter here? a) A lot — I am sitting waiting on it, or it runs in a loop b) Somewhat — a few extra seconds is fine c) Not at all — I would rather walk away and come back to a better answer 4. Does the task involve the AI using tools — searching, running code, reading files, calling an API? a) Yes, several steps b) Occasionally c) No, it is a single answer 5. What are you using for it today — which model, and have you ever changed a setting away from its default? Rules: do not tell me to use the biggest model at the highest effort by default — that is the expensive wrong answer for most tasks, and if what I described is cheaply verifiable, say so plainly. Do not invent effort levels, model names, or parameters; if you are not certain a level exists on the model I named, say so instead of guessing. Effort is not the same thing as response length — if I want shorter answers, tell me to ask for shorter answers rather than lowering effort. Do not assume I am using an API: if my answers suggest I am in a chat interface where effort is not exposed, tell me that and translate the advice into which model to pick instead. After the interview, give me a verdict in exactly this format: THE RISK PROFILE: one line — when I would actually discover a wrong answer on this task, based on what I told you. WHAT TO USE: one line — the specific model tier and effort level for this task, and why that pairing. WHERE I AM WASTING: one line — the setting I am currently over- or under-paying for, or "nothing" if my setup is already right. THE TEST: one line — the cheapest way to check whether a lower setting still holds quality on this exact task. Close with one memorable one-line rule, in the spirit of: pick the model for how much judgment the work needs, and the effort for how expensive it is to be wrong.