AI Super Simplified
evaluation

The Spec Adherence Test

A copy-paste prompt that builds you a side-by-side AI test with a right answer — so you can see which tool actually follows instructions rather than which one just looks impressive.

You are my AI evaluation partner. I want to find out which of the AI tools I use actually follows instructions precisely, versus which one just produces something that looks impressive. We are going to build a test I can run myself, with a right answer, so I am not taking any leaderboard's word for it.

Interview me first. Ask one question at a time and wait for my answer before asking the next. Do not assume anything about my setup, and do not invent details about my tools — if you need to know something, ask me.

1. Which two or three AI tools do you want to compare? Name them, or tell me you are not sure what you have access to and I will help you work it out.
2. What kind of work do you most want them to get right?
   a. Writing and editing
   b. Spreadsheets, data, and numbers
   c. Code, or building something visual
   d. Research and summarising
   e. Something else — tell me
3. How do you want to judge the result?
   a. I want a countable right answer I can verify myself
   b. I want to compare quality side by side and pick a winner
4. Roughly how much time do you want to spend on this — five minutes, half an hour, or longer?

Once I have answered all four, do all of the following.

Write me ONE test prompt I can paste into each tool unchanged. Build it so a correct answer is checkable, not a matter of taste. Give it explicit, countable requirements — exact numbers of things, exact formats, exact lengths, an exact structure — and state them plainly enough that I can verify each one myself afterwards.

Then give me a scoring sheet as a simple checklist. Every line must be something I can verify by counting or looking, never something I have to have an opinion about. Tell me what a pass, a partial, and a fail look like for each line.

Then tell me in one short paragraph what result would actually mean something, and what result would be noise. Be honest: if this test is likely to come out a tie, say so up front rather than letting me run it and feel like I learned something I did not.

Then warn me about the single most likely way this goes wrong — the way I might accidentally build a test that measures nothing at all.

End with a one-line verdict I can use: what I should conclude if the first tool wins, if the second wins, and if they tie.