Same prompt. Different AI. Touch the results.
Every week we feed one identical prompt to competing AI models and publish exactly what each one returned — interactive apps you can open and run, and written answers side by side. No benchmark charts, no cherry-picking. See the differences yourself, then steal the prompt and rerun the test.
The Planetarium Test
One prompt asked two models to build a working planetarium from first principles — real orbital math, no libraries, no data files. Both shipped a working night sky.
All matchups
The archive grows every week.
The Planetarium Test
One prompt asked two models to build a working planetarium from first principles — real orbital math, no libraries, no data files. Both shipped a working night sky.
The Email Rewrite Test
One brutal corporate email. Two models asked to rewrite it — warmer, clearer, still professional. Same words in, very different words out.
The Data Analysis Test
Same messy CSV. Two models asked to find the insight hidden in the numbers — no chart libraries, just logic and output. One found it. One got close.
The Landing Page Test
One product brief. Two models asked to build a complete, styled landing page — hero, features, CTA, the works. Fully interactive. No frameworks.
The Debugging Test
A broken Python script with three real bugs. Two models asked to find them all, explain what's wrong, and ship a fixed version. Time to working code: very different.
The Summarization Test
One dense 800-word business article. Two models asked to distill it to the five things that matter — in under 150 words. Compression reveals what each model actually understands.
The Persuasion Test
One weak argument. Two models asked to make it airtight — anticipate every objection, sharpen every claim. One came back with a case. One came back with a lecture.
The Creative Brief Test
Same vague creative brief, three AI takes. See which model asks the right questions vs. dives in and which output you'd actually use.
The Hard Feedback Test
Give all three AIs a mediocre piece of writing and ask for honest feedback. Watch which model tells you the truth and which one wraps every critique in cotton wool.
The Trip Planning Test
One prompt, one destination, three itineraries. See which AI builds a day you'd actually want to live vs. a generic tourist checklist.
The Salary Negotiation Test
You got a job offer $15K below your target. Ask all three AIs to help you negotiate. See which one gives you a script you'd actually use.
The Explain-It-Simply Test
Quantum entanglement explained to a curious 12-year-old. No jargon, no hedging, no "it's complicated." See which model can actually teach.