AI Super Simplified
Edition 289

AI Agents Can Now Finish 15.8% of Real Freelance Jobs — Up From 2.5% Eight Months Ago | Edition 289

Edition 289 — 240 real freelance jobs, $144,000 on the line, human-graded. What splits the 15.8% that pass from the 84% that fail.

By Jerry Croteau
AI Agents Can Now Finish 15.8% of Real Freelance Jobs - AI Super Simplified Edition 289

Last edition we covered what people actually ask AI chatbots — mostly requests for help, not handovers. This one's the other half: not what people ask, but what AI can actually finish on its own, for money, when a real client is grading it.

The number: 15.8%. That's the share of 240 real, paid freelance projects — worth a combined $144,000 — that AI agents completed at a quality a paying client would accept, per the newest Remote Labor Index results. Eight months ago, the best any AI could manage was 2.5%. Roughly a sixfold jump.

Take either headline: AI's ability to finish real work more than sextupled in eight months, or 84% of real, paid freelance work still comes back unusable. Both are true — and the second one is the useful story, because the line between finishing and botching isn't random. It falls in a specific, checkable place.

Who's actually keeping score

Two organizations run the Remote Labor Index, with different incentives worth knowing up front. The Center for AI Safety (CAIS) is a nonprofit focused on reducing AI risk — no product to sell. Scale AI sells human-labeled data and AI evaluation services for a living, and its research arm runs the actual grading here. Not a reason to dismiss the numbers — the underlying paper has 47 named authors and is public on arXiv — but a real commercial stake in a benchmark about how close AI is to replacing paid human work.

The setup: 358 verified freelancers, averaging 2,341 hours worked and $23,364 in career earnings, contributed 550 real projects, narrowed to a final 240 spanning 23 categories — 3D/CAD, architecture, design, video, audio, data analysis, web apps. Every price and completion time comes from the professional who did the work.

Grading is entirely human: evaluators compare each AI's deliverable against the professional's original from a “reasonable client” standpoint, settled by majority vote. An automated AI judge overestimated automation rates by 2–3x on the newest models — judging the work needs the same computer-use skill the AI workers are tested on. So they stuck with people, agreeing 94.4% of the time.

Tap a category to see the Remote Labor Index's own published verdict and evidence — no invented pass rates. · Open full-screen ↗

Where the 84% actually breaks down

Scale AI's review of failed submissions found repeat patterns, not randomness: 45.6% below professional quality, 35.7% incomplete (truncated videos, missing files), 17.6% technical (corrupt or empty files), 14.8% inconsistent across files — these overlap, since one bad submission often hits several. Example: told to swap a diamond's cut in a ring photo, an agent ignored the reference image and generated two brand-new rings from scratch — a quality, brief-following, and consistency miss, all in one project.

The newest results, published July 1, show the same shape at a higher level. CAIS's review of Fable 5's redesigned engagement ring calls it “qualitatively much better than deliverables from previous AIs” and still unprofessional — a low-effort prong design. On a renovation-render project, GPT-5.5's photorealistic bathroom render looked finished, except it was faked with an image generator instead of rendered from the actual 3D model.

That's the boundary, eight months of progress later: agents generate something plausible from a blank page far better than before, and are still unreliable at following a brief exactly, revising without breaking things, and not quietly faking work that only looks right at a glance.

ModelAutomation rate
Gemini 3 Pro1.25%
Grok 42.08%
GPT-5.22.50%
Manus 1.6 Max2.92%
Opus 4.64.17%
GPT-5.56.3%
Opus 4.88.3%
Fable 5 (newest)15.8%
Every published Remote Labor Index score to date (human-graded automation rate)

None of this supports “robots are taking freelance jobs,” and it doesn't support “AI can't do anything” either — both were true at 2.5% and neither is true at 15.8%. It supports a specific rule for anyone deciding whether to hand a real project to an agent this month: the safer bets are generate-from-scratch work — a first draft, a logo, a sound effect, a rough data pull — where “good enough to revise” is the bar. The riskier bets are anything requiring a precise spec followed exactly, editing without breaking things, or operating unfamiliar software the way a professional would. Ask which kind of job you're handing over before you trust what comes back.

Sources: Remote Labor Index paper (arXiv:2510.26787) · Scale AI, “The Remote Labor Index,” Oct 29, 2025 · CAIS, “A Significant Increase in Digital Labor Automation,” Jul 1, 2026.