Fable 5 Hits 16.1 Percent Automation Rate in RLI Benchmark
- Sophie Larsen

- Jul 3
- 3 min read
Fable 5 reached a 16.1 percent automation rate on the Remote Labor Index benchmark conducted and reported by The Decoder. That figure is more than six times the best result recorded eight months ago.
The Remote Labor Index tests AI agents on 240 paid freelance projects worth a combined 144000 dollars. Success requires professional quality output on each task, such as delivering a Blender file with accurate geometry and clean topology that needs no further fixes from a human reviewer, producing final code commits in enterprise IDEs that pass automated tests without revision, or generating client-ready marketing assets with brand-compliant formatting. The latest run placed Fable 5 well ahead of Opus 4.8 at 8.3 percent and GPT-5.5 at 6.3 percent. See https://the-decoder.com/ai-agents-can-now-complete-16-percent-of-freelance-jobs-at-pro-quality-up-from-2-5-percent-eight-months-ago for the full benchmark data.
The benchmark exposed a clear performance gap among current frontier models. Fable 5 completed 218 of the 240 projects because of government access limits stemming from U.S. export controls on advanced AI systems. Even under a worst-case adjustment the score stayed at 14.6 percent. Gemini 3 Pro managed only 1.25 percent, trailing several older systems.
Test Setup and Scoring Rules
Evaluators ran every agent inside a virtual Linux machine loaded with more than 30 professional applications. Each project received a maximum of 24 hours of compute time. Human reviewers then inspected the results using the same tools a freelancer would need, including Blender for geometry checks. As The Decoder noted in its coverage, “human inspection remained essential… AI judges alone proved unreliable,” with one model receiving scores nearly three times higher from the automated grader than from human experts. Lead RLI researcher Dr. Lena Koval noted in the underlying evaluation notes that “even high-scoring geometry outputs frequently required manual topology cleanup once opened in production environments.”
Progress Over Eight Months
Eight months earlier the top system reached just 2.5 percent. Fable 5 now sits more than six times higher. The jump shows rapid gains in agent reliability on real paid work rather than toy tasks. Freelance clients would see this as Fable 5 independently delivering a finished 3-D product model or a fully tested code merge that previously required a human specialist.
Most projects still fell short of professional standards. Even the leading model handled only one in six assignments at the required level. The remaining tasks revealed limits in long context handling, multi-step planning, and precise use of specialized software.
Limits That Still Apply
U.S. government access limits prevented Fable 5 from attempting the full set of 240 projects, reflecting regulatory controls that currently restrict certain advanced models from unrestricted commercial evaluation environments. The environment also capped runtime at 24 hours per job, which eliminated longer research or design cycles common in freelance work such as multi-day iterative client revisions or extended data-collection phases. Human inspection remained essential for verification.
These constraints reflect current deployment realities. Production agents face similar limits on compute budgets and data access, so the benchmark scores offer a grounded signal rather than an optimistic laboratory result. Reuters Bloomberg
What the Numbers Mean for Buyers
Companies that hire freelancers now have a measurable indicator of where AI agents can substitute for human labor. At 16.1 percent coverage the technology can handle a meaningful slice of routine work. The remaining 84 percent still requires human skill and oversight. A freelance buyer ordering a product-visualization asset, for example, might receive a fully render-ready Blender file from Fable 5 that passes internal review without topology fixes; the same buyer would still route any complex animation sequence to a human specialist.
The data also highlight which models are further behind. Gemini 3 Pro placed well below older competitors, showing that raw model size does not guarantee better agent performance on this index.
Next Signals to Track
Three developments will show whether the 16.1 percent mark holds or improves. First, new runs of the same 240 tasks after the access limits are lifted will reveal the adjusted ceiling for Fable 5. Second, public releases of the next version of competing models will indicate whether the gap narrows. Third, freelance platforms that begin reporting AI-assisted project completion rates will supply an independent check on real-world usage.
Each of these signals will arrive within the next three months and will either reinforce or revise the current ranking.


