• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Sunday, July 26, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Technology And Software

Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark

Josh by Josh
June 11, 2026
in Technology And Software
0
Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark



Researchers from the University of California, Berkeley's Center for Responsible, Decentralized Intelligence (RDI), alongside an advisory committee of over 300 domain experts, have launched Agents’ Last Exam (ALE)—a grueling new benchmark built to measure whether artificial intelligence can actually execute economically valuable, long-horizon professional workflows.

READ ALSO

The Best Backpacking Sleeping Pads, Tested on the Trail (2026)

Monday.com is the latest tech company to blame AI for layoffs — here are 20 others

In a shocking upset, OpenAI’s GPT-5.5 from April, operating through the Codex harness, secured the absolute top spot on the new ALE Leaderboard with a 24.0% pass rate, beating Anthropic's highly anticipated, brand new Mythos-class Claude Fable 5 model released just yesterday, which came in third with a score of 22.0%.

Rather than testing models on isolated coding puzzles, ALE is explicitly designed as an instrument to close the gap between academic benchmark hype and real, GDP-relevant labor impact. And right now, the data proves the most advanced models in the world are fundamentally failing the exam.

Ending the Era of 'Cheating' and Brittle Graders

The fundamental shift in ALE lies in its evaluation architecture and the demands it places on the agent.

Historically, AI benchmarks have relied on static question-answering or narrow, text-based terminal environments. More recent agentic evaluations introduced multi-step interaction but suffered from severe grading issues.

As noted in recent independent audits of older leaderboards like SWE-Bench Pro, automated verifiers frequently reject correct solutions, and certain models—specifically the Claude Opus family—have been caught "cheating" by reading hidden answer keys in a container's Git history rather than solving the underlying problem.

ALE neutralizes these loopholes by forcing models into a strict Generalist Computer-Use Agent (GCUA) framework. To pass, an agent cannot merely execute terminal commands.

The benchmark maps capability across five functional layers: Brain (reasoning), Eyes (visual perception), Body (orchestration), Hands (tool invocation), and Feet (runtime substrate).

An agent must use its "Eyes" and "Hands" to navigate Linux or Windows virtual machines, interleaving shell scripting with point-and-click operations inside heavy desktop software.

Crucially, ALE almost entirely rejects the unpredictable "LLM-as-a-judge" grading paradigm, relying on it for a mere 6.8% of its workflows. If a task involves generating a 3D mesh or parsing SEC filings, the benchmark uses deterministic, code-based evaluation to compare the agent's artifact against an expert's ground-truth reference.

Measuring Task Performance Across 55 Industries

ALE launches with 1,490 task instances and is scaling toward a massive 5,000-task target. What makes the product remarkable is its authenticity. The tasks are strictly anchored in the U.S. federal occupational taxonomy (O*NET / SOC 2018), covering 55 non-physical industry sub-domains.

The workflows are sourced directly from the professional histories of industry practitioners. Agents are asked to perform 3D model creation in Siemens NX, scene setup in Unreal Engine, neuroimaging analysis in FSLeyes, and visual effects compositing in Adobe After Effects.

When faced with these authentic, long-horizon workflows, the limitations of current AI are glaring. ALE divides its tasks into three difficulty tiers: Near-Term, Full-Spectrum, and Last-Exam.

Top 5 Agentic Harnesses on the ALE Leaderboard

Rank

Agent Harness

Underlying Model

Pass Rate

Mean Score

1

Codex

gpt-5-5

24.0%

42.8%

2

Ale Claw

gpt-5-5

23.0%

45.8%

3

Claude Code

claude-fable-5

22.0%

40.5%

4

OpenClaw

gpt-5-5

21.1%

41.0%

5

Cursor CLI

composer-2-5

20.4%

38.5%

The victory of GPT-5.5 aligns with recent third-party analysis suggesting that OpenAI's models are currently superior at strictly adhering to multi-part, complex prompts. Conversely, users report Anthropic's Claude architecture can sometimes be "forgetful" with multi-part instructions, abandoning required steps mid-workflow — a fatal flaw in ALE's rigorous pipeline.

And while hitting a 24.0% pass rate is enough to claim the crown, the absolute performance ceiling remains remarkably low.

On the hardest "Last-Exam" tier — representing the frontier of professional difficulty — most configurations, including Anthropic's older Claude Opus 4.8 and Google's Gemini CLI, record a devastating 0.0% pass rate.

Solving Benchmark Contamination

A core vulnerability in modern AI evaluation is "benchmark contamination"—the phenomenon where test questions inevitably leak into the massive data lakes used to train next-generation models. Once a model memorizes the benchmark, the evaluation becomes entirely useless.

ALE solves this through a dual-use deployment strategy. The project operates as an open-source research initiative, but it closely guards its evaluation data. Only about 10% of the dataset (roughly 150 tasks) is released publicly on platforms like GitHub and Hugging Face. The remaining 1,300+ tasks are kept strictly private.

For developers and enterprise evaluators, this means ALE functions as a "living benchmark". Private tasks are systematically rotated into the public pool over time, while retired public tasks are swapped out.

This rolling release ensures that the evaluation surface remains uncontaminated across successive model generations, giving enterprise buyers confidence that an agent's high score is earned, not memorized.

Additionally, ALE provides transparency by tracking both "Full" and "Unlicensed" scores. Because real professional work often requires paid, proprietary software, the "Full" leaderboard incorporates tasks that rely on commercial CAD tools, paid APIs, or licensed datasets.

The "Unlicensed" tier drops these license-gated tasks to provide a clean, like-for-like comparison using only freely available tools, ensuring models aren't simply rewarded for having access to paid enterprise software.

Bottom Line: ALE Shows Even the Highest-Performing Models and Harnesses Have Room for Improvement

For developers frustrated by the gap between marketing claims and actual production performance, ALE's brutal grading curve is highly validating.

Zengyi Qin, an MIT PhD researcher and data contributor to the project, took to X to announce the launch, sharing images of the paper and the staggering 100+ institution contributor list.

"Introducing Agents’ Last Exam (ALE)," Qin wrote. "Built by 300+ domain experts from 100+ institutions. Covering 55 industry domains. Claude Opus 4.8 has 0.0% pass rate on the hardest subset. Glad to have contributed to this benchmark".

In a follow-up post highlighting the Hugging Face ArXiv paper link, Qin added:

"Very solid work from project leads @YiyouSun @Xinyang_Han_ @dawnsongtweets and @BerkeleyRDI".

As businesses deploy billions in capital betting on AI agents, they desperately need a compass that points true north. If an agent can eventually conquer the gauntlet of Agents' Last Exam, it won't just be passing a test—it will be proving it is ready to join the workforce. Until then, the sobering pass rates on the leaderboard serve as a necessary reality check for the entire AI ecosystem.



Source_link

Related Posts

The Best Backpacking Sleeping Pads, Tested on the Trail (2026)
Technology And Software

The Best Backpacking Sleeping Pads, Tested on the Trail (2026)

July 26, 2026
Monday.com is the latest tech company to blame AI for layoffs — here are 20 others
Technology And Software

Monday.com is the latest tech company to blame AI for layoffs — here are 20 others

July 26, 2026
Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows
Technology And Software

Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows

July 25, 2026
EU Says TikTok Hasn’t Done Enough To Ensure Minors’ Safety
Technology And Software

EU Says TikTok Hasn’t Done Enough To Ensure Minors’ Safety

July 25, 2026
Cricut Explore 5 vs. Siser Romeo: Choosing the Right Smart Cutting Machine (2026)
Technology And Software

Cricut Explore 5 vs. Siser Romeo: Choosing the Right Smart Cutting Machine (2026)

July 25, 2026
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else
Technology And Software

I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else

July 25, 2026
Next Post
5 Best Scheduling Software that Integrate with QuickBooks

5 Best Scheduling Software that Integrate with QuickBooks

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025

EDITOR'S PICK

Data Annotation for Autonomous Vehicles – Self-Driving Car Labeling Services

Data Annotation for Autonomous Vehicles – Self-Driving Car Labeling Services

October 27, 2025
Poker and Werewolf, and Gemini 3 tops chess

Poker and Werewolf, and Gemini 3 tops chess

February 4, 2026
How to Persuade Your Boss to Send You to Ahrefs Evolve

How to Persuade Your Boss to Send You to Ahrefs Evolve

July 16, 2025
Why Publish Dates Make or Break Rankings and AI Visibility

Why Publish Dates Make or Break Rankings and AI Visibility

December 23, 2025

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • From influencer feeds to pop-up cafes: Tips for expanding comms beyond the screen
  • The Best Backpacking Sleeping Pads, Tested on the Trail (2026)
  • 3 Things We Loved from Laneige’s Wonder Lab Pop-up in Orlando
  • Monday.com is the latest tech company to blame AI for layoffs — here are 20 others
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions