AI benchmarks hit ceiling—custom evaluations take center stage

The gist

AI models are acing old benchmarks but flunking real-world, long-horizon tasks—so companies are ditching public leaderboards for custom, expert-driven evaluations that actually predict useful performance.

What to know

Real-World Tasks, Real Shortfalls

Traditional benchmarks missed a critical flaw: even 'PhD-level' models faltered at basic coordination and practical execution, exposing a gulf between academic scores and genuine job-readiness.

By mid-2026, what counted as a meaningful benchmark had shifted because strong scores on older tests no longer settled the question of whether AI could do economically useful work over time. Reboot argued that models would have to prove “mastery of practical real-world tasks,” not just academic-style correctness, and highlighted the mismatch bluntly: many benchmarks still looked like “multiple choice quizzes,” even as supposedly “PhD-level” systems remained unreliable at basic but valuable coordination work such as booking hotels or flights.

That change became visible in the new benchmark designs that scaled during 2026, which were built to expose the gap between headline capability and job-relevant performance. Reboot pointed to OpenAI’s GDPVal as a sign of the new standard: it “broke down 44 occupations into component tasks, hired seasoned professionals to create example prompts, and judged how far AI models are from human experts,” a structure that made persistent weaknesses on long-horizon, economically valuable tasks harder to hide behind saturated legacy scores.

Sources
Reboot

Leaderboards Lose Their Luster

As top models plateaued on legacy tests, industry insiders admitted that tiny score bumps masked models’ inability to deliver reliable, economically valuable results.

What broke first was confidence that old scores still meant much. Understanding AI argued that “conventional benchmarks like MMLU have a natural lifecycle,” and once scores crowd the ceiling, “MMLU has saturated”; worse, test-set flaws mean “it’s impossible to get a score much higher than 93%” without drifting into noise or cheating. OpenAI made the same point more bluntly in coding: Olivia from the Frontier Evals team joked that labs “increment like 0.1” on leaderboards and then claim the best model, even though that edge is “not super convincing at this point at all.”

The deeper problem was that narrow evals were not reliably measuring deployable usefulness. OpenAI said SWE-Bench Verified was itself “a cleanup of original bench academic benchmark,” after finding some agent failures came from “bad problem setups rather than just to models being dumb,” while critics of METR warned that its graph rests on “the length of time it takes humans to carry out software engineering tasks that models can successfully complete 50% of the time,” a threshold that can mislead even when raised to 80%. That mattered because, as AI Engineer noted, economic value requires reliability “95 99% of the time” so workers can “tab tab tab through” outputs instead of constantly rechecking them.

Sources
Latent SpaceUnderstanding AIAI EngineerTransformer

Long Tasks, Low Success Rates

Despite massive technical gains, frontier AIs flunked the toughest long-horizon benchmarks, with most models failing over 97% of real-world, expert-sourced tasks.

Agents’ Last Exam is hard proof that frontier systems are not yet consistently handling the kind of work companies would actually pay for. As Hugging Face Daily Papers describes it, the paper “introduces Agents’ Last Exam (ALE), a benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes”; it was “developed in collaboration with 250+ industry experts,” “covers non-physical industries defined with reference to O*NET / SOC 2018,” and “covers… 55 subfields grouped into 13 industry clusters covering 1K+ tasks,” organized around “a task taxonomy.” More specifically, ALE measures “sustained, economically valuable work,” with “1,500+ expert-sourced tasks spanning 55 occupations.”

The results are stark precisely because the tasks resemble sustained professional execution rather than short benchmark puzzles. Hugging Face Daily Papers reports that “Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is 2.6%,” while Agentic AI Weekly says Berkeley RDI, working with “300+ global experts across 55 industries” to test “real digital labor-market work,” concluded: “the age of useful agents is here. The age of truly job-ready agents is not,” with “most frontier agents we tested, including Fable 5, achieved a 0% success rate.” The hardest tasks “require sustained reasoning, deep domain expertise, and reliable execution over long” horizons.

Sources
Hugging Face Daily PapersAgentic AI Weekly

Custom Evals for Every Workflow

Enterprises now demand workflow-specific evaluations, exposing how public leaderboards overlook the unique challenges and performance needs of real business processes.

Custom evals are spreading because enterprises have learned that public benchmark strength does not tell them how a model will behave inside a specific workflow. One Useful Thing argued organizations need to know “specifically what YOUR AI is good at, not what AIs are good at on average,” while Eye on AI captured why buyers stopped trusting generic scoreboards: “a new eval come out every year… and then people realize it's no longer relevant,” citing LM1B’s 2011 news-source setup as an example of a once-prominent test that no longer predicts targeted deployment performance.

That is pushing both vendors and benchmark builders toward evaluations that mirror the business process itself, with repeated scenarios, expert review, and domain-specific comparisons. Alex Ratner framed this as an “evaluation gap” for “real enterprise work” and walked through Snorkel AI’s $3M Open Benchmarks Grant as part of fixing it, while LILT’s multilingual enterprise leaderboard showed why workflow context matters: “in coding, GPT 5.5 performs best in Spanish, Claude Opus 5.5 wins in Japanese, while Muse Spark 1.3 leads in Serbian,” a pattern no generic leaderboard would surface for a global operation.

Sources
One Useful ThingEye on AIChain of Thought | AI Agents, Infrastructure & EngineeringPR Newswire - Consumer Technology

Expert Grading Sets New Bar

The next wave of benchmarks uses curated, outcome-based grading by professionals, revealing stark performance drops when AI is tested on end-to-end, real-world tasks instead of narrow academic challenges.

What distinguishes the new benchmark wave is not just harder questions, but tighter measurement design. According to Last Week in AI, OpenAI introduced “GDPval… evaluating AI against human professionals across nine high-GDP industries and 44 occupations,” where “experienced professionals directly compared AI-generated deliverables with peer-produced” work and chose a winner, producing a verifiable win-or-tie outcome; that structure matters because it spans healthcare, finance, manufacturing, and government, with roles from software engineers and nurses to journalists and investment bankers, instead of collapsing practical ability into one generic public leaderboard score.

The same pattern appears in newer domain benchmarks that trade breadth-without-structure for curated suites with explicit grading rules. Gradient Ascent notes that Claude 4.5 Opus scores 74.4% on SWE-bench but drops to 11.0% on FeatureBench because SWE-bench “tests bug fixes within single pull requests,” while FeatureBench “evaluates end-to-end feature development”; meanwhile RISE-Video “shifts evaluation from aesthetics to cognitive understanding” using “467 human-annotated samples spanning eight categories,” and ALE adds “1,500+ expert-sourced tasks… with outcome-based, verifiable evaluation” across 55 industry domains, a design that is “longer-horizon… and harder” than Terminal-Bench or SWE-bench-Pro—“Compared to Terminal-Bench and SWE-bench-Pro… longer-horizon… and harder (the best agent passes just 25.2%, vs.”—than Terminal-Bench or SWE-bench-Pro.

Sources
Last Week in AIGradient AscentAgentic AI Weekly

Get the stories behind the trends

Deep-dive reporting and the weekly brief, in your inbox.