Markets Closed
Global Markets
S&P 500 7,411.98 ▲ +0.0% DOW 51,947.25 ▲ +0.5% NASDAQ 24,975.82 ▼ -0.6% RUSSELL 2K 2,930 ▼ -0.3% VIX 18.58 ▼ -0.6% GOLD 4,055.7 ▲ +0.2% CRUDE OIL 90.47 ▼ -1.9% EUR/USD 1.14 ▼ -0.0% BTC 64,101 ▼ -1.5% ETH 1,858.17 ▼ -1.2%
Fintech

AI agents fail job tasks in most real-work tests, Berkeley benchmark finds

UC Berkeley researchers found the best AI agent completed 26.2% of real workplace assignments, with harder tasks exposing steep limits.

Rafael Ortiz

By Rafael Ortiz · Fintech Correspondent

· 3 min read

AI agents fail job tasks in most real-work tests, Berkeley benchmark finds
Photo: PYMNTS

AI agents fail job tasks far more often than they complete them when tested on real workplace assignments, according to a new benchmark from UC Berkeley’s Center for Responsible, Decentralized Intelligence. The strongest system in the study, OpenAI’s Codex tool using GPT-5.5, finished 26.2% of the assignments correctly, a result that tempers claims that agentic AI is close to replacing broad categories of knowledge work.

The benchmark was designed to test predictions that AI agents could overtake humans across most jobs as soon as 2026 or 2027, Yiyou Sun, a core author of the work, told FrontierNews.ai. Researchers collected 1,490 assignments from more than 250 professionals in 55 industries, covering work such as drafting a legal filing, building a financial model and designing a manufacturing component.

According to the research paper, the tasks reflected projects that would ordinarily take people from hours to weeks. Systems were judged on whether they produced a complete and correct final deliverable. A partly useful answer did not earn credit if the final work product was absent or wrong.

How well do AI agents perform on real work tasks?

The headline result was uneven performance across difficulty levels. OpenAI’s Codex on GPT-5.5 led the overall ranking with a 26.2% success rate, according to the paper. On the most difficult tasks, where systems had to sustain work across many steps and avoid errors that could derail the final output, the average pass rate across tested systems fell to 2.6%.

Codex reached 8.6% on that hardest group. Anthropic’s Claude Code running on Opus 4.7 did not pass any of the hardest assignments, the paper found.

The researchers also tested whether far greater computing resources would materially improve outcomes. One version used $630 of compute and processed 763 million tokens, according to the paper, but passed 2.9% of tasks. Tokens are the text and data units that large AI models process when reading prompts, using tools and generating responses.

The study’s authors said those results suggest that adding more compute does not reliably solve the problems that appear in longer, higher-stakes assignments. For investors and executives, the finding points to a gap between impressive single-step outputs and dependable delivery of finished work.

Where did AI agents do better?

Performance improved on easier assignments. On a 59-task subset that current systems could partially address, leading models completed about 30% of tasks, according to the paper. Codex on GPT-5.5 reached 42.4% on that easier set.

The distinction is central to how the technology may be used in companies. An AI agent may answer a discrete finance or legal question, yet fail when asked to gather inputs, check intermediate results, apply judgment across several stages and produce a document or model ready for use.

An AI agent is software that uses a model to pursue a goal through multiple actions, often by calling tools, reading information and revising outputs. The Berkeley benchmark indicates that today’s agents remain more reliable on narrow, structured work than on open-ended projects that require sustained accuracy.

What does this mean for corporate AI adoption?

The findings align with a cautious pattern in corporate finance. PYMNTS Intelligence has reported that more than eight in 10 CFOs at large companies either use AI for accounts payable or are seriously considering it. A separate PYMNTS Intelligence study found that about 7% of U.S. finance chiefs have deployed AI agents in live finance operations, while another 5% are testing them before a broader rollout.

Those early uses are concentrated in repeatable work with defined rules, according to PYMNTS Intelligence. The Berkeley results suggest that companies evaluating AI agents may need to separate tasks that can be specified and checked from projects that require continuous judgment, accountability and a complete finished deliverable.

This story draws on original reporting from PYMNTS.

More from Fintech

All Fintech →