Skip to main content

AI Agents Are Flunking the Stock Market's Most Realistic Test Yet

A new benchmark that simulates actual trading workflows has every leading AI model scoring below 60%. The gap between knowing and doing has never been clearer.

The Old Way of Testing AI Is Dead

For years, we judged AI the way we judge students: give it a pile of math problems, a few essays, some code snippets, and tally the score. High number? Smart bot. That made sense when AI lived in a chat window. But now AI is being asked to do real work—work that involves messy systems, time pressure, and consequences for mistakes. Nowhere is that more true than in the stock market, where one wrong decimal can cost millions.

We're used to AI suggesting stocks or summarizing earnings calls. But behind the scenes, the real ask is heavier: parse thousands of filings, cross-reference data, execute trades, manage portfolios. The question has shifted from 'Can it answer?' to 'Can it finish the job?'

Meet RealReplicaBench, the Test That's Humiliating Every Model

Enter RealReplicaBench, built by the team behind Accio Work, an AI agent platform for e-commerce. But its implications hit finance hard. The test throws 107 realistic commercial tasks at the models: reconciling supplier invoices, managing logistics, processing customs records. The results are brutal. Not one model scored above 56.1 out of 100. That's an F. The best performer, Claude Opus 5, managed just 66 out of 107 tasks—and only when running on a specific framework. Gemini? Dead last.

Why so low? Because RealReplicaBench doesn't give partial credit. There's no 'almost' in completion. If the task isn't fully done, it's a zero.

Why Standard Benchmarks Miss the Point

Think of old benchmarks like a written driving test. You can ace questions about signaling and right-of-way, but that doesn't mean you can parallel park in downtown traffic. In the stock market, the gap is even wider. Knowing the definition of a P/E ratio is one thing; building a financial model from scratch, pulling live data, and executing a multi-leg options trade is another.

RealReplicaBench simulates the latter. It doesn't just present a problem in text. It creates a full simulated environment: a browser, a command line, APIs, file systems, and changing backend states. The agent has to click buttons, fill forms, deal with pop-ups—just like a human—while keeping track of a goal that might span hours.

The Stock Market Twist: It's All About Workflow

Here's where it hits close to home. The benchmark's tasks aren't just e-commerce; they're the same kind of multi-step, high-stakes workflows you'd find in financial operations. One task, for example, requires an agent to process 5,383 customs records, aggregate them by category and supplier, filter for the top three suppliers based on a policy, and then create dashboards and project tasks across Google Workspace, Box, and Jira. Sound familiar? It's basically a junior analyst's nightmare.

Another task involves booking a shipment from China to the U.S., calculating all costs—ocean freight, trucking, insurance, customs, bonds—and then excluding routes that exceed 30 days or have invalid port connections. The catch: the agent isn't just asked to suggest a route. It has to actually create a booking and generate a Shipment ID. If it doesn't, the task is marked incomplete. In stock market terms, that's like being asked to execute a trade and then just writing a note saying 'I would buy.' No, you have to place the order.

Why 'Looks Done' Isn't Good Enough

One of the most refreshing things about RealReplicaBench is that it doesn't trust what the AI says. Instead, a verifier reads the final state of the environment. Did the file get created? Did the booking ID appear? Did the calendar event show up? If yes, you pass. If no, you fail. This is a huge shift from the old way of evaluating AI, where you'd just look at the text output and judge if it seemed right.

In the stock market, this is exactly what you want. An AI that says 'I've placed a buy order' but hasn't actually done it is worse than useless. You need proof—a confirmation number, a timestamp, a record in the ledger. RealReplicaBench forces that kind of rigor.

What This Means for Stock Market AI

So why should you care? Because if AI can't handle these simulated e-commerce tasks, it's nowhere near ready for the stock market's even more complex, faster-moving workflows. The benchmark's 107 tasks were distilled from 1.6 million real conversations and 200,000 execution traces. That's not lab-made fluff; it's the grunt work of actual business. And the failure rate is a wake-up call.

The good news is that the benchmark also shows what's needed to improve. The Accio Work team found that the framework you use matters as much as the model. For instance, Claude Opus 5 performed better on their Accio Work framework (61% pass rate) than on others (56% and 60%). That suggests that building better 'harness'—the scaffolding that connects models to tools and data—is a key to unlocking real-world competence.

The Future Is About Getting Things Done

The takeaway is simple: we're moving from an era of 'knowing' to an era of 'doing.' The stock market doesn't pay for a model that can recite financial definitions. It pays for a model that can complete a trade, file a report, or reconcile an account without leaving the last 20% for a human to finish. RealReplicaBench is a step toward making that a measurable, testable reality.

For now, the scores are ugly. But that's not a reason to despair; it's a reason to demand better. As the team behind the benchmark says, 'The ability to build a harness will become as important as the underlying model.' For anyone building or buying AI for the stock market, that's the lesson to take home: don't settle for an agent that's 'almost' done. Hold out for one that can actually deliver.

Share this article:

Comments (0)

No comments yet. Be the first to comment!