Your AI can test 10,000 strategies a day. That's the problem.
Measured 2026-09-12 · 660,005 backtests · out-of-sample
The pitch is a loop that rewrites a losing strategy and retests it until the curve turns. Run that and you haven't found an edge — you've raised the bar you have to clear, and you've raised it on yourself. Here is that bar, drawn from 660,005 of our own backtests.
"AI trading" spans a lot: execution algorithms that shave basis points off large orders, sentiment models reading filings and calls, reinforcement learning for portfolio allocation. Those are real and this page is not about them. This page is about one specific claim — that an AI agent which generates strategies, backtests them, rewrites the losers and retests until a curve turns green has found you an edge. That claim is the one being sold hardest right now, it is the one a retail reader is most likely to buy, and it is the one where we have 660,005 of our own backtests to check it against. If you want the wider survey, this is not it; if you want to know whether the loop works, keep reading.
Every retest raises the score you need
Pre-register one idea and test it once and the bar is zero — there is nothing to correct for, because you did not go shopping. Every additional attempt raises it: 0.77 by ten tries, 1.31 by 760, 1.70 by 70,000. The curve is √(2 ln N) / √T — the selection hurdle already published in our methodology and applied to every verdict on this site.
Evaluated at a 7.7-year window. The formula is on the methodology page and it is applied to every verdict we publish.
A junior quant runs the research loop about once a month. An agent runs it while you sleep. The pitch treats that as pure gain — more shots, more winners. But the winner is picked from the pile, and the bigger the pile, the better the best one looks by luck alone. At ten tries you need 0.77. At the ~731 per asset we ran you need 1.31. At 10,000 a day for a week you need 1.70 — and at one pre-registered test you need nothing at all.
"It reads its own autopsy and rewrites itself until the curve turns" is not a description of research. It is a description of searching until you find noise shaped like signal. The honest version of that loop needs a hurdle that grows with every attempt, and an attempt counter that the agent cannot reset.
We ran the pile. Here's what got out.
903 assets x 382 indicators, run on whichever timeframes each asset has real history for — daily and weekly for most, plus 4-hour and hourly on the 37 with intraday data. That is 660,005 backtests, about 731 per asset, not the 1,379,784 a full cross would give. Already enough to need a hurdle of 1.31, and a fraction of what an always-on agent does in a day.
10 of 903 assets — 1.1% — produced a setup that beat buy-and-hold consistently and cleared the hurdle for how many times we looked. 190 more were consistent but couldn't be told apart from the luckiest of ~731 tries. 703 had nothing. Every verdict, per asset
Can you actually make money with an AI trading bot?
Every page ranking for this query poses that question. None of them answers it with a number. Here is ours, with the three things a performance claim is worthless without — the drawdown it took to get there, the costs charged, and what buy-and-hold did over the identical window.
| Measured across 541,051 out-of-sample backtests | Value |
|---|---|
| Median annual return, per setup | 1.60% |
| Median buy-and-hold return, same windows | 10.00% |
| Median maximum drawdown | −41.6% |
| Setups whose drawdown exceeded 20% | 90% |
| Setups whose drawdown exceeded 50% | 33% |
| Trading costs charged, every position change | 0.08% per side |
| Slippage modelled | none — so these are optimistic |
The typical setup in our grid earned 1.60% a year while buy-and-hold earned 10.00%, and it gave up 41.6% from peak to trough to do it. Nine in ten drew down more than 20%. One in three lost more than half. That is the honest answer to the question, and the reason no vendor publishes it: a win rate can be made to look like anything, but a drawdown paired with its benchmark cannot.
The survivors still mostly can't be levered into a win
Suppose the agent does find something real but small. The obvious move is to lever it up to something worth trading. Across the 541,051 of those backtests with at least ten trades, 5.0% beat buy-and-hold outright — and 26% would need more leverage than their own worst drawdown survives. They don't underperform the benchmark; they can't reach it at any leverage that keeps the account open.
None of this says AI can't trade. It says the loop being sold — generate, test, rewrite, retest, ship the first green curve — is the oldest mistake in quantitative finance with a faster engine bolted on. An agent that counted its attempts and raised its own bar accordingly would be genuinely new, and we'd publish that result the day someone shows the counter. Until then, ask any AI trading product one question: how many strategies did you test before you showed me this one? If it can't answer, the backtest means nothing.
Hypothetical backtests, costs included, no forward guarantee. See the disclaimer.