The idea in one breath: Ordinary look-ahead bias is a bug in your data pipeline and you can fix it. This one is not in your pipeline. A model trained on text up to 2025 has already read what happened to every liquid instrument through 2024, so a backtest over that period is asking a question it knows the answer to — and no amount of careful data handling on your side touches it.
You know the classic version. Your backtest accidentally uses the closing price to decide an entry at the open, or a symbol list that only includes companies still listed today, and the results look wonderful until they meet a live market. Both are bugs in your code and both are fixable by someone careful.
Now put a language model in the loop. You feed it a chart or a description of one, dated March 2023, and ask what it would do. You have been scrupulous: no future bars, no forward-filled indicators, a clean walk-forward. And the answer is still contaminated, because the model read the news archives, the earnings recaps, the forum posts and the retrospectives. It knows what March 2023 turned into. That knowledge is not in your dataset. It is in the weights.
Why this is worse than the ordinary kind
- It is invisible. There is no column to drop and no join to fix. The leak has no location in your code.
- It is not consistent. The model recalls famous moves better than quiet ones, so contamination is heaviest exactly where the price action was most dramatic — the periods a backtest leans on hardest.
- It survives paraphrase. Anonymising the ticker helps less than you would hope, because an instrument is identifiable from its own price path if the move was distinctive enough.
- It gets worse as models improve. Better recall of the training corpus is a feature everywhere else and a contaminant here.
What the research finds when it is controlled for
The interesting work in this area builds memory-controlled benchmarks — tests designed so the model cannot benefit from having read the outcome. The consistent finding is that apparent skill falls sharply once the memorised history is excluded. Not to zero in every case, but far enough that results reported without the control should be read as an upper bound rather than an estimate.
Two defences, and only one is practical
- Test only after the cutoff. Restrict evaluation to data published after the model’s training cutoff. Correct, and it leaves you months of history rather than years — which for most strategies is a sample too small to conclude anything from, so you have traded one problem for another.
- Keep the model out of the replay loop entirely. Let it do the things that do not require it to be ignorant — summarising a filing, extracting a symbol from an image, drafting the reasoning for a decision something else made. Let a deterministic rule make the call that gets scored. This is the practical answer for almost everyone.
The second option is not a compromise. It is a division of labour that happens to make the system testable: whatever the deterministic part decides can be replayed against historical bars honestly, because arithmetic has not read the news.
Order the steps of designing an LLM strategy test that this contamination cannot quietly inflate. First step first.
- 1Replay the deterministic decisions against full history, which is now safe because arithmetic read no news.
- 2Report the post-cutoff sample size beside the result, because it is the number that limits the claim.
- 3Establish the model’s training cutoff, and treat it as a hard boundary rather than a guideline.
- 4Restrict any evaluation the model does influence to data published after the cutoff.
- 5Move the scored decision to a deterministic rule the model does not influence.
- 6Decide which part of the system the model is allowed to be: describer, extractor, or decider.
Part of Track 15 · AI & Automated Trading — see the full syllabus.
MCP, and what a broker opening one means
MCP is a standard way for a model to call somebody else’s software. When a broker opens one, it is publishing a menu of operations an agent may perform on your account. The interesting part of any such announcement is never the protocol — it is which operations are on the menu, and whether placing an order is one of them.
Continue the track