The idea in one breath: Ask a general model to read a chart and it will produce fluent, specific, confident prices. Some of them will be approximately right and none of them were measured. It is inferring plausible numbers from context, which is a different operation from reading an axis, and the tell is that the errors are small enough to look like rounding.
Upload a screenshot to a general assistant and ask where support is. You will get a paragraph naming a level, describing the structure, and suggesting where a stop belongs. It reads exactly like analysis. The question worth asking is where the number came from.
Describing versus measuring
Measuring means reading the axis, mapping pixel positions onto prices, and reporting what is there. Describing means producing the sentence a competent analyst would most likely write about an image like this one. Those two processes agree often enough to be confusing and they fail completely differently.
A measurement error is bounded by resolution: you can be off by a pixel. A description error is bounded by nothing except plausibility, so a described level lands in the right neighbourhood and drifts — often by an amount that would not survive contact with a stop.
Why the errors are seductive
- They are specific. A wrong price with two decimal places reads as more careful than a right price rounded to the nearest figure.
- They are close. Being wrong by a fraction of a percent looks like precision, and on a tight stop it is the whole trade.
- They are confident. Nothing in the output distinguishes a number that was read from one that was inferred, because the model does not distinguish them either.
- They are consistent. Ask twice and you often get the same wrong number, which feels like corroboration and is only the same process running again.
The architecture that avoids it
The fix is not a better model. It is refusing to let the model be the source of any number that matters. Have it read what a chart states in words — the ticker in the title bar, the timeframe on the selector — and hand that off to something that fetches the actual candles. The instrument name is a label the model can read; the price is a measurement it cannot.
That is exactly what ORIN’s measured read does with an uploaded screenshot, and the restriction is enforced rather than requested: the response schema admits only transcriptions, any level the model volunteers is stripped before anything downstream sees it, and a read that mentions support, resistance or a target is rejected outright instead of edited. ORIN also ships the other kind, deliberately and separately — the Upload beta hands the whole picture to a vision model and shows you what it says. Nothing that read produces is graded, journalled, or counted in any published figure, and every screen it appears on says so. The point of this lesson is not that the second kind is forbidden; it is that you have to know which one you are holding.
Four AI-generated chart reads. Three have ordinary problems. One contains the specific failure this lesson is about — a level that was described rather than measured, and is wrong in the way that costs money.
Part of Track 15 · AI & Automated Trading — see the full syllabus.
The look-ahead problem lives in the weights
Ordinary look-ahead bias is a bug in your data pipeline and you can fix it. This one is not in your pipeline. A model trained on text up to 2025 has already read what happened to every liquid instrument through 2024, so a backtest over that period is asking a question it knows the answer to — and no amount of careful data handling on your side touches it.
Continue the track