The confluence score measures how much of the checklist agrees. Confidence is a different question — what actually happened to comparable setups — and it frequently sits well below a high score. It is built by nested shrinkage: a replay curve first, then the platform’s record, then your own.
Two different questions
A high score with a modest confidence is the system working. The score is an agreement measure computed from the chart in front of it. Confidence asks what became of setups that looked like this one, which is a question the chart cannot answer by itself.
The verdict always names whose record a confidence number came from — yours, the platform's, or a replay's — rather than describing it with an adjective. "Moderately confident" tells a reader nothing they can check.
How it is built
Three sources are blended by shrinkage, each given a weight reflecting how much it deserves to be trusted: a backtest curve from replayed history, the platform's own resolved outcomes, and your personal bins once you have enough resolved trades for a given asset class.
Shrinkage means a small sample cannot shout. Your first few resolved trades nudge the number; they do not replace it. As your record grows it takes over, per asset class rather than all at once.
Floors before a source counts
The platform curve does not activate until enough outcomes have actually resolved — a pooled floor overall and a higher one per asset class. Below those floors the observed rates are shown but are not used to move anyone's confidence.
This is why most published cells currently read as collecting. The honest state of a young dataset is empty cells, and the product shows them empty rather than filling them with a number that has not earned its place.
Why confidence sits low today
The backtest anchor comes from a large walk-forward replay of historical setups, and what it found was that roughly the same fraction of setups reached target first almost regardless of score. That result is the reason confidence publishes conservatively rather than tracking the grade.
The basis is also labelled precisely: the replay measures simulated fills, not resolved live outcomes, and the verdict says so rather than letting a replay borrow the authority of a live track record.
Frequently asked
- Why is the confidence number so much lower than the grade suggests?
- Because it is checked against outcomes rather than against the checklist. The replay anchor found hit rates that barely varied with score, so confidence publishes conservatively until live resolved evidence outranks the replay it is currently leaning on.
- Whose track record is the confidence number based on?
- The verdict names its basis every time — model, backtest, platform, personal or blended. Your own bins take over per asset class once enough of your trades have resolved, and shrinkage keeps a small personal sample from dominating before it has earned it.
Updated Sep 1, 2026