This is the first of a recurring report: what the grading engine actually saw over the period, taken straight from the event log, published whether or not it flatters us. It is also the least impressive one we will ever publish, because almost every cell in it currently reads collecting.
The format
Every edition answers the same four questions, in the same order, from the same queries. Fixing the questions in advance is the point — a report that chooses its metrics after seeing the data is a marketing asset wearing a lab coat.
- How did each grade band resolve? Broken out A+ through F, so the claim that a higher grade means a better outcome is checkable rather than assumed.
- Where was the engine most wrong? The setup family with the widest gap between grade and resolution, named.
- What did traders pass on that worked? Logged passes are decisions, and the cost of the good ones is the most useful number the journal produces.
- What changed in the engine, and did it help? Every scoring change from the changelog, with the before-and-after on the affected cells.
Question 1 — resolution by grade band
Nothing here has cleared the sample floor, so nothing here shows a number. The floor exists because a resolution rate on two hundred observations tells you almost nothing about whether an edge exists — the methodology page covers why, and what the threshold is.
Question 2 — where the engine was most wrong
Answering this requires the same resolved-outcome data as question 1, so it has the same answer: not yet. When it is answerable, the setup family named here will be named on its own Playbook page too, with the same figure. A number that appears in a blog post and not on the page it describes is a number chosen for the blog post.
Question 3 — the cost of good passes
This one is different, and worth explaining because it is the metric we most want to publish. When a trader logs a pass, ORIN keeps watching the setup. If it resolves to target, that pass had a cost — and the aggregate of those costs is a far more honest measure of a filtering tool than any hit rate.
It is also the number most likely to embarrass the product, since a tool that talks people out of winners is worse than useless. Publishing it anyway is the only reason to believe the rest.
Question 4 — engine changes
Covered in the changelog as it happens. In this period: the scoring engine has not changed, so there is no before-and-after to report.
3 posts · all of them