ORIN
Sign in
ORIN Labs · Research

Primary research on AI, charts, and how traders actually decide.

Most of what this industry states as fact traces to three academic papers, the newest from 2019. ORIN Labs runs and publishes original studies instead — with the method stated up front, the sources listed, the data open where we can open it, and every finding dated so you can see when it was last true.

Published

Why we run these studies

Five questions traders, brokers and journalists ask constantly, and the state of the public evidence for each. Every source is listed at the bottom of this page.

01

The category’s vendors publish no original research

The analysis products sold into broker platforms ship technical commentary, trade ideas, and signal reports. None of the major vendors has released a data study, a benchmark, or an annual industry report. There is product, and there is marketing, and between them there is nothing a reader can check.

What it costs a reader A trader or a compliance officer asking whether any of this works has no primary source to read — only vendor claims and the commentary those vendors sell.

02

Almost every trading statistic in circulation comes from three papers

Nearly every “day trading statistics” page traces to the same three sources: Chague, De-Losso & Giovannetti’s 2019 study of Brazilian equity futures, Barber and Odean’s Taiwan work, and Jordan & Diltz (2003). Hundreds of pages aggregate them. No major fresh primary dataset has been published in seven years.

What it costs a reader The numbers people quote about their own odds describe different markets, different instruments, and in one case a market structure two decades old.

03

The academic work on LLM trading has no practitioner translation

ArXiv now hosts live LLM trading arenas (DeepFund, AI-Trader), agent-reliability audits (TradeTrap), and strategy-design benchmarks (AlphaForgeBench, Market-Bench). It is dense, fast-moving, and written for other researchers. Researchers themselves note how little of it evaluates models on the thing traders actually do.

What it costs a reader The evidence that exists is unreadable to the people it is about, and the specific question of whether a vision model can read a chart is barely examined at all.

04

Prop-firm performance data all traces to publishers with a stake in it

The evaluation-firm niche looks well covered, but nearly every article aggregates a single dataset — FPFX Technology’s 300k+ accounts — and the sites publishing it are evaluation firms or affiliates selling challenges. The industry openly acknowledges that firms have a financial incentive not to publish real pass and payout data.

What it costs a reader Someone deciding whether to buy a challenge is reading numbers supplied, almost without exception, by the people selling it.

05

The retention claim made to brokers has never been measured in public

Analysis vendors sell brokers on the premise that embedded analytics increase trader lifetime, deposits, and volume. The only primary broker-side dataset in general circulation is a Devexperts survey of trader preferences. The premise itself has not been quantified anywhere public.

What it costs a reader A broker evaluating this category is asked to accept the central claim on assertion, with no measurement to test it against.

What we are publishing

10 studies in three groups: recurring flagship studies built on new primary data, research for firms deploying or supervising this technology, and shorter standing reference published between them. Of those, 9 are still to come — each one marked as such rather than listed as a title with nothing behind it.

Tier 1

Flagship studies · recurring primary data

Large-sample studies, repeated on a fixed schedule so the numbers can be compared year over year.

WP-01annual · benchmark studyComing soon

“Can AI Read Charts?” — The ORIN Benchmark

Question
No public head-to-head exists between vision language models and credentialed human chartists. The category’s defining question has not been measured.
Method
Frontier vision models against CMT charterholders, blind-scored on pattern identification, support and resistance placement, and trade-setup quality. The benchmark dataset and scoring rubric are open-sourced so the result can be reproduced and disputed.
Data
Generated for the study: a fixed chart sample, a published task taxonomy, and a panel of credentialed technicians scoring blind alongside the models.
WP-02annual · large-scale backtestComing soon

The Chart Pattern Efficacy Study

Question
The canonical academic work on chart-pattern efficacy dates to 2000. Since then the question has been answered mostly by assertion.
Method
Millions of detected pattern occurrences (head & shoulders, flags, double tops, etc.) across 10+ years — forward returns published by pattern, timeframe, and asset class.
Data
Public OHLCV history, with detection run by ORIN’s own engine — the same engine the product uses, so the detection criteria are the published ones.
WP-03annual · survey + telemetry reportComing soon

State of the Day Trader

Question
The profitability statistics in general circulation are between seven and twenty-three years old, and none of them describe today’s instruments or costs.
Method
Primary survey of 500–1,000 active traders (profitability, behavior, tooling, AI adoption), joined with anonymized platform telemetry as the user base grows.
Data
Survey panel year one; survey + telemetry thereafter.
Tier 2

Deployment & operating research

Research for firms deploying or supervising this technology — the questions that come up in diligence.

WP-04benchmark reportComing soon

The Retention Economics of Embedded Analysis

Question
The claim that embedded analysis increases trader lifetime is central to how this category is sold, and it has never been measured in public.
Method
Measure analytics engagement against trader lifetime, deposit frequency, and traded volume across participating brokers, with the method published before the data is collected.
Data
Anonymized, aggregated platform data contributed by participating brokerages under a published data-use agreement.
WP-05regulatory white paperPublished

AI Decision Support and the Advice Line

Question
No public analysis maps how the SEC, FINRA, the CFTC and ESMA treat AI-generated market analysis — or where decision support ends and regulated advice begins.
Method
Jurisdiction-by-jurisdiction mapping of the advice perimeter, with a compliance architecture for deploying AI analysis inside regulated platforms.
Data
Statutes, adopted rules, case law and official regulatory guidance, verified against each issuing body’s own publication.
Read the paper
WP-06prop-firm analysisComing soon

The Evaluation Science Report

Question
The available evidence suggests most evaluation failures are loss-limit breaches rather than missed profit targets — but every publisher of that evidence sells evaluations.
Method
Behavioural analysis of why traders fail evaluations: drawdown mechanics, position sizing, and decision quality under challenge rules, published by an author with no challenge fees to protect.
Data
Public firm disclosures and published evaluation datasets, extended with contributed firm data where a firm will agree to publication regardless of result.
Tier 3

Standing reference

Shorter studies and practitioner digests, published between the flagships.

WP-07large-scale backtestComing soon

The Indicator Graveyard

Method
Predictive value of RSI, MACD, Bollinger Bands and their common variants, tested at scale across assets and market regimes, reported per indicator rather than in aggregate.
WP-08behavioral studyComing soon

The Signal Fatigue Study

Method
Alerts received against alerts acted on, and how decision quality changes with alert volume — the first measurement of decision fatigue in active retail traders.
WP-09recurring digestsComing soon

The ArXiv Translation Series

Method
Plain-language digests of the academic literature — what the live LLM trading arenas actually found, what TradeTrap’s reliability results mean in practice — each linking the paper it summarises.
WP-10behavioral studyComing soon

Revenge Trading, Quantified

Method
Post-loss behaviour shifts measured in real trade data: position-size escalation, hold-time compression, and win-rate decay in the trades that follow a loss.

How we publish

Cadence
One flagship study per quarter

One study built on new primary data per quarter, with shorter reference pieces and literature digests published between them. Fewer, larger, and slower is the honest ceiling for a small research team.

Method first
The design is published before the data

Each study states its questions, its sample, and its scoring rules before the results are known. A study that picks its metrics after seeing the data can produce any conclusion you like, which is why the order matters more than the sample size.

Findings
One page per finding, dated

Every finding is stated once, in one sentence, with its basis and the date it was verified, on its own page. Quoting that sentence should carry enough with it that the claim stays true away from the paper — including when what it describes later changes.

Openness
Methodology and data in the open

Datasets, scoring rubrics and code are published wherever licensing allows, so a result can be reproduced or contradicted by someone who did not run it. Research nobody can check is an opinion with a chart on it.

Standards
A fixed methodology charter

Disclosed conflicts, no paid placement in any study, a standing corrections policy, and legal review on regulatory work. ORIN sells software into the market this lab studies; that conflict is stated on every paper rather than managed quietly.

Sources behind the gap analysis

Findings above were verified against current public output in the category, reviewed August 2026. Every figure on this page that ORIN did not measure itself traces to one of these.

R1

Chague, De-Losso & Giovannetti (2019), “Day Trading for a Living?” — Brazilian equity futures; 97% of individuals persisting 300+ days lost money. The most-cited statistic in the niche.

R2

Barber, Odean et al. — Taiwan Stock Exchange day-trading studies — Fewer than 1% of day traders predictably profitable after fees.

R3

Jordan & Diltz (2003), Financial Analysts Journal — Roughly twice as many U.S. day traders lose money as make money; ~20% more than marginally profitable.

R4

FPFX Technology dataset (300k+ evaluation accounts) — 5–10% pass prop-firm evaluations; ~7% of challenge buyers ever receive a payout; ~70% of failures from loss-limit breaches. Circulated via prop-firm and affiliate publishers.

R5

Devexperts trader survey — 38% of traders rank platform quality above fees/spreads; 47% consider their current platform outdated. The lone primary broker-side dataset in circulation.

R6

ArXiv 2024–2026 LLM-trading corpus — DeepFund and AI-Trader (live arenas), TradeTrap (agent reliability), AlphaForgeBench and Market-Bench (strategy design), FinTradeBench and related evaluation work. Dense, fast-moving, untranslated for practitioners; chart-reading evaluation largely absent.

R7

Trading Central & Autochartist public content libraries — Market commentary, signals, education, and product marketing; no original data studies or industry research identified.

ORIN is analysis software, not investment advice. Markets carry risk of loss. Read the risk disclosure.