How Comparand works and where it fails
How the answers are built, what the analog search can and cannot show, and which tests have not passed.
What the data is
Comparand covers Hyperliquid only. It has one-minute bars for 212 coins. The archive runs from 2025-04-01 to 2026-09-30. That is 548 days, and 522 of them are complete.
The other 26 days are known gaps. The processing run refused each of them, and nothing is served for them. Most of the gaps fall in August 2025. 22 of the 26 gap days are in that month.
Each complete day records the source files and the processing run that produced it. The coverage page lists every gap day and every count.
Why every answer carries its sources
Every answer has four blocks: data, prov, meter, and disclosure. The prov block lists the archive days, a digest for each day, the definitions hash, the code revisions, and the source files.
The rule is that every served number carries its archive days, its file sha, and its reduce sha.
Stale data is refused, not served. Live labels older than 120 seconds return a live_stale error. A refused call is not billed.
Where it fails: a provenance block shows what went into an answer. It does not show that the answer is right.
What the analog search is and is not
The analog search finds past moments that look most like a coin's state at one time. For each match it shows the outcome over the next 1 or 4 hours, the time that outcome became known, and the return in basis points. One basis point, or bp, is one hundredth of one percent.
Each match is one past case, with its own date and its own outcome. The matches are listed nearest first.
It is not an average. The packet format has no field for a mean, a median, a weighted value, a hit rate, or a consensus. A strict schema check rejects those fields.
It is not a forecast. A past case says what followed then. It does not say what will follow now.
Where it fails: an internal audit on 2026-10-03 found that several matches can come from the same moment. A query for 20 matches can return 20 coins from one minute.
The result that stopped us making forecasts
We tested a forecast built on the analog search. It weighted each past case by how near it was, then combined the outcomes into one expected return. The test is called analog1.
The fit window ran from 2026-06-10 to 2026-07-24. The holdout window ran from 2026-07-25 to 2026-09-24. The test covered 146 coins. The table shows the holdout results, with one column for each horizon.
| Holdout, 2026-07-25 to 2026-09-24 | 1 hour | 4 hours |
|---|---|---|
| Trades | 190,825 | 52,746 |
| Average result before costs | +0.55 bp | -1.65 bp |
| Average cost per trade | 12.43 bp | 12.65 bp |
| Average result after costs | -11.88 bp | -14.30 bp |
| Hit rate | 41.6% | 45.3% |
The forecast lost money after costs at both horizons. At 1 hour the result before costs was slightly positive, at +0.55 bp, and costs were 12.43 bp. At 4 hours the result was negative even before costs.
The fit window was negative too. At 1 hour the forecast lost 12.87 bp per trade over 140,855 trades.
Costs were venue fees and a spread estimate, taken from the config files of the run.
This is one holdout window, 146 coins, and one set of costs. It does not prove that no analog signal could ever make money. It does show that this forecast did not.
Every analog answer carries this result in its disclosure block. This result is why the product makes no forecasts. It keeps each past case and its outcome, and it adds no average.
The live cascade screen and how often it is wrong
The live screen flags liquidation cascades. A cascade here is the labels cascade flag, counted per coin in each 5-minute bar. A flag is a hit only when it lands on the same bar and in the same direction. A flag one bar early or late counts as a miss and a false flag.
We scored the picked rule on 38 archive days, from 2026-08-18 to 2026-09-24. The rule was picked on the training window, from 2026-06-03 to 2026-08-17. The other candidates were reported, but they were not used to pick.
| Test window, 38 days | Count or rate |
|---|---|
| Cascade events | 4,426 |
| Flags | 24,348 |
| Flags that were hits | 3,126 |
| Precision | 12.8% |
| Recall | 70.6% |
Precision is 12.8%. So 87 of every 100 flags are not cascades. That is 21,222 of the 24,348 flags.
Recall is 70.6%. So about 29 of every 100 cascades are missed. That is 1,300 of the 4,426 events.
The screen passed its pre-registered bar. The bar asked for precision of at least 10% and recall of at least 50%. That bar is low, and a screen can pass it while most of its flags are wrong.
Where it fails: most flags are not cascades. Check a flag against the data before you rely on it.
Backtests and why they are off
A backtest runs a market-making rule against a replay of the order book. A fill model decides when a simulated resting order fills. Until the fill model passes its test, reports say replay fidelity is not yet established.
The fill model has to pass an identity test called K1. Replay fills have to match real fills with recall of at least 0.95, precision of at least 0.95, and a bracket of at least 0.90.
K1 has failed twice. The first attempt used input that was later found to be defective. The second attempt used clean input. It reached precision of 0.968 and a bracket of 0.998, but recall was 0.801. The bar is 0.95.
Backtests stay off until the fill model passes K1.
Every backtest is set up the same way before it runs. A pre-registration stub is written from a template with 17 fixed headings. It records the pass rule and the cost stack: 1.5 bp per side for a maker fill, 4.5 bp per side for a taker fill, and 200 milliseconds of latency.
Every backtest is also counted in a per-key trial ledger before it runs. The ledger counts every trial. It also counts each cluster of similar return series once.
Where it fails: pre-registration fixes the rules before a run. It does not check the fill model. A backtest on a fill model that fails K1 cannot be trusted.
What we have not established
- Replay fidelity. It is not established on any of the 548 days.
- A fill model that passes K1. It has failed twice.
- A forecast that makes money after costs. The analog1 forecast did not, on one holdout window across 146 coins.
- A cascade screen with high precision. Precision was 12.8% on 38 test days.
- Full coverage. 26 days are gaps, and 97 days have no cohort data.
The coverage page lists the gap days and the counts behind these numbers.