How Comparand works and where it fails

How the answers are built, what the analog search can and cannot show, and which tests have not passed.

What the data is

Comparand covers Hyperliquid only. It has one-minute bars for 212 coins. The archive runs from 2025-04-01 to 2026-09-30. That is 548 days, and 522 of them are complete.

Source: derived/coverage/capability.json (span, counts); docs/COVERAGE.md; src/analogmcp/mcp_server.py (1-minute bars); AGENTS.md (Hyperliquid only)

The other 26 days are known gaps. The processing run refused each of them, and nothing is served for them. Most of the gaps fall in August 2025. 22 of the 26 gap days are in that month.

Source: derived/coverage/capability.json (known_gaps); docs/COVERAGE.md; src/analogmcp/provenance.py (known-gap days are never served)

Each complete day records the source files and the processing run that produced it. The coverage page lists every gap day and every count.

Why every answer carries its sources

Every answer has four blocks: data, prov, meter, and disclosure. The prov block lists the archive days, a digest for each day, the definitions hash, the code revisions, and the source files.

Source: src/analogmcp/envelope.py (build); src/analogmcp/service.py (prov fields)

The rule is that every served number carries its archive days, its file sha, and its reduce sha.

Source: AGENTS.md

Stale data is refused, not served. Live labels older than 120 seconds return a live_stale error. A refused call is not billed.

Source: src/analogmcp/live/serve.py (max age 120 seconds); src/analogmcp/mcp_server.py (refused, not billed)

Where it fails: a provenance block shows what went into an answer. It does not show that the answer is right.

What the analog search is and is not

The analog search finds past moments that look most like a coin's state at one time. For each match it shows the outcome over the next 1 or 4 hours, the time that outcome became known, and the return in basis points. One basis point, or bp, is one hundredth of one percent.

Source: config/disclosures/analog1.json (1h and 4h horizons); src/analogmcp/analogs/packets.py (outcome fields)

Each match is one past case, with its own date and its own outcome. The matches are listed nearest first.

It is not an average. The packet format has no field for a mean, a median, a weighted value, a hit rate, or a consensus. A strict schema check rejects those fields.

Source: src/analogmcp/analogs/packets.py (strict schema, no mean, median, weighted value, hit rate, or consensus)

It is not a forecast. A past case says what followed then. It does not say what will follow now.

Source: .planning/PROJECT.md (Why no forecasts)

Where it fails: an internal audit on 2026-10-03 found that several matches can come from the same moment. A query for 20 matches can return 20 coins from one minute.

Source: .planning/STATE.md (2026-10-03 naivety audit entry)

The result that stopped us making forecasts

We tested a forecast built on the analog search. It weighted each past case by how near it was, then combined the outcomes into one expected return. The test is called analog1.

Source: config/disclosures/analog1.json (the A1 forecast text)

The fit window ran from 2026-06-10 to 2026-07-24. The holdout window ran from 2026-07-25 to 2026-09-24. The test covered 146 coins. The table shows the holdout results, with one column for each horizon.

Source: reports/analog1/a1_1h_result.json and reports/analog1/a1_4h_result.json (hl-mm repo; windows and gross results); config/disclosures/analog1.json (trades, cost, after-cost result, hit rate)
Holdout, 2026-07-25 to 2026-09-241 hour4 hours
Trades190,82552,746
Average result before costs+0.55 bp-1.65 bp
Average cost per trade12.43 bp12.65 bp
Average result after costs-11.88 bp-14.30 bp
Hit rate41.6%45.3%
Source: config/disclosures/analog1.json (holdout trades, cost, after-cost result, hit rate); reports/analog1/a1_1h_result.json and reports/analog1/a1_4h_result.json (before-cost result)

The forecast lost money after costs at both horizons. At 1 hour the result before costs was slightly positive, at +0.55 bp, and costs were 12.43 bp. At 4 hours the result was negative even before costs.

The fit window was negative too. At 1 hour the forecast lost 12.87 bp per trade over 140,855 trades.

Source: config/disclosures/analog1.json (fit window)

Costs were venue fees and a spread estimate, taken from the config files of the run.

Source: reports/analog1/a1_1h_result.json (run arguments: fees, spread proxy) in the hl-mm repo

This is one holdout window, 146 coins, and one set of costs. It does not prove that no analog signal could ever make money. It does show that this forecast did not.

Every analog answer carries this result in its disclosure block. This result is why the product makes no forecasts. It keeps each past case and its outcome, and it adds no average.

Source: src/analogmcp/analogs/disclosure.py (disclosure text); .planning/PROJECT.md (Why no forecasts); .planning/STATE.md (owner ruling: keep outcomes, no average)

The live cascade screen and how often it is wrong

The live screen flags liquidation cascades. A cascade here is the labels cascade flag, counted per coin in each 5-minute bar. A flag is a hit only when it lands on the same bar and in the same direction. A flag one bar early or late counts as a miss and a false flag.

Source: .planning/phases/04-live-labels/04-CASC.json (spec: truth, hit)

We scored the picked rule on 38 archive days, from 2026-08-18 to 2026-09-24. The rule was picked on the training window, from 2026-06-03 to 2026-08-17. The other candidates were reported, but they were not used to pick.

Source: .planning/phases/04-live-labels/04-CASC.json (split, selection, verdict_on)
Test window, 38 daysCount or rate
Cascade events4,426
Flags24,348
Flags that were hits3,126
Precision12.8%
Recall70.6%
Source: .planning/phases/04-live-labels/04-CASC.json (test, picked rule C2_spike_ext, parameters F 50000 and p 0.8)

Precision is 12.8%. So 87 of every 100 flags are not cascades. That is 21,222 of the 24,348 flags.

Recall is 70.6%. So about 29 of every 100 cascades are missed. That is 1,300 of the 4,426 events.

The screen passed its pre-registered bar. The bar asked for precision of at least 10% and recall of at least 50%. That bar is low, and a screen can pass it while most of its flags are wrong.

Source: .planning/phases/04-live-labels/04-CASC.json (spec: pass; verdict; counts from test); the 87 and 29 figures are computed from the same counts

Where it fails: most flags are not cascades. Check a flag against the data before you rely on it.

Backtests and why they are off

A backtest runs a market-making rule against a replay of the order book. A fill model decides when a simulated resting order fills. Until the fill model passes its test, reports say replay fidelity is not yet established.

Source: .planning/ROADMAP.md (fill simulator line; replay fidelity wording)

The fill model has to pass an identity test called K1. Replay fills have to match real fills with recall of at least 0.95, precision of at least 0.95, and a bracket of at least 0.90.

Source: .planning/ROADMAP.md (K1 pass line)

K1 has failed twice. The first attempt used input that was later found to be defective. The second attempt used clean input. It reached precision of 0.968 and a bracket of 0.998, but recall was 0.801. The bar is 0.95.

Source: .planning/MORNING-2026-10-04.md (K1 attempt 2)

Backtests stay off until the fill model passes K1.

Source: .planning/ROADMAP.md (results withdrawn until a post-B1 extract passes K1)

Every backtest is set up the same way before it runs. A pre-registration stub is written from a template with 17 fixed headings. It records the pass rule and the cost stack: 1.5 bp per side for a maker fill, 4.5 bp per side for a taker fill, and 200 milliseconds of latency.

Every backtest is also counted in a per-key trial ledger before it runs. The ledger counts every trial. It also counts each cluster of similar return series once.

Where it fails: pre-registration fixes the rules before a run. It does not check the fill model. A backtest on a fill model that fails K1 cannot be trusted.

Source: src/analogmcp/bt/prereg.py (headings, cost stack, latency); src/analogmcp/bt/ledger.py (trial counts and clusters); AGENTS.md (every backtest is counted before it runs)

What we have not established

Source: derived/coverage/capability.json (replay fidelity, gaps, cohort); .planning/MORNING-2026-10-04.md (K1); config/disclosures/analog1.json (analog1); .planning/phases/04-live-labels/04-CASC.json (precision)

The coverage page lists the gap days and the counts behind these numbers.