The research question

A ranking rule can look better simply because we let later outcomes influence an earlier decision. That error is subtle. The code may never place a future price inside a feature column, yet the evaluation can still leak if we choose a rule after seeing which candidates won, label an outcome before its evidence was available, or compare rules with different amounts of follow-up.

Edition 11 established that only reconciled, verified closes belong in a win or loss denominator. Today we take the next step: when two frozen ranking rules select different candidates, does the apparent improvement survive a decision-time cutoff, a label-availability cutoff, equal costs and explicit coverage?

Prerequisites: Python 3.12 or later, a terminal, and about 35 minutes. The lab uses only the standard library. Every candidate, feature, outcome and cost is synthetic. It makes no network call, reads no credential, connects to no broker or wallet, and submits no order.

What changed in the inspected trading code

On October 8, 2026 UTC, PR #593 merged into main in Primus19/memecoine-mcp-server. We inspected head commit cbd61a70b5a521232db02c60fdf47570630746a1 and merge commit f2a62be0ec69b1ecab662922a7e818fc524db2b9. Relevant files include app/robinhood_market_quote.py, app/robinhood_momentum.py, app/asset_worker.py and tests/test_robinhood_market_quote.py.

The change adds an opt-in exact-pool fallback for a sparse primary Robinhood quote. It verifies the network, pool address and base-token identity. It requires finite price, liquidity, change, volume and transaction fields from one current response. A missing horizon is rejected rather than filled with an older value. The momentum evaluator accepts the new pair-native provenance while retaining its other requirements. Its tests also verify that adverse safety evidence or negative short-horizon momentum remains ineligible.

One inspected score combines short and longer momentum with buy-to-sell pressure:

2 × five-minute change + one-hour change + capped buy/sell ratio

That formula gives us a concrete challenger for this lab. It does not give us a production result. A reliable quote can make a candidate evaluable, but it cannot prove that the ranker predicts a profitable outcome.

What deployment evidence supports

Railway records show that the multi-asset paper worker successfully deployed commit e0e9830ff02aaca437867ad5b6a8d220288e88f7 at 02:49 UTC on October 8. That later commit contains the PR #593 change and subsequent reporting repairs. The direct PR #593 deployment was replaced by those later releases.

The read-only Railway log query returned no retained runtime lines for the successful deployment. Therefore the evidence supports deployed code identity, not a claim that a fallback quote was used, a candidate was admitted, a position was opened, or a funded outcome was realized. The repository PR describes additional observations, but this article does not promote those descriptions into independently verified trading results.

1. Freeze both clocks

The decision cutoff is when the rankers are allowed to see features. The evaluation cutoff is when outcomes are counted. These clocks answer different questions.

Every feature row in the lab exists at 2026-10-01T12:00:00+00:00. Every outcome becomes available later. Candidate B has an outcome in the fixture, but its label does not become available until October 3. Our evaluation stops on October 2, so B remains unlabeled for both rules. Storing its eventual result in the same object is safe only because the evaluator checks the availability time before reading it.

The rules are declared before evaluation. The baseline sorts on one-hour change. The challenger uses the inspected momentum formula. Each selects three candidates, with the candidate name as a deterministic tie breaker.

2. Charge equal explicit costs

For every labeled selection, the lab subtracts a synthetic cost of 0.05 units. This single figure is educational, not an estimate of any venue's fee, spread, slippage, gas or price impact. A real experiment would record those components separately and reconcile them to execution receipts.

Both rules achieve coverage of two labeled selections out of three. Equal coverage makes the small comparison easier to read, although it does not make the sample statistically useful. If one rule had lower coverage, its apparent net result would need to be interpreted alongside that missingness.

3. Run the complete synthetic lab

Save this as lab_12.py, then run python -I lab_12.py from a new directory:

from dataclasses import dataclass
from datetime import datetime
from decimal import Decimal


@dataclass(frozen=True)
class Candidate:
    name: str
    observed_at: str
    m5: Decimal
    h1: Decimal
    buys: int
    sells: int
    gross: Decimal
    label_available_at: str


rows = (
    Candidate("A", "2026-10-01T12:00:00+00:00", Decimal("1"), Decimal("15"), 10, 5,
              Decimal("0.20"), "2026-10-01T14:00:00+00:00"),
    Candidate("B", "2026-10-01T12:00:00+00:00", Decimal("0"), Decimal("14"), 5, 5,
              Decimal("0.40"), "2026-10-03T12:00:00+00:00"),
    Candidate("C", "2026-10-01T12:00:00+00:00", Decimal("-1"), Decimal("13"), 8, 4,
              Decimal("-0.15"), "2026-10-01T20:00:00+00:00"),
    Candidate("D", "2026-10-01T12:00:00+00:00", Decimal("5"), Decimal("8"), 20, 5,
              Decimal("0.12"), "2026-10-01T18:00:00+00:00"),
    Candidate("E", "2026-10-01T12:00:00+00:00", Decimal("2"), Decimal("5"), 6, 6,
              Decimal("0.04"), "2026-10-01T16:00:00+00:00"),
    Candidate("F", "2026-10-01T12:00:00+00:00", Decimal("-2"), Decimal("10"), 3, 6,
              Decimal("-0.10"), "2026-10-01T17:00:00+00:00"),
)

decision_cutoff = datetime.fromisoformat("2026-10-01T12:00:00+00:00")
evaluation_cutoff = datetime.fromisoformat("2026-10-02T12:00:00+00:00")
cost = Decimal("0.05")


def baseline(row):
    return row.h1


def challenger(row):
    pressure = min(Decimal("20"), Decimal(row.buys) / Decimal(max(row.sells, 1)))
    return Decimal("2") * row.m5 + row.h1 + pressure


def rank(score):
    eligible = tuple(row for row in rows
                     if datetime.fromisoformat(row.observed_at) <= decision_cutoff)
    return tuple(sorted(eligible, key=lambda row: (-score(row), row.name))[:3])


def evaluate(selected):
    labeled = tuple(row for row in selected
                    if datetime.fromisoformat(row.label_available_at) <= evaluation_cutoff)
    unlabeled = tuple(row for row in selected if row not in labeled)
    gross = sum((row.gross for row in labeled), Decimal("0"))
    costs = cost * len(labeled)
    net = gross - costs
    wins = sum(row.gross - cost > 0 for row in labeled)
    return labeled, unlabeled, gross, costs, net, wins


baseline_rows = rank(baseline)
challenger_rows = rank(challenger)
baseline_result = evaluate(baseline_rows)
challenger_result = evaluate(challenger_rows)

assert [row.name for row in baseline_rows] == ["A", "B", "C"]
assert [baseline(row) for row in baseline_rows] == [
    Decimal("15"), Decimal("14"), Decimal("13")]
assert [row.name for row in challenger_rows] == ["D", "A", "B"]
assert [challenger(row) for row in challenger_rows] == [
    Decimal("22"), Decimal("19"), Decimal("15")]
assert [row.name for row in baseline_result[0]] == ["A", "C"]
assert [row.name for row in challenger_result[0]] == ["D", "A"]
assert [row.name for row in baseline_result[1]] == ["B"]
assert [row.name for row in challenger_result[1]] == ["B"]
assert baseline_result[2:5] == (
    Decimal("0.05"), Decimal("0.10"), Decimal("-0.05"))
assert challenger_result[2:5] == (
    Decimal("0.32"), Decimal("0.10"), Decimal("0.22"))
assert baseline_result[5] == 1
assert challenger_result[5] == 2
assert challenger_result[4] - baseline_result[4] == Decimal("0.27")
assert all(datetime.fromisoformat(row.observed_at) <= decision_cutoff for row in rows)
assert all(datetime.fromisoformat(row.label_available_at) > decision_cutoff for row in rows)
assert rows == tuple(rows)


def ids(items):
    return ",".join(row.name for row in items)


print("decision_cutoff=2026-10-01T12:00:00+00:00 evaluation_cutoff=2026-10-02T12:00:00+00:00")
print("baseline_rank=A,B,C scores=A:15.00,B:14.00,C:13.00")
print("challenger_rank=D,A,B scores=D:22.00,A:19.00,B:15.00")
print(f"baseline_labeled={ids(baseline_result[0])} coverage=2/3 wins=1 losses=1 gross=0.05 costs=0.10 net=-0.05")
print(f"challenger_labeled={ids(challenger_result[0])} coverage=2/3 wins=2 losses=0 gross=0.32 costs=0.10 net=0.22")
print("baseline_unlabeled=B:LABEL_AFTER_CUTOFF challenger_unlabeled=B:LABEL_AFTER_CUTOFF")
print("comparison=synthetic challenger_minus_baseline_net=0.27 checks=16_passed no_orders_submitted")

Expected complete standard output:

decision_cutoff=2026-10-01T12:00:00+00:00 evaluation_cutoff=2026-10-02T12:00:00+00:00
baseline_rank=A,B,C scores=A:15.00,B:14.00,C:13.00
challenger_rank=D,A,B scores=D:22.00,A:19.00,B:15.00
baseline_labeled=A,C coverage=2/3 wins=1 losses=1 gross=0.05 costs=0.10 net=-0.05
challenger_labeled=D,A coverage=2/3 wins=2 losses=0 gross=0.32 costs=0.10 net=0.22
baseline_unlabeled=B:LABEL_AFTER_CUTOFF challenger_unlabeled=B:LABEL_AFTER_CUTOFF
comparison=synthetic challenger_minus_baseline_net=0.27 checks=16_passed no_orders_submitted

We executed this file in an isolated directory with a cleared environment. Its complete standard output matched the block above.

4. Interpret the result narrowly

The baseline selects A, B and C. B is unavailable at the evaluation cutoff, so A and C contribute one synthetic win, one loss and negative 0.05 after costs. The challenger selects D, A and B. With B excluded for the same reason, D and A contribute two wins and positive 0.22 after costs.

The difference is 0.27 units in favor of the challenger on this invented fixture. It is not evidence of production profitability, model accuracy or durable improvement. The comparison contains only two labeled selections per rule, and A appears in both. We designed the numbers to exercise the method, not to simulate a market distribution.

The essential result is procedural: B's eventual positive outcome never enters either score or metric before the evaluation cutoff. Both rules report the same 2/3 coverage. Costs are subtracted consistently. The selected identities and rule formulas are printed so another person can reproduce the comparison.

5. Why chronological evaluation matters

Scikit-learn's official TimeSeriesSplit documentation explains that ordinary cross-validation can train on future data and evaluate on the past. Its expanding-window approach keeps later observations in later test sets. Our small lab does not use scikit-learn, but it follows the same time-order principle with explicit cutoffs.

Bailey, Borwein, López de Prado and Zhu's paper on backtest overfitting studies a broader problem: selecting a strategy from many tried configurations can make an in-sample winner look stronger than it is. We do not calculate their probability-of-overfitting statistic here. With six invented rows, doing so would create false precision. The applicable lesson is to declare the two rules before examining the evaluation outcomes and to preserve a later, untouched test.

Troubleshooting

If B appears in a labeled set, compare its label_available_at with the evaluation cutoff and keep the comparison timezone-aware. If the challenger order differs, retain Decimal arithmetic and the deterministic name tie breaker. If a cost total is zero, confirm that the cost is multiplied by labeled selections, not all selected candidates.

Run without -O because optimized mode removes assertions. If you change the top-k value, update both the denominator and expected identities. Do not fill an unavailable label with zero, and do not move the evaluation cutoff after seeing a favorable late outcome.

Evidence limits and completion check

The lab checks a time-aware evaluation policy against a tiny synthetic fixture. It does not validate provider completeness, Robinhood execution, fill quality, token safety, the deployed ranker, venue costs or live returns. The inspected code establishes current identity and completeness checks. Railway establishes that a later containing commit deployed successfully. Neither source supplies the synthetic outcomes.

You are finished when the complete output matches, you can explain why B is excluded despite having an eventual outcome, and you can recompute both net totals by hand. Save the rules, feature snapshot, decision cutoff, label-availability times, evaluation cutoff and output together.

The next experiment will move from one cutoff to several chronological folds. We will keep each rule frozen, give every fold the same test duration, preserve unresolved outcomes, and check whether the apparent advantage persists rather than depending on one convenient window.

Primary technical references

Scikit-learn's official TimeSeriesSplit documentation defines chronological train and test splits and explains why ordinary cross-validation is inappropriate for time-ordered data. Bailey, Borwein, López de Prado and Zhu's paper, “The Probability of Backtest Overfitting”, presents a framework for measuring the risk created by selecting investment strategies from repeated backtests. The paper motivates a future larger-sample test; it does not validate this lab or the deployed trading system.