The research question
A rule that wins one historical comparison may be responding to one convenient market interval. Edition 12 prevented later labels from leaking into a single comparison. Today we ask the promised next question: does that apparent advantage persist across several equal-duration chronological folds when the two rules stay frozen, costs stay fixed and unavailable outcomes remain unlabeled?
This is a consistency check, not a search for the most favorable window. We declare the baseline, challenger, selection count, fold boundaries, cost and reporting fields before evaluating any outcome. Each fold represents a new decision point followed by the same 24-hour evaluation period.
Prerequisites: Python 3.12 or later, a terminal and about 40 minutes. The lab uses only the standard library. Every name, feature, label and result is synthetic. It makes no network request, reads no credential, connects to no broker or wallet, and submits no order.
What changed in the inspected trading code
On October 8, 2026 UTC, PR #601 merged into main in Primus19/memecoine-mcp-server. We inspected PR head 10f011fb358afa97ccb12f944cd3737b93287950, merge commit 539803de57b3c066349688f9d71a38269bdd7a94, and the changed files services/solana-executor/index.mjs, services/solana-executor/runner-research.test.mjs, app/learning_loop.py and their tests.
The Solana executor change adds a prospective, shadow-only 90-minute continuation comparison for canonically reconciled 30-minute funded exits. A continuation seed is created only after the relevant close is reconciled and the balance is flat. Historical points are not filled retroactively. The comparison records the actual remaining quantity, prior recovered cash, a conservative quote floor, fee stress and incomplete evidence. Duplicate seeds are rejected. The inspected code explicitly marks the research as unable to change the live rule.
The same change improves post-exit quote observations. Research quote calls omit a transaction taker, protective quotes receive priority, missed horizons are recorded and quote work remains bounded. These are useful evidence controls. They are not a claim that holding longer is profitable.
The runner research tests cover start-date gating, receipt completeness, fixed quantity, duplicate prevention, stop and trailing-stop cases, and an incomplete result when the quote floor is absent. Those tests establish intended code behavior on fixtures. They do not establish a funded result.
What deployment evidence supports
Railway shows a successful deployment of merge commit 539803de57b3c066349688f9d71a38269bdd7a94 to the Solana executor. The deployment was created at 22:14:47 UTC on October 8 and reached success at approximately 22:15:34 UTC. Later repository commits were skipped for that service because their paths did not match its deployment filters, so this remains the current successful executor code identity we inspected.
The bounded read-only log check returned 51 recent lines with no error-like line. Targeted searches for supervision, exit observation and failed exit-quote events returned no retained matching lines. That absence cannot prove that a continuation observation ran successfully. Railway supports the deployment identity and a clean bounded tail, but not a claim about funded closes, continuation outcomes, profitability or model accuracy.
For that reason, the exercise below is an explicitly synthetic lab. It borrows the production discipline of prospective observation without publishing production account data or pretending that invented returns came from Railway.
1. Freeze the experiment before opening a fold
The baseline sorts candidates by one-hour change. The challenger uses the same inspected formula introduced in Edition 12:
2 × five-minute change + one-hour change + capped buy/sell ratio
Both rules select two candidates. The cap on the ratio is 20. Candidate name is the deterministic tie breaker. The synthetic cost is 0.05 units for each labeled selection.
Nothing is reweighted after a fold. If the challenger loses fold two, we do not change its coefficients before fold three. Changing a rule after inspecting the loss would turn the later fold into another development sample rather than an honest evaluation.
2. Give every fold the same amount of time
The three decision cutoffs are October 1, 3 and 5 at 12:00 UTC. Each evaluation cutoff is exactly 24 hours later. Equal test duration matters because a longer interval has more opportunity to collect outcomes than a shorter one.
This small script uses manually declared boundaries because the goal is to expose every clock. A larger project may use a cross-validation utility, but the underlying policy must remain visible: features come from no later than the decision cutoff, labels count only when available by the evaluation cutoff, and each fold covers the same duration.
Fold three contains two eventual outcomes that arrive after its evaluation cutoff. They remain unavailable even though their values are stored in the fixture. We report coverage as 1/2 for both rules instead of treating those rows as zero, losses or late wins.
3. Run the complete synthetic lab
Save this as lab_13.py, then run python -I lab_13.py from a new directory:
from dataclasses import dataclass
from datetime import datetime, timedelta
from decimal import Decimal
@dataclass(frozen=True)
class Candidate:
name: str
observed_at: str
m5: Decimal
h1: Decimal
buys: int
sells: int
gross: Decimal
label_available_at: str
@dataclass(frozen=True)
class Fold:
number: int
decision_cutoff: str
evaluation_cutoff: str
rows: tuple[Candidate, ...]
def row(name, observed, m5, h1, buys, sells, gross, available):
return Candidate(name, observed, Decimal(m5), Decimal(h1), buys, sells,
Decimal(gross), available)
folds = (
Fold(1, "2026-10-01T12:00:00+00:00", "2026-10-02T12:00:00+00:00", (
row("A", "2026-10-01T12:00:00+00:00", "1", "15", 10, 5, "0.15", "2026-10-01T18:00:00+00:00"),
row("B", "2026-10-01T12:00:00+00:00", "-1", "14", 8, 4, "-0.05", "2026-10-01T20:00:00+00:00"),
row("C", "2026-10-01T12:00:00+00:00", "5", "8", 20, 5, "0.15", "2026-10-02T02:00:00+00:00"),
row("D", "2026-10-01T12:00:00+00:00", "0", "4", 4, 4, "0.00", "2026-10-01T22:00:00+00:00"),
)),
Fold(2, "2026-10-03T12:00:00+00:00", "2026-10-04T12:00:00+00:00", (
row("E", "2026-10-03T12:00:00+00:00", "1", "15", 10, 5, "0.15", "2026-10-03T18:00:00+00:00"),
row("F", "2026-10-03T12:00:00+00:00", "-1", "14", 8, 4, "0.10", "2026-10-03T20:00:00+00:00"),
row("G", "2026-10-03T12:00:00+00:00", "5", "8", 20, 5, "-0.10", "2026-10-04T02:00:00+00:00"),
row("K", "2026-10-03T12:00:00+00:00", "0", "4", 4, 4, "0.00", "2026-10-03T22:00:00+00:00"),
)),
Fold(3, "2026-10-05T12:00:00+00:00", "2026-10-06T12:00:00+00:00", (
row("H", "2026-10-05T12:00:00+00:00", "1", "15", 10, 5, "0.10", "2026-10-05T18:00:00+00:00"),
row("I", "2026-10-05T12:00:00+00:00", "-1", "14", 8, 4, "0.50", "2026-10-07T12:00:00+00:00"),
row("J", "2026-10-05T12:00:00+00:00", "5", "8", 20, 5, "0.50", "2026-10-07T12:00:00+00:00"),
row("L", "2026-10-05T12:00:00+00:00", "0", "4", 4, 4, "0.00", "2026-10-05T22:00:00+00:00"),
)),
)
cost = Decimal("0.05")
def baseline(candidate):
return candidate.h1
def challenger(candidate):
pressure = min(Decimal("20"),
Decimal(candidate.buys) / Decimal(max(candidate.sells, 1)))
return Decimal("2") * candidate.m5 + candidate.h1 + pressure
def rank(fold, score):
cutoff = datetime.fromisoformat(fold.decision_cutoff)
eligible = tuple(candidate for candidate in fold.rows
if datetime.fromisoformat(candidate.observed_at) <= cutoff)
return tuple(sorted(eligible, key=lambda item: (-score(item), item.name))[:2])
def evaluate(fold, selected):
cutoff = datetime.fromisoformat(fold.evaluation_cutoff)
labeled = tuple(candidate for candidate in selected
if datetime.fromisoformat(candidate.label_available_at) <= cutoff)
unavailable = tuple(candidate for candidate in selected if candidate not in labeled)
net = sum((candidate.gross - cost for candidate in labeled), Decimal("0"))
return labeled, unavailable, net
baseline_total = Decimal("0")
challenger_total = Decimal("0")
better = worse = ties = 0
unavailable_lines = []
for fold in folds:
baseline_rows = rank(fold, baseline)
challenger_rows = rank(fold, challenger)
baseline_result = evaluate(fold, baseline_rows)
challenger_result = evaluate(fold, challenger_rows)
baseline_net = baseline_result[2]
challenger_net = challenger_result[2]
delta = challenger_net - baseline_net
assert [candidate.name for candidate in baseline_rows] == {
1: ["A", "B"], 2: ["E", "F"], 3: ["H", "I"]}[fold.number]
assert [candidate.name for candidate in challenger_rows] == {
1: ["C", "A"], 2: ["G", "E"], 3: ["J", "H"]}[fold.number]
assert datetime.fromisoformat(fold.evaluation_cutoff) - datetime.fromisoformat(
fold.decision_cutoff) == timedelta(hours=24)
assert all(datetime.fromisoformat(candidate.observed_at) <=
datetime.fromisoformat(fold.decision_cutoff) for candidate in fold.rows)
baseline_total += baseline_net
challenger_total += challenger_net
better += delta > 0
worse += delta < 0
ties += delta == 0
baseline_ids = ",".join(candidate.name for candidate in baseline_rows)
challenger_ids = ",".join(candidate.name for candidate in challenger_rows)
print(
f"fold={fold.number} baseline_rank={baseline_ids} challenger_rank={challenger_ids} "
f"baseline_coverage={len(baseline_result[0])}/2 "
f"challenger_coverage={len(challenger_result[0])}/2 "
f"baseline_net={baseline_net:.2f} challenger_net={challenger_net:.2f} "
f"delta={delta:.2f}"
)
if baseline_result[1] or challenger_result[1]:
baseline_missing = ",".join(candidate.name for candidate in baseline_result[1]) or "none"
challenger_missing = ",".join(candidate.name for candidate in challenger_result[1]) or "none"
unavailable_lines.append(
f"unavailable fold={fold.number} baseline={baseline_missing} "
f"challenger={challenger_missing}"
)
assert baseline_total == Decimal("0.20")
assert challenger_total == Decimal("0.20")
assert (better, worse, ties) == (1, 1, 1)
assert unavailable_lines == ["unavailable fold=3 baseline=I challenger=J"]
print(f"aggregate baseline_net={baseline_total:.2f} "
f"challenger_net={challenger_total:.2f} "
f"delta={challenger_total - baseline_total:.2f}")
print(f"consistency challenger_better={better} challenger_worse={worse} ties={ties}")
for line in unavailable_lines:
print(line)
print("checks=18_passed data=synthetic no_orders_submitted")
Expected complete standard output:
fold=1 baseline_rank=A,B challenger_rank=C,A baseline_coverage=2/2 challenger_coverage=2/2 baseline_net=0.00 challenger_net=0.20 delta=0.20
fold=2 baseline_rank=E,F challenger_rank=G,E baseline_coverage=2/2 challenger_coverage=2/2 baseline_net=0.15 challenger_net=-0.05 delta=-0.20
fold=3 baseline_rank=H,I challenger_rank=J,H baseline_coverage=1/2 challenger_coverage=1/2 baseline_net=0.05 challenger_net=0.05 delta=0.00
aggregate baseline_net=0.20 challenger_net=0.20 delta=0.00
consistency challenger_better=1 challenger_worse=1 ties=1
unavailable fold=3 baseline=I challenger=J
checks=18_passed data=synthetic no_orders_submitted
We executed the standalone file in a new temporary directory with an empty environment, isolated Python mode and no live-service credentials. Its complete standard output matched the block above.
4. Read the fold results before the total
Fold one favors the challenger by 0.20 units after synthetic costs. Fold two reverses that result and favors the baseline by 0.20. Fold three is a tie with equal 1/2 coverage. The aggregate is 0.20 for each rule, so the challenger’s apparent single-window advantage does not persist across this fixture.
The per-fold result is more informative than the aggregate alone. A total delta of zero could hide three ties, or it could hide the exact reversal seen here. Reporting better, worse and tied folds makes instability visible.
This result is deliberately inconvenient. A lab that always confirms the challenger would teach selection, not evaluation. The correct conclusion is not that the baseline is superior. It is that three tiny synthetic folds provide no basis for promoting either rule.
5. Connect the method to established guidance
Scikit-learn’s official TimeSeriesSplit documentation explains that ordinary cross-validation can train on future data and evaluate on the past. It also states that equally spaced samples are required for comparable metrics because each test set should cover the same time duration. Our lab does not import scikit-learn, but it applies those two principles directly.
Bailey, Borwein, López de Prado and Zhu study the probability of backtest overfitting when researchers select investment strategies from repeated simulations. Their paper proposes combinatorially symmetric cross-validation for a larger formal analysis. We do not calculate that statistic here. Twelve invented candidates and two fixed rules cannot support it. The relevant lesson is narrower: repeated strategy selection can make a historical winner look stronger than it is, so rule choice and evaluation evidence must be separated.
Troubleshooting
If a rank differs, retain Decimal arithmetic, the ratio cap and candidate-name tie breaker. If fold duration fails, keep timezone-aware ISO timestamps and verify that every evaluation cutoff is 24 hours after its decision cutoff. If I or J appears as labeled, compare label_available_at with fold three’s evaluation cutoff before reading gross.
Run without -O because optimized mode removes assertions. Charge costs only to labeled selections in this lab, then report coverage beside net result. Do not replace an unavailable label with zero. Do not extend one evaluation cutoff after seeing a late favorable value. If you change a coefficient, restart with untouched future folds and record the new rule as a different experiment.
Evidence limits and completion check
The lab verifies a small evaluation procedure, not a trading strategy. It does not test provider completeness, fill quality, token safety, venue costs, the deployed continuation policy, live P&L or customer outcomes. Repository inspection supports the code description. Railway supports the deployed commit identity and bounded log statement. The numbers in the lab are original synthetic fixtures and are not funded observations.
You are finished when the complete output matches, you can explain why fold three has 1/2 coverage for both rules, and you can reproduce each net result by subtracting 0.05 from every labeled gross outcome. Preserve the frozen rules, fold boundaries, feature rows, label-availability times, costs and output together.
The next experiment will replace the hand-built folds with a prospective evaluation ledger. It will define a minimum evidence count, record fold-level dispersion and use a predeclared promotion gate that can return “keep researching” instead of forcing a winner.
Primary technical references
Scikit-learn’s official TimeSeriesSplit documentation defines chronological splits, equal-duration test sets and the optional gap between training and test data. Bailey, Borwein, López de Prado and Zhu’s paper, “The Probability of Backtest Overfitting”, develops a formal framework for strategy-selection overfitting. Neither source validates this synthetic fixture or the deployed trading system.