Research question

Can a prospective evaluation ledger stop a promising aggregate result from being promoted when the evidence is too small, incomplete, or concentrated in one fold?

Edition 13 compared one ranking rule across several chronological windows. That reduced dependence on one lucky split, but it left an important decision unspecified: how much evidence is enough? If we choose that threshold after seeing the results, the threshold becomes another tuned parameter. This edition writes the gate first, keeps unresolved observations unresolved, and allows the decision to be KEEP_RESEARCHING.

This is an explicitly synthetic lab. It does not reproduce a funded strategy, estimate future profit, or submit an order. The units below represent an invented difference between a challenger and a baseline after assumed costs. A positive number favors the challenger; a negative number favors the baseline.

Time: about 40 minutes. Prerequisites: Python 3.12 or later, a text editor, and the walk-forward ideas from Edition 13. The script uses only the standard library and makes no network, broker, exchange, or wallet calls.

What was inspected

The publication source was checked before assigning this number. The course fixture still contains Editions 1 through 6. The research fixture contains Editions 7 through 13 once each, with Edition 13 released on October 9, 2026. No October 10 entry or competing Edition 14 draft existed when this article was prepared.

The trading repository was inspected at merged commits 7439728, a15c278 and 9fc4167. PR 612 introduced a bounded, funded nomination path that records market confirmations without constructing paper positions. Its code labels that evidence MARKET_CONFIRMATION_NOT_A_FILL, limits retained confirmation state, deduplicates nomination identities, and leaves existing route, risk, receipt and profitability checks in place. The same review records explicit minimum-evidence language for later promotion. PRs 613 and 614 then repaired and bounded the Solana service health response without changing trading rules.

Read-only deployment evidence distinguishes code from operation. The funded worker and market-feed deployments for 7439728 succeeded. The first two Solana executor deployments failed readiness, while deployment 86a38118 for 9fc4167 succeeded and its build log states that the /health check passed. That is deployment evidence, not proof that a strategy earns money. No funded trade outcome is used in this lab.

Step 1: Declare the gate before the data

Write the policy down before reading the fold results. Our teaching gate requires all of the following:

  1. At least six resolved folds.
  2. Coverage of at least 80 percent in every resolved fold.
  3. A positive mean difference.
  4. A positive median difference.
  5. Positive differences in at least 67 percent of resolved folds.
  6. No single positive fold contributing more than 50 percent of all positive improvement.

The first rule prevents five observations from being described as six. The coverage rule prevents a partial comparison from being treated as complete. Mean and median answer different questions: the mean uses magnitude, while the median asks what happened in a typical fold. The positive-ratio rule requires breadth. The concentration rule catches a total dominated by one unusually favorable period.

These values are educational policy choices, not universal statistical thresholds. A real study must define its sampling unit, holding period, dependence, costs and materiality threshold. It may need far more observations.

Step 2: Create the prospective ledger

Save the complete script below as lab_14.py. Each row has a stable fold identifier, a status, a coverage fraction, and a challenger-minus-baseline difference. Unresolved rows have no difference and are excluded from outcome statistics. They remain visible in the ledger instead of being converted to zero or silently deleted.

from decimal import Decimal
from statistics import median, pstdev

policy = {
    "min_resolved": 6,
    "min_coverage": Decimal("0.80"),
    "min_positive_ratio": Decimal("0.67"),
    "max_concentration": Decimal("0.50"),
}

ledger = [
    {"fold": "F1", "status": "resolved", "coverage": "1.00", "delta": "0.30"},
    {"fold": "F2", "status": "resolved", "coverage": "1.00", "delta": "0.05"},
    {"fold": "F3", "status": "resolved", "coverage": "1.00", "delta": "0.02"},
    {"fold": "F4", "status": "resolved", "coverage": "1.00", "delta": "-0.04"},
    {"fold": "F5", "status": "resolved", "coverage": "0.75", "delta": "-0.03"},
    {"fold": "F6", "status": "unresolved", "coverage": "0.50", "delta": None},
    {"fold": "F7", "status": "unresolved", "coverage": "0.00", "delta": None},
]

resolved = [row for row in ledger if row["status"] == "resolved"]
unresolved = [row for row in ledger if row["status"] != "resolved"]
deltas = [Decimal(row["delta"]) for row in resolved]
coverages = [Decimal(row["coverage"]) for row in resolved]
positive = [value for value in deltas if value > 0]

mean_delta = sum(deltas) / len(deltas)
median_delta = median(deltas)
positive_ratio = Decimal(len(positive)) / len(deltas)
concentration = max(positive) / sum(positive)

checks = {
    "sample_size": len(resolved) >= policy["min_resolved"],
    "coverage": min(coverages) >= policy["min_coverage"],
    "mean": mean_delta > 0,
    "median": median_delta > 0,
    "positive_ratio": positive_ratio >= policy["min_positive_ratio"],
    "concentration": concentration <= policy["max_concentration"],
}
failed = [name for name, passed in checks.items() if not passed]
decision = "PROMOTE_TO_NEXT_TEST" if not failed else "KEEP_RESEARCHING"

print(
    "policy "
    f"min_resolved={policy['min_resolved']} "
    f"min_coverage={policy['min_coverage']:.2f} "
    f"min_positive_ratio={policy['min_positive_ratio']:.2f} "
    f"max_concentration={policy['max_concentration']:.2f}"
)
print(f"ledger total={len(ledger)} resolved={len(resolved)} unresolved={len(unresolved)}")
print("resolved_ids=" + ",".join(row["fold"] for row in resolved))
print("unresolved_ids=" + ",".join(row["fold"] for row in unresolved))
print(
    "metrics "
    f"mean_delta={mean_delta:.2f} median_delta={median_delta:.2f} "
    f"positive_ratio={positive_ratio:.2f} stdev={pstdev(deltas):.2f} "
    f"coverage_floor={min(coverages):.2f} concentration={concentration:.2f}"
)
print("gate " + " ".join(f"{name}={'PASS' if passed else 'FAIL'}" for name, passed in checks.items()))
print(f"decision={decision} failed=" + ",".join(failed))
print("data=synthetic external_calls=0 orders_submitted=0")

Run it in an empty directory with python -I lab_14.py. Isolated mode ignores user-level Python configuration. The complete expected output is:

policy min_resolved=6 min_coverage=0.80 min_positive_ratio=0.67 max_concentration=0.50
ledger total=7 resolved=5 unresolved=2
resolved_ids=F1,F2,F3,F4,F5
unresolved_ids=F6,F7
metrics mean_delta=0.06 median_delta=0.02 positive_ratio=0.60 stdev=0.12 coverage_floor=0.75 concentration=0.81
gate sample_size=FAIL coverage=FAIL mean=PASS median=PASS positive_ratio=FAIL concentration=FAIL
decision=KEEP_RESEARCHING failed=sample_size,coverage,positive_ratio,concentration
data=synthetic external_calls=0 orders_submitted=0

Step 3: Interpret the rejection

The aggregate result looks favorable. Mean and median are both positive. That is not enough. Only five folds are resolved, the weakest resolved fold has 75 percent coverage, only three of five folds are positive, and F1 supplies 81 percent of all positive improvement.

The correct result is not REJECT_FOREVER. It is KEEP_RESEARCHING. This distinction matters because a promotion gate controls what happens next; it does not claim to know the permanent truth about a strategy. The challenger needs more prospectively collected, sufficiently covered evidence. The failed folds and unresolved rows stay in the record.

Do not repair the outcome by reducing min_resolved to five, changing 67 percent to 60 percent, or removing F5 after running the script. Those edits would use the observed result to tune the gate. If the policy itself is wrong, version it, explain why, and evaluate the revised policy on later evidence.

Step 4: Turn the fixture into an audit trail

A production research ledger needs more than the four fields shown here. Record the source version, candidate rule version, baseline version, start and end times, cost model, eligibility rule, missing-data reason, and who approved the gate. Record every candidate tried, not only the winner. Store raw observations separately from derived metrics so a reviewer can recompute the decision.

Keep promotion separate from execution. Passing a research gate may authorize another controlled evaluation. It should not silently grant wallet access, change a risk limit, or submit an order. The inspected repository follows the same useful separation when it labels current market confirmation as not being a fill.

Troubleshooting

If median or pstdev fails, confirm that all resolved deltas are Decimal values and unresolved rows are excluded. If the standard deviation prints 0.00, check that you did not replace every delta with its mean. If concentration divides by zero, the fixture has no positive deltas; define that case as a failed concentration check instead of pretending a percentage exists. If your output order differs, remember that dictionaries preserve insertion order in current Python, but an explicit list of check names is safer when porting the example.

For live research, duplicate fold IDs, overlapping evaluation windows, revised source data and incomplete cost records are evidence problems, not formatting problems. Reject or quarantine them before calculating a score.

Evidence and limits

This lab verifies the behavior of one predeclared gate on seven invented ledger rows. It demonstrates that a positive aggregate can still fail because evidence is insufficient or concentrated. It does not establish independence between real folds, calibrate the thresholds, estimate a Sharpe ratio, prove model accuracy, or demonstrate profitable learning.

Bailey and López de Prado's Deflated Sharpe Ratio paper explains why performance estimates can be inflated by multiple testing, non-normal returns, and short samples. We do not calculate that statistic here because five synthetic fold differences are not an adequate return history and the number of tried strategies is not defined. The paper supports tracking selection and sample limitations; it does not validate this custom gate or this repository.

Completion check: You should be able to name all four failed checks, explain why unresolved rows were not changed to zero, and reproduce the exact output without credentials or network access.

Next experiment: Append later synthetic folds prospectively, freeze the policy version, and compare a simple fixed gate with an uncertainty interval that respects the declared sampling unit. The next edition should test whether the decision changes because evidence accumulated, not because the rule was rewritten after the fact.

Primary references