A research result becomes useful when another person can inspect how it was produced. A screenshot of a score is not enough. They need the question, evidence, rules, command and limitations. This companion guide supports our existing research series; it does not replace the next numbered edition.

Start with one question

Choose a narrow question, such as whether two admission rules accept different observations from the same evidence set. Decide which comparison will answer it before running either rule. Keep a simple baseline. If you change both admission and scoring at once, a different output cannot tell you which change mattered.

Write an acceptance condition for the experiment itself: another reviewer can run the same command against the same inputs and explain any remaining difference. That condition is separate from whether the candidate is better than the baseline.

Freeze inputs and their availability

Save a bounded input file and a checksum. Include source observation time, retrieval time and the time the observation became available to your decision. An observation about an earlier event may have arrived later. A replay must not give the agent information it could not have had at the decision timestamp.

For a hypothetical example, a quote observed at 10:00:00 but first available at 10:00:05 cannot support a decision made at 10:00:02. The comparison should exclude it with an explicit reason. Do not repair the replay using a later quote unless the experiment is specifically studying that repair.

Keep the comparison controlled

Save the repository commit and the baseline and candidate configuration. Run both against the frozen input. Keep fees, missing-data behavior and evaluation windows fixed. Record admitted and rejected records, including reasons. Do not retain only the candidate's successful cases.

For learned models, fit preprocessing only on the training portion. The scikit-learn documentation explains how information leaking from evaluation data into preparation can distort results: common pitfalls. Decision-time availability requires a separate check even when train and test rows are different.

Package the replay

Include a README, dependency versions, one run command and expected output. Use a small synthetic example when real inputs cannot be shared, and label it synthetic. Keep customer data, credentials and account details out of the package.

A practical README answers five questions: what is being tested, which inputs are used, how to run it, what output to expect and what the experiment cannot establish. Ask a reviewer to follow it from a fresh environment. An undocumented local file is a reproducibility failure worth fixing.

Report uncertainty

State sample size, exclusions, assumptions and differences from live execution. A reproducible comparison does not establish future profitability. External actions can produce partial fills, uncertain responses and later reconciliation work that a local example does not reproduce.

Download the research reproducibility checklist. Continue with Edition 10's frozen-evidence comparison and subscribe to the research journal.