The research question

A candidate ranker puts possible trades in an order. An eligibility rule decides which candidates reach that ranker. Change either one and the first name on the list can change. Change both together and it becomes difficult to explain why.

In Edition 9, we kept fresh observations when merging sources. Today we freeze the resulting evidence and ask: which changes come from admitting more candidates, and which come from changing their scores? We will measure eligibility, ordering and missing evidence. Outcomes remain unmeasured.

Prerequisites: Python 3.12 or later, a terminal, and a new folder. Allow about 35 minutes. The original lab uses the standard library, seven invented candidates and two hand-written scoring rules. It needs no credentials, network access or trading software. These rules are teaching examples, not trained models or replicas of the trading executor.

The inspected change and its limits

On October 6, 2026 UTC, PR #574 was merged into main in Primus19/memecoine-mcp-server. Its head was d81c2e416a882ef1756175ad6d28b90f000d7df9; its merge commit was a76736485d31b66c0193ca78c37e1d53ae679672. We inspected app/robinhood_early_mover.py and tests/test_robinhood_funded_evidence.py at that head. The current trading branch tip during this review was 5756c55667154697bfcc36f02f6a1c7369ca7bea, an economic-calendar update.

The change introduces LEGACY_EVIDENCE_WITHOUT_UNIQUE_BUYER_CHECKS_V1. It excludes four unique-buyer failure categories from the evidence requirements used when preparing a funded entry. Those categories concern unavailable buyer evidence, incomplete or stale collection, five-minute participation, and hourly persistence. Other evidence failures still count. The inspected tests cover missing safety evidence, concentration limits, incomplete cost evidence, a forged decision and restoration of the earlier buyer requirements.

This is an admission-policy change. The code does not establish that a scoring model learned a useful new relationship. Missing buyer attribution also does not establish that there were no buyers, that buyers were independent, or that wash trading was absent.

Read-only Railway deployment evidence showed the multi-asset paper worker successfully running the merge commit. A later calendar commit was skipped for this service. At 10:03 UTC on October 6, its Robinhood runtime status reported NO_VERIFIED_FUNDED_TEST and end_to_end_swap_verified=false. That is a bounded observation of this lane, not a conclusion about every trading service or later activity. We inspected source and deployment records; we did not submit a trade, verify profitability or infer fills from a successful deployment.

1. Freeze the evidence before comparing rules

Each invented row has a name, quote age, availability age, two evidence flags, momentum points and friction points. All ages are measured relative to one fixed decision instant. A negative availability age means the row becomes available after the decision. The freshness cutoff is 120 seconds, chosen for this exercise.

core_verified stands for a simplified bundle of required checks. It is not a real safety assessment. buyers_verified=False means the lab lacks the required buyer evidence. We keep that flag visible when a policy permits the row through.

The rows use frozen data classes with primitive fields and a tuple container. A SHA-256 fingerprint of canonical JSON gives us a repeatable input identity. We sort by candidate name before hashing because input-list order should not change the evidence identity. In this fixture, names are unique. A real recorder would need an explicit instrument, venue, observation and decision identity.

2. Define two separate comparisons

The strict policy requires core evidence, timely data and buyer evidence. The relaxed policy changes only the buyer requirement. Both reject unsafe or mistimed rows. We then apply each scoring rule to each eligible cohort, producing a small two-by-two comparison.

Version 1 ranks momentum points. Version 2 subtracts ten times the friction points. These are arbitrary integers in a synthetic score, not currency, fees, expected returns or recommendations. Candidate name resolves ties so equal scores produce a stable order.

Comparing versions on the strict cohort isolates the scoring change. Comparing policies with version 1 held fixed isolates the admission change. Comparing only strict/version 1 with relaxed/version 2 would combine both effects.

3. Run the complete lab

Save the following as lab_10.py in your new folder. Run python -I lab_10.py with Python 3.12 or later. The isolated interpreter ignores Python environment configuration and user site packages. The script itself never reads credentials or opens a connection.

from dataclasses import asdict, dataclass
from hashlib import sha256
import json


@dataclass(frozen=True)
class Candidate:
    name: str
    quote_age: int
    available_age: int
    core_verified: bool
    buyers_verified: bool
    momentum: int
    friction: int


rows = (
    Candidate("A", 10, 5, True, True, 90, 2),
    Candidate("B", 20, 5, True, True, 80, 0),
    Candidate("C", 15, 5, True, False, 99, 5),
    Candidate("D", 25, 5, True, False, 85, 0),
    Candidate("E", 10, 5, False, True, 100, 0),
    Candidate("F", 121, 5, True, True, 120, 0),
    Candidate("G", 10, -1, True, True, 130, 0),
)


def fingerprint(items):
    payload = [asdict(row) for row in sorted(items, key=lambda r: r.name)]
    raw = json.dumps(payload, sort_keys=True, separators=(",", ":"))
    return sha256(raw.encode("utf-8")).hexdigest()


def eligible(items, require_buyers):
    return tuple(row for row in items
                 if 0 <= row.quote_age <= 120
                 and row.available_age >= 0
                 and row.core_verified
                 and (row.buyers_verified or not require_buyers))


def rank(items, version):
    if version not in ("v1", "v2"):
        raise ValueError("unknown ranker")
    def score(row):
        return row.momentum if version == "v1" else row.momentum - 10 * row.friction
    return tuple(row.name for row in sorted(items, key=lambda r: (-score(r), r.name)))


def names(items):
    return {row.name for row in items}


strict = eligible(rows, require_buyers=True)
relaxed = eligible(rows, require_buyers=False)
added = names(relaxed) - names(strict)
common = tuple(row for row in relaxed if row.name in names(strict))
baseline_hash = fingerprint(rows)
strict_v1 = rank(strict, "v1")
strict_v2 = rank(strict, "v2")
relaxed_v1 = rank(relaxed, "v1")
relaxed_v2 = rank(relaxed, "v2")

assert names(strict) == {"A", "B"}
assert names(relaxed) == {"A", "B", "C", "D"}
assert names(strict) <= names(relaxed)
assert all(not row.buyers_verified for row in relaxed if row.name in added)
assert not {"E", "F", "G"} & names(relaxed)
assert rank(common, "v1") == strict_v1
assert strict_v1 == ("A", "B") and strict_v2 == ("B", "A")
assert relaxed_v1 == ("C", "A", "D", "B")
assert relaxed_v2 == ("D", "B", "A", "C")
assert rank(eligible(tuple(reversed(rows)), False), "v2") == relaxed_v2
assert fingerprint(tuple(reversed(rows))) == baseline_hash
assert fingerprint(rows) == baseline_hash

print("input_rows=7 same_input_after_comparison=True")
print("strict_eligible=2 relaxed_eligible=4 added=C,D")
print("added_buyer_evidence=unverified core_or_time_rejections=E,F,G")
print("strict_v1=A,B strict_v2=B,A")
print("relaxed_v1=C,A,D,B relaxed_v2=D,B,A,C")
print("common_cohort_v1=A,B ordering_unchanged=True")
print("checks=12_passed outcomes=unmeasured data=synthetic no_orders_submitted")

Expected complete standard output:

input_rows=7 same_input_after_comparison=True
strict_eligible=2 relaxed_eligible=4 added=C,D
added_buyer_evidence=unverified core_or_time_rejections=E,F,G
strict_v1=A,B strict_v2=B,A
relaxed_v1=C,A,D,B relaxed_v2=D,B,A,C
common_cohort_v1=A,B ordering_unchanged=True
checks=12_passed outcomes=unmeasured data=synthetic no_orders_submitted

We executed this code in a separate directory with a cleared environment and checked the entire output against this block. The twelve assertions cover cohort identities, inclusion, evidence flags, vetoes, shared-cohort ordering, both rankers, input permutation and unchanged input content.

4. Explain the changed first choice

The strict policy admits A and B. Version 1 chooses A first because 90 momentum points exceed 80. Version 2 gives A 70 points after its friction penalty and leaves B at 80, so B moves ahead. The eligible population stayed identical; the score formula caused this change.

The relaxed policy admits C and D as well. With version 1 unchanged, C takes first place on 99 momentum points. That change arises from a wider population. Restrict this same result to the common A/B cohort and their ordering remains A, B.

On the relaxed cohort, version 2 chooses D. Its 85 points exceed B's 80, A's 70 and C's 49. Both policy and ranker now differ from the original strict/version 1 configuration. Calling D an improved model prediction would require an outcome experiment that this script has not performed.

E lacks core evidence. F's quote is 121 seconds old. G is unavailable at the decision instant. None is admitted under either policy, even though their momentum scores are high. A ranking formula should never quietly become a way around an upstream evidence requirement.

5. Keep the denominator and uncertainty visible

Record seven inspected rows, two strict eligible rows, four relaxed eligible rows and two newly admitted rows with unverified buyer evidence. These counts explain the population shift. They do not measure predictive accuracy.

There are no outcome labels in this lab. We cannot calculate a win rate, estimate an execution-adjusted return, or choose a production policy from the output. Adding labels later requires a fixed horizon, observations available at decision time, explicit costs and rules for missing outcomes. Report coverage alongside any average. Excluding hard-to-measure cases can change the denominator and make two results incomparable.

A shared-cohort comparison answers what happens on candidates both policies admit. It says nothing about the quality of the newly admitted candidates. Keep those two questions separate in your experiment record.

Troubleshooting and extensions

An assertion failure after an edit means one of the declared expectations changed. Inspect the altered row, policy and score before updating the expected output. Do not remove the assertion merely to finish the exercise. Avoid python -O, which disables assertions.

If your output changes when you reverse the rows, inspect the tie rule and hashing order. If G becomes eligible, check the sign of available_age. If C's missing evidence disappears from your report, preserve the original flag rather than overwriting it with the admission decision.

Try giving A and B equal version 2 scores and predict which name wins the tie. Then change only F's age to 120 and predict the new cohort. Keep these extensions in separate copies, with new expected outputs. The printed counts describe the supplied fixture and must be revised when it changes.

Completion check and the next experiment

Finish when the complete output matches, you can explain all four rankings, and you can identify which newly admitted candidates have unverified evidence. Save the script, output, policy definitions and input fingerprint together. A fingerprint detects content changes; it does not prove that an input was truthful, timely or independently collected.

Next, we will attach fixed-horizon synthetic outcomes to a frozen decision record and keep those labels out of the ranking inputs. That experiment will test how missing outcomes and cost assumptions affect evaluation coverage. It will remain an educational simulation until a separately verified operational dataset supports stronger claims.

Primary technical references

Python's official data class documentation describes frozen=True, and its hashlib documentation specifies SHA-256 and digest construction. These support the record and fingerprint mechanics. The inspected repository change supplies the research question; the fictional scores, candidates and comparison are original educational work.