The research question

When two data sources describe the same pool, which observation should a research system keep? It is tempting to trust the provider called last. That shortcut makes arrival order a hidden rule: a stale response can replace a fresher one simply because it was processed later.

Edition 8 separated quote age, response age and decision-time availability. This edition continues that investigation by comparing and merging observations from different sources. We will use the journal to explain the engineering behind actual changes, the models we evaluate, what the evidence supports, and experiments readers can reproduce. The aim is a growing research record, with useful lessons even when a hypothesis fails.

Time: approximately 30 minutes. You need: Python 3.12 or later, a terminal, and a folder for the exercise. The lab uses only the standard library. All observations below are invented; no broker, wallet, market feed, credentials, or orders are involved.

What changed in our research code

On October 2, 2026 UTC, the trading repository gained a market-observation merge helper in app/emerging_crypto.py. The inspected change, identified by commit 32f11d3c4520fb8c837bc955f3810bc391b95f11, compares provider-response timestamps instead of letting later source processing replace a newer registry quote. The change was subsequently merged through PR #511.

It also includes chain in the pool identity, normalizes pool-address case for non-Solana chains, and preserves Solana identifier case. The accompanying tests exercise stale, future-dated, and missing timestamps, separate chains, and Solana case sensitivity.

Those are inspected code and test changes. They do not establish that this repair improved trading results, was deployed to every service, produced a funded fill, or made a strategy profitable. Those questions need separate operational and outcome evidence. Today we will reproduce the underlying failure mode without making any of those claims.

Step 1: Specify the observation contract

A price is only useful alongside its identity and time. Our miniature record carries a chain, pool identifier, price string, and timezone-aware response-receipt timestamp. The chain keeps equal-looking pool identifiers on different networks separate. The timestamp lets us compare responses without relying on list order.

Receipt time is when a response was received, not necessarily when its underlying quote was produced. A production adapter should preserve both timestamps when supplied, the provider identity, the instrument identity, and any applicable quote expiry. A recently received cached response can still contain old market data. This lab teaches one boundary; it does not solve every freshness problem.

Decide the comparison rules before seeing the results: reject timestamps in the future, reject observations older than 120 seconds, and retain the newest accepted observation for each exact chain/pool identity. The 120-second cutoff is a teaching parameter, not a recommended trading setting.

Step 2: Reproduce the overwrite bug

Imagine the registry observes the fictional Robinhood pool 0xABC at price 2.00 five seconds ago. A direct provider then returns 1.00 from five minutes ago. A dictionary update in source-processing order can leave 1.00 in the merged record even though 2.00 was received more recently.

A second direct response claims to arrive one minute in the future. A rule that simply chooses the largest timestamp would select this invalid record. The useful comparison therefore requires both a timestamp-validity check and an ordering rule.

We also include a fictional Base pool with the same address text and two case-distinct Solana identifiers. These expose identity bugs that a price-only demonstration would miss.

Step 3: Build and run a deterministic lab

Save this as lab_09.py, then run python lab_09.py. Keep the fixed clock; replacing it with the current clock makes this teaching fixture dependent on the day you run it.

from datetime import datetime, timedelta, timezone

now = datetime(2026, 10, 2, 12, 0, tzinfo=timezone.utc)
max_age_seconds = 120

def observation(chain, pool, price, seconds_ago):
    return {
        "chain": chain, "pool": pool, "price": price,
        "received_at": (now - timedelta(seconds=seconds_ago)).isoformat(),
    }

def identity(row):
    chain = row["chain"].lower()
    pool = row["pool"]
    return chain, pool if chain == "solana" else pool.lower()

def timestamp(row):
    stamp = datetime.fromisoformat(row["received_at"])
    if stamp.tzinfo is None:
        raise ValueError("timezone required")
    return stamp.astimezone(timezone.utc)

registry = [
    observation("robinhood", "0xABC", "2.00", 5),
    observation("base", "0xABC", "3.00", 6),
    observation("solana", "AbCd", "4.00", 7),
]
direct = [
    observation("robinhood", "0xabc", "1.00", 300),
    observation("robinhood", "0xabc", "9.00", -60),
    observation("solana", "abcd", "5.00", 8),
]

selected = {}
rejected = {"stale": 0, "future": 0}
for row in registry + direct:
    stamp = timestamp(row)
    age = (now - stamp).total_seconds()
    if age < 0:
        rejected["future"] += 1
        continue
    if age > max_age_seconds:
        rejected["stale"] += 1
        continue
    key = identity(row)
    prior = selected.get(key)
    if prior is None or stamp > timestamp(prior):
        selected[key] = row

assert selected[("robinhood", "0xabc")]["price"] == "2.00"
assert selected[("base", "0xabc")]["price"] == "3.00"
assert ("solana", "AbCd") in selected
assert ("solana", "abcd") in selected
assert rejected == {"stale": 1, "future": 1}
assert len(selected) == 4

print("selected_identities=4 rejected_stale=1 rejected_future=1")
print("robinhood_price=2.00 base_price=3.00")
print("solana_case_sensitive_identities=2")
print("checks=6_passed data=synthetic no_orders_submitted")

Expected output:

selected_identities=4 rejected_stale=1 rejected_future=1
robinhood_price=2.00 base_price=3.00
solana_case_sensitive_identities=2
checks=6_passed data=synthetic no_orders_submitted

The six assertions check the retained Robinhood price, the separate Base price, both case-distinct Solana identities, rejection counts, and the final number of identities. Matching this output demonstrates the declared behavior on these rows. It does not validate a market strategy.

Step 4: Explain why the result is reproducible

The two Robinhood direct rows are rejected for different reasons: one is too old for this exercise, and one is future-dated. The freshest valid registry observation remains. The Base row survives separately because chain is part of the key. The Solana rows survive separately because their pool identifiers differ in case.

The comparison uses >, so the earlier accepted row wins an exact timestamp tie. That is an explicit local policy, not a universal answer. A production system may need deterministic provider priority, an agreement check, or to hold conflicting observations for review. Preserve the original rows so the decision can be audited later.

This tutorial is deliberately stricter than the inspected merge helper: it adds an explicit age cutoff and rejects future observations altogether. It is an educational adaptation, not a claim that the production helper has this complete contract.

Step 5: Extend the experiment before connecting a feed

Start with a duplicate of the fresh Robinhood row and confirm that the selected identity count does not increase. Next, change the direct timestamp to two seconds ago and its price to 2.10. Predict the retained value before running the script, then update the assertion and inspect the output.

Try a timezone-naive timestamp such as 2026-10-02T12:00:00. The lab raises a clear error. A real ingestion worker would record a rejected-observation reason and continue processing other rows according to its contract. Do not silently assign a guessed timezone to an ambiguous record.

Finally, add a provider identifier and quote-production time. Design a test where receipt time is fresh but the quote-production time is old. State which timestamp bounds economic freshness and which timestamp measures transport latency. Keep those two questions separate.

Troubleshooting

If Python raises an assertion error after you edit a row, compare the expectation with the new timestamp, cutoff, and identity. Keep prices as strings here; this exercise performs no price arithmetic. If you add calculations later, specify decimal precision and units rather than silently mixing currency values.

The pool strings are illustrative identifiers. This script does not verify address syntax, contract safety, liquidity, supported routes, or buy/sell executability. A correct merge does not grant permission to trade.

Completion check and next experiment

You are finished when you can explain the old overwrite failure, reproduce the six checks, preserve chain and case semantics, and distinguish response age from quote age. Save the script and its output with a note explaining your tie and rejection policies.

Our next research question is how to test a model or candidate ranker on identical, fresh inputs. We will keep changes in data preparation separate from changes in model behavior, so an apparent improvement can be traced to the right cause.

Review Edition 8's decision-time evidence contract or start from Edition 1.

Primary technical reference

Python's official datetime documentation explains timezone-aware timestamps and ISO-format parsing. The repository repair identified above is the source of this edition's operational question; the lab is an original synthetic demonstration.