The research question

A trading decision is easy to label only after its outcome is complete. Before then, an open position is not a loss, a reported sale is not necessarily a verified close, and sale proceeds alone are not profit. If we force those records into WIN or LOSS, we give a model labels that the evidence does not support.

Edition 10 compared admission policies and rankers on one frozen evidence set. It deliberately left outcomes unmeasured. Today we add outcome records and ask: which trades have enough evidence for a verified label, and which must stay open or incomplete?

Prerequisites: Python 3.12 or later, a terminal, and about 35 minutes. The lab uses only the standard library. Its six trades, fills, prices and fees are invented. It has no network calls, credentials, broker connection, wallet or order submission.

What changed in the inspected trading code

On October 7, 2026 UTC, PR #580 merged into main in Primus19/memecoine-mcp-server. We inspected its head, 95b51fa09159f0d33184b2c830d92cf89a5d63b6, and merge commit, 0b179248cc72c4f0b84f47b5a84f2ace6e2411a3. The relevant files include app/solana_accounting.py, app/trading_digest.py, app/trading_report_status.py, services/solana-executor/index.mjs, services/solana-executor/operations.mjs and their regression tests.

The change separates several facts that an earlier report could blur:

  1. A funded position may be open and supervised.
  2. A sale may be reported while its receipt is still pending or inconsistent.
  3. Finalized token and USDC deltas may verify one fill.
  4. A complete buy, every partial sale and the final sale are required to reconcile the round trip.
  5. Reconciled USDC cash flow is reported before network fees, while network fees remain separate.

It also moves funded protective checks onto an independent five-second supervision loop. Entry and exit mutations share a serialized lane, while slower research and receipt work do not own that lane. A recent sell quote can be reused only when its input identity, amount, transaction, request identity and expiry checks agree.

Those are inspected code behaviors and tests. They do not prove an exit was executed successfully, that a model improved, or that a trading result was profitable.

What the deployment evidence supports

Railway deployed merge commit 0b179248cc72c4f0b84f47b5a84f2ace6e2411a3 successfully to the Solana early executor. In a read-only interval from 09:55:55 through 10:04:56 UTC on October 7, the service emitted 109 SOLANA_PROBE_SUPERVISION heartbeats. Every inspected heartbeat reported no supervision error and zero open funded probe positions.

This verifies that the new supervision loop was running during that interval. Because the open count was zero, it does not verify the liquidation branch, a funded close, receipt reconciliation or realized profit. The distinction matters: a healthy idle supervisor is operational evidence, not outcome evidence.

1. Declare the outcome states before scoring

Our lab permits five states. VERIFIED_CLOSED means the lifecycle contains a buy, a final sale, unique event identities, finalized receipts and a flat token quantity. OPEN means there is no final sale yet. RECEIPT_PENDING means a close was reported but at least one receipt is not finalized. MISSING_BUY means sale proceeds lack their cost basis. QUANTITY_NOT_FLAT means the recorded sells do not account for the purchased quantity.

Only VERIFIED_CLOSED receives a win or loss label. The others remain visible, with a reason, but do not enter the labeled denominator. This is not the same as deleting them. We report coverage so a reader can see how much of the intended cohort is actually evaluable.

The experiment uses a fixed evaluation cutoff. Trade C is still open at that cutoff. It may close later, but later evidence cannot be silently inserted into this evaluation snapshot. A future evaluation can advance the cutoff and create a new, versioned result.

2. Reconcile cash flow and quantity together

USDC cash flow alone is insufficient. Trade E has final-sale proceeds but no recorded buy. Treating its 0.90 USDC as profit would ignore what the position cost. Trade F has a buy and final sale, but its token deltas leave 0.1 token unaccounted for. Calling that lifecycle closed would hide remaining exposure or missing evidence.

For a complete lifecycle, gross cash flow is the sum of buy and sale USDC deltas. Network fees are summed separately and subtracted to obtain the lab's net figure. We use Decimal from strings so the accounting fixture does not introduce binary floating-point artifacts. The fee figures here are synthetic USDC estimates, not observed network charges.

3. Run the complete synthetic lab

Save the following as lab_11.py. Run it from a new folder with python -I lab_11.py. Isolated mode prevents user site packages and Python environment settings from changing the exercise.

from dataclasses import dataclass
from decimal import Decimal


@dataclass(frozen=True)
class Fill:
    event_id: str
    trade: str
    action: str
    token_delta: Decimal
    usdc_delta: Decimal
    network_fee_usdc: Decimal
    finalized: bool = True


fills = (
    Fill("a-buy", "A", "BUY", Decimal("1"), Decimal("-1.00"), Decimal("0.01")),
    Fill("a-part", "A", "PARTIAL_SELL", Decimal("-0.4"), Decimal("0.50"), Decimal("0.005")),
    Fill("a-final", "A", "FINAL_SELL", Decimal("-0.6"), Decimal("0.70"), Decimal("0.005")),
    Fill("b-buy", "B", "BUY", Decimal("1"), Decimal("-1.00"), Decimal("0.01")),
    Fill("b-final", "B", "FINAL_SELL", Decimal("-1"), Decimal("0.80"), Decimal("0.01")),
    Fill("c-buy", "C", "BUY", Decimal("1"), Decimal("-1.00"), Decimal("0.01")),
    Fill("d-buy", "D", "BUY", Decimal("1"), Decimal("-1.00"), Decimal("0.01")),
    Fill("d-final", "D", "FINAL_SELL", Decimal("-1"), Decimal("1.20"), Decimal("0.01"), False),
    Fill("e-final", "E", "FINAL_SELL", Decimal("-1"), Decimal("0.90"), Decimal("0.01")),
    Fill("f-buy", "F", "BUY", Decimal("1"), Decimal("-1.00"), Decimal("0.01")),
    Fill("f-final", "F", "FINAL_SELL", Decimal("-0.9"), Decimal("1.10"), Decimal("0.01")),
)


def reconcile(trade, rows):
    actions = [row.action for row in rows]
    if "FINAL_SELL" not in actions:
        return {"trade": trade, "state": "OPEN"}
    if "BUY" not in actions:
        return {"trade": trade, "state": "MISSING_BUY"}
    if len({row.event_id for row in rows}) != len(rows):
        return {"trade": trade, "state": "DUPLICATE_EVENT"}
    if not all(row.finalized for row in rows):
        return {"trade": trade, "state": "RECEIPT_PENDING"}
    if sum((row.token_delta for row in rows), Decimal("0")) != 0:
        return {"trade": trade, "state": "QUANTITY_NOT_FLAT"}
    gross = sum((row.usdc_delta for row in rows), Decimal("0"))
    fees = sum((row.network_fee_usdc for row in rows), Decimal("0"))
    net = gross - fees
    return {"trade": trade, "state": "VERIFIED_CLOSED", "gross": gross,
            "fees": fees, "net": net, "label": "WIN" if net > 0 else "LOSS"}


trades = tuple("ABCDEF")
results = tuple(reconcile(name, tuple(row for row in fills if row.trade == name))
                for name in trades)
verified = tuple(row for row in results if row["state"] == "VERIFIED_CLOSED")
excluded = tuple(row for row in results if row["state"] != "VERIFIED_CLOSED")
gross = sum((row["gross"] for row in verified), Decimal("0"))
fees = sum((row["fees"] for row in verified), Decimal("0"))
net = sum((row["net"] for row in verified), Decimal("0"))

assert [row["trade"] for row in verified] == ["A", "B"]
assert [row["label"] for row in verified] == ["WIN", "LOSS"]
assert verified[0]["net"] == Decimal("0.18")
assert verified[1]["net"] == Decimal("-0.22")
assert gross == Decimal("0.00")
assert fees == Decimal("0.04")
assert net == Decimal("-0.04")
assert [row["state"] for row in excluded] == [
    "OPEN", "RECEIPT_PENDING", "MISSING_BUY", "QUANTITY_NOT_FLAT"]
assert len(verified) == 2 and len(excluded) == 4
assert next(row for row in results if row["trade"] == "C")["state"] == "OPEN"
assert next(row for row in results if row["trade"] == "D")["state"] == "RECEIPT_PENDING"
assert next(row for row in results if row["trade"] == "E")["state"] == "MISSING_BUY"
assert next(row for row in results if row["trade"] == "F")["state"] == "QUANTITY_NOT_FLAT"
assert all("label" not in row for row in excluded)
assert fills == tuple(fills)

print("records=6 verified_closed=2 open=1 incomplete=3")
print("verified_labels=A:WIN,B:LOSS win_fraction_verified=1/2")
print("verified_gross_usdc=0.00 network_fees_usdc=0.04 net_usdc=-0.04")
print("excluded=C:OPEN,D:RECEIPT_PENDING,E:MISSING_BUY,F:QUANTITY_NOT_FLAT")
print("coverage=2/6 outcomes_unlabeled=4")
print("checks=15_passed data=synthetic no_orders_submitted")

Expected complete standard output:

records=6 verified_closed=2 open=1 incomplete=3
verified_labels=A:WIN,B:LOSS win_fraction_verified=1/2
verified_gross_usdc=0.00 network_fees_usdc=0.04 net_usdc=-0.04
excluded=C:OPEN,D:RECEIPT_PENDING,E:MISSING_BUY,F:QUANTITY_NOT_FLAT
coverage=2/6 outcomes_unlabeled=4
checks=15_passed data=synthetic no_orders_submitted

We executed the script in an isolated directory with a cleared environment and compared its complete output with this block. The fifteen assertions cover both verified labels, exact Decimal totals, all four excluded states, label absence for excluded records and preservation of the input fixture.

4. Read the result without overstating it

Trade A has a verified buy, partial sale and final sale. Its gross synthetic cash flow is 0.20 USDC. After 0.02 USDC of synthetic network fees, its net is 0.18, so its label is WIN.

Trade B has a verified buy and final sale. Its gross cash flow is negative 0.20 USDC, and fees reduce the result to negative 0.22, so its label is LOSS.

The two verified trades sum to zero gross cash flow and negative 0.04 after the lab's fees. The verified win fraction is one out of two. That fraction describes this tiny invented fixture only. It is not a model accuracy estimate, a production win rate or evidence of expected profitability.

Four of six records remain unlabeled, so coverage is two out of six. Reporting 1/2 without 2/6 would conceal how little of the cohort has verified outcomes. Reporting 1/6 would be worse because it would silently convert unknown states into losses.

5. Keep missing outcomes visible

An open record resembles a right-censored observation: at the evaluation cutoff, the eventual outcome is unknown. A 2016 paper by Vock and colleagues studies this problem in health records, where follow-up can end before an outcome is observed. The domain is different, but the measurement warning transfers: discarding unknown outcomes or treating them as known non-events can bias evaluation.

We are not applying the paper's inverse probability weighting method to trading data here. That would require justified assumptions about why outcomes are missing and a much larger prospective dataset. Our smaller step is to preserve the state, exclude it from binary labels, and print coverage. This prevents a false label while leaving the harder statistical question open.

Receipt-pending and incomplete-cost records are not ordinary censoring. They are evidence-quality failures. Keep their reasons separate from OPEN, because the remedy differs. An open trade needs more follow-up. A pending receipt needs finalization or reconciliation. A missing buy needs restored cost-basis evidence. A quantity mismatch needs the missing fill or remaining position reconciled.

Troubleshooting

If Trade A is not flat, check that its partial and final token deltas sum to negative one. If the gross total differs, retain the signs: buys consume USDC and sales produce USDC. If a pending receipt receives a label, ensure the finalization check occurs before arithmetic.

Run the script without -O; optimized mode removes assert statements. If you add duplicate events, give the duplicate the same event_id and predict the DUPLICATE_EVENT state before extending the expected output. In a real system, never invent a missing receipt, price, quantity or fee to make the ledger balance.

Evidence limits and completion check

The lab checks a declared policy against six synthetic records. It does not validate Solana receipt formats, token decimals, currency conversion, provider completeness, execution price, fee estimation or the deployed reconciler. The repository tests address some code paths, while the Railway heartbeat proves only that the deployed supervisor was alive and idle in the inspected window.

You are finished when the complete output matches, you can explain why each excluded trade lacks a label, and you can reconcile A and B by hand. Save the code, output, evaluation cutoff and input records together.

The next experiment will compare two scoring rules using only outcomes that were not available when the scores were produced. We will freeze the decision-time features, label-availability time and evaluation cutoff, then test whether a ranking change survives explicit cost and coverage reporting.

Primary technical references

Python's official decimal documentation explains exact decimal arithmetic and finite-number checks. Vock and colleagues' paper, “Adapting machine learning techniques to censored time-to-event health record data”, explains why incomplete follow-up leaves outcomes unknown and why naive deletion or negative labeling can bias evaluation. Its medical method is a conceptual reference for missing outcomes, not a trading-model validation.