← Selected work KV / Research 03 Paper ↗
ICAIF 2026 submission · project lead · SCBX R&D

The ticker is gone.
Can you find it anyway?

Below is one real, masked 20-day OHLCV packet from the study's public audit artifact. Its dates are relative. Its ticker is absent. A second provider still has the underlying market record.

Switch what the model sees.

Each protocol removes the obvious label while retaining a different numeric payload. The audit attacks the fields that remain, not a ticker hidden elsewhere in the packet.

Authentic packet excerptKTD-style

Authentic closing-price series from frozen packet 44C. This compact view omits OHLC and volume.

hiddenticker + calendar dateretained20 days of OHLCV + derived features
ticker + date Top-1 99.0%

98 of 99 cross-provider packets recovered jointly.

conditioned null0.011%
Audit fidelityPaper-specified representation stress test
What is actually shownReduced U.S.-equity numeric payload; not the complete KTD-Fin system
Period targetYes · within ±5 trading days

Change the clue.
Watch the field reorder.

Each tab starts from a frozen Yahoo packet and searches the study's Alpaca/IEX candidate index. The interface replays the declared single-channel matcher. It does not query a live market service.

Frozen packet
Exposure
Released identityASSET_7119 · relative days only
Answer keyKMI · 2025-01-08
Exact distanceL2 dollars
Masked packetReleased close level
released packetselected candidate
packet start26.34
candidate start26.34

This view retains raw price level.

Exact answer-key row rank#1answer-key distance 0.0405 · recovered first
Top-2 distance gap2.9459runner-up distance − best distance
What changedAt 20 days, every channel agrees.Single-packet diagnostic

The 5- and 10-day controls truncate the same frozen 20-day packet. They are packet-level prefix diagnostics, not the paper's separately sampled exposure cells. "Exact row rank" uses one ticker and one start date; the aggregate paper metric allows ±5 trading days. The compact export stores candidate close and volume paths only. For OHLC shape, the overlay shows normalized close; the export generator computed the exact distance and rank from the full return, range, and overnight-gap arrays.

Same index.
Different geometry.

The workbench above prints the search as a ledger. This stage draws it as a place: every candidate window owns one anonymous cell in rank order, the real top five stand named, and the answer key marks its exact rank. Switch the clue — the index never moves. The answer does.

Frozen packet · L20
Matching channel
Staging
Answer-key row rankof — · L20 primary
Answer-key distance
What changedexact row rank = one ticker + one start date
Named records · real top fiveof — searched

Select a record to inspect its exported path. Teal rows are the answer key.

What is exact

Counts, names, ranks, and distances come from the frozen export behind the workbench above: 102,546 candidate windows per L20 search; each packet's real top five with their real distances and exported close paths; the answer key's exact rank and distance in every packet-and-channel cell; the R3 collision class of 1,705 records and its 18 shipped sample records with true tickers and terminal dates. Nothing here recomputes the matcher — selections replay recorded outputs.

What is drawn

The carpet depicts scale, not distance. Anonymous windows take schematic grid places in rank order; their spacing carries no measurement, pillar heights are uniform, and row shading is decorative. The sweep, the traveling beacon, and the probe are theater. Shape views plot normalized close because the compact export stores close and volume only — printed shape distances were computed from full return, range, and overnight-gap arrays. This stage shows L=20 only; the 5- and 10-day cells above are prefix diagnostics of the same frozen packet, and "exact row rank" means one ticker and one start date where the aggregate paper metric allows ±5 trading days. It is a representation audit of public U.S.-equity data — not evidence of model memorization, spontaneous exploitation, personal-privacy harm, or trading advantage.

The winning clue changes by case.

For KMI, all five channels put the exact answer-key row first after 20 days. For APTV, raw price level ranks it 9,463rd, while rebased close, returns, and OHLC shape rank it first. For AZO, relative volume ranks it 7,890th, while the four price-derived channels recover it.

Cross-provider differences can break one fingerprint without breaking the others. That is why the paper reports each channel separately from the frozen composite.

Remove the price level.
The path still talks.

The left panel transforms one frozen APTV packet. The right panel reports a separate aggregate ablation over all 99 A2 packets. They share a transformation, not an estimand.

Illustrative caseFrozen APTV packet 1D2

APTV packet · original close level58.89 → 59.35

A2 aggregaten = 99 packets · L = 20

Joint Top-1 recoveryreleased levels
Frozen composite99.0%99.0 → 14.1
Price level99.0%99.0 → 5.1
Returns100.0%100.0 → 100.0
Normalized close100.0%100.0 → 100.0
OHLC shape99.0%99.0 → 99.0
Relative volume11.1%11.1 → 11.1

The aggregate fit weights price level most heavily, so its composite recovery falls after every packet is rebased. The predeclared return and normalized-close matchers do not.

A claim about a representation, not a guessing game.

  1. Separate the mechanisms.Parametric recognition, database linkage, agent-built attacks, and downstream use are different claims that need different tests.
  2. Test the released fields against another provider.Freeze the representations, candidate sources, distance functions, and null models before opening the confirmatory split.
  3. Make mitigation earn its conclusion.Reducing a release is not enough: the study tests both collision-based resistance and whether the remaining signal survives.

The packet above is a frozen exhibit. These are study-level results.

The audit evaluated cross-provider linkage on held-out packets after freezing the metric and matching rules. The graphics below summarize two findings that give the single case its context.

01 / accumulated exposureMore days, less ambiguity.

Joint recovery for the reconstructed BlindTrade trace; descriptive lengths use n = 24, while the 20-day primary cell uses n = 99.

02 / mitigation boundaryCollisions can stop lookup; usefulness still must be shown.

Sealed 600-firm study, maximum ticker Top-1 across its declared attack ensemble.

One actual R3 equivalence class

One bin pattern.
1,705 candidate records.

The six cells below are the released five-bin summaries for one frozen validation packet. Every record in the list has the exact same pattern, so the release cannot distinguish them.

M2 validation packet 1 · frozen candidate index

Showing 18 records from the collision set. The highlighted record is the answer key.

In the frozen development-packet scaling extension, which is separate from the sealed A2 split, 100/100 KTD-style packets remained unique nearest neighbors among 3.51 million ticker-date windows. Rebasing removed the level channel, but normalized-close recovery still reached 100%.

One result does not prove the next.

Move through the chain. Each stage changes the question, evaluation split, and strength of the conclusion.

A2 · one-shot sealed confirmatory splitestablished
Representation property

Can public data identify the hidden record?

Yes. Frozen cross-provider linkage recovers all three audited representations far above their conditioned nulls.

KTD-style98 / 99ticker + date
BlindTrade94 / 99ticker + date
Epistemic port79 / 375ticker
What this does not establish

It does not show that an LLM remembers, discovers, or financially exploits the identity.

WorkbenchA2 packet replay

Three selected confirmatory Yahoo packets searched against the frozen Alpaca/IEX candidate index. This is a diagnostic subset.

Aggregate recoveryA2 sealed result

One-shot confirmatory split with frozen metrics, matching rules, and conditioned nulls.

Agent constructionP1 · separate

Held-out development allocations show that tasked coding agents can build attacks. P1 does not add confirmatory A2 evidence.

Discovery + valueSeparate assays

The neutral pilot missed its gate. The downstream header assay did not establish forecast value.

This is a representation audit for public U.S.-equity data. Linkability is not proof of model memorization, spontaneous exploitation, personal-privacy harm, or trading advantage.

Read the full paper ↗