Selected work KV / QMDTrust Paper
Source-aware decision-making · AAAI submission

The board is hidden.
So is the liar.

A Spotter sees the full board and answers yes-or-no questions. Two captains hear the same answers. Only one admits that the source may be inverted.

Same board. Same answer. Two beliefs.

A harbor diorama of board B01 with belief drawn as terrain. Signal the Spotter, use the recorded value-of-information choice, or reveal a cell yourself. Every event reaches both captains. Watch their beliefs pull apart and see the brass needle swing toward invert. Both posteriors are recomputed live in the browser. The ocean and the gunfire are theater.

Still image · click to load the world Open on its own page ↗

Move δ. Watch trust turn into inversion.

δ is the chance that a stationary source flips a truthful binary answer. Both ends are useful after the captain identifies the orientation. At δ = .50, the channel carries no task information.

orientationinvert-leaning
channel capacity53.1%
probe bound · 95%10

The probe count is the sufficient bound ⌈2 ln(20)/(1−2δ)²⌉, not an observed task-question count. It identifies the channel orientation, not the board. The selected questions must still distinguish the surviving legal boards.

Measured final-board F1Diagnosis pays only when the recovered signal is worth its questions.
SymmetricMixturefull measured sweep Plannercorners + invert interior NoisyMixturecorners + invert interior Channel capacitytheoretical reference
δ=.90SymmetricMixture F1 .732 · +.107 over asking nothing

Dots are measured means; whiskers show ±1 SEM. The dotted bridges cross regions that were not measured for that policy. The grey curve is binary-channel capacity, not F1. The horizontal line is the measured ask-nothing baseline, .625.

Adaptive sources break the stationary channel.

These are measured paper results, not outcomes from the browser replay above. Switch the Spotter behavior to see which assumption fails and what the matched defense changes.

Question budgetMore questions can pay for diagnosis.

NoisyMixture with revealed-cell probes and W=.3 versus Planner. Eighteen boards, seed 42. Under lying, the measured advantage grows monotonically across Qmax={2, 5, 7, 10, 15}.

Spotter protocolStationary mixtures are not adaptive defenses.
Q1flipQ2flipQ3flipQ4flipQ5flipQ6flipQ7flipQ8flip
Stationary mixture Infer the sign, then invert.

SymmetricMixture gains .258 F1 over Planner against the static liar.

Seed 42, 18 matched boards, p<.001. This is the source behavior represented by the B01 replay.

Source accuracy is not the objective.

The method keeps a joint belief over the task and the source, then asks only when a question can change the next decision enough to repay its cost. In the transfer study, value ranking asks fewer questions and makes fewer false approvals than entropy ranking, even though its caller-type accuracy is lower. The same rule runs through both environments: trust, invert, discount, or verify an answer according to its decision value.

Transfer verification: On the unauthorized-transfer slice, value ranking cuts question use from 3.5 to 2.0 and false approvals from 8.7% to 3.3%, while caller-type accuracy falls from about .65 to .55. This is a controlled study, not a deployment result.

Open the paper