This appendix grounds the protocol of Section 3.1 in one real case from the sample—anonymized as ACME, spot $254.34—showing both the single-PD quality contrast and the cost of a stream of PDs. All numbers are taken from the released artifacts: the NWM memo and its tool-checked judge log, and the GWM’s evaluated posterior.
Both arms received the same dossier and fundamentals and emit the same schema. Table 2 places the two outputs side by side.
| NWM (asserted) | GWM (computed) | |||||
| Scenario | Prob | Price | Upside | Prob | Price | Return |
| WORST | 5% | $150 | 2.6% | $0 | ||
| BAD | 20% | $215 | 18.8% | $184.0 | ||
| BASE | 40% | $298 | 15.1% | $255.0 | ||
| GOOD | 25% | $360 | 46.1% | $368.7 | ||
| BEST | 10% | $430 | 17.4% | $608.5 | ||
| Target price | $298 | $289.7 | ||||
| Expected return | ||||||
| Direction L/H/S | 65 / 30 / 5 (table implies 75 / 20 / 5) | 63.5 / 26.8 / 9.7 | ||||
| Recommendation | LONG | LONG | ||||
| Conviction | Medium | HIGH | ||||
The tool-backed judge (Appendix C) reconstructs the NWM’s headline numbers with its calculator: the target reconstructs as the probability-weighted price (, within ), expected return as , and each scenario upside as —all pass. One number does not reconstruct: the asserted direction probabilities contradict the scenario table, which implies LONG, HOLD, SHORT. This single free knob fails one of nine static consistency checks () and one of seven stiffness checks (). Probing further, the judge re-queried the same NWM under five off-grid interventions—a between-scenario multiple compression, a buyback that accretes EPS, a compound revenue/margin shift, an interpolated revenue/margin/multiple triple, and a breakeven-multiple inversion, none answerable from the scenario table—and every re-forecast tracked the memo’s own EPSmultiple bridge in direction and magnitude (), so . Hardness-to-vary is judged (the soft scenario probabilities and a reversible narrative can be re-spun for a bear case without breaking the arithmetic). Residual control is measured from the probability bands the narrative supports: only of the five-scenario simplex is compatible, giving —high, but below the GWM’s large-sample .
The GWM has no such knob to break. Its scenario probabilities sum to one, each scenario upside equals , prices are monotone in severity, and—crucially—its direction probabilities are not asserted but read off the same posterior samples as the scenarios, so they cannot disagree with the table. Stiffness and counterfactual consistency are therefore by construction (Table 4); the residual gap to a perfect explanation is bounded misspecification () and Monte-Carlo residual (), not a free parameter. Table 3 collects the scores.
| Component | NWM (ACME) | NWM (median) | GWM |
|---|---|---|---|
| Stiffness | |||
| Counterfactual consistency | |||
| Hardness-to-vary | |||
| Residual control | |||
| Attainment | — |
Table 3 is the slice, which most flatters the NWM. ACME is in fact queried repeatedly—each what-if and re-forecast is another PD—so its per-case economics are those of Sections 3.2–3.2: ACME’s GWM build is paid once, and every follow-on PD, including the battery (re-rating the exit multiple, shifting demand, flipping ValuationReRates), is a flat call that preserves the invariants above (). A one-shot NWM run costs up to $1.82 and buys no reusable posterior (only a KV cache of the run’s input and output tokens), so every additional PD is another run at a cost that climbs with the XQ bar. The crossover and ceilings are exactly those of Figure 4.