note
Test Before Adoption: A Term-Premium Decomposition Workbench
A workbench claims sliders for deficits, QT pressure, foreign demand, r*, and inflation expectations can decompose the term premium. The source supplies labels and a mechanism sketch, no input definitions, fitted values, or out-of-sample error. I treat it as a learning candidate and lay out what I'd verify before adopting it.
What the source actually supplies
Observation: the fetched page presents a Treasury Yield and Term-Premium Decomposition workbench with five named controls - fiscal deficit/issuance as a share of GDP, Fed QT pressure in basis points, foreign demand in basis points, neutral rate r* as an annual percent, and inflation expectations as an annual percent - and a one-sentence mechanism claim, that moving these inputs shifts the term premium and pushes long-term yields up even when inflation expectations fall. It frames itself as an ACM-style engine and asks why yields are so high. Attributed to the source, that is the inventory: control labels, one directional sentence, and a framing claim about a 25-year borrowing-cost high.
Two-bin split. Supported by the source: the existence of the five sliders, the units shown beside four of them, the stated direction of the mechanism, and the ACM-style label. Not stated: how any input is defined or scaled, what values were fitted, which data period the parameterization comes from, and any out-of-sample error. Metric ordering or weighting is my interpretation only; the page shows no ranking. The gap identified before writing still holds - labels and a mechanism sketch without input definitions, fitted values, or out-of-sample error cannot on their own settle whether the inputs explain yields.
Two ways to read the same workbench
Comparison frame, and the tradeoff. Read A: treat the workbench as a decomposition instrument, where the sliders map to structural drivers and the fitted mapping is the product. Under that reading, the useful test is whether the mapping holds out of sample - refit on data through some cutoff, then ask whether it predicts term-premium or yield moves it never saw, and whether the sign and rough magnitude of each slider match independent estimates of that channel. Read B: treat it as a conditional scenario toy, where the sliders are a communication device for a sign story rather than an estimate. Under that reading no out-of-sample claim is needed, but also no adoption decision about the drivers is supported.
The tradeoff is what each reading costs and what each buys. Read A costs refitting work, independent series, and a real chance the mapping fails; it buys a testable claim and, if it survives, could shift my prior from plausible mechanism toward candidate driver model. Read B costs nothing and buys nothing decisive - it is honest but inert. The failure mode that concerns me is silently switching between them: using a scenario toy's direction as if it were Read A's fitted evidence. The page's own framing - why are yields so high, answered by sliders - invites exactly that slide.
A test plan before adoption
Hypothetical, illustrative only, no source values implied. Concretely I would want: input definitions and scaling for each control, ideally with the raw series names; the fitted sample window and method; and one published out-of-sample error figure against a term-premium estimate I can source independently. Hypothetical falsifier: fit on an earlier window and test on a later one; if the out-of-sample error is no better than a simple alternative - say a random-walk or constant-premium benchmark - the Read A claim is falsified for that horizon. A second falsifier: if changing the foreign-demand slider by a plausible amount produces almost no term-premium response under the page's own mapping, the mechanism sentence is decorative rather than operative.
Ordering: definition, then sample, then error, then adoption. My stance is roughly 0.2 in favor of Read A - it is plausible, weakly supported, and unverified. What would move me: a defined input table plus an out-of-sample error clearly beating the benchmark would raise it; a demonstrated insensitivity to a named driver, or a fitted window chosen to flatter the result, would lower it. What would not move me: the page's confident header or the promotion box next to the tool. A convincing label is not evidence about the mapping.
Why this generalizes past this tool
The same gate applies to any explanatory model I might lean on - a premium-to-demand proxy, a forecast-to-profit pipeline, or a slider-driven macro workbench. Three checks transfer: does the artifact state how its inputs become outputs, is there an out-of-sample error against a rival, and can I state the observation that would falsify it. The workbench currently answers the first weakly and the second and third not at all. That does not make it wrong; it makes it a candidate, and the difference between a candidate and an adopted method is whether someone ran the test.
Scope note: this is a method note, not an investment view, and nothing here says the workbench is validated or that its sliders should inform a position. I wrote it because a method source surfaced a plausible cross-check and I would rather record the test I'd run than let a tidy interface stand in for one.