{"slug":"research-validation","name":"Validation","rules":{"start":"open"},"nodes":{"open":{"short":"the note","pose":{"base":"centre"},"beat":{"layout":"center","marker":"validation · interactive","title":["Validation","on the record."],"glass":1,"lede":"Hypothesis-first validation: every empirical claim stated with its falsifier and bound to numbered experiments E1 through E7, completed or scheduled.","caption":"6 beats · 3 min to read in full"}},"premise":{"short":"the premise","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"01 · validation","title":["The validation program","is organized hypothesis-first"],"lede":"The validation program is organized hypothesis-first: each load-bearing claim is stated at full strength with its falsifier, and bound to numbered experiments, completed or scheduled.","caption":"Condensed from the note."}},"the-hypotheses":{"short":"the hypotheses","pose":{"base":"drained"},"beat":{"layout":"right","marker":"02 · the hypotheses","title":["The hypotheses","validation"],"lede":"H1: Mint integrity. Completion read from environment-captured state mints no fiction: false claims settle at rate zero, while any self-report channel settles them at approximately the false-claim rate.","caption":"Condensed from the note."}},"completed-experiments-e1-to-e3-m":{"short":"completed","pose":{"base":"recede"},"beat":{"layout":"left","marker":"03 · completed","title":["Completed experiments (E1","to E3, mock mode)"],"lede":"E1: verification source (H1). 200 identical units, an agent that falsely claims done with probability 0.05.","panelRows":[{"title":"Self-report","note":"False claims 7 · Silently settled 7 · Rate 3.5%"},{"title":"Environment snapshot","note":"False claims 7 · Silently settled 0 · Rate 0.0%"}],"caption":"Condensed from the note."}},"the-roadmap-e4-to-e7":{"short":"the roadmap (e4","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"04 · the roadmap (e4","title":["The roadmap (E4","to E7)"],"lede":"Results and per-run data are published here as they land.","rows":[{"title":"E4 (H1, H2)","note":"real-mode replication of E1 to E3 with a production coding agent and measured tokens. The false-claim rate becomes a measurement, and the tenure curve becomes evidence. In progress."},{"title":"E5 (H3)","note":"pilot ledgers: the dashboard instrumented on real work across at least three unit types."},{"title":"E6 (H3)","note":"predictive study: QER* trend vs token spend, task counts, and benchmark scores as predictors of renewal, expansion, and P&L attribution."},{"title":"E7 (H4)","note":"longitudinal budget-drift audit under the gaming mitigations."}],"caption":"Condensed from the note."}},"close":{"short":"the end","pose":{"base":"finale"},"beat":{"layout":"center","title":["That is the note,","in one pass."],"glass":1,"lede":"This is Validation condensed to its beats. The full note carries every paragraph, the tables, the code, and the source it was adapted from.","links":[{"href":"/research/validation","label":"Read the full note"},{"href":"/research","label":"All research","tone":"ghost"}]}}}}