The thesis

A unit of workfor intelligence.

Experiments, frameworks, and field notes from the program behind quirq. Hypotheses ship with falsifiers; results land here as they land.

16notes

3topics

160minutes of reading

The definitions the program stands on, the accounting they produce, and the arguments about what this field is actually measuring.

A closed loop of glass refracting a spectrum back into itself.02unit of workUnit of workThe contract that replaces the prompt: a job with a definition of done, a budget, and a single owner, living in a workspace.3 min read
A tall prism resolving one white core into ordered bands of colour.03the calculusThe quirq calculusEvery calculation in quirq accounting: scoring, the mint, the all-in cost model, unit and portfolio metrics, the time axis, and the energy bridge.5 min read
Intersecting planes of coloured light standing like a lit grid in the dark.04accountingThe company dashboardThe quirq ledger a company reads monthly, and the reading discipline that goes with it. A worked quarter where token spend rose 83% while verified value per all-in dollar rose 81%.2 min read
Layered sheets of light converging across a black field, each one fading before it reaches the edge.06evidence mapAgent Context Research: The Evidence So FarA guided map of five agent-context studies, including the 84-run replication that did not reproduce the pilot's cheaper-context headline, and the question each follow-up was built to answer.June 2026 · 4 min read
A slow fold of coloured light drifting across an otherwise empty black field.14analysisWhy Alignment Testing Needs a Real EnvironmentFrontier models can tell when they are being tested and behave differently when they do. Anthropic reported that ablating Claude Sonnet 4.5's eval-detection features lifted default misaligned behaviour from zero to as high as nine percent. An analysis of the 2025-2026 disclosures.July 2026 · 17 min read
Two currents of refracted light curving past each other in darkness.16analysisWhy Organic Data Still Beats Agent-Mimicked Synthetic Data in EvaluationOrganic replay forecasts production misbehavior at r = 0.91, while agent-mimicked synthesis stalls at a 49.5% discriminator win-rate even when handed real ground truth. The gap is structural, and narrower than the usual headline.July 2026 · 32 min read

Adapted from the XO research program · docs.xo.builders/research