The thesis

A unit of workfor intelligence.

Experiments, frameworks, and field notes from the program behind quirq. Hypotheses ship with falsifiers; results land here as they land.

16notes

3topics

160minutes of reading

Measured runs, one configuration against another: what we changed, what it cost in tokens, and what each result does not license us to say.

A starburst of light rays radiating from a single point in a dark void.08experimentThe Incurious Agent84 controlled runs on environmental curiosity across two coding agents and seven progressively richer environments. Agents read the surface and skip the substance: the curated memory built for the task was opened in 1 of 12 runs.June 2026 · 12 min read
A dense starburst of dispersed rays thrown outward from one bright point in a void.09experimentThe Self-Sufficient AgentWe raised the difficulty until the bare agent should have broken: harder tasks on a real 170-file service. On these tasks it didn't break. With no documentation at all, two coding agents satisfied nine of nine non-obvious functional requirements, identically.June 2026 · 9 min read
Soft ribbons of refracted colour folding over one another against black.10experimentRelevance, Not VolumeTwo operating contracts for a coding agent, matched to the same ~14 KB. The generic one left conformance at its 8% floor; the one carrying a single project-specific rule lifted it to 100% for Codex and 80% for Claude.June 2026 · 8 min read
Fine rays radiating unevenly from a single point of light, thinning into darkness.11experimentCuriosity Comparison Between AgentsA third coding agent joins the context ladder. Gemini tops the curiosity index at 100 against Claude's 60 and Codex's 23, spends 38% of its actions reading files, and is the only one of the three that explores more as the workspace gets richer.June 2026 · 11 min read
A hard white edge splitting a beam into separating bands of colour.12evaluationFable 5 vs Opus 4.8: A Coding-Agent EvaluationA head-to-head on real engineering work through a harness built to make the comparison mean something. Both models passed every gate, so pass/fail did not separate them; Fable led modestly on tool calls and tokens. Only 4 of 10 tasks ran before Fable was suspended.June 2026 · 8 min read

Adapted from the XO research program · docs.xo.builders/research