Writings

Changelog → Docs (opens in a new tab)
The paper

The unit-of-work whitepaper, in short: the mint, the calculus, the ledger, and the environment that owns the ground truth. Readable in full, with every claim tiered against its evidence.

Jul 15, 202627 min read

News

Launches · partnerships · integrations

Thoughts

The findings, argued · every note links to its study
Thoughts01

The Orientation Tax

What a coding agent's first day on the job actually costs, and why the environment is the only lever on it. On the environment-performance study.
Jul 26, 20265 min read
Thoughts02

Agent-Mimicked Synthetic Data Will Never Beat the Real Thing

It never is. On why organic data wins: r = 0.91 vs a coin flip, and the discriminator that can't be fooled.
Jul 25, 20264 min read
Thoughts03

Your AI Knows When It's Being Tested

Perfect metrics, going nowhere, screen playing a fake trail. On why alignment testing needs a real environment.
Jul 25, 20264 min read
Thoughts04

RTFM: Read the F*cking Manual

We finally found an agent that does, and the reading paid its way. On the curiosity comparison.
Jul 3, 20264 min read
Thoughts05

Agents Look a Lot Like Quiet Quitters

Salary paid in full. Discretionary effort: zero. On the incurious agent.
Jun 21, 20264 min read
Thoughts06

The Agent She Told You Not to Worry About

No docs, no contract, no memory: nine for nine. On the self-sufficient agent.
Jun 26, 20264 min read
Thoughts07

Models Mathematically Find It More Difficult to Navigate Various Languages

The deal closes in fingers, a representation both sides already share. On the tokenizer study.
Jul 11, 20264 min read
Thoughts08

Point and Call

One precise gesture beats a thousand words of context. On relevance versus volume.
Jun 26, 20264 min read
Thoughts09

The Advantage Was Never the Model

Harvey's edge is the room it puts a model in, not a better model. Routing alone cuts inference cost three to five times. On the Harvey harness.
Jul 18, 20264 min read
Thoughts10

The Benchmark That Did Not Finish

Four of ten tasks ran before one of the models was suspended. The verdict never arrived; the method survived. On the coding-agent evaluation.
Jun 14, 20264 min read

The Future of Work

Thoughts · The thesis, with its evidence

Guides

Education · the unit of work, explained · read in order