{"slug":"research-coding-model-eval-harness","name":"Fable 5 vs Opus 4.8: A Coding-Agent Evaluation","rules":{"start":"open"},"nodes":{"open":{"short":"the note","pose":{"base":"centre"},"beat":{"layout":"center","marker":"evaluation · interactive","title":["Fable 5 vs Opus 4.8:","A Coding-Agent Evaluation"],"glass":1,"lede":"A head-to-head on real engineering work through a harness built to make the comparison mean something. Both models passed every gate, so pass/fail did not separate them; Fable led modestly on tool calls and tokens. Only 4 of 10 tasks ran before Fable was suspended.","caption":"9 beats · 8 min to read in full"}},"premise":{"short":"the premise","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"01 · evaluation","title":["A head-to-head between Claude","Fable 5 and Claude Opus 4.8"],"lede":"A head-to-head between Claude Fable 5 and Claude Opus 4.8 on real engineering work, run through a reproducible harness built to make the comparison actually mean something.","caption":"Condensed from the note."}},"why-run-both-and-eyeball-it-fail":{"short":"why \"run both","pose":{"base":"drained"},"beat":{"layout":"right","marker":"02 · why \"run both","title":["Why \"run both","and eyeball it\" fails"],"lede":"Handing two models the same task and picking the nicer output gives a confident answer that is usually wrong. Four problems sink that approach, and the harness is built to remove each one:","rows":[{"title":"Contamination","note":"if a model trained on the fix, you're testing memory, not skill."},{"title":"No objective \"done\"","note":"\"I prefer this output\" is a mood, not a measurement."},{"title":"Inconsistent help","note":"nudging a struggling model contaminates what you're measuring."},{"title":"Bias","note":"knowing which model wrote what tilts the scoring."}],"caption":"Condensed from the note."}},"the-core-idea":{"short":"the core idea","pose":{"base":"centre"},"beat":{"layout":"left","marker":"03 · the core idea","title":["The core","idea"],"lede":"The model gets a real codebase at a known starting point and a description of the symptom, never the fix.","caption":"Condensed from the note."}},"what-the-harness-measures":{"short":"what the harness","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"04 · what the harness","title":["What the harness","measures"],"lede":"Gates (the objective floor). Before each run, the harness confirms the task's test fails, proof the problem is real. After the run, it re-runs the gate and the full test suite to catch anything the fix broke elsewhere.","caption":"Condensed from the note."}},"what-we-tested-it-on":{"short":"what we tested it","pose":{"base":"drained"},"beat":{"layout":"left","marker":"05 · what we tested it","title":["What we","tested it"],"lede":"We ran the comparison on Click (pallets/click), a widely used open-source Python library for building command-line tools.","rows":[{"title":"a small, localized bug fix","note":""},{"title":"a multi-file feature addition","note":""},{"title":"a debug-from-a-traceback task (the model got only the","note":"failing test output, no description)"},{"title":"a deliberately open-ended design task with no single right answer","note":""}],"caption":"Condensed from the note."}},"what-we-found":{"short":"what we found","pose":{"base":"drained"},"beat":{"layout":"right","marker":"06 · what we found","title":["What we","found"],"lede":"The efficiency edge is real but not uniform. Fable's advantage is decisive on the multi-file feature and slight on the localized fix; the two are effectively even on the open-ended task; and on the debug task Opus actually did the leaner…","figure":{"kind":"bars","categories":["localized bug fix","multi-file feature","debug from traceback","open-ended design"],"series":[{"label":"Fable 5","tone":"cost","values":[5.6,23,11.8,33.4],"display":["5.6","23","11.8","33.4"]},{"label":"Opus 4.8","tone":"cost","values":[6.2,28,10.2,34.2],"display":["6.2","28","10.2","34.2"]}],"unit":"lower is more efficient, average of 5 runs","caption":"Tool calls per task."},"caption":"Condensed from the note, which carries 5 rows."}},"built-to-reuse":{"short":"built to reuse","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"07 · built to reuse","title":["Built","to reuse"],"lede":"The harness was designed so the experiment isn't a one-off. Swapping the two models or pointing it at a different codebase is a configuration change, not a rewrite, and adding a new task just means supplying a prompt and a test.","caption":"Condensed from the note."}},"close":{"short":"the end","pose":{"base":"finale"},"beat":{"layout":"center","title":["That is the note,","in one pass."],"glass":1,"lede":"This is Fable 5 vs Opus 4.8: A Coding-Agent Evaluation condensed to its beats. The full note carries every paragraph, the tables, the code, and the source it was adapted from.","links":[{"href":"/research/coding-model-eval-harness","label":"Read the full note"},{"href":"/research","label":"All research","tone":"ghost"}]}}}}