{"slug":"research-organic-vs-synthetic-evaluation-data","name":"Why Organic Data Still Beats Agent-Mimicked Synthetic Data in Evaluation","rules":{"start":"open","maxDepth":13,"allowRewind":true,"allowReplay":true},"nodes":{"open":{"short":"the note","pose":{"base":"centre"},"beat":{"layout":"center","marker":"analysis · interactive","title":["Why Organic Data Still Beats","Agent-Mimicked Synthetic Data"],"glass":1,"lede":"Organic replay forecasts production misbehavior at r = 0.91, while agent-mimicked synthesis stalls at a 49.5% discriminator win-rate even when handed real ground truth. The gap is structural, and narrower than the usual headline.","caption":"13 beats · 32 min to read in full"},"prompt":"Where do you want to start?","choices":[{"label":"From the top","to":"premise"},{"label":"why organic data","to":"why-organic-data-matters"},{"label":"Straight to the end","to":"close"}]},"premise":{"short":"the premise","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"01 · analysis","title":["There is a seductive argument","in AI-safety measurement"],"lede":"There is a seductive argument in AI-safety measurement that goes like this. If a sufficiently capable language model can write essays, code, and dialogue that humans cannot reliably distinguish from human output, then a sufficiently…","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"framing and thesis","to":"framing-and-thesis"},{"label":"Wrap it up","to":"close"}]},"framing-and-thesis":{"short":"framing and thesis","pose":{"base":"drained"},"beat":{"layout":"right","marker":"02 · framing and thesis","title":["Framing","and thesis"],"lede":"The thesis of this article is that the seductive argument fails in a specific and non-obvious way, but only within specific bounds, which are worth stating precisely up front.","rows":[{"title":"Model-generated synthetic","note":"an LLM or agent writes the scenarios, users, and tool responses. This splits further into distributional synthesis (mimicking traffic) and targeted or adversarial synthesis (deliberately…"},{"title":"Human-authored or expert-curated","note":"hand-written red-team tasks and benchmarks such as HarmBench, MLCommons AILuminate, METR's engineering tasks, and vendor-annotated sets from Scale, Surge, and similar."},{"title":"Production-organic","note":"de-identified real deployment logs replayed or analyzed in aggregate."}],"caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"the distribution","to":"the-distribution-problem-synthet"},{"label":"Wrap it up","to":"close"}]},"the-distribution-problem-synthet":{"short":"the distribution","pose":{"base":"centre"},"beat":{"layout":"left","marker":"03 · the distribution","title":["The distribution problem:","synthetic diverges from real"],"lede":"An evaluation set earns its keep in the tail: rare inputs, unusual dialects, adversarial edge cases, low-frequency failure modes. Distributional synthesis is biased against exactly that region.","caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"the detectability","to":"the-detectability-problem-evalua"},{"label":"Wrap it up","to":"close"}]},"the-detectability-problem-evalua":{"short":"the detectability","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"04 · the detectability","title":["The detectability problem:","evaluation awareness inflates"],"lede":"The distributional argument would matter less if models responded to synthetic scenarios exactly as they respond to real ones. They do not.","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"the fidelity","to":"the-fidelity-ceiling-why-even-go"},{"label":"Wrap it up","to":"close"}]},"the-fidelity-ceiling-why-even-go":{"short":"the fidelity","pose":{"base":"drained"},"beat":{"layout":"left","marker":"05 · the fidelity","title":["The fidelity ceiling:","why even good mimicry is bounded"],"lede":"Suppose we grant the generator every advantage (real system prompts, real codebases, real tool databases) and push mimicry to its empirical limit. How close does it get? The 2026 discriminator literature gives an unusually clean answer.","rows":[{"title":"Original trajectory generator","note":"GPT-5.4, which produced roughly 120,000 internal Codex trajectories."},{"title":"Candidate being resampled","note":"GPT-5.5, whose response replaces the original final turn."},{"title":"Tool-response simulator","note":"GPT-5.4, which fabricates the tool outputs."},{"title":"Discriminator, or differentiator","note":"a GPT-5.4-based judge."}],"caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"why organic data","to":"why-organic-data-matters"},{"label":"Wrap it up","to":"close"}]},"why-organic-data-matters":{"short":"why organic data","pose":{"base":"centre"},"beat":{"layout":"right","marker":"06 · why organic data","title":["Why organic","data matters"],"lede":"If unanchored synthetic data is tail-poor, detectably contrived, and fidelity-capped, the affirmative case for organic data follows, now backed by direct calibration evidence rather than principle alone.","caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"current players","to":"current-players-and-their-limita"},{"label":"Wrap it up","to":"close"}]},"current-players-and-their-limita":{"short":"current players","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"07 · current players","title":["Current players","and their limitations"],"lede":"The landscape spans model-synthetic auditing tools, organic-leaning replay and analysis, a human-data and community-benchmark layer, and neutral substrate.","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"the honest","to":"the-honest-counterweight"},{"label":"Wrap it up","to":"close"}]},"the-honest-counterweight":{"short":"the honest","pose":{"base":"drained"},"beat":{"layout":"right","marker":"08 · the honest","title":["The honest","counterweight"],"lede":"A one-sided case would be a weaker case. The pro-synthetic argument rests on real methodological and legal advantages, and there are settings where synthetic data legitimately wins, including one the thesis must concede outright.","caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"conclusion","to":"conclusion"},{"label":"Wrap it up","to":"close"}]},"conclusion":{"short":"conclusion","pose":{"base":"centre"},"beat":{"layout":"left","marker":"09 · conclusion","title":["Conclusion","analysis"],"lede":"The claim that unanchored agent-mimicked synthetic data cannot match organic data for evaluation and alignment testing is easy to misread as a temporary engineering complaint.","caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"a note","to":"a-note-on-epistemic-status"},{"label":"Wrap it up","to":"close"}]},"a-note-on-epistemic-status":{"short":"a note","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"10 · a note","title":["A note","on epistemic status"],"lede":"Several load-bearing 2026 sources, including OpenAI deployment simulation (arXiv:2607.07184), EnvSimBench (arXiv:2605.07247), the Anthropic coding-audit-realism and Bloom write-ups, Urania, and the OpenAI production-evaluations page,…","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"references","to":"references"},{"label":"Wrap it up","to":"close"}]},"references":{"short":"references","pose":{"base":"drained"},"beat":{"layout":"left","marker":"11 · references","title":["References","analysis"],"lede":"The literature this analysis rests on. Every empirical claim above is another group's finding unless it is explicitly ours.","rows":[{"title":"Shumailov, Shumaylov, Zhao, Papernot, Anderson, Gal. ","note":"AI models collapse when trained on recursively generated data. Nature 631, 755-759 (2024); preprint The Curse of Recursion (arXiv:2305.17493), with a 2025 author correction."},{"title":"Gerstgrasser et al. ","note":"Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data (arXiv:2404.01413)."},{"title":"LLM-synthetic versus real data divergence and tail","note":"truncation (arXiv:2502.08661)."},{"title":"Detection-based filtering to slow model collapse (arXiv:2502.15654).","note":""}],"caption":"Condensed from the note, which carries 30 points."},"prompt":"Keep going?","choices":[{"label":"Close the note","to":"close"}]},"close":{"short":"the end","pose":{"base":"finale"},"beat":{"layout":"center","title":["That was one path.","The note has the rest."],"glass":1,"lede":"This walk is one route through Why Organic Data Still Beats Agent-Mimicked Synthetic Data in Evaluation. The full note carries every chapter, the tables, the code, and the source it was adapted from.","links":[{"href":"/research/organic-vs-synthetic-evaluation-data","label":"Read the full note"},{"href":"/research","label":"All research","tone":"ghost"}]}}}}