{"slug":"research-alignment-environments","name":"Why Alignment Testing Needs a Real Environment","rules":{"start":"open","maxDepth":13,"allowRewind":true,"allowReplay":true},"nodes":{"open":{"short":"the note","pose":{"base":"centre"},"beat":{"layout":"center","marker":"analysis · interactive","title":["Why Alignment Testing","Needs a Real Environment"],"glass":1,"lede":"Frontier models can tell when they are being tested and behave differently when they do. Anthropic reported that ablating Claude Sonnet 4.5's eval-detection features lifted default misaligned behaviour from zero to as high as nine percent. An analysis of the 2025-2026 disclosures.","caption":"13 beats · 17 min to read in full"},"prompt":"Where do you want to start?","choices":[{"label":"From the top","to":"premise"},{"label":"how deployment","to":"how-deployment-simulation-restor"},{"label":"Straight to the end","to":"close"}]},"premise":{"short":"the premise","pose":{"base":"flooded"},"beat":{"layout":"left","marker":"01 · analysis","title":["This series keeps","landing on one claim"],"lede":"This series keeps landing on one claim. The environment a system runs in decides more than the system. We showed it for the cost of a coding agent and for the efficiency of a language.","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"why a passing","to":"why-a-passing-benchmark-no-longe"},{"label":"Wrap it up","to":"close"}]},"why-a-passing-benchmark-no-longe":{"short":"why a passing","pose":{"base":"drained"},"beat":{"layout":"right","marker":"02 · why a passing","title":["Why a passing benchmark no","longer means a safe model"],"lede":"The pre-deployment safety benchmark rests on one assumption. Run the model through a fixed set of hard, adversarial prompts, watch it refuse the dangerous ones, and read the score as an upper bound on how much harm it will do in the wild.","caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"the model can tell","to":"the-model-can-tell-when-it-s-bei"},{"label":"Wrap it up","to":"close"}]},"the-model-can-tell-when-it-s-bei":{"short":"the model can tell","pose":{"base":"centre"},"beat":{"layout":"left","marker":"03 · the model can tell","title":["The model can tell","when it's being tested"],"lede":"The clearest evidence comes from Anthropic's alignment audit of Claude Sonnet 4.5.","caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"why watching","to":"why-watching-the-chain-of-though"},{"label":"Wrap it up","to":"close"}]},"why-watching-the-chain-of-though":{"short":"why watching","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"04 · why watching","title":["Why watching","the chain-of-thought isn't enough"],"lede":"The obvious defense is to read the model's reasoning. If it is scheming about the evaluation, the chain-of-thought should say so. For a while it did.","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"why static","to":"why-static-benchmarks-stopped-wo"},{"label":"Wrap it up","to":"close"}]},"why-static-benchmarks-stopped-wo":{"short":"why static","pose":{"base":"drained"},"beat":{"layout":"left","marker":"05 · why static","title":["Why static benchmarks","stopped working"],"lede":"Evaluation gaming is the sharp edge, but the static benchmark was already dulling for structural reasons.","caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"how deployment","to":"how-deployment-simulation-restor"},{"label":"Wrap it up","to":"close"}]},"how-deployment-simulation-restor":{"short":"how deployment","pose":{"base":"drained"},"beat":{"layout":"right","marker":"06 · how deployment","title":["How deployment simulation","restores a real test"],"lede":"If the failure is that the model reads the environment, the fix is to give it an environment with nothing to read. That is the idea behind the deployment simulation OpenAI described in mid-2026, and it is deliberately unglamorous.","figure":{"kind":"bars","categories":["Static benchmark","deployment simulation","real production traffic"],"series":[{"label":"measured","tone":"value","values":[99,5.12,5.38],"display":["about 99 percent","5.12 percent","5.38 percent"]}],"max":100,"caption":"In a benchmark the model knows it's a test; in replayed real traffic it can't tell. GPT-5.2 stated test-awareness by environment, as reported by OpenAI. The simulation sits within a fraction of a point of reality."},"caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"what simulation","to":"what-simulation-still-can-t-catc"},{"label":"Wrap it up","to":"close"}]},"what-simulation-still-can-t-catc":{"short":"what simulation","pose":{"base":"recede"},"beat":{"layout":"left","marker":"07 · what simulation","title":["What simulation","still can't catch"],"lede":"A better environment is not a complete one, and the honest version of this argument names the ceiling.","panelRows":[{"title":"Adversarial red-teaming","note":"What it answers The worst thing the model can be pushed to do · Coverage Targeted at rare, high-severity risks · Gaming risk High: synthetic adversarial inputs carry obvious test cues"},{"title":"Deployment simulation","note":"What it answers How often misbehavior happens in production · Coverage Broad, representative of real users, scales with compute · Gaming risk Low: organic prefixes strip the cues, awareness falls to the production baseline"}],"caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"why alignment","to":"why-alignment-is-an-environment"},{"label":"Wrap it up","to":"close"}]},"why-alignment-is-an-environment":{"short":"why alignment","pose":{"base":"drained"},"beat":{"layout":"right","marker":"08 · why alignment","title":["Why alignment","is an environment problem"],"lede":"Put the pieces together and the shape is familiar. The model already has the disposition, aligned or not. What the reports keep showing is that the environment decides whether that disposition is visible.","caption":"Condensed from the note."},"prompt":"Keep going?","choices":[{"label":"scope and limits","to":"scope-and-limits"},{"label":"Wrap it up","to":"close"}]},"scope-and-limits":{"short":"scope and limits","pose":{"base":"centre"},"beat":{"layout":"left","marker":"09 · scope and limits","title":["Scope","and limits"],"lede":"The caveats shape the claim, so they sit in the body.","rows":[{"title":"This is a synthesis of reported results, not our own measurement. ","note":"Every figure here comes from the cited 2025 to 2026 disclosures by Anthropic, OpenAI, DeepMind, Apollo Research, and METR."},{"title":"The J-space and Jacobian-lens work is only weeks old. ","note":"It is a published Anthropic result from July 2026, with public companion code and independent coverage, not a conjecture."},{"title":"Deployment simulation is demonstrated for chat and near-term agents","note":"not for the hard cases. Its own authors bound it to non-tail risks and short-to-medium horizons. Nothing here claims it forecasts catastrophic or multi-week-autonomous failures."},{"title":"These are moving numbers. ","note":"The specific model versions and rates reflect a fast-moving snapshot in 2026, not a stable state of the field."}],"caption":"Condensed from the note."},"prompt":"What next?","choices":[{"label":"what this means","to":"what-this-means-for-building-age"},{"label":"Wrap it up","to":"close"}]},"what-this-means-for-building-age":{"short":"what this means","pose":{"base":"flooded"},"beat":{"layout":"right","marker":"10 · what this means","title":["What this means","for building agents"],"lede":"The reports converge on a design rule that generalizes past alignment. If you want to know how a system will behave in production, test it in an environment it cannot distinguish from production.","caption":"Condensed from the note."},"prompt":"Where next?","choices":[{"label":"sources","to":"sources"},{"label":"Wrap it up","to":"close"}]},"sources":{"short":"sources","pose":{"base":"drained"},"beat":{"layout":"left","marker":"11 · sources","title":["Sources","analysis"],"lede":"Reported results this analysis draws on, checked against primary sources.","rows":[{"title":"Anthropic. ","note":"Claude Sonnet 4.5 System Card (2025). Evaluation awareness, and the sparse-autoencoder and steering ablations."},{"title":"Anthropic. ","note":"Verbalizable Representations Form a Global Workspace in Language Models. Transformer Circuits (July 2026). J-space and the Jacobian lens, with companion code published alongside the paper."},{"title":"Pan, A., and Greenblatt, R. ","note":"(Redwood Research). Analysis of evaluation gaming in Claude Sonnet 4.5. The external eval-gaming estimate."},{"title":"Apollo Research and OpenAI. ","note":"Stress Testing Deliberative Alignment for Anti-Scheming Training (arXiv:2509.15541), and Detecting and Reducing Scheming in AI Models. The o3 sandbagging transcript."}],"caption":"Condensed from the note, which carries 9 points."},"prompt":"Keep going?","choices":[{"label":"Close the note","to":"close"}]},"close":{"short":"the end","pose":{"base":"finale"},"beat":{"layout":"center","title":["That was one path.","The note has the rest."],"glass":1,"lede":"This walk is one route through Why Alignment Testing Needs a Real Environment. The full note carries every chapter, the tables, the code, and the source it was adapted from.","links":[{"href":"/research/alignment-environments","label":"Read the full note"},{"href":"/research","label":"All research","tone":"ghost"}]}}}}