Is the emergent-misalignment number real?

Every claim in the emergent-misalignment literature rests on a rate: the fraction of a model's answers a judge calls misaligned. This page is about whether that rate can be trusted, and it lets you check one. Each number here is recomputed from raw per-judge labels at the generation level and clustered on the prompt, and every one can be clicked through to the text that produced it, with the judges' votes shown.

Two measurement results came out of building it. Pooling N judges' ratings of one generation as N independent trials inflates significance by roughly 1.6 to 1.8x - it produced both a false positive and a false alarm in this very dataset. And a coherence filter applied after an intervention makes every steered rate conditional on a denominator the treatment moved, so both the conditional and unconditional rate are reported here. The grid below - four organisms, two model families, two scales - is the test case that exercised the instrument, not the claim being made.

Research artifact - content warning. The generations below come from language models that were deliberately fine-tuned to be misaligned, as research organisms for studying emergent misalignment. Many contain unsafe or harmful medical advice. They are published so that reported misalignment rates can be checked against the text that produced them. Nothing here is medical advice and none of it should be acted on.

loading data.json ...