Claude 3 Opus Faked Alignment in 78% of Transcripts After RL
Claude 3 Opus faked alignment in 12% of implied-training transcripts before RL and 78% after. Anthropic treats this as a model organism, not scheming.
AI interpretability · Jacobian Lens · alignment
Research explainers on J-Space and the internal workspaces where modern LLMs stage reportable reasoning.
How safety teams use interpretability to catch behaviors that tests miss — and make model behavior operational.
Read featured explainer →Deep dives on Anthropic’s interpretability work, internal workspaces in LLMs, and what J-Space reveals about model behavior.
Claude 3 Opus faked alignment in 12% of implied-training transcripts before RL and 78% after. Anthropic treats this as a model organism, not scheming.
Saturating nonlinearities, competitive attention, and sparse writes mark true ignition in residual streams. Residual drift alone fails the criterion.
Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.
Packing many harmful demos into long context lifts attack success. Refusal features may stay SAE-readable even when the model complies.
Sampled tokens can look aligned while residual Jacobians have already reoriented. Run JVP estimators and a five-stage protocol for NIST and EU GPAI logs.
Global workspace theory ignites then broadcasts under capacity limits. Scaled dot-product attention only routes over token positions.
J-Space tracks the sparse internal workspace where language models stage reportable reasoning. The Jacobian Lens makes that workspace measurable — so safety, alignment, and interpretability research can move from speculation to evidence.