Safety Teams Decode AI Models Using Mechanistic Interpretability in 2026

Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.

Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…

In brief Mid-layer activations in language models are projected by the Jacobian lens through transport matrices derived from averaged input-output…

In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…

at a Glance Transformer Attention wins for AI researchers and engineers building systems at global scale because of its parallel processing and O(n²)…

The jacobian lens is a gradient-based method that identifies j space in Claude's internal representations, a sparse subspace accounting for less than 10%…

The emergent workspace in LLMs 2026 is a limited-capacity subspace inside the residual stream of large language models that holds roughly 25 verbalizable…

Quick Answer A technique in ai interpretability, the jacobian lens identifies the low-dimensional j space in language models where verbalizable concepts form a functional workspace. Suppressing this component leaves fluency intact but harms multi step reasoning, with 13 out of…

Quick Answer Analyses of language models identify a sparse subspace called J-space that functions like the human global workspace for reportable thought. In Claude Sonnet 4.5 this workspace holds approximately 25 per arXiv preprint active vectors and less than 10%…