Sparse Autoencoders Use L1 Penalties to Isolate Distinct Model Features
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
AI interpretability · Jacobian Lens · alignment
Research explainers on J-Space and the internal workspaces where modern LLMs stage reportable reasoning.
How safety teams use interpretability to catch behaviors that tests miss — and make model behavior operational.
Read featured explainer →Deep dives on Anthropic’s interpretability work, internal workspaces in LLMs, and what J-Space reveals about model behavior.
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…

Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…

In brief Mid-layer activations in language models are projected by the Jacobian lens through transport matrices derived from averaged input-output…

In brief Anthropic’s interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…
J-Space tracks the sparse internal workspace where language models stage reportable reasoning. The Jacobian Lens makes that workspace measurable — so safety, alignment, and interpretability research can move from speculation to evidence.