Latent Correctness Directions Emerge from Residual Activations Without Probes

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.
Math behind J-lens/Jacobian, layer-specific behavior, experiments replicating on open models (Qwen, etc.), code/tutorials for probing J-space.

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.

In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…

The jacobian lens is a gradient-based method that identifies j space in Claude's internal representations, a sparse subspace accounting for less than 10%…