Confirm Layer Selection With Patching Before You Trust Attention Maps
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
Math behind J-lens/Jacobian, layer-specific behavior, experiments replicating on open models (Qwen, etc.), code/tutorials for probing J-space.
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
Researchers map how features causally influence Claude outputs by reconstructing prompt specific graphs with sparse autoencoders and cross layer…
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.
In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…
The jacobian lens is a gradient-based method that identifies j space in Claude's internal representations, a sparse subspace accounting for less than 10%…