Sparse Autoencoders Use L1 Penalties to Isolate Distinct Model Features
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
Math behind J-lens/Jacobian, layer-specific behavior, experiments replicating on open models (Qwen, etc.), code/tutorials for probing J-space.
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.
In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…
The jacobian lens is a gradient-based method that identifies j space in Claude's internal representations, a sparse subspace accounting for less than 10%…