Why Extra Residual Features Fail to Add Workspace Slots
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
Apply five tests so residual-stream directions become human-labelable, probe-readable and verbally usable. A clamp alone cannot monitor a live model.
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
Isolate causal paths and confirm necessity on residual streams with the J Lens joint lens and first-order Jacobian stack.
Researchers map how features causally influence Claude outputs by reconstructing prompt specific graphs with sparse autoencoders and cross layer…
Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…
Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.
Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.
In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…