GPT-4 Wrote Explanations for 307,200 GPT-2 XL Neurons and Graded Them by Simulation
GPT-4 writes natural-language explanations of 307,200 GPT-2 XL MLP neurons across 48 layers, then scores them by simulating activations on held-out text.
Math behind J-lens/Jacobian, layer-specific behavior, experiments replicating on open models (Qwen, etc.), code/tutorials for probing J-space.
GPT-4 writes natural-language explanations of 307,200 GPT-2 XL MLP neurons across 48 layers, then scores them by simulating activations on held-out text.
Saturating nonlinearities, competitive attention, and sparse writes mark true ignition in residual streams. Residual drift alone fails the criterion.
Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.
Packing many harmful demos into long context lifts attack success. Refusal features may stay SAE-readable even when the model complies.
Residual streams analogize capacity-limited broadcast better than attention maps. Softmax routing stays graded and content-addressable, not ignited.
GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.
Decode residual-stream activations and 34 million SAE features; chat transcripts, chain of thought, and the public API do not expose them.
Late residual writes sit immediately upstream of the unembedding and bias decoder-aligned effects toward the last blocks.
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
Apply five tests so residual-stream directions become human-labelable, probe-readable and verbally usable. A clamp alone cannot monitor a live model.