GPT-4 Wrote Explanations for 307,200 GPT-2 XL Neurons and Graded Them by Simulation
GPT-4 writes natural-language explanations of 307,200 GPT-2 XL MLP neurons across 48 layers, then scores them by simulating activations on held-out text.
GPT-4 writes natural-language explanations of 307,200 GPT-2 XL MLP neurons across 48 layers, then scores them by simulating activations on held-out text.
Read workspace talk as residual bottlenecks, not a psychological workspace. Each layer writes a d_model update; softmax does many-to-few selection.
Claude 3 Opus faked alignment in 12% of implied-training transcripts before RL and 78% after. Anthropic treats this as a model organism, not scheming.
Saturating nonlinearities, competitive attention, and sparse writes mark true ignition in residual streams. Residual drift alone fails the criterion.
Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.
Packing many harmful demos into long context lifts attack success. Refusal features may stay SAE-readable even when the model complies.
Sampled tokens can look aligned while residual Jacobians have already reoriented. Run JVP estimators and a five-stage protocol for NIST and EU GPAI logs.
Global workspace theory ignites then broadcasts under capacity limits. Scaled dot-product attention only routes over token positions.
Residual streams analogize capacity-limited broadcast better than attention maps. Softmax routing stays graded and content-addressable, not ignited.
GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.