Decoder-Only Transformers Route Tokens Instead of Igniting a Global Workspace
QKV attention routes tokens rather than igniting broadcast. Decoder-only transformers fail GWT; the 2017 paper dropped recurrence. Use indicator tests.
QKV attention routes tokens rather than igniting broadcast. Decoder-only transformers fail GWT; the 2017 paper dropped recurrence. Use indicator tests.
Decode residual-stream activations and 34 million SAE features; chat transcripts, chain of thought, and the public API do not expose them.
GPT-3 stacks 175 billion untied parameters while Universal Transformers reuse one block. Both broadcast. Neither creates C2 sentience.
Late residual writes sit immediately upstream of the unembedding and bias decoder-aligned effects toward the last blocks.
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
Apply five tests so residual-stream directions become human-labelable, probe-readable and verbally usable. A clamp alone cannot monitor a live model.
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
Isolate causal paths and confirm necessity on residual streams with the J Lens joint lens and first-order Jacobian stack.
Researchers map how features causally influence Claude outputs by reconstructing prompt specific graphs with sparse autoencoders and cross layer…