Knowing Both Facts Still Fails the Composed Two-Hop Query
Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.
Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.
Packing many harmful demos into long context lifts attack success. Refusal features may stay SAE-readable even when the model complies.
Sampled tokens can look aligned while residual Jacobians have already reoriented. Run JVP estimators and a five-stage protocol for NIST and EU GPAI logs.
Global workspace theory ignites then broadcasts under capacity limits. Scaled dot-product attention only routes over token positions.
Residual streams analogize capacity-limited broadcast better than attention maps. Softmax routing stays graded and content-addressable, not ignited.
GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.
QKV attention routes tokens rather than igniting broadcast. Decoder-only transformers fail GWT; the 2017 paper dropped recurrence. Use indicator tests.
Decode residual-stream activations and 34 million SAE features; chat transcripts, chain of thought, and the public API do not expose them.
GPT-3 stacks 175 billion untied parameters while Universal Transformers reuse one block. Both broadcast. Neither creates C2 sentience.
Late residual writes sit immediately upstream of the unembedding and bias decoder-aligned effects toward the last blocks.