Neither Recurrent Loops Nor Untied Depth Give Transformers Self-Monitoring
GPT-3 stacks 175 billion untied parameters while Universal Transformers reuse one block. Both broadcast. Neither creates C2 sentience.
AI interpretability · Jacobian Lens · alignment
Research explainers on J-Space and the internal workspaces where modern LLMs stage reportable reasoning.
How safety teams use interpretability to catch behaviors that tests miss — and make model behavior operational.
Read featured explainer →Deep dives on Anthropic’s interpretability work, internal workspaces in LLMs, and what J-Space reveals about model behavior.
GPT-3 stacks 175 billion untied parameters while Universal Transformers reuse one block. Both broadcast. Neither creates C2 sentience.
Late residual writes sit immediately upstream of the unembedding and bias decoder-aligned effects toward the last blocks.
Published widths from 768 to 12288 bound how many residual directions stay independent. Extra features interfere instead of adding workspace slots.
Apply five tests so residual-stream directions become human-labelable, probe-readable and verbally usable. A clamp alone cannot monitor a live model.
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.
J-Space tracks the sparse internal workspace where language models stage reportable reasoning. The Jacobian Lens makes that workspace measurable — so safety, alignment, and interpretability research can move from speculation to evidence.