Score Hidden Objectives Before Compliant Tokens Appear
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Using J-space for monitoring misalignment, detecting off-rail behavior, sandbagging, jailbreaks, scalable oversight, implications for future safety research.
Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.
Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.
In brief Mid-layer activations in language models are projected by the Jacobian lens through transport matrices derived from averaged input-output…
On July 6, 2026, Anthropic published a paper that changed how we think about what language models know but never say aloud. The jacobian lens—named after the Jacobian matrix in calculus—gives researchers a tool to read the silent, internal neural…