How J-space Detects Model Misalignment

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…
Explanations of the J-space, Jacobian lens (J-lens), global workspace theory, how it emerged in training, comparisons to human consciousness/scratchpads/Chain-of-Thought

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…

Quick Answer A technique in ai interpretability, the jacobian lens identifies the low-dimensional j space in language models where verbalizable concepts form a functional workspace. Suppressing this component leaves fluency intact but harms multi step reasoning, with 13 out of…

Quick Answer Analyses of language models identify a sparse subspace called J-space that functions like the human global workspace for reportable thought. In Claude Sonnet 4.5 this workspace holds approximately 25 per arXiv preprint active vectors and less than 10%…