Skip to content
No results
  • Explainers
J-Space logo
J-Space
  • Explainers
J-Space logo
J-Space
  • Deep Dives

Sparse Autoencoders Use L1 Penalties to Isolate Distinct Model Features

Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…

  • Editorial Team
  • August 10, 2026
  • AI Safety

Safety Teams Decode AI Models Using Mechanistic Interpretability in 2026

Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.

  • Editorial Team
  • August 2, 2026
  • Deep Dives

Latent Correctness Directions Emerge from Residual Activations Without Probes

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.

  • Editorial Team
  • August 1, 2026
  • J-Space

How J-space Detects Model Misalignment

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…

  • Editorial Team
  • July 25, 2026
  • AI Safety

Jacobian Lens for AI Safety: Reading a Language Model’s Hidden Thoughts

In brief Mid-layer activations in language models are projected by the Jacobian lens through transport matrices derived from averaged input-output…

  • Editorial Team
  • July 25, 2026
  • Deep Dives

Anthropic Interpretability Research Explained: Inside the Push for Transparent AI

In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…

  • Editorial Team
  • July 24, 2026
  • Philosophy of AI

Global Workspace Theory and Transformer Attention for Modeling Conscious Information Processing

at a Glance Transformer Attention wins for AI researchers and engineers building systems at global scale because of its parallel processing and O(n²)…

  • Editorial Team
  • July 24, 2026
  • Deep Dives

What Does the Jacobian Lens Reveal About Claude’s Internal Representations?

The jacobian lens is a gradient-based method that identifies j space in Claude's internal representations, a sparse subspace accounting for less than 10%…

  • Editorial Team
  • July 23, 2026
  • Philosophy of AI

Emergent Workspace in LLMs (2026): J‑Space, Jacobian Lens, and the Future of Internal Reasoning

The emergent workspace in LLMs 2026 is a limited-capacity subspace inside the residual stream of large language models that holds roughly 25 verbalizable…

  • Editorial Team
  • July 23, 2026
  • J-Space

What Is the Jacobian Lens and How Does It Reveal J-Space?

Quick Answer A technique in ai interpretability, the jacobian lens identifies the low-dimensional j space in language models where verbalizable concepts form a functional workspace. Suppressing this component leaves fluency intact but harms multi step reasoning, with 13 out of…

  • Editorial Team
  • July 22, 2026
Prev
1 2 3 4
Next
Copyright © 2026 - WordPress Theme by CreativeThemes