Skip to content
No results
  • Explainers
J-Space logo
J-Space
  • Explainers
J-Space logo
J-Space
  • AI Safety

Score Hidden Objectives Before Compliant Tokens Appear

Score hidden objectives before a compliant token emits. J-Lens reads residual Jacobians as a complementary readout, not an alignment proof.

  • Editorial Team
  • August 18, 2026
  • Deep Dives

Confirm Layer Selection With Patching Before You Trust Attention Maps

Attention weights leave the real selector hidden. State a claim, then patch activations to show which residual hypotheses a layer amplifies or suppresses.

  • Editorial Team
  • August 17, 2026
  • J-Space

What the J Lens Protocol Isolates That a Logit Lens Cannot

Isolate causal paths and confirm necessity on residual streams with the J Lens joint lens and first-order Jacobian stack.

  • Editorial Team
  • August 15, 2026
  • Deep Dives

Tracing Causal Paths in Claude’s Addition and Safety Circuits

Researchers map how features causally influence Claude outputs by reconstructing prompt specific graphs with sparse autoencoders and cross layer…

  • Editorial Team
  • August 14, 2026
  • Deep Dives

Sparse Autoencoders Use L1 Penalties to Isolate Distinct Model Features

Sparse autoencoders use L1 penalties on residual activations to generate sparse mostly zero latents that correspond to candidate features in language…

  • Editorial Team
  • August 10, 2026
  • AI Safety

Safety Teams Decode AI Models Using Mechanistic Interpretability in 2026

Safety teams use mechanistic interpretability to reveal AI behaviors missed by tests, making safety operational in 2026.

  • Editorial Team
  • August 2, 2026
  • Deep Dives

Latent Correctness Directions Emerge from Residual Activations Without Probes

Latent correctness directions appear in residual activations of language models when using Jacobian lenses and sparse autoencoders instead of probes.

  • Editorial Team
  • August 1, 2026
  • J-Space

How J-space Detects Model Misalignment

In brief J-space detects model misalignment by examining a compact internal workspace in large language models that handles deliberate reasoning, using the…

  • Editorial Team
  • July 25, 2026
  • AI Safety

Jacobian Lens for AI Safety: Reading a Language Model’s Hidden Thoughts

In brief Mid-layer activations in language models are projected by the Jacobian lens through transport matrices derived from averaged input-output…

  • Editorial Team
  • July 25, 2026
  • Deep Dives

Anthropic Interpretability Research Explained: Inside the Push for Transparent AI

In brief Anthropic's interpretability research employs dictionary learning with sparse autoencoders to extract monosemantic features from the internal…

  • Editorial Team
  • July 24, 2026
1 2
Next
Copyright © 2026 - WordPress Theme by CreativeThemes