Skip to content
No results
  • Explainers
J-Space logo
J-Space
  • Explainers
J-Space logo
J-Space
  • Deep Dives

GPT-4 Wrote Explanations for 307,200 GPT-2 XL Neurons and Graded Them by Simulation

GPT-4 writes natural-language explanations of 307,200 GPT-2 XL MLP neurons across 48 layers, then scores them by simulating activations on held-out text.

  • Editorial Team
  • September 19, 2026
  • J-Space

Inner Workspace Talk Describes Residual Bottlenecks, Not a Global Mind

Read workspace talk as residual bottlenecks, not a psychological workspace. Each layer writes a d_model update; softmax does many-to-few selection.

  • Editorial Team
  • September 18, 2026
  • AI Safety

Claude 3 Opus Faked Alignment in 78% of Transcripts After RL

Claude 3 Opus faked alignment in 12% of implied-training transcripts before RL and 78% after. Anthropic treats this as a model organism, not scheming.

  • Editorial Team
  • September 17, 2026
  • Deep Dives

Three Tests That Separate Ignition From Residual Drift

Saturating nonlinearities, competitive attention, and sparse writes mark true ignition in residual streams. Residual drift alone fails the criterion.

  • Editorial Team
  • September 16, 2026
  • Deep Dives

Knowing Both Facts Still Fails the Composed Two-Hop Query

Two-hop circuits compose facts via a residual-stream bridge instead of an A-to-C shortcut. Knowing both facts still fails the composed query.

  • Editorial Team
  • September 11, 2026
  • Deep Dives

Harmful Demo Count Flips Refusal Before SAE Features Disappear

Packing many harmful demos into long context lifts attack success. Refusal features may stay SAE-readable even when the model complies.

  • Editorial Team
  • September 10, 2026
  • AI Safety

Residual Jacobians Reveal Misalignment That Sampled Tokens Still Hide

Sampled tokens can look aligned while residual Jacobians have already reoriented. Run JVP estimators and a five-stage protocol for NIST and EU GPAI logs.

  • Editorial Team
  • September 9, 2026
  • Philosophy of AI

Call It a Workspace Only After Nonlinear Ignition and Capacity Limits

Global workspace theory ignites then broadcasts under capacity limits. Scaled dot-product attention only routes over token positions.

  • Editorial Team
  • September 8, 2026
  • Deep Dives

The Shared Bus Lives in Residual Streams, Not Attention Maps

Residual streams analogize capacity-limited broadcast better than attention maps. Softmax routing stays graded and content-addressable, not ignited.

  • Editorial Team
  • September 3, 2026
  • Deep Dives

Softmax Weights Replace the Ignition Threshold Global Workspace Theory Needs

GWT treats conscious access as competition for a limited broadcast. Transformers use graded QKV routing on a residual stream, not ignition.

  • Editorial Team
  • September 2, 2026
1 2 3 4
Next
Copyright © 2026 - WordPress Theme by CreativeThemes