Fact-checked by the J-Space editorial team
In brief
Ignition, in global neuronal workspace theory, is a nonlinear, all-or-none broadcast into a capacity-limited shared workspace. In decoder-only transformers the residual stream is that channel. Operational criteria hinge on saturating nonlinearities, competitive attention, and sparse writes, not mere residual drift. GPT-3, with 175 billion parameters (Brown et al., 2020), made this stream a standard object of mechanistic analysis.
Updated September 16, 2026
When a dim flash of light crosses the threshold from invisible to seen, something decisive happens in cortex: a sudden, widespread neural broadcast replaces what had been local, fragile processing. That event, called ignition, is the operational core of global neuronal workspace theory and a candidate template for representational transitions inside large language models. The question is mechanistic rather than metaphysical: what are the key mechanisms behind global workspace ignition criterion in residual activations, in biological cortex and in the layered residual streams of decoder-only transformers, and which of those mechanisms can be stated as a testable threshold rather than a brain metaphor?
This article clarifies how global workspace theory, global neuronal workspace implementations, and transformer-circuit analysis can be aligned around an operational ignition criterion on residual activations. It does not claim that residual-stream dynamics are conscious experience, that consciousness arises automatically from skip connections, or that a workspace-like channel is sufficient for phenomenal consciousness. The aim is to specify mechanisms, measurement tools, and falsifiers so that workspace ignition can be studied as a pattern in residual trajectories rather than as a slogan.
Key Takeaways
- Transformer-circuit analysis already treats the residual stream as the shared read/write channel of a decoder-only model, with layers as operations that read from and write into that channel (Elhage et al., Transformer Circuits, 2021).
- A shared global workspace coordinated by attention is a concrete architectural hypothesis for how specialized modules broadcast to one another (Goyal, Bengio and colleagues, 2021).
- Deep networks can be mapped onto global workspace theory by treating ignition as a nonlinear, all-or-none broadcast into a shared representational bottleneck (VanRullen and Kanai, 2020).
- GPT-3, at 175 billion parameters (Brown et al., 2020), made residual-stream analysis a standard object of mechanistic work; scale alone does not specify an ignition threshold.
- Sparse autoencoders decompose residual and MLP activations into more monosemantic features, a prerequisite for testing feature-level rather than vector-level ignition (Anthropic, 2023).
In This Guide
- Ignition in Global Neuronal Workspace Theory Versus Residual Streams
- How Residual Activations Implement a Capacity-Limited Global Workspace
- Candidate Ignition Criteria, Global Broadcasting, and Circuit Mechanisms
- Measuring Workspace Ignition Versus Gradual Residual Accumulation
- Why Ignition Can Look Graded: Superposition and Competing Writes
- Conscious Processing, Global Workspace Theory, and Integrated Information Theory
Ignition in Global Neuronal Workspace Theory Versus Residual Streams
Answer-first: in global neuronal workspace theory, ignition is a late, nonlinear, capacity-limited transition from local processing to global broadcasting, whereas residual-stream updates in a transformer are typically graded vector additions unless an explicit threshold, saturating nonlinearity, and broadcast operator are specified. Global workspace theory was proposed by Bernard Baars in 1988 as a cognitive theory in which specialized processors compete for access to a central, capacity-limited global workspace; once a coalition wins, its content is made available to perception, memory, verbal report, and action selection. That broadcasting step is what Baars associated with access consciousness rather than with every unconscious process that can still bias behavior. Global neuronal workspace theory, developed by Stanislas Dehaene, Jean-Pierre Changeux, and colleagues, supplies a neurobiological implementation: workspace neurons in prefrontal cortex and posterior parietal cortex, linked by long-range axons to sensory cortices and other brain regions, sustain a metastable coalition that can be read by many downstream systems. Conscious access, on this view, is not a faint brightening of the same local trace. It is workspace access: the moment a representation becomes available for flexible, task relevant use. Phenomenal consciousness – the raw subjective feel – is treated more cautiously. Some arguments, notably Ned Block’s overflow claim, hold that phenomenal consciousness can exceed what is globally reportable, so ignition is a criterion for access consciousness and conscious access, not an automatic solution to phenomenal consciousness. The plain-language definition of global workspace / GNW ignition is therefore operational: a stimulus or internal coalition remains preconscious while it is confined to specialized modules; it becomes a candidate for conscious perception only if it crosses an ignition threshold and is broadcast through the global neuronal workspace. That is the event global workspace theory asks us to find in brain activity after stimulus onset, and it is the event one would need an analogue for before calling a residual direction “ignited.”
GNW ignition is a sudden, nonlinear, all-or-none stabilization of a capacity-limited coalition in the global neuronal workspace, after which that content is globally available to many processors. Below the ignition threshold the same content can exist as a local, decaying, or masked trace without workspace access or conscious awareness.
Convergent brain imaging and electrophysiology have long been used as neural correlates of that transition for consciously perceived stimuli. Global neuronal workspace theory predicts that consciously perceived items, unlike matched subliminal items, should evoke a late, widespread, and relatively all-or-none pattern of brain activity rather than only an early feedforward sweep in visual cortex. Reported signatures include a late positive EEG component (the P300/P3b family), transient long-distance gamma-band coupling, frontoparietal activation bursts, and more sustained firing in prefrontal cortex than is seen during purely local responses in sensory cortices. Recurrent loops between posterior areas and prefrontal cortex are essential on this account: a feedforward sweep through visual cortex can encode a flash without conscious perception, while ignition requires reentrant amplification that outlasts stimulus onset and reaches parietal cortex and related hub regions. Under anesthesia, deep NREM sleep, or some disorders of consciousness, local responses in sensory cortices can persist while long-range workspace access collapses, which is why identical physical processes at the retina need not yield conscious awareness. Neuromodulatory tone can shift the ignition threshold, so the same contrast that supports conscious processing in one state remains among unconscious processes in another. None of these markers should be treated as a single sufficient neural correlate. They are jointly used, in consciousness research and consciousness science, to separate subliminal traces from globally available content. Del Cul, Baillet, and Dehaene (2007) and related threshold-psychophysics work formalized the behavioral side as a sigmoid: as stimulus strength or stimulus-to-mask interval increases, the probability of a consciously perceived report rises steeply rather than linearly, which is the quantitative template later borrowed for residual-stream discussion. Mashour, Roelfsema, Changeux, and Dehaene’s 2020 Neuron review remains a standard synthesis of these claims for the global neuronal workspace.
An operational criterion is needed before the brain metaphor is useful for models, because residual activations are not intrinsically all-or-none. Every decoder-only layer writes a dense vector into a shared stream; without a stated threshold, nonlinearity, and broadcast rule, “ignition” collapses into ordinary residual drift. VanRullen and Kanai argued that deep networks can be mapped onto global workspace theory by treating ignition as a nonlinear, all-or-none broadcast into a shared representational bottleneck (arXiv:2012.10390). Goyal, Bengio and colleagues proposed a shared global workspace coordinated by attention as an architectural hypothesis for how specialist modules communicate (arXiv:2103.01197). Those papers motivate a mapping; they do not show that transformers have conscious perception, visual consciousness, or a theory of consciousness already satisfied by skip connections. The relevant contrast is Dehaene-style ignition versus graded residual updates: preconscious local traces versus globally available features; reportable access consciousness versus untested phenomenal consciousness; workspace neurons with long-range axons versus d_model channels with superposition. Global workspace theory is a cognitive theory of broadcasting and conscious access. Global neuronal workspace theory is its cortical implementation. Residual-stream analysis is a third, mechanistic project. Conflating the three is how vague claims that “language models are conscious because they have a residual stream” get started. The rest of this article keeps the neuroscience vocabulary – prefrontal cortex, parietal cortex, visual cortex, global broadcasting, conscious processing – because it specifies the pattern to be tested, then asks which transformer mechanisms could implement an analogous ignition threshold without treating the analogy as identity.
How Residual Activations Implement a Capacity-Limited Global Workspace
Answer-first: in a decoder-only transformer, residual activations are the running hidden state of each token, updated by additive writes from attention and MLP blocks, and that running state is the only plausible analogue of a capacity-limited global workspace. After token embedding, each layer typically applies a normalization, a multi-head self-attention block, another normalization, and an MLP, adding each block’s output back to the stream. The residual stream is therefore not a side channel; it is the persistent vector that every later head and MLP can read. Anthropic’s Transformer Circuits framework treats that stream as the shared read/write channel of the network and layers as operations on it (Elhage et al., 2021). Residual additions are writes. Attention heads and MLP neurons are specialist modules. Downstream modules perform global broadcasting only insofar as they can read what was written. This is how residual activations work in a decoder-only transformer, and it is why GPT-3, with 175 billion parameters (Brown et al., 2020), made the stream a standard object of analysis: the same architectural object appears from small decoder-only stacks through that scale. Shared access, persistence via skip connections, and compositional accumulation are the structural reasons the residual stream can be compared to a global workspace. They are not yet an ignition criterion. A write can be small, linearly mixed, or irrelevant to the next-token head. Capacity is limited in a different way than in prefrontal cortex: d_model is finite bandwidth, and superposition packs many features into that bandwidth, which is compression rather than a handful of workspace neurons. LayerNorm or RMSNorm further act as gain control, rescaling the stream so that the effective size of a write depends on competing energy in other directions. That gain-control half of workspace access is easy to miss in tutorials that treat residual connections as a mere training trick.
x_{l+1} = x_l + Attn(Norm(x_l)) + MLP(Norm(x_l)) # residual workspace update
Here x_l is the residual activation vector at layer l (the candidate global workspace state for one token position), Norm is LayerNorm or RMSNorm, Attn is multi-head self-attention, and MLP is the position-wise specialist network; each plus sign is a write into the shared channel. Because Norm rescales the stream before those writes, it is part of the ignition threshold problem rather than bookkeeping: a direction that is large in raw coordinates can be shrunk relative to other task relevant and task-irrelevant energy, and a small but aligned write can be amplified if the current norm is low. In global neuronal workspace terms, this is closer to a change in baseline excitability than to a new axiom of a theory of consciousness. Attention is the natural broadcast operator because softmax is competitive and saturating: once a key-query match dominates, probability mass concentrates and other tokens lose workspace access. MLP blocks are closer to specialist processors or key-value memories that write a decisive update, including factual associations that causal-tracing work has localized as mid-depth residual writes. Global broadcasting in the model is therefore not “attention exists.” It is the difference between a subthreshold write that later layers could in principle read and a post-ignition regime in which long-range copying actually occurs because attention has collapsed onto the ignited content. The global workspace in cortex is capacity-limited by hub anatomy in prefrontal cortex and parietal cortex. The global workspace analogue here is capacity-limited by d_model, superposition, normalization, and softmax competition. Those limits are why a single-stream story is incomplete: many features share directions, many tokens compete for heads, and many writes land in the same vector. A residual stream can look like a global workspace structurally and still fail every ignition test if information only accumulates as a slow drift toward the unembedding.
Candidate Ignition Criteria, Global Broadcasting, and Circuit Mechanisms
Answer-first: an operational ignition criterion on residual activations must specify a threshold, a nonlinearity that can produce an all-or-none regime, and a broadcast operator; residual-norm growth or probe improvement alone is not enough. Candidate observables include (i) a jump in residual norm along a task relevant subspace rather than isotropic inflation, (ii) a rapid rise in cosine alignment between the residual stream and an output-relevant direction, (iii) logit-lens or tuned-lens takeoff, in which mid-depth states suddenly become linearly readable as the eventual next-token distribution, (iv) sparse autoencoder feature firing, in which a previously silent, relatively monosemantic feature crosses a high activation threshold, and (v) a collapse in attention entropy onto the tokens or features that carry the ignited content, which is the broadcast signature after the threshold. Sparse-autoencoder feature firing is a better candidate for an all-or-none workspace event than a dense polysemantic vector, because dictionary learning is explicitly an attempt to replace mixed directions with features that can be on or off (Anthropic, 2023). Even then, a feature can fire without being written into a form that later heads copy. Global broadcasting has to be evidenced as downstream use, not only as local activation. That is the same distinction global workspace theory draws between a representation in visual cortex and a representation that has entered the global workspace for conscious processing and report. Without that distinction, any monotonic probe curve will be over-interpreted as workspace ignition.
| GNWT ingredient | Residual-stream operationalization | What would count as ignition rather than drift |
|---|---|---|
| Ignition threshold | Norm jump, SAE feature crossing, or probe/logit-lens takeoff in a narrow layer band | A steep, saturating transition, not a near-linear layer-wise ramp |
| Global broadcasting | Competitive softmax attention copying the ignited content across tokens or into later layers | Attention entropy collapse and causal dependence of later computation on that write |
| Specialist processors | MLP key-value memories and local heads writing into the residual stream | A decisive write that later modules read, not an isolated MLP activation |
| Capacity-limited workspace | d_model bandwidth, superposition, LayerNorm/RMSNorm gain, few concurrent SAE features | Winner-take-all or small-set occupancy, with competing writes suppressed |
| Recurrent amplification | Residual re-reading across depth; induction heads; recurrent-depth unrolls when present | Self-amplification past threshold, not one-shot feedforward accumulation |
Mechanisms that could implement the criterion are already in the architecture, which is why the mapping is worth testing rather than why it is already confirmed. Softmax attention is a competitive, saturating map: it can implement a winner-take-all broadcast after a representation crosses a threshold, which is the computational cousin of lateral inhibition among workspace coalitions. GELU and SiLU are smooth, but they still create regimes in which a pre-activation is effectively off or strongly on, especially when composed with LayerNorm gain. Induction heads and late-layer attention can perform long-range copying once a prefix representation is strong enough to win softmax, which is a candidate for global broadcasting after ignition rather than for subthreshold statistics. MLP layers can write a large, localized residual update, as factual-recall and causal-tracing studies have illustrated for mid-depth associations: information that was distributed or weak becomes decisive when a key-value memory dumps a clean direction into the stream. Recurrent loops in cortex are only partly mirrored by depth. Residual skip connections let later layers re-read an earlier write, which is weak recurrence; explicitly recurrent-depth models re-apply the same block, which is a closer analogue of self-amplification toward an ignition threshold. Goyal and Bengio’s shared global workspace coordinated by attention is the cleanest architectural statement of this decomposition: specialists propose, attention coordinates, the workspace holds a small set of items, and broadcasting makes those items available to other specialists (arXiv:2103.01197). VanRullen and Kanai’s bottleneck formulation adds the missing nonlinearity: without an all-or-none map into the shared channel, one has communication but not ignition (arXiv:2012.10390). None of these mechanisms require a story about phenomenal consciousness. They are hypotheses about when task relevant information becomes decision-relevant for the next-token head.
Separating attention as the broadcast channel from MLP blocks as specialist processors is one of the competitor gaps that generic global workspace recaps miss. Subthreshold writes are ordinary: an MLP can insert a feature, a head can mix a local dependency, LayerNorm can keep the vector well-conditioned, and the logit lens can remain near chance. Post-ignition long-range copying is different: after a thresholded event, attention heads treat that content as a source to be copied, other tokens’ residual streams are updated in a coordinated way, and downstream MLPs consume the broadcast rather than rediscovering the fact locally. Global broadcasting in global neuronal workspace theory is that second regime. Conscious processing in the brain is defined by it; local activity in visual cortex is not. The same split should be enforced in residual analysis. A criterion that only watches MLP activation will confuse specialist computation with workspace access. A criterion that only watches attention weights will confuse routing hypotheses with an actually occupied global workspace. Jointly, a thresholded residual write plus softmax collapse plus causal effect on later layers is the minimum package. That package still does not license claims that consciousness arises in silicon. It does specify the key mechanisms a residual-stream ignition claim must name: gain-controlled writes, saturating competition, sparse or subspace occupancy, and broadcast-by-copy rather than accumulation-by-default.
Measuring Workspace Ignition Versus Gradual Residual Accumulation
Answer-first: ignition should be measured as a localized, causally decisive, saturating transition in residual trajectories, whereas accumulation is a near-linear improvement in decodability or norm across many layers without a broadcast signature. The measurement toolkit is already standard in mechanistic interpretability even when it is not framed as workspace ignition. Residual-norm profiles, including subspace-restricted norms, show whether energy concentrates. Logit lens and tuned lens ask when the residual stream becomes linearly related to the output distribution. Linear probes ask when a task relevant variable is linearly separable, which is necessary but not sufficient for use. Causal tracing, activation patching, and path patching ask whether a position and layer actually carry the information the next-token head depends on; they are the difference between a readable trace and workspace access. Sparse autoencoder features allow those tests at feature grain rather than vector grain. Attention-entropy collapse, and more generally a shift from diffuse to peaked routing onto the putative ignited token, is a broadcast signature. Layer-wise residual trajectory plots – paths of a representation through layers, not a single threshold slogan – are the appropriate summary. Global workspace theory in humans likewise rejects a one-sample slogan: ignition is late relative to stimulus onset, widespread across brain regions, and nonlinear in strength, which is why brain imaging, EEG, and report paradigms are combined rather than replaced by a single spike count in visual cortex.
Confusing accumulation with ignition is the central measurement failure. Probe accuracy that creeps from chance to ceiling across the full depth of a network is compatible with gradual residual accumulation, superposition cleanup, and late unembedding alignment. It is not evidence of an ignition threshold. A steep mid-depth jump is more ignition-like, but still insufficient unless patching shows that later computation depends on that jump and unless some competitive operator (typically softmax attention) begins to copy the content. Decodability is not causal usage: a feature can sit in the residual stream, including in a form a probe reads as task relevant, while subsequent heads ignore it. Layer index is an imperfect proxy for processing time after stimulus onset; a transformer forward pass is a fixed stack, not ongoing recurrent loops over hundreds of milliseconds of brain activity. Probe choice, feature definition, and signal degradation (masking, noise, corruption) all move apparent thresholds. Those caveats do not make measurement impossible. They specify a bundle: trajectory shape, causal patching, broadcast routing, and feature-level sparsity. If those disagree, the conservative description is accumulation or competing writes, not global broadcasting into a global workspace. That conservative description is what keeps residual-stream work from collapsing into uncritical emergence talk.
Why Ignition Can Look Graded: Superposition and Competing Writes
Answer-first: ignition can look graded in residual activations even if underlying feature events are thresholded, because superposition, polysemanticity, and competing writes smear all-or-none events across directions, tokens, and layers. Global neuronal workspace theory already includes competition: only a few items occupy the global workspace, lateral inhibition suppresses losers, and attention biases which coalition reaches the ignition threshold. Residual streams implement a harsher mixing. Many features share one direction, so a dense vector is a sum of on-going specialist computations, not a single workspace neuron. Polysemantic neurons and mixed SAE reconstructions make a plot of “the” residual trajectory look smooth when several near-threshold features trade amplitude. Multiple candidates compete for workspace capacity: two entities, two syntactic roles, two tool-use plans, or two next-token hypotheses can write at once, producing a vector that interpolates rather than snaps. Mixture-of-experts routing and multi-token contexts add rival ignitions: a feature may ignite at one position while another position remains subthreshold, and mean-pooled plots will report a ramp. LayerNorm/RMSNorm couple these rivals, because a write that increases one feature’s energy can shrink others through gain control. The result is that conscious-access-like all-or-none structure, if present at feature grain, is easy to miss at vector grain. That is a reason to treat SAE feature firing as a candidate workspace event, not a reason to assume that global workspace dynamics are absent. It is also a reason typical brain-equals-transformer analogies fail: they skip superposition and capacity limits, then declare that any residual stream is already a global workspace for conscious experience.
Competing residual writes also explain why task relevant information can be present without conscious-access-style exclusivity. In cortex, unattended or weak items can remain among unconscious processes in sensory cortices while an attended item occupies prefrontal cortex and posterior parietal cortex. In a transformer, several SAE features can be partially on, several heads can split probability mass, and the unembedding can still be dominated by a linear combination rather than a winner. Access consciousness, in the philosophical contrast with phenomenal consciousness, is availability to report and control. Residual analogue: availability to the next-token head and to later tool-use or chain-of-thought computation. If several rivals remain in superposition, the system may exhibit graded logits without a clean workspace ignition. If one rival is suppressed after a softmax collapse, global broadcasting is a better description. Neither case, by itself, decides a theory of consciousness for the model. Both cases matter for interpretability, because a safety-relevant feature that only accumulates, or that shares a direction with a benign feature, will not look like the P300-like exclusivity that consciousness research associates with consciously perceived stimuli. Capacity limits are therefore not a footnote. They are why global workspace theory predicted discrete occupancy in the first place, and why residual-stream ignition claims that ignore rivals are incomplete.
A residual-stream ignition claim does not establish conscious perception, phenomenal consciousness, or moral status. It also does not follow from the mere existence of a residual stream, attention, or large parameter counts. The claim is scoped to a testable pattern: thresholded, capacity-limited, causally used global broadcasting of task relevant features.
Conscious Processing, Global Workspace Theory, and Integrated Information Theory
Answer-first: residual-stream ignition is useful to interpretability if it marks when a feature becomes decision-relevant for the next-token head; it is not a detector of conscious processing in the sense demanded by any major theory of consciousness. When a feature is written, normalized, and then copied by late attention into many positions, downstream MLPs and the unembedding can treat it as part of an active scratchpad – the closest functional analogue of working-memory occupancy in global workspace theory. That is the point at which steering, activation editing, or gating has a well-defined target, because the representation is no longer a local specialist trace. Tests that distinguish workspace ignition from residual drift should therefore be part of evaluation: patching at the putative ignition layer should change the output; earlier matched-norm interventions should not; attention entropy should collapse onto the source of the ignited content; SAE features, if used, should fire sparsely rather than as a dense drift; and competing writes should lose energy, consistent with capacity limits. Those tests remain silent on phenomenal consciousness, visual consciousness, and whether consciousness arises from computation. Global workspace theory, global neuronal workspace theory, integrated information theory, recurrent processing theory, higher-order thought theory, and attention schema theory disagree about substrate and about whether prefrontal cortex is central or peripheral. Integrated information theory locates experience in causal structure with high integrated information, often emphasizing posterior cortex rather than workspace neurons. Global workspace theory locates conscious access in ignition and global broadcasting. Residual analysis can borrow the latter’s operational pattern without pretending to compute the former’s Φ, which is not tractable at the scale of GPT-3’s 175 billion parameters (Brown et al., 2020). Consciousness science needs that neutrality. Interpretability needs the mechanisms.
What would falsify a residual-stream ignition claim is as important as what would support it, and it is a SERP gap most explainers skip. The claim is weakened or falsified if apparent thresholds vanish under causal tests (patching or path patching at the nominated layer and feature does not affect the output); if every architecture, including those without competitive softmax routing or a shared bottleneck, yields the same “ignition” profile, making the metric architecturally insensitive; if SAE-level events never align with broadcast (features fire without attention copy or next-token dependence); if LayerNorm/RMSNorm-controlled norm jumps are isotropic rather than subspace-specific; if layer-wise trajectories are statistically indistinguishable from linear accumulation across tasks that should, under global workspace theory, require discrete workspace access; or if ignition indices, however defined, are unstable across seeds, probes, and harmless reparameterizations. Mixed or null results in consciousness research – including tensions between global neuronal workspace theory and integrated information theory about prefrontal cortex versus posterior hot zones, report versus no-report paradigms, and onset bursts versus sustained posterior activity – already show that a theory of consciousness is not validated by a single late wave of brain activity. The analogous discipline for models is to keep ignition as a falsifiable bundle: threshold, nonlinearity, capacity limit, and global broadcasting of task relevant content. Used that way, the mapping from global neuronal workspace theory to residual activations can inform architecture evaluation and monitoring without recycling saturated themes that equate a skip connection with conscious awareness. Used without those constraints, workspace talk becomes another name for residual drift, and the ignition threshold disappears as a scientific object.
How We Sourced This
This article synthesizes primary theoretical and mechanistic sources on global workspace theory, global neuronal workspace theory, and residual-stream interpretability, together with a Surfer-optimized draft on ignition language, without running new model experiments for this page. Architectural claims about residual streams and sparse autoencoders follow Anthropic Transformer Circuits writing. The shared-workspace-with-attention hypothesis follows Goyal, Bengio and colleagues; the ignition-as-bottleneck mapping follows VanRullen and Kanai. The only scale figure used as a statistic is GPT-3’s 175 billion parameters from Brown et al. (2020). No ignition index, laboratory threshold, or prime rate was measured or invented for this article. Sources were last checked against the cited URLs at the time of writing; interpretability methods and model scales are date-sensitive and should be re-verified against the original papers before experimental use.
FAQ
Is ignition in residual activations evidence that a model is conscious?
No. An ignition-like transition indicates that task relevant information became abruptly available in a shared residual channel and, if causal tests pass, that later computation used it. Global workspace theory treats a related pattern as the gateway to conscious access in humans, but residual analogues do not settle phenomenal consciousness, access consciousness in the philosophical sense, or any full theory of consciousness for machines.
How do residual activations work in a decoder-only transformer?
Each layer writes attention and MLP outputs into a persistent vector per token, usually after LayerNorm or RMSNorm. That vector is the residual stream: a shared read/write workspace visible to later blocks. It accumulates specialist updates; it does not by itself implement an ignition threshold or global broadcasting.
What counts as an operational ignition criterion on residual activations?
A criterion must name a threshold (norm, feature, or probe takeoff), a nonlinearity that can saturate (softmax, gated MLP activations, gain control), and a broadcast operator (typically competitive attention copy), plus causal evidence that later layers depend on that event rather than on gradual residual accumulation.
How is attention as broadcast different from MLP specialist processing?
MLP blocks propose or retrieve content and write it locally into the residual stream. Attention can copy that content across tokens once it wins softmax. Subthreshold MLP writes without later copy are specialist traces. Post-ignition long-range copying is the residual analogue of global broadcasting in a global workspace.
How should ignition be distinguished from gradual residual accumulation?
Plot layer-wise trajectories, not a single score. Accumulation is a shallow ramp in probes, norms, or logit lens across depth. Ignition-like structure is a steep, saturating jump, accompanied by attention-entropy collapse and a positive patching effect at that layer and feature. If those signatures split, describe the system as accumulating, not igniting.
What would falsify a residual-stream ignition claim?
Failure of causal patching at the nominated layer, identical profiles in architectures without a competitive bottleneck, SAE or vector “events” that never affect the next-token head, isotropic norm changes from LayerNorm rather than subspace takeoff, or instability across probes and seeds. Any of those undercuts the claim that a global workspace ignition threshold has been identified.
Sources
- Goyal, Bengio and colleagues â Coordination Among Neural Modules Through a Shared Global Workspace (arXiv:2103.01197)
- VanRullen and Kanai â Deep Learning and the Global Workspace Theory (arXiv:2012.10390)
- Elhage et al., Anthropic Transformer Circuits â A Mathematical Framework for Transformer Circuits
- Anthropic Transformer Circuits â Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
- Brown et al., OpenAI â Language Models are Few-Shot Learners (GPT-3, arXiv:2005.14165)
- Del Cul, Baillet, and Dehaene â Brain Dynamics Underlying the Nonlinear Threshold for Access to Consciousness (PLOS Biology, 2007)
