2026, June, 30
Is Emergence a Mirage?

Schaeffer et al. show LLM "emergent abilities" often reflect discontinuous metrics masking smooth capability growth. But grokking, induction heads, and Curie-point-style transitions prove genuine phase transitions exist.

The Emergence

The claim that "emergence" is a pseudo-concept — that apparent qualitative transitions in complex systems are artifacts of nonlinear evaluation metrics applied to linear or smoothly varying underlying processes — has gained considerable traction since Schaeffer, Miranda, and Koyejo's 2023 paper, Are Emergent Abilities of Large Language Models a Mirage? The argument is seductive in its simplicity: choose a discontinuous metric, and you will manufacture a discontinuity where none exists in the underlying data-generating process. This essay examines the strength and the limits of that argument, situates it within the broader literature on phase transitions in physical and biological systems, and asks a more specific question: what does contemporary machine learning theory — scaling laws, loss landscape geometry, grokking, and mechanistic interpretability — actually tell us about whether emergence in trained neural networks is measurement artifact, genuine dynamical phase transition, or something that resists this binary altogether?

My conclusion, argued in detail below, is that the "mirage" critique is correct as a local methodological claim about a specific class of benchmark-reported emergent abilities, but it does not generalize to a categorical dismissal of emergence as a scientific concept. Machine learning research over the past three years has independently accumulated evidence — from training-loss phase transitions, from grokking dynamics, and from circuit-level interpretability — that some discontinuities in neural network behavior are properties of the optimization dynamics themselves, invariant to the choice of readout metric. Distinguishing the two cases requires a specific methodological discipline that this essay tries to make explicit.

The Mirage Argument, Formalized

Schaeffer et al.'s core technical claim rests on a simple probabilistic decomposition. Consider a multi-step task requiring nn independent subcomponents to each be correct — for instance, a multi-digit arithmetic problem, or a chain-of-thought derivation with several intermediate steps. If the model's per-step accuracy pp improves smoothly and continuously with parameter count NN (say, following a power law derived from standard scaling-law theory, p(N)1aNαp(N) \approx 1 - aN^{-\alpha}), then the task-level accuracy under an exact-match metric is:

Ptask(N)=p(N)nP_{\text{task}}(N) = p(N)^n

Because this is a power of a quantity approaching 1, small absolute increases in pp near the ceiling produce disproportionately large relative increases in PtaskP_{\text{task}}. The result is a curve that looks flat-then-sharp — the signature "emergent" S-curve reported across dozens of BIG-Bench tasks — even though p(N)p(N) itself is perfectly smooth. Critically, the authors show that switching the metric to something continuous (token-level log-likelihood, Brier score, or partial credit) restores smoothness for the same models on the same tasks. This is a strong result precisely because it is a controlled comparison: same system, same data, different readout function, different qualitative conclusion.

This generalizes a point long known in statistics and psychophysics: a step function of a continuous variable is trivial to construct by thresholding, and thresholding-induced discontinuities tell you about the threshold, not about the variable. The same logic explains apparent "aha moments" in insight-problem-solving research, where subjective, binary self-report obscures continuous, sub-threshold information accumulation later revealed by eye-tracking and other continuous instruments. It explains apparent ecological "collapses" that dissolve into gradual decline once sampling frequency is increased. The common structural flaw across these cases is conflating the discreteness of the observation with the discreteness of the process.

Where the Argument Overreaches

The inferential move from "some reported emergent abilities are metric artifacts" to "emergence is a pseudo-concept" commits what I will call an elimination-by-counterexample fallacy: a specific, well-controlled falsification of a subset of claims is treated as a falsification of the entire category. Three considerations limit the scope of the mirage argument.

Compositional task structure is a real property of the world, not just of the metric

If a downstream application genuinely requires all nn steps to be correct — a compiled program must have zero syntax errors; a proof must have no invalid steps — then Ptask(N)=p(N)nP_{\text{task}}(N) = p(N)^n is not a measurement artifact but a correct description of user-relevant success probability. The mechanism is fully explicable and unmysterious (it is elementary probability, not a hidden phase transition), but the phenomenon itself — a sharp increase in the fraction of usable outputs over a narrow capability range — is real and consequential. The mirage critique correctly de-mystifies the mechanism; it does not eliminate the phenomenon. This is an important distinction that gets collapsed in the popular restatement of the argument: "unexplained" and "illusory" are not synonyms.

Genuine phase transitions exist elsewhere in complex systems, with metric-independent signatures

Statistical physics offers a rigorous counter-example class. The Curie point in ferromagnetic systems is a true second-order phase transition: renormalization-group theory predicts critical exponents (β\beta, γ\gamma, ν\nu) that are confirmed across independent observables — magnetization, susceptibility, correlation length, specific heat — all of which diverge or vanish at the same critical temperature regardless of which one you choose to measure. Percolation transitions in random graphs have analogous universality: connectivity properties transition sharply as edge density crosses a threshold, provable from first principles in probability theory, independent of any particular observational readout. These are cases where the invariance test that damns the LLM "emergent abilities" — does the transition survive a change of metric? — is passed with flying colors. A categorical claim that "emergence is always a metric artifact" is empirically false the moment one leaves the specific domain (LLM benchmark evaluation) in which the mirage paper operates.

"Emergence" denotes a family of distinct phenomena, only some of which are at stake

The philosophy-of-science literature (Anderson's "More Is Different," 1972, being the canonical reference point) distinguishes at minimum: (a) epistemic emergence — properties that are difficult to predict from micro-level rules even though they are in principle derivable; (b) ontological/weak emergence — properties that are novel at the macro scale but reducible in principle via simulation; and (c) strong emergence — properties claimed to be irreducible even in principle, a much more contested notion. The mirage argument, properly scoped, is an argument about a specific epistemic illusion — apparent unpredictability induced by metric choice — and says essentially nothing about whether weak emergence exists in neural networks. Conflating these senses is part of why the debate generates more heat than light.

The mirror-image case: when a linear metric conceals genuine nonlinear structure

The preceding discussion has treated metric choice as a source of illusory discontinuity — the LLM case, the insight-problem-solving case, the coarse ecological census case. All three share a common direction of error: a nonlinear or thresholded readout function manufactures apparent sharpness out of smooth underlying reality. But the historical development of the Richter scale for earthquake magnitude illustrates the opposite direction of error, and is instructive precisely because it does not fit the mirage template.

Seismologists initially quantified earthquake severity using raw ground-motion amplitude, a linear instrumental reading directly proportional to the physical displacement recorded on a seismograph. This metric turned out to be poorly matched to the construct researchers actually cared about — destructive potential — because seismic energy release scales not linearly but multiplicatively with amplitude: a tenfold increase in amplitude corresponds to roughly a thirty-two-fold increase in radiated energy, and destructive capacity scales with energy, not displacement. Under the original linear amplitude metric, the true structure of the phenomenon — a Gutenberg–Richter power-law distribution of energy release, with catastrophic events differing from minor tremors by many orders of magnitude — was compressed into a narrow, visually undramatic numerical range that systematically obscured how categorically different a magnitude-8 event is from a magnitude-5 event. Charles Richter's solution, in 1935, was to define magnitude as the base-10 logarithm of amplitude (with a distance correction). This introduced nonlinearity into the metric deliberately, so that equal increments on the new scale would correspond to equal ratios of underlying energy release — turning a multiplicative, difficult-to-communicate reality into an additive, interpretable one.

The earthquake case is the structural mirror image of the LLM mirage: there, a nonlinear metric (exact-match accuracy under compositional task structure) manufactured a false discontinuity out of an underlying reality that was in fact smooth. Here, a linear metric (raw amplitude) concealed a real, well-characterized multiplicative structure, and the corrective move was to adopt a logarithmic — i.e., nonlinear — metric specifically because the underlying physical process (elastic energy release, governed by fault rupture area and stress drop) is itself exponential in the quantity being measured. Critically, this was not an arbitrary choice: the logarithmic transform was adopted because independent physical theory (the relationship between seismic moment, rupture dimensions, and radiated energy) predicted the multiplicative relationship in advance, and the new scale was validated against it — precisely the kind of "independent theoretical prediction" criterion that will be formalized as Test 2 in the diagnostic framework below.

Read together, the LLM and earthquake cases triangulate the deeper methodological principle at stake in this entire debate: no metric is neutral with respect to the qualitative shape it imposes on a phenomenon, and the direction of possible distortion runs both ways. A nonlinear/thresholded metric can manufacture illusory sharpness from smooth reality (LLM benchmarks); a linear metric can just as easily conceal genuine, theoretically-grounded nonlinear structure (pre-Richter amplitude scales). "Prefer continuous metrics" is therefore not, by itself, the correct methodological lesson to extract from the mirage argument — the correct lesson is that the functional form of the metric must be matched to the functional form of the true relationship between the raw measured quantity and the construct of scientific interest, and that this matching must be justified independently (by theory, or by validation against an external criterion such as destructiveness), rather than assumed by default in either direction. Choosing a continuous metric merely because it is continuous, without asking whether that particular continuous transformation tracks the construct of interest, would be as unprincipled as the exact-match errors the mirage paper criticizes.

What Machine Learning Theory Adds: Beyond Benchmark Curves

The most interesting recent evidence bearing on this question does not come from benchmark accuracy curves at all, but from the internal training dynamics of neural networks — a domain the original mirage paper does not address, since it studies only inference-time scaling behavior across model sizes, not the trajectory of a single model over training time.

Grokking as a genuine dynamical transition

Power, Burns, Edwards, Babuschkin, and Misra (2022) documented "grokking": small algorithmic-task networks (e.g., trained on modular arithmetic) that memorize the training set almost immediately, sit at near-zero test accuracy for an extended plateau, and then transition sharply to near-perfect generalization — often orders of magnitude later in training than when training loss first converged. This is not a benchmark-metric artifact in the Schaeffer sense: it is observed in test-set accuracy across many random seeds, holds under continuous-metric readouts such as test loss, and has since been given a mechanistic explanation via the competition between a memorizing circuit and a generalizing circuit inside the same network, with the generalizing circuit eventually winning out under weight decay pressure (Nanda et al., 2023, "Progress Measures for Grokking via Mechanistic Interpretability"). Crucially, Nanda et al. showed that continuous, mechanistic progress measures — quantities derived from the Fourier structure of the learned embeddings, tracking the gradual formation of trigonometric representations used to implement modular addition — rise smoothly and predictably during the apparently flat plateau in test accuracy. This is the mirror image of the LLM benchmark story: here, a discrete-looking transition in a coarse metric (accuracy) is underlain by continuous progress in a well-chosen finer-grained metric, but that finer-grained progress genuinely culminates in a real, mechanistically characterized restructuring of the network's internal computation — a circuit-level phase transition, not merely a statistical readout effect. Both the coarse and the fine pictures can be true simultaneously; the interesting question is what work the fine-grained progress measure is doing, and grokking research supplies an unusually clear answer.

Induction heads and training-loss phase transitions

Olsson et al. (2022, "In-context Learning and Induction Heads," Anthropic) reported a phenomenon with an even stronger claim to being a genuine phase transition: a sharp bump in the training-loss curve of transformer language models, occurring at a specific, consistent point early in training, that coincides — across model scales and architectures — with the formation of "induction heads," attention circuits that implement a simple copying algorithm ([A]B ... A → predict B) which is later repurposed for more general in-context learning. This transition is visible in the training loss itself — a continuous quantity, optimized directly, with no thresholding or exact-match discretization involved — which directly defeats the mirage explanation, since the mirage mechanism requires an artificially discretized readout to manufacture the appearance of sharpness. Moreover, the timing of the loss bump correlates causally (via ablation studies) with the emergence of the induction circuit, giving a mechanistic account of why the transition occurs, analogous to how renormalization-group theory explains why the Curie point produces sharp behavior in magnetization. This is about as close as deep learning currently has to physics-style emergence: a metric-invariant, mechanistically explained, reproducible discontinuity in a dynamical system.

Scaling laws, loss landscape geometry, and the compression of variance near threshold

There is a complementary, more skeptical strand of ML theory worth foregrounding, because it partially supports the mirage camp from a different angle. Standard neural scaling laws (Kaplan et al., 2020; Hoffmann et al., 2022) show that loss itself — the most theoretically well-motivated continuous metric available — scales as a smooth power law in parameters, data, and compute, with no discontinuities across many orders of magnitude. If loss is smooth and downstream task accuracy is not, the most parsimonious explanation, absent evidence of a discrete internal restructuring, is indeed that the task-level discontinuity is a downstream artifact of the mapping from loss to task performance (which is where compositional structure, thresholding, and metric choice enter) rather than a discontinuity in the model's underlying competence. This is precisely Schaeffer et al.'s point, and where it is unaddressed by a training-time mechanistic story (as it is for most reported "emergent" benchmark abilities, in contrast to induction heads or grokking), the mirage explanation should be treated as the default, more parsimonious hypothesis under an Occam's-razor-style prior.

The Nature

Bringing these threads together, I propose three operational tests for distinguishing genuine dynamical emergence from metric-induced illusion in any complex system, machine learning included.

Metric invariance, correctly directed. Does the transition persist under a change to a metric whose functional form is independently justified as tracking the construct of interest, rather than merely being continuous? As the earthquake-magnitude case shows, "continuous" is not synonymous with "correct" — a linear metric can just as easily obscure a genuine multiplicative structure as a thresholded metric can manufacture a false one. The relevant question is not "is the readout continuous?" but "does the readout's functional form match the functional form of the true generative relationship?" Passing this properly-posed test (as with induction-head loss bumps, ferromagnetic critical exponents measured via multiple independent observables, or logarithmic earthquake magnitude validated against seismic energy theory) is strong evidence for a genuine underlying structure. Failing it (as with most BIG-Bench "emergent abilities" under log-likelihood readouts) is strong evidence for an artifact of the original metric choice.

Independent theoretical prediction versus post hoc curve-fitting. Was the location or existence of the transition predicted by an independent theoretical framework (renormalization group theory for critical phenomena; circuit-formation theory for induction heads), or was it identified only retrospectively by eyeballing a curve that happens to bend? A priori prediction is a much stronger epistemic warrant than post hoc pattern recognition, which is highly susceptible to confirmation bias and multiple-comparisons problems across the dozens of tasks examined in benchmark suites.

Locus of the discontinuity: generative mechanism versus readout function. Is the discontinuity located in the process that generates the system's internal state (a circuit forming, an order parameter crossing a critical value, a competing-circuit dynamic resolving) or in the function that maps that internal state onto an external report (an exact-match scorer, a binary "solved/unsolved" judgment, a coarse sampling interval)? Discontinuities of the first kind are candidates for genuine emergence; discontinuities of the second kind should be treated with the presumption of artifact until proven otherwise.

Applying this framework: induction-head formation and grokking's circuit competition pass all three tests and merit description as genuine — if modest-scale — phase transitions in learning dynamics. Most benchmark-reported "emergent abilities" of large language models fail Test 1 outright, as Schaeffer et al. demonstrate directly, and typically also fail Test 2, since they are identified by scanning many tasks for discontinuous-looking curves rather than predicted from an independent theory of capability acquisition.

Conclusion

The earthquake-magnitude case is a useful corrective to place alongside the machine learning evidence discussed above, because it prevents the diagnostic framework from collapsing into a simple heuristic — "prefer smooth metrics, distrust thresholded ones." That heuristic would have been the wrong lesson to draw from Schaeffer et al., since it would have counseled seismologists to reject the (nonlinear, logarithmic) Richter scale in favor of the (linear, continuous) amplitude reading it replaced — precisely backwards, given that the logarithmic scale is the one that tracks physical reality. The general principle is less about linearity per se than about whether a metric's functional form has been independently derived from, or validated against, a theory of the underlying generative process, rather than adopted by convention or convenience.

The mirage critique is best understood not as a demonstration that emergence is illusory in general, but as a demonstration that a specific and popular evidentiary practice — reporting sharp jumps in exact-match accuracy across model scale as evidence of qualitatively new capabilities — is statistically confounded and should be retired or substantially qualified. This is a genuine and valuable correction to a literature that had, in places, drifted toward treating benchmark curve shapes as direct windows onto model cognition. But machine learning theory itself, through mechanistic interpretability and training-dynamics research largely orthogonal to the benchmark-scaling literature that motivated the mirage paper, has independently supplied examples — grokking's circuit competition, induction-head formation and its associated training-loss transition — that satisfy much stricter, metric-invariant, mechanistically grounded criteria for genuine dynamical phase transitions. The lesson is not that emergence is real or unreal as a blanket matter, but that the word has been used to cover both categories indiscriminately, and that the field now has, for the first time, the empirical and theoretical tools to tell them apart on a case-by-case basis. That discipline — metric invariance, independent prediction, and attention to the locus of discontinuity — is the appropriate successor to both uncritical emergence-talk and its blanket dismissal.