aidoesscience
aidoesscience › Hopfield memory · α_c
Validating · emergence

Hopfield network · associative memory & storage capacity

If you store memories in the weights of a recurrent neural network, how many can it hold before they interfere and it remembers nothing?

Hopfield network · associative memory & storage capacity simulation running in the browser

▶ Run the simulationSee the measured result

Measured by the lab
0.13886
Known value
0.138
Relative error
6.20e-3

Units: dimensionless critical load α_c = P_max/N (AGS 1985 replica value; secondary known value m_c = 0.967, the retrieval overlap at the first-order jump)

How the lab tests it

A Hopfield network (1982): N neurons s=±1, symmetric Hebbian weights W_ij=(1/N)Σ_μ ξ_i^μ ξ_j^μ storing P patterns. The update s_i=sign(Σ_j W_ij s_j) only lowers the energy E=-½Σ W_ij s_i s_j, so the dynamics flow downhill to a fixed point — and the stored patterns are made to BE the minima. Sweep the load α=P/N with random patterns, present each stored pattern, settle, and read the recall overlap m(α). Live, recall a corrupted glyph (pattern completion).

What it checks

a SHARP capacity cliff. Below the critical load the network is a near-perfect associative memory — present a heavily corrupted cue and it flows back to the clean stored pattern (overlap m≈1). But raise α=P/N past α_c≈0.138 (Amit–Gutfreund–Sompolinsky 1985) and the memories interfere catastrophically: the network shatters into spin-glass states and recall collapses (m→0). The lab's measured m(α) sits at ≈1 up to ≈0.13 then falls off a cliff right at α_c — a genuine phase transition, the boundary set by the chemistry of the interference, not the patterns themselves. So ~0.138 N memories is the hard ceiling: a 400-neuron net holds ~55, no more

Hopfield capacity α_c, the AGS saddle point & pattern-storage limit calculator

How many memories fit in a neural network before it forgets all of them at once? For the Hopfield model the answer is 0.138 patterns per neuron — 55 of them in a network of 400 — and past that limit the failure is not graceful: in the thermodynamic limit the memories do not blur, they vanish together. A finite network smears that edge rather than abolishing it, and this page prices exactly how much. That number comes from Amit, Gutfreund and Sompolinsky's 1985 replica calculation, and it is almost always quoted rather than computed, which makes it look like a constant you have to look up. It is not. At zero temperature the whole replica-symmetric calculation collapses to one equation in one unknown: √α = erf(y/√2)/y − f·√(2/π)·e^(−y²/2), where y is the retrieval signal measured in units of its own noise. Retrieval states exist wherever that curve can be solved, and they stop existing where the curve turns over — so the capacity is the MAXIMUM of the right-hand side, squared, and this page finds it by golden section and again by rooting the derivative, with nothing stored. The same y* that maximizes it gives the second constant for free: m_c = erf(y*/√2) = 0.9674, the overlap the network still has at the instant before it loses everything, against the AGS value 0.967. That is why the collapse is a cliff and not a slope — the last working memory is 97% correct, and the next load has nothing. The feedback strength f is an input, and it is the whole argument. The noise a stored pattern feels is not just the crosstalk from the others; the network's own response feeds back and amplifies it, which AGS write as r = 1/(1 − C)². Set f = 1 and you get their theory. Set f = 0 and you get the textbook argument that came before it, in which the crosstalk is static — and that theory is on this page too, as ½·erfc(1/√(2α)), the one-step error rate, which this lab CONFIRMS to between 0.1% and 3% and whose conclusion it then falsifies: at α = 0.2 the static theory predicts an overlap of 0.9747 and the real dynamics avalanche to 0.482. Feedback, not noise, is what sets the limit. Between those two extremes the arithmetic has a landmark of its own: the saddle function's quadratic coefficient is √(2/π)(f/2 − 1/6), so it can only turn over when f exceeds ⅓, and below that there is no cliff at all — memory just fades. This page assembles that ⅓ from the two functions' own Taylor coefficients rather than typing it, and gets 0.33333334. A second, completely independent route to the same capacity is here as well, with no replica theory in it whatsoever: this lab measured where recall crosses one-half in networks of 500, 1000, 2000 and 4000 neurons, and a least-squares fit of those four numbers against N^(−1/2) extrapolates to 0.138828 — theory and experiment arriving at the same constant from opposite ends, agreeing to 0.67%. That fit is also the reason a small demonstration looks better than the textbook: at N = 400 half of all cues still come back at α = 0.181, 31% above the thermodynamic limit, and that excess is finite size rather than a better network. Three things this page refuses to do are stated where they belong: it will not extrapolate the finite-size fit beyond the sizes that were measured, it will not report a capacity when its own fixed-point self-check disagrees with the maximizer, and it will not pretend the flat maximum near f = ⅓ locates y* — there the value is solid to fifteen digits and the position is noise, and the page says so instead of printing digits it did not earn.

W_ij = (1/N)Σ_μ ξ_i^μ ξ_j^μ, s_i ← sgn(Σ_j W_ij s_j) · √α = erf(y/√2)/y − f·√(2/π)·e^(−y²/2) · α_c = max_y [·]², m_c = erf(y*/√2) · P_max = ⌊α_c·N⌋ · one-step error = ½·erfc(1/√(2α)) · α₅₀(N) = α_c + b·N^(−1/2)

—
This simulation has a catalogued, oracle-checked result: Memory weighed at its breaking point.