Concept

Learning rules

How a system changes its own parameters from experience — the question neuroscience asks about synapses and control engineering asks about adaptive controllers.

Also called: plasticity · adaptation

From the EE side

From the neuro side

Hebb’s rule, spike-timing-dependent plasticity, and error-driven parameter updates are all answers to one question: given an outcome, which parameters should move, and by how much? A controller that retunes itself and a synapse that strengthens are solving the same problem under different names. They differ less in arithmetic than in wiring: what each element is told, and when.

A rule with a specification

In 1959 Bernard Widrow, then working on adaptive filters, and his doctoral student Ted Hoff found the rule now called least mean squares (LMS). They published it in 1960 with Adaline, a weighted sum and threshold — the artificial neuron’s arithmetic. Each weight moves in proportion to the error times its own input:

wj←wj+2μ ϵ xjw_j \leftarrow w_j + 2\mu\,\epsilon\,x_j

Here xjx_j is the input on weight wjw_j, ϵ\epsilon the desired output minus the weighted sum, and μ\mu a small constant. As approximate steepest descent on a quadratic error surface, it came with a specification: a stability bound on μ\mu, a time constant, and a misadjustment — excess mean-square error from its one-sample gradient, as a fraction of the minimum — that grows as adaptation speeds up.

By October 1960 Widrow, acting on Hoff’s suggestion of electroplating, had weights made of “memistors”: copper electroplated onto pencil lead, conductance set by the integral of the plating current, each plating direction by the sign of error times input. Adaptive control’s MIT rule (1958) has the same shape: error times the error’s sensitivity to the parameter.

Rules found rather than derived

Hebb’s 1949 postulate, a conjecture when he made it, has two factors and no error: a connection strengthens when one cell repeatedly takes part in firing another. Bliss and Lømo reported long-term potentiation in the rabbit hippocampus in 1973. Levy and Steward saw an effect of order with trains of stimuli in 1983; timing single spikes, Markram and colleagues (1997) and Bi and Poo (1998) showed that the order of pre- and postsynaptic spikes within about 20 ms sets the sign of the change.

None of these rules knows whether a change helped. In a three-factor rule, coincidence leaves a decaying flag at the synapse, called an eligibility trace, and a third signal decides whether and which way the weight changes: a neuromodulator such as dopamine, released by neurons whose firing tracks a reward prediction error. In 2014 Yagishita and colleagues found that dopamine promoted the enlargement of striatal spines only if it arrived 0.3–2 s after glutamate paired with postsynaptic spikes.

What each weight is told, and when

Hebb’s two factors are both at the synapse. LMS needs the weight’s own input and one error, computed at the output and broadcast to every weight. That layout is where the readings meet: three-factor rules use it too, and Intel’s neuromorphic Loihi chip builds it — local spike traces, extra state per synapse for tags, reward spikes broadcast to chosen synapses.

The analogy breaks at the broadcast. Adaline’s error belongs to one element, comes from a teacher’s target, and arrives while the input is still present — and for a linear element the input alone sets the error’s sensitivity to that weight. Dopamine is shared by many neurons, carries no target and arrives late, so each synapse must remember being active without knowing whether it was responsible. The MIT rule met delay too: phase lag between error and adjustment could make it unstable. Specificity costs wires; delay costs memory.

Nearby concepts

All topics under Learning rules