Concept

Sparse coding

Representing information with few active elements at a time, rather than many slightly-active ones — cheap in energy and bandwidth, and apparently the strategy sensory systems settled on.

Also called: event-driven representation · efficient coding · sparse representation

From the EE side

From the neuro side

If most of your channels are quiet most of the time, you can spend your power budget on the ones that are not. Sensory neurons appear to have worked this out well ahead of sensor designers.

Few cells at a time

In the 1961 chapter behind efficient coding, Horace Barlow proposed that a sensory relay picks the code that spends the fewest impulses on average, and noted that cortex, with vastly more cells than the fibres feeding it, could spend fewer still. In 1996 Bruno Olshausen and David Field had a learning algorithm reconstruct natural image patches while penalising activity, so that only a few units were active for any one patch. It produced localised, oriented, bandpass features, like the simple cells of visual cortex.

One motive is energy. A binary unit carries the most bits per time slot firing in half of them, but if a spike costs ten times as much as silence, it carries the most per unit of energy firing about one slot in six. Levy and Baxter showed in 1996 that the costlier the spike, the lower the efficient firing rate, and Attwell and Laughlin’s 2001 budget for rodent grey matter predicted codes with at most 15% of neurons active at once.

Few coefficients, few measurements

Signal processing finds sparseness in signals, not cells. Natural images are sparse in a wavelet basis, which transform coding exploits. Compressed sensing, from Candès, Romberg and Tao and from Donoho in 2006, moves the saving into the measurement. A signal of NN samples with KK nonzero coefficients can be recovered exactly from on the order of Klog⁡NK \log N random measurements, and one with KK significant coefficients approximately, with an error comparable to that of keeping only its KK largest. The recovery minimises the ℓ1\ell_1 norm, the sum of absolute values — the criterion of basis pursuit, an earlier signal-processing method, and a penalty Olshausen and Field had also tried.

Where they meet, and where they part

They meet in the address-event representation, which sends each spike between chips as the binary address of the cell that fired, and nothing for silence; a scan pays for every output every frame. It came from Carver Mead’s Caltech group: Massimo Sivilotti’s 1991 thesis described an arbitered, data-driven readout for a silicon retina, and Misha Mahowald’s 1992 thesis named it. Kwabena Boahen cut its overhead in 2000, and neuromorphic hardware adopted it widely. Mahowald’s model showed the catch: far better timing than a scan at low event rates, rapidly worse past a critical one.

They part twice. The two words are used almost interchangeably, and the two properties usually travel together, but they say different things: sparse is how many units are active, event-driven is when a unit speaks. Olshausen and Field’s code is sparse but recomputed for every image; an event camera, descended from the silicon retina, reports only changes in brightness, so it is sparse only while little in the scene changes.

And the readings spend sparseness in opposite directions. Compressed sensing uses it to take fewer measurements than samples. A sparse code, as Field described it in 1994, keeps the dimensionality of its input and may even raise it. In cat primary visual cortex, by Olshausen’s estimate from anatomical counts, axons leaving for higher areas outnumber those arriving from the thalamus about 25 to 1. A sampled system pays for every channel, and compresses; cortex pays mainly for activity, and expands.

Nearby concepts

All topics under Sparse coding