ee→neuro · algorithm · 1993 · growing

Spike sorting

An electrode in the brain hears many neurons at once; spike sorting assigns each detected spike to a putative neuron by its shape and position, and population results inherit its errors.


In 1926 Edgar Adrian and Yngve Zotterman, recording from the nerve of a small frog muscle, saw impulses often a millisecond or less apart, closer than a nerve fibre’s refractory period allows, so more than one fibre was firing. They cut strips off the muscle until a single regular rhythm was left — the experiment behind rate coding. Inside a brain there is nothing to cut. An electrode in cortex records the sum of every spike within range and of many more too distant to tell apart, and the separating is done afterwards: detect each spike, describe its shape, group similar shapes, call each group a neuron. That is spike sorting, and most of what electrode recordings say about populations of neurons rests on it.

A receiver for a crowd

The pipeline is a detection-and-classification receiver. In the unsupervised method that Rodrigo Quian Quiroga and colleagues published in 2004, the signal is band-passed from 300 Hz to 6 kHz, which removes the slow field potential and keeps spikes a millisecond or two long, and a spike is declared wherever it crosses a threshold set against the noise:

Vth=4σn,σn=median⁡∣x∣0.6745V_{\text{th}} = 4\sigma_n, \qquad \sigma_n = \frac{\operatorname{median}|x|}{0.6745}

Here VthV_{\text{th}} is the threshold, xx the filtered signal and σn\sigma_n an estimate of the noise’s standard deviation; for Gaussian noise the median of ∣x∣|x| is 0.6745 standard deviations. The median is the point. A plain standard deviation counts the spikes too, so a busy channel would raise its own threshold and lose its small spikes; the median hardly notices them while they are a small fraction of the samples. It is the noise estimate of wavelet denoising, carried over. Each spike is then reduced to a few numbers — amplitude, width, principal components or wavelet coefficients — which are clustered, by hand in two-dimensional projections or by an algorithm.

The difficulty is that much of the noise is neurons. Henze and colleagues estimated that a tetrode in hippocampal area CA1 is within range of about a thousand, of which 60 to 100 should fire spikes large enough to separate. The rest sum into a continuous background. Quian Quiroga’s group simulated it as just that, small spikes superimposed at random, and noted that it gives noise and spikes similar spectra. With both in the same band, filtering by frequency has little left to offer; the separation has to come from the shape of each spike and, given more than one site, from position.

Listening from more than one place

The method is older than the probes. In 1964 George Gerstein and Wesley Clark, at MIT, burnt small holes in the insulation of a tungsten microelectrode so that it would hear several neighbouring neurons, and had a digital computer separate them by waveform: in Quian Quiroga’s account, by mean-square distance from a template chosen for each neuron.

On one wire, though, shape is a weak label. It depends on the cell’s form and on its distance and orientation from the tip, so two cells can look alike, and one can look like two, because in a burst its spikes shrink. The stereotrode of Bruce McNaughton, John O’Keefe and Carol Barnes (1983) answered with geometry: two tips side by side see each cell from two distances, and the ratio of its amplitudes on them depends on where the cell sits, not on how large its spike happens to be, so it holds through a burst. They isolated up to five units at once in rat hippocampus. The tetrode, four wires, which Charles Gray and colleagues credit to O’Keefe and Recce and to Wilson and McNaughton in 1993, did better again. In cat visual cortex Gray’s group isolated 154 cells at 28 tetrode sites, against 95 from the best single wire at each; the single wires had often lumped two or more cells into one cluster.

A Neuropixels shank has rows of sites 20 µm apart, and on probes that dense most neurons’ spikes land on 5 to 50 channels, far more dimensions than anyone can cut by hand. Kilosort (Marius Pachitariu and colleagues, 2016) models the recording as a sum of templates, one per candidate neuron, placed at its spike times, and fits it on a graphics processor by matching pursuit: find the template that best explains the data, subtract it, look again. Scoring templates against the data is in effect matched filtering, the optimal linear way to detect a known waveform, and the subtraction is what separates two neurons firing together, which a threshold detector takes for one event.

Ensembles of single neurons

What sorting gave neuroscience is the ensemble of single neurons, recorded at once, each with its own spike train. Matthew Wilson and Bruce McNaughton had it at scale in 1993: with tetrodes they recorded 73 to 148 hippocampal neurons at a time and predicted the rat’s movement through its environment from them. A year later they found that cells which had fired together when the rat was in particular places fired together more in the sleep that followed — a result about pairs, which depends on each spike having gone to the right cell. Sorting with Kilosort and curating by hand, Nicholas Steinmetz and colleagues recorded about 30,000 neurons in 42 regions of the mouse brain.

Not every question about a population needs the sort. An intracortical brain–computer interface can run on unsorted threshold crossings: Fraser and colleagues’ monkeys steered a cursor that way about as well as with sorted spikes. Trautmann and colleagues found population dynamics, and the conclusions drawn from them, much the same either way. What sorting adds is identity: one cell’s tuning, one pair’s synchrony, the same neuron next week.

What it cannot check

In vivo there is no answer key. The direct check is a second electrode inside one of the recorded cells: with a glass pipette inside a neuron and a tetrode outside it, Kenneth Harris and colleagues found error rates for manual cluster cutting of typically 0–30%, depending on the cell, its neighbours and the experience of the operator; a semi-automatic method brought them down to 0–8%. Otherwise there are simulations, hybrid data with known spikes pasted in, and indirect tests, among them the argument Adrian and Zotterman made: a cluster with two spikes closer than a refractory period cannot be one cell alone. And the sorters disagree. In Alessio Buccino and colleagues’ comparison, six sorters run on one 15-minute Neuropixels recording reported 187 to 628 units each; of 2,031 units in all, all six agreed on 33.

The errors are not neutral, and each kind leaves its own mark. A missed spike lowers a cell’s apparent rate. A merge lends one cell another’s spikes, which distorts its receptive field (Hill, Mehta and Kleinfeld, 2011). A split can turn a bursting cell, whose later spikes are smaller, into two apparent cells, one firing just after the other. Sorting by waveform alone biases tuning curves unless every spike is classified correctly (Ventura, 2009), and the misses gather where they matter. Overlapping spikes are the hardest to resolve, so two cells on one electrode lose exactly their coincident spikes, and their cross-correlation dips at zero lag, where synchrony would show (Bar-Gad and colleagues, 2001). In simulations, single-channel sorting levelled off at 8 to 10 correctly identified neurons when up to 20 were present, and slow-firing cells were twice as likely to be missed (Pedreira and colleagues, 2012).

And chronic probes move. The brain shifts relative to the shank, mostly along it, and a neuron’s spikes shrink and vanish from the sites that held its template, so that, uncorrected, its rate tracks the probe’s position. When Steinmetz and colleagues moved a Neuropixels 2.0 probe in a 50 µm triangle wave, drift correction — resampling the data across sites, as in image registration — cut the average absolute correlation between neurons’ rates and the probe’s position from 0.22 to 0.07, against 0.04 by chance. In chronic recordings a version of it let them follow 93% of well-isolated units across sessions up to 16 days apart, judged by each unit’s responses to a set of natural images. It needs dense, regular sites: Kilosort’s authors note that tetrodes, and microelectrode arrays with sites more than 40 µm apart such as the Utah array, cannot be drift-corrected this way.

Origins & further reading

  1. G. L. Gerstein & W. A. Clark, 1964. Simultaneous Studies of Firing Patterns in Several Neurons. Science. paper · doi
  2. Bruce L. McNaughton et al., 1983. The stereotrode: A new technique for simultaneous isolation of several single units in the central nervous system from multiple unit records. Journal of Neuroscience Methods. paper · doi
  3. Matthew A. Wilson & Bruce L. McNaughton, 1993. Dynamics of the Hippocampal Ensemble Code for Space. Science. paper · doi
  4. Matthew A. Wilson & Bruce L. McNaughton, 1994. Reactivation of Hippocampal Ensemble Memories During Sleep. Science. paper · doi
  5. C. M. Gray et al., 1995. Tetrodes markedly improve the reliability and yield of multiple single-unit isolation from multi-unit recordings in cat striate cortex. Journal of Neuroscience Methods. paper · doi
  6. Kenneth D. Harris et al., 2000. Accuracy of Tetrode Spike Separation as Determined by Simultaneous Intracellular and Extracellular Measurements. Journal of Neurophysiology. paper · doi
  7. Darrell A. Henze et al., 2000. Intracellular Features Predicted by Extracellular Recordings in the Hippocampus In Vivo. Journal of Neurophysiology. paper · doi
  8. Izhar Bar-Gad et al., 2001. Failure in identification of overlapping spikes from multiple neuron activity causes artificial correlations. Journal of Neuroscience Methods. paper · doi
  9. R. Quian Quiroga et al., 2004. Unsupervised Spike Detection and Sorting with Wavelets and Superparamagnetic Clustering. Neural Computation. paper · doi
  10. Valérie Ventura, 2009. Traditional waveform based spike sorting yields biased rate code estimates. Proceedings of the National Academy of Sciences. paper · doi
  11. George W. Fraser et al., 2009. Control of a brain–computer interface without spike sorting. Journal of Neural Engineering. paper · doi
  12. D. N. Hill et al., 2011. Quality Metrics to Accompany Spike Sorting of Extracellular Signals. Journal of Neuroscience. paper · doi
  13. Carlos Pedreira et al., 2012. How many neurons can we see with current spike sorting algorithms?. Journal of Neuroscience Methods. paper · doi
  14. Marius Pachitariu et al., 2016. Fast and accurate spike sorting of high-channel count probes with KiloSort. Advances in Neural Information Processing Systems 29 (NIPS 2016). paper
  15. Nicholas A. Steinmetz et al., 2019. Distributed coding of choice, action and engagement across the mouse brain. Nature. paper · doi
  16. Eric M. Trautmann et al., 2019. Accurate Estimation of Neural Population Dynamics without Spike Sorting. Neuron. paper · doi
  17. Alessio P. Buccino et al., 2020. SpikeInterface, a unified framework for spike sorting. eLife. paper · doi
  18. Nicholas A. Steinmetz et al., 2021. Neuropixels 2.0: A miniaturized high-density probe for stable, long-term brain recordings. Science. paper · doi
  19. Marius Pachitariu et al., 2024. Spike sorting with Kilosort4. Nature Methods. paper · doi
  20. E. D. Adrian & Yngve Zotterman, 1926. The impulses produced by sensory nerve-endings: Part 2. The Journal of Physiology. paper · doi
  21. Rodrigo Quian Quiroga, 2007. Spike sorting. Scholarpedia. web · doi

Concepts

Related

Updated October 4, 2026