Cochlear implants
An electrode array in the cochlea driven by a filter bank, which restores speech understanding using a couple of dozen channels where the ear has thousands.
The cochlea is tonotopic: position along it maps to frequency. So an electrode array threaded into it can, in principle, stimulate the auditory nerve in a frequency-ordered way. Split the incoming sound into bands, take the envelope of each band, and use it to modulate pulse trains on the corresponding electrode.
This is the most successful neural prosthesis by a wide margin — around a million users by 2022, most of whom understand speech well enough to use a telephone.
From one wire to many
The first auditory implant was a transformer. In Paris in February 1957 the electrophysiologist André Djourno and the otolaryngologist Charles Eyriès buried a small induction coil in the temporalis muscle of a deaf patient, wired it to a few millimetres of exposed eighth cranial nerve, and drove it from outside the head — a secondary winding under the skin, where in TMS the secondary is the tissue itself. The patient could sense environmental sounds but not understand speech, and the device failed within months. Stimulating an auditory nerve was not new; the Swedish neurosurgeon Lundberg had done it during an operation in 1950. Leaving a stimulator attached was.
News of it reached William House in Los Angeles, who in 1961, with the neurosurgeon John Doyle and his brother Jim, an electronics engineer, put a gold wire into the cochleas of two patients. Developed with the electrical engineer Jack Urban, House’s single-channel implant gave awareness of sound and an aid to lipreading but not speech. An NIH-commissioned evaluation of all thirteen implant users then in the United States, eleven of them House’s patients, reported in 1977 that none could understand speech through their prostheses. Made by 3M, House’s device became the first cochlear implant the FDA approved, in November 1984.
Multi-electrode implants were tried at Stanford in 1964, in San Francisco in 1974 and in France in 1976. Two groups working independently set out to make multichannel stimulation carry speech without lipreading: Ingeborg and Erwin Hochmair in Vienna, whose first patient was implanted in December 1977, and Graeme Clark’s team in Melbourne, whose first followed in August 1978. Clark’s design became the Nucleus implant, in 1985 the first multichannel device the FDA approved; the Hochmairs went on to found MED-EL.
Why the interleaving mattered
Several early multichannel implants stimulated electrodes simultaneously, and performance was disappointing. The reason is physical: current from one electrode spreads through the conductive fluid of the cochlea, so simultaneous stimulation on neighbouring contacts sums in the tissue unpredictably and smears the frequency separation the array was built to provide.
Continuous interleaved sampling fixed it by never stimulating two electrodes at once — pulses are staggered in time so their fields do not superpose. The improvement in speech recognition was substantial, and it came from the processor rather than from better electrodes. Blake Wilson’s group at the Research Triangle Institute tested CIS in 1989–90 on users of the Utah-designed Ineraid implant, whose electrodes were wired straight out to a plug in the skin, so a laboratory processor could drive the same electrodes as the clinical one. All seven subjects had been chosen for doing well with the clinical processor’s simultaneous analogue stimulation; all seven scored higher with CIS.
The schedule was not the only change. Wilson’s own account names two other departures from earlier processors: CIS extracted no speech features, sending every band’s envelope and leaving the brain to decide what mattered, and it pulsed far faster, around a thousand pulses per second per electrode or more. Each envelope is smoothed only to 200–400 Hz, so the voice’s periodicity survives, then compressed from up to about 100 dB of acoustic range into the roughly 10 dB that electrical hearing offers. Later comparisons against simultaneous analogue stimulation came out closer than the 1991 result; the interleaving was one fix among several, not the whole story.
Wilson’s NIH-funded work was placed in the public domain, and by 2008 every implant in wide clinical use offered CIS. In 2013 Clark, Ingeborg Hochmair and Wilson shared the Lasker–DeBakey Clinical Medical Research Award for the modern cochlear implant; two of the three are electrical engineers.
What it says about coding
The result is startling from an information point of view. The healthy cochlea has some 3,500 inner hair cells and roughly 30,000 auditory nerve fibres, with exquisite frequency selectivity and phase locking. An implant has around 12 to 22 electrodes, little or no phase information, and crude spatial selectivity — and speech comes through.
That makes every implant an experiment as well as a treatment. A CIS processor is a channel vocoder — the telephone-engineering scheme Homer Dudley described at Bell Labs in 1939 — with pulse trains on electrodes for carriers, and the experiment is how few channels the nerve can be given before speech stops getting through. Shannon and colleagues ran the acoustic version in 1995: replace each band of speech with noise carrying only that band’s envelope, and listeners with normal hearing understood simple sentences with three bands. In Friesen and colleagues’ measurements six years later, implant users improved as electrodes were added up to about seven or eight and then levelled off, while normal-hearing listeners given the same processing in noise kept improving to at least 20 channels.
Put together, what the auditory nerve actually needs to convey speech in quiet is short: a handful of independent places along the tonotopic map, each carrying an envelope that follows the voice up to a few hundred hertz, and a brain allowed months to learn the new code. Wilson’s own summary is that a surprisingly sparse representation can be enough once it clears a threshold — one that single-channel devices cleared only for a few exceptional patients.
So speech must be extremely redundant, and most of what the intact ear transmits is not required for it. That is an efficient-coding argument arrived at by prosthesis rather than by theory, and it also explains the honest failure mode: music and speech in noise, which depend on the fine spectral and temporal detail the implant discards, remain much harder for implant users than quiet speech.
Where it stalls
More electrodes have not bought more channels. In Wilson and Dorman’s 2008 review, arrays in wide clinical use carried 12 to 22 electrodes, yet no user tested had shown more than about eight effective channels with a real-time processor, against roughly 28 independent auditory filters across the speech range in normal hearing. Interleaving removed the summation of simultaneous fields, not the spread of each one: the contacts sit in conductive perilymph, some distance from the spiral ganglion cells they are meant to reach, so each excites a broad population of fibres and neighbouring populations overlap.
The failure mode has numbers. In one group of average users, casually spoken sentences scored about 70% in quiet and 27% at a +5 dB signal-to-noise ratio, common in workplaces and classrooms. For most users pitch stops rising with pulse rate above about 300 Hz, and the same group, choosing among five familiar melodies stripped of rhythm, scored 33% where chance was 20%. Outcomes on identical hardware run from the floor to the ceiling, and by Zeng’s reckoning at the millionth implant, speech recognition in quiet with one implant has not improved since the early 1990s — the basis of his call to redesign the stimulating interface itself.
The interface is also where the power goes. The implant has no battery; power and data cross the skin by induction, as in Djourno’s coil, and in 2008-era designs the link was about 40% efficient, delivering 20–40 mW through 4–10 mm of skin. The current sources need enough compliance voltage for whatever impedance each electrode presents, and series capacitors on every contact block net DC — the charge-balance rule that also governs deep brain stimulation.
Moving the contacts closer to the nerve cuts the current and the crosstalk together, which helps the energy budget and the channel count at once. All three major manufacturers built arrays that hug the modiolus in the late 1990s for that reason. It sharpened the stimulation in some cases, but Wilson and Dorman’s judgement was that a large gain in independent channels may well need a fundamentally different electrode, placement or mode of stimulation.
Origins & further reading
- A. Djourno et al., 1957. De l'excitation électrique du nerf cochléaire chez l'homme, par induction à distance, à l'aide d'un micro-bobinage inclus à demeure. Comptes rendus des séances de la Société de biologie et de ses filiales. paper
- Blake S. Wilson et al., 1991. Better speech recognition with cochlear implants. Nature. paper · doi
- Robert V. Shannon et al., 1995. Speech Recognition with Primarily Temporal Cues. Science. paper · doi
- Lendra M. Friesen et al., 2001. Speech recognition in noise as a function of the number of spectral channels: Comparison of acoustic hearing and cochlear implants. The Journal of the Acoustical Society of America. paper · doi
- Blake S. Wilson & Michael F. Dorman, 2008. Cochlear implants: A remarkable past and a brilliant future. Hearing Research. paper · doi
- Fan-Gang Zeng et al., 2008. Cochlear Implants: System Design, Integration, and Evaluation. IEEE Reviews in Biomedical Engineering. paper · doi
- Albert Mudry & Mara Mills, 2013. The Early History of the Cochlear Implant: A Retrospective. JAMA Otolaryngology–Head & Neck Surgery. paper · doi
- Blake S. Wilson, 2013. Toward better representations of sound with cochlear implants. Nature Medicine. paper · doi
- Fan-Gang Zeng, 2022. Celebrating the one millionth cochlear implant. JASA Express Letters. paper · doi
Concepts
Related