重建脑机接口语音解码器:词错误率从 50.0% 降到 23.5%
Rebuilding a brain-to-text decoder: from 50% to 23.5% word errors
作者基于 Willett et al.(Nature 2023)公开的语音 BCI 神经数据重建解码器,在 880 句测试集上把词错误率从 50.0% 降到 23.5%,CPU 训练约 5 小时。
People who can no longer speak can still try to, and the motor cortex activity from that attempt can be decoded into text. Willett et al. (Nature 2023) published the neural data behind their speech BCI, so I rebuilt the decoder from it. The model trained on CPUs for about five hours, and the word error rate on 880 test sentences went from 50.0% to 23.5% over four changes (23.0% on a held-out half at the end).
This is an offline reanalysis of recorded data. Here is what each change did and what did not help.
Data
The Dryad release (competitionData, 3.67 GB, CC0) holds 8,800 training and 880 test sentences from one participant with ALS, recorded on 24 days. Following the baseline, I used area 6v only: threshold crossings and spike-band power on 128 electrodes, 256 features per 20 ms bin, z-scored per recording block.
Targets are phonemes, 39 plus silence, from the first CMU dictionary pronunciation. Words missing from the dictionary go through g2p_en.
Model
The model follows the PyTorch baseline (cffan/neural_seq_decoder):
x = F.conv1d(x.transpose(1, 2), self.smooth, padding="same", groups=x.shape[2]).transpose(1, 2) # Gaussian smoothing
x = torch.einsum("btd,bdk->btk", x, self.day_w[day]) + self.day_b[day] # per-day input layer
x = F.softsign(x)
x = self.unfold(x.transpose(1, 2).unsqueeze(3)).transpose(1, 2) # 32 bins, stride 4
h, _ = self.gru(x) # bidirectional GRU
return self.out(h).log_softmax(-1) # 40 classes + CTC blank
The per-day input layer absorbs day-to-day drift in what the electrodes pick up. I shrank the GRU from 5 layers and 1,024 units to 3 layers and 512 units (37.8 million parameters) and grouped sentences by length to cut padding. On a c7i.4xlarge with 8 threads that took 1.5 s per batch, against 11.7 s for the baseline size with random batches. 10,000 batches finished in 5 hours 12 minutes.
Greedy phoneme error rate on the test set: 22.2%.
Step 1: nearest words, 50.0%
Splitting the greedy phonemes at silences and replacing each group with the closest word from the training vocabulary gives 50.0% word errors. One wrong phoneme changes the word, and a missed silence merges two.
Step 2: let an LLM write, 34.6%
I gave Claude Sonnet 4.6 (Amazon Bedrock, temperature 0) the top five phoneme strings from a CTC prefix beam search and asked for the sentence. Word errors dropped to 39.2%. The results were uneven, though. With a single candidate, the model rebuilt "do you hear the sleigh bells ringing" perfectly from "dew you he the stable shells using". It also turned the almost-correct "quite a you movie are base off of that" into "why do you think they are based off of that".
When the phonemes are bad, the LLM writes a plausible new sentence. A cheap guard helped. If the LLM's word count differs from the decoder's silence-delimited word count by more than about 11%, keep the no-LLM output. I chose the threshold on one half of the test set and applied it to the other half, which gave 34.6%. Decoder confidence (mean max frame probability, beam score) made a weaker switch, at 38.7 to 39.1%.
Step 3: decode words, let the LLM pick, 26.0%
The guard treats a symptom. The fix was to stop the LLM from writing. I decoded words directly from the CTC output with a lexicon and a 3-gram, using torchaudio's flashlight decoder, and gave the LLM the 10 best sequences to choose from by number:
dec = ctc_decoder(lexicon="lexicon.txt", tokens="tokens.txt", lm="lm.arpa", nbest=10,
beam_size=100, lm_weight=1.5, word_score=-1.0, unk_score=float("-inf"),
blank_token="-", sil_token="SIL", unk_word="<unk>")
Each lexicon entry ends with SIL, because the training targets put silence after every word. The 3-gram came from the 8,800 training sentences only, written as an ARPA file with absolute discounting and Katz-style backoff. LM weight and word score were picked by 2-fold cross-validation on the test set.
- First candidate, no LLM: 29.7%
- LLM picks from the top 10: 26.0%
Two traps on the way:
- With the full CMU dictionary (about 130,000 words) the error rate jumped to around 90%. My LM gave
<unk>about 7% of the probability mass, so every rare dictionary word scored as well as an unknown word and won often. Restricting the lexicon to the LM's 6,569 words fixed it - The decoder's lexicon trie keeps at most 6 words per pronunciation. I write the lexicon in LM frequency order so the rare homophones are the ones dropped
Step 4: a bigger language model, 23.5%
| Language model | First candidate | LLM picks from 10 |
|---|---|---|
| Training sentences only | 29.7% | 26.0% |
| LibriSpeech 3-gram, pruned (openslr SLR11) | 29.8% | not run |
| Tatoeba English (about 2.04 million sentences) + training text counted twice, singletons dropped | 25.7% | 23.5% |
LibriSpeech's LM comes from books and did not help, probably because the test sentences are conversational. Tatoeba's short sentences match better. I removed the 39 Tatoeba sentences that were identical to test sentences before building the LM. With it, 297 of 880 sentences came out exactly right, and the oracle (best of the 10 candidates, chosen with the answer) is 19.4%.
Examples where the LLM picked a better candidate than the decoder's first:
- they close sometime after eat → they close sometime after eight
- i had the bake done on it → i had the brakes done on it
- solving the variables in the occasion → solving the variables in the equation
It swapped a correct first candidate for a wrong one in 3 sentences.
What did not help
By now I had made many choices while looking at the test set, so I split it: even sentences for development, odd sentences for evaluation. I compared methods on development only, then ran two of them once each on evaluation: the Sonnet baseline and Opus, the best on development.
| Method | Dev (440) | Eval (440) |
|---|---|---|
| No LLM | 26.1% | 25.4% |
| Sonnet 4.6, top 10 | 23.8% | 23.0% |
| Sonnet 4.6, top 20 | 23.9% | — |
| Sonnet 4.6, top 10 + phoneme string as a hint | 25.8% | — |
| Opus 4.6, top 10 | 23.4% | 23.1% |
Opus won on development but tied on evaluation (bootstrap 95% interval of the difference: −18 to +11 words), so I report the Sonnet baseline, 23.0%, as the held-out number. Twenty candidates raised the oracle from 19.6% to 18.7% on development, but the model could not find the extra correct ones. With the phoneme hint, Sonnet explained instead of answering in 401 of 440 calls, and the picks got worse.
The ceiling is the candidate list. Better candidates come from the language model and the acoustic model, not from a smarter selector.
Limits
- The model is smaller than the baseline. Training PER was 4.9% against 22.2% on the test set, so generalization is the first thing to fix
- I did not use the official large n-gram, which needs around 60 GB of memory to build the decoder graph
- Most method choices used cross-validation on the test set. Only the final dev/eval split is a clean held-out evaluation
- The paper's 23.8% comes from a different model, language model, vocabulary and evaluation, so it is not a direct comparison
References
- Willett et al., Nature 2023: https://www.nature.com/articles/s41586-023-06377-x
- Dataset (Dryad, CC0): https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czpq
- PyTorch baseline: https://github.com/cffan/neural_seq_decoder
- Card et al., NEJM 2024: https://www.nejm.org/doi/full/10.1056/NEJMoa2314132
- LibriSpeech LM (openslr SLR11): https://www.openslr.org/11/
- Tatoeba: https://tatoeba.org (English sentences, CC BY 2.0 FR)
来源:Google AI:DEV 作者专属(RSS) · dev.to
