TY - JOUR
T1 - Latent belief reinforcement learning for online motor imagery classification
AU - Luo, Huan
AU - Lu, Na
AU - Wang, Xiaopeng
AU - Niu, Xu
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/11
Y1 - 2026/11
N2 - Motor Imagery-based Brain–Computer Interfaces (MI-BCIs) using Electroencephalography (EEG) face critical challenges in achieving low-latency, high-accuracy online classification. Traditional sliding fixed-length window methods fail to adapt to inter-trial EEG signal variability, either truncating informative segments or incorporating irrelevant data. To address the limitation, we propose Point Voting Network–Latent Belief Reinforcement Learning (PVN-LBRL), a novel framework that formulates online dynamic window MI classification as a Partially Observable Markov Decision Process (POMDP). The framework integrates: a Point Voting Network encoder for EEG feature extraction, a Gated Recurrent Unit-based latent dynamic model for belief updates, a belief decoder for explicit belief estimation, and a Q-network for adaptive halting decisions. By reformulating the action set from {“wait” vs. labels} to {“wait” vs. “halt”}, the early halting dilemma in RL exploration is addressed, increasing the initial waiting probability. The introduction of PVN enhances extracting discriminative patterns from noisy EEG slices. The latent belief updates enables effective evidence accumulation. This PVN-LBRL framework integrates the feature representation strength of deep learning with the sequential decision optimization strength of RL. Extensive experiments show that PVN-LBRL achieves a state-of-the-art ITR with a short prediction length, reaching 96.35 bits/min on the BCI C IV-2a with only a 40-sample window.
AB - Motor Imagery-based Brain–Computer Interfaces (MI-BCIs) using Electroencephalography (EEG) face critical challenges in achieving low-latency, high-accuracy online classification. Traditional sliding fixed-length window methods fail to adapt to inter-trial EEG signal variability, either truncating informative segments or incorporating irrelevant data. To address the limitation, we propose Point Voting Network–Latent Belief Reinforcement Learning (PVN-LBRL), a novel framework that formulates online dynamic window MI classification as a Partially Observable Markov Decision Process (POMDP). The framework integrates: a Point Voting Network encoder for EEG feature extraction, a Gated Recurrent Unit-based latent dynamic model for belief updates, a belief decoder for explicit belief estimation, and a Q-network for adaptive halting decisions. By reformulating the action set from {“wait” vs. labels} to {“wait” vs. “halt”}, the early halting dilemma in RL exploration is addressed, increasing the initial waiting probability. The introduction of PVN enhances extracting discriminative patterns from noisy EEG slices. The latent belief updates enables effective evidence accumulation. This PVN-LBRL framework integrates the feature representation strength of deep learning with the sequential decision optimization strength of RL. Extensive experiments show that PVN-LBRL achieves a state-of-the-art ITR with a short prediction length, reaching 96.35 bits/min on the BCI C IV-2a with only a 40-sample window.
KW - Brain-computer interface
KW - Motor imagery
KW - Online classification
UR - https://www.scopus.com/pages/publications/105034031605
U2 - 10.1016/j.patcog.2026.113503
DO - 10.1016/j.patcog.2026.113503
M3 - 文章
AN - SCOPUS:105034031605
SN - 0031-3203
VL - 179
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 113503
ER -