EDBT 2026 Demo / reviewers in the wild / expert
Yuanqing Li 0001
dblp:51/2525-1
· DBLP profile ↗
94ranked-venue papers
19as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 13 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subdomain Adaption Method for Cross-Domain Emotion Recognition and Clinical Consciousness DetectionabstractEmotion recognition and consciousness detection are critical in neuroscience. Patients with disorder of consciousness (DOC) often cannot express feelings or respond definitively to stimuli due to impaired functions, posing challenges for accurate annotation and model transfer from healthy individuals to patients with DOC. Existing domain adaptation methods struggle to capture inter-subdomain relationships, resulting in suboptimal cross-domain learning performance. This study proposes a novel Category Maximum Mean Discrepancy (CMMD) subdomain adaptation method for EEG-based emotion recognition and consciousness detection. The CMMD framework aligns subdomain distributions of the same category in target and source domains, capturing subtle EEG features in emotional states and improving cross-domain classification. On the emotion recognition benchmark datasets SEED and SEED-IV, our method outperforms existing baselines, achieving 85.67% and 74.44% accuracy. In clinical experiments on 25 DOC patients, our research explored within-group dynamics among 15 MCS and 10 UWS patients via EEG-based emotion recognition. Experiments revealed higher average emotion recognition accuracy in MCS patients (66.21%) than UWS patients (56.99%), with a positive correlation to CRS-R scores (p<0.01). Our research demonstrated that emotion recognition accuracy in all MCS patients and five UWS patients surpassed chance level. All 15 MCS patients had an average recognition accuracy of 66.21%. Although the accuracy of 5 out of 10 UWS patients ranged from 60.42% to 68.75%, barely exceeding the 59.5-60% significance threshold of chance level, this provides objective neurophysiological evidence for potential residual consciousness, suggesting physicians should avoid overlooking subtle signs of consciousness in UWS patients. Results in clinical application highlight our method's potential in understanding DOC patient subgroups and enhancing consciousness detection. Jiahui Pan 0003, Zhipeng He 0001, Junbiao Zhu, Rongming Liang, Zerong Chen, Qiuyou Xie, Jingcong Li 0001, Yuanqing Li 0001 |
IEEE Internet Things J. | 8 |
| 2025 | Test-Time Learning for Large Language ModelsabstractWhile Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts. In this paper, we propose a Test-Time Learning (TTL) paradigm for LLMs, namely TLM, which dynamically adapts LLMs to target domains using only unlabeled test data during testing. Specifically, we first provide empirical evidence and theoretical insights to reveal that more accurate predictions from LLMs can be achieved by minimizing the input perplexity of the unlabeled test data. Based on this insight, we formulate the Test-Time Learning process of LLMs as input perplexity minimization, enabling self-supervised enhancement of LLM performance. Furthermore, we observe that high-perplexity samples tend to be more informative for model optimization. Accordingly, we introduce a Sample Efficient Learning Strategy that actively selects and emphasizes these high-perplexity samples for test-time updates. Lastly, to mitigate catastrophic forgetting and ensure adaptation stability, we adopt Low-Rank Adaptation (LoRA) instead of full-parameter optimization, which allows lightweight model updates while preserving more original knowledge from the model. We introduce the AdaptEval benchmark for TTL and demonstrate through experiments that TLM improves performance by at least 20% compared to original LLMs on domain knowledge adaptation. Jinwu Hu, Zitian Zhang, Xutao Wen, Chao Shuai, Wei Luo 0006, Bin Xiao 0002, Yuanqing Li 0001, Mingkui Tan |
ICML | 8 |
| 2025 | Curse of High Dimensionality Issue in Transformer for Long Context ModelingabstractTransformer-based large language models (LLMs) excel in natural language processing tasks by capturing long-range dependencies through self-attention mechanisms. However, long-context modeling faces significant computational inefficiencies due to redundant attention computations: while attention weights are often sparse, all tokens consume equal computational resources. In this paper, we reformulate traditional probabilistic sequence modeling as a supervised learning task, enabling the separation of relevant and irrelevant tokens and providing a clearer understanding of redundancy. Based on this reformulation, we theoretically analyze attention sparsity, revealing that only a few tokens significantly contribute to predictions. Building on this, we formulate attention optimization as a linear coding problem and propose a group coding strategy, theoretically showing its ability to improve robustness against random noise and enhance learning efficiency. Motivated by this, we propose Dynamic Group Attention (DGA), which leverages the group coding to explicitly reduce redundancy by aggregating less important tokens during attention computation. Empirical results show that our DGA significantly reduces computational costs while maintaining competitive performance. Shuhai Zhang, Zeng You, Yaofo Chen, Zhiquan Wen, Qianyue Wang, Zhijie Qiu, Yuanqing Li 0001, Mingkui Tan |
ICML | 7 |
| 2025 | Decoding Musical Neural Activity in Patients With Disorders of Consciousness Through Self-Supervised Contrastive Domain GeneralizationabstractIdentifying the brain responses of patients with disorders of consciousness (DOCs), which include comas, vegetative states (VSs, also called unresponsive wakefulness syndrome (UWS)) and minimally conscious states (MCSs), based on electroencephalography (EEG) has important clinical diagnosis implications. However, due to impaired motor and cognitive abilities, patients with DOCs may not be able to express their feelings and their brain responses to different stimuli, making it difficult to correctly label data. EEG classification algorithms trained with these data cannot make reliable classifications and predictions for clinical diagnosis purposes. To identify the brain responses produced for different types of stimuli in patients with DOCs, we proposed a self-supervised contrastive domain generalization framework (SSCDG) for cross-subject EEG classification. The model was first trained with healthy-subject EEG data induced by different stimuli to learn their corresponding unsupervised representations. Then, we used these representations to train a classifier to predict the emotional states of patients with DOCs under the corresponding stimuli. SSCDG was first evaluated on the SEED dataset, and it achieved an accuracy of 87.6%, which was 1.1% higher than that of the state-of-the-art (SOTA) approaches. Moreover, the SSCDG method was utilized to categorize EEG data acquired from seventeen DOC patients, including eleven in a UWS state and six in an MCS state, with seven patients demonstrating notable accuracy in three-class EEG classification tasks. The SSCDG results indicated that the seven patients with DOCs may have shown classifiable EEG responses to the presented stimuli. Honghua Cai, Jiahui Pan 0003, Qiuyi Xiao, Jiarui Jin, Yuanqing Li 0001, Qiuyou Xie |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | EEG-Based Cross-Subject Emotion Recognition Using Sparse Bayesian Learning With Enhanced Covariance AlignmentabstractEEG (Electroencephalography)-based emotion recognition has emerged as a crucial area of research due to its potential applications in mental health, brain-computer interfaces (BCIs), and affective computing. However, the inherent variability in EEG signals across individuals, coupled with limited dataset sizes, significantly hinders the development of robust and generalizable emotion recognition models. To overcome these challenges, we propose the Sparse Bayesian Learning with Enhanced Covariance Alignment (SBLECA) algorithm. SBLECA formulates cross-subject emotion recognition as an end-to-end decoding problem, integrating spatiotemporal filtering and classification within a sparse Bayesian learning (SBL) framework. Crucially, SBLECA incorporates a novel covariance alignment technique to mitigate inter-subject variability in EEG patterns. Rigorous evaluations on two publicly available emotion datasets demonstrate that SBLECA consistently outperforms state-of-the-art methods. Furthermore, SBLECA offers valuable insights into the neural correlates of emotion through interpretable visualizations of learned spatial and temporal filters. SBLECA holds promise as a valuable EEG decoding tool to advance the development and translation of neurotechnologies and biomarkers for brain disorders. Feifei Qi, Weichen Huang, Yuanqing Li 0001, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | EmotionMIL: An End-to-End Multiple Instance Learning Framework for Emotion Recognition From EEG SignalsabstractEmotion recognition from EEG signals offers significant advantages in affective computing, as EEG more accurately reflects internal emotional states than other modalities, such as facial expressions or peripheral physiological signals. Modeling and capturing subtle affective changes over time is crucial for real-world applications to achieve better human-computer interaction. However, training such models usually requires segment-level emotion labels, which are costly and may not be feasible. Assigning the overall label to all EEG segments within a trial can lead to inaccurate model training and degraded performance, as emotions evolve continuously. This highlights the need for models capable of learning from trial-wise emotion labels while capturing temporal dynamics of emotional responses within each segment because trial-wise post-stimulus labels are more accessible. To this end, we propose EmotionMIL, an end-to-end EEG-based emotion recognition framework that leverages recent advances in deep multiple instance learning (MIL). This framework enables robust emotion recognition from weakly labeled EEG signals and identifies the most prominent emotional responses. EmotionMIL captures the temporal dynamics of emotions using a retentive self-attention mechanism, which adaptively assigns weights to EEG segments based on their relevance in predicting the overall emotion label. A pseudo-bag augmentation strategy is also introduced to enhance the model's generalization ability by generating additional pseudo-bags from the original ones. Evaluated on three benchmark datasets—DEAP, DREAMER, and SEED—EmotionMIL outperforms state-of-the-art non-MIL and MIL models in both subject-dependent and subject-independent tasks, achieving superior accuracy and F1-score. Ablation study further validates the model design, while visualization results demonstrate that EmotionMIL effectively identifies both spatial EEG patterns and temporal emotional dynamics. These findings underscore EmotionMIL's potential for robust, interpretable emotion recognition, paving the way for real-world applications in emotion-aware systems. The code is available athttps://github.com/yuty2009/emotionmil. Feifei Qi, Lingli Wang, Yanbin He, Jingang Yu, Wei Wu 0022, Zhu Liang Yu, Yuanqing Li 0001, Zhenghui Gu, Tianyou Yu |
IEEE Trans. Affect. Comput. | 8 |
| 2025 | A Multimodal Consistency-Based Self-Supervised Contrastive Learning Framework for Automated Sleep Staging in Patients With Disorders of ConsciousnessabstractSleep is a fundamental human activity, and automated sleep staging holds considerable investigational potential. Despite numerous deep learning methods proposed for sleep staging that exhibit notable performance, several challenges remain unresolved, including inadequate representation and generalization capabilities, limitations in multimodal feature extraction, the scarcity of labeled data, and the restricted practical application for patients with disorder of consciousness (DOC). This paper proposes MultiConsSleepNet, a multimodal consistency-based sleep staging network. This network comprises a unimodal feature extractor and a multimodal consistency feature extractor, aiming to explore universal representations of electroencephalograms (EEGs) and electrooculograms (EOGs) and extract the consistency of intra- and intermodal features. Additionally, self-supervised contrastive learning strategies are designed for unimodal and multimodal consistency learning to address the current situation in clinical practice where it is difficult to obtain high-quality labeled data but has a huge amount of unlabeled data. It can effectively alleviate the model's dependence on labeled data, and improve the model's generalizability for effective migration to DOC patients. Experimental results on three publicly available datasets demonstrate that MultiConsSleepNet achieves state-of-the-art performance in sleep staging with limited labeled data and effectively utilizes unlabeled data, enhancing its practical applicability. Furthermore, the proposed model yields promising results on a self-collected DOC dataset, offering a novel perspective for sleep staging research in patients with DOC. Jiahui Pan 0003, Yangzuyi Yu, Wanxin Wei, Shuyu Chen 0002, Heyi Zheng, Yanbin He, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Few-Shot Learning for Annotation-Efficient Nucleus Instance SegmentationabstractNucleus instance segmentation from histopathology images suffers from the extremely laborious and expert-dependent annotation of nucleus instances. As a promising solution to this task, annotation-efficient deep learning paradigms have recently attracted much research interest, such as weakly-/semi-supervised learning, generative adversarial learning, etc. In this paper, we propose to formulate annotation-efficient nucleus instance segmentation from the perspective of few-shot learning (FSL). Our work was motivated by that, with the prosperity of computational pathology, an increasing number of fully-annotated datasets are publicly accessible, and we hope to leverage these external datasets to assist nucleus instance segmentation on the target dataset which only has very limited annotation. To achieve this goal, we adopt the meta-learning based FSL paradigm, which however has to be tailored in two substantial aspects before adapting to our task. First, since the novel classes may be inconsistent with those of the external dataset, we extend the basic definition of few-shot instance segmentation (FSIS) to generalized few-shot instance segmentation (GFSIS). Second, to cope with the intrinsic challenges of nucleus segmentation, including touching between adjacent cells, cellular heterogeneity, etc., we further introduce a structural guidance mechanism into the GFSIS network, finally leading to a unified Structurally-Guided Generalized Few-Shot Instance Segmentation (SGFSIS) framework. Extensive experiments on a couple of publicly accessible datasets demonstrate that, SGFSIS can outperform other annotation-efficient learning baselines, including semi-supervised learning, simple transfer learning, etc., with comparable performance to fully supervised learning with around 10% annotations. Zihao Wu 0004, Jie Yang 0002, Danyi Li, Yuan Gao 0015, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001, Jin-Gang Yu |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean DiscrepancyabstractLarge language models (LLMs) such as ChatGPT have exhibited remarkable performance in generating human-like texts. However, machine-generated texts (MGTs) may carry critical risks, such as plagiarism issues and hallucination information. Therefore, it is very urgent and important to detect MGTs in many situations. Unfortunately, it is challenging to distinguish MGTs and human-written texts because the distributional discrepancy between them is often very subtle due to the remarkable performance of LLMS. In this paper, we seek to exploit \textit{maximum mean discrepancy} (MMD) to address this issue in the sense that MMD can well identify distributional discrepancies. However, directly training a detector with MMD using diverse MGTs will incur a significantly increased variance of MMD since MGTs may contain \textit{multiple text populations} due to various LLMs. This will severely impair MMD's ability to measure the difference between two samples. To tackle this, we propose a novel \textit{multi-population} aware optimization method for MMD called MMD-MP, which can \textit{avoid variance increases} and thus improve the stability to measure the distributional discrepancy. Relying on MMD-MP, we develop two methods for paragraph-based and sentence-based detection, respectively. Extensive experiments on various LLMs, \eg, GPT2 and ChatGPT, show superior detection performance of our MMD-MP. Shuhai Zhang, Yiliao Song, Yuanqing Li 0001, Bo Han 0003, Mingkui Tan |
ICLR | 4 |
| 2024 | MFASleepNet: Multi-view fusion attention-based deep neural network for automatic sleep stagingabstractSleep staging is important for both the assessment of sleep quality and the diagnosis of sleep-related disorders. Although previous studies have attempted to automatically classify sleep stages and achieved high classification performance, most automatic staging algorithms currently use only the time or frequency domain information of the data, resulting in limited feature representation capability. To address the above problems, we propose a multi-view fusion attention-based deep neural network (MFASleepNet) to classify sleep stages using single-channel EEG signals. MFASleepNet consists of a multi-scale CNN, a multi-view fusion attention module, a global external attention module and a conditional random field. Specifically, the multi-scale CNN can extract features from different frequency bands of the signal. Then the multi-view fusion attention module can fuse multiple views from different time-frequency spaces to improve the feature representation capability. Subsequently, the global external attention module first captures the global feature representation of each sample through the global context block, and then captures the common feature representation of all samples through the external attention mechanism, and outputs a preliminary sleep stage result. Finally, the conditional random field is used to learn sleep stage transition rules to improve classification performance. We evaluate the performance of our model on the Sleep-EDF dataset, and MFASleepNet achieves an accuracy of 87.2% and an F1 score of 81.1% on the Fpz-Cz channel. The experimental results show that our MFASleepNet can optimise the performance of sleep staging and achieve better staging performance than state-of-the-art methods. Zhoujie Hou, Jiahui Pan 0003, Yuanqing Li 0001 |
IJCNN | 3 |
| 2024 | Cross-Device Collaborative Test-Time AdaptationabstractIn this paper, we propose test-time Collaborative Lifelong Adaptation (CoLA), which is a general paradigm that can be incorporated with existing advanced TTA methods to boost the adaptation performance and efficiency in a multi-device collaborative manner. Specifically, we maintain and store a set of device-shared _domain knowledge vectors_, which accumulates the knowledge learned from all devices during their lifelong adaptation process. Based on this, CoLA conducts two collaboration strategies for devices with different computational resources and latency demands. 1) Knowledge reprogramming learning strategy jointly learns new domain-specific model parameters and a reweighting term to reprogram existing shared domain knowledge vectors, termed adaptation on _principal agents_. 2) Similarity-based knowledge aggregation strategy solely aggregates the knowledge stored in shared domain vectors according to domain similarities in an optimization-free manner, termed adaptation on _follower agents_. Experiments verify that CoLA is simple but effective, which boosts the efficiency of TTA and demonstrates remarkable superiority in collaborative, lifelong, and single-domain TTA scenarios, e.g., on follower agents, we enhance accuracy by over 30\% on ImageNet-C while maintaining nearly the same efficiency as standard inference. The source code is available at https://github.com/Cascol-Chen/COLA. Shuaicheng Niu, Deyu Chen, Shuhai Zhang, Yuanqing Li 0001, Mingkui Tan |
NeurIPS | 6 |
| 2024 | EPMF: Efficient Perception-Aware Multi-Sensor Fusion for 3D Semantic SegmentationabstractWe study multi-sensor fusion for 3D semantic segmentation that is important to scene understanding for many applications, such as autonomous driving and robotics. For example, for autonomous cars equipped with RGB cameras and LiDAR, it is crucial to fuse complementary information from different sensors for robust and accurate segmentation. Existing fusion-based methods, however, may not achieve promising performance due to the vast difference between the two modalities. In this work, we investigate a collaborative fusion scheme called perception-aware multi-sensor fusion (PMF) to effectively exploit perceptual information from two modalities, namely, appearance information from RGB images and spatio-depth information from point clouds. To this end, we first project point clouds to the camera coordinate using perspective projection. In this way, we can process both inputs from LiDAR and cameras in 2D space while preventing the information loss of RGB images. Then, we propose a two-stream network that consists of a LiDAR stream and a camera stream to extract features from the two modalities, separately. The extracted features are fused by effective residual-based fusion modules. Moreover, we introduce additional perception-aware losses to measure the perceptual difference between the two modalities. Last, we propose an improved version of PMF,i.e., EPMF, which is more efficient and effective by optimizing data pre-processing and network architecture under perspective projection. Specifically, we propose cross-modal alignment and cropping to obtain tight inputs and reduce unnecessary computational costs. We then explore more efficient contextual modules under perspective projection and fuse the LiDAR features into the camera stream to boost the performance of the two-stream network. Extensive experiments on benchmark data sets show the superiority of our method. For example, on nuScenes test set, our EPMF outperforms the state-of-the-art method,i.e., RangeFormer, by0.9%in mIoU. Compared to PMF, EPMF also achieves2.06× acceleration with2.0%improvement in mIoU. Our source code is available athttps://github.com/ICEORY/PMF. Mingkui Tan, Zhuangwei Zhuang, Sitao Chen, Kui Jia, Qicheng Wang, Yuanqing Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | An Affective Brain-Computer Interface Based on a Transfer Learning MethodabstractAn affective brain-computer interface (aBCI) can detect affective states based on brain signals and might assist people in improving their emotion regulation abilities. However, individual differences in emotional brain patterns make cross-subject emotion identification extremely challenging. Traditional supervised single-subject classification schemes require considerable calibration samples from new individuals to train subject-dependent models. Individuals are easily fatigued with long-term EEG collection processes, which may affect performance in subsequent online experiments. In this study, we propose a real-time aBCI system using domain-fusion-based multisource style transfer mapping (DF-MS-STM) to detect positive, neutral, and negative emotional states without the need for additional training sessions. Sixteen subjects participated in our online experiments to test the performance of our aBCI system and an average online prediction accuracy of 72.17$\pm$12.25% was obtained for three-class emotion recognition tasks in the last three experimental sessions. Our proposed algorithm significantly outperformed numerous baseline methods in terms of cross-subject emotion classification. In addition, we identified distinct brain patterns in response to different emotional stimuli based on the results of event-related spectral perturbation (ERSP) analyses. These neural patterns might provide new insights for emotional brain mechanistic studies and related aBCIs. Weichen Huang, Zijing Guan, Kendi Li, Yajun Zhou, Yuanqing Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | FBSTCNet: A Spatio-Temporal Convolutional Network Integrating Power and Connectivity Features for EEG-Based Emotion DecodingabstractElectroencephalography (EEG)-based emotion recognition plays a key role in the development of affective brain-computer interfaces (BCIs). However, emotions are complex and extracting salient EEG features underlying distinct emotional states is inherently limited by low signal-to-noise ratio (SNR) and low spatial resolution of practical EEG data, which is further compounded by the lack of effective spatio-temporal filter optimization approaches for generic EEG features. To address these challenges, this study proposes a set of neural networks termed the Filter-Bank Spatio-Temporal Convolutional Networks (FBSTCNets) for performing end-to-end multi-class emotion recognition via robust extraction of power and/or connectivity features from EEG. First, a filter bank is employed to construct a multiview spectral representation of EEG data. Next, a temporal convolutional layer, followed by a depth-wise spatial convolutional layer, performs spatio-temporal filtering, transforming EEG into latent signals with higher SNR. A feature extraction layer then extracts power and/or connectivity features from the latent signals. Finally, a fully connected layer with a cropped decoding strategy predicts the emotional state. Experimental results on two public emotion EEG datasets, SEED and SEED-IV, demonstrate that FBSTCNets outperform previous benchmark methods in decoding accuracy. Our approach provides a principled emotion decoding framework for designing high-performance spatio-temporal filtering networks tailored to specific EEG feature types. The FBSTCNet source code is available athttps://github.com/TimeSpacerRob/FBSTCNet. Weichen Huang, Yuanqing Li 0001, Wei Wu 0022 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Toward Compact and Robust Model Learning Under Dynamically Perturbed EnvironmentsabstractNetwork pruning has been widely studied to reduce the complexity of deep neural networks (DNNs) and hence speed up their inference. Unfortunately, most existing pruning methods ignore the changes in the model’s robustness before and after pruning, which makes pruned models vulnerable under dynamically perturbed environments (e.g., autonomous driving). Only a few works have explored the robustness of pruned models against adversarial attacks that significantly differ from perturbations in real-world scenarios. To bridge the gap between real-world applications and existing studies, in this work, we propose an adversarial pruning scheme, which automatically identifies and preserves robust channels to obtain robust pruned models that are suitable for practical deployment in dynamically perturbed environments. Specifically, to simulate real-world perturbations, we first employ multi-type adversarial attack samples and adversarial perturbation samples generated by an adversarial perturbation generator to create mixed noise samples. Then, we propose a plug-and-play feature scoring module and a novel contribution difference loss to evaluate the robustness of intermediate features dynamically. Next, to leverage robust intermediate features to identify robust channels, we have developed a simple but effective gating mechanism that evaluates the robustness of channels and preserves robust channels during training. Lastly, we compress the model in a layer-wise or block-wise manner. Compared to existing methods, our scheme enhances the robustness of the pruned model in a broader sense, making it better able to against dynamic perturbations in the real world. Extensive experimental results on well-known dataset benchmarks and popular network architectures demonstrate the effectiveness of our method. Hui Luo 0002, Zhuangwei Zhuang, Yuanqing Li 0001, Mingkui Tan, Cen Chen 0002, Jianlin Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | ST-SCGNN: A Spatio-Temporal Self-Constructing Graph Neural Network for Cross-Subject EEG-Based Emotion Recognition and Consciousness DetectionabstractIn this paper, a novel spatio-temporal self-constructing graph neural network (ST-SCGNN) is proposed for cross-subject emotion recognition and consciousness detection. For spatio-temporal feature generation, activation and connection pattern features are first extracted and then combined to leverage their complementary emotion-related information. Next, a self-constructing graph neural network with a spatio-temporal model is presented. Specifically, the graph structure of the neural network is dynamically updated by the self-constructing module of the input signal. Experiments based on the SEED and SEED-IV datasets showed that the model achieved average accuracies of 85.90% and 76.37%, respectively. Both values exceed the state-of-the-art metrics with the same protocol. In clinical besides, patients with disorders of consciousness (DOC) suffer severe brain injuries, and sufficient training data for EEG-based emotion recognition cannot be collected. Our proposed ST-SCGNN method for cross-subject emotion recognition was first attempted in training in ten healthy subjects and testing in eight patients with DOC. We found that two patients obtained accuracies significantly higher than chance level and showed similar neural patterns with healthy subjects. Covert consciousness and emotion-related abilities were thus demonstrated in these two patients. Our proposed ST-SCGNN for cross-subject emotion recognition could be a promising tool for consciousness detection in DOC patients. Jiahui Pan 0003, Rongming Liang, Zhipeng He 0001, Jingcong Li 0001, Yanbin He, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Electromagnetic Source Imaging via a Data-Synthesis-Based Convolutional Encoder-Decoder NetworkabstractElectromagnetic source imaging (ESI) requires solving a highly ill-posed inverse problem. To seek a unique solution, traditional ESI methods impose various forms of priors that may not accurately reflect the actual source properties, which may hinder their broad applications. To overcome this limitation, in this article, a novel data-synthesized spatiotemporally convolutional encoder-decoder network (DST-CedNet) method is proposed for ESI. The DST-CedNet recasts ESI as a machine learning problem, where discriminative learning and latent-space representations are integrated in a CedNet to learn a robust mapping from the measured electroencephalography/magnetoencephalography (E/MEG) signals to the brain activity. In particular, by incorporating prior knowledge regarding dynamical brain activities, a novel data synthesis strategy is devised to generate large-scale samples for effectively training CedNet. This stands in contrast to traditional ESI methods where the prior information is often enforced via constraints primarily aimed for mathematical convenience. Extensive numerical experiments as well as analysis of a real MEG and epilepsy EEG dataset demonstrate that the DST-CedNet outperforms several state-of-the-art ESI methods in robustly estimating source signals under a variety of source configurations. Gexin Huang, Ke Liu 0008, Jiawen Liang, Zhenghui Gu, Feifei Qi, Yuanqing Li 0001, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Prototypical multiple instance learning for predicting lymph node metastasis of breast cancer from whole-slide pathological images
Jin-Gang Yu, Zihao Wu 0004, Shule Deng, Yuanqing Li 0001, Caifeng Ou, Chunjiang He, Baiye Wang, Pusheng Zhang |
Medical Image Anal. | 5 |
| 2023 | Sparse Bayesian Learning for End-to-End EEG DecodingabstractDecoding brain activity from non-invasive electroencephalography (EEG) is crucial for brain-computer interfaces (BCIs) and the study of brain disorders. Notably, end-to-end EEG decoding has gained widespread popularity in recent years owing to the remarkable advances in deep learning research. However, many EEG studies suffer from limited sample sizes, making it difficult for existing deep learning models to effectively generalize to highly noisy EEG data. To address this fundamental limitation, this paper proposes a novel end-to-end EEG decoding algorithm that utilizes a low-rank weight matrix to encode both spatio-temporal filters and the classifier, all optimized under a principled sparse Bayesian learning (SBL) framework. Importantly, this SBL framework also enables us to learn hyperparameters that optimally penalize the model in a Bayesian fashion. The proposed decoding algorithm is systematically benchmarked on five motor imagery BCI EEG datasets ( N=192) and an emotion recognition EEG dataset ( N=45), in comparison with several contemporary algorithms, including end-to-end deep-learning-based EEG decoding algorithms. The classification results demonstrate that our algorithm significantly outperforms the competing algorithms while yielding neurophysiologically meaningful spatio-temporal patterns. Our algorithm therefore advances the state-of-the-art by providing a novel EEG-tailored machine learning tool for decoding brain activity. Feifei Qi, David P. Wipf, Tianyou Yu, Yuanqing Li 0001, Yu Zhang 0009, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Neurofeedback Training With an Electroencephalogram-Based Brain-Computer Interface Enhances Emotion RegulationabstractEmotion regulation plays a vital role in human beings daily lives by helping them deal with social problems and protects mental and physical health. However, objective evaluation of the efficacy of emotion regulation and assessment of the improvement in emotion regulation ability at the individual level remain challenging. In this study, we leveraged neurofeedback training to design a real-time EEG-based brain-computer interface (BCI) system for users to effectively regulate their emotions. Twenty healthy subjects performed 10 BCI-based neurofeedback training sessions to regulate their emotion towards a specific emotional state (positive, negative, or neutral), while their EEG signals were analyzed in real time via machine learning to predict their emotional states. The prediction results were presented as feedback on the screen to inform the subjects of their immediate emotional state, based on which the subjects could update their strategies for emotion regulation. The experimental results indicated that the subjects improved their ability to regulate these emotions through our BCI neurofeedback training. Further EEG-based spectrum analysis revealed how each emotional state was related to specific EEG patterns, which were progressively enhanced through long-term training. These results together suggested that long-term EEG-based neurofeedback training could be a promising tool for helping people with emotional or mental disorders. Weichen Huang, Wei Wu 0022, Molly V. Lucas, Haiyun Huang, Zhenfu Wen, Yuanqing Li 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Temporal Micro-Action Localization for Videofluoroscopic Swallowing StudyabstractVideofluoroscopic swallowing study (VFSS) visualizes the swallowing movement by using X-ray fluoroscopy, which is the most widely used method for dysphagia examination. To better facilitate swallowing assessment, the temporal parameter is one of the most important indicators. However, most information of that acquire is hand-crafted and elaborated, which is time-consuming and difficult to ensure objectivity and accuracy. In this article, we propose to formulate this task as a temporal action localization task and solve it using deep neural networks. However, the action of VFSS has the following characteristics such as small motion targets, small action amplitudes, large sample variances, short duration, and variations in duration. Furthermore, all existing methods often rely on daily behaviors, which makes locating and recognizing micro-actions more challenging. To address the above issues, we first collect and annotate the VFSS micro-action dataset, which includes 847 VFSS data from 71 subjects, due to the lack of benchmarks. We then introduce a coarse-to-fine mechanism to handle the short and repeated nature of micro-actions, which can significantly enhancing micro-action localization accuracy. Moreover, we propose a Variable-Size Window Generator method, which improves the model's characterization performance and addresses the issue of different action timings, leading to further improvements in localization accuracy. The results of our experiments demonstrate the superiority of our method, with significantly improved performance (46.10% vs. 37.70%). Xianghui Ruan, Zhuokun Chen, Zeng You, Yuanqing Li 0001, Zulin Dou, Mingkui Tan |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Bayesian Collaborative Learning for Whole-Slide Image ClassificationabstractWhole-slide image (WSI) classification is fundamental to computational pathology, which is challenging in extra-high resolution, expensive manual annotation, data heterogeneity, etc. Multiple instance learning (MIL) provides a promising way towards WSI classification, which nevertheless suffers from the memory bottleneck issue inherently, due to the gigapixel high resolution. To avoid this issue, the overwhelming majority of existing approaches have to decouple the feature encoder and the MIL aggregator in MIL networks, which may largely degrade the performance. Towards this end, this paper presents a Bayesian Collaborative Learning (BCL) framework to address the memory bottleneck issue with WSI classification. Our basic idea is to introduce an auxiliary patch classifier to interact with the target MIL classifier to be learned, so that the feature encoder and the MIL aggregator in the MIL classifier can be learned collaboratively while preventing the memory bottleneck issue. Such a collaborative learning procedure is formulated under a unified Bayesian probabilistic framework and a principled Expectation-Maximization algorithm is developed to infer the optimal model parameters iteratively. As an implementation of the E-step, an effective quality-aware pseudo labeling strategy is also suggested. The proposed BCL is extensively evaluated on three publicly available WSI datasets, i.e., CAMELYON16, TCGA-NSCLC and TCGA-RCC, achieving an AUC of 95.6%, 96.0% and 97.5% respectively, which consistently outperforms all the methods compared. Comprehensive analysis and discussion will also be presented for in-depth understanding of the method. To promote future work, our source code is released at: https://github.com/Zero-We/BCL. Jin-Gang Yu, Zihao Wu 0004, Shule Deng, Qihang Wu, Zhongtang Xiong, Tianyou Yu, Gui-Song Xia, Qingping Jiang, Yuanqing Li 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2022 | V2C: Visual Voice CloningabstractExisting Voice Cloning (VC) tasks aim to convert a para-graph text to a speech with desired voice specified by a ref-erence audio. This has significantly boosted the development of artificial speech applications. However, there also exist many scenarios that cannot be well reflected by these VC tasks, such as movie dubbing, which requires the speech to be with emotions consistent with the movie plots. To fill this gap, in this work we propose a new task named Vi-sual Voice Cloning (V2C), which seeks to convert a para-graph of text to a speech with both desired voice speci-fied by a reference audio and desired emotion specified by a reference video. To facilitate research in this field, we construct a dataset, V2C-Animation, and propose a strong baseline based on existing state-of-the-art (SoTA) VC techniques. Our dataset contains 10,217 animated movie clips covering a large variety of genres (e.g., Comedy, Fantasy) and emotions (e.g., happy, sad). We further design a set of evaluation metrics, named MCD-DTW-SL, which help eval-uate the similarity between ground-truth speeches and the synthesised ones. Extensive experimental results show that even SoTA VC methods cannot generate satisfying speeches for our V2C task. We hope the proposed new task together with the constructed dataset and evaluation metric will fa-cilitate the research in the field of voice cloning and broader vision-and-language community. Source code and dataset will be released in https://github.com/chenqi008/V2C. Qi Chen 0014, Mingkui Tan, Yuankai Qi, Jiaqiu Zhou, Yuanqing Li 0001, Qi Wu 0001 |
CVPR | 5 |
| 2022 | Robot control with multitasking of brain-computer interfaceabstractBrain-computer interfaces (BCI) have been extensively researched to assist people with motor paralysis in controlling external devices such as a robotic limb. However, most BCI systems required participants to focus on a single task, limiting their ability to generate other mental or physical activities. Therefore, people's performance of the BCI-based robotic control in multitasking was discussed, as eight healthy subjects performed motor-related tasks of motor imagery and two-handed balancing ball movement, while simultaneously performing visuospatial attention to asynchronously trigger “drinking” actions of a humanoid robot arm with accuracies of 90% and 87.5%, respectively. The online results indicate that the BCI-based robot control system developed for multi-task conditions has a high potential for human augmentation. Yajun Zhou, Zilin Lu, Yuanqing Li 0001 |
ICARCV | 3 |
| 2022 | Dynamic User Activity and Data Detection for Grant-Free NOMA via Weighted ℓ2, 1 MinimizationabstractGrant-free non-orthogonal multiple access (NOMA) has recently received wide attention for reducing signaling overhead and transmission latency in massive machine-type communications (mMTC). In grant-free NOMA systems, user activity and data (UAD) has to be detected, which is challenging in practice. As an emerging technique, compressive sensing (CS) shows great promise in solving this problem by exploiting the inherent sparsity nature of user activity. This paper proposes to use the weighted$\ell _{2, 1}$minimization (WL21M) to jointly detect UAD in realistic dynamic scenarios. At first, the average recoverability of the WL21M is analyzed. This analysis reveals the fact that the WL21M can improve the detection performance by means of an appropriate weighting and the incorporation of intrinsic temporal correlation. Motivated by the analysis, a collaborative hierarchical match pursuit (C-HiMP) algorithm is proposed for dynamic UAD detection. In the C-HiMP, a sequence of WL21M problems are solved in the subspaces spanned by all of the components in the hierarchical estimated support sets, where the weights are collaboratively updated by the solutions in previous time slots so that an attractive self-correction capacity is obtained. Simulation results demonstrate that the proposed C-HiMP can obtain significant performance improvements, in terms of detection accuracy and computational complexity, compared with several state-of-the-art CS-based detection algorithms. Jun Zhang 0026, Zhijing Yang, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2022 | Compressive Sensing-Based Power Allocation Optimization for Energy Harvesting IoT NodesabstractIn this paper, we address the problem of optimizing the power allocation at each time slot for energy-harvesting sensors in an IoT system, where each sensor transmits its observation via a coherent multiple access channel, and thus the observation vector received at the fusion center (FC) becomes a compressed version of the original observations. Our goal is to minimize the number of transmissions required by the high-fidelity reconstruction of the original observations in the FC. In this scenario, the power allocation, coupling with the channel effect, constitute an effective measurement matrix in compressive sensing. Because its performance decides the requirement on the number of transmissions, the goal can be achieved through constructing the measurement matrix as good as possible, or equivalently, solving the power optimization allocation problem, subject to the available energy constraints at each sensor. However, the sensors can only obtain unreliable and intermittent available energy, so that the traditional performance metric, i.e., mutual coherence (MC), of the measurement matrix cannot be directly used to guide the optimization, because an “equal-norm columns” assumption is implicitly required, but not satisfied in our scenario due to the available energy constraints. Moreover, this optimization problem is also non-convex. To overcome these obstacles, we first carry out a distortion analysis based on the generalized MC, which abandons the “equal-norm columns” assumption. The theoretical results indicate that the measurement matrix construction can be formulated as an optimization problem that not only minimizes the MC, but also minimizes the maximum and maximizes the minimum of the column norms of the effective matrix. We further transform this problem into a sequence of surrogate convex problems and iteratively find the solution. Numerical results show that the proposed framework improves the tradeoffs between reconstruction accuracy and the number of transmissions over various power allocation strategies. In some cases, where other strategies achieve a probability of exact recovery of below 0.7, the proposed framework can achieve a more than 0.9 probability. Jun Zhang 0026, Guangfei Xie, Guojun Han, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2021 | DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyabstractCategory-level 6D object pose and size estimation is to predict full pose configurations of rotation, translation, and size for object instances observed in single, arbitrary views of cluttered scenes. In this paper, we propose a new method of Dual Pose Network with refined learning of pose consistency for this task, shortened as DualPoseNet. DualPoseNet stacks two parallel pose decoders on top of a shared pose encoder, where the implicit decoder predicts object poses with a working mechanism different from that of the explicit one; they thus impose complementary supervision on the training of pose encoder. We construct the encoder based on spherical convolutions, and design a module of Spherical Fusion wherein for a better embedding of pose-sensitive features from the appearance and shape observations. Given no testing CAD models, it is the novel introduction of the implicit decoder that enables the refined pose prediction during testing, by enforcing the predicted pose consistency between the two decoders using a self-adaptive loss term. Thorough experiments on benchmarks of both category- and instance-level object pose datasets confirm efficacy of our designs. DualPoseNet outperforms existing methods with a large margin in the regime of high precision. Our code is released publicly at https://github.com/Gorilla-Lab-SCUT/DualPoseNet. Jiehong Lin, Zewei Wei, Zhihao Li 0002, Songcen Xu, Kui Jia, Yuanqing Li 0001 |
ICCV | 6 |
| 2021 | Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic Segmentationabstract3D LiDAR (light detection and ranging) semantic segmentation is important in scene understanding for many applications, such as auto-driving and robotics. For example, for autonomous cars equipped with RGB cameras and LiDAR, it is crucial to fuse complementary information from different sensors for robust and accurate segmentation. Existing fusion-based methods, however, may not achieve promising performance due to the vast difference between the two modalities. In this work, we investigate a collaborative fusion scheme called perception-aware multi-sensor fusion (PMF) to exploit perceptual information from two modalities, namely, appearance information from RGB images and spatio-depth information from point clouds. To this end, we first project point clouds to the camera coordinates to provide spatio-depth information for RGB images. Then, we propose a two-stream network to extract features from the two modalities, separately, and fuse the features by effective residual-based fusion modules. Moreover, we propose additional perception-aware losses to measure the perceptual difference between the two modalities. Extensive experiments on two benchmark data sets show the superiority of our method. For example, on nuScenes, our PMF outperforms the state-of-the-art method by 0.8% in mIoU. Zhuangwei Zhuang, Kui Jia, Qicheng Wang, Yuanqing Li 0001, Mingkui Tan |
ICCV | 5 |
| 2021 | Weighted Conditional Distribution Adaptation for Motor Imagery Classification
Fake Gu, Tianyou Yu, Yuanqing Li 0001 |
ICIG (1) | 5 |
| 2021 | Deep Unfolding With Weighted ℓ₂ Minimization for Compressive SensingabstractCompressive sensing (CS) aims to accurately reconstruct high-dimensional signals from a small number of measurements by exploiting signal sparsity and structural priors. However, signal priors utilized in existing CS reconstruction algorithms rely mainly on hand-crafted design, which often cannot offer the best sparsity-undersampling tradeoff because high-order structural priors of signals are hard to be captured in this manner. In this article, a new recovery guarantee of the unified CS reconstruction model-weighted ℓ1minimization (WL1M) is derived, which indicates universal priors could hardly lead to the optimal selection of the weights. Motivated by the analysis, we propose a deep unfolding network for the general WL1M model. The proposed deep unfolding-based WL1M (D-WL1M) integrates universal priors with learning capability so that all of the parameters, including the crucial weights, can be learned from training data. We demonstrate the proposed D-WL1M outperforms several state-of-the-art CS-based methods and deep learning-based methods by a large margin via the experiments on the Caltech-256 image data set. Jun Zhang 0026, Yuanqing Li 0001, Zhu Liang Yu, Zhenghui Gu, Yu Cheng 0010, Huoqing Gong |
IEEE Internet Things J. | 2 |
| 2021 | An EEG-Based Brain Computer Interface for Emotion Recognition and Its Application in Patients with Disorder of ConsciousnessabstractRecognizing human emotions based on electroencephalogram (EEG) signals has received a great deal of attentions. Most of the existing studies focused on offline analysis, and real-time emotion recognition using a brain computer interface (BCI) approach remains to be further investigated. In this paper, we proposed an EEG-based BCI system for emotion recognition. Specifically, two classes of video clips that represented positive and negative emotions were presented to the subjects one by one, while the EEG data were collected and processed simultaneously, and instant feedback was provided after each clip. Ten healthy subjects participated in the experiment and achieved a high average online accuracy of 91.5$\pm$6.34 percent. The experimental results demonstrated that the subjects emotions had been sufficiently evoked and efficiently recognized by our system. Clinically, patients with disorder of consciousness (DOC), such as coma, vegetative state, minimally conscious state and emergence minimally conscious state, suffer from motor impairment and generally cannot provide adequate emotion expressions. Consequently, doctors have difficulty in detecting the emotional states of these patients. Therefore, we applied our emotion recognition BCI system to patients with DOC. Eight DOC patients participated in our experiment, and three of them achieved significant online accuracy. The experimental results show that the proposed BCI system could be a promising tool to detect the emotional states of patients with DOC. Haiyun Huang, Qiuyou Xie, Jiahui Pan 0003, Yanbin He, Zhenfu Wen, Ronghao Yu, Yuanqing Li 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2021 | Spatiotemporal-Filtering-Based Channel Selection for Single-Trial EEG ClassificationabstractAchieving high classification performance in electroencephalogram (EEG)-based brain-computer interfaces (BCIs) often entails a large number of channels, which impedes their use in practical applications. Despite the previous efforts, it remains a challenge to determine the optimal subset of channels in a subject-specific manner without heavily compromising the classification performance. In this article, we propose a new method, called spatiotemporal-filtering-based channel selection (STECS), to automatically identify a designated number of discriminative channels by leveraging the spatiotemporal information of the EEG data. In STECS, the channel selection problem is cast under the framework of spatiotemporal filter optimization by incorporating a group sparsity constraints, and a computationally efficient algorithm is developed to solve the optimization problem. The performance of STECS is assessed on three motor imagery EEG datasets. Compared with state-of-the-art spatiotemporal filtering algorithms using full EEG channels, STECS yields comparable classification performance with only half of the channels. Moreover, STECS significantly outperforms the existing channel selection methods. These results suggest that this algorithm holds promise for simplifying BCI setups and facilitating practical utility. Feifei Qi, Wei Wu 0022, Zhu Liang Yu, Zhenghui Gu, Zhenfu Wen, Tianyou Yu, Yuanqing Li 0001 |
IEEE Trans. Cybern. | 7 |
| 2020 | FGN: Fully Guided Network for Few-Shot Instance SegmentationabstractFew-shot instance segmentation (FSIS) conjoins the few-shot learning paradigm with general instance segmentation, which provides a possible way of tackling instance segmentation in the lack of abundant labeled data for training. This paper presents a Fully Guided Network (FGN) for few-shot instance segmentation. FGN perceives FSIS as a guided model where a so-called support set is encoded and utilized to guide the predictions of a base instance segmentation network (i.e., Mask R-CNN), critical to which is the guidance mechanism. In this view, FGN introduces different guidance mechanisms into the various key components in Mask R-CNN, including Attention-Guided RPN, Relation-Guided Detector, and Attention-Guided FCN, in order to make full use of the guidance effect from the support set and adapt better to the inter-class generalization. Experiments on public datasets demonstrate that our proposed FGN can outperform the state-of-the-art methods. Zhibo Fan, Jin-Gang Yu, Jiarong Ou, Changxin Gao, Gui-Song Xia, Yuanqing Li 0001 |
CVPR | 7 |
| 2020 | Censoring-Aware Deep Ordinal Regression for Survival Prediction from Pathological Images
Lichao Xiao, Jin-Gang Yu, Jiarong Ou, Shule Deng, Zhenhua Yang, Yuanqing Li 0001 |
MICCAI (5) | 7 |
| 2020 | Imaging brain extended sources from EEG/MEG based on variation sparsity using automatic relevance determination
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 5 |
| 2020 | Hyperspectral Image Spectral-Spatial-Range Gabor FilteringabstractSpectral-spatial Gabor filtering, which is based on 3-D local harmonic analysis, has been a powerful spectral-spatial feature extraction tool for hyperspectral image (HSI) classification. However, existing spectral-spatial Gabor approaches are prone to oversmoothing, neglecting the existences of edges and negatively affecting the classification. In this article, we propose a new HSI Gabor filtering concept, called spectral-spatial-range Gabor filtering, which intends to restrain edge interference from disturbing local spectral-spatial harmonic components. Contributions and novelties of our work can be identified as follows: 1) an HSI filtering framework is created, which can accommodate various Gabor filtering procedures and hence offer the potential to guide the design of new Gabor filters; 2) following such a unified filtering framework and taking into consideration both local spectral-spatial harmonic characteristics and range domain variations, we develop a new concept of spectral-spatial-range Gabor filtering; and 3) utilizing this proposed Gabor prototype and elaborating mathematical derivations, we achieve a novel discriminative spectral-spatial-range Gabor filtering method, which can deal with discriminative local harmonics and edge interference simultaneously along the spectral-spatial-range domain, obtaining highly discriminative Gabor features while yielding linear computational complexity. Our novel method is evaluated on four real HSI data sets and achieves excellent performances. Lin He 0001, Chenying Liu 0001, Jun Li 0009, Yuanqing Li 0001, Shutao Li 0001, Zhu Liang Yu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Exemplar-Based Recursive Instance Segmentation With Application to Plant Image AnalysisabstractInstance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art. Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | Full-Spectrum-Knowledge-Aware Tensor Model for Energy-Resolved CT Iterative ReconstructionabstractEnergy-resolved computed tomography (ErCT) with a photon counting detector concurrently produces multiple CT images corresponding to different photon energy ranges. It has the potential to generate energy-dependent images with improved contrast-to-noise ratio and sufficient material-specific information. Since the number of detected photons in one energy bin in ErCT is smaller than that in conventional energy-integrating CT (EiCT), ErCT images are inherently more noisy than EiCT images, which leads to increased noise and bias in the subsequent material estimation. In this work, we first deeply analyze the intrinsic tensor properties of two-dimensional (2D) ErCT images acquired in different energy bins and then present a F ull- S pectrum-knowledge-aware Tensor analysis and processing (FSTensor) method for ErCT reconstruction to suppress noise-induced artifacts to obtain high-quality ErCT images and high-accuracy material images. The presented method is based on three considerations: (1) 2D ErCT images obtained in different energy bins can be treated as a 3-order tensor with three modes, i.e., width, height and energy bin, and a rich global correlation exists among the three modes, which can be characterized by tensor decomposition. (2) There is a locally piecewise smooth property in the 3-order ErCT images, and it can be captured by a tensor total variation regularization. (3) The images from the full spectrum are much better than the ErCT images with respect to noise variance and structural details and serve as external information to improve the reconstruction performance. We then develop an alternating direction method of multipliers algorithm to numerically solve the presented FSTensor method. We further utilize a genetic algorithm to tackle the parameter selection in ErCT reconstruction, instead of manually determining parameters. Simulation, preclinical and synthesized clinical ErCT results demonstrate that the presented FSTensor method leads to significant improvements over the filtered back-projection, robust principal component analysis, tensor-based dictionary learning and low-rank tensor decomposition with spatial-temporal total variation methods. Dong Zeng, Yongshuai Ge, Sui Li, Qi Xie 0002, Hao Zhang 0026, Zhaoying Bian, Qian Zhao 0002, Yuanqing Li 0001, Zongben Xu, Deyu Meng, Jianhua Ma 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2019 | A novel multi-step reinforcement learning method for solving reward hacking
Yinlong Yuan, Zhu Liang Yu, Zhenghui Gu, Xiaoyan Deng, Yuanqing Li 0001 |
Appl. Intell. | 5 |
| 2019 | MMAN: Multi-modality aggregation network for brain segmentation from MR images
Jingcong Li 0001, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 5 |
| 2019 | A novel multi-step Q-learning method to improve data efficiency for deep reinforcement learning
Yinlong Yuan, Zhu Liang Yu, Zhenghui Gu, Yao Yeboah, Wei Wu 0022, Xiaoyan Deng, Jingcong Li 0001, Yuanqing Li 0001 |
Knowl. Based Syst. | 8 |
| 2019 | Classification of symmetric positive definite matrices based on bilinear isometric Riemannian embedding
Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Pattern Recognit. | 4 |
| 2018 | A EOG-based switch and its application for "start/stop" control of a wheelchair
Yuanqing Li 0001, Shenghong He, Qiyun Huang, Zhenghui Gu, Zhu Liang Yu |
Neurocomputing | 1 |
| 2018 | Deep learning based on Batch Normalization for P300 signal detection
Mingfei Liu, Wei Wu 0022, Zhenghui Gu, Zhu Liang Yu, Feifei Qi, Yuanqing Li 0001 |
Neurocomputing | 6 |
| 2018 | Variation sparse source imaging based on conditional mean for electromagnetic extended sources
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001, Srikantan S. Nagarajan |
Neurocomputing | 5 |
| 2018 | Deterministic construction of sparse binary matrices via incremental integer optimization
Jun Zhang 0026, Zhu Liang Yu, Ling Cen, Zhenghui Gu, Zhiping Lin 0001, Yuanqing Li 0001 |
Inf. Sci. | 6 |
| 2018 | A Novel clustering method based on hybrid K-nearest-neighbor graph
Yikun Qin, Zhu Liang Yu, Chang-Dong Wang 0001, Zhenghui Gu, Yuanqing Li 0001 |
Pattern Recognit. | 5 |
| 2018 | A Novel Three-Dimensional P300 Speller Based on Stereo Visual StimuliabstractGoal: P300 spellers are among the most popular types of brain-computer interfaces (BCIs) and are extremely useful assistive devices that enable severely disabled patients to communicate. However, P300 speller performances should be further improved to translate laboratory designs into practical applications. We aimed to design a new speller paradigm that could evoke higher event-related potentials (ERPs) than traditional P300 spellers, thus improving the performance of BCI systems. Methods: We proposed a new P300 speller paradigm based on three-dimensional (3-D) stereo visual stimuli. In this paradigm, flashing buttons are presented in 3-D stereo form. We designed two experiments, one that tested a traditional two-dimensional (2-D) speller and another that tested the proposed 3-D speller. Twelve healthy volunteers participated in our experiments. We compared the ERPs elicited by the 2-D speller and the 3-D speller, and we also compared the classification accuracy, information transfer rate (ITR), and user workload between the two paradigms. Results: The 3-D P300 speller elicited higher amplitudes of P300 waveforms than the traditional 2-D P300 speller. The online experimental results showed that the classification accuracy and the ITR were significantly improved with the 3-D P300 speller. We also found that the user workload of the 3-D P300 speller was significantly lower than that of the 2-D P300 speller. Conclusion : The proposed 3-D P300 speller based on stereo visual stimuli outperformed a traditional 2-D P300 speller. This finding indicates that our 3-D paradigm offers a new method that will improve the performance of P300 BCI systems. Jun Qu, Fei Wang 0026, Zhenping Xia, Tianyou Yu, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 8 |
| 2017 | An Algorithm Combining Spatial Filtering and Temporal Down-Sampling with Applications to ERP Feature Extraction
Feifei Qi, Yuanqing Li 0001, Zhenfu Wen, Wei Wu 0022 |
ICONIP (2) | 2 |
| 2017 | An online semi-supervised P300 speller based on extreme learning machine
Zhenghui Gu, Zhu Liang Yu, Yuanqing Li 0001 |
Neurocomputing | 4 |
| 2017 | An MVPA method based on sparse representation for pattern localization in fMRI data analysis
Fangyi Wang, Yuanqing Li 0001, Zhenghui Gu |
Neurocomputing | 2 |
| 2017 | Discriminative Low-Rank Gabor Filtering for Spectral-Spatial Hyperspectral Image ClassificationabstractSpectral-spatial classification of remotely sensed hyperspectral images has attracted a lot of attention in recent years. Although Gabor filtering has been used for feature extraction from hyperspectral images, its capacity to extract relevant information from both the spectral and the spatial domains of the image has not been fully explored yet. In this paper, we present a new discriminative low-rank Gabor filtering (DLRGF) method for spectral-spatial hyperspectral image classification. A main innovation of the proposed approach is that our implementation is accomplished by decomposing the standard 3-D spectral-spatial Gabor filter into eight subfilters, which correspond to different combinations of low-pass and bandpass single-rank filters. Then, we show that only one of the subfilters (i.e., the one that performs low-pass spatial filtering and bandpass spectral filtering) is actually appropriate to extract suitable features based on the characteristics of hyperspectral images. This allows us to perform spectral-spatial classification in a highly discriminative and computationally efficient way, by significantly decreasing the computational complexity (from cubic to linear order) compared with the 3-D spectral-spatial Gabor filter. In order to theoretically prove the discriminative ability of the selected subfilter, we derive an overall classification risk bound to evaluate the discriminating abilities of the features provided by the different subfilters. Our experimental results, conducted using different hyperspectral images, indicate that the proposed DLRGF method exhibits significant improvements in terms of classification accuracy and computational performance when compared with the 3-D spectral-spatial Gabor filter and other state-of-the-art spectral-spatial classification methods. Lin He 0001, Jun Li 0009, Antonio Plaza, Yuanqing Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2016 | A new neural-dynamic control method of position and angular stabilization for autonomous quadrotor UAVsabstractQuadrotor unmanned aerial vehicles (UAVs) have been widely used or have great potential applications in military, entertainment, postal delivery, agriculture for working aloft, photographing, etc. The position and angular stabilization of quadrotor UAVs is very significant and it is a challenging work because of the nonlinear dynamic behavior. In this paper, a neural dynamic method based control system is designed and investigated by combination of Zhang dynamics and gradient dynamic (ZD-GD) methods. Quadrotor UAVs equipped with the ZD-GD controllers can realize position and angular stabilization autonomously. Computer simulation results substantiate the efficiency and accuracy of the proposed neural dynamic method based ZD-GD controllers. Besides, the performance of the controllers can be remarkably advanced. Zhijun Zhang 0003, Jianli Yu, Yuanqing Li 0001, Xiaoyan Zhang 0002 |
FUZZ-IEEE | 3 |
| 2016 | An EEG-Based brain-computer interface for emotion recognitionabstractIn this paper, an EEG-based brain-computer interface (BCI) system used for emotion recognition is proposed to detect two basic emotional states (happiness and sadness). Selection of frequency bands plays a vital role in distinguishing brain patterns associated with emotions. This paper explores a new method to select suitable subject-specific frequency bands instead of using fixed frequency bands for the emotion recognition. Common spatial pattern and support vector machine were employed to classify two emotional states. Two experiments involving six subjects were conducted to validate our method and BCI system. An average online accuracy of 74.17% for two classes was achieved. The data analysis results demonstrated that the proposed method based on subject-specific frequency bands outperformed the method based on the fixed frequency bands in terms of accuracy. Jiahui Pan 0003, Yuanqing Li 0001, Jun Wang 0002 |
IJCNN | 2 |
| 2016 | Sufficient conditions for sparse recovery by weighted ℓ1-constrained quadratic programmingabstractIn this paper, we study the performance guarantee of weighted ℓ1-constrained quadratic programming in recovering the support of a sparse signal from a few linear measurements. A new sufficient condition for the success of weighted ℓ1-constrained quadratic programming is derived. Further, we demonstrate its applications in two typical weighted ℓ1models: Firstly, a theoretical result on the modified-BPDN can be generated directly, which indicates that if a partial support information is available and more than half of the information is accurate, then modified-BPDN relaxes the sufficient condition of BPDN ensuring the support recovery for any sparse signals. Secondly, the proposed condition offers a criterion to compute the optimal weights of weighted ℓ1minimization for nonuniform sparse models. Some simulations are carried out to validate its effectiveness. Jun Zhang 0026, Yuanqing Li 0001, Zhu Liang Yu, Zhenghui Gu |
IJCNN | 2 |
| 2016 | Multimodal BCIs: Target Detection, Multidimensional Control, and Awareness Evaluation in Patients With Disorder of ConsciousnessabstractDespite rapid advances in the study of brain–computer interfaces (BCIs) in recent decades, two fundamental challenges, namely, improvement of target detection performance and multidimensional control, continue to be major barriers for further development and applications. In this paper, we review the recent progress in multimodal BCIs (also called hybrid BCIs), which may provide potential solutions for addressing these challenges. In particular, improved target detection can be achieved by developing multimodal BCIs that utilize multiple brain patterns, multimodal signals, or multisensory stimuli. Furthermore, multidimensional object control can be accomplished by generating multiple control signals from different brain patterns or signal modalities. Here, we highlight several representative multimodal BCI systems by analyzing their paradigm designs, detection/control methods, and experimental results. To demonstrate their practicality, we report several initial clinical applications of these multimodal BCI systems, including awareness evaluation/detection in patients with disorder of consciousness (DOC). As an evolving research area, the study of multimodal BCIs is increasingly requiring more synergetic efforts from multiple disciplines for the exploration of the underlying brain mechanisms, the design of new effective paradigms and means of neurofeedback, and the expansion of the clinical applications of these systems. Yuanqing Li 0001, Jiahui Pan 0003, Jinyi Long, Tianyou Yu, Fei Wang 0026, Zhu Liang Yu, Wei Wu 0022 |
Proc. IEEE | 1 |
| 2015 | STRAPS: A Fully Data-Driven Spatio-Temporally Regularized Algorithm for M/EEG Patch Source ImagingabstractFor M/EEG-based distributed source imaging, it has been established that the L2-norm-based methods are effective in imaging spatially extended sources, whereas the L1-norm-based methods are more suited for estimating focal and sparse sources. However, when the spatial extents of the sources are unknown a priori, the rationale for using either type of methods is not adequately supported. Bayesian inference by exploiting the spatio-temporal information of the patch sources holds great promise as a tool for adaptive source imaging, but both computational and methodological limitations remain to be overcome. In this paper, based on state-space modeling of the M/EEG data, we propose a fully data-driven and scalable algorithm, termed STRAPS, for M/EEG patch source imaging on high-resolution cortices. Unlike the existing algorithms, the recursive penalized least squares (RPLS) procedure is employed to efficiently estimate the source activities as opposed to the computationally demanding Kalman filtering/smoothing. Furthermore, the coefficients of the multivariate autoregressive (MVAR) model characterizing the spatial-temporal dynamics of the patch sources are estimated in a principled manner via empirical Bayes. Extensive numerical experiments demonstrate STRAPS's excellent performance in the estimation of locations, spatial extents and amplitudes of the patch sources with varying spatial extents. Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001 |
Int. J. Neural Syst. | 5 |
| 2015 | Analysis of fMRI data based on sparsity of source components in signal dictionary
Bao Feng, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 4 |
| 2015 | Probabilistic Common Spatial Patterns for Multichannel EEG AnalysisabstractCommon spatial patterns (CSP) is a well-known spatial filtering algorithm for multichannel electroencephalogram (EEG) analysis. In this paper, we cast the CSP algorithm in a probabilistic modeling setting. Specifically, probabilistic CSP (P-CSP) is proposed as a generic EEG spatio-temporal modeling framework that subsumes the CSP and regularized CSP algorithms. The proposed framework enables us to resolve the overfitting issue of CSP in a principled manner. We derive statistical inference algorithms that can alleviate the issue of local optima. In particular, an efficient algorithm based on eigendecomposition is developed for maximum a posteriori (MAP) estimation in the case of isotropic noise. For more general cases, a variational algorithm is developed for group-wise sparse Bayesian learning for the P-CSP model and for automatically determining the model size. The two proposed algorithms are validated on a simulated data set. Their practical efficacy is also demonstrated by successful applications to single-trial classifications of three motor imagery EEG data sets and by the spatio-temporal pattern analysis of one EEG data set recorded in a Stroop color naming task. Wei Wu 0022, Zhe Chen 0001, Xiaorong Gao, Yuanqing Li 0001, Emery N. Brown, Shangkai Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Spectral-Spatial Classification of Hyperspectral Images via Spatial Translation-Invariant Wavelet-Based Sparse RepresentationabstractFor hyperspectral image (HSI) classification, it is challenging to adopt the methodology of sparse-representation-based classification. In this paper, we first propose an l1-minimization-based spectral-spatial classification method for HSIs via a spatial translation-invariant wavelet (STIW)-based sparse representation (STIW-SR), wherein both the spectrum dictionary and the analyzed signal are formed with STIW features. Due to the capability of a STIW to reduce both the observation noise and the spatial nonstationarity while maintaining the ideal spectra, which is proved with our signal-interference-noise spectrum model involved, it is expected that the pixels in the same class congregate in a lower dimensional subspace, and the separations among class-specific subspaces are enhanced, thus yielding a highly discriminative sparse representation. Then, we develop an approach to evaluate the sparsity recoverability of an l1-minimization on HSIs in a probabilistic framework. This approach takes into account not only the recovery probability under the given support length of the l0-norm solution but also the apriori probability of the support length; consequently, it overcomes the inability of traditional mutual/cumulative coherence conditions to address high-coherence HSIs. This paper reveals that the higher sparsity recoverability of a STIW-SR leads to its higher classification accuracy and that the increasing coherence does not necessarily lead to a reduced sparsity recovery probability, and this paper verifies the connection between l0and l1-minimizations on HSIs. Experimental results from realworld HSIs suggest that our classification method significantly outperforms several representative spectral-spatial classifiers and support vector machines. Lin He 0001, Yuanqing Li 0001, Xiaoxin Li 0001, Wei Wu 0022 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Energy-Efficient ECG Compression on Wireless Biosensors via Minimal Coherence Sensing and Weighted ℓ1 Minimization ReconstructionabstractLow energy consumption is crucial for body area networks (BANs). In BAN-enabled ECG monitoring, the continuous monitoring entails the need of the sensor nodes to transmit a huge data to the sink node, which leads to excessive energy consumption. To reduce airtime over energy-hungry wireless links, this paper presents an energy-efficient compressed sensing (CS)-based approach for on-node ECG compression. At first, an algorithm called minimal mutual coherence pursuit is proposed to construct sparse binary measurement matrices, which can be used to encode the ECG signals with superior performance and extremely low complexity. Second, in order to minimize the data rate required for faithful reconstruction, a weighted ℓ1 minimization model is derived by exploring the multisource prior knowledge in wavelet domain. Experimental results on MIT-BIH arrhythmia database reveals that the proposed approach can obtain higher compression ratio than the state-of-the-art CS-based methods. Together with its low encoding complexity, our approach can achieve significant energy saving in both encoding process and wireless transmission. Jun Zhang 0026, Zhenghui Gu, Zhu Liang Yu, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2015 | RSTFC: A Novel Algorithm for Spatio-Temporal Filtering and Classification of Single-Trial EEGabstractLearning optimal spatio-temporal filters is a key to feature extraction for single-trial electroencephalogram (EEG) classification. The challenges are controlling the complexity of the learning algorithm so as to alleviate the curse of dimensionality and attaining computational efficiency to facilitate online applications, e.g., brain-computer interfaces (BCIs). To tackle these barriers, this paper presents a novel algorithm, termed regularized spatio-temporal filtering and classification (RSTFC), for single-trial EEG classification. RSTFC consists of two modules. In the feature extraction module, an l2 -regularized algorithm is developed for supervised spatio-temporal filtering of the EEG signals. Unlike the existing supervised spatio-temporal filter optimization algorithms, the developed algorithm can simultaneously optimize spatial and high-order temporal filters in an eigenvalue decomposition framework and thus be implemented highly efficiently. In the classification module, a convex optimization algorithm for sparse Fisher linear discriminant analysis is proposed for simultaneous feature selection and classification of the typically high-dimensional spatio-temporally filtered signals. The effectiveness of RSTFC is demonstrated by comparing it with several state-of-the-arts methods on three brain-computer interface (BCI) competition data sets collected from 17 subjects. Results indicate that RSTFC yields significantly higher classification accuracies than the competing methods. This paper also discusses the advantage of optimizing channel-specific temporal filters over optimizing a temporal filter common to all channels. Feifei Qi, Yuanqing Li 0001, Wei Wu 0022 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Neural-Dynamic-Method-Based Dual-Arm CMG Scheme With Time-Varying Constraints Applied to Humanoid RobotsabstractWe propose a dual-arm cyclic-motion-generation (DACMG) scheme by a neural-dynamic method, which can remedy the joint-angle-drift phenomenon of a humanoid robot. In particular, according to a neural-dynamic design method, first, a cyclic-motion performance index is exploited and applied. This cyclic-motion performance index is then integrated into a quadratic programming (QP)-type scheme with time-varying constraints, called the time-varying-constrained DACMG (TVC-DACMG) scheme. The scheme includes the kinematic motion equations of two arms and the time-varying joint limits. The scheme can not only generate the cyclic motion of two arms for a humanoid robot but also control the arms to move to the desired position. In addition, the scheme considers the physical limit avoidance. To solve the QP problem, a recurrent neural network is presented and used to obtain the optimal solutions. Computer simulations and physical experiments demonstrate the effectiveness and the accuracy of such a TVC-DACMG scheme and the neural network solver. Zhijun Zhang 0003, Zhijun Li 0001, Yunong Zhang, Yamei Luo, Yuanqing Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | Explorer based on brain computer interfaceabstractIn recent years, various applications which apply the hybrid brain computer interfaces (BCIs) have been studied. In this paper, we present a hybrid BCI system to operate the explorer with P300 and motor imagery, which is mainly composed of a BCI mouse, a BCI speller and an explorer. Through this system, the user can access to his computer and manipulate (open, close, copy, paste, delete) files such as documents, pictures, music, movies and so on. The system has been tested with 5 subjects, and the experimental results show that the explorer can be successfully operated according to subjects' intention with only a small number of mistakes. Lijuan Bai, Tianyou Yu, Yuanqing Li 0001 |
IJCNN | 3 |
| 2014 | Coordinated control of an intelligentwheelchair based on a brain-computer interface and speech recognitionabstractAn intelligent wheelchair is devised, which is controlled by a coordinated mechanism based on a brain-computer interface (BCI) and speech recognition. By performing appropriate activities, users can navigate the wheelchair with four steering behaviors (start, stop, turn left, and turn right). Five healthy subjects participated in an indoor experiment. The results demonstrate the efficiency of the coordinated control mechanism with satisfactory path and time optimality ratios, and show that speech recognition is a fast and accurate supplement for BCI-based control systems. The proposed intelligent wheelchair is especially suitable for patients suffering from paralysis (especially those with aphasia) who can learn to pronounce only a single sound (e.g., ‘ah’). Hongtao Wang 0001, Yuanqing Li 0001, Tian-You Yu |
J. Zhejiang Univ. Sci. C | 2 |
| 2013 | Channel selection by Rayleigh coefficient maximization based genetic algorithm for classifying single-trial motor imagery EEG
Lin He 0001, Youpan Hu, Yuanqing Li 0001, Daoli Li |
Neurocomputing | 3 |
| 2012 | Efficient wavelet networks for function learning based on adaptive wavelet neuron selectionabstractIn this study, a novel four-layer architecture of wavelet network is proposed for function learning. Compared to conventional three-layer wavelet networks, the proposed one exploits adaptive wavelet neuron selection technique according to input information, so that the widespread structural redundancy is avoided. Meanwhile, it controls the scale of problem solution. Based on the proposed architecture, two wavelet networks including single-wavelet neural network and multi-wavelet neural network are built and verified for function learning. The experimental results demonstrate that our models are remarkably superior to some of the well-established three-layer wavelet networks including Zhang's model and Pati's model in terms of both speed and accuracy. Compared with Huang's real-time neural network, the proposed models have significantly better accuracy with basically similar speed. Jun Zhang 0026, Yuanqing Li 0001 |
IET Signal Process. | 3 |
| 2011 | Simple Models for Synaptic Information Integration
Danke Zhang, Yuwei Cui, Yuanqing Li 0001, Si Wu 0001 |
ICONIP (3) | 3 |
| 2011 | A linear discriminant analysis method based on mutual information maximization
Haihong Zhang, Cuntai Guan, Yuanqing Li 0001 |
Pattern Recognit. | 3 |
| 2010 | Feature extraction with multiscale autoregression of multichannel time series for P300 speller BCIabstractP300 is one of the most studied components of event related potentials which reflects the responses of brain to events in the external environment. In this paper, we present a new method that utilizes multiresolution autoregression of multichannel time series (MAMTS) for feature extraction of P300 wave. First, it adopts multiresolution autoregression on dyadic tree to depict the characteristic of electroencephalogram (EEG) signal. Then the corresponding autoregression noise of multichannel time series is extracted as the feature. The experiment results verified the effectiveness of this new feature for P300 speller brain compute interface (BCI). Lin He 0001, Zhenghui Gu, Yuanqing Li 0001, Zhu Liang Yu |
ICASSP | 3 |
| 2010 | Robust adaptive beamformer with a large controlled mainlobeabstractMany advanced adaptive beamformers are robust against arbitrary array steering vector (ASV) mismatches within a presumed uncertainty set. Adaptive array tolerating significant steering direction error usually requires a large size of ASV uncertainty set. In such case, however, the output signal-to-interference-plus-noise ratios (SINRs) of robust methods degrade quickly with the increasing size of the uncertainty set. In this paper, we propose a new compact ASV uncertainty set which is modelled explicitly by the uncertainty on steering direction and the other arbitrary ASV errors. A robust adaptive beamformer is derived based on this new ASV uncertainty set. To eliminate the non-convex constraint on array magnitude response, we force the real part of array response to exceed unity regarding the ASVs within the uncertainty set. Furthermore, using the worst-case optimization technique, the resultant beamformer is formulated as a quadratic optimization problem with semi-infinite second-order cone (SOC) constraints. Numerical studies show that a large robust response region is easy to achieve and the resultant beamformer achieves high performance on SINR enhancement. Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001, Wee Ser, Meng Hwa Er |
ICASSP | 3 |
| 2010 | Robust response control with linear inequality matrix constraints for adaptive beamformerabstractA novel robust adaptive beamformer, with new robust constraints on array magnitude response was proposed by utilizing the autocorrelation sequence of array weight vector and the worst-case optimization technique. The proposed adaptive beamformer was formulated as a linear programming problem with second-order cone semi-infinite constraints, which can be eliminated by using the sampling technique. In this paper, we transform these semi-infinite second-order cone constraints into some norm constraints and linear matrix inequality (LMI) constraints. The advantage of this new formulation of the problem is that the sampling of the angles is avoided. The exact optimal result of the problem can be obtained instead of the approximated one provided by sampling technique. The resultant beamformer possesses superior robustness against arbitrary array imperfections and high performance on signal-to-interference-plus-noise ratio (SINR) enhancement even with a large controlled robust response region. Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001, Wee Ser, Meng Hwa Er |
ISCAS | 3 |
| 2010 | A Sparse Infrastructure of Wavelet Network for Nonparametric Regression
Jun Zhang 0026, Zhenghui Gu, Yuanqing Li 0001 |
ISNN (1) | 3 |
| 2010 | Two conditions for equivalence of 0-norm solution and 1-norm solution in sparse representationabstractIn sparse representation, two important sparse solutions, the 0-norm and 1-norm solutions, have been receiving much of attention. The 0-norm solution is the sparsest, however it is not easy to obtain. Although the 1-norm solution may not be the sparsest, it can be easily obtained by the linear programming method. In many cases, the 0-norm solution can be obtained through finding the 1-norm solution. Many discussions exist on the equivalence of the two sparse solutions. This paper analyzes two conditions for the equivalence of the two sparse solutions. The first condition is necessary and sufficient, however, difficult to verify. Although the second is necessary but is not sufficient, it is easy to verify. In this paper, we analyze the second condition within the stochastic framework and propose a variant. We then prove that the equivalence of the two sparse solutions holds with high probability under the variant of the second condition. Furthermore, in the limit case where the 0-norm solution is extremely sparse, the second condition is also a sufficient condition with probability 1. Yuanqing Li 0001, Shun-ichi Amari |
IEEE Trans. Neural Networks | 1 |
| 2009 | Voxel selection in fMRI data analysis: A sparse representation methodabstractThis paper proposes an iterative sparse representation-based algorithmfor voxel selection in functionalmagnetic resonance imaging (fMRI) data. The output of the algorithm is a sparse weight vector, of which the magnitude of each entry represents the significance of its corresponding voxel with respect to mental tasks or stimulus. To demonstrate the validity of our algorithm and illustrate its application, we apply this algorithm to the Pittsburgh Brain Activity Interpretation Competition (PBAIC) 2007 fMRI data set for selecting the voxels which are the most relevant to the tasks of the subjects. Compared with three baseline methods, general linear model (GLM)-based statistical parametric mapping (SPM), correlation method and mutual information method, our method shows satisfactory performance for voxel selection. Yuanqing Li 0001, Zhu Liang Yu, Praneeth Namburi, Cuntai Guan |
ICASSP | 1 |
| 2009 | K-hyperline clustering learning for sparse component analysis
Zhaoshui He, Andrzej Cichocki, Yuanqing Li 0001, Shengli Xie 0001, Saeid Sanei |
Signal Process. | 3 |
| 2008 | An EEG-based BCI system for 2D cursor controlabstractIn this paper, an electroencephalogram (EEG)-based brain computer interface (BCI) is proposed for two dimensional cursor control. The horizontal and vertical movements of the cursor are controlled by mu/beta rhythm and P300 potential respectively. The main advantages of this system are: (i) two almost independent control signals are produced simultaneously; (ii) the cursor can be moved from a random position to another random position in a screen. These advantages have been demonstrated in our experiment and data analysis. Yuanqing Li 0001, Chuanchu Wang, Haihong Zhang, Cuntai Guan |
IJCNN | 1 |
| 2008 | Joint feature re-extraction and classification using an iterative semi-supervised support vector machine algorithm
Yuanqing Li 0001, Cuntai Guan |
Mach. Learn. | 1 |
| 2008 | Anomaly detection in hyperspectral imagery based on maximum entropy and nonparametric estimation
Lin He 0001, Quan Pan 0001, Wei Di, Yuanqing Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2008 | A self-training semi-supervised SVM algorithm and its application in an EEG-based brain computer interface speller system
Yuanqing Li 0001, Cuntai Guan, Huiqi Li, Zhengyang Chin |
Pattern Recognit. Lett. | 1 |
| 2008 | Equivalence Probability and Sparsity of Two Sparse Solutions in Sparse RepresentationabstractThis paper discusses the estimation and numerical calculation of the probability that the 0-norm and 1-norm solutions of underdetermined linear equations are equivalent in the case of sparse representation. First, we define the sparsity degree of a signal. Two equivalence probability estimates are obtained when the entries of the 0-norm solution have different sparsity degrees. One is for the case in which the basis matrix is given or estimated, and the other is for the case in which the basis matrix is random. However, the computational burden to calculate these probabilities increases exponentially as the number of columns of the basis matrix increases. This computational complexity problem can be avoided through a sampling method. Next, we analyze the sparsity degree of mixtures and establish the relationship between the equivalence probability and the sparsity degree of the mixtures. This relationship can be used to analyze the performance of blind source separation (BSS). Furthermore, we extend the equivalence probability estimates to the small noise case. Finally, we illustrate how to use these theoretical results to guarantee a satisfactory performance in underdetermined BSS. Yuanqing Li 0001, Andrzej Cichocki, Shun-ichi Amari, Shengli Xie 0001, Cuntai Guan |
IEEE Trans. Neural Networks | 1 |
| 2007 | A Self-Training Semi-Supervised Support Vector Machine Algorithm and its Applications in Brain Computer InterfaceabstractIn this paper, we analyze the convergence of an iterative self-training semi-supervised support vector machine (SVM) algorithm, which is designed for classification in small training data case. This algorithm converges fast and has low computational burden. Its effectiveness is also demonstrated by our data analysis results. Furthermore, we illustrate that this algorithm can be used to significantly reduce training effort and improve adaptability of a brain computer interface (BCI) system, a P300-based speller. Yuanqing Li 0001, Huiqi Li, Cuntai Guan, Zhengyang Chin |
ICASSP (1) | 1 |
| 2006 | Signal processing for brain-computer interface: enhance feature extraction and classificationabstractIn this paper we present a new scheme for brain signal processing and classification for electroencephalogram based brain-computer interfaces, by emphasizing the extraction of space-time-frequency feature as well as the combination of classifiers. In particular, we use wavelet packets as a time-frequency analysis tool and employ sparse component analysis to recover source components in the brain signals. We subsequently apply multi-class common spatial pattern filters to the signals and thus obtain important space-time-frequency features for discrimination. Furthermore, a Bayesian method is developed to boost the system, by combining multiple support vector machines in a probabilistic way. We have tested the proposed scheme on real multi-class motor imagery signals, and its efficacy has been demonstrated. Haihong Zhang, Cuntai Guan, Yuanqing Li 0001 |
ISCAS | 3 |
| 2006 | An Extended EM Algorithm for Joint Feature Extraction and Classification in Brain-Computer InterfacesabstractFor many electroencephalogram (EEG)-based brain-computer interfaces (BCIs), a tedious and time-consuming training process is needed to set parameters. In BCI Competition 2005, reducing the training process was explicitly proposed as a task. Furthermore, an effective BCI system needs to be adaptive to dynamic variations of brain signals; that is, its parameters need to be adjusted online. In this article, we introduce an extended expectation maximization (EM) algorithm, where the extraction and classification of common spatial pattern (CSP) features are performed jointly and iteratively. In each iteration, the training data set is updated using all or part of the test data and the labels predicted in the previous iteration. Based on the updated training data set, the CSP features are reextracted and classified using a standard EM algorithm. Since the training data set is updated frequently, the initial training data set can be small (semi-supervised case) or null (unsupervised case). During the above iterations, the parameters of the Bayes classifier and the CSP transformation matrix are also updated concurrently. In online situations, we can still run the training process to adjust the system parameters using unlabeled data while a subject is using the BCI system. The effectiveness of the algorithm depends on the robustness of CSP feature to noise and iteration convergence, which are discussed in this article. Our proposed approach has been applied to data set IVa of BCI Competition 2005. The data analysis results show that we can obtain satisfying prediction accuracy using our algorithm in the semisupervised and unsupervised cases. The convergence of the algorithm and robustness of CSP feature are also demonstrated in our data analysis. Yuanqing Li 0001, Cuntai Guan |
Neural Comput. | 1 |
| 2006 | Probability Estimation for Recoverability Analysis of Blind Source Separation Based on Sparse RepresentationabstractAn important application of sparse representation is underdetermined blind source separation (BSS), where the number of sources is greater than the number of observations. Within the stochastic framework, this paper discusses recoverability of underdetermined BSS based on a two-stage sparse representation approach. The two-stage approach is effective when the source matrix is sufficiently sparse. The first stage of the two-stage approach is to estimate the mixing matrix, and the second is to estimate the source matrix by minimizing the 1-norms of the source vectors subject to some constraints. After estimating the mixing matrix and fixing the number of nonzero entries of a source vector, we estimate the recoverability probability (i.e., the probability that the source vector can be recovered). A general case is then considered where the number of nonzero entries of the source vector is fixed and the mixing matrix is drawn from a specific probability distribution. The corresponding probability estimate on recoverability is also obtained. Based on this result, we further estimate the recoverability probability when the sources are also drawn from a distribution (e.g., Laplacian distribution). These probability estimates not only reflect the relationship between the recoverability and sparseness of sources, but also indicate the overall performance and confidence of the two-stage sparse representation approach for solving BSS problems. Several simulation results have demonstrated the validity of the probability estimation approach. Yuanqing Li 0001, Shun-ichi Amari, Andrzej Cichocki, Cuntai Guan |
IEEE Trans. Inf. Theory | 1 |
| 2006 | Blind estimation of channel parameters and source components for EEG signals: a sparse factorization approachabstractIn this paper, we use a two-stage sparse factorization approach for blindly estimating the channel parameters and then estimating source components for electroencephalogram (EEG) signals. EEG signals are assumed to be linear mixtures of source components, artifacts, etc. Therefore, a raw EEG data matrix can be factored into the product of two matrices, one of which represents the mixing matrix and the other the source component matrix. Furthermore, the components are sparse in the time-frequency domain, i.e., the factorization is a sparse factorization in the time frequency domain. It is a challenging task to estimate the mixing matrix. Our extensive analysis and computational results, which were based on many sets of EEG data, not only provide firm evidences supporting the above assumption, but also prompt us to propose a new algorithm for estimating the mixing matrix. After the mixing matrix is estimated, the source components are estimated in the time frequency domain using a linear programming method. In an example of the potential applications of our approach, we analyzed the EEG data that was obtained from a modified Sternberg memory experiment. Two almost uncorrelated components obtained by applying the sparse factorization method were selected for phase synchronization analysis. Several interesting findings were obtained, especially that memory-related synchronization and desynchronization appear in the alpha band, and that the strength of alpha band synchronization is related to memory performance. Yuanqing Li 0001, Andrzej Cichocki, Shun-ichi Amari |
IEEE Trans. Neural Networks | 1 |
| 2005 | Blind Identification and Deconvolution for Noisy Two-Input Two-Output Channels
Yuanqing Li 0001, Andrzej Cichocki, Jianzhao Qin |
ISNN (2) | 1 |
| 2005 | ICA and Committee Machine-Based Algorithm for Cursor Control in a BCI System
Jianzhao Qin, Yuanqing Li 0001, Andrzej Cichocki |
ISNN (1) | 2 |
| 2005 | A network model for blind source extraction in various ill-conditioned cases
Yuanqing Li 0001, Jun Wang 0002 |
Neural Networks | 1 |
| 2004 | Analysis of Sparse Representation and Blind Source SeparationabstractIn this letter, we analyze a two-stage cluster-then-l(1)-optimization approach for sparse representation of a data matrix, which is also a promising approach for blind source separation (BSS) in which fewer sensors than sources are present. First, sparse representation (factorization) of a data matrix is discussed. For a given overcomplete basis matrix, the corresponding sparse solution (coefficient matrix) with minimum l(1) norm is unique with probability one, which can be obtained using a standard linear programming algorithm. The equivalence of the l(1)-norm solution and the l(0)-norm solution is also analyzed according to a probabilistic framework. If the obtained l(1)-norm solution is sufficiently sparse, then it is equal to the l(0)-norm solution with a high probability. Furthermore, the l(1)- norm solution is robust to noise, but the l(0)-norm solution is not, showing that the l(1)-norm is a good sparsity measure. These results can be used as a recoverability analysis of BSS, as discussed. The basis matrix in this article is estimated using a clustering algorithm followed by normalization, in which the matrix columns are the cluster centers of normalized data column vectors. Zibulevsky, Pearlmutter, Boll, and Kisilev (2000) used this kind of two-stage approach in underdetermined BSS. Our recoverability analysis shows that this approach can deal with the situation in which the sources are overlapped to some degree in the analyzed domain and with the case in which the source number is unknown. It is also robust to additive noise and estimation error in the mixing matrix. Finally, four simulation examples and an EEG data analysis example are presented to illustrate the algorithm's utility and demonstrate its performance. Yuanqing Li 0001, Andrzej Cichocki, Shun-ichi Amari |
Neural Comput. | 1 |
| 2004 | Blind source estimation of FIR channels for binary sources: a grouping decision approach
Yuanqing Li 0001, Andrzej Cichocki, Liqing Zhang 0001 |
Signal Process. | 1 |
| 2003 | Blind deconvolution of FIR channels with binary sources: a grouping decision approachabstractThis paper proposes a novel grouping decision approach for blind deconvolution of FIR channels with binary sources. First, necessary and sufficient conditions for recoverability are derived. For single-input systems, a new deterministic algorithm based on grouping and decision is propose to recover the source up to a delay. Then the algorithm is extended to deal with high noise case and long decaying channel case. Furthermore blind deconvolution for multi-input systems also can be carried out as with the case of single input systems. All sources can be recovered sequentially. Finally, the validity and performance of the algorithms are illustrated by several simulation examples. Yuanqing Li 0001, Andrzej Cichocki, Liqing Zhang 0001 |
ICASSP (4) | 1 |
| 2003 | Sparse Representation and Its Applications in Blind Source SeparationabstractIn this paper, sparse representation (factorization) of a data matrix is first discussed. An overcomplete basis matrix is estimated by using the K(cid:0)means method. We have proved that for the estimated overcom- plete basis matrix, the sparse solution (coefficient matrix) with minimum l1(cid:0)norm is unique with probability of one, which can be obtained using a linear programming algorithm. The comparisons of the l1(cid:0)norm so- lution and the l0(cid:0)norm solution are also presented, which can be used in recoverability analysis of blind source separation (BSS). Next, we ap- ply the sparse matrix factorization approach to BSS in the overcomplete case. Generally, if the sources are not sufficiently sparse, we perform blind separation in the time-frequency domain after preprocessing the observed data using the wavelet packets transformation. Third, an EEG experimental data analysis example is presented to illustrate the useful- ness of the proposed approach and demonstrate its performance. Two almost independent components obtained by the sparse representation method are selected for phase synchronization analysis, and their peri- ods of significant phase synchronization are found which are related to tasks. Finally, concluding remarks review the approach and state areas that require further study. Yuanqing Li 0001, Andrzej Cichocki, Shun-ichi Amari, Sergei L. Shishkin, Jianting Cao, Fanji Gu |
NIPS | 1 |
| 2000 | Blind extraction of singularly mixed source signalsabstractThis paper introduces a novel technique for sequential blind extraction of singularly mixed sources. First, a neural-network model and an adaptive algorithm for single-source blind extraction are introduced. Next, extractability analysis is presented for singular mixing matrix, and two sets of necessary and sufficient extractability conditions are derived. The adaptive algorithm and neural-network model for sequential blind extraction are then presented. The stability of the algorithm is discussed. Simulation results are presented to illustrate the validity of the adaptive algorithm and the stability analysis. The proposed algorithm is suitable for the case of nonsingular mixing matrix as well as for singular mixing matrix. Yuanqing Li 0001, Jun Wang 0002, Jacek M. Zurada |
IEEE Trans. Neural Networks Learn. Syst. | 1 |