VLDB 2026 Research / reviewers in the wild / expert
Jinpeng Li 0002
dblp:95/2448-2
· DBLP profile ↗
27ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0001-8701-2642ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OCTMamba: A lightweight ear segmentation framework for 3D portable endoscopic OCT scanner
Junming Yan, Qiong Wang 0001, Qingjie Meng, Jinpeng Li 0002 |
Expert Syst. Appl. | 5 |
| 2026 | Entropy-increasing linear attention for multi-class unsupervised anomaly detection
Tongtong Liu 0003, Hongxia Gao, Yuxuan Tan, Jinpeng Li 0002, Jinhui Zhao |
Pattern Recognit. | 4 |
| 2025 | Accurate Coronary Microvascular Segmentation with Parallel Local-Global Chainsabstract3D vessel segmentation models aid physicians with the analysis, diagnosis, and intervention of coronary microvascular disease. Existing methods for the general medical image segmentation often produce inaccurate and discontinuous results on the complex vessel patterns, particularly for tiny vessel structures. To overcome this, recent studies have combined the convolutional neural network (CNN) and transformers in a serial architecture to enhance continuity and accuracy by leveraging long-range dependencies, i.e., the topological relationships of vessel fragments. However, the sequential architectures, where the CNN and transformers bottleneck each other, limits the ability to maintain both local and global features, which are vital in our task. In this paper, we collected the largest and highest-quality CT coronary artery dataset to date, ASACA500. Based on that, we introduce twinSeg, which enables the parallel learning of the local and global chains. To facilitate effective and efficient interaction between these chains, we propose the bidirectional attention fusion module, which enables the fusion of global and local features. Extensive evaluations conducted on our in-house ASACA and a public dataset demonstrate that twinSeg achieves state-of-the-art performance across all evaluation metrics and exhibits an exceptional average symmetric surface distance (ASSD) due to its ability to model sparse and anisotropic vessel structures. Our code and pretrained model will be released after the anonymity period. Yaling Tao, Yanwu Xu 0001, Yizhou Yu, Jinpeng Li 0002 |
BIBM | 4 |
| 2024 | PGCN: Pyramidal Graph Convolutional Network for EEG Emotion RecognitionabstractEmotion recognition is essential in the diagnosis and rehabilitation of various mental diseases. In the last decade, electroencephalogram (EEG)-based emotion recognition has been intensively investigated due to its prominative accuracy and reliability, and graph convolutional network (GCN) has become a mainstream model to decode emotions from EEG signals. However, the electrode relationship, especially long-range electrode dependencies across the scalp, may be underutilized by GCNs, although such relationships have been proven to be important in emotion recognition. The small receptive field makes shallow GCNs only aggregate local nodes. On the other hand, stacking too many layers leads to over-smoothing. To solve these problems, we propose the pyramidal graph convolutional network (PGCN), which aggregates features at three levels: local, mesoscopic, and global. First, we construct a vanilla GCN based on the 3D topological relationships of electrodes, which is used to integrate two-order local features; Second, we construct several mesoscopic brain regions based on priori knowledge and employ mesoscopic attention to sequentially calculate the virtual mesoscopic centers to focus on the functional connections of mesoscopic brain regions; Finally, we fuse the node features and their 3D positions to construct a numerical relationship adjacency matrix to integrate structural and functional connections from the global perspective. Experimental results on four public datasets indicate that PGCN enhances the relationship modelling across the scalp and achieves stateof-the-art performance in both subject-dependent and subjectindependent scenarios. Meanwhile, PGCN makes an effective trade-off between enhancing network depth and receptive fields while suppressing the ensuing over-smoothing. Our codes are publicly accessible athttps://github.com/Jinminbox/PGCN. Ming Jin 0006, Changde Du, Huiguang He, Ting Cai 0001, Jinpeng Li 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | MVCNet: Multiview Contrastive Network for Unsupervised Representation Learning for 3-D CT LesionsabstractWith the renaissance of deep learning, automatic diagnostic algorithms for computed tomography (CT) have achieved many successful applications. However, they heavily rely on lesion-level annotations, which are often scarce due to the high cost of collecting pathological labels. On the other hand, the annotated CT data, especially the 3-D spatial information, may be underutilized by approaches that model a 3-D lesion with its 2-D slices, although such approaches have been proven effective and computationally efficient. This study presents a multiview contrastive network (MVCNet), which enhances the representations of 2-D views contrastively against other views of different spatial orientations. Specifically, MVCNet views each 3-D lesion from different orientations to collect multiple 2-D views; it learns to minimize a contrastive loss so that the 2-D views of the same 3-D lesion are aggregated, whereas those of different lesions are separated. To alleviate the issue of false negative examples, the uninformative negative samples are filtered out, which results in more discriminative features for downstream tasks. By linear evaluation, MVCNet achieves state-of-the-art accuracies on the lung image database consortium and image database resource initiative (LIDC-IDRI) (88.62%), lung nodule database (LNDb) (76.69%), and TianChi (84.33%) datasets for unsupervised representation learning. When fine-tuned on 10% of the labeled data, the accuracies are comparable to the supervised learning models (89.46% versus 85.03%, 73.85% versus 73.44%, 83.56% versus 83.34% on the three datasets, respectively), indicating the superiority of MVCNet in learning representations with limited annotations. Our findings suggest that contrasting multiple 2-D views is an effective approach to capturing the original 3-D information, which notably improves the utilization of the scarce and valuable annotated CT data. Penghua Zhai, Huaiwei Cong, Enwei Zhu, Gangming Zhao, Yizhou Yu, Jinpeng Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Revisiting De-Identification of Electronic Medical Records: Evaluation of Within- and Cross-Hospital GeneralizationabstractThe de-identification task aims to detect and remove the protected health information from electronic medical records (EMRs).Previous studies generally focus on the within-hospital setting and achieve great successes, while the cross-hospital setting has been overlooked.This study introduces a new de-identification dataset comprising EMRs from three hospitals in China, creating a benchmark for evaluating both within-and cross-hospital generalization.We find significant domain discrepancy between hospitals.A model with almost perfect within-hospital performance struggles when transferred across hospitals.Further experiments show that pretrained language models and some domain generalization methods can alleviate this problem.We believe that our data and findings will encourage investigations on the generalization of medical NLP models. 1 Yiyang Liu 0002, Jinpeng Li 0002, Enwei Zhu |
EMNLP | 2 |
| 2023 | Graph to Grid: Learning Deep Representations for Multimodal Emotion RecognitionabstractMultimodal emotion recognition based on electroencephalogram (EEG) and compensating physiological signals (e.g., eye tracking) has shown potential in the diagnosis and rehabilitation tracking of depression. Since the multi-channel EEG signals are generally processed as one-dimensional (1-D) graph-like features, existing approaches can only adopt underdeveloped shallow models to recognize emotions. However, these simple models have difficulty decoupling complex emotion patterns due to their limited representation capacity. To address this problem, we propose the graph-to-grid (G2G), a concise and plug-and-play module that transforms the 1-D graph-like data into the two-dimensional (2-D) grid-like data via the numerical relation coding. After that, the well developed deep models, e.g., ResNet can be used to downstream tasks. In addition, G2G simplifies the previous complex multimodal fusion into an input matrix augmentation operation, which greatly reduces the difficulty of model design and parameter tuning. Extensive results on three public datasets (SEED, SEED5 and MPED) indicate that the proposed approach achieves state-of-the-art emotion recognition accuracy in both unimodal and multimodal settings, with good cross-session generalization ability. G2G enables the development of more appropriate multimodal emotion recognition algorithms for follow-up studies. Our code is publicly available at https://github.com/Jinminbox/G2G. Ming Jin 0006, Jinpeng Li 0002 |
ACM Multimedia | 2 |
| 2023 | Dynamic Triple Reweighting Network for Automatic Femoral Head Necrosis Diagnosis from Computed TomographyabstractAvascular necrosis of the femoral head (AVNFH) is a common orthopedic disease that seriously affects the life quality of middle-aged and elderly people. Early AVNFH is difficult to diagnose due to its complex symptoms. In recent years, some works have applied deep learning algorithms to find traces of early AVNFH in X-rays or magnetic resonance imaging (MRI). However, X-rays are difficult to reflect hidden features due to the tissue overlap; MRI is sensitive but requires more time for imaging and is expensive. This study aims to develop a computer-aided diagnosis system for early AVNFH based on computed tomography (CT), which provides layer-wise features and is less costly. To achieve this, a large-scale dataset for AVNFH was collected and annotated by experienced doctors. We propose the Dynamic Triple Reweighting Network (DTRNet) that integrates the AVNFH classification and weakly-supervised localization. DTRNet incorporates nested multi-instance learning as the first and second reweighting, and structure regularization as the third reweighting to identify diseases and localize the lesion region. Since nested multi-instance learning is inapplicable in situations with few positive samples in the patch set, we propose a dynamic pseudo-package module to compensate for this limitation. Experimental results show that DTRNet is superior to the baselines in AVNFH classification. In addition, it can locate lesions to provide more information for assisting clinical decisions. The desensitized data and codes has been made available at: https://github.com/tomas-lilingfeng/DTRNet. Gangming Zhao, Yizhou Yu, Jinpeng Li 0002 |
ACM Multimedia | 4 |
| 2023 | A unified framework of medical information annotation and extraction for Chinese clinical text
Enwei Zhu, Qilin Sheng, Huanwan Yang, Yiyang Liu 0002, Ting Cai 0001, Jinpeng Li 0002 |
Artif. Intell. Medicine | 6 |
| 2023 | Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic FeaturesabstractDecoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are difficult to generalize to novel categories that have no corresponding neural data for training. The two main reasons are 1) the under-exploitation of the multimodal semantic knowledge underlying the neural data and 2) the small number of paired (stimuli-responses) training data. To overcome these limitations, this paper presents a generic neural decoding method called BraVL that uses multimodal learning of brain-visual-linguistic features. We focus on modeling the relationships between brain, visual and linguistic features via multimodal deep generative models. Specifically, we leverage the mixture-of-product-of-experts formulation to infer a latent code that enables a coherent joint generation of all three modalities. To learn a more consistent joint representation and improve the data efficiency in the case of limited brain activity data, we exploit both intra- and inter-modality mutual information maximization regularization terms. In particular, our BraVL model can be trained under various semi-supervised scenarios to incorporate the visual and textual features obtained from the extra categories. Finally, we construct three trimodal matching datasets, and the extensive experiments lead to some interesting conclusions and cognitive insights: 1) decoding novel visual categories from human brain activity is practically possible with good accuracy; 2) decoding models using the combination of visual and linguistic features perform much better than those using either of them alone; 3) visual perception may be accompanied by linguistic influences to represent the semantics of visual stimuli. Changde Du, Kaicheng Fu, Jinpeng Li 0002, Huiguang He |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Subband fusion of complex spectrogram for fake speech detection
Cunhang Fan, Jun Xue 0001, Shunbo Dong, Mingming Ding, Jiangyan Yi, Jinpeng Li 0002, Zhao Lv |
Speech Commun. | 6 |
| 2022 | Boundary Smoothing for Named Entity RecognitionabstractNeural named entity recognition (NER) models may easily encounter the over-confidence issue, which degrades the performance and calibration.Inspired by label smoothing and driven by the ambiguity of boundary annotation in NER engineering, we propose boundary smoothing as a regularization technique for span-based neural NER models.It re-assigns entity probabilities from annotated spans to the surrounding ones.Built on a simple but strong baseline, our model achieves results better than or competitive with previous stateof-the-art systems on eight well-known NER benchmarks. 1 Further empirical analysis suggests that boundary smoothing effectively mitigates over-confidence, improves model calibration, and brings flatter neural minima and more smoothed loss landscapes. Enwei Zhu, Jinpeng Li 0002 |
ACL (1) | 2 |
| 2022 | Structure Regularized Attentive Network for Automatic Femoral Head Necrosis Diagnosis and LocalizationabstractIn recent years, several works have adopted the convolutional neural network (CNN) to diagnose the avascular necrosis of the femoral head (AVNFH) based on X-ray images or magnetic resonance imaging (MRI). However, due to the tissue overlap, X-ray images are difficult to provide fine-grained features for early diagnosis. MRI, on the other hand, has a long imaging time, is more expensive, making it impractical in mass screening. Computed tomography (CT) shows layer-wise tissues, is faster to image, and is less costly than MRI. However, to our knowledge, there is no work on CT-based automated diagnosis of AVNFH. In this work, we collected and labeled a large-scale dataset for AVNFH ranking. In addition, existing end-to-end CNNs only yields the classification result and are difficult to provide more information for doctors in diagnosis. To address this issue, we propose the structure regularized attentive network (SRANet), which is able to highlight the necrotic regions during classification based on patch attention. SRANet extracts features in chunks of images, obtains weight via the attention mechanism to aggregate the features, and constrains them by a structural regularizer with prior knowledge to improve the generalization. SRANet was evaluated on our AVNFH-CT dataset. Experimental results show that SRANet is superior to CNNs for AVNFH classification, moreover, it can localize lesions and provide more information to assist doctors in diagnosis. Our codes are made public at https://github.com/tomas-lilingfeng/SRANet. Huaiwei Cong, Gangming Zhao, Junran Peng, Zheng Zhang 0006, Jinpeng Li 0002 |
BIBM | 6 |
| 2022 | CM-MLP: Cascade Multi-scale MLP with Axial Context Relation Encoder for Edge Segmentation of Medical ImageabstractThe convolutional-based methods provide good segmentation performance in the medical image segmentation task. However, those methods have the following challenges when dealing with the edges of the medical images: (1) Previous convolutional-based methods do not focus on the boundary relationship between foreground and background around the segmentation edge, which leads to the degradation of segmentation performance when the edge changes complexly. (2) The inductive bias of the convolutional layer cannot be adapted to complex edge changes and the aggregation of multiple-segmented areas, resulting in its performance improvement mostly limited to segmenting the body of segmented areas instead of the edge. To address these challenges, we propose the CM-MLP framework on MFI (Multiscale Feature Interaction) block and ACRE (axial context relation encoder) block for accurate segmentation of the edge of medical image. In the MFI block, we propose the cascade multi-scale MLP (Cascade MLP) to process all local information from the deeper layers of the network simultaneously and utilize a cascade multiscale mechanism to fuse discrete local information gradually. Then, the ACRE block is used to make the deep supervision focus on exploring the boundary relationship between foreground and background to modify the edge of the medical image. The segmentation accuracy (Dice) of our proposed CM-MLP framework reaches 96.96%, 96.76%, and 82.54% on three benchmark datasets: CVC-ClinicDB dataset, sub-Kvasir dataset, and our inhouse dataset, respectively, which significantly outperform the state-of-the-art method. The source code and trained models will be available at https://github.com/ProgrammerHyy/CM-MLP. Jinkai Lv, Quanshui Fu, Zhiwang Zhang, Yuqiang Hu, Lin Lv, Jinpeng Li 0002 |
BIBM | 7 |
| 2022 | Deep 3D Vessel Segmentation based on Cross Transformer NetworkabstractThe coronary microvascular disease poses a great threat to human health. Computer-aided analysis/diagnosis systems help physicians intervene in the disease at early stages, where 3D vessel segmentation is a fundamental step. However, there is a lack of carefully annotated dataset to support algorithm development and evaluation. On the other hand, the commonly-used U-Net structures often yield disconnected and inaccurate segmentation results, especially for small vessel structures. In this paper, motivated by the data scarcity, we first construct two large-scale vessel segmentation datasets consisting of 100 and 500 computed tomography (CT) volumes with pixel-level annotations by experienced radiologists. To enhance the U-Net, we further propose the cross transformer network (CTN) for fine-grained vessel segmentation. In CTN, a transformer module is constructed in parallel to a U-Net to learn long-distance dependencies between different anatomical regions; and these dependencies are communicated to the U-Net at multiple stages to endow it with global awareness. Experimental results on the two in-house datasets indicate that this hybrid model alleviates unexpected disconnections by considering topological information across regions. Our codes, together with the trained models are made publicly available at https://github.com/qibaolian/ctn. Chengwei Pan, Baolian Qi, Gangming Zhao, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002 |
BIBM | 7 |
| 2022 | Faint Features Tell: Automatic Vertebrae Fracture Screening Assisted by Contrastive LearningabstractLong-term vertebral fractures severely affect the life quality of patients, causing kyphotic, lumbar deformity and even paralysis. Computed tomography (CT) is a common clinical examination to screen for this disease at early stages. However, the faint radiological appearances and unspecific symptoms lead to a high risk of missed diagnosis, especially for the mild vertebral fractures. In this paper, we argue that reinforcing the faint fracture features to encourage the inter-class separability is the key to improving the accuracy. Motivated by this, we propose a supervised contrastive learning based model to estimate Genent’s Grade of vertebral fracture with CT scans. The supervised contrastive learning, as an auxiliary task, narrows the distance of features within the same class while pushing others away, enhancing the model’s capability of capturing subtle features of vertebral fractures. Our method has a specificity of 99% and a sensitivity of 85% in binary classification, and a macro-F1 of 77% in multi-class classification, indicating that contrastive learning significantly improves the accuracy of vertebrae fracture screening. Considering the lack of datasets in this field, we construct a database including 208 samples annotated by experienced radiologists. Our desensitized data and codes will be made publicly available for the community. Huaiwei Cong, Zheng Zhang 0006, Junran Peng, Jinpeng Li 0002 |
BIBM | 6 |
| 2022 | Computer-Aided Tuberculosis Diagnosis with Attribute Reasoning Assistance
Chengwei Pan, Gangming Zhao, Junjie Fang, Baolian Qi, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002, Yizhou Yu |
MICCAI (1) | 8 |
| 2022 | Predicting liver cancers using skewed epidemiological data
Jinpeng Li 0002, Yaling Tao, Huaiwei Cong, Enwei Zhu, Ting Cai 0001 |
Artif. Intell. Medicine | 1 |
| 2022 | Dynamic Domain Adaptation for Class-Aware Cross-Subject and Cross-Session EEG Emotion RecognitionabstractIt is vital to develop general models that can be shared across subjects and sessions in the real-world deployment of electroencephalogram (EEG) emotion recognition systems. Many prior studies have exploited domain adaptation algorithms to alleviate the inter-subject and inter-session discrepancies of EEG distributions. However, these methods only aligned the global domain divergence, but overlooked the local domain divergence with respect to each emotion category. This degenerates the emotion-discriminating ability of the domain invariant features. In this paper, we argue that aligning the EEG data within the same emotion categories is important for generalizable and discriminative features. Hence, we propose the dynamic domain adaptation (DDA) algorithm where the global and local divergences are disposed by minimizing the global domain discrepancy and local subdomain discrepancy, respectively. To tackle the absence of emotion labels in the target domain, we introduce a dynamic training strategy where the model focuses on optimizing the global domain discrepancy in the early training steps, and then gradually switches to the local subdomain discrepancy. The DDA algorithm is formally implemented as an unsupervised version and a semi-supervised version for different experimental settings. Based on the coarse-to-fine alignment, our model achieves the average peak accuracy of 91.08%, 92.89% on SEED, and 81.58%, 80.82% on SEED-IV in the cross-subject and cross-session scenarios, respectively. Zhunan Li, Enwei Zhu, Ming Jin 0006, Cunhang Fan, Huiguang He, Ting Cai 0001, Jinpeng Li 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | GREN: Graph-Regularized Embedding Network for Weakly-Supervised Disease Localization in X-Ray ImagesabstractLocating diseases in chest X-ray images with few careful annotations saves large human effort. Recent works approached this task with innovative weakly-supervised algorithms such as multi-instance learning (MIL) and class activation maps (CAM), however, these methods often yield inaccurate or incomplete regions. One of the reasons is the neglection of the pathological implications hidden in the relationship across anatomical regions within each image and the relationship across images. In this paper, we argue that the cross-region and cross-image relationship, as contextual and compensating information, is vital to obtain more consistent and integral regions. To model the relationship, we propose the Graph Regularized Embedding Network (GREN), which leverages the intra-image and inter-image information to locate diseases on chest X-ray images. GREN uses a pre-trained U-Net to segment the lung lobes, and then models the intra-image relationship between the lung lobes using an intra-image graph to compare different regions. Meanwhile, the relationship between in-batch images is modeled by an inter-image graph to compare multiple images. This process mimics the training and decision-making process of a radiologist: comparing multiple regions and images for diagnosis. In order for the deep embedding layers of the neural network to retain structural information (important in the localization task), we use the Hash coding and Hamming distance to compute the graphs, which are used as regularizers to facilitate training. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for weakly-supervised disease localization. Our codes are accessible online. Baolian Qi, Gangming Zhao, Changde Du, Chengwei Pan, Yizhou Yu, Jinpeng Li 0002 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Weakly Supervised Disease Localization in Chest X-rays via Looking into Image RelationsabstractLocating diseases in chest X-ray images with few careful annotations saves large human effort in annotation. Recent works tackled this problem with innovative weakly-supervised algorithms, however, the performance of these methods on X-ray analysis is not as good as in the general computer vision tasks. Different from natural images, the global structure of different chest X-rays are relatively consistent and the disease regions are relatively inconspicuous, and radiologists often need to compare multiple images to make diagnostic decisions. Inspired by this, we propose a hypothesis that the explicit modelling of the image-to-image relationship is beneficial to the learning machines, especially when the supervision is insufficient. To model the relationship, we exploit a cross-image graph method to excavate the structural relationship between X-ray images. The proposed method represents the inter-image relationship as in-batch graphs, where each X-ray image is regarded as a node and the distance between X-ray images is defined as an edge. The graphs are used as regularizers to help preserve the structural similarity between image pairs in the embedding space. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for disease localization with limited supervision, which has a very practical use in the weakly-supervised localization tasks. The code1is accessible online. Baolian Qi, Gangming Zhao, Chaowei Fang, Zhiqiang Chen 0002, Jinpeng Li 0002 |
BIBM | 6 |
| 2021 | SNAP: Shaping neural architectures progressively via information density criterion
Zhiqiang Chen 0002, Ting-Bing Xu, Weijian Liao, Zhengcheng Li, Jinpeng Li 0002, Cheng-Lin Liu 0001, Huiguang He |
Pattern Recognit. | 5 |
| 2021 | Multi-task contrastive learning for automatic CT and X-ray diagnosis of COVID-19
Jinpeng Li 0002, Gangming Zhao, Yaling Tao, Penghua Zhai, Hao Chen 0081, Huiguang He, Ting Cai 0001 |
Pattern Recognit. | 1 |
| 2020 | FOIT: Fast Online Instance Transfer for Improved EEG Emotion RecognitionabstractThe Electroencephalogram (EEG)-based emotion recognition is promising yet limited by the requirement of a large number of training data. Collecting substantial labeled samples in the training trials is the key to the generalization on the test trials. This process is time-consuming and laborious. In recent years, several studies have proposed various semisupervised learning (e.g., active learning) and transfer learning (e.g., domain adaptation, style transfer mapping) methods to alleviate the requirement on training data. However, most of them are iterative methods, which need considerable training time and are unfeasible in practice. To tackle this problem, we present the Fast Online Instance Transfer (FOIT) for improved affective Brain-computer Interface (aBCI). FOIT selects auxiliary data from historical sessions and (or) other subjects heuristically, which are then combined with the training data for supervised training. The predictions on the test trials are made by an ensemble classifier. As a one-shot algorithm, FOIT avoids the time-consuming iterations. Experimental results show that FOIT brings significant improvement in accuracy for the three-category classification (1%-8%) on the SEED dataset and four-category classification (1%-14%) on the SEED-IV dataset in the cross-subject, cross-session and cross-all scenarios. The time cost over the baselines is moderate (~35s on average for our machine). In contrast, to achieve comparative accuracies, the iterative methods require much more time (~45s - ~900s). FOIT provides a simple, fast and practically feasible solution to improve the generalization of aBCIs and allows various choices of classifiers without constraints. Our codes are available online. Jinpeng Li 0002, Hao Chen 0081, Ting Cai 0001 |
BIBM | 1 |
| 2020 | A CNN-based comparing network for the detection of steady-state visual evoked potential responses
Jiezhen Xing, Shuang Qiu 0002, Xuelin Ma, Chenyao Wu, Jinpeng Li 0002, Shengpei Wang, Huiguang He |
Neurocomputing | 5 |
| 2020 | Multisource Transfer Learning for Cross-Subject EEG Emotion RecognitionabstractElectroencephalogram (EEG) has been widely used in emotion recognition due to its high temporal resolution and reliability. Since the individual differences of EEG are large, the emotion recognition models could not be shared across persons, and we need to collect new labeled data to train personal models for new users. In some applications, we hope to acquire models for new persons as fast as possible, and reduce the demand for the labeled data amount. To achieve this goal, we propose a multisource transfer learning method, where existing persons are sources, and the new person is the target. The target data are divided into calibration sessions for training and subsequent sessions for test. The first stage of the method is source selection aimed at locating appropriate sources. The second is style transfer mapping, which reduces the EEG differences between the target and each source. We use few labeled data in the calibration sessions to conduct source selection and style transfer. Finally, we integrate the source models to recognize emotions in the subsequent sessions. The experimental results show that the three-category classification accuracy on benchmark SEED improves by 12.72% comparing with the nontransfer method. Our method facilitates the fast deployment of emotion recognition models by reducing the reliance on the labeled data amount, which has practical significance especially in fast-deployment scenarios. Jinpeng Li 0002, Shuang Qiu 0002, Yuan-Yuan Shen, Cheng-Lin Liu 0001, Huiguang He |
IEEE Trans. Cybern. | 1 |
| 2018 | Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional DataabstractThere are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework. Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He |
ACM Multimedia | 4 |