Xun Chen 0001

dblp:34/6795-1 · DBLP profile ↗
← Back
82ranked-venue papers
5as first author
63since 2021 · last 2026
0000-0002-4922-8116ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 23 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 18 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-Supervised Pre-Training for EEG denoising
Aiping Liu, Heng Cui, Xun Chen 0001
Adv. Eng. Informatics4
2026 Clinical priors-inspired privileged knowledge distillation for reliable pancreatic lesion classification
Qiaoyu Han, Xiangpeng Hu, Huizhong Gan, Xun Chen 0001, Aiping Liu
Medical Image Anal.6
2026 Prototypical Contrastive Learning With Temporal Dynamic Graph Convolutional Network for EEG-Based Emotion Recognition
abstract
Electroencephalogram (EEG) signals are inherently non-stationary and exhibit significant inter-subject variability, leading to pronounced cross-subject distribution shifts that hinder accurate emotion recognition. Although graph convolutional networks (GCNs) and domain adaptation (DA) methods have made progress in mitigating individual differences, existing approaches still face two fundamental limitations: (1) traditional GCNs rely on static functional connectivity graphs, which fail to capture the dynamic temporal evolution of neural interactions during emotional processes, and (2) most DA-based methods only emphasize global feature alignment while overlooking emotion-specific semantic structures, thereby impairing both fine-grained discriminability and cross-subject generalization. To overcome these challenges, we propose the Prototypical Contrastive Learning with Temporal Dynamic Graph Convolutional Network (PCL-TDGCN) for EEG-based emotion recognition. Specifically, we construct an adaptive global EEG pattern memory mechanism to model temporally dynamic brain networks, thereby facilitating spatiotemporal neural interactions essential for emotion representation learning. Furthermore, we design a prototypical contrastive learning strategy that incorporates: (i) intra-domain contrastive learning to enhance the discriminability of emotional state representations, and (ii) inter-domain contrastive learning to mitigate distribution shifts across domains via semantic-aware prototypical alignment. Extensive experiments on three public datasets demonstrate that the proposed PCL-TDGCN outperforms state-of-the-art methods, achieving accuracy improvements of 1.08% (SEED), 6.53% (HIED), and 0.98% (SEED-IV) in subject-dependent experiments, and 1.08% (SEED), 7.51% (HIED), and 1.99% (SEED-IV) in subject-independent scenarios, respectively.
Yi Yang 0067, Ruoning Lyu, Ze Wang 0001, Xun Chen 0001, Chin-Teng Lin, Tzyy-Ping Jung, Feng Wan 0003
IEEE Trans. Affect. Comput.6
2026 Prior-Guided Selective Parameter Fine-Tuning for Source-Free Domain Adaptive Medical Image Segmentation
abstract
Source-free domain adaptation (SFDA) transfers knowledge from pre-trained source models to the unlabeled target domain without accessing the private source data. Conventional SFDA methods for medical image segmentation typically depend on pseudo-label driven self-training with full model fine-tuning. Although these methods have shown decent performance, the underlying principles remain insufficiently explored. In this work, we investigate SFDA through the PAC-Bayesian generalization error bound, demonstrating that its generalization error is jointly constrained by model complexity and pseudo-label noise. Motivated by this, we propose PATH, a selective PArameter fine-tuning framework guided by Topological and Historical priors for SFDA medical image segmentation. Specifically, PATH identifies domain-variant and task-distinctive parameters and sparsely updates them, thereby reducing effective model complexity during adaptation. In addition, PATH estimates pseudo-label reliability by integrating topological structure and historical prediction consistency priors to suppress pseudo-label noise. Extensive experiments on cross-scanner fundus image segmentation and cross-modality abdominal multi-organ segmentation benchmarks demonstrate that PATH outperforms competing SFDA methods, achieving state-of-the-art performance.
Fanzhe Yan, Xun Chen 0001, Aiping Liu
IEEE J. Biomed. Health Informatics3
2026 Generalizable Seizure Prediction With LLMs: Converting EEG to Textual Representations
abstract
Seizure prediction through scalp electroencephalogram (EEG) holds considerable practical potential. The primary challenge faced by existing algorithms lies in the individual heterogeneity, which hinders the generalizability of models to new patients. Additionally, inconsistencies in channel settings across various epilepsy centers further limit the applicability of models to diverse datasets. To address these challenges, we incorporate large language models (LLMs) into EEG analysis and propose a novel seizure prediction method based on LLMs (SPLLM), significantly enhancing both model generalizability and applicability. Specifically, this approach reprograms LLMs by transforming EEG signals into textual representations compatible with LLMs via a single-channel pre-training strategy. The method integrates cross-domain knowledge from both text and EEG data through a cross-attention mechanism, utilizing autoregressive pretrained LLMs to capture the temporal dependencies inherent in EEG signals. Moreover, the cross-domain generalization ability of LLMs alleviates patient heterogeneity, while the single-channel pre-training strategy enables the model to adapt to diverse channel settings. On two public datasets and one private dataset, SPLLM increases the average AUC by 8.2%, and the average balanced accuracy by 8.4% compared to existing methods. Experimental results demonstrate that the proposed method not only enhances cross-patient prediction accuracy but also adapts to data from different datasets, offering a scalable solution for the clinical application of seizure prediction.
Yuchang Zhao, Aiping Liu, Chang Li 0001, Lanlan Wang, Ruobing Qian, Xun Chen 0001
IEEE J. Biomed. Health Informatics6
2026 Task-Aware Effective Connectivity Modeling for Cognitive Function Prediction
abstract
Effective connectivity (EC) derived from resting-state functional magnetic resonance imaging (rs-fMRI) has emerged as a critical tool for deepening our understanding of brain function in both health and disease. However, most studies estimate EC on an individual basis, treating it as a hidden parameter within the model and requiring retraining the model for each subject. They often overlook the valuable population-level information and limit their generalizability. Additionally, EC is typically obtained independently of downstream tasks, reducing its capacity to effectively capture task-specific variations. To address these limitations, we propose a flexible Task-Aware Effective Connectivity (TAEC) model, designed to construct individualized, task-aware, and nonlinear causal brain networks without requiring subject-specific retraining. In this framework, a Causal Discovery Module (CDM) is introduced to capture the implicit neural representation of EC by a spatial-temporal attention mechanism, producing the estimation of an individual EC. Subsequently, we propose a Task-Aware Graph Neural Network (GNN) Predictor, which incorporates a task-aware penalty to enable end-to-end prediction, enhancing task performance and the identification of task-dependent EC patterns. Extensive experiments on twelve cognitive tasks from the Human Connectome Project (HCP) dataset demonstrate that the proposed method achieves state-of-the-art performance, validating its effectiveness in task-aware effective connectivity modeling. Furthermore, the framework discovers discriminative and task-specific EC patterns, which offer additional insights into cognitive functions.
Wantong Zou, Yu Li 0027, Hu Xiang, Xun Chen 0001, Aiping Liu
IEEE J. Biomed. Health Informatics4
2026 Refine Then Fusion: Robust 3D Brain MRI Synthesis via Vision-Language Collaboration
abstract
Metadata-guided cross-modality 3D MRI synthesis aims to generate target-contrast volumes from source-modality data conditioned on clinically available metadata, which is important for enhancing clinical imaging flexibility. However, existing methods still suffer from two main limitations: 1) They neglect spatial dependencies within volumetric representations, yielding structurally ambiguous features that blur anatomical boundaries and hinder precise semantic integration. 2) They rely on conventional cross-attention between visual and textual features, limiting the precision of visual-semantic alignment, which reduces robustness across challenging conditions. To address these issues, we propose RTFSyn, a metadata-guided 3D MRI synthesis framework that achieves effective vision-language collaboration through a refine-then-fusion paradigm. The proposed RTFSyn benefits from several merits. First, we design an axis-aware visual refinement module that captures directional dependencies within volumetric features, enabling redundancy suppression and improved structural representation before fusion. Second, we propose a cross-modal adaptive fusion module that leverages pixel packing-recovery to realize efficient cross-attention for improved alignment, while text-conditioned dynamic convolution enables fine-grained semantic injection, together enhancing vision-language collaboration. Lastly, an implicit neural decoder reconstructs the target modality as a continuous function, enabling flexible high-fidelity synthesis. Under this synergistic paradigm, RTFSyn seamlessly unites robust spatial refinement with adaptive feature fusion to achieve highly precise cross-modal alignment. Extensive experiments across four multi-center datasets demonstrate that RTFSyn not only surpasses state-of-the-art methods quantitatively, but also exhibits robust performance under diverse imaging artifacts, zero-shot evaluations, and multi-dimensional clinical validations, all with favorable computational efficiency. The high fidelity, robustness, and efficiency of RTFSyn demonstrate its great potential for clinical applications.
Jinbao Wei, Wei Wei 0068, Aiping Liu, Xun Chen 0001
IEEE Trans. Medical Imaging5
2025 Rethinking Diffusion Bridge Model with Dual Alignments for Medical Image Synthesis
abstract
Medical image synthesis is crucial in clinical workflows, enabling the generation of missing modalities from available imaging data. While recent diffusion-based models show promise in medical image synthesis, they face two key limitations: progressive distribution drift from coarse intermediate samples and structural granularity loss due to missing high-frequency constraints. To address these challenges, we propose Dual Diffusion Bridge (DualDB), a framework integrating implicit distribution alignment and explicit structural constraints within a unified diffusion bridge paradigm. First, implicit distribution alignment employs optimal transport-guided adversarial learning to minimize statistical discrepancies between intermediate and target distributions, mitigating global distribution drift. Second, explicit structural alignment applies gradient-driven constraints to preserve high-frequency anatomical features, preventing structural degradation during reverse diffusion. This complementary design ensures both global statistical consistency and local anatomical precision in the synthesized results. Extensive experiments on multi-contrast MRI and MRI-CT translation show that DualDB outperforms state-of-the-art methods in quantitative performance and visual fidelity, maintaining superior anatomical accuracy even under noisy conditions.
Jinbao Wei, Shimin Tao, Aiping Liu, Xun Chen 0001
ACM Multimedia8
2025 MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis
Jing Zhang 0165, Siying Wu, Xun Chen 0001, Yunwei Ou, Xiaoyan Sun 0001
Artif. Intell. Medicine5
2025 Uncertainty-guided Fourier-based domain generalization for seizure prediction
Zhiwei Deng, Chang Li 0001, Rencheng Song, Ruobing Qian, Xun Chen 0001
Expert Syst. Appl.6
2025 Data Distillation for Sleep Stage Classification
abstract
Deep learning frameworks have been increasingly applied in the field of sleep stage classification. However, most advanced frameworks rely on huge datasets. In practical applications, data are generally continuously collected in batches. This requires continuous retraining of the model. It is expensive to store data and train models on them. Meanwhile, extensive access to EEG data may infringe upon the privacy of patients. To solve these problems, we propose a data distillation algorithm for sleep stage classification. This method compresses large EEG datasets into a small information synthetic dataset that can be used to train the network from scratch. To our knowledge, this is the first application of data distillation in this field. In the process of data distillation, we use gradient matching to optimize the synthetic dataset, which can avoid falling into the local minimum optimization. At the same time, we introduce data augmentation in the process to enhance the generalization ability of synthetic dataset. Finally, in order to improve the stability of data distillation, we use K-medoids clustering to initialize the synthesized dataset. We validated six sleep stage classification frameworks on three publicly available datasets and proved the superiority of our method. We also explored the application of our method in the field of neural architecture search, achieving robust results.
Hanfei Guo, Chang Li 0001, Hu Peng, Heyuan Qiao, Xun Chen 0001
IEEE Internet Things J.6
2025 Feature Unlearning for EEG-Based Seizure Prediction
abstract
While patient-specific seizure prediction deep learning (DL) models can deliver remarkable performance tailored to individual patients, the development of patient-independent models that offer satisfactory cross-subject performance holds greater significance and practicality. However, these patient-independent models, which leverage electroencephalogram (EEG) data from multiple patients, may give rise to privacy concerns. This is because EEG data contains sensitive information regarding individuals’ health and mental states. Consequently, from a privacy-preserving perspective, patients may desire the removal of their data information from trained models. Yet, accommodating such forgetting requests presents a formidable challenge: how to enable DL models to forget the data information of specific patients without compromising the performance for others. Although retraining a model from scratch without the data of a specific patient can somewhat address this issue, it becomes computationally prohibitive, especially with large datasets. To tackle this, we introduce an efficient machine unlearning approach called feature unlearning (FU) for seizure prediction. This method modifies the feature projection distribution of specific patients’ data within trained models to match that of models retrained from scratch. Our proposed FU method comprises two primary components: 1) feature shifting, which alters the original distribution of feature projection of specific patients in trained models and feature retaining, which mitigates the adverse effects of feature shifting on other patients, preserving their overall performance through knowledge distillation. We assess our FU method using the CHB-MIT dataset. The results demonstrate that our FU approach can effectively remove the data information of specific patients from trained DL models while maintaining the performance for other patients.
Chenghao Shao, Chang Li 0001, Rencheng Song, Guoping Xu, Xun Chen 0001
IEEE Internet Things J.5
2025 Multi-Contrast MRI Arbitrary-Scale Super-Resolution via Dynamic Implicit Network
abstract
Multi-contrast MRI super-resolution (SR) aims to restore high-resolution target image from low-resolution one, where reference image from another contrast is used to promote this task. To better meet clinical needs, current studies mainly focus on developing arbitrary-scale MRI SR solutions rather than fixed-scale ones. However, existing arbitrary-scale SR methods still suffer from the following two issues: 1) They typically rely on fixed convolutions to learn multi-contrast features, struggling to handle the feature transformations under varying scales and input image pairs, thus limiting their representation ability. 2) They simply combine the multi-contrast features as prior information, failing to fully exploit the complementary information in the texture-rich reference images. To address these issues, we propose a Dynamic Implicit Network (DINet) for multi-contrast MRI arbitrary-scale SR. DINet offers several key advantages. First, the scale-adaptive dynamic convolution facilitates dynamic feature learning based on scale factors and input image pairs, significantly enhancing the representation ability of multi-contrast features. Second, the dual-branch implicit attention enables arbitrary-scale upsampling of MR images through implicit neural representation. Following this, we propose the modulation-then-fusion block to adaptively align and fuse multi-contrast features, effectively incorporating complementary details from reference images into the target images. By jointly combining the above-mentioned modules, our proposed DINet achieves superior MRI SR performance at arbitrary scales. Extensive experiments on three datasets demonstrate that DINet significantly outperforms state-of-the-art methods, highlighting its potential for clinical applications. The code is available athttps://github.com/weijinbao1998/DINet.
Jinbao Wei, Wei Wei 0068, Aiping Liu, Xun Chen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 VDMUFusion: A Versatile Diffusion Model-Based Unsupervised Framework for Image Fusion
abstract
Image fusion facilitates the integration of information from various source images of the same scene into a composite image, thereby benefiting perception, analysis, and understanding. Recently, diffusion models have demonstrated impressive generative capabilities in the field of computer vision, suggesting significant potential for application in image fusion. The forward process in the diffusion models requires the gradual addition of noise to the original data. However, typical unsupervised image fusion tasks (e.g., infrared-visible, medical, and multi-exposure image fusion) lack ground truth images (corresponding to the original data in diffusion models), thereby preventing the direct application of the diffusion models. To address this problem, we propose a versatile diffusion model-based unsupervised framework for image fusion, termed as VDMUFusion. In the proposed method, we integrate the fusion problem into the diffusion sampling process by formulating image fusion as a weighted average process and establishing appropriate assumptions about the noise in the diffusion model. To simplify the training process, we propose a multi-task learning framework that replaces the original noise prediction network, allowing for simultaneous prediction of noise and fusion weights. Meanwhile, our method employs joint training across various fusion tasks, which significantly improves noise prediction accuracy and yields higher quality fused images compared to training on a single task. Extensive experimental results demonstrate that the proposed method delivers very competitive performance across various image fusion tasks. The code is available at https://github.com/yuliu316316/VDMUFusion.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.5
2025 Degradation-Aware Prompted Transformer for Unified Medical Image Restoration
abstract
Medical image restoration (MedIR) aims to recover high-quality images from degraded inputs, yet faces unique challenges from physics-driven degradations and multi-modal task interference. While existing all-in-one methods handle natural image degradations well, they struggle with medical scenarios due to limited degradation perception and suboptimal multi-task optimization. In response, we introduce DaPT, a Degradation-aware Prompted Transformer, which integrates dynamic prompt learning and modular expert mining for unified MedIR. First, DaPT introduces spatially compact prompts with optimal transport regularization, amplifying inter-prompt differences to capture diverse degradation patterns. Second, a mixture of experts dynamically routes inputs to specialized modules via prompt guidance, resolving task conflicts while reducing computational overhead. The synergy of prompt learning and expert mining further enables robust restoration across multi-modal medical data, offering a practical solution for clinical imaging. Extensive experiments across multiple modalities (MRI, CT, PET) and diverse degradations, covering both in-distribution and out-of-distribution scenarios, demonstrate that DaPT consistently outperforms state-of-the-art methods and generalizes reliably to unseen settings, underscoring its robustness, effectiveness, and clinical practicality. The source code will be released at https://github.com/weijinbao1998/DaPT.
Jinbao Wei, Shimin Tao, Aiping Liu, Xun Chen 0001
IEEE Trans. Image Process.6
2025 A Flexible Spatio-Temporal Architecture Design for Artifact Removal in EEG With Arbitrary Channel-Settings
abstract
Electroencephalography (EEG) data is easily contaminated by various sources, significantly affecting subsequent analyses in neuroscience and clinical applications. Therefore, effective artifact removal is a key step in EEG preprocessing. While current deep learning methods have demonstrated notable efficacy in EEG denoising, single-channel approaches primarily focus on temporal features and neglect inter-channel correlations. Meanwhile, multi-channel methods mainly prioritize spatial features but often overlook the unique temporal dependencies of individual channels. A common limitation of both single-channel and multi-channel methods is their strict requirements on the input channel setting, which restricts their practical applicability. To address these issues, we design a flexible architecture named Artifact removal Spatio-Temporal Integration Network (ASTI-Net), a dual-branch denoising model capable of handling arbitrary EEG channel settings. ASTI-Net utilizes spatio-temporal attention weighting with dual branches that capture inter-channel spatial characteristics and intra-channel temporal dependencies. Its architecture incorporates deformable convolutional operations and channel-wise temporal processing, accommodating varying numbers of EEG channels and enhancing applicability across diverse clinical and research settings. By integrating features from both branches through a fusion reconstruction module, ASTI-Net effectively restores clean multi-channel EEG. Extensive evaluation on two semi-simulated datasets, along with qualitative assessment on real task-state EEG data, validates that ASTI-Net outperforms existing artifact removal methods.
Aiping Liu, Heng Cui, Xun Chen 0001
IEEE J. Biomed. Health Informatics4
2025 EEGDfus: A Conditional Diffusion Model for Fine-Grained EEG Denoising
abstract
Electroencephalogram (EEG) signals are vital in understanding brain activity, but their weak amplitude makes them susceptible to various artifacts. Accurate denoising of EEG data is crucial as a preprocessing step to ensure precise analysis and interpretation. In recent years, the diffusion model has garnered significant attention as a promising approach in generative modeling. This model effectively addresses the issue of over-smoothing in existing deep learning methods and thus has the potential to generate more refined denoised EEG signals. However, the generation process of the standard diffusion model is highly random, limiting its direct application to EEG denoising tasks. To address this limitation, we propose a conditional diffusion model specifically designed for EEG denoising. In this model, the standard diffusion model's denoising network is replaced by a novel dual-branch network, where noisy EEG information is used as a condition to guide the generation of corresponding clean EEG signals. This dual-branch structure leverages the complementary strengths of convolutional neural network (CNN) and Transformer architectures, integrating multi-scale features to comprehensively extract information from the signal. Extensive experiments demonstrate the remarkable performance of EEGDfus in EEG denoising. We tested it on two public datasets. Testing on two public datasets, EEGdenoiseNet and SSED, demonstrated that after denoising, the average correlation coefficient increased to 0.983 and 0.992 for EOG artifact removal, respectively. The proposed model outperforms commonly used baseline models, setting a new state-of-the-art benchmark in the field of EEG denoising.
Chang Li 0001, Aiping Liu, Ruobing Qian, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2025 MIF: Multi-Shot Interactive Fusion Model for Cancer Survival Prediction Using Pathological Image and Genomic Data
abstract
Accurate cancer survival prediction is crucial for oncologists to determine therapeutic plan, which directly influences the treatment efficacy and survival outcome of patient. Recently, multimodal fusion-based prognostic methods have demonstrated effectiveness for survival prediction by fusing diverse cancer-related data from different medical modalities, e.g., pathological images and genomic data. However, these works still face significant challenges. First, most approaches attempt multimodal fusion by simple one-shot fusion strategy, which is insufficient to explore complex interactions underlying in highly disparate multimodal data. Second, current methods for investigating multimodal interactions face the capability-efficiency dilemma, which is the difficult balance between powerful modeling capability and applicable computational efficiency, thus impeding effective multimodal fusion. In this study, to encounter these challenges, we propose an innovative multi-shot interactive fusion method named MIF for precise survival prediction by utilizing pathological and genomic data. Particularly, a novel multi-shot fusion framework is introduced to promote multimodal fusion by decomposing it into successive fusing stages, thus delicately integrating modalities in a progressive way. Moreover, to address the capacity-efficiency dilemma, various affinity-based interactive modules are introduced to synergize the multi-shot framework. Specifically, by harnessing comprehensive affinity information as guidance for mining interactions, the proposed interactive modules can efficiently generate low-dimensional discriminative multimodal representations. Extensive experiments on different cancer datasets unravel that our method not only successfully achieves state-of-the-art performance by performing effective multimodal fusion, but also possesses high computational efficiency compared to existing survival prediction methods.
Yi Shi 0013, Ao Li 0001, Xun Chen 0001
IEEE J. Biomed. Health Informatics6
2025 A GAN Guided Parallel CNN and Transformer Network for EEG Denoising
abstract
Electroencephalography (EEG) signals are often contaminated with various physiological artifacts, seriously affecting the quality of subsequent analysis. Therefore, removing artifacts is an essential step in practice. As of now, deep learning-based EEG denoising methods have exhibited unique advantages over traditional methods. However, they still suffer from the following limitations. The existing structure designs have not fully taken into account the temporal characteristics of artifacts. Meanwhile, the existing training strategies usually ignore the holistic consistency between denoised EEG signals and authentic clean ones. To address these issues, we propose a GAN guided parallel CNN and transformer network, named GCTNet. The generator contains parallel CNN blocks and transformer blocks to respectively capture local and global temporal dependencies. Then, a discriminator is employed to detect and correct the holistic inconsistencies between clean and denoised EEG signals. We evaluate the proposed network on both semi-simulated and real data. Extensive experimental results demonstrate that GCTNet significantly outperforms state-of-the-art networks in various artifact removal tasks, as evidenced by its superior objective evaluation metrics. For example, in the task of removing electromyography artifacts, GCTNet achieves 11.15% reduction in RRMSE and 9.81% improvement in SNR over other methods, highlighting the potential of the proposed method as a promising solution for EEG signals in practical applications.
Jin Yin, Aiping Liu, Chang Li 0001, Ruobing Qian, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2024 Enhancing EEG artifact removal through neural architecture search with large kernels
Le Wu 0003, Aiping Liu, Chang Li 0001, Xun Chen 0001
Adv. Eng. Informatics4
2024 Online Seizure Prediction via Fine-Tuning and Test-Time Adaptation
abstract
Privacy protection has become increasingly crucial in the field of epilepsy prediction. Some latest studies introduced the source free domain adaptation (SFDA), which only utilizes a pre-trained source model for protecting the source data privacy. However, the existing SFDA methods exist two shortcomings. (1) the offline setting, which is not suitable for real-world online scenarios (2) the poor performance, which is attributed to the absence of labeled calibration data during the adaptation phase. To this end, we proposed a online seizure prediction framework based on fine-tuning and test-time adaptation (FT3A). Specifically, FT3A employs one seizure event target data to fine-tune and continuously adapt pre-trained source model to unlabeled target data stream. In addition, the adaption and prediction is performed simultaneously. On the one hand, we design the task model as a multi-head structure to increase the confidence of the model and reduce error accumulation. On the other hand, a memory bank is introduced to store a small amount of historical EEG data, which helps handle the catastrophic forgetting concern of the model during online adaptation. Extensive experiments on public CHB-MIT dataset and the private freiburg hospital dataset indicate the superiority and generality of the proposed method.
Tingting Mao, Chang Li 0001, Rencheng Song, Guoping Xu, Xun Chen 0001
IEEE Internet Things J.5
2024 Misalignment-Resistant Deep Unfolding Network for multi-modal MRI super-resolution and reconstruction
Jinbao Wei, Yu Liu 0023, Aiping Liu, Xun Chen 0001
Knowl. Based Syst.6
2024 Detecting fake information with knowledge-enhanced AutoPrompt
Xun Chen 0001, Yadang Chen, Qianmu Li
Neural Comput. Appl.1
2024 Rethinking the Effectiveness of Objective Evaluation Metrics in Multi-Focus Image Fusion: A Statistic-Based Approach
abstract
As an effective technique to extend the depth-of-field (DOF) of optical lenses, multi-focus image fusion has recently become an active topic in image processing community. However, a major problem remaining unsolved in this field is the lack of universal criteria in selecting objective evaluation metrics. Consequently, the metrics utilized in different studies often vary significantly, leading to high difficulties in achieving unbiased evaluation. To address this problem, this paper proposes a statistic-based approach for verifying the effectiveness of objective metrics in multi-focus image fusion. The core idea is to adopt statistical correlation measures to evaluate the performance consistency between a certain fusion metric and some popular full-reference image quality assessment models. In addition, a convolutional neural network (CNN)-based fusion metric is presented to measure the similarity between the source images and the fused image based on the semantic features at multiple abstraction levels. A comparative study is conducted to evaluate 20 existing fusion metrics using the proposed statistic-based approach on a large-scale, realistic and with-ground-truth multi-focus image fusion dataset recently released. Experimental results demonstrate the feasibility of the proposed approach in evaluating the effectiveness of objective metrics and the advantage of our CNN-based metric.
Yu Liu 0023, Zhengzheng Qi, Juan Cheng 0004, Xun Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Source-Free Domain Adaptation for Privacy-Preserving Seizure Prediction
abstract
Domain adaptation (DA) techniques are frequently utilized to enhance seizure prediction accuracy by leveraging the labeled electroencephalogram data of existing patients on new patients. Traditional DA methods, however, require access to the source domain while training the adaptation model, which poses a threat to sensitive patient information and privacy. To address this issue, in this article, we propose a novel Gaussian mixture modeling (GMM)-based source-free domain adaptation (GSFDA). Our method leverages the GMM joint source model and target data structure for clustering, employs uncertainty learning to minimize DA uncertainty, and uses the mixup technique to increase model robustness while reducing the impact of noisy pseudolabels. Notably, GSFDA only requires access to the source model parameters, and not the source domain, effectively safeguarding the privacy of patient information. This has substantial clinical implications for seizure prediction.
Yuchang Zhao, Chang Li 0001, Rencheng Song, Deng Liang, Xun Chen 0001
IEEE Trans. Ind. Informatics6
2024 MM-Net: A MixFormer-Based Multi-Scale Network for Anatomical and Functional Image Fusion
abstract
Anatomical and functional image fusion is an important technique in a variety of medical and biological applications. Recently, deep learning (DL)-based methods have become a mainstream direction in the field of multi-modal image fusion. However, existing DL-based fusion approaches have difficulty in effectively capturing local features and global contextual information simultaneously. In addition, the scale diversity of features, which is a crucial issue in image fusion, often lacks adequate attention in most existing works. In this paper, to address the above problems, we propose a MixFormer-based multi-scale network, termed as MM-Net, for anatomical and functional image fusion. In our method, an improved MixFormer-based backbone is introduced to sufficiently extract both local features and global contextual information at multiple scales from the source images. The features from different source images are fused at multiple scales based on a multi-source spatial attention-based cross-modality feature fusion (CMFF) module. The scale diversity of the fused features is further enriched by a series of multi-scale feature interaction (MSFI) modules and feature aggregation upsample (FAU) modules. Moreover, a loss function consisting of both spatial domain and frequency domain components is devised to train the proposed fusion model. Experimental results demonstrate that our method outperforms several state-of-the-art fusion methods on both qualitative and quantitative comparisons, and the proposed fusion model exhibits good generalization capability. The source code of our fusion method will be available at https://github.com/yuliu316316.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.5
2023 DA-Parser: A Pre-trained Domain-aware Parsing Framework for Heterogeneous Log Analysis
abstract
Automated log analysis is widely applied in modern software-intensive systems to ensure resilience and sustainability, where log parsing is a vital initial step, converting unstructured logs into structured data for downstream analysis. However, traditional log parsing algorithms are designed to process logs within a single domain. As cross-domain dependencies and interactions between sub-modules of software systems increase, these algorithms struggle to handle the challenges posed by multi-domain log inputs, which results in a significant decline in parsing accuracy when facing heterogeneous logs. Additionally, current solutions for heterogeneous log parsing require extensive manual labeling efforts. In this paper, we propose Domain-aware Parser (DA-Parser), a framework that consists of a domain-aware head to identify the source domains of heterogeneous logs and then converts the multi-domain log parsing problem into a series of single-domain parsing problems. The domain-aware head is pretrained using a corpus of logs from 16 domains, which allows for the classification of the source domains of most heterogeneous log set without additional human labeling. Source domain tags predicted by the domain-aware head serve as a constraint to limit the template extraction process to logs from the same domain. Empirical evaluation is conducted on a multi-domain dataset containing logs from 7 domains. DA-Parser can be integrated with existing single-domain algorithms and are compatible with them, achieving superior parsing accuracy with an average of 9.26% improvement compared with single-domain algorithms.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Jingyu Wang 0001, Chang Su 0001, Weinan Tian, Min Zhang 0042, Hao Yang 0006, Xun Chen 0001
COMPSAC10
2023 PanFlowNet: A Flow-Based Deep Network for Pan-sharpening
abstract
Pan-sharpening aims to generate a high-resolution multispectral (HRMS) image by integrating the spectral information of a low-resolution multispectral (LRMS) image with the texture details of a high-resolution panchromatic (PAN) image. It essentially inherits the ill-posed nature of the super-resolution (SR) task that diverse HRMS images can degrade into an LRMS image. However, existing deep learning-based methods recover only one HRMS image from the LRMS image and PAN image using a deterministic mapping, thus ignoring the diversity of the HRMS image. In this paper, to alleviate this ill-posed issue, we propose a flow-based pan-sharpening network (PanFlowNet) to directly learn the conditional distribution of HRMS image given LRMS image and PAN image instead of learning a deterministic mapping. Specifically, we first transform this unknown conditional distribution into a given Gaussian distribution by an invertible network, and the conditional distribution can thus be explicitly defined. Then, we design an invertible Conditional Affine Coupling Block (CACB) and further build the architecture of PanFlowNet by stacking a series of CACBs. Finally, the PanFlowNet is trained by maximizing the log-likelihood of the conditional distribution given a training set and can then be used to predict diverse HRMS images. The experimental results verify that the proposed PanFlowNet can generate various HRMS images given an LRMS image and a PAN image. Additionally, the experimental results on different kinds of satellite datasets also demonstrate the superiority of our PanFlowNet compared with other state-of-the-art methods both visually and quantitatively. Code is available at Github.
Xiangyong Cao, Wenzhe Xiao, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Deyu Meng
ICCV6
2023 Accurate MRI Reconstruction via Multi-Domain Recurrent Networks
abstract
In recent years, deep convolutional neural networks (CNNs) have become dominant in MRI reconstruction from undersampled k-space. However, most existing CNNs methods reconstruct the undersampled images either in the spatial domain or in the frequency domain, and neglecting the correlation between these two domains. This hinders the further reconstruction performance improvement. To tackle this issue, in this work, we propose a new multi-domain recurrent network (MDR-Net) with multi-domain learning (MDL) blocks as its basic units to reconstruct the undersampled MR image progressively. Specifically, the MDL block interactively processes the local spatial features and the global frequency information to facilitate complementary learning, leading to fine-grained features generation. Furthermore, we introduce an effective frequency-based loss to narrow the frequency spectrum gap, compensating for over-smoothness caused by the widely used spatial reconstruction loss. Extensive experiments on public fastMRI datasets demonstrate that our MDR-Net consistently outperforms other competitive methods and is able to provide more details.
Jinbao Wei, Kongqiao Wang, Xueyang Fu, Xun Chen 0001
IJCAI7
2023 Detection of Cross-Dataset Fake Audio Based on Prosodic and Pronunciation Features
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Xun Chen 0001
INTERSPEECH6
2023 TO-Rawnet: Improving RawNet with TCN and Orthogonal Regularization for Fake Audio Detection
Chenglong Wang 0001, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Shuai Zhang 0014, Ruibo Fu, Xun Chen 0001
INTERSPEECH7
2023 Biglog: Unsupervised Large-scale Pre-training for a Unified Log Representation
abstract
Automated log analysis has been widely applied in modern data-center network, performing critical tasks such as log parsing, log anomaly detection and log-based failure prediction. However, existing approaches rely on hand-crafted features or domain-specific vectors to represent logs, which are either laborious in manual efforts or ineffective facing multiple domains in a system. Furthermore, general-purpose word embeddings are not optimized for log data, thus are data-inefficient in handling complex log analysis tasks. In this paper, we present a pre-training phase for language models to understand both in-sentence and cross-sentence features of logs, resulting in a unified representation of logs that is well-suited for various downstream analysis tasks. The pre-training phase is unsupervised, utilizing 0.45 billion logs from 16 diverse domains. Experiments on 12 publicly available evaluation datasets across 3 tasks indicate superiority of our approach against existing approaches, especially in online scenarios with limited historical logs. Our approach also exhibits remarkable few-shot learning ability and domain-adaptiveness, which not only outperforms existing approaches using only 0.0025% of their required training data, but also adapts into new domains via only a few in-domain logs. We release our code and pre-trained model.
Shimin Tao, Yilun Liu 0001, Weibin Meng, Zuomin Ren, Hao Yang 0006, Xun Chen 0001, Yuming Xie, Chang Su 0001, Xiaosong Oiao, Weinan Tian, Yichen Zhu 0001
IWQoS6
2023 EEG-based seizure prediction via hybrid vision transformer and data uncertainty learning
Zhiwei Deng, Chang Li 0001, Rencheng Song, Ruobing Qian, Xun Chen 0001
Eng. Appl. Artif. Intell.6
2023 Multi-Exposure Image Fusion via Multi-Scale and Context-Aware Feature Learning
abstract
In this letter, a deep learning (DL)-based multi-exposure image fusion (MEF) method via multi-scale and context-aware feature learning is proposed, aiming to overcome the defects of existing traditional and DL-based methods. The proposed network is based on an auto-encoder architecture. First, an encoder that combines the convolutional network and Transformer is designed to extract multi-scale features and capture the global contextual information. Then, a multi-scale feature interaction (MSFI) module is devised to enrich the scale diversity of extracted features using cross-scale fusion and Atrous spatial pyramid pooling (ASPP). Finally, a decoder with a nest connection architecture is introduced to reconstruct the fused image. Experimental results show that the proposed method outperforms several representative traditional and DL-based MEF methods in terms of both visual quality and objective assessment.
Yu Liu 0023, Juan Cheng 0004, Xun Chen 0001
IEEE Signal Process. Lett.4
2023 EEG-Based Subject-Independent Emotion Recognition Using Gated Recurrent Unit and Minimum Class Confusion
abstract
Automatic emotion recognition based on electroencephalogram (EEG) has attracted rapidly increasing interests. Due to large inter-subject variabilities, subject-independent emotion recognition faces great challenges. Recently, domain adaptation methods have been successfully applied in this field due to their ability to align features from different subjects. However, since EEG signals corresponding to some emotions have similar oscillation patterns, they are often confused and aligned to the wrong categories, which limits the generalization ability of the model across subjects. Besides, almost all methods only support offline applications, which require collecting a large number of samples of new subjects. To achieve online recognition, a simpler model is needed. In this paper, a novel Gated Recurrent Unit-Minimum Class Confusion (GRU-MCC) model is proposed. Specifically, a simple feature extractor based on gated recurrent unit (GRU) is firstly applied to model the spatial dependence of multiple electrodes and obtain high-level discriminative features. Then, during training, minimum class confusion (MCC) loss is introduced to reduce the confusion between the correct and ambiguous classes for the target subject and increase the transfer gains. We conduct both offline and online experiments on two public datasets: SEED and MPED. The results indicate that our method can obtain the superior performance.
Heng Cui, Aiping Liu, Xu Zhang 0002, Xiang Chen 0004, Jun Liu 0004, Xun Chen 0001
IEEE Trans. Affect. Comput.6
2023 EEG-Based Emotion Recognition via Neural Architecture Search
abstract
With the flourishing development of deep learning (DL) and the convolution neural network (CNN), electroencephalogram-based (EEG) emotion recognition is occupying an increasingly crucial part in the field of brain-computer interface (BCI). However, currently employed architectures have mostly been designed manually by human experts, which is a time-consuming and labor-intensive process. In this paper, we proposed a novel neural architecture search (NAS) framework based on reinforcement learning (RL) for EEG-based emotion recognition, which can automatically design network architectures. The proposed NAS mainly contains three parts: search strategy, search space, and evaluation strategy. During the search process, a recurrent network (RNN) controller is used to select the optimal network structure in the search space. We trained the controller with RL to maximize the expected reward of the generated models on a validation set and force parameter sharing among the models. We evaluated the performance of NAS on the DEAP and DREAMER dataset. On the DEAP dataset, the average accuracies reached 97.94%, 97.74%, and 97.82% on arousal, valence, and dominance respectively. On the DREAMER dataset, average accuracies reached 96.62%, 96.29% and 96.61% on arousal, valence, and dominance, respectively. The experimental results demonstrated that the proposed NAS outperforms the state-of-the-art CNN-based methods.
Chang Li 0001, Zhongzhen Zhang, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
IEEE Trans. Affect. Comput.6
2023 EEG-Based Emotion Recognition via Channel-Wise Attention and Self Attention
abstract
Emotion recognition based on electroencephalography (EEG) is a significant task in the brain-computer interface field. Recently, many deep learning-based emotion recognition methods are demonstrated to outperform traditional methods. However, it remains challenging to extract discriminative features for EEG emotion recognition, and most methods ignore useful information in channel and time. This article proposes an attention-based convolutional recurrent neural network (ACRNN) to extract more discriminative features from EEG signals and improve the accuracy of emotion recognition. First, the proposed ACRNN adopts a channel-wise attention mechanism to adaptively assign the weights of different channels, and a CNN is employed to extract the spatial information of encoded EEG signals. Then, to explore the temporal information of EEG signals, extended self-attention is integrated into an RNN to recode the importance based on intrinsic similarity in EEG signals. We conducted extensive experiments on the DEAP and DREAMER databases. The experimental results demonstrate that the proposed ACRNN outperforms state-of-the-art methods.
Chang Li 0001, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Feng Wan 0003, Xun Chen 0001
IEEE Trans. Affect. Comput.7
2023 MSCAF-Net: A General Framework for Camouflaged Object Detection via Learning Multi-Scale Context-Aware Features
abstract
The aim of camouflaged object detection (COD) is to find objects that are hidden in their surrounding environment. Due to the factors like low illumination, occlusion, small size and high similarity to the background, COD is recognized to be a very challenging task. In this paper, we propose a general COD framework, termed as MSCAF-Net, focusing on learning multi-scale context-aware features. To achieve this target, we first adopt the improved Pyramid Vision Transformer (PVTv2) model as the backbone to extract global contextual information at multiple scales. An enhanced receptive field (ERF) module is then designed to refine the features at each scale. Further, a cross-scale feature fusion (CSFF) module is introduced to achieve sufficient interaction of multi-scale information, aiming to enrich the scale diversity of extracted features. In addition, inspired the mechanism of the human visual system, a dense interactive decoder (DID) module is devised to output a rough localization map, which is used to modulate the fused features obtained in the CSFF module for more accurate detection. The effectiveness of our MSCAF-Net is validated on four benchmark datasets. The results show that the proposed method significantly outperforms state-of-the-art (SOTA) COD models by a large margin. Besides, we also investigate the potential of our MSCAF-Net on some other vision tasks that are highly related to COD, such as polyp segmentation, COVID-19 lung infection segmentation, transparent object detection and defect detection. Experimental results demonstrate the high versatility of the proposed MSCAF-Net. The source code and results of our method are available athttps://github.com/yuliu316316/MSCAF-COD.
Yu Liu 0023, Haihang Li, Juan Cheng 0004, Xun Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Silent Speech Recognition Based on High-Density Surface Electromyogram Using Hybrid Neural Networks
abstract
This article presents a silent speech recognition approach based on high-density (HD) surface electromyogram (sEMG) using hybrid neural networks that support anomaly detection. In the hybrid networks, both a convolutional long short-term memory module and an autoencoder module were designed to extract discriminative spatio-temporal features and potentially identify any anomaly patterns, respectively. To verify the effectiveness of the proposed method, experimental data were recorded using HD-sEMG arrays with 64 channels from 11 subjects subvocalizing 33 Chinese words and articulating 9 anomaly patterns. The proposed method significantly outperformed other comparison methods (p< 0.05) and achieved the highest anomaly detection rate of 90.61% while maintaining a high level of target word-pattern classification accuracy of 82.30%. These findings demonstrate the effectiveness of the proposed method for improving the robustness of the SSR approach based on HD-sEMG recordings against anomaly muscular activities. This article also provides a novel solution for building practical and robust sEMG-based SSR systems with broad applications, such as instant messaging and human-computer interaction.
Xi Chen 0092, Yuanfei Xia, Le Wu 0003, Xiang Chen 0004, Xun Chen 0001, Xu Zhang 0002
IEEE Trans. Hum. Mach. Syst.6
2023 EEG-based Emotion Recognition via Transformer Neural Architecture Search
abstract
Emotion recognition based on electroencephalogram (EEG) plays an increasingly important role in the field of brain–computer interfaces. Recently, deep learning has been widely applied to EEG decoding owning to its excellent capabilities in automatic feature extraction. Transformer holds great superiority in processing time-series signals due to its long-term dependencies extraction ability. However, most existing transformer architectures are designed manually by human experts, which is a time-consuming and resource-intensive process. In this article, we propose an automatic transformer neural architectures search (TNAS) framework based on multiobjective evolution algorithm (MOEA) for the EEG-based emotion recognition. The proposed TNAS conducts the MOEA strategy that considers both accuracy and model size to discover the optimal model from well-trained supernet for the emotion recognition. We conducted extensive experiments to evaluate the performance of the proposed TNAS on the DEAP and DREAMER datasets. The experimental results showed that the proposed TNAS outperforms the state-of-the-art methods.
Chang Li 0001, Zhongzhen Zhang, Guoning Huang, Yu Liu 0023, Xun Chen 0001
IEEE Trans. Ind. Informatics6
2023 Bi-CapsNet: A Binary Capsule Network for EEG-Based Emotion Recognition
abstract
In recent years, deep learning has gained widespread attention in electroencephalogram (EEG)-based emotion recognition. However, deep learning methods are usually time-consuming with a large amount of memory usage, which obstructs their practical usage on resource-constrained devices. In this paper, we propose a binary capsule network (Bi-CapsNet) for EEG emotion recognition with low computational cost and memory usage. The Bi-CapsNet binarizes 32-bit weights and activations to 1 b, and replaces floating-point operations with efficient bitwise operations. To address the issue of function discontinuity in backward propagation, we use a continuous function to approximate the binarization process. Two popular EEG emotion databases, namely, DEAP and DREAMER, are used for performance evaluation. In comparison to its full-precision counterpart, the Bi-CapsNet achieves a $>\!25\times$reduction on the computational cost and a $>\!5\times$ reduction on the memory usage, while with only a $< $1% drop on the recognition accuracy. Compared to some state-of-the-art EEG emotion recognition methods, the proposed method obtains more competitive performance. In addition, the Bi-CapsNet is implemented on a mobile phone via an open-source binary inference framework named Bolt, and it achieves an $\sim\! 5\times$ inference acceleration in comparison to its full-precision counterpart.
Yu Liu 0023, Chang Li 0001, Juan Cheng 0004, Rencheng Song, Xun Chen 0001
IEEE J. Biomed. Health Informatics6
2023 Hyperspectral Anomaly Detection With Tensor Average Rank and Piecewise Smoothness Constraints
abstract
Anomaly detection in hyperspectral images (HSIs) has attracted considerable interest in the remote-sensing domain, which aims to identify pixels with different spectral and spatial features from their surroundings. Most of the existing anomaly detection methods convert the 3-D data cube to a 2-D matrix composed of independent spectral vectors, which destroys the intrinsic spatial correlation between the pixels and their surrounding pixels, thus leading to considerable degradation in detection performance. In this article, we develop a tensor-based anomaly detection algorithm that can effectively preserve the spatial–spectral information of the original data. We first separate the 3-D HSI data into a background tensor and an anomaly tensor. Then the tensor nuclear norm based on the tensor singular value decomposition (SVD) is exploited to characterize the global low rank existing in both the spectral and spatial directions of the background tensor. In addition, the total variation (TV) regularization is incorporated due to the piecewise smoothness. For the anomaly component, the$l_{2.1}$norm is exploited to promote the group sparsity of anomalous pixels. In order to improve the ability of the algorithm to distinguish the anomaly from the background, we design a robust background dictionary. We first split the HSI data into local clusters by leveraging their spectral similarity and spatial distance. Then we develop a simple but effective way based on the SVD to select representative pixels as atoms. The constructed background dictionary can effectively represent the background materials and eliminate anomalies. Experimental results obtained using several real hyperspectral datasets demonstrate the superiority of the proposed method compared with some state-of-the-art anomaly detection algorithms.
Jun Liu 0004, Xun Chen 0001, Wei Li 0032, Hongbin Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution Reconstruction
abstract
Magnetic resonance imaging (MRI) with high resolution (HR) provides more detailed information for accurate diagnosis and quantitative image analysis. Despite the significant advances, most existing super-resolution (SR) reconstruction network for medical images has two flaws: 1) All of them are designed in a black-box principle, thus lacking sufficient interpretability and further limiting their practical applications. Interpretable neural network models are of significant interest since they enhance the trustworthiness required in clinical practice when dealing with medical images. 2) most existing SR reconstruction approaches only use a single contrast or use a simple multi-contrast fusion mechanism, neglecting the complex relationships between different contrasts that are critical for SR improvement. To deal with these issues, in this paper, a novel Model-Guided interpretable Deep Unfolding Network (MGDUN) for medical image SR reconstruction is proposed. The Model-Guided image SR reconstruction approach solves manually designed objective functions to reconstruct HR MRI. We show how to unfold an iterative MGDUN algorithm into a novel model-guided deep unfolding network by taking the MRI observation matrix and explicit multi-contrast relationship matrix into account during the end-to-end optimization. Extensive experiments on the multi-contrast IXI dataset and BraTs 2019 dataset demonstrate the superiority of our proposed model.
Li Zhang 0104, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Zhiwei Xiong, Feng Wu 0001
ACM Multimedia5
2022 Multi-channel EEG-based emotion recognition in the presence of noisy labels
Chang Li 0001, Yimeng Hou, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
Sci. China Inf. Sci.6
2022 DSG-Fusion: Infrared and visible image fusion via generative adversarial networks and guided filter
Hongtao Huo, Jing Li 0040, Chang Li 0001, Xun Chen 0001
Expert Syst. Appl.6
2022 Superpixel-Based Noise-Robust Sparse Unmixing of Hyperspectral Image
abstract
Sparse unmixing (SU) of hyperspectral image (HSI), as a semisupervised approach, aims to find the optimal subset of the spectral library known in advance to represent each pixel in HSI. However, most of the existing SU methods cannot take full advantage of spatial information and mixed noise in HSI. To this end, we propose a superpixel-based noise-robust SU method (SNRSU) in the presence of mixed noise. First, we perform superpixel segmentation (SS) on the first principal component of HSI to extract the homogeneous regions. Then, we unmix each superpixel based on sparse representation (SR) and low-rank representation (LRR) in the maximuma posterioriframework, which can make full use of the spatial–spectral information in HSI under complex mixed noise. A number of experiments on simulated and real HSI datasets confirm the superior performance of the proposed SNRSU both qualitatively and quantitatively.
Chang Li 0001, Chenhong Sui, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
IEEE Geosci. Remote. Sens. Lett.6
2022 SF-Net: A Multi-Task Model for Brain Tumor Segmentation in Multimodal MRI via Image Fusion
abstract
Automatic segmentation of brain tumor regions from multimodal MRI scans is of great clinical significance. In this letter, we propose a “Segmentation-Fusion” multi-task model named SF-Net for brain tumor segmentation. In comparison to the widely-used multi-task model that adds a variational autoencoder (VAE) decoder to reconstruct the input data, using image fusion as an additional regularization for feature learning helps to achieve more sufficient fusion of multimodal features, which is beneficial to the multimodal image segmentation problem. To further improve the performance of the multi-task model, an uncertainty-based approach that can adaptively adjust the loss weights of different tasks during the training process is introduced for model training. Experimental results on the BraTS 2020 benchmark demonstrate that the proposed method can achieve higher segmentation accuracy than the VAE-based approach. In addition, as the by-product of the multi-task model, the image fusion results obtained are of high quality on the brain tumor regions. The source code of the proposed method is available at https://github.com/yuliu316316/SF-Net.
Yu Liu 0023, Fuhao Mu, Xun Chen 0001
IEEE Signal Process. Lett.4
2022 Video-Based Heart Rate Measurement Against Uneven Illuminations Using Multivariate Singular Spectrum Analysis
abstract
Spatially uneven illuminations are the dominant interference of video-based heart rate (HR) screening for cooperated subjects in a telehealth service. In this letter, a remote photoplethysmography (rPPG) method is introduced to stably extract pulsatile signals against uneven facial illuminations based on the multivariate singular spectrum analysis (MSSA). This method first divides the facial skins into multiple patches, where the hue channels resistant to light intensity variations are prepared from selected optimal patches. Considering the spatial correlations of heartbeats, the hue signals are then decomposed using the MSSA to reconstruct pulses. Finally, the HR is determined as the one with the highest ratio of energy around the dominant frequency from the first group of MSSA reconstructed signals. Experimental results demonstrate the effectiveness of the proposed method on the in-house BSIPL-rPPG database and the public COHFACE database, where the correlation coefficients of the estimated HRs achieve 0.95 and 0.98, respectively, outperforming those of the comparison methods.
Rencheng Song, Xiaoxue Sun, Juan Cheng 0004, Xuezhi Yang, Xun Chen 0001
IEEE Signal Process. Lett.5
2022 SOM-Net: Unrolling the Subspace-Based Optimization for Solving Full-Wave Inverse Scattering Problems
abstract
In this paper, an unrolling algorithm of the iterative subspace-based optimization method (SOM) is proposed for solving full-wave inverse scattering problems (ISPs). The unrolling network, named SOM-Net, inherently embeds the Lippmann-Schwinger physical model into the design of network structures. The SOM-Net takes the deterministic induced current and the raw permittivity image obtained from back-propagation (BP) as the input. It then updates the induced current and the permittivity successively in sub-network blocks of the SOM-Net by imitating iterations of the SOM. The final output of the SOM-Net is the full predicted induced current, from which the scattered field and the permittivity image can also be deduced analytically. The parameters of the SOM-Net are optimized in a supervised manner with the total loss to simultaneously ensure the consistency of the induced current, the scattered field, and the permittivity in the governing equations. Numerical tests on both synthetic and experimental data verify the superior performance of the proposed SOM-Net over typical ones. The results on challenging examples like scatterers with tough profiles or high permittivity demonstrate the good generalization ability of the SOM-Net. With the use of deep unrolling technology, this work builds a bridge between traditional iterative methods and deep learning methods for solving ISPs.
Yu Liu 0023, Rencheng Song, Xudong Chen 0001, Chang Li 0001, Xun Chen 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 Deep Fourier Ranking Quantization for Semi-Supervised Image Retrieval
abstract
To reduce the extreme label dependence of supervised product quantization methods, the semi-supervised paradigm usually employs massive unlabeled data to assist in regularizing deep networks, thereby improving model performance. However, the existing method focuses on the overall distribution consistency between unlabeled data and class prototypes, while ignoring subtle individual variances between unlabeled instances. Therefore, the local neighborhood structure is not fully explored, which will cause the model to easily overfit in the training set. In this paper, we introduce a new Fourier perspective to alleviate this issue by exploring the semantic relations between unlabeled instances in a self-supervised manner. Specifically, based on Fourier Transform, we first design a Phase Mixing (PM) strategy, which can manipulate the mixing area and values of the phase component between two images to control the proportion of semantic information. In this way, we can construct multi-level similarity neighbors naturally for unlabeled data. Then, a ranking quantization loss is formulated to perceive multi-level semantic variances in neighbor instances, which improves the robustness and generalization of the model. Extensive experiments in three different semi-supervised settings show that our method outperforms existing state-of-the-art methods by averaged 3.95% improvement on four datasets.
Pandeng Li, Hongtao Xie 0001, Shaobo Min, Jiannan Ge, Xun Chen 0001, Yongdong Zhang 0001
IEEE Trans. Image Process.5
2022 An Effective Photoplethysmography Heart Rate Estimation Framework Integrating Two-Level Denoising Method and Heart Rate Tracking Algorithm Guided by Finite State Machine
abstract
In order to achieve accurate heart rate (HR) estimation in complex scenes, this paper presents an effective photoplethysmography (PPG) HR estimation framework integrating two-level denoising method and HR tracking algorithm guided by finite state machine (FSM). Aiming at solving the problems of low signal-to-noise ratio and co-frequency (the noise frequency is close to the HR frequency) caused by motion artifacts, the two-level denoising method consisting of the cascaded adaptive filtering and the differential denoising guided by FSM are designed to remove motion-related noises in PPG signals. In order to solve the problem of HR tracking error caused by poor wrist contact, the HR tracking algorithm guided by FSM is proposed to obtain the global optimization capability. The results of HR estimation experiments conducted on the IEEE Signal Processing Cup database and the WeData database created by ourselves show that the proposed framework can effectively cope with the problems of low signal-to-noise ratio and co-frequency. Even if tracking errors occur due to poor wristband contact, the proposed HR tracking algorithm guided by FSM can correct them in time when the HR component appears again. The average absolute error of HR estimation on the two databases are 1.76 BPM (beats per minute) and 2.77 BPM, respectively, which is more accurate compared to other algorithms.
Jingbin Guo, Xiang Chen 0004, Xu Zhang 0002, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2022 A Joint Constrained CCA Model for Network-Dependent Brain Subregion Parcellation
abstract
Connectivity-based brain region parcellation from functional magnetic resonance imaging (fMRI) data is complicated by heterogeneity among aged and diseased subjects, particularly when the data are spatially transformed to a common space. Here, we propose a group-guided functional brain region parcellation model capable of obtaining subregions from a target region with consistent connectivity profiles across multiple subjects, even when the fMRI signals are kept in their native spaces. The model is based on a joint constrained canonical correlation analysis (JC-CCA) method that achieves group-guided parcellation while allowing the data dimension of the parcellated regions for each subject to vary. We performed extensive experiments on synthetic and real data to demonstrate the superiority of the proposed model compared to other classical methods. When applied to fMRI data of subjects with and without Parkinson's disease (PD) to estimate the subregions in the Putamen, significant between-group differences were found in the derived subregions and the connectivity patterns. Superior classification and regression results were obtained, demonstrating its potential in clinical practice.
Qinrui Ling, Aiping Liu, Yu Li 0027, Xueyang Fu, Xun Chen 0001, Martin J. McKeown, Feng Wu 0001
IEEE J. Biomed. Health Informatics5
2022 A Novel SSA-CCA Framework forMuscle Artifact Removal from Ambulatory EEG
abstract
Electroencephalography (EEG) has gained popularity in various types of biomedical applications as a signal source that can be easily acquired and conveniently analyzed. However, owing to a complex scalp electrical environment, EEG is often polluted by diverse artifacts, with electromyography artifacts being the most difficult to remove. In particular, for ambulatory EEG devices with a restricted number of channels, dealing with muscle artifacts is a challenge. In this study, we propose a simple but effective novel scheme that combines singular spectrum analysis (SSA) and canonical correlation analysis (CCA) algorithms for single-channel problems and then extend it to a fewchannel case by adding additional combining and dividing operations to channels. We evaluated our proposed framework on both semi-simulated and real-life data and compared it with some state-of-theart methods. The results demonstrate this novel framework's superior performance in both single-channel and few-channel cases. This promising approach, based on its effectiveness and low time cost, is suitable for real-world biomedical signal processing applications.
Yuheng Feng, Qingze Liu, Aiping Liu, Ruobing Qian, Xun Chen 0001
Virtual Real. Intell. Hardw.5
2021 Multi-view 3D Reconstruction with Transformers
abstract
Deep CNN-based methods have so far achieved the state of the art results in multi-view 3D object reconstruction. Despite the considerable progress, the two core modules of these methods - view feature extraction and multi-view fusion, are usually investigated separately, and the relations among multiple input views are rarely explored. Inspired by the recent great success in Transformer models, we reformulate the multi-view 3D reconstruction as a sequence-to-sequence prediction problem and propose a framework named 3D Volume Transformer. Unlike previous CNN-based methods using a separate design, we unify the feature extraction and view fusion in a single Transformer network. A natural advantage of our design lies in the exploration of view-to-view relationships using self-attention among multiple unordered inputs. On ShapeNet - a large-scale 3D reconstruction benchmark, our method achieves a new state-of-the-art accuracy in multi-view reconstruction with fewer parameters (70% less) than CNN-based methods. Experimental results also suggest the strong scaling capability of our method. Our code will be made publicly available.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Zhengxia Zou, Tianyang Shi, Tim Salcudean, Z. Jane Wang 0001, Rabab K. Ward
ICCV3
2021 Metric learning for novel motion rejection in high-density myoelectric pattern recognition
Le Wu 0003, Xu Zhang 0002, Xiang Chen 0004, Xun Chen 0001
Knowl. Based Syst.5
2021 Constrained independent vector extraction of quasi-periodic signals from multiple data sets
Rencheng Song, Juan Cheng 0004, Aiping Liu, Chang Li 0001, Xun Chen 0001
Signal Process.6
2021 Different Input Resolutions and Arbitrary Output Resolution: A Meta Learning-Based Deep Framework for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion has gained ever-increasing attention in recent years due to its great significance in a variety of vision-based applications. However, existing fusion methods suffer from some limitations in terms of the spatial resolutions of both input source images and output fused image, which prevents their practical usage to a great extent. In this paper, we propose a meta learning-based deep framework for the fusion of infrared and visible images. Unlike most existing methods, the proposed framework can accept the source images of different resolutions and generate the fused image of arbitrary resolution just with a single learned model. In the proposed framework, the features of each source image are first extracted by a convolutional network and upscaled by a meta-upscale module with an arbitrary appropriate factor according to practical requirements. Then, a dual attention mechanism-based feature fusion module is developed to combine features from different source images. Finally, a residual compensation module, which can be iteratively adopted in the proposed framework, is designed to enhance the capability of our method in detail extraction. In addition, the loss function is formulated in a multi-task learning manner via simultaneous fusion and super-resolution, aiming to improve the effect of feature learning. And, a new contrast loss inspired by a perceptual contrast enhancement approach is proposed to further improve the contrast of the fused image. Extensive experiments on widely-used fusion datasets demonstrate the effectiveness and superiority of the proposed method. The code of the proposed method is publicly available at https://github.com/yuliu316316/MetaLearning-Fusion.
Huafeng Li 0001, Yueliang Cen, Yu Liu 0023, Xun Chen 0001, Zhengtao Yu 0001
IEEE Trans. Image Process.4
2021 Interpreting Bottom-Up Decision-Making of CNNs via Hierarchical Inference
abstract
With the great success of convolutional neural networks (CNNs), interpretation of their internal network mechanism has been increasingly critical, while the network decision-making logic is still an open issue. In the bottom-up hierarchical logic of neuroscience, the decision-making process can be deduced from a series of sub-decision-making processes from low to high levels. Inspired by this, we propose the Concept-harmonized HierArchical INference (CHAIN) interpretation scheme. In CHAIN, a network decision-making process from shallow to deep layers is interpreted by the hierarchical backward inference based on visual concepts from high to low semantic levels. Firstly, we learned a general hierarchical visual-concept representation in CNN layered feature space by concept harmonizing model on a large concept dataset. Secondly, for interpreting a specific network decision-making process, we conduct the concept-harmonized hierarchical inference backward from the highest to the lowest semantic level. Specifically, the network learning for a target concept at a deeper layer is disassembled into that for concepts at shallower layers. Finally, a specific network decision-making process is explained as a form of concept-harmonized hierarchical inference, which is intuitively comparable to the bottom-up hierarchical visual recognition way. Quantitative and qualitative experiments demonstrate the effectiveness of the proposed CHAIN at both instance and class levels.
Dan Wang 0011, Xinrui Cui, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Image Process.3
2021 Hand Gesture Recognition based on Surface Electromyography using Convolutional Neural Network with Transfer Learning Method
abstract
This paper presents an effective transfer learning (TL) strategy for the realization of surface electromyography (sEMG)-based gesture recognition with high generalization and low training burden. To realize the idea of taking a well-trained model as the feature extractor of the target networks, 30 hand gestures involving various states of finger joints, elbow joint and wrist joint are selected to compose the source task, and a convolutional neural network (CNN)-based source network is designed and trained as the general gesture EMG feature extraction network. Then, two types of target networks, in the forms of CNN-only and CNN+LSTM (long short-term memory) respectively, are designed with the same CNN architecture as the feature extraction network. Finally, gesture recognition experiments on three different target gesture datasets are carried out under TL and Non-TL strategies respectively. The experimental results verify the validity of the proposed TL strategy in improving hand gesture recognition accuracy and reducing training burden. For both the CNN-only and the CNN+LSTM target networks, on the three target datasets from new users, new gestures and different collection scheme, the proposed TL strategy improves the recognition accuracy by 10%∼38%, reduces the training time to tens of times, and guarantees the recognition accuracy of more than 90% when only 2 repetitions of each gesture are used to fine-tune the parameters of target networks. The proposed TL strategy has important application value for promoting the development of myoelectric control systems.
Xiang Chen 0004, Yu Li 0027, Ruochen Hu, Xu Zhang 0002, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2021 Emotion Recognition From Multi-Channel EEG via Deep Forest
abstract
Recently, deep neural networks (DNNs) have been applied to emotion recognition tasks based on electroencephalography (EEG), and have achieved better performance than traditional algorithms. However, DNNs still have the disadvantages of too many hyperparameters and lots of training data. To overcome these shortcomings, in this article, we propose a method for multi-channel EEG-based emotion recognition using deep forest. First, we consider the effect of baseline signal to preprocess the raw artifact-eliminated EEG signal with baseline removal. Secondly, we construct 2 D frame sequences by taking the spatial position relationship across channels into account. Finally, 2 D frame sequences are input into the classification model constructed by deep forest that can mine the spatial and temporal information of EEG signals to classify EEG emotions. The proposed method can eliminate the need for feature extraction in traditional methods and the classification model is insensitive to hyperparameter settings, which greatly reduce the complexity of emotion recognition. To verify the feasibility of the proposed model, experiments were conducted on two public DEAP and DREAMER databases. On the DEAP database, the average accuracies reach to 97.69% and 97.53% for valence and arousal, respectively; on the DREAMER database, the average accuracies reach to 89.03%, 90.41%, and 89.89% for valence, arousal and dominance, respectively. These results show that the proposed method exhibits higher accuracy than the state-of-art methods.
Juan Cheng 0004, Meiyao Chen, Chang Li 0001, Yu Liu 0023, Rencheng Song, Aiping Liu, Xun Chen 0001
IEEE J. Biomed. Health Informatics7
2021 Striatal Subdivisions Estimated via Deep Embedded Clustering With Application to Parkinson's Disease
abstract
Recent fMRI connectivity-based parcellation (CBP) methods have been developed to obtain homogeneous and functionally coherent brain parcels. However, most of these studies utilize traditional clustering methods that neglect hidden nonlinear features. To enhance parcellation performance, here we propose a deep embedded connectivity-based parcellation (DECBP) framework and apply it to determine functional subdivisions of the striatum in public resting state fMRI data sets. This framework integrates fMRI connectivity features into deep embedded clustering (DEC), a deep neural network based on a stacked autoencoder. Compared to three prevalent clustering methods and their combinations with principal component analysis (PCA), the DECBP exhibited a significantly higher similarity between scans, individuals, and groups, indicating enhanced reproducibility. The generated reliable parcellations were also largely consistent with other public atlases. We further explored the functional subunits in the striatum in a data set from 23 Parkinson's disease (PD) subjects and 27 age-matched healthy controls (HC). All putaminal subregions of PD demonstrated lower interhemispheric connectivity than those of HC, which might reflect imbalance in the pathological progression of PD. Such hypo-connectivity was also observed between putaminal subregions and other brain regions, reflecting neuroimaging manifestations of the altered cortico-striato-thalamo-cortical circuit. These observed weaker couplings were associated with PD severity and duration. Our results support the utilization of the DECBP framework and suggest that abnormal connectivity in putaminal subregions may be a potential indicator of PD.
Yu Li 0027, Aiping Liu, Taomian Mi, Runyu Yang, Piu Chan, Martin J. McKeown, Xun Chen 0001, Feng Wu 0001
IEEE J. Biomed. Health Informatics7
2021 PulseGAN: Learning to Generate Realistic Pulse Waveforms in Remote Photoplethysmography
abstract
Remote photoplethysmography (rPPG) is a non-contact technique for measuring cardiac signals from facial videos. High-quality rPPG pulse signals are urgently demanded in many fields, such as health monitoring and emotion recognition. However, most of the existing rPPG methods can only be used to get average heart rate (HR) values due to the limitation of inaccurate pulse signals. In this paper, a new framework based on generative adversarial network, called PulseGAN, is introduced to generate realistic rPPG pulse signals through denoising the chrominance (CHROM) signals. Considering that the cardiac signal is quasi-periodic and has apparent time-frequency characteristics, the error losses defined in time and spectrum domains are both employed with the adversarial loss to enforce the model generating accurate pulse waveforms as its reference. The proposed framework is tested on three public databases. The results show that the PulseGAN framework can effectively improve the waveform quality, thereby enhancing the accuracy of HR, the interbeat interval (IBI) and the related heart rate variability (HRV) features. The proposed method significantly improves the quality of waveforms compared to the input CHROM signals, with the mean absolute error of AVNN (the average of all normal-to-normal intervals) reduced by 41.19%, 40.45%, 41.63%, and the mean absolute error of SDNN (the standard deviation of all NN intervals) reduced by 37.53%, 44.29%, 58.41%, in the cross-database test on the UBFC-RPPG, PURE, and MAHNOB-HCI databases, respectively. This framework can be easily integrated with other existing rPPG methods to further improve the quality of waveforms, thereby obtaining more reliable IBI features and extending the application scope of rPPG techniques.
Rencheng Song, Juan Cheng 0004, Chang Li 0001, Yu Liu 0023, Xun Chen 0001
IEEE J. Biomed. Health Informatics6
2021 Hip Landmark Detection With Dependency Mining in Ultrasound Image
abstract
Developmental dysplasia of the hip (DDH) is a common and serious disease in infants. Hip landmark detection plays a critical role in diagnosing the development of neonatal hip in the ultrasound image. However, the local confusion and the regional weakening make this task challenging. To solve these challenges, we explore the stable hip structure and the distinguishable local features to provide dependencies for hip landmark detection. In this paper, we propose a novel architecture named Dependency Mining ResNet (DM-ResNet), which investigates end-to-end dependency mining for more accurate and much faster hip landmark detection. First of all, we convert the landmark detection to the heatmap estimation by ResNet to build a strong baseline architecture for fast and accurate detection. Secondly, a dependency mining module is explored to mine the dependencies and leverage both the local and global information to decline the local confusion and strengthen the weakening region. Thirdly, we propose a simple but effective local voting algorithm (LVA) that seeks trade-off between long-range and short-range dependencies in the hip ultrasound image. Besides, a dataset with 2000 annotated hip ultrasound images is constructed in our work. It is the first public hip ultrasound dataset for open research. Experimental results show that our method achieves excellent precision in hip landmark detection (average point error of 0.719mm and successful detection rate within 1mm of 79.9%).
Hongtao Xie 0001, Chuanbin Liu 0001, Xun Chen 0001, Yongdong Zhang 0001
IEEE Trans. Medical Imaging6
2020 ECG-based multi-class arrhythmia detection using spatio-temporal attention-based convolutional recurrent neural network
Jing Zhang 0165, Aiping Liu, Xiang Chen 0004, Xu Zhang 0002, Xun Chen 0001
Artif. Intell. Medicine6
2020 Sparse unmixing of hyperspectral data with bandwise model
Chang Li 0001, Yu Liu 0023, Juan Cheng 0004, Rencheng Song, Jiayi Ma 0001, Chenhong Sui, Xun Chen 0001
Inf. Sci.7
2020 EEG-based emotion recognition using an end-to-end regional-asymmetric convolutional neural network
Heng Cui, Aiping Liu, Xu Zhang 0002, Xiang Chen 0004, Kongqiao Wang, Xun Chen 0001
Knowl. Based Syst.6
2020 Exploring the feasibility of seamless remote heart rate measurement using multiple synchronized cameras
Juan Cheng 0004, Xingmao Wang, Rencheng Song, Yu Liu 0023, Chang Li 0001, Xun Chen 0001
Multim. Tools Appl.6
2020 A Novel Postprocessing Method for Robust Myoelectric Pattern-Recognition Control Through Movement Pattern Transition Detection
abstract
Pattern-recognition-based myoelectric control systems are not yet widely available due to their limited robustness in real-life situations. Some postprocessing methods were introduced to improve the robustness in previous studies, but there is lack of investigation into movement transition phases. This article presents a novel postprocessing method based on movement pattern transition (MPT) detection. An image-based index is used to quantify the similarity of adjacent feature matrices from high-density surface electromyogram (EMG) signals. MPT detection is implemented by applying a double threshold to the calculated index. The proposed postprocessing method is used to rectify the EMG pattern recognition decisions from the classifier by incorporating the detected information. Two representative testing schemes are used to verify the robustness of the proposed method against force level variation and consecutive nonstop task performance. The proposed method achieved mean classification accuracy improvements of 7.33% and 10.91% with respect to the baseline performance of a raw classifier (without any postprocessing) in the two testing schemes. It also outperformed other common postprocessing methods (p <; 0.05). Considering both the accuracy improvement and time efficiency for rapid responses to MPT, the proposed method could be a potential option for postprocessing to enhance the robustness of myoelectric control.
Xu Zhang 0002, Le Wu 0003, Xiang Chen 0004, Xun Chen 0001
IEEE Trans. Hum. Mach. Syst.5
2020 Exploration of Chinese Sign Language Recognition Using Wearable Sensors Based on Deep Belief Net
abstract
In this paper, deep belief net (DBN) was applied into the field of wearable-sensor based Chinese sign language (CSL) recognition. Eight subjects were involved in the study, and all of the subjects finished a five-day experiment performing CSL on a target word set consisting of 150 CSL subwords. During the experiment, surface electromyography (sEMG), accelerometer (ACC), and gyroscope (GYRO) signals were collected from the participants. In order to obtain the optimal structure of the network, three different sensor fusion strategies, including data-level fusion, feature-level fusion, and decision-level fusion, were explored. In addition, for the feature-level fusion strategy, two different feature sources, which are hand-crafted features and network generated features, and two different network structures, which are fully-connected net and DBN, were also compared. The result showed that feature level fusion could achieve the best recognition accuracy among the three fusion strategies, and feature-level fusion with network generated features and DBN could achieve the best recognition accuracy. The best recognition accuracy realized in this study was 95.1% for the user-dependent test and 88.2% for the user-independent test. The significance of the study is that it applied the deep learning method into the field of wearable sensors-based CSL recognition, and according to our knowledge it's the first study comparing human engineered features with the network generated features in the correspondent field. The results from the study shed lights on the method of using network-generated features during sensor fusion and CSL recognition.
Xiang Chen 0004, Xu Zhang 0002, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2020 Zero-Shot Learning Based on Deep Weighted Attribute Prediction
abstract
In zero-shot learning, attributes play as a bridge from original images to class labels. Therefore, to achieve accurate zero-shot image classification, we mainly focus on improving attribute prediction accuracy by taking full advantage of prior information about attribute from two aspects. First, we present a new attribute classifier called deep attribute prediction (DeepAP) model by using supervised deep convolutional neural networks (DCNNs), where the attribute label information participates in the training of DCNNs. Unlike common DCNNs that are usually used to extract image features, the constructed DCNNs are used to directly predict attribute values from the original input images. Thus, the designed DeepAP model can serve as the mapping from low-level image features to high-level semantic attributes in the traditional direct attribute prediction (DAP) model. Second, another prior information about attribute, i.e., class-attribute matrix is used to mine the attribute-class correlation with sparse representation coefficients. Since the attribute-class correlation can reflect different contributions of attributes to classification, we use it to define attribute weights and incorporate the idea of weighted attributes into DeepAP to form the deep weighted attribute prediction (DWAP) model. Experiments on three real datasets show that DWAP outperforms the deep attribute network and DAP on attribute prediction and zero-shot image classification.
Xuesong Wang 0001, Chen Chen 0033, Yuhu Cheng 0001, Xun Chen 0001, Yu Liu 0023
IEEE Trans. Syst. Man Cybern. Syst.4
2019 Medical Image Fusion via Convolutional Sparsity Based Morphological Component Analysis
abstract
In this letter, a sparse representation (SR) model named convolutional sparsity based morphological component analysis (CS-MCA) is introduced for pixel-level medical image fusion. Unlike the standard SR model, which is based on single image component and overlapping patches, the CS-MCA model can simultaneously achieve multi-component and global SRs of source images, by integrating MCA and convolutional sparse representation (CSR) into a unified optimization framework. For each source image, in the proposed fusion method, the CSRs of its cartoon and texture components are first obtained by the CS-MCA model using pre-learned dictionaries. Then, for each image component, the sparse coefficients of all the source images are merged and the fused component is accordingly reconstructed using the corresponding dictionary. Finally, the fused image is calculated as the superposition of the fused cartoon and texture components. Experimental results demonstrate that the proposed method can outperform some benchmarking and state-of-the-art SR-based fusion methods in terms of both visual perception and objective assessment.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.2
2019 Sparse Group Representation Model for Motor Imagery EEG Classification
abstract
A potential limitation of a motor imagery (MI) based brain-computer interface (BCI) is that it usually requires a relatively long time to record sufficient electroencephalogram (EEG) data for robust classifier training. The calibration burden during data acquisition phase will most probably cause a subject to be reluctant to use a BCI system. To alleviate this issue, we propose a novel sparse group representation model (SGRM) for improving the efficiency of MI-based BCI by exploiting the intersubject information. Specifically, preceded by feature extraction using common spatial pattern, a composite dictionary matrix is constructed with training samples from both the target subject and other subjects. By explicitly exploiting within-group sparse and group-wise sparse constraints, the most compact representation of a test sample of the target subject is then estimated as a linear combination of columns in the dictionary matrix. Classification is implemented by calculating the class-specific representation residual based on the significant training samples corresponding to the nonzero representation coefficients. Accordingly, the proposed SGRM method effectively reduces the required training samples from the target subject due to auxiliary data available from other subjects. With two public EEG data sets, extensive experimental comparisons are carried out between SGRM and other state-of-the-art approaches. Superior classification performance of our method using 40 trials of the target subject for model calibration (Averaged accuracy = 78.2%, Kappa = 0.57 and Averaged accuracy = 77.7%, Kappa = 0.55 for the two data sets, respectively) indicates its promising potential for improving the practicality of MI-based BCI.
Yong Jiao, Yu Zhang 0009, Xun Chen 0001, Erwei Yin, Jing Jin 0001, Xingyu Wang 0004, Andrzej Cichocki
IEEE J. Biomed. Health Informatics3
2018 Bone Age Assessment with X-Ray Images Based on Contourlet Motivated Deep Convolutional Networks
abstract
Bone age assessment (BAA) is a widely performed procedure for skeletal maturity evaluation in pediatric radiology. It has various clinical applications such as diagnosis of endocrine disorders, monitoring of growth hormone therapy and prediction of final adult height for adolescents. Recent studies indicate that deep learning techniques have great potential in developing automated BAA methods with significant improvements in terms of conventional computer-assisted approaches. In this paper, we propose a multi-scale feature fusion framework for bone age assessment based on deep convolutional neural networks. In our method, the non-subsampled contourlet transform (NSCT) is firstly performed on an input left-hand radiograph to obtain its multi-scale and multi-direction representations. Then, the decomposed bands at each scale are fed to a convolutional network that contains a series of convolutional and pooling layers for feature extraction, respectively. Finally, the feature maps from different branches are concatenated and put into a regression network consisting of several fully connected layers to obtain the bone age estimation. Experimental results on a public BAA dataset demonstrate that the proposed method can achieve state-of-the-art performance.
Xun Chen 0001, Chao Zhang 0057, Yu Liu 0023
MMSP1
2017 A medical image fusion method based on convolutional neural networks
abstract
Medical image fusion technique plays an an increasingly critical role in many clinical applications by deriving the complementary information from medical images with different modalities. In this paper, a medical image fusion method based on convolutional neural networks (CNNs) is proposed. In our method, a siamese convolutional network is adopted to generate a weight map which integrates the pixel activity information from two source images. The fusion process is conducted in a multi-scale manner via image pyramids to be more consistent with human visual perception. In addition, a local similarity based strategy is applied to adaptively adjust the fusion mode for the decomposed coefficients. Experimental results demonstrate that the proposed method can achieve promising results in terms of both visual quality and objective assessment.
Yu Liu 0023, Xun Chen 0001, Juan Cheng 0004, Hu Peng
FUSION2
2017 Illumination Variation-Resistant Video-Based Heart Rate Measurement Using Joint Blind Source Separation and Ensemble Empirical Mode Decomposition
abstract
Recent studies have demonstrated that heart rate (HR) could be estimated using video data [e.g., exploring human facial regions of interest (ROIs)] under well-controlled conditions. However, in practice, the pulse signals may be contaminated by motions and illumination variations. In this paper, tackling the illumination variation challenge, we propose an illumination-robust framework using joint blind source separation (JBSS) and ensemble empirical mode decomposition (EEMD) to effectively evaluate HR from webcam videos. The framework takes the hypotheses that both facial ROI and background ROI have similar illumination variations. The background ROI is then considered as a noise reference sensor to denoise the facial signals by using the JBSS technique to extract the underlying illumination variation sources. Further, the reconstructed illumination-resisted green channel of the facial ROI is detrended and decomposed into a number of intrinsic mode functions using EEMD to estimate the HR. Experimental results demonstrated that the proposed framework could estimate HR more accurately than the state-of-the-art methods. The Bland-Altman plots showed that it led to better agreement with HR ground truth with the mean bias 1.15 beats/min (bpm), with 95% limits from -15.43 to 17.73 bpm, and the correlation coefficient 0.53. This study provides a promising solution for realistic noncontact and robust HR measurement applications.
Juan Cheng 0004, Xun Chen 0001, Lingxi Xu, Z. Jane Wang 0001
IEEE J. Biomed. Health Informatics2
2016 Image Fusion With Convolutional Sparse Representation
abstract
As a popular signal modeling technique, sparse representation (SR) has achieved great success in image fusion over the last few years with a number of effective algorithms being proposed. However, due to the patch-based manner applied in sparse coding, most existing SR-based fusion methods suffer from two drawbacks, namely, limited ability in detail preservation and high sensitivity to misregistration, while these two issues are of great concern in image fusion. In this letter, we introduce a recently emerged signal decomposition model known as convolutional sparse representation (CSR) into image fusion to address this problem, which is motivated by the observation that the CSR model can effectively overcome the above two drawbacks. We propose a CSR-based image fusion framework, in which each source image is decomposed into a base layer and a detail layer, for multifocus image fusion and multimodal image fusion. Experimental results demonstrate that the proposed fusion methods clearly outperform the SR-based methods in terms of both objective assessment and visual quality.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.2
2016 Underdetermined Joint Blind Source Separation for Two Datasets Based on Tensor Decomposition
abstract
In this letter, we aim to jointly separate the underdetermined mixtures of latent sources from two datasets, where the number of sources exceeds the number of observations in each dataset. Currently available blind source separation (BSS) methods, including joint blind source separation (JBSS) and underdetermined blind source separation (UBSS), cannot address this underdetermined problem effectively. We exploit the second-order statistics of observations and introduce a novel BSS method, termed as underdetermined joint blind source separation (UJBSS). Considering the dependence information between two datasets, the problem of jointly estimating the mixing matrices is tackled via canonical polyadic (CP) decomposition of a specialized tensor in which a set of spatial covariance matrices are stacked. Furthermore, the estimated mixing matrices are used to recover the sources from each dataset separately. Numerical results demonstrate the competitive performance of the proposed method when compared to a commonly used JBSS method, multiset canonical correlation analysis (MCCA), and the single-set UBSS method, UBSS with free active sources (UBSS-FAS).
Liang Zou, Xun Chen 0001, Z. Jane Wang 0001
IEEE Signal Process. Lett.2
2014 Time varying brain connectivity modeling using FMRI signals
abstract
Inferring brain connectivity networks has been increasingly important for understanding brain functioning. It is suggested that brain is inherently non-stationary and the dynamic patterns of brain networks may provide deeper insights into brain function. However, the majority of current models assume that brain connectivity networks have time invariant structures, neglecting the variability in brain interactions over time. To investigate time varying brain connectivity networks, a stick time varying model is presented in this paper. Simulation results demonstrate that the proposed method could improve the accuracy in estimating time-dependent connectivity patterns. It is also applied to real fMRI data set for studying time-varying resting-state brain connectivity networks.
Aiping Liu, Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
ICASSP2
2014 A Three-Step Multimodal Analysis Framework for Modeling Corticomuscular Activity With Application to Parkinson's Disease
abstract
Corticomuscular coupling analysis based on multiple datasets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding human motor control systems. A popular conventional method to assess corticomuscular coupling has been the pair-wise magnitude-squared coherence (MSC) between EEG and concomitant EMG recordings. However, there are certain limitations associated with the MSC, including the difficulty in robustly assessing group inference, only dealing with two types of datasets simultaneously and the biologically implausible assumption of pair-wise interactions. To overcome such limitations, in this paper, we propose assessing corticomuscular coupling by combining multiset canonical correlation analysis (M-CCA) and joint independent component analysis (jICA). The proposed method takes advantage of the M-CCA and jICA to ensure that the extracted components are maximally correlated across multiple datasets and meanwhile statistically independent within each dataset. Simulations were performed to illustrate the performance of the proposed method. We also applied the proposed method to concurrent EEG, EMG, and behavior data collected in a Parkinson's disease (PD) study. The results reveal highly correlated temporal patterns among the three types of signals and corresponding spatial activation patterns. In addition to the expected motor areas, the corresponding spatial activation patterns demonstrate enhanced occipital connectivity in the PD subjects, consistent with previous medical findings.
Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
IEEE J. Biomed. Health Informatics1
2013 A Joint Multimodal Group Analysis Framework for Modeling Corticomuscular Activity
abstract
Corticomuscular coupling analysis based on multiple data sets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding human motor control systems. Two probably most popular methods are the pair-wise magnitude-squared coherence (MSC) between EEG and simultaneously-recorded EMG signals, and partial least square (PLS). Unfortunately, MSC and PLS generally deal with only two types of data sets at the same time, while we may need to analyze more than two types of data sets. Moreover, it is not straightforward to extend MSC to the group level for combining results across subjects. Also, PLS can have the information mixing problem since only the variations in one data set are used to predict the other data set. To address these concerns, we propose a joint multimodal analysis framework for corticomuscular coupling analysis. The proposed framework models multiple data spaces simultaneously in a multidirectional fashion. Furthermore, to address the inter-subject variability concern in real-world medical applications, we extend the proposed framework from the individual subject level to the group level to obtain common corticomuscular coupling patterns across subjects. We apply the proposed framework to concurrent EEG, EMG and behavior data collected in a Parkinson's disease (PD) study. The results reveal several highly correlated temporal patterns among the three types of signals and their corresponding spatial activation patterns. In PD subjects, there are enhanced connections between occipital region and other regions, which is consistent with the previous medical finding. The proposed framework is a promising technique for performing multi-subject and multi-modal data analysis.
Xun Chen 0001, Xiang Chen 0004, Rabab K. Ward, Z. Jane Wang 0001
IEEE Trans. Multim.1
2012 A tridirectional method for corticomuscular coupling analysis in Parkinson's disease
abstract
Corticomuscular coupling analysis based on multiple datasets such as electroencephalography (EEG) and electromyography (EMG) signals provides a useful tool for understanding the underlying mechanisms of human motor control systems. In this work, we propose a tridirectional statistical modeling and analysis method to identify the coupling relationships between three types of datasets. Different from conventional approaches where only two datasets are considered and the interest is to interpret one dataset by another in a unidirectional fashion, the goal in this paper is to model three data spaces simultaneously in a tridirectional fashion. To address the intersubject variability concern in real-world medical applications, we further propose a group analysis framework based on the proposed method and apply it to concurrent EEG, EMG and behavior signals collected from 8 normal subjects and 9 patients with Parkinson's disease (PD) performing a dynamic motor task. The results demonstrate highly correlated temporal patterns among the three types of signals and meaningful spatial activation patterns. The proposed approach is a promising technique for performing multi-subject and multi-modal data analysis.
Xun Chen 0001, Z. Jane Wang 0001, Martin J. McKeown
MMSP1
2012 A P300-based BCI classification algorithm using median filtering and Bayesian feature extraction
abstract
A brain computer interface (BCI) system translates a person's brain activity into useful control or communication signals. In this paper, an effective P300-based BCI identification algorithm using median filtering and Bayesian classifier is proposed to improve the classification accuracy and computation efficiency of P300-based BCI. Median filtering is firstly applied to remove noises and Bayesian Linear Discriminant Analysis (BLDA) is then employed for classification. Testing on the P300 speller paradigm in dataset II of 2004 BCI Competition III, we show that a 90% average classification accuracy can be achieved and the highest accuracy is 100%. The proposed method is also computationally efficient and thus it represents a practical implementation for man-computer communication control, especially for on-line applications.
Xun Chen 0001, Rabab K. Ward
MMSP3