EDBT 2026 Demo / reviewers in the wild / expert
Juan Cheng 0004
dblp:49/6464-4
· DBLP profile ↗
22ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-1206-1698ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LST-rPPG: A long-range spatio-temporal model for high-accuracy heart rate variability measurement
Jiajie Li 0011, Juan Cheng 0004, Rencheng Song, Yu Liu 0023 |
Expert Syst. Appl. | 2 |
| 2026 | Video-Based Instantaneous Heart Rate Measurement With Enhanced Time-Frequency RepresentationsabstractRemote photoplethysmography (rPPG) for heart rate (HR) measurement based on facial videos has recently attracted increasing attention. However, most existing methods focus on average heart rate (AHR) over a period rather than instantaneous heart rate (IHR), which better reflects physical and mental states. To address this issue, we propose a novel rPPG-based method for measuring IHR values from facial videos. Our method employs the wavelet synchrosqueezed transform (WSST) to generate time-frequency representations (TFRs) of chrominance (CHROM) signals from multiple facial regions of interest (ROIs), synchronously reflecting the IHR during a video segment. Furthermore, the TransUNet is introduced to refine these TFR images, enhancing the ridge line information related to IHRs. Comprehensive comparisons and ablation studies on four public datasets (UBFC-rPPG, PURE, UBFC-Phys, and MMPD) reveal that our WSST-UNet method achieves superior performance over several typical rPPG methods, achieving mean absolute errors (MAE) of 2.34 beats per minute (bpm), 1.29 bpm, 5.03 bpm, and 6.58 bpm, respectively. The proposed method offers a promising solution for practical application in video-based IHR measurements. Juan Cheng 0004, Xiwen Luo, Rencheng Song, Yu Liu 0023 |
IEEE Trans. Multim. | 1 |
| 2025 | MambaDiff: Mamba-Enhanced Diffusion Model for 3D Medical Image SegmentationabstractAccurate 3D medical image segmentation is crucial for diagnosis and treatment. Diffusion models demonstrate promising performance in medical image segmentation tasks due to the progressive nature of the generation process and the explicit modeling of data distributions. However, the weak guidance of conditional information and insufficient feature extraction in diffusion models lead to the loss of fine-grained features and structural consistency in the segmentation results, thereby affecting the accuracy of medical image segmentation. To address this challenge, we propose a Mamba-Enhanced Diffusion Model for 3D Medical Image Segmentation. We extract multilevel semantic features from the original images using an encoder and tightly integrate them with the denoising process of the diffusion model through a Semantic Hierarchical Embedding (SHE) mechanism, to capture the intricate relationship between the noisy label and image data. Meanwhile, we design a Global-Slice Perception Mamba (GSPM) layer, which integrates multi-dimensional perception mechanisms to endow the model with comprehensive spatial reasoning and feature extraction capabilities. Experimental results show that our proposed MambaDiff achieves more competitive performance compared to prior arts with substantially fewer parameters on four public medical image segmentation datasets including BraTS 2021, BraTS 2024, LiTS and MSD Hippocampus. The source code of our method is available at https://github.com/yuliu316316/MambaDiff. Yu Liu 0023, Juan Cheng 0004, Haolin Zhan, Zhiqin Zhu |
IEEE Trans. Image Process. | 3 |
| 2025 | VDMUFusion: A Versatile Diffusion Model-Based Unsupervised Framework for Image FusionabstractImage fusion facilitates the integration of information from various source images of the same scene into a composite image, thereby benefiting perception, analysis, and understanding. Recently, diffusion models have demonstrated impressive generative capabilities in the field of computer vision, suggesting significant potential for application in image fusion. The forward process in the diffusion models requires the gradual addition of noise to the original data. However, typical unsupervised image fusion tasks (e.g., infrared-visible, medical, and multi-exposure image fusion) lack ground truth images (corresponding to the original data in diffusion models), thereby preventing the direct application of the diffusion models. To address this problem, we propose a versatile diffusion model-based unsupervised framework for image fusion, termed as VDMUFusion. In the proposed method, we integrate the fusion problem into the diffusion sampling process by formulating image fusion as a weighted average process and establishing appropriate assumptions about the noise in the diffusion model. To simplify the training process, we propose a multi-task learning framework that replaces the original noise prediction network, allowing for simultaneous prediction of noise and fusion weights. Meanwhile, our method employs joint training across various fusion tasks, which significantly improves noise prediction accuracy and yields higher quality fused images compared to training on a single task. Extensive experimental results demonstrate that the proposed method delivers very competitive performance across various image fusion tasks. The code is available at https://github.com/yuliu316316/VDMUFusion. Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Rethinking the Effectiveness of Objective Evaluation Metrics in Multi-Focus Image Fusion: A Statistic-Based ApproachabstractAs an effective technique to extend the depth-of-field (DOF) of optical lenses, multi-focus image fusion has recently become an active topic in image processing community. However, a major problem remaining unsolved in this field is the lack of universal criteria in selecting objective evaluation metrics. Consequently, the metrics utilized in different studies often vary significantly, leading to high difficulties in achieving unbiased evaluation. To address this problem, this paper proposes a statistic-based approach for verifying the effectiveness of objective metrics in multi-focus image fusion. The core idea is to adopt statistical correlation measures to evaluate the performance consistency between a certain fusion metric and some popular full-reference image quality assessment models. In addition, a convolutional neural network (CNN)-based fusion metric is presented to measure the similarity between the source images and the fused image based on the semantic features at multiple abstraction levels. A comparative study is conducted to evaluate 20 existing fusion metrics using the proposed statistic-based approach on a large-scale, realistic and with-ground-truth multi-focus image fusion dataset recently released. Experimental results demonstrate the feasibility of the proposed approach in evaluating the effectiveness of objective metrics and the advantage of our CNN-based metric. Yu Liu 0023, Zhengzheng Qi, Juan Cheng 0004, Xun Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | MM-Net: A MixFormer-Based Multi-Scale Network for Anatomical and Functional Image FusionabstractAnatomical and functional image fusion is an important technique in a variety of medical and biological applications. Recently, deep learning (DL)-based methods have become a mainstream direction in the field of multi-modal image fusion. However, existing DL-based fusion approaches have difficulty in effectively capturing local features and global contextual information simultaneously. In addition, the scale diversity of features, which is a crucial issue in image fusion, often lacks adequate attention in most existing works. In this paper, to address the above problems, we propose a MixFormer-based multi-scale network, termed as MM-Net, for anatomical and functional image fusion. In our method, an improved MixFormer-based backbone is introduced to sufficiently extract both local features and global contextual information at multiple scales from the source images. The features from different source images are fused at multiple scales based on a multi-source spatial attention-based cross-modality feature fusion (CMFF) module. The scale diversity of the fused features is further enriched by a series of multi-scale feature interaction (MSFI) modules and feature aggregation upsample (FAU) modules. Moreover, a loss function consisting of both spatial domain and frequency domain components is devised to train the proposed fusion model. Experimental results demonstrate that our method outperforms several state-of-the-art fusion methods on both qualitative and quantitative comparisons, and the proposed fusion model exhibits good generalization capability. The source code of our fusion method will be available at https://github.com/yuliu316316. Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | CCSR-Net: Unfolding Coupled Convolutional Sparse Representation for Multi-focus Image Fusion
Kecheng Zheng, Juan Cheng 0004, Yu Liu 0023 |
PRCV (10) | 2 |
| 2023 | Multi-Exposure Image Fusion via Multi-Scale and Context-Aware Feature LearningabstractIn this letter, a deep learning (DL)-based multi-exposure image fusion (MEF) method via multi-scale and context-aware feature learning is proposed, aiming to overcome the defects of existing traditional and DL-based methods. The proposed network is based on an auto-encoder architecture. First, an encoder that combines the convolutional network and Transformer is designed to extract multi-scale features and capture the global contextual information. Then, a multi-scale feature interaction (MSFI) module is devised to enrich the scale diversity of extracted features using cross-scale fusion and Atrous spatial pyramid pooling (ASPP). Finally, a decoder with a nest connection architecture is introduced to reconstruct the fused image. Experimental results show that the proposed method outperforms several representative traditional and DL-based MEF methods in terms of both visual quality and objective assessment. Yu Liu 0023, Juan Cheng 0004, Xun Chen 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | EEG-Based Emotion Recognition via Neural Architecture SearchabstractWith the flourishing development of deep learning (DL) and the convolution neural network (CNN), electroencephalogram-based (EEG) emotion recognition is occupying an increasingly crucial part in the field of brain-computer interface (BCI). However, currently employed architectures have mostly been designed manually by human experts, which is a time-consuming and labor-intensive process. In this paper, we proposed a novel neural architecture search (NAS) framework based on reinforcement learning (RL) for EEG-based emotion recognition, which can automatically design network architectures. The proposed NAS mainly contains three parts: search strategy, search space, and evaluation strategy. During the search process, a recurrent network (RNN) controller is used to select the optimal network structure in the search space. We trained the controller with RL to maximize the expected reward of the generated models on a validation set and force parameter sharing among the models. We evaluated the performance of NAS on the DEAP and DREAMER dataset. On the DEAP dataset, the average accuracies reached 97.94%, 97.74%, and 97.82% on arousal, valence, and dominance respectively. On the DREAMER dataset, average accuracies reached 96.62%, 96.29% and 96.61% on arousal, valence, and dominance, respectively. The experimental results demonstrated that the proposed NAS outperforms the state-of-the-art CNN-based methods. Chang Li 0001, Zhongzhen Zhang, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | EEG-Based Emotion Recognition via Channel-Wise Attention and Self AttentionabstractEmotion recognition based on electroencephalography (EEG) is a significant task in the brain-computer interface field. Recently, many deep learning-based emotion recognition methods are demonstrated to outperform traditional methods. However, it remains challenging to extract discriminative features for EEG emotion recognition, and most methods ignore useful information in channel and time. This article proposes an attention-based convolutional recurrent neural network (ACRNN) to extract more discriminative features from EEG signals and improve the accuracy of emotion recognition. First, the proposed ACRNN adopts a channel-wise attention mechanism to adaptively assign the weights of different channels, and a CNN is employed to extract the spatial information of encoded EEG signals. Then, to explore the temporal information of EEG signals, extended self-attention is integrated into an RNN to recode the importance based on intrinsic similarity in EEG signals. We conducted extensive experiments on the DEAP and DREAMER databases. The experimental results demonstrate that the proposed ACRNN outperforms state-of-the-art methods. Chang Li 0001, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Feng Wan 0003, Xun Chen 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | MSCAF-Net: A General Framework for Camouflaged Object Detection via Learning Multi-Scale Context-Aware FeaturesabstractThe aim of camouflaged object detection (COD) is to find objects that are hidden in their surrounding environment. Due to the factors like low illumination, occlusion, small size and high similarity to the background, COD is recognized to be a very challenging task. In this paper, we propose a general COD framework, termed as MSCAF-Net, focusing on learning multi-scale context-aware features. To achieve this target, we first adopt the improved Pyramid Vision Transformer (PVTv2) model as the backbone to extract global contextual information at multiple scales. An enhanced receptive field (ERF) module is then designed to refine the features at each scale. Further, a cross-scale feature fusion (CSFF) module is introduced to achieve sufficient interaction of multi-scale information, aiming to enrich the scale diversity of extracted features. In addition, inspired the mechanism of the human visual system, a dense interactive decoder (DID) module is devised to output a rough localization map, which is used to modulate the fused features obtained in the CSFF module for more accurate detection. The effectiveness of our MSCAF-Net is validated on four benchmark datasets. The results show that the proposed method significantly outperforms state-of-the-art (SOTA) COD models by a large margin. Besides, we also investigate the potential of our MSCAF-Net on some other vision tasks that are highly related to COD, such as polyp segmentation, COVID-19 lung infection segmentation, transparent object detection and defect detection. Experimental results demonstrate the high versatility of the proposed MSCAF-Net. The source code and results of our method are available athttps://github.com/yuliu316316/MSCAF-COD. Yu Liu 0023, Haihang Li, Juan Cheng 0004, Xun Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Bi-CapsNet: A Binary Capsule Network for EEG-Based Emotion RecognitionabstractIn recent years, deep learning has gained widespread attention in electroencephalogram (EEG)-based emotion recognition. However, deep learning methods are usually time-consuming with a large amount of memory usage, which obstructs their practical usage on resource-constrained devices. In this paper, we propose a binary capsule network (Bi-CapsNet) for EEG emotion recognition with low computational cost and memory usage. The Bi-CapsNet binarizes 32-bit weights and activations to 1 b, and replaces floating-point operations with efficient bitwise operations. To address the issue of function discontinuity in backward propagation, we use a continuous function to approximate the binarization process. Two popular EEG emotion databases, namely, DEAP and DREAMER, are used for performance evaluation. In comparison to its full-precision counterpart, the Bi-CapsNet achieves a $>\!25\times$reduction on the computational cost and a $>\!5\times$ reduction on the memory usage, while with only a $< $1% drop on the recognition accuracy. Compared to some state-of-the-art EEG emotion recognition methods, the proposed method obtains more competitive performance. In addition, the Bi-CapsNet is implemented on a mobile phone via an open-source binary inference framework named Bolt, and it achieves an $\sim\! 5\times$ inference acceleration in comparison to its full-precision counterpart. Yu Liu 0023, Chang Li 0001, Juan Cheng 0004, Rencheng Song, Xun Chen 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Multi-channel EEG-based emotion recognition in the presence of noisy labels
Chang Li 0001, Yimeng Hou, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001 |
Sci. China Inf. Sci. | 4 |
| 2022 | Superpixel-Based Noise-Robust Sparse Unmixing of Hyperspectral ImageabstractSparse unmixing (SU) of hyperspectral image (HSI), as a semisupervised approach, aims to find the optimal subset of the spectral library known in advance to represent each pixel in HSI. However, most of the existing SU methods cannot take full advantage of spatial information and mixed noise in HSI. To this end, we propose a superpixel-based noise-robust SU method (SNRSU) in the presence of mixed noise. First, we perform superpixel segmentation (SS) on the first principal component of HSI to extract the homogeneous regions. Then, we unmix each superpixel based on sparse representation (SR) and low-rank representation (LRR) in the maximuma posterioriframework, which can make full use of the spatial–spectral information in HSI under complex mixed noise. A number of experiments on simulated and real HSI datasets confirm the superior performance of the proposed SNRSU both qualitatively and quantitatively. Chang Li 0001, Chenhong Sui, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Video-Based Heart Rate Measurement Against Uneven Illuminations Using Multivariate Singular Spectrum AnalysisabstractSpatially uneven illuminations are the dominant interference of video-based heart rate (HR) screening for cooperated subjects in a telehealth service. In this letter, a remote photoplethysmography (rPPG) method is introduced to stably extract pulsatile signals against uneven facial illuminations based on the multivariate singular spectrum analysis (MSSA). This method first divides the facial skins into multiple patches, where the hue channels resistant to light intensity variations are prepared from selected optimal patches. Considering the spatial correlations of heartbeats, the hue signals are then decomposed using the MSSA to reconstruct pulses. Finally, the HR is determined as the one with the highest ratio of energy around the dominant frequency from the first group of MSSA reconstructed signals. Experimental results demonstrate the effectiveness of the proposed method on the in-house BSIPL-rPPG database and the public COHFACE database, where the correlation coefficients of the estimated HRs achieve 0.95 and 0.98, respectively, outperforming those of the comparison methods. Rencheng Song, Xiaoxue Sun, Juan Cheng 0004, Xuezhi Yang, Xun Chen 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Constrained independent vector extraction of quasi-periodic signals from multiple data sets
Rencheng Song, Juan Cheng 0004, Aiping Liu, Chang Li 0001, Xun Chen 0001 |
Signal Process. | 3 |
| 2021 | Emotion Recognition From Multi-Channel EEG via Deep ForestabstractRecently, deep neural networks (DNNs) have been applied to emotion recognition tasks based on electroencephalography (EEG), and have achieved better performance than traditional algorithms. However, DNNs still have the disadvantages of too many hyperparameters and lots of training data. To overcome these shortcomings, in this article, we propose a method for multi-channel EEG-based emotion recognition using deep forest. First, we consider the effect of baseline signal to preprocess the raw artifact-eliminated EEG signal with baseline removal. Secondly, we construct 2 D frame sequences by taking the spatial position relationship across channels into account. Finally, 2 D frame sequences are input into the classification model constructed by deep forest that can mine the spatial and temporal information of EEG signals to classify EEG emotions. The proposed method can eliminate the need for feature extraction in traditional methods and the classification model is insensitive to hyperparameter settings, which greatly reduce the complexity of emotion recognition. To verify the feasibility of the proposed model, experiments were conducted on two public DEAP and DREAMER databases. On the DEAP database, the average accuracies reach to 97.69% and 97.53% for valence and arousal, respectively; on the DREAMER database, the average accuracies reach to 89.03%, 90.41%, and 89.89% for valence, arousal and dominance, respectively. These results show that the proposed method exhibits higher accuracy than the state-of-art methods. Juan Cheng 0004, Meiyao Chen, Chang Li 0001, Yu Liu 0023, Rencheng Song, Aiping Liu, Xun Chen 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | PulseGAN: Learning to Generate Realistic Pulse Waveforms in Remote PhotoplethysmographyabstractRemote photoplethysmography (rPPG) is a non-contact technique for measuring cardiac signals from facial videos. High-quality rPPG pulse signals are urgently demanded in many fields, such as health monitoring and emotion recognition. However, most of the existing rPPG methods can only be used to get average heart rate (HR) values due to the limitation of inaccurate pulse signals. In this paper, a new framework based on generative adversarial network, called PulseGAN, is introduced to generate realistic rPPG pulse signals through denoising the chrominance (CHROM) signals. Considering that the cardiac signal is quasi-periodic and has apparent time-frequency characteristics, the error losses defined in time and spectrum domains are both employed with the adversarial loss to enforce the model generating accurate pulse waveforms as its reference. The proposed framework is tested on three public databases. The results show that the PulseGAN framework can effectively improve the waveform quality, thereby enhancing the accuracy of HR, the interbeat interval (IBI) and the related heart rate variability (HRV) features. The proposed method significantly improves the quality of waveforms compared to the input CHROM signals, with the mean absolute error of AVNN (the average of all normal-to-normal intervals) reduced by 41.19%, 40.45%, 41.63%, and the mean absolute error of SDNN (the standard deviation of all NN intervals) reduced by 37.53%, 44.29%, 58.41%, in the cross-database test on the UBFC-RPPG, PURE, and MAHNOB-HCI databases, respectively. This framework can be easily integrated with other existing rPPG methods to further improve the quality of waveforms, thereby obtaining more reliable IBI features and extending the application scope of rPPG techniques. Rencheng Song, Juan Cheng 0004, Chang Li 0001, Yu Liu 0023, Xun Chen 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Sparse unmixing of hyperspectral data with bandwise model
Chang Li 0001, Yu Liu 0023, Juan Cheng 0004, Rencheng Song, Jiayi Ma 0001, Chenhong Sui, Xun Chen 0001 |
Inf. Sci. | 3 |
| 2020 | Exploring the feasibility of seamless remote heart rate measurement using multiple synchronized cameras
Juan Cheng 0004, Xingmao Wang, Rencheng Song, Yu Liu 0023, Chang Li 0001, Xun Chen 0001 |
Multim. Tools Appl. | 1 |
| 2017 | A medical image fusion method based on convolutional neural networksabstractMedical image fusion technique plays an an increasingly critical role in many clinical applications by deriving the complementary information from medical images with different modalities. In this paper, a medical image fusion method based on convolutional neural networks (CNNs) is proposed. In our method, a siamese convolutional network is adopted to generate a weight map which integrates the pixel activity information from two source images. The fusion process is conducted in a multi-scale manner via image pyramids to be more consistent with human visual perception. In addition, a local similarity based strategy is applied to adaptively adjust the fusion mode for the decomposed coefficients. Experimental results demonstrate that the proposed method can achieve promising results in terms of both visual quality and objective assessment. Yu Liu 0023, Xun Chen 0001, Juan Cheng 0004, Hu Peng |
FUSION | 3 |
| 2017 | Illumination Variation-Resistant Video-Based Heart Rate Measurement Using Joint Blind Source Separation and Ensemble Empirical Mode DecompositionabstractRecent studies have demonstrated that heart rate (HR) could be estimated using video data [e.g., exploring human facial regions of interest (ROIs)] under well-controlled conditions. However, in practice, the pulse signals may be contaminated by motions and illumination variations. In this paper, tackling the illumination variation challenge, we propose an illumination-robust framework using joint blind source separation (JBSS) and ensemble empirical mode decomposition (EEMD) to effectively evaluate HR from webcam videos. The framework takes the hypotheses that both facial ROI and background ROI have similar illumination variations. The background ROI is then considered as a noise reference sensor to denoise the facial signals by using the JBSS technique to extract the underlying illumination variation sources. Further, the reconstructed illumination-resisted green channel of the facial ROI is detrended and decomposed into a number of intrinsic mode functions using EEMD to estimate the HR. Experimental results demonstrated that the proposed framework could estimate HR more accurately than the state-of-the-art methods. The Bland-Altman plots showed that it led to better agreement with HR ground truth with the mean bias 1.15 beats/min (bpm), with 95% limits from -15.43 to 17.73 bpm, and the correlation coefficient 0.53. This study provides a promising solution for realistic noncontact and robust HR measurement applications. Juan Cheng 0004, Xun Chen 0001, Lingxi Xu, Z. Jane Wang 0001 |
IEEE J. Biomed. Health Informatics | 1 |