Yu Liu 0023

dblp:97/2274-23 · DBLP profile ↗
← Back
70ranked-venue papers
17as first author
55since 2021 · last 2027
0000-0003-2211-3535ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 13 first-author · 29 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2027 PA-LRG: Prototype-aware and low-rank guided multi-view clustering
Fengna Yang, Zhiqin Zhu, Yiyao An, Yu Liu 0023
Expert Syst. Appl.5
2026 Adaptive Dynamic Dehazing via Instruction-Driven and Task-Feedback Closed-Loop Optimization for Diverse Downstream Task Adaptation
abstract
In real-world vision systems, haze removal is required not only to enhance image visibility but also to meet the specific needs of diverse downstream tasks. To address this challenge, we propose a novel adaptive dynamic dehazing framework that incorporates a closed-loop optimization mechanism. It enables feedback-driven refinement based on downstream task performance and user instruction–guided adjustment during inference, allowing the model to satisfy the specific requirements of multiple downstream tasks without retraining. Technically, our framework integrates two complementary and innovative mechanisms: (1) a task feedback loop that dynamically modulates dehazing outputs based on performance across multiple downstream tasks, and (2) a text instruction interface that allows users to specify high-level task preferences. This dual-guidance strategy enables the model to adapt its dehazing behavior after training, tailoring outputs in real time to the evolving needs of multiple tasks. Extensive experiments across various vision tasks demonstrate the strong effectiveness, robustness, and generalizability of our approach. These results establish a new paradigm for interactive, task-adaptive dehazing that actively collaborates with downstream applications.
Shuaitian Song, Huafeng Li 0001, Shujuan Wang, Yu Liu 0023
AAAI5
2026 LST-rPPG: A long-range spatio-temporal model for high-accuracy heart rate variability measurement
Jiajie Li 0011, Juan Cheng 0004, Rencheng Song, Yu Liu 0023
Expert Syst. Appl.4
2026 FCDFusion: A noise-robust infrared and visible image fusion method based on Feature Coupling and Disentanglement
Ting Lv, Yu Liu 0023
Knowl. Based Syst.2
2026 DPFR: Semi-supervised gland segmentation via density perturbation and feature recalibration
Jiejiang Yu, Yu Liu 0023
Medical Image Anal.2
2026 Multimodal artificial intelligence for disease diagnosis: Advances, applications, and challenges
Shaozhe Wang, Fan Zhang 0070, Yu Liu 0023, Huafeng Li 0001, Junyu Dong, David Zhang 0001
Pattern Recognit.5
2026 Infrared-assisted single-stage framework for joint restoration and fusion of visible and infrared images under hazy conditions
Huafeng Li 0001, Jiaqi Fang, Yu Liu 0023
Pattern Recognit.4
2026 A Survey on lightweight technology of neural networks for medical image segmentation
abstract
Recent advances in medical image segmentation have significantly improved segmentation accuracy. Nevertheless, the clinical deployment of large-scale segmentation networks remains constrained by challenges such as excessive parameter counts, complex architectures, and limited adaptability to diverse deployment environments. The absence of lightweight design further restricts their integration into resource-limited edge devices. To address these barriers, lightweight strategies have emerged as an effective solution. Structural optimization simplifies network architectures to reduce computational costs, while model compression techniques shrink model size without sacrificing performance. At the same time, hardware-level acceleration provides additional support for efficient inference in real-world scenarios. This review systematically summarizes recent lightweight methods for medical image segmentation from both software and hardware perspectives. Representative algorithmic approaches are highlighted, including pruning, quantization, knowledge distillation, and efficient network architectures, along with hardware-aware optimization strategies tailored for edge deployment. Moreover, we explored the mainstream approach of integrating large-scale models with lightweight technologies to achieve the optimal balance between segmentation accuracy and computational efficiency. Finally, current limitations and potential research directions are outlined to promote the translation of lightweight segmentation models into routine clinical workflows. By providing a structured reference, this review aims to support researchers and practitioners in advancing the efficient and practical application of medical image segmentation in clinical environments.
Zhiqin Zhu, Hanchen Wang 0005, Guanqiu Qi, Neal Mazur, Yu Liu 0023, Huafeng Li 0001, Baisen Cong, Litao Bai
Pattern Recognit.6
2026 Trustworthy Anatomical Rectification: A Weakly-Supervised Pre-Alignment Framework for Deformable Medical Image Registration
Yu Liu 0023
IEEE Signal Process. Lett.2
2026 PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
abstract
Electroencephalography (EEG) signals provide a promising and involuntary reflection of brain activity related to emotional states, offering significant advantages over behavioral cues such as facial expressions. However, EEG signals are often noisy, affected by artifacts, and vary across individuals, complicating emotion recognition. While multimodal approaches have used peripheral physiological signals (PPS) such as galvanic skin response to complement EEG, they often overlook the dynamic synchronization and consistent semantics between the modalities. Additionally, the temporal dynamics of emotional fluctuations across different time resolutions in PPS remain underexplored. To address these challenges, we propose PhysioSync, a novel pretraining framework leveraging temporal and cross-modal contrastive learning (CM-CL), inspired by physiological synchronization phenomena. PhysioSync incorporates cross-modal consistency alignment (CM-CA) to model dynamic relationships between EEG and complementary PPS, enabling emotion-related synchronizations across modalities. Besides, it introduces long- and short-term temporal contrastive learning (LS-TCL) to capture emotional synchronization at different temporal resolutions within modalities. After pretraining, cross-resolution and cross-modal features are hierarchically fused and fine-tuned to enhance emotion recognition. Experiments on DEAP and DREAMER datasets demonstrate PhysioSync’s advanced performance under unimodal and cross-modal conditions, highlighting its effectiveness for EEG-centered emotion recognition.
Jia Li 0013, Yu Liu 0023, Zhenzhen Hu 0004, Meng Wang 0001
IEEE Trans. Comput. Soc. Syst.3
2026 Wiener-Deconvolution-Driven Event-Based Deblurring for Low-Light Imaging
abstract
We address event-based deblurring for low-light imaging, where conventional frames suffer severe blur, noise and saturation, while events capture sharp high-frequency contrast changes with microsecond latency that can guide the recovery of lost structures. Existing event-based reconstruction methods neither explicitly model low-light noise and saturation nor enforce precise alignment between events and frames, which limits cross-modal fusion and deblurring quality. We propose the Wiener-Deconvolution-Driven Event-Based Deblurring Network (WiED-Net), which embeds the Wiener deconvolution into a deep architecture so that the physical imaging model and noise statistics are encoded in the frequency domain and high-frequency recovery is stabilized on noise dominated night data. WiED-Net adopts a two stage design. The first stage applies Wiener deconvolution in both image and feature spaces to suppress noise, recover saturated regions and reduce ringing, assisted by an eventguided cross-modal feature fusion (ECFF) module for accurate alignment. The second stage uses a multi-scale fusion module to integrate the complementary event and image branches. Training is constrained by a set of losses, including a tailored blur kernel loss that provides closed-loop regularization from physical priors. Together, these designs enable WiED-Net to recover fine details while robustly suppressing artifacts and noise, and to achieve superior quantitative and qualitative performance, achieving superior quantitative and qualitative performance with a notable improvement of 1.97 dB in PSNR and 5% in SSIM over the previous state-of-the-art methods in low-light deblurring. Code will be available at https://github.com/zhuzifeng38/WiED-Net.
Zeyu Xiao 0002, Jianlong Jin, Feng Xue 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Synergistic Prompting for Complementarity and Consistency in Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) aims to partition unlabeled multi-view data into semantically coherent groups, even when certain views are missing due to sensor failures, data collection constraints, or privacy concerns. Despite advancements in deep IMVC methods, two critical challenges remain unresolved: (i) the lack of explicit mechanisms to model cross-view complementarity and (ii) the absence of principled strategies to ensure global semantic consistency across views. To address these challenges, we propose SP-IMVC, a novel Synergistic Prompting framework that jointly models complementarity and consistency under view incompleteness. Specifically, we introduce two types of learnable prompts: the Cross-View Complementary Prompt (CVCP), which aggregates auxiliary representations from available views to enrich the semantics of the current view and mitigate information loss; and the Latent Anchor Prompt (LAP), which utilizes a global anchor prompt pool to provide adaptive semantic priors that promote globally consistent representations. These prompts are optimized jointly within a unified architecture to achieve synergistic prompting of cross-view complementarity and global semantic consistency. Extensive experiments on six public benchmarks demonstrate that SP-IMVC consistently outperforms 14 state-of-the-art IMVC approaches, particularly in scenarios with high missing-view ratios, validating the effectiveness and robustness of our synergistic prompt-guided clustering framework. The code will be released to facilitate future research.
Xiaoshuai Hao, Yingbo Tang, Peng Hao 0003, Yunfeng Diao, Guangyin Jin, Yu Liu 0023
IEEE Trans. Image Process.8
2026 STAFuse: Scene-Text Aggregation-Guided Composite Degradation-Robust Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to integrate complementary information from source images to generate high-quality fusion images that serve downstream tasks. However, the differentiated representation of image scene content, the unpredictability of degradation modes in source images, and the complexity of composite degradations pose significant challenges to constructing degradation-robust fusion models. To overcome these challenges, this study proposes STAFuse, a framework designed for composite degradation-robust image fusion that utilizes adaptive degradation-mode identification and aggregated textual prior guidance. We first introduce a scene- and degradation-aware mechanism that extracts crucial context and degradation data, converting it into a textual format to generate dynamic convolution kernels. This allows for the adaptive identification and elimination of varying degradation effects and scene disparities. Additionally, we implement an aggregated scene prior that condenses multi-source information into a fusion text, simulating an ideal scene to effectively guide the retrieval and fusion of multimodal information. Finally, we design a textual-domain supervision loss to perform auxiliary semantic supervision on the fusion model, thereby suppressing degradation effects and improving the visual fidelity of the fusion results. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art methods in terms of robustness, flexibility, and the aggregation of complementary information.
Ting Lv, Hong Jiang 0006, Yu Liu 0023
IEEE Trans. Image Process.3
2026 Tensor Wheel Decomposition: Theory and Application to Tensor Completion
abstract
Recently, tensor network (TN) decompositions have gained prominence in computer vision and contributed promising results to tensor recovery for their capability of compactly and efficiently representing high-order tensors. However, current TN topologies are rather being developed towards more intricate structures to pursue incremental improvements, resulting in a drastically increased number of TN ranks, which requires laborious hyper-parameter selection, especially for higher-order cases. In this paper, we propose a novel TN decomposition, dubbed tensor wheel (TW) decomposition, in which a high-order tensor is represented by a set of latent factors mapped into a specific wheel topology. Such a decomposition is constructed starting from analyzing the graph structure, aiming to more accurately characterize the complex interactions inside objectives while maintaining a lower hyper-parameter scale, theoretically alleviating the above deficiencies. The comprehensive analysis of the mathematical properties fully demonstrates that TW decomposition can be more potential in representation capabilities and more flexible in controlling both parameter storage and computational costs. To compute the TW-format decomposition, the sequential singular value decomposition (SVD)-based and the alternating least squares (ALS)-based learning algorithms are developed. Furthermore, to investigate the validity of TW decomposition, we provide its one numerical application, i.e., tensor completion (TC), yet develop an efficient proximal alternating minimization-based solving algorithm with guaranteed convergence. Experimental results on both synthetic and real-world data reveal that TW decomposition significantly outperforms other state-of-the-art tensor decompositions for incomplete-tensor inference, especially under solely few observations, thus substantiating the superiority and reliability of TW decomposition.
Zhong-Cheng Wu, Liang-Jian Deng, Ting-Zhu Huang, Hong-Xia Dou, Gemine Vivone, Yu Liu 0023
IEEE Trans. Image Process.6
2026 LSRNet: A Novel Interpretable Low-rank Sparse Representation Guided Fusion Network for Polarization and Intensity Images
abstract
Polarization and intensity images fusion (PIF) has extracted extensive attentions as it can generate images with clear scene information and salient texture details of the object surface that are important for downstream applications. However, existing deep learning-based PIF methods usually lack interpretability and ignore the interactions among multi-modal features. To this end, we propose a novel interpretable low-rank sparse representation guided fusion network for polarization and intensity images (termed LSRNet). Specifically, a low-rank sparse representation deep unfolding module is designed to acquire the base and detail features of the source images, with the ability of improving the interpretability of the network. In addition, a cross-modal connection complementary feature extraction module is proposed, which aims to establish dependency among features of multi-modalities to fully extract complementary features of the source images. In order to demonstrate the validity of our LSRNet and take into account shortcomings of existing datasets for PIF, a multi-scene polarization and intensity image dataset, named MSPI dataset, is constructed, which includes 1034 high-resolution aligned image pairs. According to the best of our knowledge, this is the most comprehensive dataset for PIF that with a large number of image pairs, high resolution and multiple scene types. Extensive experiments on our MSPI dataset and two publicly available datasets (i.e., 12CFC and HCP) demonstrate the superior fusion performance, generalization ability, and desirable running efficiency of our LSRNet. Our codes and dataset will be publicly available at https://github.com/thebinyang/LSRNet.
Bin Yang 0008, Licheng Liu, Yu Liu 0023, Jing Li 0040
IEEE Trans. Image Process.4
2026 Correlation-Guided Recursive Pyramid Network for Deformable Brain MRI Registration
abstract
As a key preprocessing technique in medical image analysis, deformable image registration has remained a research focus over the past decade. Recently, deep learning-based registration methods have become mainstream. Nevertheless, simultaneously handling large-scale deformations and accurate feature matching remains a persistent challenge. While pyramid architectures are widely employed to mitigate large-scale deformations, existing methods often exhibit an unbalanced focus. One group emphasizes iterative refinement to handle large deformations but relies on implicit, coarse feature interactions. Conversely, the other group concentrates on explicit matching techniques, but such static matching is often unreliable in regions with significant anatomical discrepancies. To bridge this gap, we propose a novel Correlation-Guided Recursive Pyramid Network (CRPNet). Unlike previous approaches, CRPNet addresses these challenges in a unified manner by embedding explicit correlation modeling directly into the recursive optimization. Specifically, we propose a Correlation-Guided Intra-layer Recursive Strategy (CGIRS), which enables the network to continuously refine matching accuracy through recursive feedback while preventing cross-scale error propagation. To facilitate this, we design a Spatial Correlation Module (SPCM) for accurate spatial correspondence and a Semantic Correlation Module (SECM) for high-level semantic alignment. Extensive experiments on three brain imaging datasets demonstrate that our method achieves state-of-the-art performance, particularly exhibiting exceptional robustness under extreme deformations, proving the efficacy of our method for deformable brain MRI registration. The code is available at https://github.com/ZhangWH0129/CRPNet.
Yu Liu 0023
IEEE Trans. Image Process.2
2026 Video-Based Instantaneous Heart Rate Measurement With Enhanced Time-Frequency Representations
abstract
Remote photoplethysmography (rPPG) for heart rate (HR) measurement based on facial videos has recently attracted increasing attention. However, most existing methods focus on average heart rate (AHR) over a period rather than instantaneous heart rate (IHR), which better reflects physical and mental states. To address this issue, we propose a novel rPPG-based method for measuring IHR values from facial videos. Our method employs the wavelet synchrosqueezed transform (WSST) to generate time-frequency representations (TFRs) of chrominance (CHROM) signals from multiple facial regions of interest (ROIs), synchronously reflecting the IHR during a video segment. Furthermore, the TransUNet is introduced to refine these TFR images, enhancing the ridge line information related to IHRs. Comprehensive comparisons and ablation studies on four public datasets (UBFC-rPPG, PURE, UBFC-Phys, and MMPD) reveal that our WSST-UNet method achieves superior performance over several typical rPPG methods, achieving mean absolute errors (MAE) of 2.34 beats per minute (bpm), 1.29 bpm, 5.03 bpm, and 6.58 bpm, respectively. The proposed method offers a promising solution for practical application in video-based IHR measurements.
Juan Cheng 0004, Xiwen Luo, Rencheng Song, Yu Liu 0023
IEEE Trans. Multim.5
2026 Learning Dual Modality Interactions for Event-Based Motion Deblurring
abstract
Event cameras hold great potential for motion deblurring because they capture motion information with microsecond precision, offering robustness to motion blur. However, the limited interaction between RGB frames and event streams presents a significant challenge, preventing the full utilization of the event cameras' unique advantages. To address this, we proposeDual frame-eventInteraction and introduce a multi-scaleNetwork structure, DuInt-Net. DuInt-Net aims to tackle two key challenges: (1) enhancing the representational and interaction capabilities between RGB frames and event streams, and (2) adaptively selecting richer visual features for improved motion deblurring. We introduce an event-frame joint interaction module that consists of three branches: a base branch, a global awareness attention branch, and a local enhancement attention branch. The base branch processes essential pixel-level features that retain the original structural information. The global branch integrates event data to improve large-scale motion understanding, while the local branch uses large-kernel convolutions to refine fine-grained details in RGB frames. For superior reconstruction performance, we also propose the event-guided multi-scale fusion attention module, which effectively combines local visual information and global frame-event relationships. Extensive experiments demonstrate that DuInt-Net achieves superior performance, both quantitatively and qualitatively, showcasing its superior motion deblurring capabilities.
Zeyu Xiao 0002, Zhuoyuan Li 0001, Yang Zhao 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001
IEEE Trans. Multim.4
2026 Seeing Clearly and Detecting Precisely: Perceptual Enhancement and Focus Calibration for Small-Object Detection
abstract
Small-object detection remains challenging due to limited pixel information, blurred boundaries, and weak semantic cues. Although recent advances in multiscale fusion and attention mechanisms have led to improved performance, existing methods still struggle to preserve high-frequency structural details and achieve precise localization-particularly in dense, cluttered, or low-resolution scenarios. These limitations are primarily caused by the loss of fine-grained features during downsampling and the absence of region-aware focus mechanisms. Inspired by the human visual strategy of "see clearly and detect precisely," we propose PEFC-Net, a novel framework that enhances both perceptual clarity and localization accuracy for small-object detection. To mitigate structural degradation, we introduce the hybrid structural perception (HSP) module, which jointly encodes spatial gradients and localized frequency components through wavelet-based decomposition and edge-aware refinement. To further improve region-level focus, we design the axis-aligned focus calibration (AAFC) module, which captures long-range directional context via axis-sensitive pooling and adaptively refines attention with shape-aware calibration. Extensive experiments on four challenging benchmarks-VisDrone-2019, TT100K, NWPU VHR-10, and DIOR-demonstrate that PEFC-Net consistently outperforms state-of-the-art methods, delivering robust performance under occlusion, dense distribution, and scale variation.
Zhiqin Zhu, Guanqiu Qi, Huafeng Li 0001, Yu Liu 0023
IEEE Trans. Neural Networks Learn. Syst.6
2025 UniFuse: A Unified All-In-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments
Dayong Su, Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023
ICCV5
2025 MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
abstract
Multi-sensor fusion models play a crucial role in autonomous driving perception, particularly in tasks like 3D object detection and HD map construction. These models provide essential and comprehensive static environmental information for autonomous driving systems. While camera-LiDAR fusion methods have shown promising results by integrating data from both modalities, they often depend on complete sensor inputs. This reliance can lead to low robustness and potential failures when sensors are corrupted or missing, raising significant safety concerns. To tackle this challenge, we introduce the Multi-Sensor Corruption Benchmark (MSC-Bench), the first comprehensive benchmark aimed at evaluating the robustness of multi-sensor autonomous driving perception models against various sensor corruptions. Our benchmark includes 16 combinations of corruption types that disrupt both camera and LiDAR inputs, either individually or concurrently. Extensive evaluations of six 3D object detection models and four HD map construction models reveal substantial performance degradation under adverse weather conditions and sensor failures, underscoring critical safety issues. The benchmark toolkit and affiliated code and model checkpoints have been made publicly accessible. Project website: MSC-Bench.
Xiaoshuai Hao, Guanqun Liu 0008, Yuheng Ji, Mengchuan Wei, Haimei Zhao, Lingdong Kong, Rong Yin 0001, Yu Liu 0023
ICME9
2025 VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and Results
abstract
This paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD.
Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde
VCIP11
2025 MulFS-CAP: Multimodal Fusion-Supervised Cross-Modality Alignment Perception for Unregistered Infrared-Visible Image Fusion
abstract
In this study, we propose Multimodal Fusion-supervised Cross-modality Alignment Perception (MulFS-CAP), a novel framework for single-stage fusion of unregistered infrared-visible images. Traditional two-stage methods depend on explicit registration algorithms to align source images spatially, often adding complexity. In contrast, MulFS-CAP seamlessly blends implicit registration with fusion, simplifying the process and enhancing suitability for practical applications. MulFS-CAP utilizes a shared shallow feature encoder to merge unregistered infrared-visible images in a single stage. To address the specific requirements of feature-level alignment and fusion, we develop a consistent feature learning approach via a learnable modality dictionary. This dictionary provides complementary information for unimodal features, thereby maintaining consistency between individual and fused multimodal features. As a result, MulFS-CAP effectively reduces the impact of modality variance on cross-modality feature alignment, allowing for simultaneous registration and fusion. Additionally, in MulFS-CAP, we advance a novel cross-modality alignment approach, creating a correlation matrix to detail pixel relationships between source images. This matrix aids in aligning features across infrared and visible images, further refining the fusion process. The above designs make MulFS-CAP more lightweight, effective and explicit registration-free. Experimental results from different datasets demonstrate the effectiveness of our proposed method and its superiority over the state-of-the-art two-stage methods.
Huafeng Li 0001, Zengyi Yang, Wei Jia 0001, Zhengtao Yu 0001, Yu Liu 0023
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Lightweight Efficient Rate-Adaptive Network for Compression-Aware Image Rescaling
abstract
Compression-aware image rescaling approaches convert high-resolution images to compressed low-resolution ones to fit various display devices or save bandwidth/storage. Inverse upscaling is successively performed to enlarge the low-resolution images to the original sizes with rich details. However, previous compression-aware image rescaling methods lack adaptivity to diverse compression rates, or require multiple large models with huge computational cost for adjusting. To overcome these challenges, we propose a lightweight efficient rate-adaptive network (LERAN) for compression-aware image rescaling. We design a non-invertible framework based on quality factor-driven feature modulation modules and an expandable training strategy, to achieve the adaptivity to various compression rates with only one light and efficient model. Moreover, alternative recursive blocks are presented for lighter weights with very small performance drop. During training, we also introduce a sparse low-resolution residual feature loss which promotes easier convergence of the model without adding further computational burden. Extensive experimental results demonstrate that our method significantly outperforms state-of-the-art compression-aware image rescaling approaches for different compression rates on popular benchmarks, with an all-in-one lightweight model and much faster speed. The code will be available at https://github.com/5ofwind .
Dingyi Li, Yu Liu 0023
IEEE Signal Process. Lett.3
2025 MambaDiff: Mamba-Enhanced Diffusion Model for 3D Medical Image Segmentation
abstract
Accurate 3D medical image segmentation is crucial for diagnosis and treatment. Diffusion models demonstrate promising performance in medical image segmentation tasks due to the progressive nature of the generation process and the explicit modeling of data distributions. However, the weak guidance of conditional information and insufficient feature extraction in diffusion models lead to the loss of fine-grained features and structural consistency in the segmentation results, thereby affecting the accuracy of medical image segmentation. To address this challenge, we propose a Mamba-Enhanced Diffusion Model for 3D Medical Image Segmentation. We extract multilevel semantic features from the original images using an encoder and tightly integrate them with the denoising process of the diffusion model through a Semantic Hierarchical Embedding (SHE) mechanism, to capture the intricate relationship between the noisy label and image data. Meanwhile, we design a Global-Slice Perception Mamba (GSPM) layer, which integrates multi-dimensional perception mechanisms to endow the model with comprehensive spatial reasoning and feature extraction capabilities. Experimental results show that our proposed MambaDiff achieves more competitive performance compared to prior arts with substantially fewer parameters on four public medical image segmentation datasets including BraTS 2021, BraTS 2024, LiTS and MSD Hippocampus. The source code of our method is available at https://github.com/yuliu316316/MambaDiff.
Yu Liu 0023, Juan Cheng 0004, Haolin Zhan, Zhiqin Zhu
IEEE Trans. Image Process.1
2025 VDMUFusion: A Versatile Diffusion Model-Based Unsupervised Framework for Image Fusion
abstract
Image fusion facilitates the integration of information from various source images of the same scene into a composite image, thereby benefiting perception, analysis, and understanding. Recently, diffusion models have demonstrated impressive generative capabilities in the field of computer vision, suggesting significant potential for application in image fusion. The forward process in the diffusion models requires the gradual addition of noise to the original data. However, typical unsupervised image fusion tasks (e.g., infrared-visible, medical, and multi-exposure image fusion) lack ground truth images (corresponding to the original data in diffusion models), thereby preventing the direct application of the diffusion models. To address this problem, we propose a versatile diffusion model-based unsupervised framework for image fusion, termed as VDMUFusion. In the proposed method, we integrate the fusion problem into the diffusion sampling process by formulating image fusion as a weighted average process and establishing appropriate assumptions about the noise in the diffusion model. To simplify the training process, we propose a multi-task learning framework that replaces the original noise prediction network, allowing for simultaneous prediction of noise and fusion weights. Meanwhile, our method employs joint training across various fusion tasks, which significantly improves noise prediction accuracy and yields higher quality fused images compared to training on a single task. Extensive experimental results demonstrate that the proposed method delivers very competitive performance across various image fusion tasks. The code is available at https://github.com/yuliu316316/VDMUFusion.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.2
2025 Probability Map-Guided Network for 3D Volumetric Medical Image Segmentation
abstract
3D medical images are volumetric data that provide spatial continuity and multi-dimensional information. These features provide rich anatomical context. However, their anisotropy may result in reduced image detail along certain directions. This can cause blurring or distortion between slices. In addition, global or local intensity inhomogeneities are often observed. This may be due to limitations of the imaging equipment, inappropriate scanning parameters, or variations in the patient's anatomy. This inhomogeneity may blur lesion boundaries and may also mask true features, causing the model to focus on irrelevant regions. Therefore, a probability map-guided network for 3D volumetric medical image segmentation (3D-PMGNet) is proposed. The probability maps generated from the intermediate features are used as supervisory signals to guide the segmentation process. A new probability map reconstruction method is designed, combining dynamic thresholding with local adaptive smoothing. This enhances the reliability of high-response regions while suppressing low-response noise. A learnable channel-wise temperature coefficient is introduced to adjust the probability distribution to make it closer to the true distribution; in addition, a feature fusion method based on dynamic prompt encoding is developed. The response strength of the main feature maps is dynamically adjusted, and this adjustment is achieved through the spatial position encoding derived from the probability maps. The proposed method has been evaluated on four datasets. Experimental results show that the proposed method outperforms state-of-the-art 3D medical image segmentation methods. The source codes have been publicly released at https://github.com/ZHANGZIMENG01/3D-PMGNet.
Zhiqin Zhu, Zimeng Zhang, Guanqiu Qi, Yu Liu 0023
IEEE Trans. Image Process.6
2025 Focus Affinity Perception and Super-Resolution Embedding for Multifocus Image Fusion
abstract
Despite the fact that there is a remarkable achievement on multifocus image fusion, most of the existing methods only generate a low-resolution image if the given source images suffer from low resolution. Obviously, a naive strategy is to independently conduct image fusion and image super-resolution. However, this two-step approach would inevitably introduce and enlarge artifacts in the final result if the result from the first step meets artifacts. To address this problem, in this article, we propose a novel method to simultaneously achieve image fusion and super-resolution in one framework, avoiding step-by-step processing of fusion and super-resolution. Since a small receptive field can discriminate the focusing characteristics of pixels in detailed regions, while a large receptive field is more robust to pixels in smooth regions, a subnetwork is first proposed to compute the affinity of features under different types of receptive fields, efficiently increasing the discriminability of focused pixels. Simultaneously, in order to prevent from distortion, a gradient embedding-based super-resolution subnetwork is also proposed, in which the features from the shallow layer, the deep layer, and the gradient map are jointly taken into account, allowing us to get an upsampled image with high resolution. Compared with the existing methods, which implemented fusion and super-resolution independently, our proposed method directly achieves these two tasks in a parallel way, avoiding artifacts caused by the inferior output of image fusion or super-resolution. Experiments conducted on the real-world dataset substantiate the superiority of our proposed method compared with state of the arts.
Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023, Guangming Lu 0002, Yong Xu 0001, Zhengtao Yu 0001, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 A Deep Learning Framework for Infrared and Visible Image Fusion Without Strict Registration
Huafeng Li 0001, Junyu Liu, Yu Liu 0023
Int. J. Comput. Vis.4
2024 Misalignment-Resistant Deep Unfolding Network for multi-modal MRI super-resolution and reconstruction
Jinbao Wei, Yu Liu 0023, Aiping Liu, Xun Chen 0001
Knowl. Based Syst.4
2024 Rethinking the Effectiveness of Objective Evaluation Metrics in Multi-Focus Image Fusion: A Statistic-Based Approach
abstract
As an effective technique to extend the depth-of-field (DOF) of optical lenses, multi-focus image fusion has recently become an active topic in image processing community. However, a major problem remaining unsolved in this field is the lack of universal criteria in selecting objective evaluation metrics. Consequently, the metrics utilized in different studies often vary significantly, leading to high difficulties in achieving unbiased evaluation. To address this problem, this paper proposes a statistic-based approach for verifying the effectiveness of objective metrics in multi-focus image fusion. The core idea is to adopt statistical correlation measures to evaluate the performance consistency between a certain fusion metric and some popular full-reference image quality assessment models. In addition, a convolutional neural network (CNN)-based fusion metric is presented to measure the similarity between the source images and the fused image based on the semantic features at multiple abstraction levels. A comparative study is conducted to evaluate 20 existing fusion metrics using the proposed statistic-based approach on a large-scale, realistic and with-ground-truth multi-focus image fusion dataset recently released. Experimental results demonstrate the feasibility of the proposed approach in evaluating the effectiveness of objective metrics and the advantage of our CNN-based metric.
Yu Liu 0023, Zhengzheng Qi, Juan Cheng 0004, Xun Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 ITFuse: An interactive transformer for infrared and visible image fusion
Wei Tang 0018, Fazhi He, Yu Liu 0023
Pattern Recognit.3
2024 Brain tumor segmentation in MRI with multi-modality spatial information enhancement and boundary shape correction
Zhiqin Zhu, Guanqiu Qi, Neal Mazur, Yu Liu 0023
Pattern Recognit.6
2024 HF2TNet: A Hierarchical Fusion Two-Stage Training Network for Infrared and Visible Image Fusion
abstract
In the field of infrared and visible image fusion, current algorithms often focus on complex feature extraction and sophisticated fusion mechanisms, ignoring the issues of information redundancy and feature imbalance. These limit effective information aggregation. To address these issues, this paper proposes a hierarchical fusion strategy with a two-stage training network, abbreviated as HF2TNet, which achieves effective information aggregation in a staged manner. In the initial training stage, a three-stream encoder-decoder architecture is proposed, seamlessly integrating CNN and transformer modules. This architecture extracts both global and local features from visible and infrared images, capturing their shared attributes before the fusion process. Moreover, a multi-shared attention module (MSAM) is proposed to profoundly reconstruct and augment the visible and infrared features, ensuring the preservation and enhancement of details across modalities. In the subsequent stage, HF2TNet utilizes the pre-integrated features as query inputs for the dual MSAMs. These modules interact with the previously reconstructed infrared and visible features to enhance complementary information and ensure a balanced feature fusion. Experimental results indicate HF2TNet's superior performance on standard datasets like MSRS and TNO, especially in complex scenes, demonstrating its potential in multimodal image fusion.
Ting Lv, Chuanming Ji, Hong Jiang 0006, Yu Liu 0023
IEEE Signal Process. Lett.4
2024 Video Rescaling With Recurrent Diffusion
abstract
Video rescaling helps to fit different display devices. In video rescaling systems, videos are downsampled for easier storage, transmission and preview. The downsampled videos can be upsampled with a neural network to restore the details when needed. Previous group-based video rescaling algorithms benefit from the joint downsampling and joint upsampling of multiple frames, but are restricted by the fully joint operation. In this paper, we propose a recurrent diffusion-based framework for video rescaling. We employ biased joint operation and recurrent diffusion, to make a better use of the temporal relation within different frames in each image group. We explicitly control the direction of information propagation by arranging the processing order of all frames. In biased joint operation, we concentrate on restoring one frame, i.e., the middle frame. The other frames in the group are coarsely reconstructed. Our recurrent diffusion compensates the coarse frames by gradually propagating information from the middle to borders backwardly and forwardly. The recurrent diffusion module is performed by fusing the information of adjacent frames. Biased joint operation and recurrent diffusion are jointly trained. We design several propagation variants and find that our recurrent diffusion is the best among them. It is also shown that recurrent diffusion is better than non-recurrent diffusion in terms of reconstruction quality and model size. We also adopt a high-resolution fine-tuning strategy to further improve the quality of high-resolution frames. Experimental results demonstrate the effectiveness of the proposed method in terms of visual quality, quantitative evaluations, and computational efficiency. The code will be released at https://github.com/5ofwind/RDVR.
Dingyi Li, Yu Liu 0023, Zengfu Wang, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2024 MM-Net: A MixFormer-Based Multi-Scale Network for Anatomical and Functional Image Fusion
abstract
Anatomical and functional image fusion is an important technique in a variety of medical and biological applications. Recently, deep learning (DL)-based methods have become a mainstream direction in the field of multi-modal image fusion. However, existing DL-based fusion approaches have difficulty in effectively capturing local features and global contextual information simultaneously. In addition, the scale diversity of features, which is a crucial issue in image fusion, often lacks adequate attention in most existing works. In this paper, to address the above problems, we propose a MixFormer-based multi-scale network, termed as MM-Net, for anatomical and functional image fusion. In our method, an improved MixFormer-based backbone is introduced to sufficiently extract both local features and global contextual information at multiple scales from the source images. The features from different source images are fused at multiple scales based on a multi-source spatial attention-based cross-modality feature fusion (CMFF) module. The scale diversity of the fused features is further enriched by a series of multi-scale feature interaction (MSFI) modules and feature aggregation upsample (FAU) modules. Moreover, a loss function consisting of both spatial domain and frequency domain components is devised to train the proposed fusion model. Experimental results demonstrate that our method outperforms several state-of-the-art fusion methods on both qualitative and quantitative comparisons, and the proposed fusion model exhibits good generalization capability. The source code of our fusion method will be available at https://github.com/yuliu316316.
Yu Liu 0023, Juan Cheng 0004, Z. Jane Wang 0001, Xun Chen 0001
IEEE Trans. Image Process.1
2023 CCSR-Net: Unfolding Coupled Convolutional Sparse Representation for Multi-focus Image Fusion
Kecheng Zheng, Juan Cheng 0004, Yu Liu 0023
PRCV (10)3
2023 TCCFusion: An infrared and visible image fusion method based on transformer and cross correlation
Wei Tang 0018, Fazhi He, Yu Liu 0023
Pattern Recognit.3
2023 Multi-Exposure Image Fusion via Multi-Scale and Context-Aware Feature Learning
abstract
In this letter, a deep learning (DL)-based multi-exposure image fusion (MEF) method via multi-scale and context-aware feature learning is proposed, aiming to overcome the defects of existing traditional and DL-based methods. The proposed network is based on an auto-encoder architecture. First, an encoder that combines the convolutional network and Transformer is designed to extract multi-scale features and capture the global contextual information. Then, a multi-scale feature interaction (MSFI) module is devised to enrich the scale diversity of extracted features using cross-scale fusion and Atrous spatial pyramid pooling (ASPP). Finally, a decoder with a nest connection architecture is introduced to reconstruct the fused image. Experimental results show that the proposed method outperforms several representative traditional and DL-based MEF methods in terms of both visual quality and objective assessment.
Yu Liu 0023, Juan Cheng 0004, Xun Chen 0001
IEEE Signal Process. Lett.1
2023 EEG-Based Emotion Recognition via Neural Architecture Search
abstract
With the flourishing development of deep learning (DL) and the convolution neural network (CNN), electroencephalogram-based (EEG) emotion recognition is occupying an increasingly crucial part in the field of brain-computer interface (BCI). However, currently employed architectures have mostly been designed manually by human experts, which is a time-consuming and labor-intensive process. In this paper, we proposed a novel neural architecture search (NAS) framework based on reinforcement learning (RL) for EEG-based emotion recognition, which can automatically design network architectures. The proposed NAS mainly contains three parts: search strategy, search space, and evaluation strategy. During the search process, a recurrent network (RNN) controller is used to select the optimal network structure in the search space. We trained the controller with RL to maximize the expected reward of the generated models on a validation set and force parameter sharing among the models. We evaluated the performance of NAS on the DEAP and DREAMER dataset. On the DEAP dataset, the average accuracies reached 97.94%, 97.74%, and 97.82% on arousal, valence, and dominance respectively. On the DREAMER dataset, average accuracies reached 96.62%, 96.29% and 96.61% on arousal, valence, and dominance, respectively. The experimental results demonstrated that the proposed NAS outperforms the state-of-the-art CNN-based methods.
Chang Li 0001, Zhongzhen Zhang, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
IEEE Trans. Affect. Comput.5
2023 EEG-Based Emotion Recognition via Channel-Wise Attention and Self Attention
abstract
Emotion recognition based on electroencephalography (EEG) is a significant task in the brain-computer interface field. Recently, many deep learning-based emotion recognition methods are demonstrated to outperform traditional methods. However, it remains challenging to extract discriminative features for EEG emotion recognition, and most methods ignore useful information in channel and time. This article proposes an attention-based convolutional recurrent neural network (ACRNN) to extract more discriminative features from EEG signals and improve the accuracy of emotion recognition. First, the proposed ACRNN adopts a channel-wise attention mechanism to adaptively assign the weights of different channels, and a CNN is employed to extract the spatial information of encoded EEG signals. Then, to explore the temporal information of EEG signals, extended self-attention is integrated into an RNN to recode the importance based on intrinsic similarity in EEG signals. We conducted extensive experiments on the DEAP and DREAMER databases. The experimental results demonstrate that the proposed ACRNN outperforms state-of-the-art methods.
Chang Li 0001, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Feng Wan 0003, Xun Chen 0001
IEEE Trans. Affect. Comput.5
2023 MSCAF-Net: A General Framework for Camouflaged Object Detection via Learning Multi-Scale Context-Aware Features
abstract
The aim of camouflaged object detection (COD) is to find objects that are hidden in their surrounding environment. Due to the factors like low illumination, occlusion, small size and high similarity to the background, COD is recognized to be a very challenging task. In this paper, we propose a general COD framework, termed as MSCAF-Net, focusing on learning multi-scale context-aware features. To achieve this target, we first adopt the improved Pyramid Vision Transformer (PVTv2) model as the backbone to extract global contextual information at multiple scales. An enhanced receptive field (ERF) module is then designed to refine the features at each scale. Further, a cross-scale feature fusion (CSFF) module is introduced to achieve sufficient interaction of multi-scale information, aiming to enrich the scale diversity of extracted features. In addition, inspired the mechanism of the human visual system, a dense interactive decoder (DID) module is devised to output a rough localization map, which is used to modulate the fused features obtained in the CSFF module for more accurate detection. The effectiveness of our MSCAF-Net is validated on four benchmark datasets. The results show that the proposed method significantly outperforms state-of-the-art (SOTA) COD models by a large margin. Besides, we also investigate the potential of our MSCAF-Net on some other vision tasks that are highly related to COD, such as polyp segmentation, COVID-19 lung infection segmentation, transparent object detection and defect detection. Experimental results demonstrate the high versatility of the proposed MSCAF-Net. The source code and results of our method are available athttps://github.com/yuliu316316/MSCAF-COD.
Yu Liu 0023, Haihang Li, Juan Cheng 0004, Xun Chen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 DATFuse: Infrared and Visible Image Fusion via Dual Attention Transformer
abstract
The fusion of infrared and visible images aims to generate a composite image that can simultaneously contain the thermal radiation information of an infrared image and the plentiful texture details of a visible image to detect targets under various weather conditions with a high spatial resolution of scenes. Previous deep fusion models were generally based on convolutional operations, resulting in a limited ability to represent long-range context information. In this paper, we propose a novel end-to-end model for infrared and visible image fusion via a dual attention Transformer termed DATFuse. To accurately examine the significant areas of the source images, a dual attention residual module (DARM) is designed for important feature extraction. To further model long-range dependencies, a Transformer module (TRM) is devised for global complementary information preservation. Moreover, a loss function that consists of three terms, namely, pixel loss, gradient loss, and structural loss, is designed to train the proposed model in an unsupervised manner. This can avoid manually designing complicated activity-level measurement and fusion strategies in traditional image fusion methods. Extensive experiments on public datasets reveal that our DATFuse outperforms other representative state-of-the-art approaches in both qualitative and quantitative assessments. The proposed model is also extended to address other infrared and visible image fusion tasks without fine-tuning, and the promising results demonstrate that it has good generalization ability. The source code is available athttps://github.com/tthinking/DATFuse.
Wei Tang 0018, Fazhi He, Yu Liu 0023, Yansong Duan, Tongzhen Si
IEEE Trans. Circuits Syst. Video Technol.3
2023 EEG-based Emotion Recognition via Transformer Neural Architecture Search
abstract
Emotion recognition based on electroencephalogram (EEG) plays an increasingly important role in the field of brain–computer interfaces. Recently, deep learning has been widely applied to EEG decoding owning to its excellent capabilities in automatic feature extraction. Transformer holds great superiority in processing time-series signals due to its long-term dependencies extraction ability. However, most existing transformer architectures are designed manually by human experts, which is a time-consuming and resource-intensive process. In this article, we propose an automatic transformer neural architectures search (TNAS) framework based on multiobjective evolution algorithm (MOEA) for the EEG-based emotion recognition. The proposed TNAS conducts the MOEA strategy that considers both accuracy and model size to discover the optimal model from well-trained supernet for the emotion recognition. We conducted extensive experiments to evaluate the performance of the proposed TNAS on the DEAP and DREAMER datasets. The experimental results showed that the proposed TNAS outperforms the state-of-the-art methods.
Chang Li 0001, Zhongzhen Zhang, Guoning Huang, Yu Liu 0023, Xun Chen 0001
IEEE Trans. Ind. Informatics5
2023 Bi-CapsNet: A Binary Capsule Network for EEG-Based Emotion Recognition
abstract
In recent years, deep learning has gained widespread attention in electroencephalogram (EEG)-based emotion recognition. However, deep learning methods are usually time-consuming with a large amount of memory usage, which obstructs their practical usage on resource-constrained devices. In this paper, we propose a binary capsule network (Bi-CapsNet) for EEG emotion recognition with low computational cost and memory usage. The Bi-CapsNet binarizes 32-bit weights and activations to 1 b, and replaces floating-point operations with efficient bitwise operations. To address the issue of function discontinuity in backward propagation, we use a continuous function to approximate the binarization process. Two popular EEG emotion databases, namely, DEAP and DREAMER, are used for performance evaluation. In comparison to its full-precision counterpart, the Bi-CapsNet achieves a $>\!25\times$reduction on the computational cost and a $>\!5\times$ reduction on the memory usage, while with only a $< $1% drop on the recognition accuracy. Compared to some state-of-the-art EEG emotion recognition methods, the proposed method obtains more competitive performance. In addition, the Bi-CapsNet is implemented on a mobile phone via an open-source binary inference framework named Bolt, and it achieves an $\sim\! 5\times$ inference acceleration in comparison to its full-precision counterpart.
Yu Liu 0023, Chang Li 0001, Juan Cheng 0004, Rencheng Song, Xun Chen 0001
IEEE J. Biomed. Health Informatics1
2023 YDTR: Infrared and Visible Image Fusion via Y-Shape Dynamic Transformer
abstract
Infrared and visible image fusion is aims to generate a composite image that can simultaneously describe the salient target in the infrared image and texture details in the visible image of the same scene. Since deep learning (DL) exhibits great feature extraction ability in computer vision tasks, it has also been widely employed in handling infrared and visible image fusion issue. However, the existing DL-based methods generally extract complementary information from source images through convolutional operations, which results in limited preservation of global features. To this end, we propose a novel infrared and visible image fusion method, i.e., the Y-shape dynamic Transformer (YDTR). Specifically, a dynamic Transformer module (DTRM) is designed to acquire not only the local features but also the significant context information. Furthermore, the proposed network is devised in a Y-shape to comprehensively maintain the thermal radiation information from the infrared image and scene details from the visible image. Considering the specific information provided by the source images, we design a loss function that consists of two terms to improve fusion quality: a structural similarity (SSIM) term and a spatial frequency (SF) term. Extensive experiments on mainstream datasets illustrate that the proposed method outperforms both classical and state-of-the-art approaches in both qualitative and quantitative assessments. We further extend the YDTR to address other infrared and RGB-visible images and multi-focus images without fine-tuning, and the satisfactory fusion results demonstrate that the proposed method has good generalization capability.
Wei Tang 0018, Fazhi He, Yu Liu 0023
IEEE Trans. Multim.3
2023 X-Net: a dual encoding-decoding method in medical image segmentation
Li Yin 0011, Zhiqin Zhu, Guanqiu Qi, Yu Liu 0023
Vis. Comput.6
2022 Multi-channel EEG-based emotion recognition in the presence of noisy labels
Chang Li 0001, Yimeng Hou, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
Sci. China Inf. Sci.5
2022 Superpixel-Based Noise-Robust Sparse Unmixing of Hyperspectral Image
abstract
Sparse unmixing (SU) of hyperspectral image (HSI), as a semisupervised approach, aims to find the optimal subset of the spectral library known in advance to represent each pixel in HSI. However, most of the existing SU methods cannot take full advantage of spatial information and mixed noise in HSI. To this end, we propose a superpixel-based noise-robust SU method (SNRSU) in the presence of mixed noise. First, we perform superpixel segmentation (SS) on the first principal component of HSI to extract the homogeneous regions. Then, we unmix each superpixel based on sparse representation (SR) and low-rank representation (LRR) in the maximuma posterioriframework, which can make full use of the spatial–spectral information in HSI under complex mixed noise. A number of experiments on simulated and real HSI datasets confirm the superior performance of the proposed SNRSU both qualitatively and quantitatively.
Chang Li 0001, Chenhong Sui, Rencheng Song, Juan Cheng 0004, Yu Liu 0023, Xun Chen 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 SF-Net: A Multi-Task Model for Brain Tumor Segmentation in Multimodal MRI via Image Fusion
abstract
Automatic segmentation of brain tumor regions from multimodal MRI scans is of great clinical significance. In this letter, we propose a “Segmentation-Fusion” multi-task model named SF-Net for brain tumor segmentation. In comparison to the widely-used multi-task model that adds a variational autoencoder (VAE) decoder to reconstruct the input data, using image fusion as an additional regularization for feature learning helps to achieve more sufficient fusion of multimodal features, which is beneficial to the multimodal image segmentation problem. To further improve the performance of the multi-task model, an uncertainty-based approach that can adaptively adjust the loss weights of different tasks during the training process is introduced for model training. Experimental results on the BraTS 2020 benchmark demonstrate that the proposed method can achieve higher segmentation accuracy than the VAE-based approach. In addition, as the by-product of the multi-task model, the image fusion results obtained are of high quality on the brain tumor regions. The source code of the proposed method is available at https://github.com/yuliu316316/SF-Net.
Yu Liu 0023, Fuhao Mu, Xun Chen 0001
IEEE Signal Process. Lett.1
2022 SOM-Net: Unrolling the Subspace-Based Optimization for Solving Full-Wave Inverse Scattering Problems
abstract
In this paper, an unrolling algorithm of the iterative subspace-based optimization method (SOM) is proposed for solving full-wave inverse scattering problems (ISPs). The unrolling network, named SOM-Net, inherently embeds the Lippmann-Schwinger physical model into the design of network structures. The SOM-Net takes the deterministic induced current and the raw permittivity image obtained from back-propagation (BP) as the input. It then updates the induced current and the permittivity successively in sub-network blocks of the SOM-Net by imitating iterations of the SOM. The final output of the SOM-Net is the full predicted induced current, from which the scattered field and the permittivity image can also be deduced analytically. The parameters of the SOM-Net are optimized in a supervised manner with the total loss to simultaneously ensure the consistency of the induced current, the scattered field, and the permittivity in the governing equations. Numerical tests on both synthetic and experimental data verify the superior performance of the proposed SOM-Net over typical ones. The results on challenging examples like scatterers with tough profiles or high permittivity demonstrate the good generalization ability of the SOM-Net. With the use of deep unrolling technology, this work builds a bridge between traditional iterative methods and deep learning methods for solving ISPs.
Yu Liu 0023, Rencheng Song, Xudong Chen 0001, Chang Li 0001, Xun Chen 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 MATR: Multimodal Medical Image Fusion via Multiscale Adaptive Transformer
abstract
Owing to the limitations of imaging sensors, it is challenging to obtain a medical image that simultaneously contains functional metabolic information and structural tissue details. Multimodal medical image fusion, an effective way to merge the complementary information in different modalities, has become a significant technique to facilitate clinical diagnosis and surgical navigation. With powerful feature representation ability, deep learning (DL)-based methods have improved such fusion results but still have not achieved satisfactory performance. Specifically, existing DL-based methods generally depend on convolutional operations, which can well extract local patterns but have limited capability in preserving global context information. To compensate for this defect and achieve accurate fusion, we propose a novel unsupervised method to fuse multimodal medical images via a multiscale adaptive Transformer termed MATR. In the proposed method, instead of directly employing vanilla convolution, we introduce an adaptive convolution for adaptively modulating the convolutional kernel based on the global complementary context. To further model long-range dependencies, an adaptive Transformer is employed to enhance the global semantic extraction capability. Our network architecture is designed in a multiscale fashion so that useful multimodal information can be adequately acquired from the perspective of different scales. Moreover, an objective function composed of a structural loss and a region mutual information loss is devised to construct constraints for information preservation at both the structural-level and the feature-level. Extensive experiments on a mainstream database demonstrate that the proposed method outperforms other representative and state-of-the-art methods in terms of both visual quality and quantitative evaluation. We also extend the proposed method to address other biomedical image fusion issues, and the pleasing fusion results illustrate that MATR has good generalization capability. The code of the proposed method is available at https://github.com/tthinking/MATR.
Wei Tang 0018, Fazhi He, Yu Liu 0023, Yansong Duan
IEEE Trans. Image Process.3
2021 Different Input Resolutions and Arbitrary Output Resolution: A Meta Learning-Based Deep Framework for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion has gained ever-increasing attention in recent years due to its great significance in a variety of vision-based applications. However, existing fusion methods suffer from some limitations in terms of the spatial resolutions of both input source images and output fused image, which prevents their practical usage to a great extent. In this paper, we propose a meta learning-based deep framework for the fusion of infrared and visible images. Unlike most existing methods, the proposed framework can accept the source images of different resolutions and generate the fused image of arbitrary resolution just with a single learned model. In the proposed framework, the features of each source image are first extracted by a convolutional network and upscaled by a meta-upscale module with an arbitrary appropriate factor according to practical requirements. Then, a dual attention mechanism-based feature fusion module is developed to combine features from different source images. Finally, a residual compensation module, which can be iteratively adopted in the proposed framework, is designed to enhance the capability of our method in detail extraction. In addition, the loss function is formulated in a multi-task learning manner via simultaneous fusion and super-resolution, aiming to improve the effect of feature learning. And, a new contrast loss inspired by a perceptual contrast enhancement approach is proposed to further improve the contrast of the fused image. Extensive experiments on widely-used fusion datasets demonstrate the effectiveness and superiority of the proposed method. The code of the proposed method is publicly available at https://github.com/yuliu316316/MetaLearning-Fusion.
Huafeng Li 0001, Yueliang Cen, Yu Liu 0023, Xun Chen 0001, Zhengtao Yu 0001
IEEE Trans. Image Process.3
2021 Emotion Recognition From Multi-Channel EEG via Deep Forest
abstract
Recently, deep neural networks (DNNs) have been applied to emotion recognition tasks based on electroencephalography (EEG), and have achieved better performance than traditional algorithms. However, DNNs still have the disadvantages of too many hyperparameters and lots of training data. To overcome these shortcomings, in this article, we propose a method for multi-channel EEG-based emotion recognition using deep forest. First, we consider the effect of baseline signal to preprocess the raw artifact-eliminated EEG signal with baseline removal. Secondly, we construct 2 D frame sequences by taking the spatial position relationship across channels into account. Finally, 2 D frame sequences are input into the classification model constructed by deep forest that can mine the spatial and temporal information of EEG signals to classify EEG emotions. The proposed method can eliminate the need for feature extraction in traditional methods and the classification model is insensitive to hyperparameter settings, which greatly reduce the complexity of emotion recognition. To verify the feasibility of the proposed model, experiments were conducted on two public DEAP and DREAMER databases. On the DEAP database, the average accuracies reach to 97.69% and 97.53% for valence and arousal, respectively; on the DREAMER database, the average accuracies reach to 89.03%, 90.41%, and 89.89% for valence, arousal and dominance, respectively. These results show that the proposed method exhibits higher accuracy than the state-of-art methods.
Juan Cheng 0004, Meiyao Chen, Chang Li 0001, Yu Liu 0023, Rencheng Song, Aiping Liu, Xun Chen 0001
IEEE J. Biomed. Health Informatics4
2021 PulseGAN: Learning to Generate Realistic Pulse Waveforms in Remote Photoplethysmography
abstract
Remote photoplethysmography (rPPG) is a non-contact technique for measuring cardiac signals from facial videos. High-quality rPPG pulse signals are urgently demanded in many fields, such as health monitoring and emotion recognition. However, most of the existing rPPG methods can only be used to get average heart rate (HR) values due to the limitation of inaccurate pulse signals. In this paper, a new framework based on generative adversarial network, called PulseGAN, is introduced to generate realistic rPPG pulse signals through denoising the chrominance (CHROM) signals. Considering that the cardiac signal is quasi-periodic and has apparent time-frequency characteristics, the error losses defined in time and spectrum domains are both employed with the adversarial loss to enforce the model generating accurate pulse waveforms as its reference. The proposed framework is tested on three public databases. The results show that the PulseGAN framework can effectively improve the waveform quality, thereby enhancing the accuracy of HR, the interbeat interval (IBI) and the related heart rate variability (HRV) features. The proposed method significantly improves the quality of waveforms compared to the input CHROM signals, with the mean absolute error of AVNN (the average of all normal-to-normal intervals) reduced by 41.19%, 40.45%, 41.63%, and the mean absolute error of SDNN (the standard deviation of all NN intervals) reduced by 37.53%, 44.29%, 58.41%, in the cross-database test on the UBFC-RPPG, PURE, and MAHNOB-HCI databases, respectively. This framework can be easily integrated with other existing rPPG methods to further improve the quality of waveforms, thereby obtaining more reliable IBI features and extending the application scope of rPPG techniques.
Rencheng Song, Juan Cheng 0004, Chang Li 0001, Yu Liu 0023, Xun Chen 0001
IEEE J. Biomed. Health Informatics5
2020 Sparse unmixing of hyperspectral data with bandwise model
Chang Li 0001, Yu Liu 0023, Juan Cheng 0004, Rencheng Song, Jiayi Ma 0001, Chenhong Sui, Xun Chen 0001
Inf. Sci.2
2020 Exploring the feasibility of seamless remote heart rate measurement using multiple synchronized cameras
Juan Cheng 0004, Xingmao Wang, Rencheng Song, Yu Liu 0023, Chang Li 0001, Xun Chen 0001
Multim. Tools Appl.4
2020 Zero-Shot Learning Based on Deep Weighted Attribute Prediction
abstract
In zero-shot learning, attributes play as a bridge from original images to class labels. Therefore, to achieve accurate zero-shot image classification, we mainly focus on improving attribute prediction accuracy by taking full advantage of prior information about attribute from two aspects. First, we present a new attribute classifier called deep attribute prediction (DeepAP) model by using supervised deep convolutional neural networks (DCNNs), where the attribute label information participates in the training of DCNNs. Unlike common DCNNs that are usually used to extract image features, the constructed DCNNs are used to directly predict attribute values from the original input images. Thus, the designed DeepAP model can serve as the mapping from low-level image features to high-level semantic attributes in the traditional direct attribute prediction (DAP) model. Second, another prior information about attribute, i.e., class-attribute matrix is used to mine the attribute-class correlation with sparse representation coefficients. Since the attribute-class correlation can reflect different contributions of attributes to classification, we use it to define attribute weights and incorporate the idea of weighted attributes into DeepAP to form the deep weighted attribute prediction (DWAP) model. Experiments on three real datasets show that DWAP outperforms the deep attribute network and DAP on attribute prediction and zero-shot image classification.
Xuesong Wang 0001, Chen Chen 0033, Yuhu Cheng 0001, Xun Chen 0001, Yu Liu 0023
IEEE Trans. Syst. Man Cybern. Syst.5
2019 Medical Image Fusion via Convolutional Sparsity Based Morphological Component Analysis
abstract
In this letter, a sparse representation (SR) model named convolutional sparsity based morphological component analysis (CS-MCA) is introduced for pixel-level medical image fusion. Unlike the standard SR model, which is based on single image component and overlapping patches, the CS-MCA model can simultaneously achieve multi-component and global SRs of source images, by integrating MCA and convolutional sparse representation (CSR) into a unified optimization framework. For each source image, in the proposed fusion method, the CSRs of its cartoon and texture components are first obtained by the CS-MCA model using pre-learned dictionaries. Then, for each image component, the sparse coefficients of all the source images are merged and the fused component is accordingly reconstructed using the corresponding dictionary. Finally, the fused image is calculated as the superposition of the fused cartoon and texture components. Experimental results demonstrate that the proposed method can outperform some benchmarking and state-of-the-art SR-based fusion methods in terms of both visual perception and objective assessment.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.1
2019 Video Super-Resolution Using Non-Simultaneous Fully Recurrent Convolutional Network
abstract
Video super-resolution (SR) aims at restoring fine details and enhancing visual experience for low-resolution (LR) videos. In this paper, we propose a very deep non-simultaneous fully recurrent convolutional network for video SR. To make full use of temporal information, we employ motion compensation, very deep fully recurrent convolutional layers and late fusion in our system. Residual connection is also employed in our recurrent structure for more accurate SR. Finally a new model ensemble strategy is used to combine our method with single-image SR method. Experimental results demonstrate that the proposed method is better than state-of-the-art SR methods on quantitative visual quality assessment.
Dingyi Li, Yu Liu 0023, Zengfu Wang
IEEE Trans. Image Process.2
2018 Bone Age Assessment with X-Ray Images Based on Contourlet Motivated Deep Convolutional Networks
abstract
Bone age assessment (BAA) is a widely performed procedure for skeletal maturity evaluation in pediatric radiology. It has various clinical applications such as diagnosis of endocrine disorders, monitoring of growth hormone therapy and prediction of final adult height for adolescents. Recent studies indicate that deep learning techniques have great potential in developing automated BAA methods with significant improvements in terms of conventional computer-assisted approaches. In this paper, we propose a multi-scale feature fusion framework for bone age assessment based on deep convolutional neural networks. In our method, the non-subsampled contourlet transform (NSCT) is firstly performed on an input left-hand radiograph to obtain its multi-scale and multi-direction representations. Then, the decomposed bands at each scale are fed to a convolutional network that contains a series of convolutional and pooling layers for feature extraction, respectively. Finally, the feature maps from different branches are concatenated and put into a regression network consisting of several fully connected layers to obtain the bone age estimation. Experimental results on a public BAA dataset demonstrate that the proposed method can achieve state-of-the-art performance.
Xun Chen 0001, Chao Zhang 0057, Yu Liu 0023
MMSP3
2017 A medical image fusion method based on convolutional neural networks
abstract
Medical image fusion technique plays an an increasingly critical role in many clinical applications by deriving the complementary information from medical images with different modalities. In this paper, a medical image fusion method based on convolutional neural networks (CNNs) is proposed. In our method, a siamese convolutional network is adopted to generate a weight map which integrates the pixel activity information from two source images. The fusion process is conducted in a multi-scale manner via image pyramids to be more consistent with human visual perception. In addition, a local similarity based strategy is applied to adaptively adjust the fusion mode for the decomposed coefficients. Experimental results demonstrate that the proposed method can achieve promising results in terms of both visual quality and objective assessment.
Yu Liu 0023, Xun Chen 0001, Juan Cheng 0004, Hu Peng
FUSION1
2017 Video super-resolution using motion compensation and residual bidirectional recurrent convolutional network
abstract
Video super-resolution (SR) aims at restoring finer details and enhancing visual experience. In this paper, we propose a novel method named residual recurrent convolutional network (RRCN) for video SR. In our method, motion compensation and bidirectional residual convolutional network are combined to model the spatial and temporal non-linear mappings. To leverage sufficient amount of temporal information, we employ motion compensation, bidirectional recurrent convolutional layers and late fusion in of our network. We also apply residual connections in our recurrent structure for more accurate SR. Experimental results demonstrate the superiority of the proposed method over state-of-the-art single-image and multi-frame based SR approaches in terms of both quantitative assessment and visual quality.
Dingyi Li, Yu Liu 0023, Zengfu Wang
ICIP2
2017 Image classification based on convolutional neural networks with cross-level strategy
Yu Liu 0023, Jun Yu 0001, Zengfu Wang
Multim. Tools Appl.1
2016 Automatic chessboard corner detection method
abstract
Chessboard corner detection is a necessary procedure of the popular chessboard pattern‐based camera calibration technique, in which the inner corners on a two‐dimensional chessboard are employed as calibration markers. In this study, an automatic chessboard corner detection algorithm is presented for camera calibration. In authors’ method, an initial corner set is first obtained with an improved Hessian corner detector. Then, a novel strategy that utilises both intensity and geometry characteristics of the chessboard pattern is presented to eliminate fake corners from the initial corner set. After that, a simple yet effective approach is adopted to sort the detected corners into a meaningful order. Finally, the sub‐pixel location of each corner is calculated. The proposed algorithm only requires a user input of the chessboard size, while all the other parameters can be adaptively calculated with a statistical approach. The experimental results demonstrate that the proposed method has advantages over the popular OpenCV chessboard corner detection method in terms of detection accuracy and computational efficiency. Furthermore, the effectiveness of the proposed method used for camera calibration is also verified in authors’ experiments.
Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang
IET Image Process.1
2016 Image Fusion With Convolutional Sparse Representation
abstract
As a popular signal modeling technique, sparse representation (SR) has achieved great success in image fusion over the last few years with a number of effective algorithms being proposed. However, due to the patch-based manner applied in sparse coding, most existing SR-based fusion methods suffer from two drawbacks, namely, limited ability in detail preservation and high sensitivity to misregistration, while these two issues are of great concern in image fusion. In this letter, we introduce a recently emerged signal decomposition model known as convolutional sparse representation (CSR) into image fusion to address this problem, which is motivated by the observation that the CSR model can effectively overcome the above two drawbacks. We propose a CSR-based image fusion framework, in which each source image is decomposed into a base layer and a detail layer, for multifocus image fusion and multimodal image fusion. Experimental results demonstrate that the proposed fusion methods clearly outperform the SR-based methods in terms of both objective assessment and visual quality.
Yu Liu 0023, Xun Chen 0001, Rabab K. Ward, Z. Jane Wang 0001
IEEE Signal Process. Lett.1
2015 Simultaneous image fusion and denoising with adaptive sparse representation
abstract
In this study, a novel adaptive sparse representation (ASR) model is presented for simultaneous image fusion and denoising. As a powerful signal modelling technique, sparse representation (SR) has been successfully employed in many image processing applications such as denoising and fusion. In traditional SR‐based applications, a highly redundant dictionary is always needed to satisfy signal reconstruction requirement since the structures vary significantly across different image patches. However, it may result in potential visual artefacts as well as high computational cost. In the proposed ASR model, instead of learning a single redundant dictionary, a set of more compact sub‐dictionaries are learned from numerous high‐quality image patches which have been pre‐classified into several corresponding categories based on their gradient information. At the fusion and denoising processes, one of the sub‐dictionaries is adaptively selected for a given set of source image patches. Experimental results on multi‐focus and multi‐modal image sets demonstrate that the ASR‐based fusion method can outperform the conventional SR‐based method in terms of both visual quality and objective assessment.
Yu Liu 0023, Zengfu Wang
IET Image Process.1
2015 Dense SIFT for ghost-free multi-exposure fusion
Yu Liu 0023, Zengfu Wang
J. Vis. Commun. Image Represent.1
2014 A practical algorithm for automatic chessboard corner detection
abstract
Chessboard corner detection is a fundamental work of the popular chessboard pattern-based camera calibration technique. In this paper, a fast and robust algorithm for chessboard corner detection is presented. In our method, an initial corner set is obtained with an improved Hessian corner detector. And then, a novel strategy which takes both textural and geometrical characteristics of a chessboard into consideration is employed to eliminate fake corners in the initial corner set. The proposed algorithm only requires a user-input of the total number of chessboard inner corners, while all the other parameters can be adaptively calculated with a statistical approach. Experimental results on two public data sets demonstrate that the proposed method can outperform the most commonly used OpenCV method in terms of both detection rate and computational efficiency.
Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang
ICIP1
2013 Multi-focus Image Fusion Based on Sparse Representation with Adaptive Sparse Domain Selection
abstract
Sparse representation (SR) has been widely used in many image processing applications including image fusion. As the contents vary significantly across different images, a highly redundant dictionary is always required in the sparse model, which reduces the algorithm stability and efficiency. This paper proposes a multi-focus image fusion method based on SR with adaptive sparse domain selection (SR-ASDS). Under SR-ASDS, numerous high-quality image patches are first classified into several categories according to their gradient information, and each category is applied into training a compact sub-dictionary. At the fusion process, a corresponding sub-dictionary is adaptively selected for a given pair of source image patches. Moreover, we present a general optimization framework for the merging rule design of the SR based image fusion. Numerous experiments on both clear images and the noisy ones demonstrate that the proposed method outperforms the fusion methods which use a single dictionary, in terms of several popular objective evaluation criteria.
Yu Liu 0023, Zengfu Wang
ICIG1