Juncheng Li 0003

dblp:182/7674-3 · DBLP profile ↗
← Back
48ranked-venue papers
5as first author
43since 2021 · last 2026
0000-0001-7314-6754ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-view doubly supervised knowledge distillation for diagnosis of liver cancers with imbalanced ultrasound imaging modalities
Lehang Guo, Juncheng Li 0003, Jun Wang 0024, Jun Shi 0004
Eng. Appl. Artif. Intell.3
2026 UCPNet: An Ultra-Lightweight Cross-Perception Network for Real-Time Semantic Segmentation
Guoan Xu, Juncheng Li 0003, Heyou Chang, Guangwei Gao
Image Vis. Comput.3
2026 Enhancing feature discrimination with pseudo-labels for foundation model in segmentation of 3D medical images
abstract
Development of medical image segmentation foundation models relies on large-scale samples. However, it is more time-consuming to annotate 3D medical images than 2D natural images, making it challenging to collect sufficient annotated samples. While pseudo-labeling offers a potential solution to expand the annotated dataset, it may introduce noisy labels that can create systematic biases, particularly affecting the segmentation performance of smaller anatomical structures. To this end, we propose a pseudo-label enriched segmentation framework (PESF), which integrates confidence filtering and perturbation-based curriculum learning. To begin with, our pseudo-labeling approach applies a well-pretrained foundation model to generate pseudo-labels for previously unannotated organ categories, effectively expanding the number of classes in the original dataset. Subsequently, we develop a confidence-based filtering mechanism, leveraging a feature extraction module combined with a confidence prediction module to quantitatively assess and filter out low-quality pseudo-labels, thereby minimizing the detrimental effects of noisy pseudo-labels on the model's optimization. Furthermore, a progressive sampling strategy that integrates curriculum learning with Gaussian random perturbations is proposed, systematically introducing training samples from simpler to more complex cases, thereby enhancing the model's generalization capability across organs of varying shapes and sizes. Additionally, our theoretical analysis reveals that incorporating these extra pseudo-labeled classes strengthens feature discrimination by increasing the angular margins between class decision boundaries in the embedding space. Experimental results demonstrate that PESF achieves a 6.8% improvement in the overall average Dice Similarity Coefficient (DSC) compared to the baseline SAM-Med3D on (Amos, FLARE22, WORD, BTCV), with particularly gains in challenging anatomical structures such as the pancreas and esophagus. The code is available at https://github.com/lonezhizi/PESF.
Ge Jin 0002, Qian Zhang 0013, Yong Cheng 0001, Yingwen Zhu, De Yu, Yongqi Yuan, Juncheng Li 0003, Jun Shi 0004
Neural Networks8
2026 Diffusion-based Laplacian frequency-aware network for low-light image enhancement
Juncheng Li 0003, Guangwei Gao, Chia-Wen Lin
Pattern Recognit.3
2025 Cross Paradigm Representation and Alignment Transformer for Image Deraining
abstract
Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications.
Shun Zou, Juncheng Li 0003, Guangwei Gao, Guo-Jun Qi
ACM Multimedia3
2025 Fast MRI reconstruction: A thorough survey from single-modal to multi-modal
Weiyi Lyu, Xinming Fang, Chaoyan Huang, Minhua Lu, Jun Wang 0024, Jun Shi 0004, Juncheng Li 0003
Expert Syst. Appl.7
2025 Multi-resolution based dual-channel UNet with cross clique for medical image dense prediction
Xueying Zhou, Ge Jin 0002, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Expert Syst. Appl.4
2025 A Trustworthy Curriculum Learning Guided Multi-Target Domain Adaptation Network for Autism Spectrum Disorder Classification
abstract
Domain adaptation has demonstrated success in classification of multi-center autism spectrum disorder (ASD). However, current domain adaptation methods primarily focus on classifying data in a single target domain with the assistance of one or multiple source domains, lacking the capability to address the clinical scenario of identifying ASD in multiple target domains. In response to this limitation, we propose a Trustworthy Curriculum Learning Guided Multi-Target Domain Adaptation (TCL-MTDA) network for identifying ASD in multiple target domains. To effectively handle varying degrees of data shift in multiple target domains, we propose a trustworthy curriculum learning procedure based on the Dempster-Shafer (D-S) Theory of Evidence. Additionally, a domain-contrastive adaptation method is integrated into the TCL-MTDA process to align data distributions between source and target domains, facilitating the learning of domain-invariant features. The proposed TCL-MTDA method is evaluated on 437 subjects (including 220 ASD patients and 217 NCs) from the Autism Brain Imaging Data Exchange (ABIDE). Experimental results validate the effectiveness of our proposed method in multi-target ASD classification, achieving an average accuracy of 71.46% (95% CI: 68.85% - 74.06%) across four target domains, significantly outperforming most baseline methods (p<0.05).
Jiale Dun, Jun Wang 0024, Juncheng Li 0003, Qianhui Yang, Wenlong Hang, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics3
2025 Topological GCN Guided Improved Conformer for Detection of Hip Landmarks From Ultrasound Images
abstract
The B-mode ultrasound based computer-aided diagnosis (CAD) has shown its effectiveness for diagnosis of Developmental Dysplasia of the Hip (DDH) in infants within 6 months. Hip landmark detection is a feasible way for the CAD of DDH according to the Graf's method. However, existing landmark detection algorithms mainly focus on designing special models to capture the features from hip ultrasound images, but generally ignore the important spatial relations among different landmarks. To this end, a novel weakly supervised learning-based algorithm, the Topological Graph Convolutional Network (TGCN) guided Improved Conformer (TGCN-ICF), is proposed for detecting landmarks from hip ultrasound images. The TGCN-ICF includes two subnetworks: an Improved Conformer (ICF) subnetwork to generate heatmaps and constraint vectors from ultrasound images, and a TGCN subnetwork to additionally explore topological relations among hip landmarks with the guidance of class labels for further refining and improving the detection accuracy. Moreover, a new Mutual Modulation Fusion (MMF) module is developed to fully exchange and fuse the extracted feature information from the convolutional neural network (CNN) and Transformer branches in ICF. Meanwhile, a novel Mutual Supervision Constraint (MSC) strategy is designed to provide a constraint for detection of each hip landmark. The experimental results on two real-world DDH datasets demonstrate that the TGCN-ICF outperforms all the compared algorithms, suggesting its potential applications.
Tianxiang Huang, Ge Jin 0002, Juncheng Li 0003, Jun Wang 0024, Qian Wang 0001, Jun Du 0006, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2025 High-Frequency Modulated Transformer for Multi-Contrast MRI Super-Resolution
abstract
Accelerating the MRI acquisition process is always a key issue in modern medical practice, and great efforts have been devoted to fast MR imaging. Among them, multi-contrast MR imaging is a promising and effective solution that utilizes and combines information from different contrasts. However, existing methods may ignore the importance of the high-frequency priors among different contrasts. Moreover, they may lack an efficient method to fully utilize the information from the reference contrast. In this paper, we propose a lightweight and accurate High-frequency Modulated Transformer (HFMT) for multi-contrast MRI super-resolution. The key ideas of HFMT are high-frequency prior enhancement and its fusion with global features. Specifically, we employ an enhancement module to enhance and amplify the high-frequency priors in the reference and target modalities. In addition, we utilize the Rectangle Window Transformer Block (RWTB) to capture global information in the target contrast. Meanwhile, we propose a novel cross-attention mechanism to fuse the high-frequency enhanced features with the global features sequentially, which assists the network in recovering clear texture details from the low-resolution inputs. Extensive experiments show that our proposed method can reconstruct high-quality images with fewer parameters and faster inference time.
Juncheng Li 0003, Hanhui Yang, Qiaosi Yi, Minhua Lu, Jun Shi 0004, Tieyong Zeng
IEEE Trans. Medical Imaging1
2025 Efficient Image Super-Resolution With Feature Interaction Weighted Hybrid Network
abstract
Lightweight image super-resolution aims to reconstruct high-resolution images from low-resolution images using low computational costs. However, existing methods result in the loss of middle-layer features due to activation functions. To minimize the impact of intermediate feature loss on reconstruction quality, we propose a Feature Interaction Weighted Hybrid Network (FIWHN), which comprises a series of Wide-residual Distillation Interaction Block (WDIB) as the backbone. Every third WDIB forms a Feature Shuffle Weighted Group (FSWG) by applying mutual information shuffle and fusion. Moreover, to mitigate the negative effects of intermediate feature loss, we introduce Wide Residual Weighting units within WDIB. These units effectively fuse features of varying levels of detail through a Wide-residual Distillation Connection (WRDC) and a Self-Calibrating Fusion (SCF). To compensate for global feature deficiencies, we incorporate a Transformer and explore a novel architecture to combine CNN and Transformer. We show that our FIWHN achieves a favorable balance between performance and efficiency through extensive experiments on low-level and high-level tasks.
Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jian Yang 0003, Guo-Jun Qi, Chia-Wen Lin
IEEE Trans. Multim.2
2024 Learning Coupled Dictionaries from Unpaired Data for Image Super-Resolution
abstract
The difficulty of acquiring high-resolution (HR) and low-resolution (LR) image pairs in real scenarios limits the performance of existing learning-based image super-resolution (SR) methods in the real world. To conduct training on real-world unpaired data, current methods focus on synthesizing pseudo LR images to associate unpaired images. However, the realness and diversity of pseudo LR images are vulnerable due to the large image space. In this paper, we cir-cumvent the difficulty of image generation and propose an alternative to build the connection between unpaired images in a compact proxy space. Specifically, we first construct coupled HR and LR dictionaries, and then encode HR and LR images into a common latent code space using these dictionaries. In addition, we develop an autoencoder-based framework to couple these dictionaries during optimization by reconstructing input HR and LR images. The coupled dictionaries enable our method to employ a shal-low network architecture with only 18 layers to achieve efficient image SR. Extensive experiments show that our method (DictSR) can effectively model the LR-to-HR mapping in coupled dictionaries and produces state-of-the-art performance on benchmark datasets.
Longguang Wang, Juncheng Li 0003, Yingqian Wang 0002, Qingyong Hu, Yulan Guo
CVPR2
2024 Topological GCN for Improving Detection of Hip Landmarks from B-Mode Ultrasound Images
Tianxiang Huang, Ge Jin 0002, Juncheng Li 0003, Jun Wang 0024, Jun Du 0006, Jun Shi 0004
MICCAI (5)4
2024 Enhanced dual contrast representation learning with cell separation and merging for breast cancer diagnosis
Yang Liu 0119, Yiqi Zhu, Zhehao Gu, Jinshan Pan, Juncheng Li 0003, Ming Fan 0003, Lihua Li 0002, Tieyong Zeng
Comput. Vis. Image Underst.5
2024 Few sampling meshes-based 3D tooth segmentation via region-aware graph convolutional network
Bodong Cheng, Najun Niu, Jun Wang 0024, Tieyong Zeng, Guixu Zhang, Jun Shi 0004, Juncheng Li 0003
Expert Syst. Appl.8
2024 Guest Editorial: Advanced image restoration and enhancement in the wild
abstract
Image restoration and enhancement has always been a fundamental task in computer vision and is widely used in numerous applications, such as surveillance imaging, remote sensing, and medical imaging. In recent years, remarkable progress has been witnessed with deep learning techniques. Despite the promising performance achieved on synthetic data, compelling research challenges remain to be addressed in the wild. These include: (i) degradation models for low-quality images in the real world are complicated and unknown, (ii) paired low-quality and high-quality data are difficult to acquire in the real world, and a large quantity of real data are provided in an unpaired form, (iii) it is challenging to incorporate cross-modal information provided by advanced imaging techniques (e.g. RGB-D camera) for image restoration, (iv) real-time inference on edge devices is important for image restoration and enhancement methods, and (v) it is difficult to provide the confidence or performance bounds of a learning-based method on different images/regions. This special issue invites original contributions in datasets, innovative architectures, and training methods for image restoration and enhancement to address these and other challenges. In this Special Issue, we have received 17 papers, of which 8 papers underwent the peer review process, while the rest were desk-rejected. Among these reviewed papers, 5 papers have been accepted and 3 papers have been rejected as they did not meet the criteria of IET Computer Vision. Thus, the overall submissions were of high quality, which marks the success of this Special Issue. The five eventually accepted papers can be clustered into two categories, namely video reconstruction and image super-resolution. The first category of papers aims at reconstructing high-quality videos. The papers in this category are of Zhang et al., Gu et al., and Xu et al. The second category of papers studies the task of image super-resolution. The papers in this category are of Dou et al. and Yang et al. A brief presentation of each of the paper in this special issue is as follows. Zhang et al. propose a point-image fusion network for event-based frame interpolation. Temporal information in event streams plays a critical role in this task as it provides temporal context cues complementary to images. Previous approaches commonly transform the unstructured event data to structured data formats through voxelisation and then employ advanced CNNs to extract temporal information. However, the voxelisation operation inevitably leads to information loss and introduces redundant computation. To address these limitations, the proposed method directly extracts temporal information from the events at the point level without relying on any voxelisation operation. Afterwards, a fusion module is adopted to aggregate complementary cues from both points and images for frame interpolation. Experiments on both synthetic and real-world datasets show that their method produces state-of-the-art accuracy with high efficiency. Gu et al. develop a temporal shift reconstruction network for compressive video sensing. To exploit the temporal cues between adjacent frames during the reconstruction of videos, most previous approaches commonly preform alignment between initial reconstructions. However, the estimated motions are usually too coarse to provide accurate temporal information. To remedy this, the proposed network employs stacked temporal shift reconstruction blocks to enhance the initial reconstruction progressively. Within each block, an efficient temporal shift operation is used to capture temporal structures in addition to computational overheads. Then, a bidirectional alignment module is adopted to capture the temporal dependencies in a video sequence. Different from previous methods that only extract supplementary information from the key frames, the proposed alignment module can receive temporal information from the whole video sequence via bidirectional propagations. Experiments demonstrate the superior performance of the proposed method. Qu et al. propose a lightweight video frame interpolation network with a three-scale encoding-decoding structure. Specifically, multi-scale motion information is first extracted from the input video. Then, recurrent convolutional layers are adopted to refine the resultant features. Afterwards, the resultant features are aggregated to generate high-quality interpolated frames. Experimental results on the CelebA and Helen datasets show that the proposed method outperforms state-of-the-art methods while using fewer parameters. Dou et al. introduce a decoder structure-guided CNN-Transformer network for face super-resolution. Most previous approaches follow a multi-task learning paradigm to perform landmark detection while super-resolving the low-resolution images. However, these methods require additional annotation cost, and the extracted facial prior structures are usually of low quality. To address these issues, the proposed network employs a global-local feature extraction unit to extract the global structure while capturing local texture details. In addition, a multi-state fusion module is incorporated to aggregate embeddings from different stages. Experiments show that the proposed method surpasses previous approaches by notable margins. Yang et al. study the problem of blind super-resolution and propose a method to exploit degradation information through degradation representation learning. Specifically, a generative adversarial network is employed to model the degradation process from HR images to LR images and constrain the data distribution of the synthetic LR images. Then, the learnt representation is adopted to super-resolve the input low-resolution images using a transformer-based SR network. Experiments on both synthetic and real-world datasets demonstrate the effectiveness and superiority of the proposed method. Longguang Wang received his BE and PhD degrees from Shandong University and National University of Defence Technology (NUDT) in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 40 peer-reviewed journals and conference publications (including TPAMI, TIP, CVPR, ICCV, and ECCV). He has organised three workshops at CVPR 2022 and 2023. His research interests include low-level vision and 3D vision, particularly on image restoration, image enhancement, image generation, depth estimation, point cloud understanding, and network acceleration. He received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM, and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, and the winner of 2019 ICCV-AIM. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an assistant professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany and at the Technical University of Munich, Munich, Germany. He is currently a lecturer with The University of Tokyo and a unit leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organised by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an associate editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at the University of Wurzburg, Germany. Also, he is a lecturer and a group leader at ETH Zurich, Switzerland. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI, and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of the Alexander von Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organiser of NTIRE, CLIC, AIM, Mobile AI, and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, and image/video compression, restoration, enhancement, and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from NUDT in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modelling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organised several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022, and ECCV 2022. He is a senior member of IEEE and ACM. Data sharing is not applicable to this article as no new data were created or analysed in this study. Longguang Wang received his B.E. and Ph.D. degrees from Shandong University and National University of Defense Technology in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 60 peer reviewed journal and conference publications (including TPAMI, TIP, CVPR, ICCV and ECCV). He served as a reviewer for more than 10 international journals (including TPAMI and TIP) and conferences (including CVPR, ICCV and ECCV). He has organized workshops at CVPR 2022/2023/2024. His research interests include low-level vision and 3D vision, particularly on image restoration, image generation, point cloud understanding, and network acceleration. His received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationalwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, the winner of 2019 ICCV-AIM, etc. Meanwhile, he served as a reviewer for more than 20 international journals and conferences. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an Assistant Professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany, and at the Technical University of Munich, Munich, Germany. He is currently a Lecturer with The University of Tokyo, and a Unit Leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images, with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organized by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an Associate Editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at theUniversity of Wurzburg, Germany. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of an Alexandervon Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organizer of NTIRE, CLIC, AIM, Mobile AI and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, image/video compression, restoration, enhancement and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from National University of Defense Technology (NUDT) in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modeling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organized several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022 and ECCV 2022. He is a Senior Member of IEEE and ACM.
Longguang Wang, Juncheng Li 0003, Naoto Yokoya, Radu Timofte, Yulan Guo
IET Comput. Vis.2
2024 Multi-View disentanglement-based bidirectional generalized distillation for diagnosis of liver cancers with ultrasound images
Lehang Guo, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Inf. Process. Manag.3
2024 EWT: Efficient Wavelet-Transformer for single image denoising
Juncheng Li 0003, Bodong Cheng, Guangwei Gao, Jun Shi 0004, Tieyong Zeng
Neural Networks1
2024 WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution Diffusion
abstract
The extraction of distribution from images with diverse weather conditions is crucial for enhancing the robustness of visual algorithms. When addressing image degradation caused by different weather, accurately perceiving the data distribution of weather-informed degradation becomes a fundamental challenge. However, given the highly stochastic nature, modelling weather distribution poses a formidable task. In this paper, we propose a novel multi-Weather distribution difFUsion blind restoration model, named WeaFU. Firstly, the model employs representation learning to map image distribution into a latent space. Subsequently, WeaFU utilizes a diffusion-based approach, with the assistance of Diffusion Distribution Generator (DDG), to perceive and extract corresponding weather distribution. This strategy ingeniously injects data distribution into the recovery process, significantly enhancing the robustness of the model in diverse weather scenarios. Finally, a Conditional Distribution-Aware Transformer (CDAT) is constructed to align the distribution information with pixels, thereby obtaining clear images. Extensive experiments on real and synthetic datasets demonstrate that WeaFU achieves superior performance.
Bodong Cheng, Juncheng Li 0003, Jun Shi 0004, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Circuits Syst. Video Technol.2
2024 Involution Transformer Based U-Net for Landmark Detection in Ultrasound Images for Diagnosis of Infantile DDH
abstract
The B-mode ultrasound based computer-aided diagnosis (CAD) has demonstrated its effectiveness for diagnosis of Developmental Dysplasia of the Hip (DDH) in infants, which can conduct the Graf's method by detecting landmarks in hip ultrasound images. However, it is still necessary to explore more valuable information around these landmarks to enhance feature representation for improving detection performance in the detection model. To this end, a novel Involution Transformer based U-Net (IT-UNet) network is proposed for hip landmark detection. The IT-UNet integrates the efficient involution operation into Transformer to develop an Involution Transformer module (ITM), which consists of an involution attention block and a squeeze-and-excitation involution block. The ITM can capture both the spatial-related information and long-range dependencies from hip ultrasound images to effectively improve feature representation. Moreover, an Involution Downsampling block (IDB) is developed to alleviate the issue of feature loss in the encoder modules, which combines involution and convolution for the purpose of downsampling. The experimental results on two DDH ultrasound datasets indicate that the proposed IT-UNet achieves the best landmark detection performance, indicating its potential applications.
Tianxiang Huang, Juncheng Li 0003, Jun Wang 0024, Jun Du 0006, Jun Shi 0004
IEEE J. Biomed. Health Informatics3
2024 Multimodal Co-Attention Fusion Network With Online Data Augmentation for Cancer Subtype Classification
abstract
It is an essential task to accurately diagnose cancer subtypes in computational pathology for personalized cancer treatment. Recent studies have indicated that the combination of multimodal data, such as whole slide images (WSIs) and multi-omics data, could achieve more accurate diagnosis. However, robust cancer diagnosis remains challenging due to the heterogeneity among multimodal data, as well as the performance degradation caused by insufficient multimodal patient data. In this work, we propose a novel multimodal co-attention fusion network (MCFN) with online data augmentation (ODA) for cancer subtype classification. Specifically, a multimodal mutual-guided co-attention (MMC) module is proposed to effectively perform dense multimodal interactions. It enables multimodal data to mutually guide and calibrate each other during the integration process to alleviate inter- and intra-modal heterogeneities. Subsequently, a self-normalizing network (SNN)-Mixer is developed to allow information communication among different omics data and alleviate the high-dimensional small-sample size problem in multi-omics data. Most importantly, to compensate for insufficient multimodal samples for model training, we propose an ODA module in MCFN. The ODA module leverages the multimodal knowledge to guide the data augmentations of WSIs and maximize the data diversity during model training. Extensive experiments are conducted on the public TCGA dataset. The experimental results demonstrate that the proposed MCFN outperforms all the compared algorithms, suggesting its effectiveness.
Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE Trans. Medical Imaging2
2024 Cost-Sensitive Weighted Contrastive Learning Based on Graph Convolutional Networks for Imbalanced Alzheimer's Disease Staging
abstract
Identifying the progression stages of Alzheimer's disease (AD) can be considered as an imbalanced multi-class classification problem in machine learning. It is challenging due to the class imbalance issue and the heterogeneity of the disease. Recently, graph convolutional networks (GCNs) have been successfully applied in AD classification. However, these works did not handle the class imbalance issue in classification. Besides, they ignore the heterogeneity of the disease. To this end, we propose a novel cost-sensitive weighted contrastive learning method based on graph convolutional networks (CSWCL-GCNs) for imbalanced AD staging using resting-state functional magnetic resonance imaging (rs-fMRI). The proposed method is developed on a multi-view graph constructed by the functional connectivity (FC) and high-order functional connectivity (HOFC) features of the subjects. A novel cost-sensitive weighted contrastive learning procedure is proposed to capture discriminative information from the minority classes, encouraging the samples in the minority class to provide adequate supervision. Considering the heterogeneity of the disease, the weights of the negative pairs are introduced into contrastive learning and they are computed based on the distance to class prototypes, which are automatically learned from the training data. Meanwhile, the cost-sensitive mechanism is further introduced into contrastive learning to handle the class imbalance issue. The proposed CSWCL-GCN is evaluated on 720 subjects (including 184 NCs, 40 SMC patients, 208 EMCI patients, 172 LMCI patients and 116 AD patients) from the ADNI (Alzheimer's Disease Neuroimaging Initiative). Experimental results show that the proposed CSWCL-GCN outperforms state-of-the-art methods on the ADNI database.
Jun Wang 0024, Juncheng Li 0003, Jun Shi 0004
IEEE Trans. Medical Imaging4
2024 Weakly Supervised Lesion Detection and Diagnosis for Breast Cancers With Partially Annotated Ultrasound Images
abstract
Deep learning (DL) has proven highly effective for ultrasound-based computer-aided diagnosis (CAD) of breast cancers. In an automatic CAD system, lesion detection is critical for the following diagnosis. However, existing DL-based methods generally require voluminous manually-annotated region of interest (ROI) labels and class labels to train both the lesion detection and diagnosis models. In clinical practice, the ROI labels, i.e. ground truths, may not always be optimal for the classification task due to individual experience of sonologists, resulting in the issue of coarse annotation to limit the diagnosis performance of a CAD model. To address this issue, a novel Two-Stage Detection and Diagnosis Network (TSDDNet) is proposed based on weakly supervised learning to improve diagnostic accuracy of the ultrasound-based CAD for breast cancers. In particular, all the initial ROI-level labels are considered as coarse annotations before model training. In the first training stage, a candidate selection mechanism is then designed to refine manual ROIs in the fully annotated images and generate accurate pseudo-ROIs for the partially annotated images under the guidance of class labels. The training set is updated with more accurate ROI labels for the second training stage. A fusion network is developed to integrate detection network and classification network into a unified end-to-end framework as the final CAD model in the second training stage. A self-distillation strategy is designed on this model for joint optimization to further improves its diagnosis performance. The proposed TSDDNet is evaluated on three B-mode ultrasound datasets, and the experimental results indicate that it achieves the best performance on both lesion detection and diagnosis tasks, suggesting promising application potential.
Jian Wang 0135, Shichong Zhou, Jun Wang 0024, Juncheng Li 0003, Shihui Ying, Cai Chang, Jun Shi 0004
IEEE Trans. Medical Imaging6
2024 Pseudo-Data Based Self-Supervised Federated Learning for Classification of Histopathological Images
abstract
Computer-aided diagnosis (CAD) can help pathologists improve diagnostic accuracy together with consistency and repeatability for cancers. However, the CAD models trained with the histopathological images only from a single center (hospital) generally suffer from the generalization problem due to the straining inconsistencies among different centers. In this work, we propose a pseudo-data based self-supervised federated learning (FL) framework, named SSL-FT-BT, to improve both the diagnostic accuracy and generalization of CAD models. Specifically, the pseudo histopathological images are generated from each center, which contain both inherent and specific properties corresponding to the real images in this center, but do not include the privacy information. These pseudo images are then shared in the central server for self-supervised learning (SSL) to pre-train the backbone of global mode. A multi-task SSL is then designed to effectively learn both the center-specific information and common inherent representation according to the data characteristics. Moreover, a novel Barlow Twins based FL (FL-BT) algorithm is proposed to improve the local training for the CAD models in each center by conducting model contrastive learning, which benefits the optimization of the global model in the FL procedure. The experimental results on four public histopathological image datasets indicate the effectiveness of the proposed SSL-FL-BT on both diagnostic accuracy and generalization.
Xiangmin Han, Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE Trans. Medical Imaging5
2024 Cross-Receptive Focused Inference Network for Lightweight Image Super-Resolution
abstract
Recently, Transformer-based methods have shown impressive performance in single image super-resolution (SISR) tasks due to the ability of global feature extraction. However, the capabilities of Transformers that need to incorporate contextual information to extract features dynamically are neglected. To address this issue, we propose a lightweight Cross-receptive Focused Inference Network (CFIN) that consists of a cascade of CT Blocks mixed with CNN and Transformer. Specifically, in the CT block, we first propose a CNN-based Cross-Scale Information Aggregation Module (CIAM) to enable the model to better focus on potentially helpful information to improve the efficiency of the Transformer phase. Then, we design a novel Cross-receptive Field Guided Transformer (CFGT) to enable the selection of contextual information required for reconstruction by using a modulated convolutional kernel that understands the current semantic information and exploits the information interaction within different self-attention. Extensive experiments have shown that our proposed CFIN can effectively reconstruct images using contextual information, and it can strike a good balance between computational cost and model performance as an efficient model.
Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jiantao Zhou 0001, Jian Yang 0003, Guo-Jun Qi
IEEE Trans. Multim.2
2023 Snow Mask Guided Adaptive Residual Network for Image Snow Removal
Bodong Cheng, Juncheng Li 0003, Tieyong Zeng
Comput. Vis. Image Underst.2
2023 CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-Resolution
abstract
Recently, deep convolution neural networks (CNNs) steered face super-resolution methods have achieved great progress in restoring degraded facial details by joint training with facial priors. However, these methods have some obvious limitations. On the one hand, multi-task joint learning requires additional marking on the dataset, and the introduced prior network will significantly increase the computational cost of the model. On the other hand, the limited receptive field of CNN will reduce the fidelity and naturalness of the reconstructed facial images, resulting in suboptimal reconstructed images. In this work, we propose an efficient CNN-Transformer Cooperation Network (CTCNet) for face super-resolution tasks, which uses the multi-scale connected encoder-decoder architecture as the backbone. Specifically, we first devise a novel Local-Global Feature Cooperation Module (LGCM), which is composed of a Facial Structure Attention Unit (FSAU) and a Transformer block, to promote the consistency of local facial detail and global facial structure restoration simultaneously. Then, we design an efficient Feature Refinement Module (FRM) to enhance the encoded features. Finally, to further improve the restoration of fine facial details, we present a Multi-scale Feature Fusion Unit (MFFU) to adaptively fuse the features from different stages in the encoder procedure. Extensive evaluations on various datasets have assessed that the proposed CTCNet can outperform other state-of-the-art methods significantly. Source code will be available at https://github.com/IVIPLab/CTCNet.
Guangwei Gao, Zixiang Xu, Juncheng Li 0003, Jian Yang 0003, Tieyong Zeng, Guo-Jun Qi
IEEE Trans. Image Process.3
2023 Multi-Scale Efficient Graph-Transformer for Whole Slide Image Classification
abstract
The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still cannot work well on the gigapixel WSIs due to their extremely large image sizes. To this end, we propose a novel Multi-scale Efficient Graph-Transformer (MEGT) framework for WSI classification. The key idea of MEGT is to adopt two independent efficient Graph-based Transformer (EGT) branches to process the low-resolution and high-resolution patch embeddings (i.e., tokens in a Transformer) of WSIs, respectively, and then fuse these tokens via a multi-scale feature fusion module (MFFM). Specifically, we design an EGT to efficiently learn the local-global information of patch tokens, which integrates the graph representation into Transformer to capture spatial-related information of WSIs. Meanwhile, we propose a novel MFFM to alleviate the semantic gap among different resolution patches during feature fusion, which creates a non-patch token for each branch as an agent to exchange information with another branch by cross-attention mechanism. In addition, to expedite network training, a new token pruning module is developed in EGT to reduce the redundant tokens. Extensive experiments on both TCGA-RCC and CAMELYON16 datasets demonstrate the effectiveness of the proposed MEGT.
Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics2
2023 Lightweight Real-Time Semantic Segmentation Network With Efficient Transformer and CNN
abstract
In the past decade, convolutional neural networks (CNNs) have shown prominence for semantic segmentation. Although CNN models have very impressive performance, the ability to capture global representation is still insufficient, which results in suboptimal results. Recently, Transformer achieved huge success in NLP tasks, demonstrating its advantages in modeling long-range dependency. Recently, Transformer has also attracted tremendous attention from computer vision researchers who reformulate the image processing tasks as a sequence-to-sequence prediction but resulted in deteriorating local feature details. In this work, we propose a lightweight real-time semantic segmentation network called LETNet. LETNet combines a U-shaped CNN with Transformer effectively in a capsule embedding style to compensate for respective deficiencies. Meanwhile, the elaborately designed Lightweight Dilated Bottleneck (LDB) module and Feature Enhancement (FE) module cultivate a positive impact on training from scratch simultaneously. Extensive experiments performed on challenging datasets demonstrate that LETNet achieves superior performances in accuracy and efficiency balance. Specifically, It only contains 0.95M parameters and 13.6G FLOPs but yields 72.8% mIoU at 120 FPS on the Cityscapes test set and 70.5% mIoU at 250 FPS on the CamVid test dataset using a single RTX 3090 GPU. Source code will be available athttps://github.com/IVIPLab/LETNet.
Guoan Xu, Juncheng Li 0003, Guangwei Gao, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001
IEEE Trans. Intell. Transp. Syst.2
2023 FBSNet: A Fast Bilateral Symmetrical Network for Real-Time Semantic Segmentation
abstract
Real-time semantic segmentation, which can be visually understood as the pixel-level classification task on the input image, currently has broad application prospects, especially in the fast-developing fields of autonomous driving and drone navigation. However, the huge burden of calculation together with redundant parameters are still the obstacles to its technological development. In this article, we propose a Fast Bilateral Symmetrical Network (FBSNet) to alleviate the above challenges. Specifically, FBSNet employs a symmetrical encoder-decoder structure with two branches, semantic information branch and spatial detail branch. The Semantic Information Branch (SIB) is the main branch with semantic architecture to acquire the contextual information of the input image and meanwhile acquire sufficient receptive field. While the Spatial Detail Branch (SDB) is a shallow and simple network used to establish local dependencies of each pixel for preserving details, which is essential for restoring the original resolution during the decoding phase. Meanwhile, a Feature Aggregation Module (FAM) is designed to effectively combine the output of these two branches. Experimental results of Cityscapes and CamVid show that the proposed FBSNet can strike a good balance between accuracy and efficiency. Specifically, it obtains 70.9% and 68.9% mIoU along with the inference speed of 90 fps and 120 fps on these two test datasets, respectively, with only 0.62 million parameters on a single RTX 2080Ti GPU. The code is available athttps://github.com/IVIPLab/FBSNet.
Guangwei Gao, Guoan Xu, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003
IEEE Trans. Multim.3
2023 Lightweight Feature De-redundancy and Self-calibration Network for Efficient Image Super-resolution
abstract
In recent years, thanks to the inherent powerful feature representation and learning abilities of the convolutional neural network (CNN), deep CNN-steered single image super-resolution approaches have achieved remarkable performance improvements. However, these methods are often accompanied by large consumption of computing and memory resources, which is difficult to be adopted in real-world application scenes. To handle this issue, we design an efficient Feature De-redundancy and Self-calibration Super-resolution network (FDSCSR). In particular, a Feature De-redundancy and Self-calibration Block (FDSCB) is proposed to reduce the repetitive feature information extracted by the model and further enhance the efficiency of the model. Then, based on FDSCB, a Local Feature Fusion Module is presented to elaborately utilize and fuse the feature information extracted by each FDSCB. Abundant experiments on benchmarks have demonstrated that our FDSCSR achieves superior performance with relatively less computational consumption and storage resource than other state-of-the-art approaches. The code is available at https://github.com/IVIPLab/FDSCSR .
Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Huimin Lu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution
abstract
Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under the constraints of limited parameters and calculations is also a huge challenge. To alleviate these issues, we propose a lightweight yet efficient Feature Distillation Interaction Weighted Network (FDIWN). Specifically, FDIWN utilizes a series of specially designed Feature Shuffle Weighted Groups (FSWG) as the backbone, and several novel mutual Wide-residual Distillation Interaction Blocks (WDIB) form an FSWG. In addition, Wide Identical Residual Weighting (WIRW) units and Wide Convolutional Residual Weighting (WCRW) units are introduced into WDIB for better feature distillation. Moreover, a Wide-Residual Distillation Connection (WRDC) framework and a Self-Calibration Fusion (SCF) unit are proposed to interact features with different scales more flexibly and efficiently. Extensive experiments show that our FDIWN is superior to other models to strike a good balance between model performance and efficiency. The code is available at https://github.com/IVIPLab/FDIWN.
Guangwei Gao, Juncheng Li 0003, Fei Wu 0004, Huimin Lu 0001, Yi Yu 0001
AAAI3
2022 Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive Transformer
abstract
Single-image super-resolution (SISR) has achieved significant breakthroughs with the development of deep learning. However, these methods are difficult to be applied in real-world scenarios since they are inevitably accompanied by the problems of computational and memory costs caused by the complex operations. To solve this issue, we propose a Lightweight Bimodal Network (LBNet) for SISR. Specifically, an effective Symmetric CNN is designed for local feature extraction and coarse image reconstruction. Meanwhile, we propose a Recursive Transformer to fully learn the long-term dependence of images thus the global information can be fully used to further refine texture details. Studies show that the hybrid of CNN and Transformer can build a more efficient model. Extensive experiments have proved that our LBNet achieves more prominent performance than other state-of-the-art methods with a relatively low computational cost and memory consumption. The code is available at https://github.com/IVIPLab/LBNet.
Guangwei Gao, Zhengxue Wang, Juncheng Li 0003, Yi Yu 0001, Tieyong Zeng
IJCAI3
2022 Adjustable super-resolution network via deep supervised learning and progressive self-distillation
Juncheng Li 0003, Faming Fang, Tieyong Zeng, Guixu Zhang, Xizhao Wang
Neurocomputing1
2022 Multi-Scale Grid Network for Image Deblurring With High-Frequency Guidance
abstract
It has been demonstrated that the blurring process reduces the high-frequency information of the original sharp image, so the main challenge for image deblurring is to reconstruct high-frequency information from the blurry image. In this paper, we propose a novel image deblurring framework to focus on the reconstruction of high-frequency information, which consists of two main subnetworks: a high-frequency reconstruction subnetwork (HFRSN) and a multi-scale grid subnetwork (MSGSN). The HFRSN is built to reconstruct latent high-frequency information from multiple scale blurry images. The MSGSN performs deblurring processes with high-frequency guidance at different scales simultaneously. Besides, in order to better use high-frequency information to restore sharpening images, we designed a high-frequency information aggregation (HFAG) module and a high-frequency information attention (HFAT) module in MSGSN. The HFAG module is designed to fuse high-frequency features and image features at the feature extraction stage, and the HFAT module is built to enhance the feature reconstruction stage. Extensive experiments on different datasets show the effectiveness and efficiency of our method.
Yang Liu 0289, Faming Fang, Tingting Wang 0007, Juncheng Li 0003, Yun Sheng, Guixu Zhang
IEEE Trans. Multim.4
2022 Efficient and Accurate Multi-Scale Topological Network for Single Image Dehazing
abstract
Single image dehazing is a challenging ill-posed problem that has drawn significant attention in the last few years. Recently, convolutional neural networks have achieved great success in image dehazing. However, it is still difficult for these increasingly complex models to recover accurate details from the hazy image. In this paper, we pay attention to the feature extraction and utilization of the input image itself. To achieve this, we propose a Multi-scale Topological Network (MSTN) to fully explore the features at different scales. Meanwhile, we design a Multi-scale Feature Fusion Module (MFFM) and an Adaptive Feature Selection Module (AFSM) to achieve the selection and fusion of features at different scales, so as to achieve progressive image dehazing. This topological network provides a large number of search paths that enable the network to extract abundant image features as well as strong fault tolerance and robustness. In addition, ASFM and MFFM can adaptively select important features and ignore interference information when fusing different scale representations. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods.
Qiaosi Yi, Juncheng Li 0003, Faming Fang, Aiwen Jiang, Guixu Zhang
IEEE Trans. Multim.2
2021 Structure-Preserving Deraining with Residue Channel Prior Guidance
abstract
Single image deraining is important for many high-level computer vision tasks since the rain streaks can severely degrade the visibility of images, thereby affecting the recognition and analysis of the image. Recently, many CNN-based methods have been proposed for rain removal. Although these methods can remove part of the rain streaks, it is difficult for them to adapt to real-world scenarios and restore high-quality rain-free images with clear and accurate structures. To solve this problem, we propose a Structure-Preserving Deraining Network (SPDNet) with RCP guidance. SPDNet directly generates high-quality rain-free images with clear and accurate structures under the guidance of RCP but does not rely on any rain-generating assumptions. Specifically, we found that the RCP of images contains more accurate structural information than rainy images. Therefore, we introduced it to our deraining network to protect structure information of the rain-free image. Meanwhile, a Wavelet-based Multi-Level Module (WMLM) is proposed as the backbone for learning the background information of rainy images and an Interactive Fusion Module (IFM) is designed to make full use of RCP information. In addition, an iterative guidance strategy is proposed to gradually improve the accuracy of RCP, refining the result in a progressive path. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Code: https://github.com/Joyies/SPDNet
Qiaosi Yi, Juncheng Li 0003, Qinyan Dai, Faming Fang, Guixu Zhang, Tieyong Zeng
ICCV2
2021 Lightweight Image Super-Resolution with Multi-Scale Feature Interaction Network
abstract
Recently, the single image super-resolution (SISR) approaches with deep and complex convolutional neural network structures have achieved promising performance. However, those methods improve the performance at the cost of higher memory consumption, which is difficult to be applied for some mobile devices with limited storage and computing resources. To solve this problem, we present a lightweight multi-scale feature interaction network (MSFIN). For lightweight SISR, MSFIN expands the receptive field and adequately exploits the informative features of the low-resolution observed images from various scales and interactive connections. In addition, we design a lightweight recurrent residual channel attention block (RRCAB) so that the network can benefit from the channel attention mechanism while being sufficiently lightweight. Extensive experiments on some benchmarks have confirmed that our proposed MSFIN can achieve comparable performance against the state-of-the-arts with a more lightweight model.
Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001
ICME3
2021 Feedback Network for Mutually Boosted Stereo Image Super-Resolution and Disparity Estimation
abstract
Under stereo settings, the problem of image super-resolution (SR) and disparity estimation are interrelated that the result of each problem could help to solve the other. The effective exploitation of correspondence between different views facilitates the SR performance, while the high-resolution (HR) features with richer details benefit the correspondence estimation. According to this motivation, we propose a Stereo Super-Resolution and Disparity Estimation Feedback Network (SSRDE-FNet), which simultaneously handles the stereo image super-resolution and disparity estimation in a unified framework and interact them with each other to further improve their performance. Specifically, the SSRDE-FNet is composed of two dual recursive sub-networks for left and right views. Besides the cross-view information exploitation in the low-resolution (LR) space, HR representations produced by the SR process are utilized to perform HR disparity estimation with higher accuracy, through which the HR features can be aggregated to generate a finer SR result. Afterward, the proposed HR Disparity Information Feedback (HRDIF) mechanism delivers information carried by HR disparity back to previous layers to further refine the SR image reconstruction. Extensive experiments demonstrate the effectiveness and advancement of SSRDE-FNet.
Qinyan Dai, Juncheng Li 0003, Qiaosi Yi, Faming Fang, Guixu Zhang
ACM Multimedia2
2021 Edge-guided Composition Network for Image Stitching
Qinyan Dai, Faming Fang, Juncheng Li 0003, Guixu Zhang, Aimin Zhou
Pattern Recognit.3
2021 MDCN: Multi-Scale Dense Cross Network for Image Super-Resolution
abstract
Convolutional neural networks have been proven to be of great benefit for single-image super-resolution (SISR). However, previous works do not make full use of multi-scale features and ignore the inter-scale correlation between different upsampling factors, resulting in sub-optimal performance. Instead of blindly increasing the depth of the network, we are committed to mining image features and learning the inter-scale correlation between different upsampling factors. To achieve this, we propose a Multi-scale Dense Cross Network (MDCN), which achieves great performance with fewer parameters and less execution time. MDCN consists of multi-scale dense cross blocks (MDCBs), hierarchical feature distillation block (HFDB), and dynamic reconstruction block (DRB). Among them, MDCB aims to detect multi-scale features and maximize the use of image features flow at different scales, HFDB focuses on adaptively recalibrate channel-wise feature responses to achieve feature distillation, and DRB attempts to reconstruct SR images with different upsampling factors in a single model. It is worth noting that all these modules can run independently. It means that these modules can be selectively plugged into any CNN model to improve model performance. Extensive experiments show that MDCN achieves competitive results in SISR, especially in the reconstruction task with multiple upsampling factors. The code is provided athttps://github.com/MIVRC/MDCN-PyTorch.
Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2021 Luminance-Aware Pyramid Network for Low-Light Image Enhancement
abstract
Low-light image enhancement based on deep convolutional neural networks (CNNs) has revealed prominent performance in recent years. However, it is still a challenging task since the underexposed regions and details are always imperceptible. Moreover, deep learning models are always accompanied by complex structures and enormous computational burden, which hinders their deployment on mobile devices. To remedy these issues, in this paper, we present a lightweight and efficient Luminance-aware Pyramid Network (LPNet) to reconstruct normal-light images in a coarse-to-fine strategy. The architecture is comprised of two coarse feature extraction branches and a luminance-aware refinement branch with an auxiliary subnet learning the luminance map of the input and target images. Besides, we propose a multi-scale contrast feature block (MSCFB) that involves channel split, channel shuffle strategies, and contrast attention mechanism. MSCFB is the essential component of our network, which achieves an excellent balance between image quality and model size. In this way, our method can not only brighten up low-light images with rich details and high contrast but also significantly ameliorate the execution speed. Extensive experiments demonstrate that our LPNet outperforms state-of-the-art methods both qualitatively and quantitatively.
Juncheng Li 0003, Faming Fang, Fang Li 0004, Guixu Zhang
IEEE Trans. Multim.2
2021 Multilevel Edge Features Guided Network for Image Denoising
abstract
Image denoising is a challenging inverse problem due to complex scenes and information loss. Recently, various methods have been considered to solve this problem by building a well-designed convolutional neural network (CNN) or introducing some hand-designed image priors. Different from previous works, we investigate a new framework for image denoising, which integrates edge detection, edge guidance, and image denoising into an end-to-end CNN model. To achieve this goal, we propose a multilevel edge features guided network (MLEFGN). First, we build an edge reconstruction network (Edge-Net) to directly predict clear edges from the noisy image. Then, the Edge-Net is embedded as part of the model to provide edge priors, and a dual-path network is applied to extract the image and edge features, respectively. Finally, we introduce a multilevel edge features guidance mechanism for image denoising. To the best of our knowledge, the Edge-Net is the first CNN model specially designed to reconstruct image edges from the noisy image, which shows good accuracy and robustness on natural images. Extensive experiments clearly illustrate that our MLEFGN achieves favorable performance against other methods and plenty of ablation studies demonstrate the effectiveness of our proposed Edge-Net and MLEFGN. The code is available at https://github.com/MIVRC/MLEFGN-PyTorch.
Faming Fang, Juncheng Li 0003, Yiting Yuan, Tieyong Zeng, Guixu Zhang
IEEE Trans. Neural Networks Learn. Syst.2
2020 Soft-Edge Assisted Network for Single Image Super-Resolution
abstract
The task of single image super-resolution (SISR) is a highly ill-posed inverse problem since reconstructing the highfrequency details from a low-resolution image is challenging. Most previous CNN-based super-resolution (SR) methods tend to directly learn the mapping from the low-resolution image to the high-resolution image through some complex convolutional neural networks. However, the method of blindly increasing the depth of the network is not the best choice because the performance improvement of such methods is marginal but the computational cost is huge. A more efficient method is to integrate the image prior knowledge into the model to assist the image reconstruction. Indeed, the soft-edge has been widely applied in many computer vision tasks as the role of an important image feature. In this paper, we propose a Soft-edge assisted Network (SeaNet) to reconstruct the high-quality SR image with the help of image soft-edge. The proposed SeaNet consists of three sub-nets: a rough image reconstruction network (RIRN), a soft-edge reconstruction network (Edge-Net), and an image refinement network (IRN). The complete reconstruction process consists of two stages. In Stage-I, the rough SR feature maps and the SR soft-edge are reconstructed by the RIRN and Edge-Net, respectively. In Stage-II, the outputs of the previous stages are fused and then feed to the IRN for high-quality SR image reconstruction. Extensive experiments show that our SeaNet converges rapidly and achieves excellent performance under the assistance of image soft-edge. The code is available at https://gitlab.com/junchenglee/seanet-pytorch.
Faming Fang, Juncheng Li 0003, Tieyong Zeng
IEEE Trans. Image Process.2
2019 Deep residual refining based pseudo-multi-frame network for effective single image super-resolution
abstract
Single image super‐resolution (SISR) has gained great attraction and progress in recent years. Since the SISR is an ill‐posed inverse problem, most researchers are concentrated on making efforts to learn effective and reasonable mapping functions from low‐resolution observation to its potential high‐resolution (HR) counterpart. In this study, the authors have proposed a deep residual refining based pseudo‐multi‐frame network for efficient SISR. A channel‐wise attention mechanism is employed for residual refinement. It can ease residual learning process through explicitly modelling non‐linear dependencies between channels by using global information embedding. Multiple potential HRs from different deconvolutional layers are further artificially learned, and then adaptively fused into final desired HR image. The authors call this strategy as pseudo‐multi‐frame SR. It could make full use of available redundant information possessed in hierarchical layers. They have evaluated the proposed network on several popular benchmark datasets. The experimental results have shown that the two highlights proposed can consistently boost final performance. The proposed network can outperform most of the state‐of‐the‐art methods with acceptable less parameters.
Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Bo Liu 0006, Jihua Ye, Mingwen Wang 0001
IET Image Process.3
2018 Progressive Feature Fusion Network for Realistic Image Dehazing
Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Mingwen Wang 0001
ACCV (1)3
2018 Multi-scale Residual Network for Image Super-Resolution
Juncheng Li 0003, Faming Fang, Kangfu Mei, Guixu Zhang
ECCV (8)1
2018 An Effective Single-Image Super-Resolution Model Using Squeeze-and-Excitation Networks
Kangfu Mei, Aiwen Jiang, Juncheng Li 0003, Jihua Ye, Mingwen Wang 0001
ICONIP (6)3