Yanye Lu

dblp:173/2256 · DBLP profile ↗
← Back
34ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0002-3063-8051ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021
YearPublicationVenuePosition
2026 MemTTA: Cluster-guided continual test-time adaptation for cross-domain segmentation
Chuqiao Yang, Chunlin Li 0003, Yixuan Yuan, Hanbo Tan, Xinlei Ma, Junhao Yan, Qingyuan He, Zhaoheng Xie, Hongbin Han, Yanye Lu, Wanyi Fu
Expert Syst. Appl.12
2026 Improve retinal artery/vein classification via channel coupling
Shuang Zeng, Chee Hong Lee, Boxu Xie, Ourui Fu, Hangzhou He, Lei Zhu 0012, Yanye Lu, Fangxiao Cheng
Expert Syst. Appl.8
2026 SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-Training
abstract
Medical image segmentation is a critical yet challenging task, primarily due to the difficulty of obtaining extensive datasets of high-quality, expert-annotated images. Contrastive learning presents a potential but still problematic solution to this issue. Because most existing methods focus on extracting instance-level or pixel-to-pixel representation, which ignores the characteristics between intra-image similar pixel groups. Moreover, when considering contrastive pairs generation, most SOTA methods mainly rely on manually setting thresholds, which requires a large number of gradient experiments and lacks efficiency and generalization. To address these issues, we propose a novel contrastive learning approach named SuperCL for medical image segmentation pre-training. Specifically, our SuperCL exploits the structural prior and pixel correlation of images by introducing two novel contrastive pairs generation strategies: Intra-image Local Contrastive Pairs (ILCP) Generation and Inter-image Global Contrastive Pairs (IGCP) Generation. Considering superpixel cluster aligns well with the concept of contrastive pairs generation, we utilize the superpixel map to generate pseudo masks for both ILCP and IGCP to guide supervised contrastive learning. Moreover, we also propose two modules named Average SuperPixel Feature Map Generation (ASP) and Connected Components Label Generation (CCL) to better exploit the prior structural information for IGCP. Finally, experiments on 8 medical image datasets indicate our SuperCL outperforms existing 12 methods. i.e. Our SuperCL achieves a superior performance with more precise predictions from visualization figures and 3.15%, 5.44%, 7.89% DSC higher than the previous best results on MMWHS, CHAOS, Spleen with 10% annotations. Our code is released at https://github.com/stevezs315/SuperCL.
Shuang Zeng, Lei Zhu 0012, Hangzhou He, Yanye Lu
IEEE Trans. Image Process.5
2025 V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer
abstract
Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach.
Hangzhou He, Lei Zhu 0012, Shuang Zeng, Yanye Lu
AAAI6
2025 Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
abstract
Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0x efficiency on ADE20K, +14.3% mIoU and 5.2x efficiency on VOC2012, and +9.1% mIoU and 6.6x efficiency on CityScapes.
Zhenxin Lei, Man Yao, Xinhao Luo, Yanye Lu, Bo Xu 0002, Guoqi Li 0002
AAAI5
2025 Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
Zhengjian Yao, Lujia Jin, Hangzhou He, Yanye Lu
ICCV5
2025 Auto-Regressively Generating Multi-View Consistent Images
Jialun Liu, Yanye Lu
ICCV6
2025 Universal Image Restoration Pre-training via Degradation Classification
abstract
This paper proposes the Degradation Classification Pre-Training (DCPT), which enables models to learn how to classify the degradation type of input images for universal image restoration pre-training. Unlike the existing self-supervised pre-training methods, DCPT utilizes the degradation type of the input image as an extremely weak supervision, which can be effortlessly obtained, even intrinsic in all image restoration datasets. DCPT comprises two primary stages. Initially, image features are extracted from the encoder. Subsequently, a lightweight decoder, such as ResNet18, is leveraged to classify the degradation type of the input image solely based on the features extracted in the first stage, without utilizing the input image. The encoder is pre-trained with a straightforward yet potent DCPT, which is used to address universal image restoration and achieve outstanding performance. Following DCPT, both convolutional neural networks (CNNs) and transformers demonstrate performance improvements, with gains of up to 2.55 dB in the 10D all-in-one restoration task and 6.53 dB in the mixed degradation scenarios. Moreover, previous self-supervised pretraining methods, such as masked image modeling, discard the decoder after pre-training, while our DCPT utilizes the pre-trained parameters more effectively. This superiority arises from the degradation classifier acquired during DCPT, which facilitates transfer learning between models of identical architecture trained on diverse degradation types. Source code and models are available at \url{https://github.com/MILab-PKU/dcpt}.
Lujia Jin, Zhengjian Yao, Yanye Lu
ICLR4
2025 Training-Free Test-Time Improvement for Explainable Medical Image Classification
Hangzhou He, Jiachen Tang, Lei Zhu 0012, Yanye Lu
MICCAI (14)5
2025 Decoupled dual-granularity rebalanced pyramid network for drug-target interaction prediction
abstract
Drug-target interaction (DTI) prediction is a critical task in drug discovery. However, the differences in size between drugs and proteins present significant challenges in accurately predicting binding sites. Additionally, the issue of modality imbalance, which arises from modality learning biases, undermines the contribution of multimodal representations to DTI prediction. To address these challenges, we propose DDGR-DTI, which is based on an innovative Decoupled Dual-Granularity Framework and a Rebalanced pyramid network (RPN). This framework divides the DTI task into two levels of granularity. The macro level, which decomposes it into subtasks based on modality, and the micro level further decomposes the representation within each subtask. Furthermore, the dual-stream attention module is utilized to perform fine-grained substructure-level interactions within each subtask, thereby enabling accurate identification of binding sites. Simultaneously, we employ an RPN, which effectively alleviates the bias towards the dominant modality in multimodal fusion through a hierarchical aggregation mechanism, emphasizing the synergistic advantages brought by modality balance. Benchmark results demonstrate that DDGR-DTI outperforms existing state-of-the-art models in both prediction performance and generalization ability. Availability: The source code and dataset can be found at https://github.com/ZZUzy/DDGR-DTI.
Zhiyuan Dong, Chaoyang Han, Ya-Juan Gao, Wanyi Fu, Yanye Lu
Briefings Bioinform.9
2025 ECS-Net: Extracellular space segmentation with contrastive and shape-aware loss by using cryo-electron microscopy imaging
Chuqiao Yang, Jiayi Xie, Xinrui Huang, Hanbo Tan, Qirun Li, Zeqing Tang, Xinlei Ma, Jiabin Lu, Qingyuan He, Wanyi Fu, Yixing Huang, Junhao Yan, Zhaoheng Xie, Yao Sui, Yanye Lu, Hongbin Han
Expert Syst. Appl.16
2025 Generative learning-based lightweight MRI brain tumor segmentation with missing modalities
Hangzhou He, Lei Zhu 0012, Zhaoheng Xie, Yanye Lu, Fangxiao Cheng
Expert Syst. Appl.6
2025 Novel extraction of discriminative fine-grained feature to improve retinal vessel segmentation
Shuang Zeng, Chee Hong Lee, Micky C. Nnamdi, Wenqi Shi 0002, J. Ben Tamo, Hangzhou He, May D. Wang, Lei Zhu 0012, Yanye Lu, Qiushi Ren
Image Vis. Comput.11
2025 Points-Supervised Fundus Vessel Segmentation via Shape Priors and Contrastive Learning
abstract
The performance of fully supervised methods for fundus vessel segmentation highly relies on a large number of full labels which are laborious and time-consuming to obtain. Although weak annotations relax the requirement for pixel-wise labeling, they pose challenges in learning comprehensive information about the target. Some methods use pseudo labels generated from network predictions for extra supervision, but false positive predictions in these labels may harm training. In this paper, to tackle this problem and to balance the annotation cost and supervision information, we introduce point annotations to fundus vessel segmentation and propose a novel method, called Points-based Vessel segmentation Network (PVN), to enhance the segmentation accuracy. In PVN, to avoid noise in pseudo labels, by combining proposed Point Activation Maps, shape priors of vessels are learned and used as soft supervision. Additionally, to further leverage the annotated vessel and background points, we design a novel contrastive learning method in a pixels-and-regions-mixed manner, which helps learn discriminative features by distinguishing between pixel and region samples of vessels and background. The performance of PVN is evaluated on laser speckle contrast imaging fundus images, 548 nm fundus images, and three public datasets, where PVN outperforms other point-supervised methods. Even with only 1% annotated pixels, PVN still achieves excellent performance. Our method is also flexible and easy to be combined with other frameworks. To the best of our knowledge, we are the first to propose and demonstrate the effectiveness of point annotations for fundus vessel segmentation. Our code is available at: https://github.com/kaiwenli325/PVN.
Hangzhou He, Shuang Zeng, Lei Zhu 0012, Yanye Lu
IEEE Trans. Medical Imaging7
2025 Branches Mutual Promotion for End-to-End Weakly Supervised Semantic Segmentation
abstract
End-to-end weakly supervised semantic segmentation (E2E-WSSS) aims at optimizing a segmentation model in a single-stage training process based on only image annotations. Existing methods adopt an online-trained classification branch to provide pseudo annotations for supervising the segmentation branch. However, this strategy makes the classification branch dominate the whole concurrent training process, hindering these two branches from assisting each other. In our work, we treat these two branches equally by viewing them as diverse ways to generate the segmentation map, and add interactions on both their supervision and operation to achieve mutual promotion. For this purpose, a bidirectional supervision mechanism is elaborated to force the consistency between the outputs of these two branches. Thus, the segmentation branch can also give feedback to the classification branch to enhance the quality of localization seeds. Moreover, our method also designs interaction operations between these two branches to exchange their knowledge to assist each other. Experiments indicate our work outperforms existing end-to-end weakly supervised segmentation methods. Codes are available at https://github.com/zh460045050/BMP-WSSS.
Lei Zhu 0012, Hangzhou He, Shuang Zeng, Yibao Zhang, Qiushi Ren, Yanye Lu
IEEE Trans. Neural Networks Learn. Syst.9
2024 Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class Label
abstract
Scribble-based weakly-supervised semantic segmentation using sparse scribble supervision is gaining traction as it reduces annotation costs when compared to fully annotated alternatives. Existing methods primarily generate pseudo-labels by diffusing labeled pixels to unlabeled ones with local cues for supervision. However, this diffusion process fails to exploit global semantics and class-specific cues, which are important for semantic segmentation. In this study, we propose a class-driven scribble promotion network, which utilizes both scribble annotations and pseudo-labels informed by image-level classes and global semantics for supervision. Directly adopting pseudo-labels might misguide the segmentation model, thus we design a localization rectification module to correct foreground representations in the feature space. To further combine the advantages of both supervisions, we also introduce a distance entropy loss for uncertainty reduction, which adapts per-pixel confidence weights according to the reliable region determined by the scribble and pseudo-label's boundary. Experiments on the ScribbleSup dataset with different qualities of scribble annotations outperform all the previous methods, demonstrating the superiority and robustness of our method. The code is available at https://github.com/Zxl19990529/Class-driven-Scribble-Promotion-Network.
Lei Zhu 0012, Hangzhou He, Lujia Jin, Yanye Lu
AAAI5
2024 Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
abstract
In this work, we investigate the potential of a large language model (LLM) to directly comprehend visual signals without the necessity of fine-tuning on multimodal datasets. The foundational concept of our method views an image as a linguistic entity, and translates it to a set of discrete words derived from the LLM's vocabulary. To achieve this, we present the Vision-to-Language Tokenizer; abbreviated as V2T Tokenizer, which transforms an image into a “foreign language” with the combined aid of an encoder-decoder, the LLM vocabulary, and a CLIP model. With this innovative image encoding, the LLM gains the ability not only for visual comprehension but also for image denoising and restoration in an auto-regressive fashion-crucially, without any fine-tuning. We undertake rig-orous experiments to validate our method, encompassing understanding tasks like image recognition, image captioning, and visual question answering, as well as image denoising tasks like inpainting, outpainting, deblurring, and shift restoration. Code and models are available at https://github.com/zh460045050/V2l-Tokenizer.
Lei Zhu 0012, Fangyun Wei, Yanye Lu
CVPR3
2024 Low-Rank Mixture-of-Experts for Continual Medical Image Segmentation
Lei Zhu 0012, Hangzhou He, Shuang Zeng, Qiushi Ren, Yanye Lu
MICCAI (8)7
2024 Scaling the Codebook Size of VQ-GAN to 100, 000 with a Utilization Rate of 99%
abstract
In the realm of image quantization exemplified by VQGAN, the process encodes images into discrete tokens drawn from a codebook with a predefined size. Recent advancements, particularly with LLAMA 3, reveal that enlarging the codebook significantly enhances model performance. However, VQGAN and its derivatives, such as VQGAN-FC (Factorized Codes) and VQGAN-EMA, continue to grapple with challenges related to expanding the codebook size and enhancing codebook utilization. For instance, VQGAN-FC is restricted to learning a codebook with a maximum size of 16,384, maintaining a typically low utilization rate of less than 12% on ImageNet. In this work, we propose a novel image quantization model named VQGAN-LC (Large Codebook), which extends the codebook size to 100,000, achieving an utilization rate exceeding 99%. Unlike previous methods that optimize each codebook entry, our approach begins with a codebook initialized with 100,000 features extracted by a pre-trained vision encoder. Optimization then focuses on training a projector that aligns the entire codebook with the feature distributions of the encoder in VQGAN-LC. We demonstrate the superior performance of our model over its counterparts across a variety of tasks, including image reconstruction, image classification, auto-regressive image generation using GPT, and image creation with diffusion- and flow-based generative models.
Lei Zhu 0012, Fangyun Wei, Yanye Lu
NeurIPS3
2024 One-Pot Multi-frame Denoising
Lujia Jin, Shi Zhao, Lei Zhu 0012, Qiushi Ren, Yanye Lu
Int. J. Comput. Vis.7
2024 Boosting Weakly Supervised Object Localization and Segmentation With Domain Adaption
abstract
Weakly supervised object localization (WSOL), adopting only image-level annotations to learn the pixel-level localization model, can release human resources in the annotation process. Most one-stage WSOL methods learn the localization model with multi-instance learning, making them only activate discriminative object parts rather than the whole object. In our work, we attribute this problem to the domain shift between the training and test process of WSOL and provide a novel perspective that views WSOL as a domain adaption (DA) task. Under this perspective, a DA-WSOL pipeline is elaborated to better assist WSOL with DA approaches by considering the specificities for the adaption of WSOL. Our DA-WSOL pipeline can discern the source-related and the Universum samples from other target samples based on a proposed target sampling strategy and then utilize them to solve the sample unbalancing and label unmatching between the source and target domain of WSOL. Experiments show that our pipeline outperforms SOTA methods on three WSOL benchmarks and can improve the performance of downstream weakly supervised semantic segmentation tasks.
Lei Zhu 0012, Qi She, Qiushi Ren, Yanye Lu
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 PCNet: Prior Category Network for CT Universal Segmentation Model
abstract
Accurate segmentation of anatomical structures in Computed Tomography (CT) images is crucial for clinical diagnosis, treatment planning, and disease monitoring. The present deep learning segmentation methods are hindered by factors such as data scale and model size. Inspired by how doctors identify tissues, we propose a novel approach, the Prior Category Network (PCNet), that boosts segmentation performance by leveraging prior knowledge between different categories of anatomical structures. Our PCNet comprises three key components: prior category prompt (PCP), hierarchy category system (HCS), and hierarchy category loss (HCL). PCP utilizes Contrastive Language-Image Pretraining (CLIP), along with attention modules, to systematically define the relationships between anatomical categories as identified by clinicians. HCS guides the segmentation model in distinguishing between specific organs, anatomical structures, and functional systems through hierarchical relationships. HCL serves as a consistency constraint, fortifying the directional guidance provided by HCS to enhance the segmentation model's accuracy and robustness. We conducted extensive experiments to validate the effectiveness of our approach, and the results indicate that PCNet can generate a high-performance, universal model for CT segmentation. The PCNet framework also demonstrates a significant transferability on multiple downstream tasks. The ablation experiments show that the methodology employed in constructing the HCS is of critical importance. The prompt and HCS can be accessed at https://github.com/PKU-MIPET/PCNet.
Ya-Juan Gao, Lei Zhu 0012, Wenrui Shao, Yanye Lu, Hongbin Han, Zhaoheng Xie
IEEE Trans. Medical Imaging5
2024 3-D Convolutional Neural Networks for RGB-D Salient Object Detection and Beyond
abstract
RGB-depth (RGB-D) salient object detection (SOD) recently has attracted increasing research interest, and many deep learning methods based on encoder-decoder architectures have emerged. However, most existing RGB-D SOD models conduct explicit and controllable cross-modal feature fusion either in the single encoder or decoder stage, which hardly guarantees sufficient cross-modal fusion ability. To this end, we make the first attempt in addressing RGB-D SOD through 3-D convolutional neural networks. The proposed model, named RD3D, aims at prefusion in the encoder stage and in-depth fusion in the decoder stage to effectively promote the full integration of RGB and depth streams. Specifically, RD3D first conducts prefusion across RGB and depth modalities through a 3-D encoder obtained by inflating 2-D ResNet and later provides in-depth feature fusion by designing a 3-D decoder equipped with rich back-projection paths (RBPPs) for leveraging the extensive aggregation ability of 3-D convolutions. Toward an improved model RD3D+, we propose to disentangle the conventional 3-D convolution into successive spatial and temporal convolutions and, meanwhile, discard unnecessary zero padding. This eventually results in a 2-D convolutional equivalence that facilitates optimization and reduces parameters and computation costs. Thanks to such a progressive-fusion strategy involving both the encoder and the decoder, effective and thorough interactions between the two modalities can be exploited and boost detection accuracy. As an additional boost, we also introduce channel-modality attention and its variant after each path of RBPP to attend to important features. Extensive experiments on seven widely used benchmark datasets demonstrate that RD3D and RD3D+ perform favorably against 14 state-of-the-art RGB-D SOD approaches in terms of five key evaluation metrics. Our code will be made publicly available at https://github.com/PPOLYpubki/RD3D.
Yanye Lu, Keren Fu, Qijun Zhao
IEEE Trans. Neural Networks Learn. Syst.3
2023 Discriminative ensemble meta-learning with co-regularization for rare fundus diseases diagnosis
Mengdi Gao, Hongyang Jiang 0001, Lei Zhu 0012, Mufeng Geng, Qiushi Ren, Yanye Lu
Medical Image Anal.7
2023 Background-Aware Classification Activation Map for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) relaxes the requirement of dense annotations for object localization by using image-level annotation to supervise the learning process. However, most WSOL methods only focus on forcing the object classifier to produce high activation score on object parts without considering the influence of background locations, causing excessive background activations and ill-pose background score searching. Based on this point, our work proposes a novel mechanism called the background-aware classification activation map (B-CAM) to add background awareness for WSOL training. Besides aggregating an object image-level feature for supervision, our B-CAM produces an additional background image-level feature to represent the pure-background sample. This additional feature can provide background cues for the object classifier to suppress the background activations on object localization maps. Moreover, our B-CAM also trained a background classifier with image-level annotation to produce adaptive background scores when determining the binary localization mask. Experiments indicate the effectiveness of the proposed B-CAM on four different types of WSOL benchmarks, including CUB-200, ILSVRC, OpenImages, and VOC2012 datasets.
Lei Zhu 0012, Qi She, Xiangxi Meng 0001, Mufeng Geng, Lujia Jin, Yibao Zhang, Qiushi Ren, Yanye Lu
IEEE Trans. Pattern Anal. Mach. Intell.9
2022 One-Pot Multi-Frame Denoising
Lujia Jin, Shi Zhao, Lei Zhu 0012, Yanye Lu
BMVC5
2022 Weakly Supervised Object Localization as Domain Adaption
abstract
Weakly supervised object localization (WSOL) focuses on localizing objects only with the supervision of image-level classification masks. Most previous WSOL methods follow the classification activation map (CAM) that localizes objects based on the classification structure with the multi-instance learning (MIL) mechanism. However, the MIL mechanism makes CAM only activate discriminative object parts rather than the whole object, weakening its performance for localizing objects. To avoid this problem, this work provides a novel perspective that models WSOL as a domain adaption (DA) task, where the score estimator trained on the source/image domain is tested on the target/pixel domain to locate objects. Under this perspective, a DA-WSOL pipeline is designed to better engage DA approaches into WSOL to enhance localization performance. It utilizes a proposed target sampling strategy to select different types of target samples. Based on these types of target samples, domain adaption localization (DAL) loss is elaborated. It aligns the feature distribution between the two domains by DA and makes the estimator perceive target domain cues by Universum regularization. Experiments show that our pipeline outperforms SOTA methods on multi benchmarks. Code are released at https://github.com/zh460045050/DA-WSOL_CVPR2022.
Lei Zhu 0012, Qi She, Yunfei You, Boyu Wang 0004, Yanye Lu
CVPR6
2022 Bagging Regional Classification Activation Maps for Weakly Supervised Object Localization
Lei Zhu 0012, Lujia Jin, Yunfei You, Yanye Lu
ECCV (10)5
2022 Content-Noise Complementary Learning for Medical Image Denoising
abstract
Medical imaging denoising faces great challenges, yet is in great demand. With its distinctive characteristics, medical imaging denoising in the image domain requires innovative deep learning strategies. In this study, we propose a simple yet effective strategy, the content-noise complementary learning (CNCL) strategy, in which two deep learning predictors are used to learn the respective content and noise of the image dataset complementarily. A medical image denoising pipeline based on the CNCL strategy is presented, and is implemented as a generative adversarial network, where various representative networks (including U-Net, DnCNN, and SRDenseNet) are investigated as the predictors. The performance of these implemented models has been validated on medical imaging datasets including CT, MR, and PET. The results show that this strategy outperforms state-of-the-art denoising algorithms in terms of visual quality and quantitative metrics, and the strategy demonstrates a robust generalization capability. These findings validate that this simple yet effective strategy demonstrates promising potential for medical image denoising tasks, which could exert a clinical impact in the future. Code is available at: https://github.com/gengmufeng/CNCL-denoising.
Mufeng Geng, Xiangxi Meng 0001, Jiangyuan Yu, Lei Zhu 0012, Lujia Jin, Bin Qiu, Hanjing Kong, Jianmin Yuan, Hongming Shan, Hongbin Han, Qiushi Ren, Yanye Lu
IEEE Trans. Medical Imaging16
2022 Triplet Cross-Fusion Learning for Unpaired Image Denoising in Optical Coherence Tomography
abstract
Optical coherence tomography (OCT) is a widely-used modality in clinical imaging, which suffers from the speckle noise inevitably. Deep learning has proven its superior capability in OCT image denoising, while the difficulty of acquiring a large number of well-registered OCT image pairs limits the developments of paired learning methods. To solve this problem, some unpaired learning methods have been proposed, where the denoising networks can be trained with unpaired OCT data. However, majority of them are modified from the cycleGAN framework. These cycleGAN-based methods train at least two generators and two discriminators, while only one generator is needed for the inference. The dual-generator and dual-discriminator structures of cycleGAN-based methods demand a large amount of computing resource, which may be redundant for OCT denoising tasks. In this work, we propose a novel triplet cross-fusion learning (TCFL) strategy for unpaired OCT image denoising. The model complexity of our strategy is much lower than those of the cycleGAN-based methods. During training, the clean components and the noise components from the triplet of three unpaired images are cross-fused, helping the network extract more speckle noise information to improve the denoising accuracy. Furthermore, the TCFL-based network which is trained with triplets can deal with limited training data scenarios. The results demonstrate that the TCFL strategy outperforms state-of-the-art unpaired methods both qualitatively and quantitatively, and even achieves denoising performance comparable with paired methods. Code is available at: https://github.com/gengmufeng/TCFL-OCT.
Mufeng Geng, Xiangxi Meng 0001, Lei Zhu 0012, Mengdi Gao, Zhiyu Huang, Bin Qiu, Yibao Zhang, Qiushi Ren, Yanye Lu
IEEE Trans. Medical Imaging11
2021 Learning the Superpixel in a Non-Iterative and Lifelong Manner
abstract
Superpixel is generated by automatically clustering pixels in an image into hundreds of compact partitions, which is widely used to perceive the object contours for its excel-lent contour adherence. Although some works use the Convolution Neural Network (CNN) to generate high-quality superpixel, we challenge the design principles of these net-works, specifically for their dependence on manual labels and excess computation resources, which limits their flexibility compared with the traditional unsupervised segmentation methods. We target at redefining the CNN-based superpixel segmentation as a lifelong clustering task and pro-pose an unsupervised CNN-based method called LNS-Net. The LNS-Net can learn superpixel in a non-iterative and lifelong manner without any manual labels. Specifically, a lightweight feature embedder is proposed for LNS-Net to efficiently generate the cluster-friendly features. With those features, seed nodes can be automatically assigned to cluster pixels in a non-iterative way. Additionally, our LNS-Net can adapt the sequentially lifelong learning by rescaling the gradient of weight based on both channel and spatial context to avoid overfitting. Experiments show that the proposed LNS-Net achieves significantly better performance on three benchmarks with nearly ten times lower complexity compared with other state-of-the-art methods.
Lei Zhu 0012, Qi She, Yanye Lu, Zhilin Lu 0002, Jie Hu 0019
CVPR4
2021 Unifying Nonlocal Blocks for Neural Networks
abstract
The nonlocal-based blocks are designed for capturing long-range spatial-temporal dependencies in computer vision tasks. Although having shown excellent performance, they still lack the mechanism to encode the rich, structured information among elements in an image or video. In this paper, to theoretically analyze the property of these nonlocal-based blocks, we provide a new perspective to interpret them, where we view them as a set of graph filters generated on a fully-connected graph. Specifically, when choosing the Chebyshev graph filter, a unified formulation can be derived for explaining and analyzing the existing nonlocal-based blocks (e.g., nonlocal block, nonlocal stage, double attention block). Furthermore, by concerning the property of spectral, we propose an efficient and robust spectral nonlocal block, which can be more robust and flexible to catch long-range dependencies when inserted into deep neural networks than the existing nonlocal blocks. Experimental results demonstrate the clear-cut improvements and practical applicabilities of our method on image classification, action recognition, semantic segmentation, and person re-identification tasks. Code are available at https://github.com/zh460045050/SNL_ICCV2021.
Lei Zhu 0012, Qi She, Yanye Lu, Xuejing Kang, Jie Hu 0019, Changhu Wang
ICCV4
2021 PMS-GAN: Parallel Multi-Stream Generative Adversarial Network for Multi-Material Decomposition in Spectral Computed Tomography
abstract
Spectral computed tomography is able to provide quantitative information on the scanned object and enables material decomposition. Traditional projection-based material decomposition methods suffer from the nonlinearity of the imaging system, which limits the decomposition accuracy. Inspired by the generative adversarial network, we proposed a novel parallel multi-stream generative adversarial network (PMS-GAN) to perform projection-based multi-material decomposition in spectral computed tomography. By designing the differential map and incorporating the adversarial network into loss function, the decomposition accuracy was significantly improved with robust performance. The proposed network was quantitatively evaluated by both simulation and experimental study. The results show that PMS-GAN outperformed the reference methods with certain robustness. Compared with Pix2pix-GAN, PMS-GAN increased the structural similarity index by 172% on the contrast agent Ultravist370, 11% on bones, and 71% on bone marrow, respectively, in a simulated test scenario. In an experimental test scenario, 9% and 38% improvements of the structural similarity index on the biopsy needle and on a torso phantom were observed, respectively. The proposed network demonstrates its capability of multi-material decomposition and has certain potential toward clinical applications.
Mufeng Geng, Zifeng Tian, Yunfei You, Ximeng Feng, Yan Xia 0002, Qiushi Ren, Xiangxi Meng 0001, Andreas K. Maier, Yanye Lu
IEEE Trans. Medical Imaging11
2021 Weakly Supervised Deep Learning-Based Optical Coherence Tomography Angiography
abstract
Optical coherence tomography angiography (OCTA) is a promising imaging modality for microvasculature studies. Deep learning networks have been widely applied in the field of OCTA reconstruction, benefiting from its powerful mapping capability among images. However, these existing deep learning-based methods depend on high-quality labels, which are hard to acquire considering imaging hardware limitations and practical data acquisition conditions. In this article, we proposed an unprecedented weakly supervised deep learning-based pipeline for OCTA reconstruction task, in the absence of high-quality training labels. The proposed pipeline was investigated on an in vivo animal dataset and a human eye dataset by a cross-validation strategy. Compared with supervised learning approaches, the proposed approach demonstrated similar or even better performance in the OCTA reconstruction task. These investigations indicate that the proposed weakly supervised learning strategy is well capable of performing OCTA reconstruction, and has a certain potential towards clinical applications.
Zhiyu Huang, Bin Qiu, Xiangxi Meng 0001, Yunfei You, Mufeng Geng, Gangjun Liu, Chuanqing Zhou, Andreas K. Maier, Qiushi Ren, Yanye Lu
IEEE Trans. Medical Imaging13