EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhu 0012
dblp:99/549-12
· DBLP profile ↗
32ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0003-0506-4268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improve retinal artery/vein classification via channel coupling
Shuang Zeng, Chee Hong Lee, Boxu Xie, Ourui Fu, Hangzhou He, Lei Zhu 0012, Yanye Lu, Fangxiao Cheng |
Expert Syst. Appl. | 7 |
| 2026 | SuperCL: Superpixel Guided Contrastive Learning for Medical Image Segmentation Pre-TrainingabstractMedical image segmentation is a critical yet challenging task, primarily due to the difficulty of obtaining extensive datasets of high-quality, expert-annotated images. Contrastive learning presents a potential but still problematic solution to this issue. Because most existing methods focus on extracting instance-level or pixel-to-pixel representation, which ignores the characteristics between intra-image similar pixel groups. Moreover, when considering contrastive pairs generation, most SOTA methods mainly rely on manually setting thresholds, which requires a large number of gradient experiments and lacks efficiency and generalization. To address these issues, we propose a novel contrastive learning approach named SuperCL for medical image segmentation pre-training. Specifically, our SuperCL exploits the structural prior and pixel correlation of images by introducing two novel contrastive pairs generation strategies: Intra-image Local Contrastive Pairs (ILCP) Generation and Inter-image Global Contrastive Pairs (IGCP) Generation. Considering superpixel cluster aligns well with the concept of contrastive pairs generation, we utilize the superpixel map to generate pseudo masks for both ILCP and IGCP to guide supervised contrastive learning. Moreover, we also propose two modules named Average SuperPixel Feature Map Generation (ASP) and Connected Components Label Generation (CCL) to better exploit the prior structural information for IGCP. Finally, experiments on 8 medical image datasets indicate our SuperCL outperforms existing 12 methods. i.e. Our SuperCL achieves a superior performance with more precise predictions from visualization figures and 3.15%, 5.44%, 7.89% DSC higher than the previous best results on MMWHS, CHAOS, Spleen with 10% annotations. Our code is released at https://github.com/stevezs315/SuperCL. Shuang Zeng, Lei Zhu 0012, Hangzhou He, Yanye Lu |
IEEE Trans. Image Process. | 2 |
| 2025 | V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept TokenizerabstractConcept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. Recent approaches have leveraged the knowledge of large language models to construct concept bottlenecks, with multimodal models like CLIP subsequently mapping image features into the concept feature space for classification. Despite this, the concepts produced by language models can be verbose and may introduce non-visual attributes, which hurts accuracy and interpretability. In this study, we investigate to avoid these issues by constructing CBMs directly from multimodal models. To this end, we adopt common words as base concept vocabulary and leverage auxiliary unlabeled images to construct a Vision-to-Concept (V2C) tokenizer that can explicitly quantize images into their most relevant visual concepts, thus creating a vision-oriented concept bottleneck tightly coupled with the multimodal model. This leads to our V2C-CBM which is training efficient and interpretable with high accuracy. Our V2C-CBM has matched or outperformed LLM-supervised CBMs on various visual classification benchmarks, validating the efficacy of our approach. Hangzhou He, Lei Zhu 0012, Shuang Zeng, Yanye Lu |
AAAI | 2 |
| 2025 | Training-Free Test-Time Improvement for Explainable Medical Image Classification
Hangzhou He, Jiachen Tang, Lei Zhu 0012, Yanye Lu |
MICCAI (14) | 3 |
| 2025 | Generative learning-based lightweight MRI brain tumor segmentation with missing modalities
Hangzhou He, Lei Zhu 0012, Zhaoheng Xie, Yanye Lu, Fangxiao Cheng |
Expert Syst. Appl. | 4 |
| 2025 | Novel extraction of discriminative fine-grained feature to improve retinal vessel segmentation
Shuang Zeng, Chee Hong Lee, Micky C. Nnamdi, Wenqi Shi 0002, J. Ben Tamo, Hangzhou He, May D. Wang, Lei Zhu 0012, Yanye Lu, Qiushi Ren |
Image Vis. Comput. | 10 |
| 2025 | Points-Supervised Fundus Vessel Segmentation via Shape Priors and Contrastive LearningabstractThe performance of fully supervised methods for fundus vessel segmentation highly relies on a large number of full labels which are laborious and time-consuming to obtain. Although weak annotations relax the requirement for pixel-wise labeling, they pose challenges in learning comprehensive information about the target. Some methods use pseudo labels generated from network predictions for extra supervision, but false positive predictions in these labels may harm training. In this paper, to tackle this problem and to balance the annotation cost and supervision information, we introduce point annotations to fundus vessel segmentation and propose a novel method, called Points-based Vessel segmentation Network (PVN), to enhance the segmentation accuracy. In PVN, to avoid noise in pseudo labels, by combining proposed Point Activation Maps, shape priors of vessels are learned and used as soft supervision. Additionally, to further leverage the annotated vessel and background points, we design a novel contrastive learning method in a pixels-and-regions-mixed manner, which helps learn discriminative features by distinguishing between pixel and region samples of vessels and background. The performance of PVN is evaluated on laser speckle contrast imaging fundus images, 548 nm fundus images, and three public datasets, where PVN outperforms other point-supervised methods. Even with only 1% annotated pixels, PVN still achieves excellent performance. Our method is also flexible and easy to be combined with other frameworks. To the best of our knowledge, we are the first to propose and demonstrate the effectiveness of point annotations for fundus vessel segmentation. Our code is available at: https://github.com/kaiwenli325/PVN. Hangzhou He, Shuang Zeng, Lei Zhu 0012, Yanye Lu |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Branches Mutual Promotion for End-to-End Weakly Supervised Semantic SegmentationabstractEnd-to-end weakly supervised semantic segmentation (E2E-WSSS) aims at optimizing a segmentation model in a single-stage training process based on only image annotations. Existing methods adopt an online-trained classification branch to provide pseudo annotations for supervising the segmentation branch. However, this strategy makes the classification branch dominate the whole concurrent training process, hindering these two branches from assisting each other. In our work, we treat these two branches equally by viewing them as diverse ways to generate the segmentation map, and add interactions on both their supervision and operation to achieve mutual promotion. For this purpose, a bidirectional supervision mechanism is elaborated to force the consistency between the outputs of these two branches. Thus, the segmentation branch can also give feedback to the classification branch to enhance the quality of localization seeds. Moreover, our method also designs interaction operations between these two branches to exchange their knowledge to assist each other. Experiments indicate our work outperforms existing end-to-end weakly supervised segmentation methods. Codes are available at https://github.com/zh460045050/BMP-WSSS. Lei Zhu 0012, Hangzhou He, Shuang Zeng, Yibao Zhang, Qiushi Ren, Yanye Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class LabelabstractScribble-based weakly-supervised semantic segmentation using sparse scribble supervision is gaining traction as it reduces annotation costs when compared to fully annotated alternatives. Existing methods primarily generate pseudo-labels by diffusing labeled pixels to unlabeled ones with local cues for supervision. However, this diffusion process fails to exploit global semantics and class-specific cues, which are important for semantic segmentation. In this study, we propose a class-driven scribble promotion network, which utilizes both scribble annotations and pseudo-labels informed by image-level classes and global semantics for supervision. Directly adopting pseudo-labels might misguide the segmentation model, thus we design a localization rectification module to correct foreground representations in the feature space. To further combine the advantages of both supervisions, we also introduce a distance entropy loss for uncertainty reduction, which adapts per-pixel confidence weights according to the reliable region determined by the scribble and pseudo-label's boundary. Experiments on the ScribbleSup dataset with different qualities of scribble annotations outperform all the previous methods, demonstrating the superiority and robustness of our method. The code is available at https://github.com/Zxl19990529/Class-driven-Scribble-Promotion-Network. Lei Zhu 0012, Hangzhou He, Lujia Jin, Yanye Lu |
AAAI | 2 |
| 2024 | Beyond Text: Frozen Large Language Models in Visual Signal ComprehensionabstractIn this work, we investigate the potential of a large language model (LLM) to directly comprehend visual signals without the necessity of fine-tuning on multimodal datasets. The foundational concept of our method views an image as a linguistic entity, and translates it to a set of discrete words derived from the LLM's vocabulary. To achieve this, we present the Vision-to-Language Tokenizer; abbreviated as V2T Tokenizer, which transforms an image into a “foreign language” with the combined aid of an encoder-decoder, the LLM vocabulary, and a CLIP model. With this innovative image encoding, the LLM gains the ability not only for visual comprehension but also for image denoising and restoration in an auto-regressive fashion-crucially, without any fine-tuning. We undertake rig-orous experiments to validate our method, encompassing understanding tasks like image recognition, image captioning, and visual question answering, as well as image denoising tasks like inpainting, outpainting, deblurring, and shift restoration. Code and models are available at https://github.com/zh460045050/V2l-Tokenizer. Lei Zhu 0012, Fangyun Wei, Yanye Lu |
CVPR | 1 |
| 2024 | Low-Rank Mixture-of-Experts for Continual Medical Image Segmentation
Lei Zhu 0012, Hangzhou He, Shuang Zeng, Qiushi Ren, Yanye Lu |
MICCAI (8) | 2 |
| 2024 | Scaling the Codebook Size of VQ-GAN to 100, 000 with a Utilization Rate of 99%abstractIn the realm of image quantization exemplified by VQGAN, the process encodes images into discrete tokens drawn from a codebook with a predefined size. Recent advancements, particularly with LLAMA 3, reveal that enlarging the codebook significantly enhances model performance. However, VQGAN and its derivatives, such as VQGAN-FC (Factorized Codes) and VQGAN-EMA, continue to grapple with challenges related to expanding the codebook size and enhancing codebook utilization. For instance, VQGAN-FC is restricted to learning a codebook with a maximum size of 16,384, maintaining a typically low utilization rate of less than 12% on ImageNet. In this work, we propose a novel image quantization model named VQGAN-LC (Large Codebook), which extends the codebook size to 100,000, achieving an utilization rate exceeding 99%. Unlike previous methods that optimize each codebook entry, our approach begins with a codebook initialized with 100,000 features extracted by a pre-trained vision encoder. Optimization then focuses on training a projector that aligns the entire codebook with the feature distributions of the encoder in VQGAN-LC. We demonstrate the superior performance of our model over its counterparts across a variety of tasks, including image reconstruction, image classification, auto-regressive image generation using GPT, and image creation with diffusion- and flow-based generative models. Lei Zhu 0012, Fangyun Wei, Yanye Lu |
NeurIPS | 1 |
| 2024 | One-Pot Multi-frame Denoising
Lujia Jin, Shi Zhao, Lei Zhu 0012, Qiushi Ren, Yanye Lu |
Int. J. Comput. Vis. | 4 |
| 2024 | Boosting Weakly Supervised Object Localization and Segmentation With Domain AdaptionabstractWeakly supervised object localization (WSOL), adopting only image-level annotations to learn the pixel-level localization model, can release human resources in the annotation process. Most one-stage WSOL methods learn the localization model with multi-instance learning, making them only activate discriminative object parts rather than the whole object. In our work, we attribute this problem to the domain shift between the training and test process of WSOL and provide a novel perspective that views WSOL as a domain adaption (DA) task. Under this perspective, a DA-WSOL pipeline is elaborated to better assist WSOL with DA approaches by considering the specificities for the adaption of WSOL. Our DA-WSOL pipeline can discern the source-related and the Universum samples from other target samples based on a proposed target sampling strategy and then utilize them to solve the sample unbalancing and label unmatching between the source and target domain of WSOL. Experiments show that our pipeline outperforms SOTA methods on three WSOL benchmarks and can improve the performance of downstream weakly supervised semantic segmentation tasks. Lei Zhu 0012, Qi She, Qiushi Ren, Yanye Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | PCNet: Prior Category Network for CT Universal Segmentation ModelabstractAccurate segmentation of anatomical structures in Computed Tomography (CT) images is crucial for clinical diagnosis, treatment planning, and disease monitoring. The present deep learning segmentation methods are hindered by factors such as data scale and model size. Inspired by how doctors identify tissues, we propose a novel approach, the Prior Category Network (PCNet), that boosts segmentation performance by leveraging prior knowledge between different categories of anatomical structures. Our PCNet comprises three key components: prior category prompt (PCP), hierarchy category system (HCS), and hierarchy category loss (HCL). PCP utilizes Contrastive Language-Image Pretraining (CLIP), along with attention modules, to systematically define the relationships between anatomical categories as identified by clinicians. HCS guides the segmentation model in distinguishing between specific organs, anatomical structures, and functional systems through hierarchical relationships. HCL serves as a consistency constraint, fortifying the directional guidance provided by HCS to enhance the segmentation model's accuracy and robustness. We conducted extensive experiments to validate the effectiveness of our approach, and the results indicate that PCNet can generate a high-performance, universal model for CT segmentation. The PCNet framework also demonstrates a significant transferability on multiple downstream tasks. The ablation experiments show that the methodology employed in constructing the HCS is of critical importance. The prompt and HCS can be accessed at https://github.com/PKU-MIPET/PCNet. Ya-Juan Gao, Lei Zhu 0012, Wenrui Shao, Yanye Lu, Hongbin Han, Zhaoheng Xie |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Discriminative ensemble meta-learning with co-regularization for rare fundus diseases diagnosis
Mengdi Gao, Hongyang Jiang 0001, Lei Zhu 0012, Mufeng Geng, Qiushi Ren, Yanye Lu |
Medical Image Anal. | 3 |
| 2023 | Background-Aware Classification Activation Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) relaxes the requirement of dense annotations for object localization by using image-level annotation to supervise the learning process. However, most WSOL methods only focus on forcing the object classifier to produce high activation score on object parts without considering the influence of background locations, causing excessive background activations and ill-pose background score searching. Based on this point, our work proposes a novel mechanism called the background-aware classification activation map (B-CAM) to add background awareness for WSOL training. Besides aggregating an object image-level feature for supervision, our B-CAM produces an additional background image-level feature to represent the pure-background sample. This additional feature can provide background cues for the object classifier to suppress the background activations on object localization maps. Moreover, our B-CAM also trained a background classifier with image-level annotation to produce adaptive background scores when determining the binary localization mask. Experiments indicate the effectiveness of the proposed B-CAM on four different types of WSOL benchmarks, including CUB-200, ILSVRC, OpenImages, and VOC2012 datasets. Lei Zhu 0012, Qi She, Xiangxi Meng 0001, Mufeng Geng, Lujia Jin, Yibao Zhang, Qiushi Ren, Yanye Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | One-Pot Multi-Frame Denoising
Lujia Jin, Shi Zhao, Lei Zhu 0012, Yanye Lu |
BMVC | 3 |
| 2022 | Weakly Supervised Object Localization as Domain AdaptionabstractWeakly supervised object localization (WSOL) focuses on localizing objects only with the supervision of image-level classification masks. Most previous WSOL methods follow the classification activation map (CAM) that localizes objects based on the classification structure with the multi-instance learning (MIL) mechanism. However, the MIL mechanism makes CAM only activate discriminative object parts rather than the whole object, weakening its performance for localizing objects. To avoid this problem, this work provides a novel perspective that models WSOL as a domain adaption (DA) task, where the score estimator trained on the source/image domain is tested on the target/pixel domain to locate objects. Under this perspective, a DA-WSOL pipeline is designed to better engage DA approaches into WSOL to enhance localization performance. It utilizes a proposed target sampling strategy to select different types of target samples. Based on these types of target samples, domain adaption localization (DAL) loss is elaborated. It aligns the feature distribution between the two domains by DA and makes the estimator perceive target domain cues by Universum regularization. Experiments show that our pipeline outperforms SOTA methods on multi benchmarks. Code are released at https://github.com/zh460045050/DA-WSOL_CVPR2022. Lei Zhu 0012, Qi She, Yunfei You, Boyu Wang 0004, Yanye Lu |
CVPR | 1 |
| 2022 | Bagging Regional Classification Activation Maps for Weakly Supervised Object Localization
Lei Zhu 0012, Lujia Jin, Yunfei You, Yanye Lu |
ECCV (10) | 1 |
| 2022 | Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
Ge-Peng Ji, Lei Zhu 0012, Mingchen Zhuge, Keren Fu |
Pattern Recognit. | 2 |
| 2022 | Explored Normalized Cut With Random Walk Refining Term for Image SegmentationabstractThe Normalized Cut (NCut) model is a popular graph-based model for image segmentation. But it suffers from the excessive normalization problem and weakens the small object and twig segmentation. In this paper, we propose an Explored Normalized Cut (ENCut) model that establishes a balance graph model by adopting a meaningful-loop and a k-step random walk, which reduces the energy of small salient region, so as to enhance the small object segmentation. To improve the twig segmentation, our ENCut model is further enhanced by a new Random Walk Refining Term (RWRT) that adds local attention to our model with the help of an un-supervising random walk. Finally, a move-making based strategy is developed to efficiently solve the ENCut model with RWRT. Experiments on three standard datasets indicate that our model can achieve state-of-the-art results among the NCut-based segmentation models. Lei Zhu 0012, Xuejing Kang, Lizhu Ye, Anlong Ming |
IEEE Trans. Image Process. | 1 |
| 2022 | Content-Noise Complementary Learning for Medical Image DenoisingabstractMedical imaging denoising faces great challenges, yet is in great demand. With its distinctive characteristics, medical imaging denoising in the image domain requires innovative deep learning strategies. In this study, we propose a simple yet effective strategy, the content-noise complementary learning (CNCL) strategy, in which two deep learning predictors are used to learn the respective content and noise of the image dataset complementarily. A medical image denoising pipeline based on the CNCL strategy is presented, and is implemented as a generative adversarial network, where various representative networks (including U-Net, DnCNN, and SRDenseNet) are investigated as the predictors. The performance of these implemented models has been validated on medical imaging datasets including CT, MR, and PET. The results show that this strategy outperforms state-of-the-art denoising algorithms in terms of visual quality and quantitative metrics, and the strategy demonstrates a robust generalization capability. These findings validate that this simple yet effective strategy demonstrates promising potential for medical image denoising tasks, which could exert a clinical impact in the future. Code is available at: https://github.com/gengmufeng/CNCL-denoising. Mufeng Geng, Xiangxi Meng 0001, Jiangyuan Yu, Lei Zhu 0012, Lujia Jin, Bin Qiu, Hanjing Kong, Jianmin Yuan, Hongming Shan, Hongbin Han, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Triplet Cross-Fusion Learning for Unpaired Image Denoising in Optical Coherence TomographyabstractOptical coherence tomography (OCT) is a widely-used modality in clinical imaging, which suffers from the speckle noise inevitably. Deep learning has proven its superior capability in OCT image denoising, while the difficulty of acquiring a large number of well-registered OCT image pairs limits the developments of paired learning methods. To solve this problem, some unpaired learning methods have been proposed, where the denoising networks can be trained with unpaired OCT data. However, majority of them are modified from the cycleGAN framework. These cycleGAN-based methods train at least two generators and two discriminators, while only one generator is needed for the inference. The dual-generator and dual-discriminator structures of cycleGAN-based methods demand a large amount of computing resource, which may be redundant for OCT denoising tasks. In this work, we propose a novel triplet cross-fusion learning (TCFL) strategy for unpaired OCT image denoising. The model complexity of our strategy is much lower than those of the cycleGAN-based methods. During training, the clean components and the noise components from the triplet of three unpaired images are cross-fused, helping the network extract more speckle noise information to improve the denoising accuracy. Furthermore, the TCFL-based network which is trained with triplets can deal with limited training data scenarios. The results demonstrate that the TCFL strategy outperforms state-of-the-art unpaired methods both qualitatively and quantitatively, and even achieves denoising performance comparable with paired methods. Code is available at: https://github.com/gengmufeng/TCFL-OCT. Mufeng Geng, Xiangxi Meng 0001, Lei Zhu 0012, Mengdi Gao, Zhiyu Huang, Bin Qiu, Yibao Zhang, Qiushi Ren, Yanye Lu |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Involution: Inverting the Inherence of Convolution for Visual RecognitionabstractConvolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involution-based models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at https://github.com/d-li14/involution. Jie Hu 0019, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu 0012, Tong Zhang 0001, Qifeng Chen 0001 |
CVPR | 6 |
| 2021 | Learning the Superpixel in a Non-Iterative and Lifelong MannerabstractSuperpixel is generated by automatically clustering pixels in an image into hundreds of compact partitions, which is widely used to perceive the object contours for its excel-lent contour adherence. Although some works use the Convolution Neural Network (CNN) to generate high-quality superpixel, we challenge the design principles of these net-works, specifically for their dependence on manual labels and excess computation resources, which limits their flexibility compared with the traditional unsupervised segmentation methods. We target at redefining the CNN-based superpixel segmentation as a lifelong clustering task and pro-pose an unsupervised CNN-based method called LNS-Net. The LNS-Net can learn superpixel in a non-iterative and lifelong manner without any manual labels. Specifically, a lightweight feature embedder is proposed for LNS-Net to efficiently generate the cluster-friendly features. With those features, seed nodes can be automatically assigned to cluster pixels in a non-iterative way. Additionally, our LNS-Net can adapt the sequentially lifelong learning by rescaling the gradient of weight based on both channel and spatial context to avoid overfitting. Experiments show that the proposed LNS-Net achieves significantly better performance on three benchmarks with nearly ten times lower complexity compared with other state-of-the-art methods. Lei Zhu 0012, Qi She, Yanye Lu, Zhilin Lu 0002, Jie Hu 0019 |
CVPR | 1 |
| 2021 | MT-ORL: Multi-Task Occlusion Relationship LearningabstractRetrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely occlusion boundary extraction and occlusion orientation prediction, and secondly, improper representation of occlusion orientation. In this paper, we propose a novel architecture called Occlusion-shared and Path-separated Network (OPNet), which solves the first issue by exploiting rich occlusion cues in shared high-level features and structured spatial information in task-specific low-level features. We then design a simple but effective orthogonal occlusion representation (OOR) to tackle the second issue. Our method surpasses the state-of-the-art methods by 6.1%/8.3% Boundary-AP and 6.5%/10% Orientation-AP on standard PIOD/BSDS ownership datasets. Code is available at https://github.com/fengpanhe/MT-ORL. Panhe Feng, Qi She, Lei Zhu 0012, Lin Zhang 0040, Zijian Feng, Changhu Wang, Chunpeng Li, Xuejing Kang, Anlong Ming |
ICCV | 3 |
| 2021 | Unifying Nonlocal Blocks for Neural NetworksabstractThe nonlocal-based blocks are designed for capturing long-range spatial-temporal dependencies in computer vision tasks. Although having shown excellent performance, they still lack the mechanism to encode the rich, structured information among elements in an image or video. In this paper, to theoretically analyze the property of these nonlocal-based blocks, we provide a new perspective to interpret them, where we view them as a set of graph filters generated on a fully-connected graph. Specifically, when choosing the Chebyshev graph filter, a unified formulation can be derived for explaining and analyzing the existing nonlocal-based blocks (e.g., nonlocal block, nonlocal stage, double attention block). Furthermore, by concerning the property of spectral, we propose an efficient and robust spectral nonlocal block, which can be more robust and flexible to catch long-range dependencies when inserted into deep neural networks than the existing nonlocal blocks. Experimental results demonstrate the clear-cut improvements and practical applicabilities of our method on image classification, action recognition, semantic segmentation, and person re-identification tasks. Code are available at https://github.com/zh460045050/SNL_ICCV2021. Lei Zhu 0012, Qi She, Yanye Lu, Xuejing Kang, Jie Hu 0019, Changhu Wang |
ICCV | 1 |
| 2021 | HDNet: Hybrid Distance Network for semantic segmentation
Chunpeng Li, Xuejing Kang, Lei Zhu 0012, Lizhu Ye, Panhe Feng, Anlong Ming |
Neurocomputing | 3 |
| 2020 | Dynamic Random Walk for Superpixel SegmentationabstractIn this paper, we propose a novel random walk model, called Dynamic Random Walk (DRW), which adds a new type of dynamic node to the original RW model and reduces redundant calculation by limiting the walk range. To solve the seed-lacking problem of the proposed DRW, we redefine the energy function of the original RW and use the first arrival probability among each node pair to avoid the interference for each partition. Relaxation of our DRW is performed with the help of a greedy strategy and the Weighted Random Walk Entropy(WRWE) that uses the gradient feature to approximate the stationary distribution. The proposed DRW not only can enhance the boundary adherence but also can run with linear time complexity. To extend our DRW for superpixel segmentation, a seed initialization strategy is proposed. It can evenly distribute seeds in both 2D and 3D space and generate superpixels in only one iteration. The experimental results demonstrate that our DRW is faster than existing RW models and better than the state-of-the-art superpixel segmentation algorithms with respect to both efficiency and segmentation effects. Xuejing Kang, Lei Zhu 0012, Anlong Ming |
IEEE Trans. Image Process. | 2 |
| 2019 | Adaptive Occlusion Boundary Extraction for Depth InferenceabstractIn this paper, we propose an adaptive occlusion boundary extraction method for depth inference based on an adaptive segmentation and classification. First, an Adaptive DRW is proposed to generate more precise seeds and adaptive segmentation results, which can improve the feature quality and lower the boundary imbalance degree. Then, to deal with the imbalanced classification, we design a cost-sensitive boosting method-Adaptive AdaCost to better classify the imbalanced boundary, which can further improve overall performance and lower the cumulative misclassification cost and cost upper bound. Benefited from our Adaptive DRW and AdaCost, we extract more reliable and precise occlusion boundaries and use them for depth inference. The experiment results demonstrate that the combination of our Adaptive DRW and Adaptive AdaCost can produce more precise occlusion boundaries, and the depth inference result with our occlusion boundaries can be greatly improved. Lizhu Ye, Lei Zhu 0012, Xuejing Kang, Anlong Ming |
ICIP | 2 |
| 2018 | Dynamic Random Walk for Superpixel Segmentation
Lei Zhu 0012, Xuejing Kang, Anlong Ming, Xuesong Zhang 0001 |
ACCV (6) | 1 |