VLDB 2026 Research / reviewers in the wild / expert
Lei Zhu 0003
dblp:99/549-3
· DBLP profile ↗
181ranked-venue papers
15as first author
150since 2021 · last 2026
0000-0003-3871-663XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 123 · 14 first-author · 98 since 2021Artificial intelligence and machine learning · 74 · 5 first-author · 62 since 2021Applied, interdisciplinary, general and emerging computing · 57 · 2 first-author · 47 since 2021Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image DetectionabstractThe rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, existing VFM-based approaches primarily adhere to either perceptual or generative paradigms, each with limitations: perceptual models capture high-level semantics but often miss subtle artifacts, whereas generative models emphasize fine-grained flaws yet overlook semantic inconsistency. To resolve this inherent trade-off, we introduce SynerDetect, a novel hierarchical synergistic framework that fundamentally unifies the two paradigms. SynerDetect achieves deep integration of heterogeneous forensic representations through two levels of synergy: Cross-Model Interactive Distillation (CMID) distills generative forensic signals into perceptual encoders via prompt-guided reconstruction; and Optimal Transport-Guided Discriminative Contrastive Learning (OT-DCL) structurally aligns and integrates these heterogeneous representations, consolidating them into a robust, unified detection space. SynerDetect achieves superior performance on standard benchmarks (AIGCDetectBenchmark and GenImage) and attains a notable 5.20% accuracy gain on the challenging Chameleon benchmark, whose synthetic images consistently pass the Visual Turing Test. These results unequivocally validate the robust, real-world generalization of our unified cross-paradigm framework. Shuaibo Li, Zhaohu Xing, Hongqiu Wang, Pengfei Hao, Zekai Liu, Qing Zhang 0006, Lei Zhu 0003 |
AAAI | 9 |
| 2026 | S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without SupervisionabstractRecent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks. Jin Ye 0002, Hongqiu Wang, Changkai Ji, Jiashi Lin, Ziyan Huang, Chenglong Ma 0002, Tianbin Li, Junjun He, Lei Zhu 0003 |
AAAI | 13 |
| 2026 | Toward Real-World High-Precision Image Matting and SegmentationabstractHigh-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts generalization across different categories. Furthermore, the scarcity of high-quality annotation has led to a reliance on inharmonious synthetic data, resulting in poor generalization to real-world scenarios. To this end, we propose a Foreground Consistent Learning model, dubbed as FCLM, to address the aforementioned issues. Specifically, we first introduce a Depth-Aware Distillation strategy where we transfer the depth-related knowledge for better foreground representation. Considering the data dilemma, we term the processing of synthetic data as domain adaptation problem where we propose a domain-invariant learning strategy to focus on foreground learning. To support interactive prediction, we contribute an Object-Oriented Decoder that can receive both visual and language prompts to predict the referring target. Experimental results show that our method quantitatively and qualitatively outperforms state-of-the-art methods. Haipeng Zhou, Zhaohu Xing, Hongqiu Wang, Jun Ma 0008, Ping Li 0016, Lei Zhu 0003 |
AAAI | 6 |
| 2026 | BladderSense: A Wearable Ultrasound System for Continuous Bladder Monitoring in Real-World UseabstractContinuous and precise bladder volume monitoring is essential for patients with lower urinary tract dysfunction (LUTD) to support timely voiding and effective rehabilitation management. However, existing devices lack skin conformity and cannot support reliable "stick-once, long-term daily use" in real-world scenarios. We present BladderSense, a skin-conforming wireless wearable system featuring an X-shaped flexible phased-array ultrasound probe. Its development faces three key challenges: preserving beam focusing under skin deformation, maintaining accurate volume estimation despite bladder position shifts, and enabling low-power wireless transmission despite large raw data volume. To address these obstacles, we employ a suitable-frequency deep-focus design to stabilize beam quality, introduce a dual-orthogonal array with a shared geometric anchor and develop a coordinate-encoded deep learning (DL) model, together enabling bladder tracking and end-to-end volume estimation. An envelope-extraction-based compression scheme further enables Bluetooth Low Energy (BLE) transmission, supporting continuous monitoring with intermittent (1-minute) sensing. Experiments with 10 participants show that BladderSense provides accurate, robust bladder volume estimation across bladder changes, posture transitions, and dynamic daily activities, realizing dependable "stick-once, long-term monitoring" for LUTD patients. Kaixin Chen 0002, Usman Saleh Toro, Jinyu Lin, Chang Huang, Junfan Xiang, Lu Wang 0002, Huachen Cui, Lei Zhu 0003, Kaishun Wu |
MobiSys | 8 |
| 2026 | Video Shadow Detection with Intra-and Inter-video Cooperation
Zhihao Chen 0004, Junting Zhao, Lei Zhu 0003, Huazhu Fu, Wei Feng 0005 |
Int. J. Comput. Vis. | 4 |
| 2026 | SegRap2025: A benchmark of gross tumor volume and lymph node clinical target volume Segmentation for Radiotherapy Planning of nasopharyngeal carcinoma
Litingyu Wang, Chenyuan Bian, Zijun Gao, Chunbin Gu, Xin Weng, Jianghao Wu 0001, Yicheng Wu 0001, Jin Ye 0002, Linhao Li, Yiwen Ye, Yong Xia 0001, Elias Tappeiner, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Junqiang Chen, Chuanyi Huang, Lisheng Wang, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Shichuan Zhang, Shaoting Zhang 0001, Wenjun Liao, Guotai Wang |
Medical Image Anal. | 26 |
| 2026 | Reason like a radiologist: Chain-of-thought and reinforcement learning for verifiable report generationabstractRadiology report generation is critical for efficiency, but current models often lack the structured reasoning of experts and the ability to explicitly ground findings in anatomical evidence, which limits clinical trust and explainability. This paper introduces BoxMed-RL, a unified training framework to generate spatially verifiable and explainable chest X-ray reports. BoxMed-RL advances chest X-ray report generation through two integrated phases: (1) Pretraining Phase. BoxMed-RL learns radiologist-like reasoning through medical concept learning and enforces spatial grounding with reinforcement learning. (2) Downstream Adapter Phase. Pretrained weights are frozen while a lightweight adapter ensures fluency and clinical credibility. Experiments on two widely used public benchmarks (MIMIC-CXR and IU X-Ray) demonstrate that BoxMed-RL achieves an average 7 % improvement in both METEOR and ROUGE-L metrics compared to state-of-the-art methods. An average 5 % improvement in large language model-based metrics further underscores BoxMed-RL's robustness in generating high-quality reports. Related code and training templates are publicly available at https://github.com/ayanglab/BoxMed-RL. Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Huichi Zhou, Zhengqing Yuan, Zhifan Gao, Lei Zhu 0003, Giorgos Papanastasiou, Yingying Fang, Guang Yang 0006 |
Medical Image Anal. | 7 |
| 2026 | LODNeuS: A Flexible Lightweight Neural Implicit Surface Representation With Unconstrained Viewpoint RenderingabstractNeRF-like methods learn implicit 3D neural representations from 2D multiview images, enabling the synthesis of compelling novel views. However, to capture high-fidelity geometry, prior methods often rely on large-scale networks. This dependency hampers the potential applications of neural implicit representations, such as MR visualization. To address this, we introduce LODNeuS, an implicit surface representation based on feature voxel grids. LODNeuS captures multiple LODs of implicit geometry by maintaining voxel grids paired with a set of corresponding lightweight decoders. This allows for high-quality rendering with the ability to dynamically switch between detail levels. Another challenge is that existing methods, both volumetric and surface-based, tend to train and render their representations within a confined space, without explicitly restricting the sampling points properly. This lack of constraints can result in ambiguity, artifacts, and inefficient use of computational resources. We study this effect during free viewpoint rendering using conventional methods and develop an adaptive sampling scheme that emphasizes a valid geometric space for sampling point allocation. Our experimental results show that LODNeuS can match the visual quality of existing methods while offering flexible and lightweight inference. The benefits of adaptive sampling are also demonstrated in the free viewpoint rendering subsection. Our work extends the capabilities of neural implicit representations beyond previously defined limitations, broadening the scope of potential applications. Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Standing on the Giants: Informative Messenger Prompts With Self-Adapter for Image RestorationabstractDespite the recent advances in the application of diffusion models to the realm of image restoration, their inherent stochastic nature can often lead to inaccuracies in reconstructing spatial structures and fine details. This paper introduces an innovative paradigm that harnesses the rich knowledge encapsulated in existing high-level pre-trained models to guide the diffusion process, offering a flexible and potent approach to restoration tasks. However, there are significant challenges in leveraging a pre-trained model for image restoration, including a data gap between natural clean images and degradation images, and paradigm differences that cause insufficient intermediate knowledge. To tackle these issues, we introduce informative Messenger prompts and a Self-adapter for the pre-trained model, whose appropriate information acts as explicit constraints for diffusion, enabling reliable result generation. Specifically, ourMeSa-IRsuccessfully adapts to feature exploration for degraded samples, via disseminating distinctive information from degraded instances. Furthermore, it bolsters knowledge representation through the bidirectional interchange of hierarchical information facilitated by the innovative use of messenger prompts. Experimental results demonstrate the state-of-the-art performance of our framework on five tasks in terms of perceptual and distortion metrics. We will release codes at https://github.com/Ephemeral182/MeSa-IR. Sixiang Chen, Tian Ye 0001, Yulun Zhang 0001, Haoyu Chen 0003, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Addressing Client Drift in Federated Learning via Class-Prototype Similarity Distillation and Adaptive MaskabstractFederated learning (FL) enables multiple clients to learn collaboratively in a distributed way, allowing for privacy protection. However, the real-world nonindependent and identically distributed (non-IID) data will lead to client drift, which degrades the performance of FL. Interestingly, we find that the logit difference between the local and global models increases as the model is continuously updated, which is the primary factor behind performance degradation. This is mainly due to catastrophic forgetting caused by non-IID data between clients. To alleviate this problem, we propose a new algorithm, named FedCSD, a class-prototype similarity distillation in a federated framework to align the logits of local and global models. FedCSD does not simply transfer global knowledge to local clients, as an insufficiently trained global model cannot provide reliable knowledge, i.e., class similarity information, and its wrong soft labels will mislead the optimization of local models. Concretely, FedCSD leverages the similarity between local logits and the global prototype to refine the global logits, thereby enhancing its class similarity information. Furthermore, FedCSD adopts an adaptive mask to filter out the terrible soft labels of the global models, thereby preventing them from misleading local optimization. Extensive experiments demonstrate the superiority of our method over the state-of-the-art FL approaches in various non-IID settings. Code is publicly available at https://github.com/IAMJackYan/FedCSD. Yunlu Yan, Chun-Mei Feng 0001, Mang Ye, Wangmeng Zuo, Ping Li 0016, Rick Siow Mong Goh, Lei Zhu 0003, C. L. Philip Chen |
IEEE Trans. Cybern. | 7 |
| 2026 | SegMamba-V2: Long-Range Sequential Modeling Mamba for General 3-D Medical Image SegmentationabstractThe Transformer architecture has demonstrated remarkable results in 3D medical image segmentation due to its capability of modeling global relationships. However, it poses a significant computational burden when processing high-dimensional medical images. Mamba, as a State Space Model (SSM), has recently emerged as a notable approach for modeling long-range dependencies in sequential data. Although a substantial amount of Mamba-based research has focused on natural language and 2D image processing, few studies explore the capability of Mamba on 3D medical images. In this paper, we propose SegMamba-V2, a novel 3D medical image segmentation model, to effectively capture long-range dependencies within whole-volume features at each scale. To achieve this goal, we first devise a hierarchical scale downsampling strategy to enhance the receptive field and mitigate information loss during downsampling. Furthermore, we design a novel tri-orientated spatial Mamba block that extends the global dependency modeling process from one plane to three orthogonal planes to improve feature representation capability. Moreover, we collect and annotate a large-scale dataset (named CRC-2000) with fine-grained categories to facilitate benchmarking evaluation in 3D colorectal cancer (CRC) segmentation. We evaluate the effectiveness of our SegMamba-V2 on CRC-2000 and three other large-scale 3D medical image segmentation datasets, covering various modalities, organs, and segmentation targets. Experimental results demonstrate that our Segmamba-V2 outperforms state-of-the-art methods by a significant margin, which indicates the universality and effectiveness of the proposed model on 3D medical image segmentation tasks. The code for SegMamba-V2 is publicly available at: https://github.com/ge-xing/SegMamba-V2. Zhaohu Xing, Tian Ye 0001, Du Cai, Baowen Gai, Xiao-Jian Wu, Feng Gao 0023, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | LCM-Net: LLM-Driven Cross-Modality MoE Feature Fusion Network for Cancer Survival AnalysisabstractCancer survival analysis aims to predict survival outcomes to evaluate the efficacy and prognosis of treatment. Although current approaches have designed diverse cross-modal learning methods to integrate genetic data and pathology images, they are frequently hindered by data redundancy. Pattern representation in high-dimensional genetic data remains a significant hurdle. Pathology data analysis is computationally intensive because of the giga-pixel resolution. Moreover, the heterogeneity of data types poses a barrier to extending multimodal fusion methods. To address the aforementioned issues, we propose a novel LLM-driven Cross-Modality MoE-feature Fusion Network (LCM-Net) with three innovative modules for boosting cancer survival prediction. Specifically, the Genomic Language Alignment (GLA) module integrates genomic features with learnable prompts. Utilizing large language models, it encodes genomic information into concise and semantically relevant representations. Then, we devise the Pathological Feature Refinement (PFR) module to serve as a plug-and-play component that filters out irrelevant regions in pathology images. Finally, we propose a Multimodal Expert Integration (MEI) module to effectively leverage the capabilities of different experts, integrating the processed features from both the genomic and pathological domains. Extensive experiments on five public datasets demonstrate that our approach outperforms state-of-the-art methods, and the ablation study confirms the effectiveness of the proposed modules. Our code is publicly available at https://github.com/script-Yang/LCM-Net. Sicheng Yang 0001, Haipeng Zhou, Weiming Wang 0002, Shifu Chen, Guang Yang 0006, Huazhu Fu, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 8 |
| 2026 | Temporal Prompt Learning With Depth Memory for Video Mirror DetectionabstractMirror detection in dynamic scenes plays a crucial role in ensuring safety for various applications, such as drone tracking and robot navigation. However, current mirror detection models often fail in areas with mirrors that have a similar visual and color appearance to their surrounding objects. They also struggle to generalize well in complex cases, primarily due to limited annotated datasets. In this work, we propose a novel temporal prompt learning network with depth memory (TPD-Net) to address these critical challenges. Our approach includes several key components. First, we introduce a Temporal Prompt Generator (TPG) to learn temporal prompt features. Then, we devise Multi-layer Depth-aware Adaptor (MDA) modules to progressively adapt prompt features from the TPG, thereby learning mirror-related features by embedding temporal depth information as guidance. Moreover, we further refine these mirror-related features by constructing a depth memory and a Depth Memory Read module to read the temporal depths stored in the memory, boosting video mirror detection. Experimental results on a benchmark dataset show that our TPD-Net significantly outperforms 22 state-of-the-art methods in video mirror detection tasks. Our code, models, and results are publicly available athttps://github.com/ge-xing/TPDNet. Zhaohu Xing, Tian Ye 0001, Xin Yang 0011, Sixiang Chen, Huazhu Fu, Yan Nei Law, Lei Zhu 0003 |
IEEE Trans. Multim. | 7 |
| 2026 | Enhancing representation learning with frequency-augmented feature mixture for robust tuberculosis screening
Abudouresuli Tuersun, Mireayi Tudi, Abudoukeyoumu Abula, Pahatijiang Nijiati, Saimaitikari Abudoubari, Feng Gao 0023, Xiaojian Wu, Zekai Liu, Lei Zhu 0003, Mayidili Nijiati |
Vis. Comput. | 9 |
| 2025 | PromptHaze: Prompting Real-world Dehazing via Depth Anything ModelabstractReal-world image dehazing remains a challenging task due to the diverse nature of haze degradation and the lack of large-scale paired datasets. Existing methods based on hand-crafted priors or generative priors struggle to recover accurate backgrounds and fine details from dense haze regions. In this work, we propose a novel paradigm, PromptHaze, for real-world image dehazing via the depth prompt from the Depth Anything model. By employing a prompt-by-prompt strategy, our method iteratively updates the depth prompt and progressively restores the background through a dehazing network with controllable dehazing strength. Extensive experiments on widely-used real-world dehazing benchmarks demonstrate the superiority of PromptHaze in recovering authentic backgrounds and fine details from various haze scenes, outperforming state-of-the-art methods across multiple quality metrics. Tian Ye 0001, Sixiang Chen, Haoyu Chen 0003, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003 |
AAAI | 8 |
| 2025 | Residual Diffusion Deblurring Model for Single Image Defocus DeblurringabstractDefocus deblurring is a challenging task due to the spatially varying nature of defocus blur with multiple plausible solutions of a single given image. However, most existing methods falter when faced with extensive and variable defocus blur, either ignoring it or relying on additional loss functions to enhance perceptual quality. This often results in unrealistic reconstructions and compromised generalizability. In this paper, we propose a novel Residual Diffusion Deblurring Model framework for single image defocus deblurring. Our approach integrates a pre-trained defocus map estimator and a lightweight pre-deblur module with a learnable receptive field, providing crucial posterior information to effectively address large-scale and varying shaped defocus blur. In addition, a carefully-design denoising network enables the generation of diverse reconstructions from a single input. This approach not only significantly improves the perceptual quality of defocus deblurring outputs through multi-step residual learning, but also offers a more efficient inference strategy. Experimental results demonstrate that our method achieves competitive performance on real-world defocus deblurring image datasets across both perceptual and distortion evaluation metrics. Haoxuan Feng, Haohui Zhou, Tian Ye 0001, Sixiang Chen, Lei Zhu 0003 |
AAAI | 5 |
| 2025 | AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image EnhancementabstractExisting low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation. Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding |
AAAI | 9 |
| 2025 | POSTA: A Go-to Framework for Customized Artistic Poster GenerationabstractPoster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibitions, where both clear content delivery and visual impact are essential. To address these limitations, we present POSTA: a modular framework powered by diffusion models and multimodal large language models (MLLMs) for customized artistic poster generation. The framework consists of three modules. Background Diffusion creates a themed background based on user input. Design MLLM then generates layout and typography elements that align with and complement the background style. Finally, to enhance the poster’s aesthetic appeal, ArtText Diffusion applies additional stylization to key text elements. The final result is a visually cohesive and appealing poster, with a fully modular process that allows for complete customization. To train our models, we develop the PosterArt dataset, comprising high-quality artistic posters annotated with layout, typography, and pixel-level stylized text segmentation. Our comprehensive experimental analysis demonstrates POSTA’s exceptional controllability and design diversity, outperforming existing models in both text accuracy and aesthetic quality. Haoyu Chen 0003, Wenbo Li 0002, Tian Ye 0001, Songhua Liu, Ying-Cong Chen, Lei Zhu 0003, Xinchao Wang |
CVPR | 8 |
| 2025 | SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationabstractSnowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale and lack dedicated evaluation metrics designed specifically for snowfall degradation, thus hindering the effective integration of real snowy images into model training to reduce domain gaps. To address these challenges, we first introduce RealSnow10K, a large-scale, high-quality dataset consisting of over 10,000 annotated real-world snowy images. In addition, we curate a preference dataset comprising 36,000 expert-ranked image pairs, enabling the adaptation of multimodal large language models (MLLMs) to better perceive snowy image quality through our innovative Multi-Model Preference Optimization (MMPO). Finally, we propose the SnowMaster, which employs MMPO-enhanced MLLM to perform accurate snowy image evaluation and pseudo-label filtering for semi-supervised training. Experiments demonstrate that SnowMaster delivers superior desnowing performance under real-world conditions. Jianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye 0001, Yun Liu 0002, Song Fei, Zhaohu Xing, Weiming Wang 0002, Lei Zhu 0003 |
CVPR | 10 |
| 2025 | RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View ImagesabstractThis paper presents RoGSplat, a novel approach for synthesizing high-fidelity novel views of unseen human from sparse multi-view images, while requiring no cumbersome per-subject optimization. Unlike previous methods that typically struggle with sparse views with few overlappings and are less effective in reconstructing complex human geometry, the proposed method enables robust reconstruction in such challenging conditions. Our key idea is to lift SMPL vertices to dense and reliable 3D prior points representing accurate human body geometry, and then regress human Gaussian parameters based on the points. To account for possible misalignment between SMPL model and images, we propose to predict image-aligned 3D prior points by leveraging both pixel-level features and voxel-level features, from which we regress the coarse Gaussians. To enhance the ability to capture high-frequency details, we further render depth maps from the coarse 3D Gaussians to help regress fine-grained pixel-wise Gaussians. Experiments on several benchmark datasets demonstrate that our method outperforms state-of-the-art methods in novel view synthesis and cross-dataset generalization. Our code is available at https://github.com/iSEE-Laboratory/RoGSplat. Junjin Xiao, Qing Zhang 0006, Yonewei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
CVPR | 4 |
| 2025 | Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data EngineabstractMirror detection is a challenging task because a mirror’s visual appearance varies depending on the reflected content. Due to limited annotated data, current methods failed to generalize well for detecting diverse mirror scenes. Semi-supervised learning with large-scale unlabeled data can improve generalization capabilities on mirror detection, but these methods often suffer from unreliable pseudo-labels due to distribution differences between labeled and unlabeled data, therefore affecting the learning process. To address this issue, we first collect a large-scale dataset of approximately 0.4 million mirror-related images from the internet, significantly expanding the data scale for mirror detection. To effectively exploit this unlabeled dataset, we propose the first semi-supervised framework (namely an iterative data engine) consisting of four steps: (1) mirror detection model training, (2) pseudo label prediction, (3) dual guidance scoring, and (4) selection of highly reliable pseudo labels. In each iteration of the data engine, we employ a geometric accuracy scoring approach to assess pseudo labels based on multiple segmentation metrics, and design a multi-modal agent-driven semantic scoring approach to enhance the semantic perception of pseudo labels. These two scoring approaches can effectively improve the reliability of pseudo labels by selecting unlabeled samples with higher scores. Our method demonstrates promising performance across three mirror detection tasks and exhibits strong generalization on unseen examples. Our code will be publicly available at https://github.com/ge-xing/DAM. Zhaohu Xing, Hongqiu Wang, Tian Ye 0001, Sixiang Chen, Wenxue Li 0003, Guang Liu 0006, Lei Zhu 0003 |
CVPR | 9 |
| 2025 | A Simple Data Augmentation for Feature Distribution Skewed Federated LearningabstractFederated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribution skewed FL scenario, a common non-IID situation in real-world applications where data from different clients exhibit varying underlying distributions. This variation leads to feature shift, which is a key issue of this scenario. While previous works have made notable progress, few pay attention to the data itself, i.e., the root of this issue. The primary goal of this paper is to mitigate feature shift from the perspective of data. To this end, we propose a simple yet remarkably effective input-level data augmentation method, namely FedRDN, which randomly injects the statistical information of the local distribution from the entire federation into the client’s data. This is beneficial to improve the generalization of local feature representations, thereby mitigating feature shift. Moreover, our FedRDN is a plug-and-play component, which can be seamlessly integrated into the data augmentation flow with only a few lines of code. Extensive experiments on several datasets show that the performance of various representative FL methods can be further improved by integrating our FedRDN, demonstrating its effectiveness, strong compatibility and generalizability. Code is available at https://github.com/IAMJackYan/FedRDN. Yunlu Yan, Huazhu Fu, Yuexiang Li, Jinheng Xie, Jun Ma 0008, Guang Yang 0006, Lei Zhu 0003 |
CVPR | 7 |
| 2025 | Leveraging Vision-Language Models for Referring Video Shadow DetectionabstractReferring Video Shadow Detection (RVSD) is an emerging vision-language task that segments specific shadow regions in videos based on natural language descriptions, combining the flexibility of language-guided localization with the temporal dynamics of video understanding. While initial approaches have demonstrated promising results, existing methods remain limited in their ability to fully leverage the rich semantic relationships between textual prompts and visual shadow features, particularly in complex scenes with multiple shadows or dynamically changing lighting conditions. This paper presents ShadowVLM, a novel vision-language model framework that leverages the crossmodal representation capabilities of large-scale vision-language models. By integrating multiscale image perception and multiscale video perception mechanisms, ShadowVLM achieves robust performance in complex multimodal understanding tasks. Our approach introduces three key innovations: (1) A hierarchical visual understanding module that extracts and interprets image features from multiple spatial perspectives, enhancing the model's visual perception. This design enables comprehensive image understanding at global, local, and fine-grained levels, improving the model's ability to capture diverse visual cues; (2) A query propagation mechanism for temporal consistency, which extends the model's temporal awareness via a specially designed memory bank. This mechanism ensures stable target tracking across frames, even in complex and dynamic environments; (3) The use of a large-scale vision-language model's cross-modal representation capability to align textual features with shadowed visual inputs. This alignment allows the model to acquire prior knowledge about shadowed objects, thereby improving its target recognition ability under challenging visual conditions. Extensive experiments demonstrate that ShadowVLM significantly outperforms previous state-of-the-art methods. Chenjun Liang, Luping Shan, Yunxun Liu, Qinglian Xue, Danying Xu, Lei Zhu 0003 |
CW | 9 |
| 2025 | Ultra-High-Definition Image Deraining via Dual-Tree Complex Wavelet Representation and State Space ModelsabstractUltra-High-Definition (UHD) image restoration under heavy rain conditions presents unique challenges that existing deraining methods fail to adequately address. Current approaches either neglect the rain veiling effect or suffer from significant information loss when processing high-resolution images. We propose DTHRNet, a novel framework that combines the multi-directional analysis capabilities of Dual-tree Complex Wavelet Transform (DTCWT) with the efficient long-range modeling of State Space Models (SSMs). Our method introduces three key innovations: (1) DTCWTbased downsampling that preserves directional rain streak information while mitigating detail loss, (2) Mamba-based architecture that maintains linear computational complexity for UHD processing, and (3) the first dedicated UHD heavy rain dataset (UHDHRain) for comprehensive evaluation. Experiments demonstrate that our approach significantly outperforms existing methods in both quantitative metrics and visual quality, particularly in preserving fine details while effectively removing both rain streaks and veiling effects. Songtao Zeng, Haobo Shen, Zichong Zhang, Song Fei, Lei Zhu 0003 |
CW | 10 |
| 2025 | QuantPrompt: Semantic Codebook-Guided Prompting for Mirror DetectionabstractMirror detection is inherently difficult due to the ambiguous nature of reflections and the absence of distinctive visual indicators. While existing approaches often focus on specific mirror attributes to enhance detection accuracy, they tend to overlook the complex and diverse environments in which mirrors typically appear. Recently, vector quantized models have shown promise in capturing discrete features of images, enabling a more comprehensive understanding of environmental context. Inspired by this, we propose a novel framework-QuantPrompt-which decouples environmental context from mirror-specific features for more effective mirror detection. Our framework begins with a Semantic Quantization Module, which encodes input images into discrete semantic tokens using a learnable codebook. This step compresses continuous visual features into context-rich tokens, enhancing the model's understanding of the surrounding environment. To complement this, we introduce a Semantic Reservation Module with a dynamic prompt decoder. This component learns mirror-specific features while preserving the original semantic context through a frozen encoder, allowing the model to better distinguish mirrors from their surroundings. Extensive experiments on benchmark mirror detection datasets show that QuantPrompt achieves state-of-the-art performance, demonstrating superior accuracy and generalization compared to existing methods. Songtao Zeng, Haobo Shen, Zichong Zhang, Zhaohu Xing, Lei Zhu 0003 |
CW | 10 |
| 2025 | Genhaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World DehazingabstractReal-world image dehazing is crucial for enhancing visual quality in computer vision applications. However, existing physics-based haze generation paradigms struggle to model the complexities of real-world haze and lack controllability, limiting the performance of existing baselines on real-world images. In this paper, we introduce GenHaze, a pioneering haze generation framework that enables the one-step generation of high-quality, reference-controllable hazy images. GenHaze leverages the pre-trained latent diffusion model (LDM) with a carefully designed clean-to-haze generation protocol to produce realistic hazy images. Additionally, by leveraging its fast, controllable generation of paired highquality hazy images, we illustrate that existing dehazing baselines can be unleashed in a simple and efficient manner. Extensive experiments indicate that GenHaze achieves visually convincing and quantitatively superior hazy images. It also significantly improves multiple existing dehazing models across 7 non-reference metrics with minimal fine-tuning epochs. Our work demonstrates that LDM possesses the potential to generate realistic degradations, providing an effective alternative to prior generation pipelines. Sixiang Chen, Tian Ye 0001, Yunlong Lin, Yeying Jin, Haoyu Chen 0003, Jianyu Lai, Song Fei, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003 |
ICCV | 11 |
| 2025 | GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection
Wenxue Li 0003, Tian Ye 0001, Xinyu Xiong, Jinbin Bai, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu 0003 |
ICCV | 10 |
| 2025 | Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
Wenbo Li 0002, Zhongdao Wang, Haoze Sun, Bangzhen Liu, Haoyu Chen 0003, Aoxue Li, Lei Zhu 0003 |
ICCV | 12 |
| 2025 | Toward Fair and Accurate Cross-Domain Medical Image Segmentation: a Vlm-Driven Active Domain Adaptation Paradigm
Hongqiu Wang, Xiangde Luo, Zhaohu Xing, Harry Qin, Shaozhi Wu, Lei Zhu 0003 |
ICCV | 8 |
| 2025 | Medical World Model
Zhao-Yang Wang, Qiuping Liu, Shuwen Sun, Kang Wang 0016, Rama Chellappa, Zongwei Zhou, Alan L. Yuille, Lei Zhu 0003, Jieneng Chen |
ICCV | 9 |
| 2025 | Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image SynthesisabstractWe present Meissonic, which elevates non-autoregressive text-to-image Masked Image Modeling (MIM) to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing methods in generating high-quality, high-resolution images. Extensive experiments validate Meissonic’s capabilities, demonstrating its potential as a new standard in text-to-image synthesis. Jinbin Bai, Tian Ye 0001, Wei Chow, Enxin Song, Xiangtai Li, Zhen Dong 0003, Lei Zhu 0003, Shuicheng Yan |
ICLR | 8 |
| 2025 | On the Importance of Language-driven Representation Learning for Heterogeneous Federated LearningabstractNon-Independent and Identically Distributed (Non-IID) training data significantly challenge federated learning (FL), impairing the performance of the global model in distributed frameworks. Inspired by the superior performance and generalizability of language-driven representation learning in centralized settings, we explore its potential to enhance FL for handling non-IID data. In specific, this paper introduces FedGLCL, a novel language-driven FL framework for image-text learning that uniquely integrates global language and local image features through contrastive learning, offering a new approach to tackle non-IID data in FL. FedGLCL redefines FL by avoiding separate local training models for each client. Instead, it uses contrastive learning to harmonize local image features with global textual data, enabling uniform feature learning across different local models. The utilization of a pre-trained text encoder in FedGLCL serves a dual purpose: it not only reduces the variance in local feature representations within FL by providing a stable and rich language context but also aids in mitigating overfitting, particularly to majority classes, by leveraging broad linguistic knowledge. Extensive experiments show that FedGLCL significantly outperforms state-of-the-art FL algorithms across different non-IID scenarios. Yunlu Yan, Chun-Mei Feng 0001, Wangmeng Zuo, Salman Khan 0001, Yong Liu 0026, Lei Zhu 0003 |
ICLR | 6 |
| 2025 | Federated Residual Low-Rank Adaptation of Large Language ModelsabstractLow-Rank Adaptation (LoRA) presents an effective solution for federated fine-tuning of Large Language Models (LLMs), as it substantially reduces communication overhead. However, a straightforward combination of FedAvg and LoRA results in suboptimal performance, especially under data heterogeneity. We noted this stems from both intrinsic (i.e., constrained parameter space) and extrinsic (i.e., client drift) limitations, which hinder it effectively learn global knowledge. In this work, we proposed a novel Federated Residual Low-Rank Adaption method, namely FRLoRA, to tackle above two limitations. It directly sums the weight of the global model parameters with a residual low-rank matrix product (\ie, weight change) during the global update step, and synchronizes this update for all local models. By this, FRLoRA performs global updates in a higher-rank parameter space, enabling a better representation of complex knowledge structure. Furthermore, FRLoRA reinitializes the local low-rank matrices with the principal singular values and vectors of the pre-trained weights in each round, to calibrate their inconsistent convergence, thereby mitigating client drift. Our extensive experiments demonstrate that FRLoRA consistently outperforms various state-of-the-art FL methods across nine different benchmarks in natural language understanding and generation under different FL scenarios. Yunlu Yan, Chun-Mei Feng 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026, Lei Zhu 0003 |
ICLR | 6 |
| 2025 | Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped RoadmapabstractCompared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and control framework for navigating robots in unknown and cluttered environments using a dynamically constructed starshaped roadmap. Our approach generates a starshaped representation of the surrounding free space from real-time sensor data using piece-wise polynomials. Additionally, an incremental roadmap maintaining the connectivity information is constructed, and a searching algorithm efficiently selects short-term goals on this roadmap. Importantly, this framework addresses dead-end situations with a graph updating mechanism. To ensure safe and efficient movement within the starshaped roadmap, we propose a reactive controller based on Dynamic System Modulation (DSM). This controller facilitates smooth motion within starshaped regions and their intersections, avoiding conservative and short-sighted behaviors and allowing the system to handle intricate obstacle configurations in unknown and cluttered environments. Comprehensive evaluations in both simulations and real-world experiments show that the proposed method achieves higher success rates and reduced travel times compared to other methods. It effectively manages intricate obstacle configurations, avoiding conservative and myopic behaviors. The source code will be released on website11Available at: github.com/kkkkkaiai/starshaped_roadmap. Kai Chen 0006, Haichao Liu 0003, Yulin Li 0001, Jianghua Duan, Lei Zhu 0003, Jun Ma 0008 |
ICRA | 5 |
| 2025 | GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory PredictionabstractAccurate prediction of pedestrian trajectories is crucial for improving the safety of autonomous driving. However, this task is generally nontrivial due to the inherent stochasticity of human motion, which naturally requires the predictor to generate multi-modal prediction. Previous works leverage various generative methods, such as GAN and VAE, for pedestrian trajectory prediction. Nevertheless, these methods may suffer from mode collapse and relatively low-quality results. The denoising diffusion probabilistic model (DDPM) has recently been applied to trajectory prediction due to its simple training process and powerful reconstruction ability. However, current diffusion-based methods do not fully utilize input information and usually require many denoising iterations that lead to a long inference time or an additional network for initialization. To address these challenges and facilitate the use of diffusion models in multi-modal trajectory prediction, we propose GDTS, a novel Goal-Guided Diffusion Model with Tree Sampling for multi-modal trajectory prediction. Considering the "goal-driven" characteristics of human motion, GDTS leverages goal estimation to guide the generation of the diffusion network. A two-stage tree sampling algorithm is presented, which leverages common features to reduce the inference time and improve accuracy for multi-modal prediction. Experimental results demonstrate that our proposed framework achieves comparable state-of-the-art performance with real-time inference speed in public datasets. Sheng Wang 0017, Lei Zhu 0003, Ming Liu 0001, Jun Ma 0008 |
IROS | 3 |
| 2025 | Surgical-MambaLLM: Mamba2-Enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang 0006, Kaishun Wu, Lei Zhu 0003 |
MICCAI (9) | 7 |
| 2025 | Source-Free Active Domain Adaptation for Efficient Medical Video Polyp Segmentation
Hongqiu Wang, Weiming Wang 0002, Harry Qin, Qiong Wang 0001, Lei Zhu 0003 |
MICCAI (10) | 6 |
| 2025 | Toward Medical Deepfake Detection: A Comprehensive Dataset and Novel Method
Shuaibo Li, Zhaohu Xing, Hongqiu Wang, Pengfei Hao, Zekai Liu, Lei Zhu 0003 |
MICCAI (14) | 7 |
| 2025 | FSA-Net: Fractal-Driven Synergistic Anatomy-Aware Network for Segmenting White Line of Toldt in Laparoscopic Images
Kecheng Wu, Zhaohu Xing, Zerong Cai, Feng Gao 0023, Wenxue Li 0003, Lei Zhu 0003 |
MICCAI (9) | 6 |
| 2025 | HybridMamba: A Dual-Domain Mamba for 3D Medical Image Segmentation
Weitong Wu 0003, Zhaohu Xing, Qin Peng, Lei Zhu 0003 |
MICCAI (3) | 5 |
| 2025 | Temporal Model-Based Federated Active Medical Image Classification
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Jinheng Xie, Jun Chen 0005, Mohamed Elhoseiny 0001, Kaishun Wu, Lei Zhu 0003 |
MICCAI (14) | 9 |
| 2025 | HRVVS: A High-Resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors
Xincheng Yao, Kangwei Guo, Ruiqiang Xiao, Haipeng Zhou, Haisu Tao, Lei Zhu 0003 |
MICCAI (10) | 8 |
| 2025 | CoC: Chain-of-Cancer Based on Cross-Modal Autoregressive Traction for Survival Prediction
Haipeng Zhou, Sicheng Yang 0001, Harry Qin, Lei Zhu 0003 |
MICCAI (15) | 6 |
| 2025 | AdvMIM: Adversarial Masked Image Modeling for Semi-supervised Medical Image Segmentation
Lei Zhu 0003, Jun Zhou 0014, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (16) | 1 |
| 2025 | Farther Than Mirror: Explore Pattern-Compensated Depth of Mirror with Temporal Changes for Video Mirror Detection
Zhaohu Xing, Tian Ye 0001, Sixiang Chen, Guang Liu 0006, Lei Zhu 0003 |
ACM Multimedia | 7 |
| 2025 | VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image SegmentationabstractConsistency learning with feature perturbation is a widely used strategy in semi-supervised medical image segmentation. However, many existing perturbation methods rely on dropout, and thus require a careful manual tuning of the dropout rate, which is a sensitive hyperparameter and often difficult to optimize and may lead to suboptimal regularization. To overcome this limitation, we propose VQ-Seg, the first approach to employ vector quantization (VQ) to discretize the feature space and introduce a novel and controllable Quantized Perturbation Module (QPM) that replaces dropout. Our QPM perturbs discrete representations by shuffling the spatial locations of codebook indices, enabling effective and controllable regularization. To mitigate potential information loss caused by quantization, we design a dual-branch architecture where the post-quantization feature space is shared by both image reconstruction and segmentation tasks. Moreover, we introduce a Post-VQ Feature Adapter (PFA) to incorporate guidance from a foundation model (FM), supplementing the high-level semantic information lost during quantization. Furthermore, we collect a large-scale Lung Cancer (LC) dataset comprising 828 CT scans annotated for central-type lung carcinoma. Extensive experiments on the LC dataset and other public benchmarks demonstrate the effectiveness of our method, which outperforms state-of-the-art approaches. Codes will be released. Sicheng Yang 0001, Zhaohu Xing, Lei Zhu 0003 |
NeurIPS | 3 |
| 2025 | Ad2Mix: Adversarial and Adaptive Mixup for Unsupervised Domain AdaptationabstractTransformer has recently gained tremendous popularity in unsupervised domain adaptation tasks due to its superior generalization ability. State-of-the-art methods leverage mixup to build an intermediate domain to reduce domain gap. However, such strategy becomes less effective when the domain gap becomes large, as the domain gap between intermediate domain and source domain is not minimized and the constructed intermediate domain is non informative. How to address the adaptation problem when domain gap becomes large is an important research problem in domain adaptation. In this paper, we propose an adversarial and adaptive mixup (Ad2mix) framework which gradually aligns the intermediate domain towards source domain to fully unleash the potential of both the transformer architecture and mixup to address the large domain gap problem. Specifically, we formulate a general framework for intermediate domain learning with mixup. We propose adversarial mixup with a specially designed mixup alike adversarial adaptation operation to reduce the domain gap between the intermediate domain and source domain. To construct an informative intermediate domain, unlike existing methods which utilize a Beta distribution to generate mixup coefficients to interpolate source and target data, we adaptively assign mixup coefficient for each target data instance based on their transferability and discriminativity information. Our framework creates a natural curriculum of intermediate domains from near source domain to near target domain for gradual adaptation. Extensive experimental studies and evaluations on three public domain adaptation benchmark datasets and one medical domain adaptation task demonstrate the superiority of our framework. Lei Zhu 0003, Yanyu Xu 0001, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
WACV | 1 |
| 2025 | A structure-based disentangled network with contrastive regularization for sign language recognition
Liqing Gao, Lei Zhu 0003, Lianyu Hu 0003, Liang Wang 0001, Wei Feng 0005 |
Expert Syst. Appl. | 2 |
| 2025 | Triplane-Smoothed Video Dehazing with CLIP-Enhanced Generalization
Haoyu Chen 0003, Tian Ye 0001, Lei Zhu 0003 |
Int. J. Comput. Vis. | 5 |
| 2025 | An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challengeabstractAugmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain. Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli |
Medical Image Anal. | 13 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 17 |
| 2025 | Revisiting medical image retrieval via knowledge consolidationabstractAs artificial intelligence and digital medicine increasingly permeate healthcare systems, robust governance frameworks are essential to ensure ethical, secure, and effective implementation. In this context, medical image retrieval becomes a critical component of clinical data management, playing a vital role in decision-making and safeguarding patient information. Existing methods usually learn hash functions using bottleneck features, which fail to produce representative hash codes from blended embeddings. Although contrastive hashing has shown superior performance, current approaches often treat image retrieval as a classification task, using category labels to create positive/negative pairs. Moreover, many methods fail to address the out-of-distribution (OOD) issue when models encounter external OOD queries or adversarial attacks. In this work, we propose a novel method to consolidate knowledge of hierarchical features and optimization functions. We formulate the knowledge consolidation by introducing Depth-aware Representation Fusion (DaRF) and Structure-aware Contrastive Hashing (SCH). DaRF adaptively integrates shallow and deep representations into blended features, and SCH incorporates image fingerprints to enhance the adaptability of positive/negative pairings. These blended features further facilitate OOD detection and content-based recommendation, contributing to a secure AI-driven healthcare environment. Moreover, we present a content-guided ranking to improve the robustness and reproducibility of retrieval results. Our comprehensive assessments demonstrate that the proposed method could effectively recognize OOD samples and significantly outperform existing approaches in medical image retrieval (p < 0 . 05 ). In particular, our method achieves a 5.6–38.9% improvement in mean Average Precision on the anatomical radiology dataset. • Structure-aware pairing using image fingerprints to address over-centralized issues. • A novel model to consolidate hierarchical embeddings for representation learning. • Addressing ill-posed gradient issues introduced by relaxed Hamming distance. • A self-supervised OOD detection module by evaluating image reconstruction disparity. • Content-guided ranking mechanism for robust and precise retrieval. Yang Nan 0002, Huichi Zhou, Xiaodan Xing, Giorgos Papanastasiou, Lei Zhu 0003, Zhifan Gao, Alejandro F. Frangi, Guang Yang 0006 |
Medical Image Anal. | 5 |
| 2025 | Diff-UNet: A diffusion embedded network for robust 3D medical image segmentation
Zhaohu Xing, Huazhu Fu, Guang Yang 0006, Lequan Yu, Bai Ying Lei, Lei Zhu 0003 |
Medical Image Anal. | 8 |
| 2025 | Unifying Physically-Informed Weather Priors in a Single Model for Image Restoration Across Multiple Adverse Weather ConditionsabstractImage restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather-related artifacts vary with the particle’s distance to the camera according to the established scene visibility analysis, where close and faraway regions are more affected by falling drops and fog effects, respectively. In challenging weather conditions, existing image restoration methods fall short by not accounting for the varying impact of adverse weather on different scene regions. We develop a novel unified imaging model combined with a weather-prior-based network that directly incorporates weather-specific physical imaging processes into the restoration process. This approach not only enhances visibility in both near and distant regions affected by drops but also outperforms current state-of-the-art methods by effectively mitigating artifacts such as fog. Our contributions include a comprehensive analysis of weather-related visual factors and the development of an innovative network architecture that leverages estimated occlusion and transmission to restore scene details. Experimental results on three synthetic benchmarks, including our Weather30K dataset, along with two all-weather datasets, and a real-world benchmark with challenging mixed weather conditions, show the superiority of our method against state-of-the-art methods. Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Vivim: A Video Vision Mamba for Ultrasound Video SegmentationabstractUltrasound video segmentation gains increasing attention in clinical practice due to the redundant dynamic references in video frames. However, traditional convolutional neural networks have a limited receptive field and transformer-based networks are unsatisfactory in constructing long-term dependency from the perspective of computational complexity. This bottleneck poses a significant challenge when processing longer sequences in medical video analysis tasks using available devices with limited memory. Recently, state space models (SSMs), famous by Mamba, have exhibited linear complexity and impressive achievements in efficient long sequence modeling, which have developed deep neural networks by expanding the receptive field on many vision tasks significantly. Unfortunately, vanilla SSMs failed to simultaneously capture causal temporal cues and preserve non-casual spatial information. To this end, this paper presents a Video Vision Mamba-based framework, dubbed as Vivim, for ultrasound video segmentation tasks. Our Vivim can effectively compress the long-term spatiotemporal representation into sequences at varying scales with our designed Temporal Mamba Block. We also introduce an improved boundary-aware affine constraint across frames to enhance the discriminative ability of Vivim on ambiguous lesions. Extensive experiments on thyroid segmentation in ultrasound videos, breast lesion segmentation in ultrasound videos, and polyp segmentation in colonoscopy videos demonstrate the effectiveness and efficiency of our Vivim, superior to existing methods. The code and dataset are available at: https://github.com/scott-yjyang/Vivim. Zhaohu Xing, Lequan Yu, Huazhu Fu, Chunwang Huang, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Enhancing Visual Reasoning With LLM-Powered Knowledge Graphs for Visual Question Localized-Answering in Robotic SurgeryabstractExpert surgeons often have heavy workloads and cannot promptly respond to queries from medical students and junior doctors about surgical procedures. Thus, research on Visual Question Localized-Answering in Surgery (Surgical-VQLA) is essential to assist medical students and junior doctors in understanding surgical scenarios. Surgical-VQLA aims to generate accurate answers and locate relevant areas in the surgical scene, requiring models to identify and understand surgical instruments, operative organs, and procedures. A key issue is the model's ability to accurately distinguish surgical instruments. Current Surgical-VQLA models rely primarily on sparse textual information, limiting their visual reasoning capabilities. To address this issue, we propose a framework called Enhancing Visual Reasoning with LLM-Powered Knowledge Graphs (EnVR-LPKG) for the Surgical-VQLA task. This framework enhances the model's understanding of the surgical scenario by utilizing knowledge graphs of surgical instruments constructed by the Large Language Model (LLM). Specifically, we design a Fine-grained Knowledge Extractor (FKE) to extract the most relevant information from knowledge graphs and perform contrastive learning with the extracted knowledge graphs and local image. Furthermore, we design a Multi-attention-based Surgical Instrument Enhancer (MSIE) module, which employs knowledge graphs to obtain an enhanced representation of the corresponding surgical instrument in the global scene. Through the MSIE module, the model can learn how to fuse visual features with knowledge graph text features, thereby strengthening the understanding of surgical instruments and further improving visual reasoning capabilities. Extensive experimental results on the EndoVis-17-VQLA and EndoVis-18-VQLA datasets demonstrate that our proposed method outperforms other state-of-the-art methods. We will release our code for future research. Pengfei Hao, Hongqiu Wang, Guang Yang 0006, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Cascaded Inner-Outer Clip Retformer for Ultrasound Video Object SegmentationabstractComputer-aided ultrasound (US) imaging is an important prerequisite for early clinical diagnosis and treatment. Due to the harsh ultrasound (US) image quality and the blurry tumor area, recent memory-based video object segmentation models (VOS) achieve frame-level segmentation by performing intensive similarity matching among the past frames which could inevitably result in computational redundancy. In this paper, we first build a larger annotated benchmark dataset for breast lesion segmentation in ultrasound videos, then we propose a lightweight clip-level VOS framework for achieving higher segmentation accuracy while maintaining the speed. Then an Inner-Outer Clip Retformer is proposed to extract spatial-temporal tumor features in parallel. Specifically, the proposed Outer Clip Retformer extracts the tumor movement feature from past video clips to locate the current clip tumor position, while the Inner Clip Retformer detailedly extracts current tumor features that can produce more accurate segmentation results. Then a Clip Contrastive loss function is further proposed to align the extracted tumor features along both the spatial-temporal dimensions to improve the segmentation accuracy. In addition, the Global Retentive Memory is proposed to maintain the complementary tumor features with lower computing resources which can generate coherent temporal movement features. In this way, our model can significantly improve the spatial-temporal perception ability without increasing a large number of parameters, achieving more accurate segmentation results while maintaining a faster segmentation speed. Finally, we conduct extensive experiments to evaluate our proposed model on several video object segmentation datasets, the results show that our framework outperforms state-of-the-art segmentation methods. Lei Zhu 0003, Zhaohu Xing, Baoliang Zhao, Ying Hu 0001, Faqin Lv, Qiong Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI ReconstructionabstractWhile multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities). Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Serp-Mamba: Advancing High-Resolution Retinal Vessel Segmentation With Selective State-Space ModelabstractUltra-Wide-Field Scanning Laser Ophthalmoscopy (UWF-SLO) images capture high-resolution views of the retina with typically spanning 200 degrees. Accurate segmentation of vessels in UWF-SLO images is essential for detecting and diagnosing fundus disease. Recent studies highlight that Mamba's selective State Space Model (SSM) excels in modeling long-range dependencies with linear computational complexity, making it highly suitable for preserving the continuity of elongated vessel structures, especially for high-resolution UWF images. Inspired by this, we propose the Serpentine Mamba (Serp-Mamba) network to address this challenging task. Specifically, we recognize the intricate, varied, and delicate nature of the tubular structure of vessels. Furthermore, the high-resolution of UWF-SLO images exacerbates the imbalance between the vessel and background categories. Based on the above observations, we first devise a Serpentine Interwoven Adaptive (SIA) scan mechanism, which scans UWF-SLO images along curved vessel structures in a snake-like crawling manner. This approach, consistent with vascular texture transformations, ensures the effective and continuous capture of curved vascular structure features. Second, we propose an Ambiguity-Driven Dual Recalibration (ADDR) module to address the category imbalance problem intensified by high-resolution images. Our ADDR module delineates pixels by two learnable thresholds and refines ambiguous pixels through a dual-driven strategy, thereby accurately distinguishing vessels and background regions. Experiment results on three datasets demonstrate the superior performance of our Serp-Mamba on high-resolution vessel segmentation. We also conduct a series of ablation studies to verify the impact of our designs. Our code will be released upon publication (https://github.com/whq-xxh/Serp-Mamba). Hongqiu Wang, Bin Sheng 0001, Huazhu Fu, Guang Yang 0006, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 9 |
| 2025 | DiffMIC-v2: Medical Image Classification via Improved Diffusion NetworkabstractRecently, Denoising Diffusion Models have achieved outstanding success in generative image modeling and attracted significant attention in the computer vision community. Although a substantial amount of diffusion-based research has focused on generative tasks, few studies apply diffusion models to medical diagnosis. In this paper, we propose a diffusion-based network (named DiffMIC-v2) to address general medical image classification by eliminating unexpected noise and perturbations in image representations. To achieve this goal, we first devise an improved dual-conditional guidance strategy that conditions each diffusion step with multiple granularities to enhance step-wise regional attention. Furthermore, we design a novel Heterologous diffusion process that achieves efficient visual representation learning in the latent space. We evaluate the effectiveness of our DiffMIC-v2 on four medical classification tasks with different image modalities, including thoracic diseases classification on chest X-ray, placental maturity grading on ultrasound images, skin lesion classification using dermatoscopic images, and diabetic retinopathy grading using fundus images. Experimental results demonstrate that our DiffMIC-v2 outperforms state-of-the-art methods by a significant margin, which indicates the universality and effectiveness of the proposed model on multi-class and multi-label classification tasks. DiffMIC-v2 can use fewer iterations than our previous DiffMIC to obtain accurate estimations, and also achieves greater runtime efficiency with superior results. The code will be publicly available at https://github.com/scott-yjyang/DiffMICv2. Huazhu Fu, Angelica I. Avilés-Rivero, Zhaohu Xing, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Norest-Net: Normal Estimation Neural Network for 3-D Noisy Point CloudsabstractThe widely deployed ways to capture a set of unorganized points, e.g., merged laser scans, fusion of depth images, and structure-from- , usually yield a 3-D noisy point cloud. Accurate normal estimation for the noisy point cloud makes a crucial contribution to the success of various applications. However, the existing normal estimation wisdoms strive to meet a conflicting goal of simultaneously performing normal filtering and preserving surface features, which inevitably leads to inaccurate estimation results. We propose a normal estimation neural network (Norest-Net), which regards normal filtering and feature preservation as two separate tasks, so that each one is specialized rather than traded off. For full noise removal, we present a normal filtering network (NF-Net) branch by learning from the noisy height map descriptor (HMD) of each point to the ground-truth (GT) point normal; for surface feature recovery, we construct a normal refinement network (NR-Net) branch by learning from the bilaterally defiltered point normal descriptor (B-DPND) to the GT point normal. Moreover, NR-Net is detachable to be incorporated into the existing normal estimation methods to boost their performances. Norest-Net shows clear improvements over the state of the arts in both feature preservation and noise robustness on synthetic and real-world captured point clouds. Yingkui Zhang, Mingqiang Wei, Lei Zhu 0003, Guibao Shen, Fu Lee Wang, Harry Qin, Qiong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | CCM-Net: Contrastive and Consistent Multi-Task Network for Artifact Segmentation and Quality Classification of OCTA ImagesabstractArtifacts are prevalent in Optical Coherence Tomography Angiography (OCTA) images, which probably interfere doctor’s diagnosis and greatly limit its utility. Therefore, it is desirable to segment artifacts and assess quality when using them for diagnosis. In this article, we propose an end-to-end network (named CCM-Net: C ontrastive and C onsistent M ulti-task Network) to jointly address artifact segmentation and quality classification of OCTA images. We first devise multiple Task-Specific Attention Blocks to integrate deep features at different CNN layers for segmenting artifacts and classifying the quality of the input OCTA image. In this way, the weights of different deep features can be automatically learned and are not the same for the two tasks. Moreover, we devise a contrastive loss and a consistency loss to leverage sample relations for further enhancing prediction accuracy. Specifically, given an input OCTA image, we first augment it with a color jitter and select another OCTA image with the same quality classification label. We then design a contrastive loss so that the segmentation results of the input OCTA image are similar to its enhanced OCTA image, while the segmentation results of the two selected OCTA images are not similar. Besides, we devise a consistency loss on the classification results of the three images, because we can find that these images have the same quality classification labels. Experiments on an in-house OCTA dataset (Multi-OCTA) demonstrate that the proposed CCM-Net outperforms state-of-the-art methods. Xiang-Ning Wang, Jixue Tang, Ping Li 0016, Lei Zhu 0003, Harry Qin, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | HADiff: hierarchy aggregated diffusion model for pathology image segmentation
Zhaohu Xing, Feng Gao 0023, Yuandong Tao, Zhenyan Han, Weiming Wang 0002, Lei Zhu 0003 |
Vis. Comput. | 8 |
| 2024 | Learning Diffusion Texture Priors for Image RestorationabstractDiffusion Models have shown remarkable performance in image generation tasks, which are capable of generating diverse and realistic image content. When adopting diffusion models for image restoration, the crucial challenge lies in how to preserve high-level image fidelity in the random-ness diffusion process and generate accurate background structures and realistic texture details. In this paper, we propose a general framework and develop a Diffusion Texture Prior Model (DTPM) for image restoration tasks. DTPM explicitly models high-quality texture details through the diffusion process, rather than global contextual content. In phase one of the training stage, we pretrain DTPM on approximately 55K high-quality image samples, after which we freeze most of its parameters. In phase two, we insert conditional guidance adapters into DTPM and equip it with an initial predictor, thereby facilitating its rapid adaptation to downstream image restoration tasks. Our DTPM could mitigate the randomness of traditional diffusion models by utilizing encapsulated rich and diverse texture knowledge and background structural information provided by the initial predictor during the sampling process. Tian Ye 0001, Sixiang Chen, Wenhao Chai, Zhaohu Xing, Harry Qin, Lei Zhu 0003 |
CVPR | 7 |
| 2024 | Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised LearningabstractFor image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel “Low-Res Leads the Way” (LWay) training framework, merging Supervised Pre-training with Self-supervised Learning to enhance the adaptability of SR models to real-world images. Our approach utilizes a low-resolution (LR) reconstruction network to extract degradation embeddings from LR images, merging them with super-resolved outputs for LR reconstruction. Leveraging unseen LR images for self-supervised learning guides the model to adapt its modeling space to the target domain, facili-tating fine-tuning of SR models without requiring paired high-resolution (HR) images. The integration of Discrete Wavelet Transform (DWT)further refines the focus on high-frequency details. Extensive evaluations show that our method significantly improves the generalization and de-tail restoration capabilities of SR models on unseen real-world datasets, outperforming existing methods. Our training regime is universally compatible, requiring no network architecture modifications, making it a practical solution for real-world SR applications. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Haoze Sun, Xueyi Zou, Zhensong Zhang, Youliang Yan, Lei Zhu 0003 |
CVPR | 9 |
| 2024 | Genuine Knowledge from Practice: Diffusion Test-Time Adaptation for Video Adverse Weather RemovalabstractReal-world vision tasks frequently suffer from the appearance of unexpected adverse weather conditions, including rain, haze, snow, and raindrops. In the last decade, convolutional neural networks and vision transformers have yielded outstanding results in single-weather video removal. However, due to the absence of appropriate adaptation, most of them fail to generalize to other weather conditions. Although ViWS-Net is proposed to remove ad-verse weather conditions in videos with a single set of pre-trained weights, it is seriously blinded by seen weather at train-time and degenerates when coming to unseen weather during test-time. In this work, we introduce test-time adaptation into adverse weather removal in videos, and propose the first framework that integrates test-time adaptation into the iterative diffusion reverse process. Specifically, we devise a diffusion-based network with a novel temporal noise model to efficiently explore frame-correlated information in degraded video clips at training stage. During inference stage, we introduce a proxy task named Diffusion Tubelet Self-Calibration to learn the primer distribution of test video stream and optimize the model by approx-imating the temporal noise model for online adaptation. Experimental results, on benchmark datasets, demonstrate that our Test-Time Adaptation method with Diffusion-based network(Diff- TTA) outperforms state-of-the-art methods in terms of restoring videos degraded by seen weather conditions. Its generalizable capability is validated with unseen weather conditions in synthesized and real-world videos. Angelica I. Avilés-Rivero, Yulun Zhang 0001, Harry Qin, Lei Zhu 0003 |
CVPR | 6 |
| 2024 | Teaching Tailored to Talent: Adverse Weather Restoration via Prompt Pool and Depth-Anything Constraint
Sixiang Chen, Tian Ye 0001, Kai Zhang 0008, Zhaohu Xing, Yunlong Lin, Lei Zhu 0003 |
ECCV (9) | 6 |
| 2024 | Two-Stage Video Shadow Detection via Temporal-Spatial Adaption
Xin Duan, Yu Cao 0019, Lei Zhu 0003, Gang Fu 0003, Xin Wang 0118, Ping Li 0016 |
ECCV (48) | 3 |
| 2024 | OpenIns3D: Snap and Lookup for 3D Open-Vocabulary Instance Segmentation
Zhening Huang, Xiaoyang Wu 0002, Xi Chen 0119, Hengshuang Zhao, Lei Zhu 0003, Joan Lasenby |
ECCV (62) | 5 |
| 2024 | Semi-supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regularization
Angelica I. Avilés-Rivero, Sixiang Chen, Haoyu Chen 0003, Lei Zhu 0003 |
ECCV (10) | 7 |
| 2024 | DragTraffic: Interactive and Controllable Traffic Scene Generation for Autonomous DrivingabstractEvaluating and training autonomous driving systems require diverse and scalable corner cases. However, most existing scene generation methods lack controllability, accuracy, and versatility, resulting in unsatisfactory generation results. Inspired by DragGAN in image generation, we propose DragTraffic, a generalized, interactive, and controllable traffic scene generation framework based on conditional diffusion. DragTraffic enables non-experts to generate a variety of realistic driving scenarios for different types of traffic agents through an adaptive mixture expert architecture. We employ a regression model to provide a general initial solution and a refinement process based on the conditional diffusion model to ensure diversity. User-customized context is introduced through cross-attention to ensure high controllability. Experiments on a real-world driving dataset show that DragTraffic outperforms existing methods in terms of authenticity, diversity, and freedom. Demo videos and code are available at https://chantsss.github.io/Dragtraffic/. Sheng Wang 0017, Fulong Ma, Tianshuai Hu, Qiang Qin, Yongkang Song, Lei Zhu 0003, Junwei Liang 0001 |
IROS | 7 |
| 2024 | MCGMapper: Light-Weight Incremental Structure from Motion and Visual Localization with Planar Markers and Camera GroupsabstractStructure from Motion (SfM) and visual localization in indoor texture-less scenes and industrial scenarios present prevalent yet challenging research topics. Existing SfM methods designed for natural scenes typically yield low accuracy or map-building failures due to insufficient robust feature extraction in such settings. Visual markers, with their artificially designed features, can effectively address these issues. Nonetheless, existing marker-assisted SfM methods encounter problems like slow running speed and difficulties in convergence; and also, they are governed by the strong assumption of unique marker size. In this paper, we propose a novel SfM framework that utilizes planar markers and multiple cameras with known extrinsics to capture the surrounding environment and reconstruct the marker map. In our algorithm, the initial poses of markers and cameras are calculated with Perspective-n-Points (PnP) in the front-end, while bundle adjustment methods customized for markers and camera groups are designed in the back-end to optimize the 6-DOF pose directly. Our algorithm facilitates the reconstruction of large scenes with different marker sizes, and its accuracy and speed of map building are shown to surpass existing methods. Our approach is suitable for a wide range of scenarios, including laboratories, basements, warehouses, and other industrial settings. Furthermore, we incorporate representative scenarios into simulations and also supply our datasets with pose labels to address the scarcity of quantitative ground-truth datasets in this research field. The datasets and source code are available on GitHub1. Yusen Xie, Zhenmin Huang, Kai Chen 0006, Lei Zhu 0003, Jun Ma 0008 |
IROS | 4 |
| 2024 | Diff-VPS: Video Polyp Segmentation via a Multi-task Diffusion Network with Adversarial Temporal Reasoning
Yingling Lu, Zhaohu Xing, Qiong Wang 0001, Lei Zhu 0003 |
MICCAI (6) | 5 |
| 2024 | Advancing UWF-SLO Vessel Segmentation with Source-Free Active Domain Adaptation and a Novel Multi-center Dataset
Hongqiu Wang, Xiangde Luo, Qingqing Tang, Mei Xin, Qiong Wang 0001, Lei Zhu 0003 |
MICCAI (9) | 7 |
| 2024 | Cross-conditioned Diffusion Model for Medical Image to Image Translation
Zhaohu Xing, Sicheng Yang 0001, Sixiang Chen, Tian Ye 0001, Harry Qin, Lei Zhu 0003 |
MICCAI (7) | 7 |
| 2024 | SegMamba: Long-Range Sequential Modeling Mamba for 3D Medical Image Segmentation
Zhaohu Xing, Tian Ye 0001, Guang Liu 0006, Lei Zhu 0003 |
MICCAI (8) | 5 |
| 2024 | LGRNet: Local-Global Reciprocal Network for Uterine Fibroid Segmentation in Ultrasound Videos
Angelica I. Avilés-Rivero, Guang Yang 0006, Harry Qin, Lei Zhu 0003 |
MICCAI (4) | 6 |
| 2024 | A New Perspective to Boost Performance Fairness For Medical Federated Learning
Yunlu Yan, Lei Zhu 0003, Yuexiang Li, Xinxing Xu, Rick Siow Mong Goh, Yong Liu 0026, Salman Khan 0001, Chun-Mei Feng 0001 |
MICCAI (10) | 2 |
| 2024 | Language-Driven Interactive Shadow DetectionabstractTraditional shadow detectors often identify all shadow regions of static images or video sequences. This work presents the Referring Video Shadow Detection (RVSD), which is an innovative task that rejuvenates the classic paradigm by facilitating the segmentation of particular shadows in videos based on descriptive natural language prompts. This novel RVSD not only achieves segmentation of arbitrary shadow areas of interest based on descriptions (flexibility) but also allows users to interact with visual content more directly and naturally by using natural language prompts (interactivity), paving the way for abundant applications ranging from advanced video editing to virtual reality experiences. To pioneer the RVSD research, we curated a well-annotated RVSD dataset, which encompasses 86 videos and a rich set of 15,011 paired textual descriptions with corresponding shadows. To the best of our knowledge, this dataset is the first one for addressing RVSD. Based on this dataset, we propose a Referring Shadow-Track Memory Network (RSM-Net) for addressing the RVSD task. In our RSM-Net, we devise a Twin-Track Synergistic Memory (TSM) to store intra-clip memory features and hierarchical inter-clip memory features, and then pass these memory features into a memory read module to refine features of the current video frame for referring shadow detection. We also develop a Mixed-Prior Shadow Attention (MSA) to utilize physical priors to obtain a coarse shadow map for learning more visual features by weighting it with the input video frame. Experimental results show that our RSM-Net achieves state-of-the-art performance for RVSD with a notable Overall IOU increase of 4.4%. Our code and dataset are available at https://github.com/whq-xxh/RVSD. Hongqiu Wang, Wei Wang 0401, Haipeng Zhou, Shaozhi Wu, Lei Zhu 0003 |
ACM Multimedia | 6 |
| 2024 | RainMamba: Enhanced Locality Learning with State Space Models for Video DerainingabstractThe outdoor vision systems are frequently contaminated by rain streaks and raindrops, which significantly degenerate the performance of visual tasks and multimedia applications. The nature of videos exhibits redundant temporal cues for rain removal with higher stability. Traditional video deraining methods heavily rely on optical flow estimation and kernel-based manners, which have a limited receptive field. Yet, transformer architectures, while enabling long-term dependencies, bring about a significant increase in computational complexity. Recently, the linear-complexity operator of the state space models (SSMs) has contrarily facilitated efficient long-term temporal modeling, which is crucial for rain streaks and raindrops removal in videos. Unexpectedly, its uni-dimensional sequential process on videos destroys the local correlations across the spatio-temporal dimension by distancing adjacent pixels. To address this, we present an improved SSMs-based video deraining network (RainMamba) with a novel Hilbert scanning mechanism to better capture sequence-level local information. We also introduce a difference-guided dynamic contrastive locality learning strategy to enhance the patch-level self-similarity learning ability of the proposed network. Extensive experiments on four synthesized video deraining datasets and real-world rainy videos demonstrate the superiority of our network in the removal of rain streaks and raindrops. Our code and results are available at https://github.com/TonyHongtaoWu/RainMamba. Weiming Wang 0002, Jinni Zhou, Lei Zhu 0003 |
ACM Multimedia | 6 |
| 2024 | Timeline and Boundary Guided Diffusion Network for Video Shadow DetectionabstractVideo Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}. Haipeng Zhou, Hongqiu Wang, Tian Ye 0001, Zhaohu Xing, Jun Ma 0008, Ping Li 0016, Qiong Wang 0001, Lei Zhu 0003 |
ACM Multimedia | 8 |
| 2024 | RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language ModelsabstractNatural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal results. All-in-one models, though capable of handling multiple tasks, typically support only a limited range and often produce overly smooth, low-fidelity outcomes due to their broad data distribution fitting. To address these challenges, we first define a new pipeline for restoring images with multiple degradations, and then introduce RestoreAgent, an intelligent image restoration system leveraging multimodal large language models. RestoreAgent autonomously assesses the type and extent of degradation in input images and performs restoration through (1) determining the appropriate restoration tasks, (2) optimizing the task sequence, (3) selecting the most suitable models, and (4) executing the restoration. Experimental results demonstrate the superior performance of RestoreAgent in handling complex degradation, surpassing human experts. Furthermore, the system’s modular design facilitates the fast integration of new tasks and models. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Sixiang Chen, Tian Ye 0001, Renjing Pei, Kaiwen Zhou 0001, Fenglong Song, Lei Zhu 0003 |
NeurIPS | 10 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 25 |
| 2024 | NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video ReconstructionabstractReconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips. Zixuan Gong, Guangyin Bao, Qi Zhang 0020, Zhongwei Wan, Duoqian Miao 0001, Shoujin Wang, Lei Zhu 0003, Changwei Wang 0001, Rongtao Xu, Liang Hu 0004, Yu Zhang 0133 |
NeurIPS | 7 |
| 2024 | UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New PeaksabstractUltra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality images at multiple resolutions (\textit{e.g.}, 1K, 2K, and 4K) within a single model, while maintaining computational efficiency. UltraPixel leverages semantics-rich representations of lower-resolution images in a later denoising stage to guide the whole generation of highly detailed high-resolution images, significantly reducing complexity. Specifically, we introduce implicit neural representations for continuous upsampling and scale-aware normalization layers adaptable to various resolutions. Notably, both low- and high-resolution processes are performed in the most compact space, sharing the majority of parameters with less than 3$\%$ additional parameters for high-resolution outputs, largely enhancing training and inference efficiency. Our model achieves fast training with reduced data requirements, producing photo-realistic high-resolution images and demonstrating state-of-the-art performance in extensive experiments. Wenbo Li 0002, Haoyu Chen 0003, Renjing Pei, Long Peng 0003, Fenglong Song, Lei Zhu 0003 |
NeurIPS | 9 |
| 2024 | Anchored Supervised Contrastive Learning for Long-Tailed Medical Image Regression
Zhaohu Xing, Lei Zhu 0003 |
PRCV (15) | 4 |
| 2024 | SSM-Net: Semi-supervised multi-task network for joint lesion segmentation and classification from pancreatic EUS images
Jiajia Li 0004, Lei Zhu 0003, Ping Zhang 0016, Ruhan Liu, Bin Sheng 0001 |
Artif. Intell. Medicine | 4 |
| 2024 | Towards High-Resolution Specular Highlight Detection
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Qifeng Lin, Siyuan Fan, Chunxia Xiao |
Int. J. Comput. Vis. | 3 |
| 2024 | ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection
Lei Zhu 0003, Jiaxing Shen, Huazhu Fu, Qing Zhang 0006, Liansheng Wang 0002 |
Int. J. Comput. Vis. | 2 |
| 2024 | FetusMapV2: Enhanced fetal pose estimation in 3D ultrasound
Chaoyu Chen, Xin Yang 0009, Yuhao Huang 0001, Wenlong Shi, Yan Cao 0002, Mingyuan Luo, Xindi Hu, Lei Zhu 0003, Lequan Yu, Kejuan Yue, Yuanji Zhang, Yi Xiong 0001, Dong Ni 0001, Weijun Huang |
Medical Image Anal. | 8 |
| 2024 | Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challengeabstract• This paper investigates the capacity of AI models for airway modelling on national datasets with paired clinical metadata. • We evaluated AI models against unharmonised, noisy, and out-of-distribution data, as well as the prognostication for FLD. • We found a new biomarker for mortality prediction, outperforming existing clinical measurements (FVC% and fibrosis scores). • In-depth analysis of AI models on airway modelling and prognosis, highlighting challenges and future research directions. Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway structures remains prohibitively time-consuming. While significant efforts have been made towards enhancing automatic airway modelling, current public-available datasets predominantly concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023′ (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for mortality prediction, a strong airway-derived biomarker (Hazard ratio>1.5, p < 0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers. Yang Nan 0002, Xiaodan Xing, Zeyu Tang 0001, Federico Felder, Sheng Zhang 0024, Roberta Eufrasia Ledda, Xiaoliu Ding, Feng Shi 0001, Tianyang Sun, Zehong Cao, Yun Gu, Pingyu Wang, Wen Tang 0005, Pengxin Yu, Han Kang, Junqiang Chen, Michail Mamalakis, Francesco Prinzi, Gianluca Carlini, Lisa Cuneo, Abhirup Banerjee, Zhaohu Xing, Lei Zhu 0003, Zacharia Mesbah, Dhruv Jain, Tsiry Mayet, Hongyu Yuan, Qing Lyu 0009, Abdul Qayyum 0002, Moona Mazher, Athol Wells, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 31 |
| 2024 | Cross-modal knowledge distillation for continuous sign language recognition
Liqing Gao, Lianyu Hu 0003, Jichao Feng, Lei Zhu 0003, Liang Wang 0001, Wei Feng 0005 |
Neural Networks | 5 |
| 2024 | Learning Physical-Spatio-Temporal Features for Video Shadow RemovalabstractShadow removal in a single image has received increasing attention in recent years. However, removing shadows over dynamic scenes remains largely under-explored. In this paper, we propose the first data-driven video shadow removal model, termed PSTNet, by exploiting three essential characteristics of video shadows, i.e., physical property, spatio relation, and temporal coherence. Specifically, a dedicated physical branch was established to conduct local illumination estimation, which is more applicable for scenes with complex lighting and textures, and then enhance the physical features via a mask-guided attention strategy. Then, we develop a progressive aggregation module to enhance the spatio and temporal characteristics of features maps, and effectively integrate the three kinds of features. Furthermore, to tackle the lack of datasets of paired shadow videos, we synthesize a dataset (SVSRD-85) with aid of the popular game GTAV by controlling the switch of the shadow renderer. Experiments against 9 state-of-the-art models, including image shadow removers and image/video restoration methods, show that our method improves the best SOTA in terms of RMSE error for the shadow area by 14.7%. In addition, we develop a lightweight model adaptation strategy to make our synthetic-driven model effective in real world scenes. The visual comparison on the public SBU-TimeLapse dataset verifies the generalization ability of our model in real scenes. Zhihao Chen 0004, Yefan Xiao, Lei Zhu 0003, Huazhu Fu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Overcoming Modality Bias in Question-Driven Sign Language Video TranslationabstractQuestion-Driven Sign Language Translation (QSLT) addresses the challenge of translating sign language using pertinent questions in question-answering contexts. However, the pronounced modality complexity between question text and sign video poses a predicament: the model tends to overly depend on questions to generate translations, thereby neglecting the value of visual cues. To tackle this issue, the paper presents a Gloss-Bridged Translator (GBT), which introduces sign gloss as an intermediary conduit to establish semantic connections between questions and videos. By leveraging gloss, visual features are transformed into textual counterparts, mitigating the modality imbalance between these representations. Moreover, a cross-modal contrastive learning strategy is implemented, bolstering the global contextual relevance and local semantic alignment between questions and sign language. The proposed methodology is validated through extensive experiments on the proposed QSL dataset and other public sign language datasets. The results show the efficacy of integrating questions into sign language translation. The GBT yields remarkable improvements over prevailing SLT methods, attesting to its effectiveness and rationale. Our code and dataset is available athttps://github.com/glq-1992/QSL. Liqing Gao, Fan Lyu, Lei Zhu 0003, Junfu Pu, Liang Wang 0001, Wei Feng 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning Motion-Guided Multi-Scale Memory Features for Video Shadow DetectionabstractNatural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods. Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2024 | Hybrid Masked Image Modeling for 3D Medical Image SegmentationabstractMasked image modeling (MIM) with transformer backbones has recently been exploited as a powerful self-supervised pre-training technique. The existing MIM methods adopt the strategy to mask random patches of the image and reconstruct the missing pixels, which only considers semantic information at a lower level, and causes a long pre-training time. This paper presents HybridMIM, a novel hybrid self-supervised learning method based on masked image modeling for 3D medical image segmentation. Specifically, we design a two-level masking hierarchy to specify which and how patches in sub-volumes are masked, effectively providing the constraints of higher level semantic information. Then we learn the semantic information of medical images at three levels, including: 1) partial region prediction to reconstruct key contents of the 3D image, which largely reduces the pre-training time burden (pixel-level); 2) patch-masking perception to learn the spatial relationship between the patches in each sub-volume (region-level); and 3) drop-out-based contrastive learning between samples within a mini-batch, which further improves the generalization ability of the framework (sample-level). The proposed framework is versatile to support both CNN and transformer as encoder backbones, and also enables to pre-train decoders for image segmentation. We conduct comprehensive experiments on five widely-used public medical image segmentation datasets, including BraTS2020, BTCV, MSD Liver, MSD Spleen, and BraTS2023. The experimental results show the clear superiority of HybridMIM against competing supervised methods, masked pre-training approaches, and other self-supervised methods, in terms of quantitative metrics, speed performance and qualitative observations. Zhaohu Xing, Lei Zhu 0003, Lequan Yu, Zhiheng Xing |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Cross-Modal Vertical Federated Learning for MRI ReconstructionabstractFederated learning enables multiple hospitals to cooperatively learn a shared model without privacy disclosure. Existing methods often take a common assumption that the data from different hospitals have the same modalities. However, such a setting is difficult to fully satisfy in practical applications, since the imaging guidelines may be different between hospitals, which makes the number of individuals with the same set of modalities limited. To this end, we formulate this practical-yet-challenging cross-modal vertical federated learning task, in which data from multiple hospitals have different modalities with a small amount of multi-modality data collected from the same individuals. To tackle such a situation, we develop a novel framework, namely Federated Consistent Regularization constrained Feature Disentanglement (Fed-CRFD), for boosting MRI reconstruction by effectively exploring the overlapping samples (i.e., same patients with different modalities at different hospitals) and solving the domain shift problem caused by different modalities. Particularly, our Fed-CRFD involves an intra-client feature disentangle scheme to decouple data into modality-invariant and modality-specific features, where the modality-invariant features are leveraged to mitigate the domain shift problem. In addition, a cross-client latent representation consistency constraint is proposed specifically for the overlapping samples to further align the modality-invariant features extracted from different modalities. Hence, our method can fully exploit the multi-source data from hospitals while alleviating the domain shift problem. Extensive experiments on two typical MRI datasets demonstrate that our network clearly outperforms state-of-the-art MRI reconstruction methods. Yunlu Yan, Hong Wang 0021, Yawen Huang, Nanjun He, Lei Zhu 0003, Yong Xu 0001, Yuexiang Li, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | DSMT-Net: Dual Self-Supervised Multi-Operator Transformation for Multi-Source Endoscopic Ultrasound DiagnosisabstractPancreatic cancer has the worst prognosis of all cancers. The clinical application of endoscopic ultrasound (EUS) for the assessment of pancreatic cancer risk and of deep learning for the classification of EUS images have been hindered by inter-grader variability and labeling capability. One of the key reasons for these difficulties is that EUS images are obtained from multiple sources with varying resolutions, effective regions, and interference signals, making the distribution of the data highly variable and negatively impacting the performance of deep learning models. Additionally, manual labeling of images is time-consuming and requires significant effort, leading to the desire to effectively utilize a large amount of unlabeled data for network training. To address these challenges, this study proposes the Dual Self-supervised Multi-Operator Transformation Network (DSMT-Net) for multi-source EUS diagnosis. The DSMT-Net includes a multi-operator transformation approach to standardize the extraction of regions of interest in EUS images and eliminate irrelevant pixels. Furthermore, a transformer-based dual self-supervised network is designed to integrate unlabeled EUS images for pre-training the representation model, which can be transferred to supervised tasks such as classification, detection, and segmentation. A large-scale EUS-based pancreas image dataset (LEPset) has been collected, including 3,500 pathologically proven labeled EUS images (from pancreatic and non-pancreatic cancers) and 8,000 unlabeled EUS images for model development. The self-supervised method has also been applied to breast cancer diagnosis and was compared to state-of-the-art deep learning models on both datasets. The results demonstrate that the DSMT-Net significantly improves the accuracy of pancreatic and breast cancer diagnosis. Jiajia Li 0004, Lei Zhu 0003, Ruhan Liu, Dinggang Shen, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Video-Instrument Synergistic Network for Referring Video Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation is fundamentally important for facilitating cognitive intelligence in robot-assisted surgery. Although existing methods have achieved accurate instrument segmentation results, they simultaneously generate segmentation masks of all instruments, which lack the capability to specify a target object and allow an interactive experience. This paper focuses on a novel and essential task in robotic surgery, i.e., Referring Surgical Video Instrument Segmentation (RSVIS), which aims to automatically identify and segment the target surgical instruments from each video frame, referred by a given language expression. This interactive feature offers enhanced user engagement and customized experiences, greatly benefiting the development of the next generation of surgical education systems. To achieve this, this paper constructs two surgery video datasets to promote the RSVIS research. Then, we devise a novel Video-Instrument Synergistic Network (VIS-Net) to learn both video-level and instrument-level knowledge to boost performance, while previous work only utilized video-level information. Meanwhile, we design a Graph-based Relation-aware Module (GRM) to model the correlation between multi-modal information (i.e., textual description and video frame) to facilitate the extraction of instrument-level information. Extensive experimental results on two RSVIS datasets exhibit that the VIS-Net can significantly outperform existing state-of-the-art referring segmentation methods. We will release our code and dataset for future research (https://github.com/whq-xxh/RSVIS). Hongqiu Wang, Guang Yang 0006, Harry Qin, Yike Guo, Yueming Jin, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image SegmentationabstractThis paper studies a practical Source-free unsupervised domain adaptation (SFUDA) problem, which transfers knowledge of source-trained models to the target domain, without accessing the source data. It has received increasing attention in recent years, while the prior arts focus on designing adaptation strategies, ignoring that different target samples exhibit different transfer abilities on the source model. Additionally, we observe pixel-wise class prediction is typically accompanied by ambiguity issue, i.e., prediction errors often occur between several confusing classes. In this study, we propose a dual-branch collaborative learning framework that aims to achieve reliable knowledge transfer from important samples to the rest by fully mining confident prototypes in the target data. Concretely, we first partition the target data into confident samples and uncertain samples via a new class-ranking reliability score and then utilize the latent features from the confident branch as guidance to promote the learning of the uncertain branch. For ambiguity issue, we propose a feature relabelling module, which exploits reliable prototypes in the mini-batch as well as in the target data to refine labels of uncertain features. We further deploy the proposed framework to commonly used CNN and state-of-the-art Transformer architectures and reveal the potential to promote the generalization ability of backbone models. Experimental results on both natural and medical benchmark datasets verify that our proposed approach exceeds state-of-the-art SFUDA methods with large margins, and achieves comparable performance to existing UDA methods. Yuntong Tian, Huazhu Fu, Lei Zhu 0003, Lequan Yu |
IEEE Trans. Multim. | 4 |
| 2024 | Unsupervised Fusion Feature Matching for Data Bias in Uncertainty Active LearningabstractActive learning (AL) aims to sample the most valuable data for model improvement from the unlabeled pool. Traditional works, especially uncertainty-based methods, are prone to suffer from a data bias issue, which means that selected data cannot cover the entire unlabeled pool well. Although there have been lots of literature works focusing on this issue recently, they mainly benefit from the huge additional training costs and the artificially designed complex loss. The latter causes these methods to be redesigned when facing new models or tasks, which is very time-consuming and laborious. This article proposes a feature-matching-based uncertainty that resamples selected uncertainty data by feature matching, thus removing similar data to alleviate the data bias issue. To ensure that our proposed method does not introduce a lot of additional costs, we specially design a unsupervised fusion feature matching (UFFM), which does not require any training in our novel AL framework. Besides, we also redesign several classic uncertainty methods to be applied to more complex visual tasks. We conduct rigorous experiments on lots of standard benchmark datasets to validate our work. The experimental results show that our UFFM is better than the similar unsupervised feature matching technologies, and our proposed uncertainty calculation method outperforms random sampling, classic uncertainty approaches, and recent state-of-the-art (SOTA) uncertainty approaches. Shuzhou Sun, Xiao Lin 0012, Ping Li 0016, Lei Zhu 0003, C. L. Philip Chen, Bin Sheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video CaptioningabstractNon-autoregressive video captioning methods generate visual words in parallel but often overlook semantic correlations among them, especially regarding verbs, leading to lower caption quality. To address this, we integrate action information of highlighted objects to enhance semantic connections among visual words. Our proposed Action-aware Language Skeleton Optimization Network (ALSO-Net) tackles the challenge of extracting action information across frames, improving understanding of complex context-dependent video actions and reducing sentence inconsistencies. ALSO-Net incorporates a linguistic skeleton tag generator to refine semantic correlations and a video action predictor to enhance verb prediction accuracy in video captions. We also address issues of unsatisfactory caption length and quality by jointly optimizing different levels of motion prediction loss. Experimental evaluation on prominent video captioning datasets demonstrates that ALSO-Net outperforms baseline methods by a significant margin and achieves competitive performance compared to state-of-the-art autoregressive methods with smaller model complexity and faster inference time. Shuqin Chen, Xian Zhong, Lei Zhu 0003, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Reference-Based Line Drawing Colorization Through Diffusion Model
Jiaze He, Ziruo Li, Ping Li 0016, Lei Zhu 0003, Bin Sheng 0001, Subrota K. Mondal |
CGI | 6 |
| 2023 | Masked Image Training for Generalizable Deep Image DenoisingabstractWhen capturing and storing images, devices inevitably introduce noise. Reducing this noise is a critical task called image denoising. Deep learning has become the de facto method for image denoising, especially with the emergence of Transformer-based models that have achieved notable state-of-the-art results on various image tasks. However, deep learning-based methods often suffer from a lack of generalization ability. For example, deep models trained on Gaussian noise may perform poorly when tested on other noise distributions. To address this issue, we present a novel approach to enhance the generalization performance of denoising networks, known as masked training. Our method involves masking random pixels of the input image and reconstructing the missing information during training. We also mask out the features in the self-attention layers to avoid the impact of training-testing inconsistency. Our approach exhibits better generalization ability than other deep learning models and is directly applicable to real-world scenarios. Additionally, our interpretability analysis demonstrates the superiority of our method. Haoyu Chen 0003, Jinjin Gu, Yihao Liu 0001, Salma Abdel Magid, Chao Dong 0005, Qiong Wang 0001, Hanspeter Pfister, Lei Zhu 0003 |
CVPR | 8 |
| 2023 | SCOTCH and SODA: A Transformer Video Shadow Detection FrameworkabstractShadows in videos are difficult to detect because of the large shadow deformation between frames. In this work, we argue that accounting for shadow deformation is essential when designing a video shadow detection method. To this end, we introduce the shadow deformation attention trajectory (SODA), a new type of video self-attention module, specially designed to handle the large shadow deformations in videos. Moreover, we present a new shadow contrastive learning mechanism (SCOTCH) which aims at guiding the network to learn a unified shadow representation from massive positive shadow pairs across different videos. We demonstrate empirically the effectiveness of our two contributions in an ablation study. Furthermore, we show that SCOTCH and SODA significantly outperforms existing techniques for video shadow detection. Code is available at the project page: https://lihaoliu-cambridge.github.io/scotch_and_soda/ Jean Prost, Lei Zhu 0003, Nicolas Papadakis, Pietro Liò, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
CVPR | 3 |
| 2023 | Video Dehazing via a Multi-Range Temporal Alignment Network with Physical PriorabstractVideo dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggregate temporal information. Specifically, we design a memory-based physical prior guidance module to encode the prior-related features into long-range memory. Besides, we formulate a multi-range scene radiance recovery module to capture space-time dependencies in multiple space-time ranges, which helps to effectively aggregate temporal information from adjacent frames. Moreover, we construct the first large-scale outdoor video dehazing benchmark dataset, which contains videos in various real-world scenarios. Experimental results on both synthetic and real conditions show the superiority of our proposed method. Xiaowei Hu 0001, Lei Zhu 0003, Qi Dou 0001, Jifeng Dai, Yu Qiao 0001, Pheng-Ann Heng |
CVPR | 3 |
| 2023 | Snow Removal in Video: A New Dataset and A Novel MethodabstractSnowfall is a common weather phenomenon that can severely affect computer vision tasks by obscuring objects and scenes. However, existing deep learning-based snow removal methods are designed for single images only. In this paper, we target a more complex task - video snow removal, which aims to restore the clear video from the snowy video. To facilitate this task, we propose the first high-quality video dataset, which simulates realistic physical characteristics of snow and haze using a rendering engine and augmentation techniques. We also develop a deep learning framework for video snow removal. Specifically, we propose a snow-query temporal aggregation module and a snow-aware contrastive learning loss function. The module aggregates features between video frames and removes snow effectively, while the loss function helps identify and eliminate snow features. We conduct extensive experiments and demonstrate that our proposed dataset is more realistic than previous datasets, and the models trained on it achieve better performance in real-world snowing images. Our proposed method outperforms state-of-the-art video and image-based methods on both synthetic and real snowy videos. Haoyu Chen 0003, Jinjin Gu, Xuequan Lu, Haoming Cai, Lei Zhu 0003 |
ICCV | 7 |
| 2023 | Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain StreaksabstractIn the real world, image degradations caused by rain often exhibit a combination of rain streaks and raindrops, thereby increasing the challenges of recovering the underlying clean image. Note that the rain streaks and raindrops have diverse shapes, sizes, and locations in the captured image, and thus modeling the correlation relationship between irregular degradations caused by rain artifacts is a necessary prerequisite for image deraining. This paper aims to present an efficient and flexible mechanism to learn and model degradation relationships in a global view, thereby achieving a unified removal of intricate rain scenes. To do so, we propose a Sparse Sampling Transformer based on Uncertainty-Driven Ranking, dubbed UDR-S2Former. Compared to previous methods, our UDR-S2Former has three merits. First, it can adaptively sample relevant image degradation information to model underlying degradation relationships. Second, explicit application of the uncertainty-driven ranking strategy can facilitate the network to attend to degradation features and understand the reconstruction process. Finally, experimental results show that our UDR-S2Former clearly outperforms state-of-the-art methods for all benchmarks. Sixiang Chen, Tian Ye 0001, Jinbin Bai, Erkang Chen, Lei Zhu 0003 |
ICCV | 6 |
| 2023 | Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic DataabstractThis paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address them. Specifically, given an input image, we first decompose it into the albedo, shading, and specular residue components to estimate a coarse specular-free image. Then, we further refine the coarse result to alleviate its visual artifacts such as color distortion. Finally, we adjust the tone of the refined result to match the tone of the input as closely as possible. In addition, to facilitate network training and quantitative evaluation, we present a large-scale synthetic dataset of object-level images, covering diverse objects and illumination conditions. Extensive experiments illustrate that our network is able to generalize well to unseen real object-level images, and even produce good results for scene-level images with multiple background objects and complex lighting. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Chunxia Xiao, Ping Li 0016 |
ICCV | 3 |
| 2023 | Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial BackpropagationabstractAlthough convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to restore weather videos due to the absence of temporal information. Furthermore, existing methods for removing adverse weather conditions (e.g., rain, fog, and snow) from videos can only handle one type of adverse weather. In this work, we propose the first framework for restoring videos from all adverse weather conditions by developing a video adverse-weather-component suppression network (ViWS-Net). To achieve this, we first devise a weather-agnostic video transformer encoder with multiple transformer stages. Moreover, we design a long short-term temporal modeling mechanism for weather messenger to early fuse input adjacent video frames and learn weather-specific information. We further introduce a weather discriminator with gradient reversion, to maintain the weather-invariant common information and suppress the weather-specific information in pixel features, by adversarially predicting weather types. Finally, we develop a messenger-driven video transformer decoder to retrieve the residual weather-specific feature, which is spatiotemporally aggregated with hierarchical pixel features and refined to predict the clean target frame of input videos. Experimental results, on benchmark datasets and real-world weather videos, demonstrate that our ViWS-Net outperforms current state-of-the-art methods in terms of restoring videos degraded by any weather condition. Angelica I. Avilés-Rivero, Huazhu Fu, Weiming Wang 0002, Lei Zhu 0003 |
ICCV | 6 |
| 2023 | Dynamic Interactive Relation Capturing via Scene Graph Learning for Robotic Surgical Report GenerationabstractFor robot-assisted surgery, an accurate surgical report reflects clinical operations during surgery and helps document entry tasks, post-operative analysis and follow-up treatment. It is a challenging task due to many complex and diverse interactions between instruments and tissues in the surgical scene. Although existing surgical report generation methods based on deep learning have achieved large success, they often ignore the interactive relation between tissues and instrumental tools, thereby degrading the report generation performance. This paper presents a neural network to boost surgical report generation by explicitly exploring the interactive relation between tissues and surgical instruments. To do so, we first devise a relational exploration (RE) module to model the interactive relation via graph learning, and an interaction perception (IP) module to assist the graph learning in RE module. In our IP module, we first devise a node tracking system to identify and append missing graph nodes of the current video frame for constructing graphs at RE module. Moreover, the IP module generates a global attention model to indicate the existence of the interactive relation on the whole scene of the current video frame to eliminate the graph learning at the current video frame. Furthermore, our IP module predicts a local attention model to more accurately identify the interaction relation of each graph node for assisting the graph updating at the RE module. After that, we concatenate features of all graph nodes of RE module and pass concatenated features into a transformer for generating the output surgical report. We validate the effectiveness of our method on a widely-used robotic surgery benchmark dataset, and experimental results show that our network can significantly outperform existing state-of-the-art surgical report generation methods (e.g., 7.48% and 5.43% higher for BLEU-1 and ROUGE). Hongqiu Wang, Yueming Jin, Lei Zhu 0003 |
ICRA | 3 |
| 2023 | Shifting More Attention to Breast Lesion Segmentation in Ultrasound Videos
Qian Dai, Lei Zhu 0003, Huazhu Fu, Qiong Wang 0001, Wenhao Rao, Liansheng Wang 0002 |
MICCAI (3) | 3 |
| 2023 | DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
Huazhu Fu, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Lei Zhu 0003 |
MICCAI (6) | 5 |
| 2023 | Uncertainty-Driven Dynamic Degradation Perceiving and Background Modeling for Efficient Single Image DesnowingabstractSingle-image snow removal aims to restore clean images from heterogeneous and irregular snow degradations. Recent methods utilize neural networks to remove various degradations directly. However, these approaches suffer from the limited ability to flexibly perceive complicated snow degradation patterns and insufficient representation of background structure information. To further improve the performance and generalization ability of snow removal, this paper aims to develop a novel and efficient paradigm from the perspective of degradation perceiving and background modeling. Sixiang Chen, Tian Ye 0001, Chenghao Xue, Haoyu Chen 0003, Yun Liu 0002, Erkang Chen, Lei Zhu 0003 |
ACM Multimedia | 7 |
| 2023 | Mask-Guided Progressive Network for Joint Raindrop and Rain Streak Removal in VideosabstractVideos captured in rainy weather are unavoidably corrupted by both rain streaks and raindrops in driving scenarios, and it is desirable and challenging to recover background details obscured by rain streaks and raindrops. However, existing video rain removal methods often address either video rain streak removal or video raindrop removal, thereby suffer from degraded performance when deal with both simultaneously. The bottleneck is a lack of a video dataset, where each video frame contains both rain streaks and raindrops. To address this issue, we in this work generate a synthesized dataset, namely VRDS, with 102 rainy videos from diverse scenarios, and each video frame has the corresponding rain streak map, raindrop mask, and the underlying rain-free clean image (ground truth). Moreover, we devise a mask-guided progressive video deraining network (ViMP-Net) to remove both rain streaks and raindrops of each video frame. Specifically, we develop an intensity-guided alignment block to predict the rain streak intensity map and remove the rain streaks of the input rainy video at the first stage. Then, we predict a raindrop mask and pass it into a devised mask-guided dual transformer block to learn inter-frame and intra-frame transformer features, which are then fed into a decoder for further eliminating raindrops. Experimental results demonstrate that our ViMP-Net outperforms state-of-the-art methods on our synthetic dataset and real-world rainy videos. Our code is available at https://github.com/TonyHongtaoWu/ViMP-Net. Haoyu Chen 0003, Lei Zhu 0003 |
ACM Multimedia | 5 |
| 2023 | Learning to Remove Shadows from a Single Image
Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | MNGNAS: Distilling Adaptive Combination of Multiple Searched Networks for One-Shot Neural Architecture SearchabstractRecently neural architecture (NAS) search has attracted great interest in academia and industry. It remains a challenging problem due to the huge search space and computational costs. Recent studies in NAS mainly focused on the usage of weight sharing to train a SuperNet once. However, the corresponding branch of each subnetwork is not guaranteed to be fully trained. It may not only incur huge computation costs but also affect the architecture ranking in the retraining procedure. We propose a multi-teacher-guided NAS, which proposes to use the adaptive ensemble and perturbation-aware knowledge distillation algorithm in the one-shot-based NAS algorithm. The optimization method aiming to find the optimal descent directions is used to obtain adaptive coefficients for the feature maps of the combined teacher model. Besides, we propose a specific knowledge distillation process for optimal architectures and perturbed ones in each searching process to learn better feature maps for later distillation procedures. Comprehensive experiments verify our approach is flexible and effective. We show improvement in precision and search efficiency in the standard recognition dataset. We also show improvement in correlation between the accuracy of the search algorithm and true accuracy by NAS benchmark datasets. Guhao Qiu, Ping Li 0016, Lei Zhu 0003, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Deep Texture-Aware Features for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin. Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Dual Multiscale Mean Teacher Network for Semi-Supervised Infection Segmentation in Chest CT Volume for COVID-19abstractAutomated detecting lung infections from computed tomography (CT) data plays an important role for combating coronavirus 2019 (COVID-19). However, there are still some challenges for developing AI system: 1) most current COVID-19 infection segmentation methods mainly relied on 2-D CT images, which lack 3-D sequential constraint; 2) existing 3-D CT segmentation methods focus on single-scale representations, which do not achieve the multiple level receptive field sizes on 3-D volume; and 3) the emergent breaking out of COVID-19 makes it hard to annotate sufficient CT volumes for training deep model. To address these issues, we first build a multiple dimensional-attention convolutional neural network (MDA-CNN) to aggregate multiscale information along different dimension of input feature maps and impose supervision on multiple predictions from different convolutional neural networks (CNNs) layers. Second, we assign this MDA-CNN as a basic network into a novel dual multiscale mean teacher network (DM [Formula: see text]-Net) for semi-supervised COVID-19 lung infection segmentation on CT volumes by leveraging unlabeled data and exploring the multiscale information. Our DM [Formula: see text]-Net encourages multiple predictions at different CNN layers from the student and teacher networks to be consistent for computing a multiscale consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from multiple predictions of MDA-CNN. Third, we collect two COVID-19 segmentation datasets to evaluate our method. The experimental results show that our network consistently outperforms the compared state-of-the-art methods. Liansheng Wang 0002, Jiacheng Wang 0002, Lei Zhu 0003, Huazhu Fu, Ping Li 0016, Gary Cheng 0001, Shuo Li 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 3 |
| 2023 | Uncertainty-Aware Multi-Dimensional Mutual Learning for Brain and Brain Tumor SegmentationabstractExisting segmentation methods for brain MRI data usually leverage 3D CNNs on 3D volumes or employ 2D CNNs on 2D image slices. We discovered that while volume-based approaches well respect spatial relationships across slices, slice-based methods typically excel at capturing fine local features. Furthermore, there is a wealth of complementary information between their segmentation predictions. Inspired by this observation, we develop an Uncertainty-aware Multi-dimensional Mutual learning framework to learn different dimensional networks simultaneously, each of which provides useful soft labels as supervision to the others, thus effectively improving the generalization ability. Specifically, our framework builds upon a 2D-CNN, a 2.5D-CNN, and a 3D-CNN, while an uncertainty gating mechanism is leveraged to facilitate the selection of qualified soft labels, so as to ensure the reliability of shared information. The proposed method is a general framework and can be applied to varying backbones. The experimental results on three datasets demonstrate that our method can significantly enhance the performance of the backbone network by notable margins, achieving a Dice metric improvement of 2.8% on MeniSeg, 1.4% on IBSR, and 1.3% on BraTS2020. Junting Zhao, Zhaohu Xing, Zhihao Chen 0004, Tong Han, Huazhu Fu, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Multi-Modal Learning for Predicting the Genotype of GliomaabstractThe isocitrate dehydrogenase (IDH) gene mutation is an essential biomarker for the diagnosis and prognosis of glioma. It is promising to better predict glioma genotype by integrating focal tumor image and geometric features with brain network features derived from MRI. Convolutional neural networks show reasonable performance in predicting IDH mutation, which, however, cannot learn from non-Euclidean data, e.g., geometric and network data. In this study, we propose a multi-modal learning framework using three separate encoders to extract features of focal tumor image, tumor geometrics and global brain networks. To mitigate the limited availability of diffusion MRI, we develop a self-supervised approach to generate brain networks from anatomical multi-sequence MRI. Moreover, to extract tumor-related features from the brain network, we design a hierarchical attention module for the brain network encoder. Further, we design a bi-level multi-modal contrastive loss to align the multi-modal features and tackle the domain gap at the focal tumor and global brain. Finally, we propose a weighted population graph to integrate the multi-modal features for genotype prediction. Experimental results on the testing set show that the proposed model outperforms the baseline deep learning models. The ablation experiments validate the performance of different components of the framework. The visualized interpretation corresponds to clinical knowledge with further validation. In conclusion, the proposed learning framework provides a novel approach for predicting the genotype of glioma. Yiran Wei 0002, Xi Chen 0042, Lei Zhu 0003, Lipei Zhang, Carola-Bibiane Schönlieb, Stephen J. Price, Chao Li 0031 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | S $^3$ Net: Self-Supervised Self-Ensembling Network for Semi-Supervised RGB-D Salient Object DetectionabstractRGB-D salient object detection aims to detect visually distinctive objects or regions from a pair of the RGB image and the depth image. State-of-the-art RGB-D saliency detectors are mainly based on convolutional neural networks but almost suffer from an intrinsic limitation relying on the labeled data, thus degrading detection accuracy in complex cases. In this work, we present a self-supervised self-ensembling network (S$^3$Net) for semi-supervised RGB-D salient object detection by leveraging the unlabeled data and exploring a self-supervised learning mechanism. To be specific, we first build a self-guided convolutional neural network (SG-CNN) as a baseline model by developing a series of three-layer cross-model feature fusion (TCF) modules to leverage complementary information among depth and RGB modalities and formulating an auxiliary task that predicts a self-supervised image rotation angle. After that, to further explore the knowledge from unlabeled data, we assign SG-CNN to a student network and a teacher network, and encourage the saliency predictions and self-supervised rotation predictions from these two networks to be consistent on the unlabeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our network quantitatively and qualitatively outperforms the state-of-the-art methods. Lei Zhu 0003, Xiaoqiang Wang 0007, Ping Li 0016, Xin Yang 0011, Qing Zhang 0006, Weiming Wang 0002, Carola-Bibiane Schönlieb, C. L. Philip Chen |
IEEE Trans. Multim. | 1 |
| 2023 | Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005 |
Vis. Comput. | 4 |
| 2022 | ICBNet: Iterative Context-Boundary Feedback Network for Polyp SegmentationabstractAccurate polyp segmentation from colonoscopy images, which is critical to automatic colorectal cancer diagnosis, attracts increasing attentions in recent years. Most existing deep learning-based methods adopt the one-stage processing pipeline, by usually fusing features from different levels or employing boundary-related attention. In this paper, we propose an novel Iterative Context-Boundary feedback Network, namely ICBNet, for robust and accurate polyp segmentation. By mimicking the “from-Preliminary-to-Refined” working paradigm of doctors, ICBNet adopts an iterative feedback learning strategy. Differently from other feedback methods which only use the prediction mask as a guide for foreground features, ICBNet refines encoder features with contextual and boundary-aware details from the preliminary segmentation and boundary predictions, and conducts such strategy in an iterative manner to achieve progressive improvement. Moreover, a dual-branch iterative feedback unit (IFU) is developed to enhance features under the guidance of segmentation and boundary predictions to enable the iterative learning. Extensive experiments on five widely-used polyp segmentation datasets demonstrate that the proposed ICBNet can utilize progressive refinement to effectively address the challenges of large appearance variations and obscure boundaries, and hence achieves more accurate and robust results against the state-of-the-arts methods. Yefan Xiao, Zhihao Chen 0004, Lequan Yu, Lei Zhu 0003 |
BIBM | 5 |
| 2022 | RSCFed: Random Sampling Consensus Federated Semi-supervised LearningabstractFederated semi-supervised learning (FSSL) aims to derive a global model by training fully-labeled and fully-unlabeled clients or training partially labeled clients. The existing approaches work well when local clients have in-dependent and identically distributed (IID) data but fail to generalize to a more practical FSSL setting, i.e., Non-IID setting. In this paper, we present a Random Sampling Consensus Federated learning, namely RSCFed, by con-sidering the uneven reliability among models from fully-labeled clients, fully-unlabeled clients or partially labeled clients. Our key motivation is that given models with large deviations from either labeled clients or unlabeled clients, the consensus could be reached by performing random sub-sampling over clients. To achieve it, instead of di-rectly aggregating local models, we first distill several sub-consensus models by random sub-sampling over clients and then aggregating the sub-consensus models to the global model. To enhance the robustness of sub-consensus models, we also develop a novel distance-reweighted model aggre-gation method. Experimental results show that our method outperforms state-of-the-art methods on three benchmarked datasets, including both natural and medical images. The code is available at https://github.com/XMed-Lab/RSCFed. Xiaoxiao Liang, Yiqun Lin, Huazhu Fu, Lei Zhu 0003, Xiaomeng Li 0001 |
CVPR | 4 |
| 2022 | Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior
Lei Zhu 0003, Huazhu Fu, Harry Qin, Carola-Bibiane Schönlieb, Wei Feng 0005, Song Wang 0002 |
ECCV (19) | 2 |
| 2022 | Rethinking Breast Lesion Segmentation in Ultrasound: A New Video Dataset and A Baseline Network
Qingqing Zheng, Mingshuang Li, Qiong Wang 0001, Lei Zhu 0003 |
MICCAI (4) | 7 |
| 2022 | A New Dataset and a Baseline Model for Breast Lesion Detection in Ultrasound Videos
Lei Zhu 0003, Huazhu Fu, Harry Qin, Liansheng Wang 0002 |
MICCAI (3) | 3 |
| 2022 | Joint Prediction of Meningioma Grade and Brain Invasion via Task-Aware Contrastive Learning
Tianling Liu, Wennan Liu, Lequan Yu, Tong Han, Lei Zhu 0003 |
MICCAI (3) | 6 |
| 2022 | NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation
Zhaohu Xing, Lequan Yu, Tong Han, Lei Zhu 0003 |
MICCAI (5) | 5 |
| 2022 | Reinforcement Learning Driven Intra-modal and Inter-modal Representation Learning for 3D Medical Image Classification
Zhonghang Zhu, Liansheng Wang 0002, Baptiste Magnier, Lei Zhu 0003, Lequan Yu |
MICCAI (3) | 4 |
| 2022 | Phase-based Memory Network for Video DehazingabstractVideo dehazing using deep-learning based methods has just received increasing attention in recent years. However, most existing methods tackle temporal consistency in the color domain only, which are less sensitive to small and imperceptible motions in a video, due to fog's drift and diffusion. In this work, we investigate in the frequency domain, which enables us to capture small motions effectively, and find that the phase component contains more semantic structures yet less haze information than the amplitude component of the hazy image. Based on these observations, we propose a novel phase-based memory network (PM-Net) to integrate the phase and color memory information for boosting video dehazing. Apart from the color memory from consecutive video frames, our PM-Net constructs a phase memory, which stores phase features of past video frames, and devise a cross-modal memory read (CMR) module, which fully leverages features from the color memory and the phase memory to boost features extracted from the current video frame for dehazing. Experimental results on the benchmark dataset of real hazy videos and a newly collected dataset of synthetic videos, show that the proposed PM-Net clearly outperforms the state-of-the-art image and video dehazing methods. Code is available at https://github.com/liuye123321/PM-Net. Huazhu Fu, Harry Qin, Lei Zhu 0003 |
ACM Multimedia | 5 |
| 2022 | Video Instance Lane Detection via Deep Temporal and Geometry Consistency ConstraintsabstractVideo instance lane detection is one of the most important tasks in autonomous driving.Due to the very sparse region and weak context in lane annotations, accurately detecting instance-level lanes in real-world traffic scenarios is challenging, especially for scenes with occlusion, bad weather conditions, dim or dazzling lights.Current methods mainly address this problem by integrating features of adjacent video frames to simply encourage temporal constancy for image-level lane detectors. However, most of them ignore lane shape constraint of adjacent frames and geometry consistency of individual lanes, thereby harming the performance of video instance lane detection. In this paper, we propose TGC-Net via temporal and geometry consistency constraints for reliable video instance lane detection. Specifically, we devise a temporal recurrent feature-shift aggregation module (T-RESA) to learn spatio-temporal lane features along horizontal, vertical, and temporal directions of the feature tensor. We further impose temporal consistency constraint by encouraging spatial distribution consistency among the lane features of adjacent frames. Besides, we devise two effective geometry constraints to ensure the integrity and continuity of lane predictions by leveraging pairwise point affinity loss and vanishing point guided geometric context, respectively. Extensive experiments on public benchmark dataset show that our TGC-Net quantitatively and qualitatively outperforms state-of-the-art video instance lane detectors and video object segmentation competitors. Our code and our results have been released at https://github.com/wmq12345/TGC-Net. Yujun Zhang 0002, Wei Feng 0005, Lei Zhu 0003, Song Wang 0002 |
ACM Multimedia | 4 |
| 2022 | Learning Multi-Scale Deep Image Prior for High-Quality Unsupervised Image DenoisingabstractAbstract Recent methods on image denoising have achieved remarkable progress, benefiting mostly from supervised learning on massive noisy/clean image pairs and unsupervised learning on external noisy images. However, due to the domain gap between the training and testing images, these methods typically have limited applicability on unseen images. Although several attempts have been made to avoid the domain gap issue by learning denoising from singe noisy image itself, they are less effective in handling real‐world noise because of assuming the noise corruptions are independent and zero mean. In this paper, we go step further beyond prior work by presenting a novel unsupervised image denoising framework trained from single noisy image without making any explicit assumptions on the noise statistics. Our approach is built upon the deep image prior (DIP), which enables diverse image restoration tasks. However, as is, the denoising performance of DIP will significantly deteriorate on nonzero‐mean noise and is sensitive to the number of iterations. To overcome this problem, we propose to utilize multi‐scale deep image prior by imposing DIP across different image scales under the constraint of a scale consistency. Experiments on synthetic and real datasets demonstrate that our method performs favorably against the state‐of‐the‐art methods for image denoising. Hao Jiang 0057, Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Wei-Shi Zheng 0001 |
Comput. Graph. Forum | 4 |
| 2022 | Unsupervised Intrinsic Image Decomposition Using Internal Self-Similarity CuesabstractRecent learning-based intrinsic image decomposition methods have achieved remarkable progress. However, they usually require massive ground truth intrinsic images for supervised learning, which limits their applicability on real-world images since obtaining ground truth intrinsic decomposition for natural images is very challenging. In this paper, we present an unsupervised framework that is able to learn the decomposition effectively from a single natural image by training solely with the image itself. Our approach is built upon the observations that the reflectance of a natural image typically has high internal self-similarity of patches, and a convolutional generation network tends to boost the self-similarity of an image when trained for image reconstruction. Based on the observations, an unsupervised intrinsic decomposition network (UIDNet) consisting of two fully convolutional encoder-decoder sub-networks, i.e., reflectance prediction network (RPN) and shading prediction network (SPN), is devised to decompose an image into reflectance and shading by promoting the internal self-similarity of the reflectance component, in a way that jointly trains RPN and SPN to reproduce the given image. A novel loss function is also designed to make effective the training for intrinsic decomposition. Experimental results on three benchmark real-world datasets demonstrate the superiority of the proposed method. Qing Zhang 0006, Lei Zhu 0003, Wei Sun 0007, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Boosting RGB-D Saliency Detection by Leveraging Unlabeled RGB ImagesabstractTraining deep models for RGB-D salient object detection (SOD) often requires a large number of labeled RGB-D images. However, RGB-D data is not easily acquired, which limits the development of RGB-D SOD techniques. To alleviate this issue, we present a Dual-Semi RGB-D Salient Object Detection Network (DS-Net) to leverage unlabeled RGB images for boosting RGB-D saliency detection. We first devise a depth decoupling convolutional neural network (DDCNN), which contains a depth estimation branch and a saliency detection branch. The depth estimation branch is trained with RGB-D images and then used to estimate the pseudo depth maps for all unlabeled RGB images to form the paired data. The saliency detection branch is used to fuse the RGB feature and depth feature to predict the RGB-D saliency. Then, the whole DDCNN is assigned as the backbone in a teacher-student framework for semi-supervised learning. Moreover, we also introduce a consistency loss on the intermediate attention and saliency maps for the unlabeled data, as well as a supervised depth and saliency loss for labeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our DDCNN outperforms state-of-the-art methods both quantitatively and qualitatively. We also demonstrate that our semi-supervised DS-Net can further improve the performance, even when using an RGB image with the pseudo depth map. Xiaoqiang Wang 0007, Lei Zhu 0003, Siliang Tang, Huazhu Fu, Ping Li 0016, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 2 |
| 2022 | A Blind Color Separation Model for Faithful Palette-Based Image RecoloringabstractPalette-based image recoloring provides a simple yet effective way for color adjustment, which allows users to interactively manipulate the color of an image by editing a compact color palette. While remarkable progress has been made by previous methods, they have the common limitations that may produce unfaithful image recoloring results i.e., the obtained result does not respond faithfully to the palette adjustment, and tend to induce visual artifacts such as color bleeding and distortion. To address these limitations, we in this paper present a novel color separation model for palette-based recoloring. Akin to previous methods, our color separation model is built upon the assumption that color of each pixel in an image can be formulated as a linear combination of a small set of same basis colors. However, different from previous palette-based recoloring methods which typically rely on heuristic rules to build the color separation model, we experimentally reveal the underlying relationship between the color separation and the palette-based recoloring, and summarize three specialized color separation priors that allow more faithful palette-based recoloring. Based on these priors, we devise a blind color separation model that not only does not require known palette as input as done in previous methods, but also enables more effective palette-based recoloring with much less visual artifacts. Experiments on two datasets demonstrate that our method outperforms the state-of-the-art palette-based recoloring methods. In addition, we show some applications enabled by the proposed color separation model, including automatic pattern coloring generation, green screen keying and region-controllable color transfer. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Triple-Cooperative Video Shadow DetectionabstractShadow detection in a single image has received significant research interests in recent years. However, much fewer works have been explored in shadow detection over dynamic scenes. The bottleneck is the lack of a well-established dataset with high-quality annotations for video shadow detection. In this work, we collect a new video shadow detection dataset (ViSha), which contains 120 videos with 11,685 frames, covering 60 object categories, varying lengths, and different motion/lighting conditions. All the frames are annotated with a high-quality pixel-level shadow mask. To the best of our knowledge, this is the first learning-oriented dataset for video shadow detection. Furthermore, we develop a new baseline model, named triple-cooperative video shadow detection network (TVSD-Net). It utilizes triple parallel networks in a cooperative manner to learn discriminative representations at intra-video and inter-video levels. Within the network, a dual gated co-attention module is proposed to constrain features from neighboring frames in the same video, while an auxiliary similarity loss is introduced to mine semantic information between different videos. Finally, we conduct a comprehensive study on ViSha, evaluating 12 state-of-the-art models (including single image shadow detectors, video object segmentation, and saliency detection methods). Experiments demonstrate that our model outperforms SOTA competitors. Zhihao Chen 0004, Lei Zhu 0003, Huazhu Fu, Wennan Liu, Harry Qin |
CVPR | 3 |
| 2021 | A Multi-Task Network for Joint Specular Highlight Detection and RemovalabstractSpecular highlight detection and removal are fundamental and challenging tasks. Although recent methods have achieved promising results on the two tasks by training on synthetic training data in a supervised manner, they are typically solely designed for highlight detection or removal, and their performance usually deteriorates significantly on real-world images. In this paper, we present a novel network that aims to detect and remove highlights from natural images. To remove the domain gap between synthetic training samples and real test images, and support the investigation of learning-based approaches, we first introduce a dataset with about 16K real images, each of which has the corresponding ground truths of highlight detection and removal. Using the presented dataset, we develop a multi-task network for joint highlight detection and removal, based on a new specular highlight image formation model. Experiments on the benchmark datasets and our new dataset show that our approach clearly outperforms state-of-the-art methods for both highlight detection and removal. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Ping Li 0016, Chunxia Xiao |
CVPR | 3 |
| 2021 | VIL-100: A New Dataset and A Baseline Model for Video Instance Lane DetectionabstractLane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we collect a new video instance lane detection (VIL-100) dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios. All the frames in each video are manually annotated to a high-quality instance-level lane annotation, and a set of frame-level and video-level metrics are included for quantitative performance evaluation. Moreover, we propose a new baseline model, named multi-level memory aggregation network (MMA-Net), for video instance lane detection. In our approach, the representation of current frame is enhanced by attentively aggregating both local and global memory features from other frames. Experiments on the new collected dataset show that the proposed MMA-Net outperforms state-of-the-art lane detection methods and video object segmentation methods. We release our dataset and code at https://github.com/yujun0-0/MMA-Net. Yujun Zhang 0002, Lei Zhu 0003, Wei Feng 0005, Huazhu Fu, Qingxia Li, Song Wang 0002 |
ICCV | 2 |
| 2021 | Boundary-Aware Transformers for Skin Lesion Segmentation
Jiacheng Wang 0002, Liansheng Wang 0002, Qichao Zhou, Lei Zhu 0003, Harry Qin |
MICCAI (1) | 5 |
| 2021 | From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real DataabstractSingle image dehazing is a challenging task, for which the domain shift between synthetic training data and real-world testing images usually leads to degradation of existing methods. To address this issue, we propose a novel image dehazing framework collaborating with unlabeled real data. First, we develop a disentangled image dehazing network (DID-Net), which disentangles the feature representations into three component maps, i.e. the latent haze-free image, the transmission map, and the global atmospheric light estimate, respecting the physical model of a haze process. Our DID-Net predicts the three component maps by progressively integrating features across scales, and refines each map by passing an independent refinement network. Then a disentangled-consistency mean-teacher network (DMT-Net) is employed to collaborate unlabeled real data for boosting single image dehazing. Specifically, we encourage the coarse predictions and refinements of each disentangled component to be consistent between the student and teacher networks by using a consistency loss on unlabeled real data. We make comparison with 13 state-of-the-art dehazing methods on a new collected dataset (Haze4K) and two widely-used dehazing datasets (i.e., SOTS and HazeRD), as well as on real-world hazy images. Experimental results demonstrate that our method has obvious quantitative and qualitative improvements over the existing methods. Lei Zhu 0003, Shunda Pei, Huazhu Fu, Harry Qin, Qing Zhang 0006, Wei Feng 0005 |
ACM Multimedia | 2 |
| 2021 | Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challengeabstractIntraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts). Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein |
Medical Image Anal. | 46 |
| 2021 | Global guidance network for breast lesion segmentation in ultrasound images
Cheng Xue 0003, Lei Zhu 0003, Huazhu Fu, Xiaowei Hu 0001, Xiaomeng Li 0001, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2021 | SAC-Net: Spatial Attenuation Context for Salient Object DetectionabstractThis paper presents a new deep neural network design for salient object detection by maximizing the integration of local and global image context within, around, and beyond the salient objects. Our key idea is to adaptively propagate and aggregate the image context features with variable attenuation over the entire feature maps. To achieve this, we design the spatial attenuation context (SAC) module to recurrently translate and aggregate the context features independently with different attenuation factors and then to attentively learn the weights to adaptively integrate the aggregated context features. By further embedding the module to process individual layers in a deep network, namely SAC-Net, we can train the network end-To-end and optimize the context features for detecting salient objects. Compared with 29 state-of-The-Art methods, experimental results show that our method performs favorably over all the others on six common benchmark data, both quantitatively and visually. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Tianyu Wang 0003, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Deep Sub-Region Network for Salient Object DetectionabstractSaliency detection is a fundamental and challenging task in computer vision, which aims at distinguishing the most conspicuous objects or regions in an image. Existing deep-learning methods mainly rely on the entire image to learn the global context information for saliency detection, which loses the spatial relation and results in ambiguity in predicting saliency maps. In this paper, we propose a novel deep sub-region network (DSR-Net) equipped with a sequence of sub-region dilated blocks (SRDB) by aggregating multi-scale salient context information of multiple sub-regions, such that the global context information from the whole image and local contexts from sub-regions are fused together, making the saliency prediction more accurate. Our SRDB separates the input feature map at different layers of a convolutional neural network (CNN) into different sub-regions and then designs a parallel ASPP module to refine feature maps at each sub-region. Experiments on the five widely-used saliency benchmark datasets demonstrate that our network outperforms recent state-of-the-art saliency detectors quantitatively and qualitatively on all the benchmarks. Liansheng Wang 0002, Rongzhen Chen, Lei Zhu 0003, Haoran Xie 0001, Xiaomeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Learning Gated Non-Local Residual for Single-Image Rain Streak RemovalabstractThis work presents a gated non-local deep residual learning framework for image deraining. It can avoid the over-deraining or under-deraining caused by the global residual learning in existing deraining networks, since the learned soft gate in our method adaptively adjusts the amount of global residual to be passed for generating the final derained result. To generate feature maps for global residual prediction, we develop a non-local guided attention module (NLAM), which first obtains non-local features by exploiting spatial inter-dependencies among all the feature positions of local features produced by convolutional neural network (CNN), and then leverages the attention mechanism to merge the local and non-local features based on their complementary relation. Moreover, we develop a channel-wise gated prediction module to learn a soft gate on the global residual by explicitly modelling channel inter-dependencies of the feature maps obtained from NLAM. Experiments on four deraining benchmark datasets and real-world rainy images show that our network has a quantitative and qualitative improvement over state-of-the-arts. Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Haoran Xie 0001, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Single-Image Real-Time Rain Removal Based on Depth-Guided Non-Local FeaturesabstractRain is a common weather phenomenon that affects environmental monitoring and surveillance systems. According to an established rain model (Garg and Nayar, 2007), the scene visibility in the rain varies with the depth from the camera, where objects faraway are visually blocked more by the fog than by the rain streaks. However, existing datasets and methods for rain removal ignore these physical properties, thus limiting the rain removal efficiency on real photos. In this work, we analyze the visual effects of rain subject to scene depth and formulate a rain imaging model that collectively considers rain streaks and fog. Also, we prepare a dataset called RainCityscapes on real outdoor photos. Furthermore, we design a novel real-time end-to-end deep neural network, for which we train to learn the depth-guided non-local features and to regress a residual map to produce a rain-free output image. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to show its superiority over others. Xiaowei Hu 0001, Lei Zhu 0003, Tianyu Wang 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Image Process. | 2 |
| 2021 | Enhancing Underexposed Photos Using Perceptually Bidirectional SimilarityabstractAlthough remarkable progress has been made, existing methods for enhancing underexposed photos tend to produce visually unpleasing results due to the existence of visual artifacts (e.g., color distortion, loss of details and uneven exposure). We observed that this is because they fail to ensure the perceptual consistency of visual information between the source underexposed image and its enhanced output. To obtain high-quality results free of these artifacts, we present a novel underexposed photo enhancement approach that is able to maintain the perceptual consistency. We achieve this by proposing an effective criterion, referred to as perceptually bidirectional similarity, which explicitly describes how to ensure the perceptual consistency. Particularly, we adopt the Retinex theory and cast the enhancement problem as a constrained illumination estimation optimization, where we formulate perceptually bidirectional similarity as constraints on illumination and solve for the illumination which can recover the desired artifact-free enhancement results. In addition, we describe a video enhancement framework that adopts the presented illumination estimation for handling underexposed videos. To this end, a probabilistic approach is introduced to propagate illuminations of sampled keyframes to the entire video by tackling a Bayesian Maximum A Posteriori problem. Extensive experiments demonstrate the superiority of our method over the state-of-the-art methods. Qing Zhang 0006, Yongwei Nie, Lei Zhu 0003, Chunxia Xiao, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | DNF-Net: A Deep Normal Filtering Network for Mesh DenoisingabstractThis article presents a deep normal filtering network, called DNF-Net, for mesh denoising. To better capture local geometry, our network processes the mesh in terms of local patches extracted from the mesh. Overall, DNF-Net is an end-to-end network that takes patches of facet normals as inputs and directly outputs the corresponding denoised facet normals of the patches. In this way, we can reconstruct the geometry from the denoised normals with feature preservation. Besides the overall network architecture, our contributions include a novel multi-scale feature embedding unit, a residual learning strategy to remove noise, and a deeply-supervised joint loss function. Compared with the recent data-driven works on mesh denoising, DNF-Net does not require manual input to extract features and better utilizes the training data to enhance its denoising performance. Finally, we present comprehensive experiments to evaluate our method and demonstrate its superiority over the state of the art on both synthetic and real-scanned meshes. Xianzhi Li 0001, Ruihui Li, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionabstractExisting shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow detection by leveraging unlabeled data and exploring the learning of multiple information of shadows simultaneously. To be specific, we first build a multi-task baseline model to simultaneously detect shadow regions, shadow edges, and shadow count by leveraging their complementary information and assign this baseline model to the student and teacher network. After that, we encourage the predictions of the three tasks from the student and teacher networks to be consistent for computing a consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from the predictions of the multi-task baseline model. Experimental results on three widely-used benchmark datasets show that our method consistently outperforms all the compared state-of- the-art methods, which verifies that the proposed network can effectively leverage additional unlabeled data to boost the shadow detection performance. Zhihao Chen 0004, Lei Zhu 0003, Song Wang 0002, Wei Feng 0005, Pheng-Ann Heng |
CVPR | 2 |
| 2020 | TexNet: Texture Loss Based Network for Gastric Antrum Segmentation in Ultrasound
Guohao Dong, Yaoxian Zou, Jiaming Jiao, Tianzhu Liang, Chaoyue Liu 0004, Lei Zhu 0003, Dong Ni 0001, Muqing Lin |
MICCAI (4) | 9 |
| 2020 | Shape Mask Generator: Learning to Refine Shape Priors for Segmenting Overlapping Cervical Cytoplasms
Youyi Song, Lei Zhu 0003, Bai Ying Lei, Bin Sheng 0001, Qi Dou 0001, Harry Qin, Kup-Sze Choi |
MICCAI (4) | 2 |
| 2020 | A Second-Order Subregion Pooling Network for Breast Lesion Segmentation in Ultrasound
Lei Zhu 0003, Rongzhen Chen, Huazhu Fu, Liansheng Wang 0002, Pheng-Ann Heng |
MICCAI (6) | 1 |
| 2020 | Learning to Detect Specular Highlights from Real-world ImagesabstractSpecular highlight detection is a challenging problem, and has many applications such as shiny object detection and light source estimation. Although various highlight detection methods have been proposed, they fail to disambiguate bright material surfaces from highlights, and cannot handle non-white-balanced images. Moreover, at present, there is still no benchmark dataset for highlight detection. In this paper, we present a large-scale real-world highlight dataset containing a rich variety of material categories, with diverse highlight shapes and appearances, in which each image is with an annotated ground-truth mask. Based on the dataset, we develop a deep learning-based specular highlight detection network (SHDNet) leveraging multi-scale context contrasted features to accurately detect specular highlights of varying scales. In addition, we design a binary cross-entropy (BCE) loss and an intersection-over-union edge (IoUE) loss for our network. Compared with existing highlight detection methods, our method can accurately detect highlights of different sizes, while effectively excluding the non-highlight regions, such as bright materials, non-specular as well as colored lighting, and even light sources. Gang Fu 0003, Qing Zhang 0006, Qifeng Lin, Lei Zhu 0003, Chunxia Xiao |
ACM Multimedia | 4 |
| 2020 | Direction-Aware Spatial Context Features for Shadow Detection and RemovalabstractShadow detection and shadow removal are fundamental and challenging tasks, requiring an understanding of the global image semantics. This paper presents a novel deep neural network design for shadow detection and removal by analyzing the spatial image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting and removing shadows. This design is developed into the DSC module and embedded in a convolutional neural network (CNN) to learn the DSC features at different levels. Moreover, we design a weighted cross entropy loss to make effective the training for shadow detection and further adopt the network for shadow removal by using a euclidean loss function and formulating a color transfer function to address the color and luminosity inconsistencies in the training pairs. We employed two shadow detection benchmark datasets and two shadow removal benchmark datasets, and performed various experiments to evaluate our method. Experimental results show that our method performs favorably against the state-of-the-art methods for both shadow detection and shadow removal. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Aggregating Attentional Dilated Features for Salient Object DetectionabstractThis paper presents a novel deep learning model to aggregate the attentional dilated features for salient object detection by exploring the complementary information between the global and local context in a convolutional neural network. There are two technical contributions to our network design. First, we develop an attentional dense atrous (dilated) spatial pyramid pooling (AD-ASPP) module to selectively use the local saliency cues captured by dilated convolutions with a small rate and the global saliency cues captured by dilated convolutions with a large rate. Second, taking the feature pyramid network as the backbone, we develop an aggregation network to integrate the refined features by formulating two consecutive chains of residual learning based modules: one chain from deep to shallow layers while another chain from shallow to deep layers. We evaluate our network on seven widely-used saliency detection benchmarks by comparing it against 21 state-of-the-art methods. Experimental results show that our network outperforms others on all the seven benchmark datasets. Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | CANet: Cross-Disease Attention Network for Joint Diabetic Retinopathy and Diabetic Macular Edema GradingabstractDiabetic retinopathy (DR) and diabetic macular edema (DME) are the leading causes of permanent blindness in the working-age population. Automatic grading of DR and DME helps ophthalmologists design tailored treatments to patients, thus is of vital importance in the clinical practice. However, prior works either grade DR or DME, and ignore the correlation between DR and its complication, i.e., DME. Moreover, the location information, e.g., macula and soft hard exhaust annotations, are widely used as a prior for grading. Such annotations are costly to obtain, hence it is desirable to develop automatic grading methods with only image-level supervision. In this article, we present a novel cross-disease attention network (CANet) to jointly grade DR and DME by exploring the internal relationship between the diseases with only image-level supervision. Our key contributions include the disease-specific attention module to selectively learn useful features for individual diseases, and the disease-dependent attention module to further capture the internal relationship between the two diseases. We integrate these two attention modules in a deep network to produce disease-specific and disease-dependent features, and to maximize the overall performance jointly for grading DR and DME. We evaluate our network on two public benchmark datasets, i.e., ISBI 2018 IDRiD challenge dataset and Messidor dataset. Our method achieves the best result on the ISBI 2018 IDRiD challenge dataset and outperforms other methods on the Messidor dataset. Our code is publicly available at https://github.com/xmengli999/CANet. Xiaomeng Li 0001, Xiaowei Hu 0001, Lequan Yu, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2020 | ψ-Net: Stacking Densely Convolutional LSTMs for Sub-Cortical Brain Structure SegmentationabstractSub-cortical brain structure segmentation is of great importance for diagnosing neuropsychiatric disorders. However, developing an automatic approach to segmenting sub-cortical brain structures remains very challenging due to the ambiguous boundaries, complex anatomical structures, and large variance of shapes. This paper presents a novel deep network architecture, namely Ψ -Net, for sub-cortical brain structure segmentation, aiming at selectively aggregating features and boosting the information propagation in a deep convolutional neural network (CNN). To achieve this, we first formulate a densely convolutional LSTM module (DC-LSTM) to selectively aggregate the convolutional features with the same spatial resolution at the same stage of a CNN. This helps to promote the discriminativeness of features at each CNN stage. Second, we stack multiple DC-LSTMs from the deepest stage to the shallowest stage to progressively enrich low-level feature maps with high-level context. We employ two benchmark datasets on sub-cortical brain structure segmentation, and perform various experiments to evaluate the proposed Ψ -Net. The experimental results show that our network performs favorably against the state-of-the-art methods on both benchmark datasets. Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Saliency-Aware Texture SmoothingabstractTexture smoothing aims to smooth out textures in images, while retaining the prominent structures. This paper presents a saliency-aware approach to the problem with two key contributions. First, we design a deep saliency network with guided non-local blocks (GNLBs) for learning long-range pixel dependencies by taking the predicted saliency map at former layer as the guidance image to help suppress the non-saliency regions in the shallow layer. The GNLB computes the saliency response at a position by a weighted sum of features at all positions, and enables us to produce results that outperform existing deep saliency models. Second, we formulate a joint optimization framework to take saliency information when iteratively separating textures from structures: on the texture layer, we smooth out structures with the help of the saliency information and migrate structures from the texture to structure layer, while on the structure layer, we adopt another deep model to detect edges and simultaneous sparse coding to push textures back to the texture layer. We tested our method on a rich variety of images and compared it with several state-of-the-art methods. Both visual and quantitative comparison results show that our method better preserves structures while removing the texture components. Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Depth-Attentional Features for Single-Image Rain RemovalabstractRain is a common weather phenomenon, where object visibility varies with depth from the camera and objects faraway are visually blocked more by fog than by rain streaks. Existing methods and datasets for rain removal, however, ignore these physical properties, thereby limiting the rain removal efficiency on real photos. In this work, we first analyze the visual effects of rain subject to scene depth and formulate a rain imaging model collectively with rain streaks and fog; by then, we prepare a new dataset called RainCityscapes with rain streaks and fog on real outdoor photos. Furthermore, we design an end-to-end deep neural network, where we train it to learn depth-attentional features via a depth-guided attention mechanism, and regress a residual map to produce the rain-free image output. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to demonstrate its superiority over the others. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Pheng-Ann Heng |
CVPR | 3 |
| 2019 | Deep Multi-Model Fusion for Single-Image DehazingabstractThis paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural network (CNN) features at different CNN layers and generate the attentional multi-level integrated features (AMLIF). Then, from the AMLIF, we further predict a haze-free result for an atmospheric scattering model, as well as for four haze-layer separation models, and then fuse the results together to produce the final haze-free image. To evaluate the effectiveness of our method, we compare our network with several state-of-the-art methods on two widely-used dehazing benchmark datasets, as well as on two sets of real-world hazy images. Experimental results demonstrate clear quantitative and qualitative improvements of our method over the state-of-the-arts. Zijun Deng, Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Qing Zhang 0006, Harry Qin, Pheng-Ann Heng |
ICCV | 2 |
| 2019 | ARS-Net: Adaptively Rectified Supervision Network for Automated 3D Ultrasound Image Segmentation
Chaoyue Liu 0004, Guohao Dong, Muqing Lin, Yaoxian Zou, Tianzhu Liang, Xujin He, Dong Ni 0001, Yi Xiong 0001, Lei Zhu 0003 |
MICCAI (3) | 10 |
| 2019 | Probabilistic Multilayer Regularization Network for Unsupervised 3D Brain Image Registration
Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng |
MICCAI (2) | 3 |
| 2019 | Segmentation of Overlapping Cytoplasm in Cervical Smear Images via Adaptive Shape Priors Extracted From Contour FragmentsabstractWe present a novel approach for segmenting overlapping cytoplasm of cells in cervical smear images by leveraging the adaptive shape priors extracted from cytoplasm's contour fragments and shape statistics. The main challenge of this task is that many occluded boundaries in cytoplasm clumps are extremely difficult to be identified and, sometimes, even visually indistinguishable. Given a clump where multiple cytoplasms overlap, our method starts by cutting its contour into a set of contour fragments. We then locate the corresponding contour fragments of each cytoplasm by a grouping process. For each cytoplasm, according to the grouped fragments and a set of known shape references, we construct its shape and, then, connect the fragments to form a closed contour as the segmentation result, which is explicitly constrained by the constructed shape. We further integrate the intensity and curvature information, which is complementary to the shape priors extracted from contour fragments, into our framework to improve the segmentation accuracy. We propose to iteratively conduct fragments grouping, shape constructing, and fragments connecting for progressively refining the shape priors and improving the segmentation results. We extensively evaluate the effectiveness of our method on two typical cervical smear datasets. The experimental results demonstrate that our approach is highly effective and consistently outperforms the state-of-the-art approaches. The proposed method is general enough to be applied to other similar microscopic image segmentation tasks, where heavily overlapped objects exist. Youyi Song, Lei Zhu 0003, Harry Qin, Bai Ying Lei, Bin Sheng 0001, Kup-Sze Choi |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Deep Attentive Features for Prostate Segmentation in 3D Transrectal UltrasoundabstractAutomatic prostate segmentation in transrectal ultrasound (TRUS) images is of essential importance for image-guided prostate interventions and treatment planning. However, developing such automatic solutions remains very challenging due to the missing/ambiguous boundary and inhomogeneous intensity distribution of the prostate in TRUS, as well as the large variability in prostate shapes. This paper develops a novel 3D deep neural network equipped with attention modules for better prostate segmentation in TRUS by fully exploiting the complementary information encoded in different layers of the convolutional neural network (CNN). Our attention module utilizes the attention mechanism to selectively leverage the multi-level features integrated from different layers to refine the features at each individual layer, suppressing the non-prostate noise at shallow layers of the CNN and increasing more prostate details into features at deep layers. Experimental results on challenging 3D TRUS volumes show that our method attains satisfactory segmentation performance. The proposed attention mechanism is a general strategy to aggregate multi-level deep features and has the potential to be used for other medical image segmentation tasks. The code is publicly available at https://github.com/wulalago/DAF3D. Yi Wang 0031, Dong Ni 0001, Haoran Dou, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Harry Qin, Pheng-Ann Heng, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2018 | Recurrently Aggregating Deep Features for Salient Object DetectionabstractSalient object detection is a fundamental yet challenging problem in computer vision, aiming to highlight the most visually distinctive objects or regions in an image. Recent works benefit from the development of fully convolutional neural networks (FCNs) and achieve great success by integrating features from multiple layers of FCNs. However, the integrated features tend to include non-salient regions (due to low level features of the FCN) or lost details of salient objects (due to high level features of the FCN) when producing the saliency maps. In this paper, we develop a novel deep saliency network equipped with recurrently aggregated deep features (RADF) to more accurately detect salient objects from an image by fully exploiting the complementary saliency information captured in different layers. The RADF utilizes the multi-level features integrated from different layers of a FCN to recurrently refine the features at each layer, suppressing the non-salient noise at low-level of the FCN and increasing more salient details into features at high layers. We perform experiments to evaluate the effectiveness of the proposed network on 5 famous saliency detection benchmarks and compare it with 15 state-of-the-art methods. Our method ranks first in 4 of the 5 datasets and second in the left dataset. Xiaowei Hu 0001, Lei Zhu 0003, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
AAAI | 2 |
| 2018 | Direction-Aware Spatial Context Features for Shadow DetectionabstractShadow detection is a fundamental and challenging task, since it requires an understanding of global image semantics and there are various backgrounds around shadows. This paper presents a novel network for shadow detection by analyzing image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting shadows. This design is developed into the DSC module and embedded in a CNN to learn DSC features at different levels. Moreover, a weighted cross entropy loss is designed to make the training more effective. We employ two common shadow detection benchmark datasets and perform various experiments to evaluate our network. Experimental results show that our network outperforms state-of-the-art methods and achieves 97% accuracy and 38% reduction on balance error rate. Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
CVPR | 2 |
| 2018 | Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection
Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
ECCV (6) | 1 |
| 2018 | R³Net: Recurrent Residual Refinement Network for Saliency DetectionabstractSaliency detection is a fundamental yet challenging task in computer vision, aiming at highlighting the most visually distinctive objects in an image. We propose a novel recurrent residual refinement network (R^3Net) equipped with residual refinement blocks (RRBs) to more accurately detect salient regions of an input image. Our RRBs learn the residual between the intermediate saliency prediction and the ground truth by alternatively leveraging the low-level integrated features and the high-level integrated features of a fully convolutional network (FCN). While the low-level integrated features are capable of capturing more saliency details, the high-level integrated features can reduce non-salient regions in the intermediate prediction. Furthermore, the RRBs can obtain complementary saliency information of the intermediate prediction, and add the residual into the intermediate prediction to refine the saliency maps. We evaluate the proposed R^3Net on five widely-used saliency detection benchmarks by comparing it with 16 state-of-the-art saliency detectors. Experimental results show that our network outperforms our competitors in all the benchmark datasets. Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Harry Qin, Guoqiang Han 0002, Pheng-Ann Heng |
IJCAI | 3 |
| 2018 | Deep Attentional Features for Prostate Segmentation in Ultrasound
Yi Wang 0031, Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Xuemiao Xu, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (4) | 4 |
| 2018 | High-Quality Exposure Correction of Underexposed PhotosabstractWe address the problem of correcting the exposure of underexposed photos. Previous methods have tackled this problem from many different perspectives and achieved remarkable progress. However, they usually fail to produce natural-looking results due to the existence of visual artifacts such as color distortion, loss of detail, exposure inconsistency, etc. We find that the main reason why existing methods induce these artifacts is because they break a perceptually similarity between the input and output. Based on this observation, an effective criterion, termed as perceptually bidirectional similarity (PBS) is proposed. Based on this criterion and the Retinex theory, we cast the exposure correction problem as an illumination estimation optimization, where PBS is defined as three constraints for estimating illumination that can generate the desired result with even exposure, vivid color and clear textures. Qualitative and quantitative comparisons, and the user study demonstrate the superiority of our method over the state-of-the-art methods. Qing Zhang 0006, Ganzhao Yuan, Chunxia Xiao, Lei Zhu 0003, Wei-Shi Zheng 0001 |
ACM Multimedia | 4 |
| 2018 | Non-Local Low-Rank Normal Filtering for Mesh DenoisingabstractAbstract This paper presents a non‐local low‐rank normal filtering method for mesh denoising. By exploring the geometric similarity between local surface patches on 3D meshes in the form of normal fields, we devise a low‐rank recovery model that filters normal vectors by means of patch groups. In summary, our method has the following key contributions. First, we present the guided normal patch covariance descriptor to analyze the similarity between patches. Second, we pack normal vectors on similar patches into the normal‐field patch‐group (NPG) matrix for rank analysis. Third, we formulate mesh denoising as a low‐rank matrix recovery problem based on the prior that the rank of the NPG matrix is high for raw meshes with noise, but can be significantly reduced for denoised meshes, whose normal vectors across similar patches should be more strongly correlated. Furthermore, we devise an objective function based on an improved truncated γ norm, and derive an op tim ization procedure using the alternative direction method of multipliers and iteratively re‐weighted least squares techniques. We conducted several experiments to evaluate our method using various 3D models, and compared our results against several state‐of‐the‐art methods. Experimental results show that our method consistently outperforms other methods and better preserves the fine details. Xianzhi Li 0001, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
Comput. Graph. Forum | 2 |
| 2018 | Feature-preserving ultrasound speckle reduction via L0 minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Chi-Wing Fu, Pheng-Ann Heng |
Neurocomputing | 1 |
| 2017 | A Non-local Low-Rank Framework for Ultrasound Speckle ReductionabstractSpeckle refers to the granular patterns that occur in ultrasound images due to wave interference. Speckle removal can greatly improve the visibility of the underlying structures in an ultrasound image and enhance subsequent post processing. We present a novel framework for speckle removal based on low-rank non-local filtering. Our approach works by first computing a guidance image that assists in the selection of candidate patches for non-local filtering in the face of significant speckles. The candidate patches are further refined using a low-rank minimization estimated using a truncated weighted nuclear norm (TWNN) and structured sparsity. We show that the proposed filtering framework produces results that outperform state-of-the-art methods both qualitatively and quantitatively. This framework also provides better segmentation results when used for pre-processing ultrasound images. Lei Zhu 0003, Chi-Wing Fu, Michael S. Brown, Pheng-Ann Heng |
CVPR | 1 |
| 2017 | Joint Bi-layer Optimization for Single-Image Rain Streak RemovalabstractWe present a novel method for removing rain streaks from a single input image by decomposing it into a rain-free background layer B and a rain-streak layer R. A joint optimization process is used that alternates between removing rain-streak details from B and removing non-streak details from R. The process is assisted by three novel image priors. Observing that rain streaks typically span a narrow range of directions, we first analyze the local gradient statistics in the rain image to identify image regions that are dominated by rain streaks. From these regions, we estimate the dominant rain streak direction and extract a collection of rain-dominated patches. Next, we define two priors on the background layer B, one based on a centralized sparse representation and another based on the estimated rain direction. A third prior is defined on the rain-streak layer R, based on similarity of patches to the extracted rain patches. Both visual and quantitative comparisons demonstrate that our method outperforms the state-of-the-art. Lei Zhu 0003, Chi-Wing Fu, Dani Lischinski, Pheng-Ann Heng |
ICCV | 1 |
| 2017 | Fast feature-preserving speckle reduction for ultrasound images via phase congruency
Lei Zhu 0003, Weiming Wang 0002, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Pheng-Ann Heng |
Signal Process. | 1 |
| 2016 | Ultrasound Speckle Reduction via L_0 Minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Pheng-Ann Heng |
ACCV (3) | 1 |
| 2016 | Non-Local Sparse and Low-Rank Regularization for Structure-Preserving Image SmoothingabstractAbstract This paper presents a new image smoothing method that better preserves prominent structures. Our method is inspired by the recent non‐local image processing techniques on the patch grouping and filtering. Overall, it has three major contributions over previous works. First, we employ the diffusion map as the guidance image to improve the accuracy of patch similarity estimation using the region covariance descriptor. Second, we model structure‐preserving image smoothing as a low‐rank matrix recovery problem, aiming at effectively filtering the texture information in similar patches. Lastly, we devise an objective function, namely the weighted robust principle component analysis (WRPCA), by regularizing the low rank with the weighted nuclear norm and sparsity pursuit with L1norm, and solve this non‐convex WRPCA optimization problem by adopting the alternative direction method of multipliers (ADMM) technique. We experiment our method with a wide variety of images and compare it against several state‐of‐the‐art methods. The results show that our method achieves better structure preservation and texture suppression as compared to other methods. We also show the applicability of our method on several image processing tasks such as edge detection, texture enhancement and seam carving. Lei Zhu 0003, Chi-Wing Fu, Yueming Jin, Mingqiang Wei, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 1 |
| 2015 | Morphology-preserving smoothing on polygonized isosurfaces of inhomogeneous binary volumes
Mingqiang Wei, Lei Zhu 0003, Jinze Yu 0001, Jun Wang 0039, Wai-Man Pang, Jianhuang Wu, Harry Qin, Pheng-Ann Heng |
Comput. Aided Des. | 2 |
| 2013 | Coarse-to-Fine Normal Filtering for Feature-Preserving Mesh Denoising Based on Isotropic SubneighborhoodsabstractAbstract State‐of‐theart normal filters usually denoise each face normal using its entire anisotropic neighborhood. However, enforcing these filters indiscriminately on the anisotropic neighborhood will lead to feature blurring, especially in challenging regions with shallow features. We develop a novel mesh denoising framework which can effectively preserve features with various sizes. Our idea is inspired by the observation that the underlying surface of a noisy mesh is piecewise smooth. In this regard, it is more desirable that we denoise each face normal within its piecewise smooth region (we call such a region as an isotropic subneighborhood) instead of using the anisotropic neighborhood. To achieve this, we first classify mesh faces into several types using a face normal tensor voting and then perform a normal filter to obtain a denoised coarse normal field. Based on the results of normal classification and the denoised coarse normal field, we segment the anisotropic neighborhood of every feature face into a number of isotropic subneighborhoods via local spectral clustering. Thus face normal filtering can be performed again on the isotropic subneighborhoods and produce a more accurate normal field. Extensive tests on various models demonstrate that our method can achieve better performance than state‐of‐theart normal filters, especially in challenging regions with features. Lei Zhu 0003, Mingqiang Wei, Jinze Yu 0001, Weiming Wang 0002, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 1 |