VLDB 2026 Research / reviewers in the wild / expert
Hao Zheng 0008
dblp:31/6916-8
· DBLP profile ↗
42ranked-venue papers
6as first author
35since 2021 · last 2025
0000-0001-7193-6242ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DGFamba: Learning Flow Factorized State Space for Visual Domain GeneralizationabstractDomain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DGFamba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li |
AAAI | 3 |
| 2025 | NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationabstractNight-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nighttime scenes with varying illumination. In light of this issue, we focus on investigating generalizable paradigms for night-time scene segmentation and propose an efficient fine-tuning scheme, dubbed NightAdapter, alleviating the domain gap across various scenes. Interestingly, different properties embedded in the day-time and night-time features can be characterized by the bands after discrete sine transform, which can be categorized into illumination-sensitive/-insensitive bands. Hence, our NightAdapter is powered by two appealing designs: (1) Illumination-Insensitive Band Adaptation that provides a foundation for understanding the prior, enhancing the robustness to illumination shifts; (2) Illumination-Sensitive Band Adaptation that fine-tunes the randomized frequency bands, mitigating the domain gap between the day-time and various night-time scenes. As a consequence, illumination-insensitive enhancement improves the domain invariance, while illumination-sensitive diminution strengthens the domain shift between different scenes. NightAdapter yields significant improvements over the state-of-the-art methods under various day-to-night, night-to-night, and in-domain night segmentation experiments. Source code is available at https://github.com/BiQiWHU/NightAdapter. Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
CVPR | 4 |
| 2025 | A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICCV | 4 |
| 2025 | GaussianReg: Rapid 2D/3D Registration for Emergency Surgery Via Explicit 3D Modeling with Gaussian Primitives
Weihao Yu 0004, Xiaoqing Guo, Xinyu Liu 0001, Yifan Liu 0010, Hao Zheng 0008, Yawen Huang, Yixuan Yuan |
ICCV | 5 |
| 2025 | D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (5) | 3 |
| 2025 | BrainSegDMIF: A Dynamic Fusion-enhanced SAM for Brain Lesion Segmentation
Hongming Wang, Huimin Huang 0002, Jiaxuan Jiang 0001, Hao Zheng 0008, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001, Jinping Xu |
ACM Multimedia | 7 |
| 2025 | AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 3 |
| 2025 | Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image RestorationabstractImage restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze).
Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images.
To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration.
Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution,
while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process.
A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised.
Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process.
A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model.
Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts. Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 3 |
| 2025 | Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationabstractDomain generalization aims to train models that perform robustly on unseen target domains without access to target data.
The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability.
However, the static alignment to class-level textual anchors remains insufficient to handle the dramatic distribution discrepancy from diverse domain-specific visual features.
In this work, we propose a novel cross-domain Schrödinger Bridge (SB) method, namely SBGen, to handle this challenge, which explicitly formulates the stochastic semantic evolution, to gain better generalization to unseen domains.
Technically, the proposed \texttt{SBGen} consists of three key components: (1) \emph{text-guided domain-aware feature selection} to isolate semantically aligned image tokens; (2) \emph{stochastic cross-domain evolution} to simulate the SB dynamics via a learnable time-conditioned drift; and (3) \emph{stochastic domain-agnostic interpolation} to construct semantically grounded feature trajectories.
Empirically, \texttt{SBGen} achieves state-of-the-art performance on domain generalization in both classification and segmentation. This work highlights the importance of modeling domain shifts as structured stochastic processes grounded in semantic alignment. Hao Zheng 0008, Jingjun Yi, Qi Bi, Huimin Huang 0002, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 1 |
| 2025 | Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions
Yawen Huang, Huimin Huang 0002, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | Learning Generalized Medical Image Representation by Decoupled Feature QueriesabstractMedical images are usually collected from multiple clinical centers with various types of scanners. When confronted with such significant cross-domain distribution discrepancy, a deep network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. Such channel redundancy limits the expressive capability of a representation, resulting in less preferable generalization ability. To address this fundamental yet challenging issue, we propose a novel decoupled feature as query (DFQ) framework for domain generalized medical image representation learning. Its general idea is to leverage the channel-wise decoupled deep features as queries. Particularly, a deep instance whitening transform with restricted isometry is proposed, which enforces each channel orthogonal to the rest channels after decoupling. Besides, the long-range dependency between decoupled deep and shallow features is implicitly constrained to minimize channel redundancy throughout training. Extensive experiments show its state-of-the-art performance on three medical domain generalization tasks with four modalities. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | GAD: Domain generalized diabetic retinopathy grading by grade-aware de-stylization
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
Pattern Recognit. | 3 |
| 2025 | PASS: Test-Time Prompting to Adapt Styles and Semantic Shapes in Medical Image SegmentationabstractTest-time adaptation (TTA) has emerged as a promising paradigm to handle the domain shifts at test time for medical images from different institutions without using extra training data. However, existing TTA solutions for segmentation tasks suffer from 1) dependency on modifying the source training stage and access to source priors or 2) lack of emphasis on shape-related semantic knowledge that is crucial for segmentation tasks. Recent research on visual prompt learning achieves source-relaxed adaptation by extended parameter space but still neglects the full utilization of semantic features, thus motivating our work on knowledge-enriched deep prompt learning. Beyond the general concern of image style shifts, we reveal that shape variability is another crucial factor causing the performance drop. To address this issue, we propose a TTA framework called PASS (Prompting to Adapt Styles and Semantic shapes), which jointly learns two types of prompts: the input-space prompt to reformulate the style of the test image to fit into the pretrained model and the semantic-aware prompts to bridge high-level shape discrepancy across domains. Instead of naively imposing a fixed prompt, we introduce an input decorator to generate the self-regulating visual prompt conditioned on the input data. To retrieve the knowledge representations and customize target-specific shape prompts for each test sample, we propose a cross-attention prompt modulator, which performs interaction between target representations and an enriched shape prompt bank. Extensive experiments demonstrate the superior performance of PASS over state-of-the-art methods on multiple medical image segmentation datasets. The code is available at https://github.com/EndoluminalSurgicalVision-IMR/PASS. Chuyan Zhang, Hao Zheng 0008, Xin You 0002, Yefeng Zheng 0001, Yun Gu |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Learning Generalized Medical Image Segmentation from Decoupled Feature QueriesabstractDomain generalized medical image segmentation requires models to learn from multiple source domains and generalize well to arbitrary unseen target domain. Such a task is both technically challenging and clinically practical, due to the domain shift problem (i.e., images are collected from different hospitals and scanners). Existing methods focused on either learning shape-invariant representation or reaching consensus among the source domains. An ideal generalized representation is supposed to show similar pattern responses within the same channel for cross-domain images. However, to deal with the significant distribution discrepancy, the network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. To address this issue, we propose to leverage channel-wise decoupled deep features as queries. With the aid of cross-attention mechanism, the long-range dependency between deep and shallow features can be fully mined via self-attention and then guides the learning of generalized representation. Besides, a relaxed deep whitening transformation is proposed to learn channel-wise decoupled features in a feasible way. The proposed decoupled fea- ture query (DFQ) scheme can be seamlessly integrate into the Transformer segmentation model in an end-to-end manner. Extensive experiments show its state-of-the-art performance, notably outperforming the runner-up by 1.31% and 1.98% with DSC metric on generalized fundus and prostate benchmarks, respectively. Source code is available at https://github.com/BiQiWHU/DFQ. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 3 |
| 2024 | Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsabstractMulti-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics. Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001 |
CVPR | 6 |
| 2024 | Self-Supervised Cross-Level Consistency Learning For Fundus Image ClassificationabstractThe rapid development of intelligent systems for eye disease diagnosis decreases the risk of people suffering from vision impairment. However, the superior discrimination ability of existing retinal disease diagnosis methods heavily relies on the large-scale high-quality annotations. In this work, we adapt the self-supervised technique for fundus image classification with the merits of bypassing the over-dependence of labeled data. Unlike most current self-supervised approaches, which only learn global pre-text representations from view-level, our method further incorporates the region-level representations into the learning process, since the pathological changes in fundus images are usually subtle and scattered. Specifically, we propose a novel self-supervised cross-level consistency learning scheme (S2C2L), which leverages both view-level and region-level representations of a vision Transformer to improve the robustness of extracted self-supervised representation. A diagnosis perception module (DPM) is constructed to enhance the activation of local pathological regions from both region and view levels, and a cross-level consistency loss is dedicated to align the representations from both levels. Extensive experiments on iChallenge-AMD, LAG and APTOS2019 datasets validate the state-of-the-art performance of our method for three common eye diseases. Qi Bi, Hao Zheng 0008, Xu Sun 0006, Jingjun Yi, Wentian Zhang, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICASSP | 2 |
| 2024 | Learning to Segment Multiple Organs from Multimodal Partially Labeled Datasets
Dong Wei 0004, Donghuan Lu, Jinghan Sun, Hao Zheng 0008, Yefeng Zheng 0001, Liansheng Wang 0002 |
MICCAI (9) | 5 |
| 2024 | MoME: Mixture of Multimodal Experts for Cancer Survival Prediction
Conghao Xiong, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 3 |
| 2024 | TAKT: Target-Aware Knowledge Transfer for Whole Slide Image Classification
Conghao Xiong, Yi Lin 0009, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 4 |
| 2024 | Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (10) | 3 |
| 2024 | Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationabstractThe rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as the cross-domain images share common pixel-wise content information but vary greatly in terms of the style. In this paper, we present a novel Spectral-dEcomposed Token (SET) learning framework to advance the frontier. Delving into further than existing fine-tuning token & frozen backbone paradigm, the proposed SET especially focuses on the way learning style-invariant features from these learnable tokens. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction.After the decomposition, style variation primarily impacts the token-based feature enhancement within the amplitude branch. To address this issue, we further develop an attention optimization method to bridge the gap between style-affected representation and static tokens during inference. Extensive cross-domain experiments show its state-of-the-art performance. Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 3 |
| 2024 | Samba: Severity-aware Recurrent Modeling for Cross-domain Medical Image GradingabstractDisease grading is a crucial task in medical image analysis. Due to the continuous progression of diseases, i.e., the variability within the same level and the similarity between adjacent stages, accurate grading is highly challenging.
Furthermore, in real-world scenarios, models trained on limited source domain datasets should also be capable of handling data from unseen target domains.
Due to the cross-domain variants, the feature distribution between source and unseen target domains can be dramatically different, leading to a substantial decrease in model performance.
To address these challenges in cross-domain disease grading, we propose a Severity-aware Recurrent Modeling (Samba) method in this paper.
As the core objective of most staging tasks is to identify the most severe lesions, which may only occupy a small portion of the image, we propose to encode image patches in a sequential and recurrent manner.
Specifically, a state space model is tailored to store and transport the severity information by hidden states.
Moreover, to mitigate the impact of cross-domain variants, an Expectation-Maximization (EM) based state recalibration mechanism is designed to map the patch embeddings into a more compact space.
We model the feature distributions of different lesions through the Gaussian Mixture Model (GMM) and reconstruct the intermediate features based on learnable severity bases.
Extensive experiments show the proposed Samba outperforms the VMamba baseline by an average accuracy of 23.5\%, 5.6\% and 4.1\% on the cross-domain grading of fatigue fracture, breast cancer and diabetic retinopathy, respectively.
Source code is available at \url{https://github.com/BiQiWHU/Samba}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Haolan Zhan, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 3 |
| 2024 | Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationabstractThe emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images.
Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, i.e., the variance of urban landscape and environment dependencies.
Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS.
The frequency space after Haar wavelet transformation provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information are retained in the low- and high- frequency components of the space, respectively.
To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier.
Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process.
Particularly, the proposed FADA consists of two branches, i.e., low- and high- frequency branches. The former one is able to stabilize the scene content, while the latter one learns the scene styles and eliminates its impact to DGSS.
Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs.
Source code is available at \url{https://github.com/BiQiWHU/FADA}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 3 |
| 2024 | Multi-Constraint Transferable Generative Adversarial Networks for Cross-Modal Brain Image Synthesis
Yawen Huang, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Guo-Jun Qi, Ling Shao 0001, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Affine Collaborative Normalization: A shortcut for adaptation in medical image analysis
Chuyan Zhang, Yuncheng Yang, Hao Zheng 0008, Yawen Huang, Yefeng Zheng 0001, Yun Gu |
Pattern Recognit. | 3 |
| 2023 | AirwayFormer: Structure-Aware Boundary-Adaptive Transformers for Airway Anatomical Labeling
Weihao Yu 0004, Hao Zheng 0008, Yun Gu, Fangfang Xie, Jiayuan Sun, Jie Yang 0002 |
MICCAI (7) | 2 |
| 2023 | Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu |
Medical Image Anal. | 5 |
| 2023 | Dive into the details of self-supervised learning for medical image analysis
Chuyan Zhang, Hao Zheng 0008, Yun Gu |
Medical Image Anal. | 2 |
| 2023 | Continuous cross-modal hashing
Hao Zheng 0008, Jinbao Wang 0001, Xiantong Zhen, Jingkuan Song, Feng Zheng 0001, Ke Lu 0002, Guo-Jun Qi |
Pattern Recognit. | 1 |
| 2023 | TNN: Tree Neural Network for Airway Anatomical LabelingabstractDetailed anatomical labeling of bronchial trees extracted from CT images can be used as fine-grained maps for intra-operative navigation. To cater to the sparse distribution of airway voxels and large class imbalance in 3D image space, a graph-neural-network-based method is proposed to map branches to nodes in a graph space and assign anatomical labels down to subsegmental level. To address the inherent problem of overlapping distribution of positional and morphological features, especially for subsegmental categories, the proposed method focuses on the relative position between sibling subsegments which is fixed in most cases. The hierarchical nomenclature is represented by multi-level labeling and each category is associated with one or two subtrees in the graph. Hyperedges are used to extract the representation of subtrees while a hypergraph neural network is developed to encode their intrinsic relationship through hyperedge interaction. A filter module is further designed to guide feature aggregation between nodes and hyperedges. With the proposed method, the final accuracies for segmental and subsegmental node classification can achieve 93.6% and 82.0% respectively. The corresponding code is publicly available at https://github.com/haozheng-sjtu/airway-labeling. Weihao Yu 0004, Hao Zheng 0008, Yun Gu, Fangfang Xie, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang |
IEEE Trans. Medical Imaging | 2 |
| 2023 | S3R: Shape and Semantics-Based Selective Regularization for Explainable Continual Segmentation Across Multiple SitesabstractIn clinical practice, it is desirable for medical image segmentation models to be able to continually learn on a sequential data stream from multiple sites, rather than a consolidated dataset, due to storage cost and privacy restrictions. However, when learning on a new site, existing methods struggle with a weak memorizability for previous sites with complex shape and semantic information, and a poor explainability for the memory consolidation process. In this work, we propose a novel Shape and Semantics-based Selective Regularization ( [Formula: see text]) method for explainable cross-site continual segmentation to maintain both shape and semantic knowledge of previously learned sites. Specifically, [Formula: see text] method adopts a selective regularization scheme to penalize changes of parameters with high Joint Shape and Semantics-based Importance (JSSI) weights, which are estimated based on the parameter sensitivity to shape properties and reliable semantics of the segmentation object. This helps to prevent the related shape and semantic knowledge from being forgotten. Moreover, we propose an Importance Activation Mapping (IAM) method for memory interpretation, which indicates the spatial support for important parameters to visualize the memorized content. We have extensively evaluated our method on prostate segmentation and optic cup and disc segmentation tasks. Our method outperforms other comparison methods in reducing model forgetting and increasing explainability. Our code is available at https://github.com/jingyzhang/S3R. Jingyang Zhang, Ran Gu, Peng Xue 0005, Mianxin Liu, Hao Zheng 0008, Yefeng Zheng 0001, Lei Ma 0006, Guotai Wang, Lixu Gu |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Semi-Supervised Skin Lesion Segmentation with Learning Model ConfidenceabstractSegmentation of skin lesions is important for disease diagnoses and treatment planning. Over the years, semi-supervised methods using pseudo labels have boosted the segmentation performance with limited labeled data and abundant unlabeled data. However, the unreliable targets in pseudo labels might lead to meaningless guidance for unlabeled data. In this paper, to solve this issue, we propose a novel confidence aware semi-supervised learning method based on a mean teacher scheme. Concretely, we design a confidence module to predict the model confidence guided by the True Class Probability. Then in the mean teacher framework, the student model gradually learns trustworthy targets from teacher model. To further improve the segmentation quality, we fine-tune the student model with reliable content in pseudo labels. We conduct extensive experiments on 2018 ISIC skin lesion segmentation dataset and our method outperforms other state-of-the-art semi-supervised approaches. Enmei Tu, Hao Zheng 0008, Yun Gu, Jie Yang 0002 |
ICASSP | 3 |
| 2021 | Refined Local-imbalance-based Weight for Airway Segmentation in CT
Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jiayuan Sun, Jie Yang 0002, Guang-Zhong Yang |
MICCAI (1) | 1 |
| 2021 | Learning Tubule-Sensitive CNNs for Pulmonary Airway and Artery-Vein Segmentation in CTabstractTraining convolutional neural networks (CNNs) for segmentation of pulmonary airway, artery, and vein is challenging due to sparse supervisory signals caused by the severe class imbalance between tubular targets and background. We present a CNNs-based method for accurate airway and artery-vein segmentation in non-contrast computed tomography. It enjoys superior sensitivity to tenuous peripheral bronchioles, arterioles, and venules. The method first uses a feature recalibration module to make the best use of features learned from the neural networks. Spatial information of features is properly integrated to retain relative priority of activated regions, which benefits the subsequent channel-wise recalibration. Then, attention distillation module is introduced to reinforce representation learning of tubular objects. Fine-grained details in high-resolution attention maps are passing down from one layer to its previous layer recursively to enrich context. Anatomy prior of lung context map and distance transform map is designed and incorporated for better artery-vein differentiation capacity. Extensive experiments demonstrated considerable performance gains brought by these components. Compared with state-of-the-art methods, our method extracted much more branches while maintaining competitive overall segmentation performance. Codes and models are available at http://www.pami.sjtu.edu.cn/News/56. Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu, Guang-Zhong Yang |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Alleviating Class-Wise Gradient Imbalance for Pulmonary Airway SegmentationabstractAutomated airway segmentation is a prerequisite for pre-operative diagnosis and intra-operative navigation for pulmonary intervention. Due to the small size and scattered spatial distribution of peripheral bronchi, this is hampered by a severe class imbalance between foreground and background regions, which makes it challenging for CNN-based methods to parse distal small airways. In this paper, we demonstrate that this problem is arisen by gradient erosion and dilation of the neighborhood voxels. During back-propagation, if the ratio of the foreground gradient to background gradient is small while the class imbalance is local, the foreground gradients can be eroded by their neighborhoods. This process cumulatively increases the noise information included in the gradient flow from top layers to the bottom ones, limiting the learning of small structures in CNNs. To alleviate this problem, we use group supervision and the corresponding WingsNet to provide complementary gradient flows to enhance the training of shallow layers. To further address the intra-class imbalance between large and small airways, we design a General Union loss function that obviates the impact of airway size by distance-based weights and adaptively tunes the gradient ratio based on the learning process. Extensive experiments on public datasets demonstrate that the proposed method can predict the airway structures with higher accuracy and better morphological completeness than the baselines. Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Learning Bronchiole-Sensitive Airway Segmentation CNNs by Feature Recalibration and Attention Distillation
Yulei Qin, Hao Zheng 0008, Yun Gu, Xiaolin Huang, Jie Yang 0002, Lihui Wang 0002, Yue Min Zhu |
MICCAI (1) | 2 |
| 2020 | Weakly Supervised Deep Learning for Breast Cancer Segmentation with Coarse Annotations
Hao Zheng 0008, Zhiguo Zhuang, Yulei Qin, Yun Gu, Jie Yang 0002, Guang-Zhong Yang |
MICCAI (4) | 1 |
| 2019 | AirwayNet: A Voxel-Connectivity Aware Approach for Accurate Airway Segmentation Using Convolutional Neural Networks
Yulei Qin, Hao Zheng 0008, Yun Gu, Mali Shen, Jie Yang 0002, Xiaolin Huang, Yue Min Zhu, Guang-Zhong Yang |
MICCAI (6) | 3 |
| 2019 | Varifocal-Net: A Chromosome Classification Approach Using Deep Convolutional NetworksabstractChromosome classification is critical for karyotyping in abnormality diagnosis. To expedite the diagnosis, we present a novel method named Varifocal-Net for simultaneous classification of chromosome's type and polarity using deep convolutional networks. The approach consists of one global-scale network (G-Net) and one local-scale network (L-Net). It follows three stages. The first stage is to learn both global and local features. We extract global features and detect finer local regions via the G-Net. By proposing a varifocal mechanism, we zoom into local parts and extract local features via the L-Net. Residual learning and multi-task learning strategies are utilized to promote high-level feature extraction. The detection of discriminative local parts is fulfilled by a localization subnet of the G-Net, whose training process involves both supervised and weakly supervised learning. The second stage is to build two multi-layer perceptron classifiers that exploit features of both two scales to boost classification performance. The third stage is to introduce a dispatch strategy of assigning each chromosome to a type within each patient case, by utilizing the domain knowledge of karyotyping. The evaluation results from 1909 karyotyping cases showed that the proposed Varifocal-Net achieved the highest accuracy per patient case (%) of 99.2 for both type and polarity tasks. It outperformed state-of-the-art methods, demonstrating the effectiveness of our varifocal mechanism, multi-scale feature ensemble, and dispatch strategy. The proposed method has been applied to assist practical karyotype diagnosis. Yulei Qin, Hao Zheng 0008, Xiaolin Huang, Jie Yang 0002, Yue Min Zhu, Lingqian Wu, Guang-Zhong Yang |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Simultaneous Accurate Detection of Pulmonary Nodules and False Positive Reduction Using 3D CNNsabstractAccurate detection of nodules in CT images is vital for lung cancer diagnosis, which greatly influences the patient's chance for survival. Motivated by successful application of convolutional neural networks (CNNs) on natural images, we propose a computer-aided diagnosis (CAD) system for simultaneous accurate pulmonary nodule detection and false positive reduction. To generate nodule candidates, we build a full 3D CNN model that employs 3D U-Net architecture as the backbone of a region proposal network (RPN). We adopt multi-task residual learning and online hard negative example mining strategy to accelerate the training process and improve the accuracy of nodule detection. Then, a 3D DenseNet-based model is presented to reduce false positive nodules. The densely connected structure reuses nodules' features and boosts feature propagation. Experimental results on LUNA16 datasets demonstrate the superior effectiveness of our approach over state-of-the-art methods. Yulei Qin, Hao Zheng 0008, Yue Min Zhu, Jie Yang 0002 |
ICASSP | 2 |
| 2018 | A Spatio-Temporal Fully Convolutional Network for Breast Lesion Segmentation in DCE-MRI
Hao Zheng 0008, Changsheng Lu, Enmei Tu, Jie Yang 0002, Nikola K. Kasabov |
ICONIP (7) | 2 |
| 2018 | Small Lesion Classification in Dynamic Contrast Enhancement MRI for Breast Cancer Early Detection
Hao Zheng 0008, Yun Gu, Yulei Qin, Xiaolin Huang, Jie Yang 0002, Guang-Zhong Yang |
MICCAI (2) | 1 |