EDBT 2026 Demo / reviewers in the wild / expert
Huimin Huang 0002
dblp:03/8493-2
· DBLP profile ↗
24ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0001-8987-350XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang |
Pattern Recognit. | 6 |
| 2025 | NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationabstractNight-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nighttime scenes with varying illumination. In light of this issue, we focus on investigating generalizable paradigms for night-time scene segmentation and propose an efficient fine-tuning scheme, dubbed NightAdapter, alleviating the domain gap across various scenes. Interestingly, different properties embedded in the day-time and night-time features can be characterized by the bands after discrete sine transform, which can be categorized into illumination-sensitive/-insensitive bands. Hence, our NightAdapter is powered by two appealing designs: (1) Illumination-Insensitive Band Adaptation that provides a foundation for understanding the prior, enhancing the robustness to illumination shifts; (2) Illumination-Sensitive Band Adaptation that fine-tunes the randomized frequency bands, mitigating the domain gap between the day-time and various night-time scenes. As a consequence, illumination-insensitive enhancement improves the domain invariance, while illumination-sensitive diminution strengthens the domain shift between different scenes. NightAdapter yields significant improvements over the state-of-the-art methods under various day-to-night, night-to-night, and in-domain night segmentation experiments. Source code is available at https://github.com/BiQiWHU/NightAdapter. Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
CVPR | 3 |
| 2025 | A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICCV | 3 |
| 2025 | D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (5) | 6 |
| 2025 | BrainSegDMIF: A Dynamic Fusion-enhanced SAM for Brain Lesion Segmentation
Hongming Wang, Huimin Huang 0002, Jiaxuan Jiang 0001, Hao Zheng 0008, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001, Jinping Xu |
ACM Multimedia | 3 |
| 2025 | AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 4 |
| 2025 | Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image RestorationabstractImage restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze).
Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images.
To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration.
Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution,
while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process.
A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised.
Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process.
A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model.
Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts. Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 4 |
| 2025 | Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationabstractDomain generalization aims to train models that perform robustly on unseen target domains without access to target data.
The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability.
However, the static alignment to class-level textual anchors remains insufficient to handle the dramatic distribution discrepancy from diverse domain-specific visual features.
In this work, we propose a novel cross-domain Schrödinger Bridge (SB) method, namely SBGen, to handle this challenge, which explicitly formulates the stochastic semantic evolution, to gain better generalization to unseen domains.
Technically, the proposed \texttt{SBGen} consists of three key components: (1) \emph{text-guided domain-aware feature selection} to isolate semantically aligned image tokens; (2) \emph{stochastic cross-domain evolution} to simulate the SB dynamics via a learnable time-conditioned drift; and (3) \emph{stochastic domain-agnostic interpolation} to construct semantically grounded feature trajectories.
Empirically, \texttt{SBGen} achieves state-of-the-art performance on domain generalization in both classification and segmentation. This work highlights the importance of modeling domain shifts as structured stochastic processes grounded in semantic alignment. Hao Zheng 0008, Jingjun Yi, Qi Bi, Huimin Huang 0002, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 4 |
| 2025 | Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions
Yawen Huang, Huimin Huang 0002, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image SegmentationabstractSemi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 1 |
| 2024 | Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsabstractMulti-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics. Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001 |
CVPR | 1 |
| 2023 | ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image SegmentationabstractVision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001 |
AAAI | 1 |
| 2023 | SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic SegmentationabstractSemi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001 |
CVPR | 1 |
| 2023 | MedFCT: A Frequency Domain Joint CNN-Transformer Network for Semi-supervised Medical Image SegmentationabstractSemi-supervised learning(SSL) is a data-efficient way in leveraging large-scale data without annotations and alleviating the dependence on labeled data. Mean-Teacher (MT) scheme with teacher-student model architecture has shown its effectiveness in semi-supervised medical image segmentation, where the student network learns from the teacher by minimizing pixel-wise consistency loss. However, existing MT-based SSLs still give rise to two main concerns: (1) limited learning ability of student network that neglects the union of local feature and global cues extraction which may impact the representation learning of variable objects. (2) limited knowledge-transferring ability of teacher network with only pixel-level consistency regularization that may result in inadequate and unstable guidance. To address these limitations, we propose a novel semi-supervised learning scheme, namely MedFCT, with two appealing designs: (1) A dual student architecture with parallel CNN and Transformer branches is designed for local-global feature extraction, where the full-frequency interaction between CNN and Transformer can be explored by a frequency domain cross-fusion (FDCF) module to learn complementarity of the two-paradigm features. (2) A comprehensive multi-level consistency regularization considering pixel-wise, feature-wise and class-wise information is presented to realize more effective guidance and knowledge transfer from teacher network. Experiments show that MedFCT outperforms previous state-of-the-art methods on two public medical image segmentation benchmarks. Shiao Xie, Huimin Huang 0002, Ziwei Niu, Lanfen Lin, Yen-Wei Chen 0001 |
ICME | 2 |
| 2023 | Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image SegmentationabstractSemi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 1 |
| 2023 | IS2Net: Intra-domain Semantic and Inter-domain Style Enhancement for Semi-supervised Medical Domain GeneralizationabstractDomain generalization (DG) demonstrates superior generalization ability in cross-center medical image segmentation. Despite its great success, existing fully supervised DG methods require collecting a large quantity of pixel-level annotations which is quite expensive and time-consuming. To address this challenge, several semi-supervised domain generalized (SSDG) methods have been proposed by simply coupling semi-supervised learning (SSL) with DG tasks, which give rise to two main concerns: (1) Intra-domain dubious semantic information: the quality of pseudo labels in each source domain suffers from the limited amount of labeled data and cross-domain discrepancy. (2) Inter-domain intangible style relationship: current models fail in integrating domain-level information and overlook the relationships among different domains, which degrades the generalization ability of model. In light of these two issues, we propose a novel SSDG framework, namely IS2Net, by arranging an inter-domain generalization branch and several intra-domain SSL branches in a parallel manner, powered by two appealing designs that build a positive interaction between them: (1) A style and semantic memory mechanism is designed to provide both high-quality class-wise representations for intra-domain semantic enhancement and stable domain-specific knowledge for inter-domain style relationship construction. (2) Confident pseudo labeling strategy aims at generating more reliable supervision for intra and inter domain branches, and thus facilitating the learning process of the whole framework. Extensive experiments show that IS2Net yields consistent improvements over the state-of- the-art methods in three public benchmarks. Shiao Xie, Ziwei Niu, Huimin Huang 0002, Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin |
ACM Multimedia | 3 |
| 2022 | ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image SegmentationabstractRecently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IJCAI | 1 |
| 2022 | MTL-ABS3Net: Atlas-Based Semi-Supervised Organ Segmentation Network With Multi-Task Learning for Medical ImagesabstractOrgan segmentation is one of the most important step for various medical image analysis tasks. Recently, semi-supervised learning (SSL) has attracted much attentions by reducing labeling cost. However, most of the existing SSLs neglected the prior shape and position information specialized in the medical images, leading to unsatisfactory localization and non-smooth of objects. In this paper, we propose a novel atlas-based semi-supervised segmentation network with multi-task learning for medical organs, named MTL-ABS3Net, which incorporates the anatomical priors and makes full use of unlabeled data in a self-training and multi-task learning manner. The MTL-ABS3Net consists of two components: an Atlas-Based Semi-Supervised Segmentation Network (ABS3Net) and Reconstruction-Assisted Module (RAM). Specifically, the ABS3Net improves the existing SSLs by utilizing atlas prior, which generates credible pseudo labels in a self-training manner; while the RAM further assists the segmentation network by capturing the anatomical structures from the original images in a multi-task learning manner. Better reconstruction quality is achieved by using MS-SSIM loss function, which further improves the segmentation accuracy. Experimental results from the liver and spleen datasets demonstrated that the performance of our method was significantly improved compared to existing state-of-the-art methods. Huimin Huang 0002, Qingqing Chen 0001, Lanfen Lin, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Akira Furukawa, Shuzo Kanasaki, Yen-Wei Chen 0001, Ruofeng Tong 0001, Hongjie Hu |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Graph-Based Pyramid Global Context Reasoning With a Saliency- Aware Projection for Covid-19 Lung Infections SegmentationabstractCoronavirus Disease 2019 (COVID-19) has rapidly spread in 2020, emerging a mass of studies for lung infection segmentation from CT images. Though many methods have been proposed for this issue, it is a challenging task because of infections of various size appearing in different lobe zones. To tackle these issues, we propose a Graph-based Pyramid Global Context Reasoning (Graph-PGCR) module, which is capable of modeling long-range dependencies among disjoint infections as well as adapt size variation. We first incorporate graph convolution to exploit long-term contextual information from multiple lobe zones. Different from previous average pooling or maximum object probability, we propose a saliency-aware projection mechanism to pick up infection-related pixels as a set of graph nodes. After graph reasoning, the relation-aware features are reversed back to the original coordinate space for the down-stream tasks. We further construct multiple graphs with different sampling rates to handle the size variation problem. To this end, distinct multi-scale long-range contextual patterns can be captured. Our Graph- PGCR module is plug-and-play, which can be integrated into any architecture to improve its performance. Experiments demonstrated that the proposed method consistently boost the performance of state-of-the-art backbone architectures on both of public and our private COVID-19 datasets. Huimin Huang 0002, Lanfen Lin, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICASSP | 1 |
| 2021 | Graph-BAS3Net: Boundary-Aware Semi-Supervised Segmentation Network with Bilateral Graph ConvolutionabstractSemi-supervised learning (SSL) algorithms have attracted much attentions in medical image segmentation by leveraging unlabeled data, which challenge in acquiring massive pixel-wise annotated samples. However, most of the existing SSLs neglected the geometric shape constraint in object, leading to unsatisfactory boundary and non-smooth of object. In this paper, we propose a novel boundary-aware semi-supervised medical image segmentation network, named Graph-BAS3Net, which incorporates the boundary information and learns duality constraints between semantics and geometrics in the graph domain. Specifically, the proposed method consists of two components: a multi-task learning framework BAS3Net and a graph-based cross-task module BGCM. The BAS3Net improves the existing GAN-based SSL by adding a boundary detection task, which encodes richer features of object shape and surface. Moreover, the BGCM further explores the co-occurrence relations between the semantics segmentation and boundary detection task, so that the network learns stronger semantic and geometric correspondences from both labeled and unlabeled data. Experimental results on the LiTS dataset and COVID-19 dataset confirm that our proposed Graph-BAS3Net outperforms the state-of-the-art methods in semi-supervised segmentation task. Huimin Huang 0002, Lanfen Lin, Yue Zhang 0042, Xiongwei Mao, Xiaohan Qian, Zhiyi Peng, Jianying Zhou 0006, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
ICCV | 1 |
| 2021 | 3D Graph-S2Net: Shape-Aware Self-ensembling Network for Semi-supervised Segmentation with Bilateral Graph Convolution
Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
MICCAI (2) | 1 |
| 2021 | Multi-phase Liver Tumor Segmentation with Spatial Aggregation and Uncertain Region Inpainting
Yue Zhang 0042, Chengtao Peng, Liying Peng, Huimin Huang 0002, Ruofeng Tong 0001, Lanfen Lin, Jingsong Li 0001, Yen-Wei Chen 0001, Qingqing Chen 0001, Hongjie Hu, Zhiyi Peng |
MICCAI (1) | 4 |
| 2021 | Medical Image Segmentation With Deep Atlas PriorabstractOrgan segmentation from medical images is one of the most important pre-processing steps in computer-aided diagnosis, but it is a challenging task because of limited annotated data, low-contrast and non-homogenous textures. Compared with natural images, organs in the medical images have obvious anatomical prior knowledge (e.g., organ shape and position), which can be used to improve the segmentation accuracy. In this paper, we propose a novel segmentation framework which integrates the medical image anatomical prior through loss into the deep learning models. The proposed prior loss function is based on probabilistic atlas, which is called as deep atlas prior (DAP). It includes prior location and shape information of organs, which are important prior information for accurate organ segmentation. Further, we combine the proposed deep atlas prior loss with the conventional likelihood losses such as Dice loss and focal loss into an adaptive Bayesian loss in a Bayesian framework, which consists of a prior and a likelihood. The adaptive Bayesian loss dynamically adjusts the ratio of the DAP loss and the likelihood loss in the training epoch for better learning. The proposed loss function is universal and can be combined with a wide variety of existing deep segmentation models to further enhance their performance. We verify the significance of our proposed framework with some state-of-the-art models, including fully-supervised and semi-supervised segmentation models on a public dataset (ISBI LiTS 2017 Challenge) for liver segmentation and a private dataset for spleen segmentation. Huimin Huang 0002, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2020 | UNet 3+: A Full-Scale Connected UNet for Medical Image SegmentationabstractRecently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version. Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001 |
ICASSP | 1 |