Yuexiang Li

dblp:165/6204 · DBLP profile ↗
← Back
93ranked-venue papers
14as first author
74since 2021 · last 2026
0000-0001-8076-2619ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 43 since 2021Artificial intelligence and machine learning · 41 · 3 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 41 · 9 first-author · 29 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Open-world Weakly-Supervised Object Localization
Jinheng Xie, Zhaochuan Luo, Rouyi Li, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Yang Zhang 0012, LinLin Shen, Zheng Shou 0001
Pattern Recognit.6
2026 Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang
Pattern Recognit.2
2026 MUSCLE: A New Perspective to Multi-Scale Fusion for Medical Image Classification Based on the Theory of Evidence
abstract
In the field of medical image analysis, medical image classification is one of the most fundamental and critical tasks. Current researches often rely on the off-the-shelf backbone networks derived from the field of computer vision, hoping to achieve satisfactory classification performance for medical images. However, given the characteristics of medical images, such as scattered distribution and varying sizes of lesions, features extracted with a single scale from the existing backbones often fail to perform accurate medical image classification. To this end, we propose a novel multi-scale learning paradigm, namely MUlti-SCale Learning with trusted Evidences (MUSCLE), which extracts and integrates features from different scales based on shape the theory of evidence, to generate the more comprehensive feature representation for the medical image classification task. Particularly, the proposed MUSCLE first estimates the uncertainties of features extracted from different scales/stages of the classification backbone as the evidences, and accordingly form the opinions regarding to the feature trustworthiness via a set of evidential deep neural networks. Then, these opinions on different scales of features are ensembled to yield an aggregated opinion, which can be used to adaptively tune the weights of multi-scale features for scatteredly distributed and size-varying lesions, and consequently improve the network capacity for accurate medical image classification. Our MUSCLE paradigm has been evaluated on five publicly available medical image datasets. The experimental results show that the proposed MUSCLE not only improves the accuracy of the original backbone network, but also enhances the reliability and interpretability of model decisions with the trusted evidences (https://github.com/Q4CS/MUSCLE).
Junlai Qiu, Junyue Cao, Yawen Huang, Ziwei Zhu 0005, Fubo Wang, Cheng Lu 0001, Yuexiang Li, Yefeng Zheng 0001
IEEE Trans. Medical Imaging7
2026 SORT-LFR: Revisiting SORT for Multi-Object Tracking in Low-Frame-Rate Videos
abstract
For certain applications like highway surveillance systems, only low-frame-rate videos are recorded, which presents a huge challenge to existing trackers, as objects tend to undergo far more abrupt changes in location, motion, and appearance between successive frames compared to normal frame rates. To handle the above challenges, we propose a novel approach, namely$\mathbb {SORT}$-$\mathbb {LFR}$, for$\mathbb {S}$imple$\mathbb {O}$nline and$\mathbb {R}$ealtime$\mathbb {T}$racking in$\mathbb {L}$ow-$\mathbb {F}$rame-$\mathbb {R}$ate videos, which consists of following techniques: 1) A feature-prior association strategy to improve the capability to track new objects with significant displacements; 2) A Kalman filter using acceleration in state space (accel-fused Kalman filter) to improve the motion estimation capability for non-constant velocity moving objects; 3) A detection-guided adaptive exponential moving average (DG-AEMA) feature update mechanism to enhance feature temporal modeling capability for tracked objects; 4) A trajectory-covariance threshold tuning (TCTT) method to filter out incorrect association results. Through these techniques, the proposed SORT achieves 91.8 HOTA, 92.6 MOTA and 93.9 IDF1, which surpass all state-of-the-art trackers on the public CityFlow and our private HighwayTrack datasets under the low-frame-rate setting.
Yawen Huang, Yubei Lin, Ziwei Zhu 0005, Xingming Zhang 0001, Yang Liu 0182, Yuexiang Li, Yefeng Zheng 0001
IEEE Trans. Multim.7
2025 DGFamba: Learning Flow Factorized State Space for Visual Domain Generalization
abstract
Domain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DGFamba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li
AAAI7
2025 S³-Mamba: Small-Size-Sensitive Mamba for Lesion Segmentation
abstract
Small lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small lesions. To tackle the challenges, we propose a Small-Size-Sensitive Mamba (S³-Mamba), which promotes the sensitivity to small lesions across three dimensions: channel, spatial, and training strategy. Specifically, an Enhanced Visual State Space block is designed to focus on small lesions through multiple residual connections to preserve local features, and selectively amplify important details while suppressing irrelevant ones through channel-wise attention. A Tensor-based Cross-feature Multi-scale Attention is designed to integrate input image features and intermediate-layer features with edge features and exploit the attentive support of features across multiple scales, thereby retaining spatial details of small lesions at various granularities. Finally, we introduce a novel regularized curriculum learning to automatically assess lesion size and sample difficulty, and gradually focus from easy samples to hard ones like small lesions. Extensive experiments on three medical image segmentation datasets show the superiority of our S³-Mamba, especially in segmenting small lesions.
Gui Wang, Yuexiang Li, Wenting Chen, Meidan Ding, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, LinLin Shen
AAAI2
2025 NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene Segmentation
abstract
Night-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nighttime scenes with varying illumination. In light of this issue, we focus on investigating generalizable paradigms for night-time scene segmentation and propose an efficient fine-tuning scheme, dubbed NightAdapter, alleviating the domain gap across various scenes. Interestingly, different properties embedded in the day-time and night-time features can be characterized by the bands after discrete sine transform, which can be categorized into illumination-sensitive/-insensitive bands. Hence, our NightAdapter is powered by two appealing designs: (1) Illumination-Insensitive Band Adaptation that provides a foundation for understanding the prior, enhancing the robustness to illumination shifts; (2) Illumination-Sensitive Band Adaptation that fine-tunes the randomized frequency bands, mitigating the domain gap between the day-time and various night-time scenes. As a consequence, illumination-insensitive enhancement improves the domain invariance, while illumination-sensitive diminution strengthens the domain shift between different scenes. NightAdapter yields significant improvements over the state-of-the-art methods under various day-to-night, night-to-night, and in-domain night segmentation experiments. Source code is available at https://github.com/BiQiWHU/NightAdapter.
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
CVPR7
2025 A Simple Data Augmentation for Feature Distribution Skewed Federated Learning
abstract
Federated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribution skewed FL scenario, a common non-IID situation in real-world applications where data from different clients exhibit varying underlying distributions. This variation leads to feature shift, which is a key issue of this scenario. While previous works have made notable progress, few pay attention to the data itself, i.e., the root of this issue. The primary goal of this paper is to mitigate feature shift from the perspective of data. To this end, we propose a simple yet remarkably effective input-level data augmentation method, namely FedRDN, which randomly injects the statistical information of the local distribution from the entire federation into the client’s data. This is beneficial to improve the generalization of local feature representations, thereby mitigating feature shift. Moreover, our FedRDN is a plug-and-play component, which can be seamlessly integrated into the data augmentation flow with only a few lines of code. Extensive experiments on several datasets show that the performance of various representative FL methods can be further improved by integrating our FedRDN, demonstrating its effectiveness, strong compatibility and generalizability. Code is available at https://github.com/IAMJackYan/FedRDN.
Yunlu Yan, Huazhu Fu, Yuexiang Li, Jinheng Xie, Jun Ma 0008, Guang Yang 0006, Lei Zhu 0003
CVPR3
2025 A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ICCV8
2025 Temporal Model-Based Federated Active Medical Image Classification
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Jinheng Xie, Jun Chen 0005, Mohamed Elhoseiny 0001, Kaishun Wu, Lei Zhu 0003
MICCAI (14)3
2025 D-CAM: Learning Generalizable Weakly-Supervised Medical Image Segmentation from Domain-Invariant CAM
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Huimin Huang 0002, Yuexiang Li, Shaoxin Li 0001, Xian Wu 0001, Yefeng Zheng 0001, Feiyue Huang
MICCAI (5)7
2025 AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
ACM Multimedia9
2025 Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image Restoration
abstract
Image restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze). Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images. To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration. Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution, while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process. A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised. Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process. A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model. Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts.
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
NeurIPS9
2025 Learning a Cross-Modal Schrödinger Bridge for Visual Domain Generalization
abstract
Domain generalization aims to train models that perform robustly on unseen target domains without access to target data. The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability. However, the static alignment to class-level textual anchors remains insufficient to handle the dramatic distribution discrepancy from diverse domain-specific visual features. In this work, we propose a novel cross-domain Schrödinger Bridge (SB) method, namely SBGen, to handle this challenge, which explicitly formulates the stochastic semantic evolution, to gain better generalization to unseen domains. Technically, the proposed \texttt{SBGen} consists of three key components: (1) \emph{text-guided domain-aware feature selection} to isolate semantically aligned image tokens; (2) \emph{stochastic cross-domain evolution} to simulate the SB dynamics via a learnable time-conditioned drift; and (3) \emph{stochastic domain-agnostic interpolation} to construct semantically grounded feature trajectories. Empirically, \texttt{SBGen} achieves state-of-the-art performance on domain generalization in both classification and segmentation. This work highlights the importance of modeling domain shifts as structured stochastic processes grounded in semantic alignment.
Hao Zheng 0008, Jingjun Yi, Qi Bi, Huimin Huang 0002, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001
NeurIPS7
2025 Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions
Yawen Huang, Huimin Huang 0002, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Yefeng Zheng 0001
Int. J. Comput. Vis.4
2025 Learning Generalized Medical Image Representation by Decoupled Feature Queries
abstract
Medical images are usually collected from multiple clinical centers with various types of scanners. When confronted with such significant cross-domain distribution discrepancy, a deep network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. Such channel redundancy limits the expressive capability of a representation, resulting in less preferable generalization ability. To address this fundamental yet challenging issue, we propose a novel decoupled feature as query (DFQ) framework for domain generalized medical image representation learning. Its general idea is to leverage the channel-wise decoupled deep features as queries. Particularly, a deep instance whitening transform with restricted isometry is proposed, which enforces each channel orthogonal to the rest channels after decoupling. Besides, the long-range dependency between decoupled deep and shallow features is implicitly constrained to minimize channel redundancy throughout training. Extensive experiments show its state-of-the-art performance on three medical domain generalization tasks with four modalities.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 GAD: Domain generalized diabetic retinopathy grading by grade-aware de-stylization
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001
Pattern Recognit.7
2025 Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI Reconstruction
abstract
While multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities).
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003
IEEE J. Biomed. Health Informatics3
2024 Learning Generalized Medical Image Segmentation from Decoupled Feature Queries
abstract
Domain generalized medical image segmentation requires models to learn from multiple source domains and generalize well to arbitrary unseen target domain. Such a task is both technically challenging and clinically practical, due to the domain shift problem (i.e., images are collected from different hospitals and scanners). Existing methods focused on either learning shape-invariant representation or reaching consensus among the source domains. An ideal generalized representation is supposed to show similar pattern responses within the same channel for cross-domain images. However, to deal with the significant distribution discrepancy, the network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. To address this issue, we propose to leverage channel-wise decoupled deep features as queries. With the aid of cross-attention mechanism, the long-range dependency between deep and shallow features can be fully mined via self-attention and then guides the learning of generalized representation. Besides, a relaxed deep whitening transformation is proposed to learn channel-wise decoupled features in a feasible way. The proposed decoupled fea- ture query (DFQ) scheme can be seamlessly integrate into the Transformer segmentation model in an end-to-end manner. Extensive experiments show its state-of-the-art performance, notably outperforming the runner-up by 1.31% and 1.98% with DSC metric on generalized fundus and prostate benchmarks, respectively. Source code is available at https://github.com/BiQiWHU/DFQ.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
AAAI6
2024 Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
AAAI7
2024 Going Beyond Multi-Task Dense Prediction with Synergy Embedding Models
abstract
Multi-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics.
Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001
CVPR7
2024 Self-Supervised Cross-Level Consistency Learning For Fundus Image Classification
abstract
The rapid development of intelligent systems for eye disease diagnosis decreases the risk of people suffering from vision impairment. However, the superior discrimination ability of existing retinal disease diagnosis methods heavily relies on the large-scale high-quality annotations. In this work, we adapt the self-supervised technique for fundus image classification with the merits of bypassing the over-dependence of labeled data. Unlike most current self-supervised approaches, which only learn global pre-text representations from view-level, our method further incorporates the region-level representations into the learning process, since the pathological changes in fundus images are usually subtle and scattered. Specifically, we propose a novel self-supervised cross-level consistency learning scheme (S2C2L), which leverages both view-level and region-level representations of a vision Transformer to improve the robustness of extracted self-supervised representation. A diagnosis perception module (DPM) is constructed to enhance the activation of local pathological regions from both region and view levels, and a cross-level consistency loss is dedicated to align the representations from both levels. Extensive experiments on iChallenge-AMD, LAG and APTOS2019 datasets validate the state-of-the-art performance of our method for three common eye diseases.
Qi Bi, Hao Zheng 0008, Xu Sun 0006, Jingjun Yi, Wentian Zhang, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ICASSP7
2024 A New Perspective to Boost Performance Fairness For Medical Federated Learning
Yunlu Yan, Lei Zhu 0003, Yuexiang Li, Xinxing Xu, Rick Siow Mong Goh, Yong Liu 0026, Salman Khan 0001, Chun-Mei Feng 0001
MICCAI (10)3
2024 Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang
MICCAI (10)8
2024 Learning Spectral-Decomposited Tokens for Domain Generalized Semantic Segmentation
abstract
The rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as the cross-domain images share common pixel-wise content information but vary greatly in terms of the style. In this paper, we present a novel Spectral-dEcomposed Token (SET) learning framework to advance the frontier. Delving into further than existing fine-tuning token & frozen backbone paradigm, the proposed SET especially focuses on the way learning style-invariant features from these learnable tokens. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction.After the decomposition, style variation primarily impacts the token-based feature enhancement within the amplitude branch. To address this issue, we further develop an attention optimization method to bridge the gap between style-affected representation and static tokens during inference. Extensive cross-domain experiments show its state-of-the-art performance.
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ACM Multimedia7
2024 Samba: Severity-aware Recurrent Modeling for Cross-domain Medical Image Grading
abstract
Disease grading is a crucial task in medical image analysis. Due to the continuous progression of diseases, i.e., the variability within the same level and the similarity between adjacent stages, accurate grading is highly challenging. Furthermore, in real-world scenarios, models trained on limited source domain datasets should also be capable of handling data from unseen target domains. Due to the cross-domain variants, the feature distribution between source and unseen target domains can be dramatically different, leading to a substantial decrease in model performance. To address these challenges in cross-domain disease grading, we propose a Severity-aware Recurrent Modeling (Samba) method in this paper. As the core objective of most staging tasks is to identify the most severe lesions, which may only occupy a small portion of the image, we propose to encode image patches in a sequential and recurrent manner. Specifically, a state space model is tailored to store and transport the severity information by hidden states. Moreover, to mitigate the impact of cross-domain variants, an Expectation-Maximization (EM) based state recalibration mechanism is designed to map the patch embeddings into a more compact space. We model the feature distributions of different lesions through the Gaussian Mixture Model (GMM) and reconstruct the intermediate features based on learnable severity bases. Extensive experiments show the proposed Samba outperforms the VMamba baseline by an average accuracy of 23.5\%, 5.6\% and 4.1\% on the cross-domain grading of fatigue fracture, breast cancer and diabetic retinopathy, respectively. Source code is available at \url{https://github.com/BiQiWHU/Samba}.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Haolan Zhan, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
NeurIPS7
2024 Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic Segmentation
abstract
The emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images. Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, i.e., the variance of urban landscape and environment dependencies. Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS. The frequency space after Haar wavelet transformation provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information are retained in the low- and high- frequency components of the space, respectively. To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier. Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process. Particularly, the proposed FADA consists of two branches, i.e., low- and high- frequency branches. The former one is able to stabilize the scene content, while the latter one learns the scene styles and eliminates its impact to DGSS. Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs. Source code is available at \url{https://github.com/BiQiWHU/FADA}.
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001
NeurIPS7
2024 Triplet-branch network with contrastive prior-knowledge embedding for disease grading
Yuexiang Li, Yawen Huang, Jingxin Liu 0005, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001
Artif. Intell. Medicine1
2024 Multi-Constraint Transferable Generative Adversarial Networks for Cross-Modal Brain Image Synthesis
Yawen Huang, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Guo-Jun Qi, Ling Shao 0001, Yefeng Zheng 0001
Int. J. Comput. Vis.3
2024 IDaTPA: importance degree based thread partitioning approach in thread level speculation
abstract
Abstract As an auto-parallelization technique with the level of thread on multi-core, Thread-Level Speculation (TLS) which is also called Speculative Multithreading (SpMT), partitions programs into multiple threads and speculatively executes them under conditions of ambiguous data and control dependence. Thread partitioning approach plays a key role to the performance enhancement in TLS. The existing heuristic rules-based approach (HR-based approach) which is an one-size-fits-all strategy, can not guarantee to achieve the satisfied thread partitioning. In this paper, an importancedegree basedthreadpartitioningapproach (IDaTPA) is proposed to realize the partition of irregular programs into multithreads. IDaTPA implements biasing partitioning for every procedure with a machine learning method. It mainly includes: constructing sample set, expression of knowledge, calculation of similarity, prediction model and the partition of the irregular programs is performed by the prediction model. Using IDaTPA, the subprocedures in unseen irregular programs can obtain their satisfied partition. On a generic SpMT processor (called Prophet) to perform the performance evaluation for multithreaded programs, the IDaTPA is evaluated and averagely delivers a speedup of 1.80 upon a 4-core processor. Furthermore, in order to obtain the portability evaluation of IDaTPA, we port IDaTPA to 8-core processor and obtain a speedup of 2.82 on average. Experiment results show that IDaTPA obtains a significant speedup increasement and Olden benchmarks respectively deliver a 5.75% performance improvement on 4-core and a 6.32% performance improvement on 8-core, and SPEC2020 benchmarks obtain a 38.20% performance improvement than the conventional HR-based approach.
Yuexiang Li, Zhiyong Zhang 0002, Xinyong Wang, Shuaina Huang, Yaning Su
Discov. Comput.1
2024 Improving vision transformer for medical image classification via token-wise perturbation
Yuexiang Li, Yawen Huang, Nanjun He, Kai Ma 0002, Yefeng Zheng 0001
J. Vis. Commun. Image Represent.1
2024 Anomaly detection via gating highway connection for retinal fundus images
Wentian Zhang, Jinheng Xie, Yawen Huang, Yu Zhang 0185, Yuexiang Li, Ramachandra Raghavendra, Yefeng Zheng 0001
Pattern Recognit.6
2024 Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-Spoofing
abstract
Face recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols.
Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo
IEEE Trans. Circuits Syst. Video Technol.5
2024 Cross-Modal Vertical Federated Learning for MRI Reconstruction
abstract
Federated learning enables multiple hospitals to cooperatively learn a shared model without privacy disclosure. Existing methods often take a common assumption that the data from different hospitals have the same modalities. However, such a setting is difficult to fully satisfy in practical applications, since the imaging guidelines may be different between hospitals, which makes the number of individuals with the same set of modalities limited. To this end, we formulate this practical-yet-challenging cross-modal vertical federated learning task, in which data from multiple hospitals have different modalities with a small amount of multi-modality data collected from the same individuals. To tackle such a situation, we develop a novel framework, namely Federated Consistent Regularization constrained Feature Disentanglement (Fed-CRFD), for boosting MRI reconstruction by effectively exploring the overlapping samples (i.e., same patients with different modalities at different hospitals) and solving the domain shift problem caused by different modalities. Particularly, our Fed-CRFD involves an intra-client feature disentangle scheme to decouple data into modality-invariant and modality-specific features, where the modality-invariant features are leveraged to mitigate the domain shift problem. In addition, a cross-client latent representation consistency constraint is proposed specifically for the overlapping samples to further align the modality-invariant features extracted from different modalities. Hence, our method can fully exploit the multi-source data from hospitals while alleviating the domain shift problem. Extensive experiments on two typical MRI datasets demonstrate that our network clearly outperforms state-of-the-art MRI reconstruction methods.
Yunlu Yan, Hong Wang 0021, Yawen Huang, Nanjun He, Lei Zhu 0003, Yong Xu 0001, Yuexiang Li, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics7
2024 Unsupervised Domain Adaptation for Medical Image Segmentation by Disentanglement Learning and Self-Training
abstract
Unsupervised domain adaption (UDA), which aims to enhance the segmentation performance of deep models on unlabeled data, has recently drawn much attention. In this paper, we propose a novel UDA method (namely DLaST) for medical image segmentation via disentanglement learning and self-training. Disentanglement learning factorizes an image into domain-invariant anatomy and domain-specific modality components. To make the best of disentanglement learning, we propose a novel shape constraint to boost the adaptation performance. The self-training strategy further adaptively improves the segmentation performance of the model for the target domain through adversarial learning and pseudo label, which implicitly facilitates feature alignment in the anatomy space. Experimental results demonstrate that the proposed method outperforms the state-of-the-art UDA methods for medical image segmentation on three public datasets, i.e., a cardiac dataset, an abdominal dataset and a brain dataset. The code will be released soon.
Qingsong Xie, Yuexiang Li, Nanjun He, Munan Ning, Kai Ma 0002, Guoxing Wang, Yong Lian 0001, Yefeng Zheng 0001
IEEE Trans. Medical Imaging2
2024 Adversarial Medical Image With Hierarchical Feature Hiding
abstract
Deep learning based methods for medical images can be easily compromised by adversarial examples (AEs), posing a great security flaw in clinical decision-making. It has been discovered that conventional adversarial attacks like PGD which optimize the classification logits, are easy to distinguish in the feature space, resulting in accurate reactive defenses. To better understand this phenomenon and reassess the reliability of the reactive defenses for medical AEs, we thoroughly investigate the characteristic of conventional medical AEs. Specifically, we first theoretically prove that conventional adversarial attacks change the outputs by continuously optimizing vulnerable features in a fixed direction, thereby leading to outlier representations in the feature space. Then, a stress test is conducted to reveal the vulnerability of medical images, by comparing with natural images. Interestingly, this vulnerability is a double-edged sword, which can be exploited to hide AEs. We then propose a simple-yet-effective hierarchical feature constraint (HFC), a novel add-on to conventional white-box attacks, which assists to hide the adversarial feature in the target feature distribution. The proposed method is evaluated on three medical datasets, both 2D and 3D, with different modalities. The experimental results demonstrate the superiority of HFC,i.e., it bypasses an array of state-of-the-art adversarial medical AE detectors more efficiently than competing adaptive attacks1, which reveals the deficiencies of medical reactive defense and allows to develop more robust defenses in future.
Qingsong Yao, Zecheng He, Yuexiang Li, Yi Lin 0009, Kai Ma 0002, Yefeng Zheng 0001, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging3
2024 Relational Experience Replay: Continual Learning by Adaptively Tuning Task-Wise Relationship
abstract
Continual learning is a promising machine learning paradigm to learn new tasks while retaining previously learned knowledge over streaming training data. Till now,rehearsal-basedmethods, keeping a small part of data from old tasks as a memory buffer, have shown good performance in mitigating catastrophic forgetting for previously learned knowledge. However, most of these methods typically treat each new task equally, which may not adequately consider the relationship or similarity between old and new tasks. Furthermore, these methods commonly neglect sample importance in the continual training process and result in sub-optimal performance on certain tasks. To address this challenging problem, we propose Relational Experience Replay (RER), a bi-level learning framework, to adaptively tune task-wise relationships and sample importance within each task to achieve a better ‘stability’ and ‘plasticity’ trade-off. As such, the proposed method is capable of accumulating new knowledge while consolidating previously learned old knowledge during continual learning. Extensive experiments conducted on three benchmark image datasets (CIFAR-10, CIFAR-100, and Tiny ImageNet) and two text datasets (20News and DBpedia) show that the proposed method can consistently improve the performance of all baselines and surpass current state-of-the-art methods.
Quanziang Wang, Renzhen Wang, Yuexiang Li, Dong Wei 0004, Hong Wang 0021, Kai Ma 0002, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Multim.3
2024 RCDNet: An Interpretable Rain Convolutional Dictionary Network for Single Image Deraining
abstract
As common weather, rain streaks adversely degrade the image quality and tend to negatively affect the performance of outdoor computer vision systems. Hence, removing rains from an image has become an important issue in the field. To handle such an ill-posed single image deraining task, in this article, we specifically build a novel deep architecture, called rain convolutional dictionary network (RCDNet), which embeds the intrinsic priors of rain streaks and has clear interpretability. In specific, we first establish a rain convolutional dictionary (RCD) model for representing rain streaks and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. By unfolding it, we then build the RCDNet in which every network module has clear physical meanings and corresponds to each operation involved in the algorithm. This good interpretability greatly facilitates an easy visualization and analysis of what happens inside the network and why it works well in the inference process. Moreover, taking into account the domain gap issue in real scenarios, we further design a novel dynamic RCDNet, where the rain kernels can be dynamically inferred corresponding to input rainy images and then help shrink the space for rain layer estimation with few rain maps, so as to ensure a fine generalization performance in the inconsistent scenarios of rain types between training and testing data. By end-to-end training such an interpretable network, all involved rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers and, thus, naturally leading to better deraining performance. Comprehensive experiments implemented on a series of representative synthetic and real datasets substantiate the superiority of our method, especially on its well generality to diverse testing scenarios and good interpretability for all its modules, compared with state-of-the-art single image derainers both visually and quantitatively. Code is available at https://github.com/hongwang01/DRCDNet.
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yuexiang Li, Yong Liang 0001, Yefeng Zheng 0001, Deyu Meng
IEEE Trans. Neural Networks Learn. Syst.4
2023 ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image Segmentation
abstract
Vision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001
AAAI7
2023 Combating Mode Collapse via Offline Manifold Entropy Estimation
abstract
Generative Adversarial Networks (GANs) have shown compelling results in various tasks and applications in recent years. However, mode collapse remains a critical problem in GANs. In this paper, we propose a novel training pipeline to address the mode collapse issue of GANs. Different from existing methods, we propose to generalize the discriminator as feature embedding and maximize the entropy of distributions in the embedding space learned by the discriminator. Specifically, two regularization terms, i.e., Deep Local Linear Embedding (DLLE) and Deep Isometric feature Mapping (DIsoMap), are introduced to encourage the discriminator to learn the structural information embedded in the data, such that the embedding space learned by the discriminator can be well-formed. Based on the well-learned embedding space supported by the discriminator, a non-parametric entropy estimator is designed to efficiently maximize the entropy of embedding vectors, playing as an approximation of maximizing the entropy of the generated distribution. By improving the discriminator and maximizing the distance of the most similar samples in the embedding space, our pipeline effectively reduces the mode collapse without sacrificing the quality of generated samples. Extensive experimental results show the effectiveness of our method which outperforms the GAN baseline, MaF-GAN on CelebA (9.13 vs. 12.43 in FID) and surpasses the recent state-of-the-art energy-based model on the ANIMEFACE dataset (2.80 vs. 2.26 in Inception score).
Bing Li 0024, Haoqian Wu, Hanbang Liang, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001
AAAI6
2023 SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic Segmentation
abstract
Semi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks.
Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001
CVPR6
2023 AdaptiveMix: Improving GAN Training via Feature Space Shrinkage
abstract
Due to the outstanding capability for data generation, Generative Adversarial Networks (GANs) have attracted considerable attention in unsupervised learning. However, training GANs is difficult, since the training distribution is dynamic for the discriminator, leading to unstable image representation. In this paper, we address the problem of training GANs from a novel perspective, i.e., robust image classification. Motivated by studies on robust image representation, we propose a simple yet effective module, namely AdaptiveMix, for GANs, which shrinks the regions of training data in the image representation space of the discriminator. Considering it is intractable to directly bound feature space, we propose to construct hard samples and narrow down the feature distance between hard and easy samples. The hard samples are constructed by mixing a pair of training images. We evaluate the effectiveness of our AdaptiveMix with widely-used and state-of-the-art GAN architectures. The evaluation results demonstrate that our AdaptiveMix can facilitate the training of GANs and effectively improve the image quality of generated samples. We also show that our AdaptiveMix can be further applied to image classification and Out-Of-Distribution (OOD) detection tasks, by equipping it with state-of-the-art methods. Extensive experiments on seven publicly available datasets show that our method effectively boosts the performance of baselines. The code is publicly available at https://github.com/WentianZhang-ML/AdaptiveMix.
Wentian Zhang, Bing Li 0024, Haoqian Wu, Nanjun He, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001
CVPR7
2023 Interactive Segmentation as Gaussian Process Classification
abstract
Click-based interactive segmentation (IS) aims to extract the target objects under user interaction. For this task, most of the current deep learning (DL)-based methods mainly follow the general pipelines of semantic segmentation. Albeit achieving promising performance, they do not fully and explicitly utilize and propagate the click information, inevitably leading to unsatisfactory segmentation results, even at clicked points. Against this issue, in this paper, we propose to formulate the IS task as a Gaussian process (GP)-based pixel-wise binary classification model on each image. To solve this model, we utilize amortized variational inference to approximate the intractable GP posterior in a data-driven manner and then decouple the approximated GP posterior into double space forms for efficient sampling with linear complexity. Then, we correspondingly construct a GP classification framework, named GPCIS, which is integrated with the deep kernel learning mechanism for more flexibility. The main specificities of the proposed GPCIS lie in: 1) Under the explicit guidance of the derived GP posterior, the information contained in clicks can be finely propagated to the entire image and then boost the segmentation; 2) The accuracy of predictions at clicks has good theoretical support. These merits of GPCIS as well as its good generality and high efficiency are substantiated by comprehensive experiments on several benchmarks, as compared with representative methods both quantitatively and qualitatively. Codes will be released at https://github.com/zmhhlnz/GPCIS_CVPR2023.
Hong Wang 0021, Qian Zhao 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001
CVPR4
2023 FemtoDet: An Object Detection Baseline for Energy Versus Performance Tradeoffs
abstract
Efficient detectors for edge devices are often optimized for parameters or speed count metrics, which remain in weak correlation with the energy of detectors. However, some vision applications of convolutional neural networks, such as always-on surveillance cameras, are critical for energy constraints. This paper aims to serve as a baseline by designing detectors to reach tradeoffs between energy and performance from two perspectives: 1) We extensively analyze various CNNs to identify low-energy architectures, including selecting activation functions, convolutions operators, and feature fusion structures on necks. These underappreciated details in past work seriously affect the energy consumption of detectors; 2) To break through the dilemmatic energy-performance problem, we propose a balanced detector driven by energy using discovered low-energy components named FemtoDet. In addition to the novel construction, we improve FemtoDet by considering convolutions and training strategy optimizations. Specifically, we develop a new instance boundary enhancement (IBE) module for convolution optimization to overcome the contradiction between the limited capacity of CNNs and detection tasks in diverse spatial representations, and propose a recursive warm-restart (RecWR) for optimizing training strategy to escape the sub-optimization of light-weight detectors by considering the data shift produced in popular augmentations. As a result, FemtoDet with only 68.77k parameters achieves a competitive score of 46.3 AP50 on PASCAL VOC and 1.11 W & 64.47 FPS on Qualcomm Snapdragon 865 CPU platforms. Extensive experiments on COCO and TJUDHD datasets indicate that the proposed method achieves competitive results in diverse scenes.
Peng Tu, Guo Ai, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001
ICCV4
2023 BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
abstract
Recent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using other modalities as conditions, considerable paired data, e.g., box/mask-image pairs, and fine-tuning time are required for nurturing models. As such paired data is time-consuming and labor-intensive to acquire and restricted to a closed set, this potentially becomes the bottleneck for applications in an open world. This paper focuses on the simplest form of user-provided conditions, e.g., box or scribble. To mitigate the aforementioned problem, we propose a training-free method to control objects and contexts in the synthesized images adhering to the given spatial conditions. Specifically, three spatial constraints, i.e., Inner-Box, Outer-Box, and Corner Constraints, are designed and seamlessly integrated into the denoising step of diffusion models, requiring no additional training and massive annotated layout data. Extensive experimental results demonstrate that the proposed constraints can control what and where to present in the images while retaining the ability of Diffusion models to synthesize with high fidelity and diverse concept coverage.
Jinheng Xie, Yuexiang Li, Yawen Huang, Wentian Zhang, Yefeng Zheng 0001, Zheng Shou 0001
ICCV2
2023 MEPNet: A Model-Driven Equivariant Proximal Network for Joint Sparse-View Reconstruction and Metal Artifact Reduction in CT Images
Hong Wang 0021, Dong Wei 0004, Yuexiang Li, Yefeng Zheng 0001
MICCAI (10)4
2023 Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image Segmentation
abstract
Semi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets.
Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001
ACM Multimedia7
2023 Learning Visual Prior via Generative Pre-Training
abstract
Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts many vision tasks. For example, in conditional image synthesis, spatial conditions failing to adhere to the prior can result in visually inaccurate synthetic results. This work aims to explicitly learn the visual prior and enable the customization of sampling. Inspired by advances in language modeling, we propose to learn Visual prior via Generative Pre-Training, dubbed VisorGPT. By discretizing visual locations, e.g., bounding boxes, human pose, and instance masks, into sequences, VisorGPT can model visual prior through likelihood maximization. Besides, prompt engineering is investigated to unify various visual locations and enable customized sampling of sequential outputs from the learned prior. Experimental results demonstrate the effectiveness of VisorGPT in modeling visual prior and extrapolating to novel scenes, potentially motivating that discrete visual locations can be integrated into the learning paradigm of current language models to further perceive visual world. Code is available at https://sierkinhane.github.io/visor-gpt.
Jinheng Xie, Kai Ye 0004, Yudong Li 0001, Yuexiang Li, Qinghong Lin, Yefeng Zheng 0001, LinLin Shen, Zheng Shou 0001
NeurIPS4
2023 Dynamically Masked Discriminator for GANs
abstract
Training Generative Adversarial Networks (GANs) remains a challenging problem. The discriminator trains the generator by learning the distribution of real/generated data. However, the distribution of generated data changes throughout the training process, which is difficult for the discriminator to learn. In this paper, we propose a novel method for GANs from the viewpoint of online continual learning. We observe that the discriminator model, trained on historically generated data, often slows down its adaptation to the changes in the new arrival generated data, which accordingly decreases the quality of generated results. By treating the generated data in training as a stream, we propose to detect whether the discriminator slows down the learning of new knowledge in generated data. Therefore, we can explicitly enforce the discriminator to learn new knowledge fast. Particularly, we propose a new discriminator, which automatically detects its retardation and then dynamically masks its features, such that the discriminator can adaptively learn the temporally-vary distribution of generated data. Experimental results show our method outperforms the state-of-the-art approaches.
Wentian Zhang, Bing Li 0024, Jinheng Xie, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Bernard Ghanem
NeurIPS6
2023 MIL-ViT: A multiple instance vision transformer for fundus image classification
Qi Bi, Xu Sun 0006, Kai Ma 0002, Cheng Bian, Munan Ning, Nanjun He, Yawen Huang, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001
J. Vis. Commun. Image Represent.9
2023 Nuclei segmentation with point annotations from pathology images via self-supervised learning and co-training
Yi Lin 0009, Zhiyong Qu, Hao Chen 0011, Zhongke Gao, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng
Medical Image Anal.5
2023 A deep weakly semi-supervised framework for endoscopic lesion segmentation
Hong Wang 0021, Haoqin Ji, Yuexiang Li, Nanjun He, Dong Wei 0004, Yawen Huang, Xinrong Chen, Yefeng Zheng 0001, Hongmeng Yu
Medical Image Anal.5
2023 InDuDoNet+: A deep unfolding dual domain network for metal artifact reduction in CT images
Hong Wang 0021, Yuexiang Li, Haimiao Zhang, Deyu Meng, Yefeng Zheng 0001
Medical Image Anal.2
2023 Adversarial Learning of Object-Aware Activation Map for Weakly-Supervised Semantic Segmentation
abstract
Recent years have witnessed impressive advances in the area of weakly-supervised semantic segmentation (WSSS). However, most of existing approaches are based on class activation maps (CAMs), which suffer from the under-segmentation problem (i.e., objects of interest are segmented partially). Although a number of literature works have been proposed to tackle this under-segmentation problem, we argue that these solutions built on CAMs may not be optimal for the WSSS task. Instead, in this paper we propose a network based on the object-aware activation map (OAM). The proposed network, termed OAM-Net, consists of four loss functions (foreground loss, background loss, average pixel and consistency loss) which ensure exactness, completeness, compactness and consistency of segmented objects via adversarial training. Compared to conventional CAM-based methods, our OAM-Net overcomes the under-segmentation drawback and significantly improves segmentation accuracy with negligible computational cost. A thorough comparison between OAM-Net and CAM-based approaches is carried out on the PASCAL VOC2012 dataset, and experimental results show that our network outperforms state-of-the-art approaches by a large margin. The code will be available soon.
Junliang Chen 0002, Weizeng Lu, Yuexiang Li, LinLin Shen, Jinming Duan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 Blind Super-Resolution of 3D MRI via Unsupervised Domain Transformation
abstract
High-resolution medical images can be effectively used for clinical diagnosis. However, the acquisition of high-resolution images is difficult and often limited by medical instruments. Super-resolution (SR) methods provide a solution, where high-resolution (HR) images can be reconstructed from low-resolution (LR) ones. Most of existing deep neural networks for 3D SR medical images trained in a non-blind process, where LR images are directly degraded from HR data via a pre-determined downscale method. Such approaches rely heavily on the assumed degradation model, resulting in inevitable deviations in real clinical practice. Blind super-resolution, as a more attractive research line for this field, aims to generate HR images from LR inputs containing unknown degradation. Towards generalizing SR models for diverse types of degradation, we propose a robust blind SR of 3D medical images in an unsupervised manner with domain correction and upscaling treatment. First, a CycleGAN-based architecture is implemented to generate the LR data from the source domain to the target one for domain correction. Then, an upscaling network is learned via pre-determined HR-LR couples for reconstruction. The proposed framework is able to automatically learn noisy and blurry correction kernels for unpaired 3D SR magnetic resonance images (MRI). Our method achieves better and more robust performances in reconstruction of HR images from LR MRI with multiple unknown degradation processes, and show its superiority to other state-of-the-art supervised models and cycle-consistency based methods, especially in severe distortion cases.
Hexiang Zhou, Yawen Huang, Yuexiang Li, Yi Zhou 0007, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics3
2022 Generalized Brain Image Synthesis with Transferable Convolutional Sparse Coding Networks
Yawen Huang, Feng Zheng 0001, Xu Sun 0006, Yuexiang Li, Ling Shao 0001, Yefeng Zheng 0001
ECCV (34)4
2022 Adaptive Convolutional Dictionary Network for CT Metal Artifact Reduction
abstract
Inspired by the great success of deep neural networks, learning-based methods have gained promising performances for metal artifact reduction (MAR) in computed tomography (CT) images. However, most of the existing approaches put less emphasis on modelling and embedding the intrinsic prior knowledge underlying this specific MAR task into their network designs. Against this issue, we propose an adaptive convolutional dictionary network (ACDNet), which leverages both model-based and learning-based methods. Specifically, we explore the prior structures of metal artifacts, e.g., non-local repetitive streaking patterns, and encode them as an explicit weighted convolutional dictionary model. Then, a simple-yet-effective algorithm is carefully designed to solve the model. By unfolding every iterative substep of the proposed algorithm into a network module, we explicitly embed the prior structure into a deep network , i.e., a clear interpretability for the MAR task. Furthermore, our ACDNet can automatically learn the prior for artifact-free CT images via training data and adaptively adjust the representation kernels for each input CT image based on its content. Hence, our method inherits the clear interpretability of model-based methods and maintains the powerful representation ability of learning-based methods. Comprehensive experiments executed on synthetic and clinical datasets show the superiority of our ACDNet in terms of effectiveness and model generalization. Code and supplementary material are available at https://github.com/hongwang01/ACDNet.
Hong Wang 0021, Yuexiang Li, Deyu Meng, Yefeng Zheng 0001
IJCAI2
2022 Point Beyond Class: A Benchmark for Weakly Semi-supervised Abnormality Localization in Chest X-Rays
Haoqin Ji, Yuexiang Li, Jinheng Xie, Nanjun He, Yawen Huang, Dong Wei 0004, Xinrong Chen, LinLin Shen, Yefeng Zheng 0001
MICCAI (3)3
2022 Orientation-Shared Convolution Representation for CT Metal Artifact Learning
Hong Wang 0021, Qi Xie 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001
MICCAI (6)3
2022 mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001
MICCAI (5)4
2022 A Multi-task Network with Weight Decay Skip Connection Training for Anomaly Detection in Retinal Fundus Images
Wentian Zhang, Xu Sun 0006, Yuexiang Li, Nanjun He, Feng Liu 0013, Yefeng Zheng 0001
MICCAI (2)3
2022 Aggregative Self-supervised Feature Learning from Limited Medical Images
Jiuwen Zhu, Yuexiang Li, Lian Ding, Shaohua Kevin Zhou
MICCAI (8)2
2022 Cross-Domain Gated Learning for Domain Generalization
Dapeng Du, Jiawei Chen 0009, Yuexiang Li, Kai Ma 0002, Gangshan Wu, Yefeng Zheng 0001, Limin Wang 0002
Int. J. Comput. Vis.3
2022 Mix-and-Interpolate: A Training Strategy to Deal With Source-Biased Medical Data
abstract
Till March 31st, 2021, the coronavirus disease 2019 (COVID-19) had reportedly infected more than 127 million people and caused over 2.5 million deaths worldwide. Timely diagnosis of COVID-19 is crucial for management of individual patients as well as containment of the highly contagious disease. Having realized the clinical value of non-contrast chest computed tomography (CT) for diagnosis of COVID-19, deep learning (DL) based automated methods have been proposed to aid the radiologists in reading the huge quantities of CT exams as a result of the pandemic. In this work, we address an overlooked problem for training deep convolutional neural networks for COVID-19 classification using real-world multi-source data, namely, the data source bias problem. The data source bias problem refers to the situation in which certain sources of data comprise only a single class of data, and training with such source-biased data may make the DL models learn to distinguish data sources instead of COVID-19. To overcome this problem, we propose MIx-aNd-Interpolate (MINI), a conceptually simple, easy-to-implement, efficient yet effective training strategy. The proposed MINI approach generates volumes of the absent class by combining the samples collected from different hospitals, which enlarges the sample space of the original source-biased dataset. Experimental results on a large collection of real patient data (1,221 COVID-19 and 1,520 negative CT images, and the latter consisting of 786 community acquired pneumonia and 734 non-pneumonia) from eight hospitals and health institutions show that: 1) MINI can improve COVID-19 classification performance upon the baseline (which does not deal with the source bias), and 2) MINI is superior to competing methods in terms of the extent of improvement.
Yuexiang Li, Jiawei Chen 0009, Dong Wei 0004, Yanchun Zhu, Junfeng Xiong, Yadong Gang, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics1
2022 Beyond Mutual Information: Generative Adversarial Network for Domain Adaptation Using Information Bottleneck Constraint
abstract
Medical images from multicentres often suffer from the domain shift problem, which makes the deep learning models trained on one domain usually fail to generalize well to another. One of the potential solutions for the problem is the generative adversarial network (GAN), which has the capacity to translate images between different domains. Nevertheless, the existing GAN-based approaches are prone to fail at preserving image-objects in image-to-image (I2I) translation, which reduces their practicality on domain adaptation tasks. In this regard, a novel GAN (namely IB-GAN) is proposed to preserve image-objects during cross-domain I2I adaptation. Specifically, we integrate the information bottleneck constraint into the typical cycle-consistency-based GAN to discard the superfluous information (e.g., domain information) and maintain the consistency of disentangled content features for image-object preservation. The proposed IB-GAN is evaluated on three tasks-polyp segmentation using colonoscopic images, the segmentation of optic disc and cup in fundus images and the whole heart segmentation using multi-modal volumes. We show that the proposed IB-GAN can generate realistic translated images and remarkably boost the generalization of widely used segmentation networks (e.g., U-Net).
Jiawei Chen 0009, Ziqi Zhang 0013, Xinpeng Xie, Yuexiang Li, Tao Xu 0026, Kai Ma 0002, Yefeng Zheng 0001
IEEE Trans. Medical Imaging4
2022 DICDNet: Deep Interpretable Convolutional Dictionary Network for Metal Artifact Reduction in CT Images
abstract
Computed tomography (CT) images are often impaired by unfavorable artifacts caused by metallic implants within patients, which would adversely affect the subsequent clinical diagnosis and treatment. Although the existing deep-learning-based approaches have achieved promising success on metal artifact reduction (MAR) for CT images, most of them treated the task as a general image restoration problem and utilized off-the-shelf network modules for image quality enhancement. Hence, such frameworks always suffer from lack of sufficient model interpretability for the specific task. Besides, the existing MAR techniques largely neglect the intrinsic prior knowledge underlying metal-corrupted CT images which is beneficial for the MAR performance improvement. In this paper, we specifically propose a deep interpretable convolutional dictionary network (DICDNet) for the MAR task. Particularly, we first explore that the metal artifacts always present non-local streaking and star-shape patterns in CT images. Based on such observations, a convolutional dictionary model is deployed to encode the metal artifacts. To solve the model, we propose a novel optimization algorithm based on the proximal gradient technique. With only simple operators, the iterative steps of the proposed algorithm can be easily unfolded into corresponding network modules with specific physical meanings. Comprehensive experiments on synthesized and clinical datasets substantiate the effectiveness of the proposed DICDNet as well as its superior interpretability, compared to current state-of-the-art MAR methods. Code is available at https://github.com/hongwang01/DICDNet.
Hong Wang 0021, Yuexiang Li, Nanjun He, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001
IEEE Trans. Medical Imaging2
2021 Alleviating Noisy-label Effects in Image Classification via Probability Transition Matrix
Ziqi Zhang 0013, Yuexiang Li, Hongxin Wei, Kai Ma 0002, Tao Xu 0026, Yefeng Zheng 0001
BMVC2
2021 Deep Reinforcement Exemplar Learning for Annotation Refinement
Yuexiang Li, Nanjun He, Sixiang Peng, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (8)1
2021 Triplet-Branch Network with Prior-Knowledge Embedding for Fatigue Fracture Grading
Yuexiang Li, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001
MICCAI (5)1
2021 Simultaneous Alignment and Surface Regression Using Hybrid 2D-3D Networks for 3D Coherent Layer Segmentation of Retina OCT Images
Dong Wei 0004, Donghuan Lu, Yuexiang Li, Kai Ma 0002, Liansheng Wang 0002, Yefeng Zheng 0001
MICCAI (8)4
2021 InDuDoNet: An Interpretable Dual Domain Network for CT Metal Artifact Reduction
Hong Wang 0021, Yuexiang Li, Haimiao Zhang, Jiawei Chen 0009, Kai Ma 0002, Deyu Meng, Yefeng Zheng 0001
MICCAI (6)2
2021 MIL-VT: Multiple Instance Learning Enhanced Vision Transformer for Fundus Image Classification
Kai Ma 0002, Qi Bi, Cheng Bian, Munan Ning, Nanjun He, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001
MICCAI (8)7
2021 GRAND: A large-scale dataset and benchmark for cervical intraepithelial Neoplasia grading with fine-grained lesion description
Yuexiang Li, Zhi-Hua Liu, Peng Xue 0001, Jiawei Chen 0009, Kai Ma 0002, Tianyi Qian, Yefeng Zheng 0001, Youlin Qiao
Medical Image Anal.1
2021 Anomaly Detection for Medical Images Using Self-Supervised and Translation-Consistent Features
abstract
As the labeled anomalous medical images are usually difficult to acquire, especially for rare diseases, the deep learning based methods, which heavily rely on the large amount of labeled data, cannot yield a satisfactory performance. Compared to the anomalous data, the normal images without the need of lesion annotation are much easier to collect. In this paper, we propose an anomaly detection framework, namely [Formula: see text], extracting [Formula: see text]elf-supervised and tr [Formula: see text]ns [Formula: see text]ation-consistent features for [Formula: see text]nomaly [Formula: see text]etection. The proposed SALAD is a reconstruction-based method, which learns the manifold of normal data through an encode-and-reconstruct translation between image and latent spaces. In particular, two constraints (i.e., structure similarity loss and center constraint loss) are proposed to regulate the cross-space (i.e., image and feature) translation, which enforce the model to learn translation-consistent and representative features from the normal data. Furthermore, a self-supervised learning module is engaged into our framework to further boost the anomaly detection accuracy by deeply exploiting useful information from the raw normal data. An anomaly score, as a measure to separate the anomalous data from the healthy ones, is constructed based on the learned self-supervised-and-translation-consistent features. Extensive experiments are conducted on optical coherence tomography (OCT) and chest X-ray datasets. The experimental results demonstrate the effectiveness of our approach.
He Zhao 0002, Yuexiang Li, Nanjun He, Kai Ma 0002, Leyuan Fang, Huiqi Li, Yefeng Zheng 0001
IEEE Trans. Medical Imaging2
2020 Generative Adversarial Networks for Video-to-Video Domain Adaptation
abstract
Endoscopic videos from multicentres often have different imaging conditions, e.g., color and illumination, which make the models trained on one domain usually fail to generalize well to another. Domain adaptation is one of the potential solutions to address the problem. However, few of existing works focused on the translation of video-based data. In this work, we propose a novel generative adversarial network (GAN), namely VideoGAN, to transfer the video-based data across different domains. As the frames of a video may have similar content and imaging conditions, the proposed VideoGAN has an X-shape generator to preserve the intra-video consistency during translation. Furthermore, a loss function, namely color histogram loss, is proposed to tune the color distribution of each translated frame. Two colonoscopic datasets from different centres, i.e., CVC-Clinic and ETIS-Larib, are adopted to evaluate the performance of domain adaptation of our VideoGAN. Experimental results demonstrate that the adapted colonoscopic video generated by our VideoGAN can significantly boost the segmentation accuracy, i.e., an improvement of 5%, of colorectal polyps on multicentre datasets. As our VideoGAN is a general network architecture, we also evaluate its performance with the CamVid driving video dataset on the cloudy-to-sunny translation task. Comprehensive experiments show that the domain gap could be substantially narrowed down by our VideoGAN.
Jiawei Chen 0009, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001
AAAI2
2020 Dual Adversarial Network for Deep Active Learning
Shuo Wang 0008, Yuexiang Li, Kai Ma 0002, Ruhui Ma, Haibing Guan, Yefeng Zheng 0001
ECCV (24)2
2020 Self-Supervised CycleGAN for Object-Preserving Image-to-Image Domain Adaptation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001
ECCV (20)3
2020 Self-Loop Uncertainty: A Novel Pseudo-Label for Semi-supervised Medical Image Segmentation
Yuexiang Li, Jiawei Chen 0009, Xinpeng Xie, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (1)1
2020 Revisiting Rubik's Cube: Self-supervised Learning with Volume-Wise Transformation for 3D Medical Image Segmentation
Xing Tao, Yuexiang Li, Wenhui Zhou 0001, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (4)2
2020 MI2GAN: Generative Adversarial Network for Medical Image Domain Adaptation Using Mutual Information Constraint
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (2)3
2020 Instance-Aware Self-supervised Learning for Nuclei Segmentation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001
MICCAI (5)3
2020 AGE challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
Huazhu Fu, Fei Li 0021, Xu Sun 0006, Xingxing Cao, Jingan Liao, José Ignacio Orlando, Xing Tao, Yuexiang Li, Mingkui Tan, Chenglang Yuan, Cheng Bian, Ruitao Xie, Jiongcheng Li, Xiaomeng Li 0001, Jing Wang 0023, Le Geng, Panming Li, Yanwu Xu 0001
Medical Image Anal.8
2020 Rubik's Cube+: A self-supervised feature learning framework for 3D medical image analysis
Jiuwen Zhu, Yuexiang Li, Kai Ma 0002, Shaohua Kevin Zhou, Yefeng Zheng 0001
Medical Image Anal.2
2020 Efficient and Effective Training of COVID-19 Classification Networks With Self-Supervised Dual-Track Learning to Rank
abstract
Coronavirus Disease 2019 (COVID-19) has rapidly spread worldwide since first reported. Timely diagnosis of COVID-19 is crucial both for disease control and patient care. Non-contrast thoracic computed tomography (CT) has been identified as an effective tool for the diagnosis, yet the disease outbreak has placed tremendous pressure on radiologists for reading the exams and may potentially lead to fatigue-related mis-diagnosis. Reliable automatic classification algorithms can be really helpful; however, they usually require a considerable number of COVID-19 cases for training, which is difficult to acquire in a timely manner. Meanwhile, how to effectively utilize the existing archive of non-COVID-19 data (the negative samples) in the presence of severe class imbalance is another challenge. In addition, the sudden disease outbreak necessitates fast algorithm development. In this work, we propose a novel approach for effective and efficient training of COVID-19 classification networks using a small number of COVID-19 CT exams and an archive of negative samples. Concretely, a novel self-supervised learning method is proposed to extract features from the COVID-19 and negative samples. Then, two kinds of soft-labels ('difficulty' and 'diversity') are generated for the negative samples by computing the earth mover's distances between the features of the negative and COVID-19 samples, from which data 'values' of the negative samples can be assessed. A pre-set number of negative samples are selected accordingly and fed to the neural network for training. Experimental results show that our approach can achieve superior performance using about half of the negative samples, substantially reducing model training time.
Yuexiang Li, Dong Wei 0004, Jiawei Chen 0009, Shilei Cao 0001, Yanchun Zhu, Lan Lan 0002, Tianyi Qian, Kai Ma 0002, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics1
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging65
2020 Computer-Aided Cervical Cancer Diagnosis Using Time-Lapsed Colposcopic Images
abstract
Cervical cancer causes the fourth most cancer-related deaths of women worldwide. Early detection of cervical intraepithelial neoplasia (CIN) can significantly increase the survival rate of patients. In this paper, we propose a deep learning framework for the accurate identification of LSIL+ (including CIN and cervical cancer) using time-lapsed colposcopic images. The proposed framework involves two main components, i.e., key-frame feature encoding networks and feature fusion network. The features of the original (pre-acetic-acid) image and the colposcopic images captured at around 60s, 90s, 120s and 150s during the acetic acid test are encoded by the feature encoding networks. Several fusion approaches are compared, all of which outperform the existing automated cervical cancer diagnosis systems using a single time slot. A graph convolutional network with edge features (E-GCN) is found to be the most suitable fusion approach in our study, due to its excellent explainability consistent with the clinical practice. A large-scale dataset, containing time-lapsed colposcopic images from 7,668 patients, is collected from the collaborative hospital to train and validate our deep learning framework. Colposcopists are invited to compete with our computer-aided diagnosis system. The proposed deep learning framework achieves a classification accuracy of 78.33%-comparable to that of an in-service colposcopist-which demonstrates its potential to provide assistance in the realistic clinical scenario.
Yuexiang Li, Jiawei Chen 0009, Peng Xue 0001, Jia Chang, Chunyan Chu, Kai Ma 0002, Yefeng Zheng 0001, Youlin Qiao
IEEE Trans. Medical Imaging1
2019 Self-supervised Feature Learning for 3D Medical Images by Playing a Rubik's Cube
Xinrui Zhuang, Yuexiang Li, Kai Ma 0002, Yujiu Yang 0001, Yefeng Zheng 0001
MICCAI (4)2
2019 Reverse active learning based atrous DenseNet for pathological image classification
abstract
BACKGROUND: Due to the recent advances in deep learning, this model attracted researchers who have applied it to medical image analysis. However, pathological image analysis based on deep learning networks faces a number of challenges, such as the high resolution (gigapixel) of pathological images and the lack of annotation capabilities. To address these challenges, we propose a training strategy called deep-reverse active learning (DRAL) and atrous DenseNet (ADN) for pathological image classification. The proposed DRAL can improve the classification accuracy of widely used deep learning networks such as VGG-16 and ResNet by removing mislabeled patches in the training set. As the size of a cancer area varies widely in pathological images, the proposed ADN integrates the atrous convolutions with the dense block for multiscale feature extraction. RESULTS: The proposed DRAL and ADN are evaluated using the following three pathological datasets: BACH, CCG, and UCSB. The experiment results demonstrate the excellent performance of the proposed DRAL + ADN framework, achieving patch-level average classification accuracies (ACA) of 94.10%, 92.05% and 97.63% on the BACH, CCG, and UCSB validation sets, respectively. CONCLUSIONS: The DRAL + ADN framework is a potential candidate for boosting the performance of deep learning models for partially mislabeled training datasets.
Yuexiang Li, Xinpeng Xie, LinLin Shen, Shaoxiong Liu
BMC Bioinform.1
2018 GT-Net: A Deep Learning Network for Gastric Tumor Diagnosis
abstract
Gastric cancer is one of the most common cancers, which causes the second largest number of deaths in the world. Traditional diagnosis approach requires pathologists to manually annotate the gastric tumor in gastric slice for cancer identification, which is laborious and time-consuming. In this paper, we proposed a deep learning based framework, namely GT-Net, for automatic segmentation of gastric tumor. The proposed GT-Net adopts different architectures for shallow and deep layers for better feature extraction. We evaluate the proposed framework on publicly available BOT gastric slice dataset. The experimental results show that our GT-Net performs better than state-of-the-art networks like FCN-8s, U-net, and achieved a new state-of-the-art F1 score of 90.88% for gastric tumor segmentation.
Yuexiang Li, Xinpeng Xie, Shaoxiong Liu, Xuechen Li 0001, LinLin Shen
ICTAI1
2018 Open snake model based on global guidance field for embryo vessel location
abstract
The development of vessels can provide important information about the growth status of animal embryos. It is, therefore, important to automatically locate the deformed vessel branches from the embryo images. However, very few vessel detectors can accurately locate all vessel branches when the captured images are low quality and the implied vessel shapes are complex. In this study, a new framework consisting of vessel region extraction and snake shape optimisation is proposed. The main contribution in this detector is a novel open snake model based on the global guidance field and deformation template initialisation. Experimental results on a specific application of an embryo vessel database [Database and source codes: https://github.com/wcxie/Egg‐embryro‐vessel‐location/ .] demonstrate that the proposed algorithm not only locates the vessel shape properly but also obtains the orientations of embryo vessel branches accurately. Comparison to traditional guidance fields and the active appearance model illustrates the effectiveness and competitiveness of the proposed model.
Weicheng Xie 0001, Jinming Duan 0001, LinLin Shen, Yuexiang Li, Meng Yang 0001, Guojun Lin
IET Comput. Vis.4
2018 Deep cross residual network for HEp-2 cell staining pattern classification
LinLin Shen, Xi Jia, Yuexiang Li
Pattern Recognit.3
2017 HEp-2 Specimen Image Segmentation and Classification Using Very Deep Fully Convolutional Network
abstract
Reliable identification of Human Epithelial-2 (HEp-2) cell patterns can facilitate the diagnosis of systemic autoimmune diseases. However, traditional approach requires experienced experts to manually recognize the cell patterns, which suffers from the inter-observer variability. In this paper, an automatic pattern recognition system using fully convolutional network (FCN) was proposed to simultaneously address the segmentation and classification problem of HEp-2 specimen images. The proposed system transforms the residual network (ResNet) to fully convolutional ResNet (FCRN) enabling the network to perform semantic segmentation task. A sand-clock shape residual module is proposed to effectively and economically improve the performance of FCRN. The publicly available I3A-2014 data set was used to train the FCRN model to classify HEp-2 specimen images into seven catalogs: homogeneous, speckled, nucleolar, centromere, golgi, nuclear membrane, and mitotic spindle. The proposed system achieves a mean class accuracy of 94.94% for leave-one-out tests, which outperforms the winner of ICPR 2014, i.e., 89.93%. At the same time, our model also achieves a segmentation accuracy of 89.03%, which is 19.05% higher than that of the benchmark approach, i.e., 69.98%.
Yuexiang Li, LinLin Shen, Shiqi Yu 0001
IEEE Trans. Medical Imaging1
2016 HEp-2 specimen classification with fully convolutional network
abstract
Reliable automatic system for Human Epithelial-2 (HEp-2) cell image classification can facilitate the diagnosis of systemic autoimmune diseases. In this paper, an automatic pattern recognition system using fully convolutional network (FCN) was proposed to address the HEp-2 specimen classification problem. The FCN in the proposed framework was adapted from VGG-16, which was trained with ICPR 2016 dataset to classify specimen images into seven catalogs: homogeneous, speckled, nucleolar, centromere, golgi, nuclear membrane, and mitotic spindle. The proposed system achieves a mean class accuracy of 90.89% for 5 fold-cross-validation tests using the I3A Contest Task 2 dataset, which is comparable to the winner of ICPR 2014, i.e. 89.93%. Furthermore, since the FCN was firstly developed for semantic segmentation, the proposed framework can simultaneously solve Task 4, Cell segmentation, newly suggested in I3A Contest 2016. The segmentation accuracy of the system is 87.38% on Task 4 dataset which is 17.4% higher than that of the traditional approach, Otsu, i.e. 69.98%.
Yuexiang Li, LinLin Shen, Xiande Zhou, Shiqi Yu 0001
ICPR1