Yizhe Zhang 0001

dblp:132/4966-1 · DBLP profile ↗
← Back
64ranked-venue papers
10as first author
43since 2021 · last 2026
0000-0002-8676-9763ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Decoupling Continual Semantic Segmentation
abstract
Continual Semantic Segmentation (CSS) requires learning new classes without forgetting previously acquired knowledge, addressing the fundamental challenge of catastrophic forgetting in dense prediction tasks. However, existing CSS methods typically employ single-stage encoder-decoder architectures where segmentation masks and class labels are tightly coupled, leading to interference between old and new class learning and suboptimal retention-plasticity balance. We introduce DecoupleCSS, a novel two-stage framework for CSS. By decoupling class-aware detection from class-agnostic segmentation, DecoupleCSS enables more effective continual learning, preserving past knowledge while learning new classes. The first stage leverages pre-trained text and image encoders, adapted using LoRA, to encode class-specific information and generate location-aware prompts. In the second stage, the Segment Anything Model (SAM) is employed to produce precise segmentation masks, ensuring that segmentation knowledge is shared across both new and previous classes. This approach improves the balance between retention and adaptability in CSS, achieving state-of-the-art performance across a variety of challenging tasks.
Yifu Guo, Yuquan Lu, Wentao Zhang 0005, Zishan Xu, Dexia Chen, Yizhe Zhang 0001
AAAI7
2026 Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical Segmentation
abstract
Semi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural complexity, while also neglecting the interaction between labeled and unlabeled data streams. To overcome these challenges, we propose a Bidirectional Channel-selective Semantic Interaction (BCSI) framework for semi-supervised medical image segmentation. First, we propose a Semantic-Spatial Perturbation (SSP) mechanism, which disturbs the data using two strong augmentation operations and leverages unsupervised learning with pseudo-labels from weak augmentations. Additionally, we employ consistency on the predictions from the two strong augmentations to further improve model stability and robustness. Second, to reduce noise during the interaction between labeled and unlabeled data, we propose a Channel-selective Router (CR) component, which dynamically selects the most relevant channels for information exchange. This mechanism ensures that only highly relevant features are activated, minimizing unnecessary interference. Finally, the Bidirectional Channel-wise Interaction (BCI) strategy is employed to supplement additional semantic information and enhance the representation of important channels. Experimental results on multiple benchmarking 3D medical datasets demonstrate that the proposed method outperforms existing semi-supervised approaches.
Kaiwen Huang 0002, Yizhe Zhang 0001, Yi Zhou 0007, Tianyang Xu 0001, Tao Zhou 0002
AAAI2
2026 Test-time generative augmentation for medical image segmentation
Xiao Ma 0011, Yuhui Tao, Zetian Zhang, Yuhan Zhang 0001, Xi Wang 0013, Sheng Zhang 0024, Zexuan Ji, Yizhe Zhang 0001, Qiang Chen 0004, Guang Yang 0006
Medical Image Anal.8
2026 Frequency-enhanced contextual conversion network for esophageal lesion segmentation
Ziqi Tang, Xiaotong Niu, Long Rong, Yizhe Zhang 0001, Yawei Bi, Nan Ru, Longsong Li, Ningli Chai, Tao Zhou 0002
Pattern Recognit.4
2026 Structured prompt-guided knowledge injection for medical image segmentation
Kelei He, Yizhe Zhang 0001, Yi Zhou 0007, Tao Zhou 0002, Dong Liang 0001
Pattern Recognit.3
2026 Hierarchical Vision-Language Interaction for Facial Action Unit Detection
abstract
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Inter action for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal atten tion architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language based AU details, culminating in robust and semantically en riched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.
Yong Li 0032, Yizhe Zhang 0001, Tianyi Zhang 0013, Muyun Jiang, Guosen Xie, Cuntai Guan
IEEE Trans. Affect. Comput.3
2026 Cell Instance Segmentation: The Devil Is in the Boundaries
abstract
State-of-the-art (SOTA) methods for cell instance segmentation are based on deep learning (DL) semantic segmentation approaches, focusing on distinguishing foreground pixels from background pixels. In order to identify cell instances from foreground pixels (e.g., pixel clustering), most methods decompose instance information into pixel-wise objectives, such as distances to foreground-background boundaries (distance maps), heat gradients with the center point as heat source (heat diffusion maps), and distances from the center point to foreground-background boundaries with fixed angles (star-shaped polygons). However, pixel-wise objectives may lose significant geometric properties of the cell instances, such as shape, curvature, and convexity, which require a collection of pixels to represent. To address this challenge, we present a novel pixel clustering method, called Ceb (for Cell boundaries), to leverage cell boundary features and labels to divide foreground pixels into cell instances. Starting with probability maps generated from semantic segmentation, Ceb first extracts potential foreground-foreground boundaries (i.e., boundary candidates) with a revised Watershed algorithm. For each boundary candidate, a boundary feature representation (called boundary signature) is constructed by sampling pixels from the current foreground-foreground boundary as well as the neighboring background-foreground boundaries. Next, a lightweight boundary classifier is used to predict its binary boundary label based on the corresponding boundary signature. Finally, cell instances are obtained by dividing or merging neighboring regions based on the predicted boundary labels. Extensive experiments on six datasets demonstrate that Ceb outperforms existing pixel clustering methods on semantic segmentation probability maps. Moreover, Ceb achieves highly competitive performance compared to state-of-the-art cell instance segmentation methods. The code is available at: https://github.com/pxliang/Ceb.
Peixian Liang, Yifan Ding 0001, Yizhe Zhang 0001, Jianxu Chen 0001, Hao Zheng 0006, Yejia Zhang, Guangyu Meng, Tim Weninger, Michael T. Niemier, Xiaobo Sharon Hu, Danny Ziyi Chen
IEEE Trans. Medical Imaging3
2025 Asymmetric Performance Profiling Using Foundation Models: Quantifying Reliability and Expert Capability in Medical AI
abstract
In safety-critical domains like medical imaging, where diagnostic errors have severe consequences, AI models must be evaluated beyond average accuracy. A trustworthy model must demonstrate two distinct virtues: high reliability on common, easy cases and high expert capability on challenging, ambiguous, or rare cases. Conventional aggregate metrics fail to distinguish between these, masking a model's fatal flaw-such as misclassifying an easy case-by rewarding its high volume of trivial successes. We present Hardness-Aware Model Evaluation (HaME), a framework that assesses models using an asymmet-ric cost-benefit analysis. HaME identifies challenging instances via foundation models, then evaluates the target model using novel metrics (HaPrecision, HaRecall, HaFt, HaAUC) and the Brittleness Gap (B-Gap). Our formulation uniquely penalizes “easy” errors far more than “hard” failures, while simultaneously rewarding “hard” successes. This shifts the evaluation from “av-erage performance” to “clinical trustworthiness.” Experiments on medical image classification (Dermatology, Pneumonia, Retinal OCT) and segmentation (Nuclei) reveal that HaME uncovers critical reliability gaps and expert-level specializations invisible to standard metrics.
Mingzhi Xu, Tao Zhou 0002, Qiang Chen 0004, Shuo Wang 0011, Yizhe Zhang 0001
BIBM6
2025 Textual Prototype-Guided Continual Learning for Medical Image Classification
abstract
Intelligent diagnostic systems require continual learning (CL) to learn to diagnose new diseases. However, the systems suffer from catastrophic forgetting of old knowledge when learning knowledge of new diseases. Existing CL methods that leverage pre-trained language models (PLMs) to guide visual encoders are ineffective in the medical domain due to PLMs' limited medical knowledge. Here, we propose Textual Prototype-Guided Continual Learning (TPGCL) for effective CL. TPGCL utilizes an image captioning model to generate semantically rich disease descriptions, which are then encoded into text embeddings via a text encoder to obtain the textual prototype for each class. These prototypes guide the visual encoder during CL. In addition, a gradient-weighting mechanism merges visual adapters learned from all CL stages to preserve old knowledge and prevent model growth and adapter selection issue. Extensive experiments on three medical image datasets demonstrate TPGCL's superiority in continually learning new diseases. The source code is available at https://github.com/z1968357787/TPGCL
Zhiping Zhou, Yizhe Zhang 0001, Wei-Shi Zheng 0001
BIBM3
2025 Text-Driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
Kaiwen Huang 0002, Yi Zhou 0007, Huazhu Fu, Yizhe Zhang 0001, Chen Gong 0002, Tao Zhou 0002
MICCAI (5)4
2025 Unsupervised Quality Control and Enhancement of Polyp Segmentation in Colonoscopy Videos Using Spatiotemporal Consistency
Tao Zhou 0002, Shuo Wang 0011, Yizhe Zhang 0001
MICCAI (10)5
2025 From Generalist to Specialist: Distilling a Mixture of Foundation Models for Domain-Specific Medical Image Segmentation
Qing Li 0001, Yizhe Zhang 0001, Shengxiao Yang, Qirong Li, Shuo Wang 0011, Chengyan Wang
MICCAI (2)2
2025 Hierarchical Spatio-Temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos
Jian Yang 0003, Yizhe Zhang 0001, Tao Zhou 0002
MICCAI (3)3
2025 Coherence-Based Segmentation Quality Evaluator Trained on a Large Collection of Annotated Medical Images
Ahjol Senbi, Fei Lyu 0004, Qing Li 0001, Yuhui Tao, Qiang Chen 0004, Chengyan Wang, Shuo Wang 0011, Tao Zhou 0002, Yizhe Zhang 0001
PRCV (13)11
2025 ST-MIGD: Spatial-Temporal Domain Medical Image Generation via Deformation-Based Diffusion Models
Xiao Ma 0011, Yizhe Zhang 0001, Qiang Chen 0004
PRCV (14)4
2025 WeakPolyp-SAM: Segment Anything Model-driven weakly-supervised polyp segmentation
Tao Zhou 0002, Yunqi Gu, Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Huazhu Fu
Knowl. Based Syst.5
2025 Dynamic Multi-scale Feature Integration Network for unsupervised MR-CT synthesis
Jiuming Jiang, Tao Zhou 0002, Yizhe Zhang 0001, Bin Qiu, Li Zhang 0021
Neural Networks5
2025 Dual-scale enhanced and cross-generative consistency learning for semi-supervised medical image segmentation
Yunqi Gu, Tao Zhou 0002, Yizhe Zhang 0001, Yi Zhou 0007, Kelei He, Chen Gong 0002, Huazhu Fu
Pattern Recognit.3
2025 Frequency-Aware Interaction Network for Ultrasound Image Segmentation
abstract
Accurate segmentation of medical ultrasound images is crucial for guiding treatment decisions and assessing intervention effectiveness. The challenge of segmenting lesions in ultrasound images arises from factors such as low contrast, high speckle noise, artifacts, and blurred boundaries. Furthermore, this complexity varies significantly among lesions in different cases. While methods based on Convolutional Neural Networks (CNNs) and Transformers have shown promising results in this field, each approach possesses distinct advantages and limitations. To address these challenges, we propose a novel Frequency-aware Interaction Network (FINet). At the core of our FINet lies the proposed Multi-scale Frequency-aware Self-attention (MFS) module, which effectively captures multi-scale feature information within the self-attention layer. This enables our network to model both local and global features, capitalizing on the strengths of both CNNs and Transformers. Additionally, a frequency-aware network is introduced to learn the interactions between spatial locations in the frequency domain to enhance detailed feature representation such as edges. Furthermore, we present a collaborative interactive decoder network, in which a Selective Feature Interaction (SFI) module is proposed to facilitate the semantic and boundary feature interaction, resulting in more precise segmentation outcomes. Experimental results on four medical ultrasound image datasets show the superiority of our FINet over other state-of-the-art segmentation methods. More importantly, our model achieves an excellent trade-off between performance and computational efficiency.
Tao Zhou 0002, Yizhe Zhang 0001, Shangbing Gao, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.3
2025 Uncertainty-Aware Cross-Training for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning has gained considerable popularity in medical image segmentation tasks due to its capability to reduce reliance on expert-examined annotations. Several mean-teacher (MT) based semi-supervised methods utilize consistency regularization to effectively leverage valuable information from unlabeled data. However, these methods often heavily rely on the student model and overlook the potential impact of cognitive biases within the model. Furthermore, some methods employ co-training using pseudo-labels derived from different inputs, yet generating high-confidence pseudo-labels from perturbed inputs during training remains a significant challenge. In this paper, we propose an Uncertainty-aware Cross-training framework for semi-supervised medical image Segmentation (UC-Seg). Our UC-Seg framework incorporates two distinct subnets to effectively explore and leverage the correlation between them, thereby mitigating cognitive biases within the model. Specifically, we present a Cross-subnet Consistency Preservation (CCP) strategy to enhance feature representation capability and ensure feature consistency across the two subnets. This strategy enables each subnet to correct its own biases and learn shared semantics from both labeled and unlabeled data. Additionally, we propose an Uncertainty-aware Pseudo-label Generation (UPG) component that leverages segmentation results and corresponding uncertainty maps from both subnets to generate high-confidence pseudo-labels. We extensively evaluate the proposed UC-Seg on various medical image segmentation tasks involving different modality images, such as MRI, CT, ultrasound, colonoscopy, and so on. The results demonstrate that our method achieves superior segmentation accuracy and generalization performance compared to other state-of-the-art semi-supervised methods. Our code and segmentation maps will be released at https://github.com/taozh2017/UCSeg.
Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Xiaojun Wu 0001
IEEE Trans. Image Process.4
2025 Learnable Prompting SAM-Induced Knowledge Distillation for Semi-Supervised Medical Image Segmentation
abstract
The limited availability of labeled data has driven advancements in semi-supervised learning for medical image segmentation. Modern large-scale models tailored for general segmentation, such as the Segment Anything Model (SAM), have revealed robust generalization capabilities. However, applying these models directly to medical image segmentation still exposes performance degradation. In this paper, we propose a learnable prompting SAM-induced Knowledge distillation framework (KnowSAM) for semi-supervised medical image segmentation. Firstly, we propose a Multi-view Co-training (MC) strategy that employs two distinct sub-networks to employ a co-teaching paradigm, resulting in more robust outcomes. Secondly, we present a Learnable Prompt Strategy (LPS) to dynamically produce dense prompts and integrate an adapter to fine-tune SAM specifically for medical image segmentation tasks. Moreover, we propose SAM-induced Knowledge Distillation (SKD) to transfer useful knowledge from SAM to two sub-networks, enabling them to learn from SAM's predictions and alleviate the effects of incorrect pseudo-labels during training. Notably, the predictions generated by our subnets are used to produce mask prompts for SAM, facilitating effective inter-module information exchange. Extensive experimental results on various medical segmentation tasks demonstrate that our model outperforms the state-of-the-art semi-supervised segmentation approaches. Crucially, our SAM distillation framework can be seamlessly integrated into other semi-supervised segmentation methods to enhance performance. The code will be released upon acceptance of this manuscript at https://github.com/taozh2017/KnowSAM.
Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Chen Gong 0002, Dong Liang 0001
IEEE Trans. Medical Imaging4
2025 On-the-Fly Improving Segment Anything for Medical Image Segmentation Using Auxiliary Online Learning
abstract
The current variants of the Segment Anything Model (SAM), which include the original SAM and Medical SAM, still lack the capability to produce sufficiently accurate segmentation for medical images. In medical imaging contexts, it is not uncommon for human experts to rectify segmentations of specific test samples after SAM generates its segmentation predictions. These rectifications typically entail manual or semi-manual corrections employing state-of-the-art annotation tools. Motivated by this process, we introduce a novel approach that leverages the advantages of online machine learning to enhance Segment Anything (SA) during test time. We employ rectified annotations to perform online learning, with the aim of improving the segmentation quality of SA on medical images. To ensure the effectiveness and efficiency of online learning when integrated with large-scale vision models like SAM, we propose a new method called Auxiliary Online Learning (AuxOL), which entails adaptive online-batch and adaptive segmentation fusion. Experiments conducted on eight datasets covering four medical imaging modalities validate the effectiveness of the proposed method. Our work proposes and validates a new, practical, and effective approach for enhancing SA on downstream segmentation tasks (e.g., medical image segmentation). The code is publicly available at https://sam-auxol.github.io/AuxOL/.
Tao Zhou 0002, Weidi Xie, Shuo Wang 0011, Qi Dou 0001, Yizhe Zhang 0001
IEEE Trans. Medical Imaging6
2025 Domain-Interactive Contrastive Learning and Prototype-Guided Self-Training for Cross-Domain Polyp Segmentation
abstract
Accurate polyp segmentation plays a critical role in the diagnosis and treatment of colorectal cancer from colonoscopy images. While deep learning-based polyp segmentation models have made significant progress, they often suffer from performance degradation when applied to unseen target domain datasets collected from different imaging devices. To address this challenge, unsupervised domain adaptation (UDA) methods have gained attention by leveraging labeled source data and unlabeled target data to reduce the domain gap. However, existing UDA methods primarily focus on capturing class-wise representations, neglecting domain-wise representations. Additionally, uncertainty in pseudo-labels could hinder the segmentation performance. To tackle these issues, we propose a novel Domain-interactive Contrastive Learning and Prototype-guided Self-training (DCL-PS) framework for cross-domain polyp segmentation. Specifically, domain-interactive contrastive learning (DCL) with a domain-mixed prototype updating strategy is proposed to discriminate class-wise feature representations across domains. Then, to enhance the feature extraction ability of the encoder, we present a contrastive learning-based cross-consistency training (CL-CCT) strategy, which is imposed on both the prototypes obtained by the outputs of the main decoder and perturbed auxiliary outputs. Furthermore, we propose a prototype-guided self-training (PS) strategy, which dynamically assigns a weight for each pixel during self-training, filtering out unreliable pixels and improving the quality of pseudo-labels. Experimental results demonstrate the superiority of DCL-PS in improving polyp segmentation performance in the target domain. The code is released at https://github.com/taozh2017/DCLPS.
Ziru Lu, Yizhe Zhang 0001, Yi Zhou 0007, Ye Wu 0001, Tao Zhou 0002
IEEE Trans. Medical Imaging2
2024 An Empirical Study on the Fairness of Foundation Models for Multi-Organ Image Segmentation
Qing Li 0001, Yizhe Zhang 0001, Yan Li 0064, Longyu Sun, Mengting Sun, Qirong Li, Wenyue Mao, Yinghua Chu, Shuo Wang 0011, Chengyan Wang
MICCAI (12)2
2024 TSBP: Improving Object Detection in Histology Images via Test-Time Self-guided Bounding-Box Propagation
Liang Xiao 0001, Yizhe Zhang 0001
MICCAI (4)3
2024 EndoFinder: Online Image Retrieval for Explainable Colorectal Polyp Diagnosis
Ruijie Yang, Peiyao Fu, Yizhe Zhang 0001, Zhihua Wang 0008, Quanlin Li, Pinghong Zhou, Xian Yang 0001, Shuo Wang 0011
MICCAI (10)4
2024 TextPolyp: Point-Supervised Polyp Segmentation with Text Cues
Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Tao Zhou 0002
MICCAI (11)3
2024 Combining Segment Anything Model with Domain-Specific Knowledge for Semi-Supervised Learning in Medical Image Segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Ye Wu 0001, Pengfei Gu, Shuo Wang 0011
PRCV (14)1
2024 TestFit: A plug-and-play one-pass test time method for medical image segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Yuhui Tao, Shuo Wang 0011, Ye Wu 0001, Benyuan Liu, Pengfei Gu, Qiang Chen 0004, Danny Ziyi Chen
Medical Image Anal.1
2023 MLMSA: Multi-Level and Multi-Scale Attention for Lesion Detection in Endoscopy
abstract
The advancement of deep learning techniques has significantly improved abnormality detection in gastrointestinal (GI) endoscopy. However, this imaging process comes with challenges due to the complex nature of GI abnormalities. The wide variety of abnormalities in terms of type, color, texture, shape, and scale of lesions makes it difficult to accurately detect them in different scenarios. Furthermore, the presence of multiple types of lesions within the same region create complex scenarios that complicate abnormality detection. Additionally, differentiating early-stage cancers from non-cancerous lesions is a significant challenge even for experienced professionals. The simultaneous identification of cancers, particularly early-stage ones, and non-cancerous lesions within the same region remains a challenging issue in GI endoscopy imaging. In this study, we discover that multiple types of lesions exhibit a scale-sensitive characteristic that can be leveraged by multi-level feature-based deep learning models. Hence, we propose the use of a multi-level and multi-scale attention (MLMSA) neck module in a deep learning network. The MLMSA module utilizes multiple levels of features extracted from the backbone network to generate processed multi-level features that assist the detection head. By integrating the MLMSA module into the deep learning framework, our goal is to enhance the detection and differentiation of lesions, particularly the early-stage cancer, thereby advancing the capabilities of GI endoscopy imaging. Our experiment results show that integrating the MLMSA module leads to a significant improvement in the detection of GI abnormalities, providing compelling evidence for the enhanced performance achieved through the utilization of the MLMSA module in our approach.
Shuijiao Chen, Qilei Chen, Yizhe Zhang 0001, Yu Cao 0002, Benyuan Liu
HealthCom6
2023 LAGAN: Lesion-Aware Generative Adversarial Networks for Edema Area Segmentation in SD-OCT Images
abstract
Large volume of labeled data is a cornerstone for deep learning (DL) based segmentation methods. Medical images require domain experts to annotate, and full segmentation annotations of large volumes of medical data are difficult, if not impossible, to acquire in practice. Compared with full annotations, image-level labels are multiple orders of magnitude faster and easier to obtain. Image-level labels contain rich information that correlates with the underlying segmentation tasks and should be utilized in modeling segmentation problems. In this article, we aim to build a robust DL-based lesion segmentation model using only image-level labels (normal v.s. abnormal). Our method consists of three main steps: (1) training an image classifier with image-level labels; (2) utilizing a model visualization tool to generate an object heat map for each training sample according to the trained classifier; (3) based on the generated heat maps (as pseudo-annotations) and an adversarial learning framework, we construct and train an image generator for Edema Area Segmentation (EAS). We name the proposed method Lesion-Aware Generative Adversarial Networks (LAGAN) as it combines the merits of supervised learning (being lesion-aware) and adversarial training (for image generation). Additional technical treatments, such as the design of a multi-scale patch-based discriminator, further enhance the effectiveness of our proposed method. We validate the superior performance of LAGAN via comprehensive experiments on two publicly available datasets (i.e., AI Challenger and RETOUCH).
Yuhui Tao, Xiao Ma 0011, Yizhe Zhang 0001, Zexuan Ji, Wen Fan 0003, Songtao Yuan, Qiang Chen 0004
IEEE J. Biomed. Health Informatics3
2023 Flexible Fusion Network for Multi-Modal Brain Tumor Segmentation
abstract
Automated brain tumor segmentation is crucial for aiding brain disease diagnosis and evaluating disease progress. Currently, magnetic resonance imaging (MRI) is a routinely adopted approach in the field of brain tumor segmentation that can provide different modality images. It is critical to leverage multi-modal images to boost brain tumor segmentation performance. Existing works commonly concentrate on generating a shared representation by fusing multi-modal data, while few methods take into account modality-specific characteristics. Besides, how to efficiently fuse arbitrary numbers of modalities is still a difficult task. In this study, we present a flexible fusion network (termed F$^{2}$Net) for multi-modal brain tumor segmentation, which can flexibly fuse arbitrary numbers of multi-modal information to explore complementary information while maintaining the specific characteristics of each modality. Our F$^{2}$Net is based on the encoder-decoder structure, which utilizes two Transformer-based feature learning streams and a cross-modal shared learning network to extract individual and shared feature representations. To effectively integrate the knowledge from the multi-modality data, we propose a cross-modal feature-enhanced module (CFM) and a multi-modal collaboration module (MCM), which aims at fusing the multi-modal features into the shared learning network and incorporating the features from encoders into the shared decoder, respectively. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our F$^{2}$Net over other state-of-the-art segmentation methods.
Hengyi Yang, Tao Zhou 0002, Yi Zhou 0007, Yizhe Zhang 0001, Huazhu Fu
IEEE J. Biomed. Health Informatics4
2022 Real-Time, Accurate, and Consistent Video Semantic Segmentation via Unsupervised Adaptation and Cross-Unit Deployment on Mobile Device
abstract
This demonstration showcases our innovations on efficient, accurate, and temporally consistent video semantic segmentation on mobile device. We employ our test-time unsupervised scheme, AuxAdapt, to enable the segmentation model to adapt to a given video in an online manner. More specifically, we leverage a small auxiliary network to perform weight updates and keep the large, main segmen-tation network frozen. This significantly reduces the computational cost of adaptation when compared to previous methods (e.g., Tent, DVP), and at the same time, prevents catastrophic forgetting. By running AuxAdapt, we can considerably improve the temporal consistency of video segmentation while maintaining the accuracy. We demonstrate how to efficiently deploy our adaptive video segmentation algorithm on a smartphone powered by a Snapdragon® Mobile Platform11Snapdragon is a product of Qualcomm Technologies, Inc. and/or its subsidiaries., Rather than simply running the entire algorithm on the GPU, we adopt a crossunit deployment strategy. The main network, which will be frozen during test time, will perform inferences on a highly optimized AI accelerator unit, while the small auxiliary net-work, which will be updated on the fly, will run forward passes and back-propagations on the GPU. Such a deployment scheme best utilizes the available processing power on the smartphone and enables real-time operation of our adaptive video segmentation algorithm. We provide example videos in supplementary material.
Hyojin Park 0004, Alan Yessenbayev, Tushar Singhal, Navin Kumar Adhikari, Yizhe Zhang 0001, Shubhankar Borse, Frank Mayer, Balaji Calidas, Nilesh Prasad Pandey, Fatih Porikli
CVPR5
2022 Data-Driven Deep Supervision for Skin Lesion Classification
Suraj Mishra, Yizhe Zhang 0001, Li Zhang 0021, Tianyu Zhang 0001, Xiaobo Sharon Hu, Danny Ziyi Chen
MICCAI (1)2
2022 Usable Region Estimate for Assessing Practical Usability of Medical Image Segmentation Models
Yizhe Zhang 0001, Suraj Mishra, Peixian Liang, Hao Zheng 0006, Danny Ziyi Chen
MICCAI (5)1
2022 AuxAdapt: Stable and Efficient Test-Time Adaptation for Temporally Consistent Video Semantic Segmentation
abstract
In video segmentation, generating temporally consistent results across frames is as important as achieving frame-wise accuracy. This paper presents an efficient, intuitive, and unsupervised online adaptation method, AuxAdapt, for improving the temporal consistency of most neural network models. It does not require optical flow and only takes one pass of the video. Since inconsistency mainly arises from the model’s uncertainty in its output, we propose an adaptation scheme where the model learns from its own segmentation decisions as it streams a video, which allows producing more confident and temporally consistent labeling for similarly-looking pixels across frames. For stability and efficiency, we leverage a small auxiliary segmentation network (AuxNet) to assist with this adaptation. More specifically, AuxNet readjusts the decision of the original segmentation network (Main-Net) by adding its own estimations to that of MainNet. At every frame, only AuxNet is updated via back-propagation while keeping MainNet fixed. We extensively evaluate our test-time adaptation approach on standard video benchmarks, including Cityscapes, CamVid, and KITTI. The results demonstrate that our approach provides label-wise accurate, temporally consistent, and computationally efficient adaptation.
Yizhe Zhang 0001, Shubhankar Borse, Fatih Porikli
WACV1
2022 Perceptual Consistency in Video Segmentation
abstract
In this paper, we present a novel perceptual consistency perspective on video semantic segmentation, which can capture both temporal consistency and pixel-wise correctness. Given two nearby video frames, perceptual consistency measures how much the segmentation decisions agree with the pixel correspondences obtained via matching general perceptual features. More specifically, for each pixel in one frame, we find the most perceptually correlated pixel in the other frame. Our intuition is that such a pair of pixels are highly likely to belong to the same class. Next, we assess how much the segmentation agrees with such perceptual correspondences, based on which we derive the perceptual consistency of the segmentation maps across these two frames. Utilizing perceptual consistency, we can evaluate the temporal consistency of video segmentation by measuring the perceptual consistency over consecutive pairs of segmentation maps in a video. Furthermore, given a sparsely labeled test video, perceptual consistency can be utilized to aid with predicting the pixel-wise correctness of the segmentation on an unlabeled frame. More specifically, by measuring the perceptual consistency between the predicted segmentation and the available ground truth on a nearby frame and combining it with the segmentation confidence, we can accurately assess the classification correctness on each pixel. Our experiments show that the proposed perceptual consistency can more accurately evaluate the temporal consistency of video segmentation as compared to flow-based measures. Furthermore, it can help more confidently predict segmentation accuracy on unlabeled test frames, as compared to using classification confidence alone. Finally, our proposed measure can be used as a regularizer during the training of segmentation models, which leads to more temporally consistent video segmentation while maintaining accuracy.
Yizhe Zhang 0001, Shubhankar Borse, Ying Wang 0051, Ning Bi, Xiaoyun Jiang, Fatih Porikli
WACV1
2022 H-EMD: A Hierarchical Earth Mover's Distance Method for Instance Segmentation
abstract
Deep learning (DL) based semantic segmentation methods have achieved excellent performance in biomedical image segmentation, producing high quality probability maps to allow extraction of rich instance information to facilitate good instance segmentation. While numerous efforts were put into developing new DL semantic segmentation models, less attention was paid to a key issue of how to effectively explore their probability maps to attain the best possible instance segmentation. We observe that probability maps by DL semantic segmentation models can be used to generate many possible instance candidates, and accurate instance segmentation can be achieved by selecting from them a set of "optimized" candidates as output instances. Further, the generated instance candidates form a well-behaved hierarchical structure (a forest), which allows selecting instances in an optimized manner. Hence, we propose a novel framework, called hierarchical earth mover's distance (H-EMD), for instance segmentation in biomedical 2D+time videos and 3D images, which judiciously incorporates consistent instance selection with semantic-segmentation-generated probability maps. H-EMD contains two main stages: (1) instance candidate generation: capturing instance-structured information in probability maps by generating many instance candidates in a forest structure; (2) instance candidate selection: selecting instances from the candidate set for final instance segmentation. We formulate a key instance selection problem on the instance candidate forest as an optimization problem based on the earth mover's distance (EMD), and solve it by integer linear programming. Extensive experiments on eight biomedical video or 3D datasets demonstrate that H-EMD consistently boosts DL semantic segmentation models and is highly competitive with state-of-the-art methods.
Peixian Liang, Yizhe Zhang 0001, Yifan Ding 0001, Jianxu Chen 0001, Chinedu S. Madukoma, Tim Weninger, Joshua D. Shrout, Danny Ziyi Chen
IEEE Trans. Medical Imaging2
2022 Data-Driven Deep Supervision for Medical Image Segmentation
abstract
Medical image segmentation plays a vital role in disease diagnosis and analysis. However, data-dependent difficulties such as low image contrast, noisy background, and complicated objects of interest render the segmentation problem challenging. These difficulties diminish dense prediction and make it tough for known approaches to explore data-specific attributes for robust feature extraction. In this paper, we study medical image segmentation by focusing on robust data-specific feature extraction to achieve improved dense prediction. We propose a new deep convolutional neural network (CNN), which exploits specific attributes of input datasets to utilize deep supervision for enhanced feature extraction. In particular, we strategically locate and deploy auxiliary supervision, by matching the object perceptive field (OPF) (which we define and compute) with the layer-wise effective receptive fields (LERF) of the network. This helps the model pay close attention to some distinct input data dependent features, which the network might otherwise 'ignore' during training. Further, to achieve better target localization and refined dense prediction, we propose the densely decoded networks (DDN), by selectively introducing additional network connections (the 'crutch' connections). Using five public datasets (two retinal vessel, melanoma, optic disc/cup, and spleen segmentation) and two in-house datasets (lymph node and fungus segmentation), we verify the effectiveness of our proposed approach in 2D and 3D segmentation.
Suraj Mishra, Yizhe Zhang 0001, Danny Ziyi Chen, Xiaobo Sharon Hu
IEEE Trans. Medical Imaging2
2021 HS3: Learning with Proper Task Complexity in Hierarchically Supervised Semantic Segmentation
Shubhankar Borse, Yizhe Zhang 0001, Fatih Porikli
BMVC3
2021 X-Distill: Improving Self-Supervised Monocular Depth via Cross-Task Distillation
Janarbek Matai, Shubhankar Borse, Yizhe Zhang 0001, Amin Ansari, Fatih Porikli
BMVC4
2021 InverseForm: A Loss Function for Structured Boundary-Aware Segmentation
abstract
We present a novel boundary-aware loss term for semantic segmentation using an inverse-transformation network, which efficiently learns the degree of parametric transformations between estimated and target boundaries. This plug-in loss term complements the cross-entropy loss in capturing boundary transformations and allows consistent and significant performance improvement on segmentation backbone models without increasing their size and computational complexity. We analyze the quantitative and qualitative effects of our loss function on three indoor and outdoor segmentation benchmarks, including Cityscapes, NYU-Depth-v2, and PASCAL, integrating it into the training phase of several backbone networks in both single-task and multi-task settings. Our extensive experiments show that the proposed method consistently outperforms base-lines, and even sets the new state-of-the-art on two datasets.
Shubhankar Borse, Ying Wang 0051, Yizhe Zhang 0001, Fatih Porikli
CVPR3
2021 kCBAC-Net: Deeply Supervised Complete Bipartite Networks with Asymmetric Convolutions for Medical Image Segmentation
Pengfei Gu, Hao Zheng 0006, Yizhe Zhang 0001, Chaoli Wang 0001, Danny Ziyi Chen
MICCAI (1)3
2020 An Annotation Sparsification Strategy for 3D Medical Image Segmentation via Representative Selection and Self-Training
abstract
Image segmentation is critical to lots of medical applications. While deep learning (DL) methods continue to improve performance for many medical image segmentation tasks, data annotation is a big bottleneck to DL-based segmentation because (1) DL models tend to need a large amount of labeled data to train, and (2) it is highly time-consuming and label-intensive to voxel-wise label 3D medical images. Significantly reducing annotation effort while attaining good performance of DL segmentation models remains a major challenge. In our preliminary experiments, we observe that, using partially labeled datasets, there is indeed a large performance gap with respect to using fully annotated training datasets. In this paper, we propose a new DL framework for reducing annotation effort and bridging the gap between full annotation and sparse annotation in 3D medical image segmentation. We achieve this by (i) selecting representative slices in 3D images that minimize data redundancy and save annotation effort, and (ii) self-training with pseudo-labels automatically generated from the base-models trained using the selected annotated slices. Extensive experiments using two public datasets (the HVSMR 2016 Challenge dataset and mouse piriform cortex dataset) show that our framework yields competitive segmentation results comparing with state-of-the-art DL methods using less than ~ 20% of annotated data.
Hao Zheng 0006, Yizhe Zhang 0001, Lin Yang 0003, Chaoli Wang 0001, Danny Ziyi Chen
AAAI2
2020 Unlabeled Data Guided Semi-supervised Histopathology Image Segmentation
abstract
Automatic histopathology image segmentation is crucial to disease analysis. Limited available labeled data hinders the generalizability of trained models under the fully supervised setting. Semi-supervised learning (SSL) based on generative methods has been proven to be effective in utilizing diverse image characteristics. However, it has not been well explored what kinds of generated images would be more useful for model training and how to use such images. In this paper, we propose a new data guided generative method for histopathology image segmentation by leveraging the unlabeled data distributions. First, we design an image generation module. Image content and style are disentangled and embedded in a clustering-friendly space to utilize their distributions. New images are synthesized by sampling and cross-combining contents and styles. Second, we devise an effective data selection policy for judiciously sampling the generated images: (1) to make the generated training set better cover the dataset, the clusters that are underrepresented in the original training set are covered more; (2) to make the training process more effective, we identify and oversample the images of “hard cases” in the data for which annotated training data may be scarce. Our method is evaluated on glands and nuclei datasets. We show that under both the inductive and transductive settings, our SSL method consistently boosts the performance of common segmentation models and attains state-of-the-art results.
Hao Zheng 0006, Jianxu Chen 0001, Lin Yang 0003, Yizhe Zhang 0001, Danny Ziyi Chen
BIBM5
2020 InTracker: An Integrated Detector-Tracker Framework for Cell Detection and Tracking
abstract
Automatic tracking of moving cells in time-lapse image sequences plays an important role in studying many biological processes in development and diseases. Large variations in cell appearances, limited image resolution, and various cell behaviors (e.g., division, apoptosis, deformation, clustering, and migration in or out of the imaging window) make cell tracking a challenging task. However, known cell tracking methods were designed for and tailored to specific cell image sequences and behaviors, thus having limited applicability to various cell image sequences. Aiming toward more robust cell tracking, we propose a new detector-tracker approach for detection and association based cell tracking. First, we propose a new deep learning based detector to detect cells in each image frame and assign division/non-division labels to them. Second, we carefully design an Earth Mover's Distance (EMD) based hierarchical tracker to associate detected cells through the image sequence and form moving cell trajectories. The tracker is able to correct possible detection errors made by the detector. Evaluated on several open challenge datasets, our approach outperforms state-of-the-art cell tracking methods for determining cell trajectories.
Peixian Liang, Jianxu Chen 0001, Yizhe Zhang 0001, Hao Zheng 0006, Pengfei Gu, Danny Ziyi Chen
CBMS3
2020 A Coarse-to-Fine Data Generation Method for 2D and 3D Cell Nucleus Segmentation
abstract
Cell nucleus segmentation is a fundamental task in biomedical image analysis. Generating realistic cell nucleus data with ground truth masks can help tackle difficulties such as insufficient training data for deep learning models and the need to deal with "hard" cases (e.g., tightly clumped nuclei). Known nucleus generation methods generated individual nucleus masks from parametric models or based on direct transformations of real masks. It is difficult for these methods to capture and simulate the distributions of real nuclei and interactions among hard nuclei. In this paper, we propose a new three-stage coarse-to-fine nucleus generation method for 2D and 3D nucleus segmentation. The first stage simulates the positions and sizes of nuclei; the second stage simulates the shapes of nuclei and interactions among clumped nuclei; the third stage simulates the textures of nuclei. We evaluate our method on 2D and 3D cell nucleus image datasets. Experimental results show that our new nucleus generation method considerably helps improve cell nucleus segmentation performance and outperforms known nucleus generation methods for nucleus segmentation with a small amount of training data.
Zhuo Zhao, Yizhe Zhang 0001, Hao Zheng 0006, Danny Ziyi Chen
CBMS3
2020 Structured Convolutions for Efficient Neural Network Design
abstract
In this work, we tackle model efficiency by exploiting redundancy in the implicit structure of the building blocks of convolutional neural networks. We start our analysis by introducing a general definition of Composite Kernel structures that enable the execution of convolution operations in the form of efficient, scaled, sum-pooling components. As its special case, we propose Structured Convolutions and show that these allow decomposition of the convolution operation into a sum-pooling operation followed by a convolution with significantly lower complexity and fewer weights. We show how this decomposition can be applied to 2D and 3D kernels as well as the fully-connected layers. Furthermore, we present a Structural Regularization loss that promotes neural network layers to leverage on this desired structure in a way that, after training, they can be decomposed with negligible performance loss. By applying our method to a wide range of CNN architectures, we demonstrate 'structured' versions of the ResNets that are up to 2x smaller and a new Structured-MobileNetV2 that is more efficient while staying within an accuracy loss of 1% on ImageNet and CIFAR-10 datasets. We also show similar structured versions of EfficientNet on ImageNet and HRNet architecture for semantic segmentation on the Cityscapes dataset. Our method performs equally well or superior in terms of the complexity reduction in comparison to the existing tensor decomposition and channel pruning methods.
Yash Bhalgat, Yizhe Zhang 0001, Jamie Menjay Lin, Fatih Porikli
NeurIPS2
2020 A Cross-Domain Metal Trace Restoring Network for Reducing X-Ray CT Metal Artifacts
abstract
Metal artifacts commonly appear in computed tomography (CT) images of the patient body with metal implants and can affect disease diagnosis. Known deep learning and traditional metal trace restoring methods did not effectively restore details and sinogram consistency information in X-ray CT sinograms, hence often causing considerable secondary artifacts in CT images. In this paper, we propose a new cross-domain metal trace restoring network which promotes sinogram consistency while reducing metal artifacts and recovering tissue details in CT images. Our new approach includes a cross-domain procedure that ensures information exchange between the image domain and the sinogram domain in order to help them promote and complement each other. Under this cross-domain structure, we develop a hierarchical analytic network (HAN) to recover fine details of metal trace, and utilize the perceptual loss to guide HAN to concentrate on the absorption of sinogram consistency information of metal trace. To allow our entire cross-domain network to be trained end-to-end efficiently and reduce the graphic memory usage and time cost, we propose effective and differentiable forward projection (FP) and filtered back-projection (FBP) layers based on FP and FBP algorithms. We use both simulated and clinical datasets in three different clinical scenarios to evaluate our proposed network's practicality and universality. Both quantitative and qualitative evaluation results show that our new network outperforms state-of-the-art metal artifact reduction methods. In addition, the elapsed time analysis shows that our proposed method meets the clinical time requirement.
Chengtao Peng, Bin Li 0025, Peixian Liang, Jian Zheng 0001, Yizhe Zhang 0001, Bensheng Qiu, Danny Ziyi Chen
IEEE Trans. Medical Imaging5
2020 AntVis: A web-based visual analytics tool for exploring ant movement data
abstract
We present AntVis, a web-based visual analytics tool for exploring ant movement data collected from the video recording of ants moving on tree branches. Our goal is to enable domain experts to visually explore massive ant movement data and gain valuable insights via effective visualization, filtering, and comparison. This is achieved through a deep learning framework for automatic detection, segmentation, and labeling of ants, ant movement clustering based on their trace similarity, and the design and development of five coordinated views (the movement, similarity, timeline, statistical, and attribute views) for user interaction and exploration. We demonstrate the effectiveness of AntVis with several case studies developed in close collaboration with domain experts. Finally, we report the expert evaluation conducted by an entomologist and point out future directions of this study.
Tianxiao Hu, Hao Zheng 0006, Sirou Zhu, Natalie Imirzian, Yizhe Zhang 0001, Chaoli Wang 0001, David P. Hughes, Danny Ziyi Chen
Vis. Informatics6
2019 Biomedical Image Segmentation via Representative Annotation
abstract
Deep learning has been applied successfully to many biomedical image segmentation tasks. However, due to the diversity and complexity of biomedical image data, manual annotation for training common deep learning models is very timeconsuming and labor-intensive, especially because normally only biomedical experts can annotate image data well. Human experts are often involved in a long and iterative process of annotation, as in active learning type annotation schemes. In this paper, we propose representative annotation (RA), a new deep learning framework for reducing annotation effort in biomedical image segmentation. RA uses unsupervised networks for feature extraction and selects representative image patches for annotation in the latent space of learned feature descriptors, which implicitly characterizes the underlying data while minimizing redundancy. A fully convolutional network (FCN) is then trained using the annotated selected image patches for image segmentation. Our RA scheme offers three compelling advantages: (1) It leverages the ability of deep neural networks to learn better representations of image data; (2) it performs one-shot selection for manual annotation and frees annotators from the iterative process of common active learning based annotation schemes; (3) it can be deployed to 3D images with simple extensions. We evaluate our RA approach using three datasets (two 2D and one 3D) and show our framework yields competitive segmentation results comparing with state-of-the-art methods.
Hao Zheng 0006, Lin Yang 0003, Jianxu Chen 0001, Jun Han 0010, Yizhe Zhang 0001, Peixian Liang, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
AAAI5
2019 A New Ensemble Learning Framework for 3D Biomedical Image Segmentation
abstract
3D image segmentation plays an important role in biomedical image analysis. Many 2D and 3D deep learning models have achieved state-of-the-art segmentation performance on 3D biomedical image datasets. Yet, 2D and 3D models have their own strengths and weaknesses, and by unifying them together, one may be able to achieve more accurate results. In this paper, we propose a new ensemble learning framework for 3D biomedical image segmentation that combines the merits of 2D and 3D models. First, we develop a fully convolutional network based meta-learner to learn how to improve the results from 2D and 3D models (base-learners). Then, to minimize over-fitting for our sophisticated meta-learner, we devise a new training method that uses the results of the baselearners as multiple versions of “ground truths”. Furthermore, since our new meta-learner training scheme does not depend on manual annotation, it can utilize abundant unlabeled 3D image data to further improve the model. Extensive experiments on two public datasets (the HVSMR 2016 Challenge dataset and the mouse piriform cortex dataset) show that our approach is effective under fully-supervised, semisupervised, and transductive settings, and attains superior performance over state-of-the-art image segmentation methods.
Hao Zheng 0006, Yizhe Zhang 0001, Lin Yang 0003, Peixian Liang, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
AAAI2
2019 Second-Order Non-Local Attention Networks for Person Re-Identification
abstract
Recent efforts have shown promising results for person re-identification by designing part-based architectures to allow a neural network to learn discriminative representations from semantically coherent parts. Some efforts use soft attention to reallocate distant outliers to their most similar parts, while others adjust part granularity to incorporate more distant positions for learning the relationships. Others seek to generalize part-based methods by introducing a dropout mechanism on consecutive regions of the feature map to enhance distant region relationships. However, only few prior efforts model the distant or non-local positions of the feature map directly for the person re-ID task. In this paper, we propose a novel attention mechanism to directly model long-range relationships via second-order feature statistics. When combined with a generalized DropBlock module, our method performs equally to or better than state-of-the-art results for mainstream person re-identification datasets, including Market1501, CUHK03, and DukeMTMC-reID.
Bryan Bryan, Yuan Gong 0001, Yizhe Zhang 0001, Christian Poellabauer
ICCV3
2019 Decompose-and-Integrate Learning for Multi-class Segmentation in Medical Images
Yizhe Zhang 0001, Michael T. C. Ying, Danny Ziyi Chen
MICCAI (2)1
2019 HFA-Net: 3D Cardiovascular Image Segmentation with Asymmetrical Pooling and Content-Aware Fusion
Hao Zheng 0006, Lin Yang 0003, Jun Han 0010, Yizhe Zhang 0001, Peixian Liang, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
MICCAI (2)4
2018 Collaborative Random Faces-Guided Encoders for Pose-Invariant Face Representation Learning
abstract
Learning discriminant face representation for pose-invariant face recognition has been identified as a critical issue in visual learning systems. The challenge lies in the drastic changes of facial appearances between the test face and the registered face. To that end, we propose a high-level feature learning framework called "collaborative random faces (RFs)-guided encoders" toward this problem. The contributions of this paper are three fold. First, we propose a novel supervised autoencoder that is able to capture the high-level identity feature despite of pose variations. Second, we enrich the identity features by replacing the target values of conventional autoencoders with random signals (RFs in this paper), which are unique for each subject under different poses. Third, we further improve the performance of the framework by incorporating deep convolutional neural network facial descriptors and linking discriminative identity features from different RFs for the augmented identity features. Finally, we conduct face identification experiments on Multi-PIE database, and face verification experiments on labeled faces in the wild and YouTube Face databases, where face recognition rate and verification accuracy with Receiver Operating Characteristic curves are rendered. In addition, discussions of model parameters and connections with the existing methods are provided. These experiments demonstrate that our learning system works fairly well on handling pose variations.
Ming Shao, Yizhe Zhang 0001, Yun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 Suggestive Annotation: A Deep Active Learning Framework for Biomedical Image Segmentation
Lin Yang 0003, Yizhe Zhang 0001, Jianxu Chen 0001, Danny Ziyi Chen
MICCAI (3)2
2017 Deep Adversarial Networks for Biomedical Image Segmentation Utilizing Unannotated Images
Yizhe Zhang 0001, Lin Yang 0003, Jianxu Chen 0001, Maridel Fredericksen, David P. Hughes, Danny Ziyi Chen
MICCAI (3)1
2016 Coarse-to-Fine Stacked Fully Convolutional Nets for lymph node segmentation in ultrasound images
abstract
Ultrasound as a well-established imaging modality is widely used in imaging lymph nodes for clinical diagnosis and disease analysis. Quantitative analysis of lymph node features, morphology, and relations can provide valuable information for diagnosis and immune system studies. For such analysis, it is necessary to first accurately segment the lymph node areas in ultrasound images. In this paper, we develop a new deep learning method, called Coarse-to-Fine Stacked Fully Convolutional Nets (CFS-FCN), for automatically segmenting lymph nodes in ultrasound images. Our method consists of multiple stages of FCN modules. We train the CFS-FCN model to learn the segmentation knowledge from a coarse-to-fine, simple-to-complex manner. A data set of 80 ultrasound images containing both normal and diseased lymph nodes is used in our experiments, which show that our method considerably outperforms the state-of-the-art deep learning methods for lymph node segmentation.
Yizhe Zhang 0001, Michael T. C. Ying, Lin Yang 0003, Anil T. Ahuja, Danny Ziyi Chen
BIBM1
2016 3D Segmentation of Glial Cells Using Fully Convolutional Networks and k-Terminal Cut
Lin Yang 0003, Yizhe Zhang 0001, Ian H. Guldner, Danny Ziyi Chen
MICCAI (2)2
2016 Combining Fully Convolutional and Recurrent Neural Networks for 3D Biomedical Image Segmentation
abstract
Segmentation of 3D images is a fundamental problem in biomedical image analysis. Deep learning (DL) approaches have achieved the state-of-the-art segmentation performance. To exploit the 3D contexts using neural networks, known DL segmentation methods, including 3D convolution, 2D convolution on the planes orthogonal to 2D slices, and LSTM in multiple directions, all suffer incompatibility with the highly anisotropic dimensions in common 3D biomedical images. In this paper, we propose a new DL framework for 3D image segmentation, based on a combination of a fully convolutional network (FCN) and a recurrent neural network (RNN), which are responsible for exploiting the intra-slice and inter-slice contexts, respectively. To our best knowledge, this is the first DL framework for 3D image segmentation that explicitly leverages 3D image anisotropism. Evaluating using a dataset from the ISBI Neuronal Structure Segmentation Challenge and in-house image stacks for 3D fungus segmentation, our approach achieves promising results, comparing to the known DL-based 3D segmentation approaches.
Jianxu Chen 0001, Lin Yang 0003, Yizhe Zhang 0001, Mark S. Alber, Danny Ziyi Chen
NIPS3
2015 A seeding-searching-ensemble method for gland segmentation and detection
abstract
Glands are vital tissues found throughout the human body and their structure and function are affected by many diseases. The ability to segment and detect glands among other types of tissues is important for the study of normal and disease processes and is readily visualized by pathologists in microscopic detail. In this paper, we develop a new approach for segmenting and detecting intestinal glands in H&E stained histology images, which utilizes a set of advanced image processing techniques such as graph search, ensemble, feature extraction and classification. Our method computes fast, and is able to preserve gland boundaries robustly and detect glands accurately. We tested the performance of gland detection and segmentation by analyzing a dataset of 1723 glands from digitized high-resolution clinical histology images obtained in normal and diseased intestines. The experimental results show that our method outperforms considerably the state-of-the-art methods for gland segmentation and detection tasks.
Yizhe Zhang 0001, Lin Yang 0003, John D. MacKenzie, Rageshree Ramachandran, Danny Ziyi Chen
BIBM1
2015 Fast Background Removal in 3D Fluorescence Microscopy Images Using One-Class Learning
Lin Yang 0003, Yizhe Zhang 0001, Ian H. Guldner, Danny Ziyi Chen
MICCAI (3)2
2013 Random Faces Guided Sparse Many-to-One Encoder for Pose-Invariant Face Recognition
abstract
One of the most challenging task in face recognition is to identify people with varied poses. Namely, the test faces have significantly different poses compared with the registered faces. In this paper, we propose a high-level feature learning scheme to extract pose-invariant identity feature for face recognition. First, we build a single-hidden-layer neural network with sparse constraint, to extract pose-invariant feature in a supervised fashion. Second, we further enhance the discriminative capability of the proposed feature by using multiple random faces as the target values for multiple encoders. By enforcing the target values to be unique for input faces over different poses, the learned high-level feature that is represented by the neurons in the hidden layer is pose free and only relevant to the identity information. Finally, we conduct face identification on CMU Multi-PIE, and verification on Labeled Faces in the Wild (LFW) databases, where identification rank-1 accuracy and face verification accuracy with ROC curve are reported. These experiments demonstrate that our model is superior to other state-of-the-art approaches on handling pose variations.
Yizhe Zhang 0001, Ming Shao, Edward K. Wong, Yun Fu 0001
ICCV1