Chu Han

dblp:125/7005 · DBLP profile ↗
← Back
52ranked-venue papers
4as first author
38since 2021 · last 2026
0000-0001-7557-9131ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unsupervised single-domain generalization for tissue classification via progressive domain transformation
Jiatai Lin, Yanfen Cui, Bingchao Zhao, Tianpeng Deng, Jingqi Huang, Zhenwei Shi 0002, Enming Cui, Zaiyi Liu, Chu Han
Medical Image Anal.11
2026 FKDNuSeg: Flawless knowledge distillation for lightweight and fast nuclei instance segmentation and classification
Bingchao Zhao, Jingxin Luo, Jiatai Lin, Tianpeng Deng, Zaiyi Liu, Guoqiang Han 0002, Chu Han
Medical Image Anal.8
2026 QuPaS: SAM-Based Semi-Supervised Histopathological Image Segmentation With Quantum Force Field Finetuning and Adversarial Estimation
abstract
Semi-supervised segmentation (S3) is one of the preferred choices for histopathological image segmentation tasks, while how to improve model’s learning capability for unlabeled data remains a key challenge in S3. The remarkable feature extraction abilities of Segment Anything Model (SAM) offers a potential opportunity. However, SAM’s performance on contextual complex histopathological images is not so desirable due to its limitations in finely capture structural relationships. To address this issue, we propose a novel SAM-based S3framework QuPaS, which consists of Quantum Force Field (QFF) Finetuning and Adversarial Estimation (AE). QFF covers the shortage of SAM’s limited understanding of spatial structure by simulating intermolecular forces to explore the structural topological relationships between pixel-level features. AE introduces an adversarial estimation network to align the consistency of confidence distributions between different outputs, thereby reducing the interference of incompatible semantic features on the model. Extensive experiments across three challenging histopathological segmentation scenarios have demonstrate that our QuPaS completely outperforms the state-of-the-art S3methods. Furthermore, QuPaS is able to maintain stable generalization performance on previously unseen domains. The code will be released at: https://github.com/director87/QuPaS.
Siyang Feng, Xipeng Pan, Weidong Zhang 0007, Minghua Pan, Chu Han, Rushi Lan
IEEE Trans. Medical Imaging5
2025 Weakly Supervised Gland Segmentation with Class Semantic Consistency and Purified Labels Filtration
abstract
Image-level weakly supervised semantic segmentation (WSSS) reduces the dependence on high-quality data annotation, which plays a crucial role in computational pathology. Benefit from the ability to localize the objects with only binary labels, Class Activation Map (CAM) is a widely used method to initial pseudo masks. However, due to the low contrast among different tissues in histopathological images, most existing CAM-based methods perform poorly in gland segmentation. We retrospect this process and find that class consistency and semantic consistency can guide the network to effectively distinguish confusing pixels and generate fine-grained pseudo masks. Specifically, for class consistency, we propose Consistency Correlation Attention (CCA) to encourage the network to focus on the contribution of class features to semantic dependencies. For semantic consistency, we propose Multi-scale Pyramid Fusion Pooling (MPFP) to aggregate coarse-to-fine global semantic information from CAMs at multiple spatial resolutions, thus identifying class localization. Additionally, we introduce a Purified Labels Filtration (PLF) strategy during the segmentation phase to mitigate the noisy supervision signal and improve the segmentation quality of the model. Extensive experiments show that the our method achieves new state-of-the-art results on three publicly available gland datasets. Furthermore, our method demonstrates impressive domain adaptation capability, achieving satisfactory results with only a small portion of samples when faced with unseen domain data.
Siyang Feng, Huadeng Wang, Chu Han, Zhenbing Liu, Hualong Zhang, Rushi Lan, Xipeng Pan
AAAI3
2025 IPAU: Integrating Prototype, Affinity, and Uncertainty for Weakly-Supervised Histopathology Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) reduces annotation burden by using only image-level labels for histopathology image segmentation. Current WSSS methods face the challenge of bridging the information gap between weak labels and dense prediction tasks, which often results in insufficient class activation maps (CAMs) and increased false positives. Most approaches address this by mining additional object-related information. Following this direction, we propose IPAU, a framework that integrates Prototype, Affinity, and Uncertainty to enhance WSSS. In our IPAU, Prototype-based Information Enhancement (PIE) that uses class-wise prototypes to enrich CAM generation. Affinity-based Self-Refinement (ASR) that refines CAMs into pseudo-masks using affinity correlations without extra training. And Uncertainty-Aware PseudoSupervision (UAPS) that mitigates noise by focusing learning on reliable regions. Experiments on two public histopathology WSSS datasets demonstrate that our IPAU achieves state-of-theart performance.
Jiatai Lin, Jingxuan Zhou, Zhenwei Shi 0002, Zaiyi Liu, Xiao-jing Guo, Chu Han
BIBM6
2025 Rethinking mitosis detection: Towards diverse data and feature representation for better domain generalization
Jiatai Lin, Danyi Li, Bingchao Zhao, Zhenwei Shi 0002, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu, Chu Han
Artif. Intell. Medicine11
2025 Weakly supervised histopathology tissue semantic segmentation with multi-scale voting and online noise suppression
Xipeng Pan, Hualong Zhang, Huahu Deng, Huadeng Wang, Lingqiao Li, Zhenbing Liu, Yajun An, Cheng Lu 0001, Zaiyi Liu, Chu Han, Rushi Lan
Eng. Appl. Artif. Intell.11
2025 Multi-phase feature-aligned fusion model for automated colorectal cancer segmentation in contrast-enhanced CT scans
Xuewei Kang, Suyun Li, Zhanzhu Lin, Bingjiang Qiu, Chu Han, Yun Mao, Zaiyi Liu, Xin Chen 0058
Expert Syst. Appl.8
2025 Label-efficient transformer-based framework with self-supervised strategies for heterogeneous lung tumor segmentation
Zhenbing Liu, Yanfen Cui, Xin Chen 0058, Xipeng Pan, Guanchao Ye, Guangyao Wu, Yongde Liao, Leroy Volmer, Leonard Wee, Andre Dekker, Chu Han, Zaiyi Liu, Zhenwei Shi 0002
Expert Syst. Appl.12
2025 FedBCD: Federated Ultrasound Video and Image Joint Learning for Breast Cancer Diagnosis
abstract
Ultrasonography plays an essential role in breast cancer diagnosis. Current deep learning based studies train the models on either images or videos in a centralized learning manner, lacking consideration of joint benefits between two different modality models or the privacy issue of data centralization. In this study, we propose the first decentralized learning solution for joint learning with breast ultrasound video and image, called FedBCD. To enable the model to learn from images and videos simultaneously and seamlessly in client-level local training, we propose a Joint Ultrasound Video and Image Learning (JUVIL) model to bridge the dimension gap between video and image data by incorporating temporal and spatial adapters. The parameter-efficient design of JUVIL with trainable adapters and frozen backbone further reduces the computational cost and communication burden of federated learning, finally improving the overall efficiency. Moreover, considering conventional model-wise aggregation may lead to unstable federated training due to different modalities, data capacities in different clients, and different functionalities across layers. We further propose a Fisher information matrix (FIM) guided Layer-wise Aggregation method named FILA. By measuring layer-wise sensitivity with FIM, FILA assigns higher contributions to the clients with lower sensitivity, improving personalized performance during federated training. Extensive experiments on three image clients and one video client demonstrate the benefits of joint learning architecture, especially for the ones with small-scale data. FedBCD significantly outperforms nine federated learning methods on both video-based and image-based diagnoses, demonstrating the superiority and potential for clinical practice. Code is released at https://github.com/tianpeng-deng/FedBCD.
Tianpeng Deng, Chunwang Huang, Jiatai Lin, Zhenwei Shi 0002, Bingchao Zhao, Jingqi Huang, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu, Chu Han
IEEE Trans. Medical Imaging14
2025 A Colorectal Coordinate-Driven Method for Colorectum and Colorectal Cancer Segmentation in Conventional CT Scans
abstract
Automated colorectal cancer (CRC) segmentation in medical imaging is the key to achieving automation of CRC detection, staging, and treatment response monitoring. Compared with magnetic resonance imaging (MRI) and computed tomography colonography (CTC), conventional computed tomography (CT) has enormous potential because of its broad implementation, superiority for the hollow viscera (colon), and convenience without needing bowel preparation. However, the segmentation of CRC in conventional CT is more challenging due to the difficulties presenting with the unprepared bowel, such as distinguishing the colorectum from other structures with similar appearance and distinguishing the CRC from the contents of the colorectum. To tackle these challenges, we introduce DeepCRC-SL, the first automated segmentation algorithm for CRC and colorectum in conventional contrast-enhanced CT scans. We propose a topology-aware deep learning-based approach, which builds a novel 1-D colorectal coordinate system and encodes each voxel of the colorectum with a relative position along the coordinate system. We then induce an auxiliary regression task to predict the colorectal coordinate value of each voxel, aiming to integrate global topology into the segmentation network and thus improve the colorectum's continuity. Self-attention layers are utilized to capture global contexts for the coordinate regression task and enhance the ability to differentiate CRC and colorectum tissues. Moreover, a coordinate-driven self-learning (SL) strategy is introduced to leverage a large amount of unlabeled data to improve segmentation performance. We validate the proposed approach on a dataset including 227 labeled and 585 unlabeled CRC cases by fivefold cross-validation. Experimental results demonstrate that our method outperforms some recent related segmentation methods and achieves the segmentation accuracy in DSC for CRC of 0.669 and colorectum of 0.892, reaching to the performance (at 0.639 and 0.890, respectively) of a medical resident with two years of specialized CRC imaging fellowship.
Yingda Xia, Suyun Li, Jiawen Yao, Dakai Jin, Yanting Liang, Jiatai Lin, Bingchao Zhao, Chu Han, Le Lu 0001, Ling Zhang 0002, Zaiyi Liu, Xin Chen 0058
IEEE Trans. Neural Networks Learn. Syst.10
2025 Tissue-SDG: dynamic adaptive data augmentation and multi-scale contrastive learning for generalizable tissue semantic segmentation
Jiayi Peng, Jiatai Lin, Chu Han, Zaiyi Liu
Vis. Comput.4
2024 SS-WSSS: Small-Scale Weakly Supervised Semantic Segmentation for Histopathology Image
abstract
Semantic segmentation for histopathology images is one of the fundamental tasks in computational pathology. Due to the high cost of pixel-level annotation acquisition, the weakly supervised semantic segmentation (WSSS) attempts to achieve information-intensive segmentation task for histopathology images to reduce the labeling effort of pathologists by leveraging image-level labels. However, traditional WSSS requires a large-scale training set with image-level labels, which still imposes considerable labeling costs on pathologists. To this end, this work proposes a Small-Scale Weakly Supervised Semantic Segmentation (SS-WSSS) approach to achieve the comparable performance only with small-scale weakly-labeled data to further reduce pathologist’s labeling effort. Since histopathology images can easily generate massive unlabeled data, our SS-WSSS aims to learn with the unlabeled data to bridge the information gap. First, we propose a Single-to-Multi Prototype Similarity (S2M-PS) method to generate reliable pseudo-labels for unlabeled data by measuring the similarity between single-label prototypes and multi-label feature maps. Then, we introduce a Cross-Task CoTraining (CT2) method for pseudo-supervision of models with pseudo-labels self-refinement to avoid overfitting to noisy labels. We conduct the experiment on two public datasets to demonstrate the effectiveness of our SS-WSSS. In the experiment, our method achieves comparable performance with SOTA methods only using 30% labeled data.
Jiatai Lin, Guoqiang Han 0002, Jingxuan Zhou, Zhenwei Shi 0002, Zaiyi Liu, Chu Han
BIBM6
2024 DBrAL: A Novel Uncertainty-Based Active Learning Based on Deep-Broad Learning for Medical Image Classification
Hongjiang Wu, Yuping Zhong, Guoqiang Han 0002, Jiatai Lin, Zaiyi Liu, Chu Han
ICANN (8)6
2024 Active Learning by Feature Perturbation for Medical Image Classification
Yuping Zhong, Guoqiang Han 0002, Zhenwei Shi 0002, Zaiyi Liu, Chu Han, Jiatai Lin
ICONIP (4)5
2024 FedDBL: Communication and Data Efficient Federated Deep-Broad Learning for Histopathological Tissue Classification
abstract
Histopathological tissue classification is a fundamental task in computational pathology. Deep learning (DL)-based models have achieved superior performance but centralized training suffers from the privacy leakage problem. Federated learning (FL) can safeguard privacy by keeping training samples locally, while existing FL-based frameworks require a large number of well-annotated training samples and numerous rounds of communication which hinder their viability in real-world clinical scenarios. In this article, we propose a lightweight and universal FL framework, named federated deep-broad learning (FedDBL), to achieve superior classification performance with limited training samples and only one-round communication. By simply integrating a pretrained DL feature extractor, a fast and lightweight broad learning inference system with a classical federated aggregation approach, FedDBL can dramatically reduce data dependency and improve communication efficiency. Five-fold cross-validation demonstrates that FedDBL greatly outperforms the competitors with only one-round communication and limited training samples, while it even achieves comparable performance with the ones under multiple-round communications. Furthermore, due to the lightweight design and one-round communication, FedDBL reduces the communication burden from 4.6 GB to only 138.4 KB per client using the ResNet-50 backbone at 50-round training. Extensive experiments also show the scalability of FedDBL on model generalization to the unseen dataset, various client numbers, model personalization and other image modalities. Since no data or deep model sharing across different clients, the privacy issue is well-solved and the model security is guaranteed with no model inversion attack risk. Code is available at https://github.com/tianpeng-deng/FedDBL.
Tianpeng Deng, Guoqiang Han 0002, Zhenwei Shi 0002, Jiatai Lin, Qi Dou 0001, Zaiyi Liu, Xiao-jing Guo, C. L. Philip Chen, Chu Han
IEEE Trans. Cybern.10
2024 CroMAM: A Cross-Magnification Attention Feature Fusion Model for Predicting Genetic Status and Survival of Gliomas Using Histological Images
abstract
Predicting the gene mutation status in whole slide images (WSIs) is crucial for the clinical treatment, cancer management, and research of gliomas. With advancements in CNN and Transformer algorithms, several promising models have been proposed. However, existing studies have paid little attention on fusing multi-magnification information, and the model requires processing all patches from a whole slide image. In this paper, we propose a cross-magnification attention model called CroMAM for predicting the genetic status and survival of gliomas. The CroMAM first utilizes a systematic patch extraction module to sample a subset of representative patches for downstream analysis. Next, the CroMAM applies Swin Transformer to extract local and global features from patches at different magnifications, followed by acquiring high-level features and dependencies among single-magnification patches through the application of a Vision Transformer. Subsequently, the CroMAM exchanges the integrated feature representations of different magnifications and encourage the integrated feature representations to learn the discriminative information from other magnification. Additionally, we design a cross-magnification attention analysis method to examine the effect of cross-magnification attention quantitatively and qualitatively which increases the model's explainability. To validate the performance of the model, we compare the proposed model with other multi-magnification feature fusion models on three tasks in two datasets. Extensive experiments demonstrate that the proposed model achieves state-of-the-art performance in predicting the genetic status and survival of gliomas.
Jisen Guo, Peng Xu 0004, Yuankui Wu, Yunyun Tao, Chu Han, Jiatai Lin, Zaiyi Liu, Cheng Lu 0001
IEEE J. Biomed. Health Informatics5
2024 Protecting Prostate Cancer Classification From Rectal Artifacts via Targeted Adversarial Training
abstract
Magnetic resonance imaging (MRI)-based deep neural networks (DNN) have been widely developed to perform prostate cancer (PCa) classification. However, in real-world clinical situations, prostate MRIs can be easily impacted by rectal artifacts, which have been found to lead to incorrect PCa classification. Existing DNN-based methods typically do not consider the interference of rectal artifacts on PCa classification, and do not design specific strategy to address this problem. In this study, we proposed a novel Targeted adversarial training with Proprietary Adversarial Samples (TPAS) strategy to defend the PCa classification model against the influence of rectal artifacts. Specifically, based on clinical prior knowledge, we generated proprietary adversarial samples with rectal artifact-pattern adversarial noise, which can severely mislead PCa classification models optimized by the ordinary training strategy. We then jointly exploited the generated proprietary adversarial samples and original samples to train the models. To demonstrate the effectiveness of our strategy, we conducted analytical experiments on multiple PCa classification models. Compared with ordinary training strategy, TPAS can effectively improve the single- and multi-parametric PCa classification at patient, slice and lesion level, and bring substantial gains to recent advanced models. In conclusion, TPAS strategy can be identified as a valuable way to mitigate the influence of rectal artifacts on deep learning models for PCa classification.
Lei Hu 0002, Dawei Zhou 0004, Cheng Lu 0001, Chu Han, Zhenwei Shi 0002, Qikui Zhu, Xinbo Gao 0001, Nannan Wang 0001, Zaiyi Liu
IEEE J. Biomed. Health Informatics5
2023 DBL-MPE: Deep Broad Learning for Prediction of Response to Neo-adjuvant Chemotherapy Using MRI-Based Multi-angle Maximal Enhancement Projection in Breast Cancer
Zihan Cao, Zhenwei Shi 0002, Xiaomei Huang, Chu Han, Peng Xu 0004, Zaiyi Liu
ICIC (3)4
2023 Fed-CSA: Channel Spatial Attention and Adaptive Weights Aggregation-Based Federated Learning for Breast Tumor Segmentation on MRI
Zhenwei Shi 0002, Xiaomei Huang, Chu Han, Zihan Cao, Peng Xu 0004, Zaiyi Liu
ICIC (3)4
2023 Treatment Outcome Prediction for Intracerebral Hemorrhage via Generative Prognostic Model with Imaging and Tabular Data
Wenao Ma, Cheng Chen 0013, Jill M. Abrigo, Calvin Hoi-Kwan Mak, Yuqi Gong, Nga Yan Chan, Chu Han, Zaiyi Liu, Qi Dou 0001
MICCAI (5)7
2023 Joint-phase attention network for breast cancer segmentation in DCE-MRI
Rian Huang, Zeyan Xu, Zixian Li, Yanfen Cui, Yingwen Huo, Chu Han, Xiaotang Yang, Zaiyi Liu, Yi Wang 0031
Expert Syst. Appl.8
2023 SMILE: Cost-sensitive multi-task learning for nuclear segmentation and classification with imbalanced annotations
Xipeng Pan, Jijun Cheng, Feihu Hou, Rushi Lan, Cheng Lu 0001, Lingqiao Li, Zhengyun Feng, Huadeng Wang, Changhong Liang, Zhenbing Liu, Xin Chen 0058, Chu Han, Zaiyi Liu
Medical Image Anal.12
2023 CKD-TransBTS: Clinical Knowledge-Driven Hybrid Transformer With Modality-Correlated Cross-Attention for Brain Tumor Segmentation
abstract
Brain tumor segmentation (BTS) in magnetic resonance image (MRI) is crucial for brain tumor diagnosis, cancer management and research purposes. With the great success of the ten-year BraTS challenges as well as the advances of CNN and Transformer algorithms, a lot of outstanding BTS models have been proposed to tackle the difficulties of BTS in different technical aspects. However, existing studies hardly consider how to fuse the multi-modality images in a reasonable manner. In this paper, we leverage the clinical knowledge of how radiologists diagnose brain tumors from multiple MRI modalities and propose a clinical knowledge-driven brain tumor segmentation model, called CKD-TransBTS. Instead of directly concatenating all the modalities, we re-organize the input modalities by separating them into two groups according to the imaging principle of MRI. A dual-branch hybrid encoder with the proposed modality-correlated cross-attention block (MCCA) is designed to extract the multi-modality image features. The proposed model inherits the strengths from both Transformer and CNN with the local feature representation ability for precise lesion boundaries and long-range feature extraction for 3D volumetric images. To bridge the gap between Transformer and CNN features, we propose a Trans&CNN Feature Calibration block (TCFC) in the decoder. We compare the proposed model with six CNN-based models and six transformer-based models on the BraTS 2021 challenge dataset. Extensive experiments demonstrate that the proposed model achieves state-of-the-art brain tumor segmentation performance compared with all the competitors.
Jianwei Lin, Jiatai Lin, Cheng Lu 0001, Hao Chen 0011, Bingchao Zhao, Zhenwei Shi 0002, Bingjiang Qiu, Xipeng Pan, Zeyan Xu, Biao Huang 0008, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu, Chu Han
IEEE Trans. Medical Imaging15
2023 HoVer-Trans: Anatomy-Aware HoVer-Transformer for ROI-Free Breast Cancer Diagnosis in Ultrasound Images
abstract
Ultrasonography is an important routine examination for breast cancer diagnosis, due to its non-invasive, radiation-free and low-cost properties. However, the diagnostic accuracy of breast cancer is still limited due to its inherent limitations. Then, a precise diagnose using breast ultrasound (BUS) image would be significant useful. Many learning-based computer-aided diagnostic methods have been proposed to achieve breast cancer diagnosis/lesion classification. However, most of them require a pre-define region of interest (ROI) and then classify the lesion inside the ROI. Conventional classification backbones, such as VGG16 and ResNet50, can achieve promising classification results with no ROI requirement. But these models lack interpretability, thus restricting their use in clinical practice. In this study, we propose a novel ROI-free model for breast cancer diagnosis in ultrasound images with interpretable feature representations. We leverage the anatomical prior knowledge that malignant and benign tumors have different spatial relationships between different tissue layers, and propose a HoVer-Transformer to formulate this prior knowledge. The proposed HoVer-Trans block extracts the inter- and intra-layer spatial information horizontally and vertically. We conduct and release an open dataset GDPH&SYSUCC for breast cancer diagnosis in BUS. The proposed model is evaluated in three datasets by comparing with four CNN-based models and three vision transformer models via five-fold cross validation. It achieves state-of-the-art classification performance (GDPH&SYSUCC AUC: 0.924, ACC: 0.893, Spec: 0.836, Sens: 0.926) with the best model interpretability. In the meanwhile, our proposed model outperforms two senior sonographers on the breast cancer diagnosis when only one BUS image is given (GDPH&SYSUCC-AUC ours: 0.924 vs. reader1: 0.825 vs. reader2: 0.820).
Yuhao Mo, Chu Han, Zhenwei Shi 0002, Jiatai Lin, Bingchao Zhao, Chunwang Huang, Bingjiang Qiu, Yanfen Cui, Xipeng Pan, Zeyan Xu, Xiaomei Huang, Zhenhui Li, Zaiyi Liu, Changhong Liang
IEEE Trans. Medical Imaging2
2022 Learning Pre- and Post-contrast Representation for Breast Cancer Segmentation in DCE-MRI
abstract
Breast dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) plays a considerable role in high-risk breast cancer diagnosis and image-based prognostic prediction. The accurate and robust segmentation of cancerous regions is with clinical demands. However, automatic segmentation remains challenging, due to the large variations of cancers in shape and size, and the class-imbalance issue. To tackle these problems, we offer a two-stage framework, which leverages both pre- and post-contrast images for the segmentation of breast cancer. Specifically, we first employ a breast segmentation network, which generates the breast region of interest (ROI) thus removing confounding information from thorax region in DCE-MRI. Furthermore, based on the generated breast ROI, we offer an attention network to learn both pre- and post-contrast representations for distinguishing cancerous regions from the normal breast tissue. The efficacy of our framework is evaluated on a collected dataset of 261 patients with biopsy-proven breast cancers. Experimental results demonstrate our method attains a Dice coefficient of 91.11% for breast cancer segmentation. The proposed framework provides an effective cancer segmentation solution for breast examination using DCE-MRI. The code is publicly available at https://github.com/2313595986/BreastCancerMRI.
Yingwen Huo, Yupeng Pan, Zeyan Xu, Rian Huang, Chu Han, Zaiyi Liu, Yi Wang 0031
CBMS7
2022 Multi-layer pseudo-supervision for histopathology tissue semantic segmentation using patch-level classification labels
abstract
Tissue-level semantic segmentation is a vital step in computational pathology. Fully-supervised models have already achieved outstanding performance with dense pixel-level annotations. However, drawing such labels on the giga-pixel whole slide images is extremely expensive and time-consuming. In this paper, we use only patch-level classification labels to achieve tissue semantic segmentation on histopathology images, finally reducing the annotation efforts. We propose a two-step model including a classification and a segmentation phases. In the classification phase, we propose a CAM-based model to generate pseudo masks by patch-level labels. In the segmentation phase, we achieve tissue semantic segmentation by our propose Multi-Layer Pseudo-Supervision. Several technical novelties have been proposed to reduce the information gap between pixel-level and patch-level annotations. As a part of this paper, we introduce a new weakly-supervised semantic segmentation (WSSS) dataset for lung adenocarcinoma (LUAD-HistoSeg). We conduct several experiments to evaluate our proposed model on two datasets. Our proposed model outperforms five state-of-the-art WSSS approaches. Note that we can achieve comparable quantitative and qualitative results with the fully-supervised model, with only around a 2% gap for MIoU and FwIoU. By comparing with manual labeling on a randomly sampled 100 patches dataset, patch-level labeling can greatly reduce the annotation time from hours to minutes. The source code and the released datasets are available at: https://github.com/ChuHan89/WSSS-Tissue.
Chu Han, Jiatai Lin, Jinhai Mai, Yi Wang 0031, Qingling Zhang 0006, Bingchao Zhao, Xin Chen 0058, Xipeng Pan, Zhenwei Shi 0002, Zeyan Xu, Su Yao, Lixu Yan, Xiaomei Huang, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu
Medical Image Anal.1
2022 Meta multi-task nuclei segmentation with fewer training samples
Chu Han, Huasheng Yao, Bingchao Zhao, Zhenhui Li, Zhenwei Shi 0002, Xin Chen 0058, Jinrong Qu, Rushi Lan, Changhong Liang, Xipeng Pan, Zaiyi Liu
Medical Image Anal.1
2022 SeqSeg: A sequential method to achieve nasopharyngeal carcinoma segmentation free from background dominance
Guihua Tao, Haojiang Li, Jiabin Huang 0007, Chu Han, Jiazhou Chen 0001, Guangying Ruan, Yu Hu 0004, Tingting Dan, Bin Zhang 0050, Shengfeng He, Hongmin Cai
Medical Image Anal.4
2022 PDBL: Improving Histopathological Tissue Classification With Plug-and-Play Pyramidal Deep-Broad Learning
abstract
Histopathological tissue classification is a simpler way to achieve semantic segmentation for the whole slide images, which can alleviate the requirement of pixel-level dense annotations. Existing works mostly leverage the popular CNN classification backbones in computer vision to achieve histopathological tissue classification. In this paper, we propose a super lightweight plug-and-play module, named Pyramidal Deep-Broad Learning (PDBL), for any well-trained classification backbone to improve the classification performance without a re-training burden. For each patch, we construct a multi-resolution image pyramid to obtain the pyramidal contextual information. For each level in the pyramid, we extract the multi-scale deep-broad features by our proposed Deep-Broad block (DB-block). We equip PDBL in three popular classification backbones, ShuffLeNetV2, EfficientNetb0, and ResNet50 to evaluate the effectiveness and efficiency of our proposed module on two datasets (Kather Multiclass Dataset and the LC25000 Dataset). Experimental results demonstrate the proposed PDBL can steadily improve the tissue-level classification performance for any CNN backbones, especially for the lightweight models when given a small among of training samples (less than 10%). It greatly saves the computational resources and annotation efforts. The source code is available at: https://github.com/linjiatai/PDBL.
Jiatai Lin, Guoqiang Han 0002, Xipeng Pan, Zaiyi Liu, Hao Chen 0011, Danyi Li, Xiping Jia, Zhenwei Shi 0002, Zhizhen Wang, Yanfen Cui, Haiming Li, Changhong Liang, Chu Han
IEEE Trans. Medical Imaging15
2021 Spatially-Invariant Style-Codes Controlled Makeup Transfer
abstract
Transferring makeup from the misaligned reference image is challenging. Previous methods overcome this barrier by computing pixel-wise correspondences between two images, which is inaccurate and computational-expensive. In this paper, we take a different perspective to break down the makeup transfer problem into a two-step extraction-assignment process. To this end, we propose a Style-based Controllable GAN model that consists of three components, each of which corresponds to target style-code encoding, face identity features extraction, and makeup fusion, respectively. In particular, a Part-specific Style Encoder encodes the component-wise makeup style of the reference image into a style-code in an intermediate latent space W. The style-code discards spatial information and therefore is invariant to spatial misalignment. On the other hand, the style-code embeds component-wise information, enabling flexible partial makeup editing from multiple references. This style-code, together with source identity features, is integrated into a Makeup Fusion Decoder equipped with multiple AdaIN layers to generate the final result. Our proposed method demonstrates great flexibility on makeup transfer by supporting makeup removal, shade-controllable makeup transfer, and part-specific makeup transfer, even with large spatial misalignment. Extensive experiments demonstrate the superiority of our approach over state-of-the-art methods. Code is available at https://github.com/makeuptransfer/SCGAN.
Chu Han, Hongmin Cai, Guoqiang Han 0002, Shengfeng He
CVPR2
2021 Reciprocal Learning for Semi-supervised Segmentation
Xiangyun Zeng, Rian Huang, Yuming Zhong, Chu Han, Di Lin 0002, Dong Ni 0001, Yi Wang 0031
MICCAI (2)5
2021 Fusion of multi-source retinal fundus images via automatic registration for clinical diagnosis
Tingting Dan, Yu Hu 0004, Chu Han, Zhihao Fan, Zhuobin Huang, Bin Zhang 0050, Guihua Tao, Baoyi Liu, Honghua Yu, Hongmin Cai
Neurocomputing3
2021 Fast scene labeling via structural inference
Huaidong Zhang, Chu Han, Xiaodan Zhang 0003, Yong Du 0003, Xuemiao Xu, Guoqiang Han 0002, Harry Qin, Shengfeng He
Neurocomputing2
2021 Learning a consensus affinity matrix for multi-view clustering via subspaces merging on Grassmann manifold
Wentao Rong, Enhong Zhuo, Jiazhou Chen 0001, Haiyan Wang 0005, Chu Han, Hongmin Cai
Inf. Sci.6
2021 Learning task-driving affinity matrix for accurate multi-view clustering through tensor subspace learning
Haiyan Wang 0005, Guoqiang Han 0002, Junyu Li 0001, Bin Zhang 0050, Jiazhou Chen 0001, Yu Hu 0004, Chu Han, Hongmin Cai
Inf. Sci.7
2021 Video Snapshot: Single Image Motion Expansion via Invertible Motion Embedding
abstract
Unlike images, finding the desired video content in a large pool of videos is not easy due to the time cost of loading and watching. Most video streaming and sharing services provide the video preview function for a better browsing experience. In this paper, we aim to generate a video preview from a single image. To this end, we propose two cascaded networks, the motion embedding network and the motion expansion network. The motion embedding network aims to embed the spatio-temporal information into an embedded image, called video snapshot. On the other end, the motion expansion network is proposed to invert the video back from the input video snapshot. To hold the invertibility of motion embedding and expansion during training, we design four tailor-made losses and a motion attention module to make the network focus on the temporal information. In order to enhance the viewing experience, our expansion network involves an interpolation module to produce a longer video preview with a smooth transition. Extensive experiments demonstrate that our method can successfully embed the spatio-temporal information of a video into one "live" image, which can be converted back to a video preview. Quantitative and qualitative evaluations are conducted on a large number of videos to prove the effectiveness of our proposed method. In particular, statistics of PSNR and SSIM on a large number of videos show the proposed method is general, and it can generate a high-quality video from a single image.
Qianshu Zhu, Chu Han, Guoqiang Han 0002, Tien-Tsin Wong, Shengfeng He
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Transductive Zero-Shot Action Recognition via Visually Connected Graph Convolutional Networks
abstract
With the explosive growth of action categories, zero-shot action recognition aims to extend a well-trained model to novel/unseen classes. To bridge the large knowledge gap between seen and unseen classes, in this brief, we visually associate unseen actions with seen categories in a visually connected graph, and the knowledge is then transferred from the visual features space to semantic space via the grouped attention graph convolutional networks (GAGCNs). In particular, we extract visual features for all the actions, and a visually connected graph is built to attach seen actions to visually similar unseen categories. Moreover, the proposed grouped attention mechanism exploits the hierarchical knowledge in the graph so that the GAGCN enables propagating the visual-semantic connections from seen actions to unseen ones. We extensively evaluate the proposed method on three data sets: HMDB51, UCF101, and NTU RGB + D. Experimental results show that the GAGCN outperforms state-of-the-art methods.
Yangyang Xu 0003, Chu Han, Harry Qin, Xuemiao Xu, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Neural Networks Learn. Syst.2
2020 Tensor-based Low-rank and Graph Regularized Representation Learning for Multi-view Clustering
abstract
Multi-view clustering aims to partition the data into their underlying clusters via leveraging multiple views information. To exploit cross-view information, existed approaches in tensor-based subspace learning attract much attention. In order to explore essential tensor, the most recent work mainly focuses on capturing representation tensor with sparse and low-rank constraints. However, one shortcoming is that this process may suffer from instability since it did not consider retaining local structure between samples. To tackle the issue, we introduce a novel self-expressive tensor learning method considering both global and local constraints to promote the learning of representation tensor. In particular, we construct a tensor-based subspace representation that joint low-rank and graph-regularized tensor learning to a united optimization problem. The essential global structure and high-order correlations can be naturally captured through low-rank self-expressive tensor learning. Meanwhile, the local structures can be preserved by introducing graph regularized terms on representation tensor, thus bring benefits to subsequent clustering task. An effective optimization procedure for solving the proposed model is presented. We conduct extensive experiments on text, object, and gene expression datasets. The experimental results well demonstrate that the proposed method, named by TLGRL, achieves superiority over benchmark methods.
Haiyan Wang 0005, Guoqiang Han 0002, Bin Zhang 0050, Yu Hu 0004, Chu Han, Hongmin Cai
BIBM6
2020 TENet: Triple Excitation Network for Video Salient Object Detection
Sucheng Ren, Chu Han, Xin Yang 0011, Guoqiang Han 0002, Shengfeng He
ECCV (5)2
2020 Coherence and Identity Learning for Arbitrary-length Face Video Generation
abstract
Face synthesis is an interesting yet challenging task in computer vision. It is even much harder to generate a portrait video than a single image. In this paper, we propose a novel video generation framework for synthesizing arbitrary-length face videos without any face exemplar or landmark. To overcome the synthesis ambiguity of face video, we propose a divide-and-conquer strategy to separately address the video face synthesis problem from two aspects, face identity synthesis and rearrangement. To this end, we design a cascaded network which contains three components, Identity-aware GAN (IA-GAN), Face Coherence Network, and Interpolation Network. IA-GAN is proposed to synthesize photorealistic faces with the same identity from a set of noises. Face Coherence Network is designed to re-arrange the faces generated by IA-GAN while keeping the inter-frame coherence. Interpolation Network is introduced to eliminate the discontinuity between two adjacent frames and improve the smoothness of the face video. Experimental results demonstrate that our proposed network is able to generate face video with high visual quality while preserving the identity. Statistics show that our method outperforms state-of-the-art unconditional face video generative models in multiple challenging datasets.
Shuquan Ye, Chu Han, Jiaying Lin 0001, Guoqiang Han 0002, Shengfeng He
ICPR2
2020 Example-Based Colourization Via Dense Encoding Pyramids
abstract
Abstract We propose a novel deep example‐based image colourization method called dense encoding pyramid network. In our study, we define the colourization as a multinomial classification problem. Given a greyscale image and a reference image, the proposed network leverages large‐scale data and then predicts colours by analysing the colour distribution of the reference image. We design the network as a pyramid structure in order to exploit the inherent multi‐scale, pyramidal hierarchy of colour representations. Between two adjacent levels, we propose a hierarchical decoder–encoder filter to pass the colour distributions from the lower level to higher level in order to take both semantic information and fine details into account during the colourization process. Within the network, a novel parallel residual dense block is proposed to effectively extract the local–global context of the colour representations by widening the network. Several experiments, as well as a user study, are conducted to evaluate the performance of our network against state‐of‐the‐art colourization methods. Experimental results show that our network is able to generate colourful, semantically correct and visually pleasant colour images. In addition, unlike fully automatic colourization that produces fixed colour images, the reference image of our network is flexible; both natural images and simple colour palettes can be used to guide the colourization.
Chu-Feng Xiao 0001, Chu Han, Zhuming Zhang, Harry Qin, Tien-Tsin Wong, Guoqiang Han 0002, Shengfeng He
Comput. Graph. Forum2
2020 Triple U-net: Hematoxylin-aware nuclei segmentation with progressive dense feature aggregation
Bingchao Zhao, Xin Chen 0058, Zhiwen Yu 0002, Su Yao, Lixu Yan, Zaiyi Liu, Changhong Liang, Chu Han
Medical Image Anal.10
2020 Exploring Duality in Visual Question-Driven Top-Down Saliency
abstract
Top-down, goal-driven visual saliency exerts a huge influence on the human visual system for performing visual tasks. Text generations, like visual question answering (VQA) and visual question generation (VQG), have intrinsic connections with top-down saliency, which is usually involved in both VQA and VQG processes in an unsupervised manner. However, it is shown that the regions that humans choose to look at to answer questions are very different from the unsupervised attention models. In this brief, we aim to explore the intrinsic relationship between top-down saliency and text generations, and to figure out whether an accurate saliency response benefits text generation. To this end, we propose a dual supervised network with dynamic parameter prediction. Dual-supervision explicitly exploits the probabilistic correlation between the primal task top-down saliency detection and the dual task text generation, while dynamic parameter prediction encodes the given text (i.e., question or answer) into the fully convolutional network. Extensive experiments show the proposed top-down saliency method achieves the best correlation with human attention among various baselines. In addition, the proposed model can be guided by either questions or answers, and output the counterpart. Furthermore, we show that combining human-like visual question-saliency improves the performance of both answer and question generations.
Shengfeng He, Chu Han, Guoqiang Han 0002, Harry Qin
IEEE Trans. Neural Networks Learn. Syst.2
2019 Deep Line Drawing Vectorization via Line Subdivision and Topology Reconstruction
abstract
Abstract Vectorizing line drawing is necessary for the digital workflows of 2D animation and engineering design. But it is challenging due to the ambiguity of topology, especially at junctions. Existing vectorization methods either suffer from low accuracy or cannot deal with high‐resolution images. To deal with a variety of challenging containing different kinds of complex junctions, we propose a two‐phase line drawing vectorization method that analyzes the global and local topology. In the first phase, we subdivide the lines into partial curves, and in the second phase, we reconstruct the topology at junctions. With the overall topology estimated in the two phases, we can trace and vectorize the curves. To qualitatively and quantitatively evaluate our method and compare it with the existing methods, we conduct extensive experiments on not only existing datasets but also our newly synthesized dataset which contains different types of complex and ambiguous junctions. Experimental statistics show that our method greatly outperforms existing methods in terms of computational speed and achieves visually better topology reconstruction accuracy.
Zhuming Zhang, Chu Han, Chengze Li, Tien-Tsin Wong
Comput. Graph. Forum3
2019 Age estimation via attribute-region association
Yiliang Chen, Shengfeng He, Zichang Tan, Chu Han, Guoqiang Han 0002, Harry Qin
Neurocomputing4
2019 A Learning-Based Multimodel Integrated Framework for Dynamic Traffic Flow Forecasting
Teng Zhou, Guoqiang Han 0002, Xuemiao Xu, Chu Han, Yuchang Huang, Harry Qin
Neural Process. Lett.4
2019 Deep binocular tone mapping
Zhuming Zhang, Chu Han, Shengfeng He, Xueting Liu 0001, Xinghong Hu, Tien-Tsin Wong
Vis. Comput.2
2018 TransHist: Occlusion-robust shape detection in cluttered images
abstract
Shape matching plays an important role in various computer vision and graphics applications such as shape retrieval, object detection, image editing, image retrieval, etc. However, detecting shapes in cluttered images is still quite challenging due to the incomplete edges and changing perspective. In this paper, we propose a novel approach that can efficiently identify a queried shape in a cluttered image. The core idea is to acquire the transformation from the queried shape to the cluttered image by summarising all point-to-point transformations between the queried shape and the image. To do so, we adopt a point-based shape descriptor, the pyramid of arc-length descriptor (PAD), to identify point pairs between the queried shape and the image having similar local shapes. We further calculate the transformations between the identified point pairs based on PAD. Finally, we summarise all transformations in a 4D transformation histogram and search for the main cluster. Our method can handle both closed shapes and open curves, and is resistant to partial occlusions. Experiments show that our method can robustly detect shapes in images in the presence of partial occlusions, fragile edges, and cluttered backgrounds.
Chu Han, Xueting Liu 0001, Lok Tsun Sinn, Tien-Tsin Wong
Comput. Vis. Media1
2018 Deep unsupervised pixelization
abstract
In this paper, we present a novel unsupervised learning method for pixelization. Due to the difficulty in creating pixel art, preparing the paired training data for supervised learning is impractical. Instead, we propose an unsupervised learning framework to circumvent such difficulty. We leverage the dual nature of the pixelization and depixelization, and model these two tasks in the same network in a bi-directional manner with the input itself as training supervision. These two tasks are modeled as a cascaded network which consists of three stages for different purposes. GridNet transfers the input image into multi-scale grid-structured images with different aliasing effects. PixelNet associated with GridNet to synthesize pixel arts with sharp edges and perceptually optimal local structures. DepixelNet connects the previous network and aims to recover the pixelized result to the original image. For the sake of unsupervised learning, the mirror loss is proposed to hold the reversibility of feature representations in the process. In addition, adversarial, L1, and gradient losses are involved in the network to obtain pixel arts by retaining color correctness and smoothness. We show that our technique can synthesize crisper and perceptually more appropriate pixel arts than state-of-the-art image downscaling methods. We evaluate the proposed method with extensive experiments on many images. The proposed method outperforms state-of-the-art methods in terms of visual quality and user preference.
Chu Han, Shengfeng He, Qianshu Zhu, Yinjie Tan, Guoqiang Han 0002, Tien-Tsin Wong
ACM Trans. Graph.1
2017 δ-agree AdaBoost stacked autoencoder for short-term traffic flow forecasting
Teng Zhou, Guoqiang Han 0002, Xuemiao Xu, Zhizhe Lin, Chu Han, Yuchang Huang, Harry Qin
Neurocomputing5
2016 Pyramid of arclength descriptor for generating collage of shapes
abstract
This paper tackles a challenging 2D collage generation problem, focusing on shapes: we aim to fill a given region by packing irregular and reasonably-sized shapes with minimized gaps and overlaps. To achieve this nontrivial problem, we first have to analyze the boundary of individual shapes and then couple the shapes with partially-matched boundary to reduce gaps and overlaps in the collages. Second, the search space in identifying a good coupling of shapes is highly enormous, since arranging a shape in a collage involves a position, an orientation, and a scale factor. Yet, this matching step needs to be performed for every single shape when we pack it into a collage. Existing shape descriptors are simply infeasible for computation in a reasonable amount of time. To overcome this, we present a brand new, scale- and rotation-invariant 2D shape descriptor, namely pyramid of arclength descriptor (PAD). Its formulation is locally supported, scalable, and yet simple to construct and compute. These properties make PAD efficient for performing the partial-shape matching. Hence, we can prune away most search space with simple calculation, and efficiently identify candidate shapes. We evaluate our method using a large variety of shapes with different types and contours. Convincing collage results in terms of visual quality and time performance are obtained.
Kin Chung Kwan, Lok Tsun Sinn, Chu Han, Tien-Tsin Wong, Chi-Wing Fu
ACM Trans. Graph.3