VLDB 2026 Research / reviewers in the wild / expert
Kai Hu 0002
dblp:57/6633-2
· DBLP profile ↗
55ranked-venue papers
14as first author
42since 2021 · last 2026
0000-0003-1436-9522ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image SegmentationabstractAccurate segmentation of ultra-high-resolution (UHR) images, which often exceed tens of millions of pixels, is critically important in domains such as remote sensing and biomedical imaging. However, acquiring pixel-level annotations for such high-resolution images is prohibitively expensive and labor-intensive. While semi-supervised semantic segmentation can significantly reduce the annotation burden, its extension to UHR images holds great potential for addressing the unique challenges posed by sparse supervision. To this end, we propose SSR-SAM, a retrieval-style semi-supervised segmentation framework tailored for UHR images. Leveraging the promptable paradigm of the Segment Anything Model (SAM), SSR-SAM treats locally annotated regions as prompts to retrieve semantically consistent pixels across the entire image. Building upon this retrieval-style segmentation paradigm, we further introduce prompt-level perturbation, a novel trail to deploy consistency regularization for semi-supervised segmentation. It encourages the model to learn consistency across predictions guided by diverse visual-semantic prompts, thereby enhancing generalization on unlabeled data. We evaluate SSR-SAM on three UHR datasets: Inria Aerial, BCSS, and URUR. Experimental results show that SSR-SAM achieves clear performance gains over the labeled-only supervision, with average mIoU improvements of 4.9%, 4.15%, and 2.5%, respectively. Additionally, SSR-SAM possesses zero-shot segmentation capability, exhibiting potential for general retrieval-style segmentation tasks. Zhineng Chen, Kai Hu 0002, Xieping Gao 0001 |
AAAI | 4 |
| 2026 | An effective retraining strategy for unsupervised domain adaptive medical image segmentation
Kai Hu 0002, Xiongjun Ye, Xieping Gao 0001 |
Expert Syst. Appl. | 2 |
| 2026 | CSFMIL: Whole slide image classification with two-stage cross-scale fusion
Zhineng Chen, Feng-Jung Chen, Kai Hu 0002, Xieping Gao 0001 |
Neurocomputing | 5 |
| 2026 | Dual-perspective decoupling network for kidney tumor segmentation on CT images
Xinya Gan, Yuan Zhang 0022, Xiongjun Ye, Kai Hu 0002, Xieping Gao 0001 |
Neural Networks | 6 |
| 2026 | Improving mutation pathogenicity prediction of metal-binding sites in proteins with a panoramic attention mechanism
Yuan Zhang 0022, Jiafeng Wu, Qiuye Zhao, Mingyuan Dong, Junsheng Deng, Xieping Gao 0001, Kai Hu 0002, Dapeng Xiong |
Pattern Recognit. | 7 |
| 2025 | Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationabstractCurrent knowledge distillation (KD) methods for semantic segmentation focus on guiding the student to imitate the teacher's knowledge within homogeneous architectures. However, these methods overlook the diverse knowledge contained in architectures with different inductive biases, which is crucial for enabling the student to acquire a more precise and comprehensive understanding of the data during distillation. To this end, we propose for the first time a generic knowledge distillation method for semantic segmentation from a heterogeneous perspective, named HeteroAKD. Due to the substantial disparities between heterogeneous architectures, such as CNN and Transformer, directly transferring cross-architecture knowledge presents significant challenges. To eliminate the influence of architecture-specific information, the intermediate features of both the teacher and student are skillfully projected into an aligned logits space. Furthermore, to utilize diverse knowledge from heterogeneous architectures and deliver customized knowledge required by the student, a teacher-student knowledge mixing mechanism (KMM) and a teacher-student knowledge evaluation mechanism (KEM) are introduced. These mechanisms are performed by assessing the reliability and its discrepancy between heterogeneous teacher-student knowledge. Extensive experiments conducted on three main-stream benchmarks using various teacher-student pairs demonstrate that our HeteroAKD framework outperforms state-of-the-art KD methods in facilitating distillation between heterogeneous architectures. Yanglin Huang, Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
AAAI | 2 |
| 2025 | Explicit Relational Reasoning Network for Scene Text DetectionabstractConnected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmental bottleneck that their time-consuming post-processing is difficult to eliminate. To address this issue, we introduce an explicit relational reasoning network (ERRNet) to elegantly model the component relationships without post-processing. Concretely, we first represent each text instance as multiple ordered text components, and then treat these components as objects in sequential movement. In this way, scene text detection can be innovatively viewed as a tracking problem. From this perspective, we design an end-to-end tracking decoder to achieve a CC-based method dispensing with post-processing entirely. Additionally, we observe that there is an inconsistency between classification confidence and localization quality, so we propose a Polygon Monte-Carlo method to quickly and accurately evaluate the localization quality. Based on this, we introduce a position-supervised classification loss to guide the task-aligned learning of ERRNet. Experiments on challenging benchmarks demonstrate the effectiveness of our ERRNet. It consistently achieves state-of-the-art accuracy while holding highly competitive inference speed. Zhineng Chen, Yongkun Du, Zhilong Ji, Kai Hu 0002, Jinfeng Bai, Xieping Gao 0001 |
AAAI | 5 |
| 2025 | CoDiST: Combining Unimodal Contrastive Denoising and Crossmodal Disentanglement for Spatial TranscriptomicsabstractIntegrating spatial transcriptomics (ST) with histology enables precise delineation of spatial domains in complex tissues. However, effective multimodal integration faces two primary challenges: (1) High sparsity and frequent dropout events in transcriptomic data and artifacts in histological images together obscure true biological signals, resulting in intra-modal noise. (2) Excessive emphasis on modality alignment in current fusion methods allows redundant information to dominate over modality -specific features, leading to inter-modal redundancy. To address these challenges, we propose CoDiST, a novel multi-modal representation learning framework. Specifically, CoDiST uses histological information to strengthen spatial neighborhood relationships and employs unimodal contrastive learning to enhance robustness against technical noise in both modalities. Furthermore, to overcome inter-modal redundancy, CoDiST introduces the Gene Image Disentanglement Network (GIDNet). It disentangles representations into a shared subspace and specific subspaces, achieving crossmodal semantic alignment while preserving complementary features unique to each modality. In benchmarking on multiple ST datasets, CoDiST delivers clear gains over existing methods. On the Human Breast Cancer dataset, CoDiST achieves a 0.66 Adjusted Rand Index (ARI), representing a 3.6% improvement over state-of-the-art methods. Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
BIBM | 2 |
| 2025 | Dualmnet: a Lightweight Multi-Axis Interactive and Mask-Guided Network for Intracerebral Hemorrhage SegmentationabstractIntracerebral hemorrhage (ICH) segmentation is a critical step in the treatment of hemorrhagic stroke. Although many segmentation models have been developed for this task, their large parameter counts and high computational costs hinder deployment on mobile devices. To address this challenge, we propose a lightweight multi-axis interactive and mask-guided network for ICH segmentation, named DualMNet. DualMNet integrates a Multi-Axis Interactive Attention (MIA) module, which utilizes three-axis interactions to generate an attention map, extracting the inherent characteristics of CT images and enhancing feature representation with minimal computational cost. Furthermore, to overcome the difficulties of missed and false detections caused by low contrast between hemorrhagic and normal tissues, as well as the excessively small area of some bleeding regions, we design a Mask-Guided Feature Fusion (MGFF) module. This module uses intermediate mask predictions to guide the adaptive fusion of low-level and high-level features, making the model more sensitive to hemorrhagic regions. Experimental results on two public datasets, PHY and BHSD, demonstrate that DualMNet outperforms existing lightweight models in segmentation accuracy, achieving higher Dice scores while significantly reducing model parameters. Jiayao Tang 0002, Yuan Zhang 0002, Xieping Gao 0001, Kai Hu 0002 |
BIBM | 4 |
| 2025 | DMANet: Integrating Dual Dynamic Token Mixer With Masked Separable Attention for Intracerebral Hemorrhage SegmentationabstractAccurate medical image segmentation is crucial for quantifying intracranial hemorrhage (ICH) disease and evaluating treatment. The U-shaped architecture, which is based on the fusion mechanism of convolution and attention, has made significant contributions to the field of medical image segmentation. However, there exists a semantic gap between the encoder and the decoder of the U-shaped architecture, and directly fusing the two pieces of information through skip connection cannot effectively assist the decoder in feature extraction. To address these issues, we propose a new ICH segmentation framework called DMANet. Our model is built on the fundamental principle of Dual Dynamic Token Mixer, which enables DMANet to effectively capture local and global features across diverse inputs. Additionally, we incorporate Masked Separable and Dynamic Gate (MSDG) Attention into the skip connection to extract foreground and background information from encoder features using different mask strategies, thereby reducing the semantic gap between the decoder and encoder. Quantitative analysis and visualization results on Physionet dataset demonstrate that our model outperforms previous methods and achieves accurate segmentation of ICH areas. Chenyan Wen, Kai Hu 0002 |
BIBM | 3 |
| 2025 | Rethinking Dual-Stream Super-Resolution for Enhancing Remote Sensing Object DetectionabstractDetecting small instances within complex backgrounds presents significant challenges for Remote Sensing Object Detection (RSOD). While complex deep networks can enhance feature representation, they often lead to considerable computational burdens. Previous research has proposed dual-stream learning to improve the detection capabilities of compact models. We have re-evaluated existing dual-stream learning frameworks and identified their limitations in focusing on small objects. To address this issue, we propose a Dual-Stream Object Detection (DSOD) framework, which incorporates a Feature Fusion Guidance Module (FFGM). Specifically, by integrating features from the two task streams under dual-task supervision, DSOD directs the super-resolution process towards instance-related regions. This approach enhances feature extraction with object instance-specific details, significantly improving RSOD accuracy without introducing additional computational overhead. The effectiveness of DSOD has been validated across three models on three publicly available datasets. Furthermore, ablation studies highlight the importance of DSOD in optimizing RSOD performance while maintaining computational efficiency. An Luo, Kai Hu 0002 |
ICASSP | 2 |
| 2025 | MambaST: Hexagonal State Space Modeling for Spatial Domain Identification
Kai Hu 0002, Xuefeng Cui, Fa Zhang 0001 |
ISBRA (1) | 2 |
| 2025 | Multiscale Feature Enhancement and Adaptive Receptive Field for Tiny Object Detection in Remote Sensing ImagesabstractIn the field of remote sensing, detecting tiny objects remains a significant challenge, essentially due to the insufficient and weak feature representations caused by the limited pixel coverage of these objects, which presents a pressing demand for networks possessing multiscale feature enhancement ability. Furthermore, the unique prior knowledge presented in remote sensing images is not utilized effectively or appropriately, making this information just an obstacle to detecting tiny objects. To address these issues, we propose an effective network for tiny object detection called Multiscale Feature Enhancement and Adaptive Receptive Field YOLO (MA-YOLO). Specifically, the MA-YOLO integrates two innovative plug-and-play modules: the Multiscale Feature Enhancement Module (MFEM) and the Receptive Field Adaptive Multiscale Feature Fusion Module (RFAMFFM). The MFEM is designed to extract multiscale and multi-geometric features through a multi-branch refinement design with multi-size and -shape convolution, thereby enhancing feature representation. Driven by the larger receptive fields of large-kernel convolution and the guiding property of global pooling, the RFAMFFM is employed to dynamically adjust and select the receptive field to leverage prior knowledge effectively. Our proposed network outperforms the existing comparison models on two publicly available datasets, i.e., USOD and VEDAI. Specifically, compared to the state-of-the-art model, MA-YOLO achieves significant improvements of 2.2% on the USOD dataset and of 1.8% on the VEDAI dataset for the mAP50:95 metric. Yunpeng Zeng, An Luo, Kefan Zhan, Yuan Zhang 0022, Kai Hu 0002 |
ICMR | 6 |
| 2025 | Confusion-Driven Self-Supervised Progressively Weighted Ensemble Learning for Non-Exemplar Class Incremental LearningabstractNon-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge while retaining previously acquired knowledge in scenarios where prior examples are unavailable. A prevalent strategy within NECIL mitigates knowledge forgetting by freezing the feature extractor after training on the initial task. However, this freezing mechanism does not provide explicit training to differentiate between new and old classes, resulting in overlapping feature representations. To address this challenge, we propose a **C**onfusion-driven se**L**f-supervised pr**O**gressi**V**ely weighted **E**nsemble lea**R**ning (*CLOVER*) framework for NECIL. Firstly, we introduce a confusion-driven self-supervised learning approach that enhances representation extraction by guiding the model to distinguish between highly confusable classes, thereby reducing class representation overlap. Secondly, we develop a progressively weighted ensemble learning method that gradually adjusts weights to integrate diverse knowledge more effectively, further minimizing representation overlap. Finally, extensive experiments demonstrate that our proposed method achieves state-of-the-art results on the CIFAR100, TinyImageNet, and ImageNet-Subset NECIL benchmarks. Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
NeurIPS | 1 |
| 2025 | Progressive Learning Strategy for Few-Shot Class-Incremental LearningabstractThe goal of few-shot class incremental learning (FSCIL) is to learn new concepts from a limited number of novel samples while preserving the knowledge of previously learned classes. The mainstream FSCIL framework begins with training in the base session, after which the feature extractor is frozen to accommodate novel classes. We observed that traditional base-session training approaches often lead to overfitting on challenging samples, which can lead to reduced robustness in the decision boundaries and exacerbate the forgetting phenomenon when introducing incremental data. To address this issue, we proposed the progressive learning strategy (PGLS). First, inspired by curriculum learning, we developed a covariance noise perturbation approach based on the statistical information as a difficulty measure for assessing sample robustness. We then reweighted the samples based on their robustness, initially concentrating on enhancing model stability by prioritizing robust samples and subsequently leveraging weakly robust samples to improve generalization. Second, we predefined forward compatibility for various virtual class augmentation models. Within base class training, we employed a curriculum learning strategy that progressively introduced fewer to more virtual classes in order to mitigate any adverse effects on model performance. This strategy enhances the adaptability of base classes to novel ones and alleviates forgetting problems. Finally, extensive experiments conducted on the CUB200, CIFAR100, and miniImageNet datasets demonstrate the significant advantages of our proposed method over state-of-the-art models. Kai Hu 0002, Yunjiang Wang, Yuan Zhang 0022, Xieping Gao 0001 |
IEEE Trans. Cybern. | 1 |
| 2025 | Multi-Perspective Pseudo-Label Generation and Confidence-Weighted Training for Semi-Supervised Semantic SegmentationabstractSelf-training has been shown to achieve remarkable gains in semi-supervised semantic segmentation by creating pseudo-labels using unlabeled data. This approach, however, suffers from the quality of the generated pseudo-labels, and generating higher quality pseudo-labels is the main challenge that needs to be addressed. In this paper, we propose a novel method for semi-supervised semantic segmentation based on Multi-perspective pseudo-label Generation and Confidence-weighted Training (MGCT). First, we present a multi-perspective pseudo-label generation strategy that considers both global and local semantic perspectives. This strategy prioritizes pixels in all images by the global and local predictions, and subsequently generates pseudo-labels for different pixels in stages according to the ranking results. Our pseudo-label generation method shows superior suitability for semi-supervised semantic segmentation compared to other approaches. Second, we propose a confidence-weighted training method to alleviate performance degradation caused by unstable pixels. Our training method assigns confident weights to unstable pixels, which reduces the interference of unstable pixels during training and facilitates the efficient training of the model. Finally, we validate our approach on the PASCAL VOC 2012 and Cityscapes datasets, and the results indicate that we achieve new state-of-the-art performance on both datasets in all settings. Kai Hu 0002, Zhineng Chen, Yuan Zhang 0022, Xieping Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | GDTNet: A Synergistic Dilated Transformer and CNN by Gate Attention for Abdominal Multi-organ Segmentation
Yuan Zhang 0022, Xuanya Li, Kai Hu 0002 |
MMM (4) | 5 |
| 2024 | One-to-Multiple: A Progressive Style Transfer Unsupervised Domain-Adaptive Framework for Kidney Tumor SegmentationabstractIn multi-sequence Magnetic Resonance Imaging (MRI), the accurate segmentation of the kidney and tumor based on traditional supervised methods typically necessitates detailed annotation for each sequence, which is both time-consuming and labor-intensive. Unsupervised Domain Adaptation (UDA) methods can effectively mitigate inter-domain differences by aligning cross-modal features, thereby reducing the annotation burden. However, most existing UDA methods are limited to one-to-one domain adaptation, which tends to be inefficient and resource-intensive when faced with multi-target domain transfer tasks. To address this challenge, we propose a novel and efficient One-to-Multiple Progressive Style Transfer Unsupervised Domain-Adaptive (PSTUDA) framework for kidney and tumor segmentation in multi-sequence MRI. Specifically, we develop a multi-level style dictionary to explicitly store the style information of each target domain at various stages, which alleviates the burden of a single generator in a multi-target transfer task and enables effective decoupling of content and style. Concurrently, we employ multiple cascading style fusion modules that utilize point-wise instance normalization to progressively recombine content and style features, which enhances cross-modal alignment and structural consistency. Experiments conducted on the private MSKT and public KiTS19 datasets demonstrate the superiority of the proposed PSTUDA over comparative methods in multi-sequence kidney and tumor segmentation. The average Dice Similarity Coefficients are increased by at least 1.8% and 3.9%, respectively. Impressively, our PSTUDA not only significantly reduces the floating-point computation by approximately 72% but also reduces the number of model parameters by about 50%, bringing higher efficiency and feasibility to practical clinical applications. Kai Hu 0002, Jinhao Li 0009, Yuan Zhang 0022, Xiongjun Ye, Xieping Gao 0001 |
NeurIPS | 1 |
| 2024 | Multi-view Masked Contrastive Representation Learning for Endoscopic Video AnalysisabstractEndoscopic video analysis can effectively assist clinicians in disease diagnosis and treatment, and has played an indispensable role in clinical medicine. Unlike regular videos, endoscopic video analysis presents unique challenges, including complex camera movements, uneven distribution of lesions, and concealment, and it typically relies on contrastive learning in self-supervised pretraining as its mainstream technique. However, representations obtained from contrastive learning enhance the discriminability of the model but often lack fine-grained information, which is suboptimal in the pixel-level prediction tasks. In this paper, we develop a Multi-view Masked Contrastive Representation Learning (M$^2$CRL) framework for endoscopic video pre-training. Specifically, we propose a multi-view mask strategy for addressing the challenges of endoscopic videos. We utilize the frame-aggregated attention guided tube mask to capture global-level spatiotemporal sensitive representation from the global views, while the random tube mask is employed to focus on local variations from the local views. Subsequently, we combine multi-view mask modeling with contrastive learning to obtain endoscopic video representations that possess fine-grained perception and holistic discriminative capabilities simultaneously. The proposed M$^2$CRL is pre-trained on 7 publicly available endoscopic video datasets and fine-tuned on 3 endoscopic video datasets for 3 downstream tasks. Notably, our M$^2$CRL significantly outperforms the current state-of-the-art self-supervised endoscopic pre-training methods, e.g., Endo-FM (3.5% F1 for classification, 7.5% Dice for segmentation, and 2.2% F1 for detection) and other self-supervised methods, e.g., VideoMAE V2 (4.6% F1 for classification, 0.4% Dice for segmentation, and 2.1% F1 for detection). Kai Hu 0002, Yuan Zhang 0022, Xieping Gao 0001 |
NeurIPS | 1 |
| 2024 | Cross-level collaborative context-aware framework for medical image segmentation
Chao Suo, Tianxin Zhou, Kai Hu 0002, Yuan Zhang 0022, Xieping Gao 0001 |
Expert Syst. Appl. | 3 |
| 2024 | MCNet: A multi-level context-aware network for the segmentation of adrenal gland in CT images
Jinhao Li 0009, Huying Li, Yuan Zhang 0022, Xuanya Li, Kai Hu 0002, Xieping Gao 0001 |
Neural Networks | 7 |
| 2024 | Multi-scale object equalization learning network for intracerebral hemorrhage region segmentation
Yuan Zhang 0022, Yanglin Huang, Kai Hu 0002 |
Neural Networks | 3 |
| 2023 | Variational Clustering and Denoising of Spatial TranscriptomicsabstractSpatial transcriptomics data provides a unique opportunity to investigate both gene expression and spatial structure in tissues at the same time. However, incorporating spatial information to accurately identify spatial domains is difficult due to factors such as high-dimensionality, sparsity, noise, and dropout events. To address these issues, we introduce vGraphST, a novel graph-based deep learning approach tailored for spatial transcriptomics data. Our method combines auto-encoder and contrastive learning techniques to process high-dimensional data and generate meaningful low-dimensional embeddings. Additionally, we use continuous distributions instead of discrete values in both the latent space and the denoised gene expression space. Specifically, Gaussian distributions are used to model the latent space, while zero-inflated Poisson distributions are used to model the denoised gene expression space. Experimental results demonstrate the effectiveness of vGraphST in accurately representing and analyzing spatial transcriptomics data. When compared to other methods using the DLPFC dataset, vGraphST achieves an average Adjusted Rand Index (ARI) of 0.58, demonstrating its superiority in segmenting spatial domains and recognizing biologically relevant spatiotemporal patterns. Cuiyuan Li, Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 3 |
| 2023 | Self-distillation Augmented Masked Autoencoders for Histopathological Image UnderstandingabstractSelf-supervised learning (SSL) has drawn increasing attention in histopathological image analysis in recent years. Compared to contrastive learning which is troubled with the false negative problem, i.e., semantically similar images are selected as negative samples, masked autoencoders (MAE) build SSL from a generative paradigm which is probably a more appropriate pretraining. In this paper, we introduce MAE to histopathological image understanding, and moreover, verify the effect of visible patches in this task. Specifically, a novel SD-MAE model is proposed to enable a self-distillation augmented MAE. Besides the reconstruction loss on masked image patches, SD-MAE further imposes the self-distillation loss on visible patches to enhance the representational capacity of encoder located in the shallow layers. It generates a more effective feature pre-training and benefits downstream applications. We apply SD-MAE to histopathological image classification, cell segmentation and cell detection. Experiments demonstrate that SD-MAE shows highly competitive performance compared with other SSL methods in these tasks. Code is available at https://github.com/irsLu/SD-MAE/ Zhineng Chen, Shengtian Zhou, Kai Hu 0002, Xieping Gao 0001 |
BIBM | 4 |
| 2023 | Exploiting Multi-Decision and Deep Refinement for Ultrasound Image SegmentationabstractIn this paper, we propose a novel convolutional neural network (MDR-Net) for ultrasound image segmentation by exploiting multi-decision and deep refinement of the target. Our MDR-Net consists of two main parts, i.e., a multi-decision module (MDM) and a deep refinement module (DRM). Specifically, the MDM effectively addresses the issue of inconspicuous target regions in ultrasound images by combining multi-scale features and multi-receptive field self-attention to enhance the discriminative representation of features and diagnose feature points multiple times. In addition, to alleviate the problem of blurred boundaries and severe speckle noise, the DRM progressively fuses multi-scale features and makes the fused features interact with higher-level features to refine the target details step by step. Finally, we evaluate the proposed method on two publicly available datasets, namely BUSI and UDIAT. We achieve a Dice of 0.8265 and 0.8827 on the two datasets, which are at least 2% and 1.24% higher than other state-of-the-art ultrasound image segmentation methods. Xuanya Li, Kai Hu 0002, Xieping Gao 0001 |
ICASSP | 3 |
| 2023 | Pseudo Multi-Source Domain Extension and Selective Pseudo-Labeling for Unsupervised Domain Adaptive Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) attracts extra attention in medical image processing because no additional labels are required when adapting to different distributions. In this work, we propose a novel unsupervised domain adaptation framework named as Domain Expansion and PseudoLabeling (DEPL). We extend the domain of a labeled source domain data to four different distributed domains and use adversarial learning to align the image appearance level and feature level from the four different domains to the unlabeled target domain. In addition, we propose a selective pseudolabeling mechanism, namely using strong confidence pseudolabeling to boost model performance. We evaluate our model for the MR to CT adaptation segmentation task on the public dataset MMWHS. Compared to seven other state-of-the-art segmentation methods, our DEPL achieves the best Dice similarity coefficient by 82.4%, which is at least 3.9% higher than the other UDA segmentation methods. Kai Hu 0002, Xieping Gao 0001 |
ICASSP | 3 |
| 2023 | Automatic Segmentation of Nasopharyngeal Carcinoma in CT Images Using Dual Attention and Edge DetectionabstractNasopharyngeal carcinoma (NPC) is a malignant tumor with a high incidence. Accurate segmentation of the tumor region in Computed Tomography (CT) images of NPC is the key to treatment. However, the features of uneven grayscale values and hazy boundaries of NPC regions make accurate NPC segmentation particularly challenging. To address these problems, we propose an accurate and effective NPC segmentation method using Dual Attention and Edge Detection Convolutional Neural Network (DAED-Net). Firstly, we combine a 2.5D convolutional neural network with UNet++ and propose a new backbone called Dual-dimension Dense UNet (DD-UNet), which can extract more beneficial features from 3D images. Secondly, a Dual Attention Module (DAM) is proposed to help the model better segment the target region of NPC by efficiently collecting spatial and channel attention information from feature maps. Moreover, an Edge Detection Module (EDM) is introduced in the network to enhance the segmentation of the target contours. Finally, we evaluate the proposed DAED-Net on the public MICCAI 2019 StructSeg NPC dataset from different perspectives. Numerical and visual results show that the proposed method outperforms nine state-of-the-art segmentation methods and yields more accurate NPC segmentation results. Qizhi Wang, Yuan Zhang 0022, Xuanya Li, Xiongjun Ye, Kai Hu 0002 |
ICASSP | 6 |
| 2023 | Boundary Cue Guidance and Contextual Feature Mining for Glass SegmentationabstractGlass is ubiquitous in the real world, and its perception has many applications, including robot navigation and drone tracking. However, due to the transparent property of glass, the interior of a glass area can be any surrounding scene or object, which brings challenges for computer vision. Inspired by the human senses, boundary cues are one of the crucial factors for people to judge the location of glass contours. Hence, we propose a boundary cue guidance and contextual feature mining network (BCNet) to accurately and efficiently segment glass. Specifically, we first design a multi-branch boundary extraction module (MBEM) for learning accurate boundary cues combined with multi-level encoded features. Second, we propose a boundary cue guidance module (BCGM), inject the boundary cues into the representation learning, and provide constraints with object structure semantics to guide feature extraction. Besides, we design a contextual feature mining module (CFMM) to dynamically capture the contextual information of different receptive fields for the detection of different sizes and shapes of the glass. Finally, extensive experiments on two benchmark glass datasets, GDD and GSD. The results demonstrate that our BCNet achieves state-of-the-art segmentation performance against existing methods. Qiquan Xiao, Yuan Zhang 0022, Xuanya Li, Kai Hu 0002 |
ICASSP | 4 |
| 2023 | Transwnet: Integrating Transformers into CNNS via Row and Column Attention for Abdominal Multi-Organ SegmentationabstractLearning how to model global relationships and extract local details is crucial in improving the performance of multi-organ segmentation. Most existing U-shaped structure methods use feature fusion to address these two challenges, but still lack the ability to balance capturing global relationships and local details. To address these issues, we propose a novel multi-organ segmentation framework called TransWnet to mine global relationships and local details from both intra- and inter-scale perspectives. To achieve this, we innovatively design a Row and Column Swin Transformer (RCST) module that can efficiently capture global contextual features and construct local information. Specifically, we design a parallel structure of Row and Column Attention to model the global relationships of multi-scale encoded features, and further mine local information from the global relationships through a local window mechanism. Extensive experiments on the Synapse dataset show that our method outperforms state-of-the-art approaches and achieves accurate segmentation of abdominal multi-organs. Yazhen Xie, Yanglin Huang, Yuan Zhang 0022, Xuanya Li, Xiongjun Ye, Kai Hu 0002 |
ICASSP | 6 |
| 2023 | Polyp segmentation with distraction separation
Xiongjun Ye, Kai Hu 0002, Dapeng Xiong, Yuan Zhang 0022, Xuanya Li, Xieping Gao 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Boundary-Guided and Region-Aware Network With Global Scale-Adaptive for Accurate Segmentation of Breast Tumors in Ultrasound ImagesabstractBreast ultrasound (BUS) image segmentation is a critical procedure in the diagnosis and quantitative analysis of breast cancer. Most existing methods for BUS image segmentation do not effectively utilize the prior information extracted from the images. In addition, breast tumors have very blurred boundaries, various sizes and irregular shapes, and the images have a lot of noise. Thus, tumor segmentation remains a challenge. In this article, we propose a BUS image segmentation method using a boundary-guided and region-aware network with global scale-adaptive (BGRA-GSA). Specifically, we first design a global scale-adaptive module (GSAM) to extract features of tumors of different sizes from multiple perspectives. GSAM encodes the features at the top of the network in both channel and spatial dimensions, which can effectively extract multi-scale context and provide global prior information. Moreover, we develop a boundary-guided module (BGM) for fully mining boundary information. BGM guides the decoder to learn the boundary context by explicitly enhancing the extracted boundary features. Simultaneously, we design a region-aware module (RAM) for realizing the cross-fusion of diverse layers of breast tumor diversity features, which can facilitate the network to improve the learning ability of contextual features of tumor regions. These modules enable our BGRA-GSA to capture and integrate rich global multi-scale context, multi-level fine-grained details, and semantic information to facilitate accurate breast tumor segmentation. Finally, the experimental results on three publicly available datasets show that our model achieves highly effective segmentation of breast tumors even with blurred boundaries, various sizes and shapes, and low contrast. Kai Hu 0002, Xiang Zhang 0037, Dapeng Xiong, Yuan Zhang 0022, Xieping Gao 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Hexagonal Convolutional Neural Network for Spatial Transcriptomics ClassificationabstractRecent advances in spatial transcriptomics have enabled the comprehensive measurement of transcriptional profiles while retaining the spatial contextual information. Identifying spatial domains is a critical step in the analysis of spatially resolved transcriptomics. Existing unsupervised methods perform poorly on this task owing to the large amount of noise and dropout events in the transcriptomic profiles. To address this problem, we first extend an unsupervised algorithm to a supervised learning method that can identify useful features and reduce noise hindrance. Second, inspired by the classical convolution in convolutional neural networks (CNNs), we designed a regular hexagonal convolution to compensate for the missing gene expression patterns from adjacent nodes. Compared with the graph convolution in graph neural networks (GNNs), our hexagonal convolution can preserve the relative spatial location information of different nodes in graph-structured data. Third, based on the hexagonal convolution, a novel hexagonal Convolutional Neural Network (hexCNN) is proposed for spatial transcriptomics classification. Finally, we compared the proposed hexCNN with existing methods on the DLPFC dataset. The results show that hexCNN achieves a classification accuracy of 87.2% and an average Rand index (ARI) of 78.2% (1.9% and 3.3% higher than those of GNNs). Fa Zhang 0001, Kai Hu 0002, Xuefeng Cui |
BIBM | 3 |
| 2022 | TransMixer: A Hybrid Transformer and CNN Architecture for Polyp SegmentationabstractLearning how to fully extract global representations and local features is a key factor in improving the performance of polyp segmentation. In this paper, we explore the potential of combined techniques of Transformers and convolutional neural networks (CNNs) to address the challenges of polyp segmentation. Specifically, we present TransMixer, a hybrid interaction fusion architecture of the Transformer branch and the CNN branch, which is able to enhance the local details of global representations and the global context awareness of local features. To achieve this, we first bridge the semantic gap between the Transformer branch and the CNN branch through the Interaction Fusion Module (IFM), and then make full use of both respective properties to enhance polyp feature representations. After that, we further propose the Hierarchical Attention Module (HAM) to collect polyp semantic information from high-level features to gradually guide the recovery of polyp spatial information in low-level features. Quantitative and qualitative results show that the proposed model is more robust to various complex situations compared to existing methods, and achieves state-of-the-art performance in polyp segmentation. Yanglin Huang, Donghui Tan, Yuan Zhang 0022, Xuanya Li, Kai Hu 0002 |
BIBM | 5 |
| 2022 | A Novel Convolutional Neural Network Based on Adaptive Multi-Scale Aggregation and Boundary-Aware for Lateral Ventricle Segmentation on MR imagesabstractIn this paper, we propose a novel convolutional neural network based on adaptive multi-scale feature aggregation and boundary-aware for lateral ventricle segmentation (MB-Net), which mainly includes three parts, i.e., an adaptive multi-scale feature aggregation module (AMSFM), an embedded boundary refinement module (EBRM), and a local feature extraction module (LFM). Specifically, the AMSFM is used to extract multi-scale features through the different receptive fields to effectively solve the problem of distinct target regions on magnetic resonance (MR) images. The EBRM is intended to extract boundary information to effectively solve blurred boundary problems. The LFM can make the extraction of local information based on spatial and channel attention mechanisms to solve the problem of irregular shapes. Finally, extensive experiments are conducted from different perspectives to evaluate the performance of the proposed MB-Net. Furthermore, we also verify the robustness of the model on other public datasets, i.e., COVID-SemiSeg and CHASE DB1. The results show that our MB-Net can achieve competitive results when compared with state-of-the-art methods. Xuanya Li, Kai Hu 0002 |
ICASSP | 5 |
| 2022 | TransCoop: Cooperation of Transformers and CNNs for Camouflaged Object SegmentationabstractCamouflaged object segmentation (COS) is a challenging task due to the existence of high intrinsic similarities between the object and background. To overcome this challenge, we pro-pose a new framework, called TransCoop, for COS through the cooperation of Transformers and convolutional neural net-works (CNNs). Specifically, Transformer is used to model global context and structural information for accurately positioning potential target objects. Meanwhile, a well-designed texture feature fusion module (TFFM) is used to fuse the features encoded by CNN with the shallow Transformer to fully mine the low-level features in the scene. Furthermore, we propose a noise removal module (NRM), which can eliminate the background noise of low-level features with the guidance of precise target location. Notably, our TFFM and NRM can effectively realize the interaction between Transformer and CNN features. Extensive experiments on four benchmark datasets demonstrate the superiority of our TransCoop against existing state-of-the-art methods. Fucai Wu, Xuanya Li, Yuan Zhang 0022, Kai Hu 0002 |
ICME | 4 |
| 2022 | AS-Net: Attention Synergy Network for skin lesion segmentation
Kai Hu 0002, Dapeng Xiong, Zhineng Chen |
Expert Syst. Appl. | 1 |
| 2022 | Bridge-Net: Context-involved U-net with patch-based loss weight mapping for retinal blood vessel segmentation
Yuan Zhang 0022, Zhineng Chen, Kai Hu 0002, Xuanya Li, Xieping Gao 0001 |
Expert Syst. Appl. | 4 |
| 2021 | BGRA-Net: Boundary-Guided and Region-Aware Convolutional Neural Network for the Segmentation of Breast Ultrasound ImagesabstractIn this paper, we propose a novel convolutional neural network based on boundary-guided and region-aware (BGRA-Net) for breast tumor segmentation in ultrasound images. In particular, in the encoding stage, we propose a boundary-guided module (BGM) to guide the learning of boundary features in the decoding stage by explicitly strengthening the extracted boundary information. Meanwhile, in the decoding stage, we propose a region-aware module (RAM) to integrate different levels of detailed and semantic features to improve the comprehensive representation of tumor regional features. Besides, a scale-adaptive module (SAM) is further proposed to capture the characteristics of tumors with different sizes between the encoding and decoding stages. To evaluate the effectiveness of our BGRA-Net, we conduct extensive experiments on the UDIAT dataset and compare it with eight state-of-the-art methods. The experimental results show that our BGRA-Net outperforms the state-of-the-art methods and can achieve accurate segmentation of breast tumors with ambiguous boundaries. Xiang Zhang 0037, Xuanya Li, Kai Hu 0002, Xieping Gao 0001 |
BIBM | 3 |
| 2021 | A Hybrid Feature Enhancement Method for Gl And Segmentation In Histopathology ImagesabstractAccurate and automatic gland segmentation can help pathologists diagnose the malignancy of colorectal cancers. However, it remains a challenging task because of the large morphological differences between the glands and the presence of sticky glands. In this paper, a hybrid feature enhancement network (HFE-Net) for glandular segmentation is proposed, which includes a multi-scale local feature extraction block (MSLFEB) and a global feature enhancement block (GFEB). Specifically, the MSLFEB is used to extract multiscale features through different sizes of the receptive field to reduce the loss of the local information and effectively alleviate glandular adhesion. The GFEB is used to transfer the underlying features to the decoder by considering the global semantic information. Furthermore, we design a focal and variance (FV) loss function to alleviate the class imbalance and constraint the pixels within the same instance. Finally, we evaluate the proposed method on the 2015 MICCAI GlaS challenge dataset and the CRAG colorectal adenocarcinoma dataset. The results show that our HFE-Net can achieve competitive results with fewer computing resources when compared with the state-of-the-art gland segmentation methods. Xiangjiang Wu, Xuanya Li, Kai Hu 0002, Zhineng Chen, Xieping Gao 0001 |
ICASSP | 3 |
| 2021 | ERV-Net: An efficient 3D residual neural network for brain tumor segmentation
Xuanya Li, Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
Expert Syst. Appl. | 3 |
| 2021 | Deep supervised learning using self-adaptive auxiliary loss for COVID-19 diagnosis from imbalanced CT images
Kai Hu 0002, Zhineng Chen, Xuanya Li, Yuan Zhang 0022, Xieping Gao 0001 |
Neurocomputing | 1 |
| 2021 | A hierarchical and multi-view registration of serial histopathological images
Zhineng Chen, Kai Hu 0002, Shaoping Ling, Xieping Gao 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Nuclei Segmentation in Histopathology Images Using Rotation Equivariant and Multi-level Feature Aggregation Neural NetworkabstractThe histopathological analysis is the gold standard for assessing the presence and many complex diseases, like tumors. As one of the essential part of tumors, the shape, staining, and tissue distribution of the nuclei plays an important role in tumor diagnosis. However, due to nuclei congestion and possible occlusion, nuclei segmentation remains challenging. In this paper, we propose an automatic and effective nuclei segmentation method in histopathology images based on rotation equivariant and multi-level feature aggregation neural network (REMFANet). First, considering the inherent rotation equivariant of digital pathological images, we introduce group equivariant convolutions to improve the performance of the automatic segmentation of pathological images. Second, to eliminate the semantic gap between shallow features and deep features in encoder-decoder structural models, we propose a multi-level feature aggregation strategy based on U-Net 3+. Specifically, (1) we design a new decoder module to restore pixel-level predictions more accurately; (2) we propose an improved long-skip connection mode to provide richer semantic information in the decoder; (3) we also construct a semantic enhancement block to enhance the robustness of lowlevel semantic information. Finally, we evaluate our REMFA-Net on the MoNuSeg dataset and compare the results with seven state-of-the-art methods. Experimental results demonstrate the superiority of the proposed method over other models for the nuclei segmentation in histopathology images. Xuanya Li, Kai Hu 0002, Zhineng Chen, Xieping Gao 0001 |
BIBM | 3 |
| 2020 | HMOE-Net: Hybrid Multi-scale Object Equalization Network for Intracerebral Hemorrhage Segmentation in CT ImagesabstractIn this paper, we propose a novel Hybrid Multi-scale Object Equalization Network (HMOE-Net) to segment intracerebral hemorrhage (ICH) regions. In particular, we design a shallow feature extraction network (SFENet) and a deep feature extraction network (DFENet) to solve the problem of equalization learning of hybrid multi-scale object features. The multi-level feature extraction (MLFE) blocks are presented in DFENet to explore multi-level semantic features more effectively. Furthermore, we adopt a progressive feature extraction strategy combining SFENet and DFENet to further consider the differences of various ICH regions and achieve the equalization feature learning of multi-scale objects. To verify the effectiveness of HMOE-Net, we collect a clinical ICH dataset with a total of 500 CT cases from three hospitals for the evaluation. The experimental results show that HMOE-Net is superior to six state-of-the-art methods and achieves accurate segmentation for multi-scale ICH regions. Xizhi He, Kai Chen 0027, Kai Hu 0002, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
BIBM | 3 |
| 2020 | EffiDiag: an Efficient Framework for Breast Cancer Diagnosis in Multi-Gigapixel Whole Slide ImagesabstractBreast cancer diagnosis in multi-gigapixel whole slide images (WSIs) is an important task that highly relevant to cancer grading and prognosis. In recent years, many computer-aided diagnosis methods were proposed and achieved promising performance. However, they mostly suffer from heavy computational burden that becomes a significant barrier to clinical practice. Efficient solutions are urgently demanded but still less studied. In this paper, we propose a novel framework named EffiDiag for a fast and lightweight breast cancer diagnosis. To this end, a loss-modified U-net is developed at first to enable a fast suspected cancer Region Of Interest (ROI) localization. Therefore the subsequent patch-based classification, which commonly executes at the finest magnification hundreds of thousands times per WSI for cancer identification, could be carried out on these ROIs only rather than the whole WSI for speedup. Meanwhile, a super-efficient convolutional neural network (CNN) is devised to optimize the classification speed and resource consumption per classification. Experiments on the Camelyonl6 benchmark demonstrate, by integrating the two contributions into a well-established approach, 47x inference acceleration is obtained with limited accuracy drop, yet with much less resource consumption even compared to popular lightweight networks. Junda Ren, Zhineng Chen, Kai Hu 0002, Fen Xiao, Xuanya Li, Xieping Gao 0001 |
BIBM | 4 |
| 2020 | Compressive sensing MR imaging based on adaptive tight frame and reference imageabstractCompressive sensing magnetic resonance (MR) imaging is aimed at achieving high‐quality MR image reconstruction by undersampling K‐space data. It is crucial to explore prior information since compressive sensing MR imaging relies heavily on some prior assumptions, such as signal's sparse property. In this study, in order to explore the prior information fully, an improved MR image reconstruction model based on compressive sensing theory is proposed, named reference image MR imaging with adaptive tight frame. In the proposed model, an adaptive tight frame is involved to explore the sparse prior information adapt to MR images and the similarity prior information to the target image. Meanwhile, improved adaptive weighting parameters are used to trade off the sparsity between the regions with much similarity and that of little similarity. In addition, the smoothing‐based fast iterative shrinkage‐threshold algorithm is utilised to tackle the optimisation problem so as to speed up imaging. The experimental results demonstrate that the proposed MR image reconstruction method outperforms some state‐of‐the‐art methods in terms of quantitative results. Chunhong Cao, Kai Hu 0002, Fen Xiao |
IET Image Process. | 3 |
| 2020 | Automatic segmentation of intracerebral hemorrhage in CT images using encoder-decoder convolutional neural network
Kai Hu 0002, Kai Chen 0027, Xizhi He, Yuan Zhang 0022, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
Inf. Process. Manag. | 1 |
| 2020 | Automatic segmentation of dermoscopy images using saliency combined with adaptive thresholding based on wavelet transform
Kai Hu 0002, Yuan Zhang 0022, Chunhong Cao, Fen Xiao, Xieping Gao 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Markov multiple feature random fields model for the segmentation of brain MR images
Kai Hu 0002, Xieping Gao 0001, Yuan Zhang 0022 |
Expert Syst. Appl. | 1 |
| 2019 | Automatic segmentation of retinal layer boundaries in OCT images using multiscale convolutional neural network and graph search
Kai Hu 0002, Binwei Shen, Yuan Zhang 0022, Chunhong Cao, Fen Xiao, Xieping Gao 0001 |
Neurocomputing | 1 |
| 2019 | Hyperspectral image classification via compact-dictionary-based sparse representation
Chunhong Cao, Liu Deng, Fen Xiao, Wanchun Yang, Kai Hu 0002 |
Multim. Tools Appl. | 6 |
| 2019 | Print-scan invariant text image watermarking for hardcopy document authentication
Lina Tan, Kai Hu 0002, Xinmin Zhou, Rongyuan Chen, Weijin Jiang |
Multim. Tools Appl. | 2 |
| 2018 | Multi-scale deep neural network for salient object detectionabstractSalient object detection is a fundamental problem and has been received a great deal of attention in computer vision. Recently, deep learning model became a powerful tool for image feature extraction. In this study, the authors propose a multi‐scale deep neural network (MSDNN) for salient object detection. The proposed model first extracts global high‐level features and context information over the whole source image with the recurrent convolutional neural network. Then several stacked deconvolutional layers are adopted to get the multi‐scale feature representation and obtain a series of saliency maps. Finally, the authors investigate a fusion convolution module to build a final pixel level saliency map. The proposed model is extensively evaluated on six salient object detection benchmark datasets. Results show that the authors’ deep model significantly outperforms other 12 state‐of‐the‐art approaches. Fen Xiao, Wenzheng Deng, Liangchan Peng, Chunhong Cao, Kai Hu 0002, Xieping Gao 0001 |
IET Image Process. | 5 |
| 2018 | Retinal vessel segmentation of color fundus images using multiscale convolutional neural network with an improved cross-entropy loss function
Kai Hu 0002, Xiaorui Niu, Yuan Zhang 0022, Chunhong Cao, Fen Xiao, Xieping Gao 0001 |
Neurocomputing | 1 |
| 2017 | Microcalcification diagnosis in digital mammography using extreme learning machine based on hidden Markov tree model of dual-tree complex wavelet transform
Kai Hu 0002, Xieping Gao 0001 |
Expert Syst. Appl. | 1 |