VLDB 2026 Research / reviewers in the wild / expert
Gangming Zhao
dblp:204/2958
· DBLP profile ↗
34ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0001-8441-5462ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuroGT: Biophysically grounded graph transformers for self-supervised representation learning of neuronal morphology
Pengpeng Sheng, Gangming Zhao, Jun Wu 0024 |
Medical Image Anal. | 3 |
| 2026 | MambaPTP: Exploring the Potential of Mamba for Pedestrian Trajectory PredictionabstractPedestrian Trajectory Prediction (PTP) aims to predict the future trajectory of pedestrians based on a historical trajectory. Transformer-based approaches have demonstrated unparalleled performance for PTP tasks, encoding long-term temporal dependencies and heterogeneous spatial interactions of pedestrians. However, Transformer often involves redundant information and noisy interactions from irrelevant regions by considering all available trajectory features. Recently, the structured state space model, Mamba has been proposed, which captures long-range dependency in sequences with a selective mechanism to filter out redundant information. To further tap into the potential of the novel Mamba architecture for the PTP task, in this paper, we presentMambaPTP, which predicts future trajectories based purely on Mamba mechanisms, to mitigate the noisy interactions of irrelevant trajectory features and avoid repetitive trajectory modeling, while maintaining high-performance trajectory prediction. Specifically, we propose a new Bidirectional Gating Mamba (BGM) module with bidirectional state space models, which leverages the sparse gate mechanism to select informative temporal patterns and spatial interactions. Moreover, we design a Bidirectional Trajectory Alignment (BTA) module towards aligning the predicted trajectory to the ground truth, ensuring that the model to learn the effective sparse feature representation of trajectories. We conduct extensive experiments on several mainstream pedestrian trajectory prediction datasets. The results demonstrate that the proposed MambaPTP achieves competitive performance compared to advanced Transformer-based models. We hope this paper can further inspire research in Mamba for the PTP task, leading to a tighter integration of the Mamba and PTP communities. Shuangqing Zhang, Gangming Zhao, Fan Lyu, Songping Wang, Zhang Zhang 0001, Fang Zhao 0006, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Introducing DINOv2 for Medical Image Boundary Tracking
Gangming Zhao, Jun Wu 0024, Chong Tian |
ICIG (1) | 2 |
| 2025 | DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram ReconstructionabstractDecoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related ''foreground features'' from noisy ''background features'' through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework, achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https://github.com/fchest/DMF2Mel. Cunhang Fan, Enrui Liu, Gangming Zhao, Zhao Lv |
ACM Multimedia | 6 |
| 2025 | A Mamba-advanced unsupervised cross-modality medical image segmentation via domain adaptation and task decomposition
Danyang Peng, Jun Wu 0024, Feidan Kou, Gangming Zhao, Xiaohu Li |
Knowl. Based Syst. | 7 |
| 2025 | Federated Client-Tailored Adapter for Medical Image SegmentationabstractMedical image segmentation in X-ray images is beneficial for computer-aided diagnosis and lesion localization. Existing methods mainly fall into a centralized learning paradigm, which is inapplicable in the practical medical scenario that only has access to distributed data islands. Federated Learning has the potential to offer a distributed solution but struggles with heavy training instability due to client-wise domain heterogeneity (including distribution diversity and class imbalance). In this paper, we propose a novel Federated Client-tailored Adapter (FCA) framework for medical image segmentation, which achieves stable and client-tailored adaptive segmentation without sharing sensitive local data. Specifically, the federated adapter stirs universal knowledge in off-the-shelf medical foundation models to stabilize the federated training process. In addition, we develop two client-tailored federated updating strategies that adaptively decompose the adapter into common and individual components, then globally and independently update the parameter groups associated with common client-invariant and individual client-specific units, respectively. They further stabilize the heterogeneous federated learning process and realize optimal client-tailored instead of sub-optimal global-compromised segmentation models. Extensive experiments on three large-scale datasets demonstrate the effectiveness and superiority of the proposed FCA framework for federated medical segmentation. Guyue Hu 0001, Siyuan Song, Yukun Kang, Zhu Yin, Gangming Zhao, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Dynamic Strip Convolution and Adaptive Morphology Perception Plugin for Medical Anatomy SegmentationabstractMedical anatomy segmentation is essential for computer-aided diagnosis and lesion localization in medical images. For example, segmenting individual ribs benefits localizing the lung lesions and providing vital medical measurements (such as rib spacing) for generating medical reports. Existing methods segment shape-different anatomies (such as striped ribs, bulky lungs, and angular scapula) with the same network architecture, the morphology heterogeneity is heavily overlooked. Although some shape-aware operators like deformable convolution and dynamic snake convolution have been introduced to cater to specific object morphology, they still struggle with orientation-varying strip structures, such as 24 ribs and 2 clavicles. In this paper, we propose a novel convolution plugin (DSC-AMP) for medical anatomy segmentation, which is comprised of a dynamic strip convolution (DSC) operator and an adaptive morphology perception (AMP) strategy. Specifically, the dynamic strip convolution customizes gradually varying directions and offsets for each local region, achieving dynamic striped receptive fields. Additionally, the adaptive morphology perception strategy incorporates insights from various shape-aware convolutional kernels, enabling the model to discern and integrate crucial representations corresponding to heterogeneous anatomies. Extensive experiments on two large-scale datasets demonstrate the effectiveness and superiority of the proposed approach for tackling heterogeneous medical anatomy segmentation. Guyue Hu 0001, Yukun Kang, Gangming Zhao, Zhe Jin 0001, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Self-Supervised Neuron Morphology Representation With Graph TransformerabstractEffective representation of neuronal morphology is essential for cell typing and understanding brain function. However, the complexity of neuronal morphology manifests not only in inter-class structural differences but also in intra-class variations across developmental stages and environmental conditions. Such diversity poses significant challenges for existing methods in balancing robustness and discriminative power when representing neuronal morphology. To address this, we propose SGTMorph, a hybrid Graph Transformer framework that leverages the local topological modeling capabilities of graph neural networks and the global relational reasoning strengths of Transformers to explicitly encode neuronal structural information. SGTMorph incorporates a random walk-based positional encoding scheme to facilitate effective information propagation across neuronal graphs and introduces a spatially invariant encoding mechanism to improve adaptability with diverse morphology. This integrated approach enables a robust and comprehensive representation of neuronal morphology while maintaining biological fidelity. To enable label-free feature learning, we devise a self-supervised learning strategy grounded in geometric and topological similarity metrics. Extensive experiments on five datasets demonstrate SGTMorph's superior performance in neuron morphology classification and retrieval tasks. Furthermore, Its practical utility in neuronal function research is validated through the accurate predictions of two functional features: the laminar distribution of somas and axonal projection patterns. The code is available at https://github.com/big-rain/SGTMorph. Pengpeng Sheng, Gangming Zhao |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Clinical domain knowledge-derived template improves post hoc AI explanations in pneumothorax classification
Chuan Hong, Pengtao Jiang, Gangming Zhao, Nguyen Tuan Anh Tran, Xinxing Xu, Yet Yen Yan, Nan Liu 0003 |
J. Biomed. Informatics | 4 |
| 2024 | MVCNet: Multiview Contrastive Network for Unsupervised Representation Learning for 3-D CT LesionsabstractWith the renaissance of deep learning, automatic diagnostic algorithms for computed tomography (CT) have achieved many successful applications. However, they heavily rely on lesion-level annotations, which are often scarce due to the high cost of collecting pathological labels. On the other hand, the annotated CT data, especially the 3-D spatial information, may be underutilized by approaches that model a 3-D lesion with its 2-D slices, although such approaches have been proven effective and computationally efficient. This study presents a multiview contrastive network (MVCNet), which enhances the representations of 2-D views contrastively against other views of different spatial orientations. Specifically, MVCNet views each 3-D lesion from different orientations to collect multiple 2-D views; it learns to minimize a contrastive loss so that the 2-D views of the same 3-D lesion are aggregated, whereas those of different lesions are separated. To alleviate the issue of false negative examples, the uninformative negative samples are filtered out, which results in more discriminative features for downstream tasks. By linear evaluation, MVCNet achieves state-of-the-art accuracies on the lung image database consortium and image database resource initiative (LIDC-IDRI) (88.62%), lung nodule database (LNDb) (76.69%), and TianChi (84.33%) datasets for unsupervised representation learning. When fine-tuned on 10% of the labeled data, the accuracies are comparable to the supervised learning models (89.46% versus 85.03%, 73.85% versus 73.44%, 83.56% versus 83.34% on the three datasets, respectively), indicating the superiority of MVCNet in learning representations with limited annotations. Our findings suggest that contrasting multiple 2-D views is an effective approach to capturing the original 3-D information, which notably improves the utilization of the scarce and valuable annotated CT data. Penghua Zhai, Huaiwei Cong, Enwei Zhu, Gangming Zhao, Yizhou Yu, Jinpeng Li 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | START: Automatic Sleep Staging with Attention-based Cross-modal Learning TransformerabstractAutomatic sleep staging is vital to scale up sleep assessment and diagnosis to serve millions experiencing sleep deprivation and disorders and enable longitudinal sleep monitoring in home environments. However, how to learn from multi-channel raw physiological signal inputs (e.g., EEG and EOG) to capture the sleep stage and physiological signal relations remains a big challenge. In this paper, we propose a sleep staging model, named Sleep Staging Cross-modal Transformer (START), which is a transformer-only method for sleep stage classification. Our model is capable of learning a joint representation from both EEG and EOG signals by using a cross-modal fusion strategy. Experimental results show that our model outperforms the state-of-the-art methods on two public datasets. Furthermore, our model provides considerable reductions in parameters and training time compared to previous methods. Jingpeng Sun, Rongxiao Wang, Gangming Zhao, Chen Chen 0036, Yixiao Qu, Xiyuan Hu, Yizhou Yu |
BIBM | 3 |
| 2023 | Leveraging Frequency Domain Learning in 3D Vessel SegmentationabstractCoronary microvascular disease constitutes a substantial risk to human health. Employing computer-aided analysis and diagnostic systems, medical professionals can intervene early in disease progression, with 3D vessel segmentation serving as a crucial component. Nevertheless, conventional U-Net architectures tend to yield incoherent and imprecise segmentation outcomes, particularly for small vessel structures. While models with attention mechanisms, such as Transformers and large convolutional kernels, demonstrate superior performance, their extensive computational demands during training and inference lead to increased time complexity. In this study, we leverage Fourier domain learning as a substitute for multi-scale convo-lutional kernels in 3D hierarchical segmentation models, which can reduce computational expenses while preserving global receptive fields within the network. Furthermore, a zero-parameter frequency domain fusion method is designed to improve the skip connections in U-Net architecture. Experimental results on a public dataset and an in-house dataset indicate that our novel Fourier transformation-based network achieves remarkable dice performance (84.37% on ASACA500 and 80.32% on ImageCAS) in tubular vessel segmentation tasks and substantially reduces computational requirements without compromising global receptive fields. Xinyuan Wang 0009, Chengwei Pan, Hongming Dai, Gangming Zhao, Yizhou Yu |
BIBM | 4 |
| 2023 | Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsabstractGenerating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have difficulty in generating realistic and lip-synced videos while preserving identity information. To tackle this problem, we propose a two-stage framework consisting of audio-to-landmark generation and landmark-to-video rendering procedures. First, we devise a novel Transformer-based landmark generator to infer lip and jaw landmarks from the audio. Prior landmark characteristics of the speaker's face are employed to make the generated landmarks coincide with the facial outline of the speaker. Then, a video rendering model is built to translate the generated landmarks into face images. During this stage, prior appearance information is extracted from the lower-half occluded target face and static reference images, which helps generate realistic and identity-preserving visual content. For effectively exploring the prior information of static reference images, we align static reference images with the target face's pose and expression based on motion fields. Moreover, auditory features are reused to guarantee that the generated face images are well synchronized with the audio. Extensive experiments demonstrate that our method can produce more realistic, lip-synced, and identity-preserving videos than existing person-generic talking face generation methods. Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin 0004, Guanbin Li |
CVPR | 5 |
| 2023 | BEV@DC: Bird's-Eye View Assisted Training for Depth CompletionabstractDepth completion plays a crucial role in autonomous driving, in which cameras and LiDARs are two complementary sensors. Recent approaches attempt to exploit spatial geometric constraints hidden in LiDARs to enhance image-guided depth completion. However, only low efficiency and poor generalization can be achieved. In this paper, we propose BEV@DC, a more efficient and powerful multi-modal training scheme, to boost the performance of image-guided depth completion. In practice, the proposed BEV@DC model comprehensively takes advantage of LiDARs with rich geometric details in training, employing an enhanced depth completion manner in inference, which takes only images (RGB and depth) as input. Specifically, the geometric-aware LiDAR features are projected onto a unified BEV space, combining with RGB features to perform BEV completion. By equipping a newly proposed point-voxel spatial propagation network (PV-SPN), this auxiliary branch introduces strong guidance to the original image branches via 3D dense supervision and feature consistency. As a result, our baseline model demonstrates significant improvements with the sole image inputs. Concretely, it achieves state-of-the-art on several benchmarks, e.g., ranking Top-1 on the challenging KITTI depth completion benchmark. Wending Zhou, Xu Yan 0005, Yinghong Liao, Yuankai Lin, Gangming Zhao, Shuguang Cui, Zhen Li 0026 |
CVPR | 6 |
| 2023 | Learning Locality and Isotropy in Dialogue Modeling
Han Wu 0004, Haochen Tan, Mingjie Zhan, Gangming Zhao, Shaoqing Lu, Ding Liang, Linqi Song |
ICLR | 4 |
| 2023 | Dynamic Triple Reweighting Network for Automatic Femoral Head Necrosis Diagnosis from Computed TomographyabstractAvascular necrosis of the femoral head (AVNFH) is a common orthopedic disease that seriously affects the life quality of middle-aged and elderly people. Early AVNFH is difficult to diagnose due to its complex symptoms. In recent years, some works have applied deep learning algorithms to find traces of early AVNFH in X-rays or magnetic resonance imaging (MRI). However, X-rays are difficult to reflect hidden features due to the tissue overlap; MRI is sensitive but requires more time for imaging and is expensive. This study aims to develop a computer-aided diagnosis system for early AVNFH based on computed tomography (CT), which provides layer-wise features and is less costly. To achieve this, a large-scale dataset for AVNFH was collected and annotated by experienced doctors. We propose the Dynamic Triple Reweighting Network (DTRNet) that integrates the AVNFH classification and weakly-supervised localization. DTRNet incorporates nested multi-instance learning as the first and second reweighting, and structure regularization as the third reweighting to identify diseases and localize the lesion region. Since nested multi-instance learning is inapplicable in situations with few positive samples in the patch set, we propose a dynamic pseudo-package module to compensate for this limitation. Experimental results show that DTRNet is superior to the baselines in AVNFH classification. In addition, it can locate lesions to provide more information for assisting clinical decisions. The desensitized data and codes has been made available at: https://github.com/tomas-lilingfeng/DTRNet. Gangming Zhao, Yizhou Yu, Jinpeng Li 0002 |
ACM Multimedia | 2 |
| 2023 | Graph Convolution Based Cross-Network Multiscale Feature Fusion for Deep Vessel SegmentationabstractVessel segmentation is widely used to help with vascular disease diagnosis. Vessels reconstructed using existing methods are often not sufficiently accurate to meet clinical use standards. This is because 3D vessel structures are highly complicated and exhibit unique characteristics, including sparsity and anisotropy. In this paper, we propose a novel hybrid deep neural network for vessel segmentation. Our network consists of two cascaded subnetworks performing initial and refined segmentation respectively. The second subnetwork further has two tightly coupled components, a traditional CNN-based U-Net and a graph U-Net. Cross-network multi-scale feature fusion is performed between these two U-shaped networks to effectively support high-quality vessel segmentation. The entire cascaded network can be trained from end to end. The graph in the second subnetwork is constructed according to a vessel probability map as well as appearance and semantic similarities in the original CT volume. To tackle the challenges caused by the sparsity and anisotropy of vessels, a higher percentage of graph nodes are distributed in areas that potentially contain vessels while a higher percentage of edges follow the orientation of potential nearby vessels. Extensive experiments demonstrate our deep network achieves state-of-the-art 3D vessel segmentation performance on multiple public and in-house datasets. Gangming Zhao, Kongming Liang, Chengwei Pan, Fandong Zhang, Xianpeng Wu, Xinyang Hu, Yizhou Yu |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Structure Regularized Attentive Network for Automatic Femoral Head Necrosis Diagnosis and LocalizationabstractIn recent years, several works have adopted the convolutional neural network (CNN) to diagnose the avascular necrosis of the femoral head (AVNFH) based on X-ray images or magnetic resonance imaging (MRI). However, due to the tissue overlap, X-ray images are difficult to provide fine-grained features for early diagnosis. MRI, on the other hand, has a long imaging time, is more expensive, making it impractical in mass screening. Computed tomography (CT) shows layer-wise tissues, is faster to image, and is less costly than MRI. However, to our knowledge, there is no work on CT-based automated diagnosis of AVNFH. In this work, we collected and labeled a large-scale dataset for AVNFH ranking. In addition, existing end-to-end CNNs only yields the classification result and are difficult to provide more information for doctors in diagnosis. To address this issue, we propose the structure regularized attentive network (SRANet), which is able to highlight the necrotic regions during classification based on patch attention. SRANet extracts features in chunks of images, obtains weight via the attention mechanism to aggregate the features, and constrains them by a structural regularizer with prior knowledge to improve the generalization. SRANet was evaluated on our AVNFH-CT dataset. Experimental results show that SRANet is superior to CNNs for AVNFH classification, moreover, it can localize lesions and provide more information to assist doctors in diagnosis. Our codes are made public at https://github.com/tomas-lilingfeng/SRANet. Huaiwei Cong, Gangming Zhao, Junran Peng, Zheng Zhang 0006, Jinpeng Li 0002 |
BIBM | 3 |
| 2022 | Deep 3D Vessel Segmentation based on Cross Transformer NetworkabstractThe coronary microvascular disease poses a great threat to human health. Computer-aided analysis/diagnosis systems help physicians intervene in the disease at early stages, where 3D vessel segmentation is a fundamental step. However, there is a lack of carefully annotated dataset to support algorithm development and evaluation. On the other hand, the commonly-used U-Net structures often yield disconnected and inaccurate segmentation results, especially for small vessel structures. In this paper, motivated by the data scarcity, we first construct two large-scale vessel segmentation datasets consisting of 100 and 500 computed tomography (CT) volumes with pixel-level annotations by experienced radiologists. To enhance the U-Net, we further propose the cross transformer network (CTN) for fine-grained vessel segmentation. In CTN, a transformer module is constructed in parallel to a U-Net to learn long-distance dependencies between different anatomical regions; and these dependencies are communicated to the U-Net at multiple stages to endow it with global awareness. Experimental results on the two in-house datasets indicate that this hybrid model alleviates unexpected disconnections by considering topological information across regions. Our codes, together with the trained models are made publicly available at https://github.com/qibaolian/ctn. Chengwei Pan, Baolian Qi, Gangming Zhao, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002 |
BIBM | 3 |
| 2022 | BOAT: Bilateral Local Attention Vision Transformer
Gangming Zhao, Ping Li 0001, Yizhou Yu |
BMVC | 2 |
| 2022 | OneFace: One Threshold for All
Haoyu Qin, Yichao Wu, Ding Liang, Gangming Zhao, Ke Xu 0001 |
ECCV (12) | 6 |
| 2022 | Computer-Aided Tuberculosis Diagnosis with Attribute Reasoning Assistance
Chengwei Pan, Gangming Zhao, Junjie Fang, Baolian Qi, Chaowei Fang, Dingwen Zhang, Jinpeng Li 0002, Yizhou Yu |
MICCAI (1) | 2 |
| 2022 | Mix and Reason: Reasoning over Semantic Topology with Data Mixing for Domain GeneralizationabstractDomain generalization (DG) enables generalizing a learning machine from multiple seen source domains to an unseen target one. The general objective of DG methods is to learn semantic representations that are independent of domain labels, which is theoretically sound but empirically challenged due to the complex mixture of common and domain-specific factors. Although disentangling the representations into two disjoint parts has been gaining momentum in DG, the strong presumption over the data limits its efficacy in many real-world scenarios. In this paper, we propose Mix and Reason (MiRe), a new DG framework that learns semantic representations via enforcing the structural invariance of semantic topology. MiRe consists of two key components, namely, Category-aware Data Mixing (CDM) and Adaptive Semantic Topology Refinement (ASTR). CDM mixes two images from different domains in virtue of activation maps generated by two complementary classification losses, making the classifier focus on the representations of semantic objects. ASTR introduces relation graphs to represent semantic topology, which is progressively refined via the interactions between local feature aggregation and global cross-domain relational reasoning. Experiments on multiple DG benchmarks validate the effectiveness and robustness of the proposed MiRe. Chaoqi Chen, Luyao Tang, Feng Liu 0036, Gangming Zhao, Yue Huang 0001, Yizhou Yu |
NeurIPS | 4 |
| 2022 | Diagnose Like a Radiologist: Hybrid Neuro-Probabilistic Reasoning for Attribute-Based Medical Image DiagnosisabstractDuring clinical practice, radiologists often use attributes, e.g., morphological and appearance characteristics of a lesion, to aid disease diagnosis. Effectively modeling attributes as well as all relationships involving attributes could boost the generalization ability and verifiability of medical image diagnosis algorithms. In this paper, we introduce a hybrid neuro-probabilistic reasoning algorithm for verifiable attribute-based medical image diagnosis. There are two parallel branches in our hybrid algorithm, a Bayesian network branch performing probabilistic causal relationship reasoning and a graph convolutional network branch performing more generic relational modeling and reasoning using a feature representation. Tight coupling between these two branches is achieved via a cross-network attention mechanism and the fusion of their classification results. We have successfully applied our hybrid reasoning algorithm to two challenging medical image diagnosis tasks. On the LIDC-IDRI benchmark dataset for benign-malignant classification of pulmonary nodules in CT images, our method achieves a new state-of-the-art accuracy of 95.36% and an AUC of 96.54%. Our method also achieves a 3.24% accuracy improvement on an in-house chest X-ray image dataset for tuberculosis diagnosis. Our ablation study indicates that our hybrid algorithm achieves a much better generalization performance than a pure neural network architecture under very limited training data. Gangming Zhao, Quanlong Feng, Chaoqi Chen, Yizhou Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | GREN: Graph-Regularized Embedding Network for Weakly-Supervised Disease Localization in X-Ray ImagesabstractLocating diseases in chest X-ray images with few careful annotations saves large human effort. Recent works approached this task with innovative weakly-supervised algorithms such as multi-instance learning (MIL) and class activation maps (CAM), however, these methods often yield inaccurate or incomplete regions. One of the reasons is the neglection of the pathological implications hidden in the relationship across anatomical regions within each image and the relationship across images. In this paper, we argue that the cross-region and cross-image relationship, as contextual and compensating information, is vital to obtain more consistent and integral regions. To model the relationship, we propose the Graph Regularized Embedding Network (GREN), which leverages the intra-image and inter-image information to locate diseases on chest X-ray images. GREN uses a pre-trained U-Net to segment the lung lobes, and then models the intra-image relationship between the lung lobes using an intra-image graph to compare different regions. Meanwhile, the relationship between in-batch images is modeled by an inter-image graph to compare multiple images. This process mimics the training and decision-making process of a radiologist: comparing multiple regions and images for diagnosis. In order for the deep embedding layers of the neural network to retain structural information (important in the localization task), we use the Hash coding and Hamming distance to compute the graphs, which are used as regularizers to facilitate training. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for weakly-supervised disease localization. Our codes are accessible online. Baolian Qi, Gangming Zhao, Changde Du, Chengwei Pan, Yizhou Yu, Jinpeng Li 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Weakly Supervised Disease Localization in Chest X-rays via Looking into Image RelationsabstractLocating diseases in chest X-ray images with few careful annotations saves large human effort in annotation. Recent works tackled this problem with innovative weakly-supervised algorithms, however, the performance of these methods on X-ray analysis is not as good as in the general computer vision tasks. Different from natural images, the global structure of different chest X-rays are relatively consistent and the disease regions are relatively inconspicuous, and radiologists often need to compare multiple images to make diagnostic decisions. Inspired by this, we propose a hypothesis that the explicit modelling of the image-to-image relationship is beneficial to the learning machines, especially when the supervision is insufficient. To model the relationship, we exploit a cross-image graph method to excavate the structural relationship between X-ray images. The proposed method represents the inter-image relationship as in-batch graphs, where each X-ray image is regarded as a node and the distance between X-ray images is defined as an edge. The graphs are used as regularizers to help preserve the structural similarity between image pairs in the embedding space. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for disease localization with limited supervision, which has a very practical use in the weakly-supervised localization tasks. The code1is accessible online. Baolian Qi, Gangming Zhao, Chaowei Fang, Zhiqiang Chen 0002, Jinpeng Li 0002 |
BIBM | 2 |
| 2021 | GraphFPN: Graph Feature Pyramid Network for Object DetectionabstractFeature pyramids have been proven powerful in image understanding tasks that require multi-scale features. State-of-the-art methods for multi-scale feature learning focus on performing feature interactions across space and scales using neural networks with a fixed topology. In this paper, we propose graph feature pyramid networks that are capable of adapting their topological structures to varying intrinsic image structures, and supporting simultaneous feature interactions across all scales. We first define an image specific superpixel hierarchy for each input image to represent its intrinsic image structures. The graph feature pyramid network inherits its structure from this superpixel hierarchy. Contextual and hierarchical layers are designed to achieve feature interactions within the same scale and across different scales. To make these layers more powerful, we introduce two types of local channel attention for graph neural networks by generalizing global channel attention for convolutional neural networks. The proposed graph feature pyramid network can enhance the multiscale features from a convolutional feature pyramid network.We evaluate our graph feature pyramid network in the object detection task by integrating it into the Faster R-CNN algorithm. The modified algorithm outperforms not only previous state-of-the-art feature pyramid based methods with a clear margin but also other popular detection methods on both MS-COCO 2017 validation and test datasets. Gangming Zhao, Weifeng Ge, Yizhou Yu |
ICCV | 1 |
| 2021 | Multi-scale Matching Networks for Semantic CorrespondenceabstractDeep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convolutional neural networks has not been well studied to learn discriminative pixel-level features for semantic correspondence. In this paper, we propose a multi-scale matching network that is sensitive to tiny semantic differences between neighboring pixels. We follow the coarse-to-fine matching strategy and build a top-down feature and matching enhancement scheme that is coupled with the multi-scale hierarchy of deep convolutional neural networks. During feature enhancement, intra-scale enhancement fuses same-resolution feature maps from multiple layers together via local self-attention and cross-scale enhancement hallucinates higher-resolution feature maps along the top-down pathway. Besides, we learn complementary matching details at different scales thus the overall matching score is refined by features of different semantic levels gradually. Our multi-scale matching network can be trained end-to-end easily with few additional learnable parameters. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on three popular benchmarks with high computational efficiency. The code has been released at https://github.com/wintersun661/MMNet. Dongyang Zhao, Zhenghao Ji, Gangming Zhao, Weifeng Ge, Yizhou Yu |
ICCV | 4 |
| 2021 | Cross Chest Graph for Disease Diagnosis with Structural Relational ReasoningabstractLocating lesions is important in the computer-aided diagnosis of X-ray images. However, box-level annotation is time-consuming and laborious. How to locate lesions accurately with few, or even without careful annotations is an urgent problem. Although several works have approached this problem with weakly-supervised methods, the performance needs to be improved. One obstacle is that general weakly-supervised methods have failed to consider the characteristics of X-ray images, such as the highly-structural attribute. We therefore propose the Cross-chest Graph (CCG), which improves the performance of automatic lesion detection by imitating doctor's training and decision-making process. CCG models the intra-image relationship between different anatomical areas by leveraging the structural information to simulate the doctor's habit of observing different areas. Meanwhile, the relationship between any pair of images is modeled by a knowledge-reasoning module to simulate the doctor's habit of comparing multiple images. We integrate intra-image and inter-image information into a unified end-to-end framework. Experimental results on the NIH Chest-14 database (112,120 frontal-view X-ray images with 14 diseases) demonstrate that the proposed method achieves state-of-the-art performance in weakly-supervised localization of lesions by absorbing professional knowledge in the medical field. Gangming Zhao |
ACM Multimedia | 1 |
| 2021 | Multi-task contrastive learning for automatic CT and X-ray diagnosis of COVID-19
Jinpeng Li 0002, Gangming Zhao, Yaling Tao, Penghua Zhai, Hao Chen 0081, Huiguang He, Ting Cai 0001 |
Pattern Recognit. | 2 |
| 2021 | Contralaterally Enhanced Networks for Thoracic Disease DetectionabstractIdentifying and locating diseases in chest X-rays are very challenging, due to the low visual contrast between normal and abnormal regions, and distortions caused by other overlapping tissues. An interesting phenomenon is that there exist many similar structures in the left and right parts of the chest, such as ribs, lung fields and bronchial tubes. This kind of similarities can be used to identify diseases in chest X-rays, according to the experience of broad-certificated radiologists. Aimed at improving the performance of existing detection methods, we propose a deep end-to-end module to exploit the contralateral context information for enhancing feature representations of disease proposals. First of all, under the guidance of the spine line, the spatial transformer network is employed to extract local contralateral patches, which can provide valuable context information for disease proposals. Then, we build up a specific module, based on both additive and subtractive operations, to fuse the features of the disease proposal and the contralateral patch. Our method can be integrated into both fully and weakly supervised disease detection frameworks. It achieves 33.17 AP50 on a carefully annotated private chest X-ray dataset which contains 31,000 images. Experiments on the NIH chest X-ray dataset indicate that our method achieves state-of-the-art performance in weakly-supervised disease localization. Gangming Zhao, Chaowei Fang, Guanbin Li, Licheng Jiao, Yizhou Yu |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Align, Attend and Locate: Chest X-Ray Diagnosis via Contrast Induced Attention Network With Limited SupervisionabstractObstacles facing accurate identification and localization of diseases in chest X-ray images lie in the lack of high-quality images and annotations. In this paper, we propose a Contrast Induced Attention Network (CIA-Net), which exploits the highly structured property of chest X-ray images and localizes diseases via contrastive learning on the aligned positive and negative samples. To force the attention module to focus only on sites of abnormalities, we also introduce a learnable alignment module to adjust all the input images, which eliminates variations of scales, angles, and displacements of X-ray images generated under bad scan conditions. We show that the use of contrastive attention and alignment module allows the model to learn rich identification and localization information using only a small amount of location annotations, resulting in state-of-the-art performance in NIH chest X-ray dataset. Jingyu Liu 0004, Gangming Zhao, Ming Zhang 0004, Yizhou Wang 0001, Yizhou Yu |
ICCV | 2 |
| 2018 | Rethinking ReLU to Train Better CNNsabstractMost of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the most widely used. In this paper, we argue that the designed structure with the equal ratio between these two layers may not be the best choice since it could result in the poor generalization ability. Thus, we try to investigate a more suitable method on using ReL U to explore the better network architectures. Specifically, we propose a proportional module to keep the ratio between convolution and ReLU amount to be N:m (n>m). The proportional module can be applied in almost all networks with no extra computational cost to improve the performance. Comprehensive experimental results indicate that the proposed method achieves better performance on different benchmarks with different network architectures, thus verify the superiority of our work. Gangming Zhao, Zhaoxiang Zhang 0001, He Guan, Peng Tang 0005, Jingdong Wang 0001 |
ICPR | 1 |
| 2017 | Random Shifting for CNN: a Solution to Reduce Information Loss in Down-Sampling LayersabstractDown-sampling is widely adopted in deep convolutional neural networks (DCNN) for reducing the number of network parameters while preserving the transformation invariance. However, it cannot utilize information effectively because it only adopts a fixed stride strategy, which may result in poor generalization ability and information loss. In this paper, we propose a novel random strategy to alleviate these problems by embedding random shifting in the down-sampling layers during the training process. Random shifting can be universally applied to diverse DCNN models to dynamically adjust receptive fields by shifting kernel centers on feature maps in different directions. Thus, it can generate more robust features in networks and further enhance the transformation invariance of down-sampling operators. In addition, random shifting cannot only be integrated in all down-sampling layers including strided convolutional layers and pooling layers, but also improve performance of DCNN with negligible additional computational cost. We evaluate our method in different tasks (e.g., image classification and segmentation) with various network architectures (i.e., AlexNet, FCN and DFN-MR). Experimental results demonstrate the effectiveness of our proposed method. Gangming Zhao, Jingdong Wang 0001, Zhaoxiang Zhang 0001 |
IJCAI | 1 |