Tong Zhang 0017

dblp:07/4227-17 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-8838-4963ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
abstract
Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs and have achieved strong performance on standard image datasets. However, their effectiveness is often limited when confronted with cross-domain tasks where imaging domains differ from natural images. To address this limitation, we propose Consistency-guided Multi-view Collaborative Optimization (CoMuCo), a novel fine-tuning strategy for VLMs. This strategy employs two functionally complementary expert modules to extract multi-view features, while incorporating prior knowledge-based consistency constraints and information geometry-based consensus mechanisms to enhance the robustness of feature learning. Additionally, a new cross-domain few-shot benchmark is established to help comprehensively evaluate methods on imaging domains distinct from natural images. Extensive empirical evaluations on both existing and newly proposed benchmarks suggest CoMuCo consistently outperforms current methods.
Dexia Chen, Wentao Zhang 0005, Qianjie Zhu, Weibing Li, Tong Zhang 0017
AAAI6
2026 Switch-UMamba: Dynamic scanning vision Mamba UNet for medical image segmentation
Ziyao Zhang 0003, Qiankun Ma, Tong Zhang 0017, Jie Chen 0001, Hairong Zheng, Wen Gao 0001
Medical Image Anal.3
2026 Adaptively Fine-Tuning and Ensembling Vision-Language Model for Few-Shot Image Classification
abstract
The main challenge in few-shot learning (FSL) is overfitting of the model to limited training data. Recently developed vision-language models like CLIP have been employed to alleviate the overfitting issue and achieved state-of-the-art FSL performance. However, such approaches heavily depend on good alignments between images and associated texts, and therefore may not work as expected when image-text alignment is challenging in downstream tasks, e.g., fine-grained and cross domain image classification. In this study, a new CLIP-based fine-tuning and inference framework is proposed to particularly help the model more accurately recognize visually similar classes and also work well in new imaging domains. During fine-tuning, a modified contrastive loss with adaptively weighted negative pairs is proposed to effectively separate visually similar classes and cluster each class more compactly. During inference, an instance level adaptive ensemble strategy is proposed, utilizing the visual prototypes to adaptively complement the prediction from the CLIP's image-text alignment. Extensive experimental evaluations demonstrate the superiority of the proposed framework, outperforming current state-of-the-art methods by a decent margin on twelve public datasets. The source code will be released publicly.
Baishun Dong, Xiaoyuan Guan, Wei-Shi Zheng 0001, Tong Zhang 0017
IEEE Trans. Multim.5
2025 Test-Time Adaptation of Medical Vision-Language Models with Mixture of Modality Experts
abstract
Test-time adaptation (TTA) has emerged as a promising solution to improve the robustness of deep learning models under domain shifts, particularly in real-world scenarios where source data or true labels are unavailable. In this paper, we propose a novel TTA framework tailored for medical vision-language models (VLMs) that leverages a Mixture-of-Experts (MoE) mechanism. Building upon a frozen, pre-trained BiomedCLIP backbone, our method integrates parallel MoE adapters of different medical imaging modalities in each vision MLP block, enabling expert-specific adaptation without disrupting the core model representation. At inference time, only the MoE router is optimized through an entropy-regularized objective, which is further augmented by pseudo-label guidance and adaptive scaling strategies. Additionally, we propose an entropy-aware MoE scaling policy that dynamically adjusts expert influence based on prediction uncertainty, improving model adaptability. Extensive experiments on multiple medical imaging benchmarks demonstrate that our approach achieves substantial performance improvements over existing TTA baselines, while maintaining high efficiency and parameter sparsity. Our results highlight the potential of MoE-enhanced TTA to achieve robust and generalizable medical VLMs in unseen domains without access to source data. Code is available at https://openi.pcl.ac.cn/OpenMedIA/MoME-TTA.git
Hancong Wang, Yue Yu 0001, Hairong Zheng, Tong Zhang 0017
ACM Multimedia4
2025 Semi-supervised medical image segmentation via weak-to-strong perturbation consistency and edge-aware contrastive representation
Yang Yang 0002, Guoying Sun, Tong Zhang 0017, Jingyong Su
Medical Image Anal.3
2024 TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial in many real-world applications. However, intelligent models are often trained solely on in-distribution (ID) data, leading to overconfidence when misclassifying OOD data as ID classes. In this study, we propose a new learning framework which leverage simple Jigsaw-based fake OOD data and rich semantic embeddings (`anchors') from the ChatGPT description of ID knowledge to help guide the training of the image encoder. The learning framework can be flexibly combined with existing post-hoc approaches to OOD detection, and extensive empirical evaluations on multiple OOD detection benchmarks demonstrate that rich textual representation of ID knowledge and fake OOD knowledge can well help train a visual encoder for OOD detection. With the learning framework, new state-of-the-art performance was achieved on all the benchmarks. The code is available at https://github.com/Cverchen/TagFog.
Jiankang Chen, Tong Zhang 0017, Wei-Shi Zheng 0001
AAAI2
2024 C2RG: Parameter-efficient Adaptation of 3D Vision and Language Foundation Model for Coronary CTA Report Generation
abstract
Medical report generation (MRG) is a challenging yet highly demanding task in the application of multi-modal artificial intelligence in medicine. Typically, training an MRG model requires tens of thousands of labelled radiology images and reports datasets, which could be impractical for most clinical research groups. In this study, we present C2RG, a novel 3D vision and language foundation model tailored for Coronary Computed Tomography Angiography (CTA) Report Generation. Inspired by BLIP-2’s architecture, our method integrates a self-supervised pre-trained 3D cardiac vision model (ViT-B) and a general-purpose bilingual foundation model (ChatGLM-6B), with a lightweight querying Transformer (Q-Former). We also introduce a parallel high-resolution feature extractor module and a coronary calcification evaluation loss to simultaneously encode fine-grained 3D features and constrain the accuracy of report generation. We compared our model with six state-of-the-art MRG methods on a clinical dataset with 118 subjects, comprising 453 paired 3D CTA images and radiology reports. Experimental results with extensive ablations show the efficacy of our C2RG. Codes will be open-sourced after the conference.
Zhiyu Ye, Bang Yang, Shibin Wu, Hancong Wang, Hairong Zheng, Tong Zhang 0017
BIBM10
2024 SiFT: A Serial Framework with Textual Guidance for Federated Learning
Weizhuo Zhang, Yue Yu 0001, Wei-Shi Zheng 0001, Tong Zhang 0017
MICCAI (10)5
2024 Continual Learning of Image Classes With Language Guidance From a Vision-Language Model
abstract
Current deep learning models often catastrophically forget the knowledge of old classes when continually learning new ones. State-of-the-art approaches to continual learning of image classes often require retaining a small subset of old data to partly alleviate the catastrophic forgetting issue, and their performance would be degraded sharply when no old data can be stored due to privacy or safety concerns. In this study, inspired by human learning of visual knowledge with the effective help of language, we propose a novel continual learning framework based on a pre-trained vision-language model (VLM) without retaining any old data. Rich prior knowledge of each new image class is effectively encoded by the frozen text encoder of the VLM, which is then used to guide the learning of new image classes. The output space of the frozen text encoder is unchanged over the whole process of continual learning, through which image representations of different classes become comparable during model inference even when the image classes are learned at different times. Extensive empirical evaluations on multiple image classification datasets under various settings confirm the superior performance of our method over existing ones. The source code is available athttps://github.com/Fatflower/CIL_LG_VLM/.
Wentao Zhang 0005, Yujun Huang, Weizhuo Zhang, Tong Zhang 0017, Qicheng Lao, Yue Yu 0001, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Semi-supervised Medical Image Segmentation via Feature-perturbed Consistency
abstract
Although deep convolutional neural networks have achieved satisfactory performance in many medical image segmentation tasks, a considerable annotation challenge still needs to be solved, which is expensive and time-consuming for radiologists. Most existing popular semi-supervised methods mainly impose data-level perturbations (e.g., rotation, noising) or feature-level perturbations (e.g., MC dropout) on unlabeled data. In this paper, we propose a novel semi-supervised segmentation strategy with meaningful perturbations at the feature level to leverage abundant useful information naturally embedded in the unlabeled data. Specifically, we develop a dual-task network where the segmentation head produces multiple predictions with a perturbation module, and the reconstruction head further utilizes the semantic information to enhance segmentation performance. The proposed framework subtly perturbs the network at the feature-level to generate predictions which should be similar and consistent. However, enforcing them roughly to be consistent at all pixels harms stable training and neglects much delicate information. To better utilize those predictions and estimate the uncertainty, we further propose feature-perturbed consistency to exploit reliable regions for our framework to learn from. Extensive experiments on the public BraTS2020 dataset and the 2017 ACDC dataset confirm the efficiency and effectiveness of our method. In particular, the proposed method demonstrates remarkable superiority in the segmentation of boundary regions. The project is available at https://github.com/youngyzzZ/SFPC.
Yang Yang 0002, Tong Zhang 0017, Jingyong Su
BIBM3
2023 Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
Wentao Zhang 0005, Yujun Huang, Tong Zhang 0017, Qingsong Zou, Wei-Shi Zheng 0001
MICCAI (2)3
2023 Two-Person Graph Convolutional Network for Skeleton-Based Human Interaction Recognition
abstract
Graph convolutional networks (GCNs) have been the predominant methods in skeleton-based human action recognition, including human-human interaction recognition. However, when dealing with interaction sequences, current GCN-based methods simply split the two-person skeleton into two discrete graphs and perform graph convolution separately as done for single-person action classification. Such operations ignore rich interactive information and hinder effective spatial inter-body relationship modeling. To overcome the above shortcoming, we introduce a novel unified two-person graph to represent inter-body and intra-body correlations between joints. Experimental results show accuracy improvements in recognizing both interactions and individual actions when utilizing the proposed two-person graph topology. In addition, several graph labeling strategies are designed to supervise the model to learn discriminant spatial-temporal interactive features. Finally, we propose a two-person graph convolutional network (2P-GCN). Our model outperforms state-of-the-art methods on four benchmarks of three interaction datasets: SBU, interaction subsets of NTU-RGB+D and NTU-RGB+D 120.
Zhengcen Li, Yueran Li, Linlin Tang, Tong Zhang 0017, Jingyong Su
IEEE Trans. Circuits Syst. Video Technol.4
2022 Consensus-Guided Keyword Targeting for Video Captioning
Puzhao Ji, Bang Yang, Tong Zhang 0017, Yuexian Zou
PRCV (3)3
2022 CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter
Bang Yang, Tong Zhang 0017, Yuexian Zou
PRCV (1)2
2022 GNN-Based Structural Dynamics Simulation for Modular Buildings
Tong Zhang 0017, Ying Wang 0025
PRCV (3)2
2022 OpenMedIA: Open-Source Medical Image Analysis Toolbox and Benchmark Under Heterogeneous AI Computing Platforms
Jiaxin Zhuang, Xiansong Huang, Yang Yang 0002, Jiancong Chen, Yue Yu 0001, Wei Gao 0003, Ge Li 0002, Jie Chen 0001, Tong Zhang 0017
PRCV (1)9
2021 EllipseNet: Anchor-Free Ellipse Detection for Automatic Cardiac Biometrics in Fetal Echocardiography
Jiancong Chen, Xiaoxue Zhou, Yihua He, Tong Zhang 0017
MICCAI (7)6
2020 Deformable Slice-to-Volume Registration for Motion Correction of Fetal Body and Placenta MRI
abstract
In in-utero MRI, motion correction for fetal body and placenta poses a particular challenge due to the presence of local non-rigid transformations of organs caused by bending and stretching. The existing slice-to-volume registration (SVR) reconstruction methods are widely employed for motion correction of fetal brain that undergoes only rigid transformation. However, for reconstruction of fetal body and placenta, rigid registration cannot resolve the issue of misregistrations due to deformable motion, resulting in degradation of features in the reconstructed volume. We propose a Deformable SVR (DSVR), a novel approach for non-rigid motion correction of fetal MRI based on a hierarchical deformable SVR scheme to allow high resolution reconstruction of the fetal body and placenta. Additionally, a robust scheme for structure-based rejection of outliers minimises the impact of registration errors. The improved performance of DSVR in comparison to SVR and patch-to-volume registration (PVR) methods is quantitatively demonstrated in simulated experiments and 20 fetal MRI datasets from 28-31 weeks gestational age (GA) range with varying degree of motion corruption. In addition, we present qualitative evaluation of 100 fetal body cases from 20-34 weeks GA range.
Alena Uus, Tong Zhang 0017, Laurence H. Jackson, Thomas A. Roberts, Mary A. Rutherford, Joseph V. Hajnal, Maria Deprez
IEEE Trans. Medical Imaging2
2019 Towards Whole Placenta Segmentation at Late Gestation Using Multi-view Ultrasound Images
Veronika A. M. Zimmer, Alberto Gómez 0002, Emily Skelton, Nicolas Toussaint, Tong Zhang 0017, Bishesh Khanal, Robert Wright, Yohan Noh, Alison Ho, Jacqueline Matthew, Joseph V. Hajnal, Julia A. Schnabel
MICCAI (5)5
2014 A clonal selection based approach to statistical brain voxel classification in magnetic resonance images
Tong Zhang 0017, Yong Xia 0001, David Dagan Feng
Neurocomputing1