EDBT 2026 Demo / reviewers in the wild / expert
Tong Zhang 0017
dblp:07/4227-17
· DBLP profile ↗
20ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-8838-4963ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language ModelsabstractVision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs and have achieved strong performance on standard image datasets. However, their effectiveness is often limited when confronted with cross-domain tasks where imaging domains differ from natural images. To address this limitation, we propose Consistency-guided Multi-view Collaborative Optimization (CoMuCo), a novel fine-tuning strategy for VLMs. This strategy employs two functionally complementary expert modules to extract multi-view features, while incorporating prior knowledge-based consistency constraints and information geometry-based consensus mechanisms to enhance the robustness of feature learning. Additionally, a new cross-domain few-shot benchmark is established to help comprehensively evaluate methods on imaging domains distinct from natural images. Extensive empirical evaluations on both existing and newly proposed benchmarks suggest CoMuCo consistently outperforms current methods. Dexia Chen, Wentao Zhang 0005, Qianjie Zhu, Weibing Li, Tong Zhang 0017 |
AAAI | 6 |
| 2026 | Switch-UMamba: Dynamic scanning vision Mamba UNet for medical image segmentation
Ziyao Zhang 0003, Qiankun Ma, Tong Zhang 0017, Jie Chen 0001, Hairong Zheng, Wen Gao 0001 |
Medical Image Anal. | 3 |
| 2026 | Adaptively Fine-Tuning and Ensembling Vision-Language Model for Few-Shot Image ClassificationabstractThe main challenge in few-shot learning (FSL) is overfitting of the model to limited training data. Recently developed vision-language models like CLIP have been employed to alleviate the overfitting issue and achieved state-of-the-art FSL performance. However, such approaches heavily depend on good alignments between images and associated texts, and therefore may not work as expected when image-text alignment is challenging in downstream tasks, e.g., fine-grained and cross domain image classification. In this study, a new CLIP-based fine-tuning and inference framework is proposed to particularly help the model more accurately recognize visually similar classes and also work well in new imaging domains. During fine-tuning, a modified contrastive loss with adaptively weighted negative pairs is proposed to effectively separate visually similar classes and cluster each class more compactly. During inference, an instance level adaptive ensemble strategy is proposed, utilizing the visual prototypes to adaptively complement the prediction from the CLIP's image-text alignment. Extensive experimental evaluations demonstrate the superiority of the proposed framework, outperforming current state-of-the-art methods by a decent margin on twelve public datasets. The source code will be released publicly. Baishun Dong, Xiaoyuan Guan, Wei-Shi Zheng 0001, Tong Zhang 0017 |
IEEE Trans. Multim. | 5 |
| 2025 | Test-Time Adaptation of Medical Vision-Language Models with Mixture of Modality ExpertsabstractTest-time adaptation (TTA) has emerged as a promising solution to improve the robustness of deep learning models under domain shifts, particularly in real-world scenarios where source data or true labels are unavailable. In this paper, we propose a novel TTA framework tailored for medical vision-language models (VLMs) that leverages a Mixture-of-Experts (MoE) mechanism. Building upon a frozen, pre-trained BiomedCLIP backbone, our method integrates parallel MoE adapters of different medical imaging modalities in each vision MLP block, enabling expert-specific adaptation without disrupting the core model representation. At inference time, only the MoE router is optimized through an entropy-regularized objective, which is further augmented by pseudo-label guidance and adaptive scaling strategies. Additionally, we propose an entropy-aware MoE scaling policy that dynamically adjusts expert influence based on prediction uncertainty, improving model adaptability. Extensive experiments on multiple medical imaging benchmarks demonstrate that our approach achieves substantial performance improvements over existing TTA baselines, while maintaining high efficiency and parameter sparsity. Our results highlight the potential of MoE-enhanced TTA to achieve robust and generalizable medical VLMs in unseen domains without access to source data. Code is available at https://openi.pcl.ac.cn/OpenMedIA/MoME-TTA.git Hancong Wang, Yue Yu 0001, Hairong Zheng, Tong Zhang 0017 |
ACM Multimedia | 4 |
| 2025 | Semi-supervised medical image segmentation via weak-to-strong perturbation consistency and edge-aware contrastive representation
Yang Yang 0002, Guoying Sun, Tong Zhang 0017, Jingyong Su |
Medical Image Anal. | 3 |
| 2024 | TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is crucial in many real-world applications. However, intelligent models are often trained solely on in-distribution (ID) data, leading to overconfidence when misclassifying OOD data as ID classes. In this study, we propose a new learning framework which leverage simple Jigsaw-based fake OOD data and rich semantic embeddings (`anchors') from the ChatGPT description of ID knowledge to help guide the training of the image encoder. The learning framework can be flexibly combined with existing post-hoc approaches to OOD detection, and extensive empirical evaluations on multiple OOD detection benchmarks demonstrate that rich textual representation of ID knowledge and fake OOD knowledge can well help train a visual encoder for OOD detection. With the learning framework, new state-of-the-art performance was achieved on all the benchmarks. The code is available at https://github.com/Cverchen/TagFog. Jiankang Chen, Tong Zhang 0017, Wei-Shi Zheng 0001 |
AAAI | 2 |
| 2024 | C2RG: Parameter-efficient Adaptation of 3D Vision and Language Foundation Model for Coronary CTA Report GenerationabstractMedical report generation (MRG) is a challenging yet highly demanding task in the application of multi-modal artificial intelligence in medicine. Typically, training an MRG model requires tens of thousands of labelled radiology images and reports datasets, which could be impractical for most clinical research groups. In this study, we present C2RG, a novel 3D vision and language foundation model tailored for Coronary Computed Tomography Angiography (CTA) Report Generation. Inspired by BLIP-2’s architecture, our method integrates a self-supervised pre-trained 3D cardiac vision model (ViT-B) and a general-purpose bilingual foundation model (ChatGLM-6B), with a lightweight querying Transformer (Q-Former). We also introduce a parallel high-resolution feature extractor module and a coronary calcification evaluation loss to simultaneously encode fine-grained 3D features and constrain the accuracy of report generation. We compared our model with six state-of-the-art MRG methods on a clinical dataset with 118 subjects, comprising 453 paired 3D CTA images and radiology reports. Experimental results with extensive ablations show the efficacy of our C2RG. Codes will be open-sourced after the conference. Zhiyu Ye, Bang Yang, Shibin Wu, Hancong Wang, Hairong Zheng, Tong Zhang 0017 |
BIBM | 10 |
| 2024 | SiFT: A Serial Framework with Textual Guidance for Federated Learning
Weizhuo Zhang, Yue Yu 0001, Wei-Shi Zheng 0001, Tong Zhang 0017 |
MICCAI (10) | 5 |
| 2024 | Continual Learning of Image Classes With Language Guidance From a Vision-Language ModelabstractCurrent deep learning models often catastrophically forget the knowledge of old classes when continually learning new ones. State-of-the-art approaches to continual learning of image classes often require retaining a small subset of old data to partly alleviate the catastrophic forgetting issue, and their performance would be degraded sharply when no old data can be stored due to privacy or safety concerns. In this study, inspired by human learning of visual knowledge with the effective help of language, we propose a novel continual learning framework based on a pre-trained vision-language model (VLM) without retaining any old data. Rich prior knowledge of each new image class is effectively encoded by the frozen text encoder of the VLM, which is then used to guide the learning of new image classes. The output space of the frozen text encoder is unchanged over the whole process of continual learning, through which image representations of different classes become comparable during model inference even when the image classes are learned at different times. Extensive empirical evaluations on multiple image classification datasets under various settings confirm the superior performance of our method over existing ones. The source code is available athttps://github.com/Fatflower/CIL_LG_VLM/. Wentao Zhang 0005, Yujun Huang, Weizhuo Zhang, Tong Zhang 0017, Qicheng Lao, Yue Yu 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Semi-supervised Medical Image Segmentation via Feature-perturbed ConsistencyabstractAlthough deep convolutional neural networks have achieved satisfactory performance in many medical image segmentation tasks, a considerable annotation challenge still needs to be solved, which is expensive and time-consuming for radiologists. Most existing popular semi-supervised methods mainly impose data-level perturbations (e.g., rotation, noising) or feature-level perturbations (e.g., MC dropout) on unlabeled data. In this paper, we propose a novel semi-supervised segmentation strategy with meaningful perturbations at the feature level to leverage abundant useful information naturally embedded in the unlabeled data. Specifically, we develop a dual-task network where the segmentation head produces multiple predictions with a perturbation module, and the reconstruction head further utilizes the semantic information to enhance segmentation performance. The proposed framework subtly perturbs the network at the feature-level to generate predictions which should be similar and consistent. However, enforcing them roughly to be consistent at all pixels harms stable training and neglects much delicate information. To better utilize those predictions and estimate the uncertainty, we further propose feature-perturbed consistency to exploit reliable regions for our framework to learn from. Extensive experiments on the public BraTS2020 dataset and the 2017 ACDC dataset confirm the efficiency and effectiveness of our method. In particular, the proposed method demonstrates remarkable superiority in the segmentation of boundary regions. The project is available at https://github.com/youngyzzZ/SFPC. Yang Yang 0002, Tong Zhang 0017, Jingyong Su |
BIBM | 3 |
| 2023 | Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases
Wentao Zhang 0005, Yujun Huang, Tong Zhang 0017, Qingsong Zou, Wei-Shi Zheng 0001 |
MICCAI (2) | 3 |
| 2023 | Two-Person Graph Convolutional Network for Skeleton-Based Human Interaction RecognitionabstractGraph convolutional networks (GCNs) have been the predominant methods in skeleton-based human action recognition, including human-human interaction recognition. However, when dealing with interaction sequences, current GCN-based methods simply split the two-person skeleton into two discrete graphs and perform graph convolution separately as done for single-person action classification. Such operations ignore rich interactive information and hinder effective spatial inter-body relationship modeling. To overcome the above shortcoming, we introduce a novel unified two-person graph to represent inter-body and intra-body correlations between joints. Experimental results show accuracy improvements in recognizing both interactions and individual actions when utilizing the proposed two-person graph topology. In addition, several graph labeling strategies are designed to supervise the model to learn discriminant spatial-temporal interactive features. Finally, we propose a two-person graph convolutional network (2P-GCN). Our model outperforms state-of-the-art methods on four benchmarks of three interaction datasets: SBU, interaction subsets of NTU-RGB+D and NTU-RGB+D 120. Zhengcen Li, Yueran Li, Linlin Tang, Tong Zhang 0017, Jingyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Consensus-Guided Keyword Targeting for Video Captioning
Puzhao Ji, Bang Yang, Tong Zhang 0017, Yuexian Zou |
PRCV (3) | 3 |
| 2022 | CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter
Bang Yang, Tong Zhang 0017, Yuexian Zou |
PRCV (1) | 2 |
| 2022 | GNN-Based Structural Dynamics Simulation for Modular Buildings
Tong Zhang 0017, Ying Wang 0025 |
PRCV (3) | 2 |
| 2022 | OpenMedIA: Open-Source Medical Image Analysis Toolbox and Benchmark Under Heterogeneous AI Computing Platforms
Jiaxin Zhuang, Xiansong Huang, Yang Yang 0002, Jiancong Chen, Yue Yu 0001, Wei Gao 0003, Ge Li 0002, Jie Chen 0001, Tong Zhang 0017 |
PRCV (1) | 9 |
| 2021 | EllipseNet: Anchor-Free Ellipse Detection for Automatic Cardiac Biometrics in Fetal Echocardiography
Jiancong Chen, Xiaoxue Zhou, Yihua He, Tong Zhang 0017 |
MICCAI (7) | 6 |
| 2020 | Deformable Slice-to-Volume Registration for Motion Correction of Fetal Body and Placenta MRIabstractIn in-utero MRI, motion correction for fetal body and placenta poses a particular challenge due to the presence of local non-rigid transformations of organs caused by bending and stretching. The existing slice-to-volume registration (SVR) reconstruction methods are widely employed for motion correction of fetal brain that undergoes only rigid transformation. However, for reconstruction of fetal body and placenta, rigid registration cannot resolve the issue of misregistrations due to deformable motion, resulting in degradation of features in the reconstructed volume. We propose a Deformable SVR (DSVR), a novel approach for non-rigid motion correction of fetal MRI based on a hierarchical deformable SVR scheme to allow high resolution reconstruction of the fetal body and placenta. Additionally, a robust scheme for structure-based rejection of outliers minimises the impact of registration errors. The improved performance of DSVR in comparison to SVR and patch-to-volume registration (PVR) methods is quantitatively demonstrated in simulated experiments and 20 fetal MRI datasets from 28-31 weeks gestational age (GA) range with varying degree of motion corruption. In addition, we present qualitative evaluation of 100 fetal body cases from 20-34 weeks GA range. Alena Uus, Tong Zhang 0017, Laurence H. Jackson, Thomas A. Roberts, Mary A. Rutherford, Joseph V. Hajnal, Maria Deprez |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Towards Whole Placenta Segmentation at Late Gestation Using Multi-view Ultrasound Images
Veronika A. M. Zimmer, Alberto Gómez 0002, Emily Skelton, Nicolas Toussaint, Tong Zhang 0017, Bishesh Khanal, Robert Wright, Yohan Noh, Alison Ho, Jacqueline Matthew, Joseph V. Hajnal, Julia A. Schnabel |
MICCAI (5) | 5 |
| 2014 | A clonal selection based approach to statistical brain voxel classification in magnetic resonance images
Tong Zhang 0017, Yong Xia 0001, David Dagan Feng |
Neurocomputing | 1 |