Ke Zhang 0029

dblp:20/4152-29 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-9855-003XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Fine-grained Adaptive Visual Prompt for Generative Medical Visual Question Answering
abstract
Medical Visual Question Answering (MedVQA) serves as an automated medical assistant, capable of answering patient queries and aiding physician diagnoses based on medical images and questions. Recent advancements have shown that incorporating Large Language Models (LLMs) into MedVQA tasks significantly enhances the capability for answer generation. However, for tasks requiring fine-grained organ-level precise localization, relying solely on language prompts struggles to accurately locate relevant regions within medical images due to substantial background noise. To address this challenge, we explore the use of visual prompts in MedVQA tasks for the first time and propose fine-grained adaptive visual prompts to enhance generative MedVQA. Specifically, we introduce an Adaptive Visual Prompt Creator that adaptively generates region-level visual prompts based on image characteristics of various organs, providing fine-grained references for LLMs during answer retrieval and generation from the medical domain, thereby improving the model's precise cross-modal localization capabilities on original images. Furthermore, we incorporate a Hierarchical Answer Generator with Parameter-Efficient Fine-Tuning (PEFT) techniques, significantly enhancing the model's understanding of spatial and contextual information with minimal parameter increase, promoting the alignment of representation learning with the medical space. Extensive experiments on VQA-RAD, SLAKE, and DME datasets validate the effectiveness of our proposed method, demonstrating its potential in generative MedVQA.
Ting Yu 0016, Zixuan Tong, Jun Yu 0002, Ke Zhang 0029
AAAI4
2025 CyclicAligner: Knowledge-Enhanced Cyclical Alignment for Chest X-Ray Report Generation
abstract
To reduce the diagnostic burden on radiologists, recent studies have explored automatic chest X-ray (CXR) report generation via artificial intelligence. Yet, achieving robust cross-modal alignment between medical images and textual reports remains a major challenge. In this paper, we propose CyclicAligner, a knowledge-enhanced cyclical alignment framework for CXR report generation. CyclicAligner adopts a novel cyclical training paradigm with four tightly coupled tasks to effectively learn cross-modal semantic alignment: (1) an image-to-text generation task that aligns visual semantics with clinical findings, (2) a text-to-text reconstruction task that strengthens language modeling, (3) a hybrid-to-text reconstruction task that mixes vision and language tokens for text reconstruction, and (4) a traceback-alignment task that re-encodes texts generated by the image-to-text branch for text reconstruction and aligns the reconstructed text with the reference. To further enhance cross-modal understanding, we integrate domain-specific medical entity knowledge extracted from a pre-trained encoder to enrich both vision and language tokens. Moreover, CyclicAligner jointly predicts medical tags and narrative reports within a unified auto-regressive pipeline, where the tags serve as auxiliary semantic anchors that guide the report generation. Extensive experiments on public datasets demonstrate the effectiveness of our method for clinical-coherent CXR report generation. The related code is available at https://github.com/yangyan22/CyclicAligner.
Jiamei Sun, Ke Zhang 0029, Xiangyu Tan, Zhenqi Fu
BIBM3
2025 Adapter-Enhanced Hierarchical Cross-Modal Pre-Training for Lightweight Medical Report Generation
abstract
Automatic medical report generation is an emerging field that aims to transform medical images into descriptive, clinically relevant narratives, potentially reducing the workload for radiologists significantly. Despite substantial progress, the increasing model parameter size and corresponding marginal performance gains have limited further development and application. To address this challenge, we introduce an Adapter-enhanced Hierarchical cross-modal Pre-training (AHP) strategy for lightweight medical report generation. This approach significantly reduces the pre-trained model's parameter size while maintaining superior report generation performance through our proposed spatial adapters. To further address the issue of inadequate representation of visual space details, we employ a convolutional stem combined with hierarchical injectors and extractors, fully integrating with traditional Vision Transformers to achieve more comprehensive visual representations. Additionally, our cross-modal pre-training model effectively handles the inherent complex visual-textual relationships in medical imaging. Extensive experiments on multiple datasets, including IU X-Ray, MIMIC-CXR, and bladder pathology, demonstrate our model's exceptional generalization and transfer performance in downstream medical report generation tasks, highlighting AHP's potential in significantly reducing model parameters while enhancing report generation accuracy and efficiency.
Ting Yu 0002, Wangwen Lu, Weidong Han 0001, Qingming Huang, Jun Yu 0002, Ke Zhang 0029
IEEE J. Biomed. Health Informatics7
2025 Spatio-Temporal and Retrieval-Augmented Modeling for Chest X-Ray Report Generation
abstract
Chest X-ray report generation has attracted increasing research attention. However, most existing methods neglect the temporal information and typically generate reports conditioned on a fixed number of images. In this paper, we propose STREAM: Spatio-Temporal and REtrieval-Augmented Modelling for automatic chest X-ray report generation. It mimics clinical diagnosis by integrating current and historical studies to interpret the present condition (temporal), with each study containing images from multi-views (spatial). Concretely, our STREAM is built upon an encoder-decoder architecture, utilizing a large language model (LLM) as the decoder. Overall, spatio-temporal visual dynamics are packed as visual prompts and regional semantic entities are retrieved as textual prompts. First, a token packer is proposed to capture condensed spatio-temporal visual dynamics, enabling the flexible fusion of images from current and historical studies. Second, to augment the generation with existing knowledge and regional details, a progressive semantic retriever is proposed to retrieve semantic entities from a preconstructed knowledge bank as heuristic text prompts. The knowledge bank is constructed to encapsulate anatomical chest X-ray knowledge into structured entities, each linked to a specific chest region. Extensive experiments on public datasets have shown the state-of-the-art performance of our method. Related codes and the knowledge bank are available at https://github.com/yangyan22/STREAM.
Xiaoxing You, Ke Zhang 0029, Zhenqi Fu, Xianyun Wang, Jiajun Ding, Jiamei Sun, Zhou Yu 0001, Qingming Huang, Weidong Han 0001, Jun Yu 0002
IEEE Trans. Medical Imaging3
2024 Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image Reporting
abstract
Medical image reporting focused on automatically generating the diagnostic reports from medical images has garnered growing research attention. In this task, learning cross-modal alignment between images and reports is crucial. However, the exposure bias problem in autoregressive text generation poses a notable challenge, as the model is optimized by a word-level loss function using the teacher-forcing strategy. To this end, we propose a novel Token-Mixer framework that learns to bind image and text in one embedding space for medical image reporting. Concretely, Token-Mixer enhances the cross-modal alignment by matching image-to-text generation with text-to-text generation that suffers less from exposure bias. The framework contains an image encoder, a text encoder and a text decoder. In training, images and paired reports are first encoded into image tokens and text tokens, and these tokens are randomly mixed to form the mixed tokens. Then, the text decoder accepts image tokens, text tokens or mixed tokens as prompt tokens and conducts text generation for network optimization. Furthermore, we introduce a tailored text decoder and an alternative training strategy that well integrate with our Token-Mixer framework. Extensive experiments across three publicly available datasets demonstrate Token-Mixer successfully enhances the image-text alignment and thereby attains a state-of-the-art performance. Related codes are available at https://github.com/yangyan22/Token-Mixer.
Jun Yu 0002, Zhenqi Fu, Ke Zhang 0029, Ting Yu 0016, Xianyun Wang, Hanliang Jiang, Junhui Lv, Qingming Huang, Weidong Han 0001
IEEE Trans. Medical Imaging4
2024 Attribute Prototype-Guided Iterative Scene Graph for Explainable Radiology Report Generation
abstract
The potential of automated radiology report generation in alleviating the time-consuming tasks of radiologists is increasingly being recognized in medical practice. Existing report generation methods have evolved from using image-level features to the latest approach of utilizing anatomical regions, significantly enhancing interpretability. However, directly and simplistically using region features for report generation compromises the capability of relation reasoning and overlooks the common attributes potentially shared across regions. To address these limitations, we propose a novel region-based Attribute Prototype-guided Iterative Scene Graph generation framework (AP-ISG) for report generation, utilizing scene graph generation as an auxiliary task to further enhance interpretability and relational reasoning capability. The core components of AP-ISG are the Iterative Scene Graph Generation (ISGG) module and the Attribute Prototype-guided Learning (APL) module. Specifically, ISSG employs an autoregressive scheme for structural edge reasoning and a contextualization mechanism for relational reasoning. APL enhances intra-prototype matching and reduces inter-prototype semantic overlap in the visual space to fully model the potential attribute commonalities among regions. Extensive experiments on the MIMIC-CXR with Chest ImaGenome datasets demonstrate the superiority of AP-ISG across multiple metrics.
Ke Zhang 0029, Jun Yu 0002, Jianping Fan 0007, Hanliang Jiang, Qingming Huang, Weidong Han 0001
IEEE Trans. Medical Imaging1
2024 Semi-Supervised Medical Report Generation via Graph-Guided Hybrid Feature Consistency
abstract
Medical report generation generates the corresponding report according to the given radiology image, which has been attracting increasing research interest. However, existing methods mainly adopt supervised training which rely on large amount of medical reports that are actually unavailable owing to the labor-intensive labeling process and privacy protection protocol. In the meanwhile, the intrinsic relationships between local pathological changes in the image are often ignored, which actually are important hints to high quality report generation. To this end, we propose a Relation-Aware Mean Teacher (RAMT) framework, which follows a standard mean teacher paradigm for semi-supervised report generation. The key to the encoder of the backbone network is the Graph-guided Hybrid Feature Encoding (GHFE) module, which exploits a prior disease knowledge graph to encode the intrinsic relations between pathological changes into the graph embedding and learns a word dictionary to retrieve the semantic embedding for each potential pathological change. GHFE combines the graph embedding, semantic embedding and visual features to form hybrid features, which are sent to a Transformer-based decoder for report generation. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the effectiveness of our proposed approach.
Ke Zhang 0029, Hanliang Jiang, Jian Zhang 0026, Qingming Huang, Jianping Fan 0007, Jun Yu 0002, Weidong Han 0001
IEEE Trans. Multim.1
2024 Multi-Task Paired Masking With Alignment Modeling for Medical Vision-Language Pre-Training
abstract
In recent years, the growing demand for medical imaging diagnosis has placed a significant burden on radiologists. As a solution, Medical Vision-Language Pre-training (Med-VLP) methods have been proposed to learn universal representations from medical images and reports, benefiting downstream tasks without requiring fine-grained annotations. However, existing methods have overlooked the importance of cross-modal alignment in joint image-text reconstruction, resulting in insufficient cross-modal interaction. To address this limitation, we propose a unified Med-VLP framework based on Multi-task Paired Masking with Alignment (MPMA) to integrate the cross-modal alignment task into the joint image-text reconstruction framework to achieve more comprehensive cross-modal interaction, while a Global and Local Alignment (GLA) module is designed to assist self-supervised paradigm in obtaining semantic representations with rich domain knowledge. Furthermore, we introduce a Memory-Augmented Cross-Modal Fusion (MA-CMF) module to fully integrate visual information to assist report reconstruction and fuse the multi-modal representations adequately. Experimental results demonstrate that the proposed unified approach outperforms previous methods in all downstream tasks, including uni-modal, cross-modal, and multi-modal tasks.
Ke Zhang 0029, Jun Yu 0002, Hanliang Jiang, Jianping Fan 0007, Qingming Huang, Weidong Han 0001
IEEE Trans. Multim.1
2023 Parallel spatio-temporal attention-based TCN for multivariate time series prediction
Jin Fan 0003, Ke Zhang 0029, Yipan Huang, Baiping Chen
Neural Comput. Appl.2
2022 CEKD: Cross ensemble knowledge distillation for augmented fine-grained data
Ke Zhang 0029, Jin Fan 0003, Shaoli Huang, Yongliang Qiao, Fei-wei Qin
Appl. Intell.1
2020 Multi-Order Feature Statistical Model for Fine-Grained Visual Categorization
abstract
Fine-grained visual categorization aims to learn a robust image representation modeling subtle differences from similar categories. Existing methods in this field tackle the problem by designing complex frameworks, which produce high-level features by performing first-order or second-order pooling. Despite the impressive performance achieved by these strategies, the single-order networks only carry linear or non-linear information of the last convolutional layer, neglecting the fact that features from different orders are mutually complementary. In this paper, we propose a multi-order feature statistical method (MOFS), which learns fine-grained features characterizing multiple orders. Specifically, the MOFS consists of two sub-modules: (i) a first-order module modeling both mid-level and high-level features. (ii) a covariance feature statistical module capturing high-order features. By deploying these two sub-modules on the top of existing backbone networks, MOFS simultaneously captures multi-level of discriminative patters including local, global and co-related patters. We evaluate the proposed method on three challenging benchmarks, namely CUB-200-2011, Stanford Cars, and FGVC-Aircraft. Compared with state-of-the-art methods, experiment results exhibit superior performance in recognizing fine-grained objects.
Qingtao Wang, Ke Zhang 0029, Jin Fan 0003, Shaoli Huang, Lianbo Zhang
ICPR2