EDBT 2026 Demo / reviewers in the wild / expert
Xiaodan Zhang 0003
dblp:29/2631-3
· DBLP profile ↗
25ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-7002-5447ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARE: Multimodal Analogical Reasoning for Disease Evolution-Aware Radiology Report GenerationabstractRadiology report generation from longitudinal medical data is critical for assessing disease progression and automating diagnostic workflows. While recent methods incorporate longitudinal information, they primarily rely on multimodal feature fusion, with limited capacity for explicit disease evolution modeling and temporal reasoning. To address this, we propose MARE, an end-to-end framework that formulates longitudinal radiology report generation as a multimodal analogical reasoning task. Inspired by the Abduction–Mapping–Induction paradigm, MARE models latent relational structures underlying disease evolution by aligning lesion-level visual features across time and mapping them to the textual domain for temporally coherent and clinically meaningful report generation. To mitigate the spatial misalignment caused by patient positioning or imaging variation, we introduce an Adaptive Region Alignment (ARA) module for robust temporal correspondence. Additionally, we design Dual Evolution Consistency (DEC) losses to regularize analogical reasoning by enforcing temporal coherence in both visual and textual evolution paths. Extensive experiments on the Longitudinal-MIMIC dataset demonstrate that MARE significantly outperforms state-of-the-art baselines across both natural language generation and clinical effectiveness metrics, highlighting the value of structured analogical reasoning for disease evolution-aware report generation. Qingqing Gao, Tengfei Liu 0005, Xiaodan Zhang 0003, Zhongfan Sun, Boyue Wang |
AAAI | 4 |
| 2026 | MoEA-Net: Modality-Incremental Expert Aggregation Network for Retinal Prognostic PredictionabstractAutomated analysis of temporal changes in multimodal retinal images is critical for the prognostic assessment of ophthalmic diseases. Unlike traditional single-timepoint diagnosis, tracking longitudinal changes across multiple imaging modalities introduces significant data bias challenges: (1) Imbalanced modality samples compromise the integration of knowledge within minority modalities; (2) Heterogeneous visual patterns across modalities undermine the perception of disease-relevant biomarkers. To tackle these issues, we propose a Modality-Incremental Expert Aggregation Network (MoEA-Net), which unifies the inter-modal integration and intra-modal perception for enhanced retinal prognostic prediction. Specifically, we employ the large language model (LLM) with incremental LoRA layers for specific modalities to effectively integrate knowledge from imbalanced data. Besides, we introduce a Spatiotemporal-aware Expert (SAE) module to better perceive both the anatomical structures and longitudinal changes within modalities. By progressively combining the SAE module with incremental LoRA, MoEA-Net supports continual knowledge accumulation and improves accurate reasoning. Experimental results show that MoEA-Net achieves state-of-the-art performance on subretinal fluid change and visual recovery classification tasks, validating its effectiveness. Xiaodan Zhang 0003, Yanzhao Shi, Chengxin Zheng, Wanyu Zhang |
AAAI | 2 |
| 2025 | MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report GenerationabstractThe automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans involve extensive medical entities, such as diverse anatomy regions and lesions, exhibiting highly inconsistent spatial patterns in 3D volumetric space. This leads to biased learning of medical entities in existing methods, resulting in repetitiveness and inaccuracy in generated reports. To this end, we propose a Medical Entity-balanced Prompting Network (MEPNet), which harnesses the large language model (LLM) to fairly interpret various entities for accurate brain CT report generation. By introducing the visual embedding and the learning status of medical entities as enriched clues, our method prompts the LLM to balance the learning of diverse entities, thereby enhancing reports with comprehensive findings. First, to extract visual embedding of entities, we propose Knowledge-driven Joint Attention to explore and distill entity patterns using both explicit and implicit medical knowledge. Then, a Learning Status Scorer is designed to evaluate the learning of entity visual embeddings, resulting in unique learning status for individual entities. Finally, these entity visual embeddings and status are elaborately integrated into multi-modal prompts, to guide the text generation of LLM. This process allows LLM to self-adapt the learning process for biased-fitted entities, thereby covering detailed findings in generated reports. We conduct experiments on two brain CT report generation benchmarks, showing the effectiveness in clinical accuracy and text coherence. Xiaodan Zhang 0003, Yanzhao Shi, Junzhong Ji, Chengxin Zheng, Liangqiong Qu |
AAAI | 1 |
| 2025 | A New Federated Learning Framework Against Gradient Inversion AttacksabstractFederated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety of privacy-preserving methods have been integrated into FL to thwart such attacks, such as Secure Multi-party Computing (SMC), Homomorphic Encryption (HE), and Differential Privacy (DP). Despite their ability to protect data privacy, these approaches inherently involve substantial privacy-utility trade-offs. By revisiting the key to privacy exposure in FL under GIA, which lies in the frequent sharing of model gradients that contain private data, we take a new perspective by designing a novel privacy preserve FL framework that effectively ``breaks the direct connection'' between the shared parameters and the local private data to defend against GIA. Specifically, we propose a Hypernetwork Federated Learning (HyperFL) framework that utilizes hypernetworks to generate the parameters of the local model and only the hypernetwork parameters are uploaded to the server for aggregation. Theoretical analyses demonstrate the convergence rate of the proposed HyperFL, while extensive experimental results show the privacy-preserving capability and comparable performance of HyperFL. Pengxin Guo 0001, Shuang Zeng, Xiaodan Zhang 0003, Weihong Ren, Yuyin Zhou, Liangqiong Qu |
AAAI | 4 |
| 2025 | Scale-aware Multi-head Attention with Explainability for Image Captioning
Yuanzhen Guo, Xiaodan Zhang 0003, Aozhe Jia, Boyue Wang |
PRCV (3) | 2 |
| 2025 | Intra- and Inter-Head Orthogonal Attention for Image CaptioningabstractMulti-head attention (MA), which allows the model to jointly attend to crucial information from diverse representation subspaces through its heads, has yielded remarkable achievement in image captioning. However, there is no explicit mechanism to ensure MA attends to appropriate positions in diverse subspaces, resulting in overfocused attention for each head and redundancy between heads. In this paper, we propose a novel Intra- and Inter-Head Orthogonal Attention (I2OA) to efficiently improve MA in image captioning by introducing a concise orthogonal regularization to heads. Specifically, Intra-Head Orthogonal Attention enhances the attention learning of MA by introducing orthogonal constraint to each head, which decentralizes the object-centric attention to more comprehensive content-aware attention. Inter-Head Orthogonal Attention reduces the heads redundancy by applying orthogonal constraint between heads, which enlarges the diversity of representation subspaces and improves the representation ability for MA. Moreover, the proposed I2OA is flexible to combine with various multi-head attention based image captioning methods and improve the performances without increasing model complexity and parameters. Experiments on the MS COCO dataset demonstrate the effectiveness of the proposed model. Xiaodan Zhang 0003, Aozhe Jia, Junzhong Ji, Liangqiong Qu, Qixiang Ye |
IEEE Trans. Image Process. | 1 |
| 2024 | GHCL: Gaussian heuristic curriculum learning for Brain CT report generation
Qingya Shen, Yanzhao Shi, Xiaodan Zhang 0003, Junzhong Ji |
Multim. Syst. | 3 |
| 2024 | Prior tissue knowledge-driven contrastive learning for brain CT report generation
Yanzhao Shi, Junzhong Ji, Xiaodan Zhang 0003 |
Multim. Syst. | 3 |
| 2023 | Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report GenerationabstractThe automatic Brain CT reports generation can improve the efficiency and accuracy of diagnosing cranial diseases.However, current methods are limited by 1) coarse-grained supervision: the training data in image-text format lacks detailed supervision for recognizing subtle abnormalities, and 2) coupled cross-modal alignment: visual-textual alignment may be inevitably coupled in a coarse-grained manner, resulting in tangled feature representation for report generation.In this paper, we propose a novel Pathological Graph-driven Cross-modal Alignment (PGCA) model for accurate and robust Brain CT report generation.Our approach effectively decouples the cross-modal alignment by constructing a Pathological Graph to learn finegrained visual cues and align them with textual words.This graph comprises heterogeneous nodes representing essential pathological attributes (i.e., tissue and lesion) connected by intra-and inter-attribute edges with prior domain knowledge.Through carefully designed graph embedding and updating modules, our model refines the visual features of subtle tissues and lesions and aligns them with textual words using contrastive learning.Extensive experimental results confirm the viability of our method.We believe that our PGCA model holds the potential to greatly enhance the automatic generation of Brain CT reports and ultimately contribute to improved cranial disease diagnosis. Yanzhao Shi, Junzhong Ji, Xiaodan Zhang 0003, Liangqiong Qu |
EMNLP | 3 |
| 2023 | Generic Attention-model Explainability by Weighted Relevance AccumulationabstractAttention-based Transformer models have achieved remarkable progress in multi-modal tasks, such as visual question answering. The explainability of attention-based methods has recently attracted wide interest as it can explain the inner changes of attention tokens by accumulating relevancy across attention layers. Current methods simply update relevancy by equally accumulating the token relevancy before and after the attention processes. However, the importance of token values is usually different during relevance accumulation.In this paper, we propose a weighted relevancy strategy, which takes the importance of token values into consideration, to reduce distortion when equally accumulating relevance. To evaluate our method, we propose a unified CLIP-based two-stage model, named CLIPmapper, to process Vision-and-Language tasks through CLIP encoder and a following mapper. CLIPmapper consists of self-attention, cross-attention, single-modality, and cross-modality attention, thus it is more suitable for evaluating our generic explainability method. Extensive perturbation tests on visual question answering and image captioning tasks validate that our explainability method outperforms existing methods. Yiming Huang 0002, Aozhe Jia, Xiaodan Zhang 0003, Jiawei Zhang 0002 |
MMAsia | 3 |
| 2023 | Multi-scale Superpixel based Hierarchical Attention model for brain CT classification
Xiao Song 0003, Xiaodan Zhang 0003, Junzhong Ji |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | A Survey on Brain Effective Connectivity Network LearningabstractHuman brain effective connectivity characterizes the causal effects of neural activities among different brain regions. Studies of brain effective connectivity networks (ECNs) for different populations contribute significantly to the understanding of the pathological mechanism associated with neuropsychiatric diseases and facilitate finding new brain network imaging markers for the early diagnosis and evaluation for the treatment of cerebral diseases. A deeper understanding of brain ECNs also greatly promotes brain-inspired artificial intelligence (AI) research in the context of brain-like neural networks and machine learning. Thus, how to picture and grasp deeper features of brain ECNs from functional magnetic resonance imaging (fMRI) data is currently an important and active research area of the human brain connectome. In this survey, we first show some typical applications and analyze existing challenging problems in learning brain ECNs from fMRI data. Second, we give a taxonomy of ECN learning methods from the perspective of computational science and describe some representative methods in each category. Third, we summarize commonly used evaluation metrics and conduct a performance comparison of several typical algorithms both on simulated and real datasets. Finally, we present the prospects and references for researchers engaged in learning ECNs. Junzhong Ji, Aixiao Zou, Jinduo Liu 0001, Cuicui Yang, Xiaodan Zhang 0003, Yongduan Song 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Cross-modal Contrastive Attention Model for Medical Report GenerationabstractMedical report automatic generation has gained increasing interest recently as a way to help radiologists write reports more efficiently. However, this image-to-text task is rather challenging due to the typical data biases: 1) Normal physiological structures dominate the images, with only tiny abnormalities; 2) Normal descriptions accordingly dominate the reports. Existing methods have attempted to solve these problems, but they neglect to exploit useful information from similar historical cases. In this paper, we propose a novel Cross-modal Contrastive Attention (CMCA) model to capture both visual and semantic information from similar cases, with mainly two modules: a Visual Contrastive Attention Module for refining the unique abnormal regions compared to the retrieved case images; a Cross-modal Attention Module for matching the positive semantic information from the case reports. Extensive experiments on two widely-used benchmarks, IU X-Ray and MIMIC-CXR, demonstrate that the proposed model outperforms the state-of-the-art methods on almost all metrics. Further analyses also validate that our proposed model is able to improve the reports with more accurate abnormal findings and richer descriptions. Xiao Song 0003, Xiaodan Zhang 0003, Junzhong Ji, Pengxu Wei |
COLING | 2 |
| 2022 | Sparse data augmentation based on encoderforest for brain network classification
Junzhong Ji, Zihan Wang 0003, Xiaodan Zhang 0003, Junwei Li 0008 |
Appl. Intell. | 3 |
| 2022 | Relation constraint self-attention for image captioning
Junzhong Ji, Mingzhan Wang, Xiaodan Zhang 0003, Minglong Lei, Liangqiong Qu |
Neurocomputing | 3 |
| 2021 | Weakly Guided Hierarchical Encoder-Decoder Network for Brain CT Report GenerationabstractReport-writing for Brain Computed Tomography (CT) imaging is a routine procedure for diagnosing cerebrovascular diseases, while it is time-consuming and tedious for radiologists especially in highly populated areas. Automatic report generation has the potential to alleviate radiologists’ workload and reduce the diagnose error. Currently, the development of image captioning and medical image processing has driven great achievements in medical report generation. However, there is no report generation study for the Brain CT imaging and this task faces the following challenges: First, Brain CT lesions are disperse in 3-D space, with more morphological instability. Second, the Brain CT reports are long paragraphs with similar medical term. These challenges increase the difficulty o f lesions recognition and report generation for Brain CT imaging. To cope with these challenges, we propose a weakly guided hierarchical encoder-decoder network for lesions learning and Brain CT report generation. Specifically, we propose a weakly guided attention model (WGAM) in encoder to capture the most important areas and scans gradually under the weak guidance of possible lesions areas. In addition, we propose a keywords-driven interactive recurrent network (KIRN) in decoder to generate paragraphs under the weak guidance of possible lesions keywords. Experiments on our Brain CT dataset demonstrate the effectiveness of the proposed method. Sisi Yang, Junzhong Ji, Xiaodan Zhang 0003 |
BIBM | 3 |
| 2021 | Fast scene labeling via structural inference
Huaidong Zhang, Chu Han, Xiaodan Zhang 0003, Yong Du 0003, Xuemiao Xu, Guoqiang Han 0002, Harry Qin, Shengfeng He |
Neurocomputing | 3 |
| 2021 | Divergent-convergent attention for image captioning
Junzhong Ji, Zhuoran Du, Xiaodan Zhang 0003 |
Pattern Recognit. | 3 |
| 2021 | Convolutional kernels with an element-wise weighting mechanism for identifying abnormal brain connectivity patterns
Junzhong Ji, Xinying Xing, Junwei Li 0008, Xiaodan Zhang 0003 |
Pattern Recognit. | 5 |
| 2020 | Image captioning via semantic element embedding
Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Rynson W. H. Lau, Jianbin Jiao, Qixiang Ye |
Neurocomputing | 1 |
| 2020 | Spatio-Temporal Memory Attention for Image CaptioningabstractVisual attention has been successfully applied in image captioning to selectively incorporate the most relevant areas to the language generation procedure. However, the attention in current image captioning methods is only guided by the hidden state of language model, e.g. LSTM (Long-Short Term Memory), indirectly and implicitly, and thus the attended areas are weakly relevant at different time steps. Besides the spatial relationship of attention areas, the temporal relationship in attention is crucial for image captioning according to the attention transmission mechanism of human vision. In this paper, we propose a new spatio-temporal memory attention (STMA) model to learn the spatio-temporal relationship in attention for image captioning. The STMA introduces the memory mechanism to the attention model through a tailored LSTM, where the new cell is used to memorize and propagate the attention information, and the output gate is used to generate attention weights. The attention in STMA transmits with memory adaptively and dependently, which builds strong temporal connections of attentions and learns the spatio-temporal relationship of attended areas simultaneously. Besides, the proposed STMA is flexible to combine with attention-based image captioning frameworks. Experiments on MS COCO dataset demonstrate the superiority of the proposed STMA model in exploring the spatio-temporal relationship in attention and improving the current attention-based image captioning. Junzhong Ji, Xiaodan Zhang 0003, Boyue Wang, Xinhang Song |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Forest with Cross-shaped Window Scanning Mechanism to Extract Topological FeaturesabstractDeep neural networks have been successfully applied to the classification of brain networks. However, the high-dimensional and small-scale properties of the brain network data limit their extensive applications. To solve this problem, this paper proposes a new deep forest framework with cross-shaped window scanning mechanism (DF-CWSM) to extract topological features for the classification of brain networks. The cross-shaped window scanning mechanism is designed to extract the node-level and the edge-level features respectively that have meaningful interpretations in terms of corresponding network topologies. Based on the classification framework, we firstly implement the feature transformation of brain networks by the multi-level topological feature extraction. Then a cascade forest structure is used to learn the hierarchical features layer by layer. And the results of the last level of cascade forests are integrated to make the final classification. We evaluated the proposed framework on the ABIDE I data set. Experimental results show that our proposed framework can not only achieve competitive classification performance but also accurately identify the abnormal brain regions associated with ASD. Junwei Li 0008, Junzhong Ji, Xiaodan Zhang 0003, Zihan Wang 0003 |
BIBM | 4 |
| 2017 | Delving into Salient Object Subitizing and DetectionabstractSubitizing (i.e., instant judgement on the number) and detection of salient objects are human inborn abilities. These two tasks influence each other in the human visual system. In this paper, we delve into the complementarity of these two tasks. We propose a multi-task deep neural network with weight prediction for salient object detection, where the parameters of an adaptive weight layer are dynamically determined by an auxiliary subitizing network. The numerical representation of salient objects is therefore embedded into the spatial representation. The proposed joint network can be trained end-to-end using backpropagation. Experiments show the proposed multi-task network outperforms existing multi-task architectures, and the auxiliary subitizing network provides strong guidance to salient object detection by reducing false positives and producing coherent saliency maps. Moreover, the proposed method is an unconstrained method able to handle images with/without salient objects. Finally, we show state-of-the-art performance on different salient object datasets. Shengfeng He, Jianbo Jiao, Xiaodan Zhang 0003, Guoqiang Han 0002, Rynson W. H. Lau |
ICCV | 3 |
| 2017 | Keyword-driven image captioning via Context-dependent Bilateral LSTMabstractImage captioning has recently received much attention. Existing approaches, however, are limited to describing images with simple contextual information, which typically generate one sentence to describe each image with only a single contextual emphasis. In this paper, we address this limitation from a user perspective with a novel approach. Given some keywords as additional inputs, the proposed method would generate various descriptions according to the provided guidance. Hence, descriptions with different focuses can be generated for the same image. Our method is based on a new Context-dependent Bilateral Long Short-Term Memory (CDB-LSTM) model to predict a keyword-driven sentence by considering the word dependence. The word dependence is explored externally with a bilateral pipeline, and internally with a unified and joint training process. Experiments on the MS COCO dataset demonstrate that the proposed approach not only significantly outperforms the baseline method but also shows good adaptation and consistency with various keywords. Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Pengxu Wei, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao, Rynson W. H. Lau |
ICME | 1 |
| 2015 | Rich Image Description Based on RegionsabstractAbstract Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In contrast to the previous image description methods that focus on describing the whole image, this paper presents a method of generating rich image descriptions from image regions. First, we detect regions with R-CNN (regions with convolutional neural network features) framework. We then utilize the RNN (recurrent neural networks) to generate sentences for image regions. Finally, we propose an optimization method to select one suitable region. The proposed model generates several sentence description of regions in an image, which has sufficient representative power of the whole image and contains more detailed information. Comparing to general image level description, generating more specific and accurate sentences on the different regions can satisfy more personal requirements for different people. Experimental evaluations validate the effectiveness of the proposed method. Xiaodan Zhang 0003, Xinhang Song, Xiong Lv, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao |
ACM Multimedia | 1 |