EDBT 2026 Demo / reviewers in the wild / expert
Xiaoman Zhang
dblp:155/6481
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable Brain MRI Report Generation Anchored by Lesion TopographyabstractRadiologists face increasing workloads that make accurate and timely report generation both critical and challenging. This paper presents a novel system for grounded automatic brain MRI report generation, with contributions in three key areas: First, we release RadGenome-Brain MRI, a benchmark dataset featuring multi-modal scans, expert-annotated abnormality masks, and radiology reports with region-level grounding to support fine-grained, explainable report generation. Second, we propose AutoRG-Brain, the first brain MRI report generation framework that combines automatic anomaly segmentation with a visual prompting-based language model to produce structured, anatomically grounded findings. Third, we conduct extensive quantitative and expert evaluations across segmentation and reporting tasks, and demonstrate in real clinical settings that our system significantly enhances junior radiologists' ability to detect subtle abnormalities and compose high-quality reports, narrowing the gap with senior doctors. All code, models, and datasets will be publicly released to facilitate future research and development. Jiayu Lei, Xiaoman Zhang, Chaoyi Wu, Lisong Dai, Ya Zhang 0002, Yanyong Zhang, Yanfeng Wang 0001, Weidi Xie |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation ModelsabstractMedical vision-language models often struggle with generating accurate quantitative measurements in radiology reports, leading to hallucinations that undermine clinical reliability. We introduce FactCheXcker, a modular framework that de-hallucinates radiology report measurements by leveraging an improved query-code-update paradigm. Specifically, FactCheXcker employs specialized modules and the code generation capabilities of large language models to solve measurement queries generated based on the original report. After extracting measurable findings, the results are incorporated into an updated report. We evaluate FactCheXcker on endotracheal tube placement, which accounts for an average of 78% of report measurements, using the MIMIC-CXR dataset and 11 medical reportgeneration models. Our results show that FactCheXcker significantly reduces hallucinations, improves measurement precision, and maintains the quality of the original reports. Specifically, FactCheXcker improves the performance of all 11 models and achieves an average improvement of 135.0% in reducing measurement hallucinations measured by mean absolute error. Code is available at https://github.com/rajpurkarlab/FactCheXcker. Alice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim, Pranav Rajpurkar |
CVPR | 2 |
| 2025 | Research on Dynamic Properties Evolution Algorithm of Polymer Nanomaterials Based on Heterogeneous Computing Platform
Xiaoman Zhang, Tao Liu 0029, Han Qin, Ying Guo 0028, Jingshan Pan |
ICA3PP (6) | 1 |
| 2025 | FedMPS: A Robust Differential Privacy Federated Learning Based on Local Model Partition and Sparsification for Heterogeneous IIoT DataabstractIn the emerging Industrial Internet of Things (IIoT) applications, federated learning (FL) enables model training without the need to transmit raw data directly. Nevertheless, transmitting model parameters could still reveal private information. To further protect local model parameters, differential privacy combined with FL (DPFL) has been introduced. Nonetheless, adding noise in DPFL can severely impact model performance, especially in non-independent and identically distributed (non-iid) data scenarios typical of IIoT environments. It is necessary to carefully balance privacy preservation and utility. In this article, we propose a robust DPFL scheme leveraging local model partition and sparsification (namely, FedMPS) for heterogeneous IIoT scenarios. The local model is divided into a shared part, which is sparsified before adding noise to mitigate its impact, and a private part that remains on the client. We provide a theoretical analysis of the privacy guarantees. Extensive experiments on common datasets, including Fashion-MNIST, CIFAR-10, and CIFAR-100, demonstrate that the proposed approach achieves a better privacy-utility tradeoff, with a 10%–20% improvement compared to baseline methods, and performs well especially in non-iid scenarios. Danxin Wang, Chen Zhang 0027, Xiaoman Zhang, Ming Li 0042 |
IEEE Internet Things J. | 5 |
| 2025 | Is Meta-Learning Effective for Few-Shot Hyperspectral Image Classification?abstractRecently, there has been a surge of meta-learning-based approaches for the few-shot hyperspectral image classification (FSHSIC) task. Meta-learning leverages prior knowledge to teach a base-learner how to adapt quickly to a new few-shot task, which hinges on the consistency of the prior and new tasks to guarantee validity. However, hyperspectral image classification (HSIC) is an environment-dependent task, which means that the hyperspectral features of two objects in the same category can be essentially distinctive in different environments. Consequently, whether meta-learning is a feasible solution for FSHSIC is an imperative problem to investigate, notwithstanding the promising performance shown in previous meta-learning-based approaches. To this end, this work proposes a simple multilayer perceptron (MLP)-based model named SimHSIC for FSHSIC. SimHSIC utilizes only a few labeled samples from the target HSI to train the model rapidly. Surprisingly, SimHSIC outperforms existing meta-learning-based approaches, which are built upon complex three-dimensional convolutions or transformers and need heavy training processes, in prevailing public benchmarks. On the basis of extensive experiments, we conclude that the relatively better classification performances of meta-learning-based FSHSIC solutions are due mainly to the patching of each HSI pixel with the large surroundings instead of meta-learning. The code is released on https://github.com/OrigamiSL/SimHSIC. Li Shen 0009, Yangzhu Wang, Xiaoman Zhang, Huaxin Qiu 0003, Chang Nie, Wei Li 0095 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Knowledge-Enhanced Visual-Language Pretraining for Computational Pathology
Xiao Zhou 0007, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001 |
ECCV (52) | 2 |
| 2024 | RaTEScore: A Metric for Radiology Report GenerationabstractThis paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models.RaTEScore emphasizes crucial medical entities, such as diagnostic outcomes and anatomical details.Moreover, it is robust against medical synonyms and sensitive to negation expressions.Technically, we developed a comprehensive medical NER dataset, RaTE-NER, and trained an NER model specifically for this purpose.This model enables the decomposition of complex radiological reports into constituent medical entities.The metric itself is derived by comparing the similarity of entity embeddings, obtained from a language model, based on their types and relevance to clinical significance.Our evaluations demonstrate that RaTEScore aligns more closely with human preference than existing metrics, validated both on established public benchmarks and our newly proposed RaTE-Eval benchmark. Weike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
EMNLP | 3 |
| 2024 | TKMBR: Temporal Knowledge Graph-based Multi-Behavior RecommendationabstractStriving to enhance predictive performance by leveraging auxiliary behaviors, multi-behavior recommendation models have emerged in in many different fields. These models aim to address the diversity and effectiveness of interactive behaviors. While some methods have shown promising effects, they still exhibit certain limitations, such as overlooking dynamic nature of user interactions. In this paper, we present TKMBR, a temporal knowledge graph-based framework for multi-behavior recommendation. TKMBR incorporates a temporal knowledge graph to capture the temporal dynamics of user behaviors, which allows for the identification of underlying temporal patterns and the capturing of evolving user preferences over time. To augment the understanding of user preferences, heterogeneous signals are integrated and an item-side information knowledge graph is constructed based on various user-item interactions. Moreover, contrastive learning tasks are employed to alleviate the issue of data sparsity. Evaluation on three datasets using HR and NDCG shows TKMBR’s effectiveness in improving recommendation quality. Xiaoman Zhang, Xuehua Bi, Guanglei Yu, Ruyi Cao |
IJCNN | 1 |
| 2024 | PMC-LLaMA: toward building open-source language models for medicineabstractOBJECTIVE: Recently, large language models (LLMs) have showcased remarkable capabilities in natural language understanding. While demonstrating proficiency in everyday conversations and question-answering (QA) situations, these models frequently struggle in domains that require precision, such as medical applications, due to their lack of domain-specific knowledge. In this article, we describe the procedure for building a powerful, open-source language model specifically designed for medicine applications, termed as PMC-LLaMA. MATERIALS AND METHODS: We adapt a general-purpose LLM toward the medical domain, involving data-centric knowledge injection through the integration of 4.8M biomedical academic papers and 30K medical textbooks, as well as comprehensive domain-specific instruction fine-tuning, encompassing medical QA, rationale for reasoning, and conversational dialogues with 202M tokens. RESULTS: While evaluating various public medical QA benchmarks and manual rating, our lightweight PMC-LLaMA, which consists of only 13B parameters, exhibits superior performance, even surpassing ChatGPT. All models, codes, and datasets for instruction tuning will be released to the research community. DISCUSSION: Our contributions are 3-fold: (1) we build up an open-source LLM toward the medical domain. We believe the proposed PMC-LLaMA model can promote further development of foundation models in medicine, serving as a medical trainable basic generative language backbone; (2) we conduct thorough ablation studies to demonstrate the effectiveness of each proposed component, demonstrating how different training data and model scales affect medical LLMs; (3) we contribute a large-scale, comprehensive dataset for instruction tuning. CONCLUSION: In this article, we systematically investigate the process of building up an open-source medical-specific LLM, PMC-LLaMA. Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisabstractIn this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel triplet extraction module to extract the medical-related information, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel triplet encoding module with entity translation by querying a knowledge base, to exploit the rich domain knowledge in medical field, and implicitly build relationships between medical entities in the language embedding space; Third, we propose to use a Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level, enabling the ability for medical diagnosis; Fourth, we conduct thorough experiments to validate the effectiveness of our architecture, and benchmark on numerous public benchmarks e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding. Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
ICCV | 2 |
| 2023 | Generative Gradient Inversion via Over-Parameterized Networks in Federated LearningabstractFederated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While prior research has predominantly focused on low-resolution images and small batch sizes, this study highlights the feasibility of reconstructing complex images with high resolutions and large batch sizes. The success of the proposed method is contingent on constructing an over-parameterized convolutional network, so that images are generated before fitting to the gradient matching requirement. Practical experiments demonstrate that the proposed algorithm achieves high-fidelity image recovery, surpassing state-of-the-art competitors that commonly fail in more intricate scenarios. Consequently, our study shows that local participants in a federated learning system are vulnerable to potential data leakage issues. Source code is available at https://github.com/czhang024/CI-Net. Chi Zhang 0123, Xiaoman Zhang, Ekanut Sotthiwat, Yanyu Xu 0001, Ping Liu 0004, Liangli Zhen, Yong Liu 0026 |
ICCV | 2 |
| 2023 | PMC-CLIP: Contrastive Language-Image Pre-training Using Biomedical Documents
Weixiong Lin, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
MICCAI (8) | 3 |
| 2023 | Research on remote sensing image carbon emission monitoring based on deep learning
Shaoqing Zhou, Xiaoman Zhang, Shiwei Chu |
Signal Process. | 2 |
| 2023 | Self-Supervised Tumor Segmentation With Sim2Real AdaptationabstractThis paper targets on self-supervised tumor segmentation. We make the following contributions: (i) we take inspiration from the observation that tumors are often characterised independently of their contexts, we propose a novel proxy task "layer-decomposition", that closely matches the goal of the downstream task, and design a scalable pipeline for generating synthetic tumor data for pre-training; (ii) we propose a two-stage Sim2Real training regime for unsupervised tumor segmentation, where we first pre-train a model with simulated tumors, and then adopt a self-training strategy for downstream data adaptation; (iii) when evaluating on different tumor segmentation benchmarks, e.g. BraTS2018 for brain tumor segmentation and LiTS2017 for liver tumor segmentation, our approach achieves state-of-the-art segmentation performance under the unsupervised setting. While transferring the model for tumor segmentation under a low-annotation regime, the proposed approach also outperforms all existing self-supervised approaches; (iv) we conduct extensive ablation studies to analyse the critical components in data simulation, and validate the necessity of different proxy tasks. We demonstrate that, with sufficient texture randomization in simulation, model trained on synthetic data can effortlessly generalise to datasets with real tumors. Xiaoman Zhang, Weidi Xie, Chaoqin Huang, Ya Zhang 0002, Xin Chen 0033, Qi Tian 0001, Yanfeng Wang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Wavelet J-Net: A Frequency Perspective on Convolutional Neural NetworksabstractIt is well acknowledged in image processing domain that the information can be decomposed into different frequency parts and each part has its own merits. However, existing neural networks always ignore the distinctions and straightforwardly feed all the information into neural networks together, treating them equally. In this paper, we propose a novel neural networks framework named J-Net that decomposes images into different frequency bands and then processes them sequentially. Concretely, the images have been decomposed by wavelet transformation and then the wavelet coefficients are fed into neural networks gradually in different depth according to their decomposition levels. An attention module is utilized to facilitate the fusion of neural network features and injected information, yielding significant performance gain. Furthermore, we show how does the information with different frequency impact the accuracy of neural networks. Experiments show that 5.91%, 5.32% and 2.00% accuracy improvements on Caltech 101, Caltech256 and ImageNet, respectively. Linfeng Zhang 0001, Xiaoman Zhang, Chenglong Bao, Kaisheng Ma |
IJCNN | 2 |
| 2021 | SAR: Scale-Aware Restoration Learning for 3D Tumor Segmentation
Xiaoman Zhang, Shixiang Feng, Ya Zhang 0002, Yanfeng Wang 0001 |
MICCAI (2) | 1 |