VLDB 2026 Research / reviewers in the wild / expert
Chaoyi Wu
dblp:168/6353
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable Brain MRI Report Generation Anchored by Lesion TopographyabstractRadiologists face increasing workloads that make accurate and timely report generation both critical and challenging. This paper presents a novel system for grounded automatic brain MRI report generation, with contributions in three key areas: First, we release RadGenome-Brain MRI, a benchmark dataset featuring multi-modal scans, expert-annotated abnormality masks, and radiology reports with region-level grounding to support fine-grained, explainable report generation. Second, we propose AutoRG-Brain, the first brain MRI report generation framework that combines automatic anomaly segmentation with a visual prompting-based language model to produce structured, anatomically grounded findings. Third, we conduct extensive quantitative and expert evaluations across segmentation and reporting tasks, and demonstrate in real clinical settings that our system significantly enhances junior radiologists' ability to detect subtle abnormalities and compose high-quality reports, narrowing the gap with senior doctors. All code, models, and datasets will be publicly released to facilitate future research and development. Jiayu Lei, Xiaoman Zhang, Chaoyi Wu, Lisong Dai, Ya Zhang 0002, Yanyong Zhang, Yanfeng Wang 0001, Weidi Xie |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | RadIR: A Scalable Framework for Multi-grained Medical Image Retrieval via Radiology Report Mining
Chaoyi Wu, Xiao Zhou 0004, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
MICCAI (5) | 3 |
| 2024 | Grip-Reach-Touch-Repeat: A Refined Model of Grasp to Encompass One-Handed Interaction with Arbitrary Form Factor DevicesabstractWe extend grasp models to encompass one-handed interaction with arbitrary shaped touchscreen devices. Current models focus on how objects are stably held by external forces. However, with touchscreen devices, we postulate that users do a trade-off between holding securely and exploring interactively. To verify this, we first conducted a qualitative study which asked participants to grasp 3D printed objects while considering its different interactivity. Results of the study confirm our hypothesis and reveal obvious change in postures. To further verify this trade-off and design interactions, we developed a simulation software capable of computing the stability of a grasp and its reachability. We conducted the second study based on the observed predominant grasps to validate our software with a glove. Results also confirm a consistent trade-off between stability and reachability. We conclude by discussing how this research can help designing computational tools focusing on hand-held interactions with arbitrary shaped touchscreen devices. Kaixing Zhao, Chaoyi Wu, Tao Xu 0037, Liang He 0012, Marcos Serrano, Anne Roudaut |
CHI | 2 |
| 2024 | Knowledge-Enhanced Visual-Language Pretraining for Computational Pathology
Xiao Zhou 0007, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001 |
ECCV (52) | 3 |
| 2024 | RaTEScore: A Metric for Radiology Report GenerationabstractThis paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models.RaTEScore emphasizes crucial medical entities, such as diagnostic outcomes and anatomical details.Moreover, it is robust against medical synonyms and sensitive to negation expressions.Technically, we developed a comprehensive medical NER dataset, RaTE-NER, and trained an NER model specifically for this purpose.This model enables the decomposition of complex radiological reports into constituent medical entities.The metric itself is derived by comparing the similarity of entity embeddings, obtained from a language model, based on their types and relevance to clinical significance.Our evaluations demonstrate that RaTEScore aligns more closely with human preference than existing metrics, validated both on established public benchmarks and our newly proposed RaTE-Eval benchmark. Weike Zhao, Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
EMNLP | 2 |
| 2024 | PMC-LLaMA: toward building open-source language models for medicineabstractOBJECTIVE: Recently, large language models (LLMs) have showcased remarkable capabilities in natural language understanding. While demonstrating proficiency in everyday conversations and question-answering (QA) situations, these models frequently struggle in domains that require precision, such as medical applications, due to their lack of domain-specific knowledge. In this article, we describe the procedure for building a powerful, open-source language model specifically designed for medicine applications, termed as PMC-LLaMA. MATERIALS AND METHODS: We adapt a general-purpose LLM toward the medical domain, involving data-centric knowledge injection through the integration of 4.8M biomedical academic papers and 30K medical textbooks, as well as comprehensive domain-specific instruction fine-tuning, encompassing medical QA, rationale for reasoning, and conversational dialogues with 202M tokens. RESULTS: While evaluating various public medical QA benchmarks and manual rating, our lightweight PMC-LLaMA, which consists of only 13B parameters, exhibits superior performance, even surpassing ChatGPT. All models, codes, and datasets for instruction tuning will be released to the research community. DISCUSSION: Our contributions are 3-fold: (1) we build up an open-source LLM toward the medical domain. We believe the proposed PMC-LLaMA model can promote further development of foundation models in medicine, serving as a medical trainable basic generative language backbone; (2) we conduct thorough ablation studies to demonstrate the effectiveness of each proposed component, demonstrating how different training data and model scales affect medical LLMs; (3) we contribute a large-scale, comprehensive dataset for instruction tuning. CONCLUSION: In this article, we systematically investigate the process of building up an open-source medical-specific LLM, PMC-LLaMA. Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang 0002, Weidi Xie, Yanfeng Wang 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisabstractIn this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel triplet extraction module to extract the medical-related information, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel triplet encoding module with entity translation by querying a knowledge base, to exploit the rich domain knowledge in medical field, and implicitly build relationships between medical entities in the language embedding space; Third, we propose to use a Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level, enabling the ability for medical diagnosis; Fourth, we conduct thorough experiments to validate the effectiveness of our architecture, and benchmark on numerous public benchmarks e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding. Chaoyi Wu, Xiaoman Zhang, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
ICCV | 1 |
| 2023 | PMC-CLIP: Contrastive Language-Image Pre-training Using Biomedical Documents
Weixiong Lin, Xiaoman Zhang, Chaoyi Wu, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
MICCAI (8) | 4 |
| 2023 | SA-BiSeNet: Swap attention bilateral segmentation network for real-time inland waterways segmentationabstractAbstract The technology for autonomous navigation on inland waterways is worth investigating, and navigable water surface segmentation is a key part of this technology. Semantic segmentation methods based on deep learning are able to distinguish between water surface areas and non‐water surface areas. However, existing semantic segmentation methods cannot meet the requirements of the water surface segmentation task in terms of both segmentation precision and real‐time performance. In this study, a Swap Attention Bilateral Segmentation Network (SA‐BiSeNet) is proposed to improve segmentation performance while ensuring model inference speed by better fusing the two features of the dual‐branch down‐sampling network using the attention mechanism. Specifically, an innovative Swap Attention Module is designed to model the dependency between the features of the spatial detail branch and the features of the semantic branches, thus expanding the receptive fields of the spatial detail and semantic branches to each other's global contexts. This design can effectively fuse features and thus enhance feature representation. Experiments were conducted on the inland waterway dataset USVInland to verify the performance of SA‐BiSeNet in terms of segmentation precision and inference speed, and SA‐BiSeNet achieved 93.65% Mean IoU and maintained the same level of fps as the baseline. Wenbo Zhang 0003, Chaoyi Wu, Zhenshan Bao |
IET Image Process. | 2 |
| 2022 | Boundary-Enhanced Self-supervised Learning for Brain Structure Segmentation
Feng Chang, Chaoyi Wu, Yanfeng Wang 0001, Ya Zhang 0002, Xin Chen 0033, Qi Tian 0001 |
MICCAI (1) | 2 |
| 2022 | DPANet: Dual Pooling-aggregated Attention Network for fish segmentationabstractAbstract The sustainable development of marine fisheries depends on the accurate measurement of data on fish stocks. Semantic segmentation methods based on deep learning can be applied to automatically obtain segmentation masks of fish in images to obtain measurement data. However, general semantic segmentation methods cannot accurately segment fish objects in underwater images. In this study, a Dual Pooling‐aggregated Attention Network (DPANet) to adaptively capture long‐range dependencies through an efficient and computing‐friendly manner to enhance feature representation and improve segmentation performance is proposed. Specifically, a novel pooling‐aggregate position attention module and a pooling‐aggregate channel attention module are designed to aggregate contexts in the spatial dimension and channel dimension, respectively. These two modules adopt pooling operations along the channel dimension and along the spatial dimension to aggregate information, respectively, thus reducing computational costs. In these modules, attention maps are generated by four different paths and are aggregated into one. The authors conduct extensive experiments to validate the effectiveness of the DPANet and achieve new state‐of‐the‐art segmentation performance on the well‐known fish image dataset DeepFish as well as on the underwater image dataset SUIM, achieving a Mean IoU score of 91.08% and 85.39% respectively, while significantly reducing FLOPs of attention modules by about 93%. Wenbo Zhang 0003, Chaoyi Wu, Zhenshan Bao |
IET Comput. Vis. | 2 |
| 2014 | WIFI fingerprinting indoor localization system based on spatio-temporal (S-T) metricsabstractIndoor localization has greatly leveraged applications regarding to location based service (LBS), which witnessed ever-increasing impact on human life. Among the existing localization solutions, WIFI-based received signal strength index (RSSI) fingerprinting is widely used due to desirable features such as universal availability, privacy protection, and low deployment cost. However, to build a robust, accurate RSSI fingerprinting localization system regardless of application occasions confronts two challenges. The first challenge is to construct a fine-grained and up-to-date RSSI map with reasonable labor cost in the training phase, and the second challenge is to deploy effective algorithm in the localization phase. This article illustrates the design and deployment of our indoor localization system targeting at the above mentioned problems. The overall solution is based on five spatio-temporal (S-T) metrics, to improve localization accuracy. Localization performance is evaluated in three indoor scenes at different scales, which show good accuracy with a median error of 1-2m under office environment, and 3-4m accuracy with no less than 70% probability when the environment is extremely crowded and noisy. Julie Yixuan Zhu, Jialing Xu, Anny Xijia Zheng, Jiaju He, Chaoyi Wu, Victor O. K. Li |
IPIN | 5 |