VLDB 2026 Research / reviewers in the wild / expert
Wenao Ma
dblp:226/5504
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-9952-5266ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLM) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "Visual Prompts" (VP) like bounding boxes to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap raises uncertainty about whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and utilize them to solve problems. To address this limitation, we introduce VP-Bench, aiming to assess MLLMs’ capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models’ ability to perceive VPs in natural scenes, utilizing 100K visualized prompts spanning 8 shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 21 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL-2.5 and Qwen2.5-VL). In addition, we conduct a comprehensive analysis of the factors influencing VP understanding, such as attribute variations and model scale. VP-Bench establishes a new reference framework for studying MLLMs’ ability to comprehend and resolve grounded referring questions. Mingjie Xu, Jinpeng Chen 0003, Yuzhi Zhao, Jason Chun Lok Li, Zekang Du, Mengyang Wu, Kun Li 0015, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei |
AAAI | 11 |
| 2026 | RSTFA: Efficient Training-Free Human-Preference Alignment via Rejection Sampling for Text-to-Image Diffusion ModelsabstractGiven a text-to-image diffusion model pretrained on large-scale text-image pairs, can we align the model with human pReferences without further fine-tuning? In this paper, we analyze the effect of alignment tuning in diffusion models by comparing the diffusion denoising trajectory between base and aligned models. Our findings reveal that alignment tuning primarily affects superficial stylistic aspects during denoising, rather than fundamental content, suggesting superficial alignment behaviors. Based on this discovery, we introduce a novel, training-free alignment approach (RSTFA) that leverages rejection sampling at specific stylistic timesteps, ensuring human preference alignment without fine-tuning or heavy inference overhead. We provide a theoretical analysis and derive a bias bound for our rejection-sampling alignment scheme. Empirically, we show that RSTFA better preserves sample diversity than reinforcement-learning-based tuning methods. Extensive experiments on Pick-a-Pic, COCO, HPD V2, and PartiPrompts show that our method not only achieves superior alignment with human preferences compared to state-of-the-art methods, but also reduces computational demands, establishing efficient, human-centered diffusion model alignment. Hongzheng Yang, Jason Chun Lok Li, Kun Li 0015, Wenao Ma, Mingjie Xu, Yuzhi Zhao, Lai-Man Po |
IEEE Trans. Image Process. | 4 |
| 2025 | KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented GenerationabstractZiyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jason Chun Lok Li, Zhijian Hou, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen 0003, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong 0001 |
EMNLP | 11 |
| 2024 | 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
Shizhan Gong, Yuan Zhong 0003, Wenao Ma, Jinpeng Li 0004, Zhao Wang 0006, Jingyang Zhang, Pheng-Ann Heng, Qi Dou 0001 |
Medical Image Anal. | 3 |
| 2024 | Causal Effect Estimation on Imaging and Clinical Data for Treatment Decision Support of Aneurysmal Subarachnoid HemorrhageabstractAneurysmal subarachnoid hemorrhage is a medical emergency of brain that has high mortality and poor prognosis. Causal effect estimation of treatment strategies on patient outcomes is crucial for aneurysmal subarachnoid hemorrhage treatment decision-making. However, most existing studies on treatment decision-making support of this disease are unable to simultaneously compare the potential outcomes of different treatments for a patient. Furthermore, these studies fail to harmoniously integrate the imaging data with non-imaging clinical data, both of which are useful in clinical scenarios. In this paper, we estimate the causal effect of various treatments on patients with aneurysmal subarachnoid hemorrhage by integrating plain CT with non-imaging clinical data, which is represented using structured tabular data. Specifically, we first propose a novel scheme that uses multi-modality confounders distillation architecture to predict the treatment outcome and treatment assignment simultaneously. With these distilled confounder features, we design an imaging and non-imaging interaction representation learning strategy to use the complementary information extracted from different modalities to balance the feature distribution of different treatment groups. We have conducted extensive experiments using a clinical dataset of 656 subarachnoid hemorrhage cases, which was collected from the Hospital Authority Data Collaboration Laboratory in Hong Kong. Our method shows consistent improvements on the evaluation metrics of treatment effect estimation, achieving state-of-the-art results over strong competitors. Code is released at https://github.com/med-air/TOP-aSAH. Wenao Ma, Cheng Chen 0013, Yuqi Gong, Nga Yan Chan, Meirui Jiang, Calvin Hoi-Kwan Mak, Jill M. Abrigo, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Treatment Outcome Prediction for Intracerebral Hemorrhage via Generative Prognostic Model with Imaging and Tabular Data
Wenao Ma, Cheng Chen 0013, Jill M. Abrigo, Calvin Hoi-Kwan Mak, Yuqi Gong, Nga Yan Chan, Chu Han, Zaiyi Liu, Qi Dou 0001 |
MICCAI (5) | 1 |
| 2022 | Test-Time Adaptation with Calibration of Medical Image Classification Nets for Label Distribution Shift
Wenao Ma, Cheng Chen 0013, Harry Qin, Huimao Zhang, Qi Dou 0001 |
MICCAI (3) | 1 |
| 2022 | Hierarchical deep network with uncertainty-aware semi-supervised learning for vessel segmentation
Chenxin Li, Wenao Ma, Liyan Sun, Xinghao Ding, Yue Huang 0001, Guisheng Wang, Yizhou Yu |
Neural Comput. Appl. | 2 |
| 2021 | Consistent Posterior Distributions Under Vessel-Mixing: A Regularization For Cross-Domain Retinal Artery/Vein ClassificationabstractRetinal artery/vein (A/V) classification is a critical technique for diagnosing diabetes and cardiovascular diseases. Although deep learning based methods achieve impressive results in A/V classification, the performance usually degrades when directly apply the models that trained on one dataset to another set, due to the domain shift, e.g., caused by the variations in imaging protocols. In this paper, we propose a novel method to improve cross-domain generalization for pixel-wise retinal A/V classification. That is, vessel-mixing based consistency regularization, which regularizes the models to give consistent posterior distributions for vessel-mixing samples. The proposed method achieves the state-of-the-art performance on extensive experiments for cross-domain A/V classification, which is even close to the performance of fully supervised learning on target domain in some cases. Chenxin Li, Zhehan Liang, Wenao Ma, Yue Huang 0001, Xinghao Ding |
ICIP | 4 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 21 |
| 2020 | A 3D Spatially Weighted Network for Segmentation of Brain Tissue From MRIabstractThe segmentation of brain tissue in MRI is valuable for extracting brain structure to aid diagnosis, treatment and tracking the progression of different neurologic diseases. Medical image data are volumetric and some neural network models for medical image segmentation have addressed this using a 3D convolutional architecture. However, this volumetric spatial information has not been fully exploited to enhance the representative ability of deep networks, and these networks have not fully addressed the practical issues facing the analysis of multimodal MRI data. In this paper, we propose a spatially-weighted 3D network (SW-3D-UNet) for brain tissue segmentation of single-modality MRI, and extend it using multimodality MRI data. We validate our model on the MRBrainS13 and MALC12 datasets. This unpublished model ranked first on the leaderboard of the MRBrainS13 Challenge. Liyan Sun, Wenao Ma, Xinghao Ding, Yue Huang 0001, Dong Liang 0001, John W. Paisley |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Multi-task Neural Networks with Spatial Activation for Retinal Vessel Segmentation and Artery/Vein Classification
Wenao Ma, Kai Ma 0002, Jiexiang Wang, Xinghao Ding, Yefeng Zheng 0001 |
MICCAI (1) | 1 |
| 2018 | Bindctnet: A Simple Binary Dct Network for Image ClassificationabstractConvolution neural networks play an important role in the image classification tasks. However, it is time consuming to train the network and the cost of memory resources is usually high. In this paper, a simple and effective network named BinDCTNet is presented by using the binary discrete cosine transform(BinDCT) to extract the feature-maps and a hyper-parameter to reduce dimension of the extracted feature. The proposed network has extremely low computing complexity and there is almost no parameters needed to be stored. Experiments are carried out on the hand written digit dataset MNIST and the vehicle logo VLOGO dataset. The results show that the proposed network achieves the state-of-the-art accuracy with fast speed and low memory cost, which makes it applicable on mobile and embedded devices. Xiangrui Xing, Wenao Ma, Yue Huang 0001, Delu Zeng, Xinghao Ding |
ICASSP | 3 |