VLDB 2026 Research / reviewers in the wild / expert
Zhenxuan Zhang
dblp:215/3328
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0002-2904-2848ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationabstractAutomatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to reflect the clinical reliability of generated reports. Overlap-based methods overlook fine-grained details (e.g., location, severity), diagnostic metrics are constrained by fixed vocabularies. Some diagnostic metrics are limited by fixed vocabularies or templates, reducing their ability to capture diverse clinical expressions. LLM-based metrics lack interpretable reasoning, limiting trust in clinical settings. Therefore, we propose a Granular Explainable Multi-Agent Score (GEMA-Score) in this paper, which conducts both objective quantification and subjective evaluation through a large language model-based multi-agent workflow. Our GEMA-Score parses structured reports and employs stable calculations through interactive exchanges of information among agents to assess disease diagnosis, location, severity, and uncertainty. Additionally, an LLM-based scoring agent evaluates completeness, readability, and clinical terminology while providing explanatory feedback. Extensive experiments show that GEMA-Score achieves the highest correlation with human experts on public datasets (Kendall = 0.69 on ReXVal; 0.45 on RadEvalX), demonstrating improved clinical scoring reliability. Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Zhifan Gao, Dominic C. Marshall, Yingying Fang, Guang Yang 0006 |
AAAI | 1 |
| 2026 | Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical ScoresabstractCongren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Zhang Bo, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, KinHei Lee, Zhenxuan Zhang, Xiaobing Li, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Congren Dai, Krinos Li, Huichi Zhou, Shijie Liang, Zhang Bo, Enyang Liu, Hongran An, Haosen Zhang, Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Maosong Sun 0001 |
ACL (1) | 13 |
| 2026 | 3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising
Peiyuan Jing, Chun-Wun Cheng, Zhenxuan Zhang, Liutao Yang, Thiago Lima 0001, Klaus Strobel, Antoine Leimgruber, Angelica I. Avilés-Rivero, Guang Yang 0006, Javier A. Montoya-Zegarra |
ICPR (9) | 4 |
| 2026 | Reason like a radiologist: Chain-of-thought and reinforcement learning for verifiable report generationabstractRadiology report generation is critical for efficiency, but current models often lack the structured reasoning of experts and the ability to explicitly ground findings in anatomical evidence, which limits clinical trust and explainability. This paper introduces BoxMed-RL, a unified training framework to generate spatially verifiable and explainable chest X-ray reports. BoxMed-RL advances chest X-ray report generation through two integrated phases: (1) Pretraining Phase. BoxMed-RL learns radiologist-like reasoning through medical concept learning and enforces spatial grounding with reinforcement learning. (2) Downstream Adapter Phase. Pretrained weights are frozen while a lightweight adapter ensures fluency and clinical credibility. Experiments on two widely used public benchmarks (MIMIC-CXR and IU X-Ray) demonstrate that BoxMed-RL achieves an average 7 % improvement in both METEOR and ROUGE-L metrics compared to state-of-the-art methods. An average 5 % improvement in large language model-based metrics further underscores BoxMed-RL's robustness in generating high-quality reports. Related code and training templates are publicly available at https://github.com/ayanglab/BoxMed-RL. Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Huichi Zhou, Zhengqing Yuan, Zhifan Gao, Lei Zhu 0003, Giorgos Papanastasiou, Yingying Fang, Guang Yang 0006 |
Medical Image Anal. | 3 |
| 2026 | From noisy labels to intrinsic structure: A geometric-structural dual-guided framework for noise-robust medical image segmentation
Tao Wang 0085, Zhenxuan Zhang, Yuanbo Zhou, Xinlin Zhang, Yuanbin Chen, Tao Tan 0002, Guang Yang 0006, Tong Tong 0001 |
Medical Image Anal. | 2 |
| 2026 | From Coarse to Continuous: Progressive Refinement Implicit Neural Representation for Motion-Robust Anisotropic MRI ReconstructionabstractIn motion-robust magnetic resonance imaging (MRI), slice-to-volume reconstruction is critical for recovering anatomically consistent 3D brain volumes from 2D slices, especially under accelerated acquisitions or patient motion. However, this task remains challenging due to hierarchical structural disruptions. It includes local detail loss from k-space undersampling, global structural aliasing caused by motion, and volumetric anisotropy. Therefore, we propose a progressive refinement implicit neural representation (PR-INR) framework. Our PR-INR unifies motion correction, structural refinement, and volumetric synthesis within a geometry-aware coordinate space. Specifically, a motion-aware diffusion module is first employed to generate coarse volumetric reconstructions that suppress motion artifacts and preserve global anatomical structures. Then, we introduce an implicit detail restoration module that performs residual refinement by aligning spatial coordinates with visual features. It corrects local structures and enhances boundary precision. Further, a voxel continuous-aware representation module represents the image as a continuous function over 3D coordinates. It enables accurate inter-slice completion and high-frequency detail recovery. We evaluate PR-INR on five public MRI datasets under various motion conditions (3% and 5% displacement), undersampling rates (4x and 8x) and slice resolutions (scale = 5). Experimental results demonstrate that PR-INR outperforms state-of-the-art methods in both quantitative reconstruction metrics and visual quality. It further shows generalization and robustness across diverse unseen domains. Zhenxuan Zhang, Lipei Zhang, Yanqi Cheng, Zi Wang 0005, Fanwen Wang, Haosen Zhang, Yinzhe Wu 0001, Angelica I. Avilés-Rivero, Zhifan Gao, Guang Yang 0006, Peter J. Lally |
IEEE Trans. Image Process. | 1 |
| 2026 | Cyclic Self-Supervised Diffusion for Ultra Low-Field to High-Field MRI SynthesisabstractSynthesizing high-quality images from low-field MRI holds significant potential. Low-field MRI is cheaper, more accessible, and safer, but suffers from low resolution and poor signal-to-noise ratio. This synthesis process can reduce reliance on costly acquisitions and expand data availability. However, synthesizing high-field MRI still suffers from a clinical fidelity gap. There is a need to preserve anatomical fidelity, enhance fine-grained structural details, and bridge domain gaps in image contrast. To address these issues, we propose a cyclic self-supervised diffusion (CSS-Diff) framework for high-field MRI synthesis from real low-field MRI data. Our core idea is to reformulate diffusion-based synthesis under a cycle-consistent constraint. It enforces anatomical preservation throughout the generative process rather than just relying on paired pixel-level supervision. The CSS-Diff framework further incorporates two novel processes. The slice-wise gap perception network aligns inter-slice inconsistencies via contrastive learning. The local structure correction network enhances local feature restoration through self-reconstruction of masked and perturbed patches. Extensive experiments on cross-field synthesis tasks demonstrate the effectiveness of our method, achieving state-of-the-art performance (e.g., $31.80~\pm ~2.70$ dB in PSNR, $0.943~\pm ~0.102$ in SSIM, and $0.0864~\pm ~0.0689$ in LPIPS). Beyond pixel-wise fidelity, our method also preserves fine-grained anatomical structures compared with the original low-field MRI (e.g., left cerebral white matter error drops from 12.1% to 2.1%, cortex from 4.2% to 3.7%). To conclude, our CSS-Diff can synthesize images that are both quantitatively reliable and anatomically consistent. The code is available at: https://github.com/ayanglab/CSS-Diff. Zhenxuan Zhang, Peiyuan Jing, Zi Wang 0005, Ula Briski, Coraline Beitone, Yinzhe Wu 0001, Fanwen Wang, Liutao Yang, Zhifan Gao, Zhaolin Chen, Kh Tohidul Islam, Guang Yang 0006, Peter J. Lally |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Multiple token rearrangement Transformer network with explicit superpixel constraint for segmentation of echocardiography
Wanli Ding, Heye Zhang, Xiujian Liu, Zhenxuan Zhang, Shuxin Zhuang, Zhifan Gao, Lin Xu 0008 |
Medical Image Anal. | 4 |
| 2024 | Embedding Tasks Into the Latent Space: Cross-Space Consistency for Multi-Dimensional Analysis in EchocardiographyabstractMulti-dimensional analysis in echocardiography has attracted attention due to its potential for clinical indices quantification and computer-aided diagnosis. It can utilize various information to provide the estimation of multiple cardiac indices. However, it still has the challenge of inter-task conflict. This is owing to regional confusion, global abnormalities, and time-accumulated errors. Task mapping methods have the potential to address inter-task conflict. However, they may overlook the inherent differences between tasks, especially for multi-level tasks (e.g., pixel-level, image-level, and sequence-level tasks). This may lead to inappropriate local and spurious task constraints. We propose cross-space consistency (CSC) to overcome the challenge. The CSC embeds multi-level tasks to the same-level to reduce inherent task differences. This allows multi-level task features to be consistent in a unified latent space. The latent space extracts task-common features and constrains the distance in these features. This constrains the task weight region that satisfies multiple task conditions. Extensive experiments compare the CSC with fifteen state-of-the-art echocardiographic analysis methods on five datasets (10,908 patients). The result shows that the CSC can provide left ventricular (LV) segmentation, (DSC = 0.932), keypoint detection (MAE = 3.06mm), and keyframe identification (accuracy = 0.943). These results demonstrate that our method can provide a multi-dimensional analysis of cardiac function and is robust in large-scale datasets. Zhenxuan Zhang, Chengjin Yu, Heye Zhang, Zhifan Gao |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Hands-Free Assistive Manipulator Using Augmented Reality and Tongue Drive SystemabstractA human-in-the-loop system is proposed to enable hands-free collaborative manipulation for people with physical disabilities. Studies show that the cognitive burden of interfacing with a robotic assistant decreases with increased robot autonomy. Incorporating modern advances in perception with augmented reality, this paper describes a framework for obtaining high-level intents from the user to specify manipulation tasks for execution. Augmented reality glasses provide an egocentric perspective to the robot. The glasses also provide visual feedback to users on a virtual menu showing a summary of robot affordances. The system processes the vision input to interpret the users environment. A Tongue Drive System serves as the input modality for triggering task execution by the robotic arm. Several manipulation experiments are performed with comparison to Cartesian control. The outcomes are also compared to reported state-of-the-art approaches. The results demonstrate competitive performance with minimal user input requirements. Fu-Jen Chu, Ruinian Xu, Zhenxuan Zhang, Patricio A. Vela, Maysam Ghovanloo |
IROS | 3 |