Zhengxu Tang

dblp:386/3045 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 28% Deep learning architectures and training · 28% Generative modeling · 28%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
CCS: Controllable and Constrained Sampling with Diffusion Models via Initial Noise Perturbation · NeurIPS 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of Experts · ICLR 2025
Medical and health informatics › medical imaging
medical image analysis
0.912025
Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of Experts · ICLR 2025
Medical and health informatics › biomedical natural language processing
medical question answering
0.912025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025
Medical and health informatics › clinical diagnosis
multi-modal medical diagnosis
0.912025
Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of Experts · ICLR 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025

Methods — techniques the papers use, named apart from their topics

self-verification · 1.7retrieval-augmented generation · 1.7modality-specific experts · 1.7conditional mutual information loss · 1.7chain-of-thought · 1.7diffusion ODE sampling · 0.9controller algorithm · 0.9
YearPublicationVenuePosition
2025 Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
abstract
Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report summaries. However, LLMs have demonstrated significantly varied performance across different languages in natural language question-answering tasks, potentially exacerbating healthcare disparities in Low and Middle-Income Countries (LMICs). This study introduces the first multilingual ophthalmological question-answering benchmark with manually curated questions parallel across languages, allowing for direct cross-lingual comparisons. Our evaluation of 6 popular LLMs across 7 different languages reveals substantial bias across different languages, highlighting risks for clinical deployment of LLMs in LMICs. Existing debiasing methods such as Translation Chain-of-Thought or Retrieval-augmented generation (RAG) by themselves fall short of closing this performance gap, often failing to improve performance across all languages and lacking specificity for the medical domain. To address this issue, We propose CLARA (Cross-Lingual Reflective Agentic system), a novel inference time de-biasing method leveraging retrieval augmented generation and self-verification. Our approach not only improves performance across all languages but also significantly reduces the multilingual bias gap, facilitating equitable LLM application across the globe.
David S. Restrepo, Chenwei Wu 0006, Zhengxu Tang, Zitao Shuai, Thao Nguyen Minh Phan, Jun-En Ding, Cong-Tinh Dao, Jack Gallifant, Robyn Gayle Dychiao, Jose Carlo Artiaga, André Hiroshi Bando, Carolina Pelegrini Barbosa Gracitelli, Vincenz Ferrer, Leo A. Celi, Danielle S. Bitterman, Michael G. Morley, Luis Filipe Nakayama
AAAI3
2025 SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models
abstract
Text-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through multiple events. Existing T2V benchmarks primarily focus on visual quality metrics but fail to evaluate narrative coherence over extended sequences. To bridge this gap, we present SeqBench, a comprehensive benchmark for evaluating sequential narrative coherence in T2V generation. SeqBench includes a carefully designed dataset of 320 prompts spanning various narrative complexities, with 2,560 human-annotated videos generated from 8 state-of-the-art T2V models. Additionally, we design a Dynamic Temporal Graphs (DTG)-based automatic evaluation metric, which can efficiently capture long-range dependencies and temporal ordering while maintaining computational efficiency. Our DTG-based metric demonstrates a strong correlation with human annotations. Through systematic evaluation using SeqBench, we reveal critical limitations in current T2V models: failure to maintain consistent object states across multi-action sequences, physically implausible results in multi-object scenarios, and difficulties in preserving realistic timing and ordering relationships between sequential actions. SeqBench provides the first systematic framework for evaluating narrative coherence in T2V generation and offers concrete insights for improving sequential reasoning capabilities in future models. Please refer to https://videobench.github.io/SeqBench.github.io/ for more details.
Zhengxu Tang, Zizheng Wang, Luning Wang, Zitao Shuai, Siyu Qian, Yirui Wu, Haosong Rao, Chenwei Wu 0006
CBMI1
2025 Dynamic Modeling of Patients, Modalities and Tasks via Multi-modal Multi-task Mixture of Experts
abstract
Multi-modal multi-task learning holds significant promise in tackling complex diagnostic tasks and many significant medical imaging problems. It fulfills the needs in real-world diagnosis protocol to leverage information from different data sources and simultaneously perform mutually informative tasks. However, medical imaging domains introduce two key challenges: dynamic modality fusion and modality-task dependence. The quality and amount of task-related information from different modalities could vary significantly across patient samples, due to biological and demographic factors. Traditional fusion methods apply fixed combination strategies that fail to capture this dynamic relationship, potentially underutilizing modalities that carry stronger diagnostic signals for specific patients. Additionally, different clinical tasks may require dynamic feature selection and combination from various modalities, a phenomenon we term “modality-task dependence.” To address these issues, we propose M4oE, a novel Multi-modal Multi-task Mixture of Experts framework for precise Medical diagnosis. M4oE comprises Modality-Specific (MSoE) modules and a Modality-shared Modality-Task MoE (MToE) module. With collaboration from both modules, our model dynamically decomposes and learns distinct and shared information from different modalities and achieves dynamic fusion. MToE provides a joint probability model of modalities and tasks by using experts as a link and encourages experts to learn modality-task dependence via conditional mutual information loss. By doing so, M4oE offers sample and population-level interpretability of modality contributions. We evaluate M4oE on four public multi-modal medical benchmark datasets for solving two important medical diagnostic problems including breast cancer screening and retinal disease diagnosis. Results demonstrate our method's superiority over state-of-the-art methods under different metrics of classification and segmentation tasks like Accuracy, AUROC, AUPRC, and DICE.
Chenwei Wu 0006, Zitao Shuai, Zhengxu Tang, Luning Wang, Liyue Shen
ICLR3
2025 CCS: Controllable and Constrained Sampling with Diffusion Models via Initial Noise Perturbation
abstract
Diffusion models have emerged as powerful tools for generative tasks, producing high-quality outputs across diverse domains. However, how the generated data responds to the initial noise perturbation in diffusion models remains under-explored, hindering a deeper understanding of the controllability of the sampling process. In this work, we first observe an interesting phenomenon: the relationship between the change of generation outputs and the scale of initial noise perturbation is highly linear through the diffusion ODE sampling process. We then provide both theoretical and empirical analyses to justify this linearity property of the input–output (noise → generation data) relationship. Inspired by these insights, we propose a novel **C**ontrollable and **C**onstrained **S**ampling (CCS) method, along with a new controller algorithm for diffusion models, that enables precise control over both (1) the proximity of individual samples to a target image and (2) the alignment of the sample mean with the target, while preserving high sample quality. We conduct extensive experiments comparing our proposed sampling approach with other methods in terms of both sampling controllability and generated data quality. Results show that CCS achieves significantly more precise controllability while maintaining superior sample quality and diversity, enabling practical applications such as fine-grained and robust image editing. Code: [https://github.com/efzero/diffusioncontroller](https://github.com/efzero/diffusioncontroller)
Zecheng Zhang, Zhaoxu Luo, Jason Hu, Zhengxu Tang, Guanyang Wang, Liyue Shen
NeurIPS7
2025 Enhancing AI-based diabetic retinopathy screening in low- and middle-income countries with synthetic data
abstract
AI-based DR screening is promising in low- and middle-income countries (LMICs), where limited human resources constrain access to specialist-led programs. However, current systems often degrade under real-world image-quality variations, especially with portable devices that are vital for low- and middle-income countries. This study aims to develop Retsyn, a synthetic-data augmentation framework that improves screening robustness across devices and imaging conditions. RetSyn leverages advanced diffusion models to generate synthetic retinal images with diverse device and imaging quality characteristics. To address the challenges of (1) portable device data scarcity, (2) disease and quality distribution imbalance, and (3) varying image quality, RetSyn uses class and quality-conditioned diffusion for controllable synthesis, a group-balanced loss to increase coverage of minority (quality, disease) pairs, and a Direct Preference Optimization alignment step with a small paired smartphone–tabletop set. The synthesized images are then used to augment classifier training. The effectiveness of RetSyn-generated images was evaluated by training retinal diagnosis models on a combination of real and synthetic data. RetSyn yields consistent gains in-domain and out-of-domain. On low-quality tabletop images, F1 improves from 0.781 to 0.874 (binary) and 0.607 to 0.703 (three-class), while AUROC reaches 0.982 and 0.951, respectively. On out-of-domain portable images, RetSyn attains AUROC 0.813/F1 0.703 (binary) and AUROC 0.804/F1 0.609 (three-class), exceeding group-robustness baselines such as GroupDRO (binary: AUROC 0.786/F1 0.626; three-class: AUROC 0.789/F1 0.544). RetSyn presents an effective and scalable synthetic data framework that significantly enhances the robustness and generalizability of AI-based DR screening models in LMICs. By addressing the critical challenges posed by varying image quality and device characteristics, RetSyn facilitates more reliable deployment of AI diagnostics in underserved regions. Additionally, the release of the first publicly available paired smartphone-tabletop retinal image dataset will support further research into cross-device DR screening solutions. • RetSyn generates synthetic retinal images to enhance AI robustness across varying image qualities and devices. • Group-balanced training and preference optimization enable diverse medical image synthesis with minimal cross-device paired data. • Releases the first publicly available paired smartphone-tabletop retinal dataset. • Performance improved significantly on low-quality and portable device photographs. • Achieves first clinically acceptable results for smartphone-based DR screening in LMICs.
Zitao Shuai, Chenwei Wu 0006, Zhengxu Tang, David S. Restrepo, Michael G. Morley, Luis Filipe Nakayama
J. Biomed. Informatics3