VLDB 2026 Research / reviewers in the wild / expert
Zhaoyang Wang 0004
dblp:41/8315-4
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 38% Reinforcement learning · 16% Vision and language · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.7 | 2 | 2025 | Anyprefer: An Agentic Framework for Preference Data Synthesis · ICLR 2025 CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.1 | 2 | 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models · ICLR 2025 CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › LLM agents
web agents |
1.0 | 1 | 2026 | SynthAgent: Adapting Web Agents with Synthetic Supervision · ACL (1) 2026 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning
fairness and safety |
0.9 | 1 | 2025 | MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video Preference · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Synergistic Weak-Strong Collaboration by Aligning Preferences · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › collaborative learning
model collaboration |
0.9 | 1 | 2025 | Synergistic Weak-Strong Collaboration by Aligning Preferences · ACL (1) 2025 |
Machine learning › Generative modeling › synthetic data generation
preference data synthesis |
0.9 | 1 | 2025 | Anyprefer: An Agentic Framework for Preference Data Synthesis · ICLR 2025 |
Machine learning › Reinforcement learning
preference learning |
0.9 | 1 | 2025 | Anyprefer: An Agentic Framework for Preference Data Synthesis · ICLR 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Anyprefer: An Agentic Framework for Preference Data Synthesis · ICLR 2025 |
Natural language and speech › Language models and text generation › preference optimization
self-rewarding language model |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning › collaborative learning › model collaboration
small-large model collaboration |
0.9 | 1 | 2025 | Synergistic Weak-Strong Collaboration by Aligning Preferences · ACL (1) 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video Preference · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.9 | 1 | 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models · NeurIPS 2024 |
Medical and health informatics › medical imaging › medical image analysis
medical vision-language model |
0.8 | 1 | 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model evaluation
automatic evaluation |
0.3 | 1 | 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models · ICLR 2025 |
Machine learning › Trustworthy machine learning
uncertainty and calibration |
0.3 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic supervision · 1.0imitation learning · 1.0scoring model · 0.9preference tuning · 0.9human-annotated fine-tuning · 0.9direct preference optimization · 0.9cooperative two-player markov game · 0.9consistency regularization · 0.9collaborative feedback · 0.9LLM-as-a-judge · 0.9trustworthiness evaluation · 0.8benchmark construction · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynthAgent: Adapting Web Agents with Synthetic SupervisionabstractZhaoyang Wang, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao, Saravan Rajmohan, Huaxiu Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhaoyang Wang 0004, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao 0001, Saravan Rajmohan, Huaxiu Yao |
ACL (1) | 1 |
| 2025 | Synergistic Weak-Strong Collaboration by Aligning PreferencesabstractCurrent Large Language Models excel in general reasoning yet struggle with specialized tasks requiring proprietary or domain-specific knowledge. Fine-tuning large models for every niche application is often infeasible due to black-box constraints and high computational overhead. To address this, we propose a collaborative framework that pairs a specialized weak model with a general strong model. The weak model, tailored to specific domains, produces initial drafts and background information, while the strong model leverages its advanced reasoning to refine these drafts, extending LLMs’ capabilities to critical yet specialized tasks. To optimize this collaboration, we introduce a collaborative feedback to fine-tunes the weak model, which quantifies the influence of the weak model’s contributions in the collaboration procedure and establishes preference pairs to guide preference tuning of the weak model. We validate our framework through experiments on three domains. We find that the collaboration significantly outperforms each model alone by leveraging complementary strengths. Moreover, aligning the weak model with the collaborative preference further enhances overall performance. Yizhu Jiao, Xuchao Zhang, Zhaoyang Wang 0004, Yubo Ma, Zhun Deng, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Jiawei Han 0001, Huaxiu Yao |
ACL (1) | 3 |
| 2025 | CREAM: Consistency Regularized Self-Rewarding Language ModelsabstractRecent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations for preference data. These methods commonly utilize the same LLM to act as both the policy model (which generates responses) and the reward model (which scores and ranks those responses). The ranked responses are then used as preference pairs to train the LLM via direct alignment technologies (e.g. DPO). However, it is noteworthy that throughout this process, there is no guarantee of accuracy in the rewarding and ranking, which is critical for ensuring accurate rewards and high-quality preference data. Empirical results from relatively small LLMs (e.g., 7B parameters) also indicate that improvements from self-rewarding may diminish after several iterations in certain situations, which we hypothesize is due to accumulated bias in the reward system. This bias can lead to unreliable preference data for training the LLM. To address this issue, we first formulate and analyze the generalized iterative preference fine-tuning framework for self-rewarding language model. We then introduce the regularization to this generalized framework to mitigate the overconfident preference labeling in the self-rewarding process. Based on this theoretical insight, we propose a Consistency Regularized sElf-rewarding lAnguage Model (CREAM) that leverages the consistency of rewards across different iterations to regularize the self-rewarding training, helping the model to learn from more reliable preference data. With this explicit regularization, our empirical results demonstrate the superiority of CREAM in improving both reward consistency and alignment performance. The code is publicly available at https://github.com/Raibows/CREAM. Zhaoyang Wang 0004, Weilei He, Zhiyuan Liang, Xuchao Zhang, Chetan Bansal, Huaxiu Yao |
ICLR | 1 |
| 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsabstractInterleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. Peng Xia 0005, Siwei Han, Shi Qiu 0016, Yiyang Zhou, Zhaoyang Wang 0004, Zhaorun Chen, Chenhang Cui, Mingyu Ding, Huaxiu Yao |
ICLR | 5 |
| 2025 | Anyprefer: An Agentic Framework for Preference Data SynthesisabstractHigh-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its own preference data, but this can lead to inaccuracies since the reward model shares weights with the target model, thereby amplifying inherent biases. To address these issues, we propose Anyprefer, a framework designed to synthesize high-quality preference data for aligning the target model. Anyprefer frames the data synthesis process as a cooperative two-player Markov Game, where the target model and the judge model collaborate together. Here, a series of external tools are introduced to assist the judge model in accurately rewarding the target model’s responses, mitigating biases in the rewarding process. In addition, a feedback mechanism is introduced to optimize prompts for both models, enhancing collaboration and improving data quality.
The synthesized data is compiled into a new preference dataset, Anyprefer-V1, consisting of 58K high-quality preference pairs.
Extensive experiments show that Anyprefer significantly improves model alignment performance across four main applications, covering 21 datasets, achieving average improvements of 18.55% in five natural language generation datasets, 3.66% in nine vision-language understanding datasets, 30.05% in three medical image analysis datasets, and 16.00% in four visuo-motor control tasks. Yiyang Zhou, Zhaoyang Wang 0004, Tianle Wang 0009, Shangyu Xing, Peng Xia 0005, Bo Li 0026, Zijian Zhang 0010, Zhaorun Chen, Xuchao Zhang, Chetan Bansal, Mohit Bansal, Huaxiu Yao |
ICLR | 2 |
| 2025 | MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video PreferenceabstractRecent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content hallucination, safety concerns, and generation bias. To address these limitations, we introduce MJ-BENCH-VIDEO, a large-scale video preference benchmark designed to evaluate video generation across five critical aspects: Alignment, Safety, Fineness, Coherence & Consistency, and Bias & Fairness. This benchmark further incorporates 28 fine-grained criteria to provide a comprehensive evaluation of video preference. Building upon this dataset, we propose MJ-VIDEO, a Mixture-of-Experts (MoE)-based video reward model designed to deliver fine-grained reward. MJ-VIDEO can dynamically select relevant experts to accurately judge the preference based on the input text-video pair. This architecture enables more precise and adaptable preference judgments. Through extensive benchmarking on MJ-BENCH-VIDEO, we analyze the limitations of existing video reward models and demonstrate the superior performance of MJ-VIDEO in video preference assessment, achieving 17.58% and 15.87% improvements in overall and fine-grained preference judgments, respectively. Additionally, MJ-VIDEO is able to improve the alignment performance in video generation via preference fine-tuning. Haibo Tong, Zhaoyang Wang 0004, Zhaorun Chen, Haonian Ji, Shi Qiu 0016, Siwei Han, Kexin Geng, Zhongkai Xue, Yiyang Zhou, Peng Xia 0005, Mingyu Ding, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 2 |
| 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsabstractArtificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing significant risks for future model deployment. In this paper, we introduce CARES and aim to comprehensively evaluate the Trustworthiness of Med-LVLMs across the medical domain. We assess the trustworthiness of Med-LVLMs across five dimensions, including trustfulness, fairness, safety, privacy, and robustness. CARES comprises about 41K question-answer pairs in both closed and open-ended formats, covering 16 medical image modalities and 27 anatomical regions. Our analysis reveals that the models consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies and failing to maintain fairness across different demographic groups. Furthermore, they are vulnerable to attacks and demonstrate a lack of privacy awareness. We publicly release our benchmark and code in https://github.com/richard-peng-xia/CARES. Peng Xia 0005, Juanxi Tian, Yangrui Gong, Ruibo Hou, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Zhaoyang Wang 0004, Xiao Wang 0044, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li 0010, Jimeng Sun 0001, ZongYuan Ge, Gang Li 0001, James Zou 0001, Huaxiu Yao |
NeurIPS | 12 |