VLDB 2026 Research / reviewers in the wild / expert
Hualei Zhu
dblp:280/3165
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 44% Efficient and distributed learning · 19% Planning, search and constraint satisfaction · 18% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization |
1.0 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
GUI automation |
1.0 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Natural language and speech › Language models and text generation
instruction tuning |
1.0 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
game-based evaluation |
0.9 | 1 | 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing |
0.9 | 1 | 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
0.9 | 1 | 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
compositional generalization |
0.5 | 1 | 2021 | Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization · AAAI 2021 |
Natural language and speech › Machine translation
neural machine translation |
0.5 | 1 | 2021 | Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization · AAAI 2021 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.5 | 1 | 2021 | Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization · AAAI 2021 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2026 | Importance-Aware Data Selection for Efficient LLM Instruction Tuning · AAAI 2026 |
Natural language and speech › Language models and text generation › LLM agents
multimodal large language model agent |
0.3 | 1 | 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
self-play · 1.0model instruction weakness value · 1.0importance scoring · 1.0group relative policy optimization · 1.0data distillation · 1.0reinforcement learning scenarios · 0.9interactive evaluation · 0.9semi-supervised learning · 0.5iterative back-translation · 0.5curriculum learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Importance-Aware Data Selection for Efficient LLM Instruction TuningabstractInstruction tuning plays a critical role in enhancing the performance and efficiency of Large Language Models (LLMs). Its success depends not only on the quality of the instruction data but also on the inherent capabilities of the LLM itself. Some studies suggest that even a small amount of high-quality data can achieve instruction fine-tuning results that are on par with, or even exceed, those from using a full-scale dataset. However, rather than focusing solely on calculating data quality scores to evaluate instruction data, there is a growing need to select high-quality data that maximally enhances the performance of instruction tuning for a given LLM. In this paper, we propose the Model Instruction Weakness Value (MIWV) as a novel metric to quantify the importance of instruction data in enhancing model's capabilities. The MIWV metric is derived from the discrepancies in the model’s responses when using In-Context Learning (ICL), helping identify the most beneficial data for enhancing instruction tuning performance. Our experimental results demonstrate that selecting only the top 1% of data based on MIWV can outperform training on the full dataset. Furthermore, this approach extends beyond existing research that focuses on data quality scoring for data selection, offering strong empirical evidence supporting the effectiveness of our proposed method. Tingyu Jiang, Yiyao Song, Hualei Zhu, Xiaohang Xu 0002, Kenjiro Taura, Hao Henry Wang |
AAAI | 5 |
| 2026 | Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI AgentsabstractGraphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current methodologies exhibit two fundamental limitations: (1) insufficient exploitation of cross-model synergies, and (2) over-reliance on synthetic data generation without sufficient utilization. To address these challenges, we propose Co-EPG, a self-iterative training framework for Co-Evolution of Planning and Grounding. Co-EPG establishes an iterative positive feedback loop: through this loop, the planning model explores superior strategies under grounding-based reward guidance via Group Relative Policy Optimization (GRPO), generating diverse data to optimize the grounding model. Concurrently, the optimized Grounding model provides more effective rewards for subsequent GRPO training of the planning model, fostering continuous improvement. Co-EPG thus enables iterative enhancement of agent capabilities through self-play optimization and training data distillation. On the Multimodal-Mind2Web and AndroidControl benchmarks, our framework outperforms existing state-of-the-art methods after just three iterations without requiring external data. The agent consistently improves with each iteration, demonstrating robust self-enhancement capabilities. This work establishes a novel training paradigm for GUI agents, shifting from isolated optimization to an integrated, self-driven co-evolution approach. Hualei Zhu, Tingyu Jiang, Xiaohang Xu 0002, Hao Henry Wang |
AAAI | 2 |
| 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning EvaluationabstractRecent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments. Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009 |
NeurIPS | 12 |
| 2021 | Revisiting Iterative Back-Translation from the Perspective of Compositional GeneralizationabstractHuman intelligence exhibits compositional generalization (i.e., the capacity to understand and produce unseen combinations of seen components), but current neural seq2seq models lack such ability. In this paper, we revisit iterative back-translation, a simple yet effective semi-supervised method, to investigate whether and how it can improve compositional generalization. In this work: (1) We first empirically show that iterative back-translation substantially improves the performance on compositional generalization benchmarks (CFQ and SCAN). (2) To understand why iterative back-translation is useful, we carefully examine the performance gains and find that iterative back-translation can increasingly correct errors in pseudo-parallel data. (3) To further encourage this mechanism, we propose curriculum iterative back-translation, which better improves the quality of pseudo-parallel data, thus further improving the performance. Yinuo Guo, Hualei Zhu, Zeqi Lin, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001 |
AAAI | 2 |