EDBT 2026 Demo / reviewers in the wild / expert
Shuanghe Zhu
dblp:15/4805
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 83% Vision and language · 17% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › exploration
adaptive exploration |
1.0 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning › exploration
exploration strategies |
1.0 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Machine learning › Reinforcement learning
policy optimization |
1.0 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Human-AI interaction
GUI agent |
1.0 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Human-AI interaction › GUI agent
GUI grounding |
1.0 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Computer vision › Vision and language › vision-language model › multimodal large language model
GUI agent |
0.3 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning with verifiable rewards · 2.0multi-answer generation · 2.0adaptive exploration reward · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy OptimizationabstractThe emergence of Multimodal Large Language Models (MLLMs) has propelled the development of autonomous agents that operate on Graphical User Interfaces (GUIs) using pure visual input. A fundamental challenge is robustly grounding natural language instructions. This requires a precise spatial alignment, which accurately locates the coordinates of each element, and, more critically, a correct semantic alignment, which matches the instructions to the functionally appropriate UI element. Although Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be effective at improving spatial alignment for these MLLMs, we find that inefficient exploration bottlenecks semantic alignment, which prevents models from learning difficult semantic associations. To address this exploration problem, we present Adaptive Exploration Policy Optimization (AEPO), a new policy optimization framework. AEPO employs a multi-answer generation strategy to enforce broader exploration, which is then guided by a theoretically grounded Adaptive Exploration Reward (AER) function derived from first principles of efficiency η=U/C. Our AEPO-trained models, InfiGUI-G1-3B and InfiGUI-G1-7B, establish new state-of-the-art results across multiple challenging GUI grounding benchmarks, achieving significant relative improvements of up to 9.0% against the naive RLVR baseline on benchmarks designed to test generalization and semantic understanding. Yuhang Liu 0005, Shuanghe Zhu, Congkai Xie, Xueyu Hu, Shengyu Zhang 0001, Hongxia Yang, Fei Wu 0001 |
AAAI | 3 |
| 2025 | MS-Bench: Evaluating LMMs in Ancient Manuscript Study through a Dunhuang Case StudyabstractAnalyzing ancient manuscripts has traditionally been a labor-intensive and time-consuming task for philologists. While recent advancements in LMMs have demonstrated their potential across diverse domains, their effectiveness in manuscript study remains underexplored. In this paper, we introduce MS-Bench, the first comprehensive benchmark co-developed with archaeologists, comprising 5,076 high-resolution images from 4th to 14th century and 9,982 expert-curated questions across nine sub-tasks aligned with archaeological workflows. Through four prompting strategies, we systematically evaluate 32 LMMs on their effectiveness, robustness, and cultural contextualization. Our analysis reveals scale-driven performance and reliability improvements, prompting strategies' impact on performance (CoT has two-sides effect, while visual retrieval-augmented prompts provide consistent boost), and task-specific preferences depending on LMM’s visual capabilities. Although current LMMs are not yet capable of replacing domain expertise, they demonstrate promising potential to accelerate manuscript research through future human–AI collaboration. Shuanghe Zhu, Haoxiang Wu, Hangqi Li, Shengyu Zhang 0001, Junchi Yan, Kun Kuang 0001, Huaiyong Dou, Fei Wu 0001 |
NeurIPS | 3 |