VLDB 2026 Research / reviewers in the wild / expert
Leyang Yang
dblp:355/7749
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0006-5553-8037ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 47% Language models and text generation · 37% Multi-agent systems · 16% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 87% Interaction techniques and input · 13% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Human-AI interaction
GUI agent |
2.0 | 2 | 2026 | ProBench: Benchmarking GUI Agents with Accurate Process Information · AAAI 2026 History-Aware Reasoning for GUI Agents · AAAI 2026 |
Computer vision › Vision and language › vision-language model › multimodal large language model
GUI agent |
0.9 | 1 | 2025 | PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025 |
Computer vision › Vision and language › multimodal understanding
GUI understanding |
0.9 | 1 | 2025 | MP-GUI: Modality Perception with MLLMs for GUI Understanding · CVPR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | MP-GUI: Modality Perception with MLLMs for GUI Understanding · CVPR 2025 |
Natural language and speech › Language models and text generation › LLM agents
multimodal large language model agent |
0.9 | 1 | 2025 | PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2026 | History-Aware Reasoning for GUI Agents · AAAI 2026 |
Interaction techniques and input
mobile interaction |
0.3 | 1 | 2026 | ProBench: Benchmarking GUI Agents with Accurate Process Information · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.0multimodal model · 1.0task decomposition · 0.9retrieval-augmented generation · 0.9page graph · 0.9modality perception · 0.9fusion gate · 0.9automatic data collection · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | History-Aware Reasoning for GUI Agents
Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
AAAI | 2 |
| 2026 | ProBench: Benchmarking GUI Agents with Accurate Process InformationabstractWith the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and real-world devices, has received widespread attention from the community. Contemporary benchmarks aim to evaluate the comprehensive capabilities of GUI agents in GUI operation tasks, generally determining task completion solely by inspecting the final screen state. However, GUI operation tasks consist of multiple chained steps while not all critical information is presented in the final few pages. Although a few research has begun to incorporate intermediate steps into evaluation, accurately and automatically capturing this process information still remains an open challenge. To address this weakness, we introduce ProBench, a comprehensive mobile benchmark with over 200 challenging GUI tasks covering widely-used scenarios. Remaining the traditional State-related Task evaluation, we extend our dataset to include Process-related Task and design a specialized evaluation method. A newly introduced Process Provider automatically supplies accurate process information, enabling presice assessment of agent's performance. Our evaluation of advanced GUI agents reveals significant limitations for real-world GUI scenarios. These shortcomings are prevalent across diverse models, including both large-scale generalist models and smaller, GUI-specific models. A detailed error analysis further exposes several universal problems, outlining concrete directions for future improvements. Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
AAAI | 1 |
| 2025 | MP-GUI: Modality Perception with MLLMs for GUI UnderstandingabstractGraphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUIs comprise artificially designed graphical elements arranged to convey specific semantic meanings. Current multi-modal large language models (MLLMs) already proficient in processing graphical and textual components suffer from hurdles in GUI understanding due to the lack of explicit spatial structure modeling. Moreover, obtaining high-quality spatial structure data is challenging due to privacy issues and noisy environments. To address these challenges, we present MP-GUI, a specially designed MLLM for GUI understanding. MP-GUI features three precisely specialized perceivers to extract graphical, textual, and spatial modalities from the screen as GUI-tailored visual clues, with spatial structure refinement strategy and adaptively combined via a fusion gate to meet the specific preferences of different GUI understanding tasks. To cope with the scarcity of training data, we also introduce a pipeline for automatically data collecting. Extensive experiments demonstrate that MP-GUI achieves impressive results on various GUI understanding tasks with limited data. Our codes and datasets are publicly available at https://github.com/BigTaige/MP-GUI. Weizhi Chen, Leyang Yang, Sheng Zhou 0004, Shengchu Zhao, Hanbei Zhan, Jiongchao Jin, Liangcheng Li, Zirui Shao, Jiajun Bu |
CVPR | 3 |
| 2025 | PG-Agent: An Agent Powered by Page GraphabstractGraphical User Interface (GUI) agents possess significant commercial and social value, and GUI agents powered by advanced multimodal large language models (MLLMs) have demonstrated remarkable potential. Currently, existing GUI agents usually utilize sequential episodes of multi-step operations across pages as the prior GUI knowledge, which fails to capture the complex transition relationship between pages, making it challenging for the agents to deeply perceive the GUI environment and generalize to new scenarios. Therefore, we design an automated pipeline to transform the sequential episodes into page graphs, which explicitly model the graph structure of the pages that are naturally connected by actions. To fully utilize the page graphs, we further introduce Retrieval-Augmented Generation (RAG) technology to effectively retrieve reliable perception guidelines of GUI from them, and a tailored multi-agent framework PG-Agent with task decomposition strategy is proposed to be injected with the guidelines so that it can generalize to unseen scenarios. Extensive experiments on various benchmarks demonstrate the effectiveness of PG-Agent, even with limited episodes for page graph construction. Our codes will be publicly available at https://github.com/chenwz-123/PG-Agent. Weizhi Chen, Leyang Yang, Sheng Zhou 0004, Xiaoxuan Tang, Jiajun Bu, Yong Li 0004, Wei Jiang 0041 |
ACM Multimedia | 3 |
| 2023 | A Detail Geometry Learning Network for High-Fidelity Face Reconstruction
Kehua Ma, Xitie Zhang, Suping Wu, Leyang Yang, Zhixiang Yuan |
ICANN (2) | 5 |
| 2023 | Unsupervised Shape Enhancement and Factorization Machine Network for 3D Face Reconstruction
Leyang Yang, Jianchang Gong, Xueming Wang, Xiangzheng Li, Kehua Ma |
ICANN (3) | 1 |
| 2023 | A Lightweight Grouped Low-rank Tensor Approximation Network for 3D Mesh Reconstruction From VideosabstractExisting methods for 3D mesh reconstruction from videos suffer from increasingly large parameter counts and model sizes due to encoders such as multi-hidden state recurrence. Therefore many models become complex and more difficult to be applied in practice. Based on this problem, we propose a lightweight grouped low-rank tensor approximation network for 3D mesh reconstruction from videos. Specifically, firstly we propose a generalized grouped low-rank tensor approximation algorithm, which decomposes the original high-rank tensor into multiple weighted groups with different low-rank tensors to maximize the approximation of the high-rank tensor. Then we also design a lightweight selection rearranging strategy to reduce feature redundancy and focus on local fragment features. Notably, our method can be flexibly plugged into other 3D reconstruction tasks. Experiments show that our method not only improves the performance but also reduces the parameters of the entire network by about 90% compared with the existing SOTA method. Our grouped low-rank tensor approximation method reduces the parameters by about 99.8% in a single GRU. We demonstrate the outstanding generalization of our method in other 3D reconstruction tasks, eg. 3D face reconstruction and multi-view stereo. Suping Wu, Leyang Yang |
ICME | 3 |
| 2023 | A Bi-directional Optimization Network for De-obscured 3D High-Fidelity Face Reconstruction
Xitie Zhang, Suping Wu, Zhixiang Yuan, Kehua Ma, Leyang Yang |
ICONIP (14) | 6 |