Zichen Ding 0002

dblp:245/0972-2 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-1436-0291ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Graph learning · 36% Language models and text generation · 25% Vision and language · 14%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
computer-use agent
1.012026
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents · ACL (1) 2026
Machine learning › Graph learning
graph neural network
0.912025
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning · KDD (1) 2025
Machine learning › Graph learning
graph prompt learning
0.912025
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning · KDD (1) 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
GUI agent
0.912025
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis · ACL (1) 2025
Machine learning › Graph learning › graph neural network
node classification
0.912025
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning · KDD (1) 2025
Knowledge, reasoning and agents › Multi-agent systems › multimodal agent
vision-language model agent
0.912025
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis · ACL (1) 2025
Human-AI interaction
GUI agent
0.912025
OS-ATLAS: Foundation Action Model for Generalist GUI Agents · ICLR 2025
Human-AI interaction › GUI agent
GUI grounding
0.912025
OS-ATLAS: Foundation Action Model for Generalist GUI Agents · ICLR 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.312025
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning · KDD (1) 2025
Machine learning › Graph learning
pre-trained graph model
0.312025
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning · KDD (1) 2025
Computer vision › Vision and language
vision-language model
0.312025
OS-ATLAS: Foundation Action Model for Generalist GUI Agents · ICLR 2025

Methods — techniques the papers use, named apart from their topics

hybrid validation · 2.0vision-language model training · 1.7cross-platform data synthesis · 1.7holistic framework · 1.0trajectory reward model · 0.9reverse task synthesis · 0.9reinforcement learning · 0.9combinatorial optimization · 0.9
YearPublicationVenuePosition
2026 OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
abstract
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong
ACL (1)9
2026 OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents
abstract
Bowen Yang, Kaiming Jin, Zhenyu Wu, Zhaoyang Liu, Qiushi Sun, Zehao Li, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang, Qingyun Li, Yu Qiao, Zun Wang, Zichen Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kaiming Jin, Zhaoyang Liu 0001, Qiushi Sun, JingJing Xie, Zhoumianze Liu, Fangzhi Xu, Kanzhi Cheng, Yian Wang 0003, Qingyun Li, Yu Qiao 0001, Zun Wang 0001, Zichen Ding 0002
ACL (1)15
2025 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
abstract
Graphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. Despite their utility in advancing digital automation, a critical bottleneck persists: collecting high-quality trajectory data for training. Common practices for collecting such data rely on human supervision or synthetic data generation through executing pre-defined tasks, which are either resource-intensive or unable to guarantee data quality. Moreover, these methods suffer from limited data diversity and significant gaps between synthetic data and real-world environments. To address these challenges, we propose OS-Genesis, a novel GUI data synthesis pipeline that reverses the conventional trajectory collection process. Instead of relying on pre-defined tasks, OS-Genesis enables agents first to perceive environments and perform step-wise interactions, then retrospectively derive high-quality tasks to enable trajectory-level exploration. A trajectory reward model is then employed to ensure the quality of the generated trajectories. We demonstrate that training GUI agents with OS-Genesis significantly improves their performance on highly challenging online benchmarks. In-depth analysis further validates OS-Genesis's efficiency and its superior data quality and diversity compared to existing synthesis methods. Our codes, data, and checkpoints are available at OS-Genesis Homepage.
Qiushi Sun, Kanzhi Cheng, Zichen Ding 0002, Chuanyang Jin, Yian Wang 0003, Fangzhi Xu, Chengyou Jia, Zhoumianze Liu, Ben Kao, Guohao Li 0001, Junxian He, Yu Qiao 0001, Zhiyong Wu 0003
ACL (1)3
2025 OS-ATLAS: Foundation Action Model for Generalist GUI Agents
abstract
Existing efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are often reluctant to use open-source VLMs due to their significant performance lag compared to their closed-source counterparts, particularly in GUI grounding and Out-Of-Distribution (OOD) scenarios. To facilitate future research in this area, we developed OS-Atlas—a foundational GUI action model that excels at GUI grounding and OOD agentic tasks through innovations in both data and modeling. We have invested significant engineering effort in developing an open-source toolkit for synthesizing GUI grounding data across multiple platforms, including Windows, Linux, MacOS, Android, and the web. Leveraging this toolkit, we are releasing the largest open-source cross-platform GUI grounding corpus to date, which contains over 13 million GUI elements. This dataset, combined with innovations in model training, provides a solid foundation for OS-Atlas to understand GUI screenshots and generalize to unseen interfaces. Through extensive evaluation across six benchmarks spanning three different platforms (mobile, desktop, and web), OS-Atlas demonstrates significant performance improvements over previous state-of-the-art models. Our evaluation also uncovers valuable insights into continuously improving and scaling the agentic capabilities of open-source VLMs.
Zhiyong Wu 0003, Fangzhi Xu, Yian Wang 0003, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding 0002, Paul Pu Liang, Yu Qiao 0001
ICLR8
2025 RELIEF: Reinforcement Learning Empowered Graph Feature Prompt Tuning
abstract
The advent of the "pre-train, prompt'' paradigm has recently extended its generalization ability and data efficiency to graph representation learning, following its achievements in Natural Language Processing (NLP). Initial graph prompt tuning approaches tailored specialized prompting functions for Graph Neural Network (GNN) models pre-trained with specific strategies, such as edge prediction, thus limiting their applicability. In contrast, another pioneering line of research has explored universal prompting via adding prompts to the input graph's feature space, thereby removing the reliance on specific pre-training strategies. However, the necessity to add feature prompts to all nodes remains an open question. Motivated by findings from prompt tuning research in the NLP domain, which suggest that highly capable pre-trained models need less conditioning signal to achieve desired behaviors, we advocate for strategically incorporating necessary and lightweight feature prompts to certain graph nodes to enhance downstream task performance. This introduces a combinatorial optimization problem, requiring a policy to decide 1) which nodes to prompt and 2) what specific feature prompts to attach. We then address the problem by framing the prompt incorporation process as a sequential decision-making problem and propose our method, RELIEF, which employs Reinforcement Learning (RL) to optimize it. At each step, the RL agent selects a node (discrete action) and determines the prompt content (continuous action), aiming to maximize cumulative performance gain. Extensive experiments on graph and node-level tasks with various pre-training strategies in few-shot scenarios demonstrate that our RELIEF outperforms fine-tuning and other prompt-based approaches in classification performance and data efficiency. The code is available at https://github.com/JasonZhujp/RELIEF.
Jiapeng Zhu 0002, Zichen Ding 0002, Jianxiang Yu 0001, Jiaqi Tan 0006, Xiang Li 0067, Weining Qian
KDD (1)2