EDBT 2026 Demo / reviewers in the wild / expert
Guiyu Ma
dblp:387/5030
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0000-9910-5301ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
2 papers |
Human-AI interaction · 67% Interaction techniques and input · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interaction techniques and input
mobile interaction |
1.0 | 1 | 2026 | DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking · CHI 2026 |
Human-AI interaction › automation
mobile task automation |
0.8 | 1 | 2024 | VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning · UIST 2024 |
Human-AI interaction › computer agents
mobile agent |
0.3 | 1 | 2026 | DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking · CHI 2026 |
Program synthesis and code generation
programming by demonstration |
0.2 | 1 | 2024 | VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning · UIST 2024 |
Methods — techniques the papers use, named apart from their topics
vision-based UI understanding · 1.5large language model task planning · 1.5user study · 1.0screenshot-based synthesis · 1.0multi-LLM task decomposition · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information SeekingabstractInformation seeking on mobile devices is often fragmented, trapping users in repetitive cycles of context switching and data re-entry, which increases cognitive load and disrupts workflow. Existing mobile agents provide limited cross-source integration and are largely opaque, presenting progress as a linear feed with few opportunities to intervene, steer, or take control. We present DroidRetriever, a transparent, steerable system for cross-source mobile information seeking. It accepts voice or typed input and the multi-LLM system decomposes the task, navigates to target pages, takes screenshots, and synthesizes a concise report with citation-linked screenshots. We make the process transparent through a progress dashboard combining sub-task progress and real-time exploration maps for seamless takeover. DroidRetriever also pauses on detected privacy or high-risk screens and prompts intervention. Across 35 tasks over 24 apps, experiments and user studies demonstrate improvements in coverage, transparency, and reduced workload. We release our code at https://github.com/AkimotoAyako/DroidRetriever. Yiheng Bian, Yunpeng Song, Guiyu Ma, Rongrong Zhu, Zhongmin Cai |
CHI | 3 |
| 2024 | VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task PlanningabstractMobile task automation is an emerging field that leverages AI to streamline and optimize the execution of routine tasks on mobile devices, thereby enhancing efficiency and productivity. Traditional methods, such as Programming By Demonstration (PBD), are limited due to their dependence on predefined tasks and susceptibility to app updates. Recent advancements have utilized the view hierarchy to collect UI information and employed Large Language Models (LLM) to enhance task automation. However, view hierarchies have accessibility issues and face potential problems like missing object descriptions or misaligned structures. This paper introduces VisionTasker, a two-stage framework combining vision-based UI understanding and LLM task planning, for mobile task automation in a step-by-step manner. VisionTasker firstly converts a UI screenshot into natural language interpretations using a vision-based UI understanding approach, eliminating the need for view hierarchies. Secondly, it adopts a step-by-step task planning method, presenting one interface at a time to the LLM. The LLM then identifies relevant elements within the interface and determines the next action, enhancing accuracy and practicality. Extensive experiments show that VisionTasker outperforms previous methods, providing effective UI representations across four datasets. Additionally, in automating 147 real-world tasks on an Android smartphone, VisionTasker demonstrates advantages over humans in tasks where humans show unfamiliarity and shows significant improvements when integrated with the PBD mechanism. VisionTasker is open-source and available at https://github.com/AkimotoAyako/VisionTasker. Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, Zhongmin Cai |
UIST | 4 |