Divya Tyamagundlu

dblp:72/10119 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Multi-agent systems · 33% Transfer learning and domain adaptation · 29% Language models and text generation · 29%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
autonomous agents
0.912025
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents · ICLR 2025
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
On the Effects of Data Scale on UI Control Agents · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction following
0.812024
On the Effects of Data Scale on UI Control Agents · NeurIPS 2024
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.312025
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents · ICLR 2025
Empirical software engineering
developer studies
0.212024
On the Effects of Data Scale on UI Control Agents · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

fine-tuning · 2.3few-shot learning · 2.3task parameterization · 0.9robustness analysis · 0.9
YearPublicationVenuePosition
2025 AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
abstract
Autonomous agents that execute human tasks by controlling computers can enhance human productivity and application accessibility. However, progress in this field will be driven by realistic and reproducible benchmarks. We present AndroidWorld, a fully functional Android environment that provides reward signals for 116 programmatic tasks across 20 real-world Android apps. Unlike existing interactive environments, which provide a static test set, AndroidWorld dynamically constructs tasks that are parameterized and expressed in natural language in unlimited ways, thus enabling testing on a much larger and more realistic suite of tasks. To ensure reproducibility, each task includes dedicated initialization, success-checking, and tear-down logic, which modifies and inspects the device’s system state. We experiment with baseline agents to test AndroidWorld and provide initial results on the benchmark. Our best agent can complete 30.6% of AndroidWorld's tasks, leaving ample room for future work. Furthermore, we adapt a popular desktop web agent to work on Android, which we find to be less effective on mobile, suggesting future research is needed to achieve universal, cross-platform agents. Finally, we also conduct a robustness analysis, showing that task variations can significantly affect agent performance, demonstrating that without such testing, agent performance metrics may not fully reflect practical challenges. AndroidWorld and the experiments in this paper are available at https://github.com/google-research/android_world.
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William E. Bishop, Folawiyo Campbell-Ajala, Daniel Toyama, Robert James Berry, Divya Tyamagundlu, Timothy P. Lillicrap, Oriana Riva
ICLR13
2024 On the Effects of Data Scale on UI Control Agents
abstract
Autonomous agents that control user interfaces to accomplish human tasks are emerging. Leveraging LLMs to power such agents has been of special interest, but unless fine-tuned on human-collected task demonstrations, performance is still relatively low. In this work we study whether fine-tuning alone is a viable approach for building real-world UI control agents. To this end we collect and release a new dataset, AndroidControl, consisting of 15,283 demonstrations of everyday tasks with Android apps. Compared to existing datasets, each AndroidControl task instance includes both high and low-level human-generated instructions, allowing us to explore the level of task complexity an agent can handle. Moreover, AndroidControl is the most diverse computer control dataset to date, including 14,548 unique tasks over 833 Android apps, thus allowing us to conduct in-depth analysis of the model performance in and out of the domain of the training data. Using the dataset, we find that when tested in domain fine-tuned models outperform zero and few-shot baselines and scale in such a way that robust performance might feasibly be obtained simply by collecting more data. Out of domain, performance scales significantly more slowly and suggests that in particular for high-level tasks, fine-tuning on more data alone may be insufficient for achieving robust out-of-domain performance.
William E. Bishop, Alice Li, Christopher Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, Oriana Riva
NeurIPS6