VLDB 2026 Research / reviewers in the wild / expert
Stefania Raimondo
dblp:185/0518
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Language models and text generation · 33% Efficient and distributed learning · 17% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Natural language and speech › Language models and text generation › LLM agents › web agents
web agent training |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
hyperparameter sensitivity analysis · 0.9bootstrapping · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How to Train Your LLM Web Agent: A Statistical DiagnosisabstractLarge language model (LLM) agents for web interfaces have advanced rapidly, yet open-source systems still lag behind proprietary agents. Bridging this gap is key to enabling customizable, efficient, and privacy-preserving agents. Two challenges hinder progress: the reproducibility issues in RL and LLM agent training, where results often depend on sensitive factors like seeds and decoding parameters, and the focus of prior work on single-step tasks, overlooking the complexities of web-based, multi-step decision-making.
We address these gaps by providing a statistically driven study of training LLM agents for web tasks. Our two-stage pipeline combines imitation learning from a Llama 3.3 70B teacher with on-policy fine-tuning via Group Relative Policy Optimization (GRPO) on a Llama 3.1 8B student. Through 240 configuration sweeps and rigorous bootstrapping, we chart the first compute allocation curve for open-source LLM web agents. Our findings show that dedicating one-third of compute to teacher traces and the rest to RL improves MiniWoB++ success by 6 points and closes 60\% of the gap to GPT-4o on WorkArena, while cutting GPU costs by 45\%. We introduce a principled hyperparameter sensitivity analysis, offering actionable guidelines for robust and cost-effective agent training. Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Angelard-Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia |
NeurIPS | 10 |
| 2020 | A Conversational Robot for Older Adults with Alzheimer's DiseaseabstractAmid the rising cost of Alzheimer’s disease (AD), assistive health technologies can reduce care-giving burden by aiding in assessment, monitoring, and therapy. This article presents a pilot study testing the feasibility and effect of a conversational robot in a cognitive assessment task with older adults with AD. We examine the robot interactions through dialogue and miscommunication analysis, linguistic feature analysis, and the use of a qualitative analysis, in which we report key themes that were prevalent throughout the study. While conversations were typically better with human conversation partners (being longer, with greater engagement and less misunderstanding), we found that the robot was generally well liked by participants and that it was able to capture their interest in dialogue. Miscommunication due to issues of understanding and intelligibility did not seem to deter participants from their experience. Furthermore, in automatically extracting linguistic features, we examine how non-acoustic aspects of language change across participants with varying degrees of cognitive impairment, highlighting the robot’s potential as a monitoring tool. This pilot study is an exploration of how conversational robots can be used to support individuals with AD. Chloé Pou-Prom, Stefania Raimondo, Frank Rudzicz |
ACM Trans. Hum. Robot Interact. | 2 |