VLDB 2026 Research / reviewers in the wild / expert
Santhoshi Ravichandran
dblp:412/7094
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Language models and text generation · 33% Efficient and distributed learning · 17% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Natural language and speech › Language models and text generation › LLM agents › web agents
web agent training |
0.9 | 1 | 2025 | How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
hyperparameter sensitivity analysis · 0.9bootstrapping · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How to Train Your LLM Web Agent: A Statistical DiagnosisabstractLarge language model (LLM) agents for web interfaces have advanced rapidly, yet open-source systems still lag behind proprietary agents. Bridging this gap is key to enabling customizable, efficient, and privacy-preserving agents. Two challenges hinder progress: the reproducibility issues in RL and LLM agent training, where results often depend on sensitive factors like seeds and decoding parameters, and the focus of prior work on single-step tasks, overlooking the complexities of web-based, multi-step decision-making.
We address these gaps by providing a statistically driven study of training LLM agents for web tasks. Our two-stage pipeline combines imitation learning from a Llama 3.3 70B teacher with on-policy fine-tuning via Group Relative Policy Optimization (GRPO) on a Llama 3.1 8B student. Through 240 configuration sweeps and rigorous bootstrapping, we chart the first compute allocation curve for open-source LLM web agents. Our findings show that dedicating one-third of compute to teacher traces and the rest to RL improves MiniWoB++ success by 6 points and closes 60\% of the gap to GPT-4o on WorkArena, while cutting GPU costs by 45\%. We introduce a principled hyperparameter sensitivity analysis, offering actionable guidelines for robust and cost-effective agent training. Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Angelard-Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia |
NeurIPS | 2 |