Santhoshi Ravichandran

dblp:412/7094 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 50% Language models and text generation · 33% Efficient and distributed learning · 17%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Machine learning › Efficient and distributed learning › distillation
teacher-student distillation
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents › web agents
web agent training
0.912025
How to Train Your LLM Web Agent: A Statistical Diagnosis · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

hyperparameter sensitivity analysis · 0.9bootstrapping · 0.9
YearPublicationVenuePosition
2025 How to Train Your LLM Web Agent: A Statistical Diagnosis
abstract
Large language model (LLM) agents for web interfaces have advanced rapidly, yet open-source systems still lag behind proprietary agents. Bridging this gap is key to enabling customizable, efficient, and privacy-preserving agents. Two challenges hinder progress: the reproducibility issues in RL and LLM agent training, where results often depend on sensitive factors like seeds and decoding parameters, and the focus of prior work on single-step tasks, overlooking the complexities of web-based, multi-step decision-making. We address these gaps by providing a statistically driven study of training LLM agents for web tasks. Our two-stage pipeline combines imitation learning from a Llama 3.3 70B teacher with on-policy fine-tuning via Group Relative Policy Optimization (GRPO) on a Llama 3.1 8B student. Through 240 configuration sweeps and rigorous bootstrapping, we chart the first compute allocation curve for open-source LLM web agents. Our findings show that dedicating one-third of compute to teacher traces and the rest to RL improves MiniWoB++ success by 6 points and closes 60\% of the gap to GPT-4o on WorkArena, while cutting GPU costs by 45\%. We introduce a principled hyperparameter sensitivity analysis, offering actionable guidelines for robust and cost-effective agent training.
Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Angelard-Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia
NeurIPS2