VLDB 2026 Research / reviewers in the wild / expert
Anmol Kagrecha
dblp:242/8112
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-6287-5705ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 48% Question answering and dialogue systems · 37% Reinforcement learning · 16% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 33% Algorithms and data structures · 33% Mathematical optimization · 33% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
answer aggregation |
0.9 | 1 | 2025 | SkillAggregation: Reference-free LLM-Dependent Aggregation · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
0.9 | 1 | 2025 | SkillAggregation: Reference-free LLM-Dependent Aggregation · ACL (1) 2025 |
Machine learning › Reinforcement learning
bandit |
0.4 | 1 | 2019 | Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards · NeurIPS 2019 |
Algorithms and data structures › learning algorithms
best arm identification |
0.4 | 1 | 2019 | Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards · NeurIPS 2019 |
Mathematical optimization › risk measures
conditional value at risk |
0.4 | 1 | 2019 | Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards · NeurIPS 2019 |
Algorithmic game theory and mechanism design
multi-armed bandit |
0.4 | 1 | 2019 | Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewards · NeurIPS 2019 |
Natural language and speech › Language models and text generation
evaluation of language models |
0.3 | 1 | 2025 | SkillAggregation: Reference-free LLM-Dependent Aggregation · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
unsupervised weight learning · 0.9crowd aggregation · 0.9concentration inequalities · 0.8CVaR estimator · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SkillAggregation: Reference-free LLM-Dependent AggregationabstractLarge Language Models (LLMs) are increasingly used to assess NLP tasks due to their ability to generate human-like judgments.Single LLMs were used initially, however, recent work suggests using multiple LLMs as judges yields improved performance.An important step in exploiting multiple judgements is the combination stage, aggregation.Existing methods in NLP either assign equal weight to all LLM judgments or are designed for specific tasks such as hallucination detection.This work focuses on aggregating predictions from multiple systems where no reference labels are available.A new method called SkillAggregation is proposed, which learns to combine estimates from LLM judges without needing additional data or ground truth.It extends the Crowdlayer aggregation method, developed for image classification, to exploit the judge estimates during inference.The approach is compared to a range of standard aggregation methods on HaluEval-Dialogue, TruthfulQA and Chatbot Arena tasks.SkillAggregation outperforms Crowdlayer on all tasks, and yields the best performance over all approaches on the majority of tasks. 1 Guangzhi Sun, Anmol Kagrecha, P. P. Manakul, Philip C. Woodland, Mark J. F. Gales |
ACL (1) | 2 |
| 2023 | Constrained regret minimization for multi-criterion multi-armed bandits
Anmol Kagrecha, Jayakrishnan Nair 0001, Krishna P. Jagannathan |
Mach. Learn. | 1 |
| 2022 | Statistically Robust, Risk-Averse Best Arm Identification in Multi-Armed BanditsabstractTraditional multi-armed bandit (MAB) formulations usually make certain assumptions about the underlying arms’ distributions, such as bounds on the support or their tail behaviour. Moreover, such parametric information is usually ‘baked’ into the algorithms. In this paper, we show that specialized algorithms that exploit such parametric information are prone to inconsistent learning performance when the parameter is misspecified. Our key contributions are twofold: (i) We establish fundamental performance limits ofstatistically robustMAB algorithms under the fixed-budget pure exploration setting, and (ii) We propose two classes of algorithms that are asymptotically near-optimal. Additionally, we consider a risk-aware criterion for best arm identification, where the objective associated with each arm is a linear combination of the mean and the conditional value at risk (CVaR). Throughout, we make a very mild ‘bounded moment’ assumption, which lets us work with both light-tailed and heavy-tailed distributions within a unified framework. Anmol Kagrecha, Jayakrishnan Nair 0001, Krishna P. Jagannathan |
IEEE Trans. Inf. Theory | 1 |
| 2021 | Bandit algorithms: Letting go of logarithmic regret for statistical robustnessabstractWe study regret minimization in a stochastic multi-armed bandit setting, and establish a fundamental trade-off between the regret suffered under an algorithm, and its statistical robustness. Considering broad classes of underlying arms’ distributions, we show that bandit learning algorithms with logarithmic regret are always inconsistent and that consistent learning algorithms always suffer a super-logarithmic regret. This result highlights the inevitable statistical fragility of all ‘logarithmic regret’ bandit algorithms available in the literature - for instance, if a UCB algorithm designed for 1-subGaussian distributions is used in a subGaussian setting with a mismatched variance parameter, the learning performance could be inconsistent. Next, we show a positive result: statistically robust and consistent learning performance is attainable if we allow the regret to be slightly worse than logarithmic. Specifically, we propose three classes of distribution oblivious algorithms that achieve an asymptotic regret that is arbitrarily close to logarithmic. Kumar Ashutosh, Jayakrishnan Nair 0001, Anmol Kagrecha, Krishna P. Jagannathan |
AISTATS | 3 |
| 2019 | Distribution oblivious, risk-aware algorithms for multi-armed bandits with unbounded rewardsabstractClassical multi-armed bandit problems use the expected value of an arm as a metric to evaluate its goodness. However, the expected value is a risk-neutral metric. In many applications like finance, one is interested in balancing the expected return of an arm (or portfolio) with the risk associated with that return. In this paper, we consider the problem of selecting the arm that optimizes a linear combination of the expected reward and the associated Conditional Value at Risk (CVaR) in a fixed budget best-arm identification framework. We allow the reward distributions to be unbounded or even heavy-tailed. For this problem, our goal is to devise algorithms that are entirely distribution oblivious, i.e., the algorithm is not aware of any information on the reward distributions, including bounds on the moments/tails, or the suboptimality gaps across arms. In this paper, we provide a class of such algorithms with provable upper bounds on the probability of incorrect identification. In the process, we develop a novel estimator for the CVaR of unbounded (including heavy-tailed) random variables and prove a concentration inequality for the same, which could be of independent interest. We also compare the error bounds for our distribution oblivious algorithms with those corresponding to standard non-oblivious algorithms. Finally, numerical experiments reveal that our algorithms perform competitively when compared with non-oblivious algorithms, suggesting that distribution obliviousness can be realised in practice without incurring a significant loss of performance. Anmol Kagrecha, Jayakrishnan Nair 0001, Krishna P. Jagannathan |
NeurIPS | 1 |