VLDB 2026 Research / reviewers in the wild / expert
Gaurav Tomar
dblp:93/11202 · also Gaurav Singh Tomar
· DBLP profile ↗
10ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-5060-4705ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Question answering and dialogue systems · 84% Deep learning architectures and training · 9% Language models and text generation · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Query processing and optimization · 50% Information retrieval · 50% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 50% Web and mobile security · 50% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.6 | 1 | 2022 | Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence · EMNLP 2022 |
Query processing and optimization
query rewriting |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Information retrieval › machine learning for information retrieval
reinforcement learning for retrieval |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
knowledge-grounded dialogue |
0.5 | 1 | 2021 | Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features · ACL/IJCNLP (1) 2021 |
Web and mobile security › web security
API security |
0.4 | 1 | 2020 | Thieves on Sesame Street! Model Extraction of BERT-based APIs · ICLR 2020 |
Security and privacy of machine learning
model stealing |
0.4 | 1 | 2020 | Thieves on Sesame Street! Model Extraction of BERT-based APIs · ICLR 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
state tracking |
0.2 | 1 | 2022 | Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence · EMNLP 2022 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
faithfulness |
0.1 | 1 | 2021 | Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features · ACL/IJCNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
reward function · 1.1reinforcement learning · 1.1large language model · 0.6human evaluation · 0.6controllable text generation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Investigating Content Planning for Navigating Trade-offs in Knowledge-Grounded DialogueabstractKushal Chawla, Hannah Rashkin, Gaurav Singh Tomar, David Reitter. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kushal Chawla, Hannah Rashkin, Gaurav Tomar, David Reitter |
EACL (1) | 3 |
| 2023 | Measuring Attribution in Natural Language Generation ModelsabstractAbstract Large neural models have brought a new challenge to natural language generation (NLG): It has become imperative to ensure the safety and reliability of the output of models that generate freely. To this end, we present an evaluation framework, Attributable to Identified Sources (AIS), stipulating that NLG output pertaining to the external world is to be verified against an independent, provided source. We define AIS and a two-stage annotation pipeline for allowing annotators to evaluate model output according to annotation guidelines. We successfully validate this approach on generation datasets spanning three tasks (two conversational QA datasets, a summarization dataset, and a table-to-text dataset). We provide full annotation guidelines in the appendices and publicly release the annotated data at https://github.com/google-research-datasets/AIS. Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins 0001, Dipanjan Das 0001, Slav Petrov, Gaurav Tomar, Iulia Turc, David Reitter |
Comput. Linguistics | 8 |
| 2022 | Dungeons and Dragons as a Dialog Challenge for Artificial IntelligenceabstractAI researchers have posited Dungeons and Dragons (D&D) as a challenge problem to test systems on various language-related capabilities.In this paper, we frame D&D specifically as a dialogue system challenge, where the tasks are to both generate the next conversational turn in the game and predict the state of the game given the dialogue history.We create a gameplay dataset consisting of nearly 900 games, with a total of 7,000 players, 800,000 dialogue turns, 500,000 dice rolls, and 58 million words.We automatically annotate the data with partial state information about the game play.We train a large language model (LM) to generate the next game turn, conditioning it on different information.The LM can respond as a particular character or as the player who runs the game-i.e., the Dungeon Master (DM).It is trained to produce dialogue that is either in-character (roleplaying in the fictional world) or out-of-character (discussing rules or strategy).We perform a human evaluation to determine what factors make the generated output plausible and interesting.We further perform an automatic evaluation to determine how well the model can predict the game state given the history and examine how well tracking the game state improves its ability to produce plausible conversational output. Chris Callison-Burch, Gaurav Tomar, Lara J. Martin, Daphne Ippolito, Suma Bailis, David Reitter |
EMNLP | 2 |
| 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningabstractCompared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision. Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar |
EMNLP | 7 |
| 2021 | Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable FeaturesabstractHannah Rashkin, David Reitter, Gaurav Singh Tomar, Dipanjan Das. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hannah Rashkin, David Reitter, Gaurav Tomar, Dipanjan Das 0001 |
ACL/IJCNLP (1) | 3 |
| 2020 | Thieves on Sesame Street! Model Extraction of BERT-based APIs
Kalpesh Krishna, Gaurav Tomar, Ankur P. Parikh, Nicolas Papernot, Mohit Iyyer |
ICLR | 2 |
| 2016 | Expediting Support for Social Learning with Behavior Modeling
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic |
EDM | 2 |
| 2016 | Pipeline for expediting learning analytics and student support from data in social learningabstractAn important research problem in learning analytics is to expedite the cycle of data leading to the analysis of student progress and the improvement of student support. For this goal in the context of social learning, we propose a pipeline that includes data infrastructure, learning analytics, and intervention, along with computational models for individual components. Next, we describe an example of applying this pipeline to real data in a case study, whose goal is to investigate the positive effects that goal-setting students have on their peers, which suggests ways in which we might foster these social benefits through intervention. Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic |
LAK | 2 |
| 2015 | Positive Impact of Collaborative Chat Participation in an edX MOOC
Oliver Ferschke, Diyi Yang, Gaurav Tomar, Carolyn P. Rosé |
AIED | 3 |
| 2015 | Alleviating the Negative Effect of Up and Downvoting on Help Seeking in MOOC Discussion Forums
Iris K. Howley, Gaurav Tomar, Diyi Yang, Oliver Ferschke, Carolyn P. Rosé |
AIED | 2 |