Gaurav Tomar

dblp:93/11202 · also Gaurav Singh Tomar · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-5060-4705ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 84% Deep learning architectures and training · 9% Language models and text generation · 8%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 50% Information retrieval · 50%
Network and information security
1 paper
Security and privacy of machine learning · 50% Web and mobile security · 50%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.612022
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.612022
Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence · EMNLP 2022
Query processing and optimization
query rewriting
0.612022
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022
Information retrieval › machine learning for information retrieval
reinforcement learning for retrieval
0.612022
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022
Natural language and speech › Question answering and dialogue systems
knowledge-grounded dialogue
0.512021
Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features · ACL/IJCNLP (1) 2021
Web and mobile security › web security
API security
0.412020
Thieves on Sesame Street! Model Extraction of BERT-based APIs · ICLR 2020
Security and privacy of machine learning
model stealing
0.412020
Thieves on Sesame Street! Model Extraction of BERT-based APIs · ICLR 2020
Machine learning › Deep learning architectures and training › sequence modeling
state tracking
0.212022
Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence · EMNLP 2022
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
faithfulness
0.112021
Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features · ACL/IJCNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

reward function · 1.1reinforcement learning · 1.1large language model · 0.6human evaluation · 0.6controllable text generation · 0.5
YearPublicationVenuePosition
2024 Investigating Content Planning for Navigating Trade-offs in Knowledge-Grounded Dialogue
abstract
Kushal Chawla, Hannah Rashkin, Gaurav Singh Tomar, David Reitter. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kushal Chawla, Hannah Rashkin, Gaurav Tomar, David Reitter
EACL (1)3
2023 Measuring Attribution in Natural Language Generation Models
abstract
Abstract Large neural models have brought a new challenge to natural language generation (NLG): It has become imperative to ensure the safety and reliability of the output of models that generate freely. To this end, we present an evaluation framework, Attributable to Identified Sources (AIS), stipulating that NLG output pertaining to the external world is to be verified against an independent, provided source. We define AIS and a two-stage annotation pipeline for allowing annotators to evaluate model output according to annotation guidelines. We successfully validate this approach on generation datasets spanning three tasks (two conversational QA datasets, a summarization dataset, and a table-to-text dataset). We provide full annotation guidelines in the appendices and publicly release the annotated data at https://github.com/google-research-datasets/AIS.
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins 0001, Dipanjan Das 0001, Slav Petrov, Gaurav Tomar, Iulia Turc, David Reitter
Comput. Linguistics8
2022 Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence
abstract
AI researchers have posited Dungeons and Dragons (D&D) as a challenge problem to test systems on various language-related capabilities.In this paper, we frame D&D specifically as a dialogue system challenge, where the tasks are to both generate the next conversational turn in the game and predict the state of the game given the dialogue history.We create a gameplay dataset consisting of nearly 900 games, with a total of 7,000 players, 800,000 dialogue turns, 500,000 dice rolls, and 58 million words.We automatically annotate the data with partial state information about the game play.We train a large language model (LM) to generate the next game turn, conditioning it on different information.The LM can respond as a particular character or as the player who runs the game-i.e., the Dungeon Master (DM).It is trained to produce dialogue that is either in-character (roleplaying in the fictional world) or out-of-character (discussing rules or strategy).We perform a human evaluation to determine what factors make the generated output plausible and interesting.We further perform an automatic evaluation to determine how well the model can predict the game state given the history and examine how well tracking the game state improves its ability to produce plausible conversational output.
Chris Callison-Burch, Gaurav Tomar, Lara J. Martin, Daphne Ippolito, Suma Bailis, David Reitter
EMNLP2
2022 CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning
abstract
Compared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision.
Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar
EMNLP7
2021 Increasing Faithfulness in Knowledge-Grounded Dialogue with Controllable Features
abstract
Hannah Rashkin, David Reitter, Gaurav Singh Tomar, Dipanjan Das. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hannah Rashkin, David Reitter, Gaurav Tomar, Dipanjan Das 0001
ACL/IJCNLP (1)3
2020 Thieves on Sesame Street! Model Extraction of BERT-based APIs
Kalpesh Krishna, Gaurav Tomar, Ankur P. Parikh, Nicolas Papernot, Mohit Iyyer
ICLR2
2016 Expediting Support for Social Learning with Behavior Modeling
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic
EDM2
2016 Pipeline for expediting learning analytics and student support from data in social learning
abstract
An important research problem in learning analytics is to expedite the cycle of data leading to the analysis of student progress and the improvement of student support. For this goal in the context of social learning, we propose a pipeline that includes data infrastructure, learning analytics, and intervention, along with computational models for individual components. Next, we describe an example of applying this pipeline to real data in a case study, whose goal is to investigate the positive effects that goal-setting students have on their peers, which suggests ways in which we might foster these social benefits through intervention.
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic
LAK2
2015 Positive Impact of Collaborative Chat Participation in an edX MOOC
Oliver Ferschke, Diyi Yang, Gaurav Tomar, Carolyn P. Rosé
AIED3
2015 Alleviating the Negative Effect of Up and Downvoting on Help Seeking in MOOC Discussion Forums
Iris K. Howley, Gaurav Tomar, Diyi Yang, Oliver Ferschke, Carolyn P. Rosé
AIED2