Devendra Singh Sachan

dblp:167/3916 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
5since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Question answering and dialogue systems · 33% Optimization for machine learning · 21% Information extraction and text analysis · 16%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
open-domain question answering
1.632022
Improving Passage Retrieval with Zero-Shot Question Generation · EMNLP 2022
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering · NeurIPS 2021
End-to-End Training of Neural Retrievers for Open-Domain Question Answering · ACL/IJCNLP (1) 2021
Information retrieval › document retrieval
passage retrieval
0.612022
Improving Passage Retrieval with Zero-Shot Question Generation · EMNLP 2022
Information retrieval
reranking
0.612022
Improving Passage Retrieval with Zero-Shot Question Generation · EMNLP 2022
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.512021
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering · NeurIPS 2021
Information retrieval › retrieval models
neural retrieval
0.512021
End-to-End Training of Neural Retrievers for Open-Domain Question Answering · ACL/IJCNLP (1) 2021
Machine learning › Deep learning architectures and training › recurrent neural network › bidirectional recurrent network
BiLSTM
0.412019
Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function · AAAI 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function · AAAI 2019
Natural language and speech › Information extraction and text analysis › text classification
semi-supervised text classification
0.412019
Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function · AAAI 2019
Natural language and speech › Information extraction and text analysis
text classification
0.412019
Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function · AAAI 2019
Machine learning › Optimization for machine learning › stochastic optimization
adaptive gradient methods
0.312018
Adaptive Methods for Nonconvex Optimization · NeurIPS 2018
Machine learning › Optimization for machine learning
convergence analysis
0.312018
Adaptive Methods for Nonconvex Optimization · NeurIPS 2018
Machine learning › Optimization for machine learning › non-convex optimization
non-convex stochastic optimization
0.312018
Adaptive Methods for Nonconvex Optimization · NeurIPS 2018
Natural language and speech › Language models and text generation
large language model
0.112021
End-to-End Training of Neural Retrievers for Open-Domain Question Answering · ACL/IJCNLP (1) 2021
Information retrieval › document retrieval
multi-document retrieval
0.112021
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

zero-shot question generation · 1.1pre-trained language model · 1.1expectation-maximization · 1.0end-to-end training · 1.0end-to-end differentiable training · 1.0dense retrieval · 1.0virtual adversarial training · 0.4entropy minimization · 0.4adversarial training · 0.4adam · 0.3
YearPublicationVenuePosition
2023 Questions Are All You Need to Train a Dense Passage Retriever
abstract
Abstract We introduce ART, a new corpus-level autoencoding approach for training dense retrieval models that does not require any labeled training data. Dense retrieval is a central challenge for open-domain tasks, such as Open QA, where state-of-the-art methods typically require large supervised datasets with custom hard-negative mining and denoising of positive examples. ART, in contrast, only requires access to unpaired inputs and outputs (e.g., questions and potential answer passages). It uses a new passage-retrieval autoencoding scheme, where (1) an input question is used to retrieve a set of evidence passages, and (2) the passages are then used to compute the probability of reconstructing the original question. Training for retrieval based on question reconstruction enables effective unsupervised learning of both passage and question encoders, which can be later incorporated into complete Open QA systems without any further finetuning. Extensive experiments demonstrate that ART obtains state-of-the-art results on multiple QA retrieval benchmarks with only generic initialization from a pre-trained language model, removing the need for labeled data and task-specific losses.1 Our code and model checkpoints are available at: https://github.com/DevSinghSachan/art.
Devendra Singh Sachan, Mike Lewis, Dani Yogatama, Luke Zettlemoyer, Joelle Pineau, Manzil Zaheer
Trans. Assoc. Comput. Linguistics1
2022 Improving Passage Retrieval with Zero-Shot Question Generation
abstract
We propose a simple and effective re-ranking method for improving passage retrieval in open question answering.The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage.This approach can be applied on top of any retrieval method (e.g.neural or keywordbased), does not require any domain-or taskspecific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question).When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy.We also obtain new stateof-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.1
Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Scott Yih, Joelle Pineau, Luke Zettlemoyer
EMNLP1
2021 End-to-End Training of Neural Retrievers for Open-Domain Question Answering
abstract
Devendra Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L. Hamilton, Bryan Catanzaro. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Devendra Singh Sachan, Mostofa Patwary, Mohammad Shoeybi, Neel Kant, Wei Ping, William L. Hamilton, Bryan Catanzaro
ACL/IJCNLP (1)1
2021 Do Syntax Trees Help Pre-trained Transformers Extract Information?
abstract
Much recent work suggests that incorporating syntax information from dependency trees can improve task-specific transformer models.However, the effect of incorporating dependency tree information into pre-trained transformer models (e.g., BERT) remains unclear, especially given recent studies highlighting how these models implicitly encode syntax.In this work, we systematically study the utility of incorporating dependency trees into pretrained transformers on three representative information extraction tasks: semantic role labeling (SRL), named entity recognition, and relation extraction.We propose and investigate two distinct strategies for incorporating dependency structure: a late fusion approach, which applies a graph neural network on the output of a transformer, and a joint fusion approach, which infuses syntax structure into the transformer attention layers.These strategies are representative of prior work, but we introduce additional model design elements that are necessary for obtaining improved performance.Our empirical analysis demonstrates that these syntax-infused transformers obtain state-of-the-art results on SRL and relation extraction tasks.However, our analysis also reveals a critical shortcoming of these models: we find that their performance gains are highly contingent on the availability of human-annotated dependency parses, which raises important questions regarding the viability of syntax-augmented transformers in real-world applications.1
Devendra Singh Sachan, Yuhao Zhang 0004, Peng Qi 0003, William L. Hamilton
EACL1
2021 End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering
abstract
We present an end-to-end differentiable training method for retrieval-augmented open-domain question answering systems that combine information from multiple retrieved documents when generating answers. We model retrieval decisions as latent variables over sets of relevant documents. Since marginalizing over sets of retrieved documents is computationally hard, we approximate this using an expectation-maximization algorithm. We iteratively estimate the value of our latent variable (the set of relevant documents for a given question) and then use this estimate to update the retriever and reader parameters. We hypothesize that such end-to-end training allows training signals to flow to the reader and then to the retriever better than staged-wise training. This results in a retriever that is able to select more relevant documents for a question and a reader that is trained on more accurate documents to generate an answer. Experiments on three benchmark datasets demonstrate that our proposed method outperforms all existing approaches of comparable size by 2-3% absolute exact match points, achieving new state-of-the-art results. Our results also demonstrate the feasibility of learning to retrieve to improve answer generation without explicit supervision of retrieval decisions.
Devendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer, Dani Yogatama
NeurIPS1
2019 Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function
abstract
In this paper, we study bidirectional LSTM network for the task of text classification using both supervised and semisupervised approaches. Several prior works have suggested that either complex pretraining schemes using unsupervised methods such as language modeling (Dai and Le 2015; Miyato, Dai, and Goodfellow 2016) or complicated models (Johnson and Zhang 2017) are necessary to achieve a high classification accuracy. However, we develop a training strategy that allows even a simple BiLSTM model, when trained with cross-entropy loss, to achieve competitive results compared with more complex approaches. Furthermore, in addition to cross-entropy loss, by using a combination of entropy minimization, adversarial, and virtual adversarial losses for both labeled and unlabeled data, we report state-of-theart results for text classification task on several benchmark datasets. In particular, on the ACL-IMDB sentiment analysis and AG-News topic classification datasets, our method outperforms current approaches by a substantial margin. We also show the generality of the mixed objective function by improving the performance on relation extraction task.1
Devendra Singh Sachan, Manzil Zaheer, Ruslan Salakhutdinov
AAAI1
2018 Investigating the Working of Text Classifiers
abstract
Text classification is one of the most widely studied tasks in natural language processing. Motivated by the principle of compositionality, large multilayer neural network models have been employed for this task in an attempt to effectively utilize the constituent expressions. Almost all of the reported work train large networks using discriminative approaches, which come with a caveat of no proper capacity control, as they tend to latch on to any signal that may not generalize. Using various recent state-of-the-art approaches for text classification, we explore whether these models actually learn to compose the meaning of the sentences or still just focus on some keywords or lexicons for classifying the document. To test our hypothesis, we carefully construct datasets where the training and test splits have no direct overlap of such lexicons, but overall language structure would be similar. We study various text classifiers and observe that there is a big performance drop on these datasets. Finally, we show that even simple models with our proposed regularization techniques, which disincentivize focusing on key lexicons, can substantially improve classification accuracy.
Devendra Singh Sachan, Manzil Zaheer, Ruslan Salakhutdinov
COLING1
2018 Adaptive Methods for Nonconvex Optimization
abstract
Adaptive gradient methods that rely on scaling gradients down by the square root of exponential moving averages of past squared gradients, such RMSProp, Adam, Adadelta have found wide application in optimizing the nonconvex problems that arise in deep learning. However, it has been recently demonstrated that such methods can fail to converge even in simple convex optimization settings. In this work, we provide a new analysis of such methods applied to nonconvex stochastic optimization problems, characterizing the effect of increasing minibatch size. Our analysis shows that under this scenario such methods do converge to stationarity up to the statistical limit of variance in the stochastic gradients (scaled by a constant factor). In particular, our result implies that increasing minibatch sizes enables convergence, thus providing a way to circumvent the non-convergence issues. Furthermore, we provide a new adaptive optimization algorithm, Yogi, which controls the increase in effective learning rate, leading to even better performance with similar theoretical guarantees on convergence. Extensive experiments show that Yogi with very little hyperparameter tuning outperforms methods such as Adam in several challenging machine learning tasks.
Manzil Zaheer, Sashank J. Reddi, Devendra Singh Sachan, Satyen Kale, Sanjiv Kumar
NeurIPS3