VLDB 2026 Research / reviewers in the wild / expert
Nishant Yadav
dblp:230/4155
· DBLP profile ↗
7ranked-venue papers
4as first author
5since 2021 · last 2024
0009-0005-9370-6436ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Information retrieval · 70% Data mining · 20% Query processing and optimization · 8% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 70% Graph algorithms and graph theory · 30% | |
| Artificial intelligence
2 papers |
Trustworthy machine learning · 57% Probabilistic and Bayesian machine learning · 43% |
Topics — the 14 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › similarity search
nearest neighbor search |
1.3 | 2 | 2024 | Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-Encoders · ICLR 2024 Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix Factorization · EMNLP 2022 |
Information retrieval
retrieval models |
0.8 | 1 | 2024 | Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-Encoders · ICLR 2024 |
Algorithms and data structures
clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Graph algorithms and graph theory › graph clustering
correlation clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Algorithms and data structures › clustering › clustering with queries
interactive clustering |
0.6 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Information retrieval › ranking › learning to rank
extreme multi-label ranking |
0.5 | 1 | 2021 | Session-Aware Query Auto-completion using Extreme Multi-Label Ranking · KDD 2021 |
Information retrieval › query suggestion
query auto-completion |
0.5 | 1 | 2021 | Session-Aware Query Auto-completion using Extreme Multi-Label Ranking · KDD 2021 |
Information retrieval
ranking |
0.5 | 1 | 2021 | Session-Aware Query Auto-completion using Extreme Multi-Label Ranking · KDD 2021 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic database |
0.4 | 1 | 2020 | Stochastic Package Queries in Probabilistic Databases · SIGMOD Conference 2020 |
Data mining
clustering |
0.4 | 1 | 2019 | Supervised Hierarchical Clustering with Exponential Linkage · ICML 2019 |
Data mining › clustering
hierarchical clustering |
0.4 | 1 | 2019 | Supervised Hierarchical Clustering with Exponential Linkage · ICML 2019 |
Data mining › clustering
supervised clustering |
0.4 | 1 | 2019 | Supervised Hierarchical Clustering with Exponential Linkage · ICML 2019 |
Algorithms and data structures › clustering
constrained clustering |
0.2 | 1 | 2022 | Interactive Correlation Clustering with Existential Cluster Constraints · ICML 2022 |
Information retrieval
query suggestion |
0.1 | 1 | 2021 | Session-Aware Query Auto-completion using Extreme Multi-Label Ranking · KDD 2021 |
Methods — techniques the papers use, named apart from their topics
inference algorithm · 1.1existential cluster constraints · 1.1stochastic programming · 0.9scenario summarization · 0.9monte carlo methods · 0.9sparse matrix factorization · 0.8dual-encoder initialization · 0.8adaptive retrieval · 0.8matrix factorization · 0.6CUR decomposition · 0.6sequence-to-sequence model · 0.5extreme multi-label ranking · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-EncodersabstractCross-encoder (CE) models which compute similarity by jointly encoding a query-item pair perform better than using dot-product with embedding-based models (dual-encoders) at estimating query-item relevance. Existing approaches perform k-NN search with cross-encoders by approximating the CE similarity with a vector embedding space fit either with dual-encoders (DE) or CUR matrix factorization. DE-based retrieve-and-rerank approaches suffer from poor recall as DE generalizes poorly to new domains and the test-time retrieval with DE is decoupled from the CE. While CUR-based approaches can be more accurate than the DE-based retrieve-and-rerank approach, such approaches require a prohibitively large number of CE calls to compute item embeddings, thus making it impractical for deployment at scale. In this paper, we address these shortcomings with our proposed sparse-matrix factorization based method that efficiently computes latent query and item representations to approximate CE scores and performs k-NN search with the approximate CE similarity. In an offline indexing stage, we compute item embeddings by factorizing a sparse matrix containing query-item CE scores for a set of train queries. Our method produces a high-quality approximation while requiring only a fraction of CE similarity calls as compared to CUR-based methods, and allows for leveraging DE models to initialize the embedding space while avoiding compute- and resource-intensive finetuning of DE via distillation. At test time, we keep item embeddings fixed and perform retrieval over multiple rounds, alternating between a) estimating the test query embedding by minimizing error in approximating CE scores of items retrieved thus far, and b) using the updated test query embedding for retrieving more items in the next round. Our proposed k-NN search method can achieve up to 5 and 54 improvement in k-NN recall for k=1 and 100 respectively over the widely-used DE-based retrieve-and-rerank approach. Furthermore, our proposed approach to index the items by aligning item embeddings with the CE achieves up to 100x and 5x speedup over CUR-based and dual-encoder distillation based approaches respectively while matching or improving k-NN search recall over baselines. Nishant Yadav, Nicholas Monath, Manzil Zaheer, Rob Fergus, Andrew McCallum |
ICLR | 1 |
| 2022 | Efficient Nearest Neighbor Search for Cross-Encoder Models using Matrix FactorizationabstractEfficient k-nearest neighbor search is a fundamental task, foundational for many problems in NLP.When the similarity is measured by dot-product between dual-encoder vectors or ℓ 2 -distance, there already exist many scalable and efficient search methods.But not so when similarity is measured by more accurate and expensive black-box neural similarity models, such as cross-encoders, which jointly encode the query and candidate neighbor.The cross-encoders' high computational cost typically limits their use to reranking candidates retrieved by a cheaper model, such as dual encoder or TF-IDF.However, the accuracy of such a two-stage approach is upper-bounded by the recall of the initial candidate set, and potentially requires additional training to align the auxiliary retrieval model with the cross-encoder model.In this paper, we present an approach that avoids the use of a dual-encoder for retrieval, relying solely on the cross-encoder.Retrieval is made efficient with CUR decomposition, a matrix decomposition approach that approximates all pairwise cross-encoder distances from a small subset of rows and columns of the distance matrix.Indexing items using our approach is computationally cheaper than training an auxiliary dual-encoder model through distillation.Empirically, for k > 10, our approach provides test-time recall-vs-computational cost trade-offs superior to the current widely-used methods that re-rank items retrieved using a dual-encoder or TF-IDF. Nishant Yadav, Nicholas Monath, Rico Angell, Manzil Zaheer, Andrew McCallum |
EMNLP | 1 |
| 2022 | Interactive Correlation Clustering with Existential Cluster ConstraintsabstractWe consider the problem of clustering with user feedback. Existing methods express constraints about the input data points, most commonly through must-link and cannot-link constraints on data point pairs. In this paper, we introduce existential cluster constraints: a new form of feedback where users indicate the features of desired clusters. Specifically, users make statements about the existence of a cluster having (and not having) particular features. Our approach has multiple advantages: (1) constraints on clusters can express user intent more efficiently than point pairs; (2) in cases where the users’ mental model is of the desired clusters, it is more natural for users to express cluster-wise preferences; (3) it functions even when privacy restrictions prohibit users from seeing raw data. In addition to introducing existential cluster constraints, we provide an inference algorithm for incorporating our constraints into the output clustering. Finally, we demonstrate empirically that our proposed framework facilitates more accurate clustering with dramatically fewer user feedback inputs. Rico Angell, Nicholas Monath, Nishant Yadav, Andrew McCallum |
ICML | 3 |
| 2021 | Session-Aware Query Auto-completion using Extreme Multi-Label RankingabstractQuery auto-completion (QAC) is a fundamental feature in search engines where the task is to suggest plausible completions of a prefix typed in the search bar. Previous queries in the user session can provide useful context for the user's intent and can be leveraged to suggest auto-completions that are more relevant while adhering to the user's prefix. Such session-aware QACs can be generated by recent sequence-to-sequence deep learning models; however, these generative approaches often do not meet the stringent latency requirements of responding to each user keystroke. Moreover, these generative approaches pose the risk of showing nonsensical queries. One can pre-compute a relatively small subset of relevant queries for common prefixes and rank them based on the context. However, such an approach fails when no relevant queries for the current context are present in the pre-computed set. Nishant Yadav, Rajat Sen, Daniel N. Hill, Arya Mazumdar, Inderjit S. Dhillon |
KDD | 1 |
| 2021 | Clustering-based Inference for Biomedical Entity LinkingabstractRico Angell, Nicholas Monath, Sunil Mohan, Nishant Yadav, Andrew McCallum. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Rico Angell, Nicholas Monath, Sunil Mohan, Nishant Yadav, Andrew McCallum |
NAACL-HLT | 4 |
| 2020 | Stochastic Package Queries in Probabilistic DatabasesabstractWe provide methods for in-database support of decision making under uncertainty. Many important decision problems correspond to selecting a "package" (bag of tuples in a relational database) that jointly satisfy a set of constraints while minimizing some overall "cost" function; in most real-world problems, the data is uncertain. We provide methods for specifying---via a SQL extension---and processing stochastic package queries (SPQS), in order to solve optimization problems over uncertain data, right where the data resides. Prior work in stochastic programming uses Monte Carlo methods where the original stochastic optimization problem is approximated by a large deterministic optimization problem that incorporates many "scenarios", i.e., sample realizations of the uncertain data values. For large database tables, however, a huge number of scenarios is required, leading to poor performance and, often, failure of the solver software. We therefore provide a novel ßs algorithm that, instead of trying to solve a large deterministic problem, seamlessly approximates it via a sequence of smaller problems defined over carefully crafted "summaries" of the scenarios that accelerate convergence to a feasible and near-optimal solution. Experimental results on our prototype system show that ßs can be orders of magnitude faster than prior methods at finding feasible and high-quality packages. Matteo Brucato, Nishant Yadav, Azza Abouzeid, Peter J. Haas, Alexandra Meliou |
SIGMOD Conference | 2 |
| 2019 | Supervised Hierarchical Clustering with Exponential LinkageabstractIn supervised clustering, standard techniques for learning a pairwise dissimilarity function often suffer from a discrepancy between the training and clustering objectives, leading to poor cluster quality. Rectifying this discrepancy necessitates matching the procedure for training the dissimilarity function to the clustering algorithm. In this paper, we introduce a method for training the dissimilarity function in a way that is tightly coupled with hierarchical clustering, in particular single linkage. However, the appropriate clustering algorithm for a given dataset is often unknown. Thus we introduce an approach to supervised hierarchical clustering that smoothly interpolates between single, average, and complete linkage, and we give a training procedure that simultaneously learns a linkage function and a dissimilarity function. We accomplish this with a novel Exponential Linkage function that has a learnable parameter that controls the interpolation. In experiments on four datasets, our joint training procedure consistently matches or outperforms the next best training procedure/linkage function pair and gives up to 8 points improvement in dendrogram purity over discrepant pairs. Nishant Yadav, Ari Kobren, Nicholas Monath, Andrew McCallum |
ICML | 1 |