VLDB 2026 Research / reviewers in the wild / expert
Vardaan Pahuja
dblp:188/3398
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0001-7538-8474ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Transfer learning and domain adaptation · 30% Trustworthy machine learning · 22% Knowledge representation and reasoning · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 30% Data models and query languages · 30% Data integration and cleaning · 30% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
fine-tuning |
1.4 | 2 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target Data · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
calibration |
0.8 | 1 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.8 | 1 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › calibration
logit calibration |
0.8 | 1 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.7 | 1 | 2023 | Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target Data · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation matching › feature alignment
embedding alignment |
0.5 | 1 | 2021 | A Systematic Investigation of KB-Text Embedding Alignment at Scale · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.5 | 1 | 2021 | A Systematic Investigation of KB-Text Embedding Alignment at Scale · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph embedding |
0.5 | 1 | 2021 | A Systematic Investigation of KB-Text Embedding Alignment at Scale · ACL/IJCNLP (1) 2021 |
Machine learning › Representation and self-supervised learning
text embedding |
0.5 | 1 | 2021 | A Systematic Investigation of KB-Text Embedding Alignment at Scale · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.3 | 1 | 2018 | Complex Sequential Question Answering: Towards Learning to Converse Over Linked Question Answer Pairs with a Knowledge Graph · AAAI 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning |
0.3 | 1 | 2018 | Complex Sequential Question Answering: Towards Learning to Converse Over Linked Question Answer Pairs with a Knowledge Graph · AAAI 2018 |
Data models and query languages › natural language interface
natural language interface to database |
0.3 | 1 | 2018 | Tooling Framework for Instantiating Natural Language Querying System · Proc. VLDB Endow. 2018 |
Information retrieval › query formulation
natural language querying |
0.3 | 1 | 2018 | Tooling Framework for Instantiating Natural Language Querying System · Proc. VLDB Endow. 2018 |
Machine learning › Deep learning architectures and training
foundation model |
0.2 | 1 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
pre-trained models |
0.2 | 1 | 2024 | Fine-Tuning is Fine, if Calibrated · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2023 | Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target Data · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
post-processing calibration · 0.8empirical study · 0.8gradient disentanglement · 0.7class relationship preservation · 0.7knowledge base-text embedding alignment · 0.5tooling framework · 0.3subgraph retrieval · 0.3ontology mapping · 0.3dialog models · 0.3QA models · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs
Vardaan Pahuja, Weidi Luo, Yu Gu 0016, Cheng-Hao Tu 0001, Hong-You Chen, Tanya Y. Berger-Wolf, Charles V. Stewart, Song Gao 0001, Wei-Lun Chao, Yu Su 0001 |
CIKM | 1 |
| 2024 | Fine-Tuning is Fine, if CalibratedabstractFine-tuning is arguably the most straightforward way to tailor a pre-trained model (e.g., a foundation model) to downstream applications, but it also comes with the risk of losing valuable knowledge the model had learned in pre-training. For example, fine-tuning a pre-trained classifier capable of recognizing a large number of classes to master a subset of classes at hand is shown to drastically degrade the model's accuracy in the other classes it had previously learned. As such, it is hard to further use the fine-tuned model when it encounters classes beyond the fine-tuning data. In this paper, we systematically dissect the issue, aiming to answer the fundamental question, "What has been damaged in the fine-tuned model?" To our surprise, we find that the fine-tuned model neither forgets the relationship among the other classes nor degrades the features to recognize these classes. Instead, the fine-tuned model often produces more discriminative features for these other classes, even if they were missing during fine-tuning! What really hurts the accuracy is the discrepant logit scales between the fine-tuning classes and the other classes, implying that a simple post-processing calibration would bring back the pre-trained model's capability and at the same time unveil the feature improvement over all classes. We conduct an extensive empirical study to demonstrate the robustness of our findings and provide preliminary explanations underlying them, suggesting new directions for future theoretical analysis. Zheda Mai, Arpita Chowdhury, Ping Zhang 0016, Cheng-Hao Tu 0001, Hong-You Chen, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 6 |
| 2023 | A Retrieve-and-Read Framework for Knowledge Graph Link PredictionabstractKnowledge graph (KG) link prediction aims to infer new facts based on existing facts in the KG. Recent studies have shown that using the graph neighborhood of a node via graph neural networks (GNNs) provides more useful information compared to just using the query information. Conventional GNNs for KG link prediction follow the standard message-passing paradigm on the entire KG, which leads to superfluous computation, over-smoothing of node representations, and also limits their expressive power. On a large scale, it becomes computationally expensive to aggregate useful information from the entire KG for inference. To address the limitations of existing KG link prediction frameworks, we propose a novel retrieve-and-read framework, which first retrieves a relevant subgraph context for the query and then jointly reasons over the context and the query with a high-capacity reader. As part of our exemplar instantiation for the new framework, we propose a novel Transformer-based GNN as the reader, which incorporates graph-based attention structure and cross-attention between query and context for deep fusion. This simple yet effective design enables the model to focus on salient context information relevant to the query. Empirical results on two standard KG link prediction datasets demonstrate the competitive performance of the proposed method. Furthermore, our analysis yields valuable insights for designing improved retrievers within the framework. Vardaan Pahuja, Boshi Wang, Hugo Latapie, Jayanth Srinivasa, Yu Su 0001 |
CIKM | 1 |
| 2023 | Holistic Transfer: Towards Non-Disruptive Fine-Tuning with Partial Target DataabstractWe propose a learning problem involving adapting a pre-trained source model to the target domain for classifying all classes that appeared in the source data, using target data that covers only a partial label space. This problem is practical, as it is unrealistic for the target end-users to collect data for all classes prior to adaptation. However, it has received limited attention in the literature. To shed light on this issue, we construct benchmark datasets and conduct extensive experiments to uncover the inherent challenges. We found a dilemma --- on the one hand, adapting to the new target domain is important to claim better performance; on the other hand, we observe that preserving the classification accuracy of classes missing in the target adaptation data is highly challenging, let alone improving them. To tackle this, we identify two key directions: 1) disentangling domain gradients from classification gradients, and 2) preserving class relationships. We present several effective solutions that maintain the accuracy of the missing classes and enhance the overall performance, establishing solid baselines for holistic transfer of pre-trained models with partial target data. Cheng-Hao Tu 0001, Hong-You Chen, Zheda Mai, Jike Zhong, Vardaan Pahuja, Tanya Y. Berger-Wolf, Song Gao 0001, Charles V. Stewart, Yu Su 0001, Wei-Lun Chao |
NeurIPS | 5 |
| 2021 | A Systematic Investigation of KB-Text Embedding Alignment at ScaleabstractVardaan Pahuja, Yu Gu, Wenhu Chen, Mehdi Bahrami, Lei Liu, Wei-Peng Chen, Yu Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vardaan Pahuja, Yu Gu 0016, Wenhu Chen, Mehdi Bahrami, Wei-Peng Chen, Yu Su 0001 |
ACL/IJCNLP (1) | 1 |
| 2018 | Complex Sequential Question Answering: Towards Learning to Converse Over Linked Question Answer Pairs with a Knowledge GraphabstractWhile conversing with chatbots, humans typically tend to ask many questions, a significant portion of which can be answered by referring to large-scale knowledge graphs (KG). While Question Answering (QA) and dialog systems have been studied independently, there is a need to study them closely to evaluate such real-world scenarios faced by bots involving both these tasks. Towards this end, we introduce the task of Complex Sequential QA which combines the two tasks of (i) answering factual questions through complex inferencing over a realistic-sized KG of millions of entities, and (ii) learning to converse through a series of coherently linked QA pairs. Through a labor intensive semi-automatic process, involving in-house and crowdsourced workers, we created a dataset containing around 200K dialogs with a total of 1.6M turns. Further, unlike existing large scale QA datasets which contain simple questions that can be answered from a single tuple, the questions in our dialogs require a larger subgraph of the KG. Specifically, our dataset has questions which require logical, quantitative, and comparative reasoning as well as their combinations. This calls for models which can: (i) parse complex natural language questions, (ii) use conversation context to resolve coreferences and ellipsis in utterances, (iii) ask for clarifications for ambiguous queries, and finally (iv) retrieve relevant subgraphs of the KG to answer such questions. However, our experiments with a combination of state of the art dialog and QA models show that they clearly do not achieve the above objectives and are inadequate for dealing with such complex real world settings. We believe that this new dataset coupled with the limitations of existing models as reported in this paper should encourage further research in Complex Sequential QA. Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, Sarath Chandar |
AAAI | 2 |
| 2018 | Tooling Framework for Instantiating Natural Language Querying SystemabstractRecent times have seen a growing demand for natural language querying (NLQ) interfaces to retrieve information from the structured data sources such as knowledge bases. Using this interface, business users can directly interact with a database without the knowledge of the query language or the data schema. Our earlier work describes a natural language query engine called ATHENA which has several shortcoming around ease of use and compatibility with data stores, formats and flows. In this demonstration paper, we present a tooling framework to address these challenges so that one can instantiate an NLQ system with utmost ease. Our framework makes it easy and practically applicable to all NLIDB scenarios involving different sources of structured data, file formats, and ontologies to enable natural language querying on top of them with minimal human configuration. We present the tool design and the solution to the challenges towards building such a system and demonstrate its applicability in the medical domain. Manasa Jammi, Jaydeep Sen, Ashish R. Mittal, Sagar Verma, Vardaan Pahuja, Rema Ananthanarayanan, Pranay Lohia, Hima P. Karanam, Diptikalyan Saha, Karthik Sankaranarayanan |
Proc. VLDB Endow. | 5 |
| 2017 | Joint Learning of Correlated Sequence Labeling Tasks Using Bidirectional Recurrent Neural Networks
Vardaan Pahuja, Anirban Laha, Shachar Mirkin, Vikas C. Raykar, Lili Kotlerman, Guy Lev |
INTERSPEECH | 1 |