VLDB 2026 Research / reviewers in the wild / expert
Vishwajeet Kumar
dblp:134/6673
· DBLP profile ↗
18ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-5783-5207ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Question answering and dialogue systems · 50% Knowledge representation and reasoning · 22% Information extraction and text analysis · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 81% Knowledge graphs · 19% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
table question answering |
0.5 | 1 | 2021 | Topic Transferable Table Question Answering · EMNLP (1) 2021 |
Information retrieval
multimodal retrieval |
0.5 | 1 | 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.4 | 1 | 2019 | Neural Program Induction for KBQA Without Gold Programs or Query Annotations · IJCAI 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge incorporation › knowledge-infused learning › neuro-symbolic learning › neural algorithmic reasoning
neural program induction |
0.4 | 1 | 2019 | Neural Program Induction for KBQA Without Gold Programs or Query Annotations · IJCAI 2019 |
Natural language and speech › Machine translation
annotation projection |
0.2 | 1 | 2016 | Towards Semi-Automatic Generation of Proposition Banks for Low-Resource Languages · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.2 | 1 | 2016 | Towards Semi-Automatic Generation of Proposition Banks for Low-Resource Languages · EMNLP 2016 |
Information retrieval › evaluation
benchmark dataset |
0.1 | 1 | 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021 |
Knowledge graphs › knowledge graph querying
knowledge graph question answering |
0.1 | 1 | 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021 |
Methods — techniques the papers use, named apart from their topics
neural retrieval · 0.5knowledge graph search · 0.5sparse reward · 0.4reinforcement learning · 0.4syntactic parsing · 0.2annotation projection · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graph Representation of Tables+Text and Compact Subgraph Retrieval for QA Tasks
Vishwajeet Kumar, Jaydeep Sen, Bhawna Chelani, Soumen Chakrabarti |
ECIR (1) | 1 |
| 2025 | Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5abstractArkadeep Acharya, Rudra Murthy, Vishwajeet Kumar, Jaydeep Sen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar, Jaydeep Sen |
NAACL (Long Papers) | 3 |
| 2025 | MILU: A Multi-task Indic Language Understanding BenchmarkabstractSshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, Jaydeep Sen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, V. Rudra Murthy, Jaydeep Sen |
NAACL (Long Papers) | 3 |
| 2023 | Multi-Row, Multi-Span Distant Supervision For Table+Text Question AnsweringabstractVishwajeet Kumar, Yash Gupta, Saneem Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Vishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan 0002 |
ACL (1) | 1 |
| 2022 | WARM: A Weakly (+Semi) Supervised Math Word Problem SolverabstractSolving math word problems (MWPs) is an important and challenging problem in natural language processing. Existing approaches to solving MWPs require full supervision in the form of intermediate equations. However, labeling every MWP with its corresponding equations is a time-consuming and expensive task. In order to address this challenge of equation annotation, we propose a weakly supervised model for solving MWPs by requiring only the final answer as supervision. We approach this problem by first learning to generate the equation using the problem description and the final answer, which we subsequently use to train a supervised MWP solver. We propose and compare various weakly supervised techniques to learn to generate equations directly from the problem description and answer. Through extensive experiments, we demonstrate that without using equations for supervision, our approach achieves accuracy gains of 4.5% and 32% over the current state-of-the-art weakly-supervised approach, on the standard Math23K and AllArith datasets respectively. Additionally, we curate and release new datasets of roughly 10k MWPs each in English and in Hindi (a low-resource language). These datasets are suitable for training weakly supervised models. We also present an extension of our model to semi-supervised learning and present further improvements on results, along with insights. Oishik Chatterjee, Isha Pandey, Aashish Waikar, Vishwajeet Kumar, Ganesh Ramakrishnan |
COLING | 4 |
| 2021 | Meta-Learning for Effective Multi-task and Multilingual ModellingabstractIshan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ishan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi |
EACL | 3 |
| 2021 | Topic Transferable Table Question AnsweringabstractSaneem Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen, Mustafa Canim, Soumen Chakrabarti, Alfio Gliozzo, Karthik Sankaranarayanan. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Saneem A. Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen, Mustafa Canim, Soumen Chakrabarti, Alfio Massimiliano Gliozzo, Karthik Sankaranarayanan |
EMNLP (1) | 2 |
| 2021 | Capturing Row and Column Semantics in Transformer Based Question Answering over TablesabstractMichael Glass, Mustafa Canim, Alfio Gliozzo, Saneem Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avi Sil, Feifei Pan, Samarth Bharadwaj, Nicolas Rodolfo Fauceglia. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michael R. Glass, Mustafa Canim, Alfio Massimiliano Gliozzo, Saneem A. Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avirup Sil, Feifei Pan 0002, Samarth Bharadwaj, Nicolas R. Fauceglia |
NAACL-HLT | 5 |
| 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question AnsweringabstractMultimodal IR, spanning text corpus, knowledge graph and images, called outside knowledge visual question answering (OKVQA), is of much recent interest. However, the popular data set has serious limitations. A surprisingly large fraction of queries do not assess the ability to integrate cross-modal information. Instead, some are independent of the image, some depend on speculation, some require OCR or are otherwise answerable from the image alone. To add to the above limitations, frequency-based guessing is very effective because of (unintended) widespread answer overlaps between the train and test folds. Overall, it is hard to determine when state-of-the-art systems exploit these weaknesses rather than really infer the answers, because they are opaque and their 'reasoning' process is uninterpretable. An equally important limitation is that the dataset is designed for the quantitative assessment only of the end-to-end answer retrieval task, with no provision for assessing the correct(semantic) interpretation of the input query. In response, we identify a key structural idiom in OKVQA ,viz., S3 (select, substitute and search), and build a new data set and challenge around it. Specifically, the questioner identifies an entity in the image and asks a question involving that entity which can be answered only by consulting a knowledge graph or corpus passage mentioning the entity. Our challenge consists of (i)OKVQA_S3, a subset of OKVQA annotated based on the structural idiom and (ii)S3VQA, a new dataset built from scratch. We also present a neural but structurally transparent OKVQA system, S3, that explicitly addresses our challenge dataset, and outperforms recent competitive baselines. We make our code and data available at https://s3vqa.github.io/. Aman Jain, Mayank Kothyari, Vishwajeet Kumar, Preethi Jyothi, Ganesh Ramakrishnan, Soumen Chakrabarti |
SIGIR | 3 |
| 2020 | Variational Student: Learning Compact and Sparser Networks In Knowledge Distillation FrameworkabstractThe holy grail in deep neural network research is porting the memory- and computation-intensive network models on embedded platforms with a minimal compromise in model accuracy. To this end, we propose Variational Student where we reap the benefits of compressibility of the knowledge distillation framework, and sparsity inducing abilities of variational inference (VI) techniques. Essentially, we build an accurate and sparse student network, whose sparsity is induced by the variational parameters found via optimizing a loss function based on VI, leveraging the knowledge learnt by an accurate but complex pre-trained teacher network. Further, for sparsity enhancement, we also employ a Block Sparse Regularizer on a concatenated tensor of teacher and student network weights. We benchmark our results on MLP and CNN variants and illustrate an improved performance in lowering the memory footprint up to ~ 213× without a need to retrain the teacher network. Srinidhi Hegde, Ranjitha Prasad, Ramya Hebbalaguppe, Vishwajeet Kumar |
ICASSP | 4 |
| 2019 | Cross-Lingual Training for Automatic Question GenerationabstractAutomatic question generation (QG) is a challenging problem in natural language understanding.QG systems are typically built assuming access to a large number of training instances where each instance is a question and its corresponding answer.For a new language, such training instances are hard to obtain making the QG problem even more challenging.Using this as our motivation, we study the reuse of an available large QG dataset in a secondary language (e.g.English) to learn a QG model for a primary language (e.g.Hindi) of interest.For the primary language, we assume access to a large amount of monolingual text but only a small QG dataset.We propose a cross-lingual QG model which uses the following training regime: (i) Unsupervised pretraining of language models in both primary and secondary languages and (ii) joint supervised training for QG in both languages.We demonstrate the efficacy of our proposed approach using two different primary languages, Hindi and Chinese.We also create and release a new question answering dataset for Hindi consisting of 6555 sentences. Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee, Ganesh Ramakrishnan, Preethi Jyothi |
ACL (1) | 1 |
| 2019 | Putting the Horse before the Cart: A Generator-Evaluator Framework for Question Generation from TextabstractAutomatic question generation (QG) is a useful yet challenging task in NLP.Recent neural network-based approaches represent the stateof-the-art in this task.In this work, we attempt to strengthen them significantly by adopting a holistic and novel generator-evaluator framework that directly optimizes objectives that reward semantics and structure.The generator is a sequence-to-sequence model that incorporates the structure and semantics of the question being generated.The generator predicts an answer in the passage that the question can pivot on.Employing the copy and coverage mechanisms, it also acknowledges other contextually important (and possibly rare) keywords in the passage that the question needs to conform to, while not redundantly repeating words.The evaluator model evaluates and assigns a reward to each predicted question based on its conformity to the structure of ground-truth questions.We propose two novel QG-specific reward functions for text conformity and answer conformity of the generated question.The evaluator also employs structure-sensitive rewards based on evaluation measures such as BLEU, GLEU, and ROUGE-L, which are suitable for QG.In contrast, most of the previous works only optimize the cross-entropy loss, which can induce inconsistencies between training (objective) and testing (evaluation) measures.Our evaluation shows that our approach significantly outperforms state-of-the-art systems on the widelyused SQuAD benchmark as per both automatic and human evaluation. Vishwajeet Kumar, Ganesh Ramakrishnan, Yuan-Fang Li |
CoNLL | 1 |
| 2019 | Neural Program Induction for KBQA Without Gold Programs or Query AnnotationsabstractNeural Program Induction (NPI) is a paradigm for decomposing high-level tasks such as complex question-answering over knowledge bases (KBQA) into executable programs by employing neural models. Typically, this involves two key phases: i) inferring input program variables from the high-level task description, and ii) generating the correct program sequence involving these variables. Here we focus on NPI for Complex KBQA with only the final answer as supervision, and not gold programs. This raises major challenges; namely, i) noisy query annotation in the absence of any supervision can lead to catastrophic forgetting while learning, ii) reward becomes extremely sparse owing to the noise. To deal with these, we propose a noise-resilient NPI model, Stable Sparse Reward based Programmer (SSRP) that evades noise-induced instability through continual retrospection and its comparison with current learning behavior. On complex KBQA datasets, SSRP performs at par with hand-crafted rule-based models when provided with gold program input, and in the noisy settings outperforms state-of-the-art models by a significant margin even with a noisier query annotator. Ghulam Ahmed Ansari, Amrita Saha, Vishwajeet Kumar, Mohan Bhambhani, Karthik Sankaranarayanan, Soumen Chakrabarti |
IJCAI | 3 |
| 2019 | Difficulty-Controllable Multi-hop Question Generation from Knowledge Graphs
Vishwajeet Kumar, Yuncheng Hua, Ganesh Ramakrishnan, Guilin Qi, Lianli Gao, Yuan-Fang Li |
ISWC (1) | 1 |
| 2018 | Automating Reading Comprehension by Generating Question and Answer Pairs
Vishwajeet Kumar, Kireeti Boorla, Yogesh Kumar Meena, Ganesh Ramakrishnan, Yuan-Fang Li |
PAKDD (3) | 1 |
| 2016 | Towards Semi-Automatic Generation of Proposition Banks for Low-Resource LanguagesabstractAnnotation projection based on parallel corpora has shown great promise in inexpensively creating Proposition Banks for languages for which high-quality parallel corpora and syntactic parsers are available.In this paper, we present an experimental study where we apply this approach to three languages that lack such resources: Tamil, Bengali and Malayalam.We find an average quality difference of 6 to 20 absolute F-measure points vis-avis high-resource languages, which indicates that annotation projection alone is insufficient in low-resource scenarios.Based on these results, we explore the possibility of using annotation projection as a starting point for inexpensive data curation involving both experts and non-experts.We give an outline of what such a process may look like and present an initial study to discuss its potential and challenges. Alan Akbik, Vishwajeet Kumar, Yunyao Li 0001 |
EMNLP | 2 |
| 2016 | Building Compact Lexicons for Cross-Domain SMT by Mining Near-Optimal Pattern Sets
Pankaj Singh, Ashish Kulkarni, Himanshu Ojha, Vishwajeet Kumar, Ganesh Ramakrishnan |
PAKDD (1) | 4 |
| 2015 | A Machine Assisted Human Translation System for Technical DocumentsabstractTranslation systems are known to benefit from the availability of a bilingual lexicon for a domain of interest. A system, aiming to build such a lexicon from source language corpus, often requires human assistance and is confronted by conflicting requirements of minimizing human translation effort while improving the translation quality. We present an approach that exploits redundancy in the source corpus and extracts recurring patterns which are: frequent, syntactically well-formed, and provide maximum corpus coverage. The patterns generalize over phrases and word types and our approach finds a succinct set of good patterns with high coverage. Our interactive system leverages these patterns in multiple iterations of translation and post-editing, thereby progressively generating a high quality bilingual lexicon. Vishwajeet Kumar, Ashish Kulkarni, Pankaj Singh, Ganesh Ramakrishnan, Ganesh Arnaal |
K-CAP | 1 |