Vishwajeet Kumar

dblp:134/6673 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-5783-5207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 50% Knowledge representation and reasoning · 22% Information extraction and text analysis · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 81% Knowledge graphs · 19%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
table question answering
0.512021
Topic Transferable Table Question Answering · EMNLP (1) 2021
Information retrieval
multimodal retrieval
0.512021
Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.412019
Neural Program Induction for KBQA Without Gold Programs or Query Annotations · IJCAI 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge incorporation › knowledge-infused learning › neuro-symbolic learning › neural algorithmic reasoning
neural program induction
0.412019
Neural Program Induction for KBQA Without Gold Programs or Query Annotations · IJCAI 2019
Natural language and speech › Machine translation
annotation projection
0.212016
Towards Semi-Automatic Generation of Proposition Banks for Low-Resource Languages · EMNLP 2016
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.212016
Towards Semi-Automatic Generation of Proposition Banks for Low-Resource Languages · EMNLP 2016
Information retrieval › evaluation
benchmark dataset
0.112021
Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021
Knowledge graphs › knowledge graph querying
knowledge graph question answering
0.112021
Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering · SIGIR 2021

Methods — techniques the papers use, named apart from their topics

neural retrieval · 0.5knowledge graph search · 0.5sparse reward · 0.4reinforcement learning · 0.4syntactic parsing · 0.2annotation projection · 0.2
YearPublicationVenuePosition
2025 Graph Representation of Tables+Text and Compact Subgraph Retrieval for QA Tasks
Vishwajeet Kumar, Jaydeep Sen, Bhawna Chelani, Soumen Chakrabarti
ECIR (1)1
2025 Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5
abstract
Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar, Jaydeep Sen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Arkadeep Acharya, Rudra Murthy, Vishwajeet Kumar, Jaydeep Sen
NAACL (Long Papers)3
2025 MILU: A Multi-task Indic Language Understanding Benchmark
abstract
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, Jaydeep Sen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, V. Rudra Murthy, Jaydeep Sen
NAACL (Long Papers)3
2023 Multi-Row, Multi-Span Distant Supervision For Table+Text Question Answering
abstract
Vishwajeet Kumar, Yash Gupta, Saneem Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Vishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen, Soumen Chakrabarti, Samarth Bharadwaj, Feifei Pan 0002
ACL (1)1
2022 WARM: A Weakly (+Semi) Supervised Math Word Problem Solver
abstract
Solving math word problems (MWPs) is an important and challenging problem in natural language processing. Existing approaches to solving MWPs require full supervision in the form of intermediate equations. However, labeling every MWP with its corresponding equations is a time-consuming and expensive task. In order to address this challenge of equation annotation, we propose a weakly supervised model for solving MWPs by requiring only the final answer as supervision. We approach this problem by first learning to generate the equation using the problem description and the final answer, which we subsequently use to train a supervised MWP solver. We propose and compare various weakly supervised techniques to learn to generate equations directly from the problem description and answer. Through extensive experiments, we demonstrate that without using equations for supervision, our approach achieves accuracy gains of 4.5% and 32% over the current state-of-the-art weakly-supervised approach, on the standard Math23K and AllArith datasets respectively. Additionally, we curate and release new datasets of roughly 10k MWPs each in English and in Hindi (a low-resource language). These datasets are suitable for training weakly supervised models. We also present an extension of our model to semi-supervised learning and present further improvements on results, along with insights.
Oishik Chatterjee, Isha Pandey, Aashish Waikar, Vishwajeet Kumar, Ganesh Ramakrishnan
COLING4
2021 Meta-Learning for Effective Multi-task and Multilingual Modelling
abstract
Ishan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ishan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi
EACL3
2021 Topic Transferable Table Question Answering
abstract
Saneem Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen, Mustafa Canim, Soumen Chakrabarti, Alfio Gliozzo, Karthik Sankaranarayanan. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Saneem A. Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen, Mustafa Canim, Soumen Chakrabarti, Alfio Massimiliano Gliozzo, Karthik Sankaranarayanan
EMNLP (1)2
2021 Capturing Row and Column Semantics in Transformer Based Question Answering over Tables
abstract
Michael Glass, Mustafa Canim, Alfio Gliozzo, Saneem Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avi Sil, Feifei Pan, Samarth Bharadwaj, Nicolas Rodolfo Fauceglia. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Michael R. Glass, Mustafa Canim, Alfio Massimiliano Gliozzo, Saneem A. Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avirup Sil, Feifei Pan 0002, Samarth Bharadwaj, Nicolas R. Fauceglia
NAACL-HLT5
2021 Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question Answering
abstract
Multimodal IR, spanning text corpus, knowledge graph and images, called outside knowledge visual question answering (OKVQA), is of much recent interest. However, the popular data set has serious limitations. A surprisingly large fraction of queries do not assess the ability to integrate cross-modal information. Instead, some are independent of the image, some depend on speculation, some require OCR or are otherwise answerable from the image alone. To add to the above limitations, frequency-based guessing is very effective because of (unintended) widespread answer overlaps between the train and test folds. Overall, it is hard to determine when state-of-the-art systems exploit these weaknesses rather than really infer the answers, because they are opaque and their 'reasoning' process is uninterpretable. An equally important limitation is that the dataset is designed for the quantitative assessment only of the end-to-end answer retrieval task, with no provision for assessing the correct(semantic) interpretation of the input query. In response, we identify a key structural idiom in OKVQA ,viz., S3 (select, substitute and search), and build a new data set and challenge around it. Specifically, the questioner identifies an entity in the image and asks a question involving that entity which can be answered only by consulting a knowledge graph or corpus passage mentioning the entity. Our challenge consists of (i)OKVQA_S3, a subset of OKVQA annotated based on the structural idiom and (ii)S3VQA, a new dataset built from scratch. We also present a neural but structurally transparent OKVQA system, S3, that explicitly addresses our challenge dataset, and outperforms recent competitive baselines. We make our code and data available at https://s3vqa.github.io/.
Aman Jain, Mayank Kothyari, Vishwajeet Kumar, Preethi Jyothi, Ganesh Ramakrishnan, Soumen Chakrabarti
SIGIR3
2020 Variational Student: Learning Compact and Sparser Networks In Knowledge Distillation Framework
abstract
The holy grail in deep neural network research is porting the memory- and computation-intensive network models on embedded platforms with a minimal compromise in model accuracy. To this end, we propose Variational Student where we reap the benefits of compressibility of the knowledge distillation framework, and sparsity inducing abilities of variational inference (VI) techniques. Essentially, we build an accurate and sparse student network, whose sparsity is induced by the variational parameters found via optimizing a loss function based on VI, leveraging the knowledge learnt by an accurate but complex pre-trained teacher network. Further, for sparsity enhancement, we also employ a Block Sparse Regularizer on a concatenated tensor of teacher and student network weights. We benchmark our results on MLP and CNN variants and illustrate an improved performance in lowering the memory footprint up to ~ 213× without a need to retrain the teacher network.
Srinidhi Hegde, Ranjitha Prasad, Ramya Hebbalaguppe, Vishwajeet Kumar
ICASSP4
2019 Cross-Lingual Training for Automatic Question Generation
abstract
Automatic question generation (QG) is a challenging problem in natural language understanding.QG systems are typically built assuming access to a large number of training instances where each instance is a question and its corresponding answer.For a new language, such training instances are hard to obtain making the QG problem even more challenging.Using this as our motivation, we study the reuse of an available large QG dataset in a secondary language (e.g.English) to learn a QG model for a primary language (e.g.Hindi) of interest.For the primary language, we assume access to a large amount of monolingual text but only a small QG dataset.We propose a cross-lingual QG model which uses the following training regime: (i) Unsupervised pretraining of language models in both primary and secondary languages and (ii) joint supervised training for QG in both languages.We demonstrate the efficacy of our proposed approach using two different primary languages, Hindi and Chinese.We also create and release a new question answering dataset for Hindi consisting of 6555 sentences.
Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee, Ganesh Ramakrishnan, Preethi Jyothi
ACL (1)1
2019 Putting the Horse before the Cart: A Generator-Evaluator Framework for Question Generation from Text
abstract
Automatic question generation (QG) is a useful yet challenging task in NLP.Recent neural network-based approaches represent the stateof-the-art in this task.In this work, we attempt to strengthen them significantly by adopting a holistic and novel generator-evaluator framework that directly optimizes objectives that reward semantics and structure.The generator is a sequence-to-sequence model that incorporates the structure and semantics of the question being generated.The generator predicts an answer in the passage that the question can pivot on.Employing the copy and coverage mechanisms, it also acknowledges other contextually important (and possibly rare) keywords in the passage that the question needs to conform to, while not redundantly repeating words.The evaluator model evaluates and assigns a reward to each predicted question based on its conformity to the structure of ground-truth questions.We propose two novel QG-specific reward functions for text conformity and answer conformity of the generated question.The evaluator also employs structure-sensitive rewards based on evaluation measures such as BLEU, GLEU, and ROUGE-L, which are suitable for QG.In contrast, most of the previous works only optimize the cross-entropy loss, which can induce inconsistencies between training (objective) and testing (evaluation) measures.Our evaluation shows that our approach significantly outperforms state-of-the-art systems on the widelyused SQuAD benchmark as per both automatic and human evaluation.
Vishwajeet Kumar, Ganesh Ramakrishnan, Yuan-Fang Li
CoNLL1
2019 Neural Program Induction for KBQA Without Gold Programs or Query Annotations
abstract
Neural Program Induction (NPI) is a paradigm for decomposing high-level tasks such as complex question-answering over knowledge bases (KBQA) into executable programs by employing neural models. Typically, this involves two key phases: i) inferring input program variables from the high-level task description, and ii) generating the correct program sequence involving these variables. Here we focus on NPI for Complex KBQA with only the final answer as supervision, and not gold programs. This raises major challenges; namely, i) noisy query annotation in the absence of any supervision can lead to catastrophic forgetting while learning, ii) reward becomes extremely sparse owing to the noise. To deal with these, we propose a noise-resilient NPI model, Stable Sparse Reward based Programmer (SSRP) that evades noise-induced instability through continual retrospection and its comparison with current learning behavior. On complex KBQA datasets, SSRP performs at par with hand-crafted rule-based models when provided with gold program input, and in the noisy settings outperforms state-of-the-art models by a significant margin even with a noisier query annotator.
Ghulam Ahmed Ansari, Amrita Saha, Vishwajeet Kumar, Mohan Bhambhani, Karthik Sankaranarayanan, Soumen Chakrabarti
IJCAI3
2019 Difficulty-Controllable Multi-hop Question Generation from Knowledge Graphs
Vishwajeet Kumar, Yuncheng Hua, Ganesh Ramakrishnan, Guilin Qi, Lianli Gao, Yuan-Fang Li
ISWC (1)1
2018 Automating Reading Comprehension by Generating Question and Answer Pairs
Vishwajeet Kumar, Kireeti Boorla, Yogesh Kumar Meena, Ganesh Ramakrishnan, Yuan-Fang Li
PAKDD (3)1
2016 Towards Semi-Automatic Generation of Proposition Banks for Low-Resource Languages
abstract
Annotation projection based on parallel corpora has shown great promise in inexpensively creating Proposition Banks for languages for which high-quality parallel corpora and syntactic parsers are available.In this paper, we present an experimental study where we apply this approach to three languages that lack such resources: Tamil, Bengali and Malayalam.We find an average quality difference of 6 to 20 absolute F-measure points vis-avis high-resource languages, which indicates that annotation projection alone is insufficient in low-resource scenarios.Based on these results, we explore the possibility of using annotation projection as a starting point for inexpensive data curation involving both experts and non-experts.We give an outline of what such a process may look like and present an initial study to discuss its potential and challenges.
Alan Akbik, Vishwajeet Kumar, Yunyao Li 0001
EMNLP2
2016 Building Compact Lexicons for Cross-Domain SMT by Mining Near-Optimal Pattern Sets
Pankaj Singh, Ashish Kulkarni, Himanshu Ojha, Vishwajeet Kumar, Ganesh Ramakrishnan
PAKDD (1)4
2015 A Machine Assisted Human Translation System for Technical Documents
abstract
Translation systems are known to benefit from the availability of a bilingual lexicon for a domain of interest. A system, aiming to build such a lexicon from source language corpus, often requires human assistance and is confronted by conflicting requirements of minimizing human translation effort while improving the translation quality. We present an approach that exploits redundancy in the source corpus and extracts recurring patterns which are: frequent, syntactically well-formed, and provide maximum corpus coverage. The patterns generalize over phrases and word types and our approach finds a succinct set of good patterns with high coverage. Our interactive system leverages these patterns in multiple iterations of translation and post-editing, thereby progressively generating a high quality bilingual lexicon.
Vishwajeet Kumar, Ashish Kulkarni, Pankaj Singh, Ganesh Ramakrishnan, Ganesh Arnaal
K-CAP1