VLDB 2026 Research / reviewers in the wild / expert
Ankit Gandhi
dblp:136/3925
· DBLP profile ↗
10ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0002-8286-2792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReSuMe: Retriever-Summarizer Mutual Enhancement via Reinforcement LearningabstractWe present ReSuMe, a general framework for mutual enhancement of dense retrieval systems and document summarizers through reinforcement learning. The framework jointly optimizes a language model for generating retrieval-oriented summaries and adapts the retrieval model to these summaries through alternating fine-tuning phases. We employ Group Relative Policy Optimization (GRPO) to fine-tune the language model based on retrieval relevance rather than linguistic quality alone, while the retrieval model is iteratively updated using contrastive learning on the generated summaries. This co-optimization process addresses the fundamental distribution shift problem that arises when retrieval models trained on full documents must operate on synthetic summaries during inference. By progressively reducing this distribution gap, our framework yields two key benefits: improved retrieval performance and a high-quality document summarizer optimized for retrieval tasks. We demonstrate our framework using Contriever on the MS-MARCO dataset, achieving consistent improvements of 13.2% in MRR@10 and 6.7% in Recall@100 over the baseline. The framework is model-agnostic and can be applied to enhance any dense retrieval system while simultaneously producing an effective document summarization model. Owais Makroo, Nikhil Pattisapu, Karan Gupta 0002, Ankit Gandhi, Vijay Huddar, Atul Saroop |
WWW | 4 |
| 2023 | Beyond Hard Negatives in Product Search: Semantic Matching Using One-Class Classification (SMOCC)abstractSemantic matching is an important component of a product search pipeline. Its goal is to capture the semantic intent of the search query as opposed to the syntactic matching performed by a lexical matching system. A semantic matching model captures relationships like synonyms, and also captures common behavioral patterns to retrieve relevant results by generalizing from purchase data. They however suffer from lack of availability of informative negative examples for model training. Various methods have been proposed in the past to address this issue based upon hard-negative mining and contrastive learning. Arindam Bhattacharya, Ankit Gandhi, Vijay Huddar, Ankith M. S, Aayush Moroney, Atul Saroop, Rahul Bhagat |
WSDM | 2 |
| 2021 | Spatio-Temporal Multi-graph Networks for Demand Forecasting in Online Marketplaces
Ankit Gandhi, Aakanksha, Sivaramakrishnan Kaveri, Vineet Chaoji |
ECML/PKDD (4) | 1 |
| 2016 | Weakly Supervised Learning of Heterogeneous Concepts in Videos
Sohil Shah, Kuldeep Kulkarni, Arijit Biswas, Ankit Gandhi, Om Deshmukh, Larry Davis 0001 |
ECCV (6) | 4 |
| 2016 | LIVELINET: A Multimodal Deep Recurrent Neural Network to Predict Liveliness in Educational Videos
Arjun Sharma, Arijit Biswas, Ankit Gandhi, Sonal Patil, Om Deshmukh |
EDM | 3 |
| 2016 | ViZig: Anchor Points based Non-Linear Navigation and Summarization in Educational VideosabstractInstructional videos are one of the most popular ways of teaching and learning in an online setting. However, navigation in videos is linear as compared to other instructional resources such as textbooks, where a table of topics and a multi-faceted index of different anchor points i.e., list of figures, tables aid in efficiently navigating to a desired point of interest. There is a lack of appropriate techniques and interfaces which can support such textbook-style navigation in instructional videos. This paper presents a novel approach to automatically localize and classify different anchor points in a video including figures, tables, equations, flowcharts, code snippets and charts. Our approach uses a deep convolution neural network in a semi-supervised fashion where the training data is obtained from the unconstrained Internet images. On an anchor point dataset of about 10K images, the proposed algorithm leads to a classification accuracy of 86%. Further, we designed a system ViZig that uses these localized anchor points along with a automatically generated list of topics for non-linear video navigation and studied its effectiveness in real-world. Our user studies with 18 participants establish that the proposed video navigation mechanism provides statistically significant time savings as compared to the popularly used time-synched transcript along with youtube-style timeline scrubbing. Ankit Gandhi, Arijit Biswas, Kundan Srivastava, Om Deshmukh |
IUI | 2 |
| 2015 | Topic Transition in Educational Videos Using Visually Salient Words
Ankit Gandhi, Arijit Biswas, Om Deshmukh |
EDM | 1 |
| 2015 | MMToC: A Multimodal Method for Table of Content Creation in Educational VideosabstractIn this paper we propose a multimodal method called MMToC for automatically creating a table of content for educational videos. MMToC defines and quantifies word saliency for visual words extracted from the slides and spoken words obtained from the speech transcript. The saliency scores from these two modalities are combined to obtain a ranked list of salient words. These ranked words along with their saliency scores are used to formulate a topic segmentation cost function. The cost function is optimized using a dynamic program framework to obtain the topic segments of the video. These segments are labelled with their corresponding topic names for creating the table of content. We perform experiments on 24 hours of lectures spread across 23 videos ranging over 20-75 minutes duration each. We compare the proposed method with LDA-based video segmentation approaches and show that the proposed MMToC method is significantly better (F-score improvement of 0.19 and 0.24 on two datasets). We also perform a user study to demonstrate the effectiveness of MMToC for navigating educational videos. Arijit Biswas, Ankit Gandhi, Om Deshmukh |
ACM Multimedia | 2 |
| 2013 | Decomposing Bag of Words HistogramsabstractWe aim to decompose a global histogram representation of an image into histograms of its associated objects and regions. This task is formulated as an optimization problem, given a set of linear classifiers, which can effectively discriminate the object categories present in the image. Our decomposition bypasses harder problems associated with accurately localizing and segmenting objects. We evaluate our method on a wide variety of composite histograms, and also compare it with MRF-based solutions. In addition to merely measuring the accuracy of decomposition, we also show the utility of the estimated object and background histograms for the task of image classification on the PASCAL VOC 2007 dataset. Ankit Gandhi, Karteek Alahari, C. V. Jawahar |
ICCV | 1 |
| 2013 | Detection of Cut-and-Paste in Document ImagesabstractMany documents are created by Cut-And-Paste (CAP) of existing documents. In this paper, we proposed a novel technique to detect CAP in document images. This can help in detecting unethical CAP in document image collections. Our solution is recognition free, and scalable to large collection of documents. Our formulation is also independent of the imaging process (camera based or scanner based) and does not use any language specific information for matching across documents. We model the solution as finding a mixture of homographies, and design a linear programming (LP) based solution to compute the same. Our method is presently limited by the fact that we do not support detection of CAP in documents formed by editing of the textual content. Our experiments demonstrate that without loss of generality (i.e. without assuming the number of source documents), we can correctly detect and match the CAP content in a questioned document image by simultaneously comparing with large number of images in the database. We achieve the CAP detection accuracy of as high as 90%, even when the spatial extent of the CAP content in a document image is as small as 15% of the entire image area. Ankit Gandhi, C. V. Jawahar |
ICDAR | 1 |