Lovekesh Vig

dblp:16/4559 · DBLP profile ↗
← Back
54ranked-venue papers
2as first author
21since 2021 · last 2026
0000-0001-9834-3308ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Theory of computation · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Refine and Align: Confidence Calibration Through Multi-Agent Interaction in VQA
abstract
In the context of Visual Question Answering (VQA) and Agentic AI, calibration refers to how closely an AI system's confidence in its answers reflects their actual correctness. This aspect becomes especially important when such systems operate autonomously and must make decisions under visual uncertainty. While modern VQA systems, powered by advanced vision-language models (VLMs), are increasingly used in high-stakes domains like medical diagnostics and autonomous navigation due to their improved accuracy, the reliability of their confidence estimates remains under-examined. Particularly, these systems often produce overconfident responses. To address this, we introduce AlignVQA, a debate-based multi-agent framework, in which diverse specialized VLM -- each following distinct prompting strategies -- generate candidate answers and then engage in two-stage interaction: generalist agents critique, refine and aggregate these proposals. This debate process yields confidence estimates that more accurately reflect the model’s true predictive performance. We find that more calibrated specialized agents produce better aligned confidences. Furthermore, we introduce a novel differentiable calibration-aware loss function called AlignCal designed to fine-tune the specialized agents by minimizing an upper bound on the calibration error. This objective explicitly improves the fidelity of each agent’s confidence estimates. Empirical results across multiple benchmark VQA datasets substantiate the efficacy of our approach, demonstrating substantial reductions in calibration discrepancies.
Jai Bardhan, Ishita Jain, Ramya Hebbalaguppe, Rohan Raju Dhanakshirur, Lovekesh Vig
AAAI6
2025 AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
abstract
Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan
ACL (1)7
2025 MIR: Methodology Inspiration Retrieval for Scientific Research Problems
abstract
Aniketh Garikaparthi, Manasi Patwardhan, Aditya Sanjiv Kanade, Aman Hassan, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Aniketh Garikaparthi, Manasi Patwardhan 0001, Aditya Kanade 0002, Aman Hassan, Lovekesh Vig, Arman Cohan
ACL (1)5
2025 Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
abstract
Peer review is fundamental to scientific research, but the growing volume of publications has intensified the challenges of this expertise-intensive process. While LLMs show promise in various scientific tasks, their potential to assist with peer review, particularly in identifying paper limitations, remains understudied. We first present a comprehensive taxonomy of limitation types in scientific research, with a focus on AI. Guided by this taxonomy, for studying limitations, we present LimitGen, the first comprehensive benchmark for evaluating LLMs’ capability to support early-stage feedback and complement human peer review. Our benchmark consists of two subsets: LimitGen-Syn, a synthetic dataset carefully created through controlled perturbations of high-quality papers, and LimitGen-Human, a collection of real human-written limitations. To improve the ability of LLM systems to identify limitations, we augment them with literature retrieval, which is essential for grounding identifying limitations in prior scientific findings. Our approach enhances the capabilities of LLM systems to generate limitations in research papers, enabling them to provide more concrete and constructive feedback.
Yilun Zhao 0001, Manasi Patwardhan 0001, Lovekesh Vig, Arman Cohan
ACL (1)4
2024 Generating Novel Leads for Drug Discovery Using LLMs with Logical Feedback
abstract
Large Language Models (LLMs) can be used as repositories of biological and chemical information to generate pharmacological lead compounds. However, for LLMs to focus on specific drug targets typically requires experimentation with progressively more refined prompts. Results thus become dependent not just on what is known about the target, but also on what is known about the prompt- engineering. In this paper, we separate the prompt into domain-constraints that can be written in a standard logical form and a simple text-based query. We investigate whether LLMs can be guided, not by refining prompts manually, but by refining the logical component automatically, keeping the query unchanged. We describe an iterative procedure LMLF (“Language Model with Logical Feedback”) in which the constraints are progressively refined using a logical notion of generalisation. On any iteration, newly generated instances are verified against the constraint, providing "logical-feedback" for the next iteration's refinement of the constraints. We evaluate LMLF using two well-known targets (inhibition of the Janus Kinase 2; and Dopamine Receptor D2); and two different LLMs (GPT-3 and PaLM). We show that LMLF, starting with the same logical constraints and query text, can be used to guide both LLMs to generate potential leads. We find: (a) Binding affinities of LMLF-generated molecules are skewed towards higher binding affinities than those from existing baselines; (b) LMLF results in generating molecules that are skewed towards higher binding affinities than without logical feedback; (c) Assessment by a computational chemist suggests that LMLF generated compounds may be novel inhibitors. These findings suggest that LLMs with logical feedback may provide a mechanism for generating new leads without requiring the domain-specialist to acquire sophisticated skills in prompt-engineering.
Shreyas Bhat Brahmavar, Ashwin Srinivasan 0001, Tirtharaj Dash, Sowmya Ramaswamy Krishnan, Lovekesh Vig, Arijit Roy 0003, Raviprasad Aduri
AAAI5
2024 SCM4SR: Structural Causal Model-based Data Augmentation for Robust Session-based Recommendation
abstract
With mounting privacy concerns, and movement towards a cookie-less internet, session-based recommendation (SR) models are gaining increasing popularity. The goal of SR models is to recommend top-K items to a user by utilizing information from past actions within a session. Many deep neural networks (DNN) based SR have been proposed in the literature, however, they experience performance declines in practice due to inherent biases (e.g., popularity bias) present in training data. To alleviate this, we propose an underlying neural-network (NN) based Structural Causal Model (SCM) which comprises an evolving user behavior (simulator) and recommendation model. The causal relations between the two sub-models and variables at consecutive timesteps are defined by a sequence of structural equations, whose parameters are learned using logged data. The learned SCM enables the simulation of a user's response on a counterfactual list of recommended items (slate). For this, we intervene on recommendation slates with counterfactual slates and simulate the user's response through learned SCM thereby generating counterfactual sessions to augment the training data. Through extensive empirical evaluation on simulated and real-world datasets, we show that the augmented data mitigates the impact of sparse training data and improves the performance of the SR models.
Muskan Gupta, Priyanka Gupta 0003, Jyoti Narwariya, Lovekesh Vig, Gautam Shroff
SIGIR4
2023 Do I have the Knowledge to Answer? Investigating Answerability of Knowledge Base Questions
abstract
Mayur Patidar, Prayushi Faldu, Avinash Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam -. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mayur Patidar, Prayushi Faldu, Avinash Kumar Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam
ACL (1)4
2023 Generalization of Fine Granular Extractions from Charts
Shubham Paliwal, Manasi Patwardhan 0001, Lovekesh Vig
ICDAR (2)3
2023 IKD+: Reliable Low Complexity Deep Models for Retinopathy Classification
abstract
Deep neural network (DNN) models for retinopathy have estimated predictive accuracies in the mid-to-high 90%. However, the following aspects remain unaddressed: State-of-the-art models are complex and require substantial computational infrastructure to train and deploy; The reliability of predictions can vary widely. In this paper, we focus on these aspects and propose a form of iterative knowledge distillation (IKD), called IKD+ that incorporates a tradeoff between size, accuracy and reliability. We investigate the functioning of IKD+ using two widely used techniques for estimating model calibration (Platt-scaling and temperature-scaling), using the best-performing model available, which is an ensemble of EfficientNets with approximately 100M parameters. We demonstrate that IKD+ equipped with temperature-scaling results in models that show up to approximately 500-fold decreases in the number of parameters than the original ensemble without a significant loss in accuracy. In addition, calibration scores (reliability) for the IKD+ models are as good as or better than the base model.
Shreyas Bhat Brahmavar, Rohit Rajesh, Tirtharaj Dash, Lovekesh Vig, Tanmay T. Verlekar, Tariq Mahmood Khan, Erik Meijering, Ashwin Srinivasan 0001
ICIP4
2023 Calibrating Deep Neural Networks using Explicit Regularisation and Dynamic Data Pruning
abstract
Deep neural networks (DNNS) are prone to miscalibrated predictions, often exhibiting a mismatch between the predicted output and the associated confidence scores. Contemporary model calibration techniques mitigate the problem of overconfident predictions by pushing down the confidence of the winning class while increasing the confidence of the remaining classes across all test samples. However, from a deployment perspective an ideal model is desired to (i) generate well calibrated predictions for high-confidence samples with predicted probability say > 0.95 and (ii) generate a higher proportion of legitimate high-confidence samples. To this end, we propose a novel regularization technique that can be used with classification losses, leading to state-of-the-art calibrated predictions at test time; From a deployment standpoint in safety critical applications, only high-confidence samples from a well-calibrated model are of interest, as the remaining samples have to undergo manual inspection. Predictive confidence reduction of these potentially "high-confidence samples" is a downside of existing calibration approaches. We mitigate this via proposing a dynamic traintime data pruning strategy which prunes low confidence samples every few epochs, providing an increase in confident yet calibrated samples. We demonstrate state-of-the-art calibration performance across image classification benchmarks, reducing training time without much compromise in accuracy. We provide insights into why our dynamic pruning strategy that prunes low confidence training samples leads to an increase in high-confidence samples at test time.
Rishabh Patra, Ramya Hebbalaguppe, Tirtharaj Dash, Gautam Shroff, Lovekesh Vig
WACV5
2022 Solving Visual Analogies Using Neural Algorithmic Reasoning (Student Abstract)
abstract
We consider a class of visual analogical reasoning problems that involve discovering the sequence of transformations by which pairs of input/output images are related, so as to analogously transform future inputs. This program synthesis task can be easily solved via symbolic search. Using a variation of the ‘neural analogical reasoning’ approach, we instead search for a sequence of elementary neural network transformations that manipulate distributed representations derived from a symbolic space, to which input images are directly encoded. We evaluate the extent to which our ‘neural reasoning’ approach generalises for images with unseen shapes and positions.
Atharv Sonwane, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001, Tirtharaj Dash
AAAI3
2022 Intent Detection and Discovery from User Logs via Deep Semi-Supervised Contrastive Clustering
abstract
Rajat Kumar, Mayur Patidar, Vaibhav Varshney, Lovekesh Vig, Gautam Shroff. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Mayur Patidar, Vaibhav Varshney, Lovekesh Vig, Gautam Shroff
NAACL-HLT4
2021 DKMA-ULD: Domain Knowledge augmented Multi-head Attention based Robust Universal Lesion Detection
Manu Sheoran, Meghal Dani, Lovekesh Vig
BMVC4
2021 CauSeR: Causal Session-based Recommendations for Handling Popularity Bias
abstract
Recommender Systems (RS) tend to recommend more popular items instead of the relevant long-tail items. Mitigating such popularity bias is crucial to ensure that less popular but relevant items are part of the recommendation list shown to the user. In this work, we study the phenomenon of popularity bias in session-based RS (SRS) obtained via deep learning (DL) models. We observe that DL models trained on the historical user-item interactions in session logs (having long-tailed item-click distributions) tend to amplify popularity bias. To understand the source of this bias amplification, we consider potential sources of bias at two distinct stages in the modeling process: i. the data-generation stage (user-item interactions captured as session logs), ii. the DL model training stage. We highlight that the popularity of an item has a causal effect on i. user-item interactions via conformity bias, as well as ii. item ranking from DL models via biased training process due to class (target item) imbalance. While most existing approaches in literature address only one of these effects, we consider a comprehensive causal inference framework that identifies and mitigates the effects at both stages. Through extensive empirical evaluation on simulated and real-world datasets, we show that our approach improves upon several strong baselines from literature for popularity bias and long-tailed classification. Ablation studies show the advantage of our comprehensive causal analysis to identify and handle bias in data generation as well as training stages.
Priyanka Gupta 0003, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff
CIKM4
2021 TSR-DSAW: Table Structure Recognition via Deep Spatial Association of Words
abstract
Existing methods for Table Structure Recognition (TSR) from camera-captured or scanned documents perform poorly on complextables consisting of nested rows / columns, multi-line texts and missing cell data.This is because current data-driven methods work by simply training deep models on large volumes of data and fail to generalize when an unseen table structure is encountered.In this paper, we propose to train a deep network to capture the spatial associations between different word pairs present in the table image for unravelling the table structure.We present an end-to-end pipeline, named TSR-DSAW: TSR via Deep Spatial Association of Words, which outputs a digital representation of a table image in a structured format such as HTML.Given a table image as input, the proposed method begins with the detection of all the words present in the image using a text-detection network like CRAFT which is followed by the generation of word-pairs using dynamic programming.These word-pairs are highlighted in individual images and subsequently, fed into a DenseNet-121 classifier trained to capture spatial associations such as same-row, same-column, same-cell or none.Finally, we perform post-processing on the classifier output to generate the table structure in HTML format.We evaluate our TSR-DSAW pipeline on two public table-image datasets -PubTabNet and ICDAR 2013, and demonstrate improvement over previous methods such as TableNet and DeepDeSRT.
Shubham Paliwal, Lovekesh Vig
ESANN4
2021 Data-Efficient Training of High-Resolution Images in Medical Domain
abstract
The ability of Graphical Processor Units (GPUs) to quickly train dataand compute-intensive deep networks has led to rapid advancements across diverse domains such as robotics, medical imaging and autonomous driving.However, memory constraints with GPU-based training for memory-intensive deep networks have forced researchers to adopt various workarounds: 1) resize the input image, 2) divide input image into smaller patches, or use smaller batch-sizes in order to fit both the model and batch training data into GPU memory.While these alternatives perform well when dealing with natural images, they suffer from 1) loss of highresolution information, 2) loss of global context and 3) sub-optimal batch sizes.Such issues will likely to become more pressing for domains like medical imaging, where data is scarce and images are often of very high resolution with subtle features.Therefore, in this paper, we demonstrate that training can be made more data-efficient by using a distributed training setup with high-resolution images and larger effective batch sizes, with batches being distributed across multiple nodes.The distributed GPU training framework, which partitions the data and only shares model parameters across different GPUs, gets around the memory constraints of single GPU training.We conduct a study in which experiments are performed for different image resolutions (ranging from 112 × 112 to 1024 × 1024) and different number of images per class to determine the effect of image resolutions on network performance.We illustrate our findings on two medical imaging datasets namely, SD-198 skin-lesion and NIH Chest X-rays.
Shruti Kunde, Amey Pandit, Kushagra Mahajan, Rekha Singhal, Lovekesh Vig
ESANN6
2021 Continual Learning for Multivariate Time Series Tasks with Variable Input Dimensions
abstract
We consider a sequence of related multivariate time series learning tasks, such as predicting failures for different instances of a machine from time series of multi-sensor data, or activity recognition tasks over different individuals from multiple wearable sensors. We focus on two under-explored practical challenges arising in such settings: (i) Each task may have a different subset of sensors, i.e., providing different partial observations of the underlying ‘system’. This restriction can be due to different manufacturers in the former case, and people wearing more or less measurement devices in the latter (ii) We are not allowed to store or re-access data from a task once it has been observed at the task level. This may be due to privacy considerations in the case of people, or legal restrictions placed by machine owners. Nevertheless, we would like to (a) improve performance on subsequent tasks using experience from completed tasks as well as (b) continue to perform better on past tasks, e.g., update the model and improve predictions on even the first machine after learning from subsequently observed ones. We note that existing continual learning methods do not take into account variability in input dimensions arising due to different subsets of sensors being available across tasks, and struggle to adapt to such variable input dimensions (VID) tasks. In this work, we address this shortcoming of existing methods. To this end, we learn task-specific generative models and classifiers, and use these to augment data for target tasks. Since the input dimensions across tasks vary, we propose a novel conditioning module based on graph neural networks to aid a standard recurrent neural network. We evaluate the efficacy of the proposed approach on three publicly available datasets corresponding to two activity recognition tasks (classification) and one prognostics task (regression). We demonstrate that it is possible to significantly enhance the performance on future and previous tasks while learning continuously from VID tasks without storing data.
Vibhor Gupta, Jyoti Narwariya, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff
ICDM4
2021 BCQ4DCA: Budget Constrained Deep Q-Network for Dynamic Campaign Allocation in Computational Advertising
abstract
Digital advertising companies typically conduct several advertising campaigns in parallel while being constrained by a fixed overall advertising budget. This gives rise to the problem of distributing the budget across the different campaigns dynamically so as to optimize the overall return on investment (ROI) (or some other metric) within a specified time duration. In this paper, we propose an RL formulation called BCQ4DCA for dynamic optimization of budget-constrained campaign allocation. The formulation is model-free and uses a novel cumulative reward model that is learned alongside a Deep Q-Network. We utilize a real-world Criteo user interaction dataset to evaluate BCQ4DCA in terms of conversion rate, budget utilization, cost per conversion, and ROI, finding that it outperforms current heuristic, attribution based approaches like DARNN, DNAMTA across a common time window.
Manasi Malik, Garima Gupta, Lovekesh Vig, Gautam Shroff
IJCNN3
2021 OSSR-PID: One-Shot Symbol Recognition in P&ID Sheets using Path Sampling and GCN
abstract
In this paper, we focus on recognition of line-drawn symbols in engineering drawings with only one prototypical example per symbol available for training. In particular, Piping and Instrumentation Diagrams (P&ID) are ubiquitous in several manufacturing, oil and gas enterprises for representing engineering schematics and equipment layout. There is an urgent need to extract and digitize information from P&IDs without the cost of annotating a varying set of symbols for each new use case. A robust one-shot learning approach for symbol recognition i.e., localization followed by classification, would therefore go a long way towards this goal. Our method works by sampling pixels sequentially along the different contour boundaries in the image. These sampled points form paths which are used in the prototypical line diagram to construct a graph that captures the structure of the contours. Subsequently, the prototypical graphs are fed into a Dynamic Graph Convolutional Neural Network (DGCNN) which is trained to classify graphs into one of the given symbol classes. Further, we append embeddings from a Resnet-34 network which is trained on symbol images containing sampled points to make the classification network more robust. Since, many symbols in P&ID are structurally very similar to each other, we utilize Arcface loss during DGCNN training which helps in maximizing symbol class separability by producing highly discriminative embeddings. During inference time, a similar line based sampling procedure is adopted for generating sampled points across P&ID image. The images consist of components attached on the pipeline (straight line). The sampled points segregated around the symbol regions are used for the classification task. The proposed pipeline, named OSSR-PID, is fast and gives outstanding performance for recognition of symbols on a synthetic dataset of 100 P&ID diagrams. We also compare our method against prior-work that uses full supervision (not one-shot) on a real-world private dataset of 12 P&ID sheets and obtain comparable/superior results. Remarkably, it is able to achieve such excellent performance using only one prototypical example per symbol.
Shubham Paliwal, Lovekesh Vig
IJCNN3
2021 Using Domain-Knowledge to Assist Lead Discovery in Early-Stage Drug Design
Tirtharaj Dash, Ashwin Srinivasan 0001, Lovekesh Vig, Arijit Roy 0003
ILP3
2021 Incorporating symbolic domain knowledge into graph neural networks
Tirtharaj Dash, Ashwin Srinivasan 0001, Lovekesh Vig
Mach. Learn.3
2020 MultiMBNN: Matched and Balanced Causal Inference with Neural Networks
Garima Gupta, Ranjitha Prasad, Lovekesh Vig, Gautam Shroff
ESANN5
2020 Capsule Based Neural Network Architecture to perform completeness check for Patent Eligibility Process
abstract
In the process of filing patents, attorneys need to ask many questions, to the inventors, to ascertain patent eligibility. We propose to ease up such conversation through a deep learning-based system. This system can automatically check whether all key ingredients required for checking the patent eligibility are present in technical write-up shared by the inventors. If not, the inventors can provide the missing information. We present a trainable model to identify various ingredients such as the objective, motivation, new observation, etc. from research articles. We model this as a sentence classification problem, which is a difficult task because a patent can be filed in any domain, and sentences involved can often be very long. To this end, we propose a dilated LSTM and capsule-based neural network architecture. We present experimental results of the proposed model on a real-world patent dataset covering patent applications in diverse domains in which our organization is carrying out research and innovation activities, and also three publicly available sentence classification datasets. Through empirical analysis, we show that a) Our model performs significantly better than several strong baselines on the patent dataset; b) Performing dilation operation on LSTMs allows us to capture long term dependencies; c) our model is comparable to existing state-of-art approaches on the publicly available datasets; d) Error analysis through LIME shows that the proposed approach can help patent attorneys to interpret the decisions taken by the classifier.
Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Vidya Vikas
IJCNN4
2020 Constructing generative logical models for optimisation problems using domain knowledge
Ashwin Srinivasan 0001, Lovekesh Vig, Gautam Shroff
Mach. Learn.2
2019 Regularizing Fully Convolutional Networks for Time Series Classification by Decorrelating Filters
abstract
Deep neural networks are prone to overfitting, especially in small training data regimes. Often, these networks are overparameterized and the resulting learned weights tend to have strong correlations. However, convolutional networks in general, and fully convolution neural networks (FCNs) in particular, have been shown to be relatively parameter efficient, and have recently been successfully applied to time series classification tasks. In this paper, we investigate the application of different regularizers on the correlation between the learned convolutional filters in FCNs using Batch Normalization (BN) as a regularizer for time series classification (TSC) tasks. Results demonstrate that despite orthogonal initialization of the filters, the average correlation across filters (especially for filters in higher layers) tends to increase as training proceeds, indicating redundancy of filters. To mitigate this redundancy, we propose a strong regularizer, using simple yet effective filter decorrelation. Our proposed method yields significant gains in classification accuracy for 44 diverse time series datasets from the UCR TSC benchmark repository.
Kaushal Paneri, Vishnu TV, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff
AAAI4
2019 Fusing Features based on Signal Properties and TimeNet for Time Series Classification
Arijit Ukil, Pankaj Malhotra, Soma Bandyopadhyay, Tulika Bose, Ishan Sahu, Ayan Mukherjee, Lovekesh Vig, Arpan Pal 0001, Gautam Shroff
ESANN7
2019 TableNet: Deep Learning Model for End-to-end Table Detection and Tabular Data Extraction from Scanned Document Images
abstract
With the widespread use of mobile phones and scanners to photograph and upload documents, the need for extracting the information trapped in unstructured document images such as retail receipts, insurance claim forms and financial invoices is becoming more acute. A major hurdle to this objective is that these images often contain information in the form of tables and extracting data from tabular sub-images presents a unique set of challenges. This includes accurate detection of the tabular region within an image, and subsequently detecting and extracting information from the rows and columns of the detected table. While some progress has been made in table detection, extracting the table contents is still a challenge since this involves more fine grained table structure (rows & columns) recognition. Prior approaches have attempted to solve the table detection and structure recognition problems independently using two separate models. In this paper, we propose TableNet: a novel end-to-end deep learning model for both table detection and structure recognition. The model exploits the interdependence between the twin tasks of table detection and table structure recognition to segment out the table and column regions. This is followed by semantic rule-based row extraction from the identified tabular sub-regions. The proposed model and extraction approach was evaluated on the publicly available ICDAR 2013 and Marmot Table datasets obtaining state of the art results. Additionally, we demonstrate that feeding additional semantic features further improves model performance and that the model exhibits transfer learning across datasets. Another contribution of this paper is to provide additional table structure annotations for the Marmot data, which currently only has annotations for table detection.
Shubham Paliwal, Vishwanath D, Rohit Rahul, Lovekesh Vig
ICDAR5
2019 Automatic Information Extraction from Piping and Instrumentation Diagrams
abstract
One of the most common modes of representing engineering schematics are Piping and Instrumentation diagrams (P&IDs) that describe the layout of an engineering process flow along with the interconnected process equipment. Over the years, P&ID diagrams have been manually generated, scanned and stored as image files. These files need to be digitized for purposes of inventory management and updation, and easy reference to different components of the schematics. There are several challenging vision problems associated with digitizing real world P&ID diagrams. Real world P&IDs come in several different resolutions, and often contain noisy textual information. Extraction of instrumentation information from these diagrams involves accurate detection of symbols that frequently have minute visual differences between them. Identification of pipelines that may converge and diverge at different points in the image is a further cause for concern. Due to these reasons, to the best of our knowledge, no system has been proposed for end-to-end data extraction from P&ID diagrams. However, with the advent of deep learning and the spectacular successes it has achieved in vision, we hypothesized that it is now possible to re-examine this problem armed with the latest deep learning models. To that end, we present a novel pipeline for information extraction from P&ID sheets via a combination of traditional vision techniques and state-of-the-art deep learning models to identify and isolate pipeline codes, pipelines, inlets and outlets, and for detecting symbols. This is followed by association of the detected components with the appropriate pipeline. The extracted pipeline information is used to populate a tree-like data-structure for capturing the structure of the piping schematics. We evaluated proposed method on a real world dataset of P&ID sheets obtained from an oil firm and have obtained promising results.
Rohit Rahul, Shubham Paliwal, Lovekesh Vig
ICPRAM4
2019 ConvTimeNet: A Pre-trained Deep Convolutional Neural Network for Time Series Classification
abstract
Training deep neural networks often requires careful hyper-parameter tuning and significant computational resources. In this paper, we propose ConvTimeNet (CTN): an off-the-shelf deep convolutional neural network (CNN) trained on diverse univariate time series classification (TSC) source tasks. Once trained, CTN can be easily adapted to new TSC target tasks via a small amount of fine-tuning using labeled instances from the target tasks. We note that the length of convolutional filters is a key aspect when building a pre-trained model that can generalize to time series of different lengths across datasets. To achieve this, we incorporate filters of multiple lengths in all convolutional layers of CTN to capture temporal features at multiple time scales. We consider all 65 datasets with time series of lengths up to 512 points from the UCR TSC Benchmark for training and testing transferability of CTN: We train CTN on a randomly chosen subset of 24 datasets using a multi-head approach with a different softmax layer for each training dataset, and study generalizability and transferability of the learned filters on the remaining 41 TSC datasets. We observe significant gains in classification accuracy as well as computational efficiency when using pre-trained CTN as a starting point for subsequent task-specific fine-tuning compared to existing state-of-the-art TSC approaches. We also provide qualitative insights into the working of CTN by: i) analyzing the activations and filters of first convolution layer suggesting the filters in CTN are generically useful, ii) analyzing the impact of the design decision to incorporate multiple length decisions, and iii) finding regions of time series that affect the final classification decision via occlusion sensitivity analysis.
Kathan Kashiparekh, Jyoti Narwariya, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff
IJCNN4
2019 ChartNet: Visual Reasoning over Statistical Charts using MAC-Networks
abstract
Despite the improvements in perception accuracies brought about via deep learning, developing systems combining accurate visual perception with the ability to reason over the visual percepts remains extremely challenging. A particular application area of interest from an accessibility perspective is that of reasoning over statistical charts such as bar and pie charts. To this end, we formulate the problem of reasoning over statistical charts as a classification task using MAC-Networks to give answers from a predefined vocabulary of generic answers. Additionally, we enhance the capabilities of MAC-Networks to give chart-specific answers to open-ended questions by replacing the classification layer by a regression layer to localize the textual answers present over the images. We call our network ChartNet, and demonstrate its efficacy on predicting both in vocabulary and out of vocabulary answers. To test our methods, we generated our own dataset of statistical chart images and corresponding question answer pairs. Results show that ChartNet consistently outperform other state-of-the-art methods on reasoning over these questions and may be a viable candidate for applications containing images of statistical charts.
Arindam Chowdhury, Lovekesh Vig
IJCNN4
2019 Hierarchical Capsule Based Neural Network Architecture for Sequence Labeling
abstract
Sequence Labeling is one of the most prominent tasks in NLP. The traditional text classification models do not carry context from one sentence to another and hence may not perform well on these tasks. These models lack a hierarchical structure that can aid them in dissecting the input structure at different levels to allow flow of context between sentences. In this paper, we propose a hierarchical neural network comprising of Bi-LSTMs, Dilated Convolution operation, Capsules and Conditional Random Field (CRF) to understand the discourse/ abstract structure and predict next probable label by using label history. We have performed experiments on 3 publicly available datasets through which we have demonstrated that our model has achieved state-of-art performance on these datasets.
Puneet Agarwal, Gautam Shroff, Lovekesh Vig
IJCNN4
2019 Meta-Learning for Black-Box Optimization
Vishnu TV, Pankaj Malhotra, Jyoti Narwariya, Lovekesh Vig, Gautam Shroff
ECML/PKDD (2)4
2019 Sequence and Time Aware Neighborhood for Session-based Recommendations: STAN
abstract
Recent advances in sequence-aware approaches for session-based recommendation, such as those based on recurrent neural networks, highlight the importance of leveraging sequential information from a session while making recommendations. Further, a session based k-nearest-neighbors approach (SKNN) has proven to be a strong baseline for session-based recommendations. However, SKNN does not take into account the readily available sequential and temporal information from sessions. In this work, we propose Sequence and Time Aware Neighborhood (STAN), with vanilla SKNN as its special case. STAN takes into account the following factors for making recommendations: i) position of an item in the current session, ii) recency of a past session w.r.t. to the current session, and iii) position of a recommendable item in a neighboring session. The importance of above factors for a specific application can be adjusted via controllable decay factors. Despite being simple, intuitive and easy to implement, empirical evaluation on three real-world datasets shows that STAN significantly improves over SKNN, and is even comparable to the recently proposed state-of-the-art deep learning approaches. Our results suggest that STAN can be considered as a strong baseline for evaluating session-based recommendation algorithms in future.
Diksha Garg, Priyanka Gupta 0003, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff
SIGIR4
2019 Logical Explanations for Deep Relational Machines Using Relevance Information
abstract
Our interest in this paper is in the construction of symbolic explanations for predictions made by a deep neural network. We will focus attention on deep relational machines (DRMs: a term introduced in Lodhi (2013)). A DRM is a deep network in which the input layer consists of Boolean-valued functions (features) that are defined in terms of relations provided as domain, or background, knowledge. Our DRMs differ from those in Lodhi (2013), which uses an Inductive Logic Programming (ILP) engine to first select features (we use random selections from a space of features that satisfies some approximate constraints on logical relevance and non-redundancy). But why do the DRMs predict what they do? One way of answering this was provided in recent work Ribeiro et al. (2016), by constructing readable proxies for a black-box predictor. The proxies are intended only to model the predictions of the black-box in local regions of the instance-space. But readability alone may not be enough: to be understandable, the local models must use relevant concepts in an meaningful manner. We investigate the use of a Bayes-like approach to identify logical proxies for local predictions of a DRM. As a preliminary step, we show that DRM's with our randomised propositionalization method achieve predictive performance that is comparable to the best reports in the ILP literature. Our principal results on logical explanations show: (a) Models in first-order logic can approximate the DRM's prediction closely in a small local region; and (b) Expert-provided relevance information can play the role of a prior to distinguish between logical explanations that perform equivalently on prediction alone.
Ashwin Srinivasan 0001, Lovekesh Vig, Michael Bain 0001
J. Mach. Learn. Res.2
2018 An Efficient End-to-End Neural Model for Handwritten Text Recognition
Arindam Chowdhury, Lovekesh Vig
BMVC2
2018 Automatic Conversational Helpdesk Solution using Seq2Seq and Slot-filling Models
abstract
Helpdesk is a key component of any large IT organization, where users can log a ticket about any issue they face related to IT infrastructure, administrative services, human resource services, etc. Normally, users have to assign appropriate set of labels to a ticket so that it could be routed to right domain expert who can help resolve the issue. In practice, the number of labels are very large and organized in form of a tree. It is non-trivial to describe the issue completely and attach appropriate labels unless one knows the cause of the problem and the related labels. Sometimes domain experts discuss the issue with the users and change the ticket labels accordingly, without modifying the ticket description. This results in inconsistent and badly labeled data, making it hard for supervised algorithms to learn from. In this paper, we propose a novel approach of creating a conversational helpdesk system, which will ask relevant questions to the user, for identification of the right category and will then raise a ticket on users' behalf. We use attention based seq2seq model to assign the hierarchical categories to tickets. We use a slot filling model to help us decide what questions to ask to the user, if the top-k model predictions are not consistent. We also present a novel approach to generate training data for the slot filling model automatically based on attention in the hierarchical classification model. We demonstrate via a simulated user that the proposed approach can give us a significant gain in accuracy on ticket-data without asking too many questions to users. Finally, we also show that our seq2seq model is as versatile as other approaches on publicly available datasets, as state of the art approaches.
Mayur Patidar, Puneet Agarwal, Lovekesh Vig, Gautam Shroff
CIKM3
2018 Evolutionary RL for Container Loading
Sarmimala Saikia, Richa Verma, Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001
ESANN5
2018 Automatic Chromosome Classification using Deep Attention Based Sequence Learning of Chromosome Bands
abstract
Chromosome Karyotyping refers to the task of segmenting and classifying individual chromosome images obtained from stained cell images microphotographed during the metaphase stage of cell division. The karyotyped images are useful for diagnosis of genetic disorders such as down syndrome, turner syndrome and certain types of cancers. In many hospitals and labs, a significant amount of manual effort and time is spent on segmenting and classifying the individual chromosome images. Recently, deep learning models have been applied to automate this task with promising results. An important characteristic of a chromosome is the presence of sequence of dark and light bands produced by giemsa staining which is used by cytogeneticists to manually perform karyotyping. We propose Residual Convolutional Recurrent Attention Neural Network (Res-CRANN) which exploits this property of band sequence for chromo-some classification. Res-CRANN is end-to-end trainable in which a sequence of feature vectors, extracted from the feature maps produced by convolutional layers of Residual neural networks (ResNet) is fed into Recurrent Neural Networks (RNN) and subsequently, an attention mechanism is applied on top of RNN output sequences which are further classified into one of the 24 labels. The attention mechanism after recurrent layers facilitates the network to learn to pay selective attention to the sequence of bands and relate them to different classes of chromosomes. We demonstrate the proposed architecture's efficacy on a publicly available Bioimage chromosome classification dataset and observe that our model outperforms the baseline models created using traditional deep convolutional neural network and ResNet-50 by approximately 3% Top-1 classification accuracy.
Lovekesh Vig
IJCNN3
2018 Large-Scale Assessment of Deep Relational Machines
Tirtharaj Dash, Ashwin Srinivasan 0001, Lovekesh Vig, Oghenejokpeme I. Orhobor, Ross D. King
ILP3
2018 Resolving Abstract Anaphora Implicitly in Conversational Assistants using a Hierarchically stacked RNN
abstract
Recent proliferation of conversational systems has resulted in an increased demand for more natural dialogue systems, capable of more sophisticated interactions than merely providing factual answers. This is evident from usage pattern of a conversational system deployed within our organization. Users not only expect it to perform co-reference resolution of anaphora, but also of the antecedent or posterior facts presented by users with respect to their query. Presence of such facts in a conversation sometimes modifies the answer of main query, e.g., answer to 'how many sick leave do I get?' would be different when a fact 'I am on contract' is also present. Sometimes there is a need to collectively resolve three or four such facts. In this paper, we propose a novel solution which uses hierarchical neural network, comprising of BiLSTM layer and a maxpool layer that is hierarchically stacked to first obtain a representation of each user utterance and then to obtain a representation for sequence of utterances. This representation is used to identify users' intention. We also improvise this model by using skip connections in the second network to allow better gradient flow. Our model, not only a)~resolves the antecedent and posterior facts, but also b)~performs better even on self-contained queries. It is also c)~faster to train, making it the most promising approach for use in our environment where frequent training and tuning is needed. It slightly outperforms the benchmark on a publicly available dataset, and e)~performs better than obvious baselines approaches on our datasets.
Prerna Khurana, Puneet Agarwal, Gautam Shroff, Lovekesh Vig
KDD4
2017 Hybrid BiLSTM-Siamese network for FAQ Assistance
abstract
We describe an automated assistant for answering frequently asked questions; our system has been deployed, and is currently answering HR-related queries in two different areas (leave management and health insurance) to a large number of users. The needs of a large global corporate lead us to model a frequently asked question (FAQ) to be an equivalence class of actually asked questions, for which there is a common answer (certified as being consistent with the organization's policy). When a new question is posed to our system, it finds the class of question, and responds with the answer for the class. At this point, the system is either correct (gives correct answer); or incorrect (gives wrong answer); or incomplete (says "I don't know''). We employ a hybrid deep-learning architecture in which a BiLSTM-based classifier is combined with second BiLSTM-based Siamese network in an iterative manner: Questions for which the classifier makes an error during training are used to generate a set of misclassified question-question pairs. These, along with correct pairs, are used to train the Siamese network to drive apart the (hidden) representations of the misclassified pairs. We present experimental results from our deployment showing that our iteratively trained hybrid network: (a) results in better performance than using just a classifier network, or just a Siamese network; (b) performs better than state-of-the art sentence classifiers in the two areas in which it has been deployed, in terms of both accuracy as well as precision-recall tradeoff; and (c) also performs well on a benchmark public dataset. We also observe that using question-question pairs in our hybrid network, results in marginally better performance than using question-to-answer pairs. Finally, estimates of precision and recall from the deployment of our automated assistant suggest that we can expect the burden on our HR department to drop from answering about 6000 queries a day to about 1000.
Prerna Khurana, Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001
CIKM4
2017 Learning and Knowledge Transfer with Memory Networks for Machine Comprehension
abstract
Enabling machines to read and comprehend unstructured text remains an unfulfilled goal for NLP research.Recent research efforts on the "machine comprehension" task have managed to achieve close to ideal performance on simulated data.However, achieving similar levels of performance on small real world datasets has proved difficult; major challenges stem from the large vocabulary size, complex grammar, and the frequent ambiguities in linguistic structure.On the other hand, the requirement of human generated annotations for training, in order to ensure a sufficiently diverse set of questions is prohibitively expensive.Motivated by these practical issues, we propose a novel curriculum inspired training procedure for Memory Networks to improve the performance for machine comprehension with relatively small volumes of training data.Additionally, we explore various training regimes for Memory Networks to allow knowledge transfer from a closely related domain having larger volumes of labelled data.We also suggest the use of a loss function to incorporate the asymmetric nature of knowledge transfer.Our experiments demonstrate improvements on Dailymail, CNN, and MCTest datasets.
Lovekesh Vig, Gautam Shroff
EACL (1)2
2017 TimeNet: Pre-trained deep recurrent neural network for time series classification
Pankaj Malhotra, Vishnu TV, Lovekesh Vig, Puneet Agarwal, Gautam Shroff
ESANN3
2017 Information Bottleneck Inspired Method For Chat Text Segmentation
abstract
We present a novel technique for segmenting chat conversations using the information bottleneck method (Tishby et al., 2000), augmented with sequential continuity constraints. Furthermore, we utilize critical non-textual clues such as time between two consecutive posts and people mentions within the posts. To ascertain the effectiveness of the proposed method, we have collected data from public Slack conversations and Fresco, a proprietary platform deployed inside our organization. Experiments demonstrate that the proposed method yields an absolute (relative) improvement of as high as 3.23% (11.25%). To facilitate future research, we are releasing manual annotations for segmentation on public Slack conversations.
Sunder Vishal, Lovekesh Vig, Gautam Shroff
IJCNLP(1)3
2017 An Investigation into the Role of Domain-Knowledge on the Use of Embeddings
Lovekesh Vig, Ashwin Srinivasan 0001, Michael Bain 0001
ILP1
2016 Automatic Container Code Recognition via Spatial Transformer Networks and Connected Component Region Proposals
abstract
Container identification and recognition is still performed manually or in a semi-automatic fashion in multiple ports globally. This results in errors and inefficiencies in port operations. The problem of automatic container identification and recognition is challenging as the ISO standard only prescribes the pattern of the code and does not specify other parameters such as the foreground and background colors, font type and size, orientation of characters (horizontal or vertical) so on. Additionally, the corrugated surface of container body makes the two dimensional projection of the text on three dimensional containers slanted and jagged. We propose a solution in the form of an end-to-end pipeline that uses Region Proposals generated based on Connected Components for text detection in conjunction with Spatial Transformer NeConnected Componentstworks for text recognition. We demonstrate via our experimental results that the pipeline is reliable and robust even in situations when the code characters are highly distorted and outperforms the state-of-the-art results for text detection and recognition over the containers. We achieve text coverage rate of 100% and text recognition rate of 99.64%.
Ramya Hebbalaguppe, Ehtesham Hassan, Lovekesh Vig
ICMLA5
2016 Generation of Near-Optimal Solutions Using ILP-Guided Sampling
Ashwin Srinivasan 0001, Gautam Shroff, Lovekesh Vig, Sarmimala Saikia
ILP3
2015 Anomaly detection in ECG time signals via deep long short-term memory networks
abstract
Electrocardiography (ECG) signals are widely used to gauge the health of the human heart, and the resulting time series signal is often analyzed manually by a medical professional to detect any arrhythmia that the patient may have suffered. Much work has been done to automate the process of analyzing ECG signals, but most of the research involves extensive preprocessing of the ECG data to derive vectorized features and subsequently designing a classifier to discriminate between healthy ECG signals and those indicative of an Arrhythmia. This approach requires knowledge and data of the different types of Arrhythmia for training. However, the heart is a complex organ and there are many different and new types of Arrhythmia that can occur which were not part of the original training set. Thus, it may be more prudent to adopt an anomaly detection approach towards analyzing ECG signals. In this paper, we utilize a deep recurrent neural network architecture with Long Short Term Memory (LSTM) units to develop a predictive model for healthy ECG signals. We further utilize the probability distribution of the prediction errors from these recurrent models to indicate normal or abnormal behavior. An added advantage of using LSTM networks is that the ECG signal can be directly fed into the network without any elaborate preprocessing as required by other techniques. Also, no prior information about abnormal signals is needed by the networks as they were trained only on normal data. We have used the MIT-BIH Arrhythmia Database to obtain ECG time series data for both normal periods and for periods during four different types of Arrhythmias, namely Premature Ventricular Contraction (PVC), Atrial Premature Contraction (APC), Paced Beats (PB) and Ventricular Couplet (VC). Results are promising and indicate that Deep LSTM models may be viable for detecting anomalies in ECG signals.
Sucheta Chauhan, Lovekesh Vig
DSAA2
2015 Ensemble of deep long short term memory networks for labelling origin of replication sequences
abstract
Advancement in sequence data generation technologies are churning out voluminous omics data and posing a massive challenge to annotate the biological functional features. Sequence data from the well studied model organism Saccharomyces cerevisiae has been commonly used to test and validate in silico prediction methods. DNA replication is a critical step in the cellular process and the sequence location where this process originates in the genomic landscape is generally referred as origin of replication. In this paper we investigate the application bidirectional Long Short Term (LSTM) Networks to predict origin of replication sequences. Long Short Term Memory (LSTM) networks have recently been shown to yield state of the art performance in speech recognition, and music generation. These networks are capable of learning long term patterns via the use of multiplication gates. This paper utilizes Deep bidirectional LSTM for prediction of origin of replication sequences belonging to the organism Saccharomyces cerevisiae. Results demonstrate that LSTMs outperform the commonly used machine learning classifiers such as Support Vector Machine (SVM), Random Forest (RF), Artificial Neural Network (ANN), and Hidden Markov Model (HMM). An important additional advantage of LSTMs is that they work directly on the sequences and obviate the need for hand coded features.
Urminder Singh, Sucheta Chauhan, A. Krishnamachari, Lovekesh Vig
DSAA4
2015 Long Short Term Memory Networks for Anomaly Detection in Time Series
Pankaj Malhotra, Lovekesh Vig, Gautam Shroff, Puneet Agarwal
ESANN2
2015 Parallel multi-objective multi-robot coalition formation
Manoj Agarwal, Nitin Agrawal 0002, Lovekesh Vig, Naveen Kumar 0001
Expert Syst. Appl.4
2014 Non-additive multi-objective robot coalition formation
Manoj Agarwal, Naveen Kumar 0001, Lovekesh Vig
Expert Syst. Appl.3
2011 A dual association model for the extinction of animal conditioning
Ashish Gupta 0005, Lovekesh Vig, David C. Noelle
Neurocomputing2
2006 Multi-robot coalition formation
abstract
As the community strives towards autonomous multi-robot systems, there is a need for these systems to autonomously form coalitions to complete assigned missions. Numerous coalition formation algorithms have been proposed in the software agent literature. Algorithms exist that form agent coalitions in both super additive and non-super additive environments. The algorithmic techniques vary from negotiation-based protocols in multi-agent system (MAS) environments to those based on computation in distributed problem solving (DPS) environments. Coalition formation behaviors have also been discussed in relation to game theory. Despite the plethora of MAS coalition formation literature, to the best of our knowledge none of the proposed algorithms have been demonstrated with an actual multi-robot system. There exists a discrepancy between the multi-agent algorithms and their applicability to the multi-robot domain. This paper aims to bridge that discrepancy by unearthing the issues that arise while attempting to tailor these algorithms to the multi-robot domain. A well-known multi-agent coalition formation algorithm has been studied in order to identify the necessary modifications to facilitate its application to the multi-robot domain. This paper reports multi-robot coalition formation results based upon simulation and actual robot experiments. A multi-agent coalition formation algorithm has been demonstrated on an actual robot system.
Lovekesh Vig, Julie A. Adams
IEEE Trans. Robotics1