EDBT 2026 Demo / reviewers in the wild / expert
Gautam Shroff
dblp:54/3265 · also Gautam M. Shroff
· DBLP profile ↗
51ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-0340-0283ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 9 since 2021Databases, data management, data science and information retrieval · 23 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Theory of computation · 4 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)abstractThe Abstraction and Reasoning Corpus (ARC) poses a significant challenge to artificial intelligence, demanding broad generalization and few-shot learning capabilities that remain elusive for current deep learning methods, including large language models (LLMs). While LLMs excel in program synthesis, their direct application to ARC yields limited success. To address this, we introduce ConceptSearch, a novel function-search algorithm that leverages LLMs for program generation and employs a concept-based scoring method to guide the search efficiently. Unlike simplistic pixel-based metrics like Hamming distance, ConceptSearch evaluates programs on their ability to capture the underlying transformation concept reflected in the input-output examples. We explore three scoring functions: Hamming distance, a CNN-based scoring function, and an LLM-based natural language scoring function. Experimental results demonstrate the effectiveness of ConceptSearch, achieving a significant performance improvement over direct prompting with GPT-4. Moreover, our novel concept-based scoring exhibits up to 30\% greater efficiency compared to Hamming distance, measured in terms of the number of iterations required to reach the correct solution. These findings highlight the potential of LLM-driven program search when integrated with concept-based guidance for tackling challenging generalization problems like ARC. Gautam Shroff |
AAAI | 2 |
| 2025 | ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC) (Student Abstract)abstractThe Abstraction and Reasoning Corpus (ARC) poses a significant challenge to artificial intelligence, demanding broad generalization and few-shot learning capabilities that remain elusive for current deep learning methods, including large language models (LLMs) (Chollet 2019). While LLMs excel in program synthesis, their direct application to ARC yields limited success. To address this, we introduce ConceptSearch, a novel function-search algorithm that leverages LLMs for program generation and employs a concept-based scoring method to guide the search efficiently. Experimental results demonstrate that ConceptSearch outperforms direct GPT-4 prompting, with our novel scoring function boosting efficiency by ~30% compared to the baseline Hamming distance scoring. Code at https://github.com/kksinghal/concept-search Gautam Shroff |
AAAI | 2 |
| 2024 | SCM4SR: Structural Causal Model-based Data Augmentation for Robust Session-based RecommendationabstractWith mounting privacy concerns, and movement towards a cookie-less internet, session-based recommendation (SR) models are gaining increasing popularity. The goal of SR models is to recommend top-K items to a user by utilizing information from past actions within a session. Many deep neural networks (DNN) based SR have been proposed in the literature, however, they experience performance declines in practice due to inherent biases (e.g., popularity bias) present in training data. To alleviate this, we propose an underlying neural-network (NN) based Structural Causal Model (SCM) which comprises an evolving user behavior (simulator) and recommendation model. The causal relations between the two sub-models and variables at consecutive timesteps are defined by a sequence of structural equations, whose parameters are learned using logged data. The learned SCM enables the simulation of a user's response on a counterfactual list of recommended items (slate). For this, we intervene on recommendation slates with counterfactual slates and simulate the user's response through learned SCM thereby generating counterfactual sessions to augment the training data. Through extensive empirical evaluation on simulated and real-world datasets, we show that the augmented data mitigates the impact of sparse training data and improves the performance of the SR models. Muskan Gupta, Priyanka Gupta 0003, Jyoti Narwariya, Lovekesh Vig, Gautam Shroff |
SIGIR | 5 |
| 2023 | Calibrating Deep Neural Networks using Explicit Regularisation and Dynamic Data PruningabstractDeep neural networks (DNNS) are prone to miscalibrated predictions, often exhibiting a mismatch between the predicted output and the associated confidence scores. Contemporary model calibration techniques mitigate the problem of overconfident predictions by pushing down the confidence of the winning class while increasing the confidence of the remaining classes across all test samples. However, from a deployment perspective an ideal model is desired to (i) generate well calibrated predictions for high-confidence samples with predicted probability say > 0.95 and (ii) generate a higher proportion of legitimate high-confidence samples. To this end, we propose a novel regularization technique that can be used with classification losses, leading to state-of-the-art calibrated predictions at test time; From a deployment standpoint in safety critical applications, only high-confidence samples from a well-calibrated model are of interest, as the remaining samples have to undergo manual inspection. Predictive confidence reduction of these potentially "high-confidence samples" is a downside of existing calibration approaches. We mitigate this via proposing a dynamic traintime data pruning strategy which prunes low confidence samples every few epochs, providing an increase in confident yet calibrated samples. We demonstrate state-of-the-art calibration performance across image classification benchmarks, reducing training time without much compromise in accuracy. We provide insights into why our dynamic pruning strategy that prunes low confidence training samples leads to an increase in high-confidence samples at test time. Rishabh Patra, Ramya Hebbalaguppe, Tirtharaj Dash, Gautam Shroff, Lovekesh Vig |
WACV | 4 |
| 2022 | Solving Visual Analogies Using Neural Algorithmic Reasoning (Student Abstract)abstractWe consider a class of visual analogical reasoning problems that involve discovering the sequence of transformations by which pairs of input/output images are related, so as to analogously transform future inputs. This program synthesis task can be easily solved via symbolic search. Using a variation of the ‘neural analogical reasoning’ approach, we instead search for a sequence of elementary neural network transformations that manipulate distributed representations derived from a symbolic space, to which input images are directly encoded. We evaluate the extent to which our ‘neural reasoning’ approach generalises for images with unseen shapes and positions. Atharv Sonwane, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001, Tirtharaj Dash |
AAAI | 2 |
| 2022 | A Program-Synthesis Challenge for ARC-Like Tasks
Aditya Challa, Ashwin Srinivasan 0001, Michael Bain 0001, Gautam Shroff |
ILP | 4 |
| 2022 | Intent Detection and Discovery from User Logs via Deep Semi-Supervised Contrastive ClusteringabstractRajat Kumar, Mayur Patidar, Vaibhav Varshney, Lovekesh Vig, Gautam Shroff. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mayur Patidar, Vaibhav Varshney, Lovekesh Vig, Gautam Shroff |
NAACL-HLT | 5 |
| 2021 | CauSeR: Causal Session-based Recommendations for Handling Popularity BiasabstractRecommender Systems (RS) tend to recommend more popular items instead of the relevant long-tail items. Mitigating such popularity bias is crucial to ensure that less popular but relevant items are part of the recommendation list shown to the user. In this work, we study the phenomenon of popularity bias in session-based RS (SRS) obtained via deep learning (DL) models. We observe that DL models trained on the historical user-item interactions in session logs (having long-tailed item-click distributions) tend to amplify popularity bias. To understand the source of this bias amplification, we consider potential sources of bias at two distinct stages in the modeling process: i. the data-generation stage (user-item interactions captured as session logs), ii. the DL model training stage. We highlight that the popularity of an item has a causal effect on i. user-item interactions via conformity bias, as well as ii. item ranking from DL models via biased training process due to class (target item) imbalance. While most existing approaches in literature address only one of these effects, we consider a comprehensive causal inference framework that identifies and mitigates the effects at both stages. Through extensive empirical evaluation on simulated and real-world datasets, we show that our approach improves upon several strong baselines from literature for popularity bias and long-tailed classification. Ablation studies show the advantage of our comprehensive causal analysis to identify and handle bias in data generation as well as training stages. Priyanka Gupta 0003, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff |
CIKM | 5 |
| 2021 | Complex Question Answering on knowledge graphs using machine translation and multi-task learningabstractSaurabh Srivastava, Mayur Patidar, Sudip Chowdhury, Puneet Agarwal, Indrajit Bhattacharya, Gautam Shroff. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Mayur Patidar, Sudip Chowdhury, Puneet Agarwal, Indrajit Bhattacharya, Gautam Shroff |
EACL | 6 |
| 2021 | Continual Learning for Multivariate Time Series Tasks with Variable Input DimensionsabstractWe consider a sequence of related multivariate time series learning tasks, such as predicting failures for different instances of a machine from time series of multi-sensor data, or activity recognition tasks over different individuals from multiple wearable sensors. We focus on two under-explored practical challenges arising in such settings: (i) Each task may have a different subset of sensors, i.e., providing different partial observations of the underlying ‘system’. This restriction can be due to different manufacturers in the former case, and people wearing more or less measurement devices in the latter (ii) We are not allowed to store or re-access data from a task once it has been observed at the task level. This may be due to privacy considerations in the case of people, or legal restrictions placed by machine owners. Nevertheless, we would like to (a) improve performance on subsequent tasks using experience from completed tasks as well as (b) continue to perform better on past tasks, e.g., update the model and improve predictions on even the first machine after learning from subsequently observed ones. We note that existing continual learning methods do not take into account variability in input dimensions arising due to different subsets of sensors being available across tasks, and struggle to adapt to such variable input dimensions (VID) tasks. In this work, we address this shortcoming of existing methods. To this end, we learn task-specific generative models and classifiers, and use these to augment data for target tasks. Since the input dimensions across tasks vary, we propose a novel conditioning module based on graph neural networks to aid a standard recurrent neural network. We evaluate the efficacy of the proposed approach on three publicly available datasets corresponding to two activity recognition tasks (classification) and one prognostics task (regression). We demonstrate that it is possible to significantly enhance the performance on future and previous tasks while learning continuously from VID tasks without storing data. Vibhor Gupta, Jyoti Narwariya, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff |
ICDM | 5 |
| 2021 | BCQ4DCA: Budget Constrained Deep Q-Network for Dynamic Campaign Allocation in Computational AdvertisingabstractDigital advertising companies typically conduct several advertising campaigns in parallel while being constrained by a fixed overall advertising budget. This gives rise to the problem of distributing the budget across the different campaigns dynamically so as to optimize the overall return on investment (ROI) (or some other metric) within a specified time duration. In this paper, we propose an RL formulation called BCQ4DCA for dynamic optimization of budget-constrained campaign allocation. The formulation is model-free and uses a novel cumulative reward model that is learned alongside a Deep Q-Network. We utilize a real-world Criteo user interaction dataset to evaluate BCQ4DCA in terms of conversion rate, budget utilization, cost per conversion, and ROI, finding that it outperforms current heuristic, attribution based approaches like DARNN, DNAMTA across a common time window. Manasi Malik, Garima Gupta, Lovekesh Vig, Gautam Shroff |
IJCNN | 4 |
| 2020 | MultiMBNN: Matched and Balanced Causal Inference with Neural Networks
Garima Gupta, Ranjitha Prasad, Lovekesh Vig, Gautam Shroff |
ESANN | 6 |
| 2020 | Capsule Based Neural Network Architecture to perform completeness check for Patent Eligibility ProcessabstractIn the process of filing patents, attorneys need to ask many questions, to the inventors, to ascertain patent eligibility. We propose to ease up such conversation through a deep learning-based system. This system can automatically check whether all key ingredients required for checking the patent eligibility are present in technical write-up shared by the inventors. If not, the inventors can provide the missing information. We present a trainable model to identify various ingredients such as the objective, motivation, new observation, etc. from research articles. We model this as a sentence classification problem, which is a difficult task because a patent can be filed in any domain, and sentences involved can often be very long. To this end, we propose a dilated LSTM and capsule-based neural network architecture. We present experimental results of the proposed model on a real-world patent dataset covering patent applications in diverse domains in which our organization is carrying out research and innovation activities, and also three publicly available sentence classification datasets. Through empirical analysis, we show that a) Our model performs significantly better than several strong baselines on the patent dataset; b) Performing dilation operation on LSTMs allows us to capture long term dependencies; c) our model is comparable to existing state-of-art approaches on the publicly available datasets; d) Error analysis through LIME shows that the proposed approach can help patent attorneys to interpret the decisions taken by the classifier. Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Vidya Vikas |
IJCNN | 3 |
| 2020 | Recommending in changing timesabstractRecommender systems today face major challenges in keeping up with dynamic customer preferences. Disruptions or sudden changes in the environment affect customer preferences drastically and render historical data ineffective for modeling. With businesses relying heavily on Machine Learning(ML) based recommender systems for catering to customer preferences, the accuracy of timely recommendations gains prime significance. Shruti Kunde, Amey Pandit, Rekha Singhal, Manoj Nambiar 0001, Gautam Shroff |
RecSys | 6 |
| 2020 | Constructing generative logical models for optimisation problems using domain knowledge
Ashwin Srinivasan 0001, Lovekesh Vig, Gautam Shroff |
Mach. Learn. | 3 |
| 2019 | Regularizing Fully Convolutional Networks for Time Series Classification by Decorrelating FiltersabstractDeep neural networks are prone to overfitting, especially in small training data regimes. Often, these networks are overparameterized and the resulting learned weights tend to have strong correlations. However, convolutional networks in general, and fully convolution neural networks (FCNs) in particular, have been shown to be relatively parameter efficient, and have recently been successfully applied to time series classification tasks. In this paper, we investigate the application of different regularizers on the correlation between the learned convolutional filters in FCNs using Batch Normalization (BN) as a regularizer for time series classification (TSC) tasks. Results demonstrate that despite orthogonal initialization of the filters, the average correlation across filters (especially for filters in higher layers) tends to increase as training proceeds, indicating redundancy of filters. To mitigate this redundancy, we propose a strong regularizer, using simple yet effective filter decorrelation. Our proposed method yields significant gains in classification accuracy for 44 diverse time series datasets from the UCR TSC benchmark repository. Kaushal Paneri, Vishnu TV, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff |
AAAI | 5 |
| 2019 | Retrieving Relationships from a Knowledge Graph for Question Answering
Puneet Agarwal, Maya Ramanath, Gautam Shroff |
ECIR (1) | 3 |
| 2019 | Fusing Features based on Signal Properties and TimeNet for Time Series Classification
Arijit Ukil, Pankaj Malhotra, Soma Bandyopadhyay, Tulika Bose, Ishan Sahu, Ayan Mukherjee, Lovekesh Vig, Arpan Pal 0001, Gautam Shroff |
ESANN | 9 |
| 2019 | ConvTimeNet: A Pre-trained Deep Convolutional Neural Network for Time Series ClassificationabstractTraining deep neural networks often requires careful hyper-parameter tuning and significant computational resources. In this paper, we propose ConvTimeNet (CTN): an off-the-shelf deep convolutional neural network (CNN) trained on diverse univariate time series classification (TSC) source tasks. Once trained, CTN can be easily adapted to new TSC target tasks via a small amount of fine-tuning using labeled instances from the target tasks. We note that the length of convolutional filters is a key aspect when building a pre-trained model that can generalize to time series of different lengths across datasets. To achieve this, we incorporate filters of multiple lengths in all convolutional layers of CTN to capture temporal features at multiple time scales. We consider all 65 datasets with time series of lengths up to 512 points from the UCR TSC Benchmark for training and testing transferability of CTN: We train CTN on a randomly chosen subset of 24 datasets using a multi-head approach with a different softmax layer for each training dataset, and study generalizability and transferability of the learned filters on the remaining 41 TSC datasets. We observe significant gains in classification accuracy as well as computational efficiency when using pre-trained CTN as a starting point for subsequent task-specific fine-tuning compared to existing state-of-the-art TSC approaches. We also provide qualitative insights into the working of CTN by: i) analyzing the activations and filters of first convolution layer suggesting the filters in CTN are generically useful, ii) analyzing the impact of the design decision to incorporate multiple length decisions, and iii) finding regions of time series that affect the final classification decision via occlusion sensitivity analysis. Kathan Kashiparekh, Jyoti Narwariya, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff |
IJCNN | 5 |
| 2019 | Hierarchical Capsule Based Neural Network Architecture for Sequence LabelingabstractSequence Labeling is one of the most prominent tasks in NLP. The traditional text classification models do not carry context from one sentence to another and hence may not perform well on these tasks. These models lack a hierarchical structure that can aid them in dissecting the input structure at different levels to allow flow of context between sentences. In this paper, we propose a hierarchical neural network comprising of Bi-LSTMs, Dilated Convolution operation, Capsules and Conditional Random Field (CRF) to understand the discourse/ abstract structure and predict next probable label by using label history. We have performed experiments on 3 publicly available datasets through which we have demonstrated that our model has achieved state-of-art performance on these datasets. Puneet Agarwal, Gautam Shroff, Lovekesh Vig |
IJCNN | 3 |
| 2019 | CRESA: A Deep Learning Approach to Competing Risks, Recurrent Event Survival Analysis
Garima Gupta, Vishal Sunder, Ranjitha Prasad, Gautam Shroff |
PAKDD (2) | 4 |
| 2019 | Meta-Learning for Black-Box Optimization
Vishnu TV, Pankaj Malhotra, Jyoti Narwariya, Lovekesh Vig, Gautam Shroff |
ECML/PKDD (2) | 5 |
| 2019 | Sequence and Time Aware Neighborhood for Session-based Recommendations: STANabstractRecent advances in sequence-aware approaches for session-based recommendation, such as those based on recurrent neural networks, highlight the importance of leveraging sequential information from a session while making recommendations. Further, a session based k-nearest-neighbors approach (SKNN) has proven to be a strong baseline for session-based recommendations. However, SKNN does not take into account the readily available sequential and temporal information from sessions. In this work, we propose Sequence and Time Aware Neighborhood (STAN), with vanilla SKNN as its special case. STAN takes into account the following factors for making recommendations: i) position of an item in the current session, ii) recency of a past session w.r.t. to the current session, and iii) position of a recommendable item in a neighboring session. The importance of above factors for a specific application can be adjusted via controllable decay factors. Despite being simple, intuitive and easy to implement, empirical evaluation on three real-world datasets shows that STAN significantly improves over SKNN, and is even comparable to the recently proposed state-of-the-art deep learning approaches. Our results suggest that STAN can be considered as a strong baseline for evaluating session-based recommendation algorithms in future. Diksha Garg, Priyanka Gupta 0003, Pankaj Malhotra, Lovekesh Vig, Gautam Shroff |
SIGIR | 5 |
| 2018 | Automatic Conversational Helpdesk Solution using Seq2Seq and Slot-filling ModelsabstractHelpdesk is a key component of any large IT organization, where users can log a ticket about any issue they face related to IT infrastructure, administrative services, human resource services, etc. Normally, users have to assign appropriate set of labels to a ticket so that it could be routed to right domain expert who can help resolve the issue. In practice, the number of labels are very large and organized in form of a tree. It is non-trivial to describe the issue completely and attach appropriate labels unless one knows the cause of the problem and the related labels. Sometimes domain experts discuss the issue with the users and change the ticket labels accordingly, without modifying the ticket description. This results in inconsistent and badly labeled data, making it hard for supervised algorithms to learn from. In this paper, we propose a novel approach of creating a conversational helpdesk system, which will ask relevant questions to the user, for identification of the right category and will then raise a ticket on users' behalf. We use attention based seq2seq model to assign the hierarchical categories to tickets. We use a slot filling model to help us decide what questions to ask to the user, if the top-k model predictions are not consistent. We also present a novel approach to generate training data for the slot filling model automatically based on attention in the hierarchical classification model. We demonstrate via a simulated user that the proposed approach can give us a significant gain in accuracy on ticket-data without asking too many questions to users. Finally, we also show that our seq2seq model is as versatile as other approaches on publicly available datasets, as state of the art approaches. Mayur Patidar, Puneet Agarwal, Lovekesh Vig, Gautam Shroff |
CIKM | 4 |
| 2018 | Evolutionary RL for Container Loading
Sarmimala Saikia, Richa Verma, Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001 |
ESANN | 4 |
| 2018 | KNADIA: Enterprise KNowledge Assisted DIAlogue Systems Using Deep LearningabstractIn this paper we present the design, architecture and implementation of KNADIA, a conversational dialogue system for intra-enterprise use, providing knowledge-assisted question answering and transactional assistance to employees of a large organization. KNADIA has been deployed in production in TCS, a large organization with over 380,000 employees distributed globally; the system is currently supporting a few thousand active users making hundreds of queries per day. We identify, define and distinguish two distinct classes of use-cases: virtual assistance and knowledge synthesis, which we have found to cover a variety of enterprise needs. KNADIA supports both types of conversational agents, with multiple instances of each, while presenting a common digital persona. KNADIA disambiguates which agent should respond to each query. Further, since individual components use deep learning algorithms giving probabilistic outputs, the confidence with which different components answer is often required before answering or acting. The user's dialogue context also needs to be maintained judiciously. Due to these and many other challenges, the overall architecture of KNADIA is non-trivial. We present instances of HR-assistance and technical knowledge synthesis that are in production use in TCS along with accuracy figures and key performance metrics. Finally, we suggest that many elements of our architecture are also generally applicable to other complex deep learning systems. Mahesh P. Singh, Puneet Agarwal, Ashish Chaudhary, Gautam Shroff, Prerna Khurana, Mayur Patidar, Vivek Bisht, Rachit Bansal, Prateek Sachan |
ICDE | 4 |
| 2018 | Resolving Abstract Anaphora Implicitly in Conversational Assistants using a Hierarchically stacked RNNabstractRecent proliferation of conversational systems has resulted in an increased demand for more natural dialogue systems, capable of more sophisticated interactions than merely providing factual answers. This is evident from usage pattern of a conversational system deployed within our organization. Users not only expect it to perform co-reference resolution of anaphora, but also of the antecedent or posterior facts presented by users with respect to their query. Presence of such facts in a conversation sometimes modifies the answer of main query, e.g., answer to 'how many sick leave do I get?' would be different when a fact 'I am on contract' is also present. Sometimes there is a need to collectively resolve three or four such facts. In this paper, we propose a novel solution which uses hierarchical neural network, comprising of BiLSTM layer and a maxpool layer that is hierarchically stacked to first obtain a representation of each user utterance and then to obtain a representation for sequence of utterances. This representation is used to identify users' intention. We also improvise this model by using skip connections in the second network to allow better gradient flow. Our model, not only a)~resolves the antecedent and posterior facts, but also b)~performs better even on self-contained queries. It is also c)~faster to train, making it the most promising approach for use in our environment where frequent training and tuning is needed. It slightly outperforms the benchmark on a publicly available dataset, and e)~performs better than obvious baselines approaches on our datasets. Prerna Khurana, Puneet Agarwal, Gautam Shroff, Lovekesh Vig |
KDD | 3 |
| 2018 | MobiCom'18 Panel: Hammer & Nail vis-a-vis AI / ML Applications to Networked SystemsabstractArtificial Intelligence (AI) and Machine Learning (ML) approaches, well known from IT disciplines, are beginning to excite the networking and networked systems community. Of late, we are seeing a huge excitement about applying AI and ML to networked systems. Is this merely a hype? Are there use cases and genuine applications that could lead to real deployment and practical solutions? What are the key challenges in applying AI and ML to networked systems? Can researchers and practitioners in communication networks and networked systems tap into machine learning and AI techniques to optimize network architecture, control and management, leading to increased automation in network operations? Can researchers and practitioners in the AI community explore synergy with networking researchers to optimize network architecture and design? The above are some of the questions that would be addressed during the panel discussion. The objective of the panel discussion would be to tap the minds of the global experts in order to understand the merits and limitations and the future landscape in the intersection of networking/networked systems and AI/ML. Pravin Bhagwat, Andrea J. Goldsmith, Rajeev Rastogi, Gautam Shroff |
MobiCom | 5 |
| 2017 | Hybrid BiLSTM-Siamese network for FAQ AssistanceabstractWe describe an automated assistant for answering frequently asked questions; our system has been deployed, and is currently answering HR-related queries in two different areas (leave management and health insurance) to a large number of users. The needs of a large global corporate lead us to model a frequently asked question (FAQ) to be an equivalence class of actually asked questions, for which there is a common answer (certified as being consistent with the organization's policy). When a new question is posed to our system, it finds the class of question, and responds with the answer for the class. At this point, the system is either correct (gives correct answer); or incorrect (gives wrong answer); or incomplete (says "I don't know''). We employ a hybrid deep-learning architecture in which a BiLSTM-based classifier is combined with second BiLSTM-based Siamese network in an iterative manner: Questions for which the classifier makes an error during training are used to generate a set of misclassified question-question pairs. These, along with correct pairs, are used to train the Siamese network to drive apart the (hidden) representations of the misclassified pairs. We present experimental results from our deployment showing that our iteratively trained hybrid network: (a) results in better performance than using just a classifier network, or just a Siamese network; (b) performs better than state-of-the art sentence classifiers in the two areas in which it has been deployed, in terms of both accuracy as well as precision-recall tradeoff; and (c) also performs well on a benchmark public dataset. We also observe that using question-question pairs in our hybrid network, results in marginally better performance than using question-to-answer pairs. Finally, estimates of precision and recall from the deployment of our automated assistant suggest that we can expect the burden on our HR department to drop from answering about 6000 queries a day to about 1000. Prerna Khurana, Puneet Agarwal, Gautam Shroff, Lovekesh Vig, Ashwin Srinivasan 0001 |
CIKM | 3 |
| 2017 | Learning and Knowledge Transfer with Memory Networks for Machine ComprehensionabstractEnabling machines to read and comprehend unstructured text remains an unfulfilled goal for NLP research.Recent research efforts on the "machine comprehension" task have managed to achieve close to ideal performance on simulated data.However, achieving similar levels of performance on small real world datasets has proved difficult; major challenges stem from the large vocabulary size, complex grammar, and the frequent ambiguities in linguistic structure.On the other hand, the requirement of human generated annotations for training, in order to ensure a sufficiently diverse set of questions is prohibitively expensive.Motivated by these practical issues, we propose a novel curriculum inspired training procedure for Memory Networks to improve the performance for machine comprehension with relatively small volumes of training data.Additionally, we explore various training regimes for Memory Networks to allow knowledge transfer from a closely related domain having larger volumes of labelled data.We also suggest the use of a loss function to incorporate the asymmetric nature of knowledge transfer.Our experiments demonstrate improvements on Dailymail, CNN, and MCTest datasets. Lovekesh Vig, Gautam Shroff |
EACL (1) | 3 |
| 2017 | TimeNet: Pre-trained deep recurrent neural network for time series classification
Pankaj Malhotra, Vishnu TV, Lovekesh Vig, Puneet Agarwal, Gautam Shroff |
ESANN | 5 |
| 2017 | Information Bottleneck Inspired Method For Chat Text SegmentationabstractWe present a novel technique for segmenting chat conversations using the information bottleneck method (Tishby et al., 2000), augmented with sequential continuity constraints. Furthermore, we utilize critical non-textual clues such as time between two consecutive posts and people mentions within the posts. To ascertain the effectiveness of the proposed method, we have collected data from public Slack conversations and Fresco, a proprietary platform deployed inside our organization. Experiments demonstrate that the proposed method yields an absolute (relative) improvement of as high as 3.23% (11.25%). To facilitate future research, we are releasing manual annotations for segmentation on public Slack conversations. Sunder Vishal, Lovekesh Vig, Gautam Shroff |
IJCNLP(1) | 4 |
| 2016 | Visual Bayesian fusion to navigate a data lake
Karamjit Singh, Kaushal Paneri, Aditeya Pandey, Garima Gupta, Geetika Sharma, Puneet Agarwal, Gautam Shroff |
FUSION | 7 |
| 2016 | Generation of Near-Optimal Solutions Using ILP-Guided Sampling
Ashwin Srinivasan 0001, Gautam Shroff, Lovekesh Vig, Sarmimala Saikia |
ILP | 2 |
| 2016 | Generic Framework to Predict Repeat Behavior of Customers Using Their Transaction HistoryabstractThere exists a class of problems in e-commerce and retail businesses where the shopping behavior of customers is analyzed in order to predict their repeat behavior for products or retail stores. This analysis plays a crucial role in advertisement budgeting, product placement and relevant customer targeting. Researchers have addressed this problem by using standard predictive models, which use ad hoc features. We propose a metamodel that abstracts the different dimensions of data present in transactional datasets. These dimensions can be customer, product, offer, target, marketplace and transactions. Our framework also has abstract functions for comprehensive feature set generation, and includes different machine learning algorithms to learn prediction model. Our framework works end-to-end from feature engineering to reporting repeat probabilities of customers for products (or marketplace, brand, website or storechain). Moreover, the predicted repeat behavior of customers for different products along with their transactional history is used by our offer optimization model i-Prescribe to suggest products to be offered to customers with the goal of maximizing the return on investment of given marketing budget. We prove that our abstract features work on two different data-challenge datasets, by sharing experimental results. Auon Haidar Kazmi, Gautam Shroff, Puneet Agarwal |
WI | 2 |
| 2015 | Succinctly summarizing machine usage via multi-subspace clustering of multi-sensor dataabstractModern industrial equipments of all kinds are instrumented with a large number of sensors that continuously transmit their readings wirelessly, giving rise to what is often referred to as the `industrial internet'. Such data are often explored by engineers to determine the different usage patterns and behavior of similar machines. In this paper we describe a technique to automatically summarize the usage and behavioral patterns of a collection of similar machines by a small set of rules that nevertheless cover a large fraction of the observed data. We characterize the usage and behavior of a machine over a day, by a collection of single-sensor histograms; thus each day is a point in a high-dimensional space. We first cluster days according to each sensor separately and then combine the clusters using communities in a specially constructed graph that considers common days within clusters of different sensors. In the process some clusters of a single sensor get merged. Finally, we discover rules, each comprising of memberships in clusters of possibly different sensors. Thus, we use the term multi-subspace clustering to describe such a collection of cluster-based rules. Last but not the least, we attempt to cover a large fraction of observed days with a small number of such rules. We present empirical results on voluminous (100s of GBs) real-life sensor data and also compare our technique with related work in subspace clustering and histogram summarization. Sarmimala Saikia, Gautam Shroff, Puneet Agarwal, Ashwin Srinivasan 0001 |
DSAA | 2 |
| 2015 | Predictive reliability mining for early warnings in populations of connected machinesabstractTraditional reliability analysis of complex machinery involves statistical modeling of historical data on part failures from warranty claims, using distributions from exponential family such as the Weibull or log-normal distribution. When observed failures (in one or more parts) across a population of machines exceed the number expected based on such a model, this may serve as an early warning of a potential systemic problem with the population. Of course, such early warnings rely on some exceptionally high failures having actually occurred. However, modern connected vehicles, engines and machines of all kinds are equipped with on-board electronics that transmit alerts, referred to as `diagnostic trouble codes' or DTCs over the network, whenever abnormal conditions are detected. Such DTC signals should also be able to serve as early-warning indicators, typically before actual failures are observed in large numbers. In this paper, we develop a graphical Bayesian model that augments standard reliability analysis with early-warning indicators such as DTC signals observed over the industrial Internet. We demonstrate that our augmented model can detect of potential problems earlier than that using traditional reliability analysis. Going further, we note that significant deviations from expected failure counts might often occur only in some unknown subset of the population, e.g., a particular batch, or machines manufactured at a particular plant. In such cases, deviations from expected numbers are insignificant across the full population. We present a rule mining technique that discovers such subsets efficiently even when the number of dimensions across which a subset may be defined is large. We term our approach as reliability mining since it combines the use of a Bayesian reliability model with subgroup discovery using data mining techniques. We present experimental results using synthetically simulated scenarios as well as real-life data from a major global automobile manufacturer. Karamjit Singh, Gautam Shroff, Puneet Agarwal |
DSAA | 2 |
| 2015 | Long Short Term Memory Networks for Anomaly Detection in Time Series
Pankaj Malhotra, Lovekesh Vig, Gautam Shroff, Puneet Agarwal |
ESANN | 3 |
| 2015 | Business data fusion
Surya Yadav, Gautam Shroff, Ehtesham Hassan, Puneet Agarwal |
FUSION | 2 |
| 2014 | Incremental entity fusion from linked documents
Pankaj Malhotra, Puneet Agarwal, Gautam Shroff |
FUSION | 3 |
| 2014 | Prescriptive information fusion
Gautam Shroff, Puneet Agarwal, Karamjit Singh, Auon Haidar Kazmi, Sapan Shah, Avadhut Sardeshmukh |
FUSION | 1 |
| 2013 | Email Analytics for Activity Management and Insight DiscoveryabstractEmails constitute the bulk of all official communications in any organization. Email repositories are tacit store-houses of knowledge about people, projects and processes. Mining one's own email repository can also provide interesting and valuable insights about his or her engagements and contacts along different dimensions. In this paper, we propose an email analytics framework that combines text-mining, network analysis and data analytics principles to mine email repositories for useful insights. While individuals are more attuned to looking at emails as individual items along with a history that is embedded in the trail, mining the whole collection can also lead to knowledge-discovery about similarities and dissimilarities of different engagements. This in turn can lead to valuable information like comparative status reports on various projects or deeper insights about why certain projects succeed while others don't. Given the volumes, diversity and noisy nature of e-mails, it becomes impossible for human beings to comprehend the impact of all of it unless the task is automated and approached in a structured fashion. We show that combination of text and network analytics along with temporal reasoning can provide valuable insights about task-states, actionable items, recommendations and forecasts. These insights can be exploited very effectively for project-management tasks like automated identification of bottlenecks or their causes, elimination of inefficiencies, early-warnings and suggestions about proactive measures to avoid problems. It is possible to extend the framework quite easily to analyze multiple email repositories of different users, though this work does not address the privacy or security concerns that might exist. Lipika Dey, Sameera Bharadwaja H., G. Meera, Gautam Shroff |
Web Intelligence | 4 |
| 2012 | Catching the Long-Tail: Extracting Local News Events from Twitter
Puneet Agarwal, Rajgopal Vaithiyanathan, Gautam Shroff |
ICWSM | 4 |
| 2011 | Enterprise information fusion for real-time business intelligence
Gautam Shroff, Puneet Agarwal, Lipika Dey |
FUSION | 1 |
| 2011 | A blackboard architecture for data-intensive information fusion using locality-sensitive hashing
Gautam Shroff, Puneet Agarwal, Shefali Bhat |
FUSION | 1 |
| 2009 | Experiments in distributed side-by-side software developmentabstractIn distributed side-by-side software development, a pair of distributed team members are assigned a single task and allowed to (a) work concurrently on two different computers and (b) see each others' displays. They can control when they communicate with each other, view each others' actions, and in Prasun Dewan, Puneet Agarwal, Gautam Shroff, Rajesh Hegde |
CollaborateCom | 3 |
| 2008 | Dev 2.0: model driven development in the cloudabstractIn recent years the term Web 2.0 has been used to describe the transformation of the internet from a world of publishers and readers to one of collaborators where everyone is a creator of content, and 'communities' bind participants in an ecosystem. We have also seen the success of some software as a service (SaaS) applications, such as Salesforce.com, to the extent that application development is itself available as a hosted service (Salesforce's AppExchange, Coghead, etc.). We call this 'Dev 2.0', where the line between users and developers is blurred, and an application is available to all stakeholders through the life-cycle as it evolves. Gautam Shroff |
SIGSOFT FSE | 1 |
| 2006 | An Empirical Study of Factors and their Relationships in Outsourced Software MaintenanceabstractIT outsourcing, which commenced as a cost reduction operation, has now become an essential parameter for measuring the operational efficiency of an organization. An increasing number of operational IT systems are moving into the maintenance phase, becoming potential candidates for outsourcing. Human and organizational factors, typical to maintenance activities, such as organization climate, customer attitude, engineers' attitude, etc., have a significant influence on software maintenance effort and make the task of effort estimation complex. In this paper we present the results of an empirical study, carried out to identify and study the influence of such factors on software maintenance effort. Pankaj Bhatt, Gautam Shroff, K. Williams, Arun Kumar Misra |
APSEC | 2 |
| 2006 | An influence model for factors in outsourced software maintenanceabstractThe rapid growth of the Internet in the recent past has encouraged global deployment of work by an increasing number of organizations around the world, and they are now in a better position to outsource their IT functions to specialist vendors. With the passage of time, we find more and more software systems moving into the maintenance phase. Such software systems have become an increasingly significant expenditure for businesses. Consequently, these are often potential candidates for outsourcing. Inadequate information regarding the size, complexity, reliability, maintainability, etc., of these systems often makes the task of estimating the maintenance effort a challenge. Other human and organizational factors, typical to maintenance activities, such as organization climate, customer attitude, engineers' attitude, the need for multi-location support teams, etc., make the situation even more complex. In this paper we present the results of an empirical study carried out to identify such factors and study their influence on the maintenance effort. We classify these factors in four categories, namely system baseline, maintenance team, customer's attitude and organizational climate. We also propose a model which can help a practitioner to predict and control the impact on maintenance effort, based on the strengths of these factors. Copyright © 2006 John Wiley & Sons, Ltd. Pankaj Bhatt, Gautam Shroff, C. Anantaram, Arun Kumar Misra |
J. Softw. Maintenance Res. Pract. | 2 |
| 2005 | Collaborative development of business applicationsabstractCollaborative software development models, inspired by the open source community, are also being considered for development and deployment of business applications. We describe current work and future directions towards a hosted model to support collaboration during software development. We submit that such a platform can enable rapid application development, facilitate globally located virtual teams spanning locations to work as a single unit, and also encourage re-use, especially of technical frameworks. Going forward, we foresee extending such environments to providing an end-to-end development process accessible to globally distributed virtual teams. Gautam Shroff, Anish Mehta, Puneet Agarwal, Rajesh Sinha |
CollaborateCom | 1 |
| 1996 | Transparent parallel replication of logically partitioned databasesabstractThis paper presents a protocol for efficient transaction management in an environment of replicated autonomous databases. Each replicated copy has ownership over mutually exclusive portions of the database. The protocol improves response time and throughput by exploiting parallelism although reducing the degree of transaction isolation. Most modifications to the database are assumed to be on the locally owned portion of the database, with only occasional nonlocal writes/updates. Read operations, however can access either local or nonlocal objects equally. We are able to prove that users of our parallel replicated database system can view it equivalent to that of a single database providing "degree 2" transaction isolation, i.e. the replication and parallelism is transparent to the application programmer. The protocol communication overheads are limited allowing it to be efficiently implemented over even wide area networks. Experimental results using a prototype demonstrating the performance improvements are presented. Rekha Goel, Gautam Shroff |
HiPC | 2 |