VLDB 2026 Research / reviewers in the wild / expert
Sameep Mehta
dblp:91/877
· DBLP profile ↗
75ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0002-9599-1526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 33 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSystems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as ToolsabstractRecent advancements in Large Language Models (LLMs) has lead to the development of agents capable of complex reasoning and interaction with external tools. In enterprise contexts, the effective use of such tools that are often enabled by application programming interfaces (APIs) is hindered by poor documentation, complex input or output schema, and large number of operations. These challenges make tool selection difficult and reduce the accuracy of payload formation upto 25%. We propose ACE, an automated tool creation and enrichment framework that transforms enterprise APIs into LLM-compatible tools. ACE (i) generates enriched tool specifications with parameter descriptions and examples to improve selection and invocation accuracy, and (ii) incorporates a dynamic shortlisting mechanism that filters relevant tools at runtime, reducing prompt complexity while maintaining scalability. We validate our framework on both proprietary and open-source APIs and demonstrate its integration with agentic frameworks. To the best of our knowledge, ACE is the first end-to-end framework that automates the creation, enrichment, and dynamic selection of enterprise API tools for LLM agents. Prerna Agarwal, Soujanya Soni, Rohith D. Vallam, Renuka Sindhgatta, Sameep Mehta |
AAAI | 6 |
| 2026 | DFAgent: From Natural Language Data Interactions to Reusable Agent-Ready ToolsabstractWe present DataFoundry Agent (DFAgent), a system that forges reusable, agent-ready tools from interactive data exploration, quality, and remediation tasks. Users engage with data through natural-language prompts for operations that include inspection, transformation, and visualization. These interactions automatically generate executable code snippets that are logged. From these snippets, DFAgent acts as a foundry, synthesizing a governed catalog of enriched tools exposed via the Model Context Protocol (MCP). In this way, user-derived logic for all data operations is transformed into standardized, composable tools without reimplementation. We demonstrate how diverse interactions accumulate into a reusable toolset, highlighting a paradigm that unifies natural language interaction, executable code generation, and tool foundry processes for agentic data systems. Neelamadhav Gantayat, Renuka Sindhgatta, Sambit Ghosh, Sameep Mehta, Soujanya Soni |
AAAI | 4 |
| 2026 | From Natural Language to Executable ETL Flows: The IBM DataStage AssistantabstractModern ETL (Extract, Transform, Load) tools offer graphical, no-code interfaces for workflow creation but still require users to manually identify transformation functions and configure their properties, which is time-consuming and demands prior expertise. We present the research and engineering foundations of the IBM DataStage Assistant, a deployed capability that generates complete multi-stage ETL flows directly from natural language (NL) descriptions. Our framework infers transformation functions, their properties, and transformer expressions, enabling novices to discover relevant functions and allowing experts to bypass manual configuration. The proposed framework achieves a prediction accuracy of 96.4% for flow predictions, 87.0% for properties, and 83.6% for transformer expressions. We also show a document exploration module that uses retrieval-augmented generation (RAG) over product documentation to answer tool-specific questions in NL. Implemented in IBM DataStage, this approach supports iterative, in-environment workflow design and reduces context switching. In initial studies, it achieves up to 90% time savings for novices and 50% for experts. Nitin Gupta 0005, Thomas Gschwind, Shramona Chakraborty, Sameep Mehta, Tristan Tyler, Shreya Sisodia, Ben Clermont |
AAAI | 4 |
| 2026 | ToolSmith: A Multi-Agent Framework for Enterprise Tool CreationabstractAlthough LLMs can generate tools for generic domains and tasks, they struggle with enterprise-related domains that involve proprietary APIs and data schemas. We present ToolSmith, a framework for autonomously generating and validating agent-compatible tools. Given an API specification and a Tool Specification Requirement (TSR), ToolSmith produces a tool function and verifies it through a closed-loop process: it creates natural language (NL) tests and executes the tool in a secure agent sandbox for validation. For state-changing tools, ToolSmith confirms outcomes by querying the API with parameters derived from the NL tests. If the tool fails to produce the desired output, ToolSmith generates diagnostic feedback to iteratively regenerate it. By ensuring both functional correctness and agent compatibility, ToolSmith enables reliable automation of enterprise workflows. Purna Chandra Sekhar Vakudavathu, Kushal Mukherjee, Jayachandu Bandlamudi, Renuka Sindhgatta, Sameep Mehta |
AAAI | 5 |
| 2025 | Question-guided Insights Generation for Automated Exploratory Data AnalysisabstractExploratory Data Analysis (EDA) derives meaningful insights from extensive and complex datasets. This process typically involves a series of analytical operations to identify the patterns within the data. However, the effectiveness of EDA is often limited by the user's domain knowledge and proficiency in data exploration methods. To overcome these challenges, we developed QUIS, a fully automated EDA system that uncovers insights by generating data-related questions and exploring subspaces in the dataset without prior training. QUIS allows users to control key system parameters such as beam width, beam depth, and expansion factor for subspace selection, the interestingness score for filtering valuable insights, and parameters for managing the quality and quantity of generated questions. Abhijit Manatkar, Ashlesha Akella, Krishnasuri Narayanam, Sameep Mehta |
AAAI | 4 |
| 2025 | Robust Evaluation of LLM-Generated GraphQL Queries for Web Services
Vansika Sonthalia, Manish Kesarwani, Sameep Mehta |
ICWS | 3 |
| 2024 | LLMGuard: Guarding against Unsafe LLM BehaviorabstractAlthough the rise of Large Language Models (LLMs) in enterprise settings brings new opportunities and capabilities, it also brings challenges, such as the risk of generating inappropriate, biased, or misleading content that violates regulations and can have legal concerns. To alleviate this, we present "LLMGuard", a tool that monitors user interactions with an LLM application and flags content against specific behaviours or conversation topics. To do this robustly, LLMGuard employs an ensemble of detectors. Shubh Goyal, Medha Hira, Shubham Mishra, Sukriti Goyal, Arnav Goel, Niharika Dadu, Kirushikesh D. B., Sameep Mehta, Nishtha Madaan |
AAAI | 8 |
| 2024 | Sequential API Function Calling Using GraphQL SchemaabstractAvirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Avirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta |
EMNLP | 8 |
| 2024 | LLM-powered GraphQL Generator for Data Retrieval
Balaji Ganesan, Sambit Ghosh, Nitin Gupta 0005, Manish Kesarwani, Sameep Mehta, Renuka Sindhgatta |
IJCAI | 5 |
| 2023 | Foundations and Applications in Large-scale AI Models: Pre-training, Fine-tuning, and Prompt-based LearningabstractDeep learning techniques have advanced rapidly in recent years, leading to significant progress in pre-trained and fine-tuned large-scale AI models. For example, in the natural language processing domain, the traditional "pre-train, fine-tune" paradigm is shifting towards the "pre-train, prompt, and predict" paradigm, which has achieved great success on many tasks across different application domains such as ChatGPT/BARD for Conversational AI and P5 for a unified recommendation system. Moreover, there has been a growing interest in models that combine vision and language modalities (vision-language models) which are applied to tasks like Visual Captioning/Generation. Considering the recent technological revolution, it is essential to emphasize these paradigm shifts and highlight the paradigms with the potential to solve different tasks. We thus provide a platform for academic and industrial researchers to showcase their latest work, share research ideas, discuss various challenges, and identify areas where further research is needed in pre-training, fine-tuning, and prompt-learning methods for large-scale AI models. We foster the development of a strong research community focused on solving challenges related to large-scale AI models, providing superior and impactful strategies that can change people's lives in the future. Zhiyuan Cheng 0002, Dhaval Patel 0002, Linsey Pang, Sameep Mehta, Kexin Xie, Ed H. Chi, Wei Liu 0007, Nitesh V. Chawla, James Bailey 0001 |
KDD | 4 |
| 2022 | Toward Scientific Workflows in a Serverless WorldabstractServerless computing and FaaS have gained popularity due to their ease of design, deployment, scaling and billing on clouds. However, when used to compose and orchestrate scientific workflows, they pose limitations due to cold starts, message indirection, vendor lock-in and lack of provenance support. Here, we propose a design for a Ser verless Scientific Workflow Orchestrator that overcomes these challenges using techniques like function fusion, pilot invocations and data fabrics. Aakash Khochare, Yogesh L. Simmhan, Sameep Mehta, Arvind Agarwal |
e-Science | 3 |
| 2021 | Data Quality for Machine Learning TasksabstractThe quality of training data has a huge impact on the efficiency, accuracy and complexity of machine learning tasks. Data remains susceptible to errors or irregularities that may be introduced during collection, aggregation or annotation stage. This necessitates profiling and assessment of data to understand its suitability for machine learning tasks and failure to do so can result in inaccurate analytics and unreliable decisions. While researchers and practitioners have focused on improving the quality of models, there are limited efforts towards improving the data quality. Nitin Gupta 0005, Shashank Mujumdar, Hima Patel, Satoshi Masuda, Naveen Panwar, Sambaran Bandyopadhyay, Sameep Mehta, Shanmukha C. Guttula, Shazia Afzal, Ruhi Sharma Mittal, Vitobha Munigala |
KDD | 7 |
| 2021 | 2nd International Workshop on Data Quality Assessment for Machine LearningabstractThe 2nd International Workshop on Data Quality Assessment for Machine Learning (DQAML'21) is organized in conjunction with the Special Interest Group on Knowledge Discovery and Data Mining (SIGKDD). This workshop aims to serve as a forum for the presentation of research related to data quality assessment and remediation in AI/ML pipeline. Data quality is a critical issue in the data preparation phase and involves numerous challenging problems related to detection, remediation, visualization and evaluation of data issues. The workshop aims to provide a platform to researchers and practitioners to discuss such challenges across different modalities of data like structured, time series, text and graphical. The aim is to attract perspectives from both industrial and academic circles. Hima Patel, Fuyuki Ishikawa, Laure Berti-Équille, Nitin Gupta 0005, Sameep Mehta, Satoshi Masuda, Shashank Mujumdar, Shazia Afzal, Srikanta J. Bedathur, Yasuharu Nishi |
KDD | 5 |
| 2020 | Multidimensional Analysis of Trust in News Articles (Student Abstract)abstractThe advancements in the field of Information Communication Technology have engendered revolutionary changes in the journalism industry, not only on the part of the journalists and the media personnel, but also on the people consuming these news stories, who today, are only a click away from all the updates they need. However, these advances have also exposed the prevailing venality, wearying off the trust of the public in news media. How then, does an individual discern that which, out of the countless news stories for an incident, should be trusted? This work introduces a system that presents the user a multidimensional analysis for trust in news from various media sources based on the textual content of the articles, assessment of the journalists' perspectives and the temporal diversity of the issues being covered by the media houses publishing the news articles. Our experiments on a self-collected dataset confirm that the system aids in a comprehensive analysis of trust. Maitree Leekha, Utkarsh Chawla, Ayush Agarwal, Mudit Saxena, Nishtha Madaan, Kalapriya Kannan, Sameep Mehta |
AAAI | 8 |
| 2020 | On Efficiently Processing Business Lineage QueriesabstractIn this paper, we look at the problem of retrieving business need specific l ineage i nformation f rom t he provenance graphs. A provenance graph models the events happening on various assets on a data platform. The output of a lineage query may contain large number of nodes. However, a user depending on her business role may only want a small subset of these nodes and the lineage relationships among these nodes. We formally define the notion of a class of business lineage queries wherein a business user specifies the lineage events relevant to her business need, in terms of event labels. The lineage output then consists of these events of interest, assets associated with these events and the lineage relationships among these events and assets of interest. We propose a novel framework for efficiently executing such business lineage queries and experimentally illustrate the effectiveness of the same. Rajmohan C, Sameep Mehta, Kiran Pulapa |
IEEE BigData | 3 |
| 2020 | Budgeted Batch Mode Active Learning with Generalized Cost and Utility FunctionsabstractActive learning reduces the labeling cost by actively querying labels for the most valuable data points. Typical active learning methods select the most informative examples one-at-a-time, their batch variants exist which select a set of most informative points instead of one point at a time. These points are selected in such a way that when added to the training data along with their labels, they provide maximum benefit to the underlying model. In this paper, we present a learning framework that actively selects optimal set of examples (in a batch) within a given budget, based on given utility and cost functions. The framework is generic enough to incorporate any utility and any cost function defined on a set of examples. Furthermore, we propose a novel utility function based on the Facility Location problem that considers three important characteristics of utility i.e., diversity, density and point utility. We also propose a novel cost function, by formulating the cost computation problem as an optimization problem, the solution to which turns out to be the minimum spanning tree. Thus, our framework provides the optimal batch of points within the given budget based on the cost and utility functions. We evaluate our method on several data sets and show its superior performance over baseline methods. Arvind Agarwal, Shashank Mujumdar, Nitin Gupta 0005, Sameep Mehta |
ICPR | 4 |
| 2020 | Overview and Importance of Data Quality for Machine Learning TasksabstractIt is well understood from literature that the performance of a machine learning (ML) model is upper bounded by the quality of the data. While researchers and practitioners have focused on improving the quality of models (such as neural architecture search and automated feature selection), there are limited efforts towards improving the data quality. One of the crucial requirements before consuming datasets for any application is to understand the dataset at hand and failure to do so can result in inaccurate analytics and unreliable decisions. Assessing the quality of the data across intelligently designed metrics and developing corresponding transformation operations to address the quality gaps helps to reduce the effort of a data scientist for iterative debugging of the ML pipeline to improve model performance. This tutorial highlights the importance of analysing data quality in terms of its value for machine learning applications. This tutorial surveys all the important data quality related approaches discussed in literature, focusing on the intuition behind them, highlighting their strengths and similarities, and illustrates their applicability to real-world problems. Finally we will discuss the interesting work IBM Research is doing in this space. Abhinav Jain 0001, Hima Patel, Lokesh Nagalapatti, Nitin Gupta 0005, Sameep Mehta, Shanmukha C. Guttula, Shashank Mujumdar, Shazia Afzal, Ruhi Sharma Mittal, Vitobha Munigala |
KDD | 5 |
| 2019 | Learning Convolutional Neural Networks with Deep Part EmbeddingsabstractWe propose a novel concept of Deep Part Embeddings (DPEs), which can be used to learn new Convolutional Neural Networks (CNNs) for different classes. We define DPE as a neuron of a trained CNN along with its network of filter activations that is interpretable as a part of a class that the neuron contributes to. Given a new class C, we explore the idea of combining different DPEs that intuitively constitute C, from trained CNNs (not on C), into a network that learns the class C with few training samples. An important application of our proposed framework is the ability to modify a CNN trained on n classes to learn a new class with limited training data without significantly affecting its performance on the n classes. We visually illustrate the different network architectures and extensively evaluate their performance against the baselines. Nitin Gupta 0005, Shashank Mujumdar, Prerna Agarwal, Abhinav Jain 0001, Sameep Mehta |
ICASSP | 5 |
| 2019 | Radial Loss for Learning Fine-grained Video Similarity MetricabstractIn this paper, we propose the Radial Loss which utilizes category and sub-category labels to learn an order-preserving fine-grained video similarity metric. We propose an end-to-end quadlet-based Convolutional Neural Network (CNN) combined with Long Short-term Memory (LSTM) Unit to model video similarities by learning the pairwise distance relationships between samples in a quadlet generated using the category and sub-category labels. We showcase two novel applications of learning a video similarity metric - (i) fine-grained video retrieval, (ii) fine-grained event detection, along with simultaneous shot boundary detection, and correspondingly show promising results against those of the baselines on two new fine-grained video datasets. Abhinav Jain 0001, Prerna Agarwal, Shashank Mujumdar, Nitin Gupta 0005, Sameep Mehta, Chiranjoy Chattopadhyay |
ICASSP | 5 |
| 2019 | On Efficiently Processing Workflow Provenance Queries in SparkabstractIn this paper, we look at how we can leverage Spark platform for efficiently processing fine-grained provenance queries on large volumes of workflow provenance data. Simple recursive querying based Spark solutions involve large data scanning cost and hence do not work well. We propose a novel provenance framework which is engineered to quickly determine a small volume of data containing the entire lineage of the queried data-item. This small volume of data is then recursively processed to figure out the provenance of the queried data-item. We study the effectiveness of the proposed framework on a provenance trace obtained from a financial domain text curation workflow and report our observations. We show that the proposed framework easily outperforms the naive approaches. Rajmohan C, Pranay Lohia, Siddhartha Brahma, Mauricio A. Hernández, Sameep Mehta |
ICDCS | 6 |
| 2019 | Hardening Deep Neural Networks via Adversarial Model CascadesabstractDeep neural networks (DNNs) are vulnerable to malicious inputs crafted by an adversary to produce erroneous outputs. Works on securing neural networks against adversarial examples achieve high empirical robustness on simple datasets such as MNIST. However, these techniques are inadequate when empirically tested on complex data sets such as CIFAR-10 and SVHN. Further, existing techniques are designed to target specific attacks and fail to generalize across attacks. We propose Adversarial Model Cascades (AMC) as a way to tackle the above inadequacies. Our approach trains a cascade of models sequentially where each model is optimized to be robust towards a mixture of multiple attacks. Ultimately, it yields a single model which is secure against a wide range of attacks; namely FGSM, Elastic, Virtual Adversarial Perturbations and Madry. On an average, AMC increases the model's empirical robustness against various attacks simultaneously, by a significant margin (of 6.225% for MNIST, 5.075% for SVHN and 2.65% for CIFAR-10 ). At the same time, the model's performance on non-adversarial inputs is comparable to the state-of-the-art models. Deepak Vijaykeerthy, Anshuman Suri, Sameep Mehta, Ponnurangam Kumaraguru |
IJCNN | 3 |
| 2019 | Design diagrams as ontological source
Pranay Lohia, Kalapriya Kannan, Biplav Srivastava, Sameep Mehta |
ESEC/SIGSOFT FSE | 4 |
| 2019 | A novel feature transform framework using deep neural network for multimodal floor plan retrieval
Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
Int. J. Document Anal. Recognit. | 4 |
| 2018 | On Building Efficient Temporal Indexes on Hyperledger FabricabstractWe discuss the problem of constructing efficient temporal indexes on Hyperledger Fabric, a popular Blockchain platform. The temporal nature of the data inserted by Fabric transactions can be leveraged to support various use-cases. This requires that temporal queries be processed efficiently on this data. Currently this presents significant challenges as this data is organized on file-system, is exposed via limited API and does not support temporal indexes. In a prior work [1], we presented two models for creating temporal indexes on Fabric which overcome these limitations and improve the performance of temporal queries on Fabric. The first model creates a copy of each event inserted and stores temporally close events together on Fabric. The second model keeps the event count intact but tags metadata to each event s.t. temporally close events share the same metadata. In this paper, we present variants on these two models which are better able to handle the skew present in Fabric data. We discuss the details and show that these variants significantly outperform the approaches presented in [1] when Fabric data contains skew. We also discuss the performance tradeoffs among these variants across various dimensions - data storage, query performance, event insertion time etc. Sandeep Hans, Sameep Mehta, Praveen Jayachandran |
IEEE CLOUD | 3 |
| 2018 | Secure k-NN as a Service over Encrypted Data in Multi-User SettingabstractTo securely leverage the advantages of Cloud Computing, a lot of research has happened in the area of "Secure Query Processing over Encrypted Data". As a concrete use case, many encryption schemes have been proposed for securely processing k Nearest Neighbors (SkNN) over encrypted data in the outsourced setting. Recently Zhu et al.[1] proposed a SkNN solution which claimed to satisfy following four properties: (1)Data Privacy, (2)Key Confidentiality, (3)Query Privacy, and (4)Query Controllability. However, in this paper, we present an attack which breaks the Query Controllability claim of their scheme. Further, we propose a new SkNN solution which satisfies all the four existing properties along with an additional essential property of Query Check Verification. We analyze the security of our proposed scheme and present the detailed experimental results to showcase the efficiency in real world scenario. Akshar Kaul, Sameep Mehta |
IEEE CLOUD | 3 |
| 2018 | Semantic Understanding for Contextual In-Video AdvertisingabstractWith the increasing consumer base of online video content, it is important for advertisers to understand the video context when targeting video ads to consumers. To improve the consumer experience and quality of ads, key factors need to be considered such as (i) ad relevance to video content (ii) where and how video ads are placed, and (iii) non-intrusive user experience. We propose a framework to semantically understand the video content for better ad recommendation that ensure these criteria. Rishi Madhok, Shashank Mujumdar, Nitin Gupta 0005, Sameep Mehta |
AAAI | 4 |
| 2018 | Content and Context: Two-Pronged Bootstrapped Learning for Regex-Formatted Entity ExtractionabstractRegular expressions are an important building block of rule-based information extraction systems. Regexes can encode rules to recognize instances of simple entities which can then feed into the identification of more complex cross-entity relationships. Manually crafting a regex that recognizes all possible instances of an entity is difficult since an entity can manifest in a variety of different forms. Thus, the problem of automatically generalizing manually crafted seed regexes to improve the recall of IE systems has attracted research attention. In this paper, we propose a bootstrapped approach to improve the recall for extraction of regex-formatted entities, with the only source of supervision being the seed regex. Our approach starts from a manually authored high precision seed regex for the entity of interest, and uses the matches of the seed regex and the context around these matches to identify more instances of the entity. These are then used to identify a set of diverse, high recall regexes that are representative of this entity. Through an empirical evaluation over multiple real world document corpora, we illustrate the effectiveness of our approach. Stanley Simoes, Deepak P 0001, Munu Sairamesh, Deepak Khemani, Sameep Mehta |
AAAI | 5 |
| 2018 | Model Extraction Warning in MLaaS ParadigmabstractMachine learning models deployed on the cloud are susceptible to several security threats including extraction attacks. Adversaries may abuse a model's prediction API to steal the model thus compromising model confidentiality, privacy of training data, and revenue from future query payments. This work introduces a model extraction monitor that quantifies the extraction status of models by continually observing the API query and response streams of users. We present two novel strategies that measure either the information gain or the coverage of the feature space spanned by user queries to estimate the learning rate of individual and colluding adversaries. Both approaches have low computational overhead and can easily be offered as services to model owners to warn them against state of the art extraction attacks. We demonstrate empirical performance results of these approaches for decision tree and neural network models using open source datasets and BigML MLaaS platform. Manish Kesarwani, Bhaskar Mukhoty, Vijay Arya, Sameep Mehta |
ACSAC | 4 |
| 2018 | Generating Adversarial Text Samples
Suranjana Samanta, Sameep Mehta |
ECIR | 2 |
| 2018 | Efficient Secure k-Nearest Neighbours over Encrypted Data
Manish Kesarwani, Akshar Kaul, Prasad Naldurg, Sikhar Patranabis, Sameep Mehta, Debdeep Mukhopadhyay |
EDBT | 6 |
| 2018 | Efficiently Processing Temporal Queries on Hyperledger FabricabstractIn this paper, we discuss the problem of efficiently handling temporal queries on Hyperledger Fabric, a popular implementation of Blockchain technology. The temporal nature of the data inserted by the Hyperledger Fabric transactions can be leveraged to support various use-cases. This requires that the temporal queries be processed efficiently on this data. Currently this presents significant challenges as this data is organized on file-system, is exposed to users via a limited API and does not support any temporal indexes. We present two models for overcoming these limitations and improving the performance of temporal queries on Fabric. The first model creates a copy of each event inserted by a Fabric transaction and stores temporally close events together on Fabric. The second model keeps the event count intact but tags some metadata to each event being inserted on Fabric s.t. temporally close events share the same metadata. We discuss these two models in detail and show that these two models significantly outperform the naive ways of handling temporal queries on Fabric. We also discuss the performance trade-offs for these two models across various dimensions - data storage, query performance, data ingestion time etc. Sandeep Hans, Kushagra Aggarwal, Sameep Mehta, Bapi Chatterjee, Praveen Jayachandran |
ICDE | 4 |
| 2018 | Pentuplet Loss for Simultaneous Shots and Critical Points Detection in a VideoabstractCritical events in videos amount to the set of frames where the user attention is heightened. Such events are usually fine-grained activities and do not necessarily have defined shot boundaries. Traditional approaches to the task of Shot Boundary Detection (SBD) in videos perform frame-level classification to obtain shot boundaries and fail to identify the critical shots in the video. We model the problem of identifying critical frames and shot boundaries in a video as learning an image frame similarity metric where the distance relationships between different types of video frames are modeled. We propose a novel pentuplet loss to learn the frame image similarity metric through a pentuplet-based deep learning framework. We showcase the results of our proposed framework on soccer highlight videos against state-of-the-art baselines and significantly outperform them for the task of shot boundary detection. The proposed framework shows promising results for the task of critical frame detection against human annotations on soccer highlight videos. Nitin Gupta 0005, Abhinav Jain 0001, Prerna Agarwal, Shashank Mujumdar, Sameep Mehta |
ICPR | 5 |
| 2018 | Learning an Order Preserving Image Similarity through Deep RankingabstractRecently, deep learning frameworks have been shown to learn a feature embedding that captures fine-grained image similarity using image triplets or quadruplets that consider pairwise relationships between image pairs. In real-world datasets, a class contains fine-grained categorization that exhibits within-class variability. In such a scenario, these frameworks fail to learn the relative ordering between - (i) samples belonging to the same category, (ii) samples from a different category within a class and (iii) samples belonging to a different class. In this paper, we propose the quadlet loss function, that learns an order-preserving fine-grained image similarity by learning through quadlets (query:q, positive:p, intermediate:i, negative:n) where p is sampled from the same category as q, i belongs to a fine-grained category within the class of q and n is sampled from a different class than that of q. We propose a deep quadlet network to learn the feature embedding using the quadlet loss function. We present an extensive evaluation of our proposed ranking model against state-of-the-art baselines on three datasets with fine-grained categorization. The results show significant improvement over the baselines for both order-preserving fine-grained ranking task and general image ranking task. Nitin Gupta 0005, Shashank Mujumdar, Suranjana Samanta, Sameep Mehta |
ICPR | 4 |
| 2018 | REXplore: A Sketch Based Interactive Explorer for Real Estates Using Building Floor Plan ImagesabstractThe increasing trend of using online platforms for real estate rent/sale makes automatic retrieval of similar floor plans a key requirement to help architects and buyers alike. Although sketch based image retrieval has been explored in the multimedia community, the problem of hand-drawn floor plan retrieval has been less researched in the past. In this paper, we propose REXplore (Real Estate eXplore), a novel framework that uses sketch based query mode to retrieve corresponding similar floor plan images from a repository using Cyclic Generative Adversarial Networks (Cyclic GAN) for mapping between sketch and image domain. The key contributions of our proposed approach are : (1) a novel sketch based floor plan retrieval framework using an intuitive and convenient sketch query mode; (2) A conjunction of Cyclic GANs and Convolution Neural Networks (CNNs) for the task of hand-drawn floor plan image retrieval. Extensive experimentation and comparison with baseline results authenticates our claim. Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
ISM | 4 |
| 2018 | Leveraging semantic resources in diversified query expansionabstractA search query, being a very concise grounding of user intent, could potentially have many possible interpretations. Search engines hedge their bets by diversifying top results to cover multiple such possibilities so that the user is likely to be satisfied, whatever be her intended interpretation. Diversified Query Expansion is the problem of diversifying query expansion suggestions, so that the user can specialize the query to better suit her intent, even before perusing search results. In this paper, we consider the usage of semantic resources and tools to arrive at improved methods for diversified query expansion. In particular, we develop two methods, those that leverage Wikipedia and pre-learnt distributional word embeddings respectively. Both the approaches operate on a common three-phase framework; that of first taking a set of informative terms from the search results of the initial query, then building a graph, following by using a diversity-conscious node ranking to prioritize candidate terms for diversified query expansion. Our methods differ in the second phase, with the first method Select-Link-Rank (SLR) linking terms with Wikipedia entities to accomplish graph construction; on the other hand, our second method, Select-Embed-Rank (SER), constructs the graph using similarities between distributional word embeddings. Through an empirical analysis and user study, we show that SLR ourperforms state-of-the-art diversified query expansion methods, thus establishing that Wikipedia is an effective resource to aid diversified query expansion. Our empirical analysis also illustrates that SER outperforms the baselines convincingly, asserting that it is the best available method for those cases where SLR is not applicable; these include narrow-focus search systems where a relevant knowledge base is unavailable. Our SLR method is also seen to outperform a state-of-the-art method in the task of diversified entity ranking. Adit Krishnan, Deepak P 0001, Sayan Ranu, Sameep Mehta |
World Wide Web | 4 |
| 2017 | Provenance in Context of Hadoop as a Service (HaaS) - State of the Art and Research DirectionsabstractHadoop as a service (HaaS), also known as Hadoop in the cloud, is a big data analytics framework that stores and analyzes data in the cloud using Hadoop/Spark. In this paper, we discuss the importance of providing provenance capabilities in context of Hadoop as a service (HaaS) framework. We first review the state of the art in provenance tracking in context of databases and work-flow processing, in context of cloud and in context of big data analytics frameworks like Hadoop and Spark. We next identify a number of provenance capabilities which have been developed in context of databases and workflow processing but the corresponding solutions have not been developed in context of Hadoop or Spark. We argue that developing these solutions is important so that a comprehensive provenance aware Hadoop as a Service (HaaS) can be provided on cloud. The paper ends by identifying some research challenges in developing these provenance capabilities. Sameep Mehta, Sandeep Hans, Bapi Chatterjee, Pranay Lohia, Rajmohan C |
CLOUD | 2 |
| 2017 | DANIEL: A Deep Architecture for Automatic Analysis and Retrieval of Building Floor PlansabstractAutomatically finding out existing building layouts from a repository is always helpful for an architect to ensure reuse of design and timely completion of projects. In this paper, we propose Deep Architecture for fiNdIng alikE Layouts (DANIEL). Using DANIEL, an architect can search from the existing projects repository of layouts (floor plan), and give accurate recommendation to the buyers. DANIEL is also capable of recommending the property buyers, having a floor plan image, the corresponding rank ordered list of alike layouts. DANIEL is based on the deep learning paradigm to extract both low and high level semantic features from a layout image. The key contributions in the proposed approach are: (i) novel deep learning framework to retrieve similar floor plan layouts from repository; (ii) analysing the effect of individual deep convolutional neural network layers for floor plan retrieval task; and (iii) creation of a new complex dataset ROBIN (Repository Of BuildIng plaNs), having three broad dataset categories with 510 real world floor plans.We have evaluated DANIEL by performing extensive experiments on ROBIN and compared our results with eight different state-of-the-art methods to demonstrate DANIEL's effectiveness on challenging scenarios. Nitin Gupta 0005, Chiranjoy Chattopadhyay, Sameep Mehta |
ICDAR | 4 |
| 2017 | Deep Attribute Driven Image Similarity Learning Using Limited DataabstractIn this work, we propose to derive the attribute specific similarity score for a pair of images using an existing parent deep model. As an example, given two facial images, we derive a similarity score for attributes like gender and complexion using an existing face recognition model. It is not always feasible to train a new model for each attribute, as training of deep neural network based model requires a large number of labelled samples to reliably learn the parameters. Hence, in the proposed framework a similarity score for each attribute is obtained as a weighted combination of all the hidden layer features of the parent model. The weights are attribute specific, and are estimated by minimizing the proposed triplet based hinge loss criteria over small number of labelled samples. Although generic, the proposed approach is developed in the context of a specific application to search for social media profiles of suspects of law enforcement agencies. To measure the effectiveness of our proposed approach, we have also created a social media dataset "LFW Social (LFW-S)", corresponding to the Labeled Faces in the Wild (LFW) dataset. The key motivation behind our approach is not to improve upon the existing baseline methods but to reduce the overhead of generating a labeled dataset for learning new attribute. However, it is worth noting that the learnt attribute driven models performs at par with the existing baseline models on attribute driven ranking task. Nitin Gupta 0005, Vikas Joshi, L. Venkata Subramaniam, Sameep Mehta |
ISM | 5 |
| 2017 | Coherent Visual Description of Textual InstructionsabstractText is the easiest means to record information but need not always be the best means for understanding a concept. In psychological theories, it is argued that when information is presented visually, it provides a better means to understand a concept. While techniques exist for generating text from a given image, the inverse problem that is to automatically fetch coherent images to represent a given set of instructions (sequence of text), is a hard one. In this paper, we present a novel multistage framework to convert textual instructions into coherent visual descriptions (text instructions annotated with images). The key components in the proposed approach are: (i) novel framework, which combines the text as well as image analysis to generate visual descriptions; (ii) ensure coherency across visual descriptions, using a combination of deep learning and graph based approach. Effectiveness of our proposed approach is shown through a user study on a dataset of instructions and corresponding images collected from WikiHow website. Shashank Mujumdar, Nitin Gupta 0005, Abhinav Jain 0001, Sameep Mehta |
ISM | 4 |
| 2017 | AnnoFin-A hybrid algorithm to annotate financial text
Ananda Swarup Das, Sameep Mehta, L. Venkata Subramaniam |
Expert Syst. Appl. | 2 |
| 2016 | Select, Link and Rank: Diversified Query Expansion and Entity Ranking Using Wikipedia
Adit Krishnan, Deepak P 0001, Sayan Ranu, Sameep Mehta |
WISE (1) | 4 |
| 2015 | Exploring a Scalable Solution to Identifying Events in Noisy Twitter StreamsabstractThe unprecedented use of social media through smartphones and other web-enabled mobile devices has enabled the rapid adoption of platforms like Twitter. Event detection has found many applications on the web, including breaking news identification and summarization. The recent increase in the usage of Twitter during crises has attracted researchers to focus on detecting events in tweets. However, current solutions have focused on static Twitter data. The necessity to detect events in a streaming environment during fast paced events such as a crisis presents new opportunities and challenges. In this paper, we investigate event detection in the context of real-time Twitter streams as observed in real-world crises. We highlight the key challenges in this problem: the informal nature of text, and the high-volume and high-velocity characteristics of Twitter streams. We present a novel approach to address these challenges using single-pass clustering and the compression distance to efficiently detect events in Twitter streams. Through experiments on large Twitter datasets, we demonstrate that the proposed framework is able to detect events in near real-time and can scale to large and noisy Twitter streams. Shamanth Kumar, Huan Liu 0001, Sameep Mehta, L. Venkata Subramaniam |
ASONAM | 3 |
| 2015 | Identifying Top-k Consistent News-Casters on TwitterabstractNews-casters are Twitter users who periodically pick up interesting news from online news media and spread it to their followers' network. Existing works on Twitter user analysis have only analysed a pre-defined set of users for user modeling, influence analysis and news recommendation. The problem of identifying prominent, trustworthy and consistent news-casters is unaddressed so far. In this paper, we present a framework, NCFinder, to discover top-k consistent news-casters directly from Twitter. NCFinder uses news headlines published in online news sources to periodically collect authentic news-tweets and processes them to discover news-casters, news sources and news concepts. Next, NCFinder builds a tripartite graph among news-casters, news source and news concepts and employs HITS algorithm on it to score the news-casters on daily basis. The daily score profiles of the news-casters collected over a time-period are then used to infer top-$k$ consistent news-casters. We run NCFinder from 11th Nov. to 24th Nov., 2014 and discover top-100 consistent news-casters and their profile information. Sahisnu Mazumder, Sameep Mehta, Dhaval Patel 0002 |
CIKM | 2 |
| 2015 | Entity Linking for Web Search Queries
Deepak P 0001, Sayan Ranu, Prithu Banerjee, Sameep Mehta |
ECIR | 4 |
| 2015 | Tracking Political Elections on Social Media: Applications and Experience
Danish Contractor, Bhupesh Chawda, Sameep Mehta, L. Venkata Subramaniam, Tanveer A. Faruquie |
IJCAI | 3 |
| 2015 | Being Aware of the World: Toward Using Social Media to Support the Blind With NavigationabstractThis paper lays the ground work for assistive navigation using wearable sensors and social sensors to foster situational awareness for the blind. Our system acquires social media messages to gauge the relevant aspects of an event and to create alerts. We propose social semantics that captures the parameters required for querying and reasoning an event-of-interest, such as what, where, who, when, severity, and action from the Internet of things, using an event summarization algorithm. Our approach integrates wearable sensors in the physical world to estimate user location based on metric and landmark localization. Streaming data from the cyber world are employed to provide awareness by summarizing the events around the user based on the situation awareness factor. It is illustrated using disaster and socialization event scenarios. Discovered local events are fed back using sound localization so that the user can actively participate in a social event or get early warning of any hazardous events. A feasibility evaluation of our proposed algorithm included comparing the output of the algorithm to ground truth, a survey with sighted participants about the algorithm output, and a sound localization user interface study with blind-folded sighted participants. Thus, our framework supports the navigation problem for the blind by combining the advantages of our real-time localization technologies so that the user is being made aware of the world, a necessity for independent travel. Samleo L. Joseph, Jizhong Xiao, Bhupesh Chawda, Kanika Narang, Nitendra Rajput, Sameep Mehta, L. Venkata Subramaniam |
IEEE Trans. Hum. Mach. Syst. | 7 |
| 2014 | ActMiner: Discovering Location-Specific Activities from Community-Authored Reviews
Sahisnu Mazumder, Dhaval Patel 0002, Sameep Mehta |
DaWaK | 3 |
| 2014 | Outcome aware ranking in value creation networks
Sampath Kameshwaran, Vinayaka Pandit, Sameep Mehta, Ambika Agarwal, Kashyap Dixit |
Knowl. Inf. Syst. | 3 |
| 2013 | ESTHETE: a news browsing system to visualize the context and evolution of news storiesabstractProviding the history and context(s) of a news article that emerges in the middle of an evolving news story--sometimes multiple news stories--is a complex task. The complexity of the task is compounded by the fact that different users are interested in different contexts of the article, and it is impossible to guess what a particular user is most interested in. In this paper, we introduce ESTHETE, a system that provides rich context(s) (through what we call personalized flexible context extraction), by preprocessing and storing articles in a structured representation (directed graphs) that makes it easy for the user to explore different contexts. The advantage of this approach is that the incremental computational expense in incorporating new articles as they are published is minimal. Our system is available at: http://konfrap.com/esthete. Rahul Goyal, Ravee Malla, Amitabha Bagchi, Sameep Mehta, Maya Ramanath |
CIKM | 4 |
| 2013 | Discovery and Analysis of Evolving Topical Social Discussions on Unstructured Microblogs
Kanika Narang, Seema Nagar, Sameep Mehta, L. Venkata Subramaniam, Kuntal Dey |
ECIR | 3 |
| 2013 | Efficient multifaceted screening of job applicantsabstractBuilt on top of human resources management databases within the enterprise, we present a decision support system for managing and optimizing screening activities during the hiring process in a large organization. The basic idea is to prioritize the efforts of human resource practitioners to focus on candidates that are likely of high quality, that are likely to accept a job offer if made one, and that are likely to remain with the organization for the long term. To do so, the system first individually ranks candidates along several dimensions using a keyword matching algorithm and several bipartite ranking algorithms with univariate loss trained on historical actions. Next, individual rankings are aggregated to derive a single list that is presented to the recruitment team through an interactive portal. The portal supports multiple filters that facilitate effective identification of candidates. We demonstrate the usefulness of our system on data collected from a large organization over several years with business value metrics showing greater hiring yield with less interviews. Similarly, using historical pre-hire data we demonstrate accurate identification of candidates that will have quickly left the organization. The system has been deployed as described in a large globally integrated enterprise. Sameep Mehta, Rakesh Pimplikar, Amit Singh 0003, Lav R. Varshney, Karthik Visweswariah |
EDBT | 1 |
| 2013 | An Empirical Assessment of Contemporary Online Media in Ad-Hoc Corpus Creation for Social Events
Kanika Narang, Seema Nagar, Sameep Mehta, L. Venkata Subramaniam, Kuntal Dey |
IJCNLP | 3 |
| 2013 | NLP for uncertain data at scale
Sameep Mehta, L. Venkata Subramaniam |
HLT-NAACL | 1 |
| 2013 | Topical Discussions on Unstructured Microblogs: Analysis from a Geographical Perspective
Seema Nagar, Kanika Narang, Sameep Mehta, L. Venkata Subramaniam, Kuntal Dey |
WISE (2) | 3 |
| 2013 | Towards combating rumors in social networks: Models and metricsabstractRumor is a potentially harmful social phenomenon that has been observed in all human societies in all times. Social networking sites provide a platform for the rapid interchange of information and hence, for the rapid dissemination of unsubstantiated Rudra M. Tripathy, Amitabha Bagchi, Sameep Mehta |
Intell. Data Anal. | 3 |
| 2012 | RETRAiN: A REcommendation Tool for Reconfiguration of RetAil BaNk Branch
Rakesh Pimplikar, Sameep Mehta |
ICSOC | 2 |
| 2011 | Design and Analysis of Value Creation NetworksabstractThere are many diverse domains like academic collaboration, service industry, and movies, where a group of agents are involved in a set of activities through interactions or collaborations to create value. The end result of the value creation process is two pronged: firstly, there is a cumulative value created due to the interactions and secondly, a network that captures the pattern of historical interactions between the agents. In this paper we summarize our efforts towards design and analysis of value creation networks: 1) network representation of interactions and value creations, 2) identify contribution of a node based on values created from various activities, and 3) ranking nodes based on structural properties of interactions and the resulting values. To highlight the efficacy of our proposed algorithms, we present results on IMDB and services industry data. Sampath Kameshwaran, Sameep Mehta, Vinayaka Pandit |
AAAI | 2 |
| 2011 | Simultaneously improving CSAT and profit in a retail banking organizationabstractCustomer satisfaction (CSAT) is the key driver for retention and growth in retail banking and several techniques have been applied by banks to achieve this. For instance, banks in emerging markets with high footfall in branches have gone beyond the traditional approach of segmenting customers and services to optimizing the wait time for customers visiting the bank's branch. While this approach has significantly improved service quality, it has also added a new dimension in the service quality metric : pro-actively identify and address customer needs for (i) efficient banking experience and (ii) enhancing profit by selling additional services to existing customer. In this paper we present a system that addresses the challenge involved in providing better service to retail banking customer while ensuring that a larger share of customer's wallet comes to the branch. We do this by combining predictive analytics, scheduling and process optimization techniques. Sameep Mehta, Ullas Nambiar, Vishal S. Batra, Sumit Negi, Prasad Deshpande, Gyana R. Parija |
CIKM | 1 |
| 2011 | A System for Providing Differentiated QoS in Retail Banking
Sameep Mehta, Girish Chafle, Gyana R. Parija, Vikas Kedia |
IJCAI | 1 |
| 2010 | Outcome aware ranking in interaction networksabstractIn this paper, we present a novel ranking technique that we developed in the context of an application that arose in a Service Delivery setting. We consider the problem of ranking agents of a service organization. The service agents typically need to interact with other service agents to accomplish the end goal of resolving customer requests. Their ranking needs to take into account two aspects: firstly, their importance in the network structure that arises as a result of their interactions, and secondly, the value generated by the interactions involving them. We highlight several other applications which have the common theme of ranking the participants of a value creation process based on the network structure of their interactions and the value generated by their interactions. We formally present the problem and describe the modeling technique which enables us to encode the value of interaction in the graph. Our ranking algorithm is based on extension of eigen value methods. We present experimental results on real-life, public domain datasets from the Internet Movie DataBase. This makes our experiments replicable and verifiable. Sampath Kameshwaran, Vinayaka Pandit, Sameep Mehta, Nukala Viswanadham, Kashyap Dixit |
CIKM | 3 |
| 2010 | A study of rumor control strategies on social networksabstractIn this paper we study and evaluate rumor-like methods for combating the spread of rumors on a social network. We model rumor spread as a diffusion process on a network and suggest the use of an "anti-rumor" process similar to the rumor process. We study two natural models by which these anti-rumors may arise. The main metrics we study are the belief time, i.e., the duration for which a person believes the rumor to be true and point of decline, i.e., point after which anti-rumor process dominates the rumor process. We evaluate our methods by simulating rumor spread and anti-rumor spread on a data set derived from the social networking site Twitter and on a synthetic network generated according to the Watts and Strogatz model. We find that the lifetime of a rumor increases if the delay in detecting it increases, and the relationship is at least linear. Further our findings show that coupling the detection and anti-rumor strategy by embedding agents in the network, we call them beacons, is an effective means of fighting the spread of rumor, even if these beacons do not share information. Rudra M. Tripathy, Amitabha Bagchi, Sameep Mehta |
CIKM | 3 |
| 2009 | Analyses for Service Interaction Networks with Applications to Service DeliveryabstractOne of the distinguishing features of the services industry is the high emphasis on people interacting with other people and serving customers rather than transforming physical goods like in the traditional manufacturing processes. It is evident that analysis of such interactions is an essential aspect of designing effective and efficient services delivery. In this work we focus on learning individual and team behavior of different people or agents of a service organization by studying the patterns and outcomes of historical interactions. For each past interaction, we assume that only the list of participants and an outcome indicating the overall effectiveness of the interaction are known. Note that this offers limited information on the mutual (pairwise) compatibility of different participants. We develop the notion of service interaction networks which is an abstraction of the historical data and allows one to cast practical problems in a formal setting. We identify the unique characteristics of analyzing service interaction networks when compared to traditional analyses considered in social network analysis and establish a need for new modeling and algorithmic techniques for such networks. On the algorithmic front, we develop new algorithms to infer attributes of agents individually and in team settings. Our first algorithm is based on a novel modification to the eigen-vector based centrality for ranking the agents and the second algorithm is an iterative update technique that can be applied for subsets of agents as well. One of the challenges of conducting research in this setting is the sensitive and proprietary nature of the data. Therefore, there is a need for a realistic simulator for studying service interaction networks. We present the initial version of our simulator that is geared to capture several characteristics of service interaction networks that arise in real-life. Sampath Kameshwaran, Sameep Mehta, Vinayaka Pandit, Gyana R. Parija, Sudhanshu Singh, Nukala Viswanadham |
SDM | 2 |
| 2008 | On quantifying changes in temporally evolving datasetabstractIn this paper, we present a general framework to quantify changes in temporally evolving data. We focus on changes that materialize due to evolution and interactions of features extracted from the data. The changes are captured by the following key transformations: create, merge, split, continue, and cease. First, we identify various factors which influence the importance of each transformation. These factors are then combined using a weight vector. The weight vector encapsulates domain knowledge. We evaluate our algorithm using the following datasets: DBLP, IMDB, Text and Scientific Dataset. Rohan Choudhary, Sameep Mehta, Amitabha Bagchi |
CIKM | 2 |
| 2008 | Towards Characterization of Actor Evolution and Interactions in News Corpora
Rohan Choudhary, Sameep Mehta, Amitabha Bagchi, Rahul Balakrishnan |
ECIR | 2 |
| 2008 | A visual-analytic toolkit for dynamic interaction graphsabstractIn this article we describe a visual-analytic tool for the interrogation of evolving interaction network data such as those found in social, bibliometric, WWW and biological applications. The tool we have developed incorporates common visualization paradigms such as zooming, coarsening and filtering while naturally integrating information extracted by a previously described event-driven framework for characterizing the evolution of such networks. The visual front-end provides features that are specifically useful in the analysis of interaction networks, capturing the dynamic nature of both individual entities as well as interactions among them. The tool provides the user with the option of selecting multiple views, designed to capture different aspects of the evolving graph from the perspective of a node, a community or a subset of nodes of interest. Standard visual templates and cues are used to highlight critical changes that have occurred during the evolution of the network. A key challenge we address in this work is that of scalability - handling large graphs both in terms of the efficiency of the back-end, and in terms of the efficiency of the visual layout and rendering. Two case studies based on bibliometric and Wikipedia data are presented to demonstrate the utility of the toolkit for visual knowledge discovery. Xintian Yang, Sitaram Asur, Srinivasan Parthasarathy 0001, Sameep Mehta |
KDD | 4 |
| 2008 | ReCon: A tool to Recommend dynamic server Consolidation in multi-cluster data centersabstractRenewed focus on virtualization technologies and increased awareness about management and power costs of running under-utilized servers has spurred interest in consolidating existing applications on fewer number of servers in the data center. The ability to migrate virtual machines dynamically between physical servers in real-time has also added a dynamic aspect to consolidation. However, there is a lack of planning tools that can analyze historical data collected from an existing environment and compute the potential benefits of server consolidation especially in the dynamic setting. In this paper we describe such a consolidation recommendation tool, called ReCon. Recon takes static and dynamic costs of given servers, the costs of VM migration, the historical resource consumption data from the existing environment and provides an optimal dynamic plan of VM to physical server mapping over time. We also present the results of applying the tool on historical data obtained from a large production environment. Sameep Mehta, Anindya Neogi |
NOMS | 1 |
| 2006 | Robust periodicity detection algorithmsabstractPeriodicity detection is an important pre-processing step for many time series algorithms. It provides important information about the structural properties of a time series. Feature vectors based on periodicity can be used for clustering, classification, abnormality detection, and human motion understanding. The periodicity detection task is not difficult in case of simple and uncontaminated signal. Unfortunately, most of the real datasets exhibit one or more of the following properties: i) non-stationarity, ii) interlaced cyclic patterns and iii) data contamination, which makes the period detection extremely challenging. A seemingly straightforward solution is to develop individual specialized algorithms for handling each case separately. However, determining if a time series is non-stationary or is contaminated in itself is an extremely difficult task. In this article, we propose generic algorithms which can detect periods in complex, noisy and incomplete datasets. The algorithm leverages the frequency characterization and autocorrelation structure inherent in a time series to estimate its periodicity. We extend the methods to handle non-stationary time series by tracking the candidate periods using a Kalman filter. We also address the interesting problem of finding multiple interlaced periodicities. Srinivasan Parthasarathy 0001, Sameep Mehta |
CIKM | 2 |
| 2006 | On Trajectory Representation for Scientific FeaturesabstractIn this article, we present trajectory representation algorithms for tangible features found in temporally varying scientific datasets. Rather than modeling the features as points, we take attributes like shape and extent of the feature into account. Our contention is that these attributes play an important role in understanding the temporal evolution and interactions among features. The proposed representation scheme is based on motion and shape parameters including linear velocity, angular velocity, etc. We use these parameters to segment the trajectory instead of relying on the geometry of the trajectory. We evaluate our algorithms on real datasets originating from different domains. We show the accuracy of the motion and shape parameter estimation by reconstructing the trajectories with high accuracy. Finally, we present performance and scalability results. Sameep Mehta, Srinivasan Parthasarathy 0001, Raghu Machiraju |
ICDM | 1 |
| 2005 | Mining Spatial Object Associations for Scientific Data
Hui Yang 0002, Srinivasan Parthasarathy 0001, Sameep Mehta |
IJCAI | 3 |
| 2005 | A generalized framework for mining spatio-temporal patterns in scientific dataabstractIn this paper, we present a general framework to discover spatial associations and spatio-temporal episodes for scientific datasets. In contrast to previous work in this area, features are modeled as geometric objects rather than points. We define multiple distance metrics that take into account objects' extent and thus are more robust in capturing the influence of an object on other objects in spatial neighborhood. We have developed algorithms to discover four different types of spatial object interaction (association) patterns. We also extend our approach to accommodate temporal information and propose a simple algorithm to derive spatio-temporal episodes. We show that such episodes can be used to reason about critical events. We evaluate our framework on real datasets to demonstrate its efficacy. The datasets originate from two different areas: Computational Molecular Dynamics and Computational Fluid Flow. We present results highlighting the importance of the identified patterns and episodes by using knowledge from the underlying domains. We also show that the proposed algorithms scale linearly with respect to the dataset size. Hui Yang 0002, Srinivasan Parthasarathy 0001, Sameep Mehta |
KDD | 3 |
| 2005 | Dynamic Classification of Defect Structures in Molecular Dynamics Simulation DataabstractIn this application paper we explore techniques to classify anomalous structures (defects) in data generated from Molecular Dynamics (MD) simulations of Silicon (Si) atom systems. These systems are studied to understand the processes behind the formation of various defects as they have a profound impact on the electrical and mechanical properties of Silicon. In our prior work [12, 13, 14] we presented techniques for defect detection. Here, we present a two-step dynamic classifier to classify the defects. The first step uses up to third-order shape moments to provide a smaller set of candidate defect classes. The second step assigns the correct class to the defect structure by considering the actual spatial positions of the individual atoms. The dynamic classifier is robust and scalable in the size of the atom systems. Each phase is immune to noise, which is characterized after a study of the simulation data. We also validate the proposed solutions by using a physical model and properties of lattices. We demonstrate the efficacy and correctness of our approach on several large datasets. Our approach is able to recognize previously seen defects and also identify new defects in real time. Sameep Mehta, Steve Barr, Tat-Sang Choy, Hui Yang 0002, Srinivasan Parthasarathy 0001, Raghu Machiraju, John Wilkins |
SDM | 1 |
| 2005 | Toward Unsupervised Correlation Preserving DiscretizationabstractDiscretization is a crucial preprocessing technique used for a variety of data warehousing and mining tasks. In this paper, we present a novel PCA-based unsupervised algorithm for the discretization of continuous attributes in multivariate data sets. The algorithm leverages the underlying correlation structure in the data set to obtain the discrete intervals and ensures that the inherent correlations are preserved. Previous efforts on this problem are largely supervised and consider only piecewise correlation among attributes. We consider the correlation among continuous attributes and, at the same time, also take into account the interactions between continuous and categorical attributes. Our approach also extends easily to data sets containing missing values. We demonstrate the efficacy of the approach on real data sets and as a preprocessing step for both classification and frequent itemset mining tasks. We show that the intervals are meaningful and can uncover hidden patterns in data. We also show that large compression factors can be obtained on the discretized data sets. The approach is task independent, i.e., the same discretized data set can be used for different data mining tasks. Thus, the data sets can be discretized, compressed, and stored once and can be used again and again. Sameep Mehta, Srinivasan Parthasarathy 0001, Hui Yang 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Correlation Preserving DiscretizationabstractDiscretization is a crucial preprocessing primitive for a variety of data warehousing and mining tasks. In this article we present a novel PCA-based unsupervised algorithm for the discretization of continuous attributes in multivariate datasets. The algorithm leverages the underlying correlation structure in the dataset to obtain the discrete intervals, and ensures that the inherent correlations are preserved. The approach also extends easily to datasets containing missing values. We demonstrate the efficacy of the approach on real datasets and as a preprocessing step for both classification and frequent item set mining tasks. We also show that the intervals are meaningful and can uncover hidden patterns in data. Sameep Mehta, Srinivasan Parthasarathy 0001, Hui Yang 0002 |
ICDM | 1 |
| 2004 | Detection and Visualization of Anomalous Structures in Molecular Dynamics Simulation DataabstractWe explore techniques to detect and visualize features in data from molecular dynamics (MD) simulations. Although the techniques proposed are general, we focus on silicon (Si) atomic systems. The first set of methods use 3D location of atoms. Defects are detected and categorized using local operators and statistical modeling. Our second set of exploratory techniques employ electron density data. This data is visualized to glean the defects. We describe techniques to automatically detect the salient isovalues for isosurface extraction and designing transfer functions. We compare and contrast the results obtained from both sources of data. Essentially, we find that the methods of defect (feature) detection are at least as robust as those based on the exploration of electron density for Si systems. Sameep Mehta, Kaden Hazzard, Raghu Machiraju, Srinivasan Parthasarathy 0001, John Wilkins |
IEEE Visualization | 1 |
| 2003 | Feature Mining Paradigms for Scientific DataabstractNumerical simulation is replacing experimentation as a means to gain insight into complex physical phenomena. Analyzing the data produced by such simulations is extremely challenging, given the enormous sizes of the datasets involved. In order to make efficient progress, analyzing such data must advance from current techniques that only visualize static images of the data, to novel techniques that can mine, track, and visualize the important features in the data. In this paper, we present our research on a unified framework that addresses this critical challenge in two science domains: computational fluid dynamics and molecular dynamics. We offer a systematic approach to detect the significant features in both domains, characterize and track them, and formulate hypotheses with regard to their complex evolution. Our framework includes two paradigms for feature mining, and the choice of one over the other, for a given application, can be determined based on local or global influence of relevant features in the data. Ming Jiang 0005, Tat-Sang Choy, Sameep Mehta, Matt Coatney, Steve Barr, Kaden Hazzard, David Richie, Srinivasan Parthasarathy 0001, Raghu Machiraju, David S. Thompson, John Wilkins, Boyd Gatlin |
SDM | 3 |