Elham Khabiri

dblp:46/960 · DBLP profile ↗
← Back
8ranked-venue papers in the field
4as first author
3since 2021 · last 2022
0000-0001-6170-419XORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 4 (3 first)Big Data, Cloud & Distributed Data Systems · 4 (1 first)
YearPublicationVenuePosition
2022 Big data techniques for industrial problems with little data
abstract
Technicians and maintenance managers in industrial environments would benefit from automatically extracting entities and relationships from different text data sources such as logs, event reports, and manuals. Extracting components from pieces of text and classifying them to the right failure type is not trivial in the domain specific setting where the vocabulary has specific meaning to the industry or domain, and labeled data set is very small. In this paper we address how to overcome these challenges in named entity recognition and classification of text, and present a way to improve the model iteratively and quickly. This interaction between components and related failures in the system can be represented in a knowledge graph, which enables further investigations such as Root Cause Analysis and Problem Diagnosis.
Elham Khabiri, Bhavna Agrawal, Joseph Lindquist, Anuradha Bhamidipaty
IEEE Big Data1
2021 Asset Modeling using Serverless Computing
abstract
Assets in the domain of Internet of Things (IoT) generate time-series data such as sensor readings and alerts. In addition, the assets have associated static data such as the make, model and other manufacturing information. The sensors in the asset components may have implicit relationships with each other, which are not interpretable without domain knowledge. Many problems exist which involve computation of relationships between sensors or subsystems in the asset components. Typically, the number of sensors in a real world asset may range anywhere from tens to thousands of sensors - and in this case, finding relationships between them becomes a highly computationally intensive task. In this paper, we study one such problem of anomaly detection in industrial data based on the functioning of the sensors and their interrelationships in both normal and abnormal conditions. We further demonstrate the issue of run-time and performance complexity in this problem, and present a speed-up strategy using Serverless Computing for parallelization, and demonstrate the usefulness of this method by comparing the speed-up achieved.
Srideepika Jayaraman, Chandra Reddy, Elham Khabiri, Dhaval Patel 0002, Anuradha Bhamidipaty, Jayant Kalagnanam
IEEE BigData3
2021 TableNN: Deep Learning Framework for Learning Domain Specific Tabular Data
abstract
Enterprises often have a large number of databases and other sources of tabular data with columns full of domain-specific jargon (e.g. alpha-numeric codings, undeclared abbreviations, etc) which usually require domain experts to decode. Due to the jargon-specific content of the tables, no pre-trained language model such as Wiki2Vec [21] can be applied readily to encode the cell semantics due to absence of unique jorgan words or alpha-numeric codes in the model vocabulary. We propose a deep learning based framework that is ideally suited for serverless computing environment, and that 1) uses a new tokenization method, called Cell-Masking, 2) encodes the semantics of the cells into contextual embedding that exploits the locality features in tabular data, called Cell2Vec, and 3) an attention-based neural network, called TableNN, that provides a supervised learning solution to classify cell entries into predefined column classes. We apply the proposed method on three publicly available datasets of varying data sizes, from different industries. Cell-Masking provides an order of magnitude lower loss value and quickest convergence for cell embedding generation. In Cell2Vec, we demonstrate that the inclusion of row and column context improves the quality of embeddings by better loss curve convergence and improvement in accuracy by 5.4% on the BTS dataset [3].
Pranav Sankhe, Elham Khabiri, Bhavna Agrawal
IEEE BigData2
2020 An End-to-End Context Aware Anomaly Detection System
abstract
Anomaly detection (AD) is very important across several real-world problems in the heavy industries and Internet-of-Things (IoT) domains. Traditional methods so far have categorized anomaly detection into (a) unsupervised, (b) semi-supervised and (c) supervised techniques. A relatively unexplored direction is the development of context aware anomaly detection systems which can build on top of any of these three techniques by using side information. Context can be captured from a different modality such as semantic graphs encoding grouping of sensors governed by the physics of the asset. Process flow diagrams of an operational plant depicting causal relationships between sensors can also provide useful context for ML algorithms. Capturing such semantics by itself can be pretty challenging, however, our paper mainly focuses on, (a) designing and implementing effective anomaly detection pipelines using sparse Gaussian Graphical Models with various statistical distance metrics, and (b) differentiating these pipelines by embedding contextual semantics inferred from graphs so as to obtain better KPIs in practice. The motivation for the latter of these two has been explained above, and the former in particular is well motivated by the relatively mediocre performance of highly parametric deep learning methods for small tabular datasets (compared to images) such as IoT sensor data. In contrast to such traditional automated deep learning (AutoAI) techniques, our anomaly detection system is based on developing semantics-driven industry specific ML pipelines which perform scalable computation evaluating several models to identify the best model. We benchmark our AD method against state-of-the-art AD techniques on publicly available UCI datasets. We also conduct a case study on IoT sensor and semantic data procured from a large thermal energy asset to evaluate the importance of semantics in enhancing our pipelines. In addition, we also provide explainable insights for our model which provide a complete perspective to a reliability engineer.
Bhanukiran Vinzamuri, Elham Khabiri, Anuradha Bhamidipaty, Gregory Mckim, Biren Gandhi
IEEE BigData2
2019 Industry Specific Word Embedding and its Application in Log Classification
abstract
Word, sentence and document embeddings have become the cornerstone of most natural language processing-based solutions. The training of an effective embedding depends on a large corpus of relevant documents. However, such corpus is not always available, especially for specialized heavy industries such as oil, mining, or steel. To address the problem, this paper proposes a semi-supervised learning framework to create document corpus and embedding starting from an industry taxonomy, along with a very limited set of relevant positive and negative documents. Our solution organizes candidate documents into a graph and adopts different explore and exploit strategies to iteratively create the corpus and its embedding. At each iteration, two metrics, called Coverage and Context Similarity, are used as proxy to measure the quality of the results. Our experiments demonstrate how an embedding created by our solution is more effective than the one created by processing thousands of industry-specific document pages. We also explore using our embedding in downstream tasks, such as building an industry specific classification model given labeled training data, as well as classifying unlabeled documents according to industry taxonomy terms.
Elham Khabiri, Wesley M. Gifford, Bhanukiran Vinzamuri, Dhaval Patel 0002, Pietro Mazzoleni
CIKM1
2015 Creating Diverse Product Review Summaries: A Graph Approach
Natwar Modani, Elham Khabiri, Harini Srinivasan, James Caverlee
WISE (1)2
2011 Summarizing User-Contributed Comments
Elham Khabiri, James Caverlee, Chiao-Fang Hsu
ICWSM1
2009 Analyzing and Predicting Community Preference of Socially Generated Metadata: A Case Study on Comments in the Digg Community
Elham Khabiri, Chiao-Fang Hsu, James Caverlee
ICWSM1