Alex Sim

dblp:75/4666 · also Alexander Sim · DBLP profile ↗
← Back
22ranked-venue papers in the field
1as first author
5since 2021 · last 2024
0000-0002-6295-1982ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 16Database Systems & Data Management · 6 (1 first)
YearPublicationVenuePosition
2024 TensorSearch: Parallel Similarity Search on Tensors
abstract
Existing similarity search methods, often limited to scalar or vector data, struggle to identify complex patterns found in scientific datasets, such as 2D seismic events or 3D magnetic flux ropes. We introduce TensorSearch, a novel parallel similarity search paradigm designed to identify known patterns in high-dimensional tensors. By directly employing tensor representations, TensorSearch captures intricate pattern structures more effectively than traditional vector-based approaches. Furthermore, its parallel architecture optimizes cache and I/O operations, enabling efficient processing of large-scale scientific data. Our performance evaluations demonstrate that TensorSearch outperforms state-of-the-art vector-based systems like Milvus by up to 10x, and achieves up to a remarkable 55x advantage over custom solution developed in Matlab used by the domain scientists. In these tests, TensorSearch exhibits linear scalability, supporting up to 2240 CPU cores.
Bin Dong 0002, Avinash Nayak, Verónica Rodríguez Tribaldos, Kesheng Wu, Jonathan Ajo-Franklin, Qile Zhang, Surendra Byna, Patrick Dobson, Alex Sim
IEEE Big Data10
2024 Serving Deep Learning Models from Relational Databases
Lixi Zhou, Kanchan Chowdhury, Saif Masood, Alexandre E. Eichenberger, Hong Min, Alex Sim, Kesheng Wu, Binhang Yuan, Jia Zou 0001
EDBT7
2023 Counterfactual Analysis: A Case Study on Impact of External Events on Building Energy Consumption
abstract
Energy consumption in buildings accounts for a significant portion of the global energy use. Consequently, understanding building energy use is important. Data over the past decade show that the energy intensity (Joules/sqft) of commercial buildings has decreased. While some of the improvements (decrease in energy use) are easily measurable such as the use of more energy efficient lighting, impact of other modifications such as changes to the operation of the HVAC system or changes in the usage pattern of the building potentially due to external events are difficult to quantify. Simply comparing energy consumption prior and post change is not accurate as energy use is impacted by many factors including external weather conditions. In this paper, we present a case study to quantify the impact of external events on the energy consumption of a medium-sized office building. We adopt an approach based on counterfactual analysis. Towards this end, we first build two models based on Linear Regression and k-Nearest Neighbors to predict the daily energy use given different input features related to the weather. We determine the statistical features of the weather that are most predictive of energy use. We then use the models to determine a counterfactual baseline and thereby to accurately estimate the impact of the events. The results of the counterfactual analysis provide new insights on the impact of the events on energy consumption. The update to the building cooling system resulted in more energy savings than direct yearly comparison reveals. On the other hand, the tests of a MPC-based controller for the HVAC system saved less energy than determined by the direct yearly comparison. Finally, the results show that there no gains in terms of energy savings due to remote work during the COVID-19 pandemic. An increase in airflow setting in the HVAC system corroborates this finding and further validates the underlying model and the counterfactual analyses.
Carolina Minami Oguchi, Dipak Ghosal, Alex Sim, Kesheng Wu
IEEE Big Data3
2023 Automatic Data Transformation Using Large Language Model - An Experimental Study on Building Energy Data
abstract
Existing approaches to automatic data transformation are insufficient to meet the requirements in many real-world scenarios, such as the building sector. First, there is no convenient interface for domain experts to provide domain knowledge easily. Second, they require significant training data collection overheads. Third, the accuracy suffers from complicated schema changes. To address these shortcomings, we present a novel approach that leverages the unique capabilities of large language models (LLMs) in coding, complex reasoning, and zero-shot learning to generate SQL code that transforms the source datasets into the target datasets. We demonstrate the viability of this approach by designing an LLM-based framework, termed SQLMorpher, which comprises a prompt generator that integrates the initial prompt with optional domain knowledge and historical patterns in external databases. It also implements an iterative prompt optimization mechanism that automatically improves the prompt based on flaw detection. The key contributions of this work include (1) pioneering an end-to-end LLM-based solution for data transformation, (2) developing a benchmark dataset of 105 real-world building energy data transformation problems, and (3) conducting an extensive empirical evaluation where our approach achieved 96% accuracy in all 105 problems. SQLMorpher demonstrates the effectiveness of utilizing LLMs in complex, domain-specific challenges, highlighting the potential of their potential to drive sustainable solutions.
Xuanmao Li, Guoxin Sun, Liang Zhang 0048, Lanjun Wang, Kesheng Wu, Lei Cao 0004, Erkang Zhu, Alex Sim, Teresa Wu, Jia Zou 0001
IEEE Big Data10
2021 Performance of the Gold Standard and Machine Learning in Predicting Vehicle Transactions
abstract
Logistic regression has long been the gold standard for choice modeling in the transportation field. Despite the rising popularity of machine learning (ML), few is applied to predicting the household vehicle transactions. To address the research gap, this paper presents a first use case of ML application to predicting household vehicle transaction decisions by leveraging a newly processed national panel data set. Model performances are reported for four ML models and the traditional multinomial logit model (MNL). Instead of treating the gold standard and ML models as competitors, this paper tries to use ML tools to inform the MNL model building process. We find the two gradient boosting based methods, CatBoost and LightGBM, are the best performing ML models; and improving logistic models with SHAP interpretation tools can achieve similar performance levels to the best performing ML methods.
Alina Lazar, Ling Jin 0001, Caitlin Brown, Anna Spurlock, Alex Sim, Kesheng Wu
IEEE BigData5
2020 Effective Missing Value Imputation Methods for Building Monitoring Data
abstract
To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.
Brian Cho 0001, Teresa Dayrit, Zhe Wang 0051, Tianzhen Hong, Alex Sim, Kesheng Wu
IEEE BigData6
2019 Federated Wireless Network Intrusion Detection
abstract
Wi-Fi has become the wireless networking standard that allows short- to medium-range device to connect without wires. For the last 20 year, the Wi-Fi technology has so pervasive that most devices in use today are mobile and connect to the internet through Wi-Fi. Unlike wired network, a wireless network lacks a clear boundary, which leads to significant Wi-Fi network security concerns, especially because the current security measures are prone to several types of intrusion. To address this problem, machine learning and deep learning methods have been successfully developed to identify network attacks. However, collecting data to develop models is expensive and raises privacy concerns. The goal of this paper is to evaluate a federated learning approach that would alleviate such privacy concerns. This initial work on intrusion detection is performed in a simulated environment. Once proven feasible, this process would allow edge devices to collaboratively update global anomaly detection models, without sharing sensitive training data. On a set of tests with the AWID intrusion detection data set, we show that our federated approach is effective in terms of classification accuracy, computation cost, as well as communication cost.
Burak Cetin, Alina Lazar, Jinoh Kim, Alex Sim, Kesheng Wu
IEEE BigData4
2019 Spatiotemporal Real-Time Anomaly Detection for Supercomputing Systems
abstract
The demands of increasingly large scientific application workflows lead to the need for more powerful supercomputers. As the scale of supercomputing systems have grown, the prediction of fault tolerance has become an increasingly critical area of study, since the prediction of system failures can improve performance by saving checkpoints in advance. We propose a real-time failure detection algorithm that adopts an event-based prediction model. The prediction model is a convolutional neural network that utilizes both traditional event attributes and additional spatio-temporal features. We present a case study using our proposed method with six years of reliability, availability, and serviceability event logs recorded by Mira, a Blue Gene/Q supercomputer at Argonne National Laboratory. In the case study, we have shown that our failure prediction model is not limited to predict the occurrence of failures in general. It is capable of accurately detecting specific types of critical failures such as coolant and power problems within reasonable lead time ranges. Our case study shows that the proposed method can achieve a F1score of 0.56 for general failures, 0.97 for coolant failures, and 0.86 for power failures.
Qiao Kang, Ankit Agrawal 0001, Alok N. Choudhary, Alex Sim, Kesheng Wu, Rajkumar Kettimuthu, Pete Beckman, Zhengchun Liu, Wei-keng Liao
IEEE BigData4
2019 Machine Learning for Prediction of Mid to Long Term Habitual Transportation Mode Use
abstract
Prediction of daily transportation mode use (car, public transit, or active travel) is a important task in transportation research. Unlike statistical models that impose a predetermined model structure, machine learning models are learned from the data, making them more flexible with higher prediction accuracy. However, prediction of mid-to long-term habitual modes still largely relies on traditional statistical analysis using small samples of cross-sectional data. Low interpretability of “black-box” machine learning models limits their usefulness for generating behavior insights needed for designing appropriate interventions. This paper, leveraging a set of unique longitudinal life course data, is the first use case to demonstrate machine learning methods applied for both predicting and interpreting regularly used travel modes. We combine sequence clustering and tree-based machine learning methods coupled with TreeExplainer to predict and interpret habitual travel modes using mid-to long-term predictors. Five life course clusters are derived to provide evaluation and interpretation contexts. This allows us to improve upon a recently developed TreeExplainer method to better distinguish predictor importance locally and globally; and predictor interactions across subpopulations within distinctive life history contexts. Our results demonstrate a promising step toward interpretable machine learning applications to mid-to long-term prediction of travel modes for transportation planning.
Alina Lazar, Alexandra Ballow, Ling Jin 0001, Anna Spurlock, Alex Sim, Kesheng Wu
IEEE BigData5
2019 Multidimensional Compression with Pattern Matching
abstract
Sensors typically record their measurements using more precision than the accuracy of the sensing techniques. Thus, experimental and observational data often contains noise that appears random and cannot be easily compressed. This noise increases storage requirement as well as computation time for analyses. In this work, we describe a line of research to develop data reduction techniques that preserve the key features while reduce the storage requirement. Our core observation is that the noise in such cases could be characterized by a small number of patterns based on statistical similarity. In earlier tests, this approach was shown to reduce the storage requirement by over 100-fold for one-dimensional sequences. In this work, we explore a set of different similarity measures for multidimensional sequences. During our tests with standard quality measures such as PSNR, we see that the new compression methods reduce the storage requirements over 100-fold while maintaining relatively low errors in peak signal-to-noise ratio. Thus, we believe that this is a new and effective way of constructing data reduction techniques.
Olivia Del Guercio, Rafael Orozco, Alex Sim, Kesheng Wu
DCC3
2018 Dynamic Online Performance Optimization in Streaming Data Compression
abstract
Compression is essential to high bandwidth applications such as scientific simulations and sensing applications to reduce resource burden such as storage, network transmission, and more recently I/O. Existing lossy compression methods attempt to minimize the Euclidean distance between original data and reconstructed data, which significantly limits either compression performance or reconstruction quality since original and reconstructed data sequences should be aligned. Substituting the Euclidean distance for a statistical similarity maximizes the compression performance while retaining essential data features. By implementing this methodology, IDEALEM has recently demonstrated compression ratios far exceeding 100:1, better than best-known compression methods, while preserving reconstruction quality. This work proposes an online algorithm for streaming data compression which takes account of generally concave trend of compression ratio curve, and optimizes key operation parameters. We demonstrate that the proposed algorithm successfully adapts one of the key parameters in IDEALEM to the optimal value and yields near maximum compression ratios for time series data.
J. Kade Gibson, Dongeun Lee 0001, Jaesik Choi, Alex Sim
IEEE BigData4
2018 Predicting Network Traffic Using TCP Anomalies
abstract
Accurately predicting network traffic volume is beneficial for congestion control, improving routing, allocating network resources and network optimization. Traffic congestion happens when a network device is receiving more data packets than its processing capability. The number of retransmissions per flow, packet duplication and synthetic reordering can seriously degrade the overall TCP performance. An unsupervised/supervised technique to accurately identify TCP anomalies occurring during file transfers based on passive measurements of TCP traffic collected using Tstat is proposed. This method will be validated on real large datasets collected from several data transfer nodes. The preliminary results indicate that the percentage of TCP anomalies correlate well with the average throughput in any given time window.
Alina Lazar, Kesheng Wu, Alex Sim
IEEE BigData3
2018 Detecting Anomalies in the LCLS Workflow
abstract
The Linac Coherent Light Source (LCLS) located at SLAC National Accelerator Laboratory has been essential to over 1023 publications since 2009. The LCLS produces vast quantities of data - thousands of gigabytes per experiment. The data must be analyzed and stored at large data centers to be available to the world-wide user community. Due to the vast quantities of data flowing through the network, many abnormal data transfers remain unnoticed. This work focuses on identifying network failures that could slow down the data transfer process. This work aims to develop a diagnostic tool to detect when network transfers become anomalously slow. The tool uses an algorithm based on the hampel filter to detect poor performance and alert SLAC administrators to bottlenecks in each phase of the workflow. We will describe our experience of preparing the data and modifying the hampel filter to enhance its effectiveness. We found that applying a heuristic to the algorithm in conjunction with parsing the data along key features improved performance.
Tal Shachaf, Alex Sim, Kesheng Wu, Wilko Kroeger
IEEE BigData2
2017 Feature Engineering and Classification Models for Partial Discharge Events in Power Transformers
abstract
To ensure the reliability of power transformers, they are monitored for partial discharge (PD) events, which are symptoms of trans- former failure. Our goal is to classify PDs to gain an understanding of the location of failure. We develop a small set of features and a stacking ensemble that outperform larger feature sets and other models in both accuracy and variaTo ensure the reliability of power transformers, they are monitored for partial discharge (PD) events, which are symptoms of transformer failure. Our goal is to classify PDs to gain an understanding of the location of failure. We develop a small set of features and a stacking ensemble that outperform larger feature sets and other models in both accuracy and variance.nce.
Kesheng Wu, Alex Sim, Seongwook Hwangbo
BDCAT3
2017 Data quality challenges with missing values and mixed types in joint sequence analysis
abstract
The goal of this paper is to investigate the impact of missing values in categorical time series sequences on common data analysis tasks. Being able to more effectively identify patterns in socio-demographic longitudinal data is an important component in a number of social science settings. However, performing fundamental analytical operations, such as clustering for grouping these data based on similarity patterns, is challenging due to the categorical and multi-dimensional nature of the data, and their corruption by missing and inconsistent values. To study these data quality issues, we employ longitudinal sequence data representations, a similarity measure designed for categorical and longitudinal data, together with state-of-the art clustering methodologies reliant on hierarchical algorithms. The key to quantifying the similarity and difference among data records is a distance metric. Given the categorical nature of our data, we employ an “edit” type distance using Optimal Matching (OM). Because each data record has multiple variables of different types, we investigate the impact of mixing these variables in a single similarity measure. Between variables with binary values and those with multiple nominal values, we find that the ability to overcome missing data problems is harder in the nominal domain versus the binary domain. Additionally, artificial clusters introduced by the alignment of leading missing values can be resolved by tuning the missing value substitution cost parameter.
Alina Lazar, Ling Jin 0001, Anna Spurlock, Kesheng Wu, Alex Sim
IEEE BigData5
2017 Accurate signal timing from high frequency streaming data
abstract
The goal of our study is to analyze massive high-frequency streaming sensor data to accurately locate the source of partial discharges (PD) in transformers. The PD signal is collected by ultra-high frequency sensors at a resolution of 0.4 ns per record resulting in a data streaming rate of 12 GB/s. A voltage threshold is applied to the data stream to extract 400 ns signal samples. We develop a voltage threshold method based on the Savitzky-Golay filter for signal arrival timing, and localize the PD with arrival time differences using Finite-Difference Time-Domain (FDTD) simulation. The Savitzky-Golay filter is able to preserve features better than other methods, resulting in improved signal-to-noise ratios and more accurate signal timing. FDTD accounts for the travel path of signals inside the transformer, allowing for more precise PD localization. Our resulting method localizes PDs more accurately than existing methods, particularly in high noise cases.
Kesheng Wu, Alex Sim, Seongwook Hwangbo
IEEE BigData3
2017 Expanding Statistical Similarity Based Data Reduction to Capture Diverse Patterns
abstract
We propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
DCC2
2017 Improving Statistical Similarity Based Data Reduction for Non-Stationary Data
abstract
We propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
SSDBM2
2016 Novel Data Reduction Based on Statistical Similarity
abstract
Applications such as scientific simulations and power grid monitoring are generating so much data quickly that compression is essential to reduce storage requirement or transmission capacity. To achieve better compression, one is often willing to discard some repeated information. These lossy compression methods are primarily designed to minimize the Euclidean distance between the original data and the compressed data. But this measure of distance severely limits either reconstruction quality or compression performance. We propose a new class of compression method by redefining the distance measure with a statistical concept known as exchangeability. This approach reduces the storage requirement and captures essential features, while reducing the storage requirement. In this paper, we report our design and implementation of such a compression method named IDEALEM. To demonstrate its effectiveness, we apply it on a set of power grid monitoring data, and show that it can reduce the volume of data much more than the best known compression method while maintaining the quality of the compressed data. In these tests, IDEALEM captures extraordinary events in the data, while its compression ratios can far exceed 100.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
SSDBM2
2004 DataMover: Robust Terabyte-Scale Multi-file Replication over Wide-Area Networks
Alex Sim, Junmin Gu, Arie Shoshani, Vijaya Natarajan
SSDBM1
2000 Coordinating Simultaneous Caching of File Bundles from Tertiary Storage
abstract
In a previous paper, we described a system called STAGS (Storage Access Coordination System) for High Energy and Physics (HEP) experiments. These experiments generate very large volumes of "event" data at a very high rate. The volumes of data may reach 100's of terabytes/year and therefore they are stored on robotic tape systems that are managed by a mass storage system. The data are stored as files on tapes according to a predetermined order, usually according to the order they are generated. A major bottleneck is the retrieval of subsets of these large datasets during the analysis phase. STAGS is designed to optimize the use of a disk cache, and thus minimize the number of files read from tape. In this paper, we describe an interesting problem of disk staging coordination that goes beyond the one-file-at-a-time requirement. The problem stems from the need to coordinate the simultaneous caching of groups of files that we refer to as "bundles of files". All files from a bundle need to be at the same time in the disk cache in order for the analysis application to proceed. This is a radically different problem from the case where the analysis applications need only one file at a time. In this paper, we describe the method of identifying the file bundles, and the scheduling of bundle caching in such a way that files shared between bundles are not removed from the cache unnecessarily. We describe the methodology and the policies used to determine the order of caching bundles of files, and the order of removing files from the cache when space is needed.
Arie Shoshani, Alex Sim, Luis M. Bernardo, Henrik Nordberg
SSDBM2
1999 Multidimensional Indexing and Query Coordination for Tertiary Storage Management
abstract
In many scientific domains, experimental devices or simulation programs generate large volumes of data. The volumes of data may reach hundreds of terabytes and therefore it is impractical to store them on disk systems. Rather they are stored on robotic tape systems that are managed by some mass storage system (MSS). A major bottleneck in analyzing the simulated/collected data is the retrieval of subsets from the tertiary storage system. We describe the architecture and implementation of a Storage Access Coordination System (STACS) designed to optimize the use of a disk cache, and thus minimize the number of files read from tape. We achieve this by using a specialized index to locate the relevant data on tapes, and by coordinating file caching over multiple queries. We focus on a specific application area, a high energy physics data management and analysis environment. STACS was implemented and is being incorporated in an operational system, scheduled to go online at the end of 1999. We also include the results of various tests that demonstrate the benefits and efficiency gained of using the STACS.
Arie Shoshani, Luis M. Bernardo, Henrik Nordberg, Doron Rotem, Alex Sim
SSDBM5