Phanwadee Sinthong

dblp:217/4824 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
5since 2021 · last 2023
0009-0006-4423-3860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2023 A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam
ICLR3
2023 TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
abstract
Transformers have gained popularity in time series forecasting for their ability to capture long-sequence interactions. However, their memory and compute-intensive requirements pose a critical bottleneck for long-term forecasting, despite numerous advancements in compute-aware self-attention modules. To address this, we propose TSMixer, a lightweight neural architecture exclusively composed of multi-layer perceptron (MLP) modules. TSMixer is designed for multivariate forecasting and representation learning on patched time series, providing an efficient alternative to Transformers. Our model draws inspiration from the success of MLP-Mixer models in computer vision. We demonstrate the challenges involved in adapting Vision MLP-Mixer for time series and introduce empirically validated components to enhance accuracy. This includes a novel design paradigm of attaching online reconciliation heads to the MLP-Mixer backbone, for explicitly modeling the time-series properties such as hierarchy and channel-correlations. We also propose a Hybrid channel modeling approach to effectively handle noisy channel interactions and generalization across diverse datasets, a common challenge in existing patch channel-mixing methods. Additionally, a simple gated attention mechanism is introduced in the backbone to prioritize important features. By incorporating these lightweight components, we significantly enhance the learning capability of simple MLP structures, outperforming complex Transformer models with minimal computing usage. Moreover, TSMixer's modular design enables compatibility with both supervised and masked self-supervised learning methods, making it a promising building block for time-series Foundation Models. TSMixer outperforms state-of-the-art MLP and Transformer models in forecasting by a considerable margin of 8-60%. It also outperforms the latest strong benchmarks of Patch-Transformer models (by 1-2%) with a significant reduction in memory and runtime (2-3X).
Vijay Ekambaram, Arindam Jati, Phanwadee Sinthong, Jayant Kalagnanam
KDD4
2021 Exploratory Data Analysis with Database-backed Dataframes: A Case Study on Airbnb Data
abstract
Choosing between various scalable dataframe libraries can be an overwhelming task for data scientists but it is critical because each framework deploys a different optimization technique that could affect the overall performance. Comparing each framework on a set of analytical tasks in isolation might not fully represent the unique characteristics of big data analyses. This paper describes a case study of applying PolyFrame, a database-backed dataframe library, on an end-to-end exploratory data analysis involving Airbnb data. PolyFrame is a scalable data analytics library that provides a Pandas-like dataframe interface on top of a variety of database systems. The familiarity of its interface enables data scientists to interact with large collections of data through a scale-independent data analysis experience without needing significant database or distributed systems knowledge. Throughout this case study we also highlight the scalability benefits and limitations of database-backed dataframes via a performance comparison with Pandas dataframes for each of the stages of the analysis.
Phanwadee Sinthong, Michael J. Carey 0001
IEEE BigData1
2021 PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes
abstract
In the last few years, the field of data science has been growing rapidly as various businesses have adopted statistical and machine learning techniques to empower their decision-making and applications. Scaling data analyses to large volumes of data requires the utilization of distributed frameworks. This can lead to serious technical challenges for data analysts and reduce their productivity. AFrame, a data analytics library, is implemented as a layer on top of Apache AsterixDB, addressing these issues by providing the data scientists' familiar interface, Pandas Dataframe, and transparently scaling out the evaluation of analytical operations through a Big Data management system. While AFrame is able to leverage data management facilities (e.g., indexes and query optimization) and allows users to interact with a large volume of data, the initial version only generated SQL++ queries and only operated against AsterixDB. In this work, we describe a new design that retargets AFrame's incremental query formation to other query-based database systems, making it more flexible for deployment against other data management systems with composable query languages.
Phanwadee Sinthong, Michael J. Carey 0001
Proc. VLDB Endow.1
2021 DQDF: Data-Quality-Aware Dataframes
abstract
Data quality assessment is an essential process of any data analysis process including machine learning. The process is time-consuming as it involves multiple independent data quality checks that are performed iteratively at scale on evolving data resulting from exploratory data analysis (EDA). Existing solutions that provide computational optimizations for data quality assessment often separate the data structure from its data quality which then requires efforts from users to explicitly maintain state-like information. They demand a certain level of distributed system knowledge to ensure high-level pipeline optimizations from data analysts who should instead be focusing on analyzing the data. We, therefore, propose data-quality-aware dataframes, a data quality management system embedded as part of a data analyst's familiar data structure, such as a Python dataframe. The framework automatically detects changes in datasets' metadata and exploits the context of each of the quality checks to provide efficient data quality assessment on ever-changing data. We demonstrate in our experiment that our approach can reduce the overall data quality evaluation runtime by 40-80% in both local and distributed setups with less than 10% increase in memory usage.
Phanwadee Sinthong, Dhaval Patel 0002, Nianjun Zhou, Shrey Shrivastava, Arun Iyengar, Anuradha Bhamidipaty
Proc. VLDB Endow.1
2020 Scaling Dnn-Based Video Analysis By Coarse-Grained And Fine-Grained Parallelism
abstract
Deep neural networks have been widely used in video analysis applications such as automatic metadata generation, action recognition, and video summarization. A fundamental module in the pipeline of such DNN-based applications is feature extraction. However, extracting features for videos is a major bottleneck since it is performed on every frame of each video sequentially. In addition, the long training time of these complex networks also hinders their usability. In this work, we identify fine-grained and coarse-grained parallelism techniques to speed up vital components in video analysis applications through inter-frame and intra-video parallelism. We demonstrate these techniques on the feature extraction and summarization modules. We leverage frame-level parallelism in feature extraction and intra-video parallelism to speed up video summarization and implement them in a distributed environment using Hadoop Map-Reduce framework to get a speed up of 2.67X on a 4-node setup. Furthermore, we show in our results that our approach has similar accuracy to the sequential applications.
Phanwadee Sinthong, Kanak Mahadik, Somdeb Sarkhel, Saayan Mitra
ICME1
2019 AFrame: Extending DataFrames for Large-Scale Modern Data Analysis
abstract
Analyzing the increasingly large volumes of data that are available today, possibly including the application of custom machine learning models, requires the utilization of distributed frameworks. This can result in serious productivity issues for “normal” data scientists. This paper introduces AFrame, a new scalable data analysis package powered by a Big Data management system that extends the data scientists' familiar DataFrame operations to efficiently operate on managed data at scale. AFrame is implemented as a layer on top of Apache AsterixDB, transparently scaling out the execution of DataFrame operations and machine learning model invocation through a parallel, shared-nothing big data management system. AFrame incrementally constructs SQL++ queries and leverages AsterixDB's semistructured data management facilities, user-defined function support, and live data ingestion support. In order to evaluate the proposed approach, this paper also introduces an extensible micro-benchmark for use in evaluating DataFrame performance in both single-node and distributed settings via a collection of representative analytic operations. This paper presents the architecture of AFrame, describes the underlying capabilities of AsterixDB that efficiently support modern data analytic operations, and utilizes the proposed benchmark to evaluate and compare the performance and support for largescale data analyses provided by alternative DataFrame libraries.
Phanwadee Sinthong, Michael J. Carey 0001
IEEE BigData1