Jaemoon Lee

dblp:72/1286 · also Jae-Moon Lee · DBLP profile ↗
← Back
6ranked-venue papers in the field
2as first author
6since 2021 · last 2025
0000-0002-9868-9410ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (2 first)Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2025 Machine Learning Techniques for Data Reduction of Climate Applications
Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka
PAKDD (1)3
2025 Foundation Model for Lossy Compression of Spatiotemporal Scientific Data
Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka
PAKDD (6)2
2024 Attention Based Machine Learning Methods for Data Reduction with Guaranteed Error Bounds
abstract
Scientific applications in fields such as high energy physics, computational fluid dynamics, and climate science generate vast amounts of data at high velocities. This exponential growth in data production is surpassing the advancements in computing power, network capabilities, and storage capacities. To address this challenge, data compression or reduction techniques are crucial. These scientific datasets have underlying data structures that consist of structured and block structured multidimensional meshes where each grid point corresponds to a tensor. It is important that data reduction techniques leverage strong spatial and temporal correlations that are ubiquitous in these applications. Additionally, applications such as CFD, process tensors comprising hundred plus species and their attributes at each grid point. Reduction techniques should be able to leverage interrelationships between the elements in each tensor.In this paper, we propose an attention-based hierarchical compression method utilizing a block-wise compression setup. We introduce an attention-based hyper-block autoencoder to capture inter-block correlations, followed by a block-wise encoder to capture block-specific information. A PCA-based post-processing step is employed to guarantee error bounds for each data block. Our method effectively captures both spatiotemporal and inter-variable correlations within and between data blocks. Compared to the state-of-the-art SZ3, our method achieves up to 8× higher compression ratio on the multi-variable S3D dataset. When evaluated on single-variable setups using the E3SM and XGC datasets, our method still achieves up to 3× and 2× higher compression ratio, respectively.
Xiao Li 0048, Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka
IEEE Big Data2
2024 Guaranteeing Error Bounds with Preservation of Derived Quantities in Compressive Autoencoders
abstract
Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QoI). Despite the notable performance of recent learned image/video compression approaches using neural networks, they do not guarantee reconstruction errors and cannot manage QoI. This work introduces the Guaranteed Autoencoder with Preserved QoI (GAEQ), which utilizes the interpretation that neural networks with piecewise linear units (PLUs) can be interpreted as a set of linear operators [1] . Although the operators are instance-specific, many instances share the same operator if they fall into the same region of the tessellation formed by PLUs.
Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka
DCC1
2024 Hybrid Approaches for Data Reduction of Spatiotemporal Scientific Applications
abstract
Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions.
Xiao Li 0048, Qian Gong, Jaemoon Lee, Scott Klasky, Anand Rangarajan 0001, Sanjay Ranka
DCC3
2023 Constrained Autoencoders: Incorporating equality constraints in learned scientific data compression
abstract
In scientific data compression, it is crucial to preserve Quantities of Interest (QoI) derived from the data for accurate post-analysis of scientific applications. In this work, we present Constrained Autoencoders (CAEs) where we impose linear QoI as constraints on neural network activations. We circumvent the difficulty of using standard convex optimization methods on the output predictor in the context of autoencoder-driven compression.
Jaemoon Lee, Anand Rangarajan 0001, Sanjay Ranka
DCC1