VLDB 2026 Research / reviewers in the wild / expert
Dongeun Lee 0001
dblp:62/688
· DBLP profile ↗
34ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0003-3306-1566ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 5 since 2021Computer networks · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Early Attack Identification in the WildabstractWhile characterizing network connections in their early stage is vital for providing timely responses against network threats, existing methods become less attractive due to the requirement of complete connection information (thus unable to make timely identification) or packet payload inspection (there-fore limited to unencrypted packets under no privacy regulation). To this end, this paper takes an approach ofpacket stream analysisreferencing statistical information of packet sequences, requiringneitherpacket inspectionnorcomplete connection information. To enable practical packet stream analysis, there exist several challenges, such asout-of-order packet sequencesintroduced by network dynamics andclass imbalancewith a tiny fraction of attack connections. To overcome these challenges, we design two deep sequence models: (i) abidirectional recurrent structuredesigned for greater resilience to out-of-order packet streams, and (ii) apre-training-enabled sequence-to-sequence structuredesigned for creating consistent representations from unbalanced class distributions using self-supervised learning. We evaluate the presented deep sequence models using real and synthetic network data collections for extensive experimentation. The experimental results support the feasibility of the proposed models outperforming baseline deep learning models, yielding up to 94.8% (F1 score) only with the first five packets (k=5) from the Internet traffic collection containing a substantial fraction of network flows experiencing out-of-order delivery. Dongeun Lee 0001, Kookjin Lee, Doowon Kim, Jinpyo Kim, Sangman Lee, Jinoh Kim |
IEEE Trans. Netw. | 2 |
| 2025 | Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuningabstractRecent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we introduce a novel approach to modeling transformer architectures using highly flexible non-autonomous neural ordinary differential equations (ODEs). Our proposed model parameterizes all weights of attention and feed-forward blocks through neural networks, expressing these weights as functions of a continuous layer index. Through spectral analysis of the model's dynamics, we uncover an increase in eigenvalue magnitude that challenges the weight-sharing assumption prevalent in existing theoretical studies. We also leverage the Lyapunov exponent to examine token-level sensitivity, enhancing model interpretability. Our neural ODE transformer demonstrates performance comparable to or better than vanilla transformers across various configurations and datasets, while offering flexible fine-tuning capabilities that can adapt to different architectural constraints. Anh Tong, Thanh Nguyen-Tang, Dongeun Lee 0001, Toan M. Tran, David Hall 0006, Cheongwoong Kang, Jaesik Choi |
ICLR | 3 |
| 2025 | Multilingual Game Programming to Enhance Computational Thinking in CS0
Dongeun Lee 0001, Omar el Ariss, Kaoning Hu, Kibum Kwon |
ITiCSE (1) | 1 |
| 2025 | Fostering Computational Thinking in CS1 through Multilingual Game Development
Dongeun Lee 0001, Omar el Ariss, Kaoning Hu, Kibum Kwon, Jonathan Tapia |
ITiCSE (1) | 1 |
| 2025 | Paralfetch: Fast Application Launch on Personal Computing/Communication DevicesabstractParalfetchspeeds up application launches on personal computing/communication devices, by means of: 1) accurate collection of launch-related disk read requests, 2) pre-scheduling of these requests to improve I/O throughput during prefetching, and 3) overlapping application execution with disk prefetching for hiding disk access time from the execution of the application. We implementedParalfetchunder Linux kernels on a desktop/laptop PC, a Raspberry Pi 3 board, and an Android smartphone. Tests with popular applications show thatParalfetchsignificantly reduces application launch times on flash-based drives and hard disk drives, and it outperformsGSoC Prefetch[18] andFAST[21], which are representative application prefetchers available for Linux-based systems. Junhee Ryu, Dongeun Lee 0001, Kang G. Shin, Kyungtae Kang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Operator-Learning-Inspired Modeling of Neural Ordinary Differential EquationsabstractNeural ordinary differential equations (NODEs), one of the most influential works of the differential equation-based deep learning, are to continuously generalize residual networks and opened a new field. They are currently utilized for various downstream tasks, e.g., image classification, time series classification, image generation, etc. Its key part is how to model the time-derivative of the hidden state, denoted dh(t)/dt. People have habitually used conventional neural network architectures, e.g., fully-connected layers followed by non-linear activations. In this paper, however, we present a neural operator-based method to define the time-derivative term. Neural operators were initially proposed to model the differential operator of partial differential equations (PDEs). Since the time-derivative of NODEs can be understood as a special type of the differential operator, our proposed method, called branched Fourier neural operator (BFNO), makes sense. In our experiments with general downstream tasks, our method significantly outperforms existing methods. Woojin Cho 0001, Seunghyeon Cho, Hyundong Jin, Jinsung Jeon, Kookjin Lee, Sanghyun Hong 0001, Dongeun Lee 0001, Noseong Park |
AAAI | 7 |
| 2024 | PAC-FNO: Parallel-Structured All-Component Fourier Neural Operators for Recognizing Low-Quality ImagesabstractA standard practice in developing image recognition models is to train a model on a specific image resolution and then deploy it. However, in real-world inference, models often encounter images different from the training sets in resolution and/or subject to natural variations such as weather changes, noise types and compression artifacts. While traditional solutions involve training multiple models for different resolutions or input variations, these methods are computationally expensive and thus do not scale in practice. To this end, we propose a novel neural network model, parallel-structured and all-component Fourier neural operator (PAC-FNO), that addresses the problem. Unlike conventional feed-forward neural networks, PAC-FNO operates in the frequency domain, allowing it to handle images of varying resolutions within a single model. We also propose a two-stage algorithm for training PAC-FNO with a minimal modification to the original, downstream model. Moreover, the proposed PAC-FNO is ready to work with existing image recognition models. Extensively evaluating methods with seven image recognition benchmarks, we show that the proposed PAC-FNO improves the performance of existing baseline models on images with various resolutions by up to 77.1% and various types of natural variations in the images at inference. Jinsung Jeon, Hyundong Jin, Sanghyun Hong 0001, Dongeun Lee 0001, Kookjin Lee, Noseong Park |
ICLR | 5 |
| 2024 | Parameterized Physics-informed Neural Networks for Parameterized PDEsabstractComplex physical systems are often described by partial differential equations (PDEs) that depend on parameters such as the Raynolds number in fluid mechanics. In applications such as design optimization or uncertainty quantification, solutions of those PDEs need to be evaluated at numerous points in the parameter space. While physics-informed neural networks (PINNs) have emerged as a new strong competitor as a surrogate, their usage in this scenario remains underexplored due to the inherent need for repetitive and time-consuming training. In this paper, we address this problem by proposing a novel extension, parameterized physics-informed neural networks (P$^2$INNs). P$^2$INNs enable modeling the solutions of parameterized PDEs via explicitly encoding a latent representation of PDE parameters. With the extensive empirical evaluation, we demonstrate that P$^2$INNs outperform the baselines both in accuracy and parameter efficiency on benchmark 1D and 2D parameterized PDEs and are also effective in overcoming the known “failure modes”. Woojin Cho 0001, Minju Jo, Haksoo Lim, Kookjin Lee, Dongeun Lee 0001, Sanghyun Hong 0001, Noseong Park |
ICML | 5 |
| 2023 | Fast Application Launch on Personal Computing/Communication Devices
Junhee Ryu, Dongeun Lee 0001, Kang G. Shin, Kyungtae Kang |
FAST | 2 |
| 2023 | Multiple Programming Languages for Improving Computational Thinking in CS1abstractComputational thinking can be deemed as thinking in algorithmic way, with which one can transpose given problems into computer algorithms. Since computational thinking requires abstract reasoning, it should not depend on particular programming languages. Unfortunately, introductory programming courses (CS1) often give students false impression that their goals are to teach a particular programming language. This study shares the design of new pedagogy for CS1 that removes dependency on a particular language and promotes computational thinking by teaching multiple programming languages simultaneously. Specifically, chosen programming languages range from low-level to high-level to expose students to different levels of abstraction from the details of computer architecture. Initial student survey responses from both trial and control groups show that there are significant improvements for the trial groups. Dongeun Lee 0001, Kaoning Hu, Omar el Ariss, Kibum Kwon |
SIGCSE (2) | 1 |
| 2023 | Climate modeling with neural advection-diffusion equation
Hwangyong Choi, Jeongwhan Choi 0002, Jeehyun Hwang, Kookjin Lee, Dongeun Lee 0001, Noseong Park |
Knowl. Inf. Syst. | 5 |
| 2023 | Automated, Reliable Zero-Day Malware Detection Based on Autoencoding ArchitectureabstractWhile a body of studies has been carried out for malware detection with its significance, they are often limited to known malware patterns due to the reliance on signature-based or supervised learning approaches. The semi-supervised learning approach would be an option for identifying previously unseen patterns (i.e., zero-day detection); however, our preliminary study reveals critical limitations from existing methods, including (i) the profiling-based approach using an autoencoder can provide better detection but is sensitive to the threshold setting, and (ii) one-class (OC) classification does not require a manual threshold discovery but may be limited with low detection rates. In this paper, we present a new detection method incorporating the concept of autoencoding and OC classification, designed to benefit from strong abstraction by neural networks (using an autoencoder) and the removal of the complex threshold selection (using an OC classifier). For this combined architecture, a challenge is concurrent training of the autoencoder and the OC classifier, which may cause an ill-suited learner due to no reference to malware instances. To this end, we introduce a new model selection method that discovers well-optimized models from a variety of combinations. The experimental results performed with public malware datasets (Meraz’18 and Drebin) show the effectiveness of our presented methods with up to 97.1% accuracy, comparable to the supervised learning-based detection. We also examine the impact of evading attacks using adversarial attack tools, the result of which shows resilience to malware variants with over 99% detection rates. Chiho Kim, Sang-Yoon Chang, Jonghyun Kim 0005, Dongeun Lee 0001, Jinoh Kim |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | Deep Sequence Models for Packet Stream Analysis and Early DecisionsabstractThe packet stream analysis is essential for the early identification of attack connections while in progress, enabling timely responses to protect system resources. However, there are several challenges for implementing effective analysis, including out-of-order packet sequences introduced due to network dynamics and class imbalance with a small fraction of attack connections available to characterize. To overcome these challenges, we present two deep sequence models: (i) a bidirectional recurrent structure designed for resilience to out-of-order packets, and (ii) a pre-training-enabled sequence-to-sequence structure designed for better dealing with unbalanced class distributions using self-supervised learning. We evaluate the presented models using a real network dataset created from month-long real traffic traces collected from backbone links with the associated intrusion log. The experimental results support the feasibility of the presented models with up to 94.8% in F1 score with the first five packets (k=5), outperforming baseline deep learning models. Dongeun Lee 0001, Kookjin Lee, Doowon Kim, Sangman Lee, Jinoh Kim |
LCN | 2 |
| 2021 | DPM: A Novel Training Method for Physics-Informed Neural Networks in ExtrapolationabstractWe present a method for learning dynamics of complex physical processes described by time-dependent nonlinear partial differential equations (PDEs). Our particular interest lies in extrapolating solutions in time beyond the range of temporal domain used in training. Our choice for a baseline method is physics-informed neural network (PINN) because the method parameterizes not only the solutions, but also the equations that describe the dynamics of physical processes. We demonstrate that PINN performs poorly on extrapolation tasks in many benchmark problems. To address this, we propose a novel method for better training PINN and demonstrate that our newly enhanced PINNs can accurately extrapolate solutions in time. Our method shows up to 72% smaller errors than state-of-the-art methods in terms of the standard L2-norm metric. Jungeun Kim, Kookjin Lee, Dongeun Lee 0001, Sheo Yon Jin, Noseong Park |
AAAI | 3 |
| 2021 | Zero-day Malware Detection using Threshold-free Autoencoding ArchitectureabstractThe impact of malware attacks has been getting more significant, targeting critical infrastructures as well as commodity computing devices. A body of studies has been carried out for detecting malware with its devastating impacts, but they are often limited to known malware attacks due to the nature of the signature-based and supervised machine learning approaches. The semi-supervised learning approach would be an option for identifying previously unseen types of malware attacks (i.e., zero-day detection); however, our preliminary studies suggest two limitations in this avenue: (1) one class (OC) classifiers can be limited with relatively low detection rates, and (2) the profiling-based approach (using an autoencoder) may yield better detection performance but under the assumption of the "ideal" threshold setting. In this paper, we tackle these challenges and present a new detection method, which combines the concepts of autoencoding and OC classification, to benefit from strong abstractions by neural networks (using an autoencoder) but to remove the necessity of the complex threshold selection (using an OC classifier). Our extensive experimental results with a recent malware dataset (Meras’18) show the effectiveness of our method with up to 96% accuracy for zero-day malware detection, which is comparable to the supervised learning-based detection (limited to known types of malware). The proposed method also shows the resilience to adversarial attacks, yielding better performance for identifying synthetic samples generated to evade the detection process than supervised learning algorithms. Chiho Kim, Sang-Yoon Chang, Jonghyun Kim 0005, Dongeun Lee 0001, Jinoh Kim |
IEEE BigData | 4 |
| 2021 | Climate Modeling with Neural Diffusion EquationsabstractOwing to the remarkable development of deep learning technology, there have been a series of efforts to build deep learning-based climate models. Whereas most of them utilize recurrent neural networks and/or graph neural networks, we design a novel climate model based on the two concepts, the neural ordinary differential equation (NODE) and the diffusion equation. Many physical processes involving a Brownian motion of particles can be described by the diffusion equation and as a result, it is widely used for modeling climate. On the other hand, neural ordinary differential equations (NODEs) are to learn a latent governing equation of ODE from data. In our presented method, we combine them into a single framework and propose a concept, called neural diffusion equation (NDE). Our NDE, equipped with the diffusion equation and one more additional neural network to model inherent uncertainty, can learn an appropriate latent governing equation that best describes a given climate dataset. In our experiments with two real-world and one synthetic datasets and eleven baselines, our method consistently outperforms existing baselines by non-trivial margins. Jeehyun Hwang, Jeongwhan Choi 0002, Hwangyong Choi, Kookjin Lee, Dongeun Lee 0001, Noseong Park |
ICDM | 5 |
| 2021 | A Novel Method to Solve Neural Knapsack Problemsabstract0-1 knapsack is of fundamental importance across many fields. In this paper, we present a game-theoretic method to solve 0-1 knapsack problems (KPs) where the number of items (products) is large and the values of items are not predetermined but decided by an external value assignment function (e.g., a neural network in our case) during the optimization process. While existing papers are interested in predicting solutions with neural networks for classical KPs whose objective functions are mostly linear functions, we are interested in solving KPs whose objective functions are neural networks. In other words, we choose a subset of items that maximize the sum of the values predicted by neural networks. Its key challenge is how to optimize the neural network-based non-linear KP objective with a budget constraint. Our solution is inspired by game-theoretic approaches in deep learning, e.g., generative adversarial networks. After formally defining our two-player game, we develop an adaptive gradient ascent method to solve it. In our experiments, our method successfully solves two neural network-based non-linear KPs and conventional linear KPs with 1 million items. Duanshun Li, Jing Liu 0024, Dongeun Lee 0001, Ali Seyedmazloom, Giridhar Kaushik, Kookjin Lee, Noseong Park |
ICML | 3 |
| 2021 | Large-Scale Flight Frequency Optimization with Global Convergence in the US Domestic Air Passenger MarketsabstractThe US domestic air passenger transportation is one of the largest markets worldwide.Optimally allocating flights to the US domestic airways (i.e., air routes) is essential in maximizing the revenue of airlines and many research works have been proposed to improve their market shares/profits.Most proposed methods, however, suffer from a lack of scalability; even state-of-the-art methods demonstrate their performance with only tens of routes.To address this shortcoming, we propose a novel unified framework to integrate the market share prediction model and the frequency optimization module, which significantly improves the scalability of the entire framework.By design, our proposed prediction model is concave w.r.t.flight frequency and its gradients are Lipschitz continuous.Exploiting these two properties allows us to use an alternating direction method of multipliers (ADMM)-based optimization technique, which quickly solves a large-scale frequency optimization problem with guaranteed global convergence.Our proposed method is able to solve a problem whose search space size is O(n 700 ) (vs.O(n 30 ) in existing works). Jinsung Jeon, Dongeun Lee 0001, Seunghyun Hwang, Soyoung Kang, Noseong Park, Duanshun Li, Kookjin Lee, Jing Liu 0024 |
SDM | 2 |
| 2020 | A Learning-based Data Augmentation for Network Anomaly DetectionabstractWhile machine learning technologies have been remarkably advanced over the past several years, one of the fundamental requirements for the success of learning-based approaches would be the availability of high-quality data that thoroughly represent individual classes in a problem space. Unfortunately, it is not uncommon to observe a significant degree of class imbalance with only a few instances for minority classes in many datasets, including network traffic traces highly skewed toward a large number of normal connections while very small in quantity for attack instances. A well-known approach to addressing the class imbalance problem is data augmentation that generates synthetic instances belonging to minority classes. However, traditional statistical techniques may be limited since the extended data through statistical sampling should have the same density as original data instances with a minor degree of variation. This paper takes a learning-based approach to data augmentation to enable effective network anomaly detection. One of the critical challenges for the learning-based approach is the mode collapse problem resulting in a limited diversity of samples, which was also observed from our preliminary experimental result. To this end, we present a novel "Divide-Augment-Combine" (DAC) strategy, which groups the instances based on their characteristics and augments data on a group basis to represent a subset independently using a generative adversarial model. Our experimental results conducted with two recently collected public network datasets (UNSW-NB15 and IDS-2017) show that the proposed technique enhances performances up to 21.5% for identifying network anomalies. Mohammad Al Olaimat, Dongeun Lee 0001, Youngsoo Kim 0002, Jonghyun Kim 0005, Jinoh Kim |
ICCCN | 2 |
| 2020 | Poster: Prototype of Configurable Redfish Query Proxy ModuleabstractRedfish is a next-generation API standard for the management of data center infrastructures. This rich API can flexibly obtain data using a query string from the client side. However, this feature is optional and not fully supported by many services. We implemented a prototype Redfish query processing module on Nginx, a well-known open source web server. The Redfish query processing module can run with a proxy module and work with any server-side or client-side applications. Additionally, our prototype implementation can be configured to properly utilize queries, which are supported on a backend server, and improve performance. Our implementation was evaluated on an OpenBMC server and a mockup server and showed potential for performance improvement. Chanyoung Park 0004, Yoonsue Joe, Myounghwan Yoo, Dongeun Lee 0001, Kyungtae Kang |
ICNP | 4 |
| 2018 | Dynamic Online Performance Optimization in Streaming Data CompressionabstractCompression is essential to high bandwidth applications such as scientific simulations and sensing applications to reduce resource burden such as storage, network transmission, and more recently I/O. Existing lossy compression methods attempt to minimize the Euclidean distance between original data and reconstructed data, which significantly limits either compression performance or reconstruction quality since original and reconstructed data sequences should be aligned. Substituting the Euclidean distance for a statistical similarity maximizes the compression performance while retaining essential data features. By implementing this methodology, IDEALEM has recently demonstrated compression ratios far exceeding 100:1, better than best-known compression methods, while preserving reconstruction quality. This work proposes an online algorithm for streaming data compression which takes account of generally concave trend of compression ratio curve, and optimizes key operation parameters. We demonstrate that the proposed algorithm successfully adapts one of the key parameters in IDEALEM to the optimal value and yields near maximum compression ratios for time series data. J. Kade Gibson, Dongeun Lee 0001, Jaesik Choi, Alex Sim |
IEEE BigData | 2 |
| 2018 | ClusterFetch: A Lightweight Prefetcher for Intensive Disk ReadsabstractBy overlapping disk accesses with computation-intensive operations, prefetching can reduce delays in launching an application and in loading significant amounts of data while the application is running. The key to effective prefetching is making the tradeoff between the mining accuracy of selecting relevant blocks, and the time to decide those blocks. To address this problem, we propose a new prefetcher called ClusterFetch. In its learning mode, ClusterFetch detects periods of intensive disk accesses by monitoring the speed at which read requests are queued; it re-organizes these reads and locates the file opened by the application just before each such period. During subsequent runs of the same application, ClusterFetch prefetches the data associated with the opening of a “trigger” file. Our experimental results show that ClusterFetch implemented in Linux can reduce the application launch time by up to 41.3 percent and the loading time by up to 38.2 percent, while taking up less than 200 KB of main memory. Junhee Ryu, Dongeun Lee 0001, Kang G. Shin, Kyungtae Kang |
IEEE Trans. Computers | 2 |
| 2017 | Expanding Statistical Similarity Based Data Reduction to Capture Diverse PatternsabstractWe propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall. Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu |
DCC | 1 |
| 2017 | Improving Statistical Similarity Based Data Reduction for Non-Stationary DataabstractWe propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall. Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu |
SSDBM | 1 |
| 2016 | Novel Data Reduction Based on Statistical SimilarityabstractApplications such as scientific simulations and power grid monitoring are generating so much data quickly that compression is essential to reduce storage requirement or transmission capacity. To achieve better compression, one is often willing to discard some repeated information. These lossy compression methods are primarily designed to minimize the Euclidean distance between the original data and the compressed data. But this measure of distance severely limits either reconstruction quality or compression performance. We propose a new class of compression method by redefining the distance measure with a statistical concept known as exchangeability. This approach reduces the storage requirement and captures essential features, while reducing the storage requirement. In this paper, we report our design and implementation of such a compression method named IDEALEM. To demonstrate its effectiveness, we apply it on a set of power grid monitoring data, and show that it can reduce the volume of data much more than the best known compression method while maintaining the quality of the compressed data. In these tests, IDEALEM captures extraordinary events in the data, while its compression ratios can far exceed 100. Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu |
SSDBM | 1 |
| 2016 | Improving Imprecise Compressive Sensing Models
Dongeun Lee 0001, Rafael Lima, Jaesik Choi |
UAI | 1 |
| 2015 | Learning Compressive Sensing Models for Big Spatio-Temporal DataabstractSensing devices including mobile phones and biomedical sensors generate massive amounts of spatio-temporal data. Compressive sensing (CS) can significantly reduce energy and resource consumption by shifting the complexity burden of encoding process to the decoder. CS reconstructs the compressed signals exactly with overwhelming probability when incoming data can be sparsely represented with a fixed number of components, which is one of drawbacks of CS frameworks because a real-world signal in general cannot be represented with the fixed number of components. We present the first CS framework that handles signals without the fixed sparsity assumption by incorporating the distribution of the number of principal components included in the signal recovery, which we show is naturally represented by the gamma distribution. This allows an analytic derivation of total error in our spatio-temporal Low Complexity Sampling (LCS). We show that LCS requires shorter compressed signals than existing CS frameworks to bound the same amount of error. Experiments with real-world sensor data also demonstrate that LCS outperforms existing CS frameworks. Dongeun Lee 0001, Jaesik Choi |
SDM | 1 |
| 2015 | ClusterFetch: A Lightweight Prefetcher for General WorkloadsabstractApplication loading times can be reduced by prefetching disk blocks into the buffer cache. Existing prefetching schemes for general workloads suffer from significant overheads and low accuracy. ClusterFetch is a lightweight prefetcher that identifies continuous sequences of I/O requests and identifies the files that trigger them. The next time that the same files are opened, the corresponding disk blocks are prefetched. In experiments, ClusterFetch reduced the launch time, by which we refer to the latency that first occurs when a program runs, by 15.2 to 30.9%, and loading times, meaning the delays that are incurred while additional data is loaded from the disk during program execution, by 15.9%. Haksu Jeong, Junhee Ryu, Dongeun Lee 0001, Jaemyoun Lee, Heonshik Shin, Kyungtae Kang |
ICPE | 3 |
| 2014 | Low complexity sensing for big spatio-temporal dataabstractMany large scale sensor networks produce tremendous data, typically as massive spatio-temporal data streams. We present a Low Complexity Sensing framework that, coupled with novel compressive sensing techniques, enables to reduce computational and communication overheads significantly without much compromising the accuracy of sensor readings. More specifically, our sensing framework randomly samples time-series data in the temporal dimension first, then in the spatial dimension. Under some mild conditions, our sensing framework holds the same theoretical bound of reconstruction error, but is much simpler and easier to implement than existing compressive sensing frameworks. In experiments with real world environmental data sets, we demonstrate that the proposed framework outperforms two existing compressive sensing frameworks designed for spatio-temporal data. Dongeun Lee 0001, Jaesik Choi |
IEEE BigData | 1 |
| 2014 | File-system-level flash caching for improving application launch time on logical hybrid disksabstractApplication launch time is an important performance metric to user experience in desktop environment. The launch time mostly depends on the performance of secondary storage. There is a cost-performance trade-off in using hard disk drive (HDD) or solid-state drive (SSD). Thus, application launch times can be reduced by utilizing SSDs as caches for slow HDDs. We propose a new SSD caching scheme which migrates data blocks from HDDs to SSDs. Since our scheme operates entirely in the file system level and does not require an extra layer for mapping SSD-cached data, which is essential in most other schemes, our scheme does not incur mapping overheads that cause significant burdens on main memory, CPU, and SSD cache itself. Experimental results demonstrate our scheme yields 56% of performance gain in application launch. Changhee Han 0002, Junhee Ryu, Dongeun Lee 0001, Jaemyoun Lee, Kyungtae Kang, Heonshik Shin |
IPCCC | 3 |
| 2012 | Reliable Wildfire Monitoring with Sparsely Deployed Wireless Sensor NetworksabstractThis paper proposes a reliable wildfire monitoring system based on a wireless sensor network (WSN) sparsely deployed in adverse conditions. The physical environment under consideration is characterized by asymmetric, irregular, and unreliable wireless links, inadequate Fresnel zone clearance, and routing problems, to name a few. We use reliable communication schemes on a fault-tolerant network topology, where sensory data are guaranteed to reach the base station with organized data storage and real-time visualization. Our approach has been validated experimentally for the case of peat-forest wildfire in southern Borneo where the fire breaks out frequently. Ikjune Yoon, Dong Kun Noh, Dongeun Lee 0001, Rony Teguh, Toshihisa Honma, Heonshik Shin |
AINA | 3 |
| 2010 | Low-complexity aggregation of collected images with correlated fields of view in wireless video sensor networksabstractWireless video sensor networks (WVSNs) require video data from sensor nodes to be delivered efficiently. Cameras in adjacent video nodes tend to have correlated fields of view (FoVs) or overlapping part when a sufficient number of video sensors are deployed. This paper proposes a data aggregation technique for WVSNs to remove the spatial redundancy, thereby reducing energy consumption and response time. Our approach exploits the correlation between discrete cosine transform (DCT) coefficients pairs of two intra-coded images from cameras with overlapping FoVs. Experiments show that an intermediate node en route to the base station achieves bit-rate savings up to 18.9%. This scheme is less complicated than other video and image coding techniques that exploit correlated FoV, allowing resource-constrained video sensors to operate more reliably and longer. Dongeun Lee 0001, Jonghun Lee, Yonghee Lee, Heejung Lee, Heonshik Shin |
ISCC | 1 |
| 2008 | Luminance scalable coding using H.264/AVC SVC extensions for mobile video applicationsabstractWe propose a luminance scalable video coding scheme (LSC) that uses the scalable video coding (SVC) extension of H.264/AVC to reduce the power consumption of mobile devices. The insufficient battery life of mobile hand-held devices is the main obstacle for mobile video applications. In particular, liquid crystal displays consume a large portion of the system power. Building on previous power reduction schemes, such as dynamic backlight luminance scaling, we exploit SVC which is used to adapt a video stream to the preferences of users, different mobile capabilities, and varying network conditions. Our video coding scheme can reduce the power consumed by an LCD display by more than 20%, even though it has a lower overhead than previous schemes and does not require any modification of the decoders. Heejung Lee, Dongeun Lee 0001, Yonghee Lee, Heonshik Shin |
ICME | 2 |
| 2007 | Low-Latency Routing for Energy-Harvesting Sensor Networks
Hyuntaek Kwon, Dong Kun Noh, Junu Kim, Dongeun Lee 0001, Heonshik Shin |
UIC | 5 |