Euiseong Seo

dblp:59/4753 · DBLP profile ↗
← Back
41ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-2103-8019ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning Analytics
abstract
Real-time machine learning (ML) analytics models deployed on edge servers often experience degraded inference accuracy due to data drift. Continuous learning mitigates this by periodically retraining models using newly collected data. However, retraining incurs significant computational overhead, increasing energy consumption and carbon footprint several-fold. Existing approaches for reducing carbon footprint predominantly focus on general workload scheduling strategies, such as shifting jobs to periods or regions with lower carbon intensity. However, they neglect continuous-learning-specific parameters like data labeling model, retraining threshold, and retraining hyperparameters. Consequently, these approaches miss opportunities to further reduce the carbon footprint and enhance inference accuracy in continuous learning systems under dynamic data drift. In this paper, we propose a novel Carbon-footprint-aware Continuous Learning (CCL) scheme that minimizes carbon emissions during model retraining without sacrificing inference accuracy. Distinct from prior workload scheduling approaches, CCL adaptively adjusts labeling model, retraining threshold, and retraining hyperparameters based on predictive models that estimate data drift severity and carbon intensity dynamics. Our adaptive real-time optimization approach consistently achieves a near-optimal balance between accuracy and carbon footprint under dynamic conditions. Experimental results demonstrate that CCL reduces operational carbon footprint by up to 68.5% compared to state-of-the-art carbon-agnostic methods, with negligible accuracy degradation.
Gwanjong Park, Dongho Ha, Myeongjae Jeon, Euiseong Seo
EuroSys5
2026 Near-Data Compaction for LSM Tree on Rack-Scale Disaggregated Storage
abstract
In LSM trees, background compaction tasks contend with foreground queries for CPU cycles, cache space, and also SAN bandwidth, if deployed in a disaggregated storage architecture. This study proposesNear-Data Compaction(NDC), which executes compaction on the storage node to utilize its underutilized computing resources. However, enabling NDC introduces several challenges. First, it has to support concurrent file access from both compute and storage nodes. In addition, it has to decide which compaction tasks to be executed on the storage node since the computing resources of a storage node are not unlimited. This study presentsTetherDB, an LSM tree for disaggregated storage architecture that addresses these challenges through lightweight dual-node coordination and selective NDC admission policies. Our evaluation demonstrates that TetherDB improves throughput by up to 2.1× compared to RocksDB in write-heavy workloads.
Sungho Moon, Daegyu Han, Hera Koo, Sangeun Chae, Duck-Ho Bae, Euiseong Seo, Beomseok Nam
IEEE Trans. Parallel Distributed Syst.6
2025 AWUPF Rediscovered: Atomic Writes to Unleash Pivotal Fault-Tolerance in SSDs
Jiyune Jeon, Sam H. Noh, Euiseong Seo
FAST4
2024 We Ain't Afraid of No File Fragmentation: Causes and Prevention of Its Performance Impact on Modern Flash SSDs
Yuhun Jun, Shin-Hyun Park, Jeong-Uk Kang, Sang-Hoon Kim, Euiseong Seo
FAST5
2024 Cloud Reamer: Enabling Inference Services in Training Clusters
abstract
CPU cores in GPU servers are often underutilized during DNN training. Co-locating CPU-based inference tasks with DNN training offers an opportunity to utilize these idle CPU cycles. However, three technical challenges must be addressed: avoiding disruption to training workloads, meeting different performance requirements for online and offline inference, and swiftly adjusting inference configurations based on available resources. This paper proposes Cloud Reamer, a scheme to colocate training and inference tasks on GPU servers, optimizing unused CPU cycles without disrupting training. Cloud Reamer prioritizes training tasks to minimize interference. For online inference, it allocates cores to ensure predictable performance, while for offline inference, it uses all available cores to maximize throughput. Cloud Reamer enhances online and offline inference performance by dynamically adjusting configurations based on surplus CPU resources. Evaluations show that Cloud Reamer improves inference throughput with minimal impact on training, maintaining training interference below $\mathbf{3. 2 \%}$. It meets latency requirements for 46% more requests for online inference and achieves a 61x throughput increase for offline inference compared to conventional methods.
Gwanjong Park, Junyeol Yu, Euiseong Seo
MASCOTS4
2024 An Adaptive Zone-Grouping Scheme Enabling General-Purpose File Systems on ZNS SSDs
abstract
Zoned namespace solid state drives (ZNS SSDs) provide superior capacity and life span compared to conventional namespace (CNS) SSDs at the same cost. Especially, a small-zone ZNS SSD can achieve a low write amplification factor (WAF) and performance isolation among I/O streams by providing the host with finer-grained control of the SSD's flash chips than the large-zone counterpart. To maximize these advantages, multiple zones in a small-zone SSD should be grouped and operated simultaneously. A ZNS-aware file system encapsulates the ZNS interface and transforms a ZNS SSD into a general-purpose storage to the host layer. However, the current ZNS-aware file systems were designed for the large-zone ZNS SSDs, and thus lack the zone group management feature. This paper proposes a dynamic zone group management scheme for ZNS-aware file systems that dynamically adjusts the number of zone groups in accordance with the workload characteristics and the size of each zone group to achieve high performance at a low WAF. The proposed scheme was implemented in the F2FS. Our evaluation with FIO, Filebench and db_bench showed that, across all cases, the proposed scheme improved the performance by up to 71.2% while reducing the WAF by 29.8% in comparison to the static zone grouping scheme.
Jungyun Choi, Yuhun Jun, Jinkyu Jeong, Euiseong Seo
SYSTOR5
2024 A Secure, Fast, and Resource-Efficient Serverless Platform with Function REWIND
Jaehyun Song, Bumsuk Kim, Minwoo Kwak, Byoungyoung Lee, Euiseong Seo, Jinkyu Jeong
USENIX ATC5
2023 Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving
abstract
The proportion of machine learning (ML) inference in modern cloud workloads is rapidly increasing, and graphic processing units (GPUs) are the most preferred computational accelerators for it. The massively parallel computing capability of GPUs is well-suited to the inference workloads but consumes more power than conventional CPUs. Therefore, GPU servers contribute significantly to the total power consumption of a data center. However, despite their heavy power consumption, GPU power management in cloud-scale has not yet been actively researched. In this paper, we reveal three findings about energy efficiency of ML inference clusters in the cloud. ❶ GPUs of different architectures have comparative advantages in energy efficiency to each other for a set of ML models. ❷ The energy efficiency of a GPU set may significantly vary depending on the number of active GPUs and their clock frequencies even when producing the same level of throughput. ❸ The service level objective(SLO)-blind dynamic voltage and frequency scaling (DVFS) driver of commercial GPUs maintain an immoderately high clock frequency. Based on these implications, we propose a hierarchical GPU resource management approach for cloud-scale inference services. The proposed approach consists of energy-aware cluster allocation, intra-cluster node scaling, intra-node GPU scaling and GPU clock scaling schemes considering the inference service architecture hierarchy. We evaluated our approach with its prototype implementation and cloud-scale simulation. The evaluation with real-world traces showed that the proposed schemes can save up to 28.3% of the cloud-scale energy consumption when serving five ML models with 105 servers having three different kinds of GPUs.
Junyeol Yu, Euiseong Seo
HPCA3
2023 Revitalizing Buffered I/O: Optimizing Page Reclaim and I/O Throttling
abstract
Buffered I/O is commonly used as a default mechanism in most file systems because it provides high performance by keeping recently accessed data in memory as page caches. We discovered that the long-standing code within the memory management of the Linux kernel performs unnecessary and indiscriminate operations during the page reclaiming and writing procedures, leading to a degradation in the performance of buffered I/O. Current memory management performs aging tasks, moving pages from the active state list to the inactive state list, even during page reclamation for urgent memory acquisition. It also unnecessarily holds the lock used for modifying the least recently used (LRU) list. These non-essential tasks results in performance degradation when buffered I/O is allocating pages and adding them into the LRU list. Furthermore, in write-intensive situations, the performance of writes is degraded by blindly forcing the write thread to sleep in order to maintain dirty pages at a pre-determined ratio in total memory. In this paper, we improve the memory management system by eliminating these unnecessarily or blindly performed aging, locks, and sleep tasks, thereby enhancing the performance of buffered I/O. We implemented our proposed approaches in the Linux kernel and evaluated these impacts on the performance of buffered I/O using FIO, FileBench, and YCSB on RocksDB. Our approach reduced the tail latency of random reads and improved the throughput of random writes in most workloads.
Chanu Yu, Euiseong Seo
ICCD3
2023 Energy-Harvesting-Aware Adaptive Inference of Deep Neural Networks in Embedded Systems
abstract
In energy harvesting IoT and sensor devices, energy influx is continuously changing and difficult to predict. Recently, the use of deep neural networks (DNNs), which consumes a large amount of energy, has increased in such devices. If a lightweight DNN model is used anticipating low energy influx, it may not achieve satisfactory inference accuracy in the ample energy flow condition. Conversely, using highly accurate sophisticated models may result in frequent inference failures due to energy depletion in situations with low energy influx. In this paper, for energy harvesting embedded systems that periodically perform DNN inference on sensor inputs, we propose an energy-harvesting-aware adaptive inference scheme to maximize inference accuracy while minimizing inference failures due to energy depletion in the long term. The model selector in the proposed scheme, which is a reinforcement learning (RL) agent, selects a DNN model from a model pool in consideration of the energy harvesting state and the accuracy and energy requirement of each model in the model pool. We implemented the proposed scheme on a microcontroller system and evaluated it with six different DNN applications with various types of input data. In the energy harvesting simulation with the real-world solar power traces, our approach, on average across the six workloads, was able to achieve a 65.62% reduction in inference failure rate with only a 6.08% increase in average error rate compared to the base DNN models.
Gwanjong Park, Euiseong Seo
ISLPED3
2023 A Close Look at Shared Resource Consumption in NoSQL Databases for Accurate Accounting
abstract
A NoSQL database plays an essential role in all parts of today’s multi-tenant large-scale IT services. Accurate information about per-tenant resource accounting is invaluable for optimal service management. However, there is a lack of study on the characteristics of resource consumption by multi-tenants in NoSQL database services, particularly for those resources consumed as part of activities asynchronously triggered by requests accumulated over time from multiple tenants. According to our investigation, this shared resource usage takes up a significant portion of total resource consumption in NoSQL databases.We assert that an accurate understanding of the shared resource consumption pattern is required to design correct and effective techniques for resource management. To this end, we conduct a detailed investigation of the shared resource consumption in popular NoSQL databases. Our focuses are to find out what portion of the overall resource consumption occurs in a shared manner, what type of operations cause them, and what effect the workload intensity has on the amount of shared resource consumption. We have developed a set of techniques for monitoring, recording, and analyzing the resource consumption data toward answering these questions. Our investigation revealed that the shared resource consumption can be as large as 30% of the total resource consumption, implying that resource accounting based only on directly causal events, may lead to underestimation of the true amount. We believe our study highlights the importance of shared resource accounting and provides crucial insights for building accurate resource accounting techniques.
Jaeryun Lee, Byung-Chul Tak, Euiseong Seo
NOMS3
2023 DaCapo: An On-Device Learning Scheme for Memory-Constrained Embedded Systems
abstract
The use of deep neural network (DNN) applications in microcontroller unit (MCU) embedded systems is getting popular. However, the DNN models in such systems frequently suffer from accuracy loss due to the dataset shift problem. On-device learning resolves this problem by updating the model parameters on-site with the real-world data, thus localizing the model to its surroundings. However, the backpropagation step during on-device learning requires the output of every layer computed during the forward pass to be stored in memory. This is usually infeasible in MCU devices as they are equipped only with a few KBs of SRAM. Given their energy limitation and the timeliness requirements, using flash memory to store the output of every layer is not practical either. Although there have been proposed a few research results to enable on-device learning under stringent memory conditions, they require the modification of the target models or the use of non-conventional gradient computation strategies. This paper proposes DaCapo, a backpropagation scheme that enables on-device learning in memory-constrained embedded systems. DaCapo stores only the output of certain layers, known as checkpoints, in SRAM, and discards the others. The discarded outputs are recomputed during backpropagation from the nearest checkpoint in front of them. In order to minimize the recomputation occurrences, DaCapo optimally plans the checkpoints to be stored in the SRAM area at a particular phase of the backpropagation and thus replaces the checkpoints stored in memory as the backpropagation progresses. We implemented the proposed scheme in an STM32F429ZI board and evaluated it with five representative DNN models. Our evaluation showed that DaCapo improved backpropagation time by up to 22% and saved energy consumption by up to 28% in comparison to AIfES, a machine learning platform optimized for MCU devices. In addition, our proposed approach enabled the training of MobileNet, which the MCU device had been previously unable to train.
Gwanjong Park, Euiseong Seo
ACM Trans. Embed. Comput. Syst.3
2022 Analysis and Mitigation of Data Sanitization Overhead in DAX File Systems
abstract
A direct access (DAX) file system maximizes the benefit of persistent memory(PM)’s low latency through removing the page cache layer from the file system access paths. However, this paper reveals that data block allocation of the DAX file systems in common is significantly slower than that of conventional file systems because the DAX file systems require the zero-out operation for the newly allocated blocks to prevent the leakage of old data previously stored in the allocated data blocks. The retarded block allocation significantly affects the file write performance. In addition to this revelation, this paper proposes an off-critical-path data block sanitization scheme tailored for DAX file systems. The proposed scheme detaches the zero-out operation from the latency-critical I/O path and performs that of released data blocks in the background. The proposed scheme’s design principle is universally applicable to most DAX file systems. For evaluation, we implemented our approach in Ext4-DAX and XFS-DAX. Our evaluation showed that the proposed scheme reduces the append write latency by 36.8%, and improved the performance of FileBench’s fileserver workload by 30.4%, YCSB’s workload A on RocksDB by 3.3%, and the Redis-benchmark by 7.4% on average, respectively.
Younghoon Lim, Euiseong Seo
ICCD4
2022 NoSQL Database Performance Diagnosis through System Call-level Introspection
abstract
Since its emergence, NoSQL databases have firmly established themselves as an indispensable software component of modern cloud-native applications. However, it also becomes increasingly challenging to perform critical management tasks such as troubleshooting unexpected performance problems. This is due to the ever-increasing diversity and specialization of NoSQL databases that make it difficult to observe the internal activities. To address these challenges, we have designed and built a technique for introspecting NoSQL databases. Our technique traces system call sequences of key operations under controlled workloads and filters scaling patterns from constant components. Novel algorithms are developed to uncover repeating patterns of system calls from massive amounts of traces and filter out background noise with high efficiency. The evaluation shows that our technique can greatly enhance the visibility into the NoSQL databases enabling us to diagnose performance problems or gain insights into internal activities.
Changho Seo, Yunchang Chae, Jaeryun Lee, Euiseong Seo, Byung-Chul Tak
NOMS4
2022 Dedup-for-speed: storing duplications in fast programming mode for enhanced read performance
abstract
Storage deduplication improves write latency, increases available space, and reduces the wear of storage media by eliminating redundant writes. The flash translation layer (FTL) of a flash solid state disk (SSD) easily enables deduplication in an SSD through the simple mapping of duplicated logical pages to the same physical page. Therefore, a few deduplicating FTLs have been proposed. However, the deduplication of partially duplicated files breaks the sequentiality of data storage at the flash page level, resulting in a significant degradation of the read performance. Although increasing available storage space, reducing flash write, and extending lifespan are barely perceptible to users, extended read latency is critical to user-perceived performance. In this paper, we propose a novel deduplication FTL called Dedup-for-Speed (DFS). The DFS FTL trades surplus capacity gained through inline deduplication for improved read performance by storing duplicated pages in fast flash modes, such as the pseudo-single-level-cell(pSLC) mode. The flash mode of a page is determined by its degree of deduplications. Duplicate pages are migrated to fast flash blocks during idle intervals to minimize interference with host-issued operations. Contrary to conventional deduplication schemes, DFS improves the read performance while maintaining the aforementioned benefits of deduplication. Our evaluation of six real-world traces showed that DFS improved the read latency by 16% on average and up to 34%. It also enhanced the write latency by 64% on average and up to 82%.
Jaeyong Bae, Jaehyung Park, Yuhun Jun, Euiseong Seo
SYSTOR4
2021 Z-Journal: Scalable Per-Core Journaling
Cassiano Campes, Joo Young Hwang, Jinkyu Jeong, Euiseong Seo
USENIX ATC5
2021 Idempotence-Based Preemptive GPU Kernel Scheduling for Embedded Systems
abstract
Mission-critical embedded systems simultaneously run multiple graphics-processing-unit (GPU) computing tasks with different criticality and timeliness requirements. Considerable research effort has been dedicated to supporting the preemptive priority scheduling of GPU kernels. However, hardware-supported preemption leads to lengthy scheduling delays and complicated designs, and most software approaches depend on the voluntary yielding of GPU resources from restructured kernels. We propose a preemptive GPU kernel scheduling scheme that harnesses the idempotence property of kernels. The proposed scheme distinguishes idempotent kernels through static source code analysis. If a kernel is not idempotent, then GPU kernels are transactionized at the operating system (OS) level. Both idempotent and transactionized kernels can be aborted at any point during their execution and rolled back to their initial state for reexecution. Therefore, low-priority kernel instances can be preempted for high-priority kernel instances and reexecuted after the GPU becomes available again. Our evaluation using the Rodinia benchmark suite showed that the proposed approach limits the preemption delay to 18 ms in the 99.9th percentile, with an average delay in execution time of less than 10 percent for high-priority tasks under a heavy load in most cases.
Cheolgi Kim, Hwansoo Han, Euiseong Seo
IEEE Trans. Computers5
2019 Compiler-Assisted GPU Thread Throttling for Reduced Cache Contention
abstract
Modern GPUs concurrently deploy thousands of threads to maximize thread level parallelism (TLP) for performance. For some applications, however, maximized TLP leads to significant performance degradation, as many concurrent threads compete for the limited amount of the data cache. In this paper, we propose a compiler-assisted thread throttling scheme, which limits the number of active thread groups to reduce cache contention and consequently improve the performance. A few dynamic thread throttling schemes have been proposed to alleviate cache contention by monitoring the cache behavior, but they often fail to provide timely responses to the dynamic changes in the cache behavior, as they adjust the parallelism afterwards in response to the monitored behavior. Our thread throttling scheme relies on compile-time adjustment of active thread groups to fit their memory footprints to the L1D capacity. We evaluated the proposed scheme with GPU programs that suffer from cache contention. Our approach improved the performance of original programs by 42.96% on average, and this is 8.97% performance boost in comparison to the static thread throttling schemes.
Sungin Hong, Euiseong Seo, Hwansoo Han
ICPP4
2018 A GPU Kernel Transactionization Scheme for Preemptive Priority Scheduling
abstract
Mission-critical embedded systems are becoming heavily dependent on graphics processing unit (GPU) computing. In general, such systems simultaneously run multiple tasks with different criticality and timeliness requirements. Consequently, many research efforts have been made in both hardware and software aspects to support the preemptive priority scheduling of GPU kernels. However, hardware-supported preemption leads to lengthy scheduling delays and complicated design, and most of the software approaches depend on the voluntary yield of GPU resources from restructured kernels. Exploiting the shared physical memory of CPUs and GPUs in heterogeneous system architecture (HSA), we propose an approach to transactionize GPU kernels at the operating system (OS) level. A transactionized GPU kernel can be aborted at any point during its execution and rolled back to its initial state for re-execution. By transactionizing GPU kernels, it is possible to forcibly evict low-priority kernels and immediately schedule high-priority kernels. The preempted low-priority kernel instances can be re-executed after a GPU becomes available. The proposed approach was implemented in a Samsung Exynos 5422 system on chip (SoC) with a Mali-T628 MP6 GPU for evaluation. Our evaluation using the Rodinia benchmark suite showed that the proposed approach limits the preemption delay to 18 μs in the 99.9th percentile with an average delay in execution time less than 10 % in most cases for high-priority tasks under a heavy load.
Jaehun Roh, Euiseong Seo
RTAS3
2017 Parity-Stream Separation and SLC/MLC Convertible Programming for Life Span and Performance Improvement of SSD RAIDs
Yoohyuk Lim, Cassiano Campes, Euiseong Seo
HotStorage4
2017 An enhanced DSM model for computation offloading
abstract
The distributed shared memory (DSM)-based computation offloading scheme allows collaborative multiple threads to dynamically migrate and execute across a mobile device and computing nodes. Despite this strong advantage, it misses a significant portion of the potential performance gain because the traditional DSM model is suboptimal for computation offloading. This paper proposes an enhanced DSM model that aims to enable multiple computing nodes to efficiently and reliably offload concurrent multiple threads from a mobile device. To achieve this design goal, we propose the following novel schemes: a) selective object tracking minimizes the set of objects to be monitored by the DSM layer; b) lock-thread repartitioning dynamically relocates threads and locks in order to reduce remote lock acquisitions and inter-node synchronizations; and c) thread-state checkpointing protects the data and context upon unexpected system failures. We implemented METEOR, which is a prototype based on the proposed schemes, and evaluated it with diverse applications. The evaluation showed that METEOR, with four computing nodes, improved the performance by up to 109% and reduced energy consumption by up to 52% in comparison with the previous DSM-based offloading scheme.
Yuhun Jun, Euiseong Seo
PerCom3
2017 Evaluation of Remote-I/O Support for a DSM-Based Computation Offloading Scheme
Yuhun Jun, Euiseong Seo
J. Comput. Sci. Technol.3
2014 Energy-credit scheduler: An energy-aware virtual machine scheduler for cloud systems
Nakku Kim, Jungwook Cho, Euiseong Seo
Future Gener. Comput. Syst.3
2013 Extensible Video Processing Framework in Apache Hadoop
abstract
Digital video is prominent big data spread all over the Internet. It is large not only in size but also in required processing power to extract useful information. Fast processing of excessive video reels is essential on criminal investigations, such as terrorism. This demo presents an extensible video processing framework in Apache Hadoop to parallelize video processing tasks in a cloud environment. Except for video transcending systems, there have been few systems that can perform various video processing in cloud computing environments. The framework employs FFmpeg for a video coder, and OpenCV for a image processing engine. To optimize the performance, it exploits MapReduce implementation details to minimize video image copy. Moreover, FFmpeg source code was modified and extended, to access and exchange essential data and information with Hadoop, effectively. A face tracking system was implemented on top of the framework for the demo, which traces the continuous face movements in a sequence of video frames. Since the system provides a web-based interface, people can try the system on site. In an 8-core environment with two quad-core systems, the system shows 75% of scalability.
Chungmo Ryu, Daecheol Lee, Minwook Jang, Cheolgi Kim, Euiseong Seo
CloudCom (2)5
2013 Virtual Battery: A testing tool for power-aware software
Youngjoo Woo, Seon-Yeong Park, Euiseong Seo
J. Syst. Archit.3
2013 Analysis of virtual machine live-migration as a method for power-capping
Jinkyu Jeong, Sung-hun Kim 0005, Hwanju Kim, Joonwon Lee, Euiseong Seo
J. Supercomput.5
2012 TwoB: a two-tier web browser architecture optimized for mobile network
abstract
The connection establishment phase including DNS lookups and TCP handshakes takes significantly long time during web browsing through mobile network. In this paper, we propose a novel web browser architecture that aims at improving mobile web browsing performance. Our approach delegates the connection establishment phase and HTTP header field delivery to a dedicated proxy server located at the joint point between WAN and mobile network to reduce both the number and size of packets on mobile network. Our evaluation showed that the proposed scheme reduces the number of mobile network packets by up to 52% and, consequently, shortens the average page loading time by up to 37%.
Junguk Cho, Jinkyu Jeong, Euiseong Seo
MoMM3
2012 Effectiveness Analysis of DVFS and DPM in Mobile Devices
Youngbin Seo, Jeongki Kim, Euiseong Seo
J. Comput. Sci. Technol.3
2012 A low-overhead networking mechanism for virtualized high-performance computing systems
Jae-Wan Jang, Euiseong Seo, Heeseung Jo, Jin-Soo Kim 0001
J. Supercomput.2
2012 Workload Characterization and Performance Implications of Large-Scale Blog Servers
abstract
With the ever-increasing popularity of Social Network Services (SNSs), an understanding of the characteristics of these services and their effects on the behavior of their host servers is critical. However, there has been a lack of research on the workload characterization of servers running SNS applications such as blog services. To fill this void, we empirically characterized real-world Web server logs collected from one of the largest South Korean blog hosting sites for 12 consecutive days. The logs consist of more than 96 million HTTP requests and 4.7TB of network traffic. Our analysis reveals the following: (i) The transfer size of nonmultimedia files and blog articles can be modeled using a truncated Pareto distribution and a log-normal distribution, respectively; (ii) user access for blog articles does not show temporal locality, but is strongly biased towards those posted with image or audio files. We additionally discuss the potential performance improvement through clustering of small files on a blog page into contiguous disk blocks, which benefits from the observed file access patterns. Trace-driven simulations show that, on average, the suggested approach achieves 60.6% better system throughput and reduces the processing time for file access by 30.8% compared to the best performance of the Ext4 filesystem.
Myeongjae Jeon, Youngjae Kim 0001, Jeaho Hwang, Joonwon Lee, Euiseong Seo
ACM Trans. Web5
2011 A comprehensive study of energy efficiency and performance of flash-based SSD
Seon-Yeong Park, Youngjae Kim 0001, Bhuvan Urgaonkar, Joonwon Lee, Euiseong Seo
J. Syst. Archit.5
2010 Log' version vector: Logging version vectors concisely in dynamic replication
Hyun-Gul Roh, Myeongjae Jeon, Euiseong Seo, Jin-Soo Kim 0001, Joonwon Lee
Inf. Process. Lett.3
2010 KAL: kernel-assisted non-invasive memory leak tolerance with a general-purpose memory allocator
abstract
Abstract Memory leaks are a continuing problem in the software developed with programming languages, such as C and C++. A recent approach adopted by some researchers is to tolerate leaks in the software application and to reclaim the leaked memory by use of specially constructed memory allocation routines. However, such routines replace the usual general‐purpose memory allocator and tend to be less efficient in speed and in memory utilization. We propose a new scheme which coexists with the existing memory allocation routines and which reclaims memory leaks. Our scheme identifies and reclaims leaked memory at the kernel level. There are some major advantages to our approach: (1) the application software does not need to be modified; (2) the application does not need to be suspended while leaked memory is reclaimed; (3) a remote host can be used to identify the leaked memory, thus minimizing impact on the application program's performance; and (4) our scheme does not degrade the service availability of the application while detecting and reclaiming memory leaks. We have implemented a prototype that works with the GNU C library and with the Linux kernel. Our prototype has been tested and evaluated with various real‐world applications. Our results show that the computational overhead of our approach is around 2% of that incurred by the conventional memory allocator in terms of throughput and average response time. We also verified that the prototype successfully suppressed address space expansion caused by memory leaks when the applications are run on synthetic workloads. Copyright © 2010 John Wiley & Sons, Ltd.
Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, Joonwon Lee
Softw. Pract. Exp.2
2010 Dynamic alteration schemes of real-time schedules for I/O device energy efficiency
abstract
Many I/O devices provide multiple power states known as the dynamic power management (DPM) feature. However, activating from sleep state requires significant transition time and this obstructs utilizing DPM in nonpreemptive real-time systems. This article suggests nonpreemptive real-time task scheduling schemes maximizing the effectiveness of the I/O device DPM support. First, we introduce a runtime schedulability check algorithm for nonpreemptive real-time systems that can check whether a modification from a valid schedule is still valid. By using this, we suggest three heuristic algorithms. The first algorithm reorders the execution sequence of tasks according to the similarity of their required device sets. The second one gathers dispersed short idle periods into one long idle period to extend sleeping state of I/O devices and the last one inserts an idle period between two consecutively scheduled tasks to prepare the required devices of a task right before the starting time of the task. The suggested schemes were evaluated for both the real-world task sets and the hypothetical task sets with simulation and the results showed that the suggested algorithms produced better energy efficiency than the existing comparative algorithms.
Euiseong Seo, Seon-Yeong Park, Joonwon Lee
ACM Trans. Embed. Comput. Syst.1
2009 Catching two rabbits: adaptive real-time support for embedded Linux
abstract
Abstract The trend of digital convergence makes multitasking common in many digital electronic products. Some applications in those systems have inherent real‐time properties, while many others have few or no timeliness requirements. Therefore the embedded Linux kernels, which are widely used in those devices, provide real‐time features in many forms. However, providing real‐time scheduling usually induces throughput degradation in heavy multitasking due to the increased context switches. Usually the throughput degradation becomes a critical problem, since the performance of the embedded processors is generally limited for cost, design and energy efficiency reasons. This paper proposes schemes to lessen the throughput degradation, which is from real‐time scheduling, by suppressing unnecessary context switches and applying real‐time scheduling mechanisms only when it is necessary. Also the suggested schemes enable the complete priority inheritance protocol to prevent the well‐known priority inversion problem. We evaluated the effectiveness of our approach with open‐source benchmarks. By using the suggested schemes, the throughput is improved while the scheduling latency is kept same or better in comparison with the existing approaches. Copyright © 2008 John Wiley & Sons, Ltd.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee
Softw. Pract. Exp.1
2008 Guest-Aware Priority-Based Virtual Machine Scheduling for Highly Consolidated Server
Hwanju Kim, Myeongjae Jeon, Euiseong Seo, Joonwon Lee
Euro-Par4
2008 TSB: A DVS algorithm with quick response for general purpose operating systems
Euiseong Seo, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee
J. Syst. Archit.1
2008 Energy Efficient Scheduling of Real-Time Tasks on Multicore Processors
abstract
Multicore processors deliver a higher throughput at lower power consumption than unicore processors. In the near future, they will thus be widely used in mobile real-time systems. There have been many research on energy-efficient scheduling of real-time tasks using DVS. These approaches must be modified for multicore processors, however, since normally all the cores in a chip must run at the same performance level. Thus, blindly adopting existing DVS algorithms that do not consider the restriction will result in a waste of energy. This article suggests Dynamic Repartitioning algorithm based on existing partitioning approaches of multiprocessor systems. The algorithm dynamically balances the task loads of multiple cores to optimize power consumption during execution. We also suggest Dynamic Core Scaling algorithm, which adjusts the number of active cores to reduce leakage power consumption under low load conditions. Simulation results show that Dynamic Repartitioning can produce energy savings of about 8 percent even with the best energy-efficient partitioning algorithm. The results also show that Dynamic Core Scaling can reduce energy consumption by about 26 percent under low load conditions.
Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Joonwon Lee
IEEE Trans. Parallel Distributed Syst.1
2007 Domain Level Page Sharing in Xen Virtual Machine Systems
Myeongjae Jeon, Euiseong Seo, Joonwon Lee
APPT2
2007 PABC: Power-Aware Buffer Cache Management for Low Power Consumption
abstract
Power consumed by memory systems becomes a serious issue as the size of the memory installed increases. With various low power modes that can be applied to each memory unit, the operating system can reduce the number of active memory units by collocating active pages onto a few memory units. This paper presents a memory management scheme based on this observation, which differs from other approaches in that all of the memory space is considered, while previous methods deal only with pages mapped to user address spaces. The buffer cache usually takes more than half of the total memory and the pages access patterns are different from those in user address spaces. Based on an analysis of buffer cache behavior and its interaction with the user space, our scheme achieves up to 63 percent more power reduction. Migrating a page to a different memory unit increases memory latencies, but it is shown to reduce the power consumed by an additional 4.4 percent
Min Lee, Euiseong Seo, Joonwon Lee, Jin-Soo Kim 0001
IEEE Trans. Computers2
2006 Dynamic Repartitioning of Real-Time Schedule on a Multicore Processor for Energy Efficiency
Euiseong Seo, Yongbon Koo, Joonwon Lee
EUC1