EDBT 2026 Demo / reviewers in the wild / expert
Jongmoo Choi
dblp:61/4476
· DBLP profile ↗
86ranked-venue papers
11as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 25 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSecurity and privacy · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARC: Adaptive Resource Coordination for Write Stall Mitigation in LSM-Tree
Guangxun Zhao, Yongjie Zhu, Suhwan Shin, See-hwan Yoo, Jongmoo Choi |
CCGrid | 5 |
| 2026 | FDPEmu: How to Separate Workloads for Better WAF on FDP SSDsabstractAs data-centric applications proliferate, mitigating Write Amplification Factor (WAF) in SSDs has become critical for sustaining performance and longevity. Flexible Data Placement (FDP), a recently ratified NVMe standard, addresses this by allowing hosts to guide data placement while retaining block interface compatibility. However, research on FDP is constrained by the scarcity of prototype device and the lack of emulation tools supporting its multi-stream architecture. In this paper, we propose FDPEmu, a high-fidelity FDP emulator extended from FEMU. FDPEmu addresses the architectural limitations of legacy emulators by implementing core FDP data structures, per-RUH write pointers, and isolation-aware Garbage Collection (GC). We also introduce Striding, a dynamic channel offset allocation technique, to mitigate channel contention in multi-stream environments. Validation against prototype FDP SSDs demonstrates high fidelity ($r \gt 0.89$ for skewed workloads) in capturing WAF trends, confirming the emulator as a credible research platform. Furthermore, our case studies on RocksDB and F2FS reveal that strictly separating data is not universally beneficial. We demonstrate that the effectiveness of isolation policies relies heavily on workload patterns, indicating that optimal FDP strategies must carefully balance data lifetime separation with effective resource utilization. Nakyeong Kim, Kwanghee Lee, Bryan S. Kim, See-hwan Yoo, Jaedong Lee, Jongmoo Choi |
ISPASS | 6 |
| 2025 | Storage Abstractions for SSDs: The Past, Present, and FutureabstractThis article traces the evolution of SSD (solid-state drive) interfaces, examining the transition from the block storage paradigm inherited from hard disk drives to SSD-specific standards customized to flash memory. Early SSDs conformed to the block abstraction for compatibility with the existing software storage stack, but studies and deployments show that this limits the performance potential for SSDs. As a result, new SSD-specific interface standards emerged to not only capitalize on the low latency and abundant internal parallelism of SSDs, but also include new command sets that diverge from the longstanding block abstraction. We first describe flash memory technology in the context of the block storage abstraction and the components within an SSD that provide the block storage illusion. We then describe the genealogy and relationships among academic research and industry standardization efforts for SSDs, along with some of their rise and fall in popularity. We classify these works into four evolving branches: (1) extending block abstraction with host-SSD hints/directives; (2) enhancing host-level control over SSDs; (3) offloading host-level management to SSDs; and (4) making SSDs byte-addressable. By dissecting these trajectories, the article also sheds light on the emerging challenges and opportunities, providing a roadmap for future research and development in SSD technologies. Xiangqun Zhang 0002, Janki Bhimani, Shuyi Pei, Sungjin Lee 0001, Yoon Jae Seong, Eui Jin Kim, Changho Choi, Eyee Hyun Nam, Jongmoo Choi, Bryan S. Kim |
ACM Trans. Storage | 10 |
| 2024 | Forensic Investigation of An Android Jellybean-based Car Audio Video Navigation SystemabstractRecently, in-vehicle infotainment (IVI) systems, also called car audio video navigation (AVN) systems hold a wealth of digital data valuable for forensic investigations, encompassing navigation history, call logs, and Bluetooth connections. They serve as central hubs for entertainment, communication, and navigation, storing crucial evidence for accidents, thefts, and cybercrimes. Therefore, forensic investigations of IVI systems are becoming increasingly important. In this paper, we conduct a forensic analysis of an Android Jellybean-based AVN system installed in Kia K5 2017. We first efficiently collect system logs as well as navigation logs using the log menu of an engineering mode provided by the car manufacturer company. Therefore, our data collection method does not require a chip-off technique or rooting of the AVN system. Next, we analyze the collected logs systematically and the differences between the two types of log data. Our forensic investigation method can provide insights into occupant activities and reconstruct events leading to incidents and car crimes. Jeehun Jung, Seong-je Cho, Jongmoo Choi, Minkyu Park |
ARES | 4 |
| 2024 | The Design and Implementation of a Capacity-Variant Storage System
Ziyang Jiao, Xiangqun Zhang 0002, Hojin Shin, Jongmoo Choi, Bryan S. Kim |
FAST | 4 |
| 2024 | Can Learned Indexes be Built Efficiently? A Deep Dive into Sampling Trade-offsabstractBy embedding the distribution of keys in indexing structure, learned indexes can minimize the index size and maximize the lookup performance. Yet, one of the problems in the present learned index is the long index-building time. The conventional learned index requires a complete traversal of the entire dataset, which makes it less practical than traditional index. This paper challenges the efficiency of build time to make the learned index practical. Our approach for a build time-efficient learned index is to employ sampled learning. In this paper, we present two error-bounded sampling schemes: Sample EB-PLA, and Sample EB-Histogram. Although sampling is a simple idea, there are several considerations to make it practical. For example, sampling interval, error-boundness, and index hyper-parameters are inter-related each other, presenting complicated trade-offs between build-time, index size, accuracy and lookup latency. Throughout the extensive experiments over six real-world datasets, we show that the index-building time can be efficiently reduced over an order of magnitude by our sampling schemes. The results reveal that the sampling expands the design space of learned indexes, including the build-time as well as lookup performance and index size. Our Pareto analysis shows that a learned index can be built more efficiently than a traditional index through sampling. Minguk Choi, See-hwan Yoo, Jongmoo Choi |
Proc. ACM Manag. Data | 3 |
| 2023 | Excessive SSD-Internal Parallelism Considered HarmfulabstractModern SSDs achieve high throughput by utilizing multiple independent channels and chips in parallel. However, we find that excessive parallelism inadvertently amplifies the garbage collection (GC) overhead due to the larger unit of space reclamation. Based on this observation, we design PLAN, a novel SSD parallelism management and data placement scheme that allocates different levels of parallelism to different workloads with different needs to minimize the GC overhead. We demonstrate the effectiveness of PLAN by evaluating it against other state-of-the-art designs across various real-world workloads. PLAN reduces write amplification with comparable or better performance to the other designs that are always at full parallelism. Xiangqun Zhang 0002, Shuyi Pei, Jongmoo Choi, Bryan S. Kim |
HotStorage | 3 |
| 2023 | ConfZNS : A Novel Emulator for Exploring Design Space of ZNS SSDsabstractThe ZNS (Zoned NameSpace) interface shifts much of the storage maintenance responsibility to the host from the underlying SSDs (Solid-State Drives). In addition, it opens a new opportunity to exploit the internal parallelism of SSDs at both hardware and software levels. By orchestrating the mapping between zones and SSD-internal resources and by controlling zone allocation among threads, ZNS SSDs provide a distinct performance trade-off between parallelism and isolation. To understand and explore the design space of ZNS SSDs, we present ConfZNS (Configurable ZNS), an easy-to-configure and timing-accurate emulator based on QEMU. ConfZNS allows users to investigate a variety of ZNS SSD's internal architecture and how it performs with existing host software. We validate the accuracy of ConfZNS using real ZNS SSDs and explore performance characteristics of different ZNS SSD designs with real-world applications such as RocksDB, F2FS, and Docker environment. Inho Song, Myounghoon Oh, Bryan S. Kim, See-hwan Yoo, Jaedong Lee, Jongmoo Choi |
SYSTOR | 6 |
| 2021 | LODIC: Logical Distributed Counting for Scalable File Access
Jeoungahn Park, Taeho Hwang, Jongmoo Choi, Changwoo Min, Youjip Won |
USENIX ATC | 3 |
| 2021 | Unsupervised video object segmentation with distractor-aware online adaptation
Ye Wang 0013, Jongmoo Choi, Yueru Chen, Siyang Li 0002, Qin Huang 0006, Kaitai Zhang, Ming-Sui Lee, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A New LSM-style Garbage Collection Scheme for ZNS SSDs
Gunhee Choi, Kwanghee Lee, Myunghoon Oh, Jongmoo Choi, Jhuyeong Jhin, Yongseok Oh |
HotStorage | 4 |
| 2020 | Video object tracking and segmentation with box annotation
Ye Wang 0013, Jongmoo Choi, Kaitai Zhang, Qin Huang 0006, Yueru Chen, Ming-Sui Lee, C.-C. Jay Kuo |
Signal Process. Image Commun. | 2 |
| 2019 | Human Motion Prediction via Learning Local Structure Representations and Temporal DependenciesabstractHuman motion prediction from motion capture data is a classical problem in the computer vision, and conventional methods take the holistic human body as input. These methods ignore the fact that, in various human activities, different body components (limbs and the torso) have distinctive characteristics in terms of the moving pattern. In this paper, we argue local representations on different body components should be learned separately and, based on such idea, propose a network, Skeleton Network (SkelNet), for long-term human motion prediction. Specifically, at each time-step, local structure representations of input (human body) are obtained via SkelNet’s branches of component-specific layers, then the shared layer uses local spatial representations to predict the future human pose. Our SkelNet is the first to use local structure representations for predicting the human motion. Then, for short-term human motion prediction, we propose the second network, named as Skeleton Temporal Network (Skel-TNet). Skel-TNet consists of three components: SkelNet and a Recurrent Neural Network, they have advantages in learning spatial and temporal dependencies for predicting human motion, respectively; a feed-forward network that outputs the final estimation. Our methods achieve promising results on the Human3.6M dataset and the CMU motion capture dataset, and the code is publicly available 1. Jongmoo Choi |
AAAI | 2 |
| 2019 | Design Tradeoffs for SSD Reliability
Bryan S. Kim, Jongmoo Choi, Sang Lyul Min |
FAST | 2 |
| 2019 | Age-invariant face recognition using gender specific 3D aging modeling
Sidra Riaz, Zahid Ali 0001, Unsang Park, Jongmoo Choi, Iacopo Masi, Premkumar Natarajan |
Multim. Tools Appl. | 4 |
| 2019 | Age progression by gender-specific 3D aging model
Sidra Riaz, Unsang Park, Jongmoo Choi, Premkumar Natarajan |
Mach. Vis. Appl. | 3 |
| 2019 | Learning Pose-Aware Models for Pose-Invariant Face Recognition in the WildabstractWe propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2]. Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | SPOT Poachers in Action: Augmenting Conservation Drones With Automatic Detection in Near Real TimeabstractThe unrelenting threat of poaching has led to increased development of new technologies to combat it. One such example is the use of long wave thermal infrared cameras mounted on unmanned aerial vehicles (UAVs or drones) to spot poachers at night and report them to park rangers before they are able to harm animals. However, monitoring the live video stream from these conservation UAVs all night is an arduous task. Therefore, we build SPOT (Systematic POacher deTector), a novel application that augments conservation drones with the ability to automatically detect poachers and animals in near real time. SPOT illustrates the feasibility of building upon state-of-the-art AI techniques, such as Faster RCNN, to address the challenges of automatically detecting animals and poachers in infrared images. This paper reports (i) the design and architecture of SPOT, (ii) a series of efforts towards more robust and faster processing to make SPOT usable in the field and provide detections in near real time, and (iii) evaluation of SPOT based on both historical videos and a real-world test run by the end users in the field. The promising results from the test in the field have led to a plan for larger-scale deployment in a national park in Botswana. While SPOT is developed for conservation drones, its design and novel techniques have wider application for automated detection from UAV videos. Elizabeth Bondi-Kelly, Fei Fang 0001, Mark Hamilton, Debarun Kar, Donnabell Dmello, Jongmoo Choi, Robert Hannaford, Arvind Iyer, Lucas Joppa, Milind Tambe, Ramakant Nevatia |
AAAI | 6 |
| 2018 | Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation
Ye Wang 0013, Jongmoo Choi, Yueru Chen, Qin Huang 0006, Siyang Li 0002, Ming-Sui Lee, C.-C. Jay Kuo |
ACCV (4) | 2 |
| 2017 | Utility-Based Hybrid Memory ManagementabstractWhile the memory footprints of cloud and HPC applications continue to increase, fundamental issues with DRAM scaling are likely to prevent traditional main memory systems, composed of monolithic DRAM, from greatly growing in capacity. Hybrid memory systems can mitigate the scaling limitations of monolithic DRAM by pairing together multiple memory technologies (e.g., different types of DRAM, or DRAM and non-volatile memory) at the same level of the memory hierarchy. The goal of a hybrid main memory is to combine the different advantages of the multiple memory types in a cost-effective manner while avoiding the disadvantages of each technology. Memory pages are placed in and migrated between the different memories within a hybrid memory system, based on the properties of each page. It is important to make intelligent page management (i.e., placement and migration) decisions, as they can significantly affect system performance.In this paper, we propose utility-based hybrid memory management (UH-MEM), a new page management mechanism for various hybrid memories, that systematically estimates the utility (i.e., the system performance benefit) of migrating a page between different memory types, and uses this information to guide data placement. UH-MEM operates in two steps. First, it estimates how much a single application would benefit from migrating one of its pages to a different type of memory, by comprehensively considering access frequency, row buffer locality, and memory-level parallelism. Second, it translates the estimated benefit of a single application to an estimate of the overall system performance benefit from such a migration.We evaluate the effectiveness of UH-MEM with various types of hybrid memories, and show that it significantly improves system performance on each of these hybrid memories. For a memory system with DRAM and non-volatile memory, UH-MEM improves performance by 14% on average (and up to 26%) compared to the best of three evaluated state-of-the-art mechanisms across a large number of data-intensive workloads. Yang Li 0183, Saugata Ghose, Jongmoo Choi, Onur Mutlu |
CLUSTER | 3 |
| 2017 | Local-Global Landmark Confidences for Face RecognitionabstractA key to successful face recognition is accurate and reliable face alignment using automatically-detected facial landmarks. Given this strong dependency between face recognition and facial landmark detection, robust face recognition requires knowledge of when the facial landmark detection algorithm succeeds and when it fails. Facial landmark confidence represents this measure of success. In this paper, we propose two methods to measure landmark detection confidence: local confidence based on local predictors of each facial landmark, and global confidence based on a 3D rendered face model. A score fusion approach is also introduced to integrate these two confidences effectively. We evaluate both confidence metrics on two datasets for face recognition: JANUS CS2 and IJB-A datasets. Our experiments show up to 9% improvements when face recognition algorithm integrates the local-global confidence metrics. KangGeon Kim, Feng-Ju Chang, Jongmoo Choi, Louis-Philippe Morency, Ramakant Nevatia, Gérard G. Medioni |
FG | 3 |
| 2017 | Deep 3D face identificationabstractWe propose a novel 3D face recognition algorithm using a deep convolutional neural network (DCNN) and a 3D face expression augmentation technique. The performance of 2D face recognition algorithms has significantly increased by leveraging the representational power of deep neural networks and the use of large-scale labeled training data. In this paper, we show that transfer learning from a CNN trained on 2D face images can effectively work for 3D face recognition by fine-tuning the CNN with an extremely small number of 3D facial scans. We also propose a 3D face expression augmentation technique which synthesizes a number of different facial expressions from a single 3D face scan. Our proposed method shows excellent recognition results on Bosphorus, BU-3DFE, and 3D-TEC datasets without using hand-crafted features. The 3D face identification using our deep features also scales well for large databases. Donghyun Kim 0006, Matthias Hernandez, Jongmoo Choi, Gérard G. Medioni |
IJCB | 3 |
| 2017 | Accurate 3D face reconstruction via prior constrained structure from motion
Matthias Hernandez, Tal Hassner, Jongmoo Choi, Gérard G. Medioni |
Comput. Graph. | 3 |
| 2016 | Accurate 3D face modeling and recognition from RGB-D stream in the presence of large pose changesabstractWe propose a 3D face modeling and recognition system using an RGB-D stream in the presence of large pose changes. In the previous work, all facial data points are registered with a reference to improve the accuracy of 3D face model from a low-resolution depth sequence. This registration often fails when applied to non-frontal faces. It causes inaccurate 3D face models and poor performance of matching. We address this problem by pre-aligning each input face (`frontalization') before the registration, which avoids registration failures. For each frame, our method estimates the 3D face pose, assesses the quality of data, segments the facial region, frontalizes it, and performs an accurate registration with the previous 3D model. The 3D-3D recognition system using accurate 3D models from our method outperforms other face recognition systems and shows 100% rank 1 recognition accuracy on a dataset with 30 subjects. Donghyun Kim 0006, Jongmoo Choi, Jatuporn Toy Leksut, Gérard G. Medioni |
ICIP | 2 |
| 2016 | Expression invariant 3D face modeling from an RGB-D videoabstractWe aim to reconstruct an accurate neutral 3D face model from an RGB-D video in the presence of extreme expression changes. Since each depth frame, taken by a low-cost sensor, is noisy, point clouds from multiple frames can be registered and aggregated to build an accurate 3D model. However, direct aggregation of multiple data produces erroneous results in natural interaction (e.g., talking and showing expressions). We propose to analyze facial expression from an RGB frame and neutralize the corresponding 3D point cloud if needed. We first estimate the person's expression by fitting blendshape coefficients using 2D facial landmarks for each frame and calculate an expression deformity (expression score). With the estimated expression score, we determine whether an input face is neutral or non-neutral. If the face is non-neutral, we proceed to neutralize the expression of the 3D point cloud in that frame. To neutralize the 3D point cloud of a face, we deform our generic 3D face model by applying the estimated blendshape coefficients, find displacement vectors from the deformed generic face to a neutral generic face, and apply the displacement vectors to the input 3D point cloud. After preprocessing frames in a video, we rank frames based on the expression scores and register the ranked frames into a single 3D model. Our system produces a neutral 3D face model in the presence of extreme expression changes even when neutral faces do not exist in the video. Donghyun Kim 0006, Jongmoo Choi, Jatuporn Toy Leksut, Gérard G. Medioni |
ICPR | 2 |
| 2016 | Face recognition using deep multi-pose representationsabstractWe introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks. Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni |
WACV | 7 |
| 2016 | Chip-Level RAID with Flexible Stripe Size and Parity Placement for Enhanced SSD ReliabilityabstractThe move from SLC to MLC/TLC flash memory technology is increasing SSD capacity at lower cost, but at the cost of sacrificing reliability. An approach to remedy this loss is to employ the RAID architecture with the chips that comprise SSDs. However, using the traditional RAID approach may result in negative effects as the total number of writes is increased due to the parity updates. In this paper, we describe Elastic Striping and Anywhere Parity (eSAP)-RAID, a RAID scheme that allows flexible stripe sizes and parity placement. Using performance and lifetime models that we derive of SSDs employing RAID-5 and eSAP-RAID, we show that eSAP-RAID brings about significant performance and reliability benefits by reducing parity writes compared to RAID-5. We also implement these schemes in SSDs using DiskSim with SSD Extension and validate the models using realistic workloads. We also discuss policies such as dynamic stripe sizing and selective data protection that exploits the flexible nature of eSAP. We show that through such policies particular reliability enhancement goals can be met. Eunjae Lee, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
IEEE Trans. Computers | 3 |
| 2016 | Exploiting Compression-Induced Internal Fragmentation for Power-Off Recovery in SSDabstractRecovery from sudden power-off (SPO) is one of the primary concerns among practitioners which bars the quick and wide deployment of flash storage devices. In this work, we propose Metadata Embedded Write (MEW), a novel scheme for handling the sudden power-off recovery in modern flash storage devices. Given that a large fraction of commercial SSDs employ compression technology, MEW exploits the compression-induced internal fragmentation in the data area to store rich metadata for fast and complete recovery. MEW consists of (i) a metadata embedding scheme to harbor SSD metadata in a physical page together with multiple compressed logical pages, (ii) an allocation chain based fast recovery scheme, and (iii) a light-weight metadata logging scheme which enables MEW to maintain the metadata for incompressible data, too. We performed extensive experiments to examine the performance of MEW. The performance overhead of MEW is 3 percent in the worst case, in terms of the write amplification factor, compared to the pure compression-based FTL that does not have any recovery scheme. Youjip Won, Jaehyuk Cha, Sungroh Yoon, Jongmoo Choi, Sooyong Kang |
IEEE Trans. Computers | 5 |
| 2015 | Decoupled Direct Memory Access: Isolating CPU and IO Traffic by Leveraging a Dual-Data-Port DRAMabstractMemory channel contention is a critical performance bottleneck in modern systems that have highly parallelized processing units operating on large data sets. The memory channel is contended not only by requests from different user applications (CPU access) but also by system requests for peripheral data (IO access), usually controlled by Direct Memory Access (DMA) engines. Our goal, in this work, is to improve system performance byeliminating memory channel contention between CPU accesses and IO accesses. To this end, we propose a hardware-software cooperative data transfer mechanism, Decoupled DMA (DDMA) that provides a specialized low-cost memory channel for IO accesses. In our DDMA design, main memoryhas two independent data channels, of which one is connected to the processor (CPU channel) and the other to the IO devices (IO channel), enabling CPU and IO accesses to be served on different channels. Systemsoftware or the compiler identifies which requests should be handled on the IO channel and communicates this to the DDMA engine, which then initiates the transfers on the IO channel. By doing so, our proposal increasesthe effective memory channel bandwidth, thereby either accelerating data transfers between system components, or providing opportunities to employ IO performance enhancement techniques (e.g., aggressive IO prefetching)without interfering with CPU accessesWe demonstrate the effectiveness of our DDMA framework in two scenarios: (i) CPU-GPU communication and (ii) in-memory communication (bulk datacopy/initialization within the main memory). By effectively decoupling accesses for CPU-GPU communication and in-memory communication from CPU accesses, our DDMA-based design achieves significant performanceimprovement across a wide variety of system configurations (e.g., 20% average performance improvement on a typical 2-channel 2-rank memory system). Donghyuk Lee, Lavanya Subramanian, Rachata Ausavarungnirun, Jongmoo Choi, Onur Mutlu |
PACT | 4 |
| 2015 | Convex Cut: A realtime pseudo-structure extraction algorithm for 3D point cloud dataabstractIn this paper, a realtime pseudo-structure extraction algorithm for 3D indoor point cloud data (PCD) is proposed. This algorithm is called Convex Cut (CC) because of its two main steps: cutting the PCD with arbitrary planes, and extracting convex parts. CC can be used as a preprocessing module for other existing algorithms to extract static parts in dynamic environments or to represent a principal 3D model of a given PCD. Its calculation time is 24 milliseconds for 50k PCD on a consumer PC, and it yields a precision value of 0.90 and a recall value of 0.99 on average in highly dynamic and cluttered environments. Some possible applications are explained such as simultaneous localization and mapping in dynamic environments, efficient dense map representation, robust 3D scan matching with plane features, and natural motion planning. ChangHyun Jun, Jihwan Youn, Jongmoo Choi, Gérard G. Medioni, Nakju Lett Doh |
IROS | 3 |
| 2015 | ThyNVM: enabling software-transparent crash consistency in persistent memory systemsabstractEmerging byte-addressable nonvolatile memories (NVMs) promise persistent memory, which allows processors to directly access persistent data in main memory. Yet, persistent memory systems need to guarantee a consistent memory state in the event of power loss or a system crash (i.e., crash consistency). To guarantee crash consistency, most prior works rely on programmers to (1) partition persistent and transient memory data and (2) use specialized software interfaces when updating persistent memory data. As a result, taking advantage of persistent memory requires significant programmer effort, e.g., to implement new programs as well as modify legacy programs. Use cases and adoption of persistent memory can therefore be largely limited. Jinglei Ren, Jishen Zhao, Samira Manabi Khan, Jongmoo Choi, Yongwei Wu 0001, Onur Mutlu |
MICRO | 4 |
| 2015 | Enabling Cost-Effective Flash based Caching with an Array of Commodity SSDsabstractSSD based cache solutions are being widely utilized to improve performance in network storage systems. With a goal of providing a cost-effective, high performing SSD cache solution, we propose a new caching solution called SRC (SSD RAID as a Cache) for an array of commodity SSDs. In designing SRC, we borrow both the well-known RAID technique and the log-structured approach and adopt them into the cache layer. In so doing, we explore a wide variety of design choices such as flush issue frequency, write units, forming stripes without parity, and garbage collection through copying rather than destaging that become possible as we make use of RAID and a log-structured approach at the cache level. Using an implementation in Linux under the Device Mapper framework, we quantitatively present and analyze results of the design space options that we considered in our design. Our experiments using realistic workload traces show that SRC performs at least 2 times better in terms of throughput than existing open source solutions. We also consider cost-effectiveness of SRC with a variety of SSD products. In particular, we compare SRC configured with MLC and TLC SATA SSDs and a single high-end NVMe SSD. We find that SRC configured as RAID-5 with low-cost MLC and TLC SATA SSDs generally outperforms that configured with a single high-end SSD in terms of both performance and lifetime per dollars spent. Yongseok Oh, Eunjae Lee, Choulseung Hyun, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
Middleware | 4 |
| 2015 | Amnesic cache management for non-volatile memoryabstractOne characteristic of non-volatile memory (NVM) is that, even though it supports non-volatility, its retention capability is limited. To handle this issue, previous studies have focused on refreshing or advanced error correction code (ECC). In this paper, we take a different approach that makes use of the limited retention capability to our advantage. Specifically, we employ NVM as a file cache and devise a new scheme called amnesic cache management (ACM). The scheme is motivated by our observation that most data in a cache are evicted within a short time period after they have been entered into the cache, implying that they can be written with the relaxed retention capability. This retention relaxation can enhance the overall cache performance in terms of latency and energy since the data retention capability is proportional to the write latency. In addition, to prevent the retention relaxation from degrading the hit ratio, we estimate the future reference intervals based on the inter-reference gap (IRG) model and manage data adaptively. Experimental results with real-world workloads show that our scheme can reduce write latency by up to 40% (30% on average) and save energy consumption by up to 49% (37% on average) compared with the conventional LRU based cache management scheme. Seungjae Baek, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh, Onur Mutlu |
MSST | 3 |
| 2015 | WARM: Improving NAND flash memory lifetime with write-hotness aware retention managementabstractIncreased NAND flash memory density has come at the cost of lifetime reductions. Flash lifetime can be extended by relaxing internal data retention time, the duration for which a flash cell correctly holds data. Such relaxation cannot be exposed externally to avoid altering the expected data integrity property of a flash device. Reliability mechanisms, most prominently refresh, restore the duration of data integrity, but greatly reduce the lifetime improvements from retention time relaxation by performing a large number of write operations. We find that retention time relaxation can be achieved more efficiently by exploiting heterogeneity in write-hotness, i.e., the frequency at which each page is written. We propose WARM, a write-hotness aware retention management policy for flash memory, which identifies and physically groups together write-hot data within the flash device, allowing the flash controller to selectively perform retention time relaxation with little cost. When applied alone, WARM improves overall flash lifetime by an average of 3.24× over a conventional management policy without refresh, across a variety of real I/O workload traces. When WARM is applied together with an adaptive refresh mechanism, the average lifetime improves by 12.9×, 1.21× over adaptive refresh alone. Yu Cai 0001, Saugata Ghose, Jongmoo Choi, Onur Mutlu |
MSST | 4 |
| 2015 | Incremental redundancy to reduce data retention errors in flash-based SSDsabstractAs the market becomes competitive, SSD manufacturers are making use of multi-bit cell flash memory such as MLC and TLC chips in their SSDs. However, these chips have lower data retention period and endurance than SLC chips. With the reduced data retention period and endurance level, retention errors occur more frequently. One solution for these retention errors is to employ strong ECC to increase error correction strength. However, employing strong ECC may result in waste of resources during the early stages of flash memory lifetime as it has high reliability and data retention errors are rare during this period. The other solution is to employ data scrubbing that periodically refreshes data by reading and then writing the data to new locations after correcting errors through ECC. Though it is a viable solution for the retention error problem, data scrubbing hurts performance and lifetime of SSDs as it incurs extra read and write requests. Targeting data retention errors, we propose incremental redundancy (IR) that incrementally reinforces error correction capabilities when the data retention error rate exceeds a certain threshold. This extends the time before data scrubbing should occur, providing a grace period in which the block may be garbage collected. We develop mathematical analyses that project the lifetime and performance of IR as well as when using conventional data scrubbing. Through mathematical analyses and experiments with both synthetic and real workloads, we compare the lifetime and performance of the two schemes. Results suggest that IR can be a promising solution to overcome data retention errors of contemporary multi-bit cell flash memory. In particular, our study shows that IR can extend the maximum data retention period by 5 to 10 times. Additionally, we show that IR can reduce the write amplification factor by half under real workloads. Heejin Park, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
MSST | 3 |
| 2015 | A-DRM: Architecture-aware Distributed Resource Management of Virtualized ClustersabstractVirtualization technologies has been widely adopted by large-scale cloud computing platforms. These virtualized systems employ distributed resource management (DRM) to achieve high resource utilization and energy savings by dynamically migrating and consolidating virtual machines. DRM schemes usually use operating-system-level metrics, such as CPU utilization, memory capacity demand and I/O utilization, to detect and balance resource contention. However, they are oblivious to microarchitecture-level resource interference (e.g., memory bandwidth contention between different VMs running on a host), which is currently not exposed to the operating system. Canturk Isci, Lavanya Subramanian, Jongmoo Choi, Depei Qian 0001, Onur Mutlu |
VEE | 4 |
| 2015 | Near laser-scan quality 3-D face reconstruction from a low-quality depth stream
Matthias Hernandez, Jongmoo Choi, Gérard G. Medioni |
Image Vis. Comput. | 2 |
| 2015 | iBuddy: Inverse Buddy for Enhancing Memory Allocation/Deallocation Performanceon Multi-Core SystemsabstractWe present a new buddy system for memory allocation that we call the lazy iBuddy system. This system is motivated by two observations of the widely used lazy buddy system on multi-core systems. First, most memory requests are for single page frames. However, the lazy buddy algorithm used in Linux continuously splits and coalesces memory blocks for single page frame requests even though the lazy layer is employed. Second, on multi-core systems, responses to bursty memory requests are delayed by lock contention caused by concurrent accesses of the multi-cores. The lazy iBuddy system overcomes the first problem by managing each page frame individually and coalescing pages only when an allocation of multiple page frames is requested. We devise the lazy iBuddy algorithm so that single page frame allocation can be done in O(1). The second problem is alleviated by dividing main memory into multiple buddy spaces and applying a fine-grained locking mechanism. Performance evaluation results based on various workloads on the XEON 16core with 32 GB main memory show that the lazy iBuddy system can improve memory allocation/deallocation time by up to 47 percent with an average of 35 percent compared with the lazy buddy system for the various configurations that we considered. Heekwon Park, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
IEEE Trans. Computers | 2 |
| 2015 | Design Tradeoffs of SSDs: From Energy Consumption's PerspectiveabstractIn this work, we studied the energy consumption characteristics of various SSD design parameters. We developed an accurate energy consumption model for SSDs that computes aggregate, as well as component-specific, energy consumption of SSDs in sub-msec time scale. In our study, we used five different FTLs (page mapping, DFTL, block mapping, and two different hybrid mappings) and four different channel configurations (two, four, eight, and 16 channels) under seven different workloads (from large-scale enterprise systems to small-scale desktop applications) in a combinatorial manner. For each combination of the aforementioned parameters, we examined the energy consumption for individual hardware components of an SSD (microcontroller, DRAM, NAND flash, and host interface). The following are some of our findings. First, DFTL is the most energy-efficient address-mapping scheme among the five FTLs we tested due to its good write amplification and small DRAM footprint. Second, a significant fraction of energy is being consumed by idle flash chips waiting for the completion of NAND operations in the other channels. FTL should be designed to fully exploit the internal parallelism so that energy consumption by idle chips is minimized. Third, as a means to increase the internal parallelism, increasing way parallelism (the number of flash chips in a channel) is more effective than increasing channel parallelism in terms of peak energy consumption, performance, and hardware complexity. Fourth, in designing high-performance and energy-efficient SSDs, channel switching delay, way switching delay, and page write latency need to be incorporated in an integrated manner to determine the optimal configuration of internal parallelism. Seokhei Cho, Changhyun Park, Youjip Won, Sooyong Kang, Jaehyuk Cha, Sungroh Yoon, Jongmoo Choi |
ACM Trans. Storage | 7 |
| 2014 | 3D Modeling from Wide Baseline Range Scans Using Contour CoherenceabstractRegistering 2 or more range scans is a fundamental problem, with application to 3D modeling. While this problem is well addressed by existing techniques such as ICP when the views overlap significantly at a good initialization, no satisfactory solution exists for wide baseline registration. We propose here a novel approach which leverages contour coherence and allows us to align two wide baseline range scans with limited overlap from a poor initialization. Inspired by ICP, we maximize the contour coherence by building robust corresponding pairs on apparent contours and minimizing their distances in an iterative fashion. We use the contour coherence under a multi-view rigid registration framework, and this enables the reconstruction of accurate and complete 3D models from as few as 4 frames. We further extend it to handle articulations, and this allows us to model articulated objects such as human body. Experimental results on both synthetic and real data demonstrate the effectiveness and robustness of our contour coherence based registration approach to wide baseline range scans, and to 3D modeling. Ruizhe Wang 0002, Jongmoo Choi, Gérard G. Medioni |
CVPR | 2 |
| 2014 | Learning symbolic descriptions of activities from examples in WAASabstractWe present an automatic system that learns symbolic representations of activities from examples in Wide Area Aerial Surveillance (WAAS). In the previous work, we presented an ERM (Entity Relationship Models)-based activity recognition system in which finding an activity is equivalent to sending a query, defined by SQL statements, to a Relational DataBase Management System (RDBMS). The system enables us to identify spatial and geo-spatial activities in WAAS as long as activities are carefully defined by human operators. Here, we show how to infer a structured definition of an activity from examples provided by a user. Our system randomly generates a set of possible SQL statements using a logic generator in a MCMC framework, uses a memory-based RDBMS to validate generated SQL statements with the input data/database, and selects the best answer that allows the RDBMS to explain the input positive examples while excluding negative examples. We have evaluated our system on real visual tracks. Our system can find activity definitions from input examples and associated query results including motion patterns (e.g., "loop") and geospatial activities (e.g., "parking in a lot"). Jongmoo Choi, Gérard G. Medioni |
SIGSPATIAL/GIS | 1 |
| 2014 | Design space exploration of an NVM-based memory hierarchyabstractNon-volatile memory (NVM) technologies support both byte addressability (like DRAM) and non-volatility (like disks). This characteristic makes it feasible for NVM to be employed at any layer of the memory hierarchy including CPU cache, main memory, file cache, storage, and hybrid memory. In this paper, we explore new challenges and opportunities that arise when NVM is introduced as a file cache in the memory hierarchy. One opportunity is that cache does not require long-term non-volatility since data are replaced when working sets are changed. This feature is well matched with NVM, which has limited retention time. In addition, the retention time of NVM is inverse proportional to the write latency, giving a chance to optimize the write performance. However, the limited retention time raises a new challenge that it may cause the hit ratio reduction and lead to cache performance degradation. To tackle this challenge, we propose a new inter-reference gap (IRG) based cache management scheme that writes data with different retention times according to their IRGs. Our proposal builds on the fact that block accesses of typical workloads show unique and regular patterns in terms of access intervals. Experimental results show that our scheme enhances system performance by up to 42% (33% on average), compared with the conventional LRU based cache management scheme. Seungjae Baek, Daeyeon Son, Jongmoo Choi, Sangyeun Cho |
ICCD | 4 |
| 2014 | Aerial Implicit 3D Video Stabilization Using Epipolar Geometry ConstraintabstractWe present an accurate video stabilization method on aerial videos using the epipolar geometry constraint. Most previous methods used 2D homography for stabilization, but failed to overcome the parallax problem. In this work, we propose to use dense correspondences for stabilization and the epipolar constraint to deal with the parallax effect. We start by estimating the dense correspondences between two frames. The dense correspondences are then used to estimate the epipolar geometry. The epipolar geometry has an implicit 3D constraint that can be used to improve the dense correspondences and handle the parallax. We evaluate our method on a real-life database containing three aerial image sequences. We also compare our method with the dominant 2D method for aerial video stabilization. The quantitative result demonstrates the effectiveness of our approach. Loc Huynh, Jongmoo Choi, Gérard G. Medioni |
ICPR | 2 |
| 2014 | Analytical model of SSD parallelism
Jinsoo Yoo, Youjip Won, Sooyong Kang, Jongmoo Choi, Sungroh Yoon, Jaehyuk Cha |
SIMULTECH | 4 |
| 2014 | Real-time 3-D face tracking and modeling framework for mid-res camabstractWe present a robust, real-time 3-D face tracking and modeling system providing accurate 6 degree-of-freedom head pose in the presence of large out-of-plane motion, strong expression changes, and partial occlusions. In this paper, we have extended the previous 3-D face tracking and modeling framework [10] with automatic initialization, reacquisition, and automatic pose correction. Our system first generates a 3-D face model from a single frontal image. We then extract uniformly distributed random points and track them in 2-D. Given these correspondences, the 3-D head pose is robustly estimated using a RANSAC-PnP process. As the head moves, we dynamically add new feature points to handle a large range of poses. A measure of the accumulated error over time allows an auto-correction mechanism to recover from drift when necessary. If the tracker gets lost, due to motion blur or strong occlusions, the system re-initializes. We present live demo results, which shows excellent tracking under large motion (roll: 360°, yaw: ±90°, pitch: −60° to +90°), fast movement, occlusion and facial expression variations. The system runs at 14 fps on a laptop CPU. By experiments on different datasets, our method shows state of the art results. Jongmoo Choi, Anh Tuan Tran 0001, Yann Dumortier, Gérard G. Medioni |
WACV | 1 |
| 2014 | IO Workload Characterization Revisited: A Data-Mining ApproachabstractOver the past few decades, IO workload characterization has been a critical issue for operating system and storage community. Even so, the issue still deserves investigation because of the continued introduction of novel storage devices such as solid-state drives (SSDs), which have different characteristics from traditional hard disks. We propose novel IO workload characterization and classification schemes, aiming at addressing three major issues: (i) deciding right mining algorithms for IO traffic analysis, (ii) determining a feature set to properly characterize IO workloads, and (iii) defining essential IO traffic classes state-of-the-art storage devices can exploit in their internal management. The proposed characterization scheme extracts basic attributes that can effectively represent the characteristics of IO workloads and, based on the attributes, finds representative access patterns in general workloads using various clustering algorithms. The proposed classification scheme finds a small number of representative patterns of a given workload that can be exploited for optimization either in the storage stack of the operating system or inside the storage device. Bumjoon Seo, Sooyong Kang, Jongmoo Choi, Jaehyuk Cha, Youjip Won, Sungroh Yoon |
IEEE Trans. Computers | 3 |
| 2014 | Unified security enhancement framework for the Android operating system
Seong-je Cho, Jongmoo Choi, Yeong-Ung Park |
J. Supercomput. | 4 |
| 2013 | Regularities considered harmful: forcing randomness to memory accesses to reduce row buffer conflicts for multi-core, multi-bank systemsabstractWe propose a novel kernel-level memory allocator, called M3 (M-cube, Multi-core Multi-bank Memory allocator), that has the following two features. First, it introduces and makes use of a notion of a memory container, which is defined as a unit of memory that comprises the minimum number of page frames that can cover all the banks of the memory organization, by exclusively assigning a container to a core so that each core achieves bank parallelism as much as possible. Second, it orchestrates page frame allocation so that pages that threads access are dispersed randomly across multiple banks so that each thread's access pattern is randomized. The development of M3 is based on a tool that we develop to fully understand the architectural characteristics of the underlying memory organization. Using an extension of this tool, we observe that the same application that accesses pages in a random manner outperforms one that accesses pages in a regular pattern such as sequential or same ordered accesses. This is because such randomized accesses reduces inter-thread access interference on the row-buffer in memory banks. We implement M3 in the Linux kernel version 2.6.32 on the Intel Xeon system that has 16 cores and 32GB DRAM. Performance evaluation with various workloads show that M3 improves the overall performance for memory intensive benchmarks by up to 85% with an average of about 40%. Heekwon Park, Seungjae Baek, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
ASPLOS | 3 |
| 2013 | Improving SSD reliability with RAID via Elastic Striping and Anywhere ParityabstractWhile the move from SLC to MLC/TLC flash memory technology is increasing SSD capacity at lower cost, it is being done at the cost of sacrificing reliability. An approach to remedy this loss is to employ the RAID architecture with the chips that comprise SSDs. However, using the traditional RAID approach may result in negative effects as the total number of writes may increase due to the parity updates, consequently leading to increased P/E cycles and higher bit error rates. Using a technique that we call Elastic Striping and Anywhere Parity (eSAP), we develop eSAP-RAID, a RAID scheme that significantly reduces parity writes while providing reliability better than RAID-5. We derive performance and lifetime models of SSDs employing RAID-5 and eSAP-RAID that show the benefits of eSAP-RAID. We also implement these schemes in SSDs using DiskSim with SSD Extension and validate the models using realistic workloads. Our results show that eSAP-RAID improves reliability considerably, while limiting its wear. Specifically, the expected lifetime of eSAP-RAID employing SSDs may be as long as current ECC based SSDs, while its reliability level can be maintained at the level of the early stages of current ECC based SSDs throughout its entire lifetime. Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
DSN | 3 |
| 2013 | Hybrid solid state drives for improved performance and enhanced lifetimeabstractAs the market becomes more competitive, SSD manufacturers are moving from SLC (Single-Level Cell) to MLC (Multi-Level Cell) flash memory chips that store two bits per cell as building blocks for SSDs. Recently, TLC chips, which store three bits per cell, is being considered as a viable solution due to their low cost. However, performance and lifetime of TLC chips are considerably limited and thus, pure TLC-based SSDs may not be viable as a general storage device. In this paper, we propose a hybrid SSD solution, namely HySSD, where SLC and TLC chips are used together to form an SSD solution performing in par with SLC-based products. Based on an analytical model, we propose a near optimal data distribution scheme that distributes data among the SLC and TLC chips for a given workload such that performance or lifetime may be optimized. Experiments with two types of SSDs both based on DiskSim with SSD Extension show that the analytic model approach can dynamically adjust data distribution as workloads evolve to enhance performance or lifetime. Yongseok Oh, Eunjae Lee, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
MSST | 3 |
| 2013 | VSSIM: Virtual machine based SSD simulatorabstractIn this paper, we present a virtual machine based SSD Simulator, VSSIM (Virtual SSD Simulator). VSSIM intends to address the issues of the trace driven simulation, e.g. trace re-scaling, accurate replay, etc. VSSIM operates on top of QEMU/KVM with software based SSD module. VSSIM runs in realtime and allows the user to measure both the host performance and the SSD behavior under various design choices. VSSIM can flexibly model the various hardware components, e.g. the number of channels, the number of ways, block size, page size, planes per chip, program, erase, read latency of NAND cells, channel switch delay, and way switch delay. VSSIM can also facilitate the implementation of the SSD firmware algorithms. To demonstrate the capability of VSSIM, we performed a number of case studies. The results of the simulation study deliver an important guideline in the firmware and hardware designs of future NAND based storage devices. Followings are some of the findings: (i) as the page size increases, the performance benefit of increasing the channel parallelism against increasing the way parallelism becomes less significant, (ii) due to the bi-modality in IO size distribution, FTL should be designed to handle multiple mapping granularity, (iii) hybrid mapping does not work in four or more way SSD due to severe log block fragmentation, (iv) as a performance metric, the Write Amplification Factor can be misleading, (v) compared to sequential write, random write operation can be benefited more from the channel level parallelism and therefore in multi-channel environment, it is beneficial to categorize larger fraction of IO as random. VSSIM is validated against commodity SSD, Intel X25M SSD. VSSIM models the sequential IO performance of X25M within 3% offset. Jinsoo Yoo, Youjip Won, Joongwoo Hwang, Sooyong Kang, Jongmoo Choi, Sungroh Yoon, Jaehyuk Cha |
MSST | 5 |
| 2013 | Comparing strategies for 3D face recognition from a 3D sensorabstractWe address the problem of 3D face recognition from 3D data, using different strategies. One strategy (1F-NF), explored earlier, is to match each individual frame to a set of reference frames. A second one (1F-3D) is to replace the set of reference frames by a 3D model resulting from the integration of individual frames. A third strategy (3D-3D) is to use a 3D face model inferred from multiple frames as the input probe. We show that the recognition performance using 3D model to 3D model outperforms the others, at the cost of a delay in response, due to the model building step. Jongmoo Choi, Ayush Sharma, Gérard G. Medioni |
RO-MAN | 1 |
| 2013 | Towards greener data centers with storage class memory
In Hwan Doh, Eunsam Kim, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
Future Gener. Comput. Syst. | 4 |
| 2013 | Energy-efficient and high-performance software architecture for storage class memoryabstractRecently, interest in incorporating Storage Class Memory (SCM), which blurs the distinction between memory and storage, into mainstream computing has been increasing rapidly. In this paper, we address the emerging questions regarding the use of SCM. Based on an embedded platform that employs FeRAM, a type of SCM, we present our findings. In summary, by introducing SCM, power efficiency improves while performance is degraded. We also show that such performance degradations may be removed with operating system level schemes that fully exploit the characteristics of SCM. Finally, we present permanent computing that supports lightweight system on/off capabilities by using SCM. Seungjae Baek, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Caching less for better performance: balancing cache size and update cost of flash memory cache in hybrid storage systems
Yongseok Oh, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
FAST | 2 |
| 2012 | Activity recognition in wide aerial video surveillance using entity relationship modelsabstractWe present the design and implementation of an activity recognition system in wide area aerial video surveillance using Entity Relationship Models (ERM). In this approach, finding an activity is equivalent to sending a query to a Relational DataBase Management System (RDBMS). By incorporating reference imagery and Geographic Information System (GIS) data, tracked objects can be associated with physical meanings, and several high levels of reasoning, such as traffic patterns or abnormal activity detection, can be performed. We demonstrate that different types of activities, with hierarchical structure, multiple actors, and context information, are effectively and efficiently defined and inferred using the ERM framework. We also show how visual tracks can be better interpreted as activities by using geo information. Experimental results on both real visual tracks and GPS traces validate our approach. Jongmoo Choi, Yann Dumortier, Jan Prokaj, Gérard G. Medioni |
SIGSPATIAL/GIS | 1 |
| 2012 | Real-time 3D face identification from a depth camera
Rui Min 0002, Jongmoo Choi, Gérard G. Medioni, Jean-Luc Dugelay |
ICPR | 2 |
| 2012 | Deduplication in SSDs: Model and quantitative analysisabstractIn NAND Flash-based SSDs, deduplication can provide an effective resolution of three critical issues: cell lifetime, write performance, and garbage collection overhead. However, deduplication at SSD device level distinguishes itself from the one at enterprise storage systems in many aspects, whose success lies in proper exploitation of underlying very limited hardware resources and workload characteristics of SSDs. In this paper, we develop a novel deduplication framework elaborately tailored for SSDs. We first mathematically develop an analytical model that enables us to calculate the minimum required duplication rate in order to achieve performance gain given deduplication overhead. Then, we explore a number of design choices for implementing deduplication components by hardware or software. As a result, we propose two acceleration techniques: sampling-based filtering and recency-based fingerprint management. The former selectively applies deduplication based upon sampling and the latter effectively exploits limited controller memory while maximizing the deduplication ratio. We prototype the proposed deduplication framework in three physical hardware platforms and investigate deduplication efficiency according to various CPU capabilities and hardware/software alternatives. Experimental results have shown that we achieve the duplication rate ranging from 4% to 51%, with an average of 17%, for the nine workloads considered in this work. The response time of a write request can be improved by up to 48% with an average of 15%, while the lifespan of SSDs is expected to increase up to 4.1 times with an average of 2.4 times. Choonghyun Lee, Sang Yup Lee, Ikjoon Son, Jongmoo Choi, Sungroh Yoon, Hu-ung Lee, Sooyong Kang, Youjip Won, Jaehyuk Cha |
MSST | 5 |
| 2012 | Real-time 3-D face tracking and modeling from awebcamabstractWe first infer a 3-D face model from a single frontal image using automatically extracted 2-D landmarks and deforming a generic 3-D model. Then, for any input image, we extract feature points and track them in 2-D. Given these correspondences, sometimes noisy and incorrect, we robustly estimate the 3-D head pose using PnP and a RANSAC process. As the head moves, we dynamically add new feature points to handle a large range of poses. When the tracker gets lost, due to motion blur or occlusions, the system re-initializes by matching feature points to the reference frontal image feature points. Our system runs in real-time (>;15Hz) on a standard CPU with a GPU card. We present results on stored video and will present a live demo, showing excellent tracking under large motion, fast movement, occlusion and facial expression variations. We also show comparative results with the ground truth BU head tracking dataset. Jongmoo Choi, Yann Dumortier, Sang-Il Choi, Muhammad Bilal Ahmad, Gérard G. Medioni |
WACV | 1 |
| 2011 | SSD Characterization: From Energy Consumption's Perspective
Balgeun Yoo, Youjip Won, Seokhei Cho, Sooyong Kang, Jongmoo Choi, Sungroh Yoon |
HotStorage | 5 |
| 2011 | An Empirical Study of Deploying Storage Class Memory into the I/O Path of Portable SystemsabstractWe explore the possibility of deploying storage class memory (SCM) into the I/O path as a file system metadata store and examine what effects it has on the performance of portable computing systems. In this regard, we develop a new flash memory-based file system that stores all metadata in SCM, while storing all file data in flash memory. In so doing, we make two contributions in this work. First, we present a model that analyzes the amount of SCM that is needed for specific flash memory storage capacity. Second, we present quantitative experimental results that show how much performance gains are possible by exploiting SCM in terms of I/O performance, energy efficiency and lifetime of the underlying flash memory. Compared to YAFFS, a popular flash memory-based file system, we show that system performance is improved by a maximum of around 320, 260, and 180% in terms of I/O performance, energy efficiency and lifetime of the Flash memory, respectively, for the realistic workloads that we considered. In Hwan Doh, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
Comput. J. | 2 |
| 2010 | Accurate 3D face reconstruction from weakly calibrated wide baseline images with profile contoursabstractWe propose a method to generate a highly accurate 3D face model from a set of wide-baseline images in a weakly calibrated setup. Our approach is purely data driven, and produces faithful 3D models without any pre-defined models, unlike other statistical model-based approaches. Our results do not rely upon a critical initialization step nor parameters for optimization steps. We process 5 images (including profile views), infer the accurate poses of cameras in all views, and then infer a dense 3D face model. The quality of 3D face models depends on the accuracy of estimated head-camera motion. First, we propose to use an iterative bundle adjustment approach to remove outliers in corresponding points. Contours in the profile views are matched to provide reliable correspondences that link two opposite side of views together. For dense reconstruction, we propose to use a face-specific cylindrical representation which allows us to solve a global optimization problem for N-view dense aggregation. Profile contours are used once again to provide constraints in the optimization step. Experimental results using synthetic and real images show that our method provides accurate and stable reconstruction results on wide-baseline images. We compare our method with state of the art methods, and show that it provides significantly better results in terms of both accuracy and efficiency. Yuping Lin, Gérard G. Medioni, Jongmoo Choi |
CVPR | 3 |
| 2010 | Janus-FTL: finding the optimal point on the spectrum between page and block mapping schemesabstractNAND flash memory based storage such as SSDs is gaining popularity in commodity computer systems. Some low-end SSDs use the block mapping FTL (Flash Translation Layer) that is good for sequential write patterns but poor for random ones. On the other hand, high-end SSDs tend to use the page mapping FTL that is effective for random write patterns, but whose performance degrades after successive random writes. Designing an FTL that adapts to various workload patterns and provides long-term stable performance is a challenging issue. To resolve this issue, we propose a new FTL, which we call Janus-FTL, that provides a spectrum between the block and page mapping schemes. By adapting along the spectrum, Janus-FTL can provide long-term superior write performance for various workload patterns. We also present a cost model of Janus-FTL that shows the existence of the optimal point on the spectrum for a given workload. Our experimental results show the superiority of Janus-FTL, which adapts itself along the spectrum for a given workload, over state-of-the-art hybrid mapping FTLs and the pure page mapping FTL Hunki Kwon, Eunsam Kim, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
EMSOFT | 3 |
| 2010 | 3D Face Reconstruction Using a Single or Multiple ViewsabstractWe present a 3D face reconstruction system that takes as input either one single view or several different views. Given a facial image, we first classify the facial pose into one of five predefined poses, then detect two anchor points that are then used to detect a set of predefined facial landmarks. Based on these initial steps, for a single view we apply a warping process using a generic 3D face model to build a 3D face. For multiple views, we apply sparse bundle adjustment to reconstruct 3D landmarks which are used to deform the generic 3D face model. Experimental results on the Color FERET and CMU multi-PIE databases confirm our framework is effective in creating realistic 3D face models that can be used in many computer vision applications, such as 3D face recognition at a distance. Jongmoo Choi, Gérard G. Medioni, Yuping Lin, Luciano Silva, Olga R. P. Bellon, Maurício Pamplona Segundo, Timothy C. Faltemier |
ICPR | 1 |
| 2010 | A performance model and file system space allocation scheme for SSDsabstractSolid State Drives (SSDs) are now becoming a part of main stream computers. Even though disk scheduling algorithms and file systems of today have been optimized to exploit the characteristics of hard drives, relatively little attention has been paid to model and exploit the characteristics of SSDs. In this paper, we consider the use of SSDs from the file system standpoint. To do so, we derive a performance model for the SSDs. Based on this model, we devise a file system space allocation scheme, which we call Greedy-Space, for block or hybrid mapping SSDs. From the Postmark benchmark results, we observe substantial performance improvements when employing the Greedy-Space scheme in ext3 and Reiser file systems running on three SSDs available in the market. Choulseung Hyun, Jongmoo Choi, Yongseok Oh, Donghee Lee 0001, Eunsam Kim, Sam H. Noh |
MSST | 2 |
| 2010 | Optimizations of LFS with slack space recycling and lazy indirect block updateabstractEven though the Log-structured File System (LFS) has elegant concept for superior write performance, it suffers from cleaning overhead. Specifically, when file system utilization is high and the system is busy, write performance of LFS degenerates significantly. Also, cascading update of meta-data triggered by modification of file data decreases LFS performance further. To overcome the performance drawbacks of LFS, we propose two schemes, namely Slack Space Recycling (SSR) and Lazy Indirect Block Update (LIBU). The SSR scheme writes modified data to invalid areas of used segments when on-demand cleaning is inevitable to serve incoming write requests. Also, the LIBU scheme accumulates meta-data update in memory beyond multiple segment writes without compromising consistency so as to decrease total amount of writes. From various experimental results, we observe significant performance improvements when employing the SSR and LIBU schemes for a wide utilization range. Yongseok Oh, Eunsam Kim, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
SYSTOR | 3 |
| 2009 | StaRSaC: Stable random sample consensus for parameter estimationabstractWe address the problem of parameter estimation in presence of both uncertainty and outlier noise. This is a common occurrence in computer vision: feature localization is performed with an inherent uncertainty which can be described as Gaussian, with unknown variance; feature matching in multiple images produces incorrect data points. RANSAC is the preferred method to reject outliers if the variance of the uncertainty noise is known, but fails otherwise, by producing either a tight fit to an incorrect solution, or by computing a solution which includes outliers. We thus propose a new estimator which enforces stability of the solution with respect to the uncertainty bound. We show that the variance of the estimated parameters (VoP) exhibits ranges of stability with respect to this bound. Within this range of stability, we can accurately segment the inliers, and estimate the parameters, the variance of the Gaussian noise. We show how to compute this stable range using RANSAC and a search. We validate our results by extensive tests and comparison with state of the art estimators on both synthetic and real data sets. These include line fitting, homography estimation, and fundamental matrix estimation. The proposed method outperforms all others. Jongmoo Choi, Gérard G. Medioni |
CVPR | 1 |
| 2009 | Disk schedulers for solid state driversabstractIn embedded systems and laptops, flash memory storage such as SSDs (Solid State Drive) have been gaining popularity due to its low energy consumption and durability. As SSDs are flash memory based devices, their performance behavior differs from those of magnetic disks. However, little attention has been paid on how to exploit SSDs from the disk scheduling algorithm view point. In this paper, we first describe behaviors of SSDs that inspires us to design a new disk scheduler for the Linux operating system. Specifically, read service time is almost constant in an SSD while write service time is not. Moreover, appropriate grouping of write requests eliminates any ordering-related restrictions and also maximizes write performance. From these observations, we propose two disk schedulers: IRBW-FIFO and IRBW-FIFO-RP. Both schedulers arrange write requests into bundles of an appropriate size while read requests are independently scheduled. Then, the IRBW-FIFO scheduler provides complete FIFO ordering to each bundle of write requests and each individual read requests while the IRBW-FIFO-RP scheduler gives higher priority to read requests than the bundles of write requests. We implement these schedulers in Linux 2.6.23, and results of executing our set of benchmark programs shows that performance improvements of up to 17% compared to existing Linux disk schedulers are achieved. Yongseok Oh, Eunsam Kim, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
EMSOFT | 4 |
| 2009 | Identifying Noncooperative Subjects at a Distance Using Face Images and Inferred Three-Dimensional Face ModelsabstractWe present an approach to identify noncooperative individuals at a distance from a sequence of images, using 3-D face models. Most biometric features (such as fingerprints, hand shape, iris, or retinal scans) require cooperative subjects in close proximity to the biometric system. We process images acquired with an ultrahigh-resolution video camera, infer the location of the subjects' head, use this information to crop the region of interest, build a 3-D face model, and use this 3-D model to perform biometric identification. To build the 3-D model, we use an image sequence, as natural head and body motion provides enough viewpoint variation to perform stereomotion for 3-D face reconstruction. We have conducted experiments on a 2-D and 3-D databases collected in our laboratory. First, we found that metric 3-D face models can be used for recognition by using simple scaling method even though there is no exact scale in the 3-D reconstruction. Second, experiments using a commercial 3-D matching engine suggest the feasibility of the proposed approach for recognition against 3-D galleries at a distance (3, 6, and 9 m). Moreover, we show initial 3-D face modeling results on various factors including head motion, outdoor lighting conditions, and glasses. The evaluation results suggest that video data alone, at a distance of 3 to 9 meters, can provide a 3-D face shape that supports successful face recognition. The performance of 3-D-3-D recognition with the currently generated models does not quite match that of 2-D-2-D. We attribute this to the quality of the inferred models, and this suggests a clear path for future research. Gérard G. Medioni, Jongmoo Choi, Cheng-Hao Kuo, Douglas Fidaleo |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2008 | LTFTL: lightweight time-shift flash translation layer for flash memory based embedded storageabstractFlash memory storage has been widely used in various embedded systems such as digital cameras, MP3 players, cellular phones, and DMB devices and now it applies to PCs as a form of SSDs. Characteristics of Flash memory necessitate a software layer called FTL (Flash Translation Layer) that directs modified data to new places in Flash memory and maintains a mapping between a logical sector number to a physical page. We notice that this out-of-place update scheme of the FTL allows a low-overhead time-shifting between multiple versions of storage state. From this observation, we propose LTFTL (Lightweight Time-shift FTL) that provides not only multiple versions of storage state but also an open-ended interface to traverse them. This open-ended interface can be used to support fault-resilience schemes, transactions of various granularities, and user-friendly roll-back services. Experimental results from a prototype implementation show that the proposed LTFTL can (1) provide a low-overhead time-shift capability at the user level by maintaining multiple storage states and (2) enhance the reliability/survivability of Flash memory by allowing to roll back to a previous consistent storage state at the storage system level. Kyoungmoon Sun, Seungjae Baek, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh, Sang Lyul Min |
EMSOFT | 3 |
| 2007 | Uniformity improving page allocation for flash memory file systemsabstractFlash memory is a storage medium that is becoming more and more popular. Though not yet fully embraced in traditional computing systems, Flash memory is prevalent in embedded systems, materialized as commodity appliances such as the digital camera and the MP3 player that we enjoy in our everyday lives. This paper considers an issue in file systems that use Flash memory as a storage medium and makes the following two contributions. First, we identify the cost of block cleaning as the key performance bottleneck for Flash memory analogous to the seek time in disk storage. We derive and define three performance parameters, namely, utilization, invalidity, and uniformity, from characteristics of Flash memory and present a formula for block cleaning cost based on these parameters. We show that, of these parameters, uniformity most strongly influences the cost of cleaning and that uniformity is a file system controllable parameter. This leads us to our second contribution, designing the modification-aware (MODA) page allocation scheme and analyzing how enhanced uniformity affects the block cleaning cost with various workloads. Real implementation experiments conducted on an embedded system show that the MODA scheme typically improves 20 to 30% in cleaning time compared to the traditional sequential allocation scheme that is used in YAFFS. Seungjae Baek, Seongjun Ahn, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
EMSOFT | 3 |
| 2007 | Exploiting non-volatile RAM to enhance flash file system performanceabstractNon-volatile RAM (NVRAM) such as PRAM (Phase-change RAM), FeRAM (Ferroelectric RAM), and MRAM (Magnetoresistive RAM) has characteristics of both non-volatile storage and random access memory (RAM). These forms of NVRAM are currently being developed by major semiconductor companies and are expected to be an everyday component in the near future. The advent of NVRAM may possibly bring about drastic changes to the system software landscape. In this work, we develop a new Flash memory based file system that exploits NVRAM in order to improve system performance. Specifically, we discuss the initial design and implementation of a file system that stores all metadata in NVRAM, while storing all file data in Flash memory. In so doing, we make two contributions in this work. First, we present a model that analyzes the amount of NVRAM that is needed for specific Flash memory storage capacity. Experimentally, we verify that this model represents the exact NVRAM usage in the realistic environment. Second, we present quantitative experimental results that show how much performance gains are possible by exploiting NVRAM. Compared to YAFFS, a popular Flash memory based file system, we show that this file system requires only minimal time for mounting and that the execution time improves by a maximum of 600% and an average of 437% for the realistic workloads that we considered. In Hwan Doh, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
EMSOFT | 2 |
| 2007 | Block recycling schemes and their cost-based optimization in nand flash memory based storage systemabstractFlash memory has many merits such as light weight, shock resistance, and low power consumption, but also has limitations like the erase-before-write property. To overcome such limitations and to use it efficiently as storage media in mobile systems, Flash memory based storage systems require special address mapping software called the FTL (Flash-memory Translation Layer). Like cleaning in Log-structured file system (LFS), the FTL often performs a merge operation for block recycling and its efficiency affects the performance of the storage system. To reduce the block recycling costs in NAND Flash memory based storage, we introduce another block recycling scheme that we call migration. Our cost-models and experimental results show that cost-based selection of merge or migration for each block recycling can decrease block recycling costs and, therefore, improve performance of Flash memory based storage systems. Also, we derive the macroscopic optimal migration/merge sequence minimizing block recycling costs for each migration/merge combination period. Experimental results show that the performance of Flash memory based storage can be further improved by the macroscopic optimization than the simple cost-based selection. Hunki Kwon, Choulseung Hyun, Seongjun Ahn, Jongmoo Choi, Donghee Lee 0001, Sam H. Noh |
EMSOFT | 6 |
| 2006 | A Real-time 3D IR Camera based on Hierarchical Orthogonal CodingabstractWe present a real-time 3D camera based on IR (infrared) structured light suitable for robots working in home environment. First, we implemented a HOC (hierarchical orthogonal coding) based FPGA board. The HOC gives robust depth images because the signal separation coding provides not only the separation of overlapped codes, but also a robust decision on pixel correspondence with error correction. The FPGA module can handle high computational cost of HOC based signal separation coding. Second, we implemented a compact optic system of the camera to project and receive IR structured light. The invisible IR pattern light provides users inconvenient in the home environment. Various objects and workspaces which has continuous and/or non-continuous surface are tested and, sensitivity to illumination change and processing time are analyzed. The experiment results show the robust performance for surface smoothness, color, and materials of objects used in home environment. The proposed approach opens a greater feasibility of applying structured light based depth imaging to a 3D modeling of cluttered workspace for home service robots Sukhan Lee 0001, Jongmoo Choi, Seungsub Oh, Jaehyuk Ryu, Jungrae Park |
ICRA | 2 |
| 2006 | A Real-Time Wall Detection Method for Indoor EnvironmentsabstractThis paper presents an effective and real-time approach for detecting walls in indoor environment. This approach relies on the fact that the rear of the opaque walls is not visible. Thus, to detect the walls in an indoor environment a set of hypothetical walls, based on the ceiling edges or ground level edges, are considered; and their validity is checked using point cloud, generated by a sensor. A certainty factor is calculated for each detected wall, which is updated continuously based on the newly gathered sensory information. Furthermore, the certainty of the walls can be updated using other source of information for better and more reliable wall detection. The novelty of this approach is in its capability to handle environments, with texture-less walls, in real-time. The algorithm has been implemented in simulation, and tested in real environment and has shown effective, reliable and real-time performance Hadi Moradi, Jongmoo Choi, Sukhan Lee 0001 |
IROS | 2 |
| 2006 | Caller Identification Based on Cognitive Robotic EngineabstractAn approach to identifying a caller or callers by a service robot is presented for a natural interaction with people in a home/office environment. The problem addressed specifically in this paper is how to successfully identify a caller in a cluttered environment with large uncertainties involved in the sensed audio-visual cues. The proposed approach is based on a proposition that the dependability of perceptual recognition may come unlikely from "the effort to make individual sensing perfect", but likely from "the effort to self-generate perceptual behaviors of integrating individual sensing that lead to mission accomplishment, no matter how imperfect and uncertain individual sensing may be". We implement the above proposition in terms of a novel robotic architecture, referred to here as "cognitive robotic engine (CRE)." CRE implemented for the case of a robot identifying a caller in a crowded and noisy environment, including its experimental results, are shown Sukhan Lee 0001, Hun-Sue Lee, Seungmin Baek, Jongmoo Choi, ByoungYoul Song, Young-Jo Cho |
RO-MAN | 4 |
| 2005 | Low-Dimensional Facial Image Representation Using FLD and MDS
Jongmoo Choi, Juneho Yi |
ICIC (1) | 1 |
| 2005 | A 3D IR Camera with Variable Structured Light for Home Service RobotsabstractThere has shown a significant interest in a high performance of, at the same time, a compact size and low cost of, 3D sensor, in reflection of a growing need of 3D environmental sensing for service robotics. One of the important requirements associated with such a 3D sensor is that sensing does not irritate or disturb human in any way while working in close and continuous contact with human. Furthermore, such a 3D sensor should be reliable and robust to the change of environmental illumination as service robots are required to work day and night. This paper presents a 3D IR camera with variable structured light that is human friendly and robust enough for application to home service robots. Infrared is chosen as the sensing medium in order to meet the requirement of human friendliness and robustness to illumination change. A Digital Mirror Device (DMD) is employed to generate and project variable patterns at a high speed for real-time operation. In implementation, we emphasize the integration of modular components to support real-time sensing and compactness in size. A number of real-world experimentations are conducted, including a human face, a statue, and a plastic model. The experimental results have demonstrated that the implemented 3D IR Camera is robust to illumination change, in addition to its advantage of human friendliness. Sukhan Lee 0001, Jongmoo Choi, Seungmin Baek, Byungchan Jung, Changsik Choi, Hunmo Kim, Jeongtaek Oh, Seungsub Oh, Jaekeun Na |
ICRA | 2 |
| 2005 | Signal Separation Coding for Robust Depth Imaging Based on Structured LightabstractThis paper presents an original approach to coding the light patterns for robust depth imaging based on structured light. We have discovered that the degradation of precision and robustness, seen in most conventional approaches to structured light, comes mainly from the overlapping of multiple codes in the signal received at a camera pixel, where the overlapped codes are from the neighbouring and/or, even, distant pixels of the projecting mirror array. Considering the criticality of separating the overlapped codes to precision and robustness, we propose a novel signal separation code, referred to here as “Hierarchical Orthogonal Code (HOC),” for depth imaging. HOC provides not only the separation of overlapped codes, but also a robust decision on pixel correspondence with error correction based on a contextual likelihood among the sets of separated codes from neighbouring camera pixels. The experimental results have shown that the proposed HOC significantly enhances the robustness and precision in depth imaging, compared to the best known conventional approaches. The proposed approach opens a greater feasibility of applying structured light based depth imaging to a 3D modelling of cluttered workspace for home service robots. Sukhan Lee 0001, Jongmoo Choi, Jaekeun Na, Seungsub Oh |
ICRA | 2 |
| 2005 | Effective Representation Using ICA for Face Recognition Robust to Local Distortion and Partial OcclusionabstractThe performance of face recognition methods using subspace projection is directly related to the characteristics of their basis images, especially in the cases of local distortion or partial occlusion. In order for a subspace projection method to be robust to local distortion and partial occlusion, the basis images generated by the method should exhibit a part-based local representation. We propose an effective part-based local representation method named locally salient ICA (LS-ICA) method for face recognition that is robust to local distortion and partial occlusion. The LS-ICA method only employs locally salient information from important facial parts in order to maximize the benefit of applying the idea of "recognition by parts." It creates part-based local basis images by imposing additional localization constraint in the process of computing ICA architecture I basis images. We have contrasted the LS-ICA method with other part-based representations such as LNMF (Localized Nonnegative Matrix Factorization) and LFA (Local Feature Analysis). Experimental results show that the LS-ICA method performs better than PCA, ICA architecture I, ICA architecture II, LFA, and LNMF methods, especially in the cases of partial occlusions and local distortions. Jongsun Kim, Jongmoo Choi, Juneho Yi, Matthew Turk 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Design, Implementation, and Performance Evaluation of a Detection-Based Adaptive Block Replacement SchemeabstractA new buffer replacement scheme, called DEAR (detection-based adaptive replacement), is presented for effective caching of disk blocks in the operating system. The proposed DEAR scheme automatically detects block reference patterns of applications and applies different replacement policies to different applications depending on the detected reference pattern. The detection is made by a periodic process and is based on the relationship between block attribute values, such as backward distance and frequency gathered in a period, and the forward distance observed in the next period. This paper also describes an implementation and performance measurement of the DEAR scheme in FreeBSD. The results from performance measurements of several real applications show that, compared with the LRU scheme, the proposed scheme reduces the number of disk I/Os by up to 51 percent, and the response time by up to 35 percent in the case of single application executions. For multiple application executions, the results show that the proposed scheme reduces the number of disk I/Os by up to 20 percent and the overall response time by up to 18 percent. Jongmoo Choi, Sam H. Noh, Sang Lyul Min, Eun-Yong Ha, Yookun Cho |
IEEE Trans. Computers | 1 |
| 2001 | LRFU: A Spectrum of Policies that Subsumes the Least Recently Used and Least Frequently Used PoliciesabstractEfficient and effective buffering of disk blocks in main memory is critical for better file system performance due to a wide speed gap between main memory and hard disks. In such a buffering system, one of the most important design decisions is the block replacement policy that determines which disk block to replace when the buffer is full, In this paper, we show that there exists a spectrum of block replacement policies that subsumes the two seemingly unrelated and independent Least Recently Used (LRU) and Least Frequently Used (LFU) policies. The spectrum is called the LRFU (Least Recently/Frequently Used) policy and is formed by how much more weight we give to the recent history than to the older history. We also show that there is a spectrum of implementations of the LRFU that again subsumes the LRU and LFU implementations. This spectrum is again dictated by how much weight is given to recent and older histories and the time complexity of the implementations lies between O(1) (the time complexity of LRU) and O(log(2) n) (the time complexity of LFU), where n is the number of blocks in the buffer, Experimental results from trace-driven simulations show that the performance of the LRFU is at least competitive with that of previously known policies for the workloads we considered. Donghee Lee 0001, Jongmoo Choi, Jong-Hun Kim, Sam H. Noh, Sang Lyul Min, Yookun Cho, Chong-Sang Kim |
IEEE Trans. Computers | 2 |
| 2000 | A Low-Overhead, High-Performance Unified Buffer Management Scheme That Exploits Sequential and Looping References
Jongmoo Choi, Jesung Kim, Sam H. Noh, Sang Lyul Min, Yookun Cho, Chong-Sang Kim |
OSDI | 2 |
| 2000 | Towards application/file-level characterization of block references: a case for fine-grained buffer managementabstractTwo contributions are made in this paper. First, we show that system level characterization of file block references is inadequate for maximizing buffer cache performance. We show that a finer-grained characterization approach is needed. Though application level characterization methods have been proposed, this is the first attempt, to the best of our knowledge, to consider file level characterizations. We propose an Application/File-level Characterization (AFC) scheme where we detect on-line the reference characteristics at the application level and then at the file level, if necessary. The results of this characterization are used to employ appropriate replacement policies in the buffer cache to maximize performance. The second contribution is in proposing an efficient and fair buffer allocation scheme. Application or file level resource management is infeasible unless there exists an allocation scheme that is efficient and fair. We propose the ΔHIT allocation scheme that takes away a block from the application/file where the removal results in the smallest reduction in the number of expected buffer cache hits. Both the AFC and ΔHIT schemes are on-line schemes that detect and allocate as applications execute. Experiments using trace-driven simulations show that substantial performance improvements can be made. For single application executions the hit ratio increased an average of 13 percentage points compared to the LRU policy, with a maximum increase of 59 percentage points, while for multiple application executions, the increase is an average of 12 percentage points, with a maximum of 32 percentage points for the workloads considered. Jongmoo Choi, Sam H. Noh, Sang Lyul Min, Yookun Cho |
SIGMETRICS | 1 |
| 1999 | On the Existence of a Spectrum of Policies that Subsumes the Least Recently Used (LRU) and Least Frequently Used (LFU) PoliciesabstractAbstractÐEfficient and effective buffering of disk blocks in main memory is critical for better file system performance due to a wide speed gap between main memory and hard disks. In such a buffering system, one of the most important design decisions is the block replacement policy that determines which disk block to replace when the buffer is full. In this paper, we show that there exists a spectrum of block replacement policies that subsumes the two seemingly unrelated and independent Least Recently Used (LRU) and Least Frequently Used (LFU) policies. The spectrum is called the LRFU (Least Recently/Frequently Used) policy and is formed by how much more weight we give to the recent history than to the older history. We also show that there is a spectrum of implementations of the LRFU that again subsumes the LRU and LFU implementations. This spectrum is again dictated by how much weight is given to recent and older histories and the time complexity of the implementations lies between O(1) (the time complexity of LRU) and O…log 2 n† (the time complexity of LFU), where n is the number of blocks in the buffer. Experimental results from trace-driven simulations show that the performance of the LRFU is at least competitive with that of previously known policies for the workloads we considered. Index TermsÐBuffer cache, LFU, LRU, replacement policy, trace-driven simulation. 1 Donghee Lee 0001, Jongmoo Choi, Jong-Hun Kim, Sam H. Noh, Sang Lyul Min, Yookun Cho, Chong-Sang Kim |
SIGMETRICS | 2 |
| 1999 | An Implementation Study of a Detection-Based Adaptive Block Replacement Scheme
Jongmoo Choi, Sam H. Noh, Sang Lyul Min, Yookun Cho |
USENIX ATC, General Track | 1 |