EDBT 2026 Demo / reviewers in the wild / expert
Jin-Soo Kim 0001
dblp:45/1315-1 · also Jinsoo Kim 0001
· DBLP profile ↗
67ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-1065-9362ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 54 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6Software engineering, systems software and programming languages · 5 · 1 since 2021Databases, data management, data science and information retrieval · 3Theory of computation · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZRAID: Leveraging Zone Random Write Area (ZRWA) for Alleviating Partial Parity Tax in ZNS RAIDabstractThe Zoned Namespace (ZNS) SSD is an innovative technology that aims to mitigate the block interface tax associated with conventional SSDs. However, constructing a RAID system using ZNS SSDs presents a significant challenge in managing partial parity for incomplete stripes. Previous research permanently logs partial parity in a limited number of reserved zones, which not only creates bottlenecks in throughput but also exacerbates write amplification, thereby reducing the device's lifetime. We refer to these inefficiencies as the partial parity tax. Minwook Kim, Seongyeop Jeong, Jin-Soo Kim 0001 |
ASPLOS (1) | 3 |
| 2023 | An Efficient Order-Preserving Recovery for F2FS with ZNS SSDabstractStorage devices use write buffers to improve performance, where multiple write requests are processed in parallel and completed in a random order. This may result in data loss in the event of a sudden failure. Therefore, Linux filesystems provide the fsync() system call to prevent data loss and ensure write order. However, the fsync() system call in F2FS, one of the most popular filesystems, is inefficient and insufficient for guaranteeing data consistency. Euidong Lee, Ikjoon Son, Jin-Soo Kim 0001 |
HotStorage | 3 |
| 2023 | Empowering Storage Systems Research with NVMeVirt: A Comprehensive NVMe Device EmulatorabstractThere have been drastic changes in the storage device landscape recently. At the center of the diverse storage landscape lies the NVMe interface, which allows high-performance and flexible communication models required by these next-generation device types. However, its hardware-oriented definition and specification are bottlenecking the development and evaluation cycle for new revolutionary storage devices. Furthermore, existing emulators lack the capability to support the advanced storage configurations that are currently in the spotlight. In this article, we present NVMeVirt, a novel approach to facilitate software-defined NVMe devices. A user can define any NVMe device type with custom features, and NVMeVirt allows it to bridge the gap between the host I/O stack and the virtual NVMe device in software. We demonstrate the advantages and features of NVMeVirt by realizing various storage types and configurations, such as conventional SSDs, low-latency high-bandwidth NVM SSDs, zoned namespace SSDs, and key-value SSDs with the support of PCI peer-to-peer DMA and NVMe-oF target offloading. We also make cases for storage research with NVMeVirt, such as studying the performance characteristics of database engines and extending the NVMe specification for the improved key-value SSD performance. Sang-Hoon Kim, Jaehoon Shim, Euidong Lee, Seong-Yeob Jeong, Ilkueon Kang, Jin-Soo Kim 0001 |
ACM Trans. Storage | 6 |
| 2020 | Introduction to the Special Section on Computational StorageabstractNo abstract available. Jin-Soo Kim 0001, Yang-Seok Ki, Erik Riedel |
ACM Trans. Storage | 1 |
| 2017 | Application-Aware Swapping for Mobile SystemsabstractThere has been a constant demand for memory in modern mobile systems to provide users with better experience. Swapping is one of the cost-effective software solutions to provide extra usable memory by reclaiming inactive pages and improving memory utilization. However, swapping has not been actively adopted to mobile systems since it incurs a significant amount of I/O, which in fact impairs system performance as well as user experience. In this paper, we propose a novel scheme to properly harness the swapping to mobile systems. We identify that a vast amount of I/O for swapping comes from the conflict of the traditional page-level approach of the swapping and the process-level memory management scheme tailored to mobile systems. Moreover, we find out that the current victim page selection policy is not effective due to the process-level policy. To address these problems, we revise the victim selection policy to resolve the conflict and to selectively perform swapping according to the efficacy of swapping. Evaluation using a running prototype with realistic workloads indicates that the propose scheme effectively reduces the paging traffic, thereby improving user experience as well as energy consumption. Sang-Hoon Kim, Jinkyu Jeong, Jin-Soo Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | GCMix: An Efficient Data Protection Scheme against the Paired Page InterferenceabstractIn multi-level cell (MLC) NAND flash memory, two logical pages are overlapped on a single physical page. Even after a logical page is programmed, the data can be corrupted if the programming of the coexisting logical page is interrupted. This phenomenon is called paired page interference. This article proposes a novel software technique to deal with the paired page interference without any additional hardware or extra page write. The proposed technique utilizes valid pages in the victim block during garbage collection (GC) as the backup against the interference, and pairs them with incoming pages written by the host. This approach eliminates undesirable page copy to backup pages against the interference. However, such a strategy has an adverse effect on the hot/cold separation policy, which is essential to improve the efficiency of GC. To limit the downside, we devise a metric to estimate the benefit of GCMix on-the-fly so that GCMix can be adaptively utilized only when the benefit outweighs the overhead. Evaluations using synthetic and real workloads show GCMix can effectively deal with the paired page interference, reducing the write amplification factor by up to 17.5%compared to the traditional technique, while providing comparable I/O performance. Sang-Hoon Kim, Jinhyuk Lee, Jin-Soo Kim 0001 |
ACM Trans. Storage | 3 |
| 2016 | NVMeDirect: A User-space I/O Framework for Application-specific Optimization on NVMe SSDs
Young-Sik Lee, Jin-Soo Kim 0001 |
HotStorage | 3 |
| 2016 | ActiveSort: Efficient external sorting using active SSDs in the MapReduce framework
Young-Sik Lee, Luis Cavazos Quero, Sang-Hoon Kim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 4 |
| 2016 | ForestDB: A Fast Key-Value Storage System for Variable-Length String KeysabstractIndexing key-value data on persistent storage is an important factor for NoSQL databases. Most key-value storage engines use tree-like structures for data indexing, but their performance and space overhead rapidly get worse as the key length becomes longer. This also affects the merge or compaction cost which is critical to the overall throughput. In this paper, we present ForestDB, a key-value storage engine for a single node of large-scale NoSQL databases. ForestDB uses a new hybrid indexing scheme called HB+-trie, which is a disk-based trie-like structure combined with B+-trees. It allows for efficient indexing and retrieval of arbitrary length string keys with relatively low disk accesses over tree-like structures, even though the keys are very long and randomly distributed in the key space. Our evaluation results show that ForestDB significantly outperforms the current key-value storage engine of Couchbase Server [1], LevelDB [2], and RocksDB [3], in terms of both the number of operations per second and the amount of disk writes per update operation. Jung-Sang Ahn, Chiyoung Seo, Ravi Mayuram, Rahim Yaseen, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 5 |
| 2016 | SmartLMK: A Memory Reclamation Scheme for Improving User-Perceived App Launch TimeabstractAs the mobile computing environment evolves, users demand high-quality apps and better user experience. Consequently, memory demand in mobile devices has soared. Device manufacturers have fulfilled the demand by equipping devices with more RAM. However, such a hardware approach is only a temporary solution and does not scale well in the resource-constrained mobile environment. Meanwhile, mobile systems adopt a new app life cycle and a memory reclamation scheme tailored for the life cycle. When a user leaves an app, the app is not terminated but cached in memory as long as there is enough free memory. If the free memory gets low, a victim app is terminated and the associated memory to the app is reclaimed. This process-level approach has worked well in the mobile environment. However, user experience can be impaired severely because the victim selection policy does not consider the user experience. In this article, we propose a novel memory reclamation scheme called SmartLMK . SmartLMK minimizes the impact of the process-level reclamation on user experience. The worthiness to keep an app in memory is modeled by means of user-perceived app launch time and app usage statistics. The memory footprint and impending memory demand are estimated from the history of the memory usage. Using these values and memory models, SmartLMK picks up the least valuable apps and terminates them at once. Our evaluation on a real Android-based smartphone shows that SmartLMK efficiently distinguishes the valuable apps among cached apps and keeps those valuable apps in memory. As a result, the user-perceived app launch time can be improved by up to 13.2%. Sang-Hoon Kim, Jinkyu Jeong, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Subpage programming for extending the lifetime of NAND flash memory
Jung Hoon Kim 0004, Sang-Hoon Kim, Jin-Soo Kim 0001 |
DATE | 3 |
| 2015 | Managing gpu buffers for caching more apps in mobile systemsabstractModern mobile systems cache apps actively to quickly respond to a user’s call to launch apps. Since the amount of usable memory is critical to the number of cacheable apps, it is important to maximize memory utilization. Meanwhile, modern mobile apps make use of graphics processingunits (GPUs) to accelerate their graphic operations and to provide better user experience. In resource-constrained mobile systems, GPU cannot afford its private memory but shares the main memory with CPU. It leads to a considerable amount of main memory to be allocated for GPU buffers which are used for processingGPU operations. These GPU buffers are, however, not managed effectively so that inactive GPU buffers occupy a large fraction of the memory and decrease memory utilization. This paper proposes a scheme to manage GPU buffers to increase the memory utilization in mobile systems. Our scheme identifies inactive GPU buffers by exploitingthe state of an app from a user’s perspective, and reduces their memory footprint by compressingthem. Our sophisticated design approach prevents GPU-specific issues from causing an unpleasant overhead. Our evaluation on a runningprototype with realistic workloads shows that the proposed scheme can secure up to 215.9 MB of extra memory from 1.5 GB of main memory and increase the average number of cached apps by up to 31.3%. Sejun Kwon, Sang-Hoon Kim, Jin-Soo Kim 0001, Jinkyu Jeong |
EMSOFT | 3 |
| 2015 | Boosting Quasi-Asynchronous I/O for Better Responsiveness in Mobile Devices
Daeho Jeong, Youngjae Lee, Jin-Soo Kim 0001 |
FAST | 3 |
| 2015 | Controlling physical memory fragmentation in mobile systemsabstractSince the adoption of hardware-accelerated features (e.g., hardware codec) improves the performance and quality of mobile devices, it revives the need for contiguous memory allocation. However, physical memory in mobile systems is highly fragmented due to the frequent spawn and exit of processes and the lack of proactive anti-fragmentation scheme. As a result, the memory allocation for large and contiguous I/O buffers suffer from the highly fragmented memory, thereby incurring high CPU usage and power consumption. This paper presents a proactive anti-fragmentation approach that groups pages with the same lifetime, and stores them contiguously in fixed-size contiguous regions. When a process is killed to secure free memory, a set of contiguous regions are freed and subsequent contiguous memory allocations can be easily satisfied without incurring additional overhead. Our prototype implementation on a Nexus 10 tablet with the Android kernel shows that the proposed scheme greatly alleviates fragmentation, thereby reducing the I/O buffer allocation time, associated CPU usage, and energy consumption. Sang-Hoon Kim, Sejun Kwon, Jin-Soo Kim 0001, Jinkyu Jeong |
ISMM | 3 |
| 2015 | Self-sorting SSD: Producing sorted data inside active SSDsabstractNowadays solid state drives (SSDs) are gaining popularity and are replacing magnetic hard disk drives (HDDs) in enterprise storage systems. As a result, extracting the maximum performance from SSDs is becoming crucial to deal with the increasing storage volume and performance needs. Active disks were introduced as a way to offload data-processing tasks from the host into disks freeing system resources and achieving better performance. In this work, we present an active SSD architecture called Self-Sorting SSD that targets to offload sorting operations which are commonly used in data-intensive and database environments and that require heavy data transfer. Processing sorting operations directly on the SSD reduces data transfer from/to the storage devices, increasing system performance and the lifetime of SSDs. Experiments on a real SSD platform reveal that our proposed architecture outperforms traditional external merge sort by up to 60.75%, reduces energy consumption by up to 58.86%, and eliminates all the data transfer overhead to compute sorted results. Luis Cavazos Quero, Young-Sik Lee, Jin-Soo Kim 0001 |
MSST | 3 |
| 2015 | A High-Performance Media Streaming Architecture Based on KVMabstractA media streaming server can be implemented on a virtual machine for the ease of resource management. However, simply running a media streaming server on a virtual machine has two problems, the duplicate data in file caches of virtual machines and the performance degradation caused by the virtualization overhead. In order to resolve these problems, this paper proposes a high-performance media streaming architecture based on KVM. First, we implement a shared cache among virtual machines in order to eliminate the duplicate cached data. Second, the send file operation is offloaded to the hypervisor to reduce the virtualization overhead in I/O operations. Our evaluations with D-DASH datasets show that the performance of a media streaming server in the proposed architecture is increased by up to 30% as compared to that of the conventional media streaming server that simply runs on a virtual machine. Woo-Yeong Jeong, Youngjae Lee, Jin-Soo Kim 0001 |
PDP | 3 |
| 2015 | Zombie Chasing: Efficient Flash Management Considering Dirty Data in the Buffer CacheabstractThis paper presents a novel technique, called Zombie Chasing, for efficient flash management in solid state drives (SSDs). Due to the unique characteristics of NAND flash memory, SSDs need to accurately understand the liveness of the data stored in themselves. Recently, the TRIM command has been introduced to notify SSDs of dead data caused by file deletions, which otherwise could not be tracked by SSDs. This paper goes one step further and proposes a new liveness state, called the zombie state, to denote live data that will be dead shortly due to the corresponding dirty data in the buffer cache. We also devise new zombie-aware garbage collection algorithms which utilize the information about such zombie data inside SSDs. To evaluate Zombie Chasing, we implement zombie-aware garbage collection algorithms in the prototype SSD and modify the Linux kernel and the Oracle DBMS to deliver the information on the zombie data to the prototype SSD. Through comprehensive evaluations using our in-house micro-benchmark and the TPC-C benchmark, we observe that Zombie Chasing improves SSD performance effectively by reducing garbage collection overhead. Especially, our evaluation with the TPC-C benchmark on the Oracle DBMS shows that Zombie Chasing enhances the Transactions Per Second (TPS) value by up to 22% with negligible overhead. Youngjae Lee, Jin-Soo Kim 0001, Sang-Won Lee 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 2 |
| 2014 | Accelerating External Sorting via On-the-fly Data Merge in Active SSDs
Young-Sik Lee, Luis Cavazos Quero, Youngjae Lee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
HotStorage | 4 |
| 2014 | Large-scale incremental processing with MapReduce
DaeWoo Lee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 2 |
| 2014 | Memory efficient and scalable address mapping for flash storage devices
Young-Kyoon Suh, Bongki Moon, Alon Efrat, Jin-Soo Kim 0001, Sang-Won Lee 0001 |
J. Syst. Archit. | 4 |
| 2014 | System-Wide Cooperative Optimization for NAND Flash-Based Mobile SystemsabstractNAND flash memory has become an essential storage medium for various mobile devices, but it has some idiosyncrasies, such as out-of-place updates and bulk erase operations, which impair the I/O performance of those devices. In particular, the random write performance is strongly influenced by the overhead of a Flash Translation Layer (FTL) that hides the idiosyncrasies of NAND flash memory. To reduce the FTL overhead, operating systems need to be adapted for FTL, but widely used mobile operating systems still mainly adopt algorithms designed for traditional hard disk drives. Although there have been recent studies on rearranging write patterns into a sequential form in the operating system, these approaches fail to produce sequential write patterns under complicated workloads, and FTL still suffers from significant garbage collection overhead. If the operating system can be made aware of the write patterns that FTL requires, the overhead can be alleviated even under random write workloads. In this paper, we propose a system-wide cooperative optimization scheme, where the operating system communicates with the underlying FTL and generates write patterns that FTL can exploit to reduce the overhead. The proposed scheme was implemented on a real mobile device, and the experimental results show that the proposed scheme constantly improves performance under diverse workloads. Hyotaek Shim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 2 |
| 2013 | OSSD: A case for object-based solid state drivesabstractThe notion of object-based storage devices (OSDs) has been proposed to overcome the limitations of the traditional block-level interface which hinders the development of intelligent storage devices. The main idea of OSD is to virtualize the physical storage into a pool of objects and offload the burden of space management into the storage device. We explore the possibility of adopting this idea for solid state drives (SSDs). The proposed object-based SSDs (OSSDs) allow more efficient management of the underlying flash storage, by utilizing object-aware data placement, hot/cold data separation, and QoS support for prioritized objects. We propose the software stack of OSSDs and implement an OSSD prototype using an iSCSI-based embedded storage device. Our evaluations with various scenarios show the potential benefits of the OSSD architecture. Young-Sik Lee, Sang-Hoon Kim, Jin-Soo Kim 0001, Jaesoo Lee, Chanik Park, Seung Ryoul Maeng |
MSST | 3 |
| 2013 | An empirical study of hot/cold data separation policies in solid state drives (SSDs)abstractSeparating hot data from cold data is known to allow for efficient management of NAND flash memory in Solid State Drives (SSDs). However, most of previous work has been evaluated with the trace-driven simulations under different workloads and testing conditions. The goal of this paper is to empirically study the performance, computation overhead, and memory consumption of the existing hot/cold data separation policies on a real SSD platform. After devising a general framework where a different policy can be easily plugged in, we have evaluated three hot/cold data separation policies: 2-level LRU (LRU), Multiple Bloom Filter (MBF), and Dynamic dAta Clustering (DAC). Our evaluation results show that DAC performs best, improving the performance by up to 58% in real workloads with a reasonable computation and memory overhead. Jongsung Lee 0001, Jin-Soo Kim 0001 |
SYSTOR | 2 |
| 2013 | μ*-Tree: An Ordered Index Structure for NAND Flash Memory with Adaptive Page Layout SchemeabstractAs NAND flash memory is gaining popularity as a storage medium for mobile embedded devices, many flash-aware file systems, flash-aware DBMSes, and flash translation layers (FTLs) require an flash-efficient index structure. This paper proposes a novel index structure called μ*-Tree which natively works on NAND flash memory, aiming at improving performance over B+-Tree. μ*-Tree stores all the nodes along the path from the root to the leaf into a single flash memory page in order to minimize the number of flash write operation when a node is updated. Furthermore, μ*-Tree has an adaptive page layout scheme which dynamically adjusts the page layout according to the workload characteristics on-the-fly. μ*-Tree also allows flash pages with different page layouts to coexist in the same tree. Our evaluation results with real workload traces show that μ*-Tree outperforms B+-Tree by up to 55 percent in terms of the time needed for flash operations. With a small in-memory cache of 32 KB, μ*-Tree improves the overall performance by up to five times compared to B+-Tree with the same cache size. Jung-Sang Ahn, Dongwon Kang, Da Woon Jung 0001, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 4 |
| 2012 | Extent Mapping Scheme for Flash Memory DevicesabstractFlash memory devices commonly rely on traditional address mapping schemes such as page mapping, block mapping or a hybrid of the two. Page mapping is more flexible than block mapping or hybrid mapping without being restricted by block boundaries. However, its mapping table tends to grow large quickly as the capacity of flash memory devices does. To overcome this limitation, we propose a novel mapping scheme that is fundamentally different from the existing mapping strategies. We call this new scheme Virtual Extent Trie (VET), as it manages mapping information by treating each I/O request as an extent and by using extents as basic mapping units rather than pages or blocks. By storing extents instead of individual addresses, VET consumes much less memory to store mapping information and still remains as flexible as page mapping. We observed in our experiments that VET reduced memory consumption by up to an order of magnitude in comparison with the traditional mapping schemes for several real world workloads. The VET scheme also scaled well with increasing address spaces by synthetic workloads. With a binary search mechanism, VET limits the mapping time to O(log log|U |), where U denotes the set of all possible logical addresses. Though the asymptotic mapping cost of VET is higher than the O(1) time of a page mapping scheme, the amount of increased overhead was almost negligible or low enough to be hidden by an accompanying I/O operation. Young-Kyoon Suh, Bongki Moon, Alon Efrat, Jin-Soo Kim 0001, Sang-Won Lee 0001 |
MASCOTS | 4 |
| 2012 | Parameter-Aware I/O Management for Solid State Disks (SSDs)abstractSolid state disks (SSDs) have many advantages over hard disk drives, including better reliability, performance, durability, and power efficiency. However, the characteristics of SSDs are completely different from those of hard disk drives with rotating disks. To achieve the full potential performance improvement with SSDs, operating systems or applications must understand the critical performance parameters of SSDs to fine-tune their accesses. However, the internal hardware and software organizations vary significantly among SSDs and, thus, each SSD exhibits different parameters which influence the overall performance. In this paper, we propose a methodology which can extract several essential parameters affecting the performance of SSDs, and apply the extracted parameters to SSD systems for performance improvement. The target parameters of SSDs considered in this paper are 1) the size of read/write unit, 2) the size of erase unit, 3) the size of read buffer, and 4) the size of write buffer. We modify two operating system components to optimize their operations with the SSD parameters. The experimental results show that such parameter-aware management leads to significant performance improvements for large file accesses by performing SSD-specific optimizations. Jaehong Kim 0006, Sangwon Seo, Da Woon Jung 0001, Jin-Soo Kim 0001, Jaehyuk Huh 0001 |
IEEE Trans. Computers | 4 |
| 2012 | FlashLight: A Lightweight Flash File System for Embedded SystemsabstractA very promising approach for using NAND flash memory as a storage medium is a flash file system. In order to design a higher-performance flash file system, two issues should be considered carefully. One issue is the design of an efficient index structure that contains the locations of both files and data in the flash memory. For large-capacity storage, the index structure must be stored in the flash memory to realize low memory consumption; however, this may degrade the system performance. The other issue is the design of a novel garbage collection (GC) scheme that reclaims obsolete pages. This scheme can induce considerable additional read and write operations while identifying and migrating valid pages. In this article, we present a novel flash file system that has the following features: ( i ) a lightweight index structure that introduces the hybrid indexing scheme and intra-inode index logging , and ( ii ) an efficient GC scheme that adopts a dirty list with an on-demand GC approach as well as fine-grained data separation and erase-unit data allocation . We implemented FlashLight in a Linux OS with kernel version 2.6.21 on an embedded device. The experimental results obtained using several benchmark programs confirm that FlashLight improves the performance by up to 27.4% over UBIFS by alleviating index management and GC overheads by up to 33.8%. Jaegeuk Kim, Hyotaek Shim, Seon-Yeong Park, Seung Ryoul Maeng, Jin-Soo Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2012 | A low-overhead networking mechanism for virtualized high-performance computing systems
Jae-Wan Jang, Euiseong Seo, Heeseung Jo, Jin-Soo Kim 0001 |
J. Supercomput. | 4 |
| 2011 | Cost optimized provisioning of elastic resources for application workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 3 |
| 2011 | BTS: Resource capacity estimate for time-targeted science workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Ewa Deelman, Seung Ryoul Maeng |
J. Parallel Distributed Comput. | 3 |
| 2011 | Replicated abstract data types: Building blocks for collaborative applications
Hyun-Gul Roh, Myeongjae Jeon, Jin-Soo Kim 0001, Joonwon Lee |
J. Parallel Distributed Comput. | 3 |
| 2011 | Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine SchedulingabstractIncreasing energy consumption in server consolidation environments leads to high maintenance costs for data centers. Main memory, no less than processor, is a major energy consumer in this environment. This paper proposes a technique for reducing memory energy consumption using virtual machine scheduling in multicore systems. We devise several heuristic scheduling algorithms by using a memory power simulator, which we designed and implemented. We also implement the biggest cover set first (BCSF) scheduling algorithm in the working server system. Through extensive simulation and implementation experiments, we observe the effectiveness of the memory-aware virtual machine scheduling in saving memory energy. In addition, we find out that power-aware memory management is essential to reduce the memory energy consumption. Jae-Wan Jang, Myeongjae Jeon, Hyo-Sil Kim, Heeseung Jo, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 5 |
| 2010 | HAMA: An Efficient Matrix Computation with the MapReduce FrameworkabstractVarious scientific computations have become so complex, and thus computation tools play an important role. In this paper, we explore the state-of-the-art framework providing high-level matrix computation primitives with MapReduce through the case study approach, and demonstrate these primitives with different computation engines to show the performance and scalability. We believe the opportunity for using MapReduce in scientific computation is even more promising than the success to date in the parallel systems literature. Sangwon Seo, Edward J. Yoon, Jaehong Kim 0006, Seongwook Jin, Jin-Soo Kim 0001, Seung Ryoul Maeng |
CloudCom | 5 |
| 2010 | An adaptive partitioning scheme for DRAM-based cache in Solid State DrivesabstractRecently, NAND flash-based Solid State Drives (SSDs) have been rapidly adopted in laptops, desktops, and server storage systems because their performance is superior to that of traditional magnetic disks. However, NAND flash memory has some limitations such as out-of-place updates, bulk erase operations, and a limited number of write operations. To alleviate these unfavorable characteristics, various techniques for improving internal software and hardware components have been devised. In particular, the internal device cache of SSDs has a significant impact on the performance. The device cache is used for two main purposes: to absorb frequent read/write requests and to store logical-to-physical address mapping information. In the device cache, we observed that the optimal ratio of the data buffering and the address mapping space changes according to workload characteristics. To achieve optimal performance in SSDs, the device cache should be appropriately partitioned between the two main purposes. In this paper, we propose an adaptive partitioning scheme, which is based on a ghost caching mechanism, to adaptively tune the ratio of the buffering and the mapping space in the device cache according to the workload characteristics. The simulation results demonstrate that the performance of the proposed scheme approximates the best performance. Hyotaek Shim, Bon-Keun Seo, Jin-Soo Kim 0001, Seung Ryoul Maeng |
MSST | 3 |
| 2010 | Log' version vector: Logging version vectors concisely in dynamic replication
Hyun-Gul Roh, Myeongjae Jeon, Euiseong Seo, Jin-Soo Kim 0001, Joonwon Lee |
Inf. Process. Lett. | 4 |
| 2010 | Superblock FTL: A superblock-based flash translation layer with a hybrid address translation schemeabstractIn NAND flash-based storage systems, an intermediate software layer called a Flash Translation Layer (FTL) is usually employed to hide the erase-before-write characteristics of NAND flash memory. We propose a novel superblock-based FTL scheme, which combines a set of adjacent logical blocks into a superblock. In the proposed Superblock FTL, superblocks are mapped at coarse granularity, while pages inside the superblock are mapped freely at fine granularity to any location in several physical blocks. To reduce extra storage and flash memory operations, the fine-grain mapping information is stored in the spare area of NAND flash memory. This hybrid address translation scheme has the flexibility provided by fine-grain address translation, while reducing the memory overhead to the level of coarse-grain address translation. Our experimental results show that the proposed FTL scheme significantly outperforms previous block-mapped FTL schemes with roughly the same memory overhead. Da Woon Jung 0001, Jeong-Uk Kang, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2009 | HPMR: Prefetching and pre-shuffling in shared MapReduce computation environmentabstractMapReduce is a programming model that supports distributed and parallel processing for large-scale data-intensive applications such as machine learning, data mining, and scientific simulation. Hadoop is an open-source implementation of the MapReduce programming model. Hadoop is used by many companies including Yahoo!, Amazon, and Facebook to perform various data mining on large-scale data sets such as user search logs and visit logs. In these cases, it is very common to share the same computing resources by multiple users due to practical considerations about cost, system utilization, and manageability. However, Hadoop assumes that all cluster nodes are dedicated to a single user, failing to guarantee high performance in the shared MapReduce computation environment. In this paper, we propose two optimization schemes, prefetching and pre-shuffling, which improve the overall performance under the shared environment while retaining compatibility with the native Hadoop. The proposed schemes are implemented in the native Hadoop-0.18.3 as a plug-in component called HPMR (High Performance MapReduce Engine). Our evaluation on the Yahoo!Grid platform with three different workloads and seven types of test sets from Yahoo! shows that HPMR reduces the execution time by up to 73%. Sangwon Seo, Ingook Jang, Kyungchang Woo, Inkyo Kim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
CLUSTER | 5 |
| 2009 | A methodology for extracting performance parameters in solid state disks (SSDs)abstractSolid state disks (SSDs) consisting of NAND flash memory are being widely used in laptops, desktops, and even enterprise servers. SSDs have many advantages over hard disk drives (HDDs) in terms of reliability, performance, durability, and power efficiency. Typically, the internal hardware and software organization varies significantly from SSD to SSD and thus each SSD exhibits different parameters which influence the overall performance. In this paper, we propose a methodology which can extract several essential parameters affecting the performance of SSDs. The target parameters of SSDs considered in this paper are (1) the size of read/write unit, (2) the size of erase unit, (3) the type of NAND flash memory used, (4) the size of read buffer, and (5) the size of write buffer. Obtaining these parameters will allow us to understand the internal architecture of the target SSD better and to get the most performance out of SSD by performing SSD-specific optimizations. Jaehong Kim 0006, Da Woon Jung 0001, Jin-Soo Kim 0001, Jaehyuk Huh 0001 |
MASCOTS | 3 |
| 2009 | DynaGrid: An Adaptive, Scalable, and Reliable Resource Provisioning Framework for WSRF-Compliant Applications
Eun-Kyu Byun, Jin-Soo Kim 0001 |
J. Grid Comput. | 2 |
| 2009 | Catching two rabbits: adaptive real-time support for embedded LinuxabstractAbstract The trend of digital convergence makes multitasking common in many digital electronic products. Some applications in those systems have inherent real‐time properties, while many others have few or no timeliness requirements. Therefore the embedded Linux kernels, which are widely used in those devices, provide real‐time features in many forms. However, providing real‐time scheduling usually induces throughput degradation in heavy multitasking due to the increased context switches. Usually the throughput degradation becomes a critical problem, since the performance of the embedded processors is generally limited for cost, design and energy efficiency reasons. This paper proposes schemes to lessen the throughput degradation, which is from real‐time scheduling, by suppressing unnecessary context switches and applying real‐time scheduling mechanisms only when it is necessary. Also the suggested schemes enable the complete priority inheritance protocol to prevent the well‐known priority inversion problem. We evaluated the effectiveness of our approach with open‐source benchmarks. By using the suggested schemes, the throughput is improved while the scheduling latency is kept same or better in comparison with the existing approaches. Copyright © 2008 John Wiley & Sons, Ltd. Euiseong Seo, Jinkyu Jeong, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee |
Softw. Pract. Exp. | 4 |
| 2008 | Estimating Resource Needs for Time-Constrained WorkflowsabstractWorkflow technologies have become a major vehicle for the easy and efficient development of science applications. At the same time new computing environments such as the Cloud are now available. A challenge is to determine the right amount of resources to provision for an application. This paper introduces an algorithm named balanced time scheduling (BTS), which estimates the minimum number of virtual processors required to execute a workflow within a user-specified finish time. The resource estimate of BTS is abstract, so it can be easily integrated with any resource description language or any resource provisioning system. The experimental results with a number of synthetic workflows demonstrate that BTS can estimate the computing capacity close to the optimal. The algorithm is scalable so that its turnaround time is only tens of seconds even with workflows having thousands of tasks and edges. Eun-Kyu Byun, Yang-Suk Kee, Ewa Deelman, Karan Vahi, Gaurang Mehta, Jin-Soo Kim 0001 |
eScience | 6 |
| 2008 | µ-FTL: : a memory-efficient flash translation layer supporting multiple mapping granularitiesabstractNAND flash memory is being widely adopted as a storage medium for embedded devices. FTL (Flash Translation Layer) is one of the most essential software components in NAND flash-based embedded devices as it allows to use legacy files systems by emulating the traditional block device interface on top of NAND flash memory. Yong-Goo Lee, Da Woon Jung 0001, Dongwon Kang, Jin-Soo Kim 0001 |
EMSOFT | 4 |
| 2008 | RMA: A Read Miss-Based Spin-Down Algorithm using an NV cacheabstractIt is an important issue to reduce the power consumption of a hard disk that takes a large amount of computer systempsilas power. As a new trend, an NV cache is used to make a disk spin down longer by servicing read/write requests instead of the disk. During the spin-down periods, write requests can be simply handled by write buffering, but read requests are still the main cause of initiating spin-ups because of a low hit ratio in the NV cache. Even when there is no user activity, read requests can be frequently generated by running applications and system services, hindering the spin-down. In this paper, we propose new NV cache policies: active write caching to reduce or to delay spin-ups caused by read misses during spin-down periods and a read miss-based spin-down algorithm to extend the spin-down periods, exploiting the NV cache effectively. Our policies reduce the power consumption of a hard disk by up to 50.1% with a 512 MB NV cache, compared with preceding approaches. Hyotaek Shim, Jaegeuk Kim, Da Woon Jung 0001, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICCD | 4 |
| 2008 | snapPVFS: Snapshot-Able Parallel Virtual File SystemabstractIn this paper, we propose a modified parallel virtual file system that provides snapshot functionality. Because typical file systems are exposed to various failures, taking a snapshot is a good way to enhance the reliability of file systems. The PVFS, which is one of the famous parallel file systems deployed in cluster systems, is vulnerable to system failures or users¿ mistakes; however, there is a scarcity of research on snapshots or online backup for the PVFS. Because a PVFS consists of multiple servers on a network, snapshots should be generated properly in each server in the system. Furthermore, before snapshots are generated, the status of each PVFS server must be checked to guarantee sound operation. To demonstrate our approach, we implemented two prototypes of a snapshot-able PVFS (snapPVFS). The performance measurements indicate that an administrator can take snapshots of an entire parallel file system and properly access any previous versions of files or directories in the future without serious performance degradation. Kwangho Cha, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICPADS | 2 |
| 2008 | Efficient Metadata Management for Flash File SystemsabstractNAND flash memory becomes one of the most popular storage for portable embedded systems. Although many flash-aware file systems, such as JFFS2 and YAFFS2, were proposed, the large memory consumption and the long mount delay have been serious obstacles for large-capacity NAND flash memory. In this paper, we present a new flash-aware file system called DFFS (direct flash file system) which fetches only the needed metadata on demand from flash memory. In addition, DFFS employs two novel metadata management schemes, inode embedding scheme and hybrid inode indexing scheme, to improve the performance of metadata operations. Comprehensive evaluation results using microbench- mark, postmark, and Linux kernel compilation trace, show that DFFS has comparable performance to JFFS2 and YAFFS2, while achieving a small memory footprint and instant mount time. Jaegeuk Kim, Heeseung Jo, Hyotaek Shim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ISORC | 4 |
| 2008 | TSB: A DVS algorithm with quick response for general purpose operating systems
Euiseong Seo, Seon-Yeong Park, Jin-Soo Kim 0001, Joonwon Lee |
J. Syst. Archit. | 3 |
| 2008 | A reconfigurable FTL (flash translation layer) architecture for NAND flash-based applicationsabstractIn this article, a novel FTL (flash translation layer) architecture is proposed for NAND flash-based applications such as MP3 players, DSCs (digital still cameras) and SSDs (solid-state drives). Although the basic function of an FTL is to translate a logical sector address to a physical sector address in flash memory, efficient algorithms of an FTL have a significant impact on performance as well as the lifetime. After the dominant parameters that affect the performance and endurance are categorized, the design space of the FTL architecture is explored based on a diverse workload analysis. With the proposed FTL architectural framework, it is possible to decide which configuration of FTL mapping parameters yields the best performance, depending on the differing characteristics of various NAND flash-based applications. Chanik Park, Wonmoon Cheon, Jeong-Uk Kang, Kangho Roh, Wonhee Cho 0006, Jin-Soo Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2008 | ScaleFFS: A scalable log-structured flash file system for mobile multimedia systemsabstractNAND flash memory has become one of the most popular storage media for mobile multimedia systems. A key issue in designing storage systems for mobile multimedia systems is handling large-capacity storage media and numerous large files with limited resources such as memory. However, existing flash file systems, including JFFS2 and YAFFS in particular, exhibit many limitations in addressing the storage capacity of mobile multimedia systems. In this article, we design and implement a scalable flash file system, called ScaleFFS, for mobile multimedia systems. ScaleFFS is designed to require only a small fixed amount of memory space and to provide fast mount time, even if the file system size grows to more than tens of gigabytes. The measurement results show that ScaleFFS can be instantly mounted regardless of the file system size, while achieving the same write bandwidth and up to 22% higher read bandwidth compared to JFFS2. Da Woon Jung 0001, Jaegeuk Kim, Jin-Soo Kim 0001, Joonwon Lee |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2007 | A group-based wear-leveling algorithm for large-capacity flash memory storage systemsabstractAlthough NAND flash memory has become one of the most popular storage media for portable devices, it has a serious problem with respect to lifetime. Each block of NAND flash memory has a limited number of program/erase cycles, usually 10,000–100,000, and data in a block become unreliable after the limit. For this reason, distributing erase operations evenly across the whole flash memory media is an important concern in designing flash memory storage systems. In this paper, we propose a memory-efficient group-based wear-leveling algorithm. Our group-based algorithm achieves a small memory footprint by grouping several logically sequential blocks and managing only the summary information for each group. We also propose an effective group summary structure and a method to reduce unnecessary wearleveling operations in order to enhance the wear-leveling performance. The evaluation results show that our group-based algorithm consumes only 8.75 % of memory space compared to the previous scheme that manages per-block information, while showing roughly the same wear-leveling performance. Da Woon Jung 0001, Yoon-Hee Chae, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee |
CASES | 4 |
| 2007 | FlexRPC: A flexible Remote Procedure Call facility for modern cluster file systemsabstractThe concept of Remote Procedure Call (RPC) was proposed more than 30 years ago. Although various RPC systems have been studied and implemented, the existing RPC systems lack many crucial features and flexibility required for building modern cluster file systems .This paper presents FlexRPC, a flexible user-level RPC system that enables to develop high-performance cluster file systems easily. FlexRPC ensures client-side thread-safeness and fully supports multithreaded RPC servers. Parallel and serial multicasting mechanisms allow for implementing sophisticated replication in modern cluster file systems. The remote procedure can be invoked using both UDP and TCP transports with at-most-once semantics. The concurrent call requests are handled by a set of worker threads on the client and server side where the number of workers varies dynamically according to the request rate. In addition, the semantics and the specification of remote procedures are designed to be as close as possible to SunRPC. The experimental results show that FlexRPC improves both latency and bandwidth significantly in spite of added functionalities. We also demonstrate the performance and the flexibility provided by FlexRPC by building working prototype of cluster file system called Kadoop on top of FlexRPC. Sang-Hoon Kim, Youngjae Lee, Jin-Soo Kim 0001 |
CLUSTER | 3 |
| 2007 | mu-tree: an ordered index structure for NAND flash memoryabstractAs NAND flash memory becomes increasingly popular as data storage for embedded systems, many file systems and database management systems are being built on it. They require an efficient index structure to locate a particular item quickly from a huge amount of directory entries or database records. This paper proposes μ-Tree, a new ordered index structure tailored to the characteristics of NAND flash memory. μ-Tree is a balanced tree similar to B+-Tree. In μ-Tree, however, all the nodes along the path from the root to the leaf are put together into a single flash memory page in order to minimize the number of flash write operations when a leaf node is updated. Our experimental evaluation shows that μ-Tree outperforms B+-Tree by up to 28% for traces extracted from real workloads. With a small in-memory cache of 8 Kbytes, μ-Tree improves the overall performance by up to 90% compared to B+-Tree with the same cache size. Dongwon Kang, Da Woon Jung 0001, Jeong-Uk Kang, Jin-Soo Kim 0001 |
EMSOFT | 4 |
| 2007 | Towards adaptive, scalable, and reliable resource provisioning for wsrf-compliant applicationsabstractAlthough WSRF (Web Services Resource Framework) and Java-based hosting environment have been successful in dealing with the heterogeneity of resources and the diversity of applications, the current Grid middleware has several limitations to support on-demand resource provisioning effectively. Eun-Kyu Byun, Jae-Wan Jang, Jin-Soo Kim 0001 |
HPDC | 3 |
| 2007 | A multi-channel architecture for high-performance NAND flash-based storage system
Jeong-Uk Kang, Jin-Soo Kim 0001, Chanik Park, Hyoungjun Park, Joonwon Lee |
J. Syst. Archit. | 2 |
| 2007 | DynaGrid: A dynamic service deployment and resource migration framework for WSRF-compliant applications
Eun-Kyu Byun, Jin-Soo Kim 0001 |
Parallel Comput. | 2 |
| 2007 | PABC: Power-Aware Buffer Cache Management for Low Power ConsumptionabstractPower consumed by memory systems becomes a serious issue as the size of the memory installed increases. With various low power modes that can be applied to each memory unit, the operating system can reduce the number of active memory units by collocating active pages onto a few memory units. This paper presents a memory management scheme based on this observation, which differs from other approaches in that all of the memory space is considered, while previous methods deal only with pages mapped to user address spaces. The buffer cache usually takes more than half of the total memory and the pages access patterns are different from those in user address spaces. Based on an analysis of buffer cache behavior and its interaction with the user space, our scheme achieves up to 63 percent more power reduction. Migrating a page to a different memory unit increases memory latencies, but it is shown to reduce the power consumed by an additional 4.4 percent Min Lee, Euiseong Seo, Joonwon Lee, Jin-Soo Kim 0001 |
IEEE Trans. Computers | 4 |
| 2007 | Design issues and performance comparisons in supporting the sockets interface over user-level communication architecture
Jae-Wan Jang, Jin-Soo Kim 0001 |
J. Supercomput. | 2 |
| 2007 | A runtime resolution scheme for priority boost conflict in implicit coscheduling
Jung-Lok Yu, Jin-Soo Kim 0001, Seung Ryoul Maeng |
J. Supercomput. | 2 |
| 2006 | CFLRU: a replacement algorithm for flash memoryabstractIn most operating systems which are customized for disk-based storage system, the replacement algorithm concerns only the number of memory hits. However, flash memory has different read and write cost in the aspects of time and energy so the replacement algorithm with flash memory should consider not only the hit count but also the replacement cost caused by selecting dirty victims. The replacement cost of dirty page is higher than that of clean page with regard to both access time and energy consumption. In this paper, we propose the Clean-First LRU (CFLRU) replacement algorithm that exploits the characteristics of flash memory. CFLRU splits the LRU list into the working region and the clean-first region and adopts a policy that evicts clean pages preferentially in the clean-first region until the number of page hits in the working region is preserved in a suitable level. Using the trace-driven simulation, the proposed algorithm reduces the average replacement cost by 28.4% in swap system and by 26.2% in buffer cache, compared with LRU algorithm. We also implement the CFLRU algorithm in the Linux kernel and present some optimization issues. Seon-Yeong Park, Da Woon Jung 0001, Jeong-Uk Kang, Jin-Soo Kim 0001, Joonwon Lee |
CASES | 4 |
| 2006 | A superblock-based flash translation layer for NAND flash memoryabstractIn NAND flash-based storage systems, an intermediate software layer called a flash translation layer (FTL) is usually employed to hide the erase-before-write characteristics of NAND flash memory. This paper proposes a novel superblockbased FTL scheme, which combines a set of adjacent logical blocks into a superblock. In the proposed FTL scheme, superblocks are mapped at coarse granularity, while pages inside the superblock are mapped freely at fine granularity to any location in several physical blocks. To reduce extra storage and flash memory operations, the fine-grain mapping information is stored in the spare area of NAND flash memory. This hybrid mapping technique has the flexibility provided by fine-grain address translation, while reducing the memory overhead to the level of coarse-grain address translation. Our experimental results show that the proposed FTL scheme decreases the garbage collection overhead up to 40 % compared to previous FTL schemes. Jeong-Uk Kang, Heeseung Jo, Jin-Soo Kim 0001, Joonwon Lee |
EMSOFT | 3 |
| 2006 | Runtime feasibility check for non-preemptive real-time periodic tasks
Joonwon Lee, Jin-Soo Kim 0001 |
Inf. Process. Lett. | 3 |
| 2005 | A dynamic grid services deployment mechanism for on-demand resource provisioningabstractRecently grid computing has started to leverage Web services technology by proposing OGSI-standard. OGSI standard defines the grid service, which presents unified interfaces to every participant of grid. In current Glubus Toolkit3(GT3),which is an implementation of OGSI, grid service factories should be deployed manually into resources to provide grid services. However, it is necessary to dynamically allocate proper amount of resource, since the demand for resource of service provider changes over time. In this paper, we propose a architecture to enable on-demand resource provisioning. We develop universal factory service (UFS) that provides a dynamic grid service deployment mechanism and a resource broker called door service. Through the experiments, we show that grid services can adaptively exploit resources according to the request rates. Eun-Kyu Byun, Jae-Wan Jang, Wook Jung, Jin-Soo Kim 0001 |
CCGRID | 4 |
| 2005 | Supporting the Sockets Interface over User-Level Communication Architecture: Design Issues and Performance ComparisonsabstractSince user-level communication architecture (ULC) provides only primitive operations for application programmers, many high-level communication layers have been developed on top of ULC. One of such high-level communication layers is the sockets interface and it can be supported over ULC architectures in several ways. The primary objective of this paper is to identify design issues and tradeoffs among these different approaches, and to quantitatively analyze their performance to understand the various costs associated with the communication. In this paper, we design and implement KSOVIA, a kernel-level sockets layer over VIA, and compare it with the existing approaches such as a user-level sockets layer over VIA and an IP emulation layer over VIA. Our measurement results show that using an IP emulation layer exhibits the worst performance in terms of latency and bandwidth and a user-level sockets layer is useful for latency-sensitive applications. KSOVIA is found to be effective for applications, which require high bandwidth or the full compatibility with the sockets interface. Jae-Wan Jang, Jin-Soo Kim 0001 |
ICPP | 2 |
| 2005 | Impact of Exploiting Load Imbalance on Coscheduling in Workstation ClustersabstractImplicit coscheduling is known to be an effective technique to improve the performance of parallel workloads in time-sharing clusters. However, implicit coscheduling still does not take into consideration the system behavior like load imbalance that severely affects cluster utilization. In this paper, we propose the use of global information to enhance the existing implicit coscheduling schemes. We also introduce a novel coscheduling approach - named PROC (process reordering-based coscheduling) - based on process reordering exploiting global load imbalance information to coordinate communicating processes. The results obtained from an in-depth simulation study show that our approach significantly outperforms previous ones (by up to 38.4%) by reducing the idle time (by up to 86.9%) and spin time (by up to 36.2%) caused by the load imbalance. Jung-Lok Yu, Driss Azougagh, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICPP | 3 |
| 2004 | Design and implementation of an OGSI-compliant Grid broker serviceabstractGrid computing promises the ability to share geographically and organizationally distributed resources to increase effective computational power and resource utilization. However, for the Grid computing to be successful, it is very important to provide middleware services that assist Grid users to easily interact with Grid environments. In this paper, we have designed and implemented a new general-purpose OGSI-compliant Grid resource broker service to hide the underlying complexity of the Grid resources from Grid users and to meet not only Grid user requirements but also resource owner policies. It focuses on discovering and scheduling dynamic resources scattered across multiple organizations. Furthermore, it can be integrated with various scheduling services. We also present experimental results and demonstrate the effectiveness of our Grid broker service. Youn-Seok Kim, Jung-Lok Yu, Jae-Gyoon Hahm, Jin-Soo Kim 0001, Joonwon Lee |
CCGRID | 4 |
| 2004 | Energy-aware demand paging on NAND flash-based embedded storagesabstractThe ever-increasing requirement for high-performance and huge-capacity memories of emerging embedded applications has led to the widespread adoption of SDRAM and NAND flash memory as main and secondary memories, respectively. In particular, the use of energy consuming memory, SDRAM, has become burdensome in battery-powered embedded systems. Intuitively, though demand paging can be used to mitigate the increasing requirement of main memory size, its applicability should be deliberately elaborated since NAND flash memory has asymmetric operation characteristics in terms of performance and energy consumption.In this paper, we present energy-aware demand paging technique to lower the energy consumption of embedded systems considering the characteristics of interactive embedded applications with large memory footprints. We also propose a flash memory-aware page replacement policy that can reduce the number of write and erase operations in NAND flash memory. With real-life workloads, we show the system-wide Energy·Delay can be reduced by 15~30% compared to the traditional shadowing architecture. Chanik Park, Jeong-Uk Kang, Seon-Yeong Park, Jin-Soo Kim 0001 |
ISLPED | 4 |
| 2004 | Memory management for multi-threaded software DSM systems
Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
Parallel Comput. | 2 |
| 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster SystemsabstractDemand for programming environments to exploit clusters of symmetric multiprocessors (SMPs) is increasing. In this paper, we present a new programming environment, called ParADE, to enable easy, portable, and high-performance programming on SMP clusters. It is an OpenMP programming environment on top of a multi-threaded software distributed shared memory (SDSM) system with a variant of home-based lazy release consistency protocol. To boost performance, the runtime system provides explicit message-passing primitives to make it a hybrid-programming environment. Collective communication primitives are used for the synchronization and work-sharing directives associated with small data structures, lessening the synchronization overhead and avoiding the implicit barriers of work-sharing directives. The OpenMP translator bridges the gap between the OpenMP abstraction and the hybrid programming interfaces of the runtime system. The experiments with several NAS benchmarks and applications on a Linux-based cluster show promising results that ParADE overcomes the performance problem of the conventional SDSM-based OpenMP environment. Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
SC | 2 |