VLDB 2026 Research / reviewers in the wild / expert
Seung Ryoul Maeng
dblp:m/SeungRyoulMaeng · also Seungryoul Maeng
· DBLP profile ↗
77ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 1 since 2021Software engineering, systems software and programming languages · 5Graphics, computer vision, multimedia, augmented reality and games · 4Artificial intelligence and machine learning · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
16 papers |
Memory systems · 26% Storage systems · 20% Cloud and datacenter computing · 17% | |
| Network and information security
2 papers |
Hardware security and side channels · 76% Systems and software security · 24% | |
| Databases, data mining, and information retrieval
2 papers |
Indexing and storage engines · 100% | |
| Software engineering, system software, and programming languages
5 papers |
Operating systems · 98% Requirements engineering and software design · 1% Concurrent programming · 1% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › non-volatile memory › persistent memory
atomic durability |
0.8 | 2 | 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory System · MICRO 2020 Efficient Hardware-Assisted Logging with Asynchronous and Direct-Update for Persistent Memory · MICRO 2018 |
Memory systems
non-volatile memory |
0.8 | 2 | 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory System · MICRO 2020 Efficient Hardware-Assisted Logging with Asynchronous and Direct-Update for Persistent Memory · MICRO 2018 |
Memory systems › non-volatile memory
persistent memory |
0.8 | 2 | 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory System · MICRO 2020 Efficient Hardware-Assisted Logging with Asynchronous and Direct-Update for Persistent Memory · MICRO 2018 |
Hardware security and side channels
trusted execution environments |
0.5 | 2 | 2016 | A Trusted IaaS Environment with Hardware Security Module · IEEE Trans. Serv. Comput. 2016 H-SVM: Hardware-Assisted Secure Virtual Machines under a Vulnerable Hypervisor · IEEE Trans. Computers 2015 |
Interconnection networks and networks-on-chip
flow control |
0.4 | 2 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 Transportation-network-inspired network-on-chip · HPCA 2014 |
Interconnection networks and networks-on-chip › ring network
hierarchical ring |
0.4 | 2 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 Transportation-network-inspired network-on-chip · HPCA 2014 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.4 | 1 | 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory System · MICRO 2020 |
Parallel and multicore computing
transactional memory |
0.4 | 1 | 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory System · MICRO 2020 |
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine scheduling |
0.4 | 3 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine Scheduling · IEEE Trans. Computers 2011 Virtualizing performance asymmetric multi-core systems · ISCA 2011 |
Cloud and datacenter computing
virtualization |
0.4 | 3 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 Virtualizing performance asymmetric multi-core systems · ISCA 2011 Transparent Fault Tolerance of Device Drivers for Virtual Machines · IEEE Trans. Computers 2010 |
Storage systems
flash and SSD |
0.4 | 2 | 2014 | System-Wide Cooperative Optimization for NAND Flash-Based Mobile Systems · IEEE Trans. Computers 2014 μ*-Tree: An Ordered Index Structure for NAND Flash Memory with Adaptive Page Layout Scheme · IEEE Trans. Computers 2013 |
Storage systems › logging
write-ahead logging |
0.3 | 1 | 2018 | Efficient Hardware-Assisted Logging with Asynchronous and Direct-Update for Persistent Memory · MICRO 2018 |
Indexing and storage engines › string indexing
trie index |
0.2 | 1 | 2016 | ForestDB: A Fast Key-Value Storage System for Variable-Length String Keys · IEEE Trans. Computers 2016 |
Hardware security and side channels › cryptographic hardware
hardware security module |
0.2 | 1 | 2016 | A Trusted IaaS Environment with Hardware Security Module · IEEE Trans. Serv. Comput. 2016 |
Storage systems
key-value storage |
0.2 | 1 | 2016 | ForestDB: A Fast Key-Value Storage System for Variable-Length String Keys · IEEE Trans. Computers 2016 |
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology |
0.2 | 1 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 |
Hardware security and side channels › trusted execution environments
hardware-assisted memory isolation |
0.2 | 1 | 2015 | H-SVM: Hardware-Assisted Secure Virtual Machines under a Vulnerable Hypervisor · IEEE Trans. Computers 2015 |
Systems and software security
virtualization security |
0.2 | 1 | 2015 | H-SVM: Hardware-Assisted Secure Virtual Machines under a Vulnerable Hypervisor · IEEE Trans. Computers 2015 |
Storage systems › flash and SSD
flash memory management |
0.2 | 1 | 2015 | Zombie Chasing: Efficient Flash Management Considering Dirty Data in the Buffer Cache · IEEE Trans. Computers 2015 |
Storage systems › flash and SSD › flash memory management
garbage collection |
0.2 | 1 | 2015 | Zombie Chasing: Efficient Flash Management Considering Dirty Data in the Buffer Cache · IEEE Trans. Computers 2015 |
Storage systems › flash and SSD
solid-state drive |
0.2 | 1 | 2015 | Zombie Chasing: Efficient Flash Management Considering Dirty Data in the Buffer Cache · IEEE Trans. Computers 2015 |
Operating systems › resource management
storage management |
0.2 | 1 | 2014 | System-Wide Cooperative Optimization for NAND Flash-Based Mobile Systems · IEEE Trans. Computers 2014 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.2 | 1 | 2014 | System-Wide Cooperative Optimization for NAND Flash-Based Mobile Systems · IEEE Trans. Computers 2014 |
Indexing and storage engines
tree index |
0.2 | 1 | 2013 | μ*-Tree: An Ordered Index Structure for NAND Flash Memory with Adaptive Page Layout Scheme · IEEE Trans. Computers 2013 |
Parallel and multicore computing › task scheduling
coordinated scheduling |
0.2 | 1 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 |
Electronic design automation › high-level synthesis
scheduling |
0.2 | 1 | 2013 | Demand-based coordinated scheduling for SMP VMs · ASPLOS 2013 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2012 | Locality-aware dynamic VM reconfiguration on MapReduce clouds · HPDC 2012 |
Processor architecture and microarchitecture › multicore design
heterogeneous multicore |
0.1 | 1 | 2011 | Virtualizing performance asymmetric multi-core systems · ISCA 2011 |
Parallel and multicore computing › task scheduling
memory-aware scheduling |
0.1 | 1 | 2011 | Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine Scheduling · IEEE Trans. Computers 2011 |
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation |
0.1 | 1 | 2011 | Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine Scheduling · IEEE Trans. Computers 2011 |
Methods — techniques the papers use, named apart from their topics
hardware security module · 0.5virtual channels · 0.4hardware logging · 0.4cache coherence protocol · 0.4address signatures · 0.4write cache · 0.3redo logging · 0.3asynchronous updates · 0.3trusted computing base minimization · 0.2simulation-based evaluation · 0.2credit network design · 0.2zombie-aware garbage collection · 0.2system management mode · 0.2nested paging · 0.2TRIM command · 0.2system-wide cooperative optimization · 0.2demand-based scheduling · 0.2communication-aware scheduling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | On-Demand Virtualization for Post-Copy OS Migration in Bare-Metal CloudabstractThe demand for bare-metal cloud services has increased rapidly because bare-metal cloud is cost-effective for various types of cloud workloads. However, as the bare-metal cloud does not utilize the abstraction of the virtualization layer, it misses the benefits of virtualization. One important benefits absent in the bare-metal cloud is the live migration of guest operating systems. Migrating an OS and applications in the OS as a single unit provides a convenient way to manage cloud services such as load balancing, fault management, and system maintenance. To enable live migration for bare-metal cloud, several approaches have been proposed but they have limitations; they require OS modifications or impose additional overheads for workloads. This paper suggests an on-demand virtualization technique for post-copy OS migration to improve manageability of the bare-metal cloud services. When live migration is requested, a lightweight virtualization layer is enabled in the host on the fly. After completion of the live migration, the virtualization layer is removed from the host. Therefore, the host returns to a bare-metal system for performance. To implement on-demand virtualization, we modify BitVisor to perform the post-copy migration on the x86 architecture. The elapsed time of on-demand virtualization is negligible. It takes only 20 ms to insert the virtualization layer and 30 ms to remove the one. The downtime of migration is reduced because of the post-copy migration. Jaeseong Im, Jongyul Kim 0001, Youngjin Kwon, Seung Ryoul Maeng |
IEEE Trans. Cloud Comput. | 4 |
| 2020 | Unbounded Hardware Transactional Memory for a Hybrid DRAM/NVM Memory SystemabstractPersistent memory programming requires failure atomicity. To achieve this in an efficient manner, recent proposals use hardware-based logging for atomic-durable updates and hardware transactional memory (HTM) for isolation. Although the unbounded HTMs are promising for both performance and programmability reasons, none of the previous studies satisfies the practical requirements. They either require unrealistic hard-ware overheads or do not allow transactions to exceed on-chip cache boundaries. Furthermore, it has never been possible to use both DRAM and NVM in HTM, though it is becoming a popular persistency model. To this end, this study proposes UHTM, unbounded hardware transactional memory for DRAM and NVM hybrid memory systems. UHTM combines the cache coherence protocol and address-signatures to detect conflicts in the entire memory space. This approach improves concurrency by significantly reducing the false-positive rates of previous studies. More importantly, UHTM allows both DRAM and NVM data to interact with each other in transactions without compromising the consistency guarantee. This is rendered possible by UHTM's hybrid version management that provides an undo-based log for DRAM and a redo-based log for NVM. The experimental results show that UHTM outperforms the state-of-the-art durable HTM, which is LLC-bounded, by 56% on average and up to 818%. Jungi Jeong, Jaewan Hong, Seung Ryoul Maeng, Changhee Jung, Youngjin Kwon |
MICRO | 3 |
| 2018 | Efficient Hardware-Assisted Logging with Asynchronous and Direct-Update for Persistent MemoryabstractSupporting atomic durability in emerging persistent memory requires data consistency across potential system failures. For atomic durability support in the non-volatile memory, the traditional write-ahead log (WAL) technique has been employed to guarantee the persistency of logs before actual data updates. Based on the WAL mechanism, recent studies proposed HWassisted logging techniques with undo, redo, or undo+redo principles. The HW log manager allows the overlapping of log writing and transaction execution, as long as the atomicity invariant can be satisfied. Although the efficiency of both log and data writes must be optimized, the prior work exhibit trade-offs in performance under various access patterns. The undo approach experiences performance degradation due to synchronous inplace data updates since the log contains only the old values. On the other hand, the undo+redo approach stores both old and new values, and does not require synchronous in-place data updates. However, the larger log size increases the amount of log writes. The prior redo approach demands extra NVM read bandwidth for indirectly updating in-place data from the new values in logs. To overcome the limitations of the previous approaches, this paper proposes a novel redo-based logging (ReDU), which performs direct and asynchronous in-place data update to NVM. ReDU exploits a small region of DRAM as a write-cache to remove NVM writes from the critical path. The experimental results show that the proposed logging mechanism provides the best performance under a variety of write patterns, showing 8.6%, 14.2%, and 23.6% better performance compared to the previous undo, redo, and undo+redo approaches, respectively. Jungi Jeong, Chang Hyun Park 0001, Jaehyuk Huh 0001, Seung Ryoul Maeng |
MICRO | 4 |
| 2017 | On-demand virtualization for live migration in bare metal cloudabstractThe level of demand for bare-metal cloud services has increased rapidly because such services are cost-effective for several types of workloads, and some cloud clients prefer a single-tenant environment due to the lower security vulnerability of such enviornments. However, as the bare-metal cloud does not utilize a virtualization layer, it cannot use live migration. Thus, there is a lack of manageability with the bare-metal cloud. Live migration support can improve the manageability of bare-metal cloud services significantly. Jaeseong Im, Jongyul Kim 0001, Jonguk Kim, Seongwook Jin, Seung Ryoul Maeng |
SoCC | 5 |
| 2016 | Application-Assisted Writeback for Hadoop ClustersabstractAchieving low and predictable execution time of short jobs in Hadoop clusters has gained a great attention due to their importance on system productivity and user experience. However, one major contributor that makes it challenging is diskI/O interference. We observed that disk writes unintentionally block latency-sensitive short jobs and cause unexpected high latency. Unfortunately, previous research including a disk read bandwidth throttling do not suffice to mitigate such interference. This paper proposes the application-assisted writeback that allows the Hadoop framework to control asynchronous writebacks. We applied the application-assisted writeback to optimize short jobs by preventing asynchronous writebacks when they are expected to interfere with short jobs. Our evaluation resultedin reduction on the average and 99-th percentile execution time of short jobs by 22% and 40%, respectively, without imposing non-acceptable overheads on co-running throughput-oriented batch jobs. In addition, combining the application-assisted writeback with the user-level disk bandwidth throttling can further accelerate short jobs. Jungi Jeong, DaeWoo Lee, Seung Ryoul Maeng |
CLUSTER | 3 |
| 2016 | ActiveSort: Efficient external sorting using active SSDs in the MapReduce framework
Young-Sik Lee, Luis Cavazos Quero, Sang-Hoon Kim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 5 |
| 2016 | ForestDB: A Fast Key-Value Storage System for Variable-Length String KeysabstractIndexing key-value data on persistent storage is an important factor for NoSQL databases. Most key-value storage engines use tree-like structures for data indexing, but their performance and space overhead rapidly get worse as the key length becomes longer. This also affects the merge or compaction cost which is critical to the overall throughput. In this paper, we present ForestDB, a key-value storage engine for a single node of large-scale NoSQL databases. ForestDB uses a new hybrid indexing scheme called HB+-trie, which is a disk-based trie-like structure combined with B+-trees. It allows for efficient indexing and retrieval of arbitrary length string keys with relatively low disk accesses over tree-like structures, even though the keys are very long and randomly distributed in the key space. Our evaluation results show that ForestDB significantly outperforms the current key-value storage engine of Couchbase Server [1], LevelDB [2], and RocksDB [3], in terms of both the number of operations per second and the amount of disk writes per update operation. Jung-Sang Ahn, Chiyoung Seo, Ravi Mayuram, Rahim Yaseen, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 6 |
| 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-ChipabstractA cost-efficient network-on-chip is needed in a scalable many-core systems. Recent multicore processors have leveraged a ring topology and hierarchical ring can increase scalability but presents different challenges, including higher hop count and global ring bottleneck. In this work, we describe a hierarchical ring topology that we refer to as a transportation-network-inspired network-on-chip (tNoC) that leverages principles from transportation network systems. In particular, we propose a novel hybridflow control for hierarchical ring topology to scale the topology efficiently. The flow control is hybrid in that the channels are allocated on flit granularity while the buffers are allocated on packet granularity. The hybrid flow control enables a simplified router microarchitecture (to minimize per-hop latency) as router input buffers are minimized and buffers are pushed to the edges, either at the output ports or at the hub routers that interconnect the local rings to the global ring-while still supporting virtual channels to avoid protocol deadlock. We describe a packet-quota-system (PQS) and a separate credit network that provide congestion management, support prioritized arbitration in the network, and provide support for multiflit packets. We also provide alternative designs for the credit network and PQS architectures. A detailed evaluation of a 64-core CMP shows that the tNoC improves performance by up to 21 percent compared with a baseline, buffered hierarchical ring topology while reducing NoC energy by 51 percent. Hanjoon Kim, Gwangsun Kim, Hwasoo Yeo, John Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 5 |
| 2016 | SmartLMK: A Memory Reclamation Scheme for Improving User-Perceived App Launch TimeabstractAs the mobile computing environment evolves, users demand high-quality apps and better user experience. Consequently, memory demand in mobile devices has soared. Device manufacturers have fulfilled the demand by equipping devices with more RAM. However, such a hardware approach is only a temporary solution and does not scale well in the resource-constrained mobile environment. Meanwhile, mobile systems adopt a new app life cycle and a memory reclamation scheme tailored for the life cycle. When a user leaves an app, the app is not terminated but cached in memory as long as there is enough free memory. If the free memory gets low, a victim app is terminated and the associated memory to the app is reclaimed. This process-level approach has worked well in the mobile environment. However, user experience can be impaired severely because the victim selection policy does not consider the user experience. In this article, we propose a novel memory reclamation scheme called SmartLMK . SmartLMK minimizes the impact of the process-level reclamation on user experience. The worthiness to keep an app in memory is modeled by means of user-perceived app launch time and app usage statistics. The memory footprint and impending memory demand are estimated from the history of the memory usage. Using these values and memory models, SmartLMK picks up the least valuable apps and terminates them at once. Our evaluation on a real Android-based smartphone shows that SmartLMK efficiently distinguishes the valuable apps among cached apps and keeps those valuable apps in memory. As a result, the user-perceived app launch time can be improved by up to 13.2%. Sang-Hoon Kim, Jinkyu Jeong, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2016 | A Trusted IaaS Environment with Hardware Security ModuleabstractWith the proliferation of cloud computing, security concerns about confidentiality violations of user data by the privileged domain and system administrators have been growing. This paper proposes secure cloud architecture with a hardware security module, which isolates cloud user data from potentially malicious privileged domains or cloud administrators. Within a securely isolated execution environment, the hardware security module provides essential security functionality with only restricted interfaces exposed to vulnerable management systems or cloud administrators. Such restriction prevents cloud administrators from affecting the security of guest VMs. The proposed architecture not only defends against wide attack vectors but also achieves a small TCB. This paper discusses our hardware and software implementation of the proposed cloud architecture, analyzes its security, and presents its performance results. Jinho Seol, Seongwook Jin, DaeWoo Lee, Jaehyuk Huh 0001, Seung Ryoul Maeng |
IEEE Trans. Serv. Comput. | 5 |
| 2015 | HePA: Hexagonal Platform Architecture for Smart Home ThingsabstractIn the internet era, where people are connected to each other, web server architectures have been developed and advanced. Now, we are witnessing the advent of the complex IoT (Internet of Things) era, where not only people, but also all devices on the planet are interconnected to each other and enormous amount of interactions is required. In this paper, we propose HePA (Hexagonal Platform Architecture), a platform architecture that is extremely scalable while maintaining required performance and reflecting requirements of the complex environment. We expect the HePA to become a reference architecture in this field. Sangwon Seo, Jaehong Kim 0006, Sangbae Yun, Jaehyuk Huh 0001, Seung Ryoul Maeng |
ICPADS | 5 |
| 2015 | Hardware-Assisted Secure Resource Accounting under a Vulnerable HypervisorabstractWith the proliferation of cloud computing to outsource computation in remote servers, the accountability of computational resources has emerged as an important new challenge for both cloud users and providers. Among the cloud resources, CPU and memory are difficult to verify their actual allocation, since the current virtualization techniques attempt to hide the discrepancy between physical and virtual allocations for the two resources. This paper proposes an online verifiable resource accounting technique for CPU and memory allocation for cloud computing. Unlike prior approaches for cloud resource accounting, the proposed accounting mechanism, called Hardware-assisted Resource Accounting (HRA), uses the hardware support for system management mode (SMM) and virtualization to provide secure resource accounting, even if the hypervisor is compromised. Using a secure isolated execution support of SMM, this study investigates two aspects of verifiable resource accounting for cloud systems. First, this paper presents how the hardware-assisted SMM and virtualization techniques can be used to implement the secure resource accounting mechanism even under a compromised hypervisor. Second, the paper investigates a sample-based resource accounting technique to minimize performance overheads. Using a statistical random sampling method, the technique estimates the overall CPU and memory allocation status with 99%~100% accuracies and performance degradations of 0.1%~0.5%. Seongwook Jin, Jinho Seol, Jaehyuk Huh 0001, Seung Ryoul Maeng |
VEE | 4 |
| 2015 | H-SVM: Hardware-Assisted Secure Virtual Machines under a Vulnerable HypervisorabstractWith increasing demands on cloud computing, protecting guest virtual machines (VMs) from malicious attackers has become critical to provide secure services. The current cloud security model with software-based virtualization relies on the invulnerability of the software hypervisor and its trustworthy administrator with the root permission. However, compromising the hypervisor with remote attacks or root permission grants the attackers with a full access capability to the memory and context of a guest VM. This paper proposes a HW-based approach to protect guest VMs even under an untrusted hypervisor. With the proposed mechanism, memory isolation is provided by the secure hardware, which is much less vulnerable than the software hypervisor. The proposed mechanism extends the current hardware support for memory virtualization based on nested paging with a small extra hardware cost. The hypervisor can still flexibly allocate physical memory pages to virtual machines for efficient resource management. In addition to the system design for secure virtualization, this paper presents a prototype implementation using system management mode. Although the current system management mode is not intended for security functions and thus limits the performance and complete protection, the prototype implementation proves the feasibility of the proposed design. Seongwook Jin, Jeongseob Ahn, Jinho Seol, Sanghoon Cha, Jaehyuk Huh 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 6 |
| 2015 | Zombie Chasing: Efficient Flash Management Considering Dirty Data in the Buffer CacheabstractThis paper presents a novel technique, called Zombie Chasing, for efficient flash management in solid state drives (SSDs). Due to the unique characteristics of NAND flash memory, SSDs need to accurately understand the liveness of the data stored in themselves. Recently, the TRIM command has been introduced to notify SSDs of dead data caused by file deletions, which otherwise could not be tracked by SSDs. This paper goes one step further and proposes a new liveness state, called the zombie state, to denote live data that will be dead shortly due to the corresponding dirty data in the buffer cache. We also devise new zombie-aware garbage collection algorithms which utilize the information about such zombie data inside SSDs. To evaluate Zombie Chasing, we implement zombie-aware garbage collection algorithms in the prototype SSD and modify the Linux kernel and the Oracle DBMS to deliver the information on the zombie data to the prototype SSD. Through comprehensive evaluations using our in-house micro-benchmark and the TPC-C benchmark, we observe that Zombie Chasing improves SSD performance effectively by reducing garbage collection overhead. Especially, our evaluation with the TPC-C benchmark on the Oracle DBMS shows that Zombie Chasing enhances the Transactions Per Second (TPS) value by up to 22% with negligible overhead. Youngjae Lee, Jin-Soo Kim 0001, Sang-Won Lee 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 4 |
| 2014 | Accelerating External Sorting via On-the-fly Data Merge in Active SSDs
Young-Sik Lee, Luis Cavazos Quero, Youngjae Lee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
HotStorage | 5 |
| 2014 | Transportation-network-inspired network-on-chipabstractA cost-efficient network-on-chip is needed in a scalable many-core systems. Recent multicore processors have leveraged a ring topology and hierarchical ring can increase scalability but presents different challenges, including higher hop count and global ring bottleneck. In this work, we describe a hierarchical ring topology that we refer to as a transportation-network-inspired network-on-chip (tNoC) that leverages principles from transportation network systems. In particular, we propose a novel hybrid flow control for hierarchical ring topology to scale the topology efficiently. The flow control is hybrid in that the channels are allocated on flit granularity while the buffers are allocated on packet granularity. The hybrid flow control enables a simplified router microarchitecture (to minimize per-hop latency) as router input buffers are minimized and buffers are pushed to the edges, either at the output ports or at the hub routers that interconnect the local rings to the global ring - while still supporting virtual channels to avoid protocol deadlock. We also describe a packet-quota-system (PQS) and a separate credit network that provide congestion management, support prioritized arbitration in the network, and provide support for multiflit packets. A detailed evaluation of a 64-core CMP shows that the tNoC improves performance by up to 21% compared with a baseline, buffered hierarchical ring topology while reducing NoC energy by 51%. Hanjoon Kim, Gwangsun Kim, Seung Ryoul Maeng, Hwasoo Yeo, John Kim 0001 |
HPCA | 3 |
| 2014 | Large-scale incremental processing with MapReduce
DaeWoo Lee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 3 |
| 2014 | System-Wide Cooperative Optimization for NAND Flash-Based Mobile SystemsabstractNAND flash memory has become an essential storage medium for various mobile devices, but it has some idiosyncrasies, such as out-of-place updates and bulk erase operations, which impair the I/O performance of those devices. In particular, the random write performance is strongly influenced by the overhead of a Flash Translation Layer (FTL) that hides the idiosyncrasies of NAND flash memory. To reduce the FTL overhead, operating systems need to be adapted for FTL, but widely used mobile operating systems still mainly adopt algorithms designed for traditional hard disk drives. Although there have been recent studies on rearranging write patterns into a sequential form in the operating system, these approaches fail to produce sequential write patterns under complicated workloads, and FTL still suffers from significant garbage collection overhead. If the operating system can be made aware of the write patterns that FTL requires, the overhead can be alleviated even under random write workloads. In this paper, we propose a system-wide cooperative optimization scheme, where the operating system communicates with the underlying FTL and generates write patterns that FTL can exploit to reduce the overhead. The proposed scheme was implemented on a real mobile device, and the experimental results show that the proposed scheme constantly improves performance under diverse workloads. Hyotaek Shim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 3 |
| 2013 | Demand-based coordinated scheduling for SMP VMsabstractAs processor architectures have been enhancing their computing capacity by increasing core counts, independent workloads can be consolidated on a single node for the sake of high resource efficiency in data centers. With the prevalence of virtualization technology, each individual workload can be hosted on a virtual machine for strong isolation between co-located workloads. Along with this trend, hosted applications have increasingly been multithreaded to take advantage of improved hardware parallelism. Although the performance of many multithreaded applications highly depends on communication (or synchronization) latency, existing schemes of virtual machine scheduling do not explicitly coordinate virtual CPUs based on their communication behaviors. Hwanju Kim, Jinkyu Jeong, Joonwon Lee, Seung Ryoul Maeng |
ASPLOS | 5 |
| 2013 | Isolated Mini-domain for Trusted Cloud ComputingabstractOn the cloud system, guest domains for cloud customers can be attacked by one of administrators with privilege or remote hackers who can compromise management tools. Therefore, the customers need a guarantee that their domains run on the secure environment with a protection against them. In this paper, we examine the security issues incurred by I/O model of hyper visors with a management domain, and propose an isolated mini-domain to protect the guest domains under the untrustworthy environment by addressing those issues. Jongse Park, Jinho Seol, Seung Ryoul Maeng |
CCGRID | 4 |
| 2013 | Towards Assurance of Availability in Virtualized Cloud SystemabstractCloud computing naturally shares physical resources, which is provided by virtualization technology. In contrast to the strength of virtualization, the serious security concerns raises due to the complexity of virtualization. It will place obstacles on the road to spread of cloud computing. In order to resolve the concern, many researchers focus on guaranteeing confidentiality and integrity of virtual machines except availability. We can support SLA (Service Level Agreement) as a part of assuring availability of virtual machines by SMM (System Management Mode)- based mechanism. Even if hypervisor is compromised, the mechanism can detect violations of SLA in virtualized cloud system. Seongwook Jin, Jinho Seol, Seung Ryoul Maeng |
CCGRID | 3 |
| 2013 | Secure Storage Service for IaaS Cloud UsersabstractCloud computing enables to reduce operating costs and maximize resource utilization. However, current cloud infrastructure is insufficient to guarantee the confidentiality of classified information for cloud users because one of administrators with privilege or remote hackers compromising management tools can leak the information. This paper addresses storage issues in cloud computing and proposes a secure storage service where stored user data are protected even against a malicious administrator and compromised software. We describe the architecture of proposed design and discuss the security issues of the design. Jinho Seol, Seongwook Jin, Seung Ryoul Maeng |
CCGRID | 3 |
| 2013 | OSSD: A case for object-based solid state drivesabstractThe notion of object-based storage devices (OSDs) has been proposed to overcome the limitations of the traditional block-level interface which hinders the development of intelligent storage devices. The main idea of OSD is to virtualize the physical storage into a pool of objects and offload the burden of space management into the storage device. We explore the possibility of adopting this idea for solid state drives (SSDs). The proposed object-based SSDs (OSSDs) allow more efficient management of the underlying flash storage, by utilizing object-aware data placement, hot/cold data separation, and QoS support for prioritized objects. We propose the software stack of OSSDs and implement an OSSD prototype using an iSCSI-based embedded storage device. Our evaluations with various scenarios show the potential benefits of the OSSD architecture. Young-Sik Lee, Sang-Hoon Kim, Jin-Soo Kim 0001, Jaesoo Lee, Chanik Park, Seung Ryoul Maeng |
MSST | 6 |
| 2013 | μ*-Tree: An Ordered Index Structure for NAND Flash Memory with Adaptive Page Layout SchemeabstractAs NAND flash memory is gaining popularity as a storage medium for mobile embedded devices, many flash-aware file systems, flash-aware DBMSes, and flash translation layers (FTLs) require an flash-efficient index structure. This paper proposes a novel index structure called μ*-Tree which natively works on NAND flash memory, aiming at improving performance over B+-Tree. μ*-Tree stores all the nodes along the path from the root to the leaf into a single flash memory page in order to minimize the number of flash write operation when a node is updated. Furthermore, μ*-Tree has an adaptive page layout scheme which dynamically adjusts the page layout according to the workload characteristics on-the-fly. μ*-Tree also allows flash pages with different page layouts to coexist in the same tree. Our evaluation results with real workload traces show that μ*-Tree outperforms B+-Tree by up to 55 percent in terms of the time needed for flash operations. With a small in-memory cache of 32 KB, μ*-Tree improves the overall performance by up to five times compared to B+-Tree with the same cache size. Jung-Sang Ahn, Dongwon Kang, Da Woon Jung 0001, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 5 |
| 2013 | Rigorous rental memory management for embedded systemsabstractMemory reservation in embedded systems is a prevalent approach to provide a physically contiguous memory region to its integrated devices, such as a camera device and a video decoder. Inefficiency of the memory reservation becomes a more significant problem in emerging embedded systems, such as smartphones and smart TVs. Many ways of using these systems increase the idle time of their integrated devices, and eventually decrease the utilization of their reserved memory. In this article, we propose a scheme to minimize the memory inefficiency caused by the memory reservation. The memory space reserved for a device can be rented for other purposes when the device is not active. For this scheme to be viable, latencies associated with reallocating the memory space should be minimal. Volatile pages are good candidates for such page reallocation since they can be reclaimed immediately as they are needed by the original device. We also provide two optimization techniques, lazy-migration and adaptive-activation. The former increases the lowered utilization of the rental memory by our volatile page allocations, and the latter saves active pages in the rental memory during the reallocation. We implemented our scheme on a smartphone development board with the Android Linux kernel. Our prototype has shown that the time for the return operation is less than 0.77 seconds in the tested cases. We believe that this time is acceptable to end-users in terms of transparency since the time can be hidden in application initialization time. The rental memory also brings throughput increases ranging from 2% to 200% based on the available memory and the applications' memory intensiveness. Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2012 | DaaC: device-reserved memory as an eviction-based file cacheabstractMost embedded systems require contiguous memory space to be reserved for each device, which may lead to memory under-utilization. Although several approaches have been proposed to address this issue, they have limitations of either inefficient memory usage or long latency for switching the reserved memory space between a device and general-purpose uses. Jinkyu Jeong, Hwanju Kim, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
CASES | 5 |
| 2012 | Locality-aware dynamic VM reconfiguration on MapReduce cloudsabstractCloud computing based on system virtualization, has been expanding its services to distributed data-intensive platforms such as MapReduce and Hadoop. Such a distributed platform on clouds runs in a virtual cluster consisting of a number of virtual machines. In the virtual cluster, demands on computing resources for each node may fluctuate, due to data locality and task behavior. However, current cloud services use a static cluster configuration, fixing or manually adjusting the computing capability of each virtual machine (VM). The fixed homogeneous VM configuration may not adapt to changing resource demands in individual nodes. Jongse Park, DaeWoo Lee, Bokyeong Kim, Jaehyuk Huh 0001, Seung Ryoul Maeng |
HPDC | 5 |
| 2012 | Scheduler support for video-oriented multimedia on client-side virtualizationabstractVirtualization has recently been adopted for client devices to provide strong isolation between services and efficient manageability. Even though multimedia service is not rare for the devices, the virtual machine hosting this service is not guaranteed to receive proper scheduling support from the underlying hypervisor. The quality of multimedia service is often compromised when several virtual machines compete for computing power. This paper presents a new scheduling scheme for the hypervisor to transparently identify if the workload handles multimedia and to provide proper scheduling supports. An implementation of our scheme has shown that the virtual machine hosting a video-oriented application receives propoer CPU scheduling even when other virtual machines host CPU intensive workloads. Hwanju Kim, Jinkyu Jeong, Jeaho Hwang, Joonwon Lee, Seung Ryoul Maeng |
MMSys | 5 |
| 2012 | FlashLight: A Lightweight Flash File System for Embedded SystemsabstractA very promising approach for using NAND flash memory as a storage medium is a flash file system. In order to design a higher-performance flash file system, two issues should be considered carefully. One issue is the design of an efficient index structure that contains the locations of both files and data in the flash memory. For large-capacity storage, the index structure must be stored in the flash memory to realize low memory consumption; however, this may degrade the system performance. The other issue is the design of a novel garbage collection (GC) scheme that reclaims obsolete pages. This scheme can induce considerable additional read and write operations while identifying and migrating valid pages. In this article, we present a novel flash file system that has the following features: ( i ) a lightweight index structure that introduces the hybrid indexing scheme and intra-inode index logging , and ( ii ) an efficient GC scheme that adopts a dirty list with an on-demand GC approach as well as fine-grained data separation and erase-unit data allocation . We implemented FlashLight in a Linux OS with kernel version 2.6.21 on an embedded device. The experimental results obtained using several benchmark programs confirm that FlashLight improves the performance by up to 27.4% over UBIFS by alleviating index management and GC overheads by up to 33.8%. Jaegeuk Kim, Hyotaek Shim, Seon-Yeong Park, Seung Ryoul Maeng, Jin-Soo Kim 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2012 | Reducing communication costs in collective I/O in multi-core cluster systems with non-exclusive scheduling
Kwangho Cha, Seung Ryoul Maeng |
J. Supercomput. | 2 |
| 2011 | Virtualizing performance asymmetric multi-core systemsabstractPerformance-asymmetric multi-cores consist of heterogeneous cores, which support the same ISA, but have different computing capabilities. To maximize the throughput of asymmetric multi-core systems, operating systems are responsible for scheduling threads to different types of cores. However, system virtualization poses a challenge for such asymmetric multi-cores, since virtualization hides the physical heterogeneity from guest operating systems. In this paper, we explore the design space of hypervisor schedulers for asymmetric multi-cores, which do not require asymmetry-awareness from guest operating systems. The proposed scheduler characterizes the efficiency of each virtual core, and map the virtual core to the most area-efficient physical core. In addition to the overall system throughput, we consider two important aspects of virtualizing asymmetric multi-cores: performance fairness among virtual machines and performance scalability for changing availability of fast and slow cores. Youngjin Kwon, Changdae Kim 0001, Seung Ryoul Maeng, Jaehyuk Huh 0001 |
ISCA | 3 |
| 2011 | Cost optimized provisioning of elastic resources for application workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 4 |
| 2011 | BTS: Resource capacity estimate for time-targeted science workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Ewa Deelman, Seung Ryoul Maeng |
J. Parallel Distributed Comput. | 5 |
| 2011 | Transparently bridging semantic gap in CPU management for virtualized environments
Hwanju Kim, Hyeontaek Lim, Jinkyu Jeong, Heeseung Jo, Joonwon Lee, Seung Ryoul Maeng |
J. Parallel Distributed Comput. | 6 |
| 2011 | Energy Reduction in Consolidated Servers through Memory-Aware Virtual Machine SchedulingabstractIncreasing energy consumption in server consolidation environments leads to high maintenance costs for data centers. Main memory, no less than processor, is a major energy consumer in this environment. This paper proposes a technique for reducing memory energy consumption using virtual machine scheduling in multicore systems. We devise several heuristic scheduling algorithms by using a memory power simulator, which we designed and implemented. We also implement the biggest cover set first (BCSF) scheduling algorithm in the working server system. Through extensive simulation and implementation experiments, we observe the effectiveness of the memory-aware virtual machine scheduling in saving memory energy. In addition, we find out that power-aware memory management is essential to reduce the memory energy consumption. Jae-Wan Jang, Myeongjae Jeon, Hyo-Sil Kim, Heeseung Jo, Jin-Soo Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 6 |
| 2010 | HAMA: An Efficient Matrix Computation with the MapReduce FrameworkabstractVarious scientific computations have become so complex, and thus computation tools play an important role. In this paper, we explore the state-of-the-art framework providing high-level matrix computation primitives with MapReduce through the case study approach, and demonstrate these primitives with different computation engines to show the performance and scalability. We believe the opportunity for using MapReduce in scientific computation is even more promising than the success to date in the parallel systems literature. Sangwon Seo, Edward J. Yoon, Jaehong Kim 0006, Seongwook Jin, Jin-Soo Kim 0001, Seung Ryoul Maeng |
CloudCom | 6 |
| 2010 | An adaptive partitioning scheme for DRAM-based cache in Solid State DrivesabstractRecently, NAND flash-based Solid State Drives (SSDs) have been rapidly adopted in laptops, desktops, and server storage systems because their performance is superior to that of traditional magnetic disks. However, NAND flash memory has some limitations such as out-of-place updates, bulk erase operations, and a limited number of write operations. To alleviate these unfavorable characteristics, various techniques for improving internal software and hardware components have been devised. In particular, the internal device cache of SSDs has a significant impact on the performance. The device cache is used for two main purposes: to absorb frequent read/write requests and to store logical-to-physical address mapping information. In the device cache, we observed that the optimal ratio of the data buffering and the address mapping space changes according to workload characteristics. To achieve optimal performance in SSDs, the device cache should be appropriately partitioned between the two main purposes. In this paper, we propose an adaptive partitioning scheme, which is based on a ghost caching mechanism, to adaptively tune the ratio of the buffering and the mapping space in the device cache according to the workload characteristics. The simulation results demonstrate that the performance of the proposed scheme approximates the best performance. Hyotaek Shim, Bon-Keun Seo, Jin-Soo Kim 0001, Seung Ryoul Maeng |
MSST | 4 |
| 2010 | Transparent Fault Tolerance of Device Drivers for Virtual MachinesabstractIn a consolidated server system using virtualization, physical device accesses from guest virtual machines (VMs) need to be coordinated. In this environment, a separate driver VM is usually assigned to this task to enhance reliability and to reuse existing device drivers. This driver VM needs to be highly reliable, since it handles all the I/O requests. This paper describes a mechanism to detect and recover the driver VM from faults to enhance the reliability of the whole system. The proposed mechanism is transparent in that guest VMs cannot recognize the fault and the driver VM can recover and continue its I/O operations. Our mechanism provides a progress monitoring-based fault detection that is isolated from fault contamination with low monitoring overhead. When a fault occurs, the system recovers by switching the faulted driver VM to another one. The recovery is performed without service disconnection or data loss and with negligible delay by fully exploiting the I/O structure of the virtualized system. Heeseung Jo, Hwanju Kim, Jae-Wan Jang, Joonwon Lee, Seung Ryoul Maeng |
IEEE Trans. Computers | 5 |
| 2009 | A buffer replacement algorithm exploiting multi-chip parallelism in solid state disksabstractSolid State Disks (SSDs) are superior to magnetic disks from a performance point of view due to the favorable features of NAND flash memory. Furthermore, thanks to improvement on flash memory density and adopting a multi-chip architecture, SSDs replace magnetic disks rapidly. Most previous studies have been conducted for enhancing the performance of SSDs, but these studies have been worked on the assumption that the operation unit of a host interface is the same as the operation unit of NAND flash memory, where it is needless to give consideration to partially-filled pages. In this paper, we analyze the overhead caused by the partially-filled pages, and propose a buffer replacement algorithm exploiting multi-chip parallelism to enhance the write performance. Our simulation results show that the proposed algorithm improves the write performance by up to 30% over existing approaches. Jinho Seol, Hyotaek Shim, Jaegeuk Kim, Seung Ryoul Maeng |
CASES | 4 |
| 2009 | HPMR: Prefetching and pre-shuffling in shared MapReduce computation environmentabstractMapReduce is a programming model that supports distributed and parallel processing for large-scale data-intensive applications such as machine learning, data mining, and scientific simulation. Hadoop is an open-source implementation of the MapReduce programming model. Hadoop is used by many companies including Yahoo!, Amazon, and Facebook to perform various data mining on large-scale data sets such as user search logs and visit logs. In these cases, it is very common to share the same computing resources by multiple users due to practical considerations about cost, system utilization, and manageability. However, Hadoop assumes that all cluster nodes are dedicated to a single user, failing to guarantee high performance in the shared MapReduce computation environment. In this paper, we propose two optimization schemes, prefetching and pre-shuffling, which improve the overall performance under the shared environment while retaining compatibility with the native Hadoop. The proposed schemes are implemented in the native Hadoop-0.18.3 as a plug-in component called HPMR (High Performance MapReduce Engine). Our evaluation on the Yahoo!Grid platform with three different workloads and seven types of test sets from Yahoo! shows that HPMR reduces the execution time by up to 73%. Sangwon Seo, Ingook Jang, Kyungchang Woo, Inkyo Kim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
CLUSTER | 6 |
| 2009 | FTL design exploration in reconfigurable high-performance SSD for server applicationsabstractSolid-state disks (SSDs) are becoming widely used in personal computers and are expected to replace a great portion of magnetic disks in servers and supercomputers. Although many high-speed SSDs are present in the market, both the design of hardware architecture and the details of the flash translation layer (FTL) are not well known. Meanwhile, in the systems requiring high-end storages, specially tuned SSDs can perform better than the generic ones, because the applications in such environment are usually fixed. Ji-Yong Shin, Zenglin Xia, Ningyi Xu, Xiongfei Cai, Seung Ryoul Maeng, Feng-Hsiung Hsu |
ICS | 6 |
| 2008 | Context-aware address translation for high performance SMP cluster systemabstractUser-level communication allows an application process to access the network interface directly. Bypassing the kernel requires that a user process accesses the network interface using its own virtual address which should be translated to a physical address. A small caching structure which is similar to the hardware TLB on the host processor has been used to cache the mappings between virtual and physical addresses on the network interface memory. In this study, we propose a new TLB architecture for the network interface. The proposed architecture splits an original caching structure into as many partitions as the number of processors on the SMP system and assigns a separate partition to each application process. In addition, the architecture becomes aware of user contexts and switches the content of caching structure in accordance with context switching. According to our experiments, our scheme achieves significant reduction in application execution time compared to the previous approach. Moon-Sang Lee, Joonwon Lee, Seung Ryoul Maeng |
CLUSTER | 3 |
| 2008 | RMA: A Read Miss-Based Spin-Down Algorithm using an NV cacheabstractIt is an important issue to reduce the power consumption of a hard disk that takes a large amount of computer systempsilas power. As a new trend, an NV cache is used to make a disk spin down longer by servicing read/write requests instead of the disk. During the spin-down periods, write requests can be simply handled by write buffering, but read requests are still the main cause of initiating spin-ups because of a low hit ratio in the NV cache. Even when there is no user activity, read requests can be frequently generated by running applications and system services, hindering the spin-down. In this paper, we propose new NV cache policies: active write caching to reduce or to delay spin-ups caused by read misses during spin-down periods and a read miss-based spin-down algorithm to extend the spin-down periods, exploiting the NV cache effectively. Our policies reduce the power consumption of a hard disk by up to 50.1% with a 512 MB NV cache, compared with preceding approaches. Hyotaek Shim, Jaegeuk Kim, Da Woon Jung 0001, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICCD | 5 |
| 2008 | snapPVFS: Snapshot-Able Parallel Virtual File SystemabstractIn this paper, we propose a modified parallel virtual file system that provides snapshot functionality. Because typical file systems are exposed to various failures, taking a snapshot is a good way to enhance the reliability of file systems. The PVFS, which is one of the famous parallel file systems deployed in cluster systems, is vulnerable to system failures or users¿ mistakes; however, there is a scarcity of research on snapshots or online backup for the PVFS. Because a PVFS consists of multiple servers on a network, snapshots should be generated properly in each server in the system. Furthermore, before snapshots are generated, the status of each PVFS server must be checked to guarantee sound operation. To demonstrate our approach, we implemented two prototypes of a snapshot-able PVFS (snapPVFS). The performance measurements indicate that an administrator can take snapshots of an entire parallel file system and properly access any previous versions of files or directories in the future without serious performance degradation. Kwangho Cha, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICPADS | 3 |
| 2008 | Efficient Metadata Management for Flash File SystemsabstractNAND flash memory becomes one of the most popular storage for portable embedded systems. Although many flash-aware file systems, such as JFFS2 and YAFFS2, were proposed, the large memory consumption and the long mount delay have been serious obstacles for large-capacity NAND flash memory. In this paper, we present a new flash-aware file system called DFFS (direct flash file system) which fetches only the needed metadata on demand from flash memory. In addition, DFFS employs two novel metadata management schemes, inode embedding scheme and hybrid inode indexing scheme, to improve the performance of metadata operations. Comprehensive evaluation results using microbench- mark, postmark, and Linux kernel compilation trace, show that DFFS has comparable performance to JFFS2 and YAFFS2, while achieving a small memory footprint and instant mount time. Jaegeuk Kim, Heeseung Jo, Hyotaek Shim, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ISORC | 5 |
| 2007 | A runtime resolution scheme for priority boost conflict in implicit coscheduling
Jung-Lok Yu, Jin-Soo Kim 0001, Seung Ryoul Maeng |
J. Supercomput. | 3 |
| 2006 | Home-based Cooperative Cache for parallel I/O applications
In-Chul Hwang, Seung Ryoul Maeng, Jung Wan Cho |
Future Gener. Comput. Syst. | 2 |
| 2006 | Log-Based Rollback Recovery without Checkpoints of Shared Memory in Software DSM
Seung Ryoul Maeng |
J. Supercomput. | 2 |
| 2005 | Impact of Exploiting Load Imbalance on Coscheduling in Workstation ClustersabstractImplicit coscheduling is known to be an effective technique to improve the performance of parallel workloads in time-sharing clusters. However, implicit coscheduling still does not take into consideration the system behavior like load imbalance that severely affects cluster utilization. In this paper, we propose the use of global information to enhance the existing implicit coscheduling schemes. We also introduce a novel coscheduling approach - named PROC (process reordering-based coscheduling) - based on process reordering exploiting global load imbalance information to coordinate communicating processes. The results obtained from an in-depth simulation study show that our approach significantly outperforms previous ones (by up to 38.4%) by reducing the idle time (by up to 86.9%) and spin time (by up to 36.2%) caused by the load imbalance. Jung-Lok Yu, Driss Azougagh, Jin-Soo Kim 0001, Seung Ryoul Maeng |
ICPP | 4 |
| 2004 | Cluster computing environment supporting single system imageabstractSingle system image (SSl) systems have been the mainstay of high-performance computing for many years. SSI requires the integration and aggregation of all types of resources in a cluster to present a single interface to users. We describe a cluster computing environment supporting SSI, constructed through three components: single process space (SPS), process migration, and dynamic load balancing. These components attempt to share all available resources in the cluster among all executing processes, so that the cluster operates like a single node with much more computing power. The most important goal is to combine these constructs in innovative ways for building cluster computing environment for SSI, as well as individually take an approach to improve performance or functionality. Our implementation of process migration has the capability of resolving broken pipe problems and bind errors on server socket reconstruction. We realize SPS based on block PID allocation. We also designed and implemented a dynamic load balancing scheme which resolves the limitations of our previous work by continuously tracing the job resource usage at runtime. The experimental results show that these three constructs for SSI clusters realized scalability, functionality and performance improvement. The cluster computing environment allows these constructs to cooperate implicitly so that they create a synergy effect at the SSI cluster system level and successfully provide a single system image to users and administrators. DaeWoo Lee, Seung Ryoul Maeng |
CLUSTER | 3 |
| 2003 | Improving Performance of a Dynamic Load Balancing System by Using Number of Effective TasksabstractEfficient resource usage is a key to achieving better performance in cluster systems. Previously, most research in this area has focused on balancing the load if each node to use the resources of an entire system more effectively. However, we can achieve further improvement in performance when the load balancing system considers the resource requirement according to the task being assigned. This kind of load balancing system, known as an initial job placement system, requires knowledge of the resource usage of a task in order to fit the job to the most suitable node. Since the initial placement requires that the tasks be scheduled before execution, all resource usage must be provided in terms of the prediction. This approach can severely affect the execution time when it uses an inaccurate prediction. We propose a novel load metric termed number of effective tasks in order to resolve the problem arising from inaccurate predictions. Thus, the initial job placement system can work without knowing job resource usage in priori. Simulation results show that the system incurs 11% shorter execution time that the conventional approach using historical behavior-based estimates. Jung-Lok Yu, Hojoong Kim, Seung Ryoul Maeng |
CLUSTER | 4 |
| 2003 | Lightweight Logging and Recovery for Distributed Shared Memory over Virtual Interface ArchitectureabstractAs software Distributed Shared Memory(DSM) systems become attractive on larger clusters, the focus of attention moves toward improving the reliability of systems. In this paper, we propose a lightweight logging scheme, called remote logging, and a recovery protocol for home-based DSM. Remote logging stores coherence-related data to the volatile memory of a remote node. The logging overhead can be moderated with high-speed system area network and user-level DMA operations supported by modern communication protocols. Remote logging tolerates multiple failures if the backup nodes of failed nodes are alive. It makes the reliability of DSM grow much higher. Experimental results show that our fault-tolerant DSM has low overhead compared to conventional stable logging and it can be effectively recovered from some concurrent failures. Seung Ryoul Maeng |
ISPDC | 3 |
| 2001 | Adaptive Prefetching Technique for Shared Virtual MemoryabstractThough shared virtual memory (SVM) systems promise low cost solutions for high performance computing, they suffer from long memory latencies. These latencies are usually caused by repetitive invalidation on shared data. Since shared data are accessed through synchronizations and the patterns by which threads synchronize are repetitive, a prefetching scheme based on such repetitiveness would reduce memory latencies. Based on this observation, we propose a prefetching technique which predicts future access behavior by analyzing access history per synchronization variable. Our technique was evaluated on an 8-node SVM system using the SPLASH-2 benchmark. The results show that our technique could achieve 34%-45% reduction in memory access latencies. Sang-Kwon Lee, Hee-Chul Yun, Joonwon Lee, Seung Ryoul Maeng |
CCGRID | 4 |
| 2001 | An Efficient Lock Protocol for Home-Based Lazy Release ConsistencyabstractHome-based lazy release consistency (HLRC) shows poor performance on lock based applications because of two reasons: a whole page is fetched on a page fault while actual modification is much smaller; and a home is at the fixed location while the access pattern is migratory. We present an efficient lock protocol for HLRC. In this protocol, the pages that are expected to be used by the acquirer are selectively updated using diffs. The diff accumulation problem is minimized by limiting the size of diffs to be sent for each page. Our protocol reduces the number of page faults inside critical sections because pages can be updated by applying locally stored diffs. This reduction yields the reduction of average lock waiting time and the reduction of message amount. The experiment with five applications shows that our protocol archives 2%-40% speedup against base HLRC for four applications. Hee-Chul Yun, Sang-Kwon Lee, Joonwon Lee, Seung Ryoul Maeng |
CCGRID | 4 |
| 2001 | An Efficient Implementation of Virtual Interface Architecture using Adaptive Transfer Mechanism on MyrinetabstractUser-level communication is investigated by many researchers, in order to resolve the performance degradation of cluster systems due to inefficient communication protocols. It removes the kernel intervention from the critical communication path. Intel, Microsoft and Compaq introduced the Virtual Interface Architecture (VIA), a standard for user-level communication. However, the existing VIA implementation shows low performance in transferring small messages, because it uses a single mechanism to transfer messages without regard to their message size. We implement a high performance VIA, KVIA (Kaist VIA). KVIA, based on descriptor and message size, dynamically selects a proper transfer mechanism. This implementation effectively handles not only large messages but also small messages. Thus, it can be better applied to the systems that frequently use small messages (e.g., lock protocols for software distributed shared memory). The performance of KVIA is reported using round-trip latency and one-way bandwidth. Our results show the round-trip latency of 40 micro-seconds and the maximum one-way bandwidth of 950 Mbits per second, which is about 74% of Myrinet link's peak bandwidth. Jung-Lok Yu, Moon-Sang Lee, Seung Ryoul Maeng |
ICPADS | 3 |
| 2000 | Dynamic code reservation multiple access for supporting various-bit-rate integrated traffics in slotted DS-CDMA systems
EuiHoon Jeong, Ara Ra Khil, Seung Ryoul Maeng |
Comput. Commun. | 4 |
| 2000 | Multistage ring network: An interconnection network for large scale shared memory multiprocessors
Dongho Yoo, Inyoung Park, Seung Ryoul Maeng |
J. Syst. Archit. | 3 |
| 1997 | An adaptive sequential prefetching scheme in shared-memory multiprocessorsabstractThe sequential prefetching scheme is a simple hardware controlled scheme, which exploits the sequentiality of memory accesses to predict which blocks will be read in the near future. We analyze the relationship between the sequentiality of application programs and the effectiveness of sequential prefetching on shared-memory multiprocessors. Also, we propose a simple hardware scheme which selects the prefetching degree on each miss by adding a small table (PDS: Prefetching Degree Selector) to the sequential prefetching scheme. This scheme could prefetch consecutive blocks aggressively for applications with high sequentiality and conservatively for applications with low sequentiality. Myoung Kwon Tcheun, Hyunsoo Yoon, Seung Ryoul Maeng |
ICPP | 3 |
| 1997 | Embedding of rings in 2-D meshes and tori with faulty nodes
Jinsoo Kim 0005, Seung Ryoul Maeng, Hyunsoo Yoon |
J. Syst. Archit. | 2 |
| 1997 | On the Correctness of Inside-Out Routing AlgorithmabstractRecently, a new routing algorithm called inside-out routing algorithm was proposed for routing an arbitrary permutation in the omega-based 2log/sub 2/ N stage networks. This paper discusses the problems of the inside-out routing algorithm and shows that the suggested condition for proper routing in the omega-omega network is insufficient. An extended necessary and sufficient condition for proper routing in the omega-omega network is also suggested. However, it is unknown if any permutation can be successfully routed by a heuristic algorithm which follows the condition. Thus, the rearrangeability of the omega-omega network still remains an open problem. Hyunsoo Yoon, Seung Ryoul Maeng |
IEEE Trans. Computers | 3 |
| 1996 | Drop-and-reroute: A new flow control policy for adaptive wormhole routing
Ji-Yun Kim, Hyunsoo Yoon, Seung Ryoul Maeng, Jung Wan Cho |
J. Syst. Archit. | 3 |
| 1995 | Drop-and-Reroute: A New flow Control Policy for Adaptive Wormhole Routing
Ji-Yun Kim, Hyunsoo Yoon, Seung Ryoul Maeng, Jung Wan Cho |
ICPP (1) | 3 |
| 1995 | Bit-permute multistage interconnection networks
Hyunsoo Yoon, Seung Ryoul Maeng |
Microprocess. Microprogramming | 3 |
| 1994 | A Pipelined Systolic Arrays Architecture for the Hierarchical Block-Matching AlgorithmabstractThis paper presents a pipelined architecture for the hierarchical block-matching motion estimation algorithm (HBMA). The hierarchical style leads to an enormous computation and complex data flow between hierarchy levels. Each stage of the proposed architecture consists of a systolic array for block-matching and an interpolation unit for bilinear interpolation. The interpolation unit regulates also the data flow suitable for fully synchronous operation. The performance analysis shows that the proposed architecture gains nearly linear speedup, thus making HBMA suitable for real time operation.> Hyung Chul Kim, Seung Ryoul Maeng |
ISCAS | 2 |
| 1994 | A new deadlock prevention scheme for nonminimal adaptive wormhole routing
Jai-Hoon Chung, Hyunsoo Yoon, Seung Ryoul Maeng |
Microprocess. Microprogramming | 3 |
| 1994 | MAKRS: a knowledge and belief representation system for multiple agents
Young Hoon Kim, Seung Ryoul Maeng, Jung Wan Cho |
Knowl. Based Syst. | 2 |
| 1994 | Eventor: An Authoring System for Interactive Multimedia Applications
Seong Bae Eun, Eun Suk No, Hyung Chul Kim, Hyunsoo Yoon, Seung Ryoul Maeng |
Multim. Syst. | 5 |
| 1993 | Specification of Multimedia Composition and a Visual Programming EnvironmentabstractMultimedia refers to the composition of multiple monomedia which should be synchronized temporally and spatially.Over the past few years, some trials to describe the composition and synchronization have been made in a variety of applications and these descriptions have been used as frameworks in their applications.Although conventional works have succeeded in describing the synchronizations well, we indicate that they do not deal with the interactivity required in Interactive Multimedia Applications(IMA) like coursewares and hypermedia systems.In this paper, we propose a new specification method based on Milner's Calculus of Communicating Systems(CCS) to cope with the interactivity.For showing the effectiveness of the specification, we design and implement a visual programming environment based on the specification mechanism, and propose a simple courseware as a programming example.Our approach has implications that the new specification mechanism can be adapted as a framework in various interactive applications and the visual programming environment acquires the benefits that it can handle user interactions and synchronizations with only visual expressions while additional texts for control commands should be augmented in conventional works. Seong Bae Eun, Eun Suk No, Hyung Chul Kim, Hyunsoo Yoon, Seung Ryoul Maeng |
ACM Multimedia | 5 |
| 1993 | Combining many-sorted logic and object-oriented programming
Byeong Man Kim, Kiyeol Ryu, Seung Ryoul Maeng, J. W. Cho |
Inf. Softw. Technol. | 3 |
| 1993 | Concurrency and inheritance in actor-based object-oriented languages
Kiyeol Ryu, Seung Ryoul Maeng, Jung Wan Cho |
J. Syst. Softw. | 2 |
| 1992 | A systolic array exploiting the inherent parallelisms of artificial neural networks
Jai-Hoon Chung, Hyunsoo Yoon, Seung Ryoul Maeng |
Microprocess. Microprogramming | 3 |
| 1992 | An efficient mapping of Boltzmann machine computations onto distributed-memory multiprocessors
D. H. Oh, Jong H. Nang, Hyunsoo Yoon, Seung Ryoul Maeng |
Microprocess. Microprogramming | 4 |
| 1992 | Specifying and inheriting concurrent objects
Kiyeol Ryu, Seung Ryoul Maeng |
Microprocess. Microprogramming | 2 |
| 1991 | A Systolic Array Exploiting the Inherent Parallelisms of Artificial Neural Networks
Jai-Hoon Chung, Hyunsoo Yoon, Seung Ryoul Maeng |
ICPP (1) | 3 |
| 1990 | Extended elman's recurrent neural network for syllable recognition
Yong Duk Cho, Ki Chul Kim, Hyunsoo Yoon, Seung Ryoul Maeng, Jung Wan Cho |
ICSLP | 4 |
| 1990 | Parallel simulation of multilayered neural networks on distributed-memory multiprocessors
Hyunsoo Yoon, Jong H. Nang, Seung Ryoul Maeng |
Microprocessing and Microprogramming | 3 |
| 1986 | A Parallel Execution Model of Logic Program Based on Dependency Relationship Graph
Seungbeom Kim, Seung Ryoul Maeng, Jung Wan Cho |
ICPP | 2 |