EDBT 2026 Demo / reviewers in the wild / expert
Yang-Suk Kee
dblp:81/5871
· DBLP profile ↗
26ranked-venue papers
9as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 9 first-authorDatabases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Computer networks · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Storage systems · 73% Cloud and datacenter computing · 10% Distributed systems · 9% | |
| Databases, data mining, and information retrieval
3 papers |
Transaction processing and concurrency control · 67% Query processing and optimization · 22% Indexing and storage engines · 10% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
flash and SSD |
0.6 | 3 | 2016 | SHARE Interface in Flash Storage for Relational and NoSQL Databases · SIGMOD Conference 2016 Durable write cache in flash memory SSD for relational and NoSQL databases · SIGMOD Conference 2014 Query processing on smart SSDs: opportunities and challenges · SIGMOD Conference 2013 |
Storage systems › storage devices › storage media
emerging storage devices |
0.4 | 1 | 2020 | Hybrid Data Reliability for Emerging Key-Value Storage Devices · FAST 2020 |
Storage systems
key-value storage |
0.4 | 1 | 2020 | Hybrid Data Reliability for Emerging Key-Value Storage Devices · FAST 2020 |
Transaction processing and concurrency control
atomicity |
0.2 | 1 | 2016 | SHARE Interface in Flash Storage for Relational and NoSQL Databases · SIGMOD Conference 2016 |
Storage systems
atomic writes |
0.2 | 1 | 2016 | SHARE Interface in Flash Storage for Relational and NoSQL Databases · SIGMOD Conference 2016 |
Storage systems
storage reliability |
0.2 | 1 | 2016 | SHARE Interface in Flash Storage for Relational and NoSQL Databases · SIGMOD Conference 2016 |
Query processing and optimization
analytical query processing |
0.2 | 1 | 2013 | Query processing on smart SSDs: opportunities and challenges · SIGMOD Conference 2013 |
Storage systems › computational storage
SmartSSD |
0.2 | 1 | 2013 | Query processing on smart SSDs: opportunities and challenges · SIGMOD Conference 2013 |
Cloud and datacenter computing
resource allocation |
0.1 | 2 | 2006 | Grid allocation and reservation - Improving grid resource allocation via integrated selection and binding · SC 2006 Robust Resource Allocation for Large-scale Distributed Shared Resource Environments · HPDC 2006 |
Distributed systems
fault tolerance |
0.1 | 2 | 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009 Robust Resource Allocation for Large-scale Distributed Shared Resource Environments · HPDC 2006 |
Distributed systems
grid computing |
0.1 | 2 | 2006 | Grid allocation and reservation - Improving grid resource allocation via integrated selection and binding · SC 2006 Realistic Modeling and Svnthesis of Resources for Computational Grids · SC 2004 |
Cloud and datacenter computing
workflow scheduling |
0.1 | 1 | 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault tolerance · SC 2009 |
Indexing and storage engines › storage management
storage engine |
0.1 | 1 | 2016 | SHARE Interface in Flash Storage for Relational and NoSQL Databases · SIGMOD Conference 2016 |
Transaction processing and concurrency control › OLTP
OLTP engine |
0.1 | 1 | 2014 | Durable write cache in flash memory SSD for relational and NoSQL databases · SIGMOD Conference 2014 |
Memory systems
processing-in-memory |
0.0 | 1 | 2013 | Query processing on smart SSDs: opportunities and challenges · SIGMOD Conference 2013 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2004 | Realistic Modeling and Svnthesis of Resources for Computational Grids · SC 2004 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.0 | 1 | 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster Systems · SC 2003 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster Systems · SC 2003 |
Memory systems › shared memory › distributed shared memory
software distributed shared memory |
0.0 | 1 | 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster Systems · SC 2003 |
Distributed systems
distributed coordination |
0.0 | 1 | 2006 | Grid allocation and reservation - Improving grid resource allocation via integrated selection and binding · SC 2006 |
Distributed systems › distributed resource management
resource discovery |
0.0 | 1 | 2004 | Realistic Modeling and Svnthesis of Resources for Computational Grids · SC 2004 |
Parallel and multicore computing › parallel programming models
hybrid programming models |
0.0 | 1 | 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster Systems · SC 2003 |
Methods — techniques the papers use, named apart from their topics
replication · 0.4erasure coding · 0.4tantalum capacitor durable cache · 0.4firmware features · 0.4query pushdown · 0.3integrated selection and binding · 0.1virtualized reservations · 0.1resource selection · 0.1composition operators · 0.1distribution fitting · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Hybrid Data Reliability for Emerging Key-Value Storage Devices
Rekha Pitchumani, Yang-Suk Kee |
FAST | 2 |
| 2019 | Towards building a high-performance, scale-in key-value storage systemabstractKey-value stores are widely used as storage backends, due to their simple, yet flexible interface for cache, storage, file system, and database systems. However, when used with high performance NVMe devices, their high compute requirements for data management often leave the device bandwidth under-utilized. This leads to a performance mismatch of what the device is capable of delivering and what it actually delivers, and the gains derived from high speed NVMe devices is nullified. In this paper, we introduce KV-SSD (Key-Value SSD) as a key technology in a holistic approach to overcome such performance imbalance. KV-SSD provides better scalability and performance by simplifying the software storage stack and consolidating redundancy, thereby lowering the overall CPU usage and releasing the memory to user applications. We evaluate the performance and scalability of KV-SSDs over state-of-the-art software alternatives built for traditional block SSDs. Our results show that, unlike traditional key-value systems, the overall performance ofKV-SSD scales linearly, and delivers 1.6 to 57x gains depending on the workload characteristics. Yangwook Kang, Rekha Pitchumani, Pratik Mishra, Yang-Suk Kee, Francisco Londono, Sangyoon Oh 0002, Jongyeol Lee, Daniel D. G. Lee |
SYSTOR | 4 |
| 2018 | Crail-KV: A High-Performance Distributed Key-Value Store Leveraging Native KV-SSDs over NVMe-oFabstractA Key-Value SSD (KV-SSD) is a new type of storage device that natively exposes a key-value interface. In this paper, we leverage KV-SSDs to develop new techniques to remove unnecessary layers of indirection traditionally imposed by block devices on distributed storage systems. Specifically, we extend the Crail distributed system [1] to leverage the KV-SSD's native key-value interface exposing it directly to clients through the NVMe-oF protocol. This architectural change simplifies key-value metadata management, as the metadata manager need only track key-value tuples rather than files comprised of blocks. These changes enable fewer RPCs and require less memory for metadata management, resulting in a performance improvement of up to 5x. Tim Bisson, Ke Chen 0020, Changho Choi, Vijay Balakrishnan, Yang-Suk Kee |
IPCCC | 5 |
| 2016 | SSD in-storage computing for list intersectionabstractRecently, there has been a renewed interest of in-storage computing in the context of solid state drives (SSDs), called "Smart SSDs." Smart SSDs allow application-specific code to execute inside SSDs. This allows applications to take advantage of the high internal bandwidth that Smart SSDs provide. This work studies the offloading of list intersection into Smart SSDs, because intersection is prominent in both search engines and analytics queries. Furthermore, intersection is interesting because the algorithms are more complex than plain scans; they are affected by multiple parameters, as we show, and provide lessons that can be used in other operations also. Jianguo Wang 0001, Dongchul Park, Yang-Suk Kee, Yannis Papakonstantinou, Steven Swanson |
DaMoN | 3 |
| 2016 | SHARE Interface in Flash Storage for Relational and NoSQL DatabasesabstractDatabase consistency and recoverability require guaranteeing write atomicity for one or more pages. However, contemporary database systems consider write operations non-atomic. Thus, many database storage engines have traditionally relied on either journaling or copy-on-write approaches for atomic propagation of updated pages to the storage. This reliance achieves write atomicity at the cost of various write amplifications such as redundant writes, tree-wandering, and compaction. This write amplification results in reduced performance and, for flash storage, accelerates device wear-out. In this paper, we propose a flash storage interface, SHARE. Being able to explicitly remap the address mapping inside flash storage using SHARE interface enables host-side database storage engines to achieve write atomicity without causing write amplification. We have implemented SHARE on a real SSD board, OpenSSD, and modified MySQL/InnoDB and Couchbase NoSQL storage engines to make them compatible with the extended SHARE interface. Our experimental results show that this SHARE-based MySQL/InnoDB and Couchbase configurations can significantly boost database performance. In particular, the inevitable and costly Couchbase compaction process can complete without copying any data pages. Gi-Hwan Oh, Chiyoung Seo, Ravi Mayuram, Yang-Suk Kee, Sang-Won Lee 0001 |
SIGMOD Conference | 4 |
| 2015 | Early experience with optimizing I/O performance using high-performance SSDs for in-memory cluster computingabstractThis paper describes our experience with storage optimization that utilizes cost-effective PCIe solid-state drives (SSDs) to improve the overall performance of a Spark framework. A key problem we address is the limited memory system performance. In particular, we adopt high-performance SSDs to alleviate the saturated DRAM bandwidth and its limited capacity. We utilize SSDs to store shuffle data and persisted RDDs. As a result, the overall performance improves due to the larger capacity of SSDs and the increased bandwidth provided by SSDs while alleviating memory contentions. Our experiments show that we can improve the performance of data-intensive applications by 23.1% on average, compared to the performance of the memory-only approach. To our knowledge, this is the first work to demonstrate performance optimizations using PCIe SSDs on Spark. I. Stephen Choi, Weiqing Yang, Yang-Suk Kee |
IEEE BigData | 3 |
| 2015 | Optimizing the Hadoop MapReduce Framework with high-performance storage devices
Sangwhan Moon, Jaehwan Lee 0001, Xiling Sun, Yang-Suk Kee |
J. Supercomput. | 4 |
| 2014 | Introducing SSDs to the Hadoop MapReduce FrameworkabstractSolid State Drive (SSD) cost-per-bit continues to decrease. Consequently, system architects increasingly consider replacing Hard Disk Drives (HDDs) with SSDs to accelerate Hadoop MapReduce processing. When attempting this, system architects usually realize that SSD characteristics and today's Hadoop framework exhibit mismatches that impede indiscriminate SSD integration. Hence, cost-effective SSD utilization has proved challenging within many Hadoop environments. This paper compares SSD performance to HDD performance within a Hadoop MapReduce framework. It identifies extensible best practices that can exploit SSD benefits within Hadoop frameworks when combined with high network bandwidth and increased parallel storage access. Terasort benchmark results demonstrate that SSDs presently deliver significant cost-effectiveness when they store intermediate Hadoop data, leaving HDDs to store Hadoop Distributed File System (HDFS) source data. Sangwhan Moon, Jaehwan Lee 0001, Yang-Suk Kee |
IEEE CLOUD | 3 |
| 2014 | Durable write cache in flash memory SSD for relational and NoSQL databasesabstractIn order to meet the stringent requirements of low latency as well as high throughput, web service providers with large data centers have been replacing magnetic disk drives with flash memory solid-state drives (SSDs). They commonly use relational and NoSQL database engines to manage OLTP workloads in the warehouse-scale computing environments. These modern database engines rely heavily on redundant writes and frequent cache flushes to guarantee the atomicity and durability of transactional updates. This has become a serious bottleneck of performance in both relational and NoSQL database engines. This paper presents a new SSD prototype called DuraSSD equipped with tantalum capacitors. The tantalum capacitors make the device cache inside DuraSSD durable, and additional firmware features of DuraSSD take advantage of the durable cache to support the atomicity and durability of page writes. It is the first time that a flash memory SSD with durable cache has been used to achieve an order of magnitude improvement in transaction throughput without compromising the atomicity and durability. Considering that the simple capacitors increase the total cost of an SSD no more than one percent, DuraSSD clearly provides a cost-effective means for transactional support. DuraSSD is also expected to alleviate the problem of high tail latency by minimizing write stalls. Woon-Hak Kang, Sang-Won Lee 0001, Bongki Moon, Yang-Suk Kee, Moonwook Oh |
SIGMOD Conference | 4 |
| 2013 | Enabling cost-effective data processing with smart SSDabstractThis paper explores the benefits and limitations of in-storage processing on current Solid-State Disk (SSD) architectures. While disk-based in-storage processing has not been widely adopted, due to the characteristics of hard disks, modern SSDs provide high performance on concurrent random writes, and have powerful processors, memory, and multiple I/O channels to flash memory, enabling in-storage processing with almost no hardware changes. In addition, offloading I/O tasks allows a host system to fully utilize devices' internal parallelism without knowing the details of their hardware configurations. To leverage the enhanced data processing capabilities of modern SSDs, we introduce the Smart SSD model, which pairs in-device processing with a powerful host system capable of handling data-oriented tasks without modifying operating system code. By isolating the data traffic within the device, this model promises low energy consumption, high parallelism, low host memory footprint and better performance. To demonstrate these capabilities, we constructed a prototype implementing this model on a real SATA-based SSD. Our system uses an object-based protocol for low-level communication with the host, and extends the Hadoop MapReduce framework to support a Smart SSD. Our experiments show that total energy consumption is reduced by 50% due to the low-power processing inside a Smart SSD. Moreover, a system with a Smart SSD can outperform host-side processing by a factor of two or three by efficiently utilizing internal parallelism when applications have light trafic to the device DRAM under the current architecture. Yangwook Kang, Yang-Suk Kee, Ethan L. Miller, Chanik Park |
MSST | 2 |
| 2013 | Query processing on smart SSDs: opportunities and challengesabstractData storage devices are getting "smarter." Smart Flash storage devices (a.k.a. "Smart SSD") are on the horizon and will package CPU processing and DRAM storage inside a Smart SSD, and make that available to run user programs inside a Smart SSD. The focus of this paper is on exploring the opportunities and challenges associated with exploiting this functionality of Smart SSDs for relational analytic query processing. We have implemented an initial prototype of Microsoft SQL Server running on a Samsung Smart SSD. Our results demonstrate that significant performance and energy gains can be achieved by pushing selected query processing components inside the Smart SSDs. We also identify various changes that SSD device manufacturers can make to increase the benefits of using Smart SSDs for data processing applications, and also suggest possible research opportunities for the database community. Jaeyoung Do, Yang-Suk Kee, Jignesh M. Patel, Chanik Park, Kwanghyun Park 0001, David J. DeWitt |
SIGMOD Conference | 2 |
| 2011 | Cost optimized provisioning of elastic resources for application workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Seung Ryoul Maeng |
Future Gener. Comput. Syst. | 2 |
| 2011 | BTS: Resource capacity estimate for time-targeted science workflows
Eun-Kyu Byun, Yang-Suk Kee, Jin-Soo Kim 0001, Ewa Deelman, Seung Ryoul Maeng |
J. Parallel Distributed Comput. | 2 |
| 2009 | VGrADS: enabling e-Science workflows on grids and clouds with fault toleranceabstractToday's scientific workflows use distributed heterogeneous resources through diverse grid and cloud interfaces that are often hard to program. In addition, especially for time-sensitive critical applications, predictable quality of service is necessary across these distributed resources. VGrADS' virtual grid execution system (vgES) provides an uniform qualitative resource abstraction over grid and cloud systems. We apply vgES for scheduling a set of deadline sensitive weather forecasting workflows. Specifically, this paper reports on our experiences with (1) virtualized reservations for batchqueue systems, (2) coordinated usage of TeraGrid (batch queue), Amazon EC2 (cloud), our own clusters (batch queue) and Eucalyptus (cloud) resources, and (3) fault tolerance through automated task replication. The combined effect of these techniques was to enable a new workflow planning method to balance performance, reliability and cost considerations. The results point toward improved resource selection and execution management support for a variety of e-Science applications over grids and cloud systems. Lavanya Ramakrishnan, Charles Koelbel, Yang-Suk Kee, Richard Wolski, Daniel Nurmi, Dennis Gannon, Graziano Obertelli, Asim YarKhan, Anirban Mandal, T. Mark Huang, Kiran Thyagaraja, Dmitrii Zagorodnov |
SC | 3 |
| 2008 | Grid Resource Abstraction, Virtualization, and Provisioning for Time-Targeted ApplicationsabstractAs a variety of science applications are integrated with large-scale HPDC (high performance distributed computing) technologies, timely resource allocation is revealed as a critical requirement to be considered. This paper introduces a new HPDC resource management paradigm named resource slot which defines a network of logical machines across time and space. A resource slot is not only a resource programming target but also a virtualized resource provisioning framework for a variety of resource management paradigms by encapsulating the resource management complexity. Especially, we present a resource provisioning technique named guided redundant submission (GRS), which probabilistically guarantees a timely resource slot allocation. Experimental results performed against 8 clusters in production show that about 5 redundant resources per slot can secure slot allocation with up to 36 logical machines, each cluster having an availability probability as low as 0.25 and the target success probability of slot allocation is 0.95. Yang-Suk Kee, Carl Kesselman |
CCGRID | 1 |
| 2008 | Estimating Resource Needs for Time-Constrained WorkflowsabstractWorkflow technologies have become a major vehicle for the easy and efficient development of science applications. At the same time new computing environments such as the Cloud are now available. A challenge is to determine the right amount of resources to provision for an application. This paper introduces an algorithm named balanced time scheduling (BTS), which estimates the minimum number of virtual processors required to execute a workflow within a user-specified finish time. The resource estimate of BTS is abstract, so it can be easily integrated with any resource description language or any resource provisioning system. The experimental results with a number of synthetic workflows demonstrate that BTS can estimate the computing capacity close to the optimal. The algorithm is scalable so that its turnaround time is only tens of seconds even with workflows having thousands of tasks and edges. Eun-Kyu Byun, Yang-Suk Kee, Ewa Deelman, Karan Vahi, Gaurang Mehta, Jin-Soo Kim 0001 |
eScience | 2 |
| 2008 | Enabling personal clusters on demand for batch resources using commodity softwareabstractProviding QoS (quality of service) in batch resources against the uncertainty of resource availability due to the space-sharing nature of scheduling policies is a critical capability required for high-performance computing. This paper introduces a technique called personal cluster which reserves a partition of batch resources on user's demand in a best-effort manner. A personal cluster provides a private cluster dedicated to the user during a user-specified time period by installing a user-level resource manager on the resource partition. This technique not only enables cost-effective resource utilization and efficient task management but also provides the user a uniform interface to heterogeneous resources regardless of local resource management software. A prototype implementation using a PBS batch resource manager and Globus Toolkits based on Web services shows that the overhead of instantiating a personal cluster of medium size is small, which is just about 1 minute for a personal cluster having 32 processors. Yang-Suk Kee, Carl Kesselman, Daniel Nurmi, Richard Wolski |
IPDPS | 1 |
| 2008 | Overcoming performance bottlenecks in using OpenMP on SMP clusters
Woo-Chul Jeun, Yang-Suk Kee, Soonhoi Ha, Changdon Kee |
Parallel Comput. | 2 |
| 2006 | Scalable Grid Application Scheduling via Decoupled Resource Selection and SchedulingabstractOver the past years grid infrastructures have been deployed at larger and larger scales, with envisioned deployments incorporating tens of thousands of resources. Therefore, application scheduling algorithms can become unscalable (albeit polynomial) and thus unusable in large-scale environments. One reason for unscalability is that these algorithms perform implicit resource selection. One can achieve better scalability by performing explicit resource selection independently from scheduling in a "decoupled' approach. Furthermore, we hypothesize that one can achieve similar or even better performance as with the non-decoupled approach, which we call the "one step" approach, by selecting resources judiciously. Leveraging the Virtual Grid abstraction, we demonstrate that the decoupled approach is indeed both scalable and effective in large-scale and highly heterogeneous resource environments. Anirban Mandal, Henri Casanova, Andrew A. Chien, Yang-Suk Kee, Ken Kennedy, Charles Koelbel |
CCGRID | 5 |
| 2006 | Robust Resource Allocation for Large-scale Distributed Shared Resource EnvironmentsabstractThis paper presents a new formulation of the resource selection and binding problem and proposes a new algorithm called integrated selection and binding to solve this problem. Our insight is that a resource selection algorithm should consider binding failures. Consequently, the key idea of the integrated selection and binding approach is to decompose a resource collection request into components that can be bound and composed independently and to select multiple sets of resources for each component. The integrated approach is more efficient and effective than the separate approach for competitive access to federated resources Yang-Suk Kee, Ken Yocum, Andrew A. Chien, Henri Casanova |
HPDC | 1 |
| 2006 | Grid allocation and reservation - Improving grid resource allocation via integrated selection and bindingabstractDiscovering and acquiring appropriate, complex resource collections in large-scale distributed computing environments is a fundamental challenge and is critical to application performance. This paper presents a new formulation of the resource selection problem and a new solution to the resource selection and binding problem called integrated selection and binding. Composition operators in our resource description language and efficient data organization enable our approach to allocate complex resource collections efficiently and effectively even in the presence of competition for resources. Our empirical evaluation shows that the integrated approach can produce solutions of significantly higher quality at higher success rate and lower cost than the traditional separate approach. The success rate of the integrated approach can tolerate as much as 15%-60% lower resource availability than the separate approach. Moreover, most requests have at least the 98th percentile rank and can be served in 6 seconds with a population of 1 million hosts. Yang-Suk Kee, Ken Yocum, Andrew A. Chien, Henri Casanova |
SC | 1 |
| 2005 | Efficient resource description and high quality selection for virtual gridsabstractSimple resource specification, resource selection, and effective binding are critical capabilities for grid middleware. We describe the virtual grid, an abstraction for providing these capabilities complex resource environments. Elements of the virtual grid include a novel resource description language (vgDL) and a resource selection and binding component (vgFAB), which accepts a vgDL specification and returns a virtual grid, that is, a set of selected and bound resources. The goals of vgFAB are efficiency, scalability, robustness to high resource contention, and the ability to produce results with quantifiable high quality. We present the design of vgDL, showing how it captures application-level resource abstractions using resource aggregates and connectivity amongst them. We present and evaluate a prototype implementation of vgFAB. Our results show that resource selection and binding for virtual grids of 10,000's of resources can scale up to grids with millions of resources, identifying good matches in less than one second. Further, these matches have quantifiable quality, enabling applications to have high confidence in the results. We demonstrate the effectiveness of our combined selection and binding approach in the presence of resource contention, showing that robust selection and binding can be achieved at moderate cost. Yang-Suk Kee, Dionysios Logothetis, Richard Y. Huang, Henri Casanova, Andrew A. Chien |
CCGRID | 1 |
| 2004 | Realistic Modeling and Svnthesis of Resources for Computational GridsabstractUnderstanding large Grid platform configurations and generating representative synthetic configurations is critical for Grid computing research. This paper presents an analysis of existing resource configurations and proposes a Grid platform generator that synthesizes realistic configurations of both computing and communication resources. Our key contributions include the development of statistical models for currently deployed resources and using these estimates for modeling the characteristics of future systems. Through the analysis of the configurations of 114 clusters and over 10,000 processors, we identify appropriate distributions for resource configuration parameters in many typical clusters. Using well-established statistical tests, we validate our models against a second resource collection of 191 clusters and over 10,000 processors, and show that our models effectively capture the resource characteristics found in real world resource infrastructures. These models are realized in a resource generator, which can be easily recalibrated by running it on a training sample set. Yang-Suk Kee, Henri Casanova, Andrew A. Chien |
SC | 1 |
| 2004 | Memory management for multi-threaded software DSM systems
Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
Parallel Comput. | 1 |
| 2003 | ParADE: An OpenMP Programming Environment for SMP Cluster SystemsabstractDemand for programming environments to exploit clusters of symmetric multiprocessors (SMPs) is increasing. In this paper, we present a new programming environment, called ParADE, to enable easy, portable, and high-performance programming on SMP clusters. It is an OpenMP programming environment on top of a multi-threaded software distributed shared memory (SDSM) system with a variant of home-based lazy release consistency protocol. To boost performance, the runtime system provides explicit message-passing primitives to make it a hybrid-programming environment. Collective communication primitives are used for the synchronization and work-sharing directives associated with small data structures, lessening the synchronization overhead and avoiding the implicit barriers of work-sharing directives. The OpenMP translator bridges the gap between the OpenMP abstraction and the hybrid programming interfaces of the runtime system. The experiments with several NAS benchmarks and applications on a Linux-based cluster show promising results that ParADE overcomes the performance problem of the conventional SDSM-based OpenMP environment. Yang-Suk Kee, Jin-Soo Kim 0001, Soonhoi Ha |
SC | 1 |
| 2001 | xBSP: An Efficient BSP Implementation for clanabstractVirtual Interface Architecture (VIA) is a light-weight protocol for protected user-level zero-copy communication. In spite of the high performance of VIA, the previous MPI implementation for GigaNet's cLAN revealed low communication performance. The main sources of the low performance are the discrepancy of communication model between MPI and VIA and multi-threading overhead. We propose a novel implementation of the Bulk Synchronous Parallel (BSP) programming library for VIA called xBSP for overcoming such problems. To the best of our knowledge, xBSP is the first implementation of the BSP library for VIA. xBSP demonstrates that selecting a proper library is important to exploit the features of light-weight protocols. The intensive use of RDMA operation leads to high performance, close to the native VIA performance with respect to round trip delay and bandwidth. Based on the study of the effects of multithreading, memory registration, and completion policy on performance, we could obtain an efficient BSP implementation for cLAN, which is confirmed by experimental results. Yang-Suk Kee, Soonhoi Ha |
CCGRID | 1 |