EDBT 2026 Demo / reviewers in the wild / expert
Yoonho Park
dblp:57/1574
· DBLP profile ↗
19ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0002-9837-0741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Energy-efficient computing · 26% High-performance computing · 22% Distributed systems · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Energy-efficient computing
power management |
0.8 | 2 | 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device Systems · DAC 2020 Power Aware Heterogeneous Node Assembly · HPCA 2019 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.4 | 1 | 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device Systems · DAC 2020 |
High-performance computing
application porting |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Parallel and multicore computing
programming models |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Storage systems
adaptive prefetching |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Cloud and datacenter computing
big data analytics |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems › distributed data processing
data shuffling |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Storage systems
i/o optimization |
0.3 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Distributed systems
fault tolerance |
0.3 | 2 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 Towards Optimal Resource Allocation in Partial-Fault Tolerant Applications · INFOCOM 2008 |
Hardware reliability and fault tolerance › memory reliability
DRAM errors |
0.2 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Memory systems › virtual memory management
page migration |
0.2 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Embedded and real-time systems
mobile computing |
0.1 | 1 | 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device Systems · DAC 2020 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Energy-efficient computing
datacenter power management |
0.1 | 1 | 2019 | Power Aware Heterogeneous Node Assembly · HPCA 2019 |
Cloud and datacenter computing › big data analytics
in-memory data analytics |
0.1 | 1 | 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data Analytics · IEEE Trans. Parallel Distributed Syst. 2017 |
Electronic design automation › physical design › placement
component placement |
0.1 | 1 | 2008 | Towards Optimal Resource Allocation in Partial-Fault Tolerant Applications · INFOCOM 2008 |
Cloud and datacenter computing
resource allocation |
0.1 | 1 | 2008 | Towards Optimal Resource Allocation in Partial-Fault Tolerant Applications · INFOCOM 2008 |
Data stream processing
distributed stream processing |
0.1 | 1 | 2006 | Design, implementation, and evaluation of the linear road bnchmark on the stream processing core · SIGMOD Conference 2006 |
Data stream processing
stream processing systems |
0.1 | 1 | 2006 | Design, implementation, and evaluation of the linear road bnchmark on the stream processing core · SIGMOD Conference 2006 |
Operating systems › resource management
memory management |
0.1 | 1 | 2014 | A System Software Approach to Proactive Memory-Error Avoidance · SC 2014 |
Operating systems › resource management › memory management
virtual memory |
0.0 | 1 | 1996 | Virtual Memory versus File Interface for Large, Memory-Intensive Scientific Applications · SC 1996 |
High-performance computing › scientific computing
scientific computing application |
0.0 | 1 | 1996 | Virtual Memory versus File Interface for Large, Memory-Intensive Scientific Applications · SC 1996 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.4q-learning · 0.4sorted assembly · 0.4balanced power assembly · 0.4application-aware assembly · 0.4log analysis · 0.4error prediction · 0.4prefetching · 0.3adaptive i/o · 0.3approximation algorithm · 0.1custom virtual memory policies · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use CaseabstractScientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services. Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer |
CLOUD | 9 |
| 2021 | Reinforcement Learning-Based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that utilizes reinforcement learning to increase the power efficiency of mobile device systems based on a multiprocessor system-on-a-chip (MPSoC). The proposed policy predicts a system’s characteristics and learns power management controls to adapt to the variations in the system. We consider the behavioral characteristics of systems that run on mobile devices under diverse scenarios. Therefore, the policy can flexibly manage the system power regardless of the application scenario and achieve lower energy consumption without compromising the user satisfaction. The average energy per unit quality of service (QoS) of the proposed policy is lower than that of the previous six dynamic voltage/frequency scaling governors by 31.66%. Furthermore, we reduce the runtime overhead by implementing the proposed policy as hardware. We implemented the policy on the field programmable gate array (FPGA) and construct a communication interface between the central processing units (CPUs) and the hardware of the proposed policy. Decision-making by the hardware-implemented policy is 3.92 times faster than by the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Jongho Yoon 0001, Seokhyeong Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that exploits reinforcement learning to increase power efficiency of mobile device systems. Our Q-learning-based policy predicts a system’s characteristics and learns power management controls to adapt to the system’s variations. Therefore, we can flexibly manage the system power regardless of the application scenario and can achieve lower energy per QoS compared to previous dynamic voltage/frequency scaling governors. To minimize the process overhead, we implemented our power management policy as hardware; the hardware-implemented policy reduced the average latency up to 40× compared to the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Young Hwan Kim, Seokhyeong Kang |
DAC | 3 |
| 2020 | Analysis and Solution of CNN Accuracy Reduction over Channel Loop TilingabstractOwing to the growth of the size of convolutional neural networks (CNNs), quantization and loop tiling (also called loop breaking) are mandatory to implement CNN on an embedded system. However, channel loop tiling of quantized CNNs induces unexpected errors. We explain why channel loop tiling of quantized CNNs induces the unexpected errors, and how the errors affect the accuracy of state-of-the-art CNNs. We also propose a method to recover accuracy under channel tiling by compressing and decompressing the most-significant bits of partial sums. Using the proposed method, we can recover accuracy by 12.3% with only 1% circuit area overhead and an additional 2% of power consumption. Yesung Kang, Yoonho Park, Eunji Kwon, Taeho Lim, Sangyun Oh, Mingyu Woo, Seokhyeong Kang |
DATE | 2 |
| 2020 | GRLC: grid-based run-length compression for energy-efficient CNN acceleratorabstractConvolutional neural networks (CNNs) require a huge amount of off-chip DRAM access, which accounts for most of its energy consumption. Compression of feature maps can reduce the energy consumption of DRAM access. However, previous compression methods show poor compression ratio if the feature maps are either extremely sparse or dense. To improve the compression ratio efficiently, we have exploited the spatial correlation and the distribution of non-zero activations in output feature maps. In this work, we propose a grid-based run-length compression (GRLC) and have implemented a hardware for the GRLC. Compared with a previous compression method [1], GRLC reduces 11% of the DRAM access and 5% of the energy consumption on average in VGG-16, ExtractionNet and ResNet-18. Yoonho Park, Yesung Kang, Eunji Kwon, Seokhyeong Kang |
ISLPED | 1 |
| 2019 | Power Aware Heterogeneous Node AssemblyabstractTo meet ever increasing computational requirements, supercomputers and data centers are beginning to utilize fat compute nodes with multiple hardware components such as manycore CPUs and accelerators. These components have intrinsic power variations even among same model components from same manufacturer. In this paper, we argue that node assembly techniques that consider these intrinsic power variations can achieve better power efficiency without any performance trade off on large scale supercomputing facilities and data centers. We propose three different node assembly techniques: (1) Sorted Assembly, (2) Balanced Power Assembly, and (3) Application-Aware Assembly. In Sorted Assembly, node components are categorized (or sorted) into groups according to their power efficiency, and components from the same group are assembled into a node. In Balanced Power Assembly, components are assembled to minimize node-to-node power variations. In Application-Aware Assembly, the most heavily used components by the application are selected based on the highest power efficiency. We evaluate the effectiveness and cost savings of the three techniques compared to the standard random assembly under different node counts and variability scenarios. Bilge Acun, Alper Buyuktosunoglu, Yoonho Park |
HPCA | 4 |
| 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous systemabstractProductivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience. Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
SC | 2 |
| 2018 | Predicting the Performance Impact of Increasing Memory Bandwidth for Scientific WorkflowsabstractThe disparity between the bandwidth provided by modern processors and by the main memory led to the issue known as memory wall, in which application performance becomes completely bound by memory speed. Newer technologies are trying to increase memory bandwidth to address this issue, but the fact is that the effects of increasing bandwidth to application performance still lack exploration. This paper investigates these effects for scientific workflows focusing on the definition of a performance model and on the execution of experiments to validate the rationale for the model. The main contribution is based on two observations: memory bound applications benefit more from an increase to memory bandwidth, and the effects of improving bandwidth for a particular application gradually diminish as bandwidth is increased. Nelson Mimura Gonzalez, José R. Brunheroto, Fausto Artico, Yoonho Park, Tereza Cristina M. B. Carvalho, Charles Miers, Maurício Aronne Pillon, Guilherme P. Koslovski |
SBAC-PAD | 4 |
| 2017 | Support for Power Efficient Proactive Cooling MechanismsabstractIncreasing scale of data centers and the density of server nodes pose significant challenges in producing power and energy efficient cooling infrastructures. Current fan based air cooling systems have significant inefficiencies in their operation causing oscillations in fan power consumption and temperature variations among cores. In this paper, we identify the cause these problems and propose proactive cooling mechanisms to mitigate the power peaks and temperature variations. An accurate temperature prediction model lies behind the basis of our solutions. We use a neural network-based modeling approach for predicting core temperatures of different workloads, under different core frequencies, fan speed levels, and ambient temperature. The model provides guidance for our proactive cooling mechanisms. We propose a preemptive and decoupled fan control mechanism that can remove the power peaks in fan power consumption and reduce the maximum cooling power by 53.3% on average as well as energy consumption by 22.4%. Moreover, through our decoupled fan control method and thermal-aware load balancing algorithm, we show that temperature variations in large scale platforms can be reduced from 25 C to 2 C, making cooling systems more efficient with negligible performance overhead. Bilge Acun, Yoonho Park, Laxmikant V. Kalé |
HiPC | 3 |
| 2017 | Leveraging Adaptive I/O to Optimize Collective Data Shuffling Patterns for Big Data AnalyticsabstractBig data analytics is an indispensable tool in transforming science, engineering, medicine, health-care, finance and ultimately business itself. With the explosion of data sizes and need for shorter time-to-solution, in-memory platforms such as Apache Spark gain increasing popularity. In this context, data shuffling, a particularly difficult transformation pattern, introduces important challenges. Specifically, data shuffling is a key component of complex computations that has a major impact on the overall performance and scalability. Thus, speeding up data shuffling is a critical goal. To this end, state-of-the-art solutions often rely on overlapping the data transfers with the shuffling phase. However, they employ simple mechanisms to decide how much data and where to fetch it from, which leads to sub-optimal performance and excessive auxiliary memory utilization for the purpose of prefetching. The latter aspect is a growing concern, given evidence that memory per computation unit is continuously decreasing while interconnect bandwidth is increasing. This paper contributes a novel shuffle data transfer strategy that addresses the two aforementioned dimensions by dynamically adapting the prefetching to the computation. We implemented this novel strategy in Spark, a popular inmemory data analytics framework. To demonstrate the benefits of our proposal, we run extensive experiments on an HPC cluster with large core count per node. Compared with the default Spark shuffle strategy, our proposal shows: up to 40 percent better performance with 50 percent less memory utilization for buffering and excellent weak scalability. Bogdan Nicolae, Carlos H. A. Costa, Claudia Misale, Kostas Katrinis, Yoonho Park |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | Towards Memory-Optimized Data Shuffling Patterns for Big Data AnalyticsabstractBig data analytics is an indispensable tool in transforming science, engineering, medicine, healthcare, finance and ultimately business itself. With the explosion of data sizes and need for shorter time-to-solution, in-memory platforms such as Apache Spark gain increasing popularity. However, this introduces important challenges, among which data shuffling is particularly difficult: on one hand it is a key part of the computation that has a major impact on the overall performance and scalability so its efficiency is paramount, while on the other hand it needs to operate with scarce memory in order to leave as much memory available for data caching. In this context, efficient scheduling of data transfers such that it addresses both dimensions of the problem simultaneously is non-trivial. State-of-the-art solutions often rely on simple approaches that yield sub optimal performance and resource usage. This paper contributes a novel shuffle data transfer strategy that dynamically adapts to the computation with minimal memory utilization, which we briefly underline as a series of design principles. Bogdan Nicolae, Carlos H. A. Costa, Claudia Misale, Kostas Katrinis, Yoonho Park |
CCGrid | 5 |
| 2016 | Speeding Up Stencil Computations with Kernel ConvolutionabstractA technique to speed up stencil computation is introduced. Computation and data reuse schemes are developed for its application to 1- and 3-dimensional stencils. The approach traverses the data domain fewer times than a state-of-the-art, straightforward iterative stencil implementation would. Performance results are shown for a variety of platforms, exemplifying how it can be straightforwardly applied with existing techniques and frameworks. The technique, named Aggregate Stencil-Loop Iteration (ASLI), works by applying a stencil obtained by the original stencil operator convolved with itself one or more times. This more complex operator creates new opportunities for in-register data reuse and increases the FLOPs-to-load ratio. The total number of FLOPs decreases for 1D but increases for 2D and 3D star-shaped stencils. In both scenarios, speed-up relative to the state-of-the-art is achieved. ASLI is relatively easy to implement and works synergistically with existing methods to optimize stencil computations. Guilherme Carvalho Januario, Bryan S. Rosenburg, Yoonho Park, Michael Perrone, José E. Moreira, Tereza Cristina M. B. Carvalho |
SBAC-PAD | 3 |
| 2014 | A System Software Approach to Proactive Memory-Error AvoidanceabstractToday's HPC systems use two mechanisms to address main-memory errors. Error-correcting codes make correctable errors transparent to software, while checkpoint/restart (CR) enables recovery from uncorrectable errors. Unfortunately, CR overhead will be enormous at exascale due to the high failure rate of memory. We propose a new OS-based approach that proactively avoids memory errors using prediction. This scheme exposes correctable error information to the OS, which migrates pages and off lines unhealthy memory to avoid application crashes. We analyze memory error patterns in extensive logs from a BG/P system and show how correctable error patterns can be used to identify memory likely to fail. We implement a proactive memory management system on BG/Q by extending the firmware and Linux. We evaluate our approach with a realistic workload and compare our overhead against CR. We show improved resilience with negligible performance overhead for applications. Carlos H. A. Costa, Yoonho Park, Bryan S. Rosenburg, Chen-Yong Cher, Kyung Dong Ryu |
SC | 2 |
| 2012 | FusedOS: Fusing LWK Performance with FWK Functionality in a Heterogeneous EnvironmentabstractTraditionally, there have been two approaches to providing an operating environment for high performance computing (HPC). A Full-Weight Kernel(FWK) approach starts with a general-purpose operating system and strips it down to better scale up across more cores and out across larger clusters. A Light-Weight Kernel (LWK) approach starts with a new thin kernel code base and extends its functionality by adding more system services needed by applications. In both cases, the goal is to provide end-users with a scalable HPC operating environment with the functionality and services needed to reliably run their applications. To achieve this goal, we propose a new approach, called Fused OS, that combines the FWK and LWK approaches. Fused OS provides an infrastructure capable of partitioning the resources of a multicoreheterogeneous system and collaboratively running different operating environments on subsets of the cores and memory, without the use of a virtual machine monitor. With Fused OS, HPC applications can enjoy both the performance characteristics of an LWK and the rich functionality of an FWK through cross-core system service delegation. This paper presents the Fused OS architecture and a prototype implementation on Blue Gene/Q. The Fused OS prototype leverages Linux with small modifications as a FWK and implements a user-level LWK called Compute Library (CL) by leveraging CNK. We present CL performance results demonstrating low noise and show micro-benchmarks running with performance commensurate with that provided by CNK. Yoonho Park, Eric Van Hensbergen, Marius Hillenbrand, Todd Inglett, Bryan S. Rosenburg, Kyung Dong Ryu, Robert W. Wisniewski |
SBAC-PAD | 1 |
| 2012 | Evaluation of a high-volume, low-latency market data processing system implemented with IBM middlewareabstractSUMMARY A stock market data processing system that can handle high data volumes at low latencies is critical to market makers. Such systems play a critical role in algorithmic trading, risk analysis, market surveillance, and many other related areas. The current systems tend to use specialized software and custom processors. We show that such a system can be built with general‐purpose middleware and run on commodity hardware. The middleware we use is IBM System S which includes transport technology from IBM WebSphere MQ Low Latency Messaging (LLM). Our performance evaluation consists of two parts. First, we determined the effectiveness of each system optimization that the hardware and software infrastructure makes available. These optimizations were implemented at all software levels—application, middleware, and operating system. Second, we evaluated our system on different hardware platforms. Copyright © 2011 John Wiley & Sons, Ltd. Yoonho Park, Richard King, Senthil Nathan, Wesley Most, Henrique Andrade |
Softw. Pract. Exp. | 1 |
| 2008 | Towards Optimal Resource Allocation in Partial-Fault Tolerant ApplicationsabstractWe introduce Zen, a new resource allocation framework that assigns application components to node clusters to achieve high availability for partial-fault tolerant (PFT) applications. These applications have the characteristic that under partial failures, they can still produce useful output though the output quality may be reduced. Thus, the primary goal of resource allocation for PFT applications is to prevent, delay, or minimize the impact of failures on the application output quality. This paper is the first to approach this resource allocation problem from a theoretical perspective, and obtains a series of results regarding component assignments that provide the highest service availability under the constraints imposed by the application data flow graph and the hosting clusters. We show that (1) even simple versions of this resource allocation problem are NP-Hard, (2) a 2-approximate polynomial-time algorithm works for tree topologies, and (3) a simple greedy component placement performs well in practice for general application topologies. We implement a system prototype to study the application availability achieved by Zen compared to failure-oblivious placement, replication, and Zen+replication. Our experimental results show that three PFT applications achieve significant data output quality and availability benefits using Zen. Nikhil Bansal 0001, Ranjita Bhagwan, Navendu Jain, Yoonho Park, Deepak S. Turaga, Chitra Venkatramani |
INFOCOM | 4 |
| 2006 | Design, implementation, and evaluation of the linear road bnchmark on the stream processing coreabstractStream processing applications have recently gained significant attention in the networking and database community. At the core of these applications is a stream processing engine that performs resource allocation and management to support continuous tracking of queries over collections of physically-distributed and rapidly-updating data streams. While numerous stream processing systems exist, there has been little work on understanding the performance characteristics of these applications in a distributed setup. In this paper, we examine the performance bottlenecks of streaming data applications, in particular the Linear Road stream data management benchmark, in achieving good performance in large-scale distributed environments, using the Stream Processing Core (SPC), a stream processing middleware we have developed. First, we present the design and implementation of the Linear Road benchmark on the SPC middleware. SPC has been designed to scale to tens of thousands of processing nodes, while supporting concurrent applications and multiple simultaneous queries. Second, we identify the main performance bottlenecks in the Linear Road application in achieving scalability and low query response latency. Our results show that data locality, buffer capacity, physical allocation of processing elements to infrastructure nodes, and packaging for transporting streamed data are important factors in achieving good application performance. Though we evaluate our system primarily for the Linear Road application, we believe it also provides useful insights into the overall system behavior for supporting other distributed and large-scale continuous streaming data applications. Finally, we examine how SPC can be used and tuned to enable a very efficient implementation of the Linear Road application in a distributed environment. Navendu Jain, Lisa Amini, Henrique Andrade, Richard King, Yoonho Park, Philippe Selo, Chitra Venkatramani |
SIGMOD Conference | 5 |
| 1997 | Custom virtual memory policies for an image reconstruction applicationabstractSome scientific applications have very large memory requirements, and are required to move data between primary and secondary storage during execution. The I/O between processor and disk can be done using standard file interfaces or virtual memory. The use of virtual memory, though simple and straight-forward, has been criticized because of poor performance. However some modern operating systems provide techniques to "customize" virtual memory to obtain better performance. The authors present results from an image reconstruction application where 3D structures of viruses are reconstructed from 2D electron microscope images. The platform is the NEC Cenju-3 parallel computer running the Cenju-3/DE, a Mach-based operating system. The virtual memory interface allows the user to specify custom page replacement and prefetching policies. For the largest data set, using custom virtual memory reduced the the data access time by a factor of 25 compared to unformatted Fortran I/O. Olin G. Johnson, Vasudha Govindan, Yoonho Park, Z. Hong Zhou |
HiPC | 3 |
| 1996 | Virtual Memory versus File Interface for Large, Memory-Intensive Scientific ApplicationsabstractScientific applications often require some strategy for temporary data storage to do the largest possible simulations. The use of virtual memory for temporary data storage has received criticism because of performance problems. However, modern virtual memory found in recent operating systems such as Cenju-3/DE give application writers control over virtual memory policies. We demonstrate that custom virtual memory policies can dramatically reduce virtual memory overhead and allow applications to run out-of-core efficiently. We also demonstrate that the main advantage of virtual memory, namely programming simplicity, is not lost. Yoonho Park, Ridway Scott, Stuart Sechrest |
SC | 1 |