EDBT 2026 Demo / reviewers in the wild / expert
Jae-Seung Yeom
dblp:89/1807 · also Jae Seung Yeom
· DBLP profile ↗
23ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-5464-6040ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | xAMM: "Attention" to Details Improves Cross-Platform Prediction AccuracyabstractAs computing becomes the major enabler in more and more fields, computing platforms also have become more heterogeneous than ever before to support different needs. Inevitably, high performance computing (HPC) centers and cloud vendors offer a diverse array of computing platforms to the user, often to a point where it overwhelms users as well as system managers. Therefore, a cross-platform performance prediction model, which leverages observations from one platform to predict performance on another, can be extremely valuable. However, building such a model for numerous platforms requires an enormous amount of effort to collect training data, which is often prohibitively expensive. To overcome this challenge, we propose$\times \text{AMM}^{1}$11Pronounced as “Exam”, an end-to-end Machine Learning (ML) pipeline that uses the attention mechanism, a transformative concept in generative AI, for two purposes: learning smart embeddings from raw application performance samples and constructing Abstract Machine Models (AMMs)-compact representations of machine properties. By integrating performance sample embeddings with AMMs where available, xAMM improves the accuracy of the state-of-the-art XGBoost model by 49.64 % for CPU$\rightarrow$CPU and 99.07 % for CPU$\rightarrow$GPU prediction compared to building the model using raw data, a common approach in the existing literature. Aakash Dhakal, Tanzima Z. Islam, Arunavo Dey, Daniel Nichols, Abhinav Bhatele, Tapasya Patki, Thomas Scogland, Jae-Seung Yeom |
CCGrid | 8 |
| 2025 | Predictive Execution of Workflows in a HPC+Cloud EnvironmentabstractMeeting deadlines for data-intensive workflows on HPC systems is challenging as jobs experience varying wait times before resources become available. This impact is significant in hybrid HPC+Cloud scheduling, which can lead to resource idleness, deadline violations, and higher costs. To address these issues, we propose scheduling data-intensive workflows over a combined HPC+Cloud hybrid environment in a deterministic manner by scavenging unused HPC resources. We predict resource availability (RA) of HPC systems, and exploit this prediction to dynamically split resource allocation between HPC's unused and Cloud's on-demand resources to complete a workflow by a given deadline. The deterministic resource allocation allows for preloading input data for workflow tasks, avoiding execution delays. Further, we develop an adaptive scaling algorithm that effectively backs up the targeted HPC allocation on Cloud facilities to avoid workflow execution delays in the event of incorrect RA estimation. Experiments show that our scheduling technique imposes minimal impact on HPC production jobs, saves cost for$>75 \%$workflow runs, suggests accurate budgets with a mean 7.11 % to 14.75 % cost estimation error, and finishes a mean 98 % to 99.4 % of tasks before deadlines. Subhendu Behera, Jae-Seung Yeom, Daniel Milroy, Marc Niethammer, Frank Mueller 0001 |
HiPC | 2 |
| 2025 | ModelX : A Novel Transfer Learning Approach Across Heterogeneous DatasetsabstractLeveraging an existing performance model to predict the runtime of a new application on a new system can save days and weeks of data collection time. However, knowledge transfer between High Performance Computing (HPC) systems can be challenging due to data heterogeneity caused by differences in data collection methods, architectural or application-specific individuality. This results in (1) sets of performance features that have significantly different names, orders, or the number of performance features that do not match between two datasets (heterogeneous domains), or (2) distribution shifts between datasets although their feature names match (homogeneous domains). While existing transfer learning techniques can handle mild distribution shifts, they fail to transfer knowledge when the source and target features do not match. This work introduces a novel transfer learning methodology-Cross Prediction Model (ModelX), which overcomes the large distribution discrepancy between homogeneous domains and enables transfer learning between heterogeneous domains. Extensive evaluations show that ModelX outperforms traditional transfer learning methods for all experiments using 11 HPC and 4 Machine Learning (ML) datasets. To the best of our knowledge, this is the first methodology to enable knowledge transfer between two heterogeneous domains with no matching features. Finally, we demonstrate an application of ModelX to an HPC job scheduling scenario using real-world job traces where it helps to reduce the job turnaround time of a set of jobs by 71%. Arunavo Dey, Neil Antony, Aakash Dhakal, Kowshik Thopalli, Jayaraman J. Thiagarajan, Tapasya Patki, Aniruddha Marathe, Thomas Scogland, Jae-Seung Yeom, Tanzima Z. Islam |
HPDC | 9 |
| 2025 | Pandemics in Silico: Scaling Agent-Based Simulations on Realistic Social Contact NetworksabstractPreventing the spread of infectious diseases requires implementing interventions at various levels of government and evaluating the potential impact and efficacy of those preemptive measures. Agent-based modeling can be used for detailed studies of the spread of such diseases in the presence of possible interventions. The computational cost of modeling epidemic diffusion through large social contact networks necessitates the use of parallel algorithms and resources in order to achieve quick turnaround times. In this work, we present Loimos, a scalable parallel framework for simulating epidemic diffusion. Loimos uses a hybrid of time-stepping and discrete event simulation to model disease spread, and is implemented on top of Charm++, an asynchronous, many-task runtime that enables over-decomposition and adaptive overlap of computation and communication. We demonstrate that Loimos is able to achieve significant speedups while scaling to large core counts. In particular, Loimos is able to simulate 200 days of a COVID19 outbreak on a digital twin of California in about 42 seconds, for an average of 4.6 billion traversed edges per second (TEPS), using 4096 cores on Perlmutter at NERSC. Joy Kitson, Ian J. Costello, Jiangzhuo Chen, Diego Jiménez, Stefan Hoops, Henning S. Mortveit, Esteban Meneses, Jae-Seung Yeom, Madhav V. Marathe, Abhinav Bhatele |
IPDPS | 8 |
| 2024 | Relative Performance Prediction Using Few-Shot LearningabstractHigh-performance computing system architectures are evolving rapidly, making exhaustive data collection for each architecture to build predictive performance models increasingly impractical. Concurrently, the arrival of new applications daily necessitates efficient performance prediction methods. Traditional data collection can take days or weeks, making it more efficient for scientists to leverage existing models to predict an application's performance on new architectures or use data from one application to predict another on the same architecture. The growing heterogeneity in applications and resources further complicates the exact matches needed for effective knowledge transfer. This work systematically studies various Machine Learning (ML) models to predict the relative performance of new applications on new platforms using existing data. Our findings demonstrate that few-shot learning using a few samples significantly enhances cross-platform knowledge transfer, multi-source models outperform single-source models, and Large Language Models (LLMs)-generated samples can effectively improve knowledge transfer efficacy. Arunavo Dey, Aakash Dhakal, Tanzima Z. Islam, Jae-Seung Yeom, Tapasya Patki, Daniel Nichols, Alexander Movsesyan, Abhinav Bhatele |
COMPSAC | 4 |
| 2024 | Predicting Cross-Architecture Performance of Parallel ProgramsabstractA variety of hardware architectures, both CPUs and GPUs, are used today to build supercomputers and parallel clusters. Often times, users can choose which hardware platform they want to run on. Modern scientific workflows have multiple computational tasks, and each task may be better suited for a different architecture in terms of performance. Deciding where to run an application or workflow task is not straightforward because of the complexity of applications, and hardware architectures, which makes performance predictions challenging. Hence, modeling the performance of scientific applications across a variety of architectures is important for achieving the best performance. In this paper, we present a machine learning based methodology to model the relative performance of applications across multiple architectures using hardware performance counters. Our machine learning model can predict the relative performance of an application with a mean absolute error of 0.11, and can be used effectively to make performance-aware and multi-architecture scheduling decisions, reducing makespan by up to 20%. Daniel Nichols, Alexander Movsesyan, Jae-Seung Yeom, Abhik Sarkar, Daniel Milroy, Tapasya Patki, Abhinav Bhatele |
IPDPS | 3 |
| 2024 | DYAD: Locality-aware Data Management for accelerating Deep Learning TrainingabstractDeep Learning (DL) is increasingly applied across various fields to solve complex scientific challenges in modern high-performance computing (HPC) systems that are beyond the reach of traditional algorithms. Training DL models for scientific applications involves processing multi-terabyte datasets in each epoch. The data access behavior during DL training exposes optimization opportunities to cache these datasets in near-compute storage accelerators in HPC systems, enhancing I/O throughput. However, current middleware solutions employ near-compute storage accelerators primarily as exclusive caches, which limits the effectiveness of cache access locality. To address this problem, we introduce DYAD, a system designed to maximize sample locality in the cache, thereby significantly increasing I/O throughput in HPC systems.DYAD optimizes I/O for DL training based on three key features. First, DYAD boosts inter-node access speeds by using a novel streaming RPC with RDMA protocol, achieving a 1.25x performance gain over state-of-the-art solutions. Second, DYAD further enhances inter-node access by coordinating data movement, which mitigates network congestion and increases throughput for inter-node accesses by up to 8.78x. Last, DYAD uses smart metadata caching that outperforms traditional global metadata access methods by several orders of magnitude in terms of lookup throughput. We demonstrate how DYAD accelerates large-scale DL training on a high-end HPC cluster with 512 GPUs by up to 10.82x faster epochs compared to UnifyFS by performing locality-aware caching on near-compute storage accelerators. Hariharan Devarajan, Ian Lumsden, Chen Wang 0004, Konstantia Georgouli, Thomas Scogland, Jae-Seung Yeom, Michela Taufer |
SBAC-PAD | 6 |
| 2024 | DFTracer: An Analysis-Friendly Data Flow Tracer for AI-Driven WorkflowsabstractModern HPC workflows involve intricate coupling of simulation, data analytics, and artificial intelligence (AI) applications to improve time to scientific insight. These workflows require a cohesive set of performance analysis tools to provide a comprehensive understanding of data exchange patterns in HPC systems. However, current tools are not designed to work with an AI-based I/O software stack that requires tracing at multiple levels of the application. To this end, we developed a data flow tracer called DFTracer to capture data-centric events from workflows and the I/O stack to build a detailed understanding of the data exchange within AI-driven workflows. DFTracer has the following three novel features, including a unified interface to capture trace data from different layers in the software stack, a trace format that is analysis-friendly and optimized to support efficiently loading multi-million events in a few seconds, and the capability to tag events with workflow-specific context to perform domain-centric data flow analysis for workflows. Additionally, we demonstrate that DFTracer has a $1.44 x$ smaller runtime overhead and 1.3-7.1x smaller trace size than state-of-the-art tracing tools such as Score-P, Recorder, and Darshan. Moreover, with AI-driven workflows, Score-P, Recorder, and Darshan cannot find I/O accesses from dynamically spawned processes, and their load performance of 100 M events is three orders of magnitude slower than DFTracer. In conclusion, we demonstrate that DFTracer can capture multi-level performance data, including contextual event tagging with a low overhead of 1-5% from AI-driven workflows such as MuMMI and Microsoft’s Megatron Deepspeed running on large-scale HPC systems. Hariharan Devarajan, Loïc Pottier, Kaushik Velusamy, Huihuo Zheng, Izzet Yildirim, Olga Kogiou, Weikuan Yu, Antonios Kougkas, Xian-He Sun, Jae-Seung Yeom, Kathryn Mohror |
SC | 10 |
| 2022 | Ubique: A New Model for Untangling Inter-task Data Dependence in Complex HPC WorkflowsabstractExploiting task parallelism is getting increasingly difficult for diverse and complex scientific workflows running on High Performance Computing (HPC) systems. In this paper, we argue that the difficulty rises from a void in the spectrum of existing data-transfer models for resolving inter-task data dependence within a workflow and propose a novel model to fill that gap: Ubique. The Ubique model combines the best from in-transit and in situ models in order for loosely coupled producer and consumer tasks to run concurrently and to resolve their data dependencies efficiently with little or no modifications to their codes, striking a balance between transparent optimization, productivity, and performance. Our preliminary evaluation suggests that Ubique can significantly outperform the parallel file system (PFS)-based model while offering automatic data transfer and synchronization which are the features lacking in many traditional models. It also identifies the performance characteristics of its key depending subsystems, which must be understood for further broadening its benefits. Jae-Seung Yeom, Dong H. Ahn, Ian Lumsden, Jakob Lüttgau, Silvina Caíno-Lores, Michela Taufer |
e-Science | 1 |
| 2022 | Enabling machine learning-ready HPC ensembles with Merlin
Jayson Luc Peterson, Benjamin Bay, Joe Koning, Peter B. Robinson, Jessica Semler, Jeremy White, Rushil Anirudh, Kevin Athey, Peer-Timo Bremer, Francesco Di Natale, Jim Gaffney, Sam Ade Jacobs, Bhavya Kailkhura, Bogdan Kustowski, Steve H. Langer, Brian K. Spears, Jayaraman J. Thiagarajan, Brian Van Essen, Jae-Seung Yeom |
Future Gener. Comput. Syst. | 20 |
| 2020 | Scalable Topological Data Analysis and Visualization for Evaluating Data-Driven Models in Scientific ApplicationsabstractWith the rapid adoption of machine learning techniques for large-scale applications in science and engineering comes the convergence of two grand challenges in visualization. First, the utilization of black box models (e.g., deep neural networks) calls for advanced techniques in exploring and interpreting model behaviors. Second, the rapid growth in computing has produced enormous datasets that require techniques that can handle millions or more samples. Although some solutions to these interpretability challenges have been proposed, they typically do not scale beyond thousands of samples, nor do they provide the high-level intuition scientists are looking for. Here, we present the first scalable solution to explore and analyze high-dimensional functions often encountered in the scientific data analysis pipeline. By combining a new streaming neighborhood graph construction, the corresponding topology computation, and a novel data aggregation scheme, namely topology aware datacubes, we enable interactive exploration of both the topological and the geometric aspect of high-dimensional data. Following two use cases from high-energy-density (HED) physics and computational biology, we demonstrate how these capabilities have led to crucial new insights in both applications. Shusen Liu 0001, Jim Gaffney, Jayson Luc Peterson, Peter B. Robinson, Harsh Bhatia, Valerio Pascucci, Brian K. Spears, Peer-Timo Bremer, Dan Maljovec, Rushil Anirudh, Jayaraman J. Thiagarajan, Sam Ade Jacobs, Brian Van Essen, David Hysom, Jae-Seung Yeom |
IEEE Trans. Vis. Comput. Graph. | 16 |
| 2019 | Parallelizing Training of Deep Generative Models on Massive Scientific DatasetsabstractTraining deep neural networks on large scientific data is a challenging task that requires enormous compute power, especially if no pre-trained models exist to initialize the process. We present a novel tournament method to train traditional as well as generative adversarial networks built on LBANN, a scalable deep learning framework optimized for HPC systems. LBANN combines multiple levels of parallelism and exploits some of the worlds largest supercomputers.We demonstrate our framework by creating a complex predictive model based on multi-variate data from high-energy-density physics containing hundreds of millions of images and hundreds of millions of scalar values derived from tens of millions of simulations of inertial confinement fusion. Our approach combines an HPC workflow and extends LBANN with optimized data ingestion and the new tournament-style training algorithm to produce a scalable neural network architecture using a CORAL-class supercomputer. Experimental results show that 64 trainers (1024 GPUs) achieve a speedup of 70.2× over a single trainer (16 GPUs) baseline, and an effective 109% parallel efficiency. Sam Ade Jacobs, Jim Gaffney, Tom Benson, Peter B. Robinson, Jayson Luc Peterson, Brian K. Spears, Brian Van Essen, David Hysom, Jae-Seung Yeom, Tim Moon, Rushil Anirudh, Jayaraman J. Thiagarajan, Shusen Liu 0001, Peer-Timo Bremer |
CLUSTER | 9 |
| 2018 | PADDLE: Performance Analysis Using a Data-Driven Learning EnvironmentabstractThe use of machine learning techniques to model execution time and power consumption, and, more generally, to characterize performance data is gaining traction in the HPC community. Although this signifies huge potential for automating complex inference tasks, a typical analytics pipeline requires selecting and extensively tuning multiple components ranging from feature learning to statistical inferencing to visualization. Further, the algorithmic solutions often do not generalize between problems, thereby making it cumbersome to design and validate machine learning techniques in practice. In order to address these challenges, we propose a unified machine learning framework, PADDLE, which is specifically designed for problems encountered during analysis of HPC data. The proposed framework uses an information-theoretic approach for hierarchical feature learning and can produce highly robust and interpretable models. We present user-centric workflows for using PADDLE and demonstrate its effectiveness in different scenarios: (a) identifying causes of network congestion; (b) determining the best performing linear solver for sparse matrices; and (c) comparing performance characteristics of parent and proxy application pairs. Jayaraman J. Thiagarajan, Rushil Anirudh, Bhavya Kailkhura, Tanzima Z. Islam, Abhinav Bhatele, Jae-Seung Yeom, Todd Gamblin |
IPDPS | 7 |
| 2017 | Massively Parallel Simulations of Spread of Infectious Diseases over Realistic Social NetworksabstractControlling the spread of infectious diseases in large populations is an important societal challenge. Mathematically, the problem is best captured as a certain class of reaction-diffusion processes (referred to as contagion processes) over appropriate synthesized interaction networks. Agent-based models have been successfully used in the recent past to study such contagion processes. We describe EpiSimdemics, a highly scalable, parallel code written in Charm++ that uses agent-based modeling to simulate disease spreads over large, realistic, co-evolving interaction networks. We present a new parallel implementation of EpiSimdemics that achieves unprecedented strong and weak scaling on different architectures - Blue Waters, Cori and Mira. EpiSimdemics achieves five times greater speedup than the second fastest parallel code in this field. This unprecedented scaling is an important step to support the long term vision of realtime epidemic science. Finally, we demonstrate the capabilities of EpiSimdemics by simulating the spread of influenza over a realistic synthetic social contact network spanning the continental United States (~280 million nodes and 5.8 billion social contacts). Abhinav Bhatele, Jae-Seung Yeom, Chris J. Kuhlman, Yarden Livnat, Keith R. Bisset, Laxmikant V. Kalé, Madhav V. Marathe |
CCGrid | 2 |
| 2017 | Performance modeling under resource constraints using deep transfer learningabstractTuning application parameters for optimal performance is a challenging combinatorial problem. Hence, techniques for modeling the functional relationships between various input features in the parameter space and application performance are important. We show that simple statistical inference techniques are inadequate to capture these relationships. Even with more complex ensembles of models, the minimum coverage of the parameter space required via experimental observations is still quite large. We propose a deep learning based approach that can combine information from exhaustive observations collected at a smaller scale with limited observations collected at a larger target scale. The proposed approach is able to accurately predict performance in the regimes of interest to performance analysts while outperforming many traditional techniques. In particular, our approach can identify the best performing configurations even when trained using as few as 1% of observations at the target scale. Aniruddha Marathe, Rushil Anirudh, Abhinav Bhatele, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Jae-Seung Yeom, Barry Rountree, Todd Gamblin |
SC | 7 |
| 2015 | Charm++ and MPI: Combining the Best of Both WorldsabstractCharm++ and MPI embody two distinct perspectives for writing parallel programs. While MPI provides a process-centric, user-driven model for developing parallel codes, Charm++ supports work-centric, system-driven parallel programming. One of them might be a better or more natural fit for individual modules that constitute a parallel application. In this paper, we present a framework that enables hybrid parallel programming with Charm++ and MPI, and allows programmers to develop different modules of a parallel application in these two languages while facilitating smooth interoperation. We describe the challenges in enabling interoperation between Charm++ and MPI, and present techniques for managing the control flow and resource sharing between the two. Finally, we demonstrate the benefits of interoperation between Charm++ and MPI through several case studies that use production applications and libraries, including CHARM/Chombo, EpiSimdemics, NAMD, FFTW, MPI-IO and ParMETIS. Abhinav Bhatele, Jae-Seung Yeom, Mark F. Adams, Francesco Miniati, Laxmikant V. Kalé |
IPDPS | 3 |
| 2014 | TRAM: Optimizing Fine-Grained Communication with Topological Routing and Aggregation of MessagesabstractFine-grained communication in supercomputing applications often limits performance through high communication overhead and poor utilization of network bandwidth. This paper presents Topological Routing and Aggregation Module (TRAM), a library that optimizes fine-grained communication performance by routing and dynamically combining short messages. TRAM collects units of fine-grained communication from the application and combines them into aggregated messages with a common intermediate destination. It routes these messages along a virtual mesh topology mapped onto the physical topology of the network. TRAM improves network bandwidth utilization and reduces communication overhead. It is particularly effective in optimizing patterns with global communication and large message counts, such as all-to-all and many-to-many, as well as sparse, irregular, dynamic or data dependent patterns. We demonstrate how TRAM improves performance through theoretical analysis and experimental verification using benchmarks and scientific applications. We present speedups on petascale systems of 6x for communication benchmarks and up to 4x for applications. Lukasz Wesolowski, Ramprasad Venkataraman, Abhishek Gupta 0002, Jae-Seung Yeom, Keith R. Bisset, Yanhua Sun, Pritish Jetley, Thomas Quinn 0001, Laxmikant V. Kalé |
ICPP | 4 |
| 2014 | Overcoming the Scalability Challenges of Epidemic Simulations on Blue WatersabstractModeling dynamical systems represents an important application class covering a wide range of disciplines including but not limited to biology, chemistry, finance, national security, and health care. Such applications typically involve large-scale, irregular graph processing, which makes them difficult to scale due to the evolutionary nature of their workload, irregular communication and load imbalance. EpiSimdemics is such an application simulating epidemic diffusion in extremely large and realistic social contact networks. It implements a graph-based system that captures dynamics among co-evolving entities. This paper presents an implementation of EpiSimdemics in Charm++ that enables future research by social, biological and computational scientists at unprecedented data and system scales. We present new methods for application-specific processing of graph data and demonstrate the effectiveness of these methods on a Cray XE6, specifically NCSA's Blue Waters system. Jae-Seung Yeom, Abhinav Bhatele, Keith R. Bisset, Eric J. Bohm, Abhishek Gupta 0002, Laxmikant V. Kalé, Madhav V. Marathe, Dimitrios S. Nikolopoulos, Martin Schulz 0001, Lukasz Wesolowski |
IPDPS | 1 |
| 2012 | High-Performance Interaction-Based Simulation of Gut Immunopathologies with ENteric Immunity Simulator (ENISI)abstractHere we present the ENteric Immunity Simulator (ENISI), a modeling system for the inflammatory and regulatory immune pathways triggered by microbe-immune cell interactions in the gut. With ENISI, immunologists and infectious disease experts can test and generate hypotheses for enteric disease pathology and propose interventions through experimental infection of an in silico gut. ENISI is an agent based simulator, in which individual cells move through the simulated tissues, and engage in context-dependent interactions with the other cells with which they are in contact. The scale of ENISI is unprecedented in this domain, with the ability to simulate $10^7$ cells for 250 simulated days on 576 cores in one and a half hours, with the potential to scale to even larger hardware and problem sizes. In this paper we describe the ENISI simulator for modeling mucosal immune responses to gastrointestinal pathogens. We then demonstrate the utility of ENISI by recreating an experimental infection of a mouse with Helicobacter pylori 26695. The results identify specific processes by which bacterial virulence factors do and do not contribute to pathogenesis associated with H. pylori strain 26695. These modeling results inform general intervention strategies by indicating immunomodulatory mechanisms such as those used in inflammatory bowel disease may be more appropriate therapeutically than directly targeting specific microbial populations through vaccination or by using antimicrobials. Keith R. Bisset, Md. Maksudul Alam, Josep Bassaganya-Riera, Adria Carbo, Stephen G. Eubank, Raquel Hontecillas, Stefan Hoops, Yongguo Mei, Katherine V. Wendelsdorf, Dawen Xie, Jae-Seung Yeom, Madhav V. Marathe |
IPDPS | 11 |
| 2010 | Strider: Runtime Support for Optimizing Strided Data Accesses on Multi-Cores with Explicitly Managed MemoriesabstractMulti-core processors with explicitly-managed local memories provide advanced capabilities to optimize data caching and prefetching in software. Unfortunately, these capabilities are neither easily accessible to programmers, nor exploited to their maximum potential by current language, compiler, or runtime frameworks. We present Strider, a runtime framework for optimizing compilers on multi-core processors with software- managed memories. Strider transparently optimizes grouping, decomposition, and scheduling of explicit software-managed accesses to multi-dimensional arrays in nested loops, given a high- level specification of loops and their data access patterns. In particular, Strider contributes new methods to improve temporal locality, optimize the critical path of scheduling data transfers for multi-stride accesses in regular nested parallel loops, and distribute accesses between cores. The prototype of Strider on the IBM Cell processor performs competitively to hand-optimized code and better than contemporary language frameworks, in both non-trivial parallel applications and important application kernels. Jae-Seung Yeom, Dimitrios S. Nikolopoulos |
SC | 1 |
| 2009 | A comparison of programming models for multiprocessors with explicitly managed memory hierarchiesabstractOn multiprocessors with explicitly managed memory hierarchies (EMM), software has the responsibility of moving data in and out of fast local memories. This task can be complex and error-prone even for expert programmers. Before we can allow compilers to handle this complexity for us, we must identify the abstractions that are general enough to allow us to write applications with reasonable effort, yet specific enough to exploit the vast on-chip memory bandwidth of EMM multi-processors. To this end, we compare two programming models against hand-tuned codes on the STI Cell, paying attention to programmability and performance. The first programming model, Sequoia, abstracts the memory hierarchy as private address spaces, each corresponding to a parallel task. The second, Cellgen, is a new framework which provides OpenMP-like semantics and the abstraction of a shared address space divided into private and shared data. We compare three applications programmed using these models against their hand-optimized counterparts in terms of abstractions, programming complexity, and performance. Scott Schneider 0001, Jae-Seung Yeom, Benjamin Rose, John C. Linford, Adrian Sandu, Dimitrios S. Nikolopoulos |
PPoPP | 2 |
| 2008 | Scheduling Asymmetric Parallelism on a PlayStation3 ClusterabstractUnderstanding the potential and implications of asymmetric multi-core processors for cluster computing is necessary, as these processors are rapidly becoming mainstream components in HPC environments. In this paper we evaluate a Linux cluster of Sony PlayStation3 consoles, using microbenchmarks and bioinformatics applications. We proceed to develop a model and scheduling techniques for effective execution of parallel applications on this low-cost, yet unconventional HPC platform based on the Cell/BE processor. We present an analytical formulation of layered parallelism for clusters of asymmetric multi-core multiprocessors and propose new co-scheduling heuristics for effectively executing MPI code with nested task and data parallelism on these systems. Our model has low execution time prediction error and is reliable in predicting optimal mappings of nested parallelism in MPI programs on the PS3 cluster. The presented co-scheduling heuristics reduce slack time on the accelerator cores of the PS3 and improve the performance of MPI applications by 1.7-2.7x, when compared against the native OS scheduler. Filip Blagojevic, Matthew Curtis-Maury, Jae-Seung Yeom, Scott Schneider 0001, Dimitrios S. Nikolopoulos |
CCGRID | 3 |
| 2007 | Security in All-Optical Networks: Self-Organization and Attack AvoidanceabstractWhile transparent WDM optical networks become more and more popular as the basis of the Next Generation Internet (NGI) infrastructure, such networks raise many unique security issues. The existing protection schemes which only consider unintended failures and only rely on postmortem detection and reaction are not sufficient to provide security assurance for such infrastructures which require timely protection from malicious sabotage as well as inadvertent faults. Moreover, as may have been observed from past practices, providing a particular solution specialized to each incessantly discovered new problems would not dramatically improve the situation. In order to increase the security of future networks we will need to imbed intelligence into the network such that they can continuously learn from the experience in faults and self-organize to protect themselves from potential failures caused by malicious new attacks or ordinary reliability problems. In this paper, we present results with simple vulnerability and attack scenarios in order to demonstrate how the self-organization helps to adapts against new vulnerabilities and avoid attacks. Jae-Seung Yeom, Ozan K. Tonguz, Gerardo A. Castañón |
ICC | 1 |