VLDB 2026 Research / reviewers in the wild / expert
Ian Karlin
dblp:07/4082
· DBLP profile ↗
19ranked-venue papers
2as first author
3since 2021 · last 2022
0009-0003-6602-0979ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
High-performance computing · 32% Interconnection networks and networks-on-chip · 31% Performance modeling and evaluation · 26% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 53% Concurrent programming · 35% Compilers and program optimization · 12% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip › network topology › tree networks
fat-tree network |
0.5 | 2 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree · SC 2016 |
High-performance computing
application porting |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Interconnection networks and networks-on-chip › cluster interconnect
infiniband |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Performance modeling and evaluation › benchmarking
interconnect benchmarking |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Parallel and multicore computing
programming models |
0.4 | 1 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 |
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect |
0.4 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
High-performance computing › supercomputing
supercomputer deployment |
0.3 | 1 | 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
Concurrent programming › concurrency bug detection
data race detection |
0.3 | 1 | 2017 | DataRaceBench: a benchmark suite for systematic evaluation of data race detection tools · SC 2017 |
Interconnection networks and networks-on-chip › network reconfiguration
network configuration |
0.3 | 1 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 |
Performance modeling and evaluation › simulation › communication system simulation
network simulation |
0.3 | 1 | 2017 | Predicting the performance impact of different fat-tree configurations · SC 2017 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2016 | Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-tree · SC 2016 |
Performance modeling and evaluation
benchmarking |
0.2 | 2 | 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous system · SC 2019 The design, deployment, and evaluation of the CORAL pre-exascale systems · SC 2018 |
High-performance computing › supercomputing
supercomputing systems |
0.1 | 1 | 2019 | An evaluation of the CORAL interconnects · SC 2019 |
Compilers and program optimization
domain-specific compilation |
0.1 | 1 | 2009 | Automating the generation of composed linear algebra kernels · SC 2009 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.1 | 1 | 2017 | DataRaceBench: a benchmark suite for systematic evaluation of data race detection tools · SC 2017 |
Memory systems › memory bandwidth
memory bandwidth optimization |
0.0 | 1 | 2009 | Automating the generation of composed linear algebra kernels · SC 2009 |
Methods — techniques the papers use, named apart from their topics
performance evaluation · 0.9benchmark suite design · 0.6communication benchmarking · 0.4TraceR-CODES simulation · 0.3trace analysis · 0.2empirical study · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Scalable Composition and Analysis Techniques for Massive Scientific WorkflowsabstractComposite science workflows are gaining traction to manage the combined effects of (1) extreme hardware heterogeneity in new High Performance Computing (HPC) systems and (2) growing software complexity – effects necessitated by the convergence of traditional HPC with data sciences. Composing, analyzing, and optimizing a composite workflow remains highly challenging as the component technologies are generally developed in isolation and often feature widely varying levels of performance, scalability, and interoperability. In this paper, we propose novel workflow composition and analysis techniques to create and optimize a scalable and effective composite workflow for heterogeneous HPC centers, and define the performance space of variables that impact composite workflow performance. We present PerfFlowAspect, an Aspect Oriented Programming (AOP)-based tool to perform cross-cutting performance analysis of composite workflows and better understand the impact of key performance variables on workflows. Our solution directly addresses AOP concerns that can affect workflow performance and covers the full software lifecycle, ranging from the workflow's initial composition through performance analysis and optimization. We use our science workflow composition techniques to implement the American Heart Association Molecule Screening (AHA MoleS) workflow. Through experimentation, we demonstrate that tuning a single performance variable can improve AHA MoleS workflow performance by a factor of up to 2.45x. Our evaluation suggests that our techniques can significantly enhance the ability of a multi-disciplinary research and development team to create a high performance composite workflow. Dong H. Ahn, Jeffrey Mast, Stephen Herbein, Francesco Di Natale, Daniel A. Kirshner, Sam Ade Jacobs, Ian Karlin, Daniel Milroy, Bronis R. de Supinski, Brian Van Essen, Jonathan E. Allen, Felice C. Lightstone |
e-Science | 8 |
| 2021 | On-the-Fly, Robust Translation of MPI LibrariesabstractMost parallel scientific applications rely on third-party libraries, some of which may have multiple implementations including open-source and vendor-proprietary. While sharing an application programming interface (API), many of these implementations do not have a shared application binary interface (ABI) and require recompiling applications to change the library implementation used. For many applications, recompiling is a long and complex process and sometimes not even an option when the application is shipped binary only. ABI incompatibility strikes at the heart of portability, productivity, and performance by (1) impeding application execution across different systems; (2) adding developer hours rebuilding an application; and (3) not taking advantage of host-optimized libraries.In this paper, we present a methodology and framework to solve ABI incompatibility across MPI libraries, which follow a well-defined API. The proposed framework called Wi4MPI translates the ABI dynamically from the MPI library used to build the application to a different MPI library available at run time. We show Wi4MPI works robustly on a wide spectrum of architectures, networks, and MPI libraries. Furthermore, we demonstrate its usefulness on several use cases highlighting significant portability, performance, and productivity benefits. Edgar A. León, Marc Joos, Nathan Hanford, Adrien Cotte, Tony Delforge, François Diakhaté, Vincent Ducrot, Ian Karlin, Marc Pérache |
CLUSTER | 8 |
| 2021 | Monitoring Large Scale Supercomputers: A Case Study with the Lassen SupercomputerabstractScalable management of user workloads on large-scale supercomputers remains a challenge due to the tradeoff between capturing adequate detail for analysis from various data sources and minimizing overhead. Co-designed frameworks, such as IBM’s Cluster System Management (CSM), provide a unified approach and novel insights for large-scale cluster management. This paper presents a longitudinal study and detailed analysis of a first-of-its-kind dataset collected by CSM from one of the world’s fastest supercomputers – comprised of over 1.4 million jobs on heterogenous nodes over multiple years. Furthermore, by focusing on a case study for power management, we identify the strengths and limitations of current CSM power measurement techniques in production. We present a deep dive into a large-scale scientific workflow, where finer-grained monitoring reveals power fluctuations at megawatt-levels resulting from the dynamic nature of the application, which are not captured by CSM. We make our unique datasets available to the HPC community, and discuss potential mitigation strategies by analyzing both coarse-grained and fine-grained data. Tapasya Patki, Adam Bertsch, Ian Karlin, Dong H. Ahn, Brian Van Essen, Barry Rountree, Bronis R. de Supinski, Nathan Besaw |
CLUSTER | 3 |
| 2020 | TOSS-2020: a commodity software stack for HPCabstractThe simulation environment of any HPC platform is key to the performance, portability, and productivity of scientific applications. This environment has traditionally been provided by platform vendors, presenting challenges for HPC centers and users including platform-specific software that tend to stagnate over the lifetime of the system. In this paper, we present the Tri-Laboratory Operating System Stack (TOSS), a production simulation environment based on Linux and open source software, with proprietary software components integrated as needed. TOSS, focused on mid-to-large scale commodity HPC systems, provides a common simulation environment across system architectures, reduces the learning curve on new systems, and benefits from a lineage of past experience and bug fixes. To further the scope and applicability of TOSS, we demonstrate its feasibility and effectiveness on a leadership-class supercomputer architecture. Our evaluation, relative to the vendor stack, includes an analysis of resource manager complexity, system noise, networking, and application performance. Edgar A. León, Trent D'Hooge, Nathan Hanford, Ian Karlin, Ramesh Pankajakshan, Jim Foraker, Christopher M. Chambreau, Matthew L. Leininger |
SC | 4 |
| 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous systemabstractProductivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience. Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
SC | 1 |
| 2019 | An evaluation of the CORAL interconnectsabstractThe US Department of Energy deployed the Summit and Sierra supercomputers with the latest state-of-the-art network interconnect technology in 2018 and both systems entered production in 2019. In this paper, we provide an in-depth assessment of the systems' network interconnects that are based on Enhanced Data Rate (EDR) 100 Gb/s Mellanox InfiniBand. Both systems use second-generation EDR Host Channel Adapters (HCAs) and switches with several new features such as Adaptive Routing (AR), switch-based collectives, and HCA-based tag matching. Although based on the same components, Summit's network is "non-blocking" (i.e., a fully provisioned Clos network) and Sierra's network has a 2:1 taper between the racks and aggregation switches. We evaluate the two systems' interconnects using traditional communication benchmarks as well as production applications. We find that the new Adaptive Routing dramatically improves performance but the other new features still need improvement. Christopher Zimmer 0001, Scott Atchley, Ramesh Pankajakshan, Brian E. Smith, Ian Karlin, Matthew L. Leininger, Adam Bertsch, Brian S. Ryujin, Jason Burmark, André Walker-Loud, Michael A. Clark, Olga Pearce |
SC | 5 |
| 2018 | Runtime and Memory Evaluation of Data Race Detection Tools
Pei-Hung Lin, Chunhua Liao, Markus Schordan, Ian Karlin |
ISoLA (2) | 4 |
| 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin |
SC | 27 |
| 2017 | Predicting the performance impact of different fat-tree configurationsabstractThe fat-tree topology is one of the most commonly used network topologies in HPC systems. Vendors support several options that can be configured when deploying fat-tree networks on production systems, such as link bandwidth, number of rails, number of planes, and tapering. This paper showcases the use of simulations to compare the impact of these design options on representative production HPC applications, libraries, and multi-job workloads. We present advances in the TraceR-CODES simulation framework that enable this analysis and evaluate its prediction accuracy against experiments on a production fat-tree network. In order to understand the impact of different network configurations on various anticipated scenarios, we study workloads with different communication patterns, computation-to-communication ratios, and scaling characteristics. Using multi-job workloads, we also study the impact of inter-job interference on performance and compare the cost-performance tradeoffs. Abhinav Bhatele, Louis H. Howell, David Böhme, Ian Karlin, Edgar A. León, Misbah Mubarak, Noah Wolfe, Todd Gamblin, Matthew L. Leininger |
SC | 5 |
| 2017 | DataRaceBench: a benchmark suite for systematic evaluation of data race detection toolsabstractData races in multi-threaded parallel applications are notoriously damaging while extremely difficult to detect. Many tools have been developed to help programmers find data races. However, there is no dedicated OpenMP benchmark suite to systematically evaluate data race detection tools for their strengths and limitations. Chunhua Liao, Pei-Hung Lin, Joshua Asplund, Markus Schordan, Ian Karlin |
SC | 5 |
| 2016 | Fast Multi-parameter Performance ModelingabstractTuning large applications requires a clever exploration of the design and configuration space. Especially on supercomputers, this space is so large that its exhaustive traversal via performance experiments becomes too expensive, if not impossible. Manually creating analytical performance models provides insights into optimization opportunities but is extremely laborious if done for applications of realistic size. If we must consider multiple performance-relevant parameters and their possible interactions, a common requirement, this task becomes even more complex. We build on previous work on automatic scalability modeling and significantly extend it to allow insightful modeling of any combination of application execution parameters. Multi-parameter modeling has so far been outside the reach of automatic methods due to the exponential growth of the model search space. We develop a new technique to traverse the search space rapidly and generate insightful performance models that enable a wide range of uses from performance predictions for balanced machine design to performance tuning. Alexandru Calotoiu, D. A. Beckingsale, Christopher W. Earl, Torsten Hoefler, Ian Karlin, Martin Schulz 0001, Felix Wolf 0001 |
CLUSTER | 5 |
| 2016 | System Noise Revisited: Enabling Application Scalability and Reproducibility with SMTabstractDespite significant advances in reducing system noise, the scalability and performance of scientific applications running on production commodity clusters today continue to suffer from the effects of noise. Unlike custom and expensive leadership systems, the Linux ecosystem provides a rich set of services that application developers utilize to increase productivity and to ease porting. The cost is the overhead that these services impose on a running application, negatively impacting its scalability and performance reproducibility. In this work, we propose and evaluate a simple yet effective way to isolate an application from system processes by leveraging Simultaneous Multi-Threading (SMT), a pervasive architectural feature on current systems. Our method requires no changes to the operating system or to the application. We quantify its effectiveness on a diverse set of scientific applications of interest to the U. S. Department of Energy showing performance improvements of up to 2.4 times at 16,384 tasks for a high-order finite elements shock hydrodynamics application. Finally, we provide guidance to system and application developers on how to best leverage SMT under different application characteristics and scales. Edgar A. León, Ian Karlin, Adam Moody |
IPDPS | 2 |
| 2016 | Characterizing parallel scientific applications on commodity clusters: an empirical study of a tapered fat-treeabstractUnderstanding the characteristics and requirements of applications that run on commodity clusters is key to properly configuring current machines and, more importantly, procuring future systems effectively. There are only a few studies, however, that are current and characterize realistic workloads. For HPC practitioners and researchers, this limits our ability to design solutions that will have an impact on real systems. We present a systematic study that characterizes applications with an emphasis on communication requirements. It includes cluster utilization data, identifying a representative set of applications from a U.S. Department of Energy laboratory, and characterizing their communication requirements. The driver for this work is understanding application sensitivity to a tapered fat-tree network. These results provided key insights into the procurement of our next generation commodity systems. We believe this investigation can provide valuable input to the HPC community in terms of workload characterization and requirements from a large supercomputing center. Edgar A. León, Ian Karlin, Abhinav Bhatele, Steve H. Langer, Christopher M. Chambreau, Louis H. Howell, Trent D'Hooge, Matthew L. Leininger |
SC | 2 |
| 2016 | Program optimizations: The interplay between power, performance, and energy
Edgar A. León, Ian Karlin, Ryan E. Grant, Matthew G. F. Dosanjh |
Parallel Comput. | 2 |
| 2015 | Optimizing Explicit Hydrodynamics for Power, Energy, and PerformanceabstractPractical considerations for future supercomputer designs will impose limits on both instantaneous power consumption and total energy consumption. Working within these constraints while providing the maximum possible performance, application developers will need to optimize their code for speed alongside power and energy concerns. This paper analyzes the effectiveness of several code optimizations including loop fusion, data structure transformations, and global allocations. A per component measurement and analysis of different architectures is performed, enabling the examination of code optimizations on different compute subsystems. Using an explicit hydrodynamics proxy application from the U.S. Department of Energy, LULESH, we show how code optimizations impact different computational phases of the simulation. This provides insight for simulation developers into the best optimizations to use during particular simulation compute phases when optimizing code for future supercomputing platforms. We examine and contrast both x86 and Blue Gene architectures with respect to these optimizations. Edgar A. León, Ian Karlin, Ryan E. Grant |
CLUSTER | 2 |
| 2015 | Data Layout Optimization for Portable Performance
Kamal Sharma, Ian Karlin, Jeff Keasler, James R. McGraw, Vivek Sarkar |
Euro-Par | 2 |
| 2013 | Exploring Traditional and Emerging Parallel Programming Models Using a Proxy ApplicationabstractParallel machines are becoming more complex with increasing core counts and more heterogeneous architectures. However, the commonly used parallel programming models, C/C++ with MPI and/or OpenMP, make it difficult to write source code that is easily tuned for many targets. Newer language approaches attempt to ease this burden by providing optimization features such as automatic load balancing, overlap of computation and communication, message-driven execution, and implicit data layout optimizations. In this paper, we compare several implementations of LULESH, a proxy application for shock hydrodynamics, to determine strengths and weaknesses of different programming models for parallel computation. We focus on four traditional (OpenMP, MPI, MPI+OpenMP, CUDA) and four emerging (Chapel, Charm++, Liszt, Loci) programming models. In evaluating these models, we focus on programmer productivity, performance and ease of applying optimizations. Ian Karlin, Abhinav Bhatele, Jeff Keasler, Bradford L. Chamberlain, Jonathan D. Cohen 0001, Zach DeVito, Riyaz Haque, Daniel E. Laney, Edward Luke, Felix Wang, David F. Richards, Martin Schulz 0001, Charles H. Still |
IPDPS | 1 |
| 2009 | Automating the generation of composed linear algebra kernelsabstractMemory bandwidth limits the performance of important kernels in many scientific applications. Such applications often use sequences of Basic Linear Algebra Subprograms (BLAS), and highly efficient implementations of those routines enable scientists to achieve high performance at little cost. However, tuning the BLAS in isolation misses opportunities for memory optimization that result from composing multiple subprograms. Because it is not practical to create a library of all BLAS combinations, we have developed a domain-specific compiler that generates them on demand. In this paper, we describe a novel algorithm for compiling linear algebra kernels and searching for the best combination of optimization choices. We also present a new hybrid analytic/empirical method for quickly evaluating the profitability of each optimization. We report experimental results showing speedups of up to 130% relative to the GotoBLAS on an AMD Opteron and up to 137% relative to MKL on an Intel Core 2. Geoffrey Belter, Elizabeth R. Jessup, Ian Karlin, Jeremy G. Siek |
SC | 3 |
| 2008 | Build to order linear algebra kernelsabstractThe performance bottleneck for many scientific applications is the cost of memory access inside linear algebra kernels. Tuning such kernels for memory efficiency is a complex task that reduces the productivity of computational scientists. Software libraries such as the Basic Linear Algebra Subprograms (BLAS) ameliorate this problem by providing a standard interface for which computer scientists and hardware vendors have created highly-tuned implementations. Scientific applications often require a sequence of BLAS operations, which presents further opportunities for memory optimization. However, because BLAS are tuned in isolation they do not take advantage of these opportunities. This phenomenon motivated the recent addition to the BLAS of several routines that perform sequences of operations. Unfortunately, the exact sequence of operations needed in a given situation is highly application dependent, so many more routines are needed. In this paper we present preliminary work on a domain- specific compiler that generates implementations for arbitrary sequences of basic linear algebra operations and tunes them for memory efficiency. We report experimental results for dense kernels and show speedups of 25 % to 120 % relative to sequences of calls to GotoBLAS and vendor-tuned BLAS on Intel Xeon and IBM PowerPC platforms. Jeremy G. Siek, Ian Karlin, Elizabeth R. Jessup |
IPDPS | 2 |