EDBT 2026 Demo / reviewers in the wild / expert
Sunita Chandrasekaran
dblp:28/2826
· DBLP profile ↗
24ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-3560-9428ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural AnalysisabstractAs GPU architectures rapidly evolve to meet the growing demands of exascale computing and machine learning, the performance implications of architectural innovations remain poorly understood across diverse workloads. NVIDIA Blackwell (B200) introduces significant architectural advances, including fifth-generation tensor cores, tensor memory (TMEM), a decompression engine (DE), and a dual-chip design; however, systematic methodologies for quantifying these improvements lag behind hardware development cycles. We contribute an open-source microbenchmark suite that provides practical insights into optimizing workloads to fully utilize the rich feature sets of modern GPU architectures. This work enables application developers to make informed architectural decisions and guides future GPU design directions. We study Blackwell GPUs and compare them to the H200 generation with respect to the memory subsystem, tensor core pipeline, and floating-point precisions (FP32, FP16, FP8, FP6, FP4). Our systematic evaluation of dense and sparse GEMM, transformer inference, and training workloads shows that B200 tensor core enhancements achieve 1.85x ResNet-50 and 1.55x GPT-1.3B mixed-precision training throughput, with 32 percent better energy efficiency than H200. Aaron Jarmusch, Sunita Chandrasekaran |
IPDPS | 2 |
| 2025 | LLM4VV:: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler TestsabstractThe usage of Large Language Models (LLMs) for software and test development has continued to increase since LLMs were first introduced, but only recently have the expectations of LLMs become more realistic. Verifying the correctness of code generated by LLMs is key to improving their usefulness, but there have been no comprehensive and fully autonomous solutions developed yet. Hallucinations are a major concern when LLMs are applied blindly to problems without taking the time and effort to verify their outputs, and an inability to explain the logical reasoning of LLMs leads to issues with trusting their results. To address these challenges while also aiming to effectively apply LLMs, this paper proposes a dual-LLM system (i.e. a generative LLM and a discriminative LLM) and experiments with the usage of LLMs for the generation of a large volume of compiler tests. We experimented with a number of LLMs possessing varying parameter counts and presented results using ten carefully-chosen metrics that we describe in detail in our narrative. Through our findings, it is evident that LLMs possess the promising potential to generate quality compiler tests and verify them automatically. Zachariah Sollenberger, Saieda Ali Zada, Sunita Chandrasekaran |
HiPC | 4 |
| 2025 | The Artificial Scientist: in-Transit Machine Learning of Plasma SimulationsabstractLarge-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list. Jeffrey Kelling, Vicente Bolea, Michael Bussmann, Ankush Checkervarty, Alexander Debus, Jan Ebert, Greg Eisenhauer, Vineeth Gutta, Stefan Kesselheim, Scott Klasky, Vedhas Pandit, Richard Pausch, Norbert Podhorszki, Franz Poeschel, David Rogers, Jeyhun Rustamov, Steve Schmerler, Ulrich Schramm, Klaus Steiniger, René Widera, Anna Willmann, Sunita Chandrasekaran |
IPDPS | 22 |
| 2024 | LLM4VV: Developing LLM-driven testsuite for compiler validation
Christian Munley, Aaron Jarmusch, Sunita Chandrasekaran |
Future Gener. Comput. Syst. | 3 |
| 2024 | UNNT: A novel Utility for comparing Neural Net and Tree-based modelsabstractThe use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry. Vineeth Gutta, Satish Ranganathan Ganakammal, Matthew Beyers, Sunita Chandrasekaran |
PLoS Comput. Biol. | 5 |
| 2023 | Implementing OpenMP's SIMD Directive in LLVM's GPU RuntimeabstractGPUs support three levels of parallelism: thread blocks, warps (or wavefronts) within a block, and threads within a warp. Some GPU programming models allow the use of all three of these levels, such as OpenMP offloading with the teams, parallel, and simd directives. However LLVM/OpenMP does not support simd and only uses two levels, thread blocks and all threads within a block. For codes with three explicit layers of parallelism this can decrease performance and potentially require restructuring of the application. In this work we present our design and implementation of the OpenMP simd directive in LLVM’s OpenMP GPU runtime, which includes both CPU-centric and GPU-centric execution models. We evaluate our prototype using kernels and a few proxy applications showing a performance improvement ranging from 1.3x to 3.5x depending on the benefit the kernels receives from such an optimization. Thus, this work enables real-world applications with three explicit layers of parallelism to expose to better exploit the full benefits of GPU architecture. Eric Wright, Johannes Doerfert, Shilei Tian, Barbara M. Chapman, Sunita Chandrasekaran |
ICPP | 5 |
| 2023 | Frontier: Exploring ExascaleabstractAs the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project. Scott Atchley, Christopher Zimmer 0001, Jack Lange, David E. Bernholdt, Verónica G. Vergara Larrea, Michael J. Brim, Reuben D. Budiardja, Sunita Chandrasekaran, Markus Eisenbach 0002, Thomas M. Evans 0001, Matthew Ezell, Nicholas Frontiere, Antigoni Georgiadou, Joseph Glenski, Philipp Grete, Steven P. Hamilton, John K. Holmen, Axel Huebl, Daniel A. Jacobson, Wayne Joubert, Kim H. McMahon, Elia Merzari, Stan G. Moore, Andrew Myers 0001, Stephen Nichols, Sarp Oral, Thomas Papatheodore, Danny Perez, David M. Rogers 0001, Evan Schneider, Jean-Luc Vay, P. K. Yeung |
SC | 9 |
| 2023 | Special issue on new trends in high-performance computing: Software systems and applicationsabstractHigh-performance computing (HPC) offers the computing power to continuously support the world's most important discoveries in various scientific and business domains such as chemistry, physics, biology, material science, drug discovery, and financial investment risk analysis. We are now in the exascale era with the Frontier exascale system that has been very recently revealed (June 2022). Researchers from across the HPC community have been developing software systems, tools, libraries, frameworks, application packages, and methods that can fully exploit these extremely powerful computing resources. Such extreme-scale computing will enable the solution of vastly more accurate predictive models and the analysis of massive quantities of data, producing quantum advances in areas of science and technology that are essential to the scientific community. New computational approaches such as machine learning/deep learning have also been heavily explored in recent years and have shown promising evidence for many problems that cannot be resolved by traditional computational simulation and engineering. Training deep neural networks with massive data is an extremely computing-intensive task that heavily relies on HPC power. The upcoming exascale computing era will be the essential basis for supporting new innovations in machine learning/deep learning-based exploration and will lead to new sciences in directions such as smart manufacturing, laboratory automation, and automatic programming. While the hardware architecture can generate extreme computing power, renovation in the software stack plays an essential role in effective performance delivery. The large supercomputers continue to move into the heterogeneous space, while the former fastest ARM-based system, Fugaku, and the many core Sunway Taihulight types of systems are marching towards the heterogeneous space. With systems equipped with GPUs, Advance RISC Machine (ARM) Support Vector Engine (SVEs), and many cores, there is a dire need for innovative software frameworks that can seamlessly migrate scientific code to these systems equipped with rich computing resources. We need innovation at different levels, including compiler tools and techniques, performance analysis tools, novel abstractions of the programming model, redesign of application-level algorithms, and so on. Furthermore, co-design of applications and low-level software frameworks can lead to more efficient use of the opportunities of exascale in many contexts. This special issue has selected 14 papers. We next present the summary of the papers presented in this special issue. The first paper titled “ParTransgrid: A scalable parallel pre-processing tool for unstructured-grid cell-centered Computational Fluid Dynamics (CFD) applications” by Jianqiang et al.1 proposes a parallel pre-processing tool, called ParTransgrid, that translates the general grid format such as the CFD General Notation System into an efficient distributed mesh data format for large-scale parallel computing. Experiment results reveal that ParTransgrid can be easily scaled to billion-level grid CFD applications and that the preparation time for parallel computing with hundreds of thousands of cores is reduced to a few minutes. The second paper titled “Generation of logic designs for efficiently solving ordinary differential equations on FPGAs” by Korch et al.2 proposes a framework that is able to automatically generate specific and optimized solver logic from easy-to-handle configuration files. No manual development and no special Field Programmable Gate Array (FPGA) or programming knowledge are required. The logic generated by this improved approach is up to 43 times faster than its hand-optimized High Level Synthesis (HLS) counterpart, depending on the solution method. The third paper titled “NAS Parallel Benchmarks with Compute Unified Device Architecture (CUDA) and Beyond” by Fernandes et al.3 provides a new CUDA implementation for NASA Parallel Benchmark (NPB). The performance results have shown up to 267% improvements over the best benchmark versions available. The authors also observe the best and worst design choices concerning code size and the performance tradeoff. Lastly, the authors highlight the challenges of implementing parallel CFD applications for Graphic Procesing Unit (GPUs) and how the computations impact the GPU's behavior. The fourth paper titled, “Using Ginkgo's Memory Accessor for Improving the Accuracy of Memory-Bound Low Precision Basic Linear Algebra Subprograms (BLAS)” by Quintana-Ortí et al.4 demonstrates that memory-bound applications operating on low precision data can increase their accuracy by relying on the memory accessor to perform all arithmetic operations in high precision. In particular, the authors demonstrate that memory-bound BLAS operations (including the sparse matrix-vector product) can be re-engineered with the memory accessor and that the resulting accessor-enabled BLAS routines achieve lower rounding errors while delivering the same performance as the fast low-precision BLAS. The fifth paper titled “Three Practical Workflow Schedulers for Easy Maximum Parallelism” by Rogers5 presents a complete characterization of the minimum effective task granularity for efficient scheduler usage scenarios. A separate job scheduler is implemented for three distinct workflow patterns involved in the preparation, execution, and analysis of computational chemistry simulations. It shows unique benefits, including simplicity of design, suitability for HPC centers, short startup time, and well-understood per-task overhead. All three new tools have been shown to scale to full utilization of Summit and have been made publicly available with tests and documentation. The sixth paper titled “LLAMA: The Low-Level Abstraction for Memory Access” by Bussmann et al.6 presents the Low-Level Abstraction of Memory Access (LLAMA), a C++ library that provides such a data structure abstraction layer with example implementations for multidimensional arrays of nested, structured data. LLAMA provides fully C++-compliant methods for defining and switching custom memory layouts for user-defined data types. The library is extensible with third-party allocators. LLAMA provides a novel tool for the development of high-performance C++ applications in a heterogeneous environment. The seventh paper titled “PAS: A new powerful and simple quantum computing simulator” by Wang et al.7 proposes a new powerful and simple CPU-based quantum computing simulator: PAS (Power And Simple). Compared with existing simulators, PAS introduces four novel optimization methods: efficient hybrid vectorization, fast bitwise operation, memory access filtering, and quantum tracking. Experiments were performed on the Intel Xeon E5-2670 v3 CPU and showed that PAS compared with the state-of-the-art simulator QuEST can achieve a mean speedup of 8.69x and 2.62x for the Quantum Field Theory (QFT) and Relativistic Quantum Chemistry (RQC) benchmarks, respectively. The eighth paper titled “Dynamics Signature based Anomaly Detection” by Bader et al.8 borrows the dynamics metrics and proposes the concept of Dynamics Signature (DS) in multi-dimensional feature space to efficiently distinguish the abnormal event from the normal behaviors of a variable star. Two datasets, parameterized sinusoidal dataset containing 262,440 light curves, and a real variable star-based dataset containing 462,996 light curves are used to evaluate the practical performance of the proposed DS algorithm. Experimental results show that their DS algorithm is highly accurate, sensitive to detecting weak microlensing events at very early stages, and fast enough to process 176,000 stars in less than 1 second on a commodity computer. The ninth paper titled “EESSI: A Cross-Platform Ready-To-Use Optimised Scientific Software Stack” by Röblitz et al.9 proposes the European Environment for Scientific Software Installations project that aims to provide a ready-to-use stack of scientific software installations that can be leveraged easily on a variety of platforms, ranging from personal workstations to cloud environments and supercomputer infrastructure, without making compromises with respect to performance. The authors provide a detailed overview of the project, highlight potential use cases, and demonstrate that the performance of the provided scientific software installations can be competitive with system-specific installations. The eleventh paper titled “A large scale parallel fluid-structure interaction computing platform for simulating structural responses to a detonation shock” by Yang et al.10 presents a partitioned fluid-structure interaction computing platform designed for parallel simulating structural responses to a detonation shock. The 3D numerical result of structural responses to a detonation shock is presented and analyzed. On 256 processor cores, the speedup ratio of the simulations for a detonation shock reaches 178.0 with 5.1 million mesh cells and the parallel efficiency achieves 69.5%. The results demonstrate the good potential of massively parallel simulations. Overall, a general-purpose fluid-structure interaction software platform with detonation support is proposed by integrating open source codes. The authors express their sincere gratitude and thanks to the Editor-in-Chief Dr. Rajkumar Buyya, for guiding them to organize this special issue. The authors appreciate the support from the editorial office. The authors are also thankful to all the authors who submitted their ideas to this special issue and to the reviewers for their thoughtful and critical suggestions to improve the quality of the submitted papers. Sunita Chandrasekaran, Min Si, Jidong Zhai, Lena Oden |
Softw. Pract. Exp. | 1 |
| 2022 | First Experiences in Performance Benchmarking with the New SPEChpc 2021 SuitesabstractModern High Performance Computing (HPC) sys-tems are built with innovative system architectures and novel programming models to further push the speed limit of computing. The increased complexity poses challenges for performance portability and performance evaluation. The Standard Perfor-mance Evaluation Corporation (SPEC) has a long history of producing industry-standard benchmarks for modern computer systems. SPEC's newly released SPEChpc 2021 benchmark suites, developed by the High Performance Group, are a bold attempt to provide a fair and objective benchmarking tool designed for state-of-the-art HPC systems. With the support of multiple host and accelerator programming models, the suites are portable across both homogeneous and heterogeneous architectures. Different workloads are developed to fit system sizes ranging from a few compute nodes to a few hundred compute nodes. In this work we present our first experiences in performance benchmarking the new SPEChpc2021 suites and evaluate their portability and basic performance characteristics on various popular and emerging HPC architectures, including x86 CPU, NVIDIA GPU, and AMD GPU. This study provides a first-hand experience of executing the SPEChpc 2021 suites at scale on production HPC systems, discusses real-world use cases, and serves as an initial guideline for using the benchmark suites. Holger Brunst, Sunita Chandrasekaran, Florina M. Ciorba, Nick Hagerty, Robert Henschel, Guido Juckeland, Junjie Li 0003, Verónica G. Vergara Larrea, Sandra Wienke, Miguel Zavala |
CCGRID | 2 |
| 2020 | Proposing a Machine Learning Framework for Classification of Patient Cohorts Using Genomics Data
Mauricio H. Ferrato, Erin L. Crowgey, Sunita Chandrasekaran |
AMIA | 3 |
| 2020 | Accelerating prediction of chemical shift of protein structures on GPUs: Using OpenACCabstractExperimental chemical shifts (CS) from solution and solid state magic-angle-spinning nuclear magnetic resonance (NMR) spectra provide atomic level information for each amino acid within a protein or protein complex. However, structure determination of large complexes and assemblies based on NMR data alone remains challenging due to the complexity of the calculations. Here, we present a hardware accelerated strategy for the estimation of NMR chemical-shifts of large macromolecular complexes based on the previously published PPM_One software. The original code was not viable for computing large complexes, with our largest dataset taking approximately 14 hours to complete. Our results show that serial code refactoring and parallel acceleration brought down the time taken of the software running on an NVIDIA Volta 100 (V100) Graphic Processing Unit (GPU) to 46.71 seconds for our largest dataset of 11.3 million atoms. We use OpenACC, a directive-based programming model for porting the application to a heterogeneous system consisting of x86 processors and NVIDIA GPUs. Finally, we demonstrate the feasibility of our approach in systems of increasing complexity ranging from 100K to 11.3M atoms. Eric Wright, Mauricio H. Ferrato, Alexander J. Bryer, Robert Searles, Juan R. Perilla, Sunita Chandrasekaran |
PLoS Comput. Biol. | 6 |
| 2019 | Correction to: Computational approaches for cancer 2017 workshop overviewabstractᅟ. Sunita Chandrasekaran, Eric Stahlberg |
BMC Bioinform. | 1 |
| 2019 | Analysis of OpenMP 4.5 Offloading in Implementations: Correctness and Overhead
José Monsalve Diaz, Kyle Friedline, Swaroop Pophale, Oscar R. Hernandez, David E. Bernholdt, Sunita Chandrasekaran |
Parallel Comput. | 6 |
| 2019 | pointerchain: Tracing pointers to their roots - A case study in molecular dynamics simulations
Millad Ghane, Sunita Chandrasekaran, Margaret S. Cheung |
Parallel Comput. | 2 |
| 2018 | Computational approaches for Cancer 2017 workshop overviewabstractScalable Deep Text Comprehension for Cancer Surveillance on High-Performance Computing.Reflecting on the noted areas of opportunity in working with patients, with organizations, with data and with technology, this supplement provides examples of how Sunita Chandrasekaran, Eric Stahlberg |
BMC Bioinform. | 1 |
| 2018 | Special issue on applications for the heterogeneous computing era 2017
Sunita Chandrasekaran, Antonio J. Peña |
Parallel Comput. | 1 |
| 2018 | The OpenACC data model: Preliminary study on its major challenges and implementations
Michael Wolfe, Seyong Lee, Xiaonan Tian, Rengan Xu, Barbara M. Chapman, Sunita Chandrasekaran |
Parallel Comput. | 7 |
| 2017 | Special Issue on Topics on Heterogeneous Computing
Sunita Chandrasekaran, Antonio J. Peña |
Parallel Comput. | 1 |
| 2016 | cusFFT: A High-Performance Sparse Fast Fourier Transform Algorithm on GPUsabstractThe Fast Fourier Transform (FFT) is one of the most important numerical tools widely used in many scientific and engineering applications. The algorithm performs O(nlogn) operations on n input data points in order to calculate only small number of k large coefficients, while the rest of n - k numbers are zero or negligibly small. The algorithm is clearly inefficient, when n points input data lead to only k <;Z n non-zero coefficients in the transformed domain. MIT in 2012 developed a sparse FFT (sFFT) algorithm that provides a solution to this problem. In this paper, we explore the challenges and propose effective solutions to efficiently port sFFT to massively parallel processors, such as GPUs, using CUDA. GPGPUs are being increasingly adopted as popular HPC platforms because of their tremendous computing power and remarkable cost efficiency. However, sFFT algorithm is a complex and computationally challenging memory-bound algorithm that is not straightforward to be implemented on GPUs. In this paper, we present some of the optimization strategies such as index coalescing, loop splitting, asynchronous data layout transformation, linear time selection algorithm that are required to compute sFFT on such massively parallel architectures. Our CUDA-based sFFT, cusFFT, performs over 10x faster than the state-of-the-art cuFFT library on GPUs and over 28x faster than the parallel FFTW on multicore CPUs. Cheng Wang 0001, Sunita Chandrasekaran, Barbara M. Chapman |
IPDPS | 2 |
| 2016 | Compiler transformation of nested loops for general purpose GPUsabstractSummary Manycore accelerators have the potential to significantly improve performance of scientific applications when offloading computationally intensive program portions to accelerators. Directive‐based high‐level programming models, such as OpenACC and OpenMP, are used to create applications for accelerators through annotating regions of code meant for offloading. OpenACC is an emerging directive‐based programming model for programming accelerators that typically enable inexperienced programmers to achieve portable and productive performance within applications. In this paper, we present our research in developing challenges and solutions when creating an open‐source OpenACC compiler in an industrial framework (OpenUH as a branch of Open64). We then discuss in detail techniques we developed for loop scheduling reduction operations on general purpose GPUs. The compiler is evaluated with benchmarks from the NAS Parallel Benchmarks suite and self‐written micro‐benchmarks for reduction operations. This implementation has been designed to serve as a compiler infrastructure for researchers to explore advanced compiler techniques, extend OpenACC to other programming models, and build performance tools used in conjunction with OpenACC programs. Copyright © 2015 John Wiley & Sons, Ltd. Xiaonan Tian, Rengan Xu, Yonghong Yan 0001, Sunita Chandrasekaran, Deepak Eachempati, Barbara M. Chapman |
Concurr. Comput. Pract. Exp. | 4 |
| 2013 | Portable mapping of openMP to multicore embedded systems using MCA APIsabstractMulticore embedded systems are being widely used in telecommunication systems, robotics, medical applications and more.While they offer a high-performance with low-power solution, programming in an efficient way is still a challenge. In order to exploit the capabilities that the hardware offers, software developers are expected to handle many of the low-level details of programming including utilizing DMA, ensuring cache coherency, and inserting synchronization primitives explicitly. The state-of-the-art involves solutions where the software toolchain is too vendor-specific thus tying the software to a particular hardware leaving no room-for portability. Cheng Wang 0001, Sunita Chandrasekaran, Barbara M. Chapman, Jim Holt |
LCTES | 2 |
| 2013 | C2FPGA - A dependency-timing graph design methodology
Sunita Chandrasekaran, Shilpa Shanbagh, Ramkumar Jayaraman, Douglas L. Maskell, Hui Yan Cheah |
J. Parallel Distributed Comput. | 1 |
| 2010 | A dependency graph based methodology for parallelizing HLL applications on FPGA (abstract only)abstractReconfigurable computing is an exciting new area of research. It opens up a number of new computing paradigms with the potential to significantly change the embedded computing landscape. Unfortunately, as with any new concept, there is much to be done before it can be brought into the mainstream. Reconfigurable computing needs to be made more usable by the general computing community by making the application mapping process more user transparent. We need to be able to efficiently exploit the parallelism and high communications bandwidth available in reconfigurable computing systems, such as provided by FPGA devices. In this paper we propose a framework to map C-based applications to an FPGA based system. The framework relies on dependency based analysis information obtained from a high level Intermediate Representation (IR) of the C-Application. The dependencies have a major impact on the instruction flow and the parallelism of the algorithms when implemented on FPGA based high performance computing platforms. We demonstrate the importance of these dependencies in algorithms by examining a compute intensive bio-informatics application. Targeting on a Xilinx Virtex 5 XC5VFX70T FPGA, our place and route results achieved 92% of the clock frequency achieved by using a manual design of the same application. The contribution in this paper is to provide a semi-automated approach to generate hardware for FPGA based systems through a high level data dependency graphical IR. Sunita Chandrasekaran, Shilpa Shanbagh, Douglas L. Maskell |
FPGA | 1 |
| 2008 | Capturing performance knowledge for automated analysisabstractAutomating the process of parallel performance experimentation, analysis, and problem diagnosis can enhance environments for performance-directed application development, compilation, and execution. This is especially true when parametric studies, modeling, and optimization strategies require large amounts of data to be collected and processed for knowledge synthesis and reuse. This paper describes the integration of the PerfExplorer performance data mining framework with the OpenUH compiler infrastructure. OpenUH provides auto-instrumentation of source code for performance experimentation and PerfExplorer provides automated and reusable analysis of the performance data through a scripting interface. More importantly, PerfExplorer inference rules have been developed to recognize and diagnose performance characteristics important for optimization strategies and modeling. Three case studies are presented which show our success with automation in OpenMP and MPI code tuning, parametric characterization, Pand power modeling. The paper discusses how the integration supports performance knowledge engineering across applications and feedback-based compiler optimization in general. Kevin A. Huck, Oscar R. Hernandez, Van Bui, Sunita Chandrasekaran, Barbara M. Chapman, Allen D. Malony, Lois C. McInnes, Boyana Norris |
SC | 4 |