EDBT 2026 Demo / reviewers in the wild / expert
Apala Guha
dblp:42/1625
· DBLP profile ↗
10ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Electronic design automation · 41% Processor architecture and microarchitecture · 24% GPUs and heterogeneous computing · 12% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.7 | 2 | 2019 | μIR -An intermediate representation for transforming and optimizing the microarchitecture of application accelerators · MICRO 2019 TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018 |
Electronic design automation › design representation
intermediate representation |
0.4 | 1 | 2019 | μIR -An intermediate representation for transforming and optimizing the microarchitecture of application accelerators · MICRO 2019 |
Processor architecture and microarchitecture
microarchitecture optimization |
0.4 | 1 | 2019 | μIR -An intermediate representation for transforming and optimizing the microarchitecture of application accelerators · MICRO 2019 |
GPUs and heterogeneous computing › GPU programming
dynamic parallelism |
0.3 | 1 | 2018 | TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.2 | 1 | 2016 | Chainsaw: Von-neumann accelerators to leverage fused instruction chains · MICRO 2016 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2012 | Memory optimization of dynamic binary translators for embedded systems · ACM Trans. Archit. Code Optim. 2012 |
Embedded and real-time systems › resource-constrained computing › resource-constrained embedded system
memory-constrained embedded systems |
0.1 | 1 | 2012 | Memory optimization of dynamic binary translators for embedded systems · ACM Trans. Archit. Code Optim. 2012 |
Parallel and multicore computing › parallel programming models › task parallelism
fork-join parallelism |
0.1 | 1 | 2018 | TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2018 | TAPAS: Generating Parallel Accelerators from Parallel Programs · MICRO 2018 |
Runtime systems and virtual machines › virtual machine implementation
code cache management |
0.0 | 1 | 2012 | Memory optimization of dynamic binary translators for embedded systems · ACM Trans. Archit. Code Optim. 2012 |
Methods — techniques the papers use, named apart from their topics
intermediate representation transformation · 0.4high-level synthesis · 0.3compiler intermediate representation · 0.3simulation · 0.3pipeline forwarding · 0.2LLVM-based compilation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Deepframe: A Profile-Driven Compiler for Spatial Hardware AcceleratorsabstractTracing code paths to form extended basic blocks is useful in many areas, compiler optimizations [1], improving instruction cache behavior [2] and custom-hardware offloading [3]. Prior work has been plagued by small traces, limited either by the overheads of dynamic profiling, statically available information [4], or side-exit branches [5]. In this work, we rethink what code path sequences to fuse and construct long traces for offloading to spatial accelerators, while minimizing the occurrence of side exits which limit dynamic coverage. We introduce a novel technique that recasts learning a program's execution patterns as a natural-language-processing problem, CBOW (Continuous Bag of Words). We then use a deep learning network to learn the relationships among paths. During the compilation phase, the compiler uses a sequence miner to decide what paths are likely to occur. The learning network predicts a Deepframe online, which is an extended basic block comprising a multi-path sequence (each path itself is composed of multiple basic blocks). We demonstrate the efficacy of Deepframe on spatial hardware accelerators and find the following: i) Deepframe can construct up to 5x (max: 27x) longer offload regions compared to prior approaches. ii) Surprisingly far-flung ILP (instruction-level parallelism) and MLP (memory-level parallelism) can be mined from the frames statically (5.5x increase in ILP and 10.5x increase in MLP). iii) The frames offloaded to the spatial accelerator have minimal side exits (mis-speculation) and achieve sufficient dynamic coverage to improve overall application performance (up to 9x improvement). We will be releasing open-source our end-to-end compiler prototype based on LLVM. Apala Guha, Naveen Vedula, Arrvindh Shriraman |
PACT | 1 |
| 2019 | μIR -An intermediate representation for transforming and optimizing the microarchitecture of application acceleratorsabstractCreating high quality application-specific accelerators requires us to make iterative changes to both algorithm behavior and microarchitecture, and this is a tedious and error-prone process. High-Level Synthesis (HLS) tools [5, 10] generate RTL for application accelerators from annotated software. Unfortunately, the generated RTL is challenging to change and optimize. The primary limitation of HLS is that the functionality and microarchitecture are conflated together in a single language (such as C++). Making changes to the accelerator design may require code restructuring, and microarchitecture optimizations are tied with program correctness. Amirali Sharifian, Reza Hojabr, Navid Rahimi, Sihao Liu, Apala Guha, Tony Nowatzki, Arrvindh Shriraman |
MICRO | 5 |
| 2018 | TAPAS: Generating Parallel Accelerators from Parallel ProgramsabstractHigh-level-synthesis (HLS) tools generate accelerators from software programs to ease the task of building hardware. Unfortunately, current HLS tools have limited support for concurrency, which impacts the speedup achievable with the generated accelerator. Current approaches only target fixed static patterns (e.g., pipeline, data-parallel kernels). This constraints the ability of software programmers to express concurrency. Moreover, the generated accelerator loses a key benefit of parallel hardware, dynamic asynchrony, and the potential to hide long latency and cache misses. We have developed TAPAS, an HLS toolchain for generating parallel accelerators from programs with dynamic parallelism. TAPAS is built on top of Tapir [22], [39], which embeds fork-join parallelism into the compiler's intermediate-representation. TAPAS leverages the compiler IR to identify parallelism and synthesizes the hardware logic. TAPAS provides first-class architecture support for spawning, coordinating and synchronizing tasks during accelerator execution. We demonstrate TAPAS can generate accelerators for concurrent programs with heterogeneous, nested and recursive parallelism. Our evaluation on Intel-Altera DE1-SoC and Arria-10 boards demonstrates that TAPAS generated accelerators achieve 20× the power efficiency of an Intel Xeon, while maintaining comparable performance. We also show that TAPAS enables lightweight tasks that can be spawned in '10 cycles and enables accelerators to exploit available fine-grain parallelism. TAPAS is a complete HLS toolchain for synthesizing parallel programs to accelerators and is open-sourced. Steve Margerm, Amirali Sharifian, Apala Guha, Arrvindh Shriraman, Gilles Pokam |
MICRO | 3 |
| 2016 | Chainsaw: Von-neumann accelerators to leverage fused instruction chainsabstractA central tenet behind accelerators is to partition a program execution into regions with different behavior (e.g., SIMD, Irregular, Compute-Intensive) and then use behavior-specialized architectures [1] for each region. It is unclear whether the gains in efficiency arise from recognizing that a simpler microarchitecture is sufficient for the acceleratable code region or the actual microarchitecture, or a combination of both. Many proposals [2], [3] seem to choose dataflow-based accelerators which encounters challenges with fabric utilization and static power when the available instruction parallelism is below the peak operation parallelism available [4]. In this paper, we develop, Chainsaw, a Von-Neumann based accelerator and demonstrate that many of the fundamental overheads (e.g., fetch-decode) can be amortized by adopting the appropriate instruction abstraction. The key insight is the notion of chains, which are compiler fused sequences of instructions. chains adapt to different acceleration behaviors by varying the length of the chains and the types of instructions that are fused into a chain. Chains convey the producer-consumer locality between dependent instructions, which the Chainsaw architecture then captures by temporally scheduling such operations on the same execution unit and uses pipeline registers to forward the values between dependent operations. Chainsaw is a generic multi-lane architecture (4-stage pipeline per lane) and does not require any specialized compound function units; it can be reloaded enabling it to accelerate multiple program paths. We have developed a complete LLVM-based compiler prototype and simulation infrastructure and demonstrated that a 8-lane Chainsaw is within 73% of the performance of an ideal dataflow architecture, while reducing the energy consumption by 45% compared to a 4-way OOO processor. Amirali Sharifian, Snehasish Kumar, Apala Guha, Arrvindh Shriraman |
MICRO | 3 |
| 2013 | Exascale workload characterization and architecture implicationsabstractEmerging exascale architectures bring forth new challenges related to heterogeneous systems power, energy, cost, and resilience. These new challenges require a shift from conventional paradigms in understanding how to best exploit and optimize these features and limitations. Our objective is to identify the top few dominant characteristics in a set of applications. Understanding these characteristics will allow the community to build and exploit customized architectures and tools best suited to optimize each dominant characteristic in the application domain. Every application will typically be composed of multiple characteristics and thus will use several of the customized accelerators and tools during its execution phases, with the eventual goal of using the entire system efficiently. In this poster, we describe a hybrid methodology, based on binary instrumentation, for characterizing scientific applications such as instruction mix and memory access patterns. We apply our methodology to proxy applications that are representative of a broad range of DOE scientific applications. With this empirical basis, we develop and validate statistical models that extrapolate application properties as a function of problem size. These models are then used to project the first quantitative characterization of an exascale computing workload, including computing and memory requirements. We evaluate the potential benefit of processor under memory, a radical new exascale architecture customization and understand how these new customization can impact applications. Prasanna Balaprakash, Darius Buntinas, Apala Guha, Rinku Gupta, Sri Hari Krishna Narayanan, Andrew A. Chien, Paul D. Hovland, Boyana Norris |
ISPASS | 4 |
| 2012 | Memory optimization of dynamic binary translators for embedded systemsabstractDynamic binary translators (DBTs) are becoming increasingly important because of their power and flexibility. DBT-based services are valuable for all types of platforms. However, the high memory demands of DBTs present an obstacle for embedded systems. Most research on DBT design has a performance focus, which often drives up the DBT memory demand. In this article, we present a memory-oriented approach to DBT design. We consider the class of translation-based DBTs and their sources of memory demand; cached translated code, cached auxiliary code and DBT data structures. We explore aspects of DBT design that impact these memory demand sources and present strategies to mitigate memory demand. We also explore performance optimizations for DBTs that handle memory demand by placing a limit on it, and repeatedly flush translations to stay within the limit, thereby replacing the memory demand problem with a performance degradation problem. Our optimizations that mitigate memory demand improve performance. Apala Guha, Kim M. Hazelwood, Mary Lou Soffa |
ACM Trans. Archit. Code Optim. | 1 |
| 2010 | Balancing memory and performance through selective flushing of software code cachesabstractDynamic binary translators (DBTs) are becoming increasingly important because of their power and flexibility. However, the high memory demands of DBTs present an obstacle for all platforms, and especially embedded systems. The memory demand is typically controlled by placing a limit on cached translations and forcing the DBT to flush all translations upon reaching the limit. This solution manifests as a performance inefficiency because many flushed translations require retranslation. Ideally, translations should be selectively flushed to minimize retranslations for a given memory limit. However, three obstacles exist:(1) it is difficult to predict which selections will minimize retranslation,(2) selective flushing results in greater book-keeping overheads than full flushing, and(3) the emergence of multicore processors and multi-threaded programming complicates most flushing algorithms. These issues have led to the widespread adoption of full flushing as a standard protocol. In this paper, we present a partial flushing approach aimed at reducing retranslation overhead and improving overall performance, given a fixed memory budget. Our technique applies uniformly to single-threaded and multi-threaded guest applications Apala Guha, Kim M. Hazelwood, Mary Lou Soffa |
CASES | 1 |
| 2010 | DBT path selection for holistic memory efficiency and performanceabstractDynamic binary translators(DBTs) provide powerful platforms for building dynamic program monitoring and adaptation tools. DBTs, however, have high memory demands because they cache translated code and auxiliary code to a software code cache and must also maintain data structures to support the code cache. The high memory demands make it difficult for memory-constrained embedded systems to take advantage of DBT-based tools. Previous research on DBT memory management focused on the translated code and auxiliary code only. However, we found that data structures are comparable to the code cache in size. We show that the translated code size, auxiliary code size and the data structure size interact in a complex manner, depending on the path selection (trace selection and link formation) strategy. Therefore, holistic memory efficiency (comprising translated code, auxiliary code and data structures) cannot be improved by focusing on the code cache only. In this paper, we use path selection for improving holistic memory efficiency which in turn impacts performance in memory-constrained environments. Although there has been previous research on path selection, such research only considered performance in memory-unconstrained environments. Apala Guha, Kim M. Hazelwood, Mary Lou Soffa |
VEE | 1 |
| 2007 | Reducing Exit Stub Memory Consumption in Code Caches
Apala Guha, Kim M. Hazelwood, Mary Lou Soffa |
HiPEAC | 1 |
| 2007 | Virtual Execution Environments: Support and ToolsabstractIn today's dynamic computing environments, the available resources and even underlying computation engine can change during the execution of a program. Additionally, current trends in software development favor the flexibility and cost-effectiveness of dynamically loaded components and libraries. Because of these trends, there has been increased research interest in virtual execution environments (VEEs) for delivering adaptable software suitable for today's rapidly changing, heterogeneous computing environments. In this project, we have been investigating tools and techniques to support implementation of VEEs using software dynamic translation (SDT). This paper highlights some of our recent results. One significant result is that we have developed novel translation techniques that reduce the memory and runtime overhead of SDT to negligible levels. We have also developed innovative debugging and instrumentation tools for SDT-based software environments. Together, these results make SDT-based systems viable for solving a wide range of pressing problems. The paper concludes with a discussion of how SDT may offer a solution to one such problem-inherent process variation in emerging chip multiprocessors. Apala Guha, Jason Hiser, Naveen Kumar 0002, Jing Yang 0003, Min Zhao 0009, Shukang Zhou, Bruce R. Childers, Jack W. Davidson, Kim M. Hazelwood, Mary Lou Soffa |
IPDPS | 1 |