EDBT 2026 Demo / reviewers in the wild / expert
Prakash Prabhu
dblp:32/5749
· DBLP profile ↗
10ranked-venue papers
5as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 4 first-authorSystems, architecture and hardware · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 76% Performance modeling and evaluation · 15% GPUs and heterogeneous computing · 8% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 90% Runtime systems and virtual machines · 10% |
Topics — the 7 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
pipeline parallelism |
0.8 | 1 | 2024 | Practical Performance Guarantees for Pipelined DNN Inference · ICML 2024 |
Mathematical optimization
discrete optimization |
0.8 | 1 | 2024 | Practical Performance Guarantees for Pipelined DNN Inference · ICML 2024 |
Compilers and program optimization › parallelization
automatic parallelization |
0.1 | 1 | 2012 | Speculative separation for privatization and reductions · PLDI 2012 |
Parallel and multicore computing › parallel programming models
automatic parallelization |
0.1 | 1 | 2012 | Speculative separation for privatization and reductions · PLDI 2012 |
GPUs and heterogeneous computing › GPU communication
CPU-GPU communication |
0.1 | 1 | 2011 | Automatic CPU-GPU communication management and optimization · PLDI 2011 |
Parallel and multicore computing
speculative parallelization |
0.1 | 1 | 2010 | Safe programmable speculative parallelism · PLDI 2010 |
Compilers and program optimization
parallelization |
0.0 | 1 | 2011 | Commutative set: a language extension for implicit parallel programming · PLDI 2011 |
Methods — techniques the papers use, named apart from their topics
combinatorial bounds · 1.5mixed-integer programming · 0.8mixed integer programming · 0.8speculative execution · 0.3runtime validation · 0.3semantic commutativity assertions · 0.2runtime management · 0.2compiler transformation · 0.2commutative set · 0.2value speculation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Practical Performance Guarantees for Pipelined DNN InferenceabstractWe optimize pipeline parallelism for deep neural network (DNN) inference by partitioning model graphs into $k$ stages and minimizing the running time of the bottleneck stage, including communication. We give practical and effective algorithms for this NP-hard problem, but our emphasis is on tackling the practitioner’s dilemma of deciding when a solution is good enough. To this end, we design novel mixed integer programming (MIP) relaxations for proving lower bounds. Applying these methods to a diverse testbed of 369 production models, for $k \in \\{2, 4, 8, 16, 32, 64\\}$, we empirically show that these lower bounds are strong enough to be useful in practice. Our lower bounds are substantially stronger than standard combinatorial bounds. For example, evaluated via geometric means across a production testbed with $k = 16$ pipeline stages, our MIP formulations raise the lower bound from 0.4598 to 0.9452, expressed as a fraction of the best partition found. In other words, our improved lower bounds close the optimality gap by a factor of 9.855x. Aaron Archer, Matthew Fahrbach, Kuikui Liu, Prakash Prabhu |
ICML | 4 |
| 2018 | MemoDyn: exploiting weakly consistent data structures for dynamic parallel memoizationabstractSeveral classes of algorithms for combinatorial search and optimization problems employ memoization data structures to speed up their serial convergence. However, accesses to these data structures impose dependences that obstruct program parallelization. Such programs often continue to function correctly even when queries into these data structures return a partial view of their contents. Weakening the consistency of these data structures can unleash new parallelism opportunities, potentially at the cost of additional computation. These opportunities must, therefore, be carefully exploited for overall speedup. This paper presents MemoDyn, a framework for parallelizing loops that access data structures with weakly consistent semantics. MemoDyn provides programming abstractions to express weak semantics, and consists of a parallelizing compiler and a runtime system that automatically and adaptively exploit the semantics for optimized parallel execution. Evaluation of MemoDyn shows that it achieves efficient parallelization, providing significant improvements over competing techniques in terms of both runtime performance and solution quality. Prakash Prabhu, Stephen R. Beard, Sotiris Apostolakis, Ayal Zaks, David I. August |
PACT | 1 |
| 2016 | Speculatively Exploiting Cross-Invocation ParallelismabstractAutomatic parallelization has shown promise in producing scalable multi-threaded programs for multi-core architectures. Most existing automatic techniques parallelize independent loops and insert global synchronization between loop invocations. For programs with many loop invocations, frequent synchronization often becomes the performance bottleneck. Some techniques exploit cross-invocation parallelism to overcome this problem. Using static analysis, they partition iterations among threads to avoid cross-thread dependences. However, this approach may fail if dependence pattern information is not available at compile time. To address this limitation, this work proposes SpecCross--the first automatic parallelization technique to exploit cross-invocation parallelism using speculation. With speculation, iterations from different loop invocations can execute concurrently, and the program synchronizes only on misspeculation. This allows SpecCross to adapt to dependence patterns that only manifest on particular inputs at runtime. Evaluation on eight programs shows that SpecCross achieves a geomean speedup of 3.43x over parallel execution without cross-invocation parallelization. Jialu Huang, Prakash Prabhu, Thomas B. Jablin, Soumyadeep Ghosh, Sotiris Apostolakis, Jae W. Lee, David I. August |
PACT | 2 |
| 2012 | Dynamically managed data for CPU-GPU architecturesabstractGPUs are flexible parallel processors capable of accelerating real applications. To exploit them, programmers must ensure a consistent program state between the CPU and GPU memories by managing data. Manually managing data is tedious and error-prone. In prior work on automatic CPU-GPU data management, alias analysis quality limits performance, and type-inference quality limits applicability. This paper presents Dynamically Managed Data (DyManD), the first automatic system to manage complex and recursive data-structures without static analyses. By replacing static analyses with a dynamic run-time system, DyManD overcomes the performance limitations of alias analysis and enables management for complex and recursive data-structures. DyManD-enabled GPU parallelization matches the performance of prior work equipped with perfectly precise alias analysis for 27 programs and demonstrates improved applicability on programs not previously managed automatically. Thomas B. Jablin, James A. Jablin, Prakash Prabhu, David I. August |
CGO | 3 |
| 2012 | Speculative separation for privatization and reductionsabstractAutomatic parallelization is a promising strategy to improve application performance in the multicore era. However, common programming practices such as the reuse of data structures introduce artificial constraints that obstruct automatic parallelization. Privatization relieves these constraints by replicating data structures, thus enabling scalable parallelization. Prior privatization schemes are limited to arrays and scalar variables because they are sensitive to the layout of dynamic data structures. This work presents Privateer, the first fully automatic privatization system to handle dynamic and recursive data structures, even in languages with unrestricted pointers. To reduce sensitivity to memory layout, Privateer speculatively separates memory objects. Privateer's lightweight runtime system validates speculative separation and speculative privatization to ensure correct parallel execution. Privateer enables automatic parallelization of general-purpose C/C++ applications, yielding a geomean whole-program speedup of 11.4x over best sequential execution on 24 cores, while non-speculative parallelization yields only 0.93x. Nick P. Johnson, Hanjun Kim 0001, Prakash Prabhu, Ayal Zaks, David I. August |
PLDI | 3 |
| 2011 | Interprocedural Exception Analysis for C++
Prakash Prabhu, Naoto Maeda, Gogul Balakrishnan, Franjo Ivancic, Aarti Gupta |
ECOOP | 1 |
| 2011 | Automatic CPU-GPU communication management and optimizationabstractThe performance benefits of GPU parallelism can be enormous, but unlocking this performance potential is challenging. The applicability and performance of GPU parallelizations is limited by the complexities of CPU-GPU communication. To address these communications problems, this paper presents the first fully automatic system for managing and optimizing CPU-GPU communcation. This system, called the CPU-GPU Communication Manager (CGCM), consists of a run-time library and a set of compiler transformations that work together to manage and optimize CPU-GPU communication without depending on the strength of static compile-time analyses or on programmer-supplied annotations. CGCM eases manual GPU parallelizations and improves the applicability and performance of automatic GPU parallelizations. For 24 programs, CGCM-enabled automatic GPU parallelization yields a whole program geomean speedup of 5.36x over the best sequential CPU-only execution. Thomas B. Jablin, Prakash Prabhu, James A. Jablin, Nick P. Johnson, Stephen R. Beard, David I. August |
PLDI | 2 |
| 2011 | Commutative set: a language extension for implicit parallel programmingabstractSequential programming models express a total program order, of which a partial order must be respected. This inhibits parallelizing tools from extracting scalable performance. Programmer written semantic commutativity assertions provide a natural way of relaxing this partial order, thereby exposing parallelism implicitly in a program. Existing implicit parallel programming models based on semantic commutativity either require additional programming extensions, or have limited expressiveness. This paper presents a generalized semantic commutativity based programming extension, called Commutative Set (COMMSET), and associated compiler technology that enables multiple forms of parallelism. COMMSET expressions are syntactically succinct and enable the programmer to specify commutativity relations between groups of arbitrary structured code blocks. Using only this construct, serializing constraints that inhibit parallelization can be relaxed, independent of any particular parallelization strategy or concurrency control mechanism. COMMSET enables well performing parallelizations in cases where they were inapplicable or non-performing before. By extending eight sequential programs with only 8 annotations per program on average, COMMSET and the associated compiler technology produced a geomean speedup of 5.7x on eight cores compared to 1.5x for the best non-COMMSET parallelization. Prakash Prabhu, Soumyadeep Ghosh, Yun Zhang 0005, Nick P. Johnson, David I. August |
PLDI | 1 |
| 2010 | Safe programmable speculative parallelismabstractExecution order constraints imposed by dependences can serialize computation, preventing parallelization of code and algorithms. Speculating on the value(s) carried by dependences is one way to break such critical dependences. Value speculation has been used effectively at a low level, by compilers and hardware. In this paper, we focus on the use of speculation by programmers as an algorithmic paradigm to parallelize seemingly sequential code. Prakash Prabhu, G. Ramalingam, Kapil Vaswani |
PLDI | 1 |
| 2008 | Field Flow Sensitive Pointer and Escape Analysis for Java Using Heap Array SSA
Prakash Prabhu, Priti Shankar |
SAS | 1 |