EDBT 2026 Demo / reviewers in the wild / expert
Hari K. Pyla
dblp:35/6056
· DBLP profile ↗
5ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2011 | Exploiting coarse-grain speculative parallelism · OOPSLA 2011 |
Parallel and multicore computing
speculative parallelization |
0.1 | 1 | 2011 | Exploiting coarse-grain speculative parallelism · OOPSLA 2011 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP extensions |
0.0 | 1 | 2011 | Exploiting coarse-grain speculative parallelism · OOPSLA 2011 |
Methods — techniques the papers use, named apart from their topics
surrogate code blocks · 0.1runtime system design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Transparent runtime deadlock eliminationabstractThread based concurrent programming is hard due to the potential of concurrency bugs (e.g., data races, atomicity violations, deadlocks, and order violations). While data races and atomicity violations can be ameliorated with appropriate synchronization (a non-trivial problem in itself !), deadlocks require fairly complex avoidance techniques which may fail when the order of lock acquisition is not known apriori [2]. The goal of this research is to present an efficient and practical system that transparently detects and eliminates deadlocks in real-world multi-threaded applications. Hari K. Pyla, Srinidhi Varadarajan |
PACT | 1 |
| 2011 | Exploiting coarse-grain speculative parallelismabstractSpeculative execution at coarse granularities (e.g., code-blocks, methods, algorithms) offers a promising programming model for exploiting parallelism on modern architectures. In this paper we present Anumita, a framework that includes programming constructs and a supporting runtime system to enable the use of coarse-grain speculation to improve program performance, without burdening the programmer with the complexity of creating, managing and retiring speculations. Speculations may be composed by specifying surrogate code blocks at any arbitrary granularity, which are then executed concurrently, with a single winner ultimately modifying program state. Anumita provides expressive semantics for winner selection that go beyond time to solution to include user-defined notions of quality of solution. Anumita can be used to improve the performance of hard to parallelize algorithms whose performance is highly dependent on input data. Anumita is implemented as a user-level runtime with programming interfaces to C, C++, Fortran and as an OpenMP extension. Performance results from several applications show the efficacy of using coarse-grain speculation to achieve (a) robustness when surrogates fail and (b) significant speedup over static algorithm choices. Hari K. Pyla, Calvin J. Ribbens, Srinidhi Varadarajan |
OOPSLA | 1 |
| 2010 | Avoiding deadlock avoidanceabstractThe evolution of processor architectures from single core designs with increasing clock frequencies to multi-core designs with relatively stable clock frequencies has fundamentally altered application design. Since application programmers can no longer rely on clock frequency increases to boost performance, over the last several years, there has been significant emphasis on application level threading to achieve performance gains. A core problem with concurrent programming using threads is the potential for deadlocks. Even well-written codes that spend an inordinate amount of effort in deadlock avoidance cannot always avoid deadlocks, particularly when the order of lock acquisitions is not known a priori. Furthermore, arbitrarily composing lock based codes may result in deadlock - one of the primary motivations for transactional memory. In this paper, we present a language independent runtime system called Sammati that provides automatic deadlock detection and recovery for threaded applications that use the POSIX threads (pthreads) interface - the de facto standard for UNIX systems. The runtime is implemented as a pre-loadable library and does not require either the application source code or recompiling/relinking phases, enabling its use for existing applications with arbitrary multi-threading models. Performance evaluation of the runtime with unmodified SPLASH, Phoenix and synthetic benchmark suites shows that it is scalable, with speedup comparable to baseline execution with modest memory overhead. Hari K. Pyla, Srinidhi Varadarajan |
PACT | 1 |
| 2008 | System-level, thermal-aware, fully-loaded process schedulingabstractProcessor power consumption produces significant heat and can result in higher average operating temperatures. High operating temperatures can lead to reduced reliability and at times thermal emergencies. Previous thermal-aware techniques use dynamic voltage and frequency scaling (DVFS) or multithreaded or multicore process migration to reduce thermals. However, these methods do not gracefully handle scenarios where processors are fully loaded, i.e. there are no free threads or cores for process scheduling. We propose techniques to reduce processor temperature when processors are fully loaded. We use system-level compiler support and dynamic runtime instrumentation to identify the relative thermal intensity of processes. We implement a thermal-aware process scheduling algorithm that reduces processor thermals while maintaining application throughput. We favor "cool" processes by reducing time slice allocations for "hot" processes. Results indicate that our thermal-aware scheduling can reduce processor thermals by up to 3 degrees Celsius with little to no loss in application throughput. Dong Li 0001, Hung-Ching Chang, Hari K. Pyla, Kirk W. Cameron |
IPDPS | 3 |
| 2007 | Tempest: A portable tool to identify hot spots in parallel codeabstractCompute clusters are consuming more power at higher densities than ever before. This results in increased thermal dissipation, the need for powerful cooling systems, and ultimately a reduction in system reliability as temperatures increase. Over the past several years, the research community has reacted to this problem by producing software tools such as HotSpot and Mercury to estimate system thermal characteristics and validate thermal-management techniques. While these tools are flexible and useful, they suffer several limitations. For the average user such simulation tools can be cumbersome to use. These tools may take significant time and expertise to port to different systems. Lastly, such tools produce significant detail and accuracy at the expense of execution time enough to prohibit iterative testing. We propose a fast, easy to use, accurate, portable software tool called Tempest (for temperature estimator) that leverages emergent thermal sensors to enable user profiling, evaluating, and reducing the thermal characteristics of systems and applications. In this paper, we illustrate the use of Tempest to analyze the thermal effects of various parallel benchmarks in clusters. Kirk W. Cameron, Hari K. Pyla, Srinidhi Varadarajan |
ICPP | 2 |