EDBT 2026 Demo / reviewers in the wild / expert
Akhil Langer
dblp:28/8857
· DBLP profile ↗
13ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 53% High-performance computing · 25% Cloud and datacenter computing · 6% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.5 | 2 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 Parallel Programming with Migratable Objects: Charm++ in Practice · SC 2014 |
High-performance computing
collective communication |
0.3 | 1 | 2018 | Framework for scalable intra-node collective operations using shared memory · SC 2018 |
Parallel and multicore computing › parallel computing › parallel communication
shared-memory communication |
0.3 | 1 | 2018 | Framework for scalable intra-node collective operations using shared memory · SC 2018 |
Parallel and multicore computing › parallel programming models
message passing |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
Parallel and multicore computing › parallel programming models › message passing
MPI implementation |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
High-performance computing
performance optimization at scale |
0.3 | 1 | 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1 · SC 2017 |
Energy-efficient computing › datacenter power management
power-constrained scheduling |
0.2 | 1 | 2014 | Maximizing Throughput of Overprovisioned HPC Data Centers Under a Strict Power Budget · SC 2014 |
Cloud and datacenter computing
resource management |
0.2 | 1 | 2014 | Maximizing Throughput of Overprovisioned HPC Data Centers Under a Strict Power Budget · SC 2014 |
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling |
0.2 | 1 | 2014 | Parallel Programming with Migratable Objects: Charm++ in Practice · SC 2014 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2014 | Maximizing Throughput of Overprovisioned HPC Data Centers Under a Strict Power Budget · SC 2014 |
High-performance computing › supercomputing
petascale computing |
0.1 | 1 | 2014 | Parallel Programming with Migratable Objects: Charm++ in Practice · SC 2014 |
High-performance computing
supercomputing |
0.1 | 1 | 2014 | Parallel Programming with Migratable Objects: Charm++ in Practice · SC 2014 |
Methods — techniques the papers use, named apart from their topics
instruction-level analysis · 0.3communication stack optimization · 0.3online resource management · 0.2migratable objects · 0.2adaptive runtime system · 0.2adaptive runtime · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Minimizing the usage of hardware counters for collective communication using triggered operations
Nusrat S. Islam, Gengbin Zheng, Sayantan Sur, Akhil Langer, María Jesús Garzarán |
Parallel Comput. | 4 |
| 2019 | Minimizing the usage of hardware counters for collective communication using triggered operationsabstractTriggered operations and counting events or counters are building blocks that can be used by communication libraries, such as MPI, to offload collective operations to the Host Fabric Interface (HFI) or Network Interface Card (NIC). Triggered operations can be used to schedule a network or arithmetic operation to occur in the future, when a trigger counter reaches a specified threshold. On completion of the operation, the value of a completion counter increases by one. With this mechanism, it is possible to create a chain of dependent operations, so that the execution of an operation is triggered when all its dependent operations have completed its execution. Nusrat S. Islam, Gengbin Zheng, Sayantan Sur, Akhil Langer, María Jesús Garzarán |
EuroMPI | 4 |
| 2018 | Framework for scalable intra-node collective operations using shared memory
Surabhi Jain, Rashid Kaleem, Marc Gamell, Akhil Langer, Dmitry Durnov, Alexander Sannikov, María Jesús Garzarán |
SC | 4 |
| 2017 | Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1abstractThis paper provides an in-depth analysis of the software overheads in the MPI performance-critical path and exposes mandatory performance overheads that are unavoidable based on the MPI-3.1 specification. We first present a highly optimized implementation of the MPI-3.1 standard in which the communication stack---all the way from the application to the low-level network communication API---takes only a few tens of instructions. We carefully study these instructions and analyze the root cause of the overheads based on specific requirements from the MPI standard that are unavoidable under the current MPI standard. We recommend potential changes to the MPI standard that can minimize these overheads. Our experimental results on a variety of network architectures and applications demonstrate significant benefits from our proposed changes. Kenneth Raffenetti, Abdelhalim Amer, Lena Oden, Charles Archer, Wesley Bland, Hajime Fujita 0002, Yanfei Guo, Tomislav Janjusic, Dmitry Durnov, Michael Blocksome, Min Si, Akhil Langer, Gengbin Zheng, Masamichi Takagi, Paul K. Coffman, Sayantan Sur, Alexander Sannikov, Sergey Oblomov, Michael Chuvelev, Masayuki Hatanaka, Paul F. Fischer, Thilina Ratnayaka, Matthew Otten, Misun Min, Pavan Balaji |
SC | 13 |
| 2015 | Split-and-Merge Method for Accelerating Convergence of Stochastic Linear Programs
Akhil Langer, Udatta S. Palekar |
ICORES | 1 |
| 2014 | An optimal distributed load balancing algorithm for homogeneous work unitsabstractMany parallel applications, for example, Adaptive Mesh Refinement simulations, need dynamic load balancing during the course of their execution because of dynamic variation in the computational load. We propose a novel tree-based fully distributed algorithm for load balancing homogeneous work units. The proposed algorithm achieves perfect load balance while doing minimum number of migrations of work units. Akhil Langer |
ICS | 1 |
| 2014 | Parallel Programming with Migratable Objects: Charm++ in PracticeabstractThe advent of petascale computing has introduced new challenges (e.g. Heterogeneity, system failure) for programming scalable parallel applications. Increased complexity and dynamism in science and engineering applications of today have further exacerbated the situation. Addressing these challenges requires more emphasis on concepts that were previously of secondary importance, including migratability, adaptivity, and runtime system introspection. In this paper, we leverage our experience with these concepts to demonstrate their applicability and efficacy for real world applications. Using the CHARM++ parallel programming framework, we present details on how these concepts can lead to development of applications that scale irrespective of the rough landscape of supercomputing technology. Empirical evaluation presented in this paper spans many miniapplications and real applications executed on modern supercomputers including Blue Gene/Q, Cray XE6, and Stampede. Bilge Acun, Abhishek Gupta 0002, Akhil Langer, Harshitha Menon, Eric Mikida, Xiang Ni, Michael P. Robson, Yanhua Sun, Ehsan Totoni, Lukasz Wesolowski, Laxmikant V. Kalé |
SC | 4 |
| 2014 | Maximizing Throughput of Overprovisioned HPC Data Centers Under a Strict Power BudgetabstractBuilding future generation supercomputers while constraining their power consumption is one of the biggest challenges faced by the HPC community. For example, US Department of Energy has set a goal of 20 MW for an exascale (1018 flops) supercomputer. To realize this goal, a lot of research is being done to revolutionize hardware design to build power efficient computers and network interconnects. In this work, we propose a software-based online resource management system that leverages hardware facilitated capability to constrain the power consumption of each node in order to optimally allocate power and nodes to a job. Our scheme uses this hardware capability in conjunction with an adaptive runtime system that can dynamically change the resource configuration of a running job allowing our resource manager to re-optimize allocation decisions to running jobs as new jobs arrive, or a running job terminates. We also propose a performance modeling scheme that estimates the essential power characteristics of a job at any scale. The proposed online resource manager uses these performance characteristics for making scheduling and resource allocation decisions that maximize the job throughput of the supercomputer under a given power budget. We demonstrate the benefits of our approach by using a mix of jobs with different power response characteristics. We show that with a power budget of 4:75 MW, we can obtain up to 5:2X improvement in job throughput when compared with the SLURM scheduling policy that is power-unaware. We corroborate our results with real experiments on a relatively small scale cluster, in which we obtain a 1:7X improvement. Osman Sarood, Akhil Langer, Abhishek Gupta 0002, Laxmikant V. Kalé |
SC | 2 |
| 2013 | Optimizing power allocation to CPU and memory subsystems in overprovisioned HPC systemsabstractEnergy consumption and power draw pose two major challenges to the HPC community for designing larger systems. Present day HPC systems consume as much as 10MW of electricity and this is fast becoming a bottleneck. Although energy bills will significantly increase with machine size, power consumption is a hard constraint that must be addressed. Intel's Running Average Power Limit (RAPL) toolkit is a recent feature that enables power capping of CPU and memory subsystems on modern hardware. In this paper, we use RAPL to evaluate the possibility of improving execution time efficiency of an application by capping power while adding more nodes. We profile the strong scaling of an application using different power caps for both CPU and memory subsystems. Our proposed interpolation scheme uses an application profile to optimize the number of nodes and the distribution of power between CPU and memory subsystems to minimize execution time under a strict power budget. We validate these estimates by running experiments on a 20-node (120 cores) Sandy Bridge cluster. Our experimental results closely match the model estimates and show speedups greater than 1.47X for all applications compared to not capping CPU and memory power. We demonstrate that the quality of solution that our interpolation scheme provides matches very closely to results obtained via exhaustive profiling. Osman Sarood, Akhil Langer, Laxmikant V. Kalé, Barry Rountree, Bronis R. de Supinski |
CLUSTER | 2 |
| 2013 | Parallel branch-and-bound for two-stage stochastic integer optimizationabstractMany real-world planning problems require searching for an optimal solution in the face of uncertain input. One approach to is to express them as a two-stage stochastic optimization problem where the search for an optimum in one stage is informed by the evaluation of multiple possible scenarios in the other stage. If integer solutions are required, then branch-and-bound techniques are the accepted norm. However, there has been little prior work in parallelizing and scaling branch-and-bound algorithms for stochastic optimization problems. In this paper, we explore the parallelization of a two-stage stochastic integer program solved using branch-and-bound. We present a range of factors that influence the parallel design for such problems. Unlike typical, iterative scientific applications, we encounter several interesting characteristics that make it challenging to realize a scalable design. We present two design variations that navigate some of these challenges. Our designs seek to increase the exposed parallelism while delegating sequential linear program solves to existing libraries. We evaluate the scalability of our designs using sample aircraft allocation problems for the US airfleet. It is important that these problems be solved quickly while evaluating large number of scenarios. Our attempts result in strong scaling to hundreds of cores for these datasets. We believe similar results are not common in literature, and that our experiences will feed usefully into further research on this topic. Akhil Langer, Ramprasad Venkataraman, Udatta S. Palekar, Laxmikant V. Kalé |
HiPC | 1 |
| 2012 | Performance Optimization of a Parallel, Two Stage Stochastic Linear ProgramabstractStochastic optimization is used in several high impact contexts to provide optimal solutions in the face of uncertainties. This paper explores the parallelization of two-stage stochastic resource allocation problems that seek an optimal solution in the first stage, while accounting for sudden changes in resource requirements by evaluating multiple possible scenarios in the second stage. Unlike typical scientific computing algorithms, linear programs (which are the individual grains of computation in our parallel design) have unpredictable and long execution times. This confounds both a priori load distribution as well as persistence-based dynamic load balancing techniques. We present a master-worker decomposition coupled with a pull-based work assignment scheme for load balance. We discuss some of the challenges encountered in optimizing both the master and the worker portions of the computations, and techniques to address them. Of note are cut retirement schemes for balancing memory requirements with duplicated worker computation, and scenario clustering for accelerating the evaluation of similar scenarios. We base our work in the context of a real application: the optimization of US military aircraft allocation to various cargo and personnel movement missions in the face of uncertain demands. We demonstrate scaling up to 122 cores of an intel 64 cluster, even for very small, but representative datasets. Our decision to eschew problem-specific decompositions has resulted in a parallel infrastructure that should be easily adapted to other similar problems. Similarly, we believe the techniques developed in this paper will be generally applicable to other contexts that require quick solutions to stochastic optimization problems. Akhil Langer, Ramprasad Venkataraman, Udatta S. Palekar, Laxmikant V. Kalé |
ICPADS | 1 |
| 2012 | Scalable Algorithms for Constructing Balanced Spanning Trees on System-Ranked Process Groups
Akhil Langer, Ramprasad Venkataraman, Laxmikant V. Kalé |
EuroMPI | 1 |
| 2012 | Scalable Algorithms for Distributed-Memory Adaptive Mesh RefinementabstractThis paper presents scalable algorithms and data structures for adaptive mesh refinement computations. We describe a novel mesh restructuring algorithm for adaptive mesh refinement computations that uses a constant number of collectives regardless of the refinement depth. To further increase scalability, we describe a localized hierarchical coordinate-based block indexing scheme in contrast to traditional linear numbering schemes, which incur unnecessary synchronization. In contrast to the existing approaches which take O(P) time and storage per process, our approach takes only constant time and has very small memory footprint. With these optimizations as well as an efficient mapping scheme, our algorithm is scalable and suitable for large, highly-refined meshes. We present strong-scaling experiments up to 2k ranks on Cray XK6, and 32k ranks on IBM Blue Gene/Q. Akhil Langer, Jonathan Lifflander, Phil Miller, Kuo-Chuan Pan, Laxmikant V. Kalé, Paul M. Ricker |
SBAC-PAD | 1 |