Nikunj Gupta

dblp:94/6162 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous Platforms
abstract
Multi-Agent Reinforcement Learning (MARL) enables multiple autonomous agents to learn and act in a shared environment. Independent learning (IL) is a widely used MARL paradigm that underpins many real-world applications requiring efficient training at scale. However, accelerating IL at scale is non-trivial. Existing MARL frameworks rely on single-process execution and homogeneous hardware assumptions, limiting scalability and underutilizing modern heterogeneous platforms composed of CPUs, GPUs, and FPGAs. Addressing this gap requires new execution models that increase parallelism while preserving IL training semantics. In this work, we present FAME, a framework that distributes computation across heterogeneous hardware resources while providing flexible interfaces that allow MARL practitioners to prototype and test new IL approaches. FAME is composed of: (1) high-level APIs that simplify IL algorithm development, (2) a heterogeneous IL training protocol that supports concurrent agent training on multiple diverse devices, while maintaining algorithm-agnostic training semantics, (3) automatic hardware configuration generation that optimizes system throughput without needing users to manually fine-tune their system setup, and (4) dynamic load balancing among devices with different compute and memory characteristics. We demonstrate FAME’s capabilities using three representative IL algorithms on a heterogeneous node platform consisting of CPUs, GPUs, and FPGAs. Implementations generated using FAME achieve a geometric mean end-to-end training time speedup of 7.1 × over state-of-the-art implementations and up to 2.7 × speedup over additional highly parallel baselines developed in this work.
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
HPDC2
2025 Accelerating Independent Multi-Agent Reinforcement Learning on Multi-GPU Platforms
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
Euro-Par (3)2
2025 ARC: A Runtime Engine for Accelerating Independent Multi-Agent Reinforcement Learning on Multi-Core Processors
abstract
Multi-Agent Reinforcement Learning (MARL) enables multiple agents to optimize individual or joint objectives in a shared environment, with applications spanning robotics, autonomous driving, and financial systems. Independent Learning (IL), a simple yet effective MARL approach, trains agents independently without modeling inter-agent communication or explicit coordination. This simplicity reduces computational requirements, making CPU platforms an attractive alternative to accelerators such as GPUs for smaller model architectures typical of IL. However, existing CPU-based MARL implementations rely on a Single-Learner training scheme, which sequentially trains agent networks and fails to utilize the full potential of multicore CPUs. This limits scalability and introduces inefficiencies, particularly for large-scale MARL systems. In this work, we present ARC, a lightweight runtime engine designed to accelerate IL training on multi-core CPU platforms. ARC introduces an Independent Multi-Learner training scheme that parallelizes agent model updates, maximizing hardware utilization and scalability, while preserving training semantics. By exploring and selecting optimal parallelization strategies tailored to the user's hardware, ARC ensures seamless acceleration without manual configuration. Through experiments on state-of-the-art IL algorithms, we demonstrate an increased end-to-end speedup of up to$28.2 \times$while exploring only 5% of the configuration space. We open-source ARC, supporting multiple algorithms and providing significant performance improvements, thereby facilitating the development of scalable MARL applications.
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001
ICPADS2
2025 hammer: Multi-level coordination of reinforcement learning agents via learned messaging
Nikunj Gupta, G. Srinivasaraghavan 0001, Swarup Mohalik, Matthew E. Taylor
Neural Comput. Appl.1
2023 Planning Multiple Epidemic Interventions with Reinforcement Learning
abstract
Combating an epidemic entails finding a plan that describes when and how to apply different interventions, such as mask-wearing mandates, vaccinations, school or workplace closures. An optimal plan will curb an epidemic with minimal loss of life, disease burden, and economic cost. Finding an optimal plan is an intractable computational problem in realistic settings. Policy-makers, however, would greatly benefit from tools that can efficiently search for plans that minimize disease and economic costs especially when considering multiple possible interventions over a continuous and complex action space given a continuous and equally complex state space. We formulate this problem as a Markov decision process. Our formulation is unique in its ability to represent multiple continuous interventions over any disease model defined by ordinary differential equations. We illustrate how to effectively apply state-of-the-art actor-critic reinforcement learning algorithms (PPO and SAC) to search for plans that minimize overall costs. We empirically evaluate the learning performance of these algorithms and compare their performance to hand-crafted baselines that mimic plans constructed by policy-makers. Our method outperforms baselines. Our work confirms the viability of a computational approach to support policy-makers.
Anh L. Mai, Nikunj Gupta, Azza Abouzeid, Dennis E. Shasha
IJCAI2
2020 Performance Evaluation of ParalleX Execution model on Arm-based Platforms
abstract
The HPC community shows a keen interest in creating diversity in the CPU ecosystem. The advent of Arm-based processors provides an alternative to the existing HPC ecosystem, which is primarily dominated by x86 processors. In this paper, we port an Asynchronous Many-Task runtime system based on the ParalleX model, i.e., High Performance ParalleX (HPX), and evaluate it on the Arm ecosystem with a suite of benchmarks. We wrote these benchmarks with an emphasis on vectorization and distributed scaling. We present the performance results on a variety of Arm processors and compare it with their x86 brethren from Intel. We show that the results obtained are equally good or better than their x86 brethren. Finally, we also discuss a few drawbacks of the present Arm ecosystem.
Nikunj Gupta, Rohit Ashiwal, Bine Brank, Sateesh K. Peddoju, Dirk Pleiter
CLUSTER1