EDBT 2026 Demo / reviewers in the wild / expert
Samuel Wiggins
dblp:352/9927
· DBLP profile ↗
6ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-0069-8213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Multi-agent Reinforcement Learning on Heterogeneous Platforms
Samuel Wiggins, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
Euro-Par (1) | 1 |
| 2026 | FAME: A Framework for Accelerating Independent Multi-Agent Reinforcement Learning on Heterogeneous PlatformsabstractMulti-Agent Reinforcement Learning (MARL) enables multiple autonomous agents to learn and act in a shared environment. Independent learning (IL) is a widely used MARL paradigm that underpins many real-world applications requiring efficient training at scale. However, accelerating IL at scale is non-trivial. Existing MARL frameworks rely on single-process execution and homogeneous hardware assumptions, limiting scalability and underutilizing modern heterogeneous platforms composed of CPUs, GPUs, and FPGAs. Addressing this gap requires new execution models that increase parallelism while preserving IL training semantics. In this work, we present FAME, a framework that distributes computation across heterogeneous hardware resources while providing flexible interfaces that allow MARL practitioners to prototype and test new IL approaches. FAME is composed of: (1) high-level APIs that simplify IL algorithm development, (2) a heterogeneous IL training protocol that supports concurrent agent training on multiple diverse devices, while maintaining algorithm-agnostic training semantics, (3) automatic hardware configuration generation that optimizes system throughput without needing users to manually fine-tune their system setup, and (4) dynamic load balancing among devices with different compute and memory characteristics. We demonstrate FAME’s capabilities using three representative IL algorithms on a heterogeneous node platform consisting of CPUs, GPUs, and FPGAs. Implementations generated using FAME achieve a geometric mean end-to-end training time speedup of 7.1 × over state-of-the-art implementations and up to 2.7 × speedup over additional highly parallel baselines developed in this work. Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
HPDC | 1 |
| 2025 | Accelerating Independent Multi-Agent Reinforcement Learning on Multi-GPU Platforms
Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
Euro-Par (3) | 1 |
| 2025 | ARC: A Runtime Engine for Accelerating Independent Multi-Agent Reinforcement Learning on Multi-Core ProcessorsabstractMulti-Agent Reinforcement Learning (MARL) enables multiple agents to optimize individual or joint objectives in a shared environment, with applications spanning robotics, autonomous driving, and financial systems. Independent Learning (IL), a simple yet effective MARL approach, trains agents independently without modeling inter-agent communication or explicit coordination. This simplicity reduces computational requirements, making CPU platforms an attractive alternative to accelerators such as GPUs for smaller model architectures typical of IL. However, existing CPU-based MARL implementations rely on a Single-Learner training scheme, which sequentially trains agent networks and fails to utilize the full potential of multicore CPUs. This limits scalability and introduces inefficiencies, particularly for large-scale MARL systems. In this work, we present ARC, a lightweight runtime engine designed to accelerate IL training on multi-core CPU platforms. ARC introduces an Independent Multi-Learner training scheme that parallelizes agent model updates, maximizing hardware utilization and scalability, while preserving training semantics. By exploring and selecting optimal parallelization strategies tailored to the user's hardware, ARC ensures seamless acceleration without manual configuration. Through experiments on state-of-the-art IL algorithms, we demonstrate an increased end-to-end speedup of up to$28.2 \times$while exploring only 5% of the configuration space. We open-source ARC, supporting multiple algorithms and providing significant performance improvements, thereby facilitating the development of scalable MARL applications. Samuel Wiggins, Nikunj Gupta, Grace Zgheib, Mahesh A. Iyer, Viktor Prasanna 0001 |
ICPADS | 1 |
| 2024 | A Heterogeneous Acceleration System for Attention-Based Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) is an emerging technology that has seen success in many AI applications. Multi-Actor-Attention-Critic (MAAC) is a state-of-the-art MARL algorithm that uses a Multi-Head Attention (MHA) mechanism to learn messages communicated among agents during the training process. Current implementations of MAAC using CPU and CPU-GPU platforms lack fine-grained parallelism among agents, sequentially executing each stage of the training loop, and their performance suffers from costly data movement involved in MHA communication learning. In this work, we develop the first high-throughput accelerator for MARL with attention-based communication on a CPU-FPGA heterogeneous system. We alleviate the limitations of existing implementations through a combination of data- and pipeline-parallel modules in our accelerator design and enable fine-grained system scheduling for exploiting concurrency among heterogeneous resources. Our design increases the overall system throughput by $4.6 \times$ and $4.1 \times$ compared to CPU and CPU-GPU implementations, respectively. Samuel Wiggins, Yuan Meng 0001, Mahesh A. Iyer, Viktor Prasanna 0001 |
FPL | 1 |
| 2023 | Characterizing Speed Performance of Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) has achieved significant success in large-scale AI systems and big-data applications such as smart grids, surveillance, etc. Existing advancements in MARL algorithms focus on improving the rewards obtained by introducing various mechanisms for inter-agent cooperation. However, these optimizations are usually compute- and memory-intensive, thus leading to suboptimal speed performance in end-to-end training time. In this work, we analyze the speed performance (i.e., latency-bounded throughput) as the key metric in MARL implementations. Specifically, we first introduce a taxonomy of MARL algorithms from an acceleration perspective categorized by (1) training scheme and (2) communication method. Using our taxonomy, we identify three state-of-the-art MARL algorithms - Multi-Agent Deep Deterministic Policy Gradient (MADDPG), Target-oriented Multi-agent Communication and Cooperation (ToM2C), and Networked Multi-Agent RL (NeurComm) - as target benchmark algorithms, and provide a systematic analysis of their performance bottlenecks on a homogeneous multi-core CPU platform. We justify the need for MARL latency-bounded throughput to be a key performance metric in future literature while also addressing opportunities for parallelization and acceleration. Samuel Wiggins, Yuan Meng 0001, Rajgopal Kannan, Viktor Prasanna 0001 |
DATA | 1 |