EDBT 2026 Demo / reviewers in the wild / expert
Sanmukh R. Kuppannagari
dblp:158/4879 · also Sanmukh Kuppannagari, Sanmukh Rao Kuppannagari
· DBLP profile ↗
26ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-2062-1483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | torch_cufft: Extending PyTorch FFT Capacity with Multi-GPU and Descriptor-Resident cuFFTXt ExecutionabstractLarge scientific images and spectral-learning workloads often require two-dimensional Fast Fourier Transforms (2D FFTs) that exceed single-GPU memory. This matters for Fourier Neural Operators (FNOs) and Transform Once (T1)-style models, where frequency-domain computation is central to the learning workflow. PyTorch provides convenient FFT APIs, but scaling these transforms across multiple GPUs requires lower-level libraries and careful memory-layout management. Yash Malhotra, Sanmukh R. Kuppannagari |
HPDC | 2 |
| 2026 | Achieving Low Latency Inference on High Resolution Images by Exploiting Sparsity in Vision Transformers
Changxin Li, Sanmukh R. Kuppannagari |
IPDPS | 2 |
| 2025 | Optimizing Deployment of Unstructured Group Convolutions for Low Latency InferenceabstractGroup convolutions are widely adopted in modern CNN architectures such as CondenseNet and ShuffleNet to enable efficient inference on GPUs. However, when the connection pattern between input and output channels does not exhibit regularity (unstructured group convolution), popular deep learning frameworks (e.g., PyTorch) often struggle with load balancing and data reuse issues leading to reduced performance. In this paper, we present a comprehensive optimization framework that combines a Knapsack-based partitioning approach with Integer Linear Programming (ILP) and advanced matrix reordering to optimize the deployment of unstructured group convolutions which are used in popular models such as CondenseNets that learn group connections. Specifically, we use knapsack algorithm to determine partition (group) sizes for the connections to minimize execution time and use an Integer Linear Programming (ILP) to assign connections to the partitions (group) output by the knapsack algorithm. We also employ three matrix reordering strategies-Hierarchical Clustering (HC), Iterative Clustering (IC), and Reverse Cuthill-McKee (RCM) on the matrix representing the input-output connectivity pattern to further improve the performance of our scheduling algorithm. Our experiments on ShuffleNet and CondenseNet demonstrate up to$1.9 \times$speedups over PyTorch. Furthermore, augmenting ILP with reordering achieves an additional$1.3 \times$improvement demonstrating the importance of optimizing for load balancing and data reuse. Changxin Li, Sanmukh R. Kuppannagari |
HiPC | 2 |
| 2025 | Longer Attention Span: Increasing Transformer Context Length With Sparse Graph Processing TechniquesabstractTransformers have demonstrated great success in numerous domains including natural language processing and bioinformatics. This success stems from the use of the attention mechanism by these models in order to represent and propagate pairwise interactions between individual tokens of sequential data. However, the primary limitation of this operation is its quadratic memory and time complexity in relation to the input's context length - the length of a sequence over which the interactions need to be captured. This significantly limits the length of sequences that can be inferred upon by these models. Extensive research has been conducted to reduce the number of pairwise interactions to sub-quadratic in relation to the context length by introducing sparsity into the attention mechanism through the development of sparse attention masks. However, efficient implementations that achieve “true sparsity” are lacking. In this work, we address this issue by proposing a graph computing view of attention where tokens are perceived as nodes of the graph and the attention mask determines the edges of the graph. Using this view, we develop graph processing algorithms to implement the attention mechanism. Both theoretically and empirically, we demonstrate that our algorithms only perform the needed computations, i.e., they are work optimal. We also perform extensive experimentation using popular attention masks to explore the impact of sparsity on execution time and achievable context length. Our experiments demonstrate significant speedups in execution times compared to state-of-theart attention implementations such as FlashAttention for large sequence lengths. We also demonstrate that our algorithms are able to achieve extremely long sequence lengths of as high as 160 million on a single NVIDIA A100 GPU (SXM4 80 GB). GitHub: https://github.com/KLab-AI3/Graph-Processing-Attention-IPDPS-2025 Nathaniel Tomczak, Sanmukh R. Kuppannagari |
IPDPS | 2 |
| 2024 | Exploring Algorithmic Design Choices for Low Latency CNN DeploymentabstractConvolutional Neural Networks (CNN s) have demonstrated significant success in advancing image and video processing technologies, significantly outperforming traditional methods in both accuracy and efficiency. However, deploying CNN s effectively across diverse hardware platforms often faces the challenge of latency, which can critically impact real-time processing applications. In this work, we explore algorithmic design choices aimed at reducing latency in CNN deployments. We implement five convolution algorithms using SYCL and inte-grate them into three popular CNN models: VGG 16, Resnetl0l, and Inception V 4. By replacing the standard PyTorch Conv2d function with our SY CL- based implementations, we evaluate the execution time of each convolution layer and the overall model on G PU s. Our extensive experiments benchmark the performance of these algorithms against the baseline implementations of the PyTorch and Pytorch Extension for Intel. The results demonstrate significant improvements in execution time, underscoring the potential of these algorithmic choices for achieving low latency in CNN deployments. Changxin Li, Sanmukh R. Kuppannagari |
HiPC | 2 |
| 2023 | PARAG: PIM Architecture for Real-Time Acceleration of GCNsabstractGraph Convolutional Networks (GCNs) have successfully incorporated deep learning to graph structures for social network analysis, bio-informatics, etc. The execution pattern of GCNs is a hybrid of graph processing and neural networks which poses unique and significant challenges for hardware implementation. Graph processing involves a large amount of irregular memory access with little computation whereas processing of neural networks involves a large number of operations with regular memory access. Existing graph processing and neural network accelerators are therefore inefficient for computing GCNs. This paper presents Parag, processing in memory (PIM) architecture for GCN computation. It consists of customized logic with minuscule computing units called Neural Processing Elements (NPEs) interfaced to each bank of the DRAM to support parallel graph processing and neural network computation. It utilizes the massive internal parallelism of DRAM to accelerate the GCN execution with high energy efficiency. Simulation results for inference of GCN over standard datasets show a latency and energy reduction by three orders of magnitude over a CPU implementation. When compared to a state-of-the-art PIM architecture, PARAG achieves on an average 4x reduction in latency and 4.23x reduction in the energy-delay-product (EDP). Gian Singh, Sanmukh R. Kuppannagari, Sarma B. K. Vrudhula |
HiPC | 2 |
| 2023 | Behind-the-Meter Solar Generation Disaggregation at Varying Aggregation Levels Using Consumer Mixture ModelsabstractThe increasing penetration of solar PhotoVoltaic (PV) panels in residential markets is leading to increasing solar generation hidden behind metering instruments of utility companies. Current metering infrastructure only measures the net load (sum of consumption and solar generation signals) from customers. However, it is desirable to observe solar generation separate from load consumption for grid optimizations. To enable that, we propose an unsupervised Behind-the-Meter (BTM) disaggregation model that utilizes a novel Consumer Mixture Model (CMM) for the modelling of consumption load in the disaggregation model. CMM uses consumption patterns of neighboring customers without PVs installed as features for modelling. We evaluate our model on an Australia dataset and use a load regression model and a state-of-the-art disaggregation model as baselines. We show that our model outperforms the baselines – the Mean Average Error of disaggregation results of our model was 28.37% lower than the state-of-the-art model. Additionally, we show that our model is agnostic to aggregation levels. This enables the utilities to focus on specific grid portions as needed. Chung Ming Cheung, Sanmukh R. Kuppannagari, Ajitesh Srivastava, Rajgopal Kannan, Viktor Prasanna 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2022 | NTTGen: a framework for generating low latency NTT implementations on FPGAabstractHomomorphic encryption (HE) is a promising technique to ensure the security and privacy of applications in the cloud. Number Theoretic Transform (NTT) is a key operation in HE-based applications. HE requires vastly different NTT parameters to meet the performance and security requirements of applications. The increasing compute capabilities and flexibility of FPGAs make them attractive to accelerate NTT. However, programming FPGA still involves hardware design expertise and significant development effort. To close the gap, we propose NTTGen, a framework to automatically generate low latency NTT designs targeting HE-based applications. NTTGen takes application parameters, latency and hardware resource constraints as input, determines the design parameters, and produces synthesizable Verilog code as output. Low latency NTT implementations are obtained by varying the data, pipeline and batch parallelism. NTTGen utilizes streaming permutation network to reduce the interconnect complexity between stages in the NTT computation. The framework supports two types of NTT cores to perform modular arithmetic, the key computation in NTT: a low latency and resource efficient NTT core for a specific class of prime moduli and a general purpose NTT core for other primes. We further develop a design space exploration flow to identify the hardware design parameters of an optimal design. We evaluate NTTGen by generating designs for various NTT parameters. The designs result in up to 2.9X improvement in latency over the state-of-the-art FPGA implementations. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
CF | 2 |
| 2022 | FPGA Accelerator for Homomorphic Encrypted Sparse Convolutional Neural Network InferenceabstractHomomorphic Encryption (HE) is a promising solution to the increasing concerns of privacy in machine learning. But HE-based CNN inference remains impractically slow. Pruning can significantly reduce the compute and memory footprint of CNNs. However, homomorphic encrypted Sparse Convolutional Neural Networks (SCNN) have vastly different compute and memory characteristics compared with unencrypted SCNN. Simply extending the design principles of existing SCNN accelerators may offset the potential acceleration offered by sparsity. To realize fast execution, we propose an FPGA accelerator to speedup the computation of linear layers, the main computational bottleneck in HE SCNN batch inference. First, we analyze the memory requirements of various linear layers in HE SCNN and discuss the unique challenges. Motivated by the analysis, we present a novel dataflow specially designed to optimize HE SCNN data reuse coupled with an efficient scheduling policy that minimizes on-chip SRAM access conflicts. Leveraging the proposed dataflow and scheduling algorithm, we demonstrate the first end-to-end acceleration of HE SCNN batch inference targeting CPU-FPGA heterogeneous platforms. For a batch of 8K images, our design achieves up to 5.6× speedup in inference latency compared with the CPU-only solution for widely studied 6-layer and 11-layer HE CNNs. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FCCM | 2 |
| 2022 | Bandwidth Efficient Homomorphic Encrypted Matrix Vector Multiplication Accelerator on FPGAabstractHomomorphic Encryption (HE) is a promising solution to the increasing concerns of privacy in Machine Learning (ML) as it enables computations directly on encrypted data. However, it imposes significant overhead on the compute system and remains impractically slow. Prior works have proposed efficient FPGA implementations of basic HE primitives such as number theoretic transform (NTT), key switching, etc. Composing the primitives together to realize higher level ML computation is still a challenge due to the large data transfer overhead. In this work, we propose an efficient FPGA implementation of HE Matrix Vector Multiplication$(\mathbf{M}\times \mathbf{V})$, a key kernel in HE-based Machine Learning applications. By analyzing the data reuse characteristics and the encryption overhead of HE$\mathbf{M}\times \mathbf{V}$, we show that simply using the principles of unencrypted$\mathbf{M}\times \mathbf{V}$to design accelerators for HE$\mathbf{M}\times \mathbf{V}$can lead to a significant amount of DRAM data transfers. We tackle the computation and data transfer challenges by proposing a bandwidth efficient dataflow that is specially optimized for HE$\mathbf{M}\times \mathbf{V}$. We identify highly reused data entities in HE$\mathbf{M}\times \mathbf{V}$and efficiently utilize the on-chip SRAM to reduce the DRAM data transfers. To speed up the computation of HE$\mathbf{M}\times \mathbf{V}$, we exploit three types of parallelism: partial sum parallelism, residual polynomial parallelism and coefficient parallelism. Leveraging these innovations, we demonstrate the first FPGA accelerator for HE matrix vector multiplication. Evaluation on 7 HE$\mathbf{M}\times \mathbf{V}$benchmarks shows that our FPGA accelerator is up to$3.8\times$(GeoMean$2.8\times$) faster compared to the 64-thread CPU implementation. Yang Yang 0111, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FPT | 2 |
| 2022 | Input Feature Pruning for Accelerating GNN Inference on Heterogeneous PlatformsabstractGraph Neural Networks (GNNs) are an emerging class of machine learning models which utilize structured graph information and node features to reduce high-dimensional input data to low-dimensional embeddings, from which predictions can be made. Due to the compounding effect of aggregating neighbor information, GNN inferences require raw data from many times more nodes than are targeted for prediction. Thus, on heterogeneous compute platforms, inference latency can be largely subject to the inter-device communication cost of transferring input feature data to the GPU/accelerator before computation has even begun. In this paper, we analyze the trade-off effect of pruning input features from GNN models, reducing the volume of raw data that the model works with to lower communication latency at the expense of an expected decrease in the overall model accuracy. We develop greedy and regression-based algorithms to determine which features to retain for optimal prediction accuracy. We evaluate pruned model variants and find that they can reduce inference latency by up to 80% with an accuracy loss of less than 5% compared to non-pruned models. Furthermore, we show that the latency reductions from input feature pruning can be extended under different system variables such as batch size and floating point precision. Jason Yik, Sanmukh R. Kuppannagari, Hanqing Zeng, Viktor Prasanna 0001 |
HIPC | 2 |
| 2022 | Estimating the Impact of Communication Schemes for Distributed Graph ProcessingabstractExtreme scale graph analytics is imperative for several real-world Big Data applications with the underlying graph structure containing millions or billions of vertices and edges. Since such huge graphs cannot fit into the memory of a single computer, distributed processing of the graph is required. Several frameworks have been developed for performing graph processing on distributed systems. The frameworks focus primarily on choosing the right computation model and the partitioning scheme under the assumption that such design choices will automatically reduce the communication overheads. For any computational model and partitioning scheme, communication schemes — the data to be communicated and the virtual interconnection network among the nodes — have significant impact on the performance. To analyze this impact, in this work, we identify widely used communication schemes and estimate their performance. Analyzing the trade-offs between the number of compute nodes and communication costs of various schemes on a distributed platform by brute force experimentation can be prohibitively expensive. Thus, our performance estimation models provide an economic way to perform the analyses given the partitions and the communication scheme as input. We validate our model on a local HPC cluster as well as the cloud hosted NSF Chameleon cluster. Using our estimates as well as the actual measurements, we compare the communication schemes and provide conditions under which one scheme should be preferred over the others. Tian Ye 0002, Sanmukh R. Kuppannagari, César A. F. De Rose, Sasindu Wijeratne, Rajgopal Kannan, Viktor Prasanna 0001 |
ISPDC | 2 |
| 2022 | PPOAccel: A High-Throughput Acceleration Framework for Proximal Policy OptimizationabstractReinforcement Learning (RL) is a major branch of AI that enables agents to learn optimal decision making via interaction with the environment. Proximal Policy Optimization (PPO) is the state-of-the-art policy optimization based RL algorithm which achieves superior overall performance on various benchmarks. A PPO agent iteratively optimizes its policy - a function which chooses optimal actions approximated by a DNN, with each iteration consisting of two computationally intensive phases: Sample Generation - where agents inference on its policy and interact with the environment to collect data, and Model Update - where the policy is trained using the collected data. In this paper, we develop the first high-throughput PPO accelerator on CPU-FPGA heterogeneous platform. Our unified systolic-array based design accelerates both the inference and the training of the deep neural network used in a RL algorithm, and is generalizable to various MLP and CNN models across a wide range of RL applications. We develop novel optimizations to simultaneously reduce data access and computation latencies, specifically: (a) optimal data flow mapping to systolic array, (b) novel memory-blocked data layout to enable streaming stall-free data access in both forward and backward propagations, and, (c) a systolic array compute sharing technique to mitigate load imbalance in the training of two networks. We evaluate our design on widely used robotics and gaming benchmarks, achieving 1.4×–26× and 1.3×–2.7× improvements in throughput, respectively, when compared with state-of-the-art CPU/CPU-GPU implementations. Yuan Meng 0001, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Safe Building HVAC Control via Batch Reinforcement LearningabstractIn this paper, we study safe building HVAC control via batch reinforcement learning. Random exploration in building HVAC control is infeasible due to safety considerations. However, diverse states are necessary for RL algorithms to learn useful policies. To enable safety during exploration, we propose guided exploration by adding a Gaussian noise to a hand-crafted rule-based controller. Adjusting the variance of the noise provides a tradeoff between the diversity of the dataset and the safety . We apply Conservative Q Learning (CQL) to learn a policy. CQL ensures that the trained policy stays within the policy distribution used to collect the dataset, thereby guarantees safety at deployment. To select the optimal policy during the offline training, we apply model-based performance evaluation. We use the widely adopted CityLearn testbed to evaluate the performance of our proposed method. Compared with a rule-based controller, our approach obtains $12\%\sim 35\%$ reduction in ramping, $3\%\sim 10\%$ reduction in 1-load factor, $3\%\sim 8\%$ reduction in daily peak at deployment with less than $10\%$ performance degradation during the exploration. On the contrary, the performance degradation of the state-of-the-art online reinforcement learning algorithm during exploration is around $8\%\sim 18\%$ . It also fails to surpass the performance of the rule-based controller at deployment. Chi Zhang 0022, Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2021 | BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement LearningabstractOnline interactions with the environment to collect data samples for training a Reinforcement Learning (RL) agent is not always feasible due to economic and safety concerns. The goal of Offline Reinforcement Learning is to address this problem by learning effective policies using previously collected datasets. Standard off-policy RL algorithms are prone to overestimations of the values of out-of-distribution (less explored) actions and are hence unsuitable for Offline RL. Behavior regularization, which constraints the learned policy within the support set of the dataset, has been proposed to tackle the limitations of standard off-policy algorithms. In this paper, we improve the behavior regularized offline reinforcement learning and propose BRAC+. First, we propose quantification of the out-of-distribution actions and conduct comparisons between using Kullback–Leibler divergence versus using Maximum Mean Discrepancy as the regularization protocol. We propose an analytical upper bound on the KL divergence as the behavior regularizer to reduce variance associated with sample based estimations. Second, we mathematically show that the learned Q values can diverge even using behavior regularized policy update under mild assumptions. This leads to large overestimations of the Q values and performance deterioration of the learned policy. To mitigate this issue, we add a gradient penalty term to the policy evaluation objective. By doing so, the Q values are guaranteed to converge. On challenging offline RL benchmarks, BRAC+ outperforms the baseline behavior regularized approaches by $40%\sim 87%$ and the state-of-the-art approach by $6%$. Chi Zhang 0022, Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
ACML | 2 |
| 2021 | DYNAMAP: Dynamic Algorithm Mapping Framework for Low Latency CNN InferenceabstractMost of the existing work on FPGA acceleration of Convolutional Neural Network (CNN) focuses on employing a single strategy (algorithm, dataflow, etc.) across all the layers. Such an approach does not achieve optimal latency on complex and deep CNNs. Emerging CNNs have diverse per-layer computation characteristics including parallelism, arithmetic intensity, locality, and memory footprint. Per-layer strategy selection and fine-grained tuning are required to achieve low end-to-end latency. However, specialized hardware modules dedicated to each layer limit the per-layer utilization and adversely affect end-to-end latency. In this paper, we address these problems by an algorithm-architecture co-optimization framework, DYNAMAP, consisting of (1) a unified hardware overlay that can be reused across layers, supporting dynamic mapping of all three families of popular convolution algorithms, and further allowing flexible dataflow switching to maximize hardware utilization for each layer; (2) a novel software Design Space Exploration (DSE) flow that customizes the hardware overlay and chooses optimal strategy mapping. We show that the algorithm mapping space increases exponentially with network depth, and while the optimal algorithm selection problem is NP-hard in general, by exploiting the series-parallel structure of CNN models, we demonstrate a polynomial-time solution for optimal algorithm mapping. DYNAMAP is optimized for any CNN, including those having diverse computation and memory requirements across the layers. We demonstrate DYNAMAP using two state-of-the-art CNNs - GoogleNet and Inception-V4. The generated accelerators achieve up to 2.8x and 1.4x speedups, respectively, wrt inference latency compared with the state-of-the-art FPGA implementations. Yuan Meng 0001, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FPGA | 2 |
| 2021 | FGYM: Toolkit for Benchmarking FPGA based Reinforcement Learning AlgorithmsabstractFPGA-based heterogeneous computing platforms are promising candidates to enable fast training of Reinforcement Learning (RL) agents. Typically, an RL agent for an environment is trained via interactions with a software that simulates the environment. While several toolkits exist to quickly deploy RL training on CPU or GPU, there lacks a similar toolkit for FPGAs. To ease the deployment process of RL using FPGAs, we demonstrate FGYM (FPGA-GYM) - a toolkit that generates an end-to-end interface between the simulation environments running on the CPU and agents running on the FPGA. FGYM supports a variety of environments and automatically generates the memory interface using PCIe. FGYM supports multiple levels of parallelism including vectorized agent-environment interactions and memory port aggregation. It also provides profiling results for users to identify the execution bottlenecks. Nathaniel Peura, Yuan Meng 0001, Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
FPL | 3 |
| 2021 | Performance Modeling and FPGA Acceleration of Homomorphic Encrypted ConvolutionabstractPrivacy of data is a critical concern when applying Machine Learning (ML) techniques to domains with sensitive data. Homomorphic Encryption (HE), by enabling computations on encrypted data, has emerged as a promising approach to perform inference on ML models such as Convolution Neural Network (CNN) in a privacy preserving manner. A significant portion of the total inference latency is in performing convolution over homomorphic encrypted data (HE-Convolution). For performing convolution over plaintext data, low latency accelerator designs have been proposed using algorithms such as im2col, frequency domain convolution, etc. However, developing accelerators for the HE versions of these algorithms is non-trivial. In this work, we develop a unified FPGA design that enables low latency execution of both im2col and frequency domain HE-Convolution. To enable selection of the efficient algorithm for each convolution layer of a CNN, we develop a performance model that takes the parameters about the encryption and convolution layer as input and outputs the computation and resource requirements of the two algorithms for that layer. We use the performance model to select convolution algorithm for each layer of ResNet-50 and obtain the first low latency batch-1 inference accelerator for CNN inference with HE-Convolution targeting FPGAs using HLS. We compare our design against prior techniques on CPUs and show that our accelerator achieves speedups in the range of $3.4\times\sim 6.7\times$ in latency. Tian Ye 0002, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
FPL | 2 |
| 2021 | Parallel Actors and Learners: A Framework for Generating Scalable RL ImplementationsabstractReinforcement Learning (RL) has achieved significant success in application domains such as robotics, games and health care. However, training RL agents is very time consuming. Current implementations exhibit poor performance due to challenges such as irregular memory accesses and thread-level synchronization overheads on CPU. In this work, we propose a framework for generating scalable reinforcement learning implementations on multi-core systems. Replay Buffer is a key component of RL algorithms which facilitates storage of samples obtained from environmental interactions and data sampling for the learning process. We define a new data structure for Prioritized Replay Buffer based on$K$-ary sum tree that supports asynchronous parallel insertions, sampling, and priority updates. To address the challenge of irregular memory accesses, we propose a novel data layout to store the nodes of the sum tree that reduces the number of cache misses. Additionally, we propose lazy writing mechanism to reduce thread-level synchronization over-heads of the Replay Buffer operations. Our framework employs parallel actors to concurrently collect data via environmental interactions, and parallel learners to perform stochastic gradient descent using the collected data. Our framework supports a wide range of reinforcement learning algorithms including DQN, DDPG, etc. We demonstrate the effectiveness of our framework in accelerating RL algorithms by performing experiments on CPU + GPU platform using OpenAI benchmarks. Our results show that the performance of our$K$-ary sum tree based Prioritized Replay Buffer improves the baseline implementations by around 4x~100x. Our proposed synchronization optimizations improve the performance by around 2x~4.4x compared with using a global lock. By plugging our Replay Buffer implementation into existing open source reinforcement learning frameworks, we achieve 1.19x~ 1.75x speedup for various algorithms. Chi Zhang 0022, Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
HiPC | 2 |
| 2021 | How to Avoid Zero-Spacing in Fractionally-Strided Convolution? A Hardware-Algorithm Co-Design MethodologyabstractFractionally Strided Convolution (FSC) is a key operation in popular image-based Deep Learning models, for example, back propagation in CNN training, the decoding stage of convolutional auto-encoders and generative CNNs (GAN), etc. FSC typically performs up-convolution on a 2-D grid image, i.e., expands it to a larger one, as compared to conventional (down)-convolution, resulting in more complex computation patterns. Specifically, it introduces additional interleaved zero-spacing (i.e. insertion and padding of zeros) in feature maps that impose excessive computation and memory access overheads on traditional convolution methods such as im2col. The resulting hardware under-utilization is especially severe in layers with large kernels and large strides, commonly seen in typical CNNs and Generative CNNs. In this paper, we propose a methodology to address this challenge using a multi-channel-multi-kernel parallel algorithm, kn2row, to eliminate zero-computations in FSC. We further develop a unified accelerator for kn2row-based convolution and FSC operations in High-Level Synthesis (HLS). Benefiting from the compute-reduction of kn2row, we achieve up to 14.6x improvement in effective resource utilization in typical convolutional auto-decoding layers, GAN layers and backward pass of Nature-CNN, a reinforcement learning bench-marking model. These lead to overall speedup of up to 3.8x in the complete forward or backward propagation phases of the above benchmarks. Our methodology leads up to 8x speedup and 11x better power efficiency than general-purpose processors. Compared with existing GAN accelerators, our methodology achieves higher normalized throughput with high portability. Yuan Meng 0001, Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
HiPC | 2 |
| 2020 | Crowdsourced Edge: A Novel Networking Paradigm for the Collaborative CommunityabstractEdge computing established paradigms are prone to implicate solely powerful server-like edge nodes, in static or semi-static topologies, of centrally-controlled edge networks. In this paper, leveraging upon recent technological advancements and trends, we introduce a novel networking paradigm employing resources provided by independent crowd peers, within a zone of local proximity, to establish collaborative networks for edge computing. We call this paradigm the Crowdsourced Edge. We detail the architecture and characteristics of this novel paradigm, highlighting its unique characteristics and specific challenges, while also positioning it vis-a-vis the existing edge computing concretisations. Finally, we demonstrate the Crowdsourced Edge functionality by presenting an ongoing use case regarding a video-enhanced object search. Stéphane Kuendig, Constantinos Marios Angelopoulos, Sanmukh R. Kuppannagari, José D. P. Rolim, Viktor Prasanna 0001 |
DCOSS | 3 |
| 2020 | Accelerating Proximal Policy Optimization on CPU-FPGA Heterogeneous PlatformsabstractReinforcement Learning (RL) is a technique that enables an agent to learn to behave optimally by repeatedly interacting with the environment and receiving the rewards. RL is widely used in domains such as robotics, game playing and finance. Proximal Policy Optimization (PPO) is the state-of-the-art policy optimization algorithm which achieves superior overall performance on various RL benchmarks. PPO iteratively optimizes its policy - a function which chooses optimal actions, with each iteration consisting of two computationally intensive phases: Inference phase - agents infer actions to interact with the environment and collect data, and Training phase - agents train the policy using the collected data. In this work, we develop the first high-throughput PPO accelerator on CPU-FPGA heterogeneous platform, targeting both phases of the algorithm for acceleration. We implement a systolic-array based architecture coupled with a novel memory-blocked data layout that enables streaming data access in both forward and backward propagations to achieve high-throughput performance. Additionally, we develop a novel systolic array compute sharing technique to mitigate the potential load imbalance in the training of two networks. We develop an accurate performance model of our design, based on which we perform design space exploration to obtain optimal design points. Our design is evaluated on widely used robotics benchmarks, achieving $2.1 \times - 30.5 \times$ and $2 \times - 27.5 \times$ improvements in throughput against state-of-the-art CPU and CPU-GPU implementations, respectively. Yuan Meng 0001, Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
FCCM | 2 |
| 2020 | QTAccel: A Generic FPGA based Design for Q-Table based Reinforcement Learning AcceleratorsabstractQ-Table based Reinforcement Learning (QRL) is a class of widely used algorithms in AI that work by successively improving the estimates of Q values -- quality of state-action pairs, stored in a table. They significantly outperform Neural Network based techniques when the state space is tractable. Fast learning for AI applications in several domains (e.g. robotics), with tractable 'mid-sized' Q-tables, still necessitates performing substantial rapid updates. State-of-the-art FPGA implementations of QRL do not scale with the increasing Q-Table state space, thus are not efficient for such applications. In this work, we develop a novel FPGA implementation of QRL, scalable to large state spaces and facilitating a large class of AI applications. Our pipelined architecture provides higher throughput while using significantly fewer on-chip resources and thereby supports a variety of action selection policies that covers Q-Learning and variations of bandit algorithms. Possible dependencies caused by consecutive Q value updates are handled, allowing the design to process one Q-sample every clock cycle. Additionally, we provide the first known FPGA implementation of the SARSA (State-Action-Reward-State-Action) algorithm. We evaluate our architecture for Q-Learning and SARSA algorithms and show that our designs achieve a high throughput of up to 180 million Q samples per second. Rachit Rajat, Yuan Meng 0001, Sanmukh R. Kuppannagari, Ajitesh Srivastava, Viktor Prasanna 0001, Rajgopal Kannan |
FPGA | 3 |
| 2018 | Optimal Discrete Net-Load Balancing in Smart Grids with High PV PenetrationabstractMitigating supply-demand mismatch is critical for smooth power grid operation. Traditionally, load curtailment techniques such as demand response have been used for this purpose. However, these cannot be the only component of a net-load balancing framework for smart grids with high PV penetration. These grids sometimes exhibit supply surplus, causing overvoltages. Currently, these are mitigated using voltage manipulation techniques such as Volt-Var Optimizations, which are computationally expensive, thereby increasing the complexity of grid operations. Taking advantage of recent technological developments that enable rapid selective connection of PV modules of an installation to the grid, we develop a unified net-load balancing framework that performs both load and solar curtailment. We show that when the available curtailment values are discrete, this problem is NP-hard and we develop bounded approximation algorithms. Our algorithms produce fast solutions, given the tight timing constraints required for grid operation, while ensuring that practical constraints such as fairness, network capacity limits, and so forth are satisfied. We also develop an online algorithm that performs net-load balancing using only data available for the current interval. Using both theoretical analysis and practical evaluations, we show that our net-load balancing algorithms provide solutions that are close to optimal in a small amount of time. Sanmukh R. Kuppannagari, Rajgopal Kannan, Viktor Prasanna 0001 |
ACM Trans. Sens. Networks | 1 |
| 2016 | Implementation of Learning-Based Dynamic Demand Response on a Campus Micro-Grid
Sanmukh R. Kuppannagari, Rajgopal Kannan, Charalampos Chelmis, Viktor Prasanna 0001 |
IJCAI | 1 |
| 2015 | Efficient Generation of Energy and Performance Pareto Front for FPGA Designs (Abstract Only)abstractAnalysis of trade-offs between energy efficiency and latency is essential to generate designs complying with a given set of constraints. Improvements in FPGA technologies offer a myriad choices for power and performance optimizations. Various algorithm intrinsic parameters also affect these objectives. The design space is compounded by the available choices. This requires efficient techniques to quickly explore the design space. Current techniques perform Gate/RTL level or functional level power modeling which are slow and hence not scalable. In this work we perform efficient design space exploration using a high level performance model. We develop a semi-automatic design framework to generate energy efficiency and latency trade-offs. The framework develops a performance model given a high level specification of a design with minimal user assistance. It then explores the entire design space to generate the dominating designs with respect to energy efficiency and latency metrics. We illustrate the framework using convolutional neural network which gained significance due to its application in deep learning. We simulate a few designs from the dominating set and show that the performance estimation for the dominating designs are close to the simulated results. We also show that our framework explores 6000 design points per minute on a commodity platform such as Dell workstation as opposed to state-of-the-art techniques which explore at 50 to 60 design points per minute. Sanmukh R. Kuppannagari, Viktor Prasanna 0001 |
FPGA | 1 |