Kailash Gogineni

dblp:261/7627 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0003-1865-5470ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RACER: A Caching Layer based Optimization for Accelerating Reinforcement Learning Workloads
abstract
Reinforcement Learning (RL) learns optimal decision-making policies from experiential (transition) datasets and maximizes the RL agent’s cumulative rewards. In order to improve the RL training efficiency, prior works have studied how to selectively sample certain critical transitions that ultimately lead to better policies and rewards. However, RL workloads still face significant challenges from a systems perspective, particularly when the agent iteratively accesses batches of data from transition datasets, whose growing sizes continue to challenge the memory hierarchy. This results in frequent and costly memory transfers between caches and Dynamic Random Access Memory (DRAM), which negatively impacts the overall training time.In this paper, we propose RACER, our novel caching layerbased optimization for RL training workloads. Firstly, recognizing that the RL agent repeatedly accesses large transition batches from growing datasets, we design a storage-cache that prioritizes critical transitions to fit within the hardware cache hierarchies. This design reduces the memory access times by sampling from a subset of critical transitions and minimizes the costly memory trips to DRAM. Second, we demonstrate how to smartly leverage key metrics (viz., temporal difference error and advantage weighting) to quantify the importance/relevance of transitions during policy optimization and to identify the transition data that would need to be spilled out of (or filled into) the caching system. We also introduce dynamic optimizations to our caching system that minimize the prospect of discarding the critical transitions. Our performance evaluation across three state-of-the-art RL algorithms, under various task environments, and on three different systems, demonstrates that RACER achieves significant optimization time improvements to the RL transition data sampling phase (a speedup of $6 \times$) and end-to-end training time (up to $2 \times$) with comparable rewards.
Kailash Gogineni, Yongsheng Mei, Karthikeya Gogineni, Tian Lan 0001, Guru Venkataramani
ISPASS1
2024 SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems
abstract
Reinforcement Learning (RL) is the process by which an agent learns optimal behavior through interactions with experience datasets, all of which aim to maximize the reward signal. RL algorithms often face performance challenges in real-world applications, especially when training with extensive and diverse datasets. For instance, applications like autonomous vehicles include sensory data, dy-namic traffic information (including movements of other vehicles and pedestrians), critical risk assessments, and varied agent actions. Consequently, RL training is significantly memory-bound due to sampling large experience datasets that may not fit entirely into the hardware caches and frequent data transfers needed between memory and the computation units (e.g., CPU, GPU), especially during batch updates. This bottleneck results in significant execution latencies and impacts the overall training time. To alleviate such is-sues, recently proposed memory-centric computing paradigms, like Processing-In-Memory (PIM), can address memory latency-related bottlenecks by performing the computations inside the memory devices. In this paper, we present SwiftRL, which explores the potential of real-world PIM architectures to accelerate popular RL workloads and their training phases. We adapt RL algorithms, namely Tab-ular Q-learning and SARSA, on UPMEM PIM systems and first observe their performance using two different environments and three sampling strategies. We then implement performance opti-mization strategies during RL adaptation to PIM by approximating the Q-value update function (which avoids high performance costs due to runtime instruction emulation used by runtime libraries) and incorporating certain PIM-specific routines specifically needed by the underlying algorithms. Moreover, we develop and assess a multi-agent version of Q-learning optimized for hardware and illustrate how PIM can be leveraged for algorithmic scaling with multiple agents. We experimentally evaluate RL workloads on OpenAI GYM environments using UPMEM hardware. Our results demonstrate a near-linear scaling of 15x in performance when the number of PIM cores increases by 16x (125 to 2000). We also compare our PIM implementation against Intel(R) Xeon(R) Silver 4110 CPU and NVIDIA RTX 3090 GPU and observe superior performance on the UPMEM PIM System for different implementations.
Kailash Gogineni, Sai Santosh Dayapule, Juan Gómez-Luna, Karthikeya Gogineni, Tian Lan 0001, Mohammad Sadrosadati, Onur Mutlu, Guru Venkataramani
ISPASS1
2023 Mayalok: A Cyber-Deception Hardware Using Runtime Instruction Infusion
abstract
Rapid rise in malware attacks has added significant costs to cyber operations. As adversaries evolve, there is a growing need for fast, targeted defenses that effectively guard computer systems against these cyber-attacks. Cyber-deception is an increasingly adopted defense strategy with its ability to continually engage with adversaries and deploy counter-measures proactively by manipulating the malware program execution flow to non-useful states for the attacker. This paper introduces Mayalok, a novel hardware-based cyber-deception framework to combat malware through runtime instruction infusion. Mayalok employs hardware deception primitives to transparently insert or skip malware program instructions during runtime and deliver the attackers a deceptive view of the system state. We evaluate and demonstrate the deception efficacy of the Mayalok framework on malware samples representing various attack vectors: Ransomware, InfoStealers, Buffer overflow, and Side-channels.
Preet Derasari, Kailash Gogineni, Guru Venkataramani
ASAP2
2023 AccMER: Accelerating Multi-Agent Experience Replay with Cache Locality-Aware Prioritization
abstract
Multi-Agent Experience Replay (MER) is a key component of off-policy reinforcement learning (RL) algorithms. By remembering and reusing experiences from the past, experience replay significantly improves the stability of RL algorithms and their learning efficiency. In many scenarios, multiple agents interact in a shared environment during online training under centralized training and decentralized execution (CTDE) paradigm. Current multi-agent reinforcement learning (MARL) algorithms consider experience replay with uniform sampling or based on priority weights to improve transition data sample efficiency in the sampling phase. However, moving transition data histories for each agent through the processor memory hierarchy is a performance limiter. Also, as the agents' transitions continuously renew every iteration, the finite cache capacity results in increased cache misses. To this end, we propose AccMER, that repeatedly reuses the transitions (experiences) for a window of$n$steps in order to improve the cache locality and minimize the transition data movement, instead of sampling new transitions at each step. Specifically, our optimization uses priority weights to select the transitions so that only high-priority transitions will be reused frequently, thereby improving the cache performance. Our experimental results on the Predator- Prey environment demonstrate the effectiveness of reusing the essential transitions based on the priority weights, where we observe an end-to-end training time reduction of 25.4% (for 32 agents) compared to existing prioritized MER algorithms without notable degradation in the mean reward.
Kailash Gogineni, Yongsheng Mei, Tian Lan 0001, Guru Venkataramani
ASAP1
2023 MAYAVI: A Cyber-Deception Hardware for Memory Load-Stores
abstract
Rapid evolution of security attacks presents a perpetual challenge to computer system defenders in terms of continuously upgrading their defense capabilities and being aware of adversarial tactics. Emerging technologies like cyber-deception offer the unique advantage of intelligently surveying hostile behavior while actively safeguarding sensitive assets by manipulating the malware execution flow to non-useful states or misrepresenting critical data.
Preet Derasari, Kailash Gogineni, Guru Venkataramani
ACM Great Lakes Symposium on VLSI2
2021 MPD: Moving Target Defense Through Communication Protocol Dialects
Yongsheng Mei, Kailash Gogineni, Tian Lan 0001, Guru Venkataramani
SecureComm (1)2