VLDB 2026 Research / reviewers in the wild / expert
Bhargava Gopireddy
dblp:178/3211
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorSecurity and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 38% Integrated circuit design · 24% Energy-efficient computing · 17% | |
| Network and information security
3 papers |
Hardware security and side channels · 100% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware security and side channels › side-channel attack
cache side-channel attacks |
0.7 | 2 | 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive World · IEEE Symposium on Security and Privacy 2019 Secure Hierarchy-Aware Cache Replacement Policy (SHARP): Defending Against Cache-Based Side Channel Attacks · ISCA 2017 |
Hardware security and side channels
side-channel attack |
0.5 | 2 | 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive World · IEEE Symposium on Security and Privacy 2019 MicroScope: enabling microarchitectural replay attacks · ISCA 2019 |
Memory systems › memory management › virtual memory
address translation |
0.4 | 1 | 2020 | BabelFish: Fusing Address Translations for Containers · ISCA 2020 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.4 | 1 | 2020 | BabelFish: Fusing Address Translations for Containers · ISCA 2020 |
Memory systems › memory management
virtual memory |
0.4 | 1 | 2020 | BabelFish: Fusing Address Translations for Containers · ISCA 2020 |
Hardware security and side channels
microarchitectural side channel |
0.4 | 1 | 2019 | MicroScope: enabling microarchitectural replay attacks · ISCA 2019 |
Hardware security and side channels › side-channel attack › cache side-channel attacks
prime+probe |
0.4 | 1 | 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive World · IEEE Symposium on Security and Privacy 2019 |
Integrated circuit design
3d integration |
0.4 | 1 | 2019 | Designing vertical processors in monolithic 3D · ISCA 2019 |
Integrated circuit design › 3d integration
monolithic 3d integration |
0.4 | 1 | 2019 | Designing vertical processors in monolithic 3D · ISCA 2019 |
Energy-efficient computing › low-power design
low-power processor design |
0.3 | 1 | 2018 | HetCore: TFET-CMOS Hetero-Device Architecture for CPUs and GPUs · ISCA 2018 |
Integrated circuit design › emerging device technologies
tunneling field effect transistor |
0.3 | 1 | 2018 | HetCore: TFET-CMOS Hetero-Device Architecture for CPUs and GPUs · ISCA 2018 |
Hardware security and side channels
cache replacement policy |
0.3 | 1 | 2017 | Secure Hierarchy-Aware Cache Replacement Policy (SHARP): Defending Against Cache-Based Side Channel Attacks · ISCA 2017 |
Processor architecture and microarchitecture › pipelining
pipeline design |
0.2 | 1 | 2016 | ScalCore: Designing a core for voltage scalability · HPCA 2016 |
Energy-efficient computing
voltage scaling |
0.2 | 1 | 2016 | ScalCore: Designing a core for voltage scalability · HPCA 2016 |
Cloud and datacenter computing › virtualization
container |
0.1 | 1 | 2020 | BabelFish: Fusing Address Translations for Containers · ISCA 2020 |
Cloud and datacenter computing
serverless computing |
0.1 | 1 | 2020 | BabelFish: Fusing Address Translations for Containers · ISCA 2020 |
Hardware security and side channels › microarchitectural attacks
microarchitectural replay attack |
0.1 | 1 | 2019 | MicroScope: enabling microarchitectural replay attacks · ISCA 2019 |
Hardware security and side channels › trusted execution environments
secure enclaves |
0.1 | 1 | 2019 | MicroScope: enabling microarchitectural replay attacks · ISCA 2019 |
Hardware security and side channels
trusted execution environments |
0.1 | 1 | 2019 | MicroScope: enabling microarchitectural replay attacks · ISCA 2019 |
Memory systems › memory hierarchy
cache hierarchy |
0.1 | 1 | 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive World · IEEE Symposium on Security and Privacy 2019 |
Memory systems › cache design
non-inclusive cache |
0.1 | 1 | 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive World · IEEE Symposium on Security and Privacy 2019 |
GPUs and heterogeneous computing
GPU architecture |
0.1 | 1 | 2018 | HetCore: TFET-CMOS Hetero-Device Architecture for CPUs and GPUs · ISCA 2018 |
Memory systems
cache coherence |
0.1 | 1 | 2017 | Secure Hierarchy-Aware Cache Replacement Policy (SHARP): Defending Against Cache-Based Side Channel Attacks · ISCA 2017 |
Memory systems › cache management
cache replacement |
0.1 | 1 | 2017 | Secure Hierarchy-Aware Cache Replacement Policy (SHARP): Defending Against Cache-Based Side Channel Attacks · ISCA 2017 |
Energy-efficient computing › power management
dynamic voltage and frequency scaling |
0.1 | 1 | 2016 | ScalCore: Designing a core for voltage scalability · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.0reverse engineering · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | BabelFish: Fusing Address Translations for ContainersabstractCloud computing has begun a transformation from using virtual machines to containers. Containers are attractive because multiple of them can share a single kernel, and add minimal performance overhead. Cloud providers leverage the lean nature of containers to run hundreds of them on a few cores. Furthermore, containers enable the serverless paradigm, which leads to the creation of short-lived processes.In this work, we identify that containerized environments create page translations that are extensively replicated across containers in the TLB and in page tables. The result is high TLB pressure and redundant kernel work during page table management. To remedy this situation, this paper proposes BabelFish, a novel architecture to share page translations across containers in the TLB and in page tables. We evaluate BabelFish with simulations of an 8-core processor running a set of Docker containers in an environment with conservative container co-location. On average, under BabelFish, 53% of the translations in containerized workloads and 93% of the translations in serverless workloads are shared. As a result, BabelFish reduces the mean and tail latency of containerized data-serving workloads by 11% and 18%, respectively. It also lowers the execution time of containerized compute workloads by 11%. Finally, it reduces serverless function bring-up time by 8% and execution time by 10%-55%. Dimitrios Skarlatos 0002, Umur Darbaz, Bhargava Gopireddy, Nam Sung Kim, Josep Torrellas |
ISCA | 3 |
| 2019 | Designing vertical processors in monolithic 3DabstractA processor laid out vertically in stacked layers can benefit from reduced wire delays, low energy consumption, and a small footprint. Such a design can be enabled by Monolithic 3D (M3D), a technology that provides short wire lengths, good thermal properties, and high integration. In current M3D technology, due to manufacturing constraints, the layers in the stack are asymmetric: the bottom-most one has a relatively higher performance. Bhargava Gopireddy, Josep Torrellas |
ISCA | 1 |
| 2019 | MicroScope: enabling microarchitectural replay attacksabstractThe popularity of hardware-based Trusted Execution Environments (TEEs) has recently skyrocketed with the introduction of Intel's Software Guard Extensions (SGX). In SGX, the user process is protected from supervisor software, such as the operating system, through an isolated execution environment called an enclave. Despite the isolation guarantees provided by TEEs, numerous microarchitectural side channel attacks have been demonstrated that bypass their defense mechanisms. But, not all hope is lost for defenders: many modern fine-grain, high-resolution side channels---e.g., execution unit port contention---introduce large amounts of noise, complicating the adversary's task to reliably extract secrets. Dimitrios Skarlatos 0002, Mengjia Yan 0001, Bhargava Gopireddy, Read Sprabery, Josep Torrellas, Christopher W. Fletcher |
ISCA | 3 |
| 2019 | Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive WorldabstractAlthough clouds have strong virtual memory isolation guarantees, cache attacks stemming from shared caches have proved to be a large security problem. However, despite the past effectiveness of cache attacks, their viability has recently been called into question on modern systems, due to trends in cache hierarchy design moving away from inclusive cache hierarchies. In this paper, we reverse engineer the structure of the directory in a sliced, non-inclusive cache hierarchy, and prove that the directory can be used to bootstrap conflict-based cache attacks on the last-level cache. We design the first cross-core Prime+Probe attack on non-inclusive caches. This attack works with minimal assumptions: the adversary does not need to share any virtual memory with the victim, nor run on the same processor core. We also show the first high-bandwidth Evict+Reload attack on the same hardware. We demonstrate both attacks by extracting key bits during RSA operations in GnuPG on a state-of-the-art non-inclusive Intel Skylake-X server. Mengjia Yan 0001, Read Sprabery, Bhargava Gopireddy, Christopher W. Fletcher, Roy H. Campbell, Josep Torrellas |
IEEE Symposium on Security and Privacy | 3 |
| 2018 | HetCore: TFET-CMOS Hetero-Device Architecture for CPUs and GPUsabstractTunneling Field-Effect Transistors (TFETs) attain much higher energy efficiency than CMOS at low voltages. However, their performance saturates at high voltages and, therefore, cannot replace CMOS when high performance is needed. Ideally, we desire a core that is as energy-efficient as a TFET core and provides as much performance as a CMOS core. To approach this goal, this paper judiciously integrates both TFET units and CMOS units in a single core, effectively creating a hetero-device core. We call it HetCore, and present CPU and GPU versions. In HetCore, TFETs are used in units that consume high power under CMOS, are amenable to pipelining or are not very latency sensitive, and use a sizable area. HetCore powers CMOS and TFET units at different voltage levels, so they operate optimally. However, all units are clocked at the same frequency. Our results based on simulations running standard applications show the potential of this approach, even with conservative assumptions. A HetCore CPU consumes on average 39% less energy than a CMOS CPU, while delivering an average performance that is within 10% of the CMOS CPU. In addition, under a fixed power budget, a multicore with HetCore CPUs can employ twice as many cores as a multicore with CMOS CPUs, resulting in average performance gains of 32% while, at the same time, improving the energy efficiency (ED2) by an average of 68%. Similar results are obtained with HetCore GPUs. Bhargava Gopireddy, Dimitrios Skarlatos 0002, Wenjuan Zhu 0002, Josep Torrellas |
ISCA | 1 |
| 2017 | Sthira: A Formal Approach to Minimize Voltage Guardbands under Variation in Networks-on-Chip for Energy EfficiencyabstractNetworks-on-Chip (NoCs) in chip multiprocessors are prone to within-die process variation as they span the whole chip. To tolerate variation, their voltages (Vdd) carry over-provisioned guardbands. As a result, prior work has proposed to save energy by operating at reduced Vddwhile occasionally suffering and fixing errors. Unfortunately, these proposals use heuristic controller designs that provide no error bounds guarantees. In this work, we develop a scheme that dynamically minimizes the Vddof groups of routers in a variation-prone NoC using formal control-theoretic methods. The scheme, called Sthira, saves substantial energy while guaranteeing the stability and convergence of error rates. We also enhance the scheme with a low-cost secondary network that retransmits erroneous packets for higher energy efficiency. The enhanced scheme is called Sthira+. We evaluate Sthira and Sthira+ with simulations of NoCs with 64-100 routers. In an NoC with 8 routers per Vdddomain, our schemes reduce the average energy consumptionof the NoC by 27%; in a futuristic NoC with one router per Vdd domain, Sthira+ and Sthira reduce the average energy consumption by 36% and 32%, respectively. The performance impact is negligible. These are significant savings over the state-of-the-art. We conclude that formal control is essential, and that the cheaper Sthira is more cost-effective than Sthira+. Raghavendra Pradyumna Pothukuchi, Amin Ansari, Bhargava Gopireddy, Josep Torrellas |
PACT | 3 |
| 2017 | Secure Hierarchy-Aware Cache Replacement Policy (SHARP): Defending Against Cache-Based Side Channel AttacksabstractIn cache-based side channel attacks, a spy that shares a cache with a victim probes cache locations to extract information on the victim's access patterns. For example, in evict+reload, the spy repeatedly evicts and then reloads a probe address, checking if the victim has accessed the address in between the two operations. While there are many proposals to combat these cache attacks, they all have limitations: they either hurt performance, require programmer intervention, or can only defend against some types of attacks. Mengjia Yan 0001, Bhargava Gopireddy, Thomas Shull, Josep Torrellas |
ISCA | 2 |
| 2016 | ScalCore: Designing a core for voltage scalabilityabstractUpcoming multicores need to provide increasingly stringent energy-efficient execution modes. Currently, energy efficiency is attained by lowering the voltage (Vdd) through DVFS. However, the effectiveness of DVFS is limited: designing cores for low Vddresults in energy inefficiency at nominal Vdd. Our goal is to design a core for Voltage Scalability, i.e., one that can work in high-performance mode (HPMode) at nominal Vdd, and in a very energy-efficient mode (EEMode) at low Vdd. We call this core ScalCore. To operate energy-efficiently in EEMode, ScalCore introduces two ideas. First, since logic and storage structures scale differently with Vdd, ScalCore applies two low Vdds to the pipeline: one to the logic stages (Vlogic) and a higher one to storage-intensive stages. Secondly, ScalCore further increases the low Vddof the storage-intensive stages (Vop), so that they are substantially faster than the logic ones. Then, it exploits the speed differential by either fusing storage-intensive pipeline stages or increasing the size of storage structures in the pipeline. Our simulations of 16 cores show that a design with ScalCores in EEMode is much more energy-efficient than one with conventional cores and aggressive DVFS: for approximately the same power, ScalCores reduce the average execution time of programs by 31%, the energy (E) consumed by 48%, and the ED product by 60%. In addition, dynamically switching between EEMode and HPMode based on program phases is very effective: it reduces the average execution time and ED product by a further 28% and 15%, respectively. Bhargava Gopireddy, Choungki Song, Josep Torrellas, Nam Sung Kim, Aditya Agrawal, Asit K. Mishra |
HPCA | 1 |