Shruti Padmanabha

dblp:127/9035 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Processor architecture and microarchitecture · 52% Cloud and datacenter computing · 18% Energy-efficient computing · 10%
Computer networks
1 paper
Network performance modeling · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › multicore design
heterogeneous multicore
1.152017
Mirage cores: the illusion of many out-of-order cores using in-order hardware · MICRO 2017
Exploring Fine-Grained Heterogeneity with Composite Cores · IEEE Trans. Computers 2016
DynaMOS: dynamic schedule migration for heterogeneous cores · MICRO 2015
Processor architecture and microarchitecture
out-of-order execution
0.522017
Mirage cores: the illusion of many out-of-order cores using in-order hardware · MICRO 2017
DynaMOS: dynamic schedule migration for heterogeneous cores · MICRO 2015
Cloud and datacenter computing › resource management
datacenter resource management
0.412019
Taiji: managing global user traffic for large-scale internet services at the edge · SOSP 2019
Parallel and multicore computing
load balancing
0.412019
Taiji: managing global user traffic for large-scale internet services at the edge · SOSP 2019
Cloud and datacenter computing › datacenter operations
datacenter reliability
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Distributed systems
fault tolerance
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore
0.212016
Exploring Fine-Grained Heterogeneity with Composite Cores · IEEE Trans. Computers 2016
Energy-efficient computing › energy-efficient architecture
energy-efficient microarchitecture
0.212016
Exploring Fine-Grained Heterogeneity with Composite Cores · IEEE Trans. Computers 2016
Energy-efficient computing
power management
0.112012
Composite Cores: Pushing Heterogeneity Into a Core · MICRO 2012
Performance modeling and evaluation › simulation › processor simulation
microarchitecture simulation
0.112016
Exploring Fine-Grained Heterogeneity with Composite Cores · IEEE Trans. Computers 2016
Processor architecture and microarchitecture › multicore design
big.LITTLE
0.012012
Composite Cores: Pushing Heterogeneity Into a Core · MICRO 2012

Methods — techniques the papers use, named apart from their topics

traffic management · 0.8simulation · 0.7power modeling · 0.4cycle-accurate simulation · 0.4runtime scheduling · 0.3
YearPublicationVenuePosition
2019 Taiji: managing global user traffic for large-scale internet services at the edge
abstract
We present Taiji, a new system for managing user traffic for large-scale Internet services that accomplishes two goals: 1) balancing the utilization of data centers and 2) minimizing network latency of user requests.
Tianyin Xu, Kaushik Veeraraghavan, Andrew Newell, Sonia Margulis, Pol Mauri Ruiz, Justin Meza, Kiryong Ha, Shruti Padmanabha, Kevin Cole, Dmitri Perelman
SOSP10
2018 Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently
Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankaralingam Panneerselvam, Alex Gyori, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, Tianyin Xu
OSDI9
2017 Mirage cores: the illusion of many out-of-order cores using in-order hardware
abstract
Heterogenous chip multiprocessors (Het-CMPs) offer a combination of large Out-of-Order (OoO) cores optimized for high single-threaded performance and small In-Order (InO) cores optimized for low-energy and area costs. Due to practical constraints, CMP designers must choose to either optimize for total system throughput by utilizing many InO cores or maximize single-thread execution with fewer OoO cores. We propose Mirage Cores, a novel Het-CMP design where clusters of InO cores are architected around an OoO in a manner that optimizes for both throughput and single-thread performance. The insight behind Mirage Cores is that InO cores can achieve near-OoO performance if they are provided with the dynamic instruction schedule of an OoO core. To leverage this, Mirage Cores employs an OoO core as an optimal instruction schedule generator as well as a high-performance alternative for all neighboring InO cores. We also develop intelligent runtime schedulers which orchestrate the arbitration and migration of applications between the InO cores and the central OoO. Fast and timely transfer of dynamic schedules from the OoO to InO allows Mirage Cores to create the appearance of all OoO cores to the user using underlying In-Order hardware.
Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke
MICRO1
2016 Exploring Fine-Grained Heterogeneity with Composite Cores
abstract
Heterogeneous multicore systems- comprising multiple cores with varying performance and energy characteristics-have emerged as a promising approach to increasing energy efficiency. Such systems reduce energy consumption by identifying application phases and migrating execution to the most efficient core that meets performance requirements. However, the overheads of migrating between cores limit opportunities to coarse-grained phases (hundreds of millions of instructions), reducing the potential to exploit energy efficient cores. We propose Composite Cores, an architecture that reduces migration overheads by bringing heterogeneity into a core. Composite Cores pairs a big and little compute μEngine that together achieve high performance and energy efficiency. By sharing architectural state between the μEngines, the migration overhead is reduced, enabling fine-grained migration and increasing the opportunities to utilize the little μEngine without sacrificing performance. An intelligent controller migrates the application between μEngines to maximize energy efficiency while constraining performance loss to a configurable bound. We evaluate Composite Cores using cycle accurate microarchitectural simulations and a detailed power model. Results show that, on average, Composite Cores are able to map 30 percent of the execution time to the little μEngine, achieving a 21 percent energy savings while maintaining 95 percent performance.
Andrew Lukefahr, Shruti Padmanabha, Reetuparna Das, Faissal M. Sleiman, Ronald G. Dreslinski, Thomas F. Wenisch, Scott A. Mahlke
IEEE Trans. Computers2
2015 DynaMOS: dynamic schedule migration for heterogeneous cores
abstract
InOrder (InO) cores achieve limited performance because their inability to dynamically reorder instructions prevents them from exploiting Instruction-Level-Parallelism. Conversely, Out-of-Order (OoO) cores achieve high performance by aggressively speculating past stalled instructions and creating highly optimized issue schedules. It has been observed that these issue schedules tend to repeat for sequences of instructions with predictable control and data-flow. An equally provisioned InO core can potentially achieve OoO's performance at a fraction of the energy cost if provided with an OoO schedule. In the context of a fine-grained heterogeneous multicore system composed of a big (OoO) core and a little (InO) core, we could offload recurring issue schedules from the big to the little core, to achieve energy-efficiency while maintaining performance.
Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke
MICRO1
2014 Heterogeneous microarchitectures trump voltage scaling for low-power cores
abstract
Heterogeneous architectures offer many potential avenues for improving energy efficiency in today's low-power cores. Two common approaches are dynamic voltage/frequency scaling (DVFS) and heterogeneous microarchitectures (HMs). Traditionally both approaches have incurred large switching overheads, which limit their applicability to coarse-grain program phases. However, recent research has demonstrated low-overhead mechanisms that enable switching at granularities as low as 1K instructions. The question remains, in this fine-grained switching regime, which form of heterogeneity offers better energy efficiency for a given level of performance?
Andrew Lukefahr, Shruti Padmanabha, Reetuparna Das, Ronald G. Dreslinski, Thomas F. Wenisch, Scott A. Mahlke
PACT2
2013 Trace based phase prediction for tightly-coupled heterogeneous cores
abstract
Heterogeneous multicore systems are composed of multiple cores with varying energy and performance characteristics. A controller dynamically detects phase changes in applications and migrates execution onto the most efficient core that meets the performance requirements. In this paper, we show that existing techniques that react to performance changes break down at fine-grain intervals, as performance variations between consecutive intervals are high. We propose a predictive trace-based switching controller that predicts an upcoming phase change in a program and preemptively migrates execution onto a more suitable core. This prediction is based on a phase's individual history and the current program context. Our implementation detects repeatable code sequences to build history, uses these histories to predict an phase change, and preemptively migrates execution to the most appropriate core. We compare our method to phase prediction schemes that track the frequency of code blocks touched during execution as well as traditional reactive controllers, and demonstrate significant increases in prediction accuracy at fine-granularities. For a big-little heterogeneous system that is comprised of a high performing out-of-order core (Big) and an energy-efficient, in-order core (Little), at granularities of 300 instructions, the trace based predictor can spend 28% of execution time on the Little, while targeting a maximum performance degradation of 5%. This translates to an increased energy savings of 15% on average over running only on Big, representing a 60% increase over existing techniques.
Shruti Padmanabha, Andrew Lukefahr, Reetuparna Das, Scott A. Mahlke
MICRO1
2012 Composite Cores: Pushing Heterogeneity Into a Core
abstract
Heterogeneous multicore systems -- comprised of multiple cores with varying capabilities, performance, and energy characteristics -- have emerged as a promising approach to increasing energy efficiency. Such systems reduce energy consumption by identifying phase changes in an application and migrating execution to the most efficient core that meets its current performance requirements. However, due to the overhead of switching between cores, migration opportunities are limited to coarse-grained phases (hundreds of millions of instructions), reducing the potential to exploit energy efficient cores. We propose Composite Cores, an architecture that reduces switching overheads by bringing the notion of heterogeneity within a single core. The proposed architecture pairs big and little compute μEngines that together can achieve high performance and energy efficiency. By sharing much of the architectural state between the μEngines, the switching overhead can be reduced to near zero, enabling fine-grained switching and increasing the opportunities to utilize the little μEngine without sacrificing performance. An intelligent controller switches between the μEngines to maximize energy efficiency while constraining performance loss to a configurable bound. We evaluate Composite Cores using cycle accurate micro architectural simulations and a detailed power model. Results show that, on average, the controller is able to map 25% of the execution to the little μEngine, achieving an 18% energy savings while limiting performance loss to 5%.
Andrew Lukefahr, Shruti Padmanabha, Reetuparna Das, Faissal M. Sleiman, Ronald G. Dreslinski, Thomas F. Wenisch, Scott A. Mahlke
MICRO2