Sundaram Ananthanarayanan

dblp:80/10926 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 61% Empirical software engineering · 30% Software testing · 9%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 20% Processor architecture and microarchitecture · 20% Energy-efficient computing · 20%
Artificial intelligence
1 paper
Speech recognition and synthesis · 50% Deep learning architectures and training · 50%

Topics — the 12 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
change management
0.412019
Keeping Master Green at Scale · EuroSys 2019
Software maintenance and evolution › release engineering
continuous integration
0.412019
Keeping Master Green at Scale · EuroSys 2019
Empirical software engineering
mining software repositories
0.412019
Keeping Master Green at Scale · EuroSys 2019
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Processor architecture and microarchitecture › special-purpose processor
coprocessor design
0.112012
Reliable computing with ultra-reduced instruction set co-processors · DAC 2012
Energy-efficient computing › power management
dynamic power management
0.112012
EmPower: FPGA based emulation of dynamic power management algorithms for multi-core systems on chip (abstract only) · FPGA 2012
Distributed systems
fault tolerance
0.112012
Reliable computing with ultra-reduced instruction set co-processors · DAC 2012
High-performance computing
performance optimization at scale
0.112016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Reconfigurable computing and FPGAs
FPGA-based emulation
0.012012
EmPower: FPGA based emulation of dynamic power management algorithms for multi-core systems on chip (abstract only) · FPGA 2012
Reconfigurable computing and FPGAs
FPGA prototyping
0.012012
Reliable computing with ultra-reduced instruction set co-processors · DAC 2012
Embedded and real-time systems › embedded hardware platform › MPSoC
multicore system-on-chip
0.012012
EmPower: FPGA based emulation of dynamic power management algorithms for multi-core systems on chip (abstract only) · FPGA 2012

Methods — techniques the papers use, named apart from their topics

batch dispatch · 0.5GPU-based inference · 0.5power-aware thread migration · 0.1formal proof · 0.1dynamic frequency scaling · 0.1clock gating · 0.1LLVM compiler back-end · 0.1FPGA-based emulation · 0.1
YearPublicationVenuePosition
2019 Keeping Master Green at Scale
abstract
Giant monolithic source-code repositories are one of the fundamental pillars of the back end infrastructure in large and fast-paced software companies. The sheer volume of everyday code changes demands a reliable and efficient change management system with three uncompromisable key requirements --- always green master, high throughput, and low commit turnaround time. Green refers to a master branch that always successfully compiles and passes all build steps, the opposite being red. A broken master (red) leads to delayed feature rollouts because a faulty code commit needs to be detected and rolled backed. Additionally, a red master has a cascading effect that hampers developer productivity--- developers might face local test/build failures, or might end up working on a codebase that will eventually be rolled back.
Sundaram Ananthanarayanan, Masoud Saeida Ardekani, Denis Haenikel, Balaji Varadarajan, Simon Soriano, Ali-Reza Adl-Tabatabai
EuroSys1
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML2
2013 Low cost permanent fault detection using ultra-reduced instruction set co-processors
abstract
In this paper, we propose a new, low hardware overhead solution for permanent fault detection at the micro-architecture/instruction level. The proposed technique is based on an ultra-reduced instruction set co-processor (URISC) that, in its simplest form, executes only one Turing complete instruction — the subleq instruction. Thus, any instruction on the main core can be redundantly executed on the URISC using a sequence of subleq instructions, and the results can be compared, also on the URISC, to detect faults. A number of novel software and hardware techniques are proposed to decrease the performance overhead of online fault detection while keeping the error detection latency bounded including: (i) URISC routines and hardware support to check both control and data flow instructions; (ii) checking only a subset of instructions in the code based on a novel check window criterion; and (iii) URISC instruction set extensions. Our experimental results, based on FPGA synthesis and RTL simulations, illustrate the benefits of the proposed techniques.
Sundaram Ananthanarayanan, Siddharth Garg, Hiren D. Patel
DATE1
2012 Reliable computing with ultra-reduced instruction set co-processors
abstract
This work presents a method to reliably perform computations in the presence of hard faults arising from aggressive technology scaling, and design defects from human error. Our method is based on an observation that a single Turing-complete instruction can mirror the semantics of any other instruction. One such instruction is the subleq instruction, which has been used for instructional purposes in the past. We find that the scope for using such a Turing-complete instruction is far greater, and in this paper, we present its applicability to fault tolerance. In particular, we extend a MIPS processor with a co-processor (called ultra-reduced instruction set co-processor -- URISC) that implements the subleq instruction. We use the URISC to execute sequences of subleq that are semantically equivalent to the faulty instructions. We formally prove this, and implement the translations in the back-end of the LLVM compiler. We generate binaries for our hardware prototype called MIPS-URISC, which we synthesize and execute on an Altera FPGA. Our experiments indicate the performance and area overheads, and the efficacy of the proposed approach.
Aravindkumar Rajendiran, Sundaram Ananthanarayanan, Hiren D. Patel, Mahesh Tripunitara, Siddharth Garg
DAC2
2012 EmPower: FPGA based emulation of dynamic power management algorithms for multi-core systems on chip (abstract only)
abstract
Dynamic power management for multi-core system on chip (MPSoC) platforms has become an increasingly critical design problem. In this paper, we present EmPower, an FPGA based emulation, validation and prototyping framework for dynamic power management research targeted at MPSoC platforms. EmPower supports two advanced power management features -- per-core dynamic frequency scaling and clock gating, and power-aware thread migration. We also provide two fully-functional parallel applications for benchmarking -- video encoding and software-defined radio. Our experimental results indicate that EmPower provides up to 36 -- improvement in run-time compared to cycle-accurate software simulations, and enables accurate and efficient exploration of the design space of power management algorithms.
Sundaram Ananthanarayanan, Chirag Ravishankar, Siddharth Garg, Andrew A. Kennings
FPGA1
2012 EmPower: FPGA based rapid prototyping of dynamic power management algorithms for multi-processor systems on chip
abstract
Dynamic power management for multi-core system on chip (MPSoC) platforms has become an increasingly critical design problem. In this paper, we present EmPower, an FPGA based rapid prototyping framework for dynamic power management algorithms targeted at MPSoC platforms. EmPower supports two advanced power management techniques (per-core dynamic frequency scaling and clock gating, and thread migration), enables software based expression of the power management algorithm, and provides on-chip power measurement capabilities. EmPower also includes two fully-functional parallel applications for benchmarking - video encoding and software-defined radio. We demonstrate the use of EmPower in rapid exploration of the large design space of power management algorithms using two illustrative case studies.
Chirag Ravishankar, Sundaram Ananthanarayanan, Siddharth Garg, Andrew A. Kennings
FPL2