Sunil Shukla

dblp:12/6233 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
1since 2021 · last 2021
0000-0002-9268-4096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Hardware accelerators and domain-specific architectures · 52% Emerging computing paradigms · 21% Energy-efficient computing · 18%
Software engineering, system software, and programming languages
3 papers
Runtime systems and virtual machines · 52% Compilers and program optimization · 26% Programming languages and type systems · 22%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.232021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Accelerator Design for Deep Learning Training: Extended Abstract: Invited · DAC 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.512021
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.422021
Accelerator Design for Deep Learning Training: Extended Abstract: Invited · DAC 2017
RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021
Emerging computing paradigms
approximate computing
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Hardware accelerators and domain-specific architectures
approximate computing accelerator
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Emerging computing paradigms › approximate computing
cross-layer approximate computing
0.412020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Energy-efficient computing
power management
0.312017
Accelerator Design for Deep Learning Training: Extended Abstract: Invited · DAC 2017
Electronic design automation
high-level synthesis
0.322012
And then there were none: a stall-free real-time garbage collector for reconfigurable hardware · PLDI 2012
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011
Runtime systems and virtual machines
garbage collection
0.112012
And then there were none: a stall-free real-time garbage collector for reconfigurable hardware · PLDI 2012
Compilers and program optimization › accelerator compilation
heterogeneous compilation
0.112012
A compiler and runtime for heterogeneous computing · DAC 2012
Runtime systems and virtual machines › garbage collection
real-time garbage collection
0.112012
And then there were none: a stall-free real-time garbage collector for reconfigurable hardware · PLDI 2012
Machine learning › Deep learning architectures and training
neural network inference
0.112020
Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020
Programming languages and type systems › domain-specific languages
hardware description languages
0.112011
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011
Cloud and datacenter computing › resource management › real-time resource allocation
real-time memory management
0.012012
And then there were none: a stall-free real-time garbage collector for reconfigurable hardware · PLDI 2012
GPUs and heterogeneous computing
heterogeneous architecture
0.012011
Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language · DAC 2011

Methods — techniques the papers use, named apart from their topics

quantization · 0.9pruning · 0.9mixed-precision arithmetic · 0.9custom number representation · 0.9performance modeling · 0.5runtime orchestration · 0.3hardware synthesis · 0.3logic synthesis · 0.2behavioral synthesis · 0.2
YearPublicationVenuePosition
2021 RaPiD: AI Accelerator for Ultra-low Precision Training and Inference
abstract
The growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications.
Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan
ISCA52
2020 Efficient AI System Design With Cross-Layer Approximate Computing
abstract
Advances in deep neural networks (DNNs) and the availability of massive real-world data have enabled superhuman levels of accuracy on many AI tasks and ushered the explosive growth of AI workloads across the spectrum of computing devices. However, their superior accuracy comes at a high computational cost, which necessitates approaches beyond traditional computing paradigms to improve their operational efficiency. Leveraging the application-level insight of error resilience, we demonstrate how approximate computing (AxC) can significantly boost the efficiency of AI platforms and play a pivotal role in the broader adoption of AI-based applications and services. To this end, we present RaPiD, a multi-tera operations per second (TOPS) AI hardware accelerator core (fabricated at 14-nm technology) that we built from the ground-up using AxC techniques across the stack including algorithms, architecture, programmability, and hardware. We highlight the workload-guided systematic explorations of AxC techniques for AI, including custom number representations, quantization/pruning methodologies, mixed-precision architecture design, instruction sets, and compiler technologies with quality programmability, employed in the RaPiD accelerator.
Swagath Venkataramani, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Mingu Kang, Ankur Agarwal, Jinwook Oh, Shubham Jain 0004, Tina Babinsky, Nianzheng Cao, Thomas W. Fox, Bruce M. Fleischer, George Gristede, Michael Guillorn, Howard Haynie, Hiroshi Inoue, Kazuaki Ishizaki, Michael J. Klaiber, Shih-Hsien Lo, Gary W. Maier, Silvia M. Müller, Michael Scheuermann, Eri Ogawa, Marcel Schaal, Mauricio J. Serrano, Joel Silberman, Christos Vezyrtzis, Wei Wang 0333, Fanchieh Yee, Matthew M. Ziegler, Ching Zhou, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Vijayalakshmi Srinivasan, Leland Chang, Kailash Gopalakrishnan
Proc. IEEE37
2018 Across the Stack Opportunities for Deep Learning Acceleration
abstract
The combination of growth in compute capabilities and availability of large datasets has led to a re-birth of deep learning. Deep Neural Networks (DNNs) have become state-of-the-art in a variety of machine learning tasks spanning domains across vision, speech, and machine translation. Deep Learning (DL) achieves high accuracy in these tasks at the expense of 100s of ExaOps of computation; posing significant challenges to efficient large-scale deployment in both resource-constrained environments and data centers.
Vijayalakshmi Srinivasan, Bruce M. Fleischer, Sunil Shukla, Matthew M. Ziegler, Joel Silberman, Jinwook Oh, Jungwook Choi, Silvia M. Müller, Ankur Agrawal, Tina Babinsky, Nianzheng Cao, Chia-Yu Chen, Pierce Chuang, Thomas W. Fox, George Gristede, Michael Guillorn, Howard Haynie, Michael J. Klaiber, Dongsoo Lee, Shih-Hsien Lo, Gary W. Maier, Michael Scheuermann, Swagath Venkataramani, Christos Vezyrtzis, Naigang Wang, Fanchieh Yee, Ching Zhou, Pong-Fei Lu, Brian W. Curran, Leland Chang, Kailash Gopalakrishnan
ISLPED3
2017 Accelerator Design for Deep Learning Training: Extended Abstract: Invited
abstract
Deep Neural Networks (DNNs) have emerged as a powerful and versatile set of techniques showing successes on challenging artificial intelligence (AI) problems. Applications in domains such as image/video processing, autonomous cars, natural language processing, speech synthesis and recognition, genomics and many others have embraced deep learning as the foundation. DNNs achieve superior accuracy for these applications with high computational complexity using very large models which require 100s of MBs of data storage, exaops of computation and high bandwidth for data movement. In spite of these impressive advances, it still takes days to weeks to train state of the art Deep Networks on large datasets - which directly limits the pace of innovation and adoption. In this paper, we present a multi-pronged approach to address the challenges in meeting both the throughput and the energy efficiency goals for DNN training.
Ankur Agrawal, Chia-Yu Chen, Jungwook Choi, Kailash Gopalakrishnan, Jinwook Oh, Sunil Shukla, Vijayalakshmi Srinivasan, Swagath Venkataramani, Wei Zhang 0022
DAC6
2015 Cycle-Accurate Replay and Debugging of Running FPGA Systems
abstract
Finding bugs in software that are timing dependent or caused by non-deterministic inputs is notoriously difficult. In FPGAs, the problem is much worse because the visibility into the running design tends to be very low, and existing tools either gather too little data for diagnosis, or are so intrusive that they perturb the timing and may mask the bug. This leads to long FPGA development cycles, and is exacerbated by the steady increase in complexity of FPGA designs. We present a tool - Panoptic on - that logs data and timing information at key design points, extracts it from the FPGA, and uses it for cycle-accurate replay of the entire execution in simulation. This allows interactive debugging with full visibility into the design.
Sunil Shukla, David F. Bacon
FCCM1
2014 Parallel real-time garbage collection of multiple heaps in reconfigurable hardware
abstract
Despite rapid increases in memory capacity, reconfigurable hardware is still programmed in a very low-level manner, generally without any dynamic allocation at all. This limits productivity especially as the larger chips encourage more and more complex designs to be attempted.
David F. Bacon, Perry Cheng, Sunil Shukla
ISMM3
2013 The Liquid Metal IP bridge
abstract
Programmers are increasingly turning to heterogeneous systems to achieve performance. Examples include FPGA-based systems that integrate reconfigurable architectures with conventional processors. However, the burden of managing the coding complexity that is intrinsic to these systems falls entirely on the programmer. This limits the proliferation of these systems as only highly-skilled programmers and FPGA developers can unlock their potential. The goal of the Liquid Metal project at IBM Research is to address the programming complexity attributed to heterogeneous FPGA-based systems. A feature of this work is a vertically integrated development lifecycle that appeals to skilled software developers. A primary enabler for this work is a canonical IP bridge, designed to offer a uniform communication methodology between software and hardware, and that is applicable across a wide range of platforms available off-the-shelf.
Perry Cheng, Stephen J. Fink, Rodric M. Rabbah, Sunil Shukla
ASP-DAC4
2013 QUKU: A dual-layer reconfigurable architecture
abstract
A new architecture, QUKU, is proposed for implementing stream-based algorithms on FPGAs, which combines the advantages of FPGA and Coarse Grain Reconfigurable Arrays (CGRAs). QUKU consists of a dynamically reconfigurable, coarse-grain Processing Element (PE) array with an associated softcore processor providing system support. At a coarse-grain, the PE array can be reconfigured on a cycle-by-cycle basis to change the PE functionality similarly to that in a conventional CGRA. At a fine-grain, the whole FPGA can be reconfigured statically to implement a completely different PE array that serves the target application in a better way. Advantages of the fine-grain reconfiguration include individually customized PEs, adaptable numeric format support and customizable interconnect network. A prototype CAD tool framework is also developed which facilitates programming the QUKU architecture. An example application consisting of two different image detectors is implemented to demonstrate the advantages of QUKU. QUKU provides up to 140 times speedup and 40 times improvement in area-time product compared to an implementation running on an FPGA-based softcore. The area-time product for QUKU is around 16% lower than that of a custom circuit based implementation on the same FPGA. The per-PE customization provides an area-time saving of approximately 31% compared to a homogeneous 4 × 4 array of PEs for the same application. The experimental results demonstrate that a dual layered reconfigurable architecture provides significant potential benefits in terms of flexibility, area and processing efficiency over existing reconfigurable computing architectures for DSP.
Neil W. Bergmann, Sunil Shukla, Jürgen Becker 0001
ACM Trans. Embed. Comput. Syst.2
2012 A compiler and runtime for heterogeneous computing
abstract
Heterogeneous systems show a lot of promise for extracting high-performance by combining the benefits of conventional architectures with specialized accelerators in the form of graphics processors (GPUs) and reconfigurable hardware (FPGAs). Extracting this performance often entails programming in disparate languages and models, making it hard for a programmer to work equally well on all aspects of an application. Further, relatively little attention is paid to co-execution---the problem of orchestrating program execution using multiple distinct computational elements that work seamlessly together.
Joshua S. Auerbach, David F. Bacon, Ioana Burcea, Perry Cheng, Stephen J. Fink, Rodric M. Rabbah, Sunil Shukla
DAC7
2012 And then there were none: a stall-free real-time garbage collector for reconfigurable hardware
abstract
Programmers are turning to radical architectures such as reconfigurable hardware (FPGAs) to achieve performance. But such systems, programmed at a very low level in languages with impoverished abstractions, are orders of magnitude more complex to use than conventional CPUs. The continued exponential increase in transistors, combined with the desire to implement ever more sophisticated algorithms, makes it imperative that such systems be programmed at much higher levels of abstraction. One of the fundamental high-level language features is automatic memory management in the form of garbage collection.
David F. Bacon, Perry Cheng, Sunil Shukla
PLDI3
2011 Virtualization of heterogeneous machines hardware description in a synthesizable object-oriented language
abstract
Lime is a new Java-compatible and object-oriented language designed to make programming of reconflgurable hardware significantly more accessible to skilled software developers. Lime programs may run either in software (via Java bytecodes) or in hardware (via behavioral and logic synthesis). This paper illustrates the salient synthesis-oriented features of the language using a photo-mosaic algorithm with inherent bit, pipeline, and data parallelism. The result is a virtual machine abstraction that extends across a heterogeneous architecture comprising a CPU, FPGA, and other computational structures.
Joshua S. Auerbach, David F. Bacon, Perry Cheng, Rodric M. Rabbah, Sunil Shukla
DAC5
2010 FPGA-based combined architecture for stream categorization and intrusion detection
abstract
This paper presents a working solution for the MEMOCODE 2010 design contest. The design presented in this paper is implemented in the Xilinx V5LX330 FPGA as a custom circuit. The solution implements pattern matching logic for all the mandatory and optional patterns while maintaining the required line rate of 500 Mbps.
Sunil Shukla, Rodric M. Rabbah, Martin Vorbach
MEMOCODE1
2007 Improving High Quality TTS using Circular Linear Prediction and Constant Pitch Transform
abstract
Current high quality concatenative TTS systems are based on unit selection from a database that is contextually and prosodically rich. These systems are computationally expensive and require a very large footprint. This paper presents a new method for representing speech segments that can improve the quality and scalability of concatenative TTS systems. The circular linear prediction model combined with the constant pitch transform provides a robust representation of speech signals that allows for limited prosodic movements without perceivable loss in quality. A method is presented for constraining the LSF tracks of speech segments to realize pitch modifications with minimal artifacts. The results of formal listening tests demonstrate that limited prosodic modifications can produce speech from fewer units whose quality equals or exceeds large database unit-selection systems. Additionally, this method is used to realize high quality emphasized speech.
Sunil Shukla, Thomas P. Barnwell III
ICASSP (4)1
2007 QUKU: A FPGA Based Flexible Coarse Grain Architecture Design Paradigm using Process Networks
abstract
DSP applications can be suitably represented using process network models. This paper uses a modification of Kahn process network to solve the problem of finding an optimum architectural template for coarse grain array on per application basis. By applying the model at architectural level in QUKU, better hardware efficiency is achieved for a wide domain of applications. A few widely used DSP algorithms have been presented to demonstrate the application of process network models into architectural template generation in QUKU.
Sunil Shukla, Neil W. Bergmann, Jürgen Becker 0001
IPDPS1
2006 From Equation to VHDL: Using Rewriting Logic for Automated Function Generation
abstract
This paper presents a novel tool flow combining rewriting logic with hardware synthesis. It enables the automated generation of synthesizable VHDL code from mathematical equations and the quick generation of functionally equivalent alternative implementations. The simple but powerful semantics of rewriting logic provide a natural mechanism for manipulating algebraic expressions, using a high-level of abstraction which is afterwards automatically converted into lower levels of abstraction. The design flow is validated by generating polynomial approximations for arbitrary continuous functions. The polynomial generation process is completely parameterized regarding polynomial degree, number representation parameters, word width and polynomial evaluation approaches. Different functionally equivalent implementations for the resulting polynomial approximations were generated and synthesized for a Virtex4 device
Carlos Morra, M. Sackmann, Sunil Shukla, Jürgen Becker 0001, Reiner W. Hartenstein
FPL3
2004 Single bit error correction implementation in CRC-16 on FPGA
abstract
Framing protocols employ cyclic redundancy check (CRC) to detect errors incurred during transmission. Generally whole frame is protected using CRC and upon detection of error, retransmission is requested. But certain protocols demand for single bit error correction capabilities for the header part of the frame, which often plays an important role in receiver synchronization. At a speed of 10 Gbps, header error correction implementation in hardware can be a bottleneck. This work presents a hardware efficient way of implementing CRC-16 over 16 bits of data, multiple bit error detection and single bit error correction on FPGA device.
Sunil Shukla, Neil W. Bergmann
FPT1
2002 Circular LPC modeling and constant pitch transform for accurate speech analysis and high quality speech synthesis
abstract
In this paper, Circular LPC analysis, a windowless signal modeling method for periodic signals, is re-visited as a high-resolution pitch-synchronous speech spectrum analysis tool. In addition, the Constant Pitch Transform is introduced as an alternative residual signal representation, suitable for speech coding and text-to-speech synthesis applications. Since circular LPC analysis requires exact periodicity, the original narrowband or wideband signal is upsampled to account for fractional pitch periods at original sampling rate. Hence, both methods are formulated to work on the higher sampling rates. The analysis and synthesis using individual pitch cycles are explained in detail. Furthermore, three unique circular LPC synthesis methods are presented to address the problem of initial rest. The results of preliminary experiments show that it is possible to synthesize artifact free speech using the proposed method.
Sunil Shukla, Ali Erdem Ertan, Thomas P. Barnwell III
ICASSP1