Suchit Subhaschandra

dblp:60/10774 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 52% Reconfigurable computing and FPGAs · 13% High-performance computing · 13%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › numerical linear algebra
GEMM
0.312018
A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case Study · FPGA 2018
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
matrix multiplication accelerator
0.312018
A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case Study · FPGA 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN acceleration
0.312017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks? · FPGA 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.312017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks? · FPGA 2017
Reconfigurable computing and FPGAs
FPGA accelerator
0.312017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks? · FPGA 2017
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks? · FPGA 2017
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore
0.112012
QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012
Electronic design automation
design space exploration
0.112012
QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012
GPUs and heterogeneous computing
heterogeneous architecture
0.112012
QuickIA: Exploring heterogeneous architectures on real prototypes · HPCA 2012
Hardware accelerators and domain-specific architectures
deep learning
0.112018
A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case Study · FPGA 2018
GPUs and heterogeneous computing
GPU computing
0.112017
Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks? · FPGA 2017

Methods — techniques the papers use, named apart from their topics

reduced precision arithmetic · 0.3performance comparison · 0.3prototyping · 0.1case study · 0.1
YearPublicationVenuePosition
2018 A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case Study
abstract
General Matrix to Matrix multiplication (GEMM) is the cornerstone for a wide gamut of applications in high performance computing (HPC), scientific computing (SC) and more recently, deep learning. In this work, we present a customizable matrix multiplication framework for the Intel HARPv2 CPU+FPGA platform that includes support for both traditional single precision floating point and reduced precision workloads. Our framework supports arbitrary size GEMMs and consists of two parts: (1) a simple application programming interface (API) for easy configuration and integration into existing software and (2) a highly customizable hardware template. The API provides both compile and runtime options for controlling key aspects of the hardware template including dynamic precision switching; interleaving and block size control; and fused deep learning specific operations. The framework currently supports single precision floating point (FP32), 16, 8, 4 and 2 bit Integer and Fixed Point (INT16, INT8, INT4, INT2) and more exotic data types for deep learning workloads: INT16xTernary, INT8xTernary, BinaryxBinary.
Duncan J. M. Moss, Krishnan Srivatsan, Eriko Nurvitadhi, Piotr Ratuszniak, Jaewoong Sim, Asit K. Mishra, Debbie Marr, Suchit Subhaschandra, Philip H. W. Leong
FPGA9
2017 SWiF: A Simplified Workload-Centric Framework for FPGA-Based Computing
abstract
In this paper, we introduce SWiF - Simplified Workload-intuitive Framework - a workload-centric, application programming framework designed to simplify the large-scale deployment of FPGAs in end-to-end applications. SWiF can intelligently mediate access to shared resources by orchestrating the distribution and scheduling of tasks across a heterogeneous mix of FPGA and CPU resources in order to improve utilization and maintain system requirements. We implemented SWiF atop Intel Accelerator Abstraction Layer (AAL) and deployed the resulting software stack in a datacenter with an Intel-based Xeon+FPGA server running Apache Spark. We demonstrate that by using SWiF's API, developers can flexibly and easily deploy FPGA-enabled applications and frameworks with almost no change to existing software stack. In particular, we demonstrate that by offloading through SWiF the compression workload of Spark unto FPGA, we gain a speedup of 3.2X in total job execution, and up to 5X when Spark's Resilient Distributed Datasets (RDDs) are persisted in memory.
David Ojika, Piotr Majcher, Wojciech Neubauer, Suchit Subhaschandra, Darin Acosta
FCCM4
2017 Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks?
Eriko Nurvitadhi, Ganesh Venkatesh, Jaewoong Sim, Debbie Marr, Randy Huang, Jason Ong Gee Hock, Yeong Tat Liew, Krishnan Srivatsan, Duncan J. M. Moss, Suchit Subhaschandra, Guy Boudoukh
FPGA10
2017 High performance binary neural networks on the Xeon+FPGA™ platform
abstract
Convolutional neural networks (CNNs) are deployed in a wide range of image recognition, scene segmentation and object detection applications. Achieving state of the art accuracy in CNNs often results in large models and complex topologies that require significant compute resources to complete in a timely manner. Binarised neural networks (BNNs) have been proposed as an optimised variant of CNNs, which constrain the weights and activations to +1 or -1 and thus offer compact models and lower computational complexity per operation. This paper presents a high performance BNN accelerator on the Intel®Xeon+FPGA™ platform. The proposed accelerator is designed to take advantage of the Xeon+FPGA system in a way that a specialised FPGA architecture can be targeted for the most compute intensive parts of the BNN whilst other parts of the topology can be handled by the Xeon™ CPU. The implementation is evaluated by comparing the raw compute performance and energy efficiency for key layers in standard CNN topologies against an Nvidia Titan X Pascal GPU and other published FPGA BNN accelerators. The results show that our single-package integrated Arria™ 10 FPGA accelerator coupled with a high-end Xeon CPU can offer comparable performance and better energy efficiency than a high-end discrete Titan X GPU card. In addition, our solution delivers the best performance compared to previous BNN FPGA implementations.
Duncan J. M. Moss, Eriko Nurvitadhi, Jaewoong Sim, Asit K. Mishra, Debbie Marr, Suchit Subhaschandra, Philip H. W. Leong
FPL6
2017 Customizable FPGA OpenCL matrix multiply design template for deep neural networks
abstract
Deep neural networks (DNNs) have gained popularity for their state-of-the-art accuracy and relative ease of use. DNNs rely on a growing variety of matrix multiply operations (i.e., dense to sparse, FP32 to N-bit). We propose an OpenCL-based matrix multiply design template, which enables automated design exploration to generate optimized FPGA matrix accelerators for DNN applications. Given the desired matrix operations (e.g., sparsity, data types), our template rapidly produces performance and area estimates for a variety of design variants and/or FPGA platforms. Upon identifying compelling design points and target platforms, FPGA implementations can then be generated automatically using the Intel® OpenCL™ FPGA SDK. We show the effectiveness of the template with a comparison to hand-tuned RTL, a design space exploration, and a DNN case study.
Jack Yinger, Eriko Nurvitadhi, Davor Capalija, Andrew C. Ling, Debbie Marr, Krishnan Srivatsan, Duncan J. M. Moss, Suchit Subhaschandra
FPT8
2012 QuickIA: Exploring heterogeneous architectures on real prototypes
abstract
Over the last decade, homogeneous multi-core processors emerged and became the de-facto approach for offering high parallelism, high performance and scalability for a wide range of platforms. We are now at an interesting juncture where several critical factors (smaller form factor devices, power challenges, need for specialization, etc) are guiding architects to consider heterogeneous chips and platforms for the next decade and beyond. Exploring heterogeneous architectures is challenging since it involves re-evaluating architecture options, OS implications and application development. In this paper, we describe these research challenges and then introduce a heterogeneous prototype platform called QuickIA that enables rapid exploration of heterogeneous architectures employing multiple generations of Intel processors for evaluating the implications of asymmetry and FPGAs to experiment with specialized processors or accelerators. We also show example case studies using the QuickIA research prototype to highlight its value in conducting heterogeneous architecture, OS and applications research.
Bhushan Chitlur, Ganapati Srinivasa, Scott Hahn, Dheeraj Reddy, David A. Koufaty, Paul Brett, Abirami Prabhakaran, Li Zhao 0002, Nelson Ijih, Suchit Subhaschandra, Sabina Grover, Xiaowei Jiang, Ravi R. Iyer 0001
HPCA11