EDBT 2026 Demo / reviewers in the wild / expert
Shrihari Sridharan
dblp:42/8626
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-4259-5263ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 57% Memory systems · 32% GPUs and heterogeneous computing · 6% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
edge accelerator |
0.8 | 1 | 2024 | Ev-Edge: Efficient Execution of Event-based Vision Algorithms on Commodity Edge Platforms · DAC 2024 |
Hardware accelerators and domain-specific architectures › vision accelerator
event-based vision accelerator |
0.8 | 1 | 2024 | Ev-Edge: Efficient Execution of Event-based Vision Algorithms on Commodity Edge Platforms · DAC 2024 |
Memory systems
in-memory computing |
0.4 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures
matrix-vector multiplication |
0.4 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Memory systems
non-volatile memory |
0.4 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Memory systems
processing-in-memory |
0.4 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Distributed systems › edge computing
edge computing platforms |
0.2 | 1 | 2024 | Ev-Edge: Efficient Execution of Event-based Vision Algorithms on Commodity Edge Platforms · DAC 2024 |
Machine learning › Deep learning architectures and training › neural network inference
DNN inference |
0.1 | 1 | 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and Challenges · Proc. IEEE 2020 |
Methods — techniques the papers use, named apart from their topics
matrix-vector multiplication · 0.9in-memory computing · 0.9analog computing · 0.9sparse frame encoding · 0.8precision selection · 0.8network mapping · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Ev-Edge: Efficient Execution of Event-based Vision Algorithms on Commodity Edge PlatformsabstractEvent cameras have emerged as a promising sensing modality for autonomous navigation systems, owing to their high temporal resolution, high dynamic range and negligible motion blur. To achieve the best performance across a range of vision tasks using the asynchronous temporal event streams from such sensors, recent research has shown that a mix of Artificial Neural Networks (ANNs), Spiking Neural Networks (SNNs), as well as hybrid SNN-ANN algorithms are desirable. However, we observe that such workloads achieve poor utilization and performance on commodity edge platforms which feature heterogeneous processing elements such as CPUs, GPUs and neural accelerators. This is due to the mismatch between the irregular nature of event streams and diverse characteristics of algorithms on the one hand and the underlying hardware platform on the other. We propose Ev-Edge, a framework that contains three key optimizations to boost the performance of event-based vision algorithms on edge platforms: (1) An Event2Sparse Frame converter directly transforms raw event streams into sparse frames, enabling the use of sparse libraries with minimal encoding overheads (2) A Dynamic Sparse Frame Aggregator merges sparse frames at runtime by trading off the temporal granularity of events and computational demand, thereby improving hardware utilization, and (3) A Network Mapper maps concurrently executing tasks to different processing elements while also selecting layer precision while considering both compute and communication overheads. On several state-of-art networks for a range of autonomous navigation tasks, Ev-Edge achieves 1.28x-2.05x improvements in latency and 1.23x-2.15x in energy over an all-GPU implementation on the NVIDIA Jetson Xavier AGX platform for single-task execution scenarios. Ev-Edge also achieves 1.43x-1.81x latency improvements over round-robin scheduling methods in multi-task execution scenarios. Shrihari Sridharan, Surya Selvam, Kaushik Roy 0001, Anand Raghunathan |
DAC | 1 |
| 2023 | X-Former: In-Memory Acceleration of TransformersabstractTransformers have achieved great success in a wide variety of natural language processing (NLP) tasks due to the self-attention mechanism, which assigns an importance score for every word relative to other words in a sequence. However, these models are very large, often reaching hundreds of billions of parameters, and therefore require a large number of dynamic random access memory (DRAM) accesses. Hence, traditional deep neural network (DNN) accelerators such as graphical processing units (GPUs) and tensor processing units (TPUs) face limitations in processing Transformers efficiently. In-memory accelerators based on nonvolatile memory (NVM) promise to be an effective solution to this challenge, since they provide high storage density while performing massively parallel matrix–vector multiplications (MVMs) within memory arrays. However, attention score computations, which are frequently used in Transformers unlike convolutional neural networks (CNNs) and recurrent neural network (RNNs), require MVMs where both the operands change dynamically for each input. As a result, conventional NVM-based accelerators incur high write latency and write energy when used for Transformers and further suffer from the low endurance of most NVM technologies. To address these challenges, we present X-Former, a hybrid in-memory hardware accelerator that consists of both NVM and CMOS processing elements to execute transformer workloads efficiently. To improve the hardware utilization of X-Former, we also propose a sequence blocking dataflow, which overlaps the computations of the two processing elements and reduces execution time. Across several benchmarks, we show that X-Former achieves up to$69.8\times $and$13\times $improvements in latency and energy over a NVIDIA GeForce GTX 1060 GPU and up to$24.1\times $and$7.95\times $improvements in latency and energy over a state-of-the-art in-memory NVM accelerator. Shrihari Sridharan, Jacob R. Stevens, Kaushik Roy 0001, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | TxSim: Modeling Training of Deep Neural Networks on Resistive Crossbar SystemsabstractDeep neural networks (DNNs) have gained tremendous popularity in recent years due to their ability to achieve superhuman accuracy in a wide variety of machine learning tasks. However, the compute and memory requirements of DNNs have grown rapidly, creating a need for energy-efficient hardware. Resistive crossbars have attracted significant interest in the design of the next generation of DNN accelerators due to their ability to natively execute massively parallel vector-matrix multiplications within dense memory arrays. However, crossbar-based computations face a major challenge due to device and circuit-level nonidealities, which manifest as errors in the vector-matrix multiplications and eventually degrade DNN accuracy. To address this challenge, there is a need for tools that can model the functional impact of nonidealities on DNN training and inference. Existing efforts toward this goal are either limited to inference or are too slow to be used for large-scale DNN training. We propose TxSim, a fast and customizable modeling framework to functionally evaluate DNN training on crossbar-based hardware considering the impact of nonidealities. The key features of TxSim that differentiate it from prior efforts are: 1) it comprehensively models nonidealities during all training operations (forward propagation, backward propagation, and weight update) and 2) it achieves computational efficiency by mapping crossbar evaluations to well-optimized Basic Linear Algebra Subprograms (BLAS) routines and incorporates speedup techniques to further reduce simulation time with minimal impact on accuracy. TxSim achieves 6×- 108× improvement in simulation speed over prior works, and thereby makes it feasible to evaluate the training of large-scale DNNs on crossbars. Our experiments using TxSim reveal that the accuracy degradation in DNN training due to nonidealities can be substantial (3%-36.4%) for large-scale DNNs and data sets, underscoring the need for further research in mitigation techniques. We also analyze the impact of various device and circuit-level parameters and the associated nonidealities to provide key insights that can guide the design of crossbar-based DNN training accelerators. Sourjya Roy, Shrihari Sridharan, Shubham Jain 0004, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Resistive Crossbars as Approximate Hardware Building Blocks for Machine Learning: Opportunities and ChallengesabstractTraditional computing systems based on the von Neumann architecture are fundamentally bottlenecked by data transfers between processors and memory. The emergence of data-intensive workloads, such as machine learning (ML), creates an urgent need to address this bottleneck by designing computing platforms that utilize the principle of colocated memory and processing units. Such an approach, known as “in-memory computing,” can potentially eliminate data movement costs by computing inside the memory array itself. Crossbars based on resistive nonvolatile memory (NVM) devices have shown immense promise in serving as the building blocks of in-memory computing systems for ML workloads. This is because their high density can lead to higher on-chip storage capacity, while they can also perform massively parallel, in situ matrix-vector multiplication (MVM) operations, thereby accelerating the main computational kernel of ML workloads. However, resistive crossbar-based analog computing is inherently approximate due to the device- and circuit-level nonidealities. Furthermore, the area and energy costs of peripheral circuits for conversions between the analog and digital domains can greatly diminish the intrinsic efficiency of crossbar-based MVM computation. We present a comprehensive overview of the emerging paradigm of computing using NVM crossbars for accelerating ML workloads. We describe the design principles of resistive crossbars, including the devices and associated circuits that constitute them. We discuss intrinsic approximations arising from the device and circuit characteristics and study their functional impact on the MVM operation. Next, we present an overview of spatial architectures that exploit the high storage density of NVM crossbars. Furthermore, we elaborate on software frameworks that effectively capture device-circuit-architecture characteristics to evaluate the performance of large-scale deep neural networks (DNNs) using resistive crossbar-based hardware. Finally, we discuss open challenges and future research directions that need to be explored in order to realize the vision of resistive crossbars as the building blocks of future computing platforms. Indranil Chakraborty, Mustafa Fayez Ali, Aayush Ankit, Shubham Jain 0004, Sourjya Roy, Shrihari Sridharan, Amogh Agrawal, Anand Raghunathan, Kaushik Roy 0001 |
Proc. IEEE | 6 |