Sriseshan Srikanth

dblp:190/7549 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-9203-8011ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 SlimSLAM: An Adaptive Runtime for Visual-Inertial Simultaneous Localization and Mapping
abstract
Simultaneous localization and mapping (SLAM) algorithms track an agent's movements through an unknown environment. SLAM must be fast and accurate to avoid adverse effects such as motion sickness in AR/VR headsets and navigation errors in autonomous robots and drones. However, accurate SLAM is computationally expensive and target platforms are often highly constrained. Therefore, to maintain real-time functionality, designers must either pay a large up-front cost to design specialized accelerators or reduce the algorithm's functionality, resulting in poor pose estimation.
Armand Behroozi, Vlad Fruchter, Lavanya Subramanian, Sriseshan Srikanth, Scott A. Mahlke
ASPLOS (3)5
2022 Scalable Energy-Efficient Microarchitectures With Computational Error Tolerance Via Redundant Residue Number Systems
abstract
Due to high leakage current and threshold voltage, Dennard scaling has reached its limit on conventional semiconductor technology. Energy reduction at the transistor level by simply lowering supply voltage has proven to be infeasible for these devices (e.g., MOSFETs). Some recently proposed millivolt switch techniques aim to mitigate these issues, by maintaining a high on/off ratio of drain currents with a much lower supply voltage. However,$V_{dd}$reduction is constrained by high intermittent error probabilities in millivolt switches. Energy-efficient microarchitectures that are computationally error-tolerant are therefore urgently needed. This article systematically leverages the error correction and checkpointing properties of Redundant Residue Number Systems (RRNS) by varying the number of non-redundant ($n$) and redundant ($r$) residues. The state-of-the-art of RRNS microarchitecture is confined to a fixed configuration point within such a($n$n,$r$r)-RRNSdesign plane, as it supports single error correction alone. Being able to efficiently handle resilience in this($n$n,$r$r)-RRNSplane significantly improves reliability, allowing further${V_{dd}}$reduction to save energy. To this end, first, we propose a scalable RRNS microarchitecture that simultaneously supports both, error-correction, as well as checkpointing with restart capabilities upon detecting uncorrectable errors. Second, we design a novel RRNS-based adaptive checkpointing&restart mechanisms that automatically guarantees reliability while minimizing the energy-delay product (EDP). To the best of our knowledge, these are the first set of checkpointing mechanisms targeting the RRNS infrastructure. Moreover, these mechanisms optimize the usage efficiency of memory capacity. Third, we systematically explore the RRNS design space to find the best ($n$,$r$) configuration point. For similar reliability when compared to a conventional binary core without computationally error-tolerant (runs at high$V_{dd}$), the proposed RRNS scalable microarchitecture reduces EDP by 53 percent on average for memory-intensive workloads and by 67 percent on average for non-memory-intensive workloads.
Bobin Deng, Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook
IEEE Trans. Computers2
2021 SortCache: Intelligent Cache Management for Accelerating Sparse Data Workloads
abstract
Sparse data applications have irregular access patterns that stymie modern memory architectures. Although hyper-sparse workloads have received considerable attention in the past, moderately-sparse workloads prevalent in machine learning applications, graph processing and HPC have not. Where the former can bypass the cache hierarchy, the latter fit in the cache. This article makes the observation that intelligent, near-processor cache management can improve bandwidth utilization for data-irregular accesses, thereby accelerating moderately-sparse workloads. We propose SortCache, a processor-centric approach to accelerating sparse workloads by introducing accelerators that leverage the on-chip cache subsystem, with minimal programmer intervention.
Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook
ACM Trans. Archit. Code Optim.1
2020 MetaStrider: Architectures for Scalable Memory-centric Reduction of Sparse Data Streams
abstract
Reduction is an operation performed on the values of two or more key-value pairs that share the same key. Reduction of sparse data streams finds application in a wide variety of domains such as data and graph analytics, cybersecurity, machine learning, and HPC applications. However, these applications exhibit low locality of reference, rendering traditional architectures and data representations inefficient. This article presents MetaStrider, a significant algorithmic and architectural enhancement to the state-of-the-art, SuperStrider. Furthermore, these enhancements enable a variety of parallel, memory-centric architectures that we propose, resulting in demonstrated performance that scales near-linearly with available memory-level parallelism.
Sriseshan Srikanth, Anirudh Jain, Joseph M. Lennon, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook
ACM Trans. Archit. Code Optim.1
2018 Memory System Design for Ultra Low Power, Computationally Error Resilient Processor Microarchitectures
abstract
Dennard scaling ended a decade ago. Energy reduction by lowering supply voltage has been limited because of guard bands and a subthreshold slope of over 60mV/decade in MOSFETs. On the other hand, newly-proposed logic devices maintain a high on/off ratio for drain currents even at significantly lower operating voltages. However, such ultra low power technology would eventually suffer from intermittent errors in logic as a result of operating close to the thermal noise floor. Computational error correction mitigates this issue by efficiently correcting stochastic bit errors that may occur in computational logic operating at low signal energies, thereby allowing for energy reduction by lowering supply voltage to tens of millivolts. Cores based on a Redundant Residual Number System (RRNS), which represents a number using a tuple of smaller numbers, are a promising candidate for implementing energyefficient computational error correction. However, prior RRNS core microarchitectures abstract away the memory hierarchy and do not consider the power-performance impact of RNS-based memory addressing. When compared with a non-error-correcting core addressing memory in binary, naive RNS-based memory addressing schemes cause a slowdown of over 3x/2x for inorder/out-of-order cores respectively. In this paper, we analyze RNS-based memory access pattern behavior and provide solutions in the form of novel schemes and the resulting design space exploration, thereby, extending and enabling a tangible, ultra low power RRNS based architecture.
Sriseshan Srikanth, Paul G. Rabbat, Eric R. Hein, Bobin Deng, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank
HPCA1
2018 Extending Moore's Law via Computationally Error-Tolerant Computing
abstract
Dennard scaling has ended. Lowering the voltage supply ( V dd ) to sub-volt levels causes intermittent losses in signal integrity, rendering further scaling (down) no longer acceptable as a means to lower the power required by a processor core. However, it is possible to correct the occasional errors caused due to lower V dd in an efficient manner and effectively lower power. By deploying the right amount and kind of redundancy, we can strike a balance between overhead incurred in achieving reliability and energy savings realized by permitting lower V dd . One promising approach is the Redundant Residue Number System (RRNS) representation. Unlike other error correcting codes, RRNS has the important property of being closed under addition, subtraction and multiplication, thus enabling computational error correction at a fraction of an overhead compared to conventional approaches. We use the RRNS scheme to design a Computationally-Redundant, Energy-Efficient core, including the microarchitecture, Instruction Set Architecture (ISA) and RRNS centered algorithms. From the simulation results, this RRNS system can reduce the energy-delay-product by about 3× for multiplication intensive workloads and by about 2× in general, when compared to a non-error-correcting binary core.
Bobin Deng, Sriseshan Srikanth, Eric R. Hein, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank
ACM Trans. Archit. Code Optim.2