Shreesha Srinath

dblp:39/7863 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Electronic design automation · 44% Processor architecture and microarchitecture · 16% Parallel and multicore computing · 14%
Computer networks
2 papers
Physical-layer communications · 50% Wireless networking · 31% Content delivery and video streaming · 19%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
0.932018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Improving high-level synthesis with decoupled data structure optimization · DAC 2016
Electronic design automation › hardware/software co-design
co-simulation
0.512021
Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021
Electronic design automation › hardware verification and test
hardware verification
0.512021
Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021
Electronic design automation › hardware verification and test
processor verification
0.512021
Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021
GPUs and heterogeneous computing › GPU programming
dynamic parallelism
0.312018
An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018
Reconfigurable computing and FPGAs
FPGA accelerator
0.312018
An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018
Electronic design automation
physical design
0.312018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Parallel and multicore computing › parallel programming models
task parallelism
0.312018
An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing
0.312018
An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018
Processor architecture and microarchitecture
chip multiprocessor
0.312017
Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017
Processor architecture and microarchitecture
pipelining
0.312017
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Parallel and multicore computing › parallel programming models
task-based programming
0.312017
Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017
Physical-layer communications › channel coding › error control coding
unequal error protection
0.322013
Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013
Design and implementation of an "approximate" communication system for wireless media applications · SIGCOMM 2010
Reconfigurable computing and FPGAs
latency-insensitive interface
0.212016
Improving high-level synthesis with decoupled data structure optimization · DAC 2016
Compilers and program optimization
loop optimization
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Processor architecture and microarchitecture › adaptive architecture
adaptive microarchitecture
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Processor architecture and microarchitecture
instruction set architecture
0.212014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Physical-layer communications
error protection
0.212013
Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013
Content delivery and video streaming
multimedia delivery
0.212013
Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013
Wireless networking
wireless network protocols
0.212013
Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013
GPUs and heterogeneous computing › GPU microarchitecture
SIMT architecture
0.212013
Microarchitectural mechanisms to exploit value structure in SIMT architectures · ISCA 2013
Embedded and real-time systems › model-based design
code generation
0.112010
Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010
Reconfigurable computing and FPGAs
FPGA arithmetic
0.112010
Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010
Electronic design automation › high-level synthesis
hardware generation
0.112010
Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010
Integrated circuit design › digital circuit design › arithmetic circuit design
multiplier design
0.112010
Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010
Integrated circuit design
system-on-chip
0.112018
A modular digital VLSI flow for high-productivity SoC design · DAC 2018
Electronic design automation › high-level synthesis
scheduling
0.112017
Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017
Multimedia systems and quality of experience
video quality assessment
0.122013
Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013
Design and implementation of an "approximate" communication system for wireless media applications · SIGCOMM 2010
Energy-efficient computing
power-performance tradeoff
0.112014
Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014
Hardware accelerators and domain-specific architectures
data-parallel accelerator
0.012013
Microarchitectural mechanisms to exploit value structure in SIMT architectures · ISCA 2013

Methods — techniques the papers use, named apart from their topics

software-defined radio · 0.5compiler transformation · 0.4task-based computation model · 0.3systemc · 0.3globally asynchronous locally synchronous clocking · 0.3continuation passing · 0.3microarchitectural template · 0.3instruction set hint · 0.3hazard resolution · 0.3out-of-order execution · 0.2decoupled data structure · 0.2simulation · 0.2
YearPublicationVenuePosition
2021 Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation
abstract
The study on verification trends in the semiconductor industry shows that the design complexity is increasing, fewer companies achieve first silicon success and need more spins before production, companies hire more verification engineers, and 53% of the whole hardware-design-cycle is spent on the design verification [18]. The cost of a respin is high, and more than 40% of the cases that contribute to it are post-fabrication functional bug exposures [16]. The study also shows that 65% of verification engineers’ time is spent on debug, test creation, and simulation [17]. This paper presents a set of tools for RISC-V processor verification engineers that help to expose more bugs before production and increase the productivity of time spent on debugging, test creation and simulation. We present Logic Fuzzer (LF), a novel tool that expands the verification space exploration without the creation of additional verification tests. The LF randomizes the states or control signals of the design-under-test at the places that do not affect functionality. It brings the processor execution outside its normal flow to increase the number of microarchitectural states exercised by the tests. We also present Dromajo, the state of the art processor verification framework for RISC-V cores. Dromajo is an RV64GC emulator that was designed specifically for co-simulation purposes. It can boot Linux, handle external stimuli, such as interrupts and debug requests on the fly, and can be integrated into existing testbench infrastructure with minimal effort. We evaluate the effectiveness of the tools on three RISC-V cores: CVA6, BlackParrot, and BOOM. Dromajo by itself found a total of nine bugs. The enhancement of Dromajo with the Logic Fuzzer increases the exposed bug count to thirteen without creating additional verification tests.
Nursultan Kabylkas, Tommy Thorn, Shreesha Srinath, Polychronis Xekalakis, Jose Renau
MICRO3
2018 A modular digital VLSI flow for high-productivity SoC design
abstract
A high-productivity digital VLSI flow for designing complex SoCs is presented. The flow includes high-level synthesis tools, an object-oriented library of synthesizable SystemC and C++ components, and a modular VLSI physical design approach based on fine-grained globally asynchronous locally synchronous (GALS) clocking. The flow was demonstrated on a 16nm FinFET testchip targeting machine learning and computer vision.
Brucek Khailany, Evgeni Khmer, Rangharajan Venkatesan, Jason Clemons, Joel S. Emer, Matthew Fojtik, Alicia Klinefelter, Michael Pellauer, Nathaniel Ross Pinckney, Sophia Shao, Shreesha Srinath, Christopher Torng, Sam Likun Xi, Yanqing Zhang 0002, Brian Zimmer
DAC11
2018 An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware
abstract
In this paper, we propose ParallelXL, an architectural framework for building application-specific parallel accelerators with low manual effort. The framework introduces a task-based computation model with explicit continuation passing to support dynamic parallelism in addition to static parallelism. In contrast, today's high-level design frameworks for accelerators focus on static data-level or thread-level parallelism that can be identified and scheduled at design time. To realize the new computation model, we develop an accelerator architecture that efficiently handles dynamic task generation and scheduling as well as load balancing through work stealing. The architecture is general enough to support many dynamic parallel constructs such as fork-join, data-dependent task spawning, and arbitrary nesting and recursion of tasks, as well as static parallel patterns. We also introduce a design methodology that includes an architectural template that allows easily creating parallel accelerators from high-level descriptions. The proposed framework is studied through an FPGA prototype as well as detailed simulations. Evaluation results show that the framework can generate high-performance accelerators targeting FPGAs for a wide range of parallel algorithms and achieve an average of 4.0x speedup over an eight-core out-of-order processor (24.1x over a single core), while being 11.8x more energy efficient.
Tao Chen 0045, Shreesha Srinath, Christopher Batten, G. Edward Suh
MICRO2
2017 Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis
Steve Dai, Ritchie Zhao, Gai Liu, Shreesha Srinath, Udit Gupta 0001, Christopher Batten, Zhiru Zhang
FPGA4
2017 Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs
abstract
Task-based parallel programming frameworks offer compelling productivity and performance benefits for modern chip multi-processors (CMPs). At the same time, CMPs also provide packed-SIMD units to exploit fine-grain data parallelism. Two fundamental challenges make using packed-SIMD units with task-parallel programs particularly difficult: (1) the intra-core parallel abstraction gap; and (2) inefficient execution of irregular tasks. To address these challenges, we propose augmenting CMPs with intra-core loop-task accelerators (LTAs). We introduce a lightweight hint in the instruction set to elegantly encode loop-task execution and an LTA microarchitectural template that can be configured at design time for different amounts of spatial/temporal decoupling to efficiently execute both regular and irregular loop tasks. Compared to an in-order CMP baseline, CMP+LTA results in an average speedup of 4.2X (1.8X area normalized) and similar energy efficiency. Compared to an out-of-order CMP baseline, CMP+LTA results in an average speedup of 2.3X (1.5X area normalized) and also improves energy efficiency by 3.2X. Our work suggests augmenting CMPs with lightweight LTAs can improve performance and efficiency on both regular and irregular loop-task parallel programs with minimal software changes.
Ji Kim, Shunning Jiang, Christopher Torng, Moyang Wang, Shreesha Srinath, Berkin Ilbeyi, Khalid Al-Hawaj, Christopher Batten
MICRO5
2016 Improving high-level synthesis with decoupled data structure optimization
abstract
Existing high-level synthesis (HLS) tools are mostly effective on algorithm-dominated programs that only use primitive data structures such as fixed size arrays and queues. However, many widely used data structures such as priority queues, heaps, and trees feature complex member methods with data-dependent work and irregular memory access patterns. These methods can be inlined to their call sites, but this does not address the aforementioned issues and may further complicate conventional HLS optimizations, resulting in a low-performance hardware implementation. To overcome this deficiency, we propose a novel HLS architectural template in which complex data structures are decoupled from the algorithm using a latency-insensitive interface. This enables overlapped execution of the algorithm and data structure methods, as well as parallel and out-of-order execution of independent methods on multiple decoupled lanes. Experimental results across a variety of real-life benchmarks show our approach is capable of achieving very promising speedups without causing significant area overhead.
Ritchie Zhao, Gai Liu, Shreesha Srinath, Christopher Batten, Zhiru Zhang
DAC3
2016 Experiences using a novel Python-based hardware modeling framework for computer architecture test chips
abstract
This poster will describe a taped-out 2×2mm 1.3 M-transistor test chip in IBM 130 nm designed using our new Python-based hardware modeling framework. The goal of our tapeout was to demonstrate the ability of this framework to enable Agile hardware design flows.
Christopher Torng, Moyang Wang, Bharath Sudheendra, Nagaraj Murali, Suren Jayasuriya, Shreesha Srinath, Taylor Pritchard, Robin Ying, Christopher Batten
Hot Chips Symposium6
2014 Architectural Specialization for Inter-Iteration Loop Dependence Patterns
abstract
Hardware specialization is an increasingly common technique to enable improved performance and energy efficiency in spite of the diminished benefits of technology scaling. This paper proposes a new approach called explicit loop specialization (XLOOPS) based on the idea of elegantly encoding inter-iteration loop dependence patterns in the instruction set. XLOOPS supports a variety of inter-iteration data-and control-dependence patterns for both single and nested loops. The XLOOPS hardware/software abstraction requires only lightweight changes to a general-purpose compiler to generate XLOOPS binaries and enables executing these binaries on: (1) traditional micro architectures with minimal performance impact, (2) specialized micro architectures to improve performance and/or energy efficiency, and (3) adaptive micro architectures that can seamlessly migrate loops between traditional and specialized execution to dynamically trade-off performance vs. Energy efficiency. We evaluate XLOOPS using a vertically integrated research methodology and show compelling performance and energy efficiency improvements compared to both simple and complex general-purpose processors.
Shreesha Srinath, Berkin Ilbeyi, Mingxing Tan, Gai Liu, Zhiru Zhang, Christopher Batten
MICRO1
2013 Microarchitectural mechanisms to exploit value structure in SIMT architectures
abstract
SIMT architectures improve performance and efficiency by exploiting control and memory-access structure across data-parallel threads. Value structure occurs when multiple threads operate on values that can be compactly encoded, e.g., by using a simple function of the thread index. We characterize the availability of control, memory-access, and value structure in typical kernels and observe ample amounts of value structure that is largely ignored by current SIMT architectures. We propose three microarchitectural mechanisms to exploit value structure based on compact affine execution of arithmetic, branch, and memory instructions. We explore these mechanisms within the context of traditional SIMT microarchitectures (GP-SIMT), found in general-purpose graphics processing units, as well as fine-grain SIMT microarchitectures (FG-SIMT), a SIMT variant appropriate for compute-focused data-parallel accelerators. Cycle-level modeling of a modern GP-SIMT system and a VLSI implementation of an eight-lane FG-SIMT execution engine are used to evaluate a range of application kernels. When compared to a baseline without compact affine execution, our approach can improve GP-SIMT cycle-level performance by 4-17% and can improve FG-SIMT absolute performance by 20-65% and energy efficiency up to 30% for a majority of the kernels.
Ji Kim, Christopher Torng, Shreesha Srinath, Derek Lockhart, Christopher Batten
ISCA3
2013 Design and Implementation of an "Approximate" Communication System for Wireless Media Applications
abstract
All practical wireless communication systems are prone to errors. At the symbol level, such wireless errors have a well-defined structure: When a receiver decodes a symbol erroneously, it is more likely that the decoded symbol is a good “approximation” of the transmitted symbol than a randomly chosen symbol among all possible transmitted symbols. Based on this property, we define approximate communication, a method that exploits this error structure to natively provide unequal error protection to data bits. Unlike traditional [forward error correction (FEC)-based] mechanisms of unequal error protection that consume additional network and spectrum resources to encode redundant data, the approximate communication technique achieves this property at the PHY layer without consuming any additional network or spectrum resources (apart from a minimal signaling overhead). Approximate communication is particularly useful to media delivery applications that can benefit significantly from unequal error protection of data bits. We show the usefulness of this method to such applications by designing and implementing an end-to-end media delivery system, called Apex. Our Software Defined Radio (SDR)-based experiments reveal that Apex can improve video quality by 5–20 dB [peak signal-to-noise ratio (PSNR)] across a diverse set of wireless conditions when compared to traditional approaches. We believe that mechanisms such as Apex can be a cornerstone in designing future wireless media delivery systems under any error-prone channel condition.
Sayandeep Sen, Tan Zhang, Syed Gilani, Shreesha Srinath, Suman Banerjee 0001, Sateesh Addepalli
IEEE/ACM Trans. Netw.4
2010 Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks
abstract
The introduction of asymmetric embedded multiplier blocks in recent Xilinx FPGAs complicates the design of larger multiplier sizes. The two different input bitwidths of the embedded multipliers lead to two different shifting factors for the partial products that must be summed. This makes even the most straightforward multiplier design less intuitive. In this thesis, I present a methodology and set of equations to automatically generate Verilog hardware description code for arbitrary multiplier sizes composed of arbitrarily-sized asymmetric embedded multiplier cores. The presented technique also uses intelligent rearrangement of the multiplier block outputs into partial product terms to reduce the overall delay of the circuit. Multipliers created with this generator are faster and use fewer DSP blocks than either those created using Xilinx Core Generator or those created by simply using the ?*? operator in Verilog. It also uses fewer LUTs than those created using the ?*? operator. Finally, the presented generator can create multipliers larger than possible with Core Generator, and is limited only by the number of available embedded multipliers.
Shreesha Srinath, Katherine Compton
FPGA1
2010 Design and implementation of an "approximate" communication system for wireless media applications
abstract
All practical wireless communication systems are prone to errors. At the symbol level such wireless errors have a well-defined structure: when a receiver decodes a symbol erroneously, it is more likely that the decoded symbol is a good "approximation" of the transmitted symbol than a randomly chosen symbol among all possible transmitted symbols. Based on this property, we define approximate communication, a method that exploits this error structure to natively provide unequal error protection to data bits. Unlike traditional (FEC-based) mechanisms of unequal error protection that consumes additional network and spectrum resources to encode redundant data, the approximate communication technique achieves this property at the PHY layer without consuming any additional network or spectrum resources (apart from a minimal signaling overhead). Approximate communication is particularly useful to media delivery applications that can benefit significantly from unequal error protection of data bits. We show the usefulness of this method to such applications by designing and implementing an end-to-end media delivery system, called Apex. Our Software Defined Radio (SDR)-based experiments reveal that Apex can improve video quality by 5 to 20 dB (PSNR) across a diverse set of wireless conditions, when compared to traditional approaches. We believe that mechanisms such as Apex can be a cornerstone in designing future wireless media delivery systems under any error-prone channel condition.
Sayandeep Sen, Syed Gilani, Shreesha Srinath, Stephen Schmitt, Suman Banerjee 0001
SIGCOMM3