EDBT 2026 Demo / reviewers in the wild / expert
Shreesha Srinath
dblp:39/7863
· DBLP profile ↗
12ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
Electronic design automation · 44% Processor architecture and microarchitecture · 16% Parallel and multicore computing · 14% | |
| Computer networks
2 papers |
Physical-layer communications · 50% Wireless networking · 31% Content delivery and video streaming · 19% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.9 | 3 | 2018 | A modular digital VLSI flow for high-productivity SoC design · DAC 2018 Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017 Improving high-level synthesis with decoupled data structure optimization · DAC 2016 |
Electronic design automation › hardware/software co-design
co-simulation |
0.5 | 1 | 2021 | Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021 |
Electronic design automation › hardware verification and test
hardware verification |
0.5 | 1 | 2021 | Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021 |
Electronic design automation › hardware verification and test
processor verification |
0.5 | 1 | 2021 | Effective Processor Verification with Logic Fuzzer Enhanced Co-simulation · MICRO 2021 |
GPUs and heterogeneous computing › GPU programming
dynamic parallelism |
0.3 | 1 | 2018 | An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 1 | 2018 | An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018 |
Electronic design automation
physical design |
0.3 | 1 | 2018 | A modular digital VLSI flow for high-productivity SoC design · DAC 2018 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.3 | 1 | 2018 | An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018 |
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing |
0.3 | 1 | 2018 | An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable Hardware · MICRO 2018 |
Processor architecture and microarchitecture
chip multiprocessor |
0.3 | 1 | 2017 | Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017 |
Processor architecture and microarchitecture
pipelining |
0.3 | 1 | 2017 | Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017 |
Parallel and multicore computing › parallel programming models
task-based programming |
0.3 | 1 | 2017 | Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017 |
Physical-layer communications › channel coding › error control coding
unequal error protection |
0.3 | 2 | 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013 Design and implementation of an "approximate" communication system for wireless media applications · SIGCOMM 2010 |
Reconfigurable computing and FPGAs
latency-insensitive interface |
0.2 | 1 | 2016 | Improving high-level synthesis with decoupled data structure optimization · DAC 2016 |
Compilers and program optimization
loop optimization |
0.2 | 1 | 2014 | Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014 |
Processor architecture and microarchitecture › adaptive architecture
adaptive microarchitecture |
0.2 | 1 | 2014 | Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2014 | Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014 |
Physical-layer communications
error protection |
0.2 | 1 | 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013 |
Content delivery and video streaming
multimedia delivery |
0.2 | 1 | 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013 |
Wireless networking
wireless network protocols |
0.2 | 1 | 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013 |
GPUs and heterogeneous computing › GPU microarchitecture
SIMT architecture |
0.2 | 1 | 2013 | Microarchitectural mechanisms to exploit value structure in SIMT architectures · ISCA 2013 |
Embedded and real-time systems › model-based design
code generation |
0.1 | 1 | 2010 | Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010 |
Reconfigurable computing and FPGAs
FPGA arithmetic |
0.1 | 1 | 2010 | Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010 |
Electronic design automation › high-level synthesis
hardware generation |
0.1 | 1 | 2010 | Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010 |
Integrated circuit design › digital circuit design › arithmetic circuit design
multiplier design |
0.1 | 1 | 2010 | Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocks · FPGA 2010 |
Integrated circuit design
system-on-chip |
0.1 | 1 | 2018 | A modular digital VLSI flow for high-productivity SoC design · DAC 2018 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2017 | Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis · FPGA 2017 |
Multimedia systems and quality of experience
video quality assessment |
0.1 | 2 | 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media Applications · IEEE/ACM Trans. Netw. 2013 Design and implementation of an "approximate" communication system for wireless media applications · SIGCOMM 2010 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 1 | 2014 | Architectural Specialization for Inter-Iteration Loop Dependence Patterns · MICRO 2014 |
Hardware accelerators and domain-specific architectures
data-parallel accelerator |
0.0 | 1 | 2013 | Microarchitectural mechanisms to exploit value structure in SIMT architectures · ISCA 2013 |
Methods — techniques the papers use, named apart from their topics
software-defined radio · 0.5compiler transformation · 0.4task-based computation model · 0.3systemc · 0.3globally asynchronous locally synchronous clocking · 0.3continuation passing · 0.3microarchitectural template · 0.3instruction set hint · 0.3hazard resolution · 0.3out-of-order execution · 0.2decoupled data structure · 0.2simulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Effective Processor Verification with Logic Fuzzer Enhanced Co-simulationabstractThe study on verification trends in the semiconductor industry shows that the design complexity is increasing, fewer companies achieve first silicon success and need more spins before production, companies hire more verification engineers, and 53% of the whole hardware-design-cycle is spent on the design verification [18]. The cost of a respin is high, and more than 40% of the cases that contribute to it are post-fabrication functional bug exposures [16]. The study also shows that 65% of verification engineers’ time is spent on debug, test creation, and simulation [17]. This paper presents a set of tools for RISC-V processor verification engineers that help to expose more bugs before production and increase the productivity of time spent on debugging, test creation and simulation. We present Logic Fuzzer (LF), a novel tool that expands the verification space exploration without the creation of additional verification tests. The LF randomizes the states or control signals of the design-under-test at the places that do not affect functionality. It brings the processor execution outside its normal flow to increase the number of microarchitectural states exercised by the tests. We also present Dromajo, the state of the art processor verification framework for RISC-V cores. Dromajo is an RV64GC emulator that was designed specifically for co-simulation purposes. It can boot Linux, handle external stimuli, such as interrupts and debug requests on the fly, and can be integrated into existing testbench infrastructure with minimal effort. We evaluate the effectiveness of the tools on three RISC-V cores: CVA6, BlackParrot, and BOOM. Dromajo by itself found a total of nine bugs. The enhancement of Dromajo with the Logic Fuzzer increases the exposed bug count to thirteen without creating additional verification tests. Nursultan Kabylkas, Tommy Thorn, Shreesha Srinath, Polychronis Xekalakis, Jose Renau |
MICRO | 3 |
| 2018 | A modular digital VLSI flow for high-productivity SoC designabstractA high-productivity digital VLSI flow for designing complex SoCs is presented. The flow includes high-level synthesis tools, an object-oriented library of synthesizable SystemC and C++ components, and a modular VLSI physical design approach based on fine-grained globally asynchronous locally synchronous (GALS) clocking. The flow was demonstrated on a 16nm FinFET testchip targeting machine learning and computer vision. Brucek Khailany, Evgeni Khmer, Rangharajan Venkatesan, Jason Clemons, Joel S. Emer, Matthew Fojtik, Alicia Klinefelter, Michael Pellauer, Nathaniel Ross Pinckney, Sophia Shao, Shreesha Srinath, Christopher Torng, Sam Likun Xi, Yanqing Zhang 0002, Brian Zimmer |
DAC | 11 |
| 2018 | An Architectural Framework for Accelerating Dynamic Parallel Algorithms on Reconfigurable HardwareabstractIn this paper, we propose ParallelXL, an architectural framework for building application-specific parallel accelerators with low manual effort. The framework introduces a task-based computation model with explicit continuation passing to support dynamic parallelism in addition to static parallelism. In contrast, today's high-level design frameworks for accelerators focus on static data-level or thread-level parallelism that can be identified and scheduled at design time. To realize the new computation model, we develop an accelerator architecture that efficiently handles dynamic task generation and scheduling as well as load balancing through work stealing. The architecture is general enough to support many dynamic parallel constructs such as fork-join, data-dependent task spawning, and arbitrary nesting and recursion of tasks, as well as static parallel patterns. We also introduce a design methodology that includes an architectural template that allows easily creating parallel accelerators from high-level descriptions. The proposed framework is studied through an FPGA prototype as well as detailed simulations. Evaluation results show that the framework can generate high-performance accelerators targeting FPGAs for a wide range of parallel algorithms and achieve an average of 4.0x speedup over an eight-core out-of-order processor (24.1x over a single core), while being 11.8x more energy efficient. Tao Chen 0045, Shreesha Srinath, Christopher Batten, G. Edward Suh |
MICRO | 2 |
| 2017 | Dynamic Hazard Resolution for Pipelining Irregular Loops in High-Level Synthesis
Steve Dai, Ritchie Zhao, Gai Liu, Shreesha Srinath, Udit Gupta 0001, Christopher Batten, Zhiru Zhang |
FPGA | 4 |
| 2017 | Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programsabstractTask-based parallel programming frameworks offer compelling productivity and performance benefits for modern chip multi-processors (CMPs). At the same time, CMPs also provide packed-SIMD units to exploit fine-grain data parallelism. Two fundamental challenges make using packed-SIMD units with task-parallel programs particularly difficult: (1) the intra-core parallel abstraction gap; and (2) inefficient execution of irregular tasks. To address these challenges, we propose augmenting CMPs with intra-core loop-task accelerators (LTAs). We introduce a lightweight hint in the instruction set to elegantly encode loop-task execution and an LTA microarchitectural template that can be configured at design time for different amounts of spatial/temporal decoupling to efficiently execute both regular and irregular loop tasks. Compared to an in-order CMP baseline, CMP+LTA results in an average speedup of 4.2X (1.8X area normalized) and similar energy efficiency. Compared to an out-of-order CMP baseline, CMP+LTA results in an average speedup of 2.3X (1.5X area normalized) and also improves energy efficiency by 3.2X. Our work suggests augmenting CMPs with lightweight LTAs can improve performance and efficiency on both regular and irregular loop-task parallel programs with minimal software changes. Ji Kim, Shunning Jiang, Christopher Torng, Moyang Wang, Shreesha Srinath, Berkin Ilbeyi, Khalid Al-Hawaj, Christopher Batten |
MICRO | 5 |
| 2016 | Improving high-level synthesis with decoupled data structure optimizationabstractExisting high-level synthesis (HLS) tools are mostly effective on algorithm-dominated programs that only use primitive data structures such as fixed size arrays and queues. However, many widely used data structures such as priority queues, heaps, and trees feature complex member methods with data-dependent work and irregular memory access patterns. These methods can be inlined to their call sites, but this does not address the aforementioned issues and may further complicate conventional HLS optimizations, resulting in a low-performance hardware implementation. To overcome this deficiency, we propose a novel HLS architectural template in which complex data structures are decoupled from the algorithm using a latency-insensitive interface. This enables overlapped execution of the algorithm and data structure methods, as well as parallel and out-of-order execution of independent methods on multiple decoupled lanes. Experimental results across a variety of real-life benchmarks show our approach is capable of achieving very promising speedups without causing significant area overhead. Ritchie Zhao, Gai Liu, Shreesha Srinath, Christopher Batten, Zhiru Zhang |
DAC | 3 |
| 2016 | Experiences using a novel Python-based hardware modeling framework for computer architecture test chipsabstractThis poster will describe a taped-out 2×2mm 1.3 M-transistor test chip in IBM 130 nm designed using our new Python-based hardware modeling framework. The goal of our tapeout was to demonstrate the ability of this framework to enable Agile hardware design flows. Christopher Torng, Moyang Wang, Bharath Sudheendra, Nagaraj Murali, Suren Jayasuriya, Shreesha Srinath, Taylor Pritchard, Robin Ying, Christopher Batten |
Hot Chips Symposium | 6 |
| 2014 | Architectural Specialization for Inter-Iteration Loop Dependence PatternsabstractHardware specialization is an increasingly common technique to enable improved performance and energy efficiency in spite of the diminished benefits of technology scaling. This paper proposes a new approach called explicit loop specialization (XLOOPS) based on the idea of elegantly encoding inter-iteration loop dependence patterns in the instruction set. XLOOPS supports a variety of inter-iteration data-and control-dependence patterns for both single and nested loops. The XLOOPS hardware/software abstraction requires only lightweight changes to a general-purpose compiler to generate XLOOPS binaries and enables executing these binaries on: (1) traditional micro architectures with minimal performance impact, (2) specialized micro architectures to improve performance and/or energy efficiency, and (3) adaptive micro architectures that can seamlessly migrate loops between traditional and specialized execution to dynamically trade-off performance vs. Energy efficiency. We evaluate XLOOPS using a vertically integrated research methodology and show compelling performance and energy efficiency improvements compared to both simple and complex general-purpose processors. Shreesha Srinath, Berkin Ilbeyi, Mingxing Tan, Gai Liu, Zhiru Zhang, Christopher Batten |
MICRO | 1 |
| 2013 | Microarchitectural mechanisms to exploit value structure in SIMT architecturesabstractSIMT architectures improve performance and efficiency by exploiting control and memory-access structure across data-parallel threads. Value structure occurs when multiple threads operate on values that can be compactly encoded, e.g., by using a simple function of the thread index. We characterize the availability of control, memory-access, and value structure in typical kernels and observe ample amounts of value structure that is largely ignored by current SIMT architectures. We propose three microarchitectural mechanisms to exploit value structure based on compact affine execution of arithmetic, branch, and memory instructions. We explore these mechanisms within the context of traditional SIMT microarchitectures (GP-SIMT), found in general-purpose graphics processing units, as well as fine-grain SIMT microarchitectures (FG-SIMT), a SIMT variant appropriate for compute-focused data-parallel accelerators. Cycle-level modeling of a modern GP-SIMT system and a VLSI implementation of an eight-lane FG-SIMT execution engine are used to evaluate a range of application kernels. When compared to a baseline without compact affine execution, our approach can improve GP-SIMT cycle-level performance by 4-17% and can improve FG-SIMT absolute performance by 20-65% and energy efficiency up to 30% for a majority of the kernels. Ji Kim, Christopher Torng, Shreesha Srinath, Derek Lockhart, Christopher Batten |
ISCA | 3 |
| 2013 | Design and Implementation of an "Approximate" Communication System for Wireless Media ApplicationsabstractAll practical wireless communication systems are prone to errors. At the symbol level, such wireless errors have a well-defined structure: When a receiver decodes a symbol erroneously, it is more likely that the decoded symbol is a good “approximation” of the transmitted symbol than a randomly chosen symbol among all possible transmitted symbols. Based on this property, we define approximate communication, a method that exploits this error structure to natively provide unequal error protection to data bits. Unlike traditional [forward error correction (FEC)-based] mechanisms of unequal error protection that consume additional network and spectrum resources to encode redundant data, the approximate communication technique achieves this property at the PHY layer without consuming any additional network or spectrum resources (apart from a minimal signaling overhead). Approximate communication is particularly useful to media delivery applications that can benefit significantly from unequal error protection of data bits. We show the usefulness of this method to such applications by designing and implementing an end-to-end media delivery system, called Apex. Our Software Defined Radio (SDR)-based experiments reveal that Apex can improve video quality by 5–20 dB [peak signal-to-noise ratio (PSNR)] across a diverse set of wireless conditions when compared to traditional approaches. We believe that mechanisms such as Apex can be a cornerstone in designing future wireless media delivery systems under any error-prone channel condition. Sayandeep Sen, Tan Zhang, Syed Gilani, Shreesha Srinath, Suman Banerjee 0001, Sateesh Addepalli |
IEEE/ACM Trans. Netw. | 4 |
| 2010 | Automatic generation of high-performance multipliers for FPGAs with asymmetric multiplier blocksabstractThe introduction of asymmetric embedded multiplier blocks in recent Xilinx FPGAs complicates the design of larger multiplier sizes. The two different input bitwidths of the embedded multipliers lead to two different shifting factors for the partial products that must be summed. This makes even the most straightforward multiplier design less intuitive. In this thesis, I present a methodology and set of equations to automatically generate Verilog hardware description code for arbitrary multiplier sizes composed of arbitrarily-sized asymmetric embedded multiplier cores. The presented technique also uses intelligent rearrangement of the multiplier block outputs into partial product terms to reduce the overall delay of the circuit. Multipliers created with this generator are faster and use fewer DSP blocks than either those created using Xilinx Core Generator or those created by simply using the ?*? operator in Verilog. It also uses fewer LUTs than those created using the ?*? operator. Finally, the presented generator can create multipliers larger than possible with Core Generator, and is limited only by the number of available embedded multipliers. Shreesha Srinath, Katherine Compton |
FPGA | 1 |
| 2010 | Design and implementation of an "approximate" communication system for wireless media applicationsabstractAll practical wireless communication systems are prone to errors. At the symbol level such wireless errors have a well-defined structure: when a receiver decodes a symbol erroneously, it is more likely that the decoded symbol is a good "approximation" of the transmitted symbol than a randomly chosen symbol among all possible transmitted symbols. Based on this property, we define approximate communication, a method that exploits this error structure to natively provide unequal error protection to data bits. Unlike traditional (FEC-based) mechanisms of unequal error protection that consumes additional network and spectrum resources to encode redundant data, the approximate communication technique achieves this property at the PHY layer without consuming any additional network or spectrum resources (apart from a minimal signaling overhead). Approximate communication is particularly useful to media delivery applications that can benefit significantly from unequal error protection of data bits. We show the usefulness of this method to such applications by designing and implementing an end-to-end media delivery system, called Apex. Our Software Defined Radio (SDR)-based experiments reveal that Apex can improve video quality by 5 to 20 dB (PSNR) across a diverse set of wireless conditions, when compared to traditional approaches. We believe that mechanisms such as Apex can be a cornerstone in designing future wireless media delivery systems under any error-prone channel condition. Sayandeep Sen, Syed Gilani, Shreesha Srinath, Stephen Schmitt, Suman Banerjee 0001 |
SIGCOMM | 3 |