Gade Narayana Sri Harsha

dblp:151/4523 · also Sri Harsha Gade · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
4since 2021 · last 2022
0000-0002-8799-2932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2022 A Novel Hybrid Cache Coherence with Global Snooping for Many-core Architectures
abstract
Cache coherence ensures correctness of cached data in multi-core processors. Traditional implementations of existing protocols make them unscalable for many core architectures. While snoopy coherence requires unscalable ordered networks, directory coherence is weighed down by high area and energy overheads. In this work, we propose Wireless-enabled Share-aware Hybrid (WiSH) to provide scalable coherence in many core processors. WiSH implements a novel Snoopy over Directory protocol using on-chip wireless links and hierarchical, clustered Network-on-Chip to achieve low-overhead and highly efficient coherence. A local directory protocol maintains coherence within a cluster of cores, while coherence among such clusters is achieved through global snoopy protocol. The ordered network for global snooping is provided through low-latency and low-energy broadcast wireless links. The overheads are further reduced through share-aware cache segmentation to eliminate coherence for private blocks. Evaluations show that WiSH reduces traffic by and runtime by , while requiring smaller storage and lower energy as compared to existing hierarchical and hybrid coherence protocols. Owing to its modularity, WiSH provides highly efficient and scalable coherence for many core processors.
Gade Narayana Sri Harsha, Sujay Deb
ACM Trans. Design Autom. Electr. Syst.1
2021 Topology Agnostic Virtual Channel Assignment and Protocol Level Deadlock Avoidance in a Network-on-Chip
abstract
A Virtual Channel (VC) is a Time Division Multiplexed (TDM) slice of a physical channel/link. A crucial step of interconnect synthesis is to assign VCs to traffic that avoids deadlocks while meeting Power, Performance and Area (PPA) objectives. For a Network on Chip (NoC), VC assignment is tightly coupled with topology generation and routing. Inefficient VC assignment can lead to NoCs which may be an order of magnitude inferior in terms of PPA. However, both VC assignment and topology generation are combinatorially hard problems. Thus, combining VC assignment with topology generation makes it difficult to efficiently solve either.In this paper, we present a topology agnostic VC assignment approach for statically routed NoCs. This segregation enables us to solve the VC assignment problem first and then subsequently generate topology. By solving VC assignment first, we reduce the problem space of topology generation leading to NoCs with tighter PPA. We model the VC assignment problem as a Traffic Conflict Graph (TCG), capturing a global view of Quality of Service (QoS), Head-of-Line (HoL) conflicts, burstiness and external protocol level dependencies (e.g. PCIe root-complex). We apply combinatorial optimization techniques on TCG to arrive at an efficient VC assignment, that ensures deadlock free designs. The algorithm has been implemented in production NoC Synthesis tool. Results obtained on more than a dozen multi-million gate System-on-Chip (SoC) demonstrate an average improvement of 30% across metrics such as routers, resizers, power/clock domain converters, latencies etc. while meeting the performance vis-à-vis hand-tuned NoCs. The hand-tuned designs have been manually optimized over several man-months by expert designers, while the tool achieves better results in a few minutes.
Anup Gangwar, Ravishankar Sreedharan, Ambica Prasad, Nitin Kumar Agarwal 0001, Gade Narayana Sri Harsha
DAC5
2021 An Automated Traffic Generation Framework for Performance Evaluation of Networks-on-Chip for Real World Use Cases
abstract
Networks-on-Chip (NoCs) are fast becoming the defacto interconnection fabric for System-on-Chip (SoC) architectures. Early stage performance modeling is a critical step in NoC design for design space exploration and achieving efficient topologies within Time-To-Market (TTM) demands. In this paper, we propose an automatic traffic generation framework for early and comprehensive analysis of NoCs. Using behavioral input and five independent control parameters, it generates a wide range of spatio-temporal traffic stimuli, including bursty traffic while honoring all design constraints. Generated profiles include use-case scenarios representative of actual SoC traffic and plethora of scenarios for out-of-bound performance analysis. Evaluation on production level SoCs shows that the profiles generated are within 5% margin of actual traffic, enables quick iterations from early stages and shrinks TTM from months to few weeks. Final evaluation using RTL testbenches meets performance in all cases.
Gade Narayana Sri Harsha, Anup Gangwar, Ambica Prasad, Nitin Kumar Agarwal 0001, Ravishankar Sreedharan
ISPASS1
2021 Design Space Optimization of Shared Memory Architecture in Accelerator-rich Systems
abstract
Shared memory architectures, as opposed to private-only memories, provide a viable alternative to meet the ever-increasing memory requirements of multi-accelerator systems to achieve high performance under stringent area and energy constraints. However, an impulsive memory sharing degrades performance due to network contention and latency to access shared memory. We propose the Accelerator Shared Memory (ASM) framework to provide an optimal private/shared memory configuration and shared data allocation under a system’s resource and network constraints. Evaluations show ASM provides up to 34.35% and 31.34% improvement in performance and energy, respectively, over baseline systems.
Mitali Sinha, Gade Narayana Sri Harsha, Pramit Bhattacharyya, Sujay Deb
ACM Trans. Design Autom. Electr. Syst.2
2020 Automated Synthesis of Custom Networks-on-Chip for Real World Applications
abstract
Network-on-Chip (NoC), using a packetized communication model presents a scalable interconnect infrastructure for System-on-Chip (SoC) architectures that meets its Performance, Power and Area (PPA) objectives. A typical NoC consists of building blocks such as routers, resizers and Power and Clock Domain Converters (PCDC). Hand crafting a NoC that meets PPA requirements within Time-to-Market (TTM) constraints is difficult if not intractable for real world systems.
Anup Gangwar, Nitin Kumar Agarwal 0001, Ravishankar Sreedharan, Ambica Prasad, Gade Narayana Sri Harsha
ICCAD5
2019 Millimeter wave wireless interconnects in deep submicron chips: Challenges and opportunities
Gade Narayana Sri Harsha, Shobha Sundar Ram, Sujay Deb
Integr.1
2019 Energy Efficient Chip-to-Chip Wireless Interconnection for Heterogeneous Architectures
abstract
Heterogeneous multichip architectures have gained significant interest in high-performance computing clusters to cater to a wide range of applications. In particular, heterogeneous systems with multiple multicore CPUs, GPUs, and memory have become common to meet application requirements. The shared resources like interconnection network in such systems pose significant challenges due to the diverse traffic requirements of CPUs and GPUs. Especially, the performance and energy consumption of inter-chip communication have remained a major bottleneck due to limitations imposed by off-chip wired links. To overcome these challenges, we propose a wireless interconnection network to provide energy-efficient, high-performance communication in heterogeneous multi-chip systems. Interference-free communication between GPUs and memory modules is achieved through directional wireless links, while omnidirectional wireless interfaces connect cores in the CPUs with other components in the system. Besides providing low-energy, high-bandwidth inter-chip communication, the wireless interconnection scales efficiently with system size to provide high performance across multiple chips. The proposed inter-chip wireless interconnection is evaluated on two system sizes with multiple CPU and multiple GPU chips, along with main memory modules. On a system with 4 CPU and 4 GPU chips, application runtime is sped up by 3.94×, packet energy is reduced by 94.4%, and packet latency is reduced by 58.34% as compared to baseline system with wired inter-chip interconnection.
Gade Narayana Sri Harsha, M. Meraj Ahmed, Sujay Deb, Amlan Ganguly
ACM Trans. Design Autom. Electr. Syst.1
2018 Data-flow Aware CNN Accelerator with Hybrid Wireless Interconnection
abstract
Deep convolution neural networks (CNNs) are computationally intensive machine learning algorithms with a large amount of data that impose various challenges for their hardware implementation. To meet the high computing demands of CNNs, many accelerator designs are proposed that revolve around achieving high parallelization, increasing on-chip data reuse and efficient memory hierarchy, etc. However, very few works have attempted to address the communication challenges in these massively parallel accelerators architectures, which is the most anticipated performance bottleneck. Traditional interconnections like bus, crossbar and even Network-on-Chip ( N o C) topologies like mesh fail to achieve the peak performance required by the large number of processing elements on accelerators. In this work, we address the communication bottlenecks of accelerators by extensively studying the application data-flow. We propose an efficient accelerator architecture that employs broadcast enabled low latency wireless links along with traditional wired links to efficiently support the data-flow of accelerators and achieve high communication performance. Evaluation of the proposed design shows that it achieves 28 % latency reduction, 19x bandwidth improvement and 35% network energy saving as compared to baseline wired networks.
Mitali Sinha, Gade Narayana Sri Harsha, Wazir Singh, Sujay Deb
ASAP2
2018 A Utilization Aware Robust Channel Access Mechanism for Wireless NoCs
abstract
Wireless Network-on-Chip (WNoC) has been proposed to overcome long-distance communication bottlenecks of wired NoCs. Token passing mechanism has generally been adapted to allocate the wireless channel among Wireless Interfaces (WIs). In this work, we propose a comparator based controller to provide a flexible and efficient channel allocation scheme. It utilizes a comparator attached to the antenna, along with modifications to header flit to perform channel allocation along with power gating WIs to save energy. Evaluation of proposed scheme on CPU/GPU system shows 53% reduction in token passes and 9% energy saving as compared to timer based approach.
Gade Narayana Sri Harsha, Sidhartha Sankar Rout, Mitali Sinha, Hemanta Kumar Mondal, Wazir Singh, Sujay Deb
ISCAS1
2018 On-Chip Wireless Channel Propagation: Impact of Antenna Directionality and Placement on Channel Performance
abstract
Long range, low latency wireless links in Networks-on-Chip (NoCs) have been shown to be the most promising solution to provide high performance intra/inter-chip communication in many core era. Significant advancements have been made in design of both Wireless NoC (WNoC) topologies and transceiver circuits to support wireless communication at chip level. However, a comprehensive understanding of wireless physical layer and its impact on performance is still lacking. There is still a lot of scope for thorough analysis of the affects of intra-chip wireless channel and antenna characteristics on signal transmission and link reliability in WNoCs. To this end, we analyse signal propagation through wireless channel by accurately modelling the intra-chip environment. We analyse the effects of antenna placement across chip plane and its directionality on the signal loss, delay and dispersion properties. The analysis shows that directional antenna exhibits better delay characteristics, while omnidirectional antennas have low loss for signal transmission in the channel. Furthermore, the placement of antenna shows considerable impact on channel characteristics due to reflections from chip edges and constructive or destructive interference between the multiple signal components. This work provides crucial insights into propagation characteristics of on-chip wireless links for better design of transceiver components and their performance.
Gade Narayana Sri Harsha, Sidhartha Sankar Rout, Sujay Deb
NOCS1
2017 HyWin: Hybrid Wireless NoC with Sandboxed Sub-Networks for CPU/GPU Architectures
abstract
Heterogeneous System Architectures (HSA) that integrate cores of different architectures (CPU, GPU, etc.) on single chip are gaining significance for many class of applications to achieve high performance. Networks-on-Chip (NoCs) in HSA are monopolized by high volume GPU traffic, penalizing CPU application performance. In addition, building efficient interfaces between systems of different specifications while achieving optimal performance is a demanding task. Homogeneous NoCs, widely used for many core systems, fall short in meeting these communication requirements. To achieve high performance interconnection in HSA, we propose HyWin topology using mm-wave wireless links. The proposed topology implements sandboxed heterogeneous sub-networks, each designed to match needs of a processing subsystem, which are then interconnected at second level using wireless network. The sandboxed sub-networks avoid conflict of network requirements, while providing optimal performance for their respective subsystems. The long range wireless links provide low latency and low energy inter-subsystem network to provide easy access to memory controllers, lower level caches across the entire system. By implementing proposed topology for CPU/GPU HSA, we show that it improves application performance by 29 percent and reduces latency by 50 percent, while reducing energy consumption by 64.5 percent and area by 17.39 percent as compared to baseline mesh.
Gade Narayana Sri Harsha, Sujay Deb
IEEE Trans. Computers1
2017 Adaptive Multi-Voltage Scaling with Utilization Prediction for Energy-Efficient Wireless NoC
abstract
Networks-on-Chip (NoCs) are fast becoming the de-facto communication infrastructures in chip multi-processors for large-scale applications. Wireless NoCs (WNoCs) offer a promising solution to reduce the long-distance communication bottlenecks of conventional NoCs by augmenting them with single hop, long-range wireless links. However, power consumption in routers and network elements still remains considerably high at ultra-deep submicron technologies. Analysis of network resources for several benchmarks shows that, utilization is application dependent and the desired performance can be achieved even without operating all resources at maximum specifications. In this work, we propose an energy-efficient WNoC architecture using Adaptive Multi-Voltage Scaling (AMS) to dynamically vary supply voltage for NoC routers and Wireless Interfaces (WIs) without adversely impacting performance. The proposed scheme uses a probabilistic model to predict router utilization during different application phases and scales voltage accordingly. It further reduces network energy by power-gating WIs that are not engaged in active communication to minimize their power consumption. We present detailed utilization estimation procedure, AMS control mechanism, and its hardware implementation. It saves up to 56 percent in network packet energy consumption and 62.50 percent power consumption in WIs for 256 core system as compared to baseline architectures without incurring significant performance penalty and area overheads.
Hemanta Kumar Mondal, Gade Narayana Sri Harsha, Shashwat Kaushik, Sujay Deb
IEEE Trans. Sustain. Comput.2
2016 Adaptive multi-voltage scaling in wireless NoC for high performance low power applications
Hemanta Kumar Mondal, Gade Narayana Sri Harsha, Raghav Kishore, Sujay Deb
DATE2
2015 Design of signal-matched critically sampled FIR rational filterbank
abstract
Wavelet transform is used for efficient signal analysis in various applications. The traditional wavelet system is implemented using integer decimation factors, although frequency tiling offered by rational decimation may better adapt to signal characteristics. In this paper, we propose a design methodology for signal-matched filterbank (FB) with rational decimation factors that achieves perfect reconstruction with FIR filters. We have applied the proposed design on some real world signals. With the proposed design, we obtain a more compressible transform domain representation than the dyadic standard wavelet transforms.
Anupriya Gogna, Gade Narayana Sri Harsha, Anubha Gupta
ICASSP2
2015 Achievable Performance Enhancements with mm-Wave Wireless Interconnects in NoC
abstract
On-chip wireless links have been shown to overcome the performance limitations of wired interconnects in Networks-on-chip (NoCs). However actual performance gains obtained are largely dependent on efficient data transmission between on-chip antennas. An analysis of on-chip wireless channel shows that propagation is highly affected by different components of the chip. In this work, we include the effects of chip environment on wireless propagation to obtain a more realistic performance evaluation of Wireless NoC (WiNoC). Using these, we derive the latency and energy characteristics of WiNoC and quantify the achievable performance. Results presented show wireless received signal with on-chip effects considered and compare them with that of a wired link.
Gade Narayana Sri Harsha, Sujay Deb
NOCS1