Rahul Boyapati

dblp:03/8857 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0003-1094-485XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Energy-efficient computing · 46% Memory systems · 27% Interconnection networks and networks-on-chip · 16%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing
leakage power reduction
0.512021
A Voting Approach for Adaptive Network-on-Chip Power-Gating · IEEE Trans. Computers 2021
Energy-efficient computing
power gating
0.512021
A Voting Approach for Adaptive Network-on-Chip Power-Gating · IEEE Trans. Computers 2021
Energy-efficient computing
power management
0.512021
A Voting Approach for Adaptive Network-on-Chip Power-Gating · IEEE Trans. Computers 2021
Interconnection networks and networks-on-chip
in-network computing
0.412019
Active-Routing: Compute on the Way for Near-Data Processing · HPCA 2019
Memory systems › processing-in-memory
near-data processing
0.412019
Active-Routing: Compute on the Way for Near-Data Processing · HPCA 2019
Memory systems
processing-in-memory
0.412019
Active-Routing: Compute on the Way for Near-Data Processing · HPCA 2019
Distributed systems › communication optimization
approximate communication
0.312017
APPROX-NoC: A Data Approximation Framework for Network-On-Chip Architectures · ISCA 2017
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing
0.112021
A Voting Approach for Adaptive Network-on-Chip Power-Gating · IEEE Trans. Computers 2021
Memory systems › memory access optimization
memory-level parallelism
0.112019
Active-Routing: Compute on the Way for Near-Data Processing · HPCA 2019
Emerging computing paradigms
approximate computing
0.112017
APPROX-NoC: A Data Approximation Framework for Network-On-Chip Architectures · ISCA 2017

Methods — techniques the papers use, named apart from their topics

voting-based power-gating policy · 0.5synthetic and real workload evaluation · 0.5architecture simulation · 0.4
YearPublicationVenuePosition
2021 A Voting Approach for Adaptive Network-on-Chip Power-Gating
abstract
Scalable Networks-on-Chip (NoCs) have become the standard interconnection mechanisms in large-scale multicore architectures. These NoCs consume a large fraction of the on-chip power budget, where the static portion is becoming dominant as technology scales down to sub-10nm node. Therefore, it is essential to reduce static power so as to achieve power- and energy-efficient computing. Power-Gating as an effective static power saving technique can be used to power off inactive routers for static power saving. However, packet deliveries in irregular power-gated networks suffer from detour or waiting time overhead to either route around or wake up power-gated routers. In this article, we proposeFly-Over (Flov), a voting approach for dynamic router power-gating in a light-weight and distributed manner, which includesFlovrouter microarchitecture, adaptive power-gating policy, and low-latency dynamic routing algorithms. We evaluateFlovusing synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full-system evaluations show thatFlovreduces the power consumption of NoC by 31 and 20 percent, respectively, on average across several benchmarks, compared to the baseline and the state-of-the-art while maintaining the similar performance.
Jiayi Huang 0001, Shilpa Bhosekar, Rahul Boyapati, Byul Hur, Ki Hwan Yum, Eun Jung Kim 0001
IEEE Trans. Computers3
2019 Active-Routing: Compute on the Way for Near-Data Processing
abstract
The explosion of data availability and the demand for faster data analysis have led to the emergence of applications exhibiting large memory footprint and low data reuse rate. These workloads, ranging from neural networks to graph processing, expose compute kernels that operate over myriads of data. Significant data movement requirements of these kernels impose heavy stress on modern memory subsystems and communication fabrics. To mitigate the worsening gap between high CPU computation density and deficient memory bandwidth, solutions like memory networks and near-data processing designs are being architected to improve system performance substantially. In this work, we examine the idea of mapping compute kernels to the memory network so as to leverage in-network computing in data-flow style, by means of near-data processing. We propose Active-Routing, an in-network compute architecture that enables computation on the way for near-data processing by exploiting patterns of aggregation over intermediate results of arithmetic operators. The proposed architecture leverages the massive memory-level parallelism and network concurrency to optimize the aggregation operations along a dynamically built Active-Routing Tree. Our evaluations show that Active-Routing can achieve upto 7X speedup with an average of 60% performance improvement, and reduce the energy-delay product by 80% across various benchmarks compared to the state-of-the-art processing-in-memory architecture.
Jiayi Huang 0001, Ramprakash Reddy Puli, Pritam Majumder, Sungkeun Kim, Rahul Boyapati, Ki Hwan Yum, Eun Jung Kim 0001
HPCA5
2017 Packet coalescing exploiting data redundancy in GPGPU architectures
abstract
General Purpose Graphics Processing Units (GPGPUs) are becoming a cost-effective hardware approach for parallel computing. Many executions on the GPGPUs place heavy stress on the memory system, creating network bottlenecks near memory controllers. We observe that data redundancy in communication traffic is common-place across a wide range of GPGPU applications. To exploit the data redundancy, we propose a packet coalescing mechanism to alleviate the network bottlenecks by directly reducing the traffic volume. The key idea is to coalesce multiple packets into one without increasing the packet size when they carry redundant cache blocks. To ensure that the coalesced packets are delivered to their respective destinations, we adopt multicast routing for the interconnection network of GPGPUs. Our coalescing approach yields 15% IPC improvement (up to 112%) in a large-scale GPGPU with 2D mesh across various GPGPU applications, by reducing average memory access time (AMAT) by 15.5% (up to 65.2%) and obtaining network bandwidth savings by 13% (up to 37%). Also, our coalescing approach achieves 7% IPC improvement in the NVIDIA Fermi architecture with the crossbar.
Kyung Hoon Kim, Rahul Boyapati, Jiayi Huang 0001, Yuho Jin, Ki Hwan Yum, Eun Jung Kim 0001
ICS2
2017 Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-Chip
abstract
Scalable Networks-on-Chip (NoCs) have become the de facto interconnection mechanism in large scale Chip Multiprocessors. Not only are NoCs devouring a large fraction of the on-chip power budget but static NoC power consumption is becoming the dominant component as technology scales down. Hence reducing static NoC power consumption is critical for energy-efficient computing. Previous research has proposed to power-gate routers attached to inactive cores so as to save static power, but requires centralized control and global network knowledge. In this paper, we propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers, which encompasses FLOV router architecture, handshake protocols, and a partition-based dynamic routing algorithm to maintain network functionalities. With simple modifications to the baseline router architecture, FLOV can facilitate FLOV links over power-gated routers. Then we present two handshake protocols for FLOV routers, restricted FLOV that can power-gate routers under restricted conditions and generalized FLOV with more power saving capability. The proposed routing algorithm provides best-effort minimal path routing without the necessity for global network information. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. Our full system evaluations show that FLOV reduces the total and static energy consumption by 18% and 22% respectively, on average across several benchmarks, compared to state-of-the-art NoC power-gating mechanism while keeping the performance degradation minimal.
Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001
IPDPS1
2017 APPROX-NoC: A Data Approximation Framework for Network-On-Chip Architectures
Rahul Boyapati, Jiayi Huang 0001, Pritam Majumder, Ki Hwan Yum, Eun Jung Kim 0001
ISCA1
2016 POSTER: Fly-Over: A Light-Weight Distributed Power-Gating Mechanism For Energy-Efficient Networks-on-Chip
abstract
Reducing static NoC power consumption is becoming critical for energy-efficient computing as technology scales down since NoCs are devouring a large fraction of the on-chip power budget. We propose Fly-Over (FLOV), a light-weight distributed mechanism for power-gating routers. With simple modifications to the baseline router architecture, FLOV links are facilitated over power-gated routers. A Handshake protocol that allows seamless router power-gating in addition to a dynamic routing algorithm, that provides best-effort minimal path without the necessity for global network information, maintain normal NoC functionality. We evaluate our schemes using synthetic workloads as well as real workloads from PARSEC 2.1 benchmark suite. The results show that FLOV can achieve on average 19.2% latency reduction and 15.9% total energy savings.
Rahul Boyapati, Jiayi Huang 0001, Kyung Hoon Kim, Ki Hwan Yum, Eun Jung Kim 0001
PACT1
2010 Efficient lookahead routing and header compression for multicasting in networks-on-chip
abstract
As technology advanced, Chip Multi-processor (CMP) architectures have emerged as a viable solution for designing processors. Networks-on-Chip (NOCs) provide a scalable communication method for CMP architectures as the number of cores is increasing. Although there has been significant research on NOC designs for unicast traffic, the research on the multicast router design is still in infancy stage. Considering that one-to-many (multicast) and one-to-all (broadcast) traffic are more common in CMP applications, it is important to design a router providing efficient multicasting. In this paper, we propose an efficient lookahead routing with limited area overhead for a recently proposed multicast routing algorithm, Recursive Partitioning Multicast (RPM) [17]. Also, we present a novel compression scheme for a multicast packet header that becomes a big overhead in large networks. Comprehensive simulation results show that with our route computation logic design, providing lookahead routing in the multicast router only costs less than 20% area overhead and this percentage keeps decreasing with larger network sizes. Compared with the basic lookahead routing design, our design can save area by over 50%. With header compression and lookahead multicast routing, the network performance is improved by 22% in a (16 x 16) network on average.
Lei Wang 0041, Poornachandran Kumar, Rahul Boyapati, Ki Hwan Yum, Eun Jung Kim 0001
ANCS3