Rahul Tripathy

dblp:350/7105 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Accurate Analytical Modeling for NoCs with Hybrid Arbitration under High Traffic Injection
abstract
Analytical performance modeling of Networks-on-Chip (NoC) are important for fast design space exploration and quick pre-silicon evaluation. Existing NoC performance analysis techniques assume certain micro-architectural details (e.g., a particular arbitration technique) to be homogeneous across the entire NoC. However, emerging NoC architectures may have hybrid arbitration across the NoC to ensure high throughput. Moreover, existing analytical models estimating performance of NoCs with finite buffers fail to analyze the performance of the NoC accurately under high traffic injection which occur in several modern-day server as well as client applications. In this work, we propose a performance analysis technique for NoCs with hybrid arbitration under high traffic injection. We propose a novel transformation to accurately compute the waiting time of the queues under hybrid arbitration. We also develop a technique to compute the effective arrival statistics to the queues when the desired injection rate is high. Thorough experimental evaluation with a wide range of injection rates at the queues of an industrial NoC show that our proposed analytical model incurs only 7% error on average and 3 orders of speed-up with respect to cycle-accurate simulation under high traffic injection.
Rahul Tripathy, Mohammad Majharul Islam, Riad Akram, Raid Ayoub, Sumit K. Mandal
DATE1
2026 FlexNoC: Fast and Flexible Analysis for NoCs with Arbitrary Topologies and Hybrid Arbitration
abstract
Performance analysis of Network-on-Chips(NoC) plays a crucial role in design space exploration of SoCs, but traditional cycle-accurate NoC simulation often limits the ability to explore a large design space efficiently due to their notoriously slow execution. There exist several lightweight performance analysis techniques to reduce the design-space exploration time for NoCs, but all of them lack flexibility. In this work, we present FlexNoC - an end-to-end fast and flexible NoC performance analysis framework based on analytical modeling grounded on queuing theory. FlexNoC considers NoCs with irregular topologies and hybrid arbitration which no existing NoC performance analysis framework considers. We establish a Domain-Specific Language (DSL) to describe an NoC with any topology. Specifically, we extend the DOT language using ANTLR-based grammar to support custom NoC primitives such as injectors, queues, servers, arbiters, sinks, and splits. The DSL enables user-defined network components and their interconnections. The queuing theory based analytical model which is the backbone of the framework incorporates hybrid arbitration, along with finite buffers to accurately capture complex interactions between queues present in any given NoC. FlexNoC is accurate in NoC performance estimation and three orders of magnitude faster than cycle accurate NoC simulation - offering an efficient platform for rapid design space exploration and early-stage NoC performance optimization. Moreover, we demonstrate that FlexNoC, through rapid design space exploration, unlocks new insights regarding arbitration techniques at router ports.
Anuparna Ganguly, Rahul Tripathy, Tyson Loveless, Mohammad Majharul Islam, Sumit K. Mandal
ISPASS2
2025 Interconnect Performance Estimation for ML Accelerators via Lightweight Analytical Model
abstract
Machine learning (ML) algorithms are being used in real time in various applications now-a-days. Since state-of-theart high performing ML algorithms are computationally intensive, there exist accelerators which reduce latency and/or improve energy-efficiency of computer systems executing ML algorithms. Performance estimation frameworks have been constructed to perform extensive design space exploration while designing these ML data flow accelerators. While the performance estimation of compute elements of data flow accelerators consists of high-level models, the performance estimation of communication elements of the accelerators still relies on cycle accurate simulations. However, cycle accurate simulations of the interconnect (communication elements) are notoriously slow and it increases the execution time of the performance estimation frameworks, hindering design space exploration. Existing analytical model-based interconnect performance estimation techniques are not applicable for the interconnects of data flow accelerators since they exhibit specific communication patterns. Therefore, we develop an end-to-end framework to estimate interconnect performance for ML data flow accelerators. Extensive experimental evaluations on different ML algorithms show that our proposed framework estimates the interconnect performance with less than 5% error.
Rahul Tripathy, Sumit K. Mandal
ISPASS1