Karim Soliman

dblp:146/9602 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-5296-9743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Reactive Deadlock Avoidance Based on Focus Routing Graph Classification for Triplet-Based Architecture Network-on-Chip
abstract
The implementation of Network-on-Chip (NoC) architectures presents considerable advantages in performance relative to traditional bus-based systems. However, the sophisticated nature of NoC designs demands careful oversight of shared resources to mitigate potential performance issues. In this context, well-structured routing algorithms are essential, as they facilitate improved traffic management and minimize congestion. Furthermore, mechanisms for deadlock prevention and avoidance are integral to routing algorithms, ensuring continuous packet transmission and ultimately enhancing network performance. This paper introduces a novel reactive deadlock avoidance method for Triplet-Based Architecture Inter-Core NoC (TriBA-cNoC). that classifies routing based on Focus Routing Graph (FRG) size to address routing-level deadlocks caused by the combination of deterministic routing and TriBA-cNoC’s inherent network characteristics. Compared to proactive techniques, it improves downstream buffer utilization and reduces power consumption. Furthermore, two shortest-path routing algorithms are introduced: DM4T-M, which incorporates a round-robin selection mechanism to alleviate congestion on critical paths and minimize hot-node formation. TSR, a novel two-stage distributed routing algorithm, addresses the computational overhead associated with output port selection in previous algorithms. Simulation results obtained using gem5 show that the proposed approach, which integrates routing-level reactive deadlock avoidance with the proposed routing algorithms, yields improvements in latency, throughput, buffer utilization, and power consumption. TSR and DM4T-M achieve latency reductions of up to 34.89% and 28.2%, respectively. Throughput increases of up to 16.94% and 7.81% are observed for TSR and DM4T-M, respectively. Moreover, the proposed approaches enhance buffer utilization by up to 13.9% and 10.44%, while reducing power consumption by up to 9.64%.
Karim Soliman, Chunfeng Li, Feng Shi 0009
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Semi-adaptive distributed approach for triplet-based architecture inter-core communication Network-on-Chip
Karim Soliman, Chunfeng Li, Feng Shi 0009
Integr.1
2025 A High Scalability Memory NoC with Shared-Inside Hierarchical-Groupings for Triplet-Based Many-Core Architecture
abstract
Innovative processor architecture designs are shifting towards Many-Core Architectures (MCAs) to meet the future demands of high-performance computing as the limits of Moore’s Law have almost been reached. Many-core processors utilize shared memory hierarchies to achieve high-speed memory systems, improving memory access efficiency. However, as the number of cores multiplies, the scalability of this system is significantly constrained by the increased proportion of long-distance and Non-Uniform Memory Access (NUMA). Improving the scalability of MCAs is crucial for achieving large/super-scale general-purpose many-core processors. This work proposes a high-scalability memory Network-on-Chip (NoC) for Triplet-Based Many-Core Architecture (TriBA), named TriBA-mNoC. TriBA-mNoC maintains a consistent core-to-core spacing as the network scale increases, effectively preventing increased long-distance memory access latency. Moreover, it leverages an inherent advantage of shared-inside hierarchical-groupings, alleviating common NUMA issues in the NoC design. Evaluations of static network characteristics show that TriBA-mNoC outperforms most classical NoCs in network diameter, average distance, and cost. TriBA-mNoC can be integrated with TriBA in the same silicon die with a tile-like floorplan, forming a novel NoC called TriBA-NoC, which can combine the strengths of both networks to maximize the architecture performance. We evaluated the memory access performance and scalability of TriBA-NoC using the mathematical evaluation models and actual simulations with real traffic (PARSEC 3.0 and SPLASH-2) at different network scales. The mathematical evaluation results indicate that TriBA-NoC achieves an aggregate speedup of approximately 3x compared with 2D-Mesh for a similar number of cores. Furthermore, TriBA-NoC’s single-core speedup efficiency remains stable as the number of cores increases under the same cache hit ratio, whereas 2D-Mesh experiences a rapid decline, highlighting TriBA-NoC’s exceptional scalability. Finally, the actual traffic simulation results show that TriBA-NoC achieves an average memory access latency and time reduction of 25.90% to 40.50% and 5.61% to 31.69%, respectively, compared with 2D-Mesh.
Chunfeng Li, Feng Shi 0009, Karim Soliman
ACM Trans. Archit. Code Optim.4
2024 NxtSPR: A deadlock-free shortest path routing dedicated to relaying for Triplet-Based many-core Architecture
Chunfeng Li, Karim Soliman, Feng Shi 0009
Parallel Comput.2