EDBT 2026 Demo / reviewers in the wild / expert
Ke Wang 0030
dblp:181/2613-30
· DBLP profile ↗
22ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-7189-9293ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 17 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning-Assisted Adaptive Hybrid Interconnection Design for Chiplet-Based Heterogeneous Systems
Md. Tareq Mahmud, Ke Wang 0030, Ahmed Louri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | ReSafe-ViT: A Reconfigurable and Reliable Accelerator for Vision Transformer InferenceabstractVision Transformers (ViTs) have gained significant traction in various vision-related tasks due to their high accuracy by leveraging the attention mechanism to process image patches as tokens. While numerous hardware accelerators have been proposed to enhance ViT performance, the reliability of ViTs remains an underexplored yet critical challenge. In ViTs, transient errors caused by Bit Flip Attack(BFA) or environment can result in substantial accuracy degradation, posing a serious reliability concern. This paper presents a hardware solution to address this type of reliability problem against transient faults arising from technology scaling, process variation, and fault attacks. We introducing a sensitivity-based policy adjustment, and proposing a flexible and high-performance accelerator, ReSafe-ViT, tailored for robust ViT inference. First, the proposed design consists of a dynamic control mechanism that selects and deploys the optimal error correction method with minimized overhead. Second, it includes a flexible hardware architecture that adapts to the computational and communication requirements of diverse ViT models and datasets and can be reconfigured to perform selected dynamic error control mechanisms. Experimental results show our proposed architecture can reduce accuracy loss by up to 87.9% in the presence of transient errors with minimal timing, power, and chip area overheads. Chunyuan Shen, Ke Wang 0030, Li Yang 0009 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Bal-DGCN: A Hardware Acceleration Framework for Balanced Computational Efficiency in DGCNsabstractDynamic Graph Convolutional Networks (DGCNs) have emerged as a powerful approach for analyzing dynamic graph-structured data across various applications, such as social networks, recommendation systems, and other domains. Typically, each DGCN layer comprises two distinct modules: a Graph Convolutional Network (GCN) module and a Recurrent Neural Network (RNN) module with unique data communication and computation patterns. Although several customized DGCN accelerators have achieved significant improvements, they primarily focus on either enhancing computational efficiency through specialized Processing Element (PE) designs or reducing redundant memory access via algorithm-level optimizations. Many have not exploited data reuse, which is becoming increasingly critical as the size of input dynamic graphs grows. This growth leads to larger and denser intermediate matrices during matrix multiplications, resulting in a substantial increase in memory access. Additionally, current DGCN accelerators ignore the sparsity during inference processing, leading to uneven workload distribution across the hardware platform and resulting in hardware underutilization. The excessive memory access and hardware underutilization ultimately degrade performance and energy efficiency. To address these challenges, this paper proposes Bal-DGCN, a DGCN hardware accelerator that consists of three key innovations: a unified PE array, a versatile interconnection fabric, and a dynamic control policy. Specifically, Bal-DGCN implements a unified PE array to efficiently perform key computations across the GCN and RNN modules, improving computational efficiency. To manage data communication among PEs, Bal-DGCN integrates a versatile interconnection fabric that supports various on-chip data communication patterns, thereby improving data reuse efficiency. Additionally, Bal-DGCN introduces a dynamic control policy that selectively allocates computational workloads to specific hardware resources to mitigate workload imbalance, maximizing hardware utilization. By significantly improving computational efficiency, data reuse, and hardware utilization, Bal-DGCN achieves substantial gains in both performance and energy efficiency. Evaluation results show that Bal-DGCN achieves an average reduction of 76% in execution time and 72% in energy consumption for DGCN inference compared to existing approaches. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | A High-Performance and Flexible Accelerator for Dynamic Graph Convolutional NetworksabstractDynamic Graph Convolutional Networks (DGCNs) have been applied to various dynamic graph-related applications, such as social networks, to achieve high inference accuracy. Typically, each DGCN layer consists of two distinct modules: a Graph Convolutional Network (GCN) module that captures spatial information, and a Recurrent Neural Network (RNN) module that extracts temporal information from input dynamic graphs. The different functionalities of these modules pose significant challenges for hardware platforms, particularly in achieving high-performance and energy-efficient inference processing. To this end, this paper introduces HiFlex, a high-performance and flexible accelerator designed for DGCN inference. At the architecture level, HiFlex implements multiple homogeneous processing elements (PEs) to perform main computations for GCN and RNN modules, along with a versatile interconnection fabric to optimize data communication and enhance on-chip data reuse efficiency. The flexible interconnection fabric can be dynamically configured to provide various on-chip topologies, supporting point-to-point and multicast communication patterns needed for GCN and RNN processing. At the algorithm level, HiFlex introduces a dynamic control policy that partitions, allocates, and configures hardware resources for distinct modules based on their computational requirements. Evaluation results using real-world dynamic graphs demonstrate that HiFlex achieves, on average, a 38% reduction in execution time and a 42 % decrease in energy consumption for DGCN inference, compared to state-of-the-art approaches such as ES-DGCN, ReaDy, and RACE. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
DATE | 2 |
| 2025 | Learning-Enabled Denial-of-Service (DoS) Attack Detection and Mitigation for Chiplet-Based Hybrid Interconnection Network
Md. Tareq Mahmud, Ke Wang 0030 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | FORT-GCN: A Fault-Tolerant and Adaptive Accelerator Design for Efficient Graph Convolutional Network InferenceabstractHardware reliability has emerged as a paramount concern for machine learning accelerators, as transient errors and permanent failures occurring during inference can severely compromise accuracy, performance, and service availability. Although fault resilience in traditional machine learning, such as Deep Neural Networks (DNNs), has been extensively studied, graph convolutional networks (GCNs) present unique reliability challenges due to their irregular computation patterns and dynamic data dependencies. Traditional fault mitigation approaches, including hardware redundancy, recomputation, and Hamming code protection, suffer from prohibitive latency and power overheads when applied to GCN accelerators. This article presents FORT-GCN, a holistic hardware architecture co-optimized for GCN-specific fault resilience. Our solution integrates three key innovations, namely permanent fault tolerance through a novel robust processing element design with runtime reconfiguration and defect-adaptive interconnects, transient error resilience via lightweight selective error correction unit design, and a fault-aware adaptive controller design that dynamically adjusts fault protection strategies based on operational faults and graph characteristics. Experimental evaluation demonstrates 35.4% improvement in fault robustness compared to conventional error-correction and redundancy-based approaches, with minimal timing, area, and power overheads. Ke Wang 0030, Yingnan Zhao 0001, Ahmed Louri |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2025 | HS-GCN: A High-Performance, Sustainable, and Scalable Chiplet-Based Accelerator for Graph Convolutional Network InferenceabstractGraph Convolutional Networks (GCNs) have been proposed to extend machine learning techniques for graphrelated applications. A typical GCN model consists of multiple layers, each including an aggregation phase, which is communication-intensive, and a combination phase, which is computation-intensive. As the size of real-world graphs increases exponentially, current customized accelerators face challenges in efficiently performing GCN inference due to limited on-chip buffers and other hardware resources for both data computation and communication, which degrades performance and energy efficiency. Additionally, scaling current monolithic designs to address the aforementioned challenges will introduce significant cost-effectiveness issues in terms of power, area, and yield. To this end, we propose HS-GCN, a high-performance, sustainable, and scalable chiplet-based accelerator for GCN inference with muchimproved energy efficiency. Specifically, HS-GCN integrates multiple reconfigurable chiplets, each of which can be configured to perform the main computations of either the aggregation phase or the combination phase, including Sparse-dense matrix multiplication (SpMM) and General matrix-matrix multiplication (GeMM). HS-GCN implements an active interposer with a flexible interconnection fabric to connect chiplets and other hardware components for efficient data communication. Additionally, HS-GCN introduces two system-level control algorithms that dynamically determine the computation order and corresponding dataflow based on the input graphs and GCN models. These selections are used to further configure the chiplet array and interconnection fabric for much-improved performance and energy efficiency. Evaluation results using real-world graphs demonstrate that HS-GCN achieves significant speedups of 26.7×, 11.2×, 3.9×, 4.7×, 3.1×, along with substantial memory access savings of 94%, 89%, 64%, 85%, 54%, and energy savings of 87%, 84%, 49%, 78%, 41% on average, as compared to HyGCN, AWB-GCN, GCNAX, I-GCN, and SGCN, respectively. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | An Efficient Hardware Accelerator Design for Dynamic Graph Convolutional Network (DGCN) InferenceabstractDynamic graph convolutional networks (DGCNs) have been increasingly used to extend machine learning techniques to applications that involve graph-structured data with temporal changes. A typical DGCN model is comprised of graph convolutional network (GCN) layers to capture spatial information, followed by recurrent network (RNN) layers for temporal information. Designing a highperformance and energy-efficient DGCN accelerator is challenging due to the distinct computation and communication requirements of the GCN and RNN layers. Specifically, the computation of GCN layers can be abstracted as Sparse-dense and General Matrix-matrix Multiplication (SpMM and GeMM), while RNN layers involve extensive element-wise addition and Hadamard product in addition to SpMM and GeMM. For data communication, GCN layers necessitate irregular data memory access due to the unstructured distribution of vertices involved in graphs, whereas RNN layers exhibit a predictable memory access pattern. We propose E-DGCN, a highperformance and energy-efficient accelerator design for improved DGCN inference. The proposed E-DGCN comprises reconfigurable processing elements that efficiently support diverse types of data computations required by GCN and RNN layers, a flexible on-chip interconnection design with an adaptive dataflow to improve data reuse during DGCN inference, and a lightweight vertex caching algorithm to leverage data locality and reduce off-chip memory access while processing temporal information. Experimental results show that the E-DGCN achieves 2.2x speed-up and 2.6x energy savings on average as compared to existing DGCN accelerators. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
DAC | 2 |
| 2024 | OPT-GCN: A Unified and Scalable Chiplet-Based Accelerator for High-Performance and Energy-Efficient GCN ComputationabstractAs the size of real-world graphs continues to grow at an exponential rate, performing the Graph Convolutional Network (GCN) inference efficiently is becoming increasingly challenging. Prior works that employ a unified computing engine with a predefined computation order lack the necessary flexibility and scalability to handle diverse input graph datasets. In this paper, we introduce OPT-GCN, a chiplet-based accelerator design that performs GCN inference efficiently while providing flexibility and scalability through an architecture-algorithm co-design. On the architecture side, the proposed design integrates a unified computing engine in each chiplet and an active interposer, both of which are adaptable to efficiently perform the GCN inference and facilitate data communication. On the algorithm side, we propose dynamic scheduling and mapping algorithms to optimize memory access and on-chip computations for diverse GCN applications. Experimental results show that the proposed design provides a memory access reduction by a factor of 11.3×, 3.4×, 1.4× energy savings of 15.2×, 3.7×, 1.6× on average compared to HyGCN, AWB-GCN, and GCNAX, respectively. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | Morph-GCNX: A Universal Architecture for High-Performance and Energy-Efficient Graph Convolutional Network AccelerationabstractWhile current Graph Convolutional Networks (GCNs) accelerators have achieved notable success in a wide range of application domains, these GCN accelerators can not support various intra- and inter- GCN dataflows or adapt to diverse GCN applications. In this paper, we propose Morph-GCNX, a flexible GCN accelerator architecture for high-performance and energy-efficient GCN execution. The proposed design consists of a flexible Processing Element (PE) array that can be partitioned at runtime and adapt to the computational needs of different layers within a GCN or multiple concurrent GCNs. The proposed Morph-GCNX also consists of a morphable interconnection design to support a wide range of GCN dataflows with various parallelization and data reuse strategies for GCN execution. We also propose a hardware-application co-exploration technique that explores the GCN and hardware design spaces to identify the best PE partition, workload allocation, dataflow, and interconnection configurations, with the goal of improving overall performance and energy. Simulation results show that the proposed Morph-GCNX architecture achieves 18.8×, 2.9×, 1.9×, 1.8×, and 2.5× better performance, reduces DRAM accesses by a factor of 10.8×, 3.7×, 2.2×, 2.5×, and 1.3×, and improves energy consumption by 13.2×, 5.6×, 2.1×, 2.5×, and 1.3×, as compared to prior designs including HyGCN, AWB-GCN, LW-GCN, GCoD, and GCNAX, respectively. Ke Wang 0030, Hao Zheng 0005, Ahmed Louri |
IEEE Trans. Sustain. Comput. | 1 |
| 2023 | FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential EquationsabstractPartial Differential Equations (PDEs) are widely employed to describe natural phenomena in many science and engineering fields. Many PDEs do not have analytical solutions, hence, numerical methods have become prevalent for approximating PDE solutions. The most widely used numerical method is the Finite Difference Method (FDM), which requires fine grids and high-precision numerical iterations that are both compute- and memory-intensive. PDE-solving accelerators have been proposed in the literature, however, they usually focus on specific types of PDEs with rigid grid sizes which limits their broader applicability. Besides, they rarely provided insight into the optimizations of parallel computing and data accesses for solving PDEs, which hinders further improvements in performance and energy efficiency. Hao Zheng 0005, Ke Wang 0030 |
ISCA | 4 |
| 2023 | GShuttle: Optimizing Memory Access Efficiency for Graph Convolutional Neural Network Accelerators
Ke Wang 0030, Hao Zheng 0005, Ahmed Louri |
J. Comput. Sci. Technol. | 2 |
| 2022 | AGAPE: Anomaly Detection with Generative Adversarial Network for Improved Performance, Energy, and Security in Manycore SystemsabstractThe security of manycore systems has become increasingly critical. In system-on-chips (SoCs), Hardware Trojans (HTs) manipulate the functionalities of the routing components to saturate the on-chip network, degrade performance, and result in the leakage of sensitive data. Existing HT detection techniques, including runtime monitoring and state-of-the-art learning-based methods, are unable to timely and accurately identify the implanted HTs, due to the increasingly dynamic and complex nature of on-chip communication behaviors. We propose AGAPE, a novel Generative Adversarial Network (GAN)-based anomaly detection and mitigation method against HTs for secured on-chip communication. AGAPE learns the distribution of the multivariate time series of a number of NoC attributes captured by on-chip sensors under both HT-free and HT-infected working conditions. The proposed GAN can learn the potential latent interactions among different runtime attributes concurrently, accurately distinguish abnormal attacked situations from normal SoC behaviors, and identify the type and location of the implanted HTs. Using the detection results, we apply the most suitable protection techniques to each type of detected HTs instead of simply isolating the entire HT-infected router, with the aim to mitigate security threats as well as reducing performance loss. Simulation results show that AGAPE enhances the HT detection accuracy by 19%, reduces network latency and power consumption by 39% and 30%, respectively, as compared to state-of-the-art security designs. Ke Wang 0030, Hao Zheng 0005, Yuan Li 0029, Ahmed Louri |
DATE | 1 |
| 2022 | FSA: An Efficient Fault-tolerant Systolic Array-based DNN Accelerator ArchitectureabstractWith the advent of Deep Neural Network (DNN) accelerators, permanent faults are increasingly becoming a serious challenge for DNN hardware accelerator, as they can severely degrade DNN inference accuracy. The State-of-the-art works address this issue by adding homogeneous redundant Processing Elements (PEs) to the DNN accelerator’s central computing array, or bypassing faulty PEs directly. However, such designs induce inference loss, extra hardware cost, and performance overhead. Moreover, current designs are able to only deal with a limited number of faults due to costs. In this paper, we propose FSA, a Fault-tolerant Systolic Array-based DNN accelerator with the goal of maintaining DNN inference accuracy in the presence of permanent faults. The key feature of the proposed FSA is a unified re-computing module (RCM) that dynamically recalculates the required DNN computations that are supposed to be accomplished by faulty PEs with minimal latency and power consumption. Simulation results show that the proposed FSA reduces inference accuracy loss by 46%, improves execution time by 23%, and reduces energy consumption by 35% on average, as compared to existing designs. Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri |
ICCD | 2 |
| 2022 | Ascend: A Scalable and Energy-Efficient Deep Neural Network Accelerator With Photonic InterconnectsabstractThe complexity and size of recent deep neural network (DNN) models have increased significantly in pursuit of high inference accuracy. Chiplet-based accelerator is considered a viable scaling approach to provide substantial computation capability and on-chip memory for efficient process of such DNN models. However, communication using metallic interconnects in prior chiplet-based accelerators poses a major challenge to system performance, energy efficiency, and scalability. Photonic interconnects can adequately support communication across chiplets due to features such as distance-independent latency, high bandwidth density, and high energy efficiency. Furthermore, the salient ease of broadcast property makes photonic interconnects suitable for DNN inference which often incurs prevalent broadcast communication. In this paper, we propose a scalable chiplet-based DNN accelerator with photonic interconnects named ASCEND. ASCEND introduces (1) a novel photonic network that supports seamless intra- and inter- chiplet broadcast communication, and flexible mapping of diverse convolution layers, and (2) a tailored dataflow that exploits the ease of broadcast property and maximizes parallelism by simultaneously processing computations with shared input data. Simulation results using multiple DNN models show that ASCEND achieves 71% and 67% reduction in execution time and energy consumption, respectively, as compared to other state-of-the-art chiplet-based DNN accelerators with metallic or photonic interconnects. Yuan Li 0029, Ke Wang 0030, Hao Zheng 0005, Ahmed Louri, Avinash Karanth |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | SGCNAX: A Scalable Graph Convolutional Neural Network Accelerator With Workload BalancingabstractGraph Convolutional Neural Networks (GCNs) have emerged as promising tools for graph-based machine learning applications. Given that GCNs are both compute- and memory-intensive, this constitutes a major challenge for the underlying hardware to efficiently process large-scale GCNs. In this paper, we introduce SGCNAX, a scalable GCN accelerator architecture for the high-performance and energy-efficient acceleration of GCNs. Unlike prior GCN accelerators that either employ limited loop optimization techniques, or determine the design variables based on random sampling, we systematically explore the loop optimization techniques for GCN acceleration and propose a flexible GCN dataflow that adapts to different GCN configurations to achieve optimal efficiency. We further propose two hardware-based techniques to address the workload imbalance problem caused by the unbalanced distribution of zeros in GCNs. Specifically, SGCNAX exploits an outer-product-based computation architecture that mitigates the intra-PE (Processing Elements) workload imbalance, and employs a group-and-shuffle approach to mitigate the inter-PE workload imbalance. Simulation results show that SGCNAX performs 9.2x, 1.6x and 1.2x better, and reduces DRAM accesses by a factor of 9.7x, 2.9x and 1.2x compared to HyGCN, AWB-GCN, and GCNAX, respectively. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | SecureNoC: A Learning-Enabled, High-Performance, Energy-Efficient, and Secure On-Chip Communication Framework DesignabstractWe propose SecureNoC, a learning-based framework to enhance NoC security against Hardware Trojan (HT) attacks while holistically improving performance and power. The proposed framework enhances NoC security with several architectural innovations, namely a per-router HT detector, multi-function bypass channels (MBCs), and a lightweight data encryption design. Specifically, the threat detector uses an artificial neural network for runtime HT detection with high accuracy. The MBCs consist of a router bypass route and reconfigurable channel buffers which can efficiently isolate malicious nodes and reduce power consumption. The proposed data encryption design adapts to diverse traffic patterns and dynamically deploys novel lightweight encryption techniques for desired security goals with improved latency. Additionally, to balance the trade-offs and handle the dynamic interactions of the proposed dynamic designs, a proactive deep-Q-learning (DQL) control policy is proposed to simultaneously provide optimized NoC security, performance, and power consumption. Simulation studies using PARSEC benchmarks show that the proposed SecureNoC achieves 36% higher HT detection accuracy over state-of-the-art NoC security techniques while reducing network latency by 39% and energy consumption by 46%. Ke Wang 0030, Hao Zheng 0005, Yuan Li 0029, Ahmed Louri |
IEEE Trans. Sustain. Comput. | 1 |
| 2021 | Adapt-NoC: A Flexible Network-on-Chip Design for Heterogeneous Manycore ArchitecturesabstractThe increased computational capability in heterogeneous manycore architectures facilitates the concurrent execution of many applications. This requires, among other things, a flexible, high-performance, and energy-efficient communication fabric capable of handling a variety of traffic patterns needed for running multiple applications at the same time. Such stringent requirements are posing a major challenge for current Network-on-Chips (NoCs) design. In this paper, we propose Adapt-NoC, a flexible NoC architecture, along with a reinforcement learning (RL)-based control policy, that can provide efficient communication support for concurrent application execution. Adapt-NoC can dynamically allocate several disjoint regions of the NoC, called subNoCs, with different sizes and locations for the concurrently running applications. Each of the dynamically-allocated subNoCs is capable of adapting to a given topology such as a mesh, cmesh, torus, or tree thus tailoring the topology to satisfy application's needs in terms of performance and power consumption. Moreover, we explore the use of RL to design an efficient control policy which optimizes the subNoC topology selection for a given application. As such, Adapt-NoC can not only provide several topology choices for concurrently running applications, but can also optimize the selection of the most suitable topology for a given application with the aim of improving performance and energy efficiency. We evaluate Adapt-NoC using both GPU and CPU benchmark suites. Simulation results show that the proposed Adapt-NoC can achieve up to 34% latency reduction, 10% overall execution time reduction and 53% NoC energy-efficiency improvement when compared to prior work. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
HPCA | 2 |
| 2020 | A Versatile and Flexible Chiplet-based System Design for Heterogeneous Manycore ArchitecturesabstractHeterogeneous manycore architectures are deployed to simultaneously run multiple and diverse applications. This requires various computing capabilities (CPUs, GPUs, and accelerators), and an efficient network-on-chip (NoC) architecture to concurrently handle diverse application communication behavior. However, supporting the concurrent communication requirements of diverse applications is challenging due to the dynamic application mapping, the complexity of handling distinct communication patterns and limited on-chip resources. In this paper, we propose Adapt-NoC, a versatile and flexible NoC architecture for chiplet-based manycore architectures, consisting of adaptable routers and links. Adapt-NoC can dynamically allocate disjoint regions of the NoC, called subNoCs, for concurrently-running applications, each of which can be optimized for different communication behavior. The adaptable routers and links are capable of providing various subNoC topologies, satisfying different latency and bandwidth requirements of various traffic patterns (e.g. all-to-all, one-to-many). Full system simulation shows that AdaptNoC can achieve 31% latency reduction, 24% energy saving and 10% execution time reduction on average, when compared to prior designs. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
DAC | 2 |
| 2020 | CURE: A High-Performance, Low-Power, and Reliable Network-on-Chip Design Using Reinforcement LearningabstractWe propose CURE, a deep reinforcement learning (DRL)-based NoC design framework that simultaneously reduces network latency, improves energy-efficiency, and tolerates transient errors and permanent faults. CURE has several architectural innovations and a DRL-based hardware controller to manage design complexity and optimize trade-offs. First, in CURE, we propose reversible multi-function adaptive channels (RMCs) to reduce NoC power consumption and network latency. Second, we implement a new fault-secure adaptive error correction hardware in each router to enhance reliability for both transient errors and permanent faults. Third, we propose a router power-gating and bypass design that powers off NoC components to reduce power and extend chip lifespan. Further, for the complex dynamic interactions of these techniques, we propose using DRL to train a proactive control policy to provide improved fault-tolerance, reduced power consumption, and improved performance. Simulation using the PARSEC benchmark shows that CURE reduces end-to-end packet latency by 39 percent, improves energy efficiency by 92 percent, and lowers static and dynamic power consumption by 24 and 38 percent, respectively, over conventional solutions. Using mean-time-to-failure, we show that CURE is 7.7× more reliable than the conventional NoC design. Ke Wang 0030, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | High-performance, Energy-efficient, Fault-tolerant Network-on-Chip Design Using Reinforcement LearninabstractNetwork-on-Chips (NoCs) are becoming the standard communication fabric for multi-core and system on a chip (SoC) architectures. As technology continues to scale, transistors and wires on the chip are becoming increasingly vulnerable to various fault mechanisms, especially timing errors, resulting in exacerbation of energy efficiency and performance for NoCs. Typical techniques for handling timing errors are reactive in nature, responding to the faults after their occurrence. They rely on error detection/correction techniques which have resulted in excessive power consumption and degraded performance, since the error detection/correction hardware is constantly enabled. On the other hand, indiscriminately disabling error handling hardware can induce more errors and intrusive retransmission traffic. Therefore, the challenge is to balance the trade-offs among error rate, packet retransmission, performance, and energy. In this paper, we propose a proactive fault-tolerant mechanism to optimize energy efficiency and performance with reinforcement learning (RL). First, we propose a new proactive error handling technique comprised of a dynamic scheme for enabling per-router error detection/correction hardware and an effective retransmission mechanism. Second, we propose the use of RL to train the dynamic control policy with the goals of providing increased fault-tolerance, reduced power consumption and improved performance as compared to conventional techniques. Our evaluation indicates that, on average, end-to-end packet latency is lowered by 55%, energy efficiency is improved by 64%, and retransmission caused by faults is reduced by 48% over the reactive error correction techniques. Ke Wang 0030, Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
DATE | 1 |
| 2019 | IntelliNoC: a holistic design framework for energy-efficient and reliable on-chip communication for manycoresabstractAs technology scales, Network-on-Chips (NoCs), currently being used for on-chip communication in manycore architectures, face several problems including high network latency, excessive power consumption, and low reliability. Simultaneously addressing these problems is proving to be difficult due to the explosion of the design space and the complexity of handling many trade-offs. In this paper, we propose IntelliNoC, an intelligent NoC design framework which introduces architectural innovations and uses reinforcement learning to manage the design complexity and simultaneously optimize performance, energy-efficiency, and reliability in a holistic manner. IntelliNoC integrates three NoC architectural techniques: (1) multifunction adaptive channels (MFACs) to improve energy-efficiency; (2) adaptive error detection/correction and re-transmission control to enhance reliability; and (3) a stress-relaxing bypass feature which dynamically powers off NoC components to prevent overheating and fatigue. To handle the complex dynamic interactions induced by these techniques, we train a dynamic control policy using Q-learning, with the goal of providing improved fault-tolerance and performance while reducing power consumption and area overhead. Simulation using PARSEC benchmarks shows that our proposed IntelliNoC design improves energy-efficiency by 67% and mean-time-to-failure (MTTF) by 77%, and decreases end-to-end packet latency by 32% and area requirements by 25% over baseline NoC architecture. Ke Wang 0030, Ahmed Louri, Avinash Karanth, Razvan C. Bunescu |
ISCA | 1 |