Yingnan Zhao 0001

dblp:92/4786-1 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2026
0009-0005-5776-6239ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Bal-DGCN: A Hardware Acceleration Framework for Balanced Computational Efficiency in DGCNs
abstract
Dynamic Graph Convolutional Networks (DGCNs) have emerged as a powerful approach for analyzing dynamic graph-structured data across various applications, such as social networks, recommendation systems, and other domains. Typically, each DGCN layer comprises two distinct modules: a Graph Convolutional Network (GCN) module and a Recurrent Neural Network (RNN) module with unique data communication and computation patterns. Although several customized DGCN accelerators have achieved significant improvements, they primarily focus on either enhancing computational efficiency through specialized Processing Element (PE) designs or reducing redundant memory access via algorithm-level optimizations. Many have not exploited data reuse, which is becoming increasingly critical as the size of input dynamic graphs grows. This growth leads to larger and denser intermediate matrices during matrix multiplications, resulting in a substantial increase in memory access. Additionally, current DGCN accelerators ignore the sparsity during inference processing, leading to uneven workload distribution across the hardware platform and resulting in hardware underutilization. The excessive memory access and hardware underutilization ultimately degrade performance and energy efficiency. To address these challenges, this paper proposes Bal-DGCN, a DGCN hardware accelerator that consists of three key innovations: a unified PE array, a versatile interconnection fabric, and a dynamic control policy. Specifically, Bal-DGCN implements a unified PE array to efficiently perform key computations across the GCN and RNN modules, improving computational efficiency. To manage data communication among PEs, Bal-DGCN integrates a versatile interconnection fabric that supports various on-chip data communication patterns, thereby improving data reuse efficiency. Additionally, Bal-DGCN introduces a dynamic control policy that selectively allocates computational workloads to specific hardware resources to mitigate workload imbalance, maximizing hardware utilization. By significantly improving computational efficiency, data reuse, and hardware utilization, Bal-DGCN achieves substantial gains in both performance and energy efficiency. Evaluation results show that Bal-DGCN achieves an average reduction of 76% in execution time and 72% in energy consumption for DGCN inference compared to existing approaches.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
IEEE Trans. Parallel Distributed Syst.1
2025 A High-Performance and Flexible Accelerator for Dynamic Graph Convolutional Networks
abstract
Dynamic Graph Convolutional Networks (DGCNs) have been applied to various dynamic graph-related applications, such as social networks, to achieve high inference accuracy. Typically, each DGCN layer consists of two distinct modules: a Graph Convolutional Network (GCN) module that captures spatial information, and a Recurrent Neural Network (RNN) module that extracts temporal information from input dynamic graphs. The different functionalities of these modules pose significant challenges for hardware platforms, particularly in achieving high-performance and energy-efficient inference processing. To this end, this paper introduces HiFlex, a high-performance and flexible accelerator designed for DGCN inference. At the architecture level, HiFlex implements multiple homogeneous processing elements (PEs) to perform main computations for GCN and RNN modules, along with a versatile interconnection fabric to optimize data communication and enhance on-chip data reuse efficiency. The flexible interconnection fabric can be dynamically configured to provide various on-chip topologies, supporting point-to-point and multicast communication patterns needed for GCN and RNN processing. At the algorithm level, HiFlex introduces a dynamic control policy that partitions, allocates, and configures hardware resources for distinct modules based on their computational requirements. Evaluation results using real-world dynamic graphs demonstrate that HiFlex achieves, on average, a 38% reduction in execution time and a 42 % decrease in energy consumption for DGCN inference, compared to state-of-the-art approaches such as ES-DGCN, ReaDy, and RACE.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
DATE1
2025 FORT-GCN: A Fault-Tolerant and Adaptive Accelerator Design for Efficient Graph Convolutional Network Inference
abstract
Hardware reliability has emerged as a paramount concern for machine learning accelerators, as transient errors and permanent failures occurring during inference can severely compromise accuracy, performance, and service availability. Although fault resilience in traditional machine learning, such as Deep Neural Networks (DNNs), has been extensively studied, graph convolutional networks (GCNs) present unique reliability challenges due to their irregular computation patterns and dynamic data dependencies. Traditional fault mitigation approaches, including hardware redundancy, recomputation, and Hamming code protection, suffer from prohibitive latency and power overheads when applied to GCN accelerators. This article presents FORT-GCN, a holistic hardware architecture co-optimized for GCN-specific fault resilience. Our solution integrates three key innovations, namely permanent fault tolerance through a novel robust processing element design with runtime reconfiguration and defect-adaptive interconnects, transient error resilience via lightweight selective error correction unit design, and a fault-aware adaptive controller design that dynamically adjusts fault protection strategies based on operational faults and graph characteristics. Experimental evaluation demonstrates 35.4% improvement in fault robustness compared to conventional error-correction and redundancy-based approaches, with minimal timing, area, and power overheads.
Ke Wang 0030, Yingnan Zhao 0001, Ahmed Louri
ACM Trans. Embed. Comput. Syst.2
2025 HS-GCN: A High-Performance, Sustainable, and Scalable Chiplet-Based Accelerator for Graph Convolutional Network Inference
abstract
Graph Convolutional Networks (GCNs) have been proposed to extend machine learning techniques for graphrelated applications. A typical GCN model consists of multiple layers, each including an aggregation phase, which is communication-intensive, and a combination phase, which is computation-intensive. As the size of real-world graphs increases exponentially, current customized accelerators face challenges in efficiently performing GCN inference due to limited on-chip buffers and other hardware resources for both data computation and communication, which degrades performance and energy efficiency. Additionally, scaling current monolithic designs to address the aforementioned challenges will introduce significant cost-effectiveness issues in terms of power, area, and yield. To this end, we propose HS-GCN, a high-performance, sustainable, and scalable chiplet-based accelerator for GCN inference with muchimproved energy efficiency. Specifically, HS-GCN integrates multiple reconfigurable chiplets, each of which can be configured to perform the main computations of either the aggregation phase or the combination phase, including Sparse-dense matrix multiplication (SpMM) and General matrix-matrix multiplication (GeMM). HS-GCN implements an active interposer with a flexible interconnection fabric to connect chiplets and other hardware components for efficient data communication. Additionally, HS-GCN introduces two system-level control algorithms that dynamically determine the computation order and corresponding dataflow based on the input graphs and GCN models. These selections are used to further configure the chiplet array and interconnection fabric for much-improved performance and energy efficiency. Evaluation results using real-world graphs demonstrate that HS-GCN achieves significant speedups of 26.7×, 11.2×, 3.9×, 4.7×, 3.1×, along with substantial memory access savings of 94%, 89%, 64%, 85%, 54%, and energy savings of 87%, 84%, 49%, 78%, 41% on average, as compared to HyGCN, AWB-GCN, GCNAX, I-GCN, and SGCN, respectively.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
IEEE Trans. Sustain. Comput.1
2024 An Efficient Hardware Accelerator Design for Dynamic Graph Convolutional Network (DGCN) Inference
abstract
Dynamic graph convolutional networks (DGCNs) have been increasingly used to extend machine learning techniques to applications that involve graph-structured data with temporal changes. A typical DGCN model is comprised of graph convolutional network (GCN) layers to capture spatial information, followed by recurrent network (RNN) layers for temporal information. Designing a highperformance and energy-efficient DGCN accelerator is challenging due to the distinct computation and communication requirements of the GCN and RNN layers. Specifically, the computation of GCN layers can be abstracted as Sparse-dense and General Matrix-matrix Multiplication (SpMM and GeMM), while RNN layers involve extensive element-wise addition and Hadamard product in addition to SpMM and GeMM. For data communication, GCN layers necessitate irregular data memory access due to the unstructured distribution of vertices involved in graphs, whereas RNN layers exhibit a predictable memory access pattern. We propose E-DGCN, a highperformance and energy-efficient accelerator design for improved DGCN inference. The proposed E-DGCN comprises reconfigurable processing elements that efficiently support diverse types of data computations required by GCN and RNN layers, a flexible on-chip interconnection design with an adaptive dataflow to improve data reuse during DGCN inference, and a lightweight vertex caching algorithm to leverage data locality and reduce off-chip memory access while processing temporal information. Experimental results show that the E-DGCN achieves 2.2x speed-up and 2.6x energy savings on average as compared to existing DGCN accelerators.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
DAC1
2024 OPT-GCN: A Unified and Scalable Chiplet-Based Accelerator for High-Performance and Energy-Efficient GCN Computation
abstract
As the size of real-world graphs continues to grow at an exponential rate, performing the Graph Convolutional Network (GCN) inference efficiently is becoming increasingly challenging. Prior works that employ a unified computing engine with a predefined computation order lack the necessary flexibility and scalability to handle diverse input graph datasets. In this paper, we introduce OPT-GCN, a chiplet-based accelerator design that performs GCN inference efficiently while providing flexibility and scalability through an architecture-algorithm co-design. On the architecture side, the proposed design integrates a unified computing engine in each chiplet and an active interposer, both of which are adaptable to efficiently perform the GCN inference and facilitate data communication. On the algorithm side, we propose dynamic scheduling and mapping algorithms to optimize memory access and on-chip computations for diverse GCN applications. Experimental results show that the proposed design provides a memory access reduction by a factor of 11.3×, 3.4×, 1.4× energy savings of 15.2×, 3.7×, 1.6× on average compared to HyGCN, AWB-GCN, and GCNAX, respectively.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 FSA: An Efficient Fault-tolerant Systolic Array-based DNN Accelerator Architecture
abstract
With the advent of Deep Neural Network (DNN) accelerators, permanent faults are increasingly becoming a serious challenge for DNN hardware accelerator, as they can severely degrade DNN inference accuracy. The State-of-the-art works address this issue by adding homogeneous redundant Processing Elements (PEs) to the DNN accelerator’s central computing array, or bypassing faulty PEs directly. However, such designs induce inference loss, extra hardware cost, and performance overhead. Moreover, current designs are able to only deal with a limited number of faults due to costs. In this paper, we propose FSA, a Fault-tolerant Systolic Array-based DNN accelerator with the goal of maintaining DNN inference accuracy in the presence of permanent faults. The key feature of the proposed FSA is a unified re-computing module (RCM) that dynamically recalculates the required DNN computations that are supposed to be accomplished by faulty PEs with minimal latency and power consumption. Simulation results show that the proposed FSA reduces inference accuracy loss by 46%, improves execution time by 23%, and reduces energy consumption by 35% on average, as compared to existing designs.
Yingnan Zhao 0001, Ke Wang 0030, Ahmed Louri
ICCD1