EDBT 2026 Demo / reviewers in the wild / expert
Hao Zheng 0005
dblp:31/6916-5
· DBLP profile ↗
34ranked-venue papers
3as first author
32since 2021 · last 2026
0000-0003-4391-2774ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 3 first-author · 28 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Graph Neural Network Training via Geometric OptimizationabstractWafer-scale computing has emerged as an alternative solution to sustain performance scaling in the post-Moore era, driven by recent technology advancements such as chiplet integration. This enables considerable computing and storage capabilities on a single chip, making it capable of accommodating large machine learning models and datasets. Recent efforts have heralded the promise of wafer-scale architectures for deep learning inference and training. However, scaling the training of Graph Neural Networks in wafer-scale architecture remains a challenge and is relatively unexplored due to irregularities in gradient propagation as well as physical constraints from flat on-chip topologies. In this paper, we propose Aster, a topology-aware framework designed to efficiently support GNN training on arbitrary wafer-scale architectures. The proposed framework, as opposed to the current application or topology-specific heuristics, can be generalized to support any network topology and irregular GNN datasets. Specifically, we mathematically formulate commonly-seen network topologies in their geometric representation and prioritize communication efficiency during GNN workload partitioning and mapping. Based on the geometric representation, we propose a quadratic assignment problem solver to efficiently map irregular dataflows to a flat topology with reduced communication distance. The simulation results show that Aster can achieve performance speedup by$2.91 \times, 1.50 \times, 1.84 \times$, and$1.58 \times$in Mesh and speedup by$3.84 \times, 1.56 \times, 2.05 \times$, and$1.49 \times$in Torus on average compared to Mini-cut [1], ScalaGraph [2], ChunkV [3], and Chunk-E [4], respectively. Fangzhou Ye, Lingxiang Yin, Hao Zheng 0005 |
HPCA | 3 |
| 2026 | TensorPrism: Rethinking Sparse High-Order Tensor Acceleration via Co-Occurrence Graph
Fangzhou Ye, Shilin Tian, Amir Ghazizadeh Ahsaei, Hao Zheng 0005 |
ISCA | 4 |
| 2025 | I-DGNN: A Graph Dissimilarity-based Framework for Designing Scalable and Efficient DGNN AcceleratorsabstractDynamic Graph Neural Networks (DGNNs) have recently been used in numerous application domains, comprehending the intricate dynamics of time-evolving graph data. Despite their theoretical advancements, effectively implementing scalable DGNNs continues to be a formidable challenge due to the constantly evolving graph data and heterogeneous computation kernels. Recent efforts attempted to either exploit the graph data reuse to reduce memory access or eliminate the redundant computations between consecutive graph snapshots to scale the DGNN acceleration. These efforts are still falling short. In prior work, each graph snapshot, regardless of its size and connectivity, passes through the entire DGNN computation pipeline from layer to layer. Consequently, substantial intermediate data is generated throughout the DGNN computation, which leads to excessive offchip memory access. To address this crucial challenge, we argue that the computations between evolving graph snapshots should be decoupled from the DGNN execution pipeline. In this paper, we propose I-DGNN, a theoretical, architectural, and algorithmic framework with the aim of designing scalable and efficient accelerators for DGNN execution with improved performance and energy efficiency. On the theory side, the key idea is to identify essential computations between consecutive graph snapshots and encapsulate them as a separate kernel independent from the DGNN model. Specifically, the proposed one-pass DGNN computing model extracts the process of graph update as a chained matrix multiplication between evolving graphs through rigorous mathematical derivations. Consequently, consecutive snapshots utilize a onepass computation kernel instead of passing through the entire DGNN execution pipeline, thereby eliminating the costly data movement of intermediate results across DGNN layers. On the architecture side, we propose a unified accelerator architecture that can be dynamically configured to support the computation characteristics of the proposed I-DGNN computing model with improved data and pipeline parallelism. On the algorithm side, we propose a new dataflow and mapping tailored for I-DGNN to further improve the data locality of inter-kernel data across the DGNN pipeline. Simulation results show that the proposed accelerator achieves $\mathbf{6 5. 9 \%}, \mathbf{7 1. 1 \%}$, and $\mathbf{5 8. 8 \%}$ reductions in execution time and $88.4 \%, 87.0 \%$, and $\mathbf{8 5. 9 \%}$ improvements in energy efficiency on average across multiple DGNN datasets compared to state-of-the-art-accelerators [1]–[3]. Hao Zheng 0005, Ahmed Louri |
HPCA | 2 |
| 2025 | Multitask Contrastive Learning using Task-Wise Training and Partitioned Embedding SpaceabstractMany real-world computer vision tasks require learning to associate multiple properties of different modalities with the same image. Multi-task learning enables a single model to learn these properties simultaneously by leveraging shared knowledge across related tasks to enhance generalization with single-modal data. On the other hand, contrastive learning effectively captures robust multi-modal features by aligning similar representations and distinguishing dissimilar ones. However, state-of-the-art methods struggle with combining these two learning approaches due to the difficulty in optimizing both shared and task-specific objectives. In this paper, we introduce a Multi-Task Contrastive Learning (MTCL) framework that partitions the embedding space to support both classification and regression tasks within a multi-task paradigm. By batching samples with tasks and structuring the embedding space to accommodate diverse task-specific requirements, our method retains the advantages of contrastive learning while addressing the unique challenges of multi-task learning. We evaluate our approach on three benchmark multi-task datasets—Zappos50K, CUB200, and MEDIC. We also introduce a multi-task Vehicles dataset that includes orientation. On the benchmark datasets, our model shows 24.5%, 17.2%, and 30.0% increase in overall classification accuracy compared to the SOTA methods. M. Shifat Hossain, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICMLA | 3 |
| 2025 | Attr-RAG: Attribution-Guided Retrieval-Augmented Generation for Scientific Experiment DesignabstractEvidence-based science depends on the iterative integration of experimentation, a process traditionally driven by slow and error-prone human effort. This has inspired the vision of an automated "robot scientist" capable of conducting end-to-end experimentation. While Large Language Models (LLMs) can generate procedural instructions, they often struggle to accurately describe scientific experiments due to the limited availability of high-quality, domain-specific examples in their training data. Retrieval-Augmented Generation (RAG) helps bridge this gap by allowing LLMs to access up-to-date external information. However, despite being effective for short questions, RAG struggles with long-form scientific experimental queries due to information loss from chunk fragmentation and retrieval of irrelevant information. In this paper, we propose Attr-RAG, an attribution-guided RAG framework to remove irrelevant or misleading context and retaining only complete, relevant information. Unlike traditional RAG methods that rely solely on vector similarity, Attr-RAG introduces a refinement stage using occlusion-based attribution to identify which retrieved chunks truly influence the LLM’s response. This attribution-guided filtering ensures that only contextually coherent chunks are used for accurate and grounded final answer generation. Attr-RAG demonstrated superior performance in 9 out of 10 chemistry lab experiment tasks of the ChemEx dataset and outperformed baselines across most quantitative evaluation metrics. In qualitative evaluations conducted by state-of-the-art LLM judges (GPT-4o, Gemini 2.5, and Grok 3), the top mean scores of 27.8, 27.1, and 22.9, respectively, were achieved across six key evaluation criteria. Fazle Rahat, M. Shifat Hossain, Arvind Ramanathan, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICMLA | 5 |
| 2025 | DiTile-DGNN: An Efficient Accelerator for Distributed Dynamic Graph Neural Network InferenceabstractDynamic Graph Neural Networks (DGNNs) have recently emerged as a promising model for learning complex temporal and spatial relationships in evolving graphs.The performance of DGNNs is enabled by the simultaneous integration of both graph neural networks (GNNs) and recurrent neural networks (RNNs).Despite the theoretical advancements, the design space of such complex models has significantly exploded due to the combinatorial challenges of heterogeneous computation kernels and intricate data dependency (i.e., intra-and inter-snapshot data dependency).This makes the computations of DGNN hard to scale, posing significant challenges in parallelism, data reuse, and communication.To address this challenge, we propose DiTile-DGNN, an efficient accelerator for large-scale DGNN execution.The proposed DiTile-DGNN consists of a redundancy-free parallelism strategy, workload balance optimization, and a reconfigurable accelerator architecture.Specifically, we propose a redundancy-free framework that can efficiently find an efficient parallelism strategy that can fully eliminate the data redundancy between graph snapshots while minimizing the communication complexity.Additionally, we propose a workload balance optimization for DGNN models to enhance resource utilization and eliminate synchronization overhead between snapshots.Lastly, we propose a reconfigurable accelerator architecture, with a flexible interconnect, that can be dynamically configured in support of various DGNN dataflows.Our simulations demonstrate that DiTile-DGNN achieves 48.4%, 56.1%, 23.2%, and 36.1% reductions in execution time and 83.4%, 84.0%, 75.6%, and 71.4% improvements in energy efficiency compared to state-of-the-art accelerators, including ReaDy [20], DGNN-Booster [8], RACE [51], and MEGA [12], on average across multiple DGNN datasets. Hao Zheng 0005, Ahmed Louri |
ISCA | 2 |
| 2025 | Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkabstractSparse Matrix Dense Matrix Multiplication (SpMM) is a fundamental computation kernel across various domains, including scientific computing, machine learning, and graph processing.Despite extensive research, existing approaches optimize SpMM using loop transformations and linear algebra principles, which (1) poorly handle unstructured sparsity patterns, (2) rely on empirical methods to explore data reuse opportunities, and (3) enforce rigid coordinate alignment, compromising data locality.In this paper, we demonstrate that these limitations stem from the fundamental matrix representation and traditional dataflows of SpMM (e.g., inner-product, outer-product, and Gustavson).We propose Aquila, a graph transformation framework that reformulates SpMM computations as a graph optimization problem, leveraging graph theory to reinterpret tiling and dataflow.First, on the theoretical side, we introduce vertex decomposition and adaptive depth traversal (ADT) to enable non-contiguous tiling, where nonzero elements from discontinuous rows and columns are clustered by connectivity rather than following matrix dimensionality.This approach quantifies data reuse and improves data locality beyond traditional loop transformations while maintaining output equivalence.Second, on the algorithm side, we develop a pull-after-push (PaP) dataflow that simultaneously enhances the dense matrix data reuse while eliminating synchronization issues in output matrix accumulation.Third, building on our theoretical approach and dataflow, we present a versatile accelerator architecture that handles a variety of SpMM kernels with diverse data sizes and sparsity patterns in a unified architecture.Additionally, we introduce a bidirectional fiber tree (BFT) format to support the proposed graph-oriented dataflow in contrast to traditional column or row-major access.Evaluation across diverse sparse datasets shows Aquila achieves speedups of 4.3×, 3.4×, 3.7×, 2.9×, and 2.7× in execution time and up to 4.8× * Both authors contributed equally to this research. Amir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye, Fan Yao 0001, Hao Zheng 0005 |
MICRO | 6 |
| 2025 | GAMMA: Gated Multi-hop Message Passing for Homophily-Agnostic Node Representation in GNNsabstractThe success of Graph Neural Networks (GNNs) leverages the homophily principle, where connected nodes share similar features and labels. However, this assumption breaks down in heterophilic graphs, where same-class nodes are often distributed across distant neighborhoods rather than immediate connections. Recent attempts expand the receptive field through multi-hop aggregation schemes that explicitly preserve intermediate representations from each hop distance. While effective at capturing heterophilic patterns, these methods require separate weight matrices per hop and feature concatenation, causing parameters to scale linearly with hop count. This leads to high computational complexity and GPU memory consumption. We propose Gated Multi-hop Message Passing (GAMMA), where nodes assess how relevant the aggregated information is from their k-hop neighbors. This assessment occurs through multiple refinement steps where the node compares each hop's embedding with its current representation, allowing it to focus on the most informative hops. During the forward pass, GAMMA finds the optimal mix of multi-hop information local to each node using a single feature vector without needing separate representations for each hop, thereby maintaining dimensionality comparable to single hop GNNs. In addition, we propose a weight sharing scheme that leverages a unified transformation for aggregated features from multiple hops so the global heterophilic patterns specific to each hop are learned during training. As such, GAMMA captures both global (per-hop) and local (per-node) heterophily patterns without high computation and memory overhead. Experiments show GAMMA matches or exceeds state-of-the-art heterophilic GNN accuracy, achieving up to $\approx20\times$ faster inference. Our code is publicly available at \url{https://github.com/amir-ghz/GAMMA}. Amir Ghazizadeh Ahsaei, Rickard Ewetz, Hao Zheng 0005 |
NeurIPS | 3 |
| 2024 | Towards Area-Efficient Path-Based In-Memory Computing using Graph IsomorphismsabstractIn-memory computing has attracted significant attention due to its potential to alleviate the issues caused by the von Neumann bottleneck. Path-based computing is a recently proposed in-memory computing paradigm for evaluating Boolean functions using nanoscale crossbars. Unlike state-of-the-art paradigms that use expensive WRITE operations to execute functions, path-based computing only relies on READ operations, which translates into benefits of low power consumption and low computational delay. Unfortunately, path-based computing comes with the penalty of substantial area overhead. In this paper, we introduce the ISO framework, a hardware-software solution for minimizing the area overhead of path-based computing systems. The framework is based on mapping computation to in-memory kernels using an intermediate k-LUT representation. The k-LUTs facilitate reusing hardware resources that realize the same computational structures. The reuse is performed by detecting identical subfunctions using isomorphic graphs. We also present a set of program instruction and scheduling algorithms to facilitate the hardware reuse. We have evaluated our proposed ISO framework on the 10 ISCAS85 benchmarks. Our experimental evaluation indicates that our proposed architecture improves energy consumption, latency, and area by $1.30\times, 76.59\times$, and $2.79\times$ on the average compared with previous state-of-the-art methods for path-based computing. Sven Thijssen, Muhammad Rashedul Haq Rashed, Hao Zheng 0005, Sumit Kumar Jha 0001, Rickard Ewetz |
ASPDAC | 3 |
| 2024 | VITA: ViT Acceleration for Efficient 3D Human Mesh Recovery via Hardware-Algorithm Co-DesignabstractVision Transformers (ViTs) have emerged as a promising solution to enable efficient 3D Human Mesh Recovery (HMR) in augmented and virtual reality (AR/VR) applications. Despite many advancements in algorithm design, it remains a challenge to efficiently accelerate ViT-based HMR due to high computational complexity, substantial memory footprint, and compromised data locality. In this paper, we propose VITA, a hardware and algorithm co-design framework for ViT-based HMR with improved performance and energy efficiency. Specifically, on the algorithm side, we propose an average pooling model to replace conventional multi-head attention, which is further optimized with improved data locality. On the hardware side, we propose an accelerator architecture that can efficiently support various dataflows and computations demanded by pooling, normalization, and convolution operations. We evaluate the proposed VITA, and the evaluation result shows that the proposed VITA design can achieve 5.05× and 69.12× speedups on average over the state-of-the-art GPUs and CPUs on HMR tasks. Shilin Tian, Chase Szafranski, Fan Yao 0001, Ahmed Louri, Chen Chen 0001, Hao Zheng 0005 |
DAC | 7 |
| 2024 | EGMA: Enhancing Data Reuse and Workload Balancing in Message Passing GNN Acceleration via Gram Matrix OptimizationabstractGraph Neural Networks (GNNs) have been widely used to handle intricate graph-related problems, in which complex vertex and edge operations are performed in the form of message passing between vertices. Such complex GNN operations are highly dependent on the graph structure and can no longer be characterized as sparse-dense or general matrix multiplications. Consequently, current matrix-based data reuse and workload balancing optimizations have limited applicability to Message Passing-based GNN acceleration. In this paper, we leverage the mathematical insights from Gram Matrix to simultaneously exploit data reuse and workload balancing opportunities for message passing-based GNN accelerations. Upon this insight, we further propose a novel accelerator, named EGMA, that can efficiently facilitate a wide range of GNN models with improved data reuse and workload balance. Consequently, EGMA can achieve performance speedup by 1.57×, 1.72×, and 1.43× and energy reduction by 38.19%, 34.02%, and 24.54% on average compared to Betty, FlowGNN, and ReGNN, respectively. Fangzhou Ye, Lingxiang Yin, Amir Ghazizadeh Ahsaei, Hao Zheng 0005 |
DAC | 4 |
| 2024 | CircuitSeer: RTL Post-PnR Delay Prediction via Coupling Functional and Structural RepresentationabstractRegister transfer level (RTL) optimization is a critical design phase that ensures timing closure and performance. Although machine learning (ML) has been utilized to quickly predict post-synthesis delay metrics, estimating post-place and route (PnR) delay remains a significant challenge. This is due to the distinct functionality-preserving characteristics of logic synthesis and the structure-dependent aspect of physical design. Furthermore, Logic Synthesis heavily restructures the netlist, resulting in substantial structural disparities that hinder capturing the post-synthesis netlist structure. Sanjay Gandham, Joe Walston, Sourav Samanta, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin, Stelios Diamantidis |
ICCAD | 5 |
| 2024 | Aurora: A Versatile and Flexible Accelerator for Graph Neural NetworksabstractGraph Neural Networks (GNNs) are pervasive across many application domains, driven by the growing demands to comprehend non-euclidean data. However, it remains a challenge to efficiently accelerate GNN applications due to varying graph data structures and diversified models. The non-euclidean data structure exhibits large variations in vertex and edge entities, and complex GNN models involve various computation characteristics and dataflows across different execution phases such as aggregation and vertex update. To address such diverse computation and communication demands, prior accelerators managed to tailor the compute engine for each GNN execution phase. Unfortunately, such limited adaptability could result in resource under-utilization and additional data movement.In this paper, we propose a versatile and flexible GNN accelerator that can efficiently facilitate the execution of various GNN models. The proposed architecture can be dynamically configured to multiple sub-accelerators, and each could be optimized for different GNN execution stages, including edge update, aggregation, and vertex update. To achieve this, the proposed architecture consists of four salient designs - a flexible processing element (PE) architecture, a unified Network-on-Chip (NoC), a degree-aware mapping, and a partitioning heuristic. The proposed unified processing element can be configured to support various fundamental GNN computations, such as vector-vector multiplication, vector-matrix multiplication, and scalar operations. The proposed NoC design can be dynamically configured to align with varying graph connectivity by dynamically bridging long-distance communications. The degree-aware mapping is proposed to avoid the unbalanced communication problem caused by high-degree vertices at the aggregation phase. Lastly, the proposed partitioning approach could efficiently allocate the computing resources to different GNN execution phases, improving the inter-phase parallelism. As such, the proposed design can increase hardware utilization with much-improved performance and energy efficiency. Simulation results show that our proposed Aurora accelerator achieves 85%, 66%, 47%, 28%, and 38% execution time reduction, and 89%, 77%, 42%, 69%, and 71% energy consumption reduction on average of multiple GNN dataset when compared to state-of-the-art accelerators HyGCN [1], AWB-GCN [2], GCNAX [3], ReGNN [4], and FlowGNN [5], respectively. Hao Zheng 0005, Ahmed Louri |
IPDPS | 2 |
| 2024 | MetaLeak: Uncovering Side Channels in Secure Processor Architectures Exploiting MetadataabstractMicroarchitectural side channels raise severe security concerns. Recent studies indicate that microarchitecture security should be examined holistically (rather than separately) in systems. Although the effects of performance optimizations on side channels are widely studied, the impacts of integrating security mechanisms intended for other threats on microarchitecture security are not well explored. In this paper, we perform the first side channel exploration in secure processor architectures that offer data confidentiality and integrity protection through hardware. We investigate microarchitecture security in the design space of secure processors and identify unique properties in the underlying metadata management schemes, which can be leveraged for new information leakage attacks. We present MetaLeak, an end-to-end side channel attack framework that exploits timing variations due to metadata maintenance to exfiltrate program secrets in secure processors. Particularly, we present two variants of the attack: MetaLeak-T that exploits the sharing of integrity tree metadata, and MetaLeak-C that manipulates counter metadata states. Our evaluation first shows highly accurate covert communication using the security metadata that can operate across cores and sockets without explicit data sharing. We further perform extensive side channel case studies on state-of-the-art secure architecture designs as well as the SGX processors. Our results show that MetaLeak can successfully exfiltrate program secrets (up to $97 \%$ accuracy) from image-processing application and cryptographic software running in enclave. Our study indicates that the fundamental metadata mechanism is the root cause of the leakage, which necessitates the use of leakage-taming techniques in future secure processors. This work highlights the need to synergistically understand microarchitecture security, as new security mechanisms are integrated. Md Hafizul Islam Chowdhuryy, Hao Zheng 0005, Fan Yao 0001 |
ISCA | 2 |
| 2024 | SCALE: A Structure-Centric Accelerator for Message Passing Graph Neural NetworksabstractMessage passing paradigm has been widely used in developing complex Graph Neural Network (GNN) models, allowing for concise representations of edge and vertex-wise operations. Despite its pivotal role in theoretical advancement, the respective expression of edge and vertex operations, along with evolving GNN variants and datasets, has inevitably led to enormous computational complexity due to heterogeneous computation kernels. In particular, such inconsistent computation characteristics present new challenges in leveraging intermediate data reuse, ensuring both edge and vertex-wise workload balance, and sustaining system scalability. In this paper, we propose a structurecentric accelerator, SCALE, that can support a variety of message passing GNN models with improved parallelism, data reuse, and scalability. The central idea is to find latent similarities among GNN primitives such as shared dataflow structure, rather than strictly adhering to heterogeneous model structure. This serves as a hinge to homogenize inconsistencies in various GNN computation kernels. To accomplish this concept, SCALE consists of three unique designs, a novel systolic array-like architecture, a degree and vertex-aware scheduling, and a coherent dataflow tailored for fused graph and neural operations. The proposed systolic array-like architecture can support varying dataflows such as all-reduce, of distinct GNN operations improving parallelism, data reuse, and throughput. The degree and vertex-aware scheduling can remedy the workload imbalance encountered in vertex and edge-wise operations. Moreover, the proposed dataflow can unify the data movement of both graph and neural operators without extra communication and storage overheads. Our simulation results show that SCALE achieves 1.82× speedup and 38.9% energy reduction on average over the state-of-the-art GNN accelerators [1]–[4]. Lingxiang Yin, Sanjay Gandham, Mingjie Lin, Hao Zheng 0005 |
MICRO | 4 |
| 2024 | Versa-DNN: A Versatile Architecture Enabling High-Performance and Energy-Efficient Multi-DNN AccelerationabstractEmerging applications utilize numerous Deep Neural Networks (DNNs) to address multiple tasks simultaneously. As these applications continue to expand, there is a growing need for off-chip memory access optimization and innovative architectures that can adapt to diverse computation, memory, and communication requirements of various DNN models. To address these challenges, we propose Versa-DNN, a versatile DNN accelerator that can provide efficient computation, memory, and communication support for the simultaneous execution of multiple DNNs. Versa-DNN features three unique designs: a flexible off-chip memory access optimization strategy, adaptable communication fabrics, and a communication and computational aware scheduling algorithm. The proposed off-chip memory optimization strategy can improve performance and energy efficiency by increasing hardware utilization, eliminating excess data duplication, and reducing off-chip memory accesses. The adaptable communication fabrics consist of distributed buffers, processing elements, and a flexible Network-on-Chip (NoC), which can dynamically morph and fission to support distinct communication and computation needs for simultaneously running DNN models. Furthermore, the proposed scheduling policy manages the simultaneous execution of multiple DNN models with improved performance and energy efficiency. Simulation results using several DNN models, show that the proposed Versa-DNN architecture achieves 41%, 238%, 392% throughput speedup and 30%, 59%, 63% energy reduction on average for different workloads when compared to state-of-the-art accelerators such as Planaria, Herald, and AI-MT, respectively. Hao Zheng 0005, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Morph-GCNX: A Universal Architecture for High-Performance and Energy-Efficient Graph Convolutional Network AccelerationabstractWhile current Graph Convolutional Networks (GCNs) accelerators have achieved notable success in a wide range of application domains, these GCN accelerators can not support various intra- and inter- GCN dataflows or adapt to diverse GCN applications. In this paper, we propose Morph-GCNX, a flexible GCN accelerator architecture for high-performance and energy-efficient GCN execution. The proposed design consists of a flexible Processing Element (PE) array that can be partitioned at runtime and adapt to the computational needs of different layers within a GCN or multiple concurrent GCNs. The proposed Morph-GCNX also consists of a morphable interconnection design to support a wide range of GCN dataflows with various parallelization and data reuse strategies for GCN execution. We also propose a hardware-application co-exploration technique that explores the GCN and hardware design spaces to identify the best PE partition, workload allocation, dataflow, and interconnection configurations, with the goal of improving overall performance and energy. Simulation results show that the proposed Morph-GCNX architecture achieves 18.8×, 2.9×, 1.9×, 1.8×, and 2.5× better performance, reduces DRAM accesses by a factor of 10.8×, 3.7×, 2.2×, 2.5×, and 1.3×, and improves energy consumption by 13.2×, 5.6×, 2.1×, 2.5×, and 1.3×, as compared to prior designs including HyGCN, AWB-GCN, LW-GCN, GCoD, and GCNAX, respectively. Ke Wang 0030, Hao Zheng 0005, Ahmed Louri |
IEEE Trans. Sustain. Comput. | 2 |
| 2023 | Venus: A Versatile Deep Neural Network Accelerator Architecture Design for Multiple ApplicationsabstractDeep Neural Network (DNN) applications are pervasive. However as demands for these applications continue to increase, so is the challenges for designing flexible and scalable architectures for multi-application implementation. Such accelerators require innovative architecture with flexible Network-on-Chips (NoCs), parallelism exploitation, and better on-chip memory organization to adequately support the diverse computation, memory, and communication needs. In this paper, we propose Venus, a versatile DNN accelerator design that can provide efficient communication and computation support for multi-applications. Venus is a tile-based architecture with a distributed buffer where each tile consists of an array of processing elements (PEs) and a portion of the distributed buffer. The other salient feature of Venus is a flexible Network-on-Chip (NoC) that can dynamically adapt to the communication needs of various running applications thus maximizing data reuse, reducing DRAM accesses, and supporting multiple dataflows with an overall aim of better execution time and better energy efficiency. Simulation results show that our proposed Venus design outperforms state-of-art accelerators (NVDLA [1], ShiDianNao [2], Eyeriss [3], Planaria [4], Simba [5]). Hao Zheng 0005, Ahmed Louri |
DAC | 2 |
| 2023 | OCMGen: Extended Design Space Exploration with Efficient FPGA Memory InferenceabstractDeep learning applications demand high memory storage and computational power to operate on millions of parameters. Field Programmable Gate Arrays (FPGAs), with high compute resources and the ability to store data on-chip in their distributed memory components such as Block RAM (BRAM) and Ultra RAM (URAM), are good candidates to deploy such memory-intensive applications [1]. However, without careful tailoring of the hardware design for a target device, current synthesis tools (e.g., Xilinx Vivado) can severely underutilize these RAM primitives reducing the usable on-chip memory (OCM). Consequently, this forces the accelerator to perform more frequent expensive off-chip accesses, limiting its performance. Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin |
FCCM | 3 |
| 2023 | Exploring Architecture, Dataflow, and Sparsity for GCN Accelerators: A Holistic FrameworkabstractRecent years have seen an increasing number of Graph Convolutional Network (GCN) models employed in various real-world applications. However, designing efficient architectures for GCN acceleration remains challenging due to the varied sparsity across graph datasets. Despite significant efforts, very few of the existing works have considered a holistic view of the entire GCN accelerator design, and therefore, the dynamic interactions between architecture, dataflow (i.e., data reuse and parallelization strategies), and compression format are not well studied Lingxiang Yin, Jun Wang 0001, Hao Zheng 0005 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | SAGA: Sparsity-Agnostic Graph Convolutional Network Acceleration with Near-Optimal Workload BalanceabstractGraph Convolutional Networks (GCNs) have shown much promise in resolving sophisticated scientific problems with non-Euclidean data, such as traffic prediction, disease classification, and many others. However, the irregular sparsity of real-world graphs remains a major challenge toward efficient GCN acceleration. In this paper, we propose SAGA, a Sparsity-Agnostic Graph Convolutional Accelerator with near-optimal workload balance. Specifically, it consists of two unique features, an NZ-based scheduling, and a novel accelerator architecture. Unlike conventional GCN accelerators with uneven distribution of sparse matrix, the proposed NZ-based scheduling leverages the metadata encoded in the compression format to enable even distribution of sparse matrix at runtime, thus achieving near-optimal workload balancing. In addition, the proposed architecture, including a task scheduler, an accumulation table, and a partial row accumulation unit, can support the proposed NZ-based scheduling without data preprocessing and reformatting with low overheads. We prototyped the proposed design through FPGAs, and our evaluation results show that SAGA achieves up to$\mathbf{1.56}\times$speedup and$\mathbf{2.05}\times$energy savings on average as compared to the prior art [1]. Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin |
ICCAD | 3 |
| 2023 | Path-Based Processing using In-Memory Systolic Arrays for Accelerating Data-Intensive ApplicationsabstractThe next wave of scientific discovery is predicated on unleashing beyond-exascale simulation capabilities using in-memory computing. Path-based computing is a promising in-memory logic style for accelerating Boolean logic with deterministic precision. However, existing studies on path-based computing are limited to executing small combinational circuits. In this paper, we propose a framework called PSYS to accelerate data-intensive scientific computing applications using path-based in-memory systolic arrays. The approach leverages path-based computing for multiplying known constants with an unknown operand, which substantially reduces the computational complexity compared with general purpose multiplication of two unknown operands. The systolic arrays minimize data movement by storing the matrix elements using non-volatile memory and performing processing in-place. The framework decomposes unstructured computations to the systolic arrays while considering the non-regular computational patterns of the applications. Our experimental evaluations employ applications from the domains of engineering, physics, and mathematics. The experimental results demonstrate that compared with the state-of-the-art, the PSYS framework improves energy and latency by a factor of 101x and 23x, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICCAD | 4 |
| 2023 | ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible InterconnectsabstractLarge-scale deep learning models are widely deployed in many application domains with remarkable performance improvements. However, training these models with immense parameters calls for unprecedented computing and communication capabilities. Recently, chiplet-based architectures have shown much promise in scaling Deep Neural Network (DNN) inference, but their applications in the training phase remain unexplored and challenging. In this paper, we posit, beyond scaling computing capability, chiplet-based architectures could also be leveraged to enable new optimization opportunities for existing parallel training algorithms (e.g., Ring and Tree-based all-reduce). Specifically, we aim to explore a variety of topological characteristics, along with the interposer technology, to sustain the performance scaling of parallel training in chiplet-based systems. We propose ARIES, a versatile chiplet-based communication architecture supporting various parallel training algorithms using a flexible interconnect design. The proposed design can adapt to various collective operations such as reduce and gather across a wide diversity of training algorithms. Moreover, such flexibility is also leveraged to further enhance existing all-reduce algorithms depending on the latency and bandwidth requirements of the DNN model and dataset size. Simulation results show that the proposed ARIES can achieve up to 3.92× speedup in execution time and 38.8% reduction in Network-on-Chip (NoC) energy consumption when compared to prior work. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Ahmed Louri, Hao Zheng 0005 |
ICCAD | 4 |
| 2023 | Polyform: A Versatile Architecture for Multi-DNN Execution via Spatial and Temporal AccelerationabstractContemporary applications and cloud workloads often comprise multiple Deep Neural Network (Multi-DNN) models. These models exhibit significant variations in computation, memory, and communication characteristics. For such heterogeneous workloads, a static and rigid hardware accelerator can no longer provide efficient and high-performance execution. To this end, we propose a versatile accelerator, called Polyform, to support the concurrent execution of different DNN models with the goal of improving energy and performance efficiency. Specifically, Polyform features two unique designs from both hardware and scheduling standpoints. On the hardware level, we have designed a flexible interconnection network that facilitates the formation of multiple sub-accelerators. Our design allows for spatial resource partitioning, including bandwidth and computation, while also providing effective communication support for various parallelism choices. On the scheduling level, Polyform employs a novel two-stage Genetic Algorithm (GA) to explore and identify the optimal configurations such as task orders, partition size, dataflow styles (e.g., weight or output stationary), and bandwidth. Our simulation shows that Polyform achieves remarkable results compared to prior work, including up to 77.8% energy reduction and a 2.79× improvement in throughput as compared to prior work [1]–[3]. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Shilin Tian, Ahmed Louri, Hao Zheng 0005 |
ICCD | 5 |
| 2023 | FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential EquationsabstractPartial Differential Equations (PDEs) are widely employed to describe natural phenomena in many science and engineering fields. Many PDEs do not have analytical solutions, hence, numerical methods have become prevalent for approximating PDE solutions. The most widely used numerical method is the Finite Difference Method (FDM), which requires fine grids and high-precision numerical iterations that are both compute- and memory-intensive. PDE-solving accelerators have been proposed in the literature, however, they usually focus on specific types of PDEs with rigid grid sizes which limits their broader applicability. Besides, they rarely provided insight into the optimizations of parallel computing and data accesses for solving PDEs, which hinders further improvements in performance and energy efficiency. Hao Zheng 0005, Ke Wang 0030 |
ISCA | 3 |
| 2023 | GShuttle: Optimizing Memory Access Efficiency for Graph Convolutional Neural Network Accelerators
Ke Wang 0030, Hao Zheng 0005, Ahmed Louri |
J. Comput. Sci. Technol. | 3 |
| 2022 | AGAPE: Anomaly Detection with Generative Adversarial Network for Improved Performance, Energy, and Security in Manycore SystemsabstractThe security of manycore systems has become increasingly critical. In system-on-chips (SoCs), Hardware Trojans (HTs) manipulate the functionalities of the routing components to saturate the on-chip network, degrade performance, and result in the leakage of sensitive data. Existing HT detection techniques, including runtime monitoring and state-of-the-art learning-based methods, are unable to timely and accurately identify the implanted HTs, due to the increasingly dynamic and complex nature of on-chip communication behaviors. We propose AGAPE, a novel Generative Adversarial Network (GAN)-based anomaly detection and mitigation method against HTs for secured on-chip communication. AGAPE learns the distribution of the multivariate time series of a number of NoC attributes captured by on-chip sensors under both HT-free and HT-infected working conditions. The proposed GAN can learn the potential latent interactions among different runtime attributes concurrently, accurately distinguish abnormal attacked situations from normal SoC behaviors, and identify the type and location of the implanted HTs. Using the detection results, we apply the most suitable protection techniques to each type of detected HTs instead of simply isolating the entire HT-infected router, with the aim to mitigate security threats as well as reducing performance loss. Simulation results show that AGAPE enhances the HT detection accuracy by 19%, reduces network latency and power consumption by 39% and 30%, respectively, as compared to state-of-the-art security designs. Ke Wang 0030, Hao Zheng 0005, Yuan Li 0029, Ahmed Louri |
DATE | 2 |
| 2022 | Adapt-Flow: A Flexible DNN Accelerator Architecture for Heterogeneous Dataflow ImplementationabstractDeep neural networks (DNNs) have been widely applied to various application domains. DNN computation is memory and compute-intensive requiring excessive memory access and a large number of computations. To efficiently implement these applications, several data reuse and parallelism exploitation strategies, called dataflows, have been proposed. Studies have shown that many DNN applications benefit from a heterogeneous dataflow strategy where the dataflow type changes from layer to layer. Unfortunately, very few existing DNN architectures can simultaneously accommodate multiple dataflows due to their limited hardware flexibility. In this paper, we propose a flexible DNN accelerator architecture, called Adapt-Flow, which has the capability of supporting multiple dataflow selections for each DNN layer at runtime. Specifically, the proposed Adapt-Flow architecture consists of (1) a flexible interconnect, (2) a dataflow selection algorithm, and (3) a dataflow mapping technique. The flexible interconnect provides dynamic support for various traffic patterns required by different dataflows. The proposed dataflow selection algorithm selects the optimal dataflow strategy for a given DNN layer with the aim of much improved performance. And the dataflow mapping technique efficiently maps the dataflow amenable to the flexible interconnect. Simulation studies show that the proposed Adapt-Flow architecture reduces execution time by 46%, 78%, 26%, and energy consumption by 45%, 80%, 25% as compared to NVDLA, ShiDianNao, and Eyeriss respectively. Hao Zheng 0005, Ahmed Louri |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Ascend: A Scalable and Energy-Efficient Deep Neural Network Accelerator With Photonic InterconnectsabstractThe complexity and size of recent deep neural network (DNN) models have increased significantly in pursuit of high inference accuracy. Chiplet-based accelerator is considered a viable scaling approach to provide substantial computation capability and on-chip memory for efficient process of such DNN models. However, communication using metallic interconnects in prior chiplet-based accelerators poses a major challenge to system performance, energy efficiency, and scalability. Photonic interconnects can adequately support communication across chiplets due to features such as distance-independent latency, high bandwidth density, and high energy efficiency. Furthermore, the salient ease of broadcast property makes photonic interconnects suitable for DNN inference which often incurs prevalent broadcast communication. In this paper, we propose a scalable chiplet-based DNN accelerator with photonic interconnects named ASCEND. ASCEND introduces (1) a novel photonic network that supports seamless intra- and inter- chiplet broadcast communication, and flexible mapping of diverse convolution layers, and (2) a tailored dataflow that exploits the ease of broadcast property and maximizes parallelism by simultaneously processing computations with shared input data. Simulation results using multiple DNN models show that ASCEND achieves 71% and 67% reduction in execution time and energy consumption, respectively, as compared to other state-of-the-art chiplet-based DNN accelerators with metallic or photonic interconnects. Yuan Li 0029, Ke Wang 0030, Hao Zheng 0005, Ahmed Louri, Avinash Karanth |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | SGCNAX: A Scalable Graph Convolutional Neural Network Accelerator With Workload BalancingabstractGraph Convolutional Neural Networks (GCNs) have emerged as promising tools for graph-based machine learning applications. Given that GCNs are both compute- and memory-intensive, this constitutes a major challenge for the underlying hardware to efficiently process large-scale GCNs. In this paper, we introduce SGCNAX, a scalable GCN accelerator architecture for the high-performance and energy-efficient acceleration of GCNs. Unlike prior GCN accelerators that either employ limited loop optimization techniques, or determine the design variables based on random sampling, we systematically explore the loop optimization techniques for GCN acceleration and propose a flexible GCN dataflow that adapts to different GCN configurations to achieve optimal efficiency. We further propose two hardware-based techniques to address the workload imbalance problem caused by the unbalanced distribution of zeros in GCNs. Specifically, SGCNAX exploits an outer-product-based computation architecture that mitigates the intra-PE (Processing Elements) workload imbalance, and employs a group-and-shuffle approach to mitigate the inter-PE workload imbalance. Simulation results show that SGCNAX performs 9.2x, 1.6x and 1.2x better, and reduces DRAM accesses by a factor of 9.7x, 2.9x and 1.2x compared to HyGCN, AWB-GCN, and GCNAX, respectively. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | SecureNoC: A Learning-Enabled, High-Performance, Energy-Efficient, and Secure On-Chip Communication Framework DesignabstractWe propose SecureNoC, a learning-based framework to enhance NoC security against Hardware Trojan (HT) attacks while holistically improving performance and power. The proposed framework enhances NoC security with several architectural innovations, namely a per-router HT detector, multi-function bypass channels (MBCs), and a lightweight data encryption design. Specifically, the threat detector uses an artificial neural network for runtime HT detection with high accuracy. The MBCs consist of a router bypass route and reconfigurable channel buffers which can efficiently isolate malicious nodes and reduce power consumption. The proposed data encryption design adapts to diverse traffic patterns and dynamically deploys novel lightweight encryption techniques for desired security goals with improved latency. Additionally, to balance the trade-offs and handle the dynamic interactions of the proposed dynamic designs, a proactive deep-Q-learning (DQL) control policy is proposed to simultaneously provide optimized NoC security, performance, and power consumption. Simulation studies using PARSEC benchmarks show that the proposed SecureNoC achieves 36% higher HT detection accuracy over state-of-the-art NoC security techniques while reducing network latency by 39% and energy consumption by 46%. Ke Wang 0030, Hao Zheng 0005, Yuan Li 0029, Ahmed Louri |
IEEE Trans. Sustain. Comput. | 2 |
| 2021 | Adapt-NoC: A Flexible Network-on-Chip Design for Heterogeneous Manycore ArchitecturesabstractThe increased computational capability in heterogeneous manycore architectures facilitates the concurrent execution of many applications. This requires, among other things, a flexible, high-performance, and energy-efficient communication fabric capable of handling a variety of traffic patterns needed for running multiple applications at the same time. Such stringent requirements are posing a major challenge for current Network-on-Chips (NoCs) design. In this paper, we propose Adapt-NoC, a flexible NoC architecture, along with a reinforcement learning (RL)-based control policy, that can provide efficient communication support for concurrent application execution. Adapt-NoC can dynamically allocate several disjoint regions of the NoC, called subNoCs, with different sizes and locations for the concurrently running applications. Each of the dynamically-allocated subNoCs is capable of adapting to a given topology such as a mesh, cmesh, torus, or tree thus tailoring the topology to satisfy application's needs in terms of performance and power consumption. Moreover, we explore the use of RL to design an efficient control policy which optimizes the subNoC topology selection for a given application. As such, Adapt-NoC can not only provide several topology choices for concurrently running applications, but can also optimize the selection of the most suitable topology for a given application with the aim of improving performance and energy efficiency. We evaluate Adapt-NoC using both GPU and CPU benchmark suites. Simulation results show that the proposed Adapt-NoC can achieve up to 34% latency reduction, 10% overall execution time reduction and 53% NoC energy-efficiency improvement when compared to prior work. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
HPCA | 1 |
| 2020 | A Versatile and Flexible Chiplet-based System Design for Heterogeneous Manycore ArchitecturesabstractHeterogeneous manycore architectures are deployed to simultaneously run multiple and diverse applications. This requires various computing capabilities (CPUs, GPUs, and accelerators), and an efficient network-on-chip (NoC) architecture to concurrently handle diverse application communication behavior. However, supporting the concurrent communication requirements of diverse applications is challenging due to the dynamic application mapping, the complexity of handling distinct communication patterns and limited on-chip resources. In this paper, we propose Adapt-NoC, a versatile and flexible NoC architecture for chiplet-based manycore architectures, consisting of adaptable routers and links. Adapt-NoC can dynamically allocate disjoint regions of the NoC, called subNoCs, for concurrently-running applications, each of which can be optimized for different communication behavior. The adaptable routers and links are capable of providing various subNoC topologies, satisfying different latency and bandwidth requirements of various traffic patterns (e.g. all-to-all, one-to-many). Full system simulation shows that AdaptNoC can achieve 31% latency reduction, 24% energy saving and 10% execution time reduction on average, when compared to prior designs. Hao Zheng 0005, Ke Wang 0030, Ahmed Louri |
DAC | 1 |
| 2019 | An Energy-Efficient Network-on-Chip Design using Reinforcement LearningabstractThe design space for energy-efficient Network-on-Chips (NoCs) has expanded significantly comprising a number of techniques. The simultaneous application of these techniques to yield maximum energy efficiency requires the monitoring of a large number of system parameters which often results in substantial engineering efforts and complicated control policies. This motivates us to explore the use of reinforcement learning (RL) approach that automatically learns an optimal control policy to improve NoC energy efficiency. First, we deploy power-gating (PG) and dynamic voltage and frequency scaling (DVFS) to simultaneously reduce both static and dynamic power. Second, we use RL to automatically explore the dynamic interactions among PG, DVFS, and system parameters, learn the critical system parameters contained in the router and cache, and eventually evolve optimal per-router control policies that significantly improve energy efficiency. Moreover, we introduce an artificial neural network (ANN) to efficiently implement the large state-action table required by RL. Simulation results using PARSEC benchmark show that the proposed RL approach improves power consumption by 26%, while improving system performance by 7%, as compared to a combined PG and DVFS design without RL. Additionally, the ANN design yields 67% area reduction, as compared to a conventional RL implementation. Hao Zheng 0005, Ahmed Louri |
DAC | 1 |