EDBT 2026 Demo / reviewers in the wild / expert
Amir Ghazizadeh Ahsaei
dblp:364/0256 · also Amir Ghazizadeh
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0008-3499-9197ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TensorPrism: Rethinking Sparse High-Order Tensor Acceleration via Co-Occurrence Graph
Fangzhou Ye, Shilin Tian, Amir Ghazizadeh Ahsaei, Hao Zheng 0005 |
ISCA | 3 |
| 2025 | Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkabstractSparse Matrix Dense Matrix Multiplication (SpMM) is a fundamental computation kernel across various domains, including scientific computing, machine learning, and graph processing.Despite extensive research, existing approaches optimize SpMM using loop transformations and linear algebra principles, which (1) poorly handle unstructured sparsity patterns, (2) rely on empirical methods to explore data reuse opportunities, and (3) enforce rigid coordinate alignment, compromising data locality.In this paper, we demonstrate that these limitations stem from the fundamental matrix representation and traditional dataflows of SpMM (e.g., inner-product, outer-product, and Gustavson).We propose Aquila, a graph transformation framework that reformulates SpMM computations as a graph optimization problem, leveraging graph theory to reinterpret tiling and dataflow.First, on the theoretical side, we introduce vertex decomposition and adaptive depth traversal (ADT) to enable non-contiguous tiling, where nonzero elements from discontinuous rows and columns are clustered by connectivity rather than following matrix dimensionality.This approach quantifies data reuse and improves data locality beyond traditional loop transformations while maintaining output equivalence.Second, on the algorithm side, we develop a pull-after-push (PaP) dataflow that simultaneously enhances the dense matrix data reuse while eliminating synchronization issues in output matrix accumulation.Third, building on our theoretical approach and dataflow, we present a versatile accelerator architecture that handles a variety of SpMM kernels with diverse data sizes and sparsity patterns in a unified architecture.Additionally, we introduce a bidirectional fiber tree (BFT) format to support the proposed graph-oriented dataflow in contrast to traditional column or row-major access.Evaluation across diverse sparse datasets shows Aquila achieves speedups of 4.3×, 3.4×, 3.7×, 2.9×, and 2.7× in execution time and up to 4.8× * Both authors contributed equally to this research. Amir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye, Fan Yao 0001, Hao Zheng 0005 |
MICRO | 1 |
| 2025 | GAMMA: Gated Multi-hop Message Passing for Homophily-Agnostic Node Representation in GNNsabstractThe success of Graph Neural Networks (GNNs) leverages the homophily principle, where connected nodes share similar features and labels. However, this assumption breaks down in heterophilic graphs, where same-class nodes are often distributed across distant neighborhoods rather than immediate connections. Recent attempts expand the receptive field through multi-hop aggregation schemes that explicitly preserve intermediate representations from each hop distance. While effective at capturing heterophilic patterns, these methods require separate weight matrices per hop and feature concatenation, causing parameters to scale linearly with hop count. This leads to high computational complexity and GPU memory consumption. We propose Gated Multi-hop Message Passing (GAMMA), where nodes assess how relevant the aggregated information is from their k-hop neighbors. This assessment occurs through multiple refinement steps where the node compares each hop's embedding with its current representation, allowing it to focus on the most informative hops. During the forward pass, GAMMA finds the optimal mix of multi-hop information local to each node using a single feature vector without needing separate representations for each hop, thereby maintaining dimensionality comparable to single hop GNNs. In addition, we propose a weight sharing scheme that leverages a unified transformation for aggregated features from multiple hops so the global heterophilic patterns specific to each hop are learned during training. As such, GAMMA captures both global (per-hop) and local (per-node) heterophily patterns without high computation and memory overhead. Experiments show GAMMA matches or exceeds state-of-the-art heterophilic GNN accuracy, achieving up to $\approx20\times$ faster inference. Our code is publicly available at \url{https://github.com/amir-ghz/GAMMA}. Amir Ghazizadeh Ahsaei, Rickard Ewetz, Hao Zheng 0005 |
NeurIPS | 1 |
| 2024 | EGMA: Enhancing Data Reuse and Workload Balancing in Message Passing GNN Acceleration via Gram Matrix OptimizationabstractGraph Neural Networks (GNNs) have been widely used to handle intricate graph-related problems, in which complex vertex and edge operations are performed in the form of message passing between vertices. Such complex GNN operations are highly dependent on the graph structure and can no longer be characterized as sparse-dense or general matrix multiplications. Consequently, current matrix-based data reuse and workload balancing optimizations have limited applicability to Message Passing-based GNN acceleration. In this paper, we leverage the mathematical insights from Gram Matrix to simultaneously exploit data reuse and workload balancing opportunities for message passing-based GNN accelerations. Upon this insight, we further propose a novel accelerator, named EGMA, that can efficiently facilitate a wide range of GNN models with improved data reuse and workload balance. Consequently, EGMA can achieve performance speedup by 1.57×, 1.72×, and 1.43× and energy reduction by 38.19%, 34.02%, and 24.54% on average compared to Betty, FlowGNN, and ReGNN, respectively. Fangzhou Ye, Lingxiang Yin, Amir Ghazizadeh Ahsaei, Hao Zheng 0005 |
DAC | 3 |
| 2023 | ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible InterconnectsabstractLarge-scale deep learning models are widely deployed in many application domains with remarkable performance improvements. However, training these models with immense parameters calls for unprecedented computing and communication capabilities. Recently, chiplet-based architectures have shown much promise in scaling Deep Neural Network (DNN) inference, but their applications in the training phase remain unexplored and challenging. In this paper, we posit, beyond scaling computing capability, chiplet-based architectures could also be leveraged to enable new optimization opportunities for existing parallel training algorithms (e.g., Ring and Tree-based all-reduce). Specifically, we aim to explore a variety of topological characteristics, along with the interposer technology, to sustain the performance scaling of parallel training in chiplet-based systems. We propose ARIES, a versatile chiplet-based communication architecture supporting various parallel training algorithms using a flexible interconnect design. The proposed design can adapt to various collective operations such as reduce and gather across a wide diversity of training algorithms. Moreover, such flexibility is also leveraged to further enhance existing all-reduce algorithms depending on the latency and bandwidth requirements of the DNN model and dataset size. Simulation results show that the proposed ARIES can achieve up to 3.92× speedup in execution time and 38.8% reduction in Network-on-Chip (NoC) energy consumption when compared to prior work. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Ahmed Louri, Hao Zheng 0005 |
ICCAD | 2 |
| 2023 | Polyform: A Versatile Architecture for Multi-DNN Execution via Spatial and Temporal AccelerationabstractContemporary applications and cloud workloads often comprise multiple Deep Neural Network (Multi-DNN) models. These models exhibit significant variations in computation, memory, and communication characteristics. For such heterogeneous workloads, a static and rigid hardware accelerator can no longer provide efficient and high-performance execution. To this end, we propose a versatile accelerator, called Polyform, to support the concurrent execution of different DNN models with the goal of improving energy and performance efficiency. Specifically, Polyform features two unique designs from both hardware and scheduling standpoints. On the hardware level, we have designed a flexible interconnection network that facilitates the formation of multiple sub-accelerators. Our design allows for spatial resource partitioning, including bandwidth and computation, while also providing effective communication support for various parallelism choices. On the scheduling level, Polyform employs a novel two-stage Genetic Algorithm (GA) to explore and identify the optimal configurations such as task orders, partition size, dataflow styles (e.g., weight or output stationary), and bandwidth. Our simulation shows that Polyform achieves remarkable results compared to prior work, including up to 77.8% energy reduction and a 2.79× improvement in throughput as compared to prior work [1]–[3]. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Shilin Tian, Ahmed Louri, Hao Zheng 0005 |
ICCD | 2 |