VLDB 2026 Research / reviewers in the wild / expert
Lingxiang Yin
dblp:348/4782
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0005-9279-9220ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling Graph Neural Network Training via Geometric OptimizationabstractWafer-scale computing has emerged as an alternative solution to sustain performance scaling in the post-Moore era, driven by recent technology advancements such as chiplet integration. This enables considerable computing and storage capabilities on a single chip, making it capable of accommodating large machine learning models and datasets. Recent efforts have heralded the promise of wafer-scale architectures for deep learning inference and training. However, scaling the training of Graph Neural Networks in wafer-scale architecture remains a challenge and is relatively unexplored due to irregularities in gradient propagation as well as physical constraints from flat on-chip topologies. In this paper, we propose Aster, a topology-aware framework designed to efficiently support GNN training on arbitrary wafer-scale architectures. The proposed framework, as opposed to the current application or topology-specific heuristics, can be generalized to support any network topology and irregular GNN datasets. Specifically, we mathematically formulate commonly-seen network topologies in their geometric representation and prioritize communication efficiency during GNN workload partitioning and mapping. Based on the geometric representation, we propose a quadratic assignment problem solver to efficiently map irregular dataflows to a flat topology with reduced communication distance. The simulation results show that Aster can achieve performance speedup by$2.91 \times, 1.50 \times, 1.84 \times$, and$1.58 \times$in Mesh and speedup by$3.84 \times, 1.56 \times, 2.05 \times$, and$1.49 \times$in Torus on average compared to Mini-cut [1], ScalaGraph [2], ChunkV [3], and Chunk-E [4], respectively. Fangzhou Ye, Lingxiang Yin, Hao Zheng 0005 |
HPCA | 2 |
| 2025 | Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkabstractSparse Matrix Dense Matrix Multiplication (SpMM) is a fundamental computation kernel across various domains, including scientific computing, machine learning, and graph processing.Despite extensive research, existing approaches optimize SpMM using loop transformations and linear algebra principles, which (1) poorly handle unstructured sparsity patterns, (2) rely on empirical methods to explore data reuse opportunities, and (3) enforce rigid coordinate alignment, compromising data locality.In this paper, we demonstrate that these limitations stem from the fundamental matrix representation and traditional dataflows of SpMM (e.g., inner-product, outer-product, and Gustavson).We propose Aquila, a graph transformation framework that reformulates SpMM computations as a graph optimization problem, leveraging graph theory to reinterpret tiling and dataflow.First, on the theoretical side, we introduce vertex decomposition and adaptive depth traversal (ADT) to enable non-contiguous tiling, where nonzero elements from discontinuous rows and columns are clustered by connectivity rather than following matrix dimensionality.This approach quantifies data reuse and improves data locality beyond traditional loop transformations while maintaining output equivalence.Second, on the algorithm side, we develop a pull-after-push (PaP) dataflow that simultaneously enhances the dense matrix data reuse while eliminating synchronization issues in output matrix accumulation.Third, building on our theoretical approach and dataflow, we present a versatile accelerator architecture that handles a variety of SpMM kernels with diverse data sizes and sparsity patterns in a unified architecture.Additionally, we introduce a bidirectional fiber tree (BFT) format to support the proposed graph-oriented dataflow in contrast to traditional column or row-major access.Evaluation across diverse sparse datasets shows Aquila achieves speedups of 4.3×, 3.4×, 3.7×, 2.9×, and 2.7× in execution time and up to 4.8× * Both authors contributed equally to this research. Amir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye, Fan Yao 0001, Hao Zheng 0005 |
MICRO | 2 |
| 2024 | EGMA: Enhancing Data Reuse and Workload Balancing in Message Passing GNN Acceleration via Gram Matrix OptimizationabstractGraph Neural Networks (GNNs) have been widely used to handle intricate graph-related problems, in which complex vertex and edge operations are performed in the form of message passing between vertices. Such complex GNN operations are highly dependent on the graph structure and can no longer be characterized as sparse-dense or general matrix multiplications. Consequently, current matrix-based data reuse and workload balancing optimizations have limited applicability to Message Passing-based GNN acceleration. In this paper, we leverage the mathematical insights from Gram Matrix to simultaneously exploit data reuse and workload balancing opportunities for message passing-based GNN accelerations. Upon this insight, we further propose a novel accelerator, named EGMA, that can efficiently facilitate a wide range of GNN models with improved data reuse and workload balance. Consequently, EGMA can achieve performance speedup by 1.57×, 1.72×, and 1.43× and energy reduction by 38.19%, 34.02%, and 24.54% on average compared to Betty, FlowGNN, and ReGNN, respectively. Fangzhou Ye, Lingxiang Yin, Amir Ghazizadeh Ahsaei, Hao Zheng 0005 |
DAC | 2 |
| 2024 | CircuitSeer: RTL Post-PnR Delay Prediction via Coupling Functional and Structural RepresentationabstractRegister transfer level (RTL) optimization is a critical design phase that ensures timing closure and performance. Although machine learning (ML) has been utilized to quickly predict post-synthesis delay metrics, estimating post-place and route (PnR) delay remains a significant challenge. This is due to the distinct functionality-preserving characteristics of logic synthesis and the structure-dependent aspect of physical design. Furthermore, Logic Synthesis heavily restructures the netlist, resulting in substantial structural disparities that hinder capturing the post-synthesis netlist structure. Sanjay Gandham, Joe Walston, Sourav Samanta, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin, Stelios Diamantidis |
ICCAD | 4 |
| 2024 | SCALE: A Structure-Centric Accelerator for Message Passing Graph Neural NetworksabstractMessage passing paradigm has been widely used in developing complex Graph Neural Network (GNN) models, allowing for concise representations of edge and vertex-wise operations. Despite its pivotal role in theoretical advancement, the respective expression of edge and vertex operations, along with evolving GNN variants and datasets, has inevitably led to enormous computational complexity due to heterogeneous computation kernels. In particular, such inconsistent computation characteristics present new challenges in leveraging intermediate data reuse, ensuring both edge and vertex-wise workload balance, and sustaining system scalability. In this paper, we propose a structurecentric accelerator, SCALE, that can support a variety of message passing GNN models with improved parallelism, data reuse, and scalability. The central idea is to find latent similarities among GNN primitives such as shared dataflow structure, rather than strictly adhering to heterogeneous model structure. This serves as a hinge to homogenize inconsistencies in various GNN computation kernels. To accomplish this concept, SCALE consists of three unique designs, a novel systolic array-like architecture, a degree and vertex-aware scheduling, and a coherent dataflow tailored for fused graph and neural operations. The proposed systolic array-like architecture can support varying dataflows such as all-reduce, of distinct GNN operations improving parallelism, data reuse, and throughput. The degree and vertex-aware scheduling can remedy the workload imbalance encountered in vertex and edge-wise operations. Moreover, the proposed dataflow can unify the data movement of both graph and neural operators without extra communication and storage overheads. Our simulation results show that SCALE achieves 1.82× speedup and 38.9% energy reduction on average over the state-of-the-art GNN accelerators [1]–[4]. Lingxiang Yin, Sanjay Gandham, Mingjie Lin, Hao Zheng 0005 |
MICRO | 1 |
| 2023 | OCMGen: Extended Design Space Exploration with Efficient FPGA Memory InferenceabstractDeep learning applications demand high memory storage and computational power to operate on millions of parameters. Field Programmable Gate Arrays (FPGAs), with high compute resources and the ability to store data on-chip in their distributed memory components such as Block RAM (BRAM) and Ultra RAM (URAM), are good candidates to deploy such memory-intensive applications [1]. However, without careful tailoring of the hardware design for a target device, current synthesis tools (e.g., Xilinx Vivado) can severely underutilize these RAM primitives reducing the usable on-chip memory (OCM). Consequently, this forces the accelerator to perform more frequent expensive off-chip accesses, limiting its performance. Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin |
FCCM | 2 |
| 2023 | Exploring Architecture, Dataflow, and Sparsity for GCN Accelerators: A Holistic FrameworkabstractRecent years have seen an increasing number of Graph Convolutional Network (GCN) models employed in various real-world applications. However, designing efficient architectures for GCN acceleration remains challenging due to the varied sparsity across graph datasets. Despite significant efforts, very few of the existing works have considered a holistic view of the entire GCN accelerator design, and therefore, the dynamic interactions between architecture, dataflow (i.e., data reuse and parallelization strategies), and compression format are not well studied Lingxiang Yin, Jun Wang 0001, Hao Zheng 0005 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | SAGA: Sparsity-Agnostic Graph Convolutional Network Acceleration with Near-Optimal Workload BalanceabstractGraph Convolutional Networks (GCNs) have shown much promise in resolving sophisticated scientific problems with non-Euclidean data, such as traffic prediction, disease classification, and many others. However, the irregular sparsity of real-world graphs remains a major challenge toward efficient GCN acceleration. In this paper, we propose SAGA, a Sparsity-Agnostic Graph Convolutional Accelerator with near-optimal workload balance. Specifically, it consists of two unique features, an NZ-based scheduling, and a novel accelerator architecture. Unlike conventional GCN accelerators with uneven distribution of sparse matrix, the proposed NZ-based scheduling leverages the metadata encoded in the compression format to enable even distribution of sparse matrix at runtime, thus achieving near-optimal workload balancing. In addition, the proposed architecture, including a task scheduler, an accumulation table, and a partial row accumulation unit, can support the proposed NZ-based scheduling without data preprocessing and reformatting with low overheads. We prototyped the proposed design through FPGAs, and our evaluation results show that SAGA achieves up to$\mathbf{1.56}\times$speedup and$\mathbf{2.05}\times$energy savings on average as compared to the prior art [1]. Sanjay Gandham, Lingxiang Yin, Hao Zheng 0005, Mingjie Lin |
ICCAD | 2 |
| 2023 | ARIES: Accelerating Distributed Training in Chiplet-Based Systems via Flexible InterconnectsabstractLarge-scale deep learning models are widely deployed in many application domains with remarkable performance improvements. However, training these models with immense parameters calls for unprecedented computing and communication capabilities. Recently, chiplet-based architectures have shown much promise in scaling Deep Neural Network (DNN) inference, but their applications in the training phase remain unexplored and challenging. In this paper, we posit, beyond scaling computing capability, chiplet-based architectures could also be leveraged to enable new optimization opportunities for existing parallel training algorithms (e.g., Ring and Tree-based all-reduce). Specifically, we aim to explore a variety of topological characteristics, along with the interposer technology, to sustain the performance scaling of parallel training in chiplet-based systems. We propose ARIES, a versatile chiplet-based communication architecture supporting various parallel training algorithms using a flexible interconnect design. The proposed design can adapt to various collective operations such as reduce and gather across a wide diversity of training algorithms. Moreover, such flexibility is also leveraged to further enhance existing all-reduce algorithms depending on the latency and bandwidth requirements of the DNN model and dataset size. Simulation results show that the proposed ARIES can achieve up to 3.92× speedup in execution time and 38.8% reduction in Network-on-Chip (NoC) energy consumption when compared to prior work. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Ahmed Louri, Hao Zheng 0005 |
ICCAD | 1 |
| 2023 | Polyform: A Versatile Architecture for Multi-DNN Execution via Spatial and Temporal AccelerationabstractContemporary applications and cloud workloads often comprise multiple Deep Neural Network (Multi-DNN) models. These models exhibit significant variations in computation, memory, and communication characteristics. For such heterogeneous workloads, a static and rigid hardware accelerator can no longer provide efficient and high-performance execution. To this end, we propose a versatile accelerator, called Polyform, to support the concurrent execution of different DNN models with the goal of improving energy and performance efficiency. Specifically, Polyform features two unique designs from both hardware and scheduling standpoints. On the hardware level, we have designed a flexible interconnection network that facilitates the formation of multiple sub-accelerators. Our design allows for spatial resource partitioning, including bandwidth and computation, while also providing effective communication support for various parallelism choices. On the scheduling level, Polyform employs a novel two-stage Genetic Algorithm (GA) to explore and identify the optimal configurations such as task orders, partition size, dataflow styles (e.g., weight or output stationary), and bandwidth. Our simulation shows that Polyform achieves remarkable results compared to prior work, including up to 77.8% energy reduction and a 2.79× improvement in throughput as compared to prior work [1]–[3]. Lingxiang Yin, Amir Ghazizadeh Ahsaei, Shilin Tian, Ahmed Louri, Hao Zheng 0005 |
ICCD | 1 |