Shan Huang 0010

dblp:06/4186-10 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0005-2012-8540ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MARCA-v2: Mamba Accelerator With Complementary State-Space Model Sparsity and Reconfigurable Architecture
abstract
Large Language models with state space model (SSM) especially Mamba have demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm performance. Current research on Mamba focuses primarily on integrating it with various application scenarios. However, there is limited research on optimizing for Mamba processing. Therefore, we profile the processing carefully and identify three main challenges in Mamba computations: (1) Large memory access overhead of element-wise operations in SSM. Based onMARCAarchitecture, the time proportion of SSM is still the bottleneck when the sequence length reaches 2048, accounting for 62.52% of the total runtime. Within the SSM, the memory access overhead of element-wise operations account for 97.17%, leading to consuming 96.56% of the time. (2) Inefficient sparse element-wise execution onMARCAarchitecture. SOTA architectures likeMARCApropose a reconfigurable reduction tree to accelerate dense element-wise operations but lack effective sparse support for sparse execution. When 30% of the elementwise operations are skipped, these skipped operations are still mapped to the PE array and trigger redundant execution cycles, incurring 1.43× performance gap with the ideal. (3) Large area overhead for nonlinear function unit. Exponential function and SiLU are two main nonlinear functions in SSM. Previous methods design specific unit for acceleration, leading to 38% and 18% area overheads of the PE. In response to these challenges, we propose a new Mamba accelerator with complementary state space model (SSM) sparsity and reconfigurable architecture,MARCA-v2, based onMARCA, to support fast and energy-efficient Mamba computations. Three novel techniques are as follows: (1) Complementary SSM sparsity with column-wise granularity. We first profile the numerical distributions of activations in SSM and propose a column-wise complementary static sparsity for SSM computation. To further enable lightweight and hardware-friendly sparse computation, we propose a δ-bitmap encoding scheme for compressed storage and introduce two abstractions for sparse element-wise operations. (2) Lightweight sparse element-wise architecture. Based on the hardware friendly sparsity algorithm andMARCAarchitecture, we design and integrate a lightweight Metadata Processing Unit (Meta-PU) into the existing pipeline, which decodes the sparsity metadata and dynamically generates control signals to guide PE arrays. The overall architecture can efficiently support both dense and sparse operations, maximizing speed and energy efficiency. (3) Reusable nonlinear function unit based on reconfigurable PE arrays. We decompose the exponential function and SiLU into several element-wise operations. Thus, the reconfigurable PEs are fully reused to execute nonlinear functions with negligible accuracy loss. We conduct extensive experiments on Mamba model families with different sizes. Experimental results show that in theprefillstage,MARCA-v2achieves 1.77-10.87×, 1.03- 1.08×, and 4.78-9.10× speedup and 8.29-33.47×/1.03-1.08×/4.78-9.10× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively. In thedecodestage,MARCA-v2achieves 0.88-7.65×/1.00-1.01×/1.19-1.64× speedup and 3.11-27.01×/1.00-1.01×/1.19-1.64× energy efficiency improvement compared with Mamba-GPU,MARCAand Spada, respectively.
Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Ningyi Xu, Guohao Dai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 ViDA: Video Diffusion Transformer Acceleration with Differential Approximation and Adaptive Dataflow
abstract
Recent advancements in Video Diffusion Transformer (VDiT) models have greatly promoted the development of video generation, as exemplified by Sora of OpenAI. However, there are still two challenges for VDiT: 1) There is still existing large inter-frame redundant computation. Previous works on reducing computation based on inter-frame similarity simply consider the Act-W operators. The remaining Act-Act operators still dominate the execution of VDiT (about 57%). 2) Operational intensity varies greatly, leading to under-utilization. There is a massive gap between the operational intensity of Act-W and Act-Act operators in VDiT with multiple frames. Previous works with the static hardware architecture and dataflow lead to under-utilization (<36.42%).
Li Ding 0012, Jun Liu 0117, Shan Huang 0010, Guohao Dai 0001
ASP-DAC3
2025 LLSM: LLM-enhanced Logic Synthesis Model with EDA-guided CoT Prompting, Hybrid Embedding and AIG-tailored Acceleration
abstract
Machine learning-based methods have shown promising results in the field of Electronic Design Automation (EDA) like logic synthesis result prediction, enabling a shift-left in the overall EDA flow. Designers should fully optimize their Register Transfer Level (RTL) designs early because remedying low-quality RTL in downstream synthesis stages is extremely challenging. However, previous works mainly start modeling from the netlist level or layout level and apply Graph Neural Networks (GNNs) to make predictions.
Shan Huang 0010, Jinhao Li 0006, Jiancai Ye, Ningyi Xu, Guohao Dai 0001
ASP-DAC1
2025 Speculative Decoding for Verilog: Speed and Quality, All in One
abstract
The rapid advancement of large language models (LLMs) has revolutionized code generation tasks across various programming languages. However, the unique characteristics of programming languages, particularly those like Verilog with specific syntax and lower representation in training datasets, pose significant challenges for conventional tokenization and decoding approaches. In this paper, we introduce a novel application of speculative decoding for Verilog code generation, showing that it can improve both inference speed and output quality, effectively achieving speed and quality all in one. Unlike standard LLM tokenization schemes, which often fragment meaningful code structures, our approach aligns decoding stops with syntactically significant tokens, making it easier for models to learn the token distribution. This refinement addresses inherent tokenization issues and enhances the model’s ability to capture Verilog’s logical constructs more effectively. Our experimental results show that our method achieves up to a $5.05 \times$ speedup in Verilog code generation and increases pass@10 functional accuracy on RTLLM by up to $\mathbf{1 7. 1 9 \%}$ compared to conventional training strategies. These findings highlight speculative decoding as a promising approach to bridge the quality gap in code generation for specialized programming languages.
Changran Xu, Yi Liu 0081, Yunhao Zhou, Shan Huang 0010, Ningyi Xu, Qiang Xu 0001
DAC4
2025 DyLGNN: Efficient LM-GNN Fine-Tuning with Dynamic Node Partitioning, Low-Degree Sparsity, and Asynchronous Sub-Batch
abstract
Text-Attributed Graphs (TAGs) tasks involve both textual node information and graph topological structure. The top-k method, using Language Models (LMs) for text encoding and Graph Neural Networks (GNNs) for graph processing, offers the best accuracy while balancing memory and training time. However, challenges still exist: (1) Static sampling of k neighbors reduces performance. Using a fixed k can result in sampling too few or too many nodes, leading to a 3.2% accuracy loss across datasets. (2) Time-consuming processing for non-trainable nodes. After partitioning all nodes into with-gradient trainable and without-gradient non-trainable sets, the number of non-trainable nodes is ~9-10 x larger than trainable nodes, resulting in nearly 70% of the total time. (3) Time-consuming data movement. For processing non-trainable nodes, after the text strings are tokenized into tokens on the CPU side, the data movement from host memory to GPU takes 30%-40% of the time. In this paper, we propose DyLGNN, an efficient end-to-end LM-GNN fine-tuning framework through three innovations: (1) Heuristic Node Partitioning. We propose an algorithm that dynamically and adaptively selects “important” nodes to participate in the training process for downstream tasks. Compared to the static top-k method, we reduce the training memory usage by 24.0%. (2) Low-Degree Sparse Attention. We point out that the embedding of low-degree nodes has minimal impact on the final results (e.g. ~1.5% accuracy loss), therefore, We perform sparse attention computation on low-degree nodes to further reduce the computation caused by “unimportant” nodes, achieving an average 1.27 x speedup. (3) Asynchronous Sub-batch Pipeline. Within the top-k framework, we analyze the time breakdown of the LM inference component. Leveraging our heuristic node partitioning, which effectively minimizes memory demands, we can asynchronously execute data movement and computation, thereby overlapping the time required for data movement. This improves GPU utilization and results in an average 1.1x speedup. We conduct experiments on several common graph datasets, and by combining the three methods mentioned above, DyLGNN achieves a 22.0% reduction in memory usage and a 1.3x end-to-end speedup compared to the top-k strategy.
Jinhao Li 0006, Shan Huang 0010, Jiancai Ye, Ningyi Xu, Guohao Dai 0001
DATE4
2025 DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
abstract
Circuit representation learning has become pivotal in electronic design automation, enabling critical tasks such as testability analysis, logic reasoning, power estimation, and SAT solving. However, existing models face significant challenges in scaling to large circuits due to limitations like over-squashing in graph neural networks and the quadratic complexity of transformer-based models. To address these issues, we introduce \textbf{DeepGate4}, a scalable and efficient graph transformer specifically designed for large-scale circuits. DeepGate4 incorporates several key innovations: (1) an update strategy tailored for circuit graphs, which reduce memory complexity to sub-linear and is adaptable to any graph transformer; (2) a GAT-based sparse transformer with global and local structural encodings for AIGs; and (3) an inference acceleration CUDA kernel that fully exploit the unique sparsity patterns of AIGs. Our extensive experiments on the ITC99 and EPFL benchmarks show that DeepGate4 significantly surpasses state-of-the-art methods, achieving 15.5\% and 31.1\% performance improvements over the next-best models. Furthermore, the Fused-DeepGate4 variant reduces runtime by 35.1\% and memory usage by 46.8\%, making it highly efficient for large-scale circuit analysis. These results demonstrate the potential of DeepGate4 to handle complex EDA tasks while offering superior scalability and efficiency.
Shan Huang 0010, Jianyuan Zhong, Zhengyuan Shi, Guohao Dai 0001, Ningyi Xu, Qiang Xu 0001
ICLR2
2025 Enabling Efficient Sparse Multiplications on GPUs With Heuristic Adaptability
abstract
Sparse matrix-vector/matrix multiplication, namely SpMMul, has become a fundamental operation during model inference in various domains. Previous studies have explored numerous optimizations to accelerate it. However, to enable efficient end-to-end inference, the following challenges remain unsolved: 1) incomplete design space and time-consuming preprocessing. Previous methods optimize SpMMul in limited loops and neglect the potential space exploration for further optimization, resulting in >30% waste of computing power. In addition, the preprocessing overhead in SparseTIR and DTC-SpMM is$1000\times $larger than sparse computing; 2) incompatibility between static dataflow and dynamic input. A static dataflow can not always be efficient to all input, leading to >80% performance loss; and 3) simplistic algorithm performance analysis. Previous studies primarily analyze performance from algorithmic advantages, without considering other aspects like hardware and data features. To tackle the above challenges, we present DA-SpMMul, a Data-Aware heuristic GPU implementation for SpMMul in multiplatforms. DA-SpMMul creatively proposes: 1) complete design space based on theoretical computations and nontrivial implementations without preprocessing. We propose three orthogonal design principles based on theoretical computations and provide nontrivial implementations on standard formats, eliminating the complex preprocessing; 2) feature-enabled adaptive algorithm selection mechanism. We design a heuristic model to enable algorithm selection considering various features; and 3) comprehensive algorithm performance analysis. We extract the features from multiple perspectives and present a comprehensive performance analysis of all algorithms. DA-SpMMul supports PyTorch on both NVIDIA and AMD and achieves an average speedup of$3.33\times $and$3.02\times $over NVIDIA cuSPARSE, and$12.05\times $and$8.32\times $over AMD rocSPARSE for sparse matrix-vector multiplication and sparse matrix-matrix multiplication, and up to$1.48\times $speedup against the state-of-the-art open-source algorithm. Integrated with graph neural network framework, PyG, DA-SpMMul achieves up to$1.22\times $speedup on GCN inference.
Shan Huang 0010, Jinhao Li 0006, Guyue Huang, Yuan Xie 0001, Yu Wang 0002, Guohao Dai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 MARCA: Mamba Accelerator with Reconfigurable Architecture
abstract
State space model (SSM) especially Mamba has demonstrated remarkable capabilities in various domains. Compared to Transformers, Mamba reduces the quadratic computational complexity and achieves a higher algorithm accuracy (e.g., the accuracy of Mamba-2.8b is higher than OPT-6.7b). However, challenges still exist in accelerating Mamba computations. (1) Incompatibility between element-wise operations and Tensor Core. Linear operations (matrix multiplications) and element-wise operations are the two dominating operations in Mamba. The time proportion of element-wise operations escalates significantly (e.g., >60% with 2048 input length). These operations do not need reduction, which is not compatible with the existing Tensor Core-based architectures (e.g., 1/16 normalized speed). (2) Large area overhead for nonlinear function unit. The optimized nonlinear function unit like exponential unit still occupies >30% of the processing element (PE) area. (3) Large memory access but limited data sharing for element-wise operations. Linear and element-wise operations in Mamba exhibit large compute intensity variance (e.g., ~3 orders of magnitude) and large read/write ratio variance (e.g., >3 orders). Due to the limited data sharing in element-wise operations, it is useless to apply the existed methods like tiling to element-wise operations.
Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Li Ding 0012, Ningyi Xu, Guohao Dai 0001
ICCAD2
2024 Fast and Efficient 2-bit LLM Inference on GPU: 2/4/16-bit in a Weight Matrix with Asynchronous Dequantization
abstract
Large language models (LLMs) have demonstrated impressive abilities in various domains while the inference cost is expensive. Many previous studies exploit quantization methods to reduce LLM inference cost by reducing latency and memory consumption. Applying 2-bit single-precision weight quantization brings >3% accuracy loss, so the state-of-the-art methods use mixed-precision methods for LLMs (e.g. Llama2-7b, etc.) to improve the accuracy. However, challenges still exist: (1) Uneven distribution in weight matrix. Weights are quantized by groups, while some groups contain weights with large range. Previous methods apply inter-weight mixed-precision quantization and neglect the range difference inside each weight matrix, resulting in >2.7% accuracy loss (e.g. LLM-MQ and APTQ). (2) Large speed degradation by adding sparse outliers. Reserving sparse outliers improves accuracy but slows down the speed affected by the outlier ratio (e.g. 1.5% outliers resulting in >30% speed degradation in SpQR). (3) Time-consuming dequantization operations on GPUs. Mainstream methods require a dequantization operation to perform computation on the quantized weights, and the 2-order dequantization operation is applied because scales of groups are also quantized. These dequantization operations lead to >50% execution time.
Jinhao Li 0006, Shan Huang 0010, Jun Liu 0117, Yaoxiu Lian, Guohao Dai 0001
ICCAD4