Xiao Shi 0001

dblp:179/9306-1 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-1152-0055ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 7 first-author · 13 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language Models
abstract
Large Language Models (LLMs) are achieving unprecedented performance across diverse tasks, benefiting from autoregressive generation. However, this left-to-right decoding paradigm inherently limits contextual understanding quality. Diffusion-based LLMs (dLLMs) offer a promising alternative by iteratively refining sequences via denoising, enabling stronger bidirectional context modeling and improved generation quality. However, dLLMs face two main challenges: redundant computation and memory overhead in multi-step denoising, and excessive inference cost from over-denoising under fixed-step schedules. To address these issues, we propose dLLM-OPU, an FPGA overlay processor to accelerate dLLMs. Our solution features two key innovations: (1) a Region-Adaptive Caching for Dynamic Column Sparsity Framework that exploits temporal locality for selective recomputation without model retraining, and (2) a Token Entropy-based Early Stopping strategy that dynamically terminates the denoising process based on token-level convergence metrics. We implement these innovations through a specialized sparse processing element (PE) array that maximizes top-k sparsity utilization by minimizing idle cycles via row-column concatenation, complemented by an efficient cache management system that reduces memory access latency and a flexible entropybased decoding unit. Implemented on a U200 FPGA, dLLM-OPU achieves $2.2 \times-5.1 \times$ speedup and $7.6 \times-20.3 \times$ energy efficiency over RTX4090 in LLaDA.
Yangbo Wei, Shaoqiang Lu, Junhong Qian, Lei He 0001, Dongge Qin, Xiao Shi 0001
ASP-DAC6
2026 DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Dongge Qin, Shiji Gao, Yizhi Ding, Xiao Shi 0001, Lei He 0001
ASPLOS (2)9
2026 Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001
ISCAS6
2026 Scalable Yield Analysis of SRAM and Analog Circuits Using Multi-Kernel Sparse Representation
abstract
With the advancement of technology nodes and the increasingly stringent requirements for stability, general yield analysis of customized circuits in the early stages of design has become a key bottleneck in manufacturing. In this article, we propose a multi-kernel sparse representation-based classification (MKSRC) method to enhance the efficiency and scalability of failure probability estimation by classifying tail samples. It employs class-balanced sampling to address data imbalance issues and utilizes multi-kernel features with adaptive kernel weights to enhance the accuracy and robustness of the classifier. Experimental results on 32-bit SRAM columns and analog circuits demonstrate that the proposed MKSRC method achieves higher classification accuracy and efficiency compared to other state-of-the-art methods, particularly in scenarios with limited training data. Compared to SOTA yield estimation methods, the MKSRC method achieves an average 2.57–3.29× improvement in both accuracy and efficiency, highlighting its ability to provide efficient and scalable yield analysis solutions for both SRAM and analog circuits.
Liangji Wu, Zhongxi Guo, Xiao Shi 0001, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.5
2025 MoE-OPU: An FPGA Overlay Processor Leveraging Expert Parallelism for MoE-based Large Language Models
abstract
The advent of Large Language Models (LLMs) like DeepSeek, empowered by the Mixture-of-Experts (MoE) architecture, has driven significant advancements across diverse applications. However, a critical challenge arises during inference: Only a small fraction of experts are activated, causing severe token allocation imbalances among experts. This inefficiency poses substantial storage and computational burdens on resource-constrained devices, exacerbated by the lack of optimization strategies that integrate expert usage-aware parameter pruning and parallel scheduling, ultimately leading to suboptimal resource utilization. To address these limitations, we propose MoE-OPU, an FPGA-based overlay processor that optimizes parallel MoE inference through three key innovations. First, we introduce N:M sparsity (1:4/2:4/4:8/6:8/8:8) in the MLP layers and mixed-precision quantization (BF16/FP8/INT4) guided by expert activation frequency, reducing the parameter size by up to 2.76× while maintaining model accuracy (only 1.53% average drop after fine-tuning). Second, a lightweight prediction network dynamically predicts next-layer "hot" experts by analyzing historical activation patterns and current hidden states, achieving an average prediction hit rate of 83.4%. Third, a reconfigurable multi-core architecture maximizes the utilization of HBM bandwidth via a systolic array that natively supports sparse and mixed-precision computations, coupled with parallel concatenation to balance compute and memory efficiency. Experimental results on a Xilinx V80 FPGA with the DeepSeek-V2-lite model demonstrate that MoE-OPU outperforms the NVIDIA A100 GPU, delivering a 6.78× higher token throughput. Compared to RTX 4090 and U200 FPGA, MoE-OPU achieves 13.37× and 7.85× improvements, respectively. These advancements highlight the potential of algorithm-hardware co-design for scalable deployment of MoE-based LLMs on edge devices.
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Xiao Shi 0001, Lei He 0001
ICCAD5
2025 Robust optimization algorithm of RF MEMS switches considering uncertainties
Hao Yan 0002, Yaning Jia, Chuangyuan Zeng, Xiaoping Liao, Xiao Shi 0001
Integr.5
2024 LVF2: A Statistical Timing Model based on Gaussian Mixture for Yield Estimation and Speed Binning
abstract
As transistor size continues to scale down, process variation has become an essential factor determining semiconductor yield and economic return. The Liberty Variation Format (LVF) is the current industrial standard that expresses statistical timing behaviors based on single Gaussian model. However, it loses accuracy when the timing distribution is non-Gaussian due to growing process variations. This paper proposes a novel LVF2 distribution model that combines two weighted skewed-normal (SN) distributions, which better captures the multi-Gaussian timing distribution while maintaining backward compatibility with LVF. Experiments using TSMC 22nm standard cells show that, compared to LVF, LVF2 reduces binning error by 7.74X in delay and 9.56X in transition time, and reduces 3σ-yield error by 4.79X and 7.18X in delay and transition time, respectively. The error reduction for path delay is diminished due to Central Limit Theorem (CLT). But it is still 2X for a typical circuit path with 8 Fanout-of-4 (FO4) inverter delays.
Junzhuo Zhou, Haoxuan Xia, Leilei Jin, Xiao Shi 0001, Wei W. Xing, Ting-Jung Lin, Lei He 0001
DAC6
2024 A Novel Cross-Perturbation for Single Domain Generalization
abstract
Single domain generalization aims to enhance the ability of the model to generalize to unknown domains when trained on a single source domain. However, the limited diversity in the training data hampers the learning of domain-invariant features, resulting in compromised generalization performance. To address this, data perturbation (augmentation) has emerged as a crucial method to increase data diversity. Nevertheless, existing perturbation methods often focus on either image-level or feature-level perturbations independently, neglecting their synergistic effects. To overcome these limitations, we propose CPerb, a simple yet effective cross-perturbation method. Specifically, CPerb utilizes both horizontal and vertical operations. Horizontally, it applies image-level and feature-level perturbations to enhance the diversity of the training data, mitigating the issue of limited diversity in single-source domains. Vertically, it introduces multi-route perturbation to learn domain-invariant features from different perspectives of samples with the same semantic category, thereby enhancing the generalization capability of the model. Additionally, we propose MixPatch, a novel feature-level perturbation method that exploits local image style information to further diversify the training data. Extensive experiments on various benchmark datasets validate the effectiveness of our method.
Dongjia Zhao, Lei Qi 0001, Xiao Shi 0001, Yinghuan Shi, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2023 A Novel Delay Calibration Method Considering Interaction between Cells and Wires
abstract
In the advanced technology, the accuracy of cell and wire delay modeling are the key metrics for timing analysis. However, when the supply voltage decreases to the near-threshold regime, the complicated process variation effect causes the cell delay and the wire delay hard to model. Most researchers study cell or wire delay separately, ignoring the coefficients between them. In this paper, we propose an N-sigma delay model by characterizing different sigma levels$\mathbf{(-3\sigma {to}+3\sigma)}$of the cell and wire delay distribution. The N-sigma cell delay model is represented by the first four moments and calibrated by the operating conditions (input slew, output load). Meanwhile, based on the Elmore model, the wire delay variability is calculated by considering the effect of drive and load cells. The delay models are verified through the ISCAS85 benchmarks and the functional units of PULPino processor with TSMC 28 nm technology. Compared to the SPICE results, the average errors for estimating the$+/-\mathbf{3\sigma}$cell delay are 2.1 % and 2.7% and those of the wire delay are 2.4% and 1.6%, respectively. The errors of path delay analysis keep below 6.6% and the speed is 103X over SPICE MC simulations.
Leilei Jin, Wenjie Fu 0003, Hao Yan 0002, Xiao Shi 0001, Longxing Shi
DATE5
2023 Token Packing for Transformers with Variable-Length Inputs
abstract
Transformer-based models has achieved remarkable success in extensive tasks for natural language processing. To face the variable-length sentences in human language, popular deep learning frameworks rely on zero padding for batch processing, which introduces significant computation and memory overhead. Existing works attempt to eliminate padding redundancy but results in low hardware efficiency due to the mismatch between the variable shape of operations and fixed shape of processing elements (PEs). This paper proposes a reconfigurable systolic array with token packing in three folds to boost hardware efficiency. First, matrix multiplications for different tokens can be packed along the array columns to improve spatial efficiency. Meanwhile, for temporal efficiency, we develop a coarse-grained pipeline for attention, where stages can run on different parts of the array at the same time. We further exploit the masking redundancy in the Transformer decoder with runtime reconfigurable inter-PE connection and buffer switching. Applied to GPT, our FPGA design has achieved 1.16× higher normalized throughput and 1.94× better runtime MAC utilization over the state-of-the-art GPU performance for variable-length input sequences from GLUE and SQuAD dataset.
Tiandong Zhao, Siyuan Miao, Shaoqiang Lu, Jialin Cao, Xiao Shi 0001, Kun Wang 0005, Lei He 0001
FPL6
2023 An efficient SRAM yield analysis method based on scaled-sigma adaptive importance sampling with meta-model accelerated
Liang Pang 0002, Mengyun Yao, Xiao Shi 0001, Hao Yan 0002, Longxing Shi
Integr.5
2022 A Compact High-Dimensional Yield Analysis Method using Low-Rank Tensor Approximation
abstract
“Curse of dimensionality” has become the major challenge for existing high-sigma yield analysis methods. In this article, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with the circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm and calibrated iteratively with a bootstrap-assisted adaptive sampling method. We also develop a novel global sensitivity analysis approach to generate a reduced LRTA meta-model which is more compact. It further accelerates the procedure of model calibration and yield estimation. Experiments on memory and analog circuits validate that the proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Chengzhen Xuan, Lei He 0001, Longxing Shi
ACM Trans. Design Autom. Electr. Syst.1
2021 An Adaptive Delay Model for Timing Yield Estimation under Wide-Voltage Range
abstract
Yield analysis for wide-voltage circuit design is a strong nonlinear integration problem. The most challenging task is how to accurately estimate the yield of long-tail distribution. This paper proposes an adaptive delay model to substitute expensive transistor-level simulation for timing yield estimation. We use the Low-Rank Tensor Approximation (LRTA) to model the delay variation from a large number of process parameters. Moreover, an adaptive nonlinear sampling algorithm is adopted to calibrate the model iteratively, which can capture the larger variability of delay distribution for different voltage regions. The proposed method is validated on benchmark circuits of TAU15 in 45nm free PDK. The experiment results show that our method achieves 20-100X speedup compared to Monte Carlo simulation at the same accuracy level.
Hao Yan 0002, Xiao Shi 0001, Chengzhen Xuan, Peng Cao 0002, Longxing Shi
ASP-DAC2
2021 Late Breaking Results: Novel Discrete Dynamic Filled Function Algorithm for Acyclic Graph Partitioning
abstract
A parallel simulation that partitions a large circuit into sub-circuits is widely used to reduce simulation runtime. To achieve higher simulation throughput, we shall consider signal directions, and thus the final partitioning solution must be acyclic. In this paper, we model a circuit as a directed graph and consider acyclic graph partitioning to minimize edge cuts. This problem differs from the traditional partitioning problem because of the additional acyclicity constraint. Unlike traditional heuristics that tend to be trapped in local minima, especially for large graphs, we present a novel discrete dynamic filled function algorithm for the acyclic graph partitioning problem. Our algorithm can guarantee convergence and effectively move from one discrete local minimizer to another better one. Experimental results show that our algorithm achieves 8% average cutsize reduction over the state-of-the-art works in a comparable runtime.
Jianli Chen, Jiarui Chen, Xiao Shi 0001, Lichong Sun, Jun Yu 0010
DAC3
2020 TYMER: A Yield-based Performance Model for Timing-speculation SRAM
abstract
In low power designs, timing-speculative techniques are proposed to boost the SRAM frequency and throughput. This paper proposes TYMER, a unified yield-based performance model for timing-speculative SRAM. In TYMER, the first sub-model evaluates access-time yield at different worldline enable time for a general 6T SRAM under low supply voltages, while the second one uses the yield results to estimate optimal sensing time and the overall read latency for speculative SRAM. TYMER is not only compared with simulation results but also the measurements from 28nm fabricated speculative SRAM chips. Both cases show precise evaluation results under different operating conditions.
Shan Shen, Liang Pang 0002, Tianxiang Shao, Xiao Shi 0001, Longxing Shi
DAC5
2020 Time-Division Multiplexing Based System-Level FPGA Routing for Logic Verification
abstract
Multi-FPGA prototyping is widely used for modern VLSI verification, but the limited number of inter-FPGA connections in a multi-FPGA system may cause routing failures. As a result, the time-division multiplexing (TDM) technique is adopted to increase its resource utilization by transmitting multiple signals through the same routing channel. Due to the large signal delay between FPGA pairs, however, the performance of such a system greatly depends on the inter-FPGA routing quality. In this paper, we propose a TDM-based system-level routing algorithm to simultaneously minimize the maximum TDM (signal multiplexing) ratio and runtime, considering the crucial ratio constraints. By weighting the routing edges, we first model the net routing as a Steiner minimum tree (SMT) problem and solve it with an approximation algorithm with the performance bound 2(1 - 1/1), where l is the number of leaves in an optimal SMT. Then, a timing-driven assignment method is presented to evenly distribute the TDM ratio to routing signals, followed by a novel reassignment algorithm to efficiently handle unbalanced net groups. Finally, a ratio-aware refinement technique is employed to further improve the solution quality. Compared with the top-3 winners at the 2019 CAD Contest at ICCAD based on the contest benchmarks, experiment results show that our proposed algorithm achieves the best runtime and TDM ratio while satisfying all TDM constraints.
Zhifeng Lin, Xiao Shi 0001, Jianli Chen, Jun Yu 0010, Yao-Wen Chang
DAC3
2020 A Non-Gaussian Adaptive Importance Sampling Method for High-Dimensional and Multi-Failure-Region Yield Analysis
abstract
Rare-event yield analysis is challenging for high-dimensional circuit cases. In this paper, we propose a non-Gaussian adaptive importance sampling (NGAIS) method. In order to approximate the failure region in high-dimensional space, we model it as a mixture of von Mises-Fisher distributions. We formulate the parameter estimation problem as a maximum likelihood estimation problem, and then solve with expectation-maximization algorithm. Experiments on bit cell, amplifier and SRAM column circuit validate that the proposed NGAIS method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Chuwen Li, Jianli Chen, Longxing Shi, Lei He 0001
ICCAD1
2020 An Efficient Adaptive Importance Sampling Method for SRAM and Analog Yield Analysis
abstract
Performance failure has become a major threat for various memory and analog circuits. It is challenging to estimate the extremely small failure probability when failed samples are distributed in multiple disjoint regions. In this article, we propose an adaptive importance sampling (AIS) algorithm. AIS has several iterations of sampling region adjustments, while existing methods predecide a static sampling distribution. We design two adaptive frameworks based on resampling and population Metropolis-Hastings (MH) to iteratively search for failure regions. The experimental results of the AIS method exhibit better efficiency and higher accuracy. For SRAM bit cell with single failure region, the AIS method uses 2-$27{\times }$ fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure regions, the AIS method is $90{\times }$ faster than Monte Carlo and 7-23 ${\times }$ over other methods. For charge pump circuit and $C^{2}MOS$ master-slave latch circuit, the AIS method can reach 6-$18{\times }$ and 4-$6{\times }$ speedup over other methods, respectively.
Xiao Shi 0001, Hao Yan 0002, Longxing Shi, Lei He 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Meta-Model based High-Dimensional Yield Analysis using Low-Rank Tensor Approximation
abstract
"Curse of dimensionality" has become the major challenge for existing high-sigma yield analysis methods. In this paper, we develop a meta-model using Low-Rank Tensor Approximation (LRTA) to substitute expensive SPICE simulation. The polynomial degree of our LRTA model grows linearly with circuit dimension. This makes it especially promising for high-dimensional circuit problems. Our LRTA meta-model is solved efficiently with a robust greedy algorithm, and calibrated iteratively with an adaptive sampling method. Experiments on bit cell and SRAM column validate that proposed LRTA method outperforms other state-of-the-art approaches in terms of accuracy and efficiency.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001
DAC1
2019 Efficient Yield Analysis for SRAM and Analog Circuits using Meta-Model based Importance Sampling Method
abstract
Performance failure has become the major threat to the robustness and reliability of various memory and analog circuits. It is challenging to accurately estimate the extremely small failure probability when failed samples are distributed in multiple disjoint failure regions. In this paper, we develop a novel meta-model based importance sampling (MIS) method. MIS utilizes Gaussian Process meta-model to construct quasi-optimal importance sampling distribution, and performs Markov Chain Monte Carlo (MCMC) simulation to generate new samples from the proposed distribution. By updating our global Importance Sampling estimator in an iterated framework, MIS leads to better efficiency and higher accuracy. For SRAM bit cell with single failure region, MIS uses 4-6X fewer samples and reaches better accuracy when compared to several recent methods. For a two-stage amplifier circuit with multiple failure schemes, MIS is 213X faster than MC without compromising accuracy, while other methods fail to cover all failure regions in our experiment.
Xiao Shi 0001, Hao Yan 0002, Qiancun Huang, Longxing Shi, Lei He 0001
ICCAD1
2019 Adaptive Clustering and Sampling for High-Dimensional and Multi-Failure-Region SRAM Yield Analysis
abstract
Statistical circuit simulation is exhibiting increasing importance for memory circuits under process variation. It is challenging to accurately estimate the extremely low failure probability as it becomes a high-dimensional and multi-failure-region problem. In this paper, we develop an Adaptive Clustering and Sampling (ACS) method. ACS proceeds iteratively to cluster samples and adjust sampling distribution, while most existing approaches pre-decide a static sampling distribution. By adaptively searching in multiple cone-shaped subspaces, ACS obtains better accuracy and efficiency. This result is validated by our experiments. For SRAM bit cell with single failure region, ACS requires 3-5X fewer samples and achieves better accuracy compared with existing approaches. For 576-dimensional SRAM column circuit with multiple failure regions, ACS is 2050X faster than MC without compromising accuracy, while other methods fail to converge to correct failure probability in our experiment.
Xiao Shi 0001, Hao Yan 0002, Xiaofen Xu, Longxing Shi, Lei He 0001
ISPD1
2018 A fast and robust failure analysis of memory circuits using adaptive importance sampling method
abstract
Performance failure has become a growing concern for the robustness and reliability of memory circuits. It is challenging to accurately estimate the extremely small failure probability when failed samples are distributed in multiple disjoint failure regions. In this paper, we develop an adaptive importance sampling (AIS) method. AIS has several iterations of sampling region adjustments, while existing methods pre-decide a static sampling distribution. By iteratively searching for failure regions, AIS may lead to better efficiency and accuracy. This is validated by our experiments. For SRAM cell with single failure region, AIS uses 5-10X fewer samples and reaches better accuracy when compared to several recent methods. For sense amplifier circuit with multiple failure regions, AIS is 4369X faster than MC without compromising accuracy, while other methods fail to cover all failure regions in our experiment.
Xiao Shi 0001, Jun Yang 0006, Lei He 0001
DAC1