EDBT 2026 Demo / reviewers in the wild / expert
Xinming Yan
dblp:166/3729
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-7661-6165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Low Bit-Width LLM Acceleration via Symmetric Lookup Format and Compute-in-Decoding Paradigm
Zihan Zou, Jiaming Lin, Xinming Yan, Shikuang Chen, Chen Zhang 0025, Xilong Kang, Hao Cai 0001, Bo Liu 0019 |
IEEE Trans. Computers | 5 |
| 2025 | TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device DeploymentabstractVision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios. Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 2 |
| 2025 | S-DMA: Sparse Diffusion Models Acceleration via Spatiality-Aware Prediction and Dimension-Adaptive DataflowabstractDiffusion Models (DMs) have demonstrated remarkable performance in a variety of image generation tasks.However, their complex architectures and intensive computations result in significant overhead and latency, posing challenges for hardware deployment.To address these issues, researchers have explored the sparsity in DMs to reduce computational workloads, including semantic sparsity in image generation and spatial sparsity in local editing.Unfortunately, existing sparsity prediction methods face critical limitations in deployment: 1) additional prediction overheads offset the benefits of sparsity; 2) convolution and general matrix multiplication (GEMM) exhibit distinct sparsity patterns, which current co-design frameworks struggle to process.In this paper, we introduce S-DMA, a software-hardware co-design framework that unifies efficient sparsity prediction while supporting various sparse operators.First, we propose a spatiality-aware similarity computation method that leverages the local similarity of images, reducing the computational complexity of sparsity prediction from O(𝑁 2 ) to O(N ).Second, we implement NAND-based similarity for sparsity prediction, which minimizes the computational overheads and ensures adaptability to different sparsity schemes.Finally, a dedicated hardware architecture is designed to efficiently leverage the algorithm optimizations.A NAND-based sparsity prediction processing unit is designed to adaptively handle the sparsity patterns.Additionally, a sparsity-aware reduction network and a dimension-adaptive Bo Liu is the corresponding author. Zihan Zou, Xinming Yan, Guang Yang 0036, Hao Cai 0001, Bo Liu 0019 |
MICRO | 2 |
| 2025 | PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision TransformerabstractVision Transformers (ViTs) have achieved outstanding performance in visual applications. However, ViT’s increasing parameters and computation overhead limit its deployment on hardware. Previous ViT accelerators focus on optimizing the core attention mechanism of ViTs due to the high overheads of language transformer-based neural networks. Nevertheless, linear layers are the actual bottleneck of ViT inference, accounting for larger than 90% FLOPs on numerous ViTs because of the short and fixed token length of ViTs. To this end, we propose PF2A-ViT, an algorithm and accelerator co-design framework, comprehensively accelerating both core attention and linear layers. At the algorithm level, a parameter-free and feature-aware dynamic token pruning (PF2ATP) strategy is proposed to reduce the dimension of feature maps and dynamically adjust the pruning ratio according to the complexity of the feature without complex execution of subnetworks. Meanwhile, a mixed-precision quantization strategy combines Hessian trace and parameter-aware signal-to-quantization-noise ratio to boost the deployment efficiency of PF2ATP. In addition, a cluster-regroup-based weight encoding strategy is proposed to compensate for the bit-wise redundant information of the quantization strategy. At the hardware level, a token pruning module based on bitonic sorters is designed to fully leverage PF2ATP. Simultaneously, a 3D-PE array with reconfigurable 4/8-bit processing elements (PE) is designed to implement the mixed-precision quantization strategy and equipped with greedy bit-wise compensation decoders to exploit the encoding strategy. Extensive experiments on multiple ViTs demonstrate the achievements of PF2A-ViT: (1) Maximally realize 3.89× speedup, 5.54× energy efficiency compared to state-of-the-art ViT accelerators. (2) Occupying a 2.25 mm area and consuming 76 mW power in 28-nm technology. Zihan Zou, Xinming Yan, Chen Zhang 0025, Shikuang Chen, Guang Yang 0036, Han Yan 0014, Hao Cai 0001, Bo Liu 0019 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Improving ELM-Based Time Series Classification by Diversified Shapelets Selection
Qifa Sun, Qiuyan Yan, Xinming Yan, Wei Chen 0036 |
QSHINE | 3 |