Jingkui Yang

dblp:399/1709 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result Reuse
abstract
Mixture-of-Experts (MoE) models significantly reduce computation in large language models by activating only a subset of experts per input token, but they introduce severe memory bottlenecks due to the large number of expert parameters. Existing offloading and prefetching strategies either incur accuracy loss, prohibitively high memory traffic, or high decoding overhead, limiting deployment on resource-constrained hardware. In this work, we present EARTH, a hardware–software co-design that addresses these challenges through three key innovations. First, we propose a dual-entropy encoding scheme that decomposes each expert into a high-information base and a delta component, enabling compact storage while preserving accuracy via adaptive precision management. Second, we introduce a delta-aware speculative prefetching and reuse mechanism that preloads base components of predicted experts and selectively fetches deltas, reusing previously computed delta patterns to reduce memory traffic and redundant computation. Third, we design a hardware accelerator that is co-designed to efficiently support this encoding and prefetching strategy, optimizing execution order, parallelism, and memory utilization. Across representative MoE workloads, EARTH reduces data movement overhead, improves prefetch efficiency, and achieves up to 2.10× speedup compared to state-of-the-art baselines, while maintaining high model accuracy.
Fangxin Liu, Ning Yang 0012, Jingkui Yang, Zongwu Wang, Chenyang Guan, Yu Feng 0007, Li Jiang 0002, Haibing Guan
ASPLOS (2)3
2026 Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix Multiplication
Jingkui Yang, Fangxin Liu, Ning Yang 0012, Chenyang Guan, Zongwu Wang, Mei Wen, Li Jiang 0002, Haibing Guan
ISCA1
2026 Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen
J. Syst. Archit.7
2025 SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level Sparsity
abstract
To more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs.
Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008
DATE1
2025 Secure Token Pruning Mechanism and Accelerator for Vision Transformer
abstract
Vision Transformers (ViTs) are vulnerable to adversarial patch attacks, posing serious challenges to their deployment in security-critical applications. Existing defense strategies improve robustness but at the cost of significant computational overhead, limiting their practicality. To address this, we propose STEM, a lightweight and parallelizable token pruning defense mechanism that enhances robustness while reducing inference cost. STEM integrates two components: Block-wise Selective Fusion (BSF), which fuses redundant tokens and flags suspicious ones, and Identify-Prune (IP), which identifies and prunes adversarial tokens based on multi-head attention statistics.To support STEM efficiently in hardware, we further design STEMA, a scalable accelerator featuring a dedicated Security Core operating in parallel with the ViT backbone. STEMA includes specialized engines for token fusion, top-k selection, and metadata management, occupying only 5% of chip area.Experimental results show that STEM achieves state-of-the-art (SOTA) defense performance and demonstrates excellent efficiency. Specifically, STEM improves robustness by up to 4.7× compared to token pruning methods and achieves 46.8×–105.8× runtime improvements over defense baselines. STEMA further delivers up to 71× speedups and 397× energy efficiency improvements, demonstrating its suitability for secure and efficient ViT deployment, especially in edge environments.
Qiuran Li, Jingkui Yang, Fanjin Xu, Jinjin Shao
ICCAD2
2025 SmartBlock: Adaptive Block Floating Point Quantization for Efficient DNN Acceleration
abstract
Deep Neural Networks (DNNs) have achieved remarkable success as model sizes continue to grow, driving the need for optimizations in both computational and energy efficiency. Block Floating Point (BFP) quantization has emerged as an effective model compression technique, offering a favorable trade-off between model accuracy and hardware cost. However, the frequent use of floating-point (FP) accumulation across BFP blocks remains a significant bottleneck, limiting further improvements in energy efficiency. State-of-the-art (SotA) accelerators mitigate this issue by introducing low-overhead accumulators with a narrower dynamic range ahead of the FP accumulator to handle a small range of values. While this approach reduces the activation of power-hungry alignment and format conversion units, it increases the complexity of the processing elements (PEs), thereby limiting the overall energy savings.
Xin Ju 0005, Jingkui Yang, Mei Wen, Minjin Tang, Zhaoyun Chen, Yang Shi 0008
ICPP2
2024 MSA2: An Efficient Sparsity-Aware Accelerator for Matrix Multiplication with Multi-core Systolic Arrays
Minjin Tang, Mei Wen, Junzhong Shen, Jingkui Yang, Zeyu Xue, Zili Shao
ICA3PP (3)4