Yongkui Yang

dblp:154/0702 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0003-1159-3115ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A 14.3-ENOB, 600-kHz BW Third-Order NS-SAR ADC with Enhanced EF-CIFF Filter
Changze Yu, Yishan Wang, Yongkui Yang
ISCAS5
2026 A Lightweight PUF-Based Secure Group Communication Scheme for Low Altitude Network With Dynamic Group Membership
abstract
Low Altitude Network (LAN) has emerged as a critical infrastructure for applications such as surveillance and emergency response. Group communication for LAN offers enhanced energy efficiency and reduced network overhead. However, existing group communication protocols encounter difficulties in managing the rekeying process efficiently for a dynamic group or require computationally expensive public key primitives for shared secret handshaking to overcome this challenge. Moreover, the security of most of the existing protocols relies primarily on the safekeeping of some secrets on the group members' devices. To overcome these limitations, we propose a novel physical unclonable function (PUF)-based lightweight secure group communication protocol. The proposed protocol utilizes a combination of the device's PUF and the one-time pad (OTP) to eliminate secure key storage at both group verifier and prover nodes, and achieves perfect forward secrecy (PFS) by eliminating dependence on static long-term secrets. The proposed protocol also supports efficient group key renewal by using a full binary tree as a secret vault for sharing and updating distributed secrets with the Chinese Remainder Theorem (CRT). Meantime, this data structure also reduces the computation and communication complexity for key renewal to$O(\log _{2}N)$at the cluster head and$O(1)$at the sensor nodes. A comparative analysis shows that the proposed protocol surpasses related protocols in terms of security features and overheads in computation, communication, as well as secret storage requirements. The proposed protocol was also validated by formal security analyses and a physical LAN implementation using Ultra96-V2 boards as cluster nodes.
Harishma Boyapally, Wenye Liu, Yongkui Yang, Chip-Hong Chang
IEEE Trans. Mob. Comput.4
2026 A Memory-Optimized Constant Geometry NTT-Based Polynomial Multiplier With a Conflict-Free Bit-Reverse Reordering Method
abstract
The number theoretic transform (NTT) and its inverse (INTT) are frequently employed to accelerate polynomial multiplication, which represents both the core computational paradigm and the performance bottleneck in lattice-based cryptography (LBC). As an important variant in the NTT algorithm family, the constant geometry (CG) NTT has received significant attention from many researchers due to its simple and consistent memory access pattern. However, it suffers from two inherent and unsatisfactory limitations: its storage capacity requirement for$\boldsymbol {N}$polynomial coefficients will exceed$\boldsymbol {N}$to ensure conflict-free read and write; and it inevitably requires data relocation in memory between NTT and INTT. To address these challenges, this work proposes a modified ping-pong memory structure with reduced size and a novel conflict-free bit-reverse reordering method. For the modified ping-pong memory structure, each PE is allocated only two memory banks, and the overall storage requirement is reduced to$1.5\boldsymbol {N}$. The proposed bit-reverse reordering algorithm, through simple logic and low-resource overhead, effectively avoids read–write conflicts and achieves seamless data relocation across the memory without stalling or additional space. Furthermore, we also introduce a lightweight twiddle factor memory access scheme, which, combined with the above techniques, drives down the resource consumption of both control logic and memory to a compact level. Finally, the FPGA implementation results of the proposed architecture demonstrate significant performance advantages over existing works in terms of area efficiency.
Jinyang Hu, Yongkui Yang, Enyi Yao
IEEE Trans. Very Large Scale Integr. Syst.2
2026 Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
abstract
While recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation (TM)—centered on high-volume data movement with minimal computation—remains underexplored. This work addresses that gap by introducing the TM unit (TMU): a reconfigurable, near-memory hardware block designed to execute data-movement-intensive (DMI) operators efficiently. The TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling support for both a wide range of coarse- and fine-grained tensor transformations. The proposed architecture integrates the TMU alongside a TPU within a high-throughput AI system-on-chip (SoC), leveraging double buffering and output forwarding to improve pipeline utilization. The TMU, synthesized under the SMIC 40-nm standard cell library, occupies only$0.019~\mathrm {\text {mm}^{2}}$while supporting over 10 representative DMI operators. Benchmarking shows that the TMU alone achieves up to$82.42\times $and$11.06\times $operator-level latency reduction over ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 22.89% reduction in end-to-end inference latency, demonstrating the effectiveness in reducing inference latency and the scalability of the TMU architecture across diverse tensor operators.
Weiyu Zhou, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Zhuoyu Wu, Anupam Chattopadhyay
IEEE Trans. Very Large Scale Integr. Syst.5
2025 AttenPU: An Area Efficient Attention Processor with Reconfigurable FP8 Precision and Dataflow
abstract
Efficient numerical representation is crucial for deep learning accelerators, especially for large language models (LLMs). The 8-bit-floating-point (FP8) data representation achieves higher precision and fewer quantization efforts than integer, which has been proven inevitable in attention-based accelerators for LLMs. Therefore, area-efficient design techniques for FP8 play a central role in lowering LLMs chip’s budget. This paper presents AttenPU, which is built upon reconfigurable FP8 units and supports E4M3 for inference and E5M2 for training. Bidirectional dataflow is exploited to enable AttenPU to interact with FP32 coprocessor to reduce latency. The design achieves a low FP8-to-INT8 area ratio of 1.63×, an area efficiency of 193.5 GFLOPS/mm2, with an 87.05% reduction in the latency of RTX3090 GPU.
Qiawei Zheng, Zheng Wang 0027, Zhuoyu Wu, Zhihao Du, Chao Chen 0022, Yongkui Yang, Wenqi Fang, Anupam Chattopadhyay
ACM Great Lakes Symposium on VLSI8
2025 SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and O(T) Complexity
Shihao Zou, Yongkui Yang
ICML5
2025 HSCIM: A High Security Compute-In-Memory Architecture with PUF based on TST-MRAM
abstract
With the rapid development of the Internet of Things (IOT) in the decades, compute-in-memory (CIM) architecture which addresses the Von-Neumann bottleneck are drawing significant attention with a gradually improving demand of high security. Physical unclonable function (PUF) emerges as a promising candidate attributed to its dependence on physical properties of devices instead of traditional key systems. Toggle spin torques magnetoresistive random access memory (TST-MRAM) is considered as satisfying storage technology due to its non-volatility and low power consumption. In this paper, a high security compute-in-memory (HSCIM) architecture based on TST-MRAM is proposed, aiming to generating PUF signals for encryption and guaranteeing security during computation. Simulation results show that with the proposed architecture, a 16-bit multiply-and-accumulate (MAC) operation along with data encryption and decryption can be performed within four clock cycles, achieving an efficient utilization of device resource.
Junyi Mai, Feilan Zhao, Zhanhong Huang, Yongkui Yang, Enyi Yao
ISCAS4
2025 A High-Density eDRAM Macro With Programmable Sense Amplifier and TG-Shifter for Logical-Instruction-Based In-Memory Computing
abstract
Embedded DRAM (eDRAM) has been widely adopted as on-chip cache memory in modern processors due to its high density. In this article, we propose a 2T gain-cell eDRAM-based macro that functions not only as traditional cache memory but also as an in-memory computing unit capable of performing logic operations. Furthermore, this eDRAM macro features in situ storing, completely eliminating the need for external memory or register access during computation. The sense amplifier in this macro is equipped with a programmable voltage reference, enabling support for various Boolean logic operations, includingand/nand,or/nor, andnot. In addition, the macro integrates a transmission-gate (TG)-based shifter cluster to perform data shifting, which is commonly required in general computations. To enhance functionality, we design an instruction set that supports compound logic computations, allowing Boolean logic, shifting, and in situ storage to be executed within a single instruction. We validated this eDRAM macro in a 32-kb bitcell array using the 40-nm logic CMOS technology. Compared with state-of-the-art designs, our macro achieves a relatively high density of 729.2 kb/mm2and a competitive logic energy of 14.1 fJ/bit.
Kunyao Lai, Enyi Yao, Yongkui Yang
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Guser: A GPGPU Power Stressmark Generator
abstract
Power stress mark is crucial for estimating Thermal Design Power (TDP) of GPGPUs to ensure efficient power control. This paper proposes Guser, the first systematic methodology to generate GPGPU power stressmarks. It features three ideas: Instruction Power Analysis (IPA) to analyze the power behavior of PTX instructions; Pipeline-based Instruction Grouping (PIG) to classify all PTX instructions into a number of groups; and Quantifying the Importance of power impact factors (QIF) to select a small number but important adjustable knobs. We adopt Optuna, an advanced BO algorithm, to generate stressmarks that maximize power consumption. We evaluate Guser on two real GPGPUs Tesla T4 (Turing) and Tesla A10 (Ampere). The experimental results show that the Guser-generated stressmark for T4 (T4-stresser) consumes 109.3 watts and the one for A10 (A10-stresser) consumes 238.7 watts, which are 48.7% and 73% higher than those consumed by the stressmarks generated by the state-of-the-art approach for CPUs, respectively. Moreover, T4-stresser and A10-stresser consume significantly more power than that consumed by any benchmarks in three benchmark suites: Cactus, Rodinia, and Parboil.
Yalong Shan, Yongkui Yang, Xuehai Qian, Zhibin Yu 0001
HPCA2
2024 Low-latency Buffering for Mixed-precision Neural Network Accelerator with MulTAP and FQPipe
abstract
Previous work has proposed precision scalable accelerators to handle mixed-precision neural network (NN) inferences on the edge, which focus on designing reconfigurable MAC arrays while leaving the issue of time-costly data buffering procedure less discussed. Besides, integer-only inference is incapable of handling emerging NN models with various non-linear activation functions. In this work, we propose a mixed-precision NN accelerator supporting int8, int16 and fp32 arithmetic with two buffering techniques namely MulTAP and FQPipe, which jointly facilitate low-latency data movement. Experiment results show that MulTAP and FQPipe boost the baseline NN accelerator with 7.7 × and 1.5 × in speed respectively, which leads to the application performance of 473.9 (int8) and 252.5 (int16) inferences per second (IPS) on YOLOv3-Tiny. Post-layout netlist with SMIC 40nm standard-cell technology demonstrates a design with an area of 26.96mm2and a power estimate of 1.83W.
Zheng Wang 0027, Wenhui Ou, Weiyu Zhou, Yongkui Yang, Chao Chen 0022
ISCAS6
2024 SpGesture: Source-Free Domain-adaptive sEMG-based Gesture Recognition with Jaccard Attentive Spiking Neural Network
abstract
Surface electromyography (sEMG) based gesture recognition offers a natural and intuitive interaction modality for wearable devices. Despite significant advancements in sEMG-based gesture recognition models, existing methods often suffer from high computational latency and increased energy consumption. Additionally, the inherent instability of sEMG signals, combined with their sensitivity to distribution shifts in real-world settings, compromises model robustness. To tackle these challenges, we propose a novel SpGesture framework based on Spiking Neural Networks, which possesses several unique merits compared with existing methods: (1) Robustness: By utilizing membrane potential as a memory list, we pioneer the introduction of Source-Free Domain Adaptation into SNN for the first time. This enables SpGesture to mitigate the accuracy degradation caused by distribution shifts. (2) High Accuracy: With a novel Spiking Jaccard Attention, SpGesture enhances the SNNs' ability to represent sEMG features, leading to a notable rise in system accuracy. To validate SpGesture's performance, we collected a new sEMG gesture dataset which has different forearm postures, where SpGesture achieved the highest accuracy among the baselines ($89.26\%$). Moreover, the actual deployment on the CPU demonstrated a latency below 100ms, well within real-time requirements. This impressive performance showcases SpGesture's potential to enhance the applicability of sEMG in real-world scenarios. The code is available at https://github.com/guoweiyu/SpGesture/.
Weiyu Guo, Ying Sun 0006, Yijie Xu, Ziyue Qiao, Yongkui Yang, Hui Xiong 0001
NeurIPS5
2024 Falic: An FPGA-Based Multi-Scalar Multiplication Accelerator for Zero-Knowledge Proof
abstract
In this paper, we propose Falic, a novel FPGA-based accelerator to accelerate multi-scalar multiplication (MSM), the most time-consuming phase of zk-SNARK proof generation. Falic innovates three techniques. First, it leverages globally asynchronous locally synchronous (GALS) strategy to build multiple small and lightweight MSM cores to parallelize the independent inner product computation on different portions of the scalar vector and point vector. Second, each MSM core contains just one large-integer modular multiplier (LIMM) that is multiplexed to perform the point additions (PADDs) generated during MSM. We strike a balance between the throughput and hardware cost by batching the appropriate number of PADDs and selecting the computation graph of PADD with proper parallelism degree. Finally, the performance is further improved by a simple cache structure that enables the computation reuse. We implement Falic on two different FPGAs with different hardware resources, i.e., the Xilinx U200 and Xilinx U250. Compared to the prior FPGA-based accelerator, Falic improves the MSM throughput by$3.9\boldsymbol{\times}$. Experimental results also show that Falic achieves a throughput speedup of up to$1.62\boldsymbol{\times}$and saves as much as$8.5\boldsymbol{\times}$energy compared to an RTX 2080Ti GPU.
Yongkui Yang, Zhenyan Lu, Jingwei Zeng, Xingguo Liu, Xuehai Qian, Zhibin Yu 0001
IEEE Trans. Computers1
2024 An Ising Model-Based Parallel Tempering Processing Architecture for Combinatorial Optimization
abstract
Combinatorial optimization problems (COPs) are prevalent in various domains and present formidable challenges for modern computers. Searching for the ground state of the Ising model emerges as a promising approach to solve these problems. Recent studies have proposed some annealing processing architectures based on the Ising model, aimed at accelerating the solution of COPs. However, most of them suffer from low solution accuracy and inefficient parallel processing. This article presents a novel parallel tempering processing architecture (PTPA) based on the fully-connected Ising model to address these issues. The proposed modified parallel tempering algorithm supports multi-spin concurrent updates per replica and employs an efficient multi-replica swap scheme, with fast speed and high accuracy. Furthermore, an independent pipelined spin update architecture is designed for each replica, which supports replica scalability while enabling efficient parallel processing. The PTPA prototype is implemented on FPGA with 8 replicas, each with 1,024 fully-connected spins. It supports up to 64 spins for concurrent updates per replica and operates at 200 MHz. Different concurrency strategies are considered to further improve the efficiency of solving COPs. In the test of various G-set problems, PTPA achieves 3.2× faster solution speed along with 0.27% better average cut accuracy compared to a state-of-the-art FPGA-based Ising machine.
Yang Zhang 0120, Xiangrui Wang, Gaopeng Fan, Yuan Cao 0003, Yiqiu Liu, Yongkui Yang, Enyi Yao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 COMPACT: Co-processor for Multi-mode Precision-adjustable Non-linear Activation Functions
abstract
Non-linear activation functions imitating neuron behaviors are ubiquitous in machine learning algorithms for time series signals while also demonstrating significant gain in precision for conventional vision-based deep learning networks. State-of-the-art implementation of such functions on GPU-like devices incurs a large physical cost, whereas edge devices adopt either linear interpolation or simplified linear functions leading to degraded precision. In this work, we design COMPACT, a co-processor with adjustable precision for multiple non-linear activation functions including but not limited to exponent, sigmoid, tangent, logarithm, and mish. Benchmarking with state-of-the-arts, COMPACT achieves a 26% reduction in the absolute error on a 1.6x widen approximation range taking advantage of the triple decomposition technique inspired by Hajduk's formula of Padé approximation. A SIMD-ISA-based vector co-processor has been implemented on FPGA which leads to a 30% reduction in execution latency but the area overhead nearly remains the same with related designs. Furthermore, COMPACT is adjustable to 46% latency improvement when the maximum absolute error is tolerant to the order of 1E-3.
Wenhui Ou, Zhuoyu Wu, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang
DATE5
2023 A Network-on-Chip-Based Annealing Processing Architecture for Large-Scale Fully Connected Ising Model
abstract
Combinatorial optimization problems are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with a number of spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored some hardware implementation methods of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a novel network-on-chip-based annealing processing architecture (NoCAPA) for a large-scale Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. Several techniques are developed to further increase convergence speed and reduce hardware resource consumption, including a dynamic multithread parallel update algorithm, a router with merge and deflection abilities, and a unique multiply-accumulate operation. The prototype is implemented in FPGA with the maximum operation frequency of 200MHz, achieving up to$120.5\times $faster than conventional simulated annealing method when solving the max-cut problem while supporting high scalability.
Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yongkui Yang, Enyi Yao
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002
ECCV (32)3
2021 OR-ML: Enhancing Reliability for Machine Learning Accelerator with Opportunistic Redundancy
abstract
Reliability plays a central role in deep sub-micron and nanometre IC fabrication technology and has recently been reported to be one of the key issues affecting the inference phase of neural networks. State-of-the-art machine learning (ML) accelerators exploit massively computing parallelism observed in neural networks to achieve high energy efficiency. The topology of ML engines' computing fabric, which constitutes large arrays of processing elements (PEs), has been increasing dramatically to incorporate the huge size and heterogeneity of the rapid evolving ML algorithm. However, it is commonly observed that activations of zero value lead to reduced PE utilization. In this work, we present a novel and low-cost approach to enhance the reliability of generic ML accelerators by Qpportunistically exploring the chances of runtime Redundancy provided by neighbouring PEs, named as OR-ML. In contrast to conventional redundancy techniques, the proposed technique introduces no additional computing resources, therefore significantly reduces the implementation overhead and achieves obvious level of protection. The design prototype is evaluated using emulated fault injection on FPGA, executing mainstream neural networks for objectionclassification and detection.
Zheng Wang 0027, Wenxuan Chen, Chao Chen 0022, Yongkui Yang, Zhibin Yu 0001
DATE5
2021 CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window Filtering
abstract
Memory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering.
Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001
ACM Great Lakes Symposium on VLSI5