VLDB 2026 Research / reviewers in the wild / expert
Enyi Yao
dblp:166/2914
· DBLP profile ↗
20ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-0019-2263ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | pHNSW: PCA-Based Filtering to Accelerate HNSW Approximate Nearest Neighbor SearchabstractHierarchical Navigable Small World (HNSW) has demonstrated impressive accuracy and low latency for high-dimensional nearest neighbor searches. However, its high computational demands and irregular, large-volume data access patterns present significant challenges to search efficiency. To address these challenges, we introduce pHNSW, an algorithm-hardware co-optimized solution that accelerates HNSW through Principal Component Analysis (PCA) filtering. On the algorithm side, we apply PCA filtering to reduce the dimensionality of the dataset, thereby lowering the volume of neighbor access and decreasing the computational load for distance calculations. On the hardware side, we design the pHNSW processor with custom instructions to optimize search throughput and energy efficiency. In the experiments, we synthesized the pHNSW processor RTL design with a 65nm technology node and evaluated it using DDR4 and HBM1.0 DRAM standards. The results show that pHNSW boosts Queries per Second (QPS) by $14.47 \times \sim 21.37 \times$ on a CPU and $5.37 \times \sim 8.46 \times$ on a GPU, while reducing energy consumption by up to $57.4 \%$ compared to standard HNSW implementation. Guangyi Zeng, Paul Delestrac, Enyi Yao, Simei Yang |
ASP-DAC | 4 |
| 2026 | A Multi-Precision Tensor Processing Unit for Accelerating Matrix Computations
Zongfan Wu, Dong Jiang 0002, Enyi Yao |
ISCAS | 4 |
| 2026 | A Memory-Optimized Constant Geometry NTT-Based Polynomial Multiplier With a Conflict-Free Bit-Reverse Reordering MethodabstractThe number theoretic transform (NTT) and its inverse (INTT) are frequently employed to accelerate polynomial multiplication, which represents both the core computational paradigm and the performance bottleneck in lattice-based cryptography (LBC). As an important variant in the NTT algorithm family, the constant geometry (CG) NTT has received significant attention from many researchers due to its simple and consistent memory access pattern. However, it suffers from two inherent and unsatisfactory limitations: its storage capacity requirement for$\boldsymbol {N}$polynomial coefficients will exceed$\boldsymbol {N}$to ensure conflict-free read and write; and it inevitably requires data relocation in memory between NTT and INTT. To address these challenges, this work proposes a modified ping-pong memory structure with reduced size and a novel conflict-free bit-reverse reordering method. For the modified ping-pong memory structure, each PE is allocated only two memory banks, and the overall storage requirement is reduced to$1.5\boldsymbol {N}$. The proposed bit-reverse reordering algorithm, through simple logic and low-resource overhead, effectively avoids read–write conflicts and achieves seamless data relocation across the memory without stalling or additional space. Furthermore, we also introduce a lightweight twiddle factor memory access scheme, which, combined with the above techniques, drives down the resource consumption of both control logic and memory to a compact level. Finally, the FPGA implementation results of the proposed architecture demonstrate significant performance advantages over existing works in terms of area efficiency. Jinyang Hu, Yongkui Yang, Enyi Yao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | An SRAM Compute-in-Memory based NTT Accelerator for CRYSTALS-KYBERabstractAmong various post-quantum cryptography (PQC) proposed by researchers, lattice-based cryptography is considered to be one of the most promising post-quantum public key cryptography systems. It has significant advantages over other PQC schemes in terms of security, computational efficiency, and versatility in designing public key encryption, digital signatures, and cryptographic negotiation protocols. Efficient implementations of the number theoretic transform (NTT) operations are crucial for many lattice-based encryption algorithms. This article presents an NTT hardware acceleration structure based on SRAM in-memory computing technology, which optimizes the key butterfly structure in the NTT algorithm by improving the 6T-SRAM array structure and incorporating near-memory computing structures. Compared with existing technologies, this accelerator can significantly reduce area and power consumption. Jinyang Hu, Xinyuan Pang, Dong Jiang 0002, Gaopeng Fan, Enyi Yao |
ISCAS | 5 |
| 2025 | HSCIM: A High Security Compute-In-Memory Architecture with PUF based on TST-MRAMabstractWith the rapid development of the Internet of Things (IOT) in the decades, compute-in-memory (CIM) architecture which addresses the Von-Neumann bottleneck are drawing significant attention with a gradually improving demand of high security. Physical unclonable function (PUF) emerges as a promising candidate attributed to its dependence on physical properties of devices instead of traditional key systems. Toggle spin torques magnetoresistive random access memory (TST-MRAM) is considered as satisfying storage technology due to its non-volatility and low power consumption. In this paper, a high security compute-in-memory (HSCIM) architecture based on TST-MRAM is proposed, aiming to generating PUF signals for encryption and guaranteeing security during computation. Simulation results show that with the proposed architecture, a 16-bit multiply-and-accumulate (MAC) operation along with data encryption and decryption can be performed within four clock cycles, achieving an efficient utilization of device resource. Junyi Mai, Feilan Zhao, Zhanhong Huang, Yongkui Yang, Enyi Yao |
ISCAS | 6 |
| 2025 | An Efficient and Flexible Hybrid Implementation of Pair and Triplet-Based STDP LearningabstractSpike-Timing-Dependent Plasticity (STDP) is a biologically inspired mathematical rule in which the timing difference between pre-synaptic and post-synaptic spikes dictates changes in synaptic strength. Due to its biological plausibility, STDP is widely employed in Spiking Neural Networks as a key training mechanism, facilitating efficient brain-like learning in neuromorphic systems. Two common variants, pair-based STDP and triplet-based STDP, are suited for diverse neural computing scenarios. However, improper use of these variants may degrade network performance or lead to undesired outcomes. In this paper, we propose a flexible hybrid STDP hardware implementation that supports two STDP learning rules and two computational schemes within a unified framework. We demonstrate its application by mapping the computational schemes to the MNIST classification task. The proposed architecture is validated on a Xilinx XA Spartan-7 FPGA. Compared to state-of-the-art approaches, our architecture improves power efficiency with a 21.5% reduction in resource consumption. Zongfan Wu, Zhibin Luo, Junyi Mai, Enyi Yao |
ISCAS | 4 |
| 2025 | A Spin Scale-Aware Self-Adaptive Ising Annealing Processing Architecture for Combinatorial Optimization ProblemsabstractThe Ising annealing processor has emerged as a promising approach to accelerate the discovery of the optimal solutions for a wide range of combinatorial optimization problems (COPs), by mapping various COPs into a unified Ising model. However, fixed computational strategies and inflexible architectures make previous designs suffer from a low hardware resource utilization rate when the numbers of the total required and real-time flipped spins vary across different COPs and iteration steps. In this paper, a novel spin scale-aware self-adaptive Ising annealing processing architecture (AIAPA) is proposed to address this problem, with an adaptive computational strategy, a custom instruction set, multi-traffic mode routers, and a fully-pipelined computing array. It can dynamically adapt to the varying scenarios during the Ising annealing process to maximize the performance of limited hardware resources. Its prototype, supporting 65k fully-connected spins, is implemented on an FPGA platform, operates at a clock frequency of 188 MHz. The AIAPA achieves up to a 24.22 times faster annealing speed compared to the state-of-the-art FPGA design on the max-cut optimization problem while maintaining a high convergence accuracy. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Longyuan Kang, Simei Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A High-Density eDRAM Macro With Programmable Sense Amplifier and TG-Shifter for Logical-Instruction-Based In-Memory ComputingabstractEmbedded DRAM (eDRAM) has been widely adopted as on-chip cache memory in modern processors due to its high density. In this article, we propose a 2T gain-cell eDRAM-based macro that functions not only as traditional cache memory but also as an in-memory computing unit capable of performing logic operations. Furthermore, this eDRAM macro features in situ storing, completely eliminating the need for external memory or register access during computation. The sense amplifier in this macro is equipped with a programmable voltage reference, enabling support for various Boolean logic operations, includingand/nand,or/nor, andnot. In addition, the macro integrates a transmission-gate (TG)-based shifter cluster to perform data shifting, which is commonly required in general computations. To enhance functionality, we design an instruction set that supports compound logic computations, allowing Boolean logic, shifting, and in situ storage to be executed within a single instruction. We validated this eDRAM macro in a 32-kb bitcell array using the 40-nm logic CMOS technology. Compared with state-of-the-art designs, our macro achieves a relatively high density of 729.2 kb/mm2and a competitive logic energy of 14.1 fJ/bit. Kunyao Lai, Enyi Yao, Yongkui Yang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A Parallel Tempering Processing Architecture with Multi-Spin Update for Fully-Connected Ising ModelsabstractCombinatorial optimization problems (COPs) are notoriously difficult to solve for classic Von-Neumann computers, which are ubiquitous in various domains. As a state-of-the-art hardware acceleration scheme for COPs, Ising machines are one of the promising research directions for the next generation of computing, but still suffer from the low solution accuracy and speed due to the high complexity of the fully-connected Ising model. In this work, a novel parallel tempering processing architecture (PTPA) is proposed with the modified parallel tempering algorithm, aimed at reducing search time and improving the solution quality. Several techniques are developed to further reduce hardware overhead and enhance parallelism, including the independent pipelined spin update architecture, approximated probability equations, and compact random number generators. Its prototype is implemented on FPGA with eight replicas, each replica containing 1,024 fully-connected spins and at most 64 concurrent update spins. The proposed design achieves an average cut accuracy of 99.43% within 1ms solution time on various G-set problems. Compared with the CPU-based parallel tempering implementation, it enhances the speed of solving the max-cut problems by 5,160 times. Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Zhanhong Huang, Gaopeng Fan, Enyi Yao |
DATE | 6 |
| 2024 | A Trusted Inference Mechanism for Edge Computing Based on Post-Quantum EncryptionabstractEdge computing is a computing framework that offers fewer computing resources compared to cloud computing but brings enterprise applications closer to data sources like Internet of Things (IoT) devices or local edge servers. This proximity to data sources brings significant business benefits such as faster insights, improved response time, and better bandwidth availability. Consequently, edge computing technologies have witnessed widespread adoption. Edge computing faces significant security challenges due to factors like limited cost, volume, and power consumption, as well as the potential risk of quantum computers breaking conventional cryptographic systems. Limitation of computing power in edge computing nodes further complicates the system architecture to achieve secure data transmission. To overcome these challenges, this paper proposes a lightweight trusted inference mechanism specifically designed for edge computing workloads. The primary objective of this mechanism is to establish a balance between security and bandwidth by incorporating a post-quantum encryption to the edge computing devices. This is achieved through sharing convolution and encryption hardware at the edge while encrypting and transmitting intermediate results of inference computations to the cloud, thereby minimizing resource usage and transmission bandwidth requirements. The FPGA platform is utilized to implement a template hardware design, which demonstrated comparable resource consumption to the stateof-art work. Simulation results indicate that our mechanism achieves a 45.3% reduction in LUTs and a 55.5% reduction in DFFs within the hardware, while introducing only a marginal software slowdown of 0.39%. Yukang Huang, Junyi Mai, Wanling Jiang, Enyi Yao |
ISCAS | 4 |
| 2024 | An Ising Model-Based Parallel Tempering Processing Architecture for Combinatorial OptimizationabstractCombinatorial optimization problems (COPs) are prevalent in various domains and present formidable challenges for modern computers. Searching for the ground state of the Ising model emerges as a promising approach to solve these problems. Recent studies have proposed some annealing processing architectures based on the Ising model, aimed at accelerating the solution of COPs. However, most of them suffer from low solution accuracy and inefficient parallel processing. This article presents a novel parallel tempering processing architecture (PTPA) based on the fully-connected Ising model to address these issues. The proposed modified parallel tempering algorithm supports multi-spin concurrent updates per replica and employs an efficient multi-replica swap scheme, with fast speed and high accuracy. Furthermore, an independent pipelined spin update architecture is designed for each replica, which supports replica scalability while enabling efficient parallel processing. The PTPA prototype is implemented on FPGA with 8 replicas, each with 1,024 fully-connected spins. It supports up to 64 spins for concurrent updates per replica and operates at 200 MHz. Different concurrency strategies are considered to further improve the efficiency of solving COPs. In the test of various G-set problems, PTPA achieves 3.2× faster solution speed along with 0.27% better average cut accuracy compared to a state-of-the-art FPGA-based Ising machine. Yang Zhang 0120, Xiangrui Wang, Gaopeng Fan, Yuan Cao 0003, Yiqiu Liu, Yongkui Yang, Enyi Yao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | DCAP: A Scalable Decoupled-Clustering Annealing Processor for Large-Scale Traveling Salesman ProblemsabstractThe Traveling Salesman Problem (TSP) is one of the most well-known NP-hard combinatorial optimization problems (COPs). Many social production problems can be effectively represented as instances of TSPs. However, solving large-scale TSPs remains a significant challenge for conventional Von Neumann computers. Many studies have proposed annealing processors to address large-scale COPs, but most of them focus on unconstrained problems, such as the Maxcut problem. In this paper, a scalable decoupled-clustering annealng processor (DCAP) for efficiently handling large-scale TSPs is presented. A decoupled hierarchical clustering algorithm is proposed for higher convergence speed and improved scalability. Several techniques have been developed in hardware to minimize area overhead and processing time, including a modified spin connection topology for the Ising model, an area-efficient random threshold generator, a one-step spin update scheme and a dynamic prediction method. The DCAP prototype is implemented on FPGA with an operating frequency of 125MHz. We tested our design on various TSP instances from the TSPLIB. Results show that our design outperforms the CPU- and GPU-based Neuro-Ising scheme by achieving maximum speedups of$780\times $and a 42% improvement in accuracy. With multi-chip interconnection, DCAP is able to handle problems of scale up to 85900 cities. Zhanhong Huang, Yang Zhang 0120, Xiangrui Wang, Dong Jiang 0002, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | An Annealing Processor based on 1k-Spin Fully-Connected Ising Model for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) find extensive applications in industrial and social scenarios such as transportation and communication. As the size of NP-hard COPs increases, it becomes impossible to obtain the optimal solution using an enumerative method. Recently, Ising model based annealing processors have received increasing attention due to their potential for rapidly converging to the near-optimal solutions after mapping the problem to them. This paper presents a novel annealing processor (AP) with 1024 fully-connected spins based on a modified Ising model annealing algorithm, which is more suitable for hardware implementation compared to conventional simulated annealing (SA) algorithm. The prototype is implemented using FPGA with the operation frequency up to 100MHz. We tested our design on various G-set problems with an average cut accuracy of 99.19% achieved. The proposed design outperforms the conventional CPU-based method by achieving a max speedup of 2204x for G51. Zhanhong Huang, Xiangrui Wang, Dong Jiang 0002, Yukang Huang, Enyi Yao |
ISCAS | 5 |
| 2023 | A Scalable Annealing Processing Architecture for Fully-Connected Ising ModelsabstractCombinational Optimization Problems (COPs) are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with numerous spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored the hardware implementation of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a scalable annealing processing architecture for Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. The prototype is synthesized using FPGA with the maximum operation frequency of 270MHz, achieving about 32 times faster than conventional simulated annealing method when solving the max-cut problem. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yukang Huang, Enyi Yao |
ISCAS | 5 |
| 2023 | A Network-on-Chip-Based Annealing Processing Architecture for Large-Scale Fully Connected Ising ModelabstractCombinatorial optimization problems are prevalent in many different fields. Most of these problems are NP-hard and challenging for computers with conventional Von-Neumann architecture. Ising machines with a number of spins have the potential to solve these problems by emulating the natural annealing process of solid matter. Recent research has explored some hardware implementation methods of Ising machines to accelerate the convergence process of such problems at room temperature. However, most of them are suffering from low scalability and low parallel processing capability due to the huge hardware cost and high complexity. In this paper, a novel network-on-chip-based annealing processing architecture (NoCAPA) for a large-scale Ising processor is described to address these issues with a NoC computing paradigm, a distributed storage scheme, and a fully pipelined structure design. Several techniques are developed to further increase convergence speed and reduce hardware resource consumption, including a dynamic multithread parallel update algorithm, a router with merge and deflection abilities, and a unique multiply-accumulate operation. The prototype is implemented in FPGA with the maximum operation frequency of 200MHz, achieving up to$120.5\times $faster than conventional simulated annealing method when solving the max-cut problem while supporting high scalability. Dong Jiang 0002, Xiangrui Wang, Zhanhong Huang, Yongkui Yang, Enyi Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2017 | Hardware architecture for large parallel array of Random Feature Extractors applied to image recognition
Aakash Patil, Shanlan Shen, Enyi Yao, Arindam Basu |
Neurocomputing | 3 |
| 2017 | VLSI Extreme Learning Machine: A Design Space ExplorationabstractIn this paper, we describe a compact low-power high-performance hardware implementation of extreme learning machine for machine learning applications. Mismatches in current mirrors are used to perform the vector-matrix multiplication that forms the first stage of this classifier and is the most computationally intensive. Both regression and classification (on UCI data sets) are demonstrated and a design space tradeoff between speed, power, and accuracy is explored. Our results indicate that for a wide set of problems, σ VTin the range of 15-25 mV gives optimal results. An input weight matrix rotation method to extend the input dimension and hidden layer size beyond the physical limits imposed by the chip is also described. This allows us to overcome a major limit imposed on most hardware machine learners. The chip is implemented in a 0.35-μm CMOS process and occupies a die area of around 5 mm × 5 mm. Operating from a 1 V power supply, it achieves an energy efficiency of 0.47 pJ/MAC at a classification rate of 31.6 kHz. Enyi Yao, Arindam Basu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Pulse-based feature extraction for hardware-efficient neural recording systemsabstractCurrent brain-machine interfaces have two machine learners-one for spike sorting and the second for intention decoding that acts on the sorted spatio-temporal spike train. In this paper, we propose a pulse-based feature extractor that can enable these two machine learners to be combined into one. We show from simulations and measurements that the information about the spike shape is still retained in the pulse counts-hence, the circuit can also be used as a traditional feature extractor. The proposed circuit also has the advantage of sharing several blocks with spike detector designs reducing system level cost. Fabricated in 65nm CMOS and operating from Vdd = 1V, the feature extractor dissipates roughly 2μW of power for an input spike rate of 100Hz. Aritra Bhaduri, Enyi Yao, Arindam Basu |
ISCAS | 2 |
| 2015 | A 128 channel 290 GMACs/W machine learning based co-processor for intention decoding in brain machine interfacesabstractA machine learning co-processor in 0.35μm CMOS for motor intention decoding in the brain-machine interfaces is presented in this paper. Using Extreme Learning Machine algorithm, time delayed sample based feature dimension enhancement, low-power analog processing and massive parallelism, it achieves an energy efficiency of 290 GMACs/W at a classification rate of 50 Hz. A portable external unit based on the proposed co-processor is verified with neural data recorded in monkey finger movements experiment, achieving a decoding accuracy of 99.3%. With time-delayed feature dimension enhancement, the classification accuracy can be increased by 5% with limited number of input channels. Yi Chen 0012, Enyi Yao, Arindam Basu |
ISCAS | 2 |
| 2015 | A 1 V, compact, current-mode neural spike detector with detection probability estimator in 65 nm CMOSabstractIn this paper, we describe a novel low power, compact, current-mode spike detector circuit for real-time neural recording systems where neural spikes or action potentials (AP) are of interest. Such a circuit can enable massive compression of data facilitating wireless transmission. This design operates by approximating the popularly used nonlinear energy operator (NEO) through standard current mode analog blocks that can operate at low voltages. To reduce sensitivity of threshold setting, this work uses a current-mode oscillator based detection probability estimator (DPE) to reject false positives caused by the background noise. The circuit is implemented in a 65 nm CMOS process and occupies 200 μm × 150 μm of chip area. Operating from a 1 V power supply, it consumes about 88 nW of static power and 10 nJ of dynamic energy per input spike. Enyi Yao, Arindam Basu |
ISCAS | 1 |