VLDB 2026 Research / reviewers in the wild / expert
Yueting Li 0001
dblp:17/8346-1
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0001-8874-6269ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: Algorithm-Hardware Co-Design of a Sparsity-Aware Dense-Sparse Scheme for DNN AcceleratorsabstractDeep neural networks (DNNs) in modern applications have increased the demand for energy-efficient DNN inference solutions, especially on resource-constrained platforms. However, the growing model capacity of DNNs incurs significant memory traffic and energy consumption. To address these challenges, we propose a novel solution that presents an algorithm-hardware co-design for reconfigurable DNN acceleration. This design exploits value- and bit-level sparsity to minimize memory footprint and enhance computational efficiency. To achieve this, the proposed algorithm leverages a static dense-sparse storage format, along with a dynamic bit-processing scheme that removes non-contributing bits. Building on this algorithm, a flexible processing element array is designed to perform LUT-based shift-accumulate operations, with fine-grained per-layer configurability. Experimental results show that this design yields 13–24% storage savings across the evaluated DNN models, while delivering up to 8.4× effective sparsity. Based on post-implementation FPGA results (from our RTL design), the proposed accelerator delivers 1.41× lower LUT usage than state-of-the-art design at similar throughput. Yueting Li 0001, Terry Tao Ye, Weisheng Zhao 0001 |
DATE | 1 |
| 2026 | ReNN-RV: Run-Time PE Reconfiguration for DNN Inference Acceleration With Custom RISC-V ISAabstractDeep neural network (DNN) accelerators integrated with RISC-V Instruction Set Architecture (ISA) extensions have enabled efficient computing on resource-constrained platforms. However, their specialization in regular compute patterns limits effectiveness on irregular workloads, making it challenging to achieve high throughput and energy efficiency. To tackle these challenges, we present ReNN-RV, which integrates a computation-aware RISC-V ISA extension with an instructiondriven processing pipeline to efficiently accelerate run-time reconfigurable processing elements (RePEs). The computation-aware ISA employs configurable opcodes and custom encodings to support fine-grained task scheduling, while an instructiondriven pipeline implements it with minimal control complexity. Moreover, theRePEaccelerator provides seamless switching between multiply-accumulate (MAC) and non-MAC operations by configuring a path multiplexer to realize multiple operators at run time. Experimental results demonstrate that ReNN-RV achieves average reductions of 14.6× in cycle count and 15.3× in execution time across representative DNN workloads compared with the baseline RISC-V design. On average, ReNN-RV outperforms state-of-the-art designs by 10.1× for energy efficiency and 10.3× for computational throughput. Yueting Li 0001, Terry Tao Ye, Ngai Wong 0001, Zhenhua Zhu 0002, Yongfu Li 0002, Weisheng Zhao 0001 |
IEEE Trans. Computers | 1 |
| 2026 | PPD: A Portable and Highly Parallel Dispatching System for Deep Learning
Wendong Xu, Yuhao Ji, Yueting Li 0001, Yuxuan Zhao 0001, Zhengwu Liu, Bei Yu 0001, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | A Custom RISC-V ISA with Scalable Processing Units for Efficient Neural Network InferenceabstractA customized RISC-V ISA with integrated digital accelerators offers a promising solution to improve energy efficiency in neural network inference.However, it often requires multiple instructions per accelerator operation, which limits computational efficiency during deep neural network inference.To overcome the instruction overhead, this design introduces a dedicated instruction set that enables scalable and fine-grained accelerator control.By incorporating the pattern-driven instruction mode, this design exploits the neural layer regularity to support efficient instruction iteration.Furthermore, this digital accelerator leverages hardware reuse for logic operations, forming a fusion-style architecture that integrates reconfigurable components.Experimental results demonstrate that the custom RISC-V ISA achieves an average runtime speedup of 8.26× and reduces the instruction count by 14.71×.This design also yields an average 8.73× reduction in cycles per instruction across MobileNetV2, ResNet50, VGG19, EfficientNet, and DenseNet-BC, validating its effectiveness across representative benchmarks.Additionally, it improves average energy efficiency by 1.74×, outperforming state-of-the-art designs. Yueting Li 0001, Wanshuang Lin, Wendong Xu, Ngai Wong 0001, Weisheng Zhao 0001 |
CF | 1 |
| 2025 | ACSNN: A 61.25 TOPS/W, 1.65 ns delay SNN Processor that combines CIM-inspired Synapse and Asynchronous ArchitectureabstractSpiking Neural Networks (SNNs) offer their biological plausibility and dynamic sensitivity which have gained significant attention in real-time systems. However, designing SNN processors with high throughput and real-time response remains challenging due to high power consumption, high area costs and routing competition. In this paper, we present ACSNN, a processor that integrates a CIM-inspired synapse array, Leaky Integrate-and-Fire (LIF) neurons and asynchronous architecture to achieve high parallelism, high energy and area efficiency while declining routing competition. Experimental results demonstrate an impressive power efficiency of 61.25 TOPS/W and a high peak throughput of 2,415 GOPS with 1.65 ns minimum compute delay, highlighting superior performance and power efficiency compared to state-of-the-art designs. Tingran Chen, Yuxuan Ran, Yueting Li 0001, Wang Kang 0001, Biao Pan |
ISCAS | 4 |
| 2024 | APIM: An Antiferromagnetic MRAM-Based Processing-In-Memory System for Efficient Bit-Level Operations of Quantized Convolutional Neural NetworksabstractQuantized Convolutional Neural Network (QCNN) is an attractive approach that reduces hardware overheads, especially for energy-constrained systems. However, existing QCNNs still require non-trivial hardware resources and memory capacity in order not to compromise model accuracy. To address this issue, we propose an antiferromagnetic magnetic random-access memory (ARAM)-based processing-in-memory (PIM) system, leveraging bit-level sparsity. Three optimization techniques are proposed to optimize hardware resource utilization while preserving CNN accuracy. Firstly, the ARAM-based memory subsystem allows dynamic adaptation of variable bit-width across CNN layers. Secondly, the bit-level accelerator employs the bit-fusion format engineered for processing data from the ARAM subsystem. Thirdly, a customized data path within the RISC-V core guarantees efficient instruction processing to the ARAM-based memory subsystem and bit-level accelerator, enabling optimal bit-level data transmission and computation. Experimental results demonstrate that this design remarkably reduces data movement by 50%-83% across existing CNNs. Compared to state-of-the-art designs, it enhances throughput and latency by an average of 5x and 10x, respectively. In addition, this design achieves speedups between 1.63x and 2.96x, outstripping other designs in AlexNet, VGG16, and ResNet18 benchmarks. Yueting Li 0001, Daoqian Zhu, Jinhao Li 0007, Ao Du, Yue Zhang 0010, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded SystemsabstractConvolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs. Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | Toward Energy-Efficient Sparse Matrix-Vector Multiplication with near STT-MRAM Computing ArchitectureabstractSparse Matrix-Vector Multiplication (SpMV) is one of the vital computational primitives used in modern workloads. SpMV performs memory access, leading to unnecessary data transmission, massive data access, and redundant multiplicative accumulators. Therefore, we propose the near spin-transfer torque magnetic random access memory (STT-MRAM) processing architecture from three optimization perspectives. These optimizations include (1) the NMP controller receives the instruction through the AXI4 bus to implement the SpMV operation in the following steps, identifies valid data, and encodes the index depending on the kernel size, (2) the NMP controller uses high-level synthesis dataflow in the shared buffer for achieving better performance throughput while do not consume bus bandwidth, and (3) the configurable MACs are implemented in the NMP core without matching step entirely during the multiplication. Using these optimizations, the NMP architecture can access the pipelined STT-MRAM (read bandwidth is 26.7GB/s). The experimental simulation results show that this design achieves up to 66x and 28x speedup compared with state-of-the-art ones and 69x speedup without sparse optimization. Yueting Li 0001, He Zhang 0011, Hao Cai 0001, Shuqin Lv, Renguang Liu, Weisheng Zhao 0001 |
ASP-DAC | 1 |
| 2023 | Experimental Demonstration of STT-MRAM-based Nonvolatile Instantly On/Off System for IoT Applications: Case StudiesabstractEnergy consumption has been a big challenge for electronic devices, particularly for battery-powered Internet of Things (IoT) equipment. To address such a challenge, on the one hand, low-power electronic design methodologies and novel power management techniques have been proposed, such as nonvolatile memories and instantly on/off systems; on the other hand, the energy harvesting technology by collecting signals from human activity or the environment has attracted widespread attention in the IoT area. However, the system with self-powered energy harvesting may suffer frequent energy failures or fluctuating energy conditions, which degrade system reliability and user experience. Therefore, how to make the system under unreliable power inputs operate correctly and efficiently is one of the most critical issues for energy harvesting technology. In this article, we built an instantly on/off system based on nonvolatile STT-MRAM for IoT applications, which can instantly power on/off under different conditions of the harvested energy. The system powers on and operates normally when the harvested energy is enough (over the preset threshold); otherwise, the system powers off and stores the operational data back to the nonvolatile STT-MRAM. We described implementations of the hardware/software co-designed architecture (with image acquisition as an example) based on the commercialized 32 MB STT-MRAM, and we experimentally demonstrated the system functionality and efficiency under five typical energy harvesting scenarios, including radio frequency, thermal, solar, piezoelectric, and WIFI. Our experimental results show that the power consumption and data restore time were reduced by 15.1% and 714 times, respectively, in comparison with the DRAM-based counterpart. Yueting Li 0001, Wang Kang 0001, Kunyu Zhou, Keni Qiu, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Work-in-Progress: Toward Energy-efficient Near STT-MRAM Processing Architecture for Neural NetworksabstractThe size of parameters in artificial neural network (NN) applications grows quickly from a handful to the GB-level. The data transmission poses a key challenge for NN, and either neuron is removed or data compression reduces pressure on memory access but cannot successfully decrease data traffic. Therefore, we propose the near spin-transfer-torque magnetic random processing architecture for developing energy-efficient NNs. Our approach provides system architects with a preliminary scheme to obtain real-time transmission that near memory controller directly compresses non-zero elements, and encodes the corresponding index depending on the kernel size. Furthermore, it adjusts the number of multiplication accumulators and avoids unnecessary hardware overheads during computation. The preliminary experimental results demonstrated this design verified with weights that currently achieve up to 3.05x speedup and 29.6% power compared with the unoptimized one. Yueting Li 0001, Bingluo Zhao, Jun Wang 0041, Weisheng Zhao 0001 |
CODES+ISSS | 1 |