VLDB 2026 Research / reviewers in the wild / expert
Qiufeng Li
dblp:141/3803
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASPA: Reassigning DDR5 Parity BandwidthabstractRecent memory advancements, such as DDR5, HBM3, and emerging memory-centric accelerators, primarily focus on increasing bandwidth capacity, yet often omit the significance of effective bandwidth utilization, i.e., bandwidth efficiency. Motivated by the suboptimal channel allocation in DDR5, where parity accounts for 25% bandwidth overhead, we propose ASPA, an efficiency-oriented solution that reallocates parity bandwidth to boost regular data transfer. The objective of ASPA is to enhance bandwidth efficiency without compromising reliability while maintaining low hardware overhead. In particular, we leverage existing CRC (Cyclic Redundancy Check) units in DRAM chips to opportunistically generate second-tier parity for the existing ECC (Error Correction Code) parity. For bulk-sized memory accesses, only the small second-tier parity is transmitted. This reduces the parity bandwidth consumption, allowing data chips to reuse the freed parity bandwidth for data transfer. Our observation indicates that the protection capability is sufficient when the second-tier parity (with a 64-bit size) is used exclusively for error detection. Furthermore, by exploiting underutilized resources in high-performance memory systems, ASPA is implemented with negligible hardware overhead. Qiufeng Li, Yanan Guo 0002, Weidong Cao 0001, Xin Xin 0008 |
HPCA | 2 |
| 2026 | Frieren: A Fault-Tolerant Reconfigurable Energy-Efficient Computing Architecture With Enhanced Reliability in Harsh EnvironmentsabstractIn harsh environments such as space, strong radiation effects often induce single-event effects that threaten the reliability of computing systems. Meanwhile, edge artificial intelligence (AI) processors deployed in these conditions must not only tolerate faults but also operate under stringent resource constraints, while still ensuring efficient task execution. Achieving high-performance and energy-efficient computation with adaptive reliability in such harsh conditions is therefore of great importance. This work presents Frieren, a fault-tolerant and reconfigurable computing architecture for reliable operation in harsh environments. A 22 nm system-on-chip (SoC) prototype is implemented to validate Frieren and evaluate its resilience to soft errors. Frieren operates in three primary modes: (1) a high-throughput computation engine mode, (2) a multi-core mode featuring adaptive dual-core lockstep (DCLS) for fault tolerance and programmable parallel computing, and (3) a JTAG-assisted scan-chain-based fault injection (FI) mode. The first two modes fully share processing elements and memory resources, ensuring zero data movement during mode transitions, while the third mode supports pre-deployment reliability evaluation by emulating transient faults. Both irradiation and hardware-level FI experiments are conducted to verify reliability, confirming the robustness of Frieren. Radiation tests of the SoC indicate that DCLS can correct up to about 83% of RISC-V errors, while customized parallel computing in multi-core mode achieves a 17.77× latency reduction. Moreover, the SoC delivers up to 17.18 TOPS/W in computation engine mode and 1.92 TOPS/W in multi-core mode, demonstrating an energy-efficient and resilient platform for AI deployment under harsh conditions. In real workloads, the SoC achieves peak energy efficiencies of 14.72 TOPS/W on SuperYOLO and 12.33 TOPS/W on DROID-SLAM. Qiufeng Li, Weirong Dong, Mingqiang Huang, Hao Yu 0001, Yiyu Shi 0001, Hiromitsu Awano, Takashi Sato 0001, Mehdi Saligane, Longyang Lin, Masanori Hashimoto |
IEEE Trans. Computers | 2 |
| 2025 | Invited Paper: Multi-Agent Generative Synthesis for Analog/RF Circuit: from Scalable Topology Generation to Efficient Inverse DesignabstractThe exponential growth of information and computational workloads has created unprecedented demands for high-productivity development of computer hardware built on foundational semiconductor integrated circuits (ICs). Yet, the lack of effective design automation techniques makes developing analog ICs–indispensable in ubiquitous computer systems–a significant bottleneck for overall design productivity and cost efficiency across the IC ecosystem. Excitingly, recent advances in generative AI present transformative opportunities to tackle the complexity and large-scale challenges of modern analog/radio-frequency (RF) IC design. This work introduces a first-of-its-kind multi-agent generative synthesis framework for analog/RF circuits. Specifically, our approach formulates analog synthesis as a multi-stage generative AI problem: first generating a novel topology conditioned on high-level textual descriptions, and then producing high-quality device parameters to meet given design specifications for the generated circuit topology. This novel paradigm offers distinct advantages over traditional methods, including controllable novelty in topology generation and significantly improved inverse design efficiency. Beyond topology generation and inverse design, our method enables broader applications such as synthetic dataset generation and privacy-preserving data sharing–emerging challenges in data-driven electronic design automation (EDA) due to the computationally intensive and confidential nature of analog/RF circuit design. This work paves the way for next-generation generative AI-driven multi-agent synthesis in analog/RF EDA. Shikai Wang, Qiufeng Li, Houbo He, Taiyun Chi |
ICCAD | 2 |
| 2025 | A Scalable External Memory Access and On-Chip Storage Architecture for Edge-AI Accelerators : - Multi-Path Rolling Data Refresh and Layer-Wise Bank Allocation -abstractFor resource-constrained AI accelerators applied in edge computing, achieving high power efficiency in neural network (NN) model computation is crucial. However, current designs often overlook the efficiency of off-chip/on-chip data interaction, leading to high latency, which in turn results in suboptimal power efficiency during computation. Additionally, inefficient memory bank allocation further exacerbates latency by causing underutilization of storage resources, thereby contributing to higher overall latency and energy consumption. To address these challenges, this paper proposes a scalable multi-path rolling data refresh and layer-wise bank allocation architecture. The rolling data refresh mechanism enables efficient data interaction between off-chip and on-chip storage, reducing latency and minimizing the area overhead of on-chip memories. The layer-wise bank allocation optimizes on-chip memory utilization according to specific application requirements, improving memory efficiency. A case study on a 28nm AI accelerator demonstrates a 30.6% reduction in area, achieves a power efficiency of 7.36–10.28 TOPS/W, and reduces external memory access by 2.63% to 37.24% on VGG16 and ViT-Small. Huizi Zhang, Qiufeng Li, Yuan Liang 0004, Zhenzhe Chen, Jinjun Xiong, Mingqiang Huang, Longyang Lin, Masanori Hashimoto |
ISLPED | 3 |
| 2025 | Unicorn-CIM: Unconvering the Vulnerability and Improving the Resilience of High-Precision Compute-in-MemoryabstractCompute-in-memory (CIM) architecture has been widely explored to address the von Neumann bottleneck in accelerating deep neural networks (DNNs). However, its reliability remains largely understudied, particularly in the emerging domain of floating-point (FP) CIM, which is crucial for speeding up high-precision inference and on-device training. This paper introduces Unicorn-CIM, a framework to uncover the vulnerability and improve the resilience of high-precision CIM, built on static random-access memory (SRAM)-based FP CIM architecture. Through the development of fault injection and extensive characterizations across multiple DNNs, Unicorn-CIM reveals how soft errors manifest in FP operations and impact overall model performance. Specifically, we find that high-precision DNNs are extremely sensitive to errors in the exponent part of FP numbers. Building on this insight, Unicorn-CIM develops an efficient algorithm-hardware co-design method that optimizes model exponent distribution through fine-tuning and incorporates a lightweight Error Correcting Code (ECC) scheme to safeguard high-precision DNNs on FP CIM. Comprehensive experiments show that our approach introduces just an 8.98% minimal logic overhead on the exponent processing path while providing robust error protection and maintaining model accuracy. This work paves the way for developing more reliable and efficient CIM hardware. Qiufeng Li |
ISLPED | 1 |
| 2024 | How accurately can soft error impact be estimated in black-box/white-box cases? - a case study with an edge AI SoC -abstractArtificial intelligence (AI) edge devices often feature numerous storage units and sequential logic circuits, making them vulnerable to soft errors. For reliable and critical edge AI applications, assessing System-on-Chip (SoC) reliability in advance is essential. Here, there are two cases: a self-designed SoC (white-box), or a commercial off-the-shelf (COTS) chip (black-box). This study uses alpha particle irradiation results on our 22nm AI SoC as a golden reference to estimate soft error impacts, injecting faults across the entire chip in the white-box case and into the accessible memory and registers in the black-box case. The results demonstrate a high degree of consistency between the white-box case and golden reference, meaning that pre-silicon reliability assessment is feasible. As for the black-box case, the proportion of memory in the SoC remains unchanged and is still significantly larger than that of registers, and hence the simulation results between black-box and white-box are not substantially different. Qiufeng Li, Longyang Lin, Wang Liao 0001, Liuyao Dai, Hao Yu 0001, Masanori Hashimoto |
DAC | 2 |
| 2023 | Agile Hardware and Software Co-Design for RISC-V-Based Multi-Precision Deep Learning MicroprocessorabstractRecent network architecture search (NAS) has been widely applied to simplify deep learning neural networks, which typically result in a multi-precision network. Many multi-precision accelerators have been developed as well to support computing multi-precision networks manually. A software-hardware interface is thereby needed to automatically map multi-precision networks onto multi-precision accelerators. In this paper, we have developed an agile hardware and software co-design for RISC-V-based multi-precision deep learning microprocessor. We have designed custom RISC-V instructions with a framework to automatically compile multi-precision CNN networks onto multi-precision CNN accelerators, demonstrated on FPGA. Experiments show that with NAS optimized multi-precision CNN models (LeNet, VGG16, ResNet, MobileNet), the RISC-V core with multi-precision accelerators can reach the highest throughput in 2,4,8-bit precisions respectively on a Xilinx ZCU102 FPGA. Zicheng He, Qiufeng Li, Hao Yu 0001 |
ASP-DAC | 3 |
| 2022 | A High Throughput Multi-bit-width 3D Systolic Accelerator for NAS Optimized Deep Neural Networks on FPGAabstractNeural architecture search (NAS) optimized multi-bit-width convolutional neural network (CNN) maintains the balance between network performance and efficiency, thus enlightening a promising method for accurate yet energy-efficient edge computing. In this work, we propose a high throughput three-dimensional (3D) systolic accelerator for NAS optimized CNNs, in which the input feature matrix, weight matrix and output feature matrix are delivering vertically, horizontally and perpendicularly through the systolic array respectively. With 3D systolic data flow, the processing time and logic resources consumption can be both reduced compared to the classical non-stationary systolic array. Besides, Booth-based multi-bit-width (INT2/4/8) multiply-add-accumulation (MAC) unit is developed within the 3D systolic accelerator. Deployed on FPGA platform Xilinx ZCU102, peek performance of the convolutional layer can reach as high as 2775 GOPS for INT2, 1650 GOPS for INT4, and 816 GOPS for INT8 respectively. The average performance on accelerating full NAS VGG16 network is 647 GOPS. Mingqiang Huang, Yucen Liu, Shuxin Yang, Kai Li 0024, Junyi Luo, Zhengke Yang, Qiufeng Li, Hao Yu 0001, Changhai Man |
FPGA | 8 |
| 2020 | Acoustic emission source localization method for high-speed train bogie
Xincheng Wei, Lixia Huang, Qiufeng Li |
Multim. Tools Appl. | 6 |
| 2014 | Wavelet transform-based feature extraction for ultrasonic flaw signal classification
Peng Yang 0007, Qiufeng Li |
Neural Comput. Appl. | 2 |