Yifeng Song

dblp:94/11182 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 A High-Speed FPGA Implementation for IVF-PQ Index Construction
abstract
The Inverted File with Product Quantization (IVF-PQ) is a widely used method for Approximate Nearest Neighbor Search (ANNS), playing a critical role in AI-driven applications such as search engines, recommendation systems, and advertising platforms. With the advent of Large Language Models (LLMs), the demand for efficient and real-time index construction has significantly increased, especially for edge-side personal applications. In this paper, we propose a scalable and high-speed FPGA implementation of IVF-PQ index construction, significantly reducing indexing latency and making it feasible for edge scenarios. First, we optimize the original index construction algorithm by introducing batch-mode centroid updates and replacing floating-point division with hardware-efficient operations, while maintaining competitive recall performance (with less than 5% degradation and up to 12.5% improvement compared to the original algorithm). Next, based on the modified algorithm, we design a flexible and scalable hardware architecture that supports two distance metrics (L2 and Inner Product), six PQ configurations, and input data with up to 1024 dimensions, all without necessitating hardware recompilation. Our implementation maximizes computational efficiency through finely tuned parallelism and dataflow, ensuring full pipeline utilization. Finally, we implement our design in Verilog and evaluate it on the Xilinx XCU280-FSVH2892-2L-E FPGA platform. Experimental results show that our accelerator achieves up to$30\times $speedup over a high-end server CPU (Intel Xeon Gold 6248R), reducing the indexing time from hours to minutes.
Yifeng Song, Yuan Du, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 An Efficient FPGA Implementation of Approximate Nearest Neighbor Search
abstract
Approximate nearest neighbor search (ANNS) plays an important role in modern artificial intelligence (AI) systems, being extensively utilized in search engines, advertising, and recommendation systems. With the advent of large language models (LLMs), ANNS is increasingly finding applications in edge scenarios such as personal assistants. The demand for efficient and fast ANNS solutions is, therefore, more pressing than ever. In this article, we propose a scalable and efficient field-programmable gate array (FPGA) implementation of ANNS based on the inverted file with product quantization (IVF-PQ) algorithm, thus marking the first hardware implementation supporting up to 1024-D datasets. First, we devise a novel architecture for the Top-Kmodule, capable of processing multiple input data streams simultaneously and linearly increasing throughput. Second, we adjust the data precision in several parts of our design, thus achieving obvious performance improvement without losing much recall. Moreover, we introduce a flexible distance calculation (Distance Cal) module that can be reused for various computational tasks at different query stages. We code our design in Verilog and implement it on Xilinx Alveo U280. The experimental results show that our search latency can be as low as 0.0071 ms at a 94% recall, while the power is 19.80 W. Compared to the state-of-the-art application-specified integrated circuit (ASIC) implementations, our design delivers a$4.5\times $speedup in latency and a 20% reduction in energy consumption.
Yifeng Song, Chenjie Liu, Danyang Zhu, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2023 A High-Speed FPGA-Based Hardware Implementation for Leighton-Micali Signature
abstract
Due to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many digital signature schemes that have resistance to quantum computing are studied and standardized by several influential international organizations. The Leighton-Micali signature (LMS) protocol, one of the hash-based signature schemes, is standardized by both the Internet Engineering Task Force (IETF) and the National Institute of Standards and Technology (NIST) due to its well-studied security and relatively small signature size. However, the heavy computation load and high latency of LMS limits its practical applications. In this paper, for the first time, we propose a full hardware implementation of LMS to accelerate all the three stages:$key~generation$,$signature~generation$, and$verification$. Considering the scalability requirement and the characteristic of the parameter sets of LMS, we extract the coarse-grained basic logic, a hash group, and build a reconfigurable architecture for all available parameters by carefully designing the parallelism degree while achieving low latency and high hardware utilization efficiency. Then, we devise a fusion architecture for$key~generation$and$signature~generation$based on the hash group module. Moreover, for the$signature~verification$stage, we propose a separate architecture by applying the hash group module along with an efficient depth-first Merkle tree module. We code our designs with Verilog language in parameterized style and implement them on a Xilinx XCVU7P FPGA platform. The experimental results show that significant improvements are obtained for different parameter sets by the proposed designs when compared to state-of-the-art works.
Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 High-Speed and Scalable FPGA Implementation of the Key Generation for the Leighton-Micali Signature Protocol
abstract
Due to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many new digital signature schemes that have resistance to quantum computing are being presented for Post-Quantum Cryptography (PQC) standardization. The Leighton-Micali signature (LMS), a kind of hash-based signature scheme, is selected as a promising candidate for the PQC signature protocols by the Internet Engineering Task Force (IETF) because of its small private and public key sizes. However, the low-efficiency in key generation forms the bottleneck in practical applications. In this paper, we propose a high-speed architecture for the key generation to accelerate the LMS for the first time. The architecture is delicately devised to be scalable, supporting all the parameter sets for the LMS. The degree of parallelism is carefully designed to achieve low latency and high hardware utilization efficiency. Moreover, the control flow is well managed to accommodate different parameter sets with constant power for the consideration of anti-power analysis attacks. We code our design with Verilog language and implement it on the Xilinx Zynq UltraScale+ FPGA. The experimental results show that, compared with the optimal software implementation running on an Intel(R) Core(TM) i7-6850K 3.60GHz CPU with threading enabled, the new design achieves 55x to 2091x speedup in different parameter configurations.
Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001
ISCAS1
2020 GPR-based Subsurface Object Detection and Reconstruction Using Random Motion and DepthNet
abstract
Ground Penetrating Radar (GPR) is one of the most important non-destructive evaluation (NDE) devices to detect the subsurface objects (i.e. rebars, utility pipes) and reveal the underground scene. One of the biggest challenges in GPR based inspection is the subsurface targets reconstruction. In order to address this issue, this paper presents a 3D GPR migration and dielectric prediction system to detect and reconstruct underground targets. This system is composed of three modules: 1) visual inertial fusion (VIF) module to generate the pose information of GPR device, 2) deep neural network module (i.e., DepthNet) which detects B-scan of GPR image, extracts hyperbola features to remove the noise in B-scan data and predicts dielectric to determine the depth of the objects, 3) 3D GPR migration module which synchronizes the pose information with GPR scan data processed by DepthNet to reconstruct and visualize the 3D underground targets. Our proposed DepthNet processes the GPR data by removing the noise in B-scan image as well as predicting depth of subsurface objects. For DepthNet model training and testing, we collect the real GPR data in the concrete test pit at Geophysical Survey System Inc. (GSSI) and create the synthetic GPR data by using gprMax3.0 simulator. The dataset we create includes 350 labeled GPR images. The DepthNet achieves an average accuracy of 92.64% for B-scan feature detection and an 0.112 average error for underground target depth prediction. In addition, the experimental results verify that our proposed method improve the migration accuracy and performance in generating 3D GPR image compared with the traditional migration methods.
Jinglun Feng, Haiyan Wang 0019, Yifeng Song, Jizhong Xiao
ICRA4