VLDB 2026 Research / reviewers in the wild / expert
Wei Yuan 0006
dblp:67/2268-6
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-9357-5716ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Top-k Sorter through Runtime Lane Selection of Priority Queues on a Directed Ring
Huawen Liang, Wei Yuan 0006, Qizhe Wu, Xi Jin 0002 |
ISCAS | 2 |
| 2025 | Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsabstractGeneral matrix-matrix multiplication (GEMM), serving as a cornerstone of AI computations, has positioned tensor processing engines (TPEs) as increasingly critical components within existing GPUs and domain-specific architectures (DSA). Our analysis identifies that the prevailing architectures primarily focus on dataflow or operand reuse strategies, when considering the combination of matrix multiplication with multiply-accumulator (MAC) itself, it provides greater optimization space for the design of TPEs. This work introduces a novel perspective on matrix multiplication from a hardware standpoint, focusing on the bit-weight dimension of MACs. Through this lens, we propose a finer-grained TPE notation, using matrix triple loops as an example, introducing new methods and ideas for designing and optimizing PE microarchitecture. Based on the new notation and transformations, we propose four optimization techniques that achieve varying degrees of improvement in timing, area, and power consumption. We implement our design in RTL using the SMIC-28nm process. Applying our methods to four classic TPE architectures (include systolic array [20], 3D-Cube [27], multiplier-adder tree [48], and 2D-Matrix [30]), we achieved area efficiency improvements of $1.27 \times, 1.28 \times, 1.56 \times$, and $1.44 \times$, and $1.04 \times, 1.56 \times, 1.49 \times$, and $1.20 \times$ for energy efficiency respectively. When applied to a bit-slice architecture, we achieved a $12.10 \times$ improvement in energy efficiency and $2.85 \times$ in area efficiency compared to Laconic [38]. Our Verilog HDL code, along with timing, area, and power reports for circuit synthesis in URL: https://github.com/wqzustc/High-Performance-Tensor-Processing-Engines. Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 0002, Zerong He, Linfeng Tao, Letian Zhao, Zhaoxi Zeng, Wei Yuan 0006, Xi Jin 0002 |
HPCA | 10 |
| 2025 | SageSC: Accelerating GraphSAGE Minibatch Inference on Memory-Intensive GraphsabstractGraph neural networks demonstrate excellent performance on node classification tasks in graph datasets. For inference tasks on memory-intensive graphs, the storage burden, memory access bottlenecks, and load imbalance issues arise. The minibatch inference proposed in GraphSAGE is an effective method for minimizing these problems. However, minibatch inference introduces new challenges: while it facilitates subsequent computation, the irregular random memory access pressure shifts to the minibatch construction phase, creating performance bottlenecks in the system. In this work, to address the aforementioned challenges, we propose a novel scattered minibatch construction and aggregation (SMCA) algorithm to optimize sampling, batch construction, and aggregation computations for minimizing their latency. This method distributes memoryintensive workloads and exploits the parallelism between memory groups. Evaluation results show that the proposed accelerator SageSC achieves speedups ranging from 3x to 96x compared to CPU/GPU baselines, especially on memory-intensive graphs, while outperforming existing state-of-the-art designs. Yuchen Gui, Wei Yuan 0006, Qizhe Wu, Huawen Liang, Letian Zhao, Linfeng Tao, Zhongguang Xu, Xi Jin 0002 |
ICCD | 2 |
| 2025 | MHE-TPE: Multi-Operand High-Radix Encoder for Mixed-Precision Fixed-Point Tensor Processing Engines
Qizhe Wu, Jinyi Zhou, Zhanhe Hu, Zhichen Zeng 0002, Huawen Liang, Jiuru Zhu, Linfeng Tao, Xin Zhang 0176, Zekang Cheng, Letian Zhao, Wei Yuan 0006, Xi Jin 0002 |
MICRO | 11 |
| 2025 | GHVSA: Graph-based high-dimensional vector search accelerator
Wei Yuan 0006, Huawen Liang, Xi Jin 0002 |
J. Syst. Archit. | 1 |
| 2025 | FANNS: An FPGA-Based Approximate Nearest-Neighbor Search AcceleratorabstractApproximate nearest-neighbor search (ANNS) based on high-dimensional vectors has been extensively utilized in data science and neural networks. However, deploying ANNS in production systems requires minimal redundant computation, high recall rates, and low on-chip memory usage, which existing hardware accelerators fail to offer. We propose FANNS, a solution for ANNS based on high-dimensional vectors that can eliminate redundant computations and reuse on-chip data. Extensive evaluations show that FANNS achieves an average of$184.1\times $,$33.0\times $,$2.9\times $, and$2.5\times $better energy efficiency than CPUs, GPUs, and two state-of-the-art ANNS architectures, i.e., DF-GAS and Vstore, respectively. Wei Yuan 0006, Xi Jin 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | A FPGA-HBM-Based Hardware Streaming Accelerator for GNN SamplingabstractSampling takes a long time during GNN training and inference, especially in large-scale graph datasets, so accelerating the sampling process is of great value. Present GNN samplers are mainly focused on the CPU or GPU side, with fewer sampling accelerator implementations on hardware platforms such as FPGAs, and most of the existing accelerators focus on the entire GNN computation process rather than on sampling. Sampling GNNs on the FPGA side is difficult because of the intensive random memory access. In this work, we propose a FPGA-HBM-based streaming sampler working in the time domain to accelerate traditional FPGA sampling, which realizes GNN sampling while reading data in bus bursts and can satisfy both playback and non-playback sampling modes. In addition, the fast loading of node feature vectors is also realized by combining the features of HBM. The proposed hardware achieves a speedup of$2\times$to$20\times$relative to traditional FPGA-based node index sampling and$5\times$to$18\times$relative to feature vector loading at the CPU side. Yuchen Gui, Qizhe Wu, Wei Yuan 0006, Huawen Liang, Xi Jin 0002 |
ASAP | 3 |
| 2024 | RingTK: A Ring, Parallel and High Performance Top-K Sorter on FPGAabstractGetting the K largest/smallest elements from$N$inputs is one of the essential operations in many applications. In this article, we propose RingTK, a ring, parallel, and high performance Top-K sorter implemented on FPGA. We use a priority queue as the basic processing unit and design a Top-K sorter based on a ring topology with a global maximum module(GMM) to obtain good scalability and high performance. Based on the ring topology, we design Ring Multiplexers (RMUX) and modify the GMM to enable RingTK to efficiently handle different K-sizes and concurrent tasks. We design a encoder to flexibly deal with different data formats and max/min Top-K tasks. Finally, we implement the proposed architecture on the Xilinx XCVU37P FPGA. The results show that the proposed architecture has good scalability, and the throughput and parallelism are approximately linear. We can achieve a throughput of 35.38GB/s with 32 PQs that exceed existing literature with K=8160. Huawen Liang, Qizhe Wu, Wei Yuan 0006, Teng Tian, Xi Jin 0002 |
FCCM | 3 |
| 2022 | FP-GNN: Adaptive FPGA accelerator for Graph Neural Networks
Teng Tian, Letian Zhao, Qizhe Wu, Wei Yuan 0006, Xi Jin 0002 |
Future Gener. Comput. Syst. | 5 |
| 2022 | QEGCN: An FPGA-based accelerator for quantized GCNs with edge-level parallelism
Wei Yuan 0006, Teng Tian, Qizhe Wu, Xi Jin 0002 |
J. Syst. Archit. | 1 |
| 2021 | A Gather Accelerator for GNNs on FPGA PlatformabstractGraph Neural Networks (GNNs) have emerged as the state-of-the-art deep learning model for representation learning on graphs. GNNs mainly include two phases with different execution patterns. The Gather phase, depends on the structure of the graph, presenting a sparse and irregular execution pattern. The Apply phase, acts like other neural networks, showing a dense and regular execution pattern. It is challenging to accelerate GNNs, due to irregular data communication to gather information within the graph. To address this challenge, hardware acceleration for Gather phase is critical. The purpose of this research is to design and implement an FPGA-based accelerator for Gather phase. It achieves excellent performance on acceleration and energy efficiency. Evaluation is performed using a Xilinx VCU128 FPGA with three commonly-used datasets. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA P100 GPU, our work achieves on average 101.28× speedup with 75.27× dynamic energy reduction and average 12.27× speedup with 45.56× dynamic energy reduction, respectively. Wei Yuan 0006, Teng Tian, Huawen Liang, Xi Jin 0002 |
ICPADS | 1 |