Ruiqi Tang

dblp:303/3585 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STP: Semantic-Triggered Prefetching for Event-Driven Workloads
abstract
Hardware prefetchers such as SPP rely on address-delta history and perform poorly on event-driven workloads, where event-type switches invalidate recent patterns. On an HFT order-book engine, SPP achieves only 8.0% L2 prefetch accuracy with 85.9% late prefetches; on B+Tree under uniform-random access, it degrades IPC by 6.7–8.1%. We present STP (Semantic-Triggered Prefetcher), which exposes event type through one non-privileged x86 hint instruction, SETHINT imm8, and combines two lightweight mechanisms: an Event Footprint Table (EFT) that replays high-frequency missed lines at type transitions, and a PC-Localized Working Set (PLWS) that gates next-line prefetching by per-PC miss rate with asymmetric feedback throttling. In gem5 on one HFT and four B+Tree settings, STP achieves up to 10 × higher L2 prefetch accuracy with 3–10 × fewer requests (up to 90% less bandwidth). On HFT, STP improves IPC by 36.1% over NoPF and by 5.4% over SPP (0.543 vs. 0.515). Under low locality, STP limits impact to − 2.0% or +0.8%, where SPP drops 6.7–8.1%. Hardware cost is 5.4 KB, reducible to ~3.1 KB.
Shichen Peng, Yupeng Gui, Han He, Zhengyang Cao, Ruiqi Tang, Xuanpeng Zhu, Xiaoyang Zeng, Yibo Fan
ACM Great Lakes Symposium on VLSI6
2026 A Multiplier-Free Similarity Metric for Token Merging in Vision Transformers
Yupeng Gui, Ruiqi Tang, Shichen Peng, Han He, Leilei Huang, Yibo Fan
ISCAS2
2026 SGFormer-RGCN: Accurate Performance Prediction for High-Level Synthesis Design Space Exploration
Ruiqi Tang, Chenglong Xiao
ISCAS1
2025 Fusion of global and adaptive local information for few-shot image classification
Ting Xiao 0002, Yiqing Xia, Ruiqi Tang, Wenli Du, Zhe Wang 0002
Pattern Recognit.3
2023 Liberator: A Data Reuse Framework for Out-of-Memory Graph Computing on GPUs
abstract
Graph analytics are widely used including recommender systems, scientific computing, and data mining. Meanwhile, GPU has become the major accelerator for such applications. However, the graph size increases rapidly and often exceeds the GPU memory, incurring severe performance degradation due to frequent data transfers between the main memory and GPUs. To relieve this problem, we focus on the utilization of data in GPUs by taking advantage of the data reuse across iterations. In our studies, we deeply analyze the memory access patterns of graph applications at different granularities. We have found that the memory footprint is accessed with a roughly sequential scan without a hotspot, which infers an extremely long reuse distance. Based on our observation, we propose a novel framework, calledLiberator, to exploit the data reuse within GPU memory. InLiberator, GPU memory is reserved for the data potentially accessed across iterations to avoid excessive data transfer between the main memory and GPUs. For the data not existing in GPU memory, a Merged and Aligned memory access manner is employed to improve the transmission efficiency. We also further optimize the framework by parallel processing of data in GPU memory and data in the main memory. We have implemented a prototype of theLiberatorframework and conducted a series of experiments on performance evaluation. The experimental results show thatLiberatorcan significantly reduce the data transfer overhead, which achieves an average of 2.7x speedup over a state-of-the-art approach.
Ruiqi Tang, Xiaoli Gong, Wenwen Wang 0001, Jin Zhang 0003, Pen-Chung Yew
IEEE Trans. Parallel Distributed Syst.2
2021 Ascetic: Enhancing Cross-Iterations Data Efficiency in Out-of-Memory Graph Processing on GPUs
abstract
Graph analytics are widely used in real-world applications, and GPUs are major accelerators for such applications. However, as graph sizes become significantly larger than the capacity of GPU memory, the performance can degrade significantly due to the heavy overhead required in moving a large amount of graph data between CPU main memory and GPU memory.
Ruiqi Tang, Kailun Wang, Xiaoli Gong, Jin Zhang 0003, Wenwen Wang 0001, Pen-Chung Yew
ICPP1