VLDB 2026 Research / reviewers in the wild / expert
Yiquan Lin
dblp:139/1255
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0006-1087-6016ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I-POP: Ignite Positive PrefetchersabstractHardware prefetching is a well-established technique for bridging the processor-memory performance gap. To improve cache miss coverage, modern processors often integrate multiple prefetchers. However, multi-prefetcher systems without proper management often suffer from suboptimal performance due to a surge of useless prefetches. Several techniques have been proposed to select appropriate prefetchers for issuing requests, but they all face limitations. Specifically, existing (1) static schemes lack feedback regulation mechanisms and suffer from inflexible prefetcher selections; (2) reinforcement learning (RL)based schemes incur high overhead and suffer from adjustment lag; and (3) performance-counter-based schemes rely on inefficient runtime metrics that fail to accurately and clearly reflect a prefetcher's true impact on performance. In this paper, we propose I-POP, a high-performance and lowoverhead prefetcher management scheme for multi-prefetcher systems. I-POP introduces a novel runtime metric, Prefetch Effectiveness (PE), which aggregates each prefetch request's beneficial and harmful effects to precisely quantify the impact of a prefetcher on performance, effectively overcoming the limitations of prior metrics. To compute and leverage this metric, I-POP incorporates two key components: the Metric Collector, which periodically calculates each prefetcher's PE, and the Control Engine, which dynamically manages all prefetchers based on their PE values. Specifically, I-POP ignites (enables) prefetchers with positive PE values, adaptively tuning their aggressiveness, and disables those with non-positive PE. We evaluated I-POP on numerous workloads, and the results show I-POP outperforms two state-of-the-art approaches, Bandit and Alecto, by$\mathbf{4. 2 \%}$and 3.5 % across three benchmark suites in a single-core system, and 6.6 % and 8.6 % in a 16 -core system, while incurring only 1.46 KB of storage overhead. Yiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu, Shishun Cai, Jiarong Ye, Zonghui Wang, Wenzhi Chen |
HPCA | 1 |
| 2025 | OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageabstractObject storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources. Zhen Jin 0008, Yiquan Chen, Mingxu Liang, Guoju Fang, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi, Zhenhua He, Shishun Cai, Wenzhi Chen |
ASPLOS (2) | 10 |
| 2024 | CINDA: Don't Ignore Instructions When Cloning Memory Access BehaviorabstractExisting workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error. Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen |
CCGrid | 9 |
| 2024 | Performance Characterization of SmartNIC NVMe-over-Fabrics Target OffloadingabstractThe NVMe-over-Fabrics (NVMe-oF) is gaining popularity in cloud data centers as a remote storage protocol for accessing NVMe storage devices across servers. With the rapid increase in throughput of the NVMe storage devices, the NVMe-oF stack consumes a significant amount of valuable CPU resources. To release these valuable computing power to other tasks, many smartNICs now support NVMe-oF Target offloading. However, this emerging NVMe-oF offloading scheme's performance has not been fully investigated. Jiexiong Xu, Yiquan Chen, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenzhi Chen |
SYSTOR | 6 |
| 2024 | PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region SizesabstractHardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%). Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | External knowledge document retrieval strategy based on intention-guided and meta-learning for task-oriented dialogues
Hongtu Xie, Yiquan Lin, Guoqian Wang |
Adv. Eng. Informatics | 3 |
| 2022 | Work-in-Progress: Lark: A Learned Secondary Index Toward LSM-tree for Resource-Constrained Embedded Storage SystemsabstractLSM-tree-based key-value stores are popular in embedded storage systems. With the growing demands of data analysis, the secondary index is created to support non-primary-key lookups. However, the lookup efficiency and space consumption of secondary index remain for further optimization. Inspired by the learned index, this paper presents Lark, a learned secondary index toward LSM-tree for resource-constrained embedded storage systems. Lark employs machine learning to speed up the non-primary-key queries and compress secondary indexes. Our preliminary evaluations show that, in comparison with traditional secondary index schemes, Lark achieves better lookup performance with less space consumption. Jianan Yuan, Shangyu Wu, Yiquan Lin, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
CODES+ISSS | 4 |
| 2022 | GCIM: Toward Efficient Processing of Graph Convolutional Networks in 3D-Stacked MemoryabstractGraph convolutional networks (GCNs) have become a powerful deep learning approach for graph-structured data. Different from traditional neural networks such as convolutional neural networks, GCNs handle irregular input graph data, and GCNs are both computation-bound and memory-bound. How to efficiently utilize the underlying computation and memory resource becomes a critical issue. The emerging 3D-stacked computation-in-memory (CIM) architecture can reduce the data movement between computing logic and memory, thereby presenting a promising solution for the processing of GCNs. An unsolved key challenge is how to allocate GCNs to take advantage of fast near-data processing of the 3D-stacked CIM architecture. This article presents GCIM, a software–hardware co-design approach to exploit the efficient processing of GCNs on the CIM architecture. At the level of hardware design, GCIM integrates lightweight computing units near memory banks to fully exploit bank-level bandwidth and parallelism. At the level of software design, a locality-aware data mapping algorithm is proposed to partition the input graph and achieve workload balancing. GCIM is evaluated through a set of representative GCN models and standard graph datasets. The experimental results show that GCIM can significantly reduce the processing latency and data movement overhead compared with representative schemes. Jiaxian Chen, Yiquan Lin, Kaoyi Sun, Jiexin Chen, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Imaging Experiment of Airborne UHF Ultra-wideband Synthetic Aperture RadarabstractThe ultrahigh frequency ultra-wideband synthetic aperture radar (UHF UWB SAR) has the well foliage penetrating and high-resolution imaging, which can be used to detect the concealed area under the foliage in forests. This paper presents an airborne UHF UWB SAR experiment and imaging results. During the winter, an airborne campaign has been carried out in Shanxi Province in China, and the raw data was collected. In this experiment, the SAR system was integrated onboard a CESSNA-172 airplane. The antenna was fixed on the suspension arm of the right wing of the airplane, while the other part of the SAR system was placed on the back seat of this airplane. The experimental results have been obtained from the collected raw data, which proved the imaging performance of the airborne UHF UWB SAR system as well as the validity of the imaging method. Hongtu Xie, Guoqian Wang, Jun Hu 0003, Keqing Duan, Zengping Chen, Shiyou Xu, Yiquan Lin, Nannan Zhu, Bin Xi, Daoxiang An |
IGARSS | 7 |