VLDB 2026 Research / reviewers in the wild / expert
Bingyao Li 0001
dblp:231/3378-1
· DBLP profile ↗
9ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-6281-6799ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OASIS: Object-Aware Page Management for Multi-GPU SystemsabstractThe ever-growing need for high-performance computing has driven the popularity of employing multi-GPU systems. Modern multi-GPU systems employ unified virtual memory (UVM) to manage page placement and migration. However, the page management in UVM is application object agnostic. In this paper, we characterize the page access behaviors in relation to the application objects, and reveal that the beneficial page management policy varies according to (i) the different data objects within the same application, and (ii) the different execution phases of the same object. This motivates the need for dynamic and proactive page management in multi-GPU systems. To this end, we propose OASIS, which dynamically identifies object patterns during the execution and proactively determines the appropriate page management policies for these objects at runtime. Experimental results show that OASIS improves the performance over uniformly adopting on-touch migration, access counter-based migration, and duplication by an average of $\mathbf{6 4 \%}$, $\mathbf{3 5 \%}$, and $\mathbf{4 2 \%}$, respectively. Moreover, OASIS achieves a $\mathbf{1 2 \%}$ performance improvement over the state-of-the-art technique (i.e., GRIT) while having significantly lower design complexity. Bingyao Li 0001, M. Tarek Ibn Ziad, Lieven Eeckhout, Jun Yang 0002, Aamer Jaleel, Xulong Tang |
HPCA | 2 |
| 2024 | GRIT: Enhancing Multi-GPU Performance with Fine-Grained Dynamic Page PlacementabstractMulti-GPU systems have become popular to cater to the growing demands for high parallelism and large memory capacity. However, the delivered performance is constrained by the non-uniform memory access (NUMA) overhead arising from data sharing and communication across multiple GPUs. Recent multi-GPUs employ unified virtual memory (UVM) to simplify the programming effort. In UVM-enabled multi-GPUs, three popular page placement schemes are adopted to mitigate the NUMA overheads: i) on-touch page migration, ii) access counter-based migration, and iii) page duplication. However, we observe that the preferred page placement scheme varies across i) different applications, ii) different pages of the same application, and iii) even different execution phases of a single page, making it challenging to find a “one-size-fits-all” page placement scheme. To this end, we propose GRIT, which dynamically and automatically determines the appropriate page placement schemes at runtime in a fine-grained manner to enhance multi-GPU performance and scalability. Experimental results indicate that GRIT achieves an average of 60%, 49%, and 29% performance improvements over uniformly adopting on-touch migration, access counter-based migration, and page duplication, respectively. Bingyao Li 0001, Aamer Jaleel, Jun Yang 0002, Xulong Tang |
HPCA | 2 |
| 2024 | STAR: Sub-Entry Sharing-Aware TLB for Multi-Instance GPUabstractNVIDIA's Multi-Instance GPU (MIG) technology enables partitioning GPU computing power and memory into sep-arate hardware instances, providing complete isolation including compute resources, caches, and memory. However, prior work identifies that MIG does not partition the last-level TLB (i.e., L3 TLB), which remains shared among all instances. To enhance TLB reach, NVIDIA GPUs reorganized the TLB structure with 16 sub-entries in each L3 TLB entry that have a one-to-one mapping to the address translations for 16 pages of size 64 KB located within the same 1 MB aligned range. Our comprehensive investigation of address translation efficiency in MIG identifies two main issues caused by L3 TLB sharing interference: (i) it results in performance degradation for co-running applications, and (ii) TLB sub-entries are not fully utilized before eviction. Based on this observation, we propose STAR to improve the utilization of TLB sub-entries through dynamic sharing of TLB entries across multiple base addresses. STAR evaluates TLB entries based on their sub-entry utilization to optimize address translation storage, dynamically adjusting between a shared and non-shared state to cater to current demand. We show that STAR improves overall performance by an average of 28.7% across various multi-tenant workloads. Bingyao Li 0001, Lieven Eeckhout, Jun Yang 0002, Aamer Jaleel, Xulong Tang |
MICRO | 1 |
| 2023 | Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUsabstractUnified Virtual Memory (UVM) is a promising feature in CPU-GPU heterogeneous systems that allows data structures to be accessed by both CPU and GPUs through unified pointers without explicit data copying. However, the delivered performance of UVM significantly relies on the efficiency of address translation. The current GPU thread block (TB) management is not aware of the translation process and heavily thrashes the per-streaming multiprocessor (SM) private Translation Look-ahead Buffers (TLBs). In this paper, we conduct a comprehensive characterization of 10 GPU benchmarks and quantify the translation reuses among the thread blocks. Our observation reveals that there exists substantial translation reuse within TBs rather than across the TBs. Moreover, the inter-TB interference significantly enlarges the intra-TB translation reuse distances. To this end, we propose a translation-aware TB scheduling and lightweight GPU L1 TLB partitioning to effectively mitigate the contention. Experimental results show that our proposed approach improves the L1 TLB hit rate, and this improvement translates to, on average, a 12.5% execution time reduction. Bingyao Li 0001, Xulong Tang |
DAC | 1 |
| 2023 | Trans-FW: Short Circuiting Page Table Walk in Multi-GPU Systems via Remote ForwardingabstractMulti-GPU systems have become a popular platform to meet the ever-growing application demands. However, employing multiple GPUs does not guarantee proportional performance improvements. While prior works have extensively studied the optimizations to mitigate the non-uniform memory accesses (NUMA) overheads, the address translation process also plays an important role in shaping the overall execution performance. In this paper, we investigate the address translation process in multi-GPU systems under unified virtual memory (UVM). We specifically focus on the efficiency of page table walk and identify three major latency penalties: i) queuing for available page table walk threads, ii) memory accesses for page walk cache misses, and iii) handling page faults. Based on our observations, we propose Trans-FW, which short circuits the page table walk by leveraging substantial translation sharing and eager remote translation forwarding. Experimental results on 10 representative multi-GPU applications show that our proposed approach improves the overall performance by 53.8% on average. Bingyao Li 0001, Jieming Yin, Anup Holey, Youtao Zhang, Jun Yang 0002, Xulong Tang |
HPCA | 1 |
| 2023 | IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE InvalidationsabstractMulti-GPU systems have emerged as a desirable platform to deliver high computing capabilities and large memory capacity to accommodate large dataset sizes. However, naively employing multi-GPU incurs non-scalable performance. One major reason is that execution efficiency suffers expensive address translations in multi-GPU systems. The data-sharing nature of GPU applications requires page migration between GPUs to mitigate non-uniform memory access overheads. Unfortunately, frequent page migration incurs substantial page table invalidation overheads to ensure translation coherence. A comprehensive investigation of multi-GPU address translation efficiency identifies two significant bottlenecks caused by page table invalidation requests: (i) increased latency for demand TLB miss requests and (ii) increased waiting latency for performing page migrations. Based on observations, we propose IDYLL, which reduces the number of page table invalidations by maintaining an “in-PTE" directory and reduces invalidation latency by batching multiple invalidation requests to exploit spatial locality. We show that IDYLL improves overall performance by 69.9% on average. Bingyao Li 0001, Yanan Guo 0002, Aamer Jaleel, Jun Yang 0002, Xulong Tang |
MICRO | 1 |
| 2021 | Improving Address Translation in Multi-GPUs via Sharing and Spilling aware TLB DesignabstractIn recent years, the ever-growing application complexity and input dataset sizes have driven the popularity of multi-GPU systems as a desirable computing platform for many application domains. While employing multiple GPUs intuitively exposes substantial parallelism for the application acceleration, the delivered performance rarely scales with the number of GPUs. One of the major challenges behind is the address translation efficiency. Many prior works focus on CPUs or single GPU execution scenarios while the address translation in multi-GPU systems receives little attention. In this paper, we conduct a comprehensive investigation of the address translation efficiency in both “single-application-multi-GPU” and “multi-application-multi-GPU” execution paradigms. Based on our observations, we propose a new TLB hierarchy design, called least-TLB, tailored for multi-GPU systems and effectively improves the TLB performance with minimal hardware overheads. Experimental results on 9 single-application workloads and 10 multi-application workloads indicate the proposed least-TLB improves the performances, on average, by 23.5% and 16.3%, respectively. Bingyao Li 0001, Jieming Yin, Youtao Zhang, Xulong Tang |
MICRO | 1 |
| 2018 | An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 1 |
| 2018 | Correction to: An Efficient Retrieval Method for Astronomical Catalog Time Series Data
Bingyao Li 0001, Ce Yu, Xiaoteng Hu, Jian Xiao 0001, Shanjiang Tang, Lianmeng Li, Bin Ma 0022 |
ICA3PP (1) | 1 |