VLDB 2026 Research / reviewers in the wild / expert
Ying Li 0049
dblp:22/1805-49
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0005-2737-0583ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDPAT: Hierarchical Distributed Page Address Translation for Wafer-Scale GPUsabstractA Wafer-scale GPU connects a large number of chiplets via a high-bandwidth, low-latency interposer-based network, promising to overcome the communication bottleneck of traditional multi-GPU systems. While prior work has prototyped wafer-scale GPUs to demonstrate technical feasibility, scaling to massive chiplet counts creates new bottlenecks: virtual-tophysical address translation becomes severely constrained by massive concurrent requests and long multi-hop network latencies. We propose HDPAT, a hardware-accelerated distributed address translation system that addresses this challenge through three complementary techniques: (1) Concentric caching converts near-IOMMU chiplets into hierarchical translation caches based on their distance to the IOMMU. A lightweight rotation mechanism ensures that there is always a nearby chiplet that can provide translation caching. (2) The redirection table further reduces the burden of IOMMU by delegating translations to caching chiplets, and (3) Prefetching proactively delivers potentially needed address translation into the chiplet to improve translation cache hit rate. Experimental results on 14 representative workloads show that HDPAT improves overall performance by an average of$1.57 \times$. Daoxuan Xu, Ying Li 0049, Yuwei Sun, Jie Ren 0015, Yifan Sun 0002 |
HPCA | 2 |
| 2025 | TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU SystemsabstractDeep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation.The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability.With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy.Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference.However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies.While an alternative solution is to test on GPU simulators, they are often too slow for these large-scale systems and depend on profiling details collected from real distributed systems to initiate the simulation.To address these challenges, we present TrioSim, a novel lightweight simulator for DNNs on multi-GPU systems.TrioSim combines performance modeling techniques and simulation methods to achieve high flexibility, high simulation speed, and * Part of this work was done while Yuhui Bao and Pranav Vaid were interns at Lightmatter. Ying Li 0049, Yuhui Bao, Gongyu Wang, Xinxin Mei, Pranav Vaid, Anandaroop Ghosh, Adwait Jog, Darius Bunandar, Ajay Joshi, Yifan Sun 0002 |
ISCA | 1 |
| 2023 | A Regression-based Model for End-to-End Latency Prediction for DNN Execution on GPUsabstractDeep neural networks (DNNs) have become increasingly popular in many domains as they reduce the requirement for human effort. However, today’s DNN applications suffer from high computational complexity and sub-optimal device utilization. To solve this problem, researchers have been proposing new system design solutions, which require performance models to help them with pre-product concept validation. This paper discusses how to build a simple, yet accurate, performance model for DNNs on GPUs. Our observations demonstrate prevalent linear relationships between the GPU execution times and operation counts of DNNs layers. Our proposed linear-regression-based execution time predictor can make predictions with an error rate of 28%.11This material is based upon work supported in part by the Google Research Scholar Award and William & Mary. This work was performed in part using the computing facilities at William & Mary and Google Cloud. This work was done while Jog was with William & Mary. Jog is currently with the University of Virginia. Ying Li 0049, Yifan Sun 0002, Adwait Jog |
ISPASS | 1 |
| 2023 | Path Forward Beyond Simulators: Fast and Accurate GPU Execution Time Prediction for DNN WorkloadsabstractToday, DNNs’ high computational complexity and sub-optimal device utilization present a major roadblock to democratizing DNNs. To reduce the execution time and improve device utilization, researchers have been proposing new system design solutions, which require performance models (especially GPU models) to help them with pre-product concept validation. Currently, researchers have been utilizing simulators to predict execution time, which provides high flexibility and acceptable accuracy, but at the cost of a long simulation time. Simulators are becoming increasingly impractical to model today’s large-scale systems and DNNs, urging us to find alternative lightweight solutions. Ying Li 0049, Yifan Sun 0002, Adwait Jog |
MICRO | 1 |