Taolei Wang

dblp:304/5252 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 OpScope: Exploiting Operation-Driven Visual Scope for QoS-Stable Cloud Gaming
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
IWQoS1
2025 Accelerating Large-Scale Out-of-GPU-Core GNN Training with Two-Level Historical Caching
Jing Wang 0055, Taolei Wang, Juntao Huang, Xinkai Wang 0003, Marius Kreutzer, Chao Li 0009, Minyi Guo
APPT2
2025 CGO: Cloud Game Orchestration via Resource Preception and CODEC Optimization
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
NPC (2)1
2024 CoCG: Fine-grained Cloud Game Co-location on Heterogeneous Platform
abstract
Cloud games have received widespread attention and exponential growth recently as a key technology for building metaverse. Unlike general tasks in the cloud, the scene-complex, latency-critical, and interaction-intensive features make it challenging for cloud game co-deployment on heterogeneous platforms. Game-grained resource allocation leads to low resource effectiveness. Although previous work tries to explore individual game partitioning methods, they still face the problem of inefficient game hosting decisions and ultimately QoS violations. In this paper, we propose a fine-grained game characteristic and scheduling strategy to co-locate games together for high resource usage effectiveness. First, we fully explore the relationships between game scenes and resource usage behaviors by breaking the cloud game into stages with multiple frames and clustering them. We adopt machine learning methods to predict game resource consumption in real-time. To further improve multi-game parallelism, we co-locate games in a complementary way and steal time from the loading stage to avoid oversubscribing. The evaluation shows that our work increased the throughput of the cloud game deployments by 23.7% with low overhead compared to previous work.
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
IPDPS1
2023 Fargraph+: Excavating the parallelism of graph processing workload on RDMA-based far memory system
Jing Wang 0055, Chao Li 0009, Taolei Wang, Junyi Mei, Lu Zhang 0049, Pengyu Wang 0003, Minyi Guo
J. Parallel Distributed Comput.4
2023 Optimizing GPU-Based Graph Sampling and Random Walk for Efficiency and Scalability
abstract
Graph sampling and random walk algorithms are playing increasingly important roles today because they can significantly reduce graph size while preserving structural information, thus enabling computationally intensive tasks on large-scale graphs. Current frameworks designed for graph sampling and random walk tasks are generally not efficient in terms of memory requirement and throughput. Not to mention that some of them result in biased results. To solve the above problems, we introduce Skywalker+, a high-performance graph sampling and random walk framework on multiple GPUs supporting multiple algorithms. Skywalker+ makes four key contributions: First, it realizes highly paralleled alias method on GPUs. Second, it applies finely adjusted workload-balancing techniques and locality-aware execution modes to present a highly efficient execution engine. Third, it optimizes the GPU memory usage with efficient buffering and data compression schemes. Last, it scales to multi-GPU to further enhance the system throughput. Abundant experiments show that Skywalker+ exhibits significant advantage over the baselines both in performance and utility.
Pengyu Wang 0003, Chao Li 0009, Jing Wang 0055, Taolei Wang, Lu Zhang 0049, Xiaofeng Hou, Minyi Guo
IEEE Trans. Computers5
2023 DRAGON: Dynamic Recurrent Accelerator for Graph Online Convolution
abstract
Despite the extraordinary applicative potentiality that dynamic graph inference may entail, its practical-physical implementation has been a topic seldom explored in literature. Although graph inference through neural networks has received plenty of algorithmic innovation, its transfer to the physical world has not found similar development. This is understandable since the most preeminent Euclidean acceleration techniques from CNN have little implication in the non-Euclidean nature of relational graphs. Instead of coping with the challenges arising from forcing naturally sparse structures into more inflexible stochastic arrangements, in DRAGON, we embrace this characteristic in order to promote acceleration. Inspired by high-performance computing approaches like Parallel Multi-moth Flame Optimization for Link Prediction (PMFO-LP), we propose and implement a novel efficient architecture, capable of producing similar speed-up and performance than baseline but at a fraction of its hardware requirements and power consumption. We leverage the hidden parallelistic capacity of our previously developed static graph convolutional processor ACE-GCN and expanded it with RNN structures, allowing the deployment of a multi-processing network referenced around a common pool of proximity-based centroids. Experimental results demonstrate outstanding acceleration. In comparison with the fastest CPU-based software implementation available in the literature, DRAGON has achieved roughly 191× speed-up. Under the largest configuration and dataset, DRAGON was also able to overtake a more power-hungry PMFO-LP by almost 1.59× in speed, and at around 89.59% in power efficiency. More importantly than raw acceleration, we demonstrate the unique functional qualities of our approach as a flexible and fault-tolerant solution that makes it an interesting alternative for an anthology of applicative scenarios.
José Romero Hung, Chao Li 0009, Taolei Wang, Jinyang Guo 0001, Pengyu Wang 0003, Chuanming Shao, Jing Wang 0055, Guoyong Shi, Xiangwen Liu
ACM Trans. Design Autom. Electr. Syst.3
2022 HyFarM: Task Orchestration on Hybrid Far Memory for High Performance Per Bit
abstract
Tapping into secondary memory resources, i.e., far memory (FM), has shown huge potential to improve the cost-efficiency of data centers. Recent advances in both storage-based vertical FM and network-based horizontal FM have raised new questions about leveraging hybrid FM tiers to achieve the best performance per bit of memory. It is still unclear how to efficiently place tasks when far memory access is enabled.In this work, we propose HyFarM, a novel task management strategy for hybrid FM clusters. We analyze FM sensitivity and cooperatively co-locate tasks to enable high utilization and scalability. Further, by tapping into dynamic memory adaption within and across servers, our strategy allows one to consistently deliver high performance on memory-intensive tasks. We evaluate our design with a heavily instrumented testbench. Compared with the state-of-the-art designs, HyFarM respectively improves memory utilization and the overall performance per bit (PPB) by up to 17.6% and 20.5%, with minor overhead.
Jing Wang 0055, Chao Li 0009, Junyi Mei, Taolei Wang, Pengyu Wang 0003, Lu Zhang 0049, Minyi Guo, Dongbai Chen, Xiangwen Liu
ICCD5
2022 Excavating the Potential of Graph Workload on RDMA-based Far Memory Architecture
abstract
Disaggregated architecture brings new opportunities to memory -consuming applications like graph processing. It allows one to outspread memory access pressure from local to far memory, providing an attractive alternative to disk-based processing. Although existing works on general-purpose far mem-ory platforms show great potentials for application expansion, it is unclear how graph processing applications could benefit from disaggregated architecture, and how different optimization methods influence the overall performance. In this paper, we take the first step to analyze the impact of graph processing workload on disaggregated architecture by extending the GridGraph framework on top of the RDMA-based far memory system. We design Fargraph, a far memory coordi-nation strategy for enhancing graph processing workload. Specif-ically, Fargraph reduces the overall data movement through a well-crafted, graph-aware data segment offloading mechanism. In addition, we use optimal data segment splitting and asynchronous data buffering to achieve graph iteration-friendly far memory access. We show that Fargraph achieves near-oracle performance for typical in-local-memory graph processing systems. Fargraph shows up to 8.3 x speedup compared to Fastswap, the state-of-the-art, general-purpose far memory platform.
Jing Wang 0055, Chao Li 0009, Taolei Wang, Lu Zhang 0049, Pengyu Wang 0003, Junyi Mei, Minyi Guo
IPDPS3
2021 Skywalker: Efficient Alias-Method-Based Graph Sampling and Random Walk on GPUs
abstract
Graph sampling and random walk operations, capturing the structural properties of graphs, are playing an important role today as we cannot directly adopt computing-intensive algorithms on large-scale graphs. Existing system frameworks for these tasks are not only spatially and temporally inefficient, but many also lead to biased results. This paper presents Skywalker, a high-throughput, quality-preserving random walk and sampling framework based on GPUs. Skywalker makes three key contributions: first, it takes the first step to realize efficient biased sampling with the alias method on a GPU. Second, it introduces well-crafted load-balancing techniques to effectively utilize the massive parallelism of GPUs. Third, it accelerates alias table construction and reduce the GPU memory requirement with efficient memory management scheme. We show that Skywalker greatly outperforms the state-of-the-art CPU-based and GPU-based baselines, in a wide spectrum of workload scenarios.
Pengyu Wang 0003, Chao Li 0009, Jing Wang 0055, Taolei Wang, Lu Zhang 0049, Jingwen Leng, Quan Chen 0002, Minyi Guo
PACT4