Dalin Wang

dblp:270/3779 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2023 RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding Columns
abstract
Embedding columns are important for deep recommendation models to achieve high accuracy, but they can be very time-consuming during inference. Machine learning (ML) compilers are used broadly in real businesses to optimize ML models automatically. Unfortunately, no existing work uses compilers to automatically accelerate the heavy embedding column computations during recommendation model inferences. To fill this gap, we propose RECom, the first ML compiler that aims at optimizing the massive embedding columns in recommendation models on the GPU. RECom addresses three major challenges. First, generating an efficient schedule on the GPU for the massive operators within embedding columns is difficult. Existing solutions usually lead to numerous small kernels and also lack inter-subgraph parallelism. We adopt a novel codegen strategy that fuses massive embedding columns into a single kernel and maps each column into a separate thread block on the GPU. Second, the complex shape computations under dynamic shape scenarios impede further graph optimizations. We develop a symbolic expression-based module to reconstruct all shape computations. Third, ML frameworks inevitably introduce redundant computations due to robustness considerations. We develop a subgraph optimization module that performs graph-level simplifications based on the entire embedding column context. Experiments on both in-house and open-source models show that RECom can achieve 6.61X and 1.91X over state-of-the-art baselines in terms of end-to-end inference latency and throughput, respectively. RECom's source code is publicly available at https://github.com/AlibabaResearch/recom.
Zaifeng Pan, Zhen Zheng, Feng Zhang 0007, Hao Liang 0003, Dalin Wang, Xiafei Qiu, Wei Lin 0016, Xiaoyong Du 0001
ASPLOS (4)6
2023 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach
abstract
Compiler optimization plays an increasingly important role to boost the performance of machine learning models for data processing and management. With increasingly complex data, the dynamic tensor shape phenomenon emerges for ML models. However, existing ML compilers either can only handle static shape models or expose a series of performance problems for both operator fusion optimization and code generation in dynamic shape scenes. This paper tackles the main challenges of dynamic shape optimization: the fusion optimization without shape value, and code generation supporting arbitrary shapes. To tackle the fundamental challenge of the absence of shape values, it systematically abstracts and excavates the shape information and designs a cross-level symbolic shape representation. With the insight that what fusion optimization relies upon is tensor shape relationships between adjacent operators rather than exact shape values, it proposes the dynamic shape fusion approach based on shape information propagation. To generate code that adapts to arbitrary shapes efficiently, it proposes a compile-time and runtime combined code generation approach. Finally, it presents a complete optimization pipeline for dynamic shape models and implements an industrial-grade ML compiler, named BladeDISC. The extensive evaluation demonstrates that BladeDISC outperforms PyTorch, TorchScript, TVM, ONNX Runtime, XLA, Torch Inductor (dynamic shape), and TensorRT by up to 6.95×, 6.25×, 4.08×, 2.04×, 2.06×, 7.92×, and 4.16× (3.54×, 3.12×, 1.95×, 1.47×, 1.24×, 2.93×, and 1.46× on average) in terms of end-to-end inference speedup on the A10 and T4 GPU, respectively. BladeDISC's source code is publicly available at https://github.com/alibaba/BladeDISC.
Zhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu 0004, Wenyi Zhao, Tianyou Guo, Xiafei Qiu, Minmin Sun, Feng Zhang 0007, Xiaoyong Du 0001, Jidong Zhai, Wei Lin 0016
Proc. ACM Manag. Data3
2022 Exploring Query Processing on CPU-GPU Integrated Edge Device
abstract
Huge amounts of data have been generated on edge devices every day, which requires efficient data analytics and management. However, due to the limited computing capacity of these edge devices, query processing at the edge faces tremendous pressure. Fortunately, in recent years, hardware vendors have integrated heterogeneous coprocessors, such as GPUs, into the edge device, which can provide much more computing power. Furthermore, the CPU-GPU integrated edge device has shown significant benefits in a variety of situations. Therefore, the exploration of query processing on such CPU-GPU integrated edge devices becomes an urgent need. In this article, we develop a fine-grained query processing engine, called FineQuery, which can perform efficient query processing on CPU-GPU integrated edge devices. Particularly, FineQuery can take advantage of both architectural features of edge devices and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experiments show that on TPC-H workloads, FineQuery reduces 42.81% latency and improves 2.39× bandwidth utilization on average compared to the implementation of using only GPU or CPU. Furthermore, query processing at the edge can bring significant performance-per-cost benefits and energy efficiency. On average, FineQuery at the edge brings 21× performance-per-cost ratio and 4× energy efficiency compared with processing the data on a discrete GPU platform.
Jiesong Liu, Feng Zhang 0007, Hourun Li, Dalin Wang, Weitao Wan, Xiaokun Fang, Jidong Zhai, Xiaoyong Du 0001
IEEE Trans. Parallel Distributed Syst.4
2021 FineQuery: Fine-Grained Query Processing on CPU-GPU Integrated Architectures
abstract
Using heterogeneous coprocessors, such as GPUs, to accelerate complicated SQL queries has been proved to be effective in the database domain. Previous works show that taking advantage of the high parallelism and computing capacity of heterogeneous coprocessors can bring significant performance improvements. However, in the discrete memory architecture, the advantages of heterogeneous coprocessors will be weakened due to the low PCI-e bandwidth and high latency. Fortunately, hardware vendors have proposed a novel integration architecture design, which integrates CPU and GPU on the same chip. This integrated architecture allows the GPU and CPU to share the same unified memory, taking new opportunities for fine-grained collaboration between the GPU and CPU to optimize SQL queries. In this paper, we propose a query processing engine, called FineQuery, to optimize the execution of SQL queries on the integrated architecture. FineQuery can take advantage of both architectural features and query characteristics by performing fine-grained workload scheduling between the CPU and the GPU. Experimental results show that 1) on the integration architecture, FineQuery can reduce the latency by 25.30% and increase the bandwidth utilization by 39.46% on average. 2) FineQuery on the integrated architecture achieves $13.74\times$ the performance-per-cost ratio and $6.87\times$ energy efficiency over query processing on the discrete GPU platform.
Dalin Wang, Feng Zhang 0007, Weitao Wan, Hourun Li, Xiaoyong Du 0001
CLUSTER1
2021 TADOC: Text analytics directly on compression
Feng Zhang 0007, Jidong Zhai, Xipeng Shen, Dalin Wang, Zheng Chen 0023, Onur Mutlu, Xiaoyong Du 0001
VLDB J.4