VLDB 2026 Research / reviewers in the wild / expert
Yue Zhao 0011
dblp:48/76-11
· DBLP profile ↗
11ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0003-4676-3612ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 52% GPUs and heterogeneous computing · 22% Storage systems · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 59% Machine learning and data management · 41% | |
| Software engineering, system software, and programming languages
1 paper |
Software maintenance and evolution · 50% Compilers and program optimization · 50% | |
| Theoretical computer science
1 paper |
Coding theory · 100% |
Topics — the 17 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › compiler optimization
computation graph optimization |
0.4 | 1 | 2020 | HARP: holistic analysis for refactoring Python-based analytics programs · ICSE 2020 |
Software maintenance and evolution
refactoring |
0.4 | 1 | 2020 | HARP: holistic analysis for refactoring Python-based analytics programs · ICSE 2020 |
High-performance computing
performance optimization at scale |
0.4 | 1 | 2020 | Enabling Runtime SpMV Format Selection through an Overhead Conscious Method · IEEE Trans. Parallel Distributed Syst. 2020 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.4 | 1 | 2020 | Enabling Runtime SpMV Format Selection through an Overhead Conscious Method · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems › storage management
storage format selection |
0.4 | 1 | 2020 | Enabling Runtime SpMV Format Selection through an Overhead Conscious Method · IEEE Trans. Parallel Distributed Syst. 2020 |
Machine learning and data management
learned database components |
0.3 | 1 | 2018 | Bridging the gap between deep learning and sparse matrix format selection · PPoPP 2018 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.3 | 1 | 2018 | Bridging the gap between deep learning and sparse matrix format selection · PPoPP 2018 |
High-performance computing › sparse linear algebra
sparse matrix format selection |
0.3 | 1 | 2018 | Bridging the gap between deep learning and sparse matrix format selection · PPoPP 2018 |
GPUs and heterogeneous computing › GPU scheduling
GPU kernel scheduling |
0.3 | 1 | 2017 | EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPU · PPoPP 2017 |
GPUs and heterogeneous computing
GPU sharing |
0.3 | 1 | 2017 | EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPU · PPoPP 2017 |
Embedded and real-time systems › real-time scheduling
preemptive scheduling |
0.3 | 1 | 2017 | EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPU · PPoPP 2017 |
Data mining
clustering |
0.2 | 1 | 2015 | Yinyang K-Means: A Drop-In Replacement of the Classic K-Means with Consistent Speedup · ICML 2015 |
Data mining › clustering › k-means clustering
k-means acceleration |
0.2 | 1 | 2015 | Yinyang K-Means: A Drop-In Replacement of the Classic K-Means with Consistent Speedup · ICML 2015 |
Data mining › clustering
k-means clustering |
0.2 | 1 | 2015 | Yinyang K-Means: A Drop-In Replacement of the Classic K-Means with Consistent Speedup · ICML 2015 |
Coding theory
error-correcting codes |
0.2 | 1 | 2014 | Implementation of Decoders for LDPC Block Codes and LDPC Convolutional Codes Based on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 |
Coding theory › error-correcting codes › LDPC codes
LDPC convolutional codes |
0.2 | 1 | 2014 | Implementation of Decoders for LDPC Block Codes and LDPC Convolutional Codes Based on GPUs · IEEE Trans. Parallel Distributed Syst. 2014 |
Performance modeling and evaluation › performance prediction
execution time prediction |
0.1 | 1 | 2020 | Enabling Runtime SpMV Format Selection through an Overhead Conscious Method · IEEE Trans. Parallel Distributed Syst. 2020 |
Methods — techniques the papers use, named apart from their topics
static analysis · 0.9speculative analysis · 0.9computation graph analysis · 0.9deep learning · 0.7cross-architecture model migration · 0.7two-stage lazy-and-light scheme · 0.4regression model · 0.4neural network-based time series prediction · 0.4thread hierarchy optimization · 0.4parallel decoding · 0.4message passing · 0.4software framework · 0.3priority-based preemptive scheduling · 0.3lower and upper bound maintenance · 0.2center clustering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | HARP: holistic analysis for refactoring Python-based analytics programsabstractModern machine learning programs are often written in Python, with the main computations specified through calls to some highly optimized libraries (e.g., TensorFlow, PyTorch). How to maximize the computing efficiency of such programs is essential for many application domains, which has drawn lots of recent attention. This work points out a common limitation in existing efforts: they focus their views only on the static computation graphs specified by library APIs, but leave the influence from the hosting Python code largely unconsidered. The limitation often causes them to miss the big picture and hence many important optimization opportunities. This work proposes a new approach named HARP to address the problem. HARP enables holistic analysis that spans across computation graphs and their hosting Python code. HARP achieves it through a set of novel techniques: analytics-conscious speculative analysis to circumvent Python complexities, a unified representation augmented computation graphs to capture all dimensions of knowledge related with the holistic analysis, and conditioned feedback mechanism to allow risk-controlled aggressive analysis. Refactoring based on HARP gives 1.3--3X and 2.07X average speedups on a set of TensorFlow and PyTorch programs. Yue Zhao 0011, Xipeng Shen |
ICSE | 2 |
| 2020 | Enabling Runtime SpMV Format Selection through an Overhead Conscious MethodabstractSparse matrix-vector multiplication (SpMV) is an important kernel and its performance is critical for many applications. Storage format selection is to select the best format to store a sparse matrix; it is essential for SpMV performance. Prior studies have focused on predicting the format that helps SpMV run fastest, but have ignored the runtime prediction and format conversion overhead. This work shows that the runtime overhead makes the predictions from previous solutions frequently sub-optimal and sometimes inferior regarding the end-to-end time. It proposes a new paradigm for SpMV storage selection, an overhead-conscious method. Through carefully designed regression models and neural network-based time series prediction models, the method captures the influence imposed on the overall program performance by the overhead and the benefits of format prediction and conversions. The method employs a novel two-stage lazy-and-light scheme to help control the possible negative effects of format predictions, and at the same time, maximize the overall format conversion benefits. Experiments show that the technique outperforms previous techniques significantly. It improves the overall performance of applications by 1.21X to 1.53X, significantly larger than the 0.83X to 1.25X upper-bound speedups overhead-oblivious methods could give. Yue Zhao 0011, Xipeng Shen |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Overhead-Conscious Format Selection for SpMV-Based ApplicationsabstractSparse matrix vector multiplication (SpMV) is an important kernel in many applications and is often the major performance bottleneck. The storage format of sparse matrices critically affects the performance of SpMV. Although there have been previous studies on selecting the appropriate format for a given matrix, they have ignored the influence of runtime prediction overhead and format conversion overhead. For many common uses of SpMV, such overhead is part of the execution times and may outweigh the benefits of new formats. Ignoring them makes the predictions from previous solutions frequently suboptimal and sometimes inferior. On the other hand, the overhead is difficult to consider, as it, along with the benefits of having a new format, varies from matrix to matrix, and from application to application. This work proposes a solution. It first explores the pros and cons of various possible treatments to the overhead in the format selection problem. It then presents an explicit approach which involves several regression models for capturing the influence of the overhead and benefits of format conversions on the overall program performance. It proposes a two-stage lazy-and-light scheme to help control the risks in the format predictions and at the same time maximize the overall format conversion benefits. Experiments show that the technique outperforms previous techniques significantly. It improves the overall performance of applications by 1.14X to 1.43X, significantly larger than the 0.82X to 1.24X upperbound speedups overhead-oblivious methods could give. Yue Zhao 0011, Xipeng Shen, Graham Yiu |
IPDPS | 1 |
| 2018 | Bridging the gap between deep learning and sparse matrix format selectionabstractThis work presents a systematic exploration on the promise and special challenges of deep learning for sparse matrix format selection---a problem of determining the best storage format for a matrix to maximize the performance of Sparse Matrix Vector Multiplication (SpMV). It describes how to effectively bridge the gap between deep learning and the special needs of the pillar HPC problem through a set of techniques on matrix representations, deep learning structure, and cross-architecture model migrations. The new solution cuts format selection errors by two thirds, and improves SpMV performance by 1.73X on average over the state of the art. Yue Zhao 0011, Jiajia Li 0001, Chunhua Liao, Xipeng Shen |
PPoPP | 1 |
| 2017 | POSTER: Bridging the Gap Between Deep Learning and Sparse Matrix Format SelectionabstractIn this work, we conduct a systematic exploration on the promise and challenges of deep learning for the sparse matrix format selection. We propose a set of novel techniques to solve special challenges to deep learning, including input matrix representations, a late-merging deep neural network structure design, and the use of transfer learning to alleviate cross-architecture portability issues. Yue Zhao 0011, Jiajia Li 0001, Chunhua Liao, Xipeng Shen |
PACT | 1 |
| 2017 | EffiSha: A Software Framework for Enabling Effficient Preemptive Scheduling of GPUabstractModern GPUs are broadly adopted in many multitasking environments, including data centers and smartphones. However, the current support for the scheduling of multiple GPU kernels (from different applications) is limited, forming a major barrier for GPU to meet many practical needs. This work for the first time demonstrates that on existing GPUs, efficient preemptive scheduling of GPU kernels is possible even without special hardware support. Specifically, it presents EffiSha, a pure software framework that enables preemptive scheduling of GPU kernels with very low overhead. The enabled preemptive scheduler offers flexible support of kernels of different priorities, and demonstrates significant potential for reducing the average turnaround time and improving the system overall throughput of programs that time share a modern GPU. Guoyang Chen, Yue Zhao 0011, Xipeng Shen, Huiyang Zhou |
PPoPP | 2 |
| 2017 | POSTER: An Infrastructure for HPC Knowledge Sharing and ReuseabstractThis paper presents a prototype infrastructure for addressing the barriers for effective accumulation, sharing, and reuse of the various types of knowledge for high performance parallel computing. Yue Zhao 0011, Chunhua Liao, Xipeng Shen |
PPoPP | 1 |
| 2016 | Towards Ontology-Based Program AnalysisabstractProgram analysis is fundamental for program optimizations, debugging, and many other tasks. But developing program analyses has been a challenging and error-prone process for general users. Declarative program analysis has shown the promise to dramatically improve the productivity in the development of program analyses. Current declarative program analysis is however subject to some major limitations in supporting cooperations among analysis tools, guiding program optimizations, and often requires much effort for repeated program preprocessing. In this work, we advocate the integration of ontology into declarative program analysis. As a way to standardize the definitions of concepts in a domain and the representation of the knowledge in the domain, ontology offers a promising way to address the limitations of current declarative program analysis. We develop a prototype framework named PATO for conducting program analysis upon ontology-based program representation. Experiments on six program analyses confirm the potential of ontology for complementing existing declarative program analysis. It supports multiple analyses without separate program preprocessing, promotes cooperative Liveness analysis between two compilers, and effectively guides a data placement optimization for Graphic Processing Units (GPU). Yue Zhao 0011, Guoyang Chen, Chunhua Liao, Xipeng Shen |
ECOOP | 1 |
| 2015 | Yinyang K-Means: A Drop-In Replacement of the Classic K-Means with Consistent SpeedupabstractThis paper presents Yinyang K-means, a new algorithm for K-means clustering. By clustering the centers in the initial stage, and leveraging efficiently maintained lower and upper bounds between a point and centers, it more effectively avoids unnecessary distance calculations than prior algorithms. It significantly outperforms classic K-means and prior alternative K-means algorithms consistently across all experimented data sets, cluster numbers, and machine configurations. The consistent, superior performance—plus its simplicity, user-control of overheads, and guarantee in producing the same clustering results as the standard K-means does—makes Yinyang K-means a drop-in replacement of the classic K-means with an order of magnitude higher performance. Yufei Ding 0001, Yue Zhao 0011, Xipeng Shen, Madan Musuvathi, Todd Mytkowicz |
ICML | 2 |
| 2014 | Implementation of Decoders for LDPC Block Codes and LDPC Convolutional Codes Based on GPUsabstractIn this paper, efficient LDPC block-code decoders/simulators which run on graphics processing units (GPUs) are proposed. We also implement the decoder for the LDPC convolutional code (LDPCCC). The LDPCCC is derived from a predesigned quasi-cyclic LDPC block code with good error performance. Compared to the decoder based on the randomly constructed LDPCCC code, the complexity of the proposed LDPCCC decoder is reduced due to the periodicity of the derived LDPCCC and the properties of the quasicyclic structure. In our proposed decoder architecture, Γ (Γ is a multiple of a warp) codewords are decoded together, and hence, the messages of Γ codewords are also processed together. Since all the Γ codewords share the same Tanner graph, messages of the Γ distinct codewords corresponding to the same edge can be grouped into one package and stored linearly. By optimizing the data structures of the messages used in the decoding process, both the read and write processes can be performed in a highly parallel manner by the GPUs. In addition, a thread hierarchy minimizing the divergence of the threads is deployed, and it can maximize the efficiency of the parallel execution. With the use of a large number of cores in the GPU to perform the simple computations simultaneously, our GPU-based LDPC decoder can obtain hundreds of times speedup compared with a serial CPU-based simulator and over 40 times speedup compared with an eight-thread CPU-based simulator. Yue Zhao 0011, Francis C. M. Lau 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Efficient Decoding of QC-LDPC Codes Using GPUs
Yue Zhao 0011, Xu Chen 0018, Chiu-Wing Sham, Wai Man Tam, Francis C. M. Lau 0002 |
ICA3PP (1) | 1 |