VLDB 2026 Research / reviewers in the wild / expert
Qidong Zhao
dblp:251/8359
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-0872-1246ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpecProto: A Parallelizing Compiler for Speculative Decoding of Large Protocol Buffers DataabstractProtobuf is a widely used data serialization format, especially in cloud environments. However, existing compilers generate only serial decoders, limiting scalability for large datasets. While parallel parsing has been studied for textual formats (e.g., XML), parallel decoding of binary formats like Protobuf remains unexplored, which present unique opportunities. Chales Hong, Dhruv Parmar, Zhijia Zhao 0001, Qidong Zhao, Xu Liu 0001 |
ASPLOS (2) | 6 |
| 2026 | Triton-Sanitizer: A Fast and Device-Agnostic Memory Sanitizer for Triton with Rich Diagnostic ContextabstractMemory access errors remain one of the most pervasive bugs in GPU programming. Existing GPU sanitizers such as compute-sanitizer detect memory access errors by instrumenting every memory instruction in low-level IRs or binaries, which imposes high overhead and provides minimal memory access error diagnostic context for fixing problems. We present Triton-Sanitizer, the first device-agnostic memory sanitizer designed for Triton, a domain-specific language for developing portable, efficient GPU kernels for deep learning workloads. Triton-Sanitizer leverages Triton's tile-oriented semantics to construct symbolic expressions for memory addresses and masks, verifies them with an SMT solver, and selectively falls back to eager simulation for indirect accesses. This hybrid analysis enables precise detection of memory access errors without false positives while avoiding the cost of per-access instrumentation. Beyond detection, Triton-Sanitizer generates rich diagnostic reports that attribute violations to the tensors nearest to the violated addresses, track the complete call path, and expose the symbolic operations responsible for incorrect addresses. Evaluated on seven widely used open-source repositories of Triton kernels, Triton-Sanitizer uncovered 24 previously unknown memory access errors, of which 8 have already been fixed and upstreamed by us. Compared to compute-sanitizer, Triton-Sanitizer achieves speedups ranging from 1.07× to 14.66×, with an average improvement of 1.62×, demonstrating its ability to enhance performance, precision, and usability in memory access error detection. Hao Wu 0077, Qidong Zhao, Songqing Chen, Yueming Hao, Tony C. W. Liu, Adnan Aziz, Keren Zhou 0001 |
ASPLOS (2) | 2 |
| 2025 | DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning WorkloadsabstractEffective performance optimization of deep learning models requires comprehensive profiling across heterogeneous computing environments, yet existing tools fail to bridge the semantic gap between high-level operations and low-level execution. This paper presents DeepContext, a novel profiling system that correlates program contexts across Python code, deep learning frameworks, C/C++ libraries, and GPU execution. DeepContext features a framework-agnostic shim layer that seamlessly correlates the behavior of the deep learning framework with hardware performance metrics. Furthermore, DeepContext provides an automated performance analyzer that offers actionable optimization guidance based on its holistic view of the entire software stack of deep learning applications. DeepContext works for mainstream deep learning frameworks and runs on modern CPU+GPU architectures with low overhead. Our evaluation demonstrates that DeepContext uncovers previously hidden performance bottlenecks in real-world deep-learning applications. Guided by DeepContext, we are able to fix multiple performance issues, achieving speed-ups between 1.06× and 1.66×. Qidong Zhao, Hao Wu 0077, Yueming Hao, Zilingfeng Ye, Jiajia Li 0001, Xu Liu 0001, Keren Zhou 0001 |
ASPLOS (3) | 1 |
| 2024 | DrPy: Pinpointing Inefficient Memory Usage in Multi-Layer Python ApplicationsabstractPython has become an increasingly popular programming language, especially in the areas of data analytics and machine learning. Many modern Python packages employ a multi-layer design: the Python layer manages various packages and expresses high-level algorithms; the native layer is written in C/C++/Fortran/CUDA for efficient computation. Typically, each layer manages its own computation and memory and exposes APIs for cross-layer interactions. Without holistic optimization, performance inefficiencies can exist at the boundary between layers. In this paper, we develop DrPy, a novel profiler that pinpoints such memory inefficiencies across layers in Python applications. Unlike existing tools, DrPy takes a hybrid and fine-grained approach to track memory objects and their usage in both Python and native layers. DrPy correlates the behavior of memory objects across layers and builds an object flow graph to pinpoint memory inefficiencies. In addition, DrPy captures rich information associated with object flow graphs, such as call paths and source code attribution to guide intuitive code optimization. Guided by DrPy, we are able to optimize many Python applications with non-trivial performance improvement. Many optimization patches have been validated by application developers and committed to application repositories. Jinku Cui, Qidong Zhao, Yueming Hao, Xu Liu 0001 |
CGO | 2 |
| 2024 | EasyView: Bringing Performance Profiles into Integrated Development EnvironmentsabstractDynamic program performance analysis (also known as profiling) is well-known for its powerful capabilities of identifying performance inefficiencies in software packages. Although a large number of profiling techniques are developed in academia and industry, very few of them are widely used by software developers in their regular software developing activities. There are three major reasons. First, the profiling tools (also known as profilers) are disjoint from the coding environments such as IDEs and editors; frequently switching focus between them significantly complicates the entire cycle of software development. Second, mastering various tools to interpret their analysis results requires substantial efforts; even worse, many tools have their own design of graphical user interfaces (GUI) for data presentation, which steepens the learning curves. Third, most existing profilers expose few interfaces to support user-defined analysis, which makes the tools less customizable to fulfill diverse user demands. We develop EasyView, a general solution to integrate the interpretation and visualization of various profiling results in the coding environments, which bridges software developers closer with profilers during the code development cycle. The novelty of EasyView lies in its significant improvement on the usability of profilers. EasyView not only provides deep insights to support intuitive analysis and optimization in a simple interface, but also enhances user experiences in using the profilers effectively and efficiently in the IDEs. Our evaluation shows that EasyView is able to support various profilers for different languages and provide unique insights into performance inefficiencies in different domains. Our user studies show that EasyView can largely improve the usability of profilers in software development cycles via facilitating performance debugging efforts. Qidong Zhao, Milind Chabbi, Xu Liu 0001 |
CGO | 1 |
| 2024 | Autost: Training-Free Neural Architecture Search For Spiking TransformersabstractSpiking Transformers have gained considerable attention because they achieve both the energy efficiency of Spiking Neural Networks (SNNs) and the high capacity of Transformers. However, the existing Spiking Transformer architectures, derived from Artificial Neural Networks (ANNs), exhibit a notable architectural gap, resulting in suboptimal performance compared to their ANN counterparts. Manually discovering optimal architectures is time-consuming. To address these limitations, we introduce AutoST, a training-free NAS method for Spiking Transformers, to rapidly identify high-performance Spiking Transformer architectures. Unlike existing training-free NAS methods, which struggle with the non-differentiability and high sparsity inherent in SNNs, we propose to utilize Floating-Point Operations (FLOPs) as a performance metric, which is independent of model computations and training dynamics, leading to a stronger correlation with performance. Our extensive experiments show that AutoST models outperform state-of-the-art manually or automatically designed SNN architectures on static and neuromorphic datasets. Full code, model, and data are released for reproduction.1 Qidong Zhao, Jinku Cui, Xu Liu 0001, Dongkuan Xu |
ICASSP | 2 |
| 2024 | Purpose Enhanced Reasoning through Iterative Prompting: Uncover Latent Robustness of ChatGPT on Code Comprehension
Qidong Zhao, Dongkuan Xu, Xu Liu 0001 |
IJCAI | 2 |
| 2023 | DroidPerf: Profiling Memory Objects on Android DevicesabstractOptimizing performance inefficiencies in memory hierarchies is well-known for native languages, such as C and C++. There are few studies, however, on exploring memory inefficiencies in Android Runtime (ART). Running in ART, managed languages, such as Java and Kotlin, employ various abstractions, such as runtime support, ahead-of-time (AOT) compilation, and garbage collection (GC), which hide important execution details from the plain source code. Qidong Zhao, Shuyin Jiao, Xu Liu 0001 |
MobiCom | 2 |
| 2022 | OJXPERF: Featherlight Object Replica Detection for Java ProgramsabstractMemory bloat is an important source of inefficiency in complex production software, especially in software written in managed languages such as Java. Prior approaches to this problem have focused on identifying objects that outlive their life span. Few studies have, however, looked into whether and to what extent myriad objects of the same type are identical. A quantitative assessment of identical objects with code-level attribution can assist developers in refactoring code to eliminate object bloat, and favor reuse of existing object(s). The result is reduced memory pressure, reduced allocation and garbage collection, enhanced data locality, and reduced re-computation, all of which result in superior performance. Hao Xu 0048, Qidong Zhao, Pengfei Su 0001, Milind Chabbi, Shuyin Jiao, Xu Liu 0001 |
ICSE | 3 |
| 2021 | A multi-objective decomposition-based ant colony optimisation algorithm with negative pheromoneabstractExisting ant colony algorithms only have one kind of pheromone. They use non-dominated solutions to update it while not making use of dominated solutions, which can provide valuable information for guiding the subsequent foraging process. To make full use of dominated solutions, we create a new kind of pheromone temporarily called a negative pheromone and propose a new ant colony optimisation algorithm called NMOACO/D, which combines MOEA/D-ACO with the negative pheromone. Many experiments have been carried out in this study to compare NMOACO/D with MOEA/D-ACO and other algorithms on several bi-objective travelling salesman problems. We demonstrate that NMOACO/D outperforms the MOEA/D-ACO and six different recently proposed related algorithms on all nine test instances. We also evaluate the effect of negative pheromone on the performance of the NMOACO/D. The results in this paper show that correctly making use of the information related to dominated solutions can further improve the ant colony algorithm performance. Jiaxu Ning, Qidong Zhao, Yunfei Feng |
J. Exp. Theor. Artif. Intell. | 2 |
| 2021 | A Novel Fireworks Algorithm for the Protein-Ligand Docking on the AutoDock
Zhuoran Liu 0002, Dingde Jiang, Changsheng Zhang 0001, Haitong Zhao, Qidong Zhao, Bin Zhang 0001 |
Mob. Networks Appl. | 5 |
| 2020 | DrCCTProf: a fine-grained call path profiler for ARM-based clustersabstractARM is an attractive CPU architecture for exascale systems because of its energy efficiency. As a recent entry into the HPC paradigm, ARM lags in its software stack, especially in the performance tooling aspect. Notably, there is a lack of fine-grained measurement tools to analyze fully optimized HPC binary executables on ARM processors. In this paper, we introduce DRCCTPROF — a fine-grained call path profiling framework for binaries running on ARM architectures. The unique ability of DRCCTPROF is to obtain full calling context at any and every machine instruction that executes, which provides more detailed diagnostic feedback for performance optimization and correctness tools. Furthermore, DRCCTPROF not only associates any instruction with source code along the call path, but also associates memory access instructions back to the constituent data object. Finally, DRCCTPROF incurs moderate overhead and provides a compact view to visualize the profiles collected from parallel executions. Qidong Zhao, Xu Liu 0001, Milind Chabbi |
SC | 1 |