EDBT 2026 Demo / reviewers in the wild / expert
Qianchao Zhu
dblp:304/5600
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0001-5021-2912ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative DecodingabstractAutoregressive decoding inherently limits the inference throughput of Large Language Model (LLM) due to its sequential dependency.Speculative decoding mitigates this by verifying multiple predicted tokens in parallel, but its efficiency remains constrained by what we identify as verification heterogeneity-the uneven difficulty of verifying different speculative candidates.In practice, a small subset of high-confidence predictions accounts for most successful verifications, yet existing methods treat all candidates uniformly, leading to redundant computation.We present HeteroSpec, a heterogeneity-adaptive speculative decoding framework that allocates verification effort in proportion to candidate uncertainty.Het-eroSpec estimates verification complexity using a lightweight entropy-based quantifier, partitions candidates via a data-driven stratification policy, and dynamically tunes speculative depth and pruning thresholds through coordinated optimization.Across five benchmarks and four LLMs, HeteroSpec delivers an average 4.24× decoding speedup over state-of-the-art methods such as EAGLE-3, while preserving exact output distributions.Crucially, HeteroSpec requires no model retraining and remains compatible with other inference optimizations, making it a practical direction for improving speculative decoding efficiency. Siran Liu, Qianchao Zhu, Zane Cao, Yongchao He |
ACL (1) | 3 |
| 2026 | Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model TrainingabstractTraining large language models (LLMs) with increasingly long and varying sequence lengths introduces severe load imbalance challenges in large-scale data-parallel training. Recent frameworks attempt to mitigate these issues through data reorganization or hybrid parallel strategies. However, they often overlook how computational and communication costs scale with sequence length, resulting in suboptimal performance. We identify three critical challenges: (1) varying computation-to-communication ratios across sequences of different lengths in distributed attention, (2) mismatch between static NIC-GPU affinity and dynamic parallel workloads, and (3) distinct optimal partitioning strategies required for quadratic attention versus linear components. Chang Chen 0001, Tiancheng Chen, Jiangfei Duan, Qianchao Zhu, Zerui Wang, Qinghao Hu 0004, Peng Sun 0006, Chao Yang 0002, Torsten Hoefler |
EuroSys | 4 |
| 2025 | StructILU: Dependency-Preserving Incomplete LU with Hierarchical Parallelism for Structured Grid PDEs on GPUsabstractThe Incomplete LU (ILU) computation is a crucial component for solving large-scale sparse linear systems arising from partial differential equations (PDEs), many of which are discretized on structured grids.However, due to inherent loop-carried data dependencies in ILU computation, implementing it on GPUs with massive computing units poses significant challenges.Existing methods either experience Hao Luo 0015, Qianchao Zhu, Xiaochen Hao, Chunxi Lei, Chengdi Ma, Yun Liang 0001, Chao Yang 0002 |
ICS | 2 |
| 2024 | Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningabstractEfficiently training large language models (LLMs) necessitates the adoption of hybrid parallel methods, integrating multiple communications collectives within distributed partitioned graphs. Overcoming communication bottlenecks is crucial and is often achieved through communication and computation overlaps. However, existing overlap methodologies tend to lean towards either fine-grained kernel fusion or limited operation scheduling, constraining performance optimization in heterogeneous training environments. Chang Chen 0001, Qianchao Zhu, Jiangfei Duan, Peng Sun 0006, Xingcheng Zhang, Chao Yang 0002 |
ASPLOS (3) | 3 |
| 2024 | FreeStencil: A Fine-Grained Solver Compiler with Graph and Kernel Optimizations on Structured Meshes for Modern GPUsabstractParallel numerical solvers for partial differential equations (PDEs) on structured meshes are critical in various scientific computing applications. However, current PDE solver frameworks suffer from programming or performance issues on modern accelerators such as GPUs. The complexity in programming significantly imposes a heavy burden on the development of new solver algorithms and incurs performance challenges, particularly for operator-based frameworks with redundant implementations or those that require cross-architecture optimization. In this paper, we propose FreeStencil, a linear solver compiler that emphasizes programming and optimization in fine grain. For programming, we utilized modular abstraction to implement matrix-free stencil computations in a fine-grained manner to avoid redundancy. For performance, we enabled graph optimizations towards solver iterations and employed multi-level tiling with fine-grained hardware awareness for high-performance GPU code generation. Experimental results demonstrate that FreeStencil achieves up to a 3.29x speedup (average 2.32x) on multiple GPU platforms for typical applications. Qianchao Zhu |
ICPP | 1 |
| 2023 | Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured GridabstractPartial differential equation (PDE) solvers are extensively utilized across numerous scientific and engineering fields. However, achieving high performance and scalability often necessitates intricate and low-level programming, particularly when leveraging deterministic sparsity patterns in structured grids. In this paper, we propose an innovative domain-specific language (DSL), Mat2Stencil, with its compiler, for PDE solvers on structured grids. Mat2Stencil introduces a structured sparse matrix abstraction, facilitating modular, flexible, and easy-to-use expression of solvers across a broad spectrum, encompassing components such as Jacobi or Gauss-Seidel preconditioners, incomplete LU or Cholesky decompositions, and multigrid methods built upon them. Our DSL compiler subsequently generates matrix-free code consisting of generalized stencils through multi-stage programming. The code allows spatial loop-carried dependence in the form of quasi-affine loops, in addition to the Jacobi-style stencil’s embarrassingly parallel on spatial dimensions. We further propose a novel automatic parallelization technique for the spatially dependent loops, which offers a compile-time deterministic task partitioning for threading, calculates necessary inter-thread synchronization automatically, and generates an efficient multi-threaded implementation with fine-grained synchronization. Implementing 4 benchmarking programs, 3 of them being the pseudo-applications in NAS Parallel Benchmarks with 6.3% lines of code and 1 being matrix-free High Performance Conjugate Gradients with 16.4% lines of code, we achieve up to 1.67× and on average 1.03× performance compared to manual implementations. Huanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu 0003 |
Proc. ACM Program. Lang. | 3 |
| 2022 | EasyView: Enabling and Scheduling Tensor Views in Deep Learning CompilersabstractIn recent years, memory-intensive operations are becoming dominant in efficiency of running novel neural networks. Just-in-time operator fusion on accelerating devices like GPU proves an effective method for optimizing memory-intensive operations, and suits the numerous varying model structures. In particular, we find memory-intensive operations on tensor views are ubiquitous in neural network implementations. Tensors are the de facto representation for numerical data in deep learning areas, while tensor views cover a bunch of sophisticated syntax, which allow various interpretations on the underlying tensor data without memory copy. The support of views in deep learning compilers could greatly enlarge operator fusion scope, and appeal to optimizing novel neural networks. Nevertheless, mainstream solutions in state-of-the-art deep learning compilers exhibit imperfections either in view syntax representations or operator fusion. In this article, we propose EasyView, which enables and schedules tensor views in an end-to-end workflow from neural networks onto devices. Aiming at maximizing memory utilization and reducing data movement, we categorize various view contexts in high-level language, and lower views in accordance with different scenarios. Reference-semantic in terms of views are kept in the lowering from native high-level language features to intermediate representations. Based on the reserved reference-semantics, memory activities related to data dependence of read and write are tracked for further compute and memory optimization. Besides, ample operator fusion is applied to memory-intensive operations with views. In our tests, the proposed work could get average 5.63X, 2.44X, and 4.67X speedup compared with the XLA, JAX, and TorchScript, respectively for hotspot Python functions. In addition, operation fusion with views could bring 8.02% performance improvement in end-to-end neural networks. Lijuan Jiang, Qianchao Zhu, Shengen Yan, Xingcheng Zhang, Dahua Lin, Wenjing Ma, Zhouyang Li, Minxi Jin, Chao Yang 0002 |
ICPP | 3 |
| 2021 | Enabling and scaling the HPCG benchmark on the newest generation Sunway supercomputer with 42 million heterogeneous coresabstractWe study and evaluate performance optimization techniques for the HPCG benchmark on the newest generation Sunway supercomputer. Specifically, a two-level blocking scheme is proposed to expose adequate parallelism in the symmetric Gauss-Seidel kernel while keeping a fast convergence rate, a fine-grained kernel fusion technique is developed to alleviate the bandwidth load on local storage with small capacity, and a low overhead thread collaboration method is presented to efficiently move data between threads and hide its cost with data transfer operations. Test results show that the optimized HPCG code is able to exploit 73.0% of the theoretical memory bandwidth, and scale to over 42 million heterogeneous cores with 95.5% weak-scaling efficiency and 5.91 Pflop/s performance. We also study how the performance can be improved if the specific rules of HPCG are not fully obeyed, and design dependency preserving parallelization and vectorization methods, further boosting performance to 27.6 Pflop/s. Qianchao Zhu, Hao Luo 0015, Chao Yang 0002, Mingshuo Ding, Wanwang Yin, Xinhui Yuan |
SC | 1 |