EDBT 2026 Demo / reviewers in the wild / expert
Heng Shi 0005
dblp:180/3315-5
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-1779-9221ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingabstractRecent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be transformed to matrix multiplications. However, this transformation introduces undesired sparsity into the kernel matrix, leading to significant redundant computation. Qiqi Gu 0002, Chenpeng Wu, Heng Shi 0005, Jianguo Yao 0002 |
PPoPP | 3 |
| 2025 | Postiz: Extending Post-increment Addressing for Loop Optimization and Code Size ReductionabstractMemory access instructions with auto-addressing modes are prevalent in various Instruction Set Architectures (ISAs), yet their use in compilers remains limited. Existing methods address code optimization in one of two ways: they either focus on reducing code size, but are constrained to basic block-level optimizations and may not fully exploit architectural benefits, or they optimize loop performance, often neglecting the advantages of post-increment instructions and focusing primarily on innermost loops while leaving outer loops unoptimized. To address these shortcomings and meet the needs of real-world Machine Learning (ML) applications, we introduce Postiz, a novel post-increment loop optimization technique. Postiz extends post-increment optimizations beyond traditional limits, incorporating enhancements for inner loops, cross-loop regions, and nested loop structures. Through a profitability analysis, Postiz optimizes code judiciously, leveraging architectural advantages and reducing code size without compromising improvement made by other optimizations. Our experiments show that Postiz is effective, achieving an optimization coverage of 98.04% on MobileNet and BERT benchmarks. In comparison to default LLVM optimization, Postiz generates approximately four times more post-increment instructions. Moreover, it reduces code size by an average of 9.45% across various platforms. These improvements represent significant advancements over current methods, showcasing Postiz’s potential to enhance compiler optimizations in a meaningful way. Enming Fan, Xiaofeng Guan, Heng Shi 0005, Hao Zhou 0009, Jianguo Yao 0002 |
CGO | 4 |
| 2025 | Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor CoresabstractThe escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to enhance efficiency without compromising model accuracy. Structured sparsity emerges as a compelling strategy to address these challenges by leveraging the emerging sparse computing hardware. Prior works mainly focus on the sparsity in model parameters, neglecting the inherent sparse patterns in activations. This oversight can lead to additional computational costs associated with activations, potentially resulting in suboptimal performance. Chenpeng Wu, Qiqi Gu 0002, Heng Shi 0005, Jianguo Yao 0002, Haibing Guan |
EuroSys | 3 |
| 2024 | Boost Linear Algebra Computation Performance via Efficient VNNI UtilizationabstractIntel's Vector Neural Network Instruction (VNNI) provides higher efficiency on calculating dense linear algebra (DLA) computations than conventional SIMD instructions. However, existing auto-vectorizers frequently deliver suboptimal utilization of VNNI by either failing to recognize VNNI's unique computation pattern at the innermost loops/basic blocks, or producing inferior code through constrained and rudimentary peephole optimizations/pattern matching techniques. Auto-tuning frameworks might generate proficient code but are hampered by the necessity for sophisticated pattern templates and extensive search processes. Hao Zhou 0009, Qiukun Han, Heng Shi 0005, Yalin Zhang 0004, Jianguo Yao 0002 |
ASPLOS (3) | 3 |
| 2024 | SPHINX: Search Space-Pruning Heterogeneous Task Scheduling for Deep Neural NetworksabstractGiven the tendency of increasingly heterogeneous AI systems and the large workload scale of deep neural networks (DNNs), there is an urgent demand for model scheduling to improve execution performance in heterogeneous computational systems. However, this is very challenging because the task scheduling under the high-dimensional search space is an NP-hard problem. Existing works either schedule under naive search spaces without simplifications or oversimplifies the optimisation, which is hard to strike a balance between efficiency and optimality. Bowen Yuchi, Heng Shi 0005, Guoqing Bao |
ICPP | 2 |
| 2024 | UFront: Toward A Unified MLIR Frontend for Deep LearningabstractAutomatic code generation for ML systems has gained popularity with the advent of compiler techniques like Multi-Level Intermediate Representation (Multi-Level IR, or MLIR). State-of-the-art MLIR frontends, including IREE-TF, Torch-MLIR, and ONNX-MLIR, aim to bridge the gap between ML frameworks and low-level hardware architectures through MLIR's progressive lowering pipeline. However, existing MLIR frontends encounter challenges such as inflexible high-level IR conversion, limited higher-level optimization opportunities, and reduced compatibility and efficiency, leading to software fragmentation and restricting their practical applications within the MLIR ecosystem. To address these challenges, we introduce UFront, a unified MLIR frontend employing a two-stage operator-to-operator compilation workflow. Unlike traditional frontends that compile model source code into binaries step by step with different MLIR transform passes, UFront decouples the process into two distinct stages. It first performs instantaneous model tracing, delegates traced computing nodes as standard Deep Neural Network (DNN) operators and transforms models written in different frameworks into unified high-level IR without relying on MLIR passes, enhancing conversion flexibility. Meanwhile, it performs high-level graph optimizations such as constant folding and operator fusion to produce more efficient high-level IR. In the second stage, UFront directly converts high-level IR into standard TOSA IR using proposed lowering patterns, eliminating transform redundancies and ensuring lower-level compatibility with existing ML compiler backends. This two-stage compilation approach enables consistent end-to-end code generation and optimization of various DNN models written in different formats within a single workflow. Extensive experiments on popular DNN models written in various frameworks demonstrate that UFront exhibits higher compatibility, faster end-to-end compilation, and is capable of producing more efficient binary execution compared to SOTA works. Guoqing Bao, Heng Shi 0005, Chengyi Cui, Yalin Zhang 0004, Jianguo Yao 0002 |
ASE | 2 |