EDBT 2026 Demo / reviewers in the wild / expert
Pengyu Mu
dblp:252/1581
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HybEdge: Explicit hybrid architecture for edge discontinuity detection
Jianhang Zhou, Long Xing, Pengyu Mu |
Expert Syst. Appl. | 4 |
| 2025 | Deep Learning Operators Performance Tuning for Changeable Sized Input Data on Tensor Accelerate HardwareabstractThe operator library is the fundamental infrastructure of deep learning acceleration hardware. Automatically generating the library and tuning its performance is promising because the manual development by well-trained and skillful programmers is costly in terms of both time and money. Tensor hardware has the best computing efficiency for deep learning applications, but the operator library programs are hard to tune because the tensor hardware primitives have many limitations. Otherwise, the performance is difficult to be fully explored. The recent advancement in LLM exacerbates this problem because the size of input data is not fixed. Therefore, mapping the computing tasks of operators to tensor hardware units is a significant challenge when the shape of the input tensor is unknown before the runtime. We propose DSAT, a deep learning operator performance autotuning technique for changeable-sized input data on tensor hardware. To match the input tensor's undetermined shape, we choose a group of abstract computing units as the basic building blocks of operators for changeable-sized input tensor shapes. We design a group of programming tuning rules to construct a large exploration space of the variant implementation of the operator programs. Based on these rules, we construct an intermediate representation of computing and memory access to describe the computing process and use it to map the abstract computing units to tensor primitives. To speed up the tuning process, we narrow down the optimization space by predicting the actual hardware resource requirement and providing an optimized cost model for performance prediction. DSAT achieves performance comparable to the vendor's manually tuned operator libraries. Compared to state-of-the-art deep learning compilers, it improves the performance of inference by 13% on average and decreases the tuning time by an order of magnitude. Pengyu Mu, Yi Liu 0013, Rui Wang 0014, Hangcheng An, Qianhe Zhao, Hailong Yang 0002, Chenhao Xie 0001, Zhongzhi Luan, Chunye Gong, Depei Qian 0001 |
IEEE Trans. Computers | 1 |
| 2024 | ViTa: Optimizing Vision Transformer Operators on Domain-Specific AcceleratorsabstractDomain-specific devices accelerate the inference of deep learning models, whose performance is sensitive to hardware resources. Large artificial intelligence models, especially those involving self-attention, have extensive computational demands that challenge efficient execution on accelerators. Due to restricted search space on resource-constrained hardware, some approaches require significant engineering effort to develop platform-specific optimization code or find suboptimal programs with search-based automatic compilers.This paper proposes ViTa, a framework for deploying vision Transformer models on resource-constrained hardware. First, ViTa adopts an analytical approach to modify the model structure, reducing computational load without sacrificing accuracy. Secondly, we divide computations into tiles and map them to computing units to enhance memory access and computational efficiency. Providing an abstract hardware layer to guide tiling and mapping can improve the model deployment efficiency. Experiments demonstrate that ViTa can reduce memory bandwidth by 43% and FLOPs by 56% while maintaining accuracy. Compared with the Vendor library, the inference time on various accelerators is reduced by an average of 24%. Pengyu Mu, Yi Liu 0013, Rui Wang 0014 |
ISPA | 1 |
| 2024 | No tricks no bluff, focusing on localizing crisp boundaries in image media
Jianhang Zhou, Pengyu Mu, Long Xing, Mingsi Sun |
Neurocomputing | 4 |
| 2023 | HAOTuner: A Hardware Adaptive Operator Auto-Tuner for Dynamic Shape Tensor CompilersabstractDeep learning compilers with auto-tuners have the ability to generate high-performance programs, particularly tensor programs on accelerators. However, the performance of these tensor programs is shape-sensitive and hardware resource-sensitive. When the tensor shape is only known at runtime instead of compile time, auto-tuners must tune the tensor programs for every possible shape, leading to significant time and cost overhead. Additionally, if a tensor program tuned for one device is deployed on a different device, the performance may not be as optimal as before. To address these challenges, we propose HAOTuner, a hardware-adaptive deep learning operator auto-tuner specifically designed for dynamic shape tensors. We leverage the concept of micro-kernels as the unit of task allocation and have observed that the size of the micro-kernel greatly impacts performance. In HAOTuner, we determine the size of micro-kernels based not only on the tensor shapes but also on the available hardware resources. Specifically, we present an algorithm to select hardware-friendly micro-kernels as candidates, reducing the tuning time. We also design a cost model that is sensitive to hardware resources to support various hardware architectures. Furthermore, we provide a model transfer solution to enable fast deployment of the cost model on different hardware platforms. We evaluate HAOTuner on six different types of GPUs. The experiments demonstrate that HAOTuner surpasses the state-of-the-art dynamic shape tensor auto-tuner in terms of running time by an average of 26% and tuning time by 25%. Moreover, HAOTuner outperforms the state-of-the-art compiler with padding in terms of running time by an average of 39% and tuning time by 6×. Pengyu Mu, Yi Liu 0013, Rui Wang 0014, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
IEEE Trans. Computers | 1 |