EDBT 2026 Demo / reviewers in the wild / expert
Jiaao He
dblp:249/2660
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-8578-5158ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DeltaFormer: Unlock the state space of TransformerabstractIn recent years, large language models with Transformer architecture as the core have made breakthrough progress in many fields. At the same time, there are also some weaknesses in the large language model that have prompted people to reflect, among which the most fundamental one is the reflection on the Transformer architecture. The Transformer architecture has high parallelism and can fully utilize the computing power of GPUs, thus replacing models such as LSTM in the past few years. However, high parallelism is not a free lunch, as it fundamentally limits the performance of models. Especially, the problems that logarithmic precision Transformer architecture can solve are strictly limited to the $TC^0$. And there are many important issues that are usually considered out of $TC^0$, such as Python code evaluation, entity tracking, chess, and other state tracking tasks. Meanwhile, some recent state space methods based on Delta Rule have been able to break through the $TC^0$ architecture, but they are limited by fixed size state spaces and perform poorly on many tasks. To this end, we have re-examined the Transformer from the perspective of a state space with kernel functions and propose an improved Transformer called DeltaFormer. We have theoretically and practically demonstrated that the proposed new architecture can break through the limitation of the inherent $TC^0$ expressivity of Transformers and verified that it is not weaker than standard Transformer in language modeling tasks. We hope our work can provide inspiration for designing more expressive models. Tenglong Ao, Jiaao He, Jianqiao Lu, Mingwu Zheng |
NeurIPS | 3 |
| 2025 | UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-TilingabstractLong-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication. Zan Zong, Yuyang Jin 0001, Kinman Lei, Jiaao He, Qigang Yang, Jidong Zhai |
SC | 5 |
| 2025 | HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender Models
Jiaao He, Shengqi Chen 0001, Kezhao Huang, Jidong Zhai |
USENIX ATC | 1 |
| 2025 | mTuner: Accelerating Parameter-Efficient Fine-Tuning on Multi-GPU Servers with Elastic Tensor
Kezhao Huang, Siqi Zhu, Mingshu Zhai, Liyan Zheng 0001, Kinman Lei, Jiaao He, Yuyang Jin 0001, Jidong Zhai |
USENIX ATC | 6 |
| 2024 | POSTER: Pattern-Aware Sparse Communication for Scalable Recommendation Model TrainingabstractRecommendation models are an important category of deep learning models whose size is growing enormous. They consist of a sparse part with TBs of memory footprint and a dense part that demands PFLOPs of computing capability to train. Unfortunately, the high sparse communication cost to re-organize data for different parallel strategies of the two parts impedes the scalability in training. Jiaao He, Shengqi Chen 0001, Jidong Zhai |
PPoPP | 1 |
| 2023 | SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, Jidong Zhai |
USENIX ATC | 2 |
| 2022 | Efficiently emulating high-bitwidth computation with low-bitwidth hardwareabstractDomain-Specific Accelerators (DSAs) are being rapidly developed to support high-performance domain-specific computation. Although DSAs provide massive computation capability, they often only support limited native data types. To mitigate this problem, previous works have explored software emulation for certain data types, which provides some compensation for hardware limitations. However, how to efficiently design more emulated data types and choose a high-performance one without hurting correctness or precision for a given application still remains an open problem. Zixuan Ma, Haojie Wang 0004, Guanyu Feng, Chen Zhang 0001, Jiaao He, Shengqi Chen 0001, Jidong Zhai |
ICS | 6 |
| 2022 | FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsabstractThe current trend in deep learning is to scale models to extremely large sizes with the objective of increasing their accuracy. Mixture-of-Expert (MoE) is the most popular pre-trained model that makes feasible the training of models with parameters beyond trillion-scale. Thanks to the dynamic activation of experts, i.e., shallow layers specialized in certain domains, it allows for sparse training of bigger models, removing the linearity between model size and computation. However, different from traditional deep learning models, it draws huge challenges to the efficiency of these training systems, including dynamic load imbalance, inefficient synchronous execution mode, and congested all-to-all communication. Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang 0004, Fuwen Luo, Shangfeng Shi |
PPoPP | 1 |
| 2022 | BaGuaLu: targeting brain scale pretrained models with over 37 million coresabstractLarge-scale pretrained AI models have shown state-of-the-art accuracy in a series of important applications. As the size of pretrained AI models grows dramatically each year in an effort to achieve higher accuracy, training such models requires massive computing and memory capabilities, which accelerates the convergence of AI and HPC. However, there are still gaps in deploying AI applications on HPC systems, which need application and system co-design based on specific hardware features. Zixuan Ma, Jiaao He, Jiezhong Qiu, Huanqi Cao, Yuanwei Wang, Zhenbo Sun, Liyan Zheng 0001, Haojie Wang 0004, Shizhi Tang, Tianyu Zheng, Junyang Lin, Guanyu Feng, Zeqiang Huang, Aohan Zeng, Jianwei Zhang 0012, Runxin Zhong, Tianhui Shi, Jie Tang 0001, Hongxia Yang, Xin Liu 0086, Jidong Zhai |
PPoPP | 2 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Tsinghua UniversityabstractIn this article we present our results from the SC19 Student Cluster Competition Reproducibility Challenge. The challenge entails reproducing the article entitled “Computing Planetary Interior Normal Modes with A Highly Parallel Polynomial Filtering Eigensolver” presented at SC'18, which proposes a parallel polynomial filtered Lanczos algorithm to directly calculate the planetary normal modes of heterogeneous planets. The proposed algorithm showed excellent performance with relatively low memory consumption and high parallel efficiency. In this work, we reproduce the scaling tests in that article on a cluster using Intel Cascade Lake architecture and use the proposed algorithm to illustrate specific normal modes of Mars. We compare the results obtained on our cluster with those in the original article. We also design a new metric to better analyze the results. In addition, we use the profiling tool Intel VTune Amplifier to explain our discoveries. Our results demonstrate that the given models show great scalability, which is similar to the original article. The required normal modes of Mars are also successfully calculated and visualized. Chen Zhang 0001, Chenggang Zhao, Jiaao He, Shengqi Chen 0001, Liyan Zheng 0001, Kezhao Huang, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Prague: High-Performance Heterogeneity-Aware Asynchronous Decentralized TrainingabstractDistributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. However, its performance is bounded by the slowest worker among all workers. For this reason, it is significantly slower in heterogeneous settings. AD-PSGD, a newly proposed synchronization method which provides numerically fast convergence and heterogeneity tolerance, suffers from deadlock issues and high synchronization overhead. Is it possible to get the best of both worlds --- designing a distributed training method that has both high performance like All-Reduce in homogeneous environment and good heterogeneity tolerance like AD-PSGD? Qinyi Luo, Jiaao He, Youwei Zhuo, Xuehai Qian |
ASPLOS | 2 |
| 2019 | Student Cluster Competition 2018, Team Tsinghua University: Reproducing performance of multi-physics simulations of the Tsunamigenic 2004 Sumatra megathrust earthquake on the Intel Skylake Architecture
Jiaao He, Chenggang Zhao, Jiping Yu, Xinjian Yu, Liyan Zheng 0001, Chenyao Lou, Shizhi Tang, Jidong Zhai |
Parallel Comput. | 1 |