EDBT 2026 Demo / reviewers in the wild / expert
Qifan Xu
dblp:164/9657
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unbalanced Vibration Suppression of Active Magnetic Bearing System based on Parameter IdentificationabstractThe suspension system with active magnetic bearing needs to solve the problem of rotor vibration caused by unbalanced interference. In this paper, the theoretical modeling and control methods are carried out for the multi-degree-of-freedom coupling effect and dynamic unbalanced vibration suppression problem of the active magnetic bearing-rigid rotating shaft system. Firstly, a four-degree-of-freedom dynamic model of a rigid rotor with unbalanced mass is established, and the dynamic decoupling of cross-coupling terms is realized by introducing the state feedback decoupling control strategy. Then, aiming at the dynamic unbalanced vibration caused by the unbalanced mass of the system, a current compensation algorithm based on parameter identification is studied to suppress the unbalanced vibration of the rotating shaft under the feedback decoupling framework. Both the simulation and experimental results verify the effectiveness of the decoupling of the system and the suppression of dynamic unbalanced vibration of the system. Yuanhao Du, Qifan Xu, Wenfei Yu, Wei Hua 0001 |
IECON | 2 |
| 2025 | MUsculo-Skeleton-Aware (MUSA) deep learning for anatomically guided head-and-neck CT deformable registration
Hengjie Liu, Elizabeth McKenzie, Di Xu 0003, Qifan Xu, Robert K. Chin, Dan Ruan, Ke Sheng |
Medical Image Anal. | 4 |
| 2023 | An Efficient 2D Method for Training Super-Large Deep Learning ModelsabstractSince the rise of Transformer [22] and BERT [6], large language models [7], [12] have been proposed and shown unprecedented performance in tasks like translation, classification, and text generation. However, due to the memory constraint, model parallelism must be used to split the model across multiple processors. Inter-layer partition, intra-layer partition, and sparse activation are the major approaches to achieve model parallelism. Among them, inter-layer partition [10], [11] often requires the model to be explicitly expressed as a stack of sub-modules, the number of which equals to the number of processors, and would introduce either gradient staleness or bubble overhead; while the sparse activation [12] is primarily designed for Google TPU cluster and hard to deploy on GPU servers, intra-layer partition [17], especially Megatron-LM [18], can be easily deployed on GPU servers and has been adopted in subsequent works like Turing-NLG and M6. Though as pioneers of intra-layer parallelism, they still show memory redundancy and sub-optimal communication efficiency, which reveals the space for further improvements. In this work, we leverage SUMMA [21] and propose Optimus, a highly efficient and scalable paradigm for training super-large language models. In Optimus, activations and gradients are partitioned and distributed along processors all the way through forward and backward propagations, with hardly any memory redundancy. The isoefficiency of communication in pure model parallelism improves from W ~ p3for Megatron-LM, to $W\sim {(\sqrt p \log p)^3}$ for our Optimus. This framework is implemented with open-source deep learning framework, PyTorch, and consolidates existing techniques such as mixed precision training [13], activation checkpointing [5], and data parallelism. In experiments on TACC Frontera supercomputers, Optimus shows 1.48× the speed for training, 1.78× speed for inference, and 8× the maximum batch size over Megatron-LM on 64 GPUs in pure model parallelism; and 1.73× speed for training, 2.32× speed for inference with data parallelism size equaling 2 on 128 GPUs. In pure model parallelism, Optimus surpasses Megatron-LM in weak scaling efficiency by a great margin, and shows an extraordinary increasing strong scaling efficiency. Optimus would facilitate the scaling of language models and serve as a strong thrust in the space exploration of artificial intelligence. Qifan Xu, Yang You 0001 |
IPDPS | 1 |
| 2022 | Tesseract: Parallelize the Tensor Parallelism EfficientlyabstractTogether with the improvements in state-of-the-art accuracies of various tasks, deep learning models are getting significantly larger. However, it is extremely difficult to implement these large models because limited GPU memory makes it impossible to fit large models into a single GPU or even a GPU server. Besides, it is highly necessary to reduce the training time for large models. Previous methods like Megatron-LM implemented a 1-Dimensional distributed method to use GPUs to speed up the training. However, these methods have a high communication overhead and a low scaling efficiency on large-scale clusters. To solve these problems, we propose Tesseract, highly scalable tensor parallelism with a novel design. It increases efficiency by reducing communication overhead and lowers the memory required for each GPU. By introducing the novel dimension into tensor parallelism, Tesseract greatly increases the memory capacity of tensor parallelism. Concretely, this new dimension furthermore increases the degree of tensor parallelism. Compared to previous 1-D and 2-D methods, Tesseract manages to reduce the communication cost on each layer, resulting in speedups of 1.38x and 1.53x respectively with strong scaling. In weak scaling experiments, Tesseract achieves a maximum of 4.0/1.7 times inference speedup and 3.4/1.7 times throughput improvement compared to 1-D/2-D methods, respectively. By introducing Tesseract, we offer a more efficient and scalable way to implement large deep learning models with limited GPU resources. Qifan Xu, Zhengda Bian, Yang You 0001 |
ICPP | 2 |