EDBT 2026 Demo / reviewers in the wild / expert
Yanqi Hao
dblp:29/9331
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Capricorn: Efficient In-Memory Checkpointing for MoE Model Training with Dynamicity AwarenessabstractMixture-of-Experts (MoE) has been extensively adopted for its incredible capability to expand model scale with a sub-linear increase in computational requirement. Training MoE models requires substantial computing nodes and extended periods, necessitating reliable distributed training systems. Checkpointing is a common approach to enhance training reliability by periodically saving model states. Current checkpointing optimizations focus on hiding checkpoint overhead in model training computations. However, these approaches overlook the dynamicity inherent in distributed MoE training, leading to an inefficient checkpointing mechanism. In this paper, we propose Capricorn, a dynamicity-aware in-memory checkpointing approach for efficient MoE model training. We observe that the dynamicity impacts computation durations at both the layer and iteration levels. At the layer level, different model layers exhibit various computation durations, while at the iteration level, the computation time of the same layer differs across iterations. To adapt to the layer-level dynamicity, Capricorn employs online profiling at the granularity of individual layers. Based on the profiling results, it strategically partitions checkpoints into chunks and schedules checkpointing communication to overlap with model computations. To deal with the dynamicity across iterations, Capricorn speculatively activates the profiling and partitioning processes utilizing the temporal locality of the experts' load. It can produce an optimal activation for low runtime overhead with high checkpoint partition accuracy. For mainstream MoE models, Capricorn achieves up to$1.56 \times$and$5.98 \times$end-to-end training speedup over Gemini and TorchSnapshot respectively under per-iteration checkpointing. Wenqian Xie, Zhiquan Lai, Yanqi Hao, Dongsheng Li 0001 |
CLUSTER | 6 |
| 2025 | Adversarial Transfer Learning-Based Hybrid Recurrent Network for Air Quality PredictionabstractAir quality modeling and forecasting has become a key problem in environmental protection. The existing prediction models typically require large‐scale and high‐quality historical data to achieve better performance. However, insufficient data volume and significant differences between data distribution across different regions will definitely reduce the effectiveness of the model reuse. To address the above issues, we propose a novel hybrid recurrent network based on domain adversarial transfer to achieve a stronger generalization ability when training air quality data from multisource domains. The proposed model mainly consists of three fundamental modules, i.e., feature extractor, regression predictor, and domain classifier. One‐dimensional convolutional neural networks (1D‐CNNs) are used to extract temporal feature of data from source and target stations. Bi‐directional gated recurrent unit (bi‐GRU) and bi‐directional long short‐term memory (bi‐LSTM) are utilized to learn temporal dependencies pattern of multivariate time series data. Two adversarial transfer strategies are employed to ensure that our model is capable of finding domain invariant representations automatically. Experiments with different number of source domains are conducted to demonstrate the effectiveness of the proposed domain transfer strategies. The experimental results also show that our composite model has superior performance for forecasting air quality in various regions. As further evidence, the adversarial training method could promote the positive transfer and alleviate the negative effect of irrelevant source data. Besides, our model exhibits preferable generalization capability as more robust prediction results are achieved on both unseen target domains and original source domains. Yanqi Hao, Chuan Luo 0001, Tianrui Li 0001, Junbo Zhang 0004, Hongmei Chen 0001 |
Int. J. Intell. Syst. | 1 |
| 2025 | Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model ParallelismabstractDeep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01–1.48 × over the fastest baseline, and speedups of up to 1.95 × over Megatron. Zhiquan Lai, Dongsheng Li 0001, Yanqi Hao, Ke-shi Ge, Xiaoge Deng, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | HSDP: Accelerating Large-scale Model Training via Efficient Sharded Data ParallelismabstractLarge deep neural network (DNN) models have demonstrated exceptional performance across diverse downstream tasks. Sharded data parallelism (SDP) has been widely used to reduce the memory footprint of model states. In a DNN training cluster, a device usually has multiple inter-device links that connect to other devices, like NVLink and InfiniBand. However, existing SDP approaches employ a single link at any given time, encountering challenges in efficient training due to significant communication overheads. We observe that the inter-device links can work independently without affecting each other. To reduce the fatal communication overhead of distributed training of large DNNs, this paper introduces HSDP, an efficient SDP training approach that enables the simultaneous utilization of multiple inter-device links. HSDP partitions models in a novel fine-grained manner and orchestrates the communication processes of partitioned parameters while considering inter-device links. This design enables concurrent communication execution and reduces communication overhead. To further optimize the training performance of HSDP, we propose a HSDP planner. The HSDP planner first abstracts the model partition and execution of HSDP into a communication parallel strategy, and builds a cost model to estimate the performance of each strategy. We then formulate the strategy searching as an optimization problem and solve it with an off-the-shelf solver. Evaluations on representative DNN workloads demonstrate that HSDP achieves up to 1.30× speedup compared to the state-of-the-art SDP training approaches. Yanqi Hao, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001 |
ISPA | 1 |
| 2011 | OrthoNets: simultaneous visual analysis of orthologs and their interaction neighborhoods across different organismsabstractMOTIVATION: Protein interaction networks contain a wealth of biological information, but their large size often hinders cross-organism comparisons. We present OrthoNets, a Cytoscape plugin that displays protein-protein interaction (PPI) networks from two organisms simultaneously, highlighting orthology relationships and aggregating several types of biomedical annotations. OrthoNets also allows PPI networks derived from experiments to be overlaid on networks extracted from public databases, supporting the identification and verification of new interactors. Any newly identified PPIs can be validated by checking whether their orthologs interact in another organism. AVAILABILITY: OrthoNets is freely available at http://wodaklab.org/orthonets/. Yanqi Hao, Anna Merkoulovitch, James Vlasblom, Shuye Pu, Andrei L. Turinsky, Denitza Roudeva, Brian Turner, Jack Greenblatt, Shoshana J. Wodak |
Bioinform. | 1 |