Xuemei Peng

dblp:253/0754 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-2705-2877ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient MoE Inference on Single Consumer-grade GPU with Dynamic Expert Caching
Rui Zhang 0003, Boxuan Yang, Rongji Wang, Xuemei Peng, Zeyi Wen
IPDPS4
2025 KALE: Knowledge Aggregation for Label-free Model Enhancement
abstract
Large foundation models have demonstrated remarkable success in natural language processing and computer vision. Applying the large models to downstream tasks often requires fine-tuning, in order to boost the predictive accuracy. However, the fine-tuning process relies heavily on labeled data and extensive training. This dependency makes fine-tuning impractical for niche applications, such as rare object detection or specialized medical tasks. To overcome these limitations, we propose KALE: Knowledge Aggregation for Label-free model Enhancement, a label-free method for model enhancement, leveraging knowledge aggregation via model fusion and adaptive representation alignment. Our method is powered by a carefully designed joint self-cooperative optimization function that considers (i) multi-granularity optimization (task-specific and layer-specific), (ii) self and cooperative supervision integration, and (iii) mitigation of error accumulation caused by entropy minimization. Additionally, we introduce a class cardinality-aware sample filtering to ensure the stability of the fusion process. We also design a lightweight representation alignment technique to refine the fusion coefficient in a few shots for quality enhancement. We evaluate our method on multiple image classification datasets using ViT-B/32 and ViT-L/14 backbones. Experimental results demonstrate that our label-free method consistently outperforms state-of-the-art unsupervised approaches, including TURTLE and supervised full fine-tuning, in terms of average performance. Specifically, compared to TURTLE, our method achieves average improvements of 20.7% with ViT-B/32 and 19.5% with ViT-L/14. Furthermore, on the challenging SUN397 dataset, our method surpasses supervised full fine-tuning by 4% and 2.3% with ViT-B/32 and ViT-L/14, respectively.
Yuebin Xu, Xuemei Peng, Zeyi Wen
CIKM2
2025 Accelerating Multi-Output GBDTs with GPUs
abstract
Gradient Boosted Decision Trees (GBDTs) have demonstrated good performance in many data science competitions. This paper studies multidimensional output GBDT (GBDT-MO) training which can handle complex dependencies between input features and multidimensional outputs. Due to the multiple output dimensionality, the training time and memory consumption of GBDT-MO are significantly higher than those of the single-dimensional output GBDT training, which makes training GBDT-MO challenging, especially for large-scale datasets. In this paper, we propose a novel GPU-accelerated GBDT-MO training system to speed up the training while maintaining competitive model quality. Our system leverages the GPU to efficiently construct decision trees, especially by dynamically choosing efficient histogram building methods at different stages of the training, enabling scalable and efficient GBDT-MO training. Furthermore, our system supports multi-GPU training on a single machine by partitioning features across GPUs and using communication-efficient synchronization strategies, allowing further scalability for high-dimensional datasets. We evaluate the performance of our proposed GBDT-MO system on several datasets. Compared with CPU-based implementations, our system achieves a speedup ranging from 30 × to 190 ×. Moreover, our system outperforms the state-of-the-art GPU-based baselines by a speedup ranging from 1.7 × to 170 ×, while maintaining strong predictive performance compared with the state-of-the-art methods.
Hanfeng Liu, Xuemei Peng, Zeyi Wen
ICPP2
2024 APF-DQN: Adaptive Objective Pathfinding via Improved Deep Reinforcement Learning Among Building Fire Hazard
Qiuhan Xu, Xuemei Peng
ICANN (9)5
2023 Design and Blocking Analysis of Locking Protocols for Real-Time DAG Tasks Under Federated Scheduling
abstract
Real-time systems require locking protocols to coordinate access to shared resources. With the booming revolution of parallel processing technology in real-time systems, there has been some work addressing the problem of extending classic locking protocols for sequential real-time tasks to parallel tasks. However, it may not be most effective to trivially follow the progress mechanisms and queue orders designed for sequential tasks since the intrastructure information within a parallel task is not taken into consideration. This article investigates the design of locking protocols for parallel tasks using a novel mechanism—longest normal Section first (LNSF)—to consider the impact of normal sections on blocking behavior in parallel tasks and further improve real-time performance. LNSF is then implemented in a locking protocol for parallel tasks named POMIP, and associated blocking analysis techniques are presented. Empirical evaluations show that our proposed analysis dominated other state-of-the-art analysis—in best cases, the acceptance ratio of the task set can be improved by around 17%.
Yang Wang 0082, Xuemei Peng, Dong Ji, Nan Guan, Wang Yi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Task Allocation for Real-time Earth Observation Service with LEO Satellites
abstract
Traditional Earth observation (EO) services using satellites mainly observe relatively large-scale objects for applications with no or weak real-time requirements. The rapid development of Low-Earth-orbit (LEO) satellites opens new opportunities to provide EO services for a much wider range of applications by collecting the observation and communication capability of many LEO satellites. The challenge is how to select and coordinate the LEO satellites to accomplish the EO task subject to strong real-time constraints. In this work, we present a holistic solution that precisely models the observation service of a single LEO satellite and allocates the work of a periodic real-time EO task to a group of LEO satellites to meet the real-time requirements. Experiments were conducted to evaluate how the parameters of the LEO satellites and the ground stations impact the satisfiability of real-time requirements. The results provide valuable guidelines for designing LEO satellites and ground stations to provide real-time object observation services.
Mingsong Lv, Xuemei Peng, Nan Guan
RTSS2
2019 Response Time Analysis of Typed DAG Tasks for G-FP Scheduling
Xuemei Peng, Meiling Han, Qingxu Deng
SETTA1