EDBT 2026 Demo / reviewers in the wild / expert
Yong Peng 0006
dblp:58/4461-6
· DBLP profile ↗
14ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0001-5803-7437ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperWeave: QoS-aware GPU Overcommitment for Deep Learning Training Jobs
Yong Peng 0006, Lujia Yin, Miao Zhang 0037 |
IWQoS | 3 |
| 2026 | Deep Model Fusion: A SurveyabstractDeep model fusion/merging is an emerging technique that integrates parameters or predictions from multiple deep learning (DL) models into a unified framework. It combines the abilities of different models to compensate for the biases and errors of an individual model, improving overall performance. However, deep model fusion, especially on large-scale DL models such as large language models (LLMs) and foundation models, faces several challenges, including high computational cost and interference between different heterogeneous models. In order to understand it better, we present a comprehensive survey to summarize the recent progress. We categorize existing model fusion methods as fourfold: 1) weight average (WA) averages the parameters of multiple models to obtain results closer to the optimal solution; 2) considering that direct averaging of models often yields suboptimal results, "mode connectivity" connects networks via paths of nonincreasing loss in weight spaces before the fusion. Along these paths, initial models are transformed into forms with consistent functions and better fusion effects; 3) similarly, for models with poor direct fusion results, "alignment" matches the corresponding units and merges these models, thus fully exploiting the corresponding relationships between the models; and 4) in addition to the above-mentioned methods of parameter fusion, "ensemble learning" fuses the outputs of multiple models in the inference stage to improve the accuracy and robustness of networks. In addition, we analyze the challenges of deep model fusion and illuminate the possible research directions in the future. Yong Peng 0006, Miao Zhang 0037, Liang Ding 0006, Han Hu 0003, Li Shen 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Hypernetwork Aggregation for Decentralized Personalized Federated LearningabstractPersonalized Federated Learning (PFL) meets each user’s personalized needs while still facing the high communication costs due to the large amount of data transmission and frequent communication. Decentralized PFL (DPFL) as an alternative discards the central server in PFL, which reduces the pressure of communication and the risk of server failure by using peer-to-peer communication.Nevertheless, DPFL still suffers from the significant communication pressure due to the transmission of a large number of model parameters, especially numerous nodes. To address the issues, we propose a novel personalized framework, DFedHP, in which each client utilizes a hypernetwork to generate the shared part of model parameters and train the personalized parameters separately. The number of parameters in a hypernetwork is much smaller than those in a typical local network, so hypernetwork aggregation reduces communication costs and the risk of privacy leakage. Furthermore, DFedHP can seamlessly integrate into existing DPFL algorithms as a plugin to boost their efficacy. At last, extensive experiments on various data heterogeneous environments demonstrate that DFedHP can reduce communication costs, accelerate convergence rate, and improve generalization performance compared with state-of-the-art (SOTA) baselines. Yong Peng 0006, Mengyao Du, Fuhui Sun, Li Shen 0008 |
IJCAI | 2 |
| 2025 | Multi-agent reinforcement learning for task offloading with hybrid decision space in multi-access edge computing
Miao Zhang 0037, Quanjun Yin, Lujia Yin, Yong Peng 0006 |
Ad Hoc Networks | 5 |
| 2025 | FlexHMB: A Flexible HMB Design Toward Bufferless Mobile FlashabstractSince the last decade, NAND flash has been widely adopted in mobile devices (e.g., smartphones) as the main storage media. Unlike enterprise solid-state drives and hard disk drives, mobile devices organize NAND flash in a DRAMless form, called mobile flash, which lacks an internal DRAM buffer to accommodate write requests due to the space constraints of mobile devices. Instead, it allocates Single-Level-Cell (SLC) flash blocks as a write buffer to accelerate write I/O bursts. However, we observe that this well-known design not only fails to consistently deliver high write performance but also compromises mobile flash capacity and longevity. The former is caused by the severe SLC reclamation interference, while the latter stems from its low density and the occupation of scarce over-provisioning blocks. To address these challenges, we propose FlexHMB, which utilizes the mobile flash controller to manage a high-performance and flexible write buffer allocated from the host-side main memory. Specifically, FlexHMB leverages the host memory buffer (HMB) feature to replace the traditional SLC buffer with a DRAM-based one. By doing so, FlexHMB shifts mobile flash towards bufferless architecture and eliminates penalties from the SLC buffer. To avoid the potential reliability issue arising from the volatility of DRAM, we design a write transaction mechanism to guarantee order consistency. While a large buffer can deliver higher performance, it may lead to competition with mobile applications for memory resources. Taking this into consideration, we design a flexible and dynamic resizing mechanism for the write buffer to make a balance between efficient request handling and memory utilization. The evaluation results show that FlexHMB reduces write latency by 93.54% compared to the SLC write buffer when memory usage is constrained. Yong Peng 0006, Shaocong Sun, Lujia Yin, Qiao Li 0001, Jie Zhang 0048 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | ACT: Action-assoCiated and Target-Related Representations for Object Navigation
Youkai Wang, Yue Hu 0016, Wansen Wu, Ting Liu 0018, Yong Peng 0006 |
MMM (1) | 5 |
| 2024 | Combine Intent Recognition with Behavior Modeling in Teaching Competition Military Simulation Platform
Shuilin Li, Chuan Ai, Yong Peng 0006 |
SIMULTECH | 4 |
| 2024 | Reinforcement learning from suboptimal demonstrations based on Reward Relabeling
Yong Peng 0006, Junjie Zeng 0004, Yue Hu 0016, Quanjun Yin |
Expert Syst. Appl. | 1 |
| 2024 | Visual Grounding With Dual Knowledge DistillationabstractVisual grounding is a task that seeks to predict the specific location of an object or region described by a linguistic expression within an image. Despite the recent success, existing methods still suffer from two problems. First, most methods use independently pre-trained unimodal feature encoders for extracting expressive feature embeddings, thus resulting in a significant semantic gap between unimodal embeddings and limiting the effective interaction of visual-linguistic contexts. Second, existing attention-based approaches equipped with the global receptive field have a tendency to neglect the local information present in the images. This limitation restricts the semantic understanding required to distinguish between referred objects and the background, consequently leading to inadequate localization performance. Inspired by the recent advance in knowledge distillation, in this paper, we propose a DUal knowlEdge disTillation (DUET) method for visual grounding models to bridge the cross-modal semantic gap and improve localization performance simultaneously. Specifically, we utilize the CLIP model as the teacher model to transfer the semantic knowledge to a student model, in which the vision and language modalities are linked into a unified embedding space. Besides, we design a self-distillation method for the student model to acquire localization knowledge by performing the region-level contrastive learning to make the predicted region close to the positive samples. To this end, this work further proposes a Semantics-Location Aware sampling mechanism to generate high-quality self-distillation samples. Extensive experiments on five datasets and ablation studies demonstrate the state-of-the-art performance of DUET and its orthogonality with different student models, thereby making DUET adaptable to a wide range of visual grounding architectures. Our code are available on DUET. Wansen Wu, Meng Cao 0002, Yue Hu 0016, Yong Peng 0006, Long Qin 0004, Quanjun Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | FedHiSyn: A Hierarchical Synchronous Federated Learning Framework for Resource and Data HeterogeneityabstractFederated Learning (FL) enables training a global model without sharing the decentralized raw data stored on multiple devices to protect data privacy. Due to the diverse capacity of the devices, FL frameworks struggle to tackle the problems of straggler effects and outdated models. In addition, the data heterogeneity incurs severe accuracy degradation of the global model in the FL training process. To address aforementioned issues, we propose a hierarchical synchronous FL framework, i.e., FedHiSyn. FedHiSyn first clusters all available devices into a small number of categories based on their computing capacity. After a certain interval of local training, the models trained in different categories are simultaneously uploaded to a central server. Within a single category, the devices communicate the local updated model weights to each other based on a ring topology. As the efficiency of training in the ring topology prefers devices with homogeneous resources, the classification based on the computing capacity mitigates the impact of straggler effects. Besides, the combination of the synchronous update of multiple categories and the device communication within a single category help address the data heterogeneity issue while achieving high accuracy. We evaluate the proposed framework based on MNIST, EMNIST, CIFAR10 and CIFAR100 datasets and diverse heterogeneous settings of devices. Experimental results show that FedHiSyn outperforms six baseline methods, e.g., FedAvg, SCAFFOLD, and FedAT, in terms of training accuracy and efficiency. Yue Hu 0016, Miao Zhang 0037, Ji Liu 0003, Quanjun Yin, Yong Peng 0006, Dejing Dou |
ICPP | 6 |
| 2022 | Efficient Flow-Based Scheduling for Geo-Distributed Simulation Tasks in Collaborative Edge and Cloud EnvironmentsabstractEdge computing is a good complement to cloud computing for deploying large-scale geo-distributed simulation applications, which are very sensitive to the communication delay among different simulation components (also called tasks in this paper) and users. We mainly focus on the efficient scheduling of simulation components in collaborative edge and cloud environments. As components should be deployed jointly with the consideration of capacity constraints of hosts, it is actually an NP-complete multi-dimensional bin packing problem. Meanwhile, dynamic changes of component and host states require the low deployment latency of scheduling algorithms. Unfortunately, most of the existing schedulers for modern clusters are queue-based, in which tasks are scheduled sequentially, thus lacking the ability to process tightly coupled tasks jointly. Other batching-based placement algorithms are usually time-consuming. This paper describes Pond, a novel flow-based scheduler with the awareness of interactions among tasks and users as well as heterogeneous multi-dimensional resources. First, characteristics of distributed simulation tasks are analysed and the scheduling problem is formulated as a min-cost max-flow (MCMF) problem over the flow network by mapping the communication overhead among tasks and users to the costs of arcs in the network. Considering the inherent defects of existing flow-based schedulers in dealing with multi-dimensional resources, a new method based on dominant resource is proposed and some problem specific heuristics are also designed. Extensive simulation experiments based on Alibaba production trace and some random synthetic parameters are conducted. Results show that Pond can reduce the average communication cost for each task significantly in a quite low deployment latency compared with some baselines. Miao Zhang 0037, Yong Peng 0006, Jiancheng Zhu, Quanjun Yin |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | A discrete PSO-based static load balancing algorithm for distributed simulations in a cloud environment
Miao Zhang 0037, Yong Peng 0006, Quanjun Yin, Xu Xie 0005 |
Future Gener. Comput. Syst. | 2 |
| 2017 | Learning real-time search on c-space GVDs
Quanjun Yin, Long Qin 0004, Yong Peng 0006 |
Frontiers Comput. Sci. | 3 |
| 2015 | Scheduling parallel jobs with tentative runs and consolidation in the cloud
Xiaocheng Liu, Yabing Zha, Quanjun Yin, Yong Peng 0006, Long Qin 0004 |
J. Syst. Softw. | 4 |