EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Chen 0005
dblp:34/10195-5
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-9905-5823ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile DevicesabstractLarge language models (LLMs) have emerged as a cornerstone for advancing AI technologies. It revolutionizes the way we interact with devices, websites, and information, and paves the way for the development of highly intuitive and capable virtual assistants. Training of today's LLMs happens in cloud data centers due to the requirement of enormous data and a significant amount of computing power. Despite extensive research in mobile edge computing, fine-tuning pre-trained LLMs using resource-constrained devices like commodity smartphones remains highly under-explored. In this paper, we propose Confidant, a practical collaborative training framework that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. To this end, Confidant partitions an LLM into several sub-models, allowing each of them to fit in the memory of a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. In specific, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. To ensure resilient distributed training, a hybrid fault tolerance mechanism is devised to proactively manage potential device and network failures. We fully implemented Confidant in C++/Python, and built a cross-framework adapter, enabling collaborative training on a variety of mobile platforms. Experimental results show that Confidant excels in achieving computation-, memory-efficient, and robust customization of LLMs - it manages to train state-of-the-art billion-sized LLMs including BERT, GPT-2, Phi2, and LLaMA3, and fine-tunes Phi2-2.7B on Alpaca in just 40.1 hours using three consumer-grade mobile devices. Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Yuyang Qin, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu |
MobiCom | 1 |
| 2025 | Demo: Customizing Transformer-based LLMs via Collaborative Training on Mobile DevicesabstractDespite large language models (LLMs) being an essential part of our lives, training of LLMs still needs to be done in cloud data centers due to the large requirements of data and computing power, leaving fine-tuning pre-trained LLMs on resource-constrained mobile devices remains highly under-explored. In this demo, we present Confidant, a practical collaborative training system that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. Confidant partitions an LLM into several sub-models, deploying each of them to a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. Specifically, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. A hybrid fault tolerance mechanism is also devised to proactively manage potential device and network failures. By building a cross-framework adapter and fully implementing Confidant on smartphones and laptops, we present the demo of collaborative training on a variety of mobile platforms. Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu |
MobiCom | 1 |
| 2024 | Latency-minimizing Semantic Communication with Dynamic Model PartitioningabstractSemantic communication is an emerging communication approach that aims to enhance efficient transmission by conveying the essential semantic meaning of the information while eliminating redundancy. In the current deep learning (DL)-based semantic communication systems, the encoder and decoder at the sender and receiver persist without modification after deployment, irrespective of variations in device computing power and channel bandwidth. This lack of adaptability may result in a decline in performance. To overcome this issue, we introduce an adaptive semantic communication approach aimed at minimizing end-to-end latency by leveraging a dynamic model partitioning mechanism. This mechanism dynamically splits the overall model into the encoder and decoder components, with the partitioning points adapting to changing communication and computing resources. Furthermore, we present a training method referred to as scheduled random partition point training to ensure that changes in the partitioning points do not adversely impact the performance of downstream tasks. Our experimental results affirm the effectiveness of these methods in terms of reducing latency and improving task performance. Yuxuan Yan, Yuhao Chen 0005, Qianqian Yang 0002, Zhiguo Shi 0001 |
ICC | 2 |
| 2024 | FTPipeHD: A Fault-Tolerant Pipeline-Parallel Distributed Training Approach for Heterogeneous Edge DevicesabstractWith the increasing proliferation of Internet-of-Things (IoT) devices, there is a growing trend towards distributing the power of deep learning (DL) among edge devices rather than centralizing it at the cloud. To deploy deep and complex models at edge devices with limited resources, model partitioning of deep neural network (DNN) models has been widely studied. However, most of the existing literature only considers distributing the inference model while still training the model at the cloud. In this paper, we propose FTPipeHD, a novel DNN training approach that trains DNN models across distributed heterogeneous devices with the fault-tolerance mechanism. To accelerate the training with the time-varying computing power of each device, we optimize the partition points dynamically according to real-time computing capacities. We also propose a novel weight redistribution approach that replicates the weights to both the neighboring nodes and the central node periodically, which combats the failure of multiple devices during training while incurring limited communication costs. Our numerical results demonstrate that FTPipeHD is 6.8 times faster in training than the state-of-the-art method when the computing capacity of the best device is 10 times greater than the worst one. It is also shown that the proposed method is able to accelerate the training even with the existence of device failures. Yuhao Chen 0005, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Mohsen Guizani |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | AccEPT: An Acceleration Scheme for Speeding up Edge Pipeline-Parallel TrainingabstractIt is usually infeasible to fit and train an entire large deep neural network (DNN) model using a single edge device due to the limited resources. To facilitate intelligent applications across edge devices, researchers have proposed partitioning a large model into several sub-models, and deploying each of them to a different edge device to collaboratively train a DNN model. However, the communication overhead caused by the large amount of data transmitted from one device to another during training, as well as the sub-optimal partition point due to the inaccurate latency prediction of computation at each edge device can significantly slow down training. In this paper, we propose AccEPT, an acceleration scheme for accelerating the edge collaborative pipeline-parallel training. In particular, we propose a light-weight adaptive latency predictor to accurately estimate the computation latency of each layer at different devices, which also adapts to unseen devices through continuous learning. Therefore, the proposed latency predictor leads to better model partitioning which balances the computation loads across participating devices. Moreover, we propose a bit-level computation-efficient data compression scheme to compress the data to be transmitted between devices during training. Our numerical results demonstrate that our proposed acceleration approach is able to significantly speed up edge pipeline parallel training up to 3 times faster in the considered experimental settings Yuhao Chen 0005, Yuxuan Yan, Qianqian Yang 0002, Yuanchao Shu, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Deep Joint Source-Channel Coding for Wireless Image Transmission with Entropy-Aware Adaptive Rate ControlabstractAdaptive rate control for deep joint source and channel coding (JSCC) is considered as an effective approach to transmit sufficient information in scenarios with limited communication resources. We propose a deep JSCC scheme for wireless image transmission with entropy-aware adaptive rate control, using a single deep neural network to support multiple rates and automatically adjust the rate based on the feature maps of the input image and their entropy, as well as the channel conditions. In particular, we maximize the entropy of the feature maps to increase the average information carried by each transmitted symbol during the training. We further decide which feature maps should be activated based on their entropy, which improves the efficiency of the transmitted symbols. We also propose a pruning module to remove less important pixels in the activated feature maps in order to further improve transmission efficiency. The experimental results demonstrate that our proposed scheme learns an effective rate control strategy that reduces the required channel bandwidth while preserving the quality of the reconstructed images. Weixuan 'Vincent' Chen, Yuhao Chen 0005, Qianqian Yang 0002, Chongwen Huang, Qian Wang 0030, Zhaoyang Zhang 0001 |
GLOBECOM | 2 |
| 2023 | The Model Inversion Eavesdropping Attack in Semantic Communication SystemsabstractIn recent years, semantic communication has been a popular research topic for its superiority in communication efficiency. As semantic communication relies on deep learning to extract meaning from raw messages, it is vulnerable to attacks targeting deep learning models. In this paper, we introduce the model inversion eavesdropping attack (MIEA) to reveal the risk of privacy leaks in the semantic communication system. In MIEA, the attacker first eavesdrops the signal being transmitted by the semantic communication system and then performs model inversion attack to reconstruct the raw message, where both the white-box and black-box settings are considered. Evaluation results show that MIEA can successfully reconstruct the raw message with good quality under different channel conditions. We then propose a defense method based on random permutation and substitution to defend against MIEA in order to achieve secure semantic communication. Our experimental results demonstrate the effectiveness of the proposed defense method in preventing MIEA. Yuhao Chen 0005, Qianqian Yang 0002, Zhiguo Shi 0001, Jiming Chen 0001 |
GLOBECOM | 1 |
| 2019 | iLoc: A Low-Cost Low-Power Outdoor Localization System for Internet of ThingsabstractNode location information is very important to many novel applications of Internet of Things (IoT). Typically, IoT nodes are resource-constrained, and thus costly and energy-hungry localization techniques fall short. In this paper, we present iLoc, a low-cost, low-power and wide-area localization system for IoT applications. iLoc is built on the emerging LoRa technology and overcomes the disadvantage of many short-range localization techniques. Central to iLoc is a mobile anchor node comprising of a simplified LoRa gateway and a smartphone. To locate an IoT node, the anchor node moves around, during which the LoRa gateway receives its locations from the smartphone, and communicates with the IoT node for the information of time of flight (ToF) as well as received signal strength indication (RSSI). In order to obtain a better distance estimation, both RSSI and ToF are integrated in the regression analysis of distance between the anchor node and the IoT node. We further design an iterative localization algorithm by judiciously deciding the locations of the anchor node step by step. The LoRa gateway and tag we prototype cost less than 10 and 5 dollars, respectively. We conduct extensive experiments and the results demonstrate that iLoc achieves an average localization error of 1.33m and power consumption of 0.25mAh in an open environment. Yuhao Chen 0005, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001 |
GLOBECOM | 2 |