Shengyuan Ye

dblp:337/7674 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-8867-0655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 6 · 5 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Mu Yuan, Xiaowen Chu 0001, Weijie Hong, Xu Chen 0004
INFOCOM1
2026 Resource-Efficient Personal Large Language Models Fine-Tuning With Collaborative Edge Computing
abstract
Large language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. Other studies focus on exploiting the potential of edge devices through resource management optimization, yet are ultimately bottlenecked by the resource wall of individual devices. To tackle these challenges, we proposePAC+, a resource efficient collaborative edge AI framework for in-situ personal LLMs fine-tuning.PAC+breaks the resource wall of personal LLMs fine-tuning with a sophisticated algorithm-system co-design. (1) Algorithmically,PAC+implements a personal LLMs fine-tuning technique that is efficient in terms of parameters, time, and memory. It utilizes Parallel Adapters to circumvent the need for a full backward pass through the LLM backbone. Additionally, an activation cache mechanism further streamlining the process by negating the necessity for repeated forward passes across multiple epochs. (2) Systematically,PAC+leverages edge devices in close proximity, pooling them as a collective resource for in-situ personal LLMs fine-tuning, utilizing a hybrid data and pipeline parallelism to orchestrate distributed training. The use of the activation cache eliminates the need for forward pass through the LLM backbone, enabling exclusive fine-tuning of the Parallel Adapters using data parallelism. Extensive evaluation of the prototype implementation demonstrates thatPAC+significantly outperforms existing collaborative edge training systems, achieving up to a$9.7\times$end-to-end speedup. Furthermore, compared to mainstream LLM fine-tuning algorithms,PAC+reduces memory footprint by up to$88.16\%$.
Shengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng, Jiangsu Du, Xiaowen Chu 0001, Guoliang Xing, Xu Chen 0004
IEEE Trans. Parallel Distributed Syst.1
2025 Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
Shengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian, Xiaowen Chu 0001, Jian Tang 0008, Xu Chen 0004
INFOCOM1
2025 Grape: Efficient Spatiotemporal Prediction Services with Stale Sensing Streams
abstract
Emerging cyber-physical systems have embraced a large number of IoT devices spanning geo-distributed, which generate and consume massive volumes of data continuously. Accurate and timely spatiotemporal predictions (STP) over these streaming sensor data are critical and, in growing demand, ubiquitous across various edge scenarios such as traffic flow forecasting. Towards that, recent advanced systems have developed sophisticated optimizations among STP pipelines, aiming at optimal prediction performance. However, based on our empirical studies in real-world settings, we identify a previously overlooked bottleneck of end-to-end STP performance: data staleness. To mitigate this issue, in this work, we investigate a new task, namely stream interception, which deliberately terminates the acceptance of incoming sensor data and anticipates model execution with imputed missing features. We propose a novel dynamic interception strategy to determine the time slot to exit waiting and present Grape, an STP system that implements it with practical system designs. Extensive evaluations on real-world traces show that Grape can strike a superior tradeoff between prediction accuracy and serving latency, achieving 1.69-1.90× speedup against traditional all-waiting baselines across various STP services with high prediction accuracy on par with offline optimal cases.
Liekang Zeng, Shengyuan Ye, Mu Yuan, Di Duan, Xu Chen 0004, Guoliang Xing
RTSS3
2025 Revisiting Location Privacy in MEC-Enabled Computation Offloading
abstract
Mobile Edge Computing (MEC) revolutionizes real-time applications by extending cloud capabilities to network edges, enabling efficient computation offloading from mobile devices. In recent years, the location privacy concern within MEC offloading has been recognized, prompting the proposal of various methodologies to mitigate this concern. However, this paper demonstrates that the prevailing privacy protection methods exhibit vulnerabilities. First, we analyze the shortcomings of current methodologies through both system modeling and evaluation metrics. Then, we introduce a Learning-based Trajectory Reconstruction Attack (LTRA) to expose the weaknesses, achieving up to 91.2% reconstruction accuracy against the state-of-the-art protection method. Further, based onw-event differential privacy, we propose an ℓ-trajectory differentially private mechanism, i.e., OffloadingBD. Compared to the existing works, OffloadingBD provides more flexible and enhanced protection with sound privacy theoretical guarantee. Lastly, we conduct extensive experiments to evaluate LTRA and OffloadingBD. The experiment results show that LTRA has good generalization ability and OffloadingBD showcases a superior balance between privacy and utility compared with baselines.
Wenzhong Ou, Bei Ouyang, Shengyuan Ye, Liekang Zeng, Lin Chen 0002, Xu Chen 0004
IEEE Trans. Inf. Forensics Secur.4
2025 Resource-Efficient Collaborative Edge Transformer Inference With Hybrid Model Parallelism
abstract
Transformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users' privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and proposeGalaxy+, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration.Galaxy+introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity and memory-aware parallelism planning for fully exploiting the resource potential. To mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments,Galaxy+devises a tile-based fine-grained overlapping of communication and computation. Furthermore, a fault-tolerant re-scheduling mechanism is developed to address device-level resource dynamics, ensuring stable and low-latency inference. Extensive evaluation based on prototype implementation demonstrates thatGalaxy+remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving a$1.2\times$to$4.24\times$end-to-end latency reduction. Besides,Galaxy+can adapt to device-level resource dynamics, swiftly rescheduling and restoring inference in the presence of unexpected straggler devices.
Shengyuan Ye, Bei Ouyang, Jiangsu Du, Liekang Zeng, Tianyi Qian, Wenzhong Ou, Xiaowen Chu 0001, Deke Guo, Yutong Lu, Xu Chen 0004
IEEE Trans. Mob. Comput.1
2025 Co-Designing Transformer Architectures for Distributed Inference With Low Communication
abstract
Transformer models have shown significant success in a wide range of tasks. However, the massive resources required for its inference prevent deployment on a single device with relatively constrainted resources, thus leaving a high threshold of integrating their advancements. Observing scenarios such as smart home applications on edge devices and cloud deployment on commodity hardware, it is promising to distribute Transformer inference across multiple devices. Unfortunately, due to the tightly-coupled feature of Transformer model, existing model parallelism approaches necessitate frequent communication to resolve data dependencies, making them unacceptable for distributed inference, especially under relatively weak interconnection. In this paper, we propose DeTransformer, a communication-efficient distributed Transformer inference system. The key idea of DeTransformer involves the co-design of Transformer architecture to reduce the communication during distributed inference. In detail, DeTransformer is based on a novel block parallelism approach, which restructures the original Transformer layer with a single block to the decoupled layer with multiple sub-blocks. Thus, it can exploit model parallelism between sub-blocks. Next, DeTransformer contains an adaptive execution approach that strikes a trade-off among communication capability, computing power and memory budget over multiple devices. It incorporates a two-phase planning for execution, namely static planning and runtime planning. The static planning runs offline, containing a profiling procedure and a weight placement strategy before execution. The runtime planning dynamically determines the optimal parallel computing strategy from an expertly crafted search space based on real-time requests. Notably, this execution approach can adapt to heterogeneous devices by distributing workload based on devices’ computing capabilities. We conduct experiments for both auto-regressive and auto-encoder tasks of Transformer models. Experimental results show that DeTransformer can reduce distributed inference latency by up to 2.81× compared to the SOTA approach on 4 devices, while effectively maintaining task accuracy and a consistent model size.
Jiangsu Du, Yuanxin Wei, Shengyuan Ye, Jiazhi Jiang, Xu Chen 0004, Dan Huang 0001, Yutong Lu
IEEE Trans. Parallel Distributed Syst.3
2024 Communication-Efficient Model Parallelism for Distributed In-Situ Transformer Inference
abstract
Transformer models have shown significant success in a wide range of tasks. Meanwhile, massive resources required by its inference prevent scenarios with resource-constrained devices from in-situ deployment, leaving a high threshold of integrating its advances. Observing that these scenarios, e.g. smart home of edge computing, are usually comprise a rich set of trusted devices with untapped resources, it is promising to distribute Transformer inference onto multiple devices. However, due to the tightly-coupled feature of Transformer model, existing model parallelism approaches necessitate frequent communication to resolve data dependencies, making them unacceptable for distributed inference, especially under weak interconnect of edge scenarios. In this paper, we propose DeTransformer, a communication-efficient distributed in-situ Transformer inference system for edge scenarios. DeTransformer is based on a novel block parallelism approach, with the key idea of restructuring the original Trans-former layer with a single block to the decoupled layer with multi-ple sub-blocks and exploit model parallelism between sub-blocks. Next, DeTransformer contains an adaptive placement approach to automatically select the optimal placement strategy by striking a trade-off among communication capability, computing power and memory budget. Experimental results show that DeTransformer can reduce distributed inference latency by up to 2.81 x compared to the SOTA approach on 4 devices, while effectively maintaining task accuracy and a consistent model size.
Yuanxin Wei, Shengyuan Ye, Jiazhi Jiang, Xu Chen 0004, Dan Huang 0001, Jiangsu Du, Yutong Lu
DATE2
2024 Pluto and Charon: A Time and Memory Efficient Collaborative Edge AI Framework for Personal LLMs Fine-tuning
abstract
Large language models (LLMs) have unlocked a plethora of powerful applications at the network edge, such as intelligent personal assistants. Data privacy and security concerns have prompted a shift towards edge-based fine-tuning of personal LLMs, away from cloud reliance. However, this raises issues of computational intensity and resource scarcity, hindering training efficiency and feasibility. While current studies investigate parameter-efficient fine-tuning (PEFT) techniques to mitigate resource constraints, our analysis indicates that these techniques are not sufficiently resource-efficient for edge devices. Other studies focus on exploiting the potential of edge devices through resource management optimization, yet are ultimately bottlenecked by the resource wall of individual devices.
Bei Ouyang, Shengyuan Ye, Liekang Zeng, Tianyi Qian, Xu Chen 0004
ICPP2
2024 Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
abstract
Transformer-based models have unlocked a plethora of powerful intelligent applications at the edge, such as voice assistant in smart home. Traditional deployment approaches offload the inference workloads to the remote cloud server, which would induce substantial pressure on the backbone network as well as raise users’ privacy concerns. To address that, in-situ inference has been recently recognized for edge intelligence, but it still confronts significant challenges stemming from the conflict between intensive workloads and limited on-device computing resources. In this paper, we leverage our observation that many edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources and propose Galaxy, a collaborative edge AI system that breaks the resource walls across heterogeneous edge devices for efficient Transformer inference acceleration. Galaxy introduces a novel hybrid model parallelism to orchestrate collaborative inference, along with a heterogeneity-aware parallelism planning for fully exploiting the resource potential. Furthermore, Galaxy devises a tile-based fine-grained overlapping of communication and computation to mitigate the impact of tensor synchronizations on inference latency under bandwidth-constrained edge environments. Extensive evaluation based on prototype implementation demonstrates that Galaxy remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 2.5× end-to-end latency reduction.
Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu 0001, Yutong Lu, Xu Chen 0004
INFOCOM1
2024 Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
abstract
On-device Deep Neural Network (DNN) training has been recognized as crucial for privacy-preserving machine learning at the edge. However, the intensive training workload and limited onboard computing resources pose significant challenges to the availability and efficiency of model training. While existing works address these challenges through native resource management optimization, we instead leverage our observation that edge environments usually comprise a rich set of accompanying trusted edge devices with idle resources beyond a single terminal. We propose Asteroid, a distributed edge training system that breaks the resource walls across heterogeneous edge devices for efficient model training acceleration. Asteroid adopts a hybrid pipeline parallelism to orchestrate distributed training, along with a judicious parallelism planning for maximizing throughput under certain resource constraints. Furthermore, a fault-tolerant yet lightweight pipeline replay mechanism is developed to tame the device-level dynamics for training robustness and performance stability. We implement Asteroid on heterogeneous edge devices with both vision and language models, demonstrating up to 12.2× faster training than conventional parallelism methods and 2.1× faster than state-of-the-art hybrid parallelism methods through evaluations. Furthermore, Asteroid can recover training pipeline 14× faster than baseline methods while preserving comparable throughput despite unexpected device exiting and failure.
Shengyuan Ye, Liekang Zeng, Xiaowen Chu 0001, Guoliang Xing, Xu Chen 0004
MobiCom1
2024 MIX3D: A Mixed Representation for Communication-Efficient Distributed 3DGS Training
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a prominent technique in novel view synthesis. The superior performance of 3DGS has catalyzed an increasing number of 3DGS- based applications in edge scenarios, where 3DGS is utilized for various purposes, such as scene representation, comprehension, and generation. Meanwhile, these edge applications also serve as primary sources of scene observations for producing 3DGS models. However, the intensive computation involved in 3DGS training and the massive number of 3D Gaussian primitives required for high-resolution scene repre-sentation hinder the effectiveness of in-situ 3DGS training on off-the-shelf edge devices, whether using standalone training or Data-Distributed-Parallel (DDP) training. To address this issue, this work proposes MIX3D, a novel mixed representation for communication-efficient distributed 3DGS training in edge scenarios. MIX3D features a global sparse sub-model and various local dense sub-models, where the sparse sub-model encodes coarse-grained appearance for the entire scene, and each dense sub-model targets fine-grained details for a specific region of the scene. Extensive evaluations on a four-device edge cluster demonstrate the effectiveness of our developed distributed 3DGS training workflow based on MIX3D, achieving reductions in training time up to 86.6% compared to vanilla DDP training and an average speedup of 3.767x over standalone training.
Ke Luo 0001, Kongyange Zhao, Shengyuan Ye, Tao Ouyang, Xu Chen 0004
MSN3
2024 MEGA: Mesh-Aligned 3DGS Towards Geometry-Preserving Online Reconstruction
Ke Luo 0001, Shengyuan Ye, Tao Ouyang, Zhi Zhou 0006
NPC (1)2
2022 Eco-FL: Adaptive Federated Learning with Efficient Edge Collaborative Pipeline Training
abstract
Federated Learning (FL) has been a promising paradigm in distributed machine learning that enables in-situ model training and global model aggregation. While it can well preserve private data for end users, to apply it efficiently on IoT devices yet suffer from their inherent variants: their available computing resources are typically constrained, heterogeneous, and changing dynamically. Existing works deploy FL on IoT devices by pruning a sparse model or adopting a tiny counterpart, which alleviates the workload but may have negative impacts on model accuracy. To address these issues, we propose Eco-FL, a novel Edge Collaborative pipeline based Federated Learning framework. On the client side, each IoT device collaborates with trusted available devices in proximity to perform pipeline training, enabling local training acceleration with efficient augmented resource orchestration. On the server side, Eco-FL adopts a novel grouping-based hierarchical architecture that combines synchronous intra-group aggregation and asynchronous inter-group aggregation, where a heterogeneity-aware dynamic grouping strategy that jointly considers response latency and data distribution is developed. To tackle the resource fluctuation during the runtime, Eco-FL further applies an adaptive scheduling policy to judiciously adjust workload allocation and client grouping at different levels. Extensive experimental results using both prototype and simulation show that, compared to state-of-the-art methods, Eco-FL can upgrade the training accuracy by up to 26.3%, reduce the local training time by up to 61.5%, and improve the local training throughput by up to 2.6 ×.
Shengyuan Ye, Liekang Zeng, Qiong Wu 0009, Ke Luo 0001, Qingze Fang, Xu Chen 0004
ICPP1