Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yan Zhuang 0004

dblp:02/5194-4 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0006-0702-5093ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 50% Transfer learning and domain adaptation · 25% Deep learning architectures and training · 25%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
mixture of experts
1.822026
Resource-Efficient LLM Customization on Mobile Devices Through Proxy Submodel Tuning · IEEE Trans. Netw. 2026
LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning · SenSys 2024
Machine learning › Transfer learning and domain adaptation › model adaptation
model customization
1.822026
Resource-Efficient LLM Customization on Mobile Devices Through Proxy Submodel Tuning · IEEE Trans. Netw. 2026
LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning · SenSys 2024
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-device learning
1.822026
Resource-Efficient LLM Customization on Mobile Devices Through Proxy Submodel Tuning · IEEE Trans. Netw. 2026
LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning · SenSys 2024
Machine learning › Efficient and distributed learning › model compression
submodel extraction
1.012026
Resource-Efficient LLM Customization on Mobile Devices Through Proxy Submodel Tuning · IEEE Trans. Netw. 2026
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
adapter tuning
0.812024
LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning · SenSys 2024

Methods — techniques the papers use, named apart from their topics

post-training submodel extraction · 1.8expert merging · 1.8adapter tuning · 1.0proxy submodel extraction · 0.8
YearPublicationVenuePosition
2026 Resource-Efficient LLM Customization on Mobile Devices Through Proxy Submodel Tuning
abstract
Considering limited on-device resources, current practices are attempting to deploy a system-level mixture-of-experts (MoE)-based foundation LLM on a mobile device to serve multiple apps and support mobile intelligence. However, mobile apps are hard to customize their services that require fine-tuning adapters associated with the LLM using private in-app data. The difficulty arises due to both the limited on-device resources and the restricted control that apps have over the foundation LLM. To address this issue, in this work, we propose LiteMoE, a novel proxy submodel tuning framework that supports mobile apps to efficiently fine-tune customized adapters on devices using proxy submodels. The key technique behind LiteMoE is a post-training submodel extraction method, whereby without additional retraining, we can identify and reserve critical experts, match and merge moderate experts, to extract a lightweight and effective proxy submodel from the foundation LLM for a specific app. To further enhance scalability and adaptability, LiteMoE incorporates adapter reuse and continuous tuning mechanisms to handle multi-task requirements and evolving user preferences. We implemented a prototype of LiteMoE and evaluated it over various MoE-based LLMs and mobile computing tasks. The results show that with LiteMoE, mobile apps are able to fine-tune customized adapters on resource-limited devices, achieving 12.7% accuracy improvement and 6.6× memory reduction compared with operating the original foundation LLM.
Yan Zhuang 0004, Chen Gong 0006, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen
IEEE Trans. Netw.1
2024 Nebula: An Edge-Cloud Collaborative Learning Framework for Dynamic Edge Environments
abstract
To bring the great power of modern DNNs into mobile computing and distributed systems, current practices primarily employ one of the two learning paradigms: cloud-based learning or on-device learning. Despite their distinct advantages, neither of these two paradigms could effectively deal with highly dynamic edge environments reflected in quick data distribution shifts and on-device resource fluctuations. In this work, we propose Nebula, an edge-cloud collaborative learning framework to enable rapid model adaptation for changing edge environments. To achieve this, we first propose a new block-level model decomposition scheme to decompose the large cloud model into multiple combinable modules. With this design, we can agilely derive personalized sub-models with compact sizes for edge devices, and quickly aggregate the updated sub-models to integrate new knowledge learned on the edge into the cloud model. We further propose an end-to-end learning framework that incorporates the modular model design into an efficient model adaptation pipeline, including an offline on-cloud model prototyping and training stage, and an online edge-cloud collaborative adaptation stage. Extensive experiments demonstrate that Nebula improves model performance (e.g., 18.89% accuracy increase) and resource efficiency (e.g., 7.12 × communication cost reduction) in adapting models to dynamic edge environments.
Yan Zhuang 0004, Zhenzhe Zheng 0001, Yunfeng Shao 0001, Bingshuai Li, Fan Wu 0006, Guihai Chen
ICPP1
2024 LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning
abstract
Considering limited on-device resources, current practices are attempting to deploy a system-level mixture-of-experts (MoE)-based foundation LLM shared by multiple mobile apps on a device to support mobile intelligence. However, mobile apps are hard to customize their services that require tuning adapters associated with the LLM using private in-app data. The difficulty arises due to both the limited on-device resources and the restricted control that apps have over the foundation LLM. To address this issue, in this work, we propose LiteMoE, a novel proxy submodel tuning framework that supports mobile apps to efficiently fine-tune customized adapters on devices using proxy submodels. The key technique behind LiteMoE is a post-training submodel extraction method, whereby without additional re-training, we can identify and reserve critical experts, match and merge moderate experts, to extract a lightweight and effective proxy submodel from the foundation LLM for a certain app. We implemented a prototype of LiteMoE and evaluated it over various MoE-based LLMs and mobile computing tasks. The results show that with LiteMoE, mobile apps are able to fine-tune customized adapters on resource-limited devices, achieving 12.7% accuracy improvement and 6.6× memory reduction compared with operating the original foundation LLM.
Yan Zhuang 0004, Zhenzhe Zheng 0001, Fan Wu 0006, Guihai Chen
SenSys1