Yuxuan Yan

dblp:225/8902 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang 0002, Ting Cao 0003, Yifan Yang 0004, Qianqian Yang 0002, Yuanchao Shu, Yuqing Yang 0001, Lili Qiu
NSDI1
2025 Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices
abstract
Large language models (LLMs) have emerged as a cornerstone for advancing AI technologies. It revolutionizes the way we interact with devices, websites, and information, and paves the way for the development of highly intuitive and capable virtual assistants. Training of today's LLMs happens in cloud data centers due to the requirement of enormous data and a significant amount of computing power. Despite extensive research in mobile edge computing, fine-tuning pre-trained LLMs using resource-constrained devices like commodity smartphones remains highly under-explored. In this paper, we propose Confidant, a practical collaborative training framework that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. To this end, Confidant partitions an LLM into several sub-models, allowing each of them to fit in the memory of a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. In specific, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. To ensure resilient distributed training, a hybrid fault tolerance mechanism is devised to proactively manage potential device and network failures. We fully implemented Confidant in C++/Python, and built a cross-framework adapter, enabling collaborative training on a variety of mobile platforms. Experimental results show that Confidant excels in achieving computation-, memory-efficient, and robust customization of LLMs - it manages to train state-of-the-art billion-sized LLMs including BERT, GPT-2, Phi2, and LLaMA3, and fine-tunes Phi2-2.7B on Alpaca in just 40.1 hours using three consumer-grade mobile devices.
Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Yuyang Qin, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu
MobiCom2
2025 Demo: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices
abstract
Despite large language models (LLMs) being an essential part of our lives, training of LLMs still needs to be done in cloud data centers due to the large requirements of data and computing power, leaving fine-tuning pre-trained LLMs on resource-constrained mobile devices remains highly under-explored. In this demo, we present Confidant, a practical collaborative training system that allows modern LLMs to be fine-tuned across multiple off-the-shelf mobile devices. Confidant partitions an LLM into several sub-models, deploying each of them to a mobile device. Multiple mobile devices then collaborate to train the LLM by employing a novel pipeline parallel training approach. Specifically, Confidant encompasses a memory-aware dynamic model partitioning and intra-device multi-processor scheduler to minimize the training time across heterogeneous platforms. A hybrid fault tolerance mechanism is also devised to proactively manage potential device and network failures. By building a cross-framework adapter and fully implementing Confidant on smartphones and laptops, we present the demo of collaborative training on a variety of mobile platforms.
Yuhao Chen 0005, Yuxuan Yan, Shuowei Ge, Qianqian Yang 0002, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001, Yuanchao Shu
MobiCom3
2025 EAR-Mapping: Edge-Assisted Real-Time Dense Mapping with Low Bandwidth Requirements
abstract
3D reconstruction plays a critical role in applications such as augmented reality (AR) and robotic systems. However, implicit neural representations (INRs), widely used in modern 3D reconstruction systems, demand substantial communication and computational resources, making the reconstruction process excessively slow and costly. In this paper, we introduce EAR-Mapping, a novel edge-assisted online 3D reconstruction framework designed for latency-sensitive mobile applications. EAR-Mapping incorporates an innovative sampling mechanism that seamlessly integrates explicit and implicit methods, enabling selective processing of camera data to maximize reconstruction performance. In addition, we utilize a value-based representation module to maximize computation resource efficiency. Finally, we design a framework that minimizes communication overhead through ROI-based data transmission. Our prototype implementation on a mobile-edge testbed demonstrates that EAR-Mapping achieves up to a 1.2x reduction in reconstruction latency and a 3.5x reduction in bandwidth usage, offering a significant advancement in the efficiency of 3D reconstruction for mobile-edge systems.
Yubin Dai, Bin Qian 0002, Yuxuan Yan, Minglei Zhao, Yangkun Liu, Yuanchao Shu
MobiSys3
2025 A Fine-grained Hemispheric Asymmetry Network for accurate and interpretable EEG-based emotion classification
abstract
In this work, we propose a Fine-grained Hemispheric Asymmetry Network (FG-HANet), an end-to-end deep learning model that leverages hemispheric asymmetry features within 2-Hz narrow frequency bands for accurate and interpretable emotion classification over raw EEG data. In particular, the FG-HANet extracts features not only from original inputs but also from their mirrored versions, and applies Finite Impulse Response (FIR) filters at a granularity as fine as 2-Hz to acquire fine-grained spectral information. Furthermore, to guarantee sufficient attention to hemispheric asymmetry features, we tailor a three-stage training pipeline for the FG-HANet to further boost its performance. We conduct extensive evaluations on two public datasets, SEED and SEED-IV, and experimental results well demonstrate the superior performance of the proposed FG-HANet, i.e. 97.11% and 85.70% accuracy, respectively, building a new state-of-the-art. Our results also reveal the hemispheric dominance under different emotional states and the hemisphere asymmetry within 2-Hz frequency bands in individuals. These not only align with previous findings in neuroscience but also provide new insights into underlying emotion generation mechanisms.
Ruofan Yan, Yuxuan Yan, Xu Niu, Jibin Wu
Neural Networks3
2025 Self-Correcting Clustering
abstract
The incorporation of target distribution significantly enhances the success of deep clustering. However, most of the related deep clustering methods suffer from two drawbacks: (1) manually-designed target distribution functions with uncertain performance and (2) cluster misassignment accumulation. To address these issues, aSelf-CorrectingClustering (Self-CC) framework is proposed. In Self-CC, a robust target distribution solver (RTDS) is designed to automatically predict the target distribution and alleviate the adverse influence of misassignments. Specifically, RTDS divides the high confidence samples selected according to the cluster assignments predicted by a clustering module into labeled samples with correct pseudo labels and unlabeled samples of possible misassignments by modeling its training loss distribution. With the divided data, RTDS can be trained in a semi-supervised way. The critical hyperparameter which controls the semi-supervised training process can be set adaptively by estimating the distribution property of misassignments in the pseudo-label space with the support of a theoretical analysis. The target distribution can be predicted by the well-trained RTDS automatically, optimizing the clustering module and correcting misassignments in the cluster assignments. The clustering module and RTDS mutually promote each other forming a positive feedback loop. Extensive experiments on four benchmark datasets demonstrate the effectiveness of the proposed Self-CC.
Hanxuan Wang, Zixuan Wang 0012, Yuxuan Yan, Gustavo Carneiro 0001, Zhen Wang 0004
IEEE Trans. Knowl. Data Eng.4
2024 Deep Online Probability Aggregation Clustering
Yuxuan Yan, Ruofan Yan
ECCV (53)1
2024 Latency-minimizing Semantic Communication with Dynamic Model Partitioning
abstract
Semantic communication is an emerging communication approach that aims to enhance efficient transmission by conveying the essential semantic meaning of the information while eliminating redundancy. In the current deep learning (DL)-based semantic communication systems, the encoder and decoder at the sender and receiver persist without modification after deployment, irrespective of variations in device computing power and channel bandwidth. This lack of adaptability may result in a decline in performance. To overcome this issue, we introduce an adaptive semantic communication approach aimed at minimizing end-to-end latency by leveraging a dynamic model partitioning mechanism. This mechanism dynamically splits the overall model into the encoder and decoder components, with the partitioning points adapting to changing communication and computing resources. Furthermore, we present a training method referred to as scheduled random partition point training to ensure that changes in the partitioning points do not adversely impact the performance of downstream tasks. Our experimental results affirm the effectiveness of these methods in terms of reducing latency and improving task performance.
Yuxuan Yan, Yuhao Chen 0005, Qianqian Yang 0002, Zhiguo Shi 0001
ICC1
2024 Prediction of human initial operation situation in confined space with a multi-task deep neural network
Mingyue Yin, Jianguang Li, Silu Wang, Yuxuan Yan
Eng. Appl. Artif. Intell.4
2024 AccEPT: An Acceleration Scheme for Speeding up Edge Pipeline-Parallel Training
abstract
It is usually infeasible to fit and train an entire large deep neural network (DNN) model using a single edge device due to the limited resources. To facilitate intelligent applications across edge devices, researchers have proposed partitioning a large model into several sub-models, and deploying each of them to a different edge device to collaboratively train a DNN model. However, the communication overhead caused by the large amount of data transmitted from one device to another during training, as well as the sub-optimal partition point due to the inaccurate latency prediction of computation at each edge device can significantly slow down training. In this paper, we propose AccEPT, an acceleration scheme for accelerating the edge collaborative pipeline-parallel training. In particular, we propose a light-weight adaptive latency predictor to accurately estimate the computation latency of each layer at different devices, which also adapts to unseen devices through continuous learning. Therefore, the proposed latency predictor leads to better model partitioning which balances the computation loads across participating devices. Moreover, we propose a bit-level computation-efficient data compression scheme to compress the data to be transmitted between devices during training. Our numerical results demonstrate that our proposed acceleration approach is able to significantly speed up edge pipeline parallel training up to 3 times faster in the considered experimental settings
Yuhao Chen 0005, Yuxuan Yan, Qianqian Yang 0002, Yuanchao Shu, Shibo He, Zhiguo Shi 0001, Jiming Chen 0001
IEEE Trans. Mob. Comput.2
2022 Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Peking University
abstract
Hidayetoluet al.(2019) proposed a novel memory-centric computation system, MemXCT. As a challenge at SC20, we reproduce the computational efficiency of MemXCT on our Azure cloud cluster. Our experiments evaluate the overall performance and the strong scalability with real datasets and verify part of the conclusions in the original article.
Zejia Fan, Zhewen Hao, Yueyang Pan, Pengcheng Xu 0005, Yuxuan Yan, Fangyuan Yang, Zhenxin Fu, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.6
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Peking University
abstract
Shi et al. (2018) proposed a highly parallel polynomial filtering eigensolver for the computation of planetary normal modes. As a challenge at the Student Cluster Competition in The International Conference for High Performance Computing, Networking, Storage and Analysis (SC19), we reproduce the computational efficiency of the polynomial filtering eigensolver on our Intel Xeon machine. We present the weak scalability, scaling of runtime with model size (in a fixed interval) and the strong scalability results in this report.
Yihua Cheng, Zejia Fan, Jing Mai, Yifan Wu 0005, Pengcheng Xu 0005, Yuxuan Yan, Zhenxin Fu, Yun Liang 0001
IEEE Trans. Parallel Distributed Syst.6
2019 Understanding and Detecting Overlay-based Android Malware at Market Scales
abstract
As a key UI feature of Android, overlay enables one app to draw over other apps by creating an extra View layer on top of the host View. While greatly facilitating user interactions with multiple apps at the same time, it is often exploited by malicious apps (malware) to attack users. To combat this threat, prior countermeasures concentrate on restricting the capabilities of overlays at the OS level, while barely seeing adoption by Android due to the concern of sacrificing overlays' usability. To address this dilemma, a more pragmatic approach is to enable the early detection of overlay-based malware at the app market level during the app review process, so that all the capabilities of overlays can stay unchanged. Unfortunately, little has been known about the feasibility and effectiveness of this approach for lack of understanding of malicious overlays in the wild. To fill this gap, in this paper we perform the first large-scale comparative study of overlay characteristics in benign and malicious apps using static and dynamic analyses. Our results reveal a set of suspicious overlay properties strongly correlated with the malice of apps, including several novel features. Guided by the study insights, we build OverlayChecker, a system that is able to automatically detect overlay-based malware at market scales. OverlayChecker has been adopted by one of the world's largest Android app stores to check around 10K newly submitted apps per day. It can efficiently (within 2 minutes per app) detect nearly all (96%) overlay-based malware using a single commodity server.
Yuxuan Yan, Zhenhua Li 0001, Qi Alfred Chen, Christo Wilson, Tianyin Xu, Ennan Zhai, Yong Li 0008, Yunhao Liu 0001
MobiSys1