Chunlin Tian

dblp:194/2903 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
abstract
Federated fine-tuning enables privacypreserving LLM adaptation but faces a critical bottleneck: the disparity between LLMs' high memory demands and edge devices' limited capacity.To break the memory barrier, we propose Chain Federated Fine-Tuning (CHAINFED), an innovative paradigm that forgoes end-to-end updates in favor of a sequential, layer-by-layer manner.It first trains the initial adapter to convergence, freezes its weights, and then proceeds to the next.This iterative train-and-freeze process forms an optimization chain, gradually enhancing the model's task-specific proficiency.CHAINFED further integrates three core techniques: 1) Dynamic Layer Co-Tuning to bridge semantic gaps between sequentially tuned layers and facilitate information flow; 2) Globally Perceptive Optimization to endow each adapter with foresight beyond its local objective; 3) Function-Oriented Adaptive Tuning to automatically identify the optimal fine-tuning starting point.Extensive experiments on multiple benchmarks demonstrate the superiority of CHAINFED over existing methods, boosting average accuracy by up to 46.46%.
Yebo Wu, Jingguang Li, Chunlin Tian, Kahou Tam, Zhijiang Guo, Li Li 0064
ACL (1)3
2026 Floe: Federated Specialization for Real-Time LLM-SLM Inference
abstract
Deploying large language models (LLMs) in realtime systems is challenging due to their high resource demands and privacy concerns. We propose Floe a hybrid federated learning framework designed for latency-sensitive, resourceconstrained environments. Floeombines a cloud-based blackbox LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy ensures efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Experiments demonstrate that Floenhances user privacy and personalization, while significantly improving model performance and reducing inference latency on edge devices under real-time constraints, compared to baseline approaches.
Chunlin Tian, Kahou Tam, Yebo Wu, Shuaihang Zhong, Li Li 0064, Nicholas D. Lane, Cheng-Zhong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2025 CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Chunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 0064, Yuanzhe Zhao, Minglei Zhang, Cheng-Zhong Xu 0001
USENIX ATC1
2024 Less is More: Efficient Brain-Inspired Learning for Autonomous Driving Trajectory Prediction
abstract
Accurately and safely predicting the trajectories of surrounding vehicles is essential for fully realizing autonomous driving (AD). This paper presents the Human-Like Trajectory Prediction model (HLTP++), which emulates human cognitive processes to improve trajectory prediction in AD. HLTP++ incorporates a novel teacher-student knowledge distillation framework. The “teacher” model, equipped with an adaptive visual sector, mimics the dynamic allocation of attention human drivers exhibit based on factors like spatial orientation, proximity, and driving speed. On the other hand, the “student” model focuses on real-time interaction and human decision-making, drawing parallels to the human memory storage mechanism. Furthermore, we improve the model’s efficiency by introducing a new Fourier Adaptive Spike Neural Network (FA-SNN), allowing for faster and more precise predictions with fewer parameters. Evaluated using the NGSIM, HighD, and MoCAD benchmarks, HLTP++ demonstrates superior performance compared to existing models, which reduces the predicted trajectory error with over 11% on the NGSIM dataset and 25% on the HighD datasets. Moreover, HLTP++ demonstrates strong adaptability in challenging environments with incomplete input data. This marks a significant stride in the journey towards fully AD systems.
Haicheng Liao, Yongkang Li 0003, Zhenning Li 0001, Chengyue Wang 0001, Guofa Li, Chunlin Tian, Zilin Bian, Kaiqun Zhu, Zhiyong Cui, Jia Hu 0003
ECAI6
2024 Ranking-based Client Imitation Selection for Efficient Federated Learning
abstract
Federated Learning (FL) enables multiple devices to collaboratively train a shared model while ensuring data privacy. The selection of participating devices in each training round critically affects both the model performance and training efficiency, especially given the vast heterogeneity in training capabilities and data distribution across devices. To deal with these challenges, we introduce a novel device selection solution called FedRank, which is based on an end-to-end, ranking-based model that is pre-trained by imitation learning against state-of-the-art analytical approaches. It not only considers data and system heterogeneity at runtime but also adaptively and efficiently chooses the most suitable clients for model training. Specifically, FedRank views client selection in FL as a ranking problem and employs a pairwise training strategy for the smart selection process. Additionally, an imitation learning-based approach is designed to counteract the cold-start issues often seen in state-of-the-art learning-based approaches. Experimental results reveal that FedRank boosts model accuracy by 5.2% to 56.9%, accelerates the training convergence up to $2.01 \times$ and saves the energy consumption up to 40.1%.
Chunlin Tian, Xinpeng Qin, Li Li 0064, Cheng-Zhong Xu 0001
ICML1
2024 FedGCS: A Generative Framework for Efficient Client Selection in Federated Learning via Gradient-based Optimization
Zhiyuan Ning 0001, Chunlin Tian, Meng Xiao 0001, Wei Fan 0010, Pengyang Wang, Li Li 0064, Pengfei Wang 0008, Yuanchun Zhou
IJCAI2
2024 FedMG: A Federated Multi-Global Optimization Framework for Autonomous Driving Control
abstract
Control is a critical module of autonomous driving systems, which ensures safety and enhances the human-machine interface. Due to the diverse control demands dictated by different driving scenarios, autonomous vehicles require a data-intensive, adaptive, and intelligent controller. To speed up the control process and improve the performance in different scenarios, we introduce a novelty federated learning framework FedMG, which efficiently coordinates diverse vehicles to train a collaboratively models while preserving data privacy to tune the control process. Through detailed analysis of driving scenarios, vehicles are clustered to different groups based on driving scenarios to seek a balance between data quality and communication efficiency. It enables the consolidation of several global models, each optimized for peak performance, thereby enhancing the overall system’s effectiveness. Extensive experiments with different numbers of vehicles and a variety of driving scenarios demonstrate the effectiveness of FedMG. The framework significantly reduces cumulative driving errors, achieving reductions ranging from 5.42% to 76.43%, while improving user comfort, with improvements ranging from 2.23% to 34.61% over baselines.
Jialiang Ma, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001
IWQoS2
2024 GreenLLM: Towards Efficient Large Language Model via Energy-aware Pruning
abstract
This paper proposes GreenLLM, a framework that effectively deploys generative Large Language Models (LLMs) on resource-limited edge devices to well meet the memory and timing constraints with minimized energy consumption. Specifically, GreenLLM employs an energy estimation scheme based on physical hardware to guide a pruning-ratio generator incorporating space, weight, and power (SWaP) constraints for optimal pruning ratio. For each layer, we employ a dependency-aware energy-efficient Pruner in a task-agnostic manner, maximally preserving most of the LLM functionality. Finally, we use downstream datasets to fine-tune the pruned model to recover performance.
Chunlin Tian, Xinpeng Qin, Li Li 0064
IWQoS1
2024 Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
abstract
Federated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, intensive memory footprint during the training process severely bottlenecks the deployment of FL on resource-limited mobile devices in real-world cases. Thus, a framework that can effectively reduce the memory footprint while guaranteeing training efficiency and model accuracy is crucial for FL.In this paper, we propose SmartFreeze, a framework that effectively reduces the memory footprint by conducting the training in a progressive manner. Instead of updating the full model in each training round, SmartFreeze divides the shared model into blocks consisting of a specified number of layers. It first trains the front block with a well-designed output module, safely freezes it after convergence, and then triggers the training of the next one. This process iterates until the whole model has been successfully trained. In this way, the backward computation of the frozen blocks and the corresponding memory space for storing the intermediate outputs and gradients are effectively saved. Except for the progressive training framework, SmartFreeze consists of the following two core components: a pace controller and a participant selector. The pace controller is designed to effectively monitor the training progress of each block at runtime and safely freezes them after convergence while the participant selector selects the right devices to participate in the training for each block by jointly considering the memory capacity, the statistical and system heterogeneity. Extensive experiments are conducted to evaluate the effectiveness of SmartFreeze on both simulation and hardware testbeds. The results demonstrate that SmartFreeze effectively reduces average memory usage by up to 82%. Moreover, it simultaneously improves the model accuracy by up to 83.1% and accelerates the training process up to 2.02 ×.
Yebo Wu, Li Li 0064, Chunlin Tian, Chang Tao, Wang Cong, Cheng-Zhong Xu 0001
IWQoS3
2024 Heterogeneity-Aware Coordination for Federated Learning via Stitching Pre-trained blocks
abstract
Federated learning (FL) coordinates multiple devices to collaboratively train a shared model while preserving data privacy. However, large memory footprint and high energy consumption during the training process excludes the low-end devices from contributing to the global model with their own data, which severely deteriorates the model performance in real-world scenarios. In this paper, we propose FedStitch, a hierarchical coordination framework for heterogeneous federated learning with pre-trained blocks. Unlike the traditional approaches that train the global model from scratch, for a new task, FedStitch composes the global model via stitching pre-trained blocks. Specifically, each participating client selects the most suitable block based on their local data from the candidate pool composed of blocks from pre-trained models. The server then aggregates the optimal block for stitching. This process iterates until a new stitched network is generated. Except for the new training paradigm, FedStitch consists of the following three core components: 1) an RL-weighted aggregator, and 2) a search space optimizer deployed on the server side, and 3) a local energy optimizer deployed on each participating client. The RL-weighted aggregator helps to select the right block in the non-IID scenario, while the search space optimizer continuously reduces the size of the candidate block pool during stitching. Meanwhile, the local energy optimizer is designed to minimize the energy consumption of each client while guaranteeing the overall training progress. The results demonstrate that compared to existing approaches, FedStitch improves the model accuracy up to 20.93%. At the same time, it achieves up to 8.12× speedup, reduces the memory footprint up to 79.5%, and achieves 89.41% energy saving at most during the learning procedure.
Shichen Zhan, Yebo Wu, Chunlin Tian, Li Li 0064
IWQoS3
2024 When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language Models
abstract
As autonomous driving systems increasingly become part of daily transportation, the ability to accurately anticipate and mitigate potential traffic accidents is paramount. Traditional accident anticipation models primarily utilizing dashcam videos are adept at predicting when an accident may occur but fall short in localizing the incident and identifying involved entities. Addressing this gap, this study introduces a novel framework that integrates Large Language Models (LLMs) to enhance predictive capabilities across multiple dimensions-what, when, and where accidents might occur. We develop an innovative chain-based attention mechanism that dynamically adjusts to prioritize high-risk elements within complex driving scenes. This mechanism is complemented by a three-stage model that processes outputs from smaller models into detailed multimodal inputs for LLMs, thus enabling a more nuanced understanding of traffic dynamics. Empirical validation on the DAD, CCD, and A3D datasets demonstrates superior performance in Average Precision (AP) and Mean Time-To-Accident (mTTA), establishing new benchmarks for accident prediction technology. Our approach not only advances the technological framework for autonomous driving safety but also enhances human-AI interaction, making predictive insights generated by autonomous systems more intuitive and actionable.
Haicheng Liao, Yongkang Li 0003, Chengyue Wang 0001, Yanchen Guan, Kahou Tam, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001
ACM Multimedia6
2024 CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus Attentions
abstract
Accurately and promptly predicting accidents among surrounding traffic agents from camera footage is crucial for the safety of autonomous vehicles (AVs). This task presents substantial challenges stemming from the unpredictable nature of traffic accidents, their long-tail distribution, the intricacies of traffic scene dynamics, and the inherently constrained field of vision of onboard cameras. To address these challenges, this study introduces a novel accident anticipation framework for AVs, termed CRASH. It seamlessly integrates five components: object detector, feature extractor, object-aware module, context-aware module, and multi-layer fusion. Specifically, we develop the object-aware module to prioritize high-risk objects in complex and ambiguous environments by calculating the spatial-temporal relationships between traffic agents. In parallel, the context-aware is also devised to extend global visual information from the temporal to the frequency domain using the Fast Fourier Transform (FFT) and capture fine-grained visual features of potential objects and broader context cues within traffic scenes. To capture a wider range of visual cues, we further propose a multi-layer fusion that dynamically computes the temporal dependencies between different scenes and iteratively updates the correlations between different visual features for accurate and timely accident prediction. Evaluated on real-world datasets-Dashcam Accident Dataset (DAD), Car Crash Dataset (CCD), and AnAn Accident Detection (A3D) datasets-our model surpasses existing top baselines in critical evaluation metrics like Average Precision (AP) and mean Time-To-Accident (mTTA). Importantly, its robustness and adaptability are particularly evident in challenging driving scenarios with missing or limited training data, demonstrating significant potential for application in real-world autonomous driving systems.
Haicheng Liao, Huanming Shen, Chengyue Wang 0001, Chunlin Tian, Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001
ACM Multimedia5
2024 HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning
abstract
Adapting Large Language Models (LLMs) to new tasks through fine-tuning has been made more efficient by the introduction of Parameter-Efficient Fine-Tuning (PEFT) techniques, such as LoRA. However, these methods often underperform compared to full fine-tuning, particularly in scenarios involving complex datasets. This issue becomes even more pronounced in complex domains, highlighting the need for improved PEFT approaches that can achieve better performance. Through a series of experiments, we have uncovered two critical insights that shed light on the training and parameter inefficiency of LoRA. Building on these insights, we have developed HydraLoRA, a LoRA framework with an asymmetric structure that eliminates the need for domain expertise. Our experiments demonstrate that HydraLoRA outperforms other PEFT approaches, even those that rely on domain knowledge during the training and inference phases. Our anonymous codes are submitted with the paper and will be publicly available. Code is available: https://github.com/Clin0212/HydraLoRA.
Chunlin Tian, Zhijiang Guo, Li Li 0064, Cheng-Zhong Xu 0001
NeurIPS1
2024 FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management
abstract
Federated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, one fundamental and prevailing challenge that hinders the deployment of FL on mobile devices is the memory limitation. This paper proposes FedHybrid, a novel framework that effectively reduces the memory footprint during the training process while guaranteeing the model accuracy and the overall training progress. Specifically, FedHybrid first selects the participating devices for each training round by jointly evaluating their memory budget, computing capability, and data diversity. After that, it judiciously analyzes the computational graph and generates an execution plan for each selected client in order to meet the corresponding memory budget while minimizing the training delay through employing a hybrid of recomputation and compression techniques according to the characteristic of each tensor. During the local training process, FedHybrid carries out the execution plan with a well-designed activation compression technique to effectively achieve memory reduction with minimum accuracy loss. We conduct extensive experiments to evaluate FedHybrid on both simulation and off-the-shelf mobile devices. The experiment results demonstrate that FedHybrid achieves up to a 39.1% increase in model accuracy and a 15.5X reduction in wall clock time under various memory budgets compared with the baselines.
Kahou Tam, Chunlin Tian, Li Li 0064, Haikai Zhao, Cheng-Zhong Xu 0001
SenSys2
2024 Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting
abstract
Federated Learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. Ever-increasing model complexity coupled with limited memory resources on the participating devices severely bottlenecks the deployment of FL in real-world scenarios. Thus, a framework that can effectively break the memory wall while jointly taking into account the hardware and statistical heterogeneity in FL is urgently required. In this article, we proposeSmartSplita framework that effectively reduces the memory footprint on the device side while guaranteeing the training progress and model accuracy for heterogeneous FL through model splitting. Towards this end,SmartSplitemploys a hierarchical structure to adaptively guide the overall training process. In each training round, the central manager, hosted on the server, dynamically selects the participating devices and sets the cutting layer by jointly considering the memory budget, training capacity, and data distribution of each device. The MEC manager, deployed within the edge server, proceeds to split the local model and perform training of the server-side portion. Meanwhile, it fine-tunes the splitting points based on the time-evolving statistical importance. The on-device manager, embedded inside each mobile device, continuously monitors the local training status while employing cost-aware checkpointing to match the runtime dynamic memory budget. Extensive experiments on representative datasets are conducted on both commercial off-the-shelf mobile device testbeds. The experimental results show thatSmartSplitexcels in FL training on highly memory-constrained mobile SoCs, offering up to a 94% peak latency reduction and 100-fold memory savings. It enhances accuracy performance by 1.49%-57.18% and adaptively adjusts to dynamic memory budgets through cost-aware recomputation
Chunlin Tian, Li Li 0064, Kahou Tam, Yebo Wu, Cheng-Zhong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2022 HARMONY: Heterogeneity-Aware Hierarchical Management for Federated Learning System
abstract
Federated learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. However, despite its emerging applications in many areas, real-world deployment of on-device FL is challenging due to wildly diverse training capability and data distribution across heterogeneous edge devices, which highly impact both model performance and training efficiency. This paper proposes Harmony, a high-performance FL framework with heterogeneity-aware hierarchical management of training devices and training data. Unlike previous work that mainly focuses on heterogeneity in either training capability or data distribution, Harmony adopts a hierarchical structure to jointly handle both heterogeneities in a unified manner. Specifically, the two core components of Harmony are a global coordinator hosted by the central server and a local coordinator deployed on each participating device. Without accessing the raw data, the global coordinator first selects the participants, and then further reorganizes their training samples based on the accurate estimation of the runtime training capability and data distribution of each device. The local coordinator keeps monitoring the local training status and conducts efficient training with guidance from the global coordinator. We conduct extensive experiments to evaluate Harmony using both hardware and simulation testbeds on representative datasets. The experimental results show that Harmony improves the accuracy performance by 1.67% - 27.62%. In addition, Harmony effectively accelerates the training process up to $3.29\times$ and $1.84\times$ on average, and saves energy up to 88.41% and 28.04% on average.
Chunlin Tian, Li Li 0064, Jun Wang 0001, Cheng-Zhong Xu 0001
MICRO1