Fenglong Cai

dblp:357/6886 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0000-7520-9632ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Edge-LLM: A Collaborative Framework for Large Language Model Serving in Edge Computing
abstract
The rapid advancement and extensive implementation of Large Language Models (LLMs) are milestones in the realm of artificial intelligence. Although Parameter-Efficient Transfer Learning (PETL), a.k.a. Adapter, methods have reduced the barrier for fine-tuning and inference on LLMs, it becomes a challenge to efficiently deploy and fine-tuning different adapter models needed for massive AI applications. With the popularity of SoC chips, the computing power of edge devices has improved significantly. To meet the computational resources required by LLM applications and improve quality of service (QoS), we propose Edge-LLM, a server-node collaboration framework for large-scale language model serving, to efficiently utilize edge resources to accelerate LLM fine-tuning and inference in resource-constrained scenarios. In the framework, we implement an adaptive quantization strategy, FM cache mechanism, and value density first (VDF) scheduling algorithm to reduce GPU overhead and accelerate LLM computation. The experimental results demonstrate that Edge-LLM can significantly improve overall computational speed by a factor of 17, decrease the number of tasks experiencing timeouts by 63%, and reduce GPU overhead by up to 43%.
Fenglong Cai, Dong Yuan 0001, Li-Zhen Cui 0001
ICWS1
2024 AutoMP: A Tool to Automate Performance Testing for Model Placement on GPUs
abstract
As AI applications are widely deployed in various fields, the computing cost of AI is skyrocketing. The high cost not only puts pressure on the environment but also poses challenges for researchers entering the field of deep learning. Green AI is gradually gaining attention, aiming to reduce computing costs and make AI application deployment more efficient and environmentally friendly. However, due to the wide variety of deep learning models, there are many challenges in efficiently deploying multiple models on GPUs. An unreasonable model deployment strategy will lead to insufficient utilization of GPU computing resources. We address some of these challenges through AutoMP, a tool that automates performance testing for model placement on GPUs. Through AutoMP, researchers can flexibly initiate a large number of experiments to study which models are suitable for inference tasks on the same GPU, thereby making GPU utilization more efficient. AutoMP provides a user-friendly visual interface and an API to meet users’ needs in different scenarios. AutoMP also provides a complete experimental analysis tool that generates visual charts of experimental data from multiple dimensions and gives experimental conclusions to assist researchers in making decisions. To date, AutoMP has been used by a large number of users for empirical research on deep learning model placement. The cumulative number of experiments has exceeded 10,000. The flexibility and extensibility of AutoMP and our own experience using it show that this tool plays a vital role in promoting Green AI.
Wei He 0020, Fenglong Cai, Wei Guo 0017, Li-Zhen Cui 0001
ISPA3
2024 FastPTM: Fast weights loading of pre-trained models for parallel inference service provisioning
Fenglong Cai, Dong Yuan 0001, Wei He 0020, Wei Guo 0017, Li-Zhen Cui 0001
Parallel Comput.1
2023 ParaTra: A Parallel Transformer Inference Framework for Concurrent Service Provision in Edge Computing
abstract
Edge computing has been widely used to deploy and service deep learning applications. Equipped with GPUs, edge nodes can process concurrent incoming inference requests of the deep learning model. However, existing methods for inference tasks do not allow efficient parallel handling of user requests. This paper investigates the popular Transformer deep learning model and develops ParaTra, a parallel transformer inference framework for providing parallel inference services to users. In the framework, the Transformer model is partitioned and deployed in users’ devices and the edge node to efficiently utilize their processing power. The concurrent inference tasks with different sizes are dynamically packaged in a scheduling queue and sent in batch to an encoder-decoder pipeline for processing. ParaTra can significantly reduce the overheads of parallel processing and the usage of GPU memory. Experiment results show that ParaTra can save up to 37.1% of GPU memory usage and improve 8.4 times of processing speed.
Fenglong Cai, Dong Yuan 0001, Mengwei Xie, Wei He 0020, Lanju Kong, Wei Guo 0017, Yali Jiang 0004, Li-Zhen Cui 0001
ICWS1