Tianjian Gong

dblp:413/8092 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%
Computer networks
1 paper
Edge and fog computing · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Edge and fog computing
edge inference
0.912025
ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models · IEEE Trans. Mob. Comput. 2025
Cloud and datacenter computing
serverless computing
0.912025
ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models · IEEE Trans. Mob. Comput. 2025
Cloud and datacenter computing › serverless computing
serverless inference
0.912025
ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models · IEEE Trans. Mob. Comput. 2025

Methods — techniques the papers use, named apart from their topics

fractional rounding · 2.6expert intervention · 2.6deep reinforcement learning · 2.6
YearPublicationVenuePosition
2025 ExpertDRL: Request Dispatching and Instance Configuration for Serverless Edge Inference With Foundation Models
abstract
The prevalence of the pre-training & fine-tuning paradigm enables machine learning models to quickly adapt to various downstream tasks by fine-tuning pre-trained foundation models (FMs), greatly facilitating various IoT applications that rely on model inference in dynamic edge serverless environments. Efficiently dispatching inference requests and configuring instances to batch inference requests can significantly enhance resource efficiency. However, existing serverless inference solutions are tailored for traditional models, make coarse-grained request dispatching and instance configuration decisions, fail to exploit the shared model backbone characteristics of the FM and capture delayed rewards in dynamic environments, and ignore communication latency between edge sites, resulting in high costs and constraint violations. In this paper, we leverage our insight that fine-grained batch inference requests can effectively exploit the shared model backbone feature of FM to save monetary costs. We propose an algorithm that incorporates deep reinforcement learning (DRL) and expert intervention for fine-grained request dispatching and instance configuration, where the DRL component outputs fractional solutions as guidance, while the expert intervention module integrates our insights—batching reduces monetary costs at the expense of increased inference latency, whereas higher configurations shorten inference latency. This module rounds fractional solutions and adjusts instance configurations to search for optimal solutions while satisfying constraints, with theoretical guarantees rigorously proved. Finally, we conducted our experiments on an OpenFaas-based platform and simulator, and extensive trace-driven evaluation results show that ExpertDRL can save costs by up to 85.14% and improve request acceptance ratio by up to 26.93%, compared to the state-of-the-art solution.
Yue Zeng 0002, Junlong Zhou, Zhihao Qu, Song Guo 0001, Tianjian Gong
IEEE Trans. Mob. Comput.6