EDBT 2026 Demo / reviewers in the wild / expert
Yanying Lin
dblp:286/6725
· DBLP profile ↗
16ranked-venue papers
9as first author
15since 2021 · last 2026
0000-0002-4809-9543ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersabstractServing Large Language Models (LLMs) in production faces significant challenges from highly variable request patterns and severe resource fragmentation in serverless clusters. Current systems rely on static pipeline configurations that struggle to adapt to dynamic workload conditions, leading to substantial inefficiencies. Yanying Lin, Chengzhi Lu, Cheng-Zhong Xu 0001, Kejiang Ye |
EuroSys | 1 |
| 2026 | DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries
Yanying Lin, Baicheng Chen, Cheng-Zhong Xu 0001, Kejiang Ye |
ISCA | 1 |
| 2026 | Connex: Endpoint Mobility Primitives for Dynamic LLM ServingabstractModern LLM serving systems increasingly adopt elastic inference pipelines where stages frequently join, leave, and migrate across nodes. However, existing GPU communication frameworks like NCCL assume static topologies, causing routing failures and P99 latency spikes during worker transitions that violate sub-millisecond tail latency requirements. We present Connex, a communication system that elevates endpoint mobility from exceptional failure to first-class primitive. Rather than optimizing individual mechanisms in isolation, Connex defines a mobility contract that the communication layer enforces whenever workers join, leave, or migrate while token streams, activations, or KV transfers are in flight. The contract is realized through three cooperating mechanisms: (1) epoch-based routing that bounds staleness without global coordination, (2) explicit handover protocols that preserve stream ordering and provide exactly-once delivery across migrations, and (3) credit-based backpressure with traffic-class isolation that prevents churn-induced interference with latency-critical paths. Evaluation on a 5-node GPU cluster under synthetic and production-derived churn shows that Connex reduces P99 tail spikes by up to 85% compared to NCCL-based baselines, achieves sub-second cutover, and maintains 100% goodput at moderate loads where baselines collapse to 0–28%, while incurring less than 5% steady-state overhead. Yanying Lin, Vincent Liu 0001, Cheng-Zhong Xu 0001, Kejiang Ye |
SIGCOMM | 1 |
| 2026 | Workload-Adapted Resource Allocation for LLM Distributed Serving in Serverless ClustersabstractLarge language models increasingly rely on pipeline parallelism for distributed inference, but existing systems face critical challenges in serverless environments: heterogeneous request distributions across pipeline stages and unpredictable workload patterns requiring rapid elasticity. Traditional static resource allocation fails to address pipeline-specific bottlenecks and cold start delays inherent in serverless architectures. We propose QUART, a workload-adapted resource allocation system for LLM distributed serving in serverless clusters. QUART introduces pipeline-aware resource management through: (1) latency-aware critical stage identification using coefficient of variation (CV)-based burst propagation analysis, (2) dynamic replica allocation with proportional-integral-derivative (PID) control for congested stages, and (3) hierarchical parameter caching with copy-on-write mechanisms enabling sub-second serverless scaling. The system addresses serverless-specific challenges through cache-aware scheduling that maintains model parameters in memory, eliminating disk I/O overhead during rapid scaling events. Evaluation with real-world workloads shows QUART reduces average response latency by up to 87.1% compared to existing serverless inference systems while achieving 2.37x improvement in goodput. Yanying Lin, Shutian Luo, Haiying Shen, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource EfficiencyabstractThis paper presents a comprehensive analysis of diffusion model serving challenges in production cloud environments. We examine the unique computational patterns and resource requirements that distinguish diffusion model serving from traditional ML workloads, revealing fundamental systemlevel challenges from their multi-stage pipeline architectures. Our analysis is based on a dataset collected from a commercial image generation service processing 3.5 million requests across 300+ GPUs of production operation. Yanying Lin, Shuaipeng Wu, Shutian Luo, Hong Xu 0001, Haiying Shen, Chong Ma 0005, Cheng-Zhong Xu 0001, Lin Qu, Kejiang Ye |
SoCC | 1 |
| 2025 | Rock: Serving Multimodal Models in Cloud with Heterogeneous-Aware Resource Orchestration for Thousands of LoRA AdaptersabstractIn this paper, we present ROCK, a novel system for efficiently serving thousands of LoRA adapters for multimodal models in cloud environments. Through extensive analysis of production workloads, we identify key challenges in current cloud-based image generation services: extreme request burstiness (up to$90 \times$normal rates), heterogeneous task characteristics, and inefficient adapter management that wastes 40 % of GPU memory and increases delays by$3 x$during peak times. ROCK addresses these challenges through a three-layer architecture that decouples hardware, adapters, and requests. Our system features dynamic heterogeneous queues that match tasks to appropriate resources based on multidimensional feature vectors, and a multilevel orchestration framework that intelligently manages adapter placement across heterogeneous storage. Experiments on a 64-GPU testbed demonstrate that ROCK reduces average response latency by$16-26 \%$, and achieves an 84.1 % cache hit rate for LoRA adapters-outperforming traditional approaches while reducing adapter update frequency by up to 77 %. Shuaipeng Wu, Yanying Lin, Wenyan Chen 0001, Chong Ma 0005, Cheng-Zhong Xu 0001, Kejiang Ye |
CLUSTER | 2 |
| 2025 | IAE-LoRa: Interference-Aware and Energy-Efficient LoRa Optimization Using Reinforcement Learning
Hongyu Tian, Kaitong Zheng, Yanying Lin, Changhao Yuan, Kejiang Ye |
ICIC (15) | 3 |
| 2025 | Understanding Serverless Inference in Mobile-Edge Networks: A Benchmark ApproachabstractAlthough the emerging serverless paradigm has the potential to become a dominant way of deploying cloud-service tasks across millions of mobile and IoT devices, the overhead characteristics of executing these tasks on such a volume of mobile devices remain largely unclear. To address this issue, this paper conducts a deep analysis based on the OpenFaaS platform—a popular open-source serverless platform for mobile edge environments—to investigate the overhead of performing deep learning inference tasks on mobile devices. To thoroughly evaluate the inference overhead, we develop a performance benchmark, namedESBench, whereby a set of comprehensive experiments are conducted with respect to a bunch of simulated mobile devices associated with an edge cluster. Our investigation reveals that the performance of deep learning inference tasks is significantly influenced by the model size and resource contention in mobile devices, leading to up to$3\times$degradation in performance. Moreover, we observe that the network environment can negatively impact the performance of mobile inference, increasing the CPU overhead under poor network conditions. Based on our findings, we further propose some recommendations for designing efficient serverless platforms and resource management strategies as well as for deploying serverless computing in the mobile edge environment. Yanying Lin, Shuaipeng Wu, Kenneth B. Kent, Kejiang Ye, Yang Wang 0006 |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | Serving LLM in Distributed GPU Cluster With Fine-Grain Pipeline ConstraintsabstractAs Large Language Models (LLMs) continue to advance, their parameter sizes are growing exponentially-far outpacing hardware capabilities. This widening gap necessitates distributed computing through pipeline parallelism for efficient inference. However, the uneven distribution of requests across pipeline stages creates significant performance bottlenecks in real-world deployments. To address this challenge, we presentPlanck, a performance optimization framework specifically designed for distributed LLM inference.Planck implements fine-grained control through two key mechanisms: a progressive SLO allocation strategy that dynamically adjusts time constraints based on workload patterns, and stage-specific performance controllers that prevent bottlenecks before they cascade through the system. By intelligently balancing resources across pipeline stages,Planck effectively eliminates queue buildup-essentially preventing traffic congestion before it forms. Evaluation using diverse workloads in real cloud environments demonstrates thatPlanck reduces P99 tail latency by up to 18% and decreases the longest queue lengths by as much as 47.8% across pipeline stages, significantly improving both system responsiveness and resource utilization. Yanying Lin, Shuaipeng Wu, Chengzhi Lu, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | EINS: Edge-Cloud Deep Model Inference with Network-Efficiency Schedule in ServerlessabstractModel inference in edge is often regarded as an effective method to alleviate high latency and enhance data privacy in edge-cloud collaborative computing environment. In this paper, we demonstrate that optimizing network communication in edge-cloud environment with limited bandwidth can enhance model inference performance. We first analyze network bottlenecks and the characteristics in model inference, then design a serverless inference system - EINS, to support collaborative optimization of network transmission and inference performance in edge-cloud environment. This system identifies concurrent network communication bottlenecks in multi-model deployment, dynamically scales capacity, and optimizes placement strategies and model transfer sequences. Real-world workload evaluation reveal that EINS can achieve a 5.7x throughput improvement and an average reduction of 62% latency in model instance startup. Yanying Lin, Wenyan Chen 0001, Yingfei Tang, Xu Duan, Kejiang Ye |
CSCWD | 2 |
| 2024 | QUART: Latency-Aware FaaS System for Pipelining Large Model InferenceabstractPipeline parallelism is a key mechanism to ensure the performance of large model serving systems. These systems need to deal with unpredictable online workloads with low latency and high good put. However, due to the specific characteristics of large models and resource constraints in pipeline parallelism, existing systems struggle to balance resource allocation across pipeline stages. The primary challenge resides in the differential distribution of requests across various stages of the pipeline. We propose QUART, a large model serving system that focuses on optimizing the performance of key stages in pipeline parallelism. QUART dynamically identifies the key stages of the pipeline and introduces an innovative two-level model parameter caching system based on forks to achieve rapid scaling of key stages within seconds. In evaluations with real-world request workloads, QUART reduces average response latency by up to 87.1%) and increases good put by 2.37x compared to the baseline. The experiments demonstrate that QUART effectively reduces tail latency and the average queue length of the pipeline. Yanying Lin, Yingfei Tang, Shutian Luo, Haiying Shen, Cheng-Zhong Xu 0001, Kejiang Ye |
ICDCS | 1 |
| 2024 | Planck: Optimizing LLM Inference Performance in Pipeline Parallelism with Fine-Grained SLO ConstraintabstractPipeline parallelism is an important strategy for improving inference performance in Large Language Models (LLMs). However, we find that different stages of LLM pipelines exhibit distinct performance and request characteristics, posing challenges to system performance in online inference scenarios. To address this issue, we propose Planck, a performance optimization framework tailored for LLM pipeline inference. By balancing request traffic, queue length, and execution time at each stage, Planck introduces a progressive SLO (Service Level Objective) allocation method and a stage instance performance controller. Planck fine-grainedly allocates SLOs to each pipeline stage and dynamically adjusts according to request distribution to control queue length. Through optimizing queue lengths across different stages of the model pipeline, Planck effectively reduces waiting time and tail latency. Evaluations conducted on a real cloud cluster using diverse workloads demonstrate that Planck effectively reduces P99 latency and queue length for each pipeline stage. Yanying Lin, Shuaipeng Wu, Chengzhi Lu, Cheng-Zhong Xu 0001, Kejiang Ye |
ICWS | 1 |
| 2023 | FLASH: Low-Latency Serverless Model Inference with Multi-Core Parallelism in EdgeabstractLow response latency holds a pivotal role in the landscape of edge deep learning model inference, yet the constrained resources within edge computing environments often limit its full potential. Within the scope of this paper, we substantiate that even amid the resource limitations inherent to edge computing environments, it remains feasible to curtail model response latency by elevating multi-core parallel efficiency. Our research encompasses a comprehensive analysis of the parallel acceleration effects observed across models featuring diverse parameter magnitudes during the inference process. This analysis culminates in the development of FLASH, an online deep model inference system tailored for Serverless edge inference, strategically optimized through the utilization of multi-core parallelism. FLASH exhibits dynamic adaptability by modulating the number of CPU cores within computational instances in accordance with traffic request loads. It also employs a dynamic scaling mechanism to finely adjust model placement, ultimately facilitating inference acceleration and mitigating the concomitant cold start overhead. Empirical experimentation conducted across a spectrum of burst-level workloads serves to underscore FLASH’s capacity, resulting in an average reduction in response latency by 33% and a maximum reduction of 75%, while concurrently realizing a throughput enhancement of 2.94x. Yanying Lin, Yingfei Tang, Wei Song 0008, Kejiang Ye |
ICPADS | 2 |
| 2023 | A Novel Multimodal Deep Learning Framework for Encrypted Traffic ClassificationabstractTraffic classification is essential for cybersecurity maintenance and network management, and has been widely used in QoS (Quality of Service) guarantees, intrusion detection, and other tasks. Recently, with the emergence of SSL/TLS encryption protocols in the modern Internet environment, the traditional payload-based classification methods are no longer effective. Some researchers have used machine learning methods to model the flow features of encrypted traffics (e.g. message type, length sequence, statistical features, etc.), and achieved good results in some cases. However, these high-level hand-designed features cannot be used for more fine-grained operations and may lead to the loss of important information, thus affecting the classification accuracy. To overcome this limitation, in this paper, we designed a novel multimodal deep learning framework for encrypted traffic classification called PEAN. PEAN uses the raw bytes and length sequence as the input, and uses the self-attention mechanism to learn the deep relationship among network packets in a biflow. Furthermore, unsupervised pre-training was introduced to enhance PEAN’s ability to characterize network packets. Experiments on a real trace set captured in a large data center demonstrate the effectiveness of PEAN, which achieves better results than the state-of-the-art methods. Kejiang Ye, Yishen Hu, Yanying Lin, Cheng-Zhong Xu 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Serverless Computing: State-of-the-Art, Challenges and OpportunitiesabstractServerless computing is growing in popularity by virtue of its lightweight and simplicity of management. It achieves these merits by reducing the granularity of the computing unit to the function level. Specifically, serverless allows users to focus squarely on the function itself while leaving other cumbersome management and scheduling issues to the platform provider, who is responsible for striking a balance between high-performance scheduling and low resource cost. In this article, we conduct a comprehensive survey of serverless computing with a particular focus on its infrastructure characteristics. Whereby some existing challenges are identified, and the associated cutting-edge solutions are analyzed. With these results, we further investigate some typical open-source frameworks and study how they address the identified challenges. Given the great advantages of serverless computing, it is expected that its deployment would dominate future cloud platforms. As such, we also envision some promising research opportunities that need to be further explored in the future. We hope that our work in this article can inspire those researchers and practitioners who are engaged in related fields to appreciate serverless computing, thereby setting foot in this promising area and making great contributions to its development. Yongkang Li 0003, Yanying Lin, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | LBNN: Perceiving the State Changes of a Core Telecommunications Network via Linear Bayesian Neural NetworkabstractThe core network is the most basic facility in the entire telecommunications network, which is consists of large number of routers, switches and firewalls. Network management like re-planning routes or adjusting policies is very important to avoid failures. However, the timing of intervention is very challenging. Too early intervention will incur unnecessary overheads, and too late intervention will cause serious disaster. In this paper, we analyzed a large data set from a real-world core telecommunications network and proposed Linear Bayesian Neural Networks (LBNN)11Code available at https://github.com/YanyingLin/Lbnn to perceive the core network state changes and make decisions about network intervention. In particular, we considered three aspects of complexity, including the weight of the mutual effect between devices, the dependence on the time dimension of the network states, and the randomness of the network state changes. The entire model is extended to a probability model based on the Bayesian framework to better capture the randomness and variability of the data. Experimental results on real-world data set show that LBNN achieves very high detection accuracy, with an average of 92.1%. Yanying Lin, Kejiang Ye, Naitian Deng, Tailin Wu, Cheng-Zhong Xu 0001 |
ICPADS | 1 |