Cunchi Lv

dblp:335/1396 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0001-7089-0315ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
abstract
Serverless computing, with its ease of management, auto-scaling, and cost-effectiveness, is widely adopted by deep learning (DL) applications. DL workloads, especially with large language models, require substantial GPU resources to ensure QoS. However, it is prone to produce GPU fragments (e.g., 15%-94%) in serverless DL systems due to the dynamicity of workloads and coarse-grained static GPU allocation mechanisms, gradually eroding the profits offered by serverless elasticity. Different from classical serverless systems that only scale horizontally, we present introspective elasticity (IE), a fine-grained and adaptive two-dimensional co-scaling mechanism to support GPU resourcing-on-demand for serverless DL tasks. Based on this insight, we build Dilu, a cross-layer and GPU-based serverless DL system with IE support. First, Dilu provides multi-factor profiling for DL tasks with efficient pruning search methods. Second, Dilu adheres to the resourcing-complementary principles in scheduling to improve GPU utilization with QoS guarantees. Third, Dilu adopts an adaptive 2D co-scaling method to enhance the elasticity of GPU provisioning in real time. Evaluations show that it can dynamically adjust the resourcing of various DL functions with low GPU fragmentation (10%-46% GPU defragmentation), high throughput (up to 1.8× inference and 1.1× training throughput increment) and QoS guarantees (11%-71% violation rate reduction), compared to the SOTA baselines.
Cunchi Lv, Xiao Shi 0003, Zhengyu Lei, Jinyue Huang, Wenting Tan, Xiaohui Zheng
ASPLOS (1)1
2025 FZeroTC: fully zero-shot text classification for simultaneously discovering and labeling unseen classes
Dongsheng Duan, Cunchi Lv, Yangxi Li
Knowl. Inf. Syst.2
2024 SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
Cunchi Lv, Xiao Shi 0003, Wenting Tan
NPC (1)1
2024 An anomaly aware network embedding framework for unsupervised anomalous link detection
Dongsheng Duan, Lingling Tong, Jie Lu 0009, Cunchi Lv, Yangxi Li
Data Min. Knowl. Discov.5
2023 Chitu: Accelerating Serverless Workflows with Asynchronous State Replication Pipelines
abstract
Serverless workflows are characterized as multi-stage computing, while downstream functions require accessing intermediate states or the output of upstream functions for running. The workflow's performance can be easily affected due to the inefficiency of data access. Studies accelerate data access with various policies, such as direct and indirect methods. However, these methods may fail due to various limitations such as resource availability.
Zhengyu Lei, Xiao Shi 0003, Cunchi Lv
SoCC3
2022 TrainFlow: A Lightweight, Programmable ML Training Framework via Serverless Paradigm
Wenting Tan, Xiao Shi 0003, Zhengyu Lei, Cunchi Lv
NPC5