VLDB 2026 Research / reviewers in the wild / expert
Ziyi Wang 0002
dblp:160/2171-2
· DBLP profile ↗
12ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-8174-0593ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TailorLLM: Collaborative End-Cloud Inference of Large and Small Language Models Based on Low-Rank AdaptationabstractWith the rapid expansion of large language model inference service users, cloud computing resource costs have become a critical challenge for service providers. Although utilizing end-device resources for auxiliary inference provides new possibilities to reduce cloud computing costs, existing solutions struggle to achieve an ideal balance across multi-task accuracy, end-to-end latency, and cloud computing costs. Ziyi Wang 0002, Haonan Jin, Lanshan Zhang |
EuroSys | 2 |
| 2026 | IntentVQA: A QoE-Aware Cloud-End Collaborative System for Video QA
Yishuo Zhang, Lanshan Zhang, Yaya Wei, Ziyan Zhong, Xiaohui Xie, Ziyi Wang 0002, Wendong Wang 0003 |
IWQoS | 6 |
| 2026 | Task-Aware Cloud-End Offloading for Vision-Language Model Serving via Dynamic Modality-Specific Adapter SchedulingabstractLarge-scale vision-language models enable powerful cross-modal understanding and generation, driving rapidly growing demand for online inference services. However, cloud-centric serving often suffers from high latency, rising costs, and network dependency, while purely on-device deployment is constrained by limited memory and reduced accuracy on complex tasks. To address this accuracy–latency–cost trilemma, we propose ShiftVL, a task-aware end–cloud serving framework that shifts suitable execution to the end device with a cloud fallback. ShiftVL serves high-frequency requests on an end-side small VLM enhanced with ViTexLoRA, a modality-disentangled parameter-efficient tuning method that preserves cross-modal alignment, while routing low-frequency or complex requests to a cloud-hosted large VLM for higher accuracy. Under tight device budgets, ShiftVL employs a predictive adapter scheduler that combines LRU-style caching with imitation learning to pre-load task-specific adapters. Experiments with InternVL models show that ShiftVL reduces cloud cost by up to 76.3% and latency by up to 42.9% while maintaining high multi-task accuracy, demonstrating its practicality for real-world vision-language model serving. Ziyi Wang 0002, Yaya Wei, Ziyan Zhong, Lanshan Zhang |
WWW | 2 |
| 2026 | Cooperative task offloading and resource allocation for sequential constraint tasks in satellite edge computing networks
Xiangyang Gong, Ziyi Wang 0002, Xirong Que |
Ad Hoc Networks | 3 |
| 2025 | Edge-Assisted Adaptive Configuration for Serverless-Based Video AnalyticsabstractThe growth of video volumes and increased DNN capabilities have led to a growing desire for video analytics, which demands intensive computation resources. Traditional resource provisioning strategies, such as configuring a cluster per peak utilization, lead to low resource efficiency. Serverless computing is a promising way to avoid wasteful resource provisioning since video analytics regularly encounters bursty input workloads and fine-grained video content dynamics. For serverless-based video analytics, the application configuration (frame rate, detection model, and computation resources) will impact several metrics, such as computation cost and analytics accuracy. In this paper, we investigate the joint configuration adjustment problem for video knobs and computation resources provided by the serverless platform. We propose an algorithm that can efficiently adapt configurations for video streams to address two key challenges in serverless-based video analytics systems, including the complex relationships between the configurations and the key performance metrics, and the dynamically best configuration. Our adaptive configuration adjustment algorithm is developed based on Markov approximation to minimize the computation cost. To guarantee the accuracy, we then design the keyframe selection algorithm based on the secretary algorithm to identify significant changes in video content. We have developed a prototype over AWS Lambda and conducted extensive experiments with real-world video streams. The results show that our algorithm can greatly reduce the computation cost under the constraint of target accuracy. Ziyi Wang 0002, Songyu Zhang, Wei Cheng 0008, Wendong Wang 0003, Yong Cui 0001 |
IEEE Trans. Netw. | 1 |
| 2024 | Ptu: Pre-Trained Model for Network Traffic UnderstandingabstractNetwork traffic understanding is crucial to providing high-quality network services and protecting network security. However, due to the growing complexity of networks and the rising proportion of encrypted traffic, existing methods for network traffic understanding face severe challenges. Traditional approaches rely on manually designed features or require a large amount of labeled data, while pre-trained models offer new possibilities. Nevertheless, existing pre-trained models have the following limitations: (1) Their inputs only contain features from the packet content, neglecting temporal information about network dynamics. (2) Their pre-training targets only focus on static characteristics of the data stream without understanding the process of the network transmission. This paper presents the Pre-trained model for network Traffic Understanding (PTU), an innovative model that employs self-supervised pre-training to address the challenges of network traffic understanding. In PTU, we design a traffic representation scheme that integrates static packet content and network dynamics into a unified input space. Furthermore, we propose a pre-training method that includes four tailored pre-training targets. This approach enables PTU to capture both static and dynamic characteristics of network traffic from massive amounts of unlabeled data, thereby achieving enhanced performance in downstream tasks through fine-tuning. Extensive experiments confirm PTU's state-of-the-art (SOTA) performance. In traffic classification tasks, PTU achieves an F1 score of over$\mathbf{0. 9 9}$and secures a more than$\mathbf{1 0 \%}$improvement in accuracy in the most challenging task of encrypted application classification. Lingfeng Peng, Xiaohui Xie, Sijiang Huang, Ziyi Wang 0002, Yong Cui 0001 |
ICNP | 4 |
| 2024 | Resolving Spurious Temporal Location Dependency for Video Corpus Moment RetrievalabstractVideo Corpus Moment Retrieval aims to retrieve the relevant video from a large corpus and localize the corresponding moment within the target video based on a specific query. Existing methods have achieved promising accuracy and efficiency through dedicated model designs. However, we suggest these methods overly exploit dataset biases instead of semantics. This will lead to exaggerated performances on biased datasets and implicit a significant deficiency in generalizability, which is an important metric yet not considered in existing studies. In this paper, we observe the degradation caused by a spurious dependency and design a model to mitigate this harm. Specifically, we generate an Out-Of-Distributed (OOD) test set from a widely used TV Retrieval dataset, revealing the existing models' erroneous dependency on the temporal locations of target moments. Therefore, we utilize a theoretical Structural Causal Model (SCM) to dig into the roots of this dependency by constructing causal paths for the models. Furthermore, we propose a concrete Clip Location Deconfounding Model (CLDM) to disentangle the confounded video features into the content part and the location confounder part, then produce results with causal intervention. Experiments show that CLDM significantly alleviates the impact brought by dataset biases thus providing advanced generalizability among existing works. Yishuo Zhang, Lanshan Zhang, Zhizhen Zhang, Ziyi Wang 0002, Xiaohui Xie, Wendong Wang 0003 |
SMC | 4 |
| 2023 | Edge-Assisted Adaptive Configuration for Serverless-Based Video AnalyticsabstractThe growth of video volumes and increased DNN capabilities have led to a growing desire for video analytics, which demands intensive computation resources. Traditional resource provisioning strategies, such as configuring a cluster per peak utilization, lead to low resource efficiency. Serverless computing is a promising way to avoid wasteful resource provisioning since video analytics regularly encounters bursty input workloads and finegrained video content dynamics. For serverless-based video analytics, the application configuration (frame rate, detection model, and computation resources) will impact several metrics, such as computation cost and analytics accuracy. In this paper, we investigate the joint configuration adjustment problem for video knobs and computation resources provided by the serverless platform. We propose an algorithm that can efficiently adapt configurations for video streams to address two key challenges in serverless-based video analytics systems, including the complex relationships between the configurations and the key performance metrics, and the dynamically best configuration. Our algorithm is developed based on Markov approximation to minimize the computation cost within an accuracy constraint. We have developed a prototype over AWS Lambda and conducted extensive experiments with real-world video streams. The results show that our algorithm can greatly reduce the computation cost under the constraint of target accuracy. Ziyi Wang 0002, Songyu Zhang, Zhixiong Wu, Yong Cui 0001 |
ICDCS | 1 |
| 2023 | Edge-Assisted Real-Time Video Analytics With Spatial-Temporal Redundancy SuppressionabstractDriven by plummeting camera prices and advances of video inference algorithms, video cameras are deployed ubiquitously and organizations usually rely on live video analytics to retrieve key information, such as the locations and identities of target objects. However, analyzing real-time video poses severe challenges to today’s network and computation systems. To balance accuracy, bandwidth usage, and latency, we present EVA, an edge-assisted real-time video analytics framework, which coordinates computationally weak cameras with more powerful edge servers to enable video analytics under the accuracy and latency requirements of applications. EVA treats the region where a target object is located as a fine-grained transmission unit and exploits the redundancies in both spatial and temporal domains to reduce the bandwidth usage. Based on the framework, we design an adaptive offloading algorithm, which coordinates the recognition process between the camera and the server. To adapt to complex environments, we then design a threshold adjustment algorithm to tune the confidence threshold dynamically. Experiments on real-world video feeds show that compared to several recent baselines on multiple video genres, EVA maintains high accuracy while reducing bandwidth usage by up to 90%. Ziyi Wang 0002, Zhizhen Zhang, Yishuo Zhang, Wei Cheng 0008, Wendong Wang 0003, Yong Cui 0001 |
IEEE Internet Things J. | 1 |
| 2022 | MultiLive: Adaptive Bitrate Control for Low-Delay Multi-Party Interactive Live StreamingabstractIn multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates tosend; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we designMultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We also deploy the algorithm on a commercial live streaming platform that provides such services for more than 2300 users. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-$5\times $improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing. Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015 |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | MultiLive: Adaptive Bitrate Control for Low-delay Multi-party Interactive Live StreamingabstractIn multi-party interactive live streaming, each user can act as both the sender and the receiver of a live video stream. Designing adaptive bitrate (ABR) algorithm for such applications poses three challenges: (i) due to the interaction requirement among the users, the playback buffer has to be kept small to reduce the end-to-end delay; (ii) the algorithm needs to decide what is the bitrate to receive and what is the set of bitrates to send; (iii) the delay and quality requirements between each pair of users may differ, for instance, depending on whether the pair is interacting directly with each other. To address these challenges, we first develop a quality of experience (QoE) model for multi-party live streaming applications. Based on this model, we design MultiLive, an adaptive bitrate control algorithm for the multi-party scenario. MultiLive models the many-to-many ABR selection problem as a non-linear programming problem. Solving the non-linear programming equation yields the target bitrate for each pair of sender-receiver. To alleviate system errors during the modeling and measurement process, we update the target bitrate through the buffer feedback adjustment. To address the throughput limitation of the uplink, we cluster the ideal streams into a few groups, and aggregate these streams through scalable video coding for transmissions. We conduct extensive trace-driven simulations to evaluate the algorithm. The experimental results show that MultiLive outperforms the fixed bitrate algorithm, with 2-5× improvement in average QoE. Furthermore, the end-to-end delay is reduced to around 100 ms, much lower than the 400 ms threshold recommended for video conferencing. Ziyi Wang 0002, Yong Cui 0001, Xin Wang 0001, Wei Tsang Ooi, Yi Li 0015 |
INFOCOM | 1 |
| 2018 | Task Scheduling with Optimized Transmission Time in Collaborative Cloud-Edge LearningabstractDeep learning has been applied in many recent advanced applications in the field of transportation, finance and medicine. These applications require significant computation resources and large-scale training samples. Cloud becomes a natural choice for conducting these learning tasks due to its abundant resources. However, deeper penetration of deep learning techniques in mission critical applications, like driverless car, calls for stricter time requirement to guarantee its interaction and larger amount of dataset for training to guarantee its accuracy, which cannot be easily satisfied by the cloud and makes the network transmission become the bottleneck. Edge learning emerges to be a promising direction to reduce data transmission time by processing and compressing the raw data at the edge of the network, while brings the concern of accuracy reduction at the meantime. To balance this tradeoff under cloud-edge architecture, we study a task scheduling problem for reducing weighted transmission time which takes learning accuracy into consideration. We also propose efficient scheduling algorithms which are able to achieve up to 50% reduction in makespan with extensive trace-driven simulations. Yutao Huang, Yifei Zhu 0001, Xiaoyi Fan 0001, Xiaoqiang Ma, Fangxin Wang 0001, Jiangchuan Liu, Ziyi Wang 0002, Yong Cui 0001 |
ICCCN | 7 |