VLDB 2026 Research / reviewers in the wild / expert
Yepeng Zhang
dblp:355/8240
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Movie101v2: Improved Movie Narration BenchmarkabstractAutomatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.Unlike standard video captioning, it involves not only describing key visual details but also inferring plots that unfold across multiple movie shots, presenting distinct and complex challenges.To advance this field, we introduce Movie101v2, a largescale, bilingual dataset with enhanced data quality specifically designed for movie narration.Revisiting the task, we propose breaking down the ultimate goal of automatic movie narration into three progressive stages, offering a clear roadmap with corresponding evaluation metrics.Based on our new benchmark, we baseline a range of large vision-language models and conduct an in-depth analysis of the challenges in movie narration generation.Our findings highlight that achieving applicable movie narration generation is a fascinating goal that requires significant research. Zihao Yue, Yepeng Zhang, Qin Jin |
ACL (1) | 2 |
| 2025 | STGNN-Based Microservice Autoscaling and Placement for Fast QoS Recovery at the EdgeabstractEdge computing enables low-latency data processing by placing resources closer to end users, and integrating microservices enhances scalability and responsiveness. However, maintaining QoS in dynamic, resource-constrained edge environments remains challenging. Existing methods often rely on local load monitoring, overlooking cross-layer load propagation, which leads to delayed autoscaling and prolonged QoS violations. To address this, we propose a Load Propagation-aware Autoscaling and Placement (LPAP) approach for proactive QoS management. LPAP anticipates performance degradation by dynamically adjusting scaling and placement based on both direct and indirect caller impacts and varying resource consumption patterns. It employs a Spatio-Temporal Graph Neural Network (STGNN) to estimate influence coefficients that quantify the impact of invocation edges on microservice performance. These coefficients guide an improved Actor-Critic algorithm that incorporates them into the Temporal Difference (TD) error and reward function to accelerate learning. Experiments show that LPAP effectively reduces QoS degradation duration and improves response times under dynamic workloads. Yepeng Zhang |
ICWS | 3 |
| 2025 | RL-Based Hybrid CPU Scaling for Soft Deadline Constrained Tasks in Container CloudsabstractExisting CPU scaling approaches have limitations that can lead to inefficient resource allocation and increased penalty costs for tasks with soft deadlines running in container clouds. First, quota allocation based approaches overlook the gap between the obtainable CPU time and allocated quota, causing inefficient CPU utilization and unexpected task behaviors. Second, core allocation based approaches ignore workload dynamics within decision intervals, potentially increasing contention for CPU time among tasks on the same core. Third, existing approaches lack strategies to allocate more resources to critical tasks that incur higher penalty costs when the node’s capacity is insufficient. This article proposes a reinforcement learning based hybrid CPU scaling approach that allocates quota and cores jointly, aiming to minimize penalty costs for timeouts. Based on the embedding generated from a fine-grained CPU demand series, we allocate CPU quotas and determine a dynamic workload-aware core sharing scheme using an attention mechanism that combines respective demands and global criticality regarding penalty costs. Additionally, we integrate the resource gap, CPU time contention, and penalty costs into the reward function to update our model online. The experimental results show the proposed approach achieves state-of-the-art performance. Yepeng Zhang, Huadong Ma |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | FaaSCTDO: Collaborative Task-Data Orchestration for Serverless WorkflowsabstractThe use of Function-as-a-Service (FaaS) platforms for executing complex serverless workflows has gained significant popularity. However, the stateless nature of FaaS requires functions to rely on remote storage to store their state, leading to performance overhead and reduced efficiency. Previous approaches using excess memory resources as caches can cause resource contention and decreased task execution performance. Moreover, the impact of workflow orchestration approaches on latency has not been fully explored. To address these challenges, we propose FaaSCTDO, a framework for collaborative task and data orchestration in serverless workflows. FaaSCTDO provides developers with comprehensive lifecycle management capabilities, treating data as orchestratable objects. It groups tasks and data with high data correlation together and allocates them on the same server node, leveraging data locality to expedite function access to data. Neng Yang, Yepeng Zhang |
CLOUD | 3 |
| 2023 | Joint Optimization of CPU Scaling and Core Sharing in Container CloudsabstractExisting approaches to scale CPU quota in container clouds lack a consideration of the gap between the promised quota and actually obtainable amount of CPU resources. However, this gap is noticeable with Completely Fair Scheduler (CFS), which is used in Linux kernel to allocate time slices to threads. CFS shares cores among containers regardless of their quotas, causing containers with more threads to exhaust their quotas earlier due to parallel execution across multiple cores. This leaves these cores idle but unavailable for containers with no threads on them, resulting in insufficient CPU time for those containers. As a result, the execution of tasks in these containers becomes slower, potentially exceeding their deadlines and incurring penalty costs. In this paper, we propose a joint optimization approach for CPU scaling and core sharing, which uses Particle Swarm Optimization (PSO) algorithm to iteratively make decisions with the goal of minimizing penalty costs for missing deadlines. In each iteration, a core sharing scheme that aims at reducing idle CPU time will be determined based on the candidate scaling plan represented by each particle’s position. By considering the impact of following periods, we evaluate the penalty costs incurred by the interdependent scaling and core sharing decisions based on the obtainable CPU time for each container, which drives the co-optimization of both aspects of the decisions. The experimental results in a real environment show the proposed approach achieves state-of-the-art performance. Yepeng Zhang, Huadong Ma |
ICPADS | 1 |
| 2023 | Graph-Based Root Cause Localization in Microservice Systems with Protection MechanismsabstractService anomalies are difficult to locate accurately due to their propagation through service dependencies in microservice systems. Besides, the protection mechanisms are introduced into the microservice systems to ensure the stable operation of services. However, the existing approaches ignore the impact of protection mechanisms on the root cause localization of abnormal services. Specifically, the circuit breaking and rate limiting mechanisms can refuse service requests and thus change the way of anomaly propagation. Moreover, the different service request frequencies and latency make service dependencies change dynamically, resulting in the different probabilities of anomaly propagation among services. In this paper, we propose a novel framework named MicroGBPM to locate the root cause of abnormal services. We model the anomaly propagation among services as a dynamically constructed service attributed graph with metrics and traces when a failure occurs. To eliminate the impact of the protection mechanisms, we design a two-stage dynamic calibration strategy to adjust the probability of anomaly propagation among services. Then, we propose a random walking approach to calculate the root cause results by using the PageRank algorithm. The experimental results show that MicroGBPM improves the accuracy of root cause localization compared to other approaches in the microservice systems with protection mechanisms. Neng Yang, Yepeng Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 4 |