Yanze Zhang

dblp:355/9702 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LightTrace: A Versatile Ebpf-Enabled Toolkit for Lightweight Distributed Tracing
abstract
Distributed tracing is widely employed for troubleshooting distributed systems such as microservices. Existing tracing systems typically improve one or more of data completeness, non-intrusiveness, or lightweight operation through various data generation strategies. However, no current solution optimizes all three aspects simultaneously, which can introduce significant overhead in I/O-intensive environments that demand both high completeness and minimal intrusion. In this paper, we introduce LightTrace, a novel, eBPF-enabled toolkit that optimizes the data transmission mechanism in distributed tracing. LightTrace integrates seamlessly with mainstream tracing systems and leverages eBPF to reduce end-to-end latency and lower overall system overhead. LightTrace is implemented using a combination of kernel-level eBPF and user-space Golang components. Our evaluation demonstrates that LightTrace decreases the average latency overhead by up to 22.8 % and improves the peak throughput of microservice systems by between$\mathbf{1 2. 3 \%}$and$\mathbf{1 8. 6 \%}$. Furthermore, LightTrace's adaptability across diverse tracing platforms underscores its versatility in various microservice environments.
Yanze Zhang, Kanye Ye Wang, Shufan Gong, Huanghuang Liang, Chuang Hu, Xiaobo Zhou 0002
ICPADS1
2025 Adaptive Deadlock Avoidance for Decentralized Multi-Agent Systems via CBF-Inspired Risk Measurement
abstract
Decentralized safe control plays an important role in multi-agent systems given the scalability and robustness without reliance on a central authority. However, without an explicit global coordinator, the decentralized control methods are often prone to deadlock - a state where the system reaches equilibrium, causing the robots to stall. In this paper, we propose a generalized decentralized framework that unifies the Control Lyapunov Function (CLF) and Control Barrier Function (CBF) to facilitate efficient task execution and ensure deadlock-free trajectories for the multi-agent systems. As the agents approach the deadlock-related undesirable equilibrium, the framework can detect the equilibrium and drive agents away before that happens. This is achieved by a secondary deadlock resolution design with an auxiliary CBF to prevent the multi-agent systems from converging to the undesirable equilibrium. To avoid dominating effects due to the deadlock resolution over the original task-related controllers, a deadlock indicator function using CBF-inspired risk measurement is proposed and encoded in the unified framework for the agents to adaptively determine when to activate the deadlock resolution. This allows the agents to follow their original control tasks and seamlessly unlock or deactivate deadlock resolution as necessary, effectively improving task efficiency. We demonstrate the effectiveness of the proposed method through theoretical analysis, numerical simulations, and real-world experiments.
Yanze Zhang, Yiwei Lyu 0002, Siwon Jo, Yupeng Yang
ICRA1
2025 Computationally and Sample Efficient Safe Reinforcement Learning Using Adaptive Conformal Prediction
abstract
Safety is a critical concern in learning-enabled autonomous systems especially when deploying these systems in real-world scenarios. An important challenge is accurately quantifying the uncertainty of unknown models to generate provably safe control policies that facilitate the gathering of informative data, thereby achieving both safe and optimal policies. Additionally, the selection of the data-driven model can significantly impact both the real-time implementation and the uncertainty quantification process. In this paper, we propose a provably sample efficient episodic safe learning framework that remains robust across various model choices with quantified uncertainty for online control tasks. Specifically, we first employ Quadrature Fourier Features (QFF) for kernel function approximation of Gaussian Processes (GPs) to enable efficient approximation of unknown dynamics. Then the Adaptive Conformal Prediction (ACP) is used to quantify the uncertainty from online observations and combined with the Control Barrier Functions (CBF) to characterize the uncertainty-aware safe control constraints under learned dynamics. Finally, an optimism-based exploration strategy is integrated with ACP-based CBFs for safe exploration and near-optimal safe nonlinear control. Theoretical proofs and simulation results are provided to demonstrate the effectiveness and efficiency of the proposed framework.
Yanze Zhang
ICRA2
2025 T3Set: A Multimodal Dataset with Targeted Suggestions for LLM-based Virtual Coach in Table Tennis Training
abstract
Coaching is critical for learning table tennis skills.However, amateur table tennis players often lack access to professional coaches due to high costs and a limited number of coaches.While recent multimodal large language models show promise as virtual coaches, most of the existing approaches merely rely on video analysis, which is not comprehensive enough.In table tennis, many important kinematic details (e.g., strength, acceleration) cannot be captured by videos.They can only be tracked using sensors.To address this gap, we present T3Set (Table Tennis Training Set), a multimodal dataset that synchronizes inertial measurement unit (IMU) data from sensors mounted on 32 players' rackets with video recordings.The sensor data has 16 dimensions and a sample rate of 100Hz.This dataset covers 7 fundamental techniques across 380 training rounds, totaling 8655 annotated strokes, with 8395 targeted suggestions from coaches.The key features of T3Set include (1) temporal alignment between sensor data, video data, and text data.(2) high-quality targeted suggestions which are consistent with predefined suggestion taxonomy.Based on T3Set, we propose a novel two-stage framework that effectively integrates motion perception with generative reasoning as a virtual coach.Our method quantitatively outperforms baseline methods.The dataset, code, and documentation are available at
Yanze Zhang, Xiao Xie, Hui Zhang 0051, Jiachen Wang 0001, Yingcai Wu
KDD (2)4
2025 Featherlight Stateful WebAssembly for Serverless Inference Workflows
abstract
In serverless inference, complex prediction tasks are executed as workflows, relying on efficient state transfer across multiple functions. Serverless platforms typically deploy each function in a separate stateless container, depending on external processes for state management, which often results in suboptimal system utilization and increased latency. We introduce WasmFlow, a novel framework designed for serverless inference that ensures low latency and high throughput. This is achieved through process-level virtualization using WebAssembly. WasmFlow operates functions on a per-thread basis within compact WebAssembly modules, significantly reducing startup times and memory usage. The framework has two key features. (1) Efficient Memory Sharing: WasmFlow facilitates direct and rapid state transfer between functions using threads within the WebAssembly runtime. This is enabled through lightweight, lock-free, zero-copy intra-process communication, complemented by effective inter-process RPC. (2) System Optimizations: We further optimize WasmFlow with an advanced synchronization technique between functions, an affinity-aware workflow scheduler, and adaptive request batching. Implemented and integrated within the Kubernetes ecosystem, WasmFlow's performance was evaluated using synthetic workloads and realworld Azure traces, including typical serverless workflows and ML models. Our results demonstrate that WasmFlow dramatically outperforms existing serverless frameworks. It reduces P90 end-to-end latency by 74x and 78x, increases function density by n1.7x and 223x compared to Faasm and SPRIGHT, and improves system throughput by 12.3x and 8.8x over Knative and WasmEdge, respectively.
Xingguo Pang, Yanze Zhang, Zhuofu Chen, Zhijun Ding, Dazhao Cheng, Xiaobo Zhou 0002
IEEE Trans. Parallel Distributed Syst.3
2024 Integrating Online Learning and Connectivity Maintenance for Communication-Aware Multi-Robot Coordination
abstract
This paper proposes a novel data-driven control strategy for maintaining connectivity in networked multi-robot systems. Existing approaches often rely on a predetermined communication model specifying whether pairwise robots can communicate given their relative distance to guide the connectivity-aware control design, which may not capture real-world communication conditions. To relax that assumption, we present the concept of Data-driven Connectivity Barrier Certificates, which utilize Control Barrier Functions (CBF) and Gaussian Processes (GP) to characterize the admissible control space for pairwise robots based on communication performance observed online. This allows robots to maintain a satisfying level of pairwise communication quality (measured by the received signal strength) while in motion. Then we propose a Data-driven Connectivity Maintenance (DCM) algorithm that combines (1) online learning of the communication signal strength and (2) a bi-level optimization-based control framework for the robot team to enforce global connectivity of the realistic multi-robot communication graph and minimally deviate from their task-related motions. We provide theoretical proofs to justify the properties of our algorithm and demonstrate its effectiveness through simulations with up to 20 robots.
Yupeng Yang, Yiwei Lyu 0002, Yanze Zhang, Ian Gao
IROS3
2024 Expeditious High-Concurrency MicroVM SnapStart in Persistent Memory with an Augmented Hypervisor
Xingguo Pang, Yanze Zhang, Dazhao Cheng, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002
USENIX ATC2