EDBT 2026 Demo / reviewers in the wild / expert
Wuyang Zhang
dblp:194/1520
· DBLP profile ↗
16ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent SpaceabstractThe rapid proliferation of Large Language Models (LLMs) has led to a fragmented and inefficient ecosystem, a state of ``model lock-in'' where seamlessly integrating novel models remains a significant bottleneck. Current routing frameworks require exhaustive, costly retraining, hindering scalability and adaptability. We introduce ZeroRouter, a new paradigm for LLM routing that breaks this lock-in. Our approach is founded on a universal latent space, a model-agnostic representation of query difficulty that fundamentally decouples the characterization of a query from the profiling of a model. This allows for zero-shot onboarding of new models without full-scale retraining. ZeroRouter features a context-aware predictor that maps queries to this universal space and a dual-mode optimizer that balances accuracy, cost, and latency. Our framework consistently outperforms all baselines, delivering higher accuracy at lower cost and latency. Wuyang Zhang, Ziyang Tao, Yanyong Zhang |
AAAI | 2 |
| 2025 | SDN Controller Design for LEO Satellite Networks with Three-Phase Traffic EngineeringabstractWith recent advancements in Low Earth Orbit (LEO) satellite constellation launch technology, the cost of deploying LEO satellites is rapidly decreasing. Leveraging these constellations, LEO satellite networks provide the potential for low-latency communications, enabling latency-sensitive applications such as financial trading and tele-surgeries. However, the large scale of these networks introduces significant management challenges, particularly in adapting to dynamic environments and responding to network events. Maintaining efficient operation requires continuous updates to forwarding rules in a constantly shifting topology, where frequent satellite handovers further complicate real-time decision-making.To address these challenges, we propose an SDN design for LEO satellite networks, incorporating a three-phase traffic engineering scheme to achieve optimal path assignments with minimal control-loop latency. The key idea is to take advantage of the trajectories of satellite movements and perform path assignment ahead of time. We evaluate the proposed controller design extensively through simulations. The results demonstrate that our three-phase traffic engineering scheme significantly improves network performance. Compared to a purely reactive controller, our design achieves up to twice the residual bandwidth while deriving path assignments that attain 95% of the optimal minimal residual bandwidth. Zibin Chen, Wuyang Zhang |
ICCCN | 2 |
| 2025 | CalibWorkflow: A General MLLM-Guided Workflow for Centimeter-Level Cross-Sensor CalibrationabstractExtrinsic calibration is a fundamental step in sensor fusion systems. However, existing methods often lack generalization capabilities when facing diverse hardware configurations, sensor poses, and environmental conditions, hindering their large-scale deployment. To address this limitation, we propose a general extrinsic calibration method, CalibWorkflow. Our core innovation lies in positioning multimodal large language models (MLLMs) as ''visual guides'' for the calibration process, leveraging their powerful vision-language understanding capabilities to guide parameter search and refinement. This reliance on visual scene understanding, rather than specific geometric features or sensor characteristics, enables the method to generalize effectively across diverse hardware and environmental conditions. Specifically, CalibWorkflow employs a three-stage calibration pipeline: initial parameter search, coarse optimization, and fine optimization. First, it utilizes the MLLM to assess the visual consistency between the projected point cloud and the image, rapidly determining an initial range for the extrinsic parameters. Next, the MLLM serves as a differential evaluator, giving simple ''better'' or ''worse'' feedback on parameter changes to guide the search through the parameter space. Finally, the method refines the calibration by matching edge features and performing non-linear optimization. Extensive experiments are conducted across six diverse scenarios and four heterogeneous sensor combinations. CalibWorkflow achieves state-of-the-art sub-degree and centimeter-level accuracy on four datasets and demonstrates highly competitive performance on others. These results thoroughly validate the generalization and robustness when facing various scenarios. Codes will be available. Wuyang Zhang, Guoliang You, Xiaomeng Chu, Wenhao Yu 0010, Yifan Duan, Yanyong Zhang |
ACM Multimedia | 2 |
| 2025 | UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous DrivingabstractThe rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary nature of these autonomous systems and closed-source GPU drivers hinder fine-grained control over GPU executions, often resulting in missed deadlines that compromise vehicle performance. To address this, we present UrgenGo, a non-intrusive, urgency-aware GPU scheduling system that operates without access to application source code. UrgenGo implicitly prioritizes GPU executions through transparent kernel launch manipulation, employing task-level stream binding, delayed kernel launching, and batched kernel launch synchronization. We conducted extensive real-world evaluations in collaboration with a self-driving startup, developing 11 GPU-bound task chains for a realistic autonomous navigation application and implementing our system on a self-driving bus. Our results show a significant 61% reduction in the overall deadline miss ratio, compared to the state-of-the-art GPU scheduler that requires source code modifications. Hanqi Zhu, Wuyang Zhang, Ziyang Tao, Xinrui Lin, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
MobiCom | 2 |
| 2025 | UniSense: Spatial-Uncertainty-Aware Collaborative Sensing for Autonomous DrivingabstractVehicle-to-vehicle collaborative perception faces fundamental deployment barriers: raw LiDAR data sharing requires over 300 Mbps per vehicle - far exceeding V2X network capacities, while network delays of 80-200ms create dangerous temporal misalignments at highway speeds. We present UniSense, a distributed collaborative perception system that enables efficient and reliable multi-vehicle perception through uncertainty-driven sensor data exchange. Instead of sharing raw sensor data, vehicles exchange compact uncertainty maps that identify regions requiring additional perceptual information. Our key innovations include: (1) a lightweight uncertainty quantification pipeline that runs in real-time on automotive hardware, identifying perception-critical regions while reducing bandwidth requirements by more than 10×, (2) a bandwidth-aware protocol that dynamically adapts data sharing based on network conditions and perception uncertainty, and (3) a selective motion compensation scheme that maintains temporal consistency. We evaluate UniSense through a year-long deployment with 16 roadside LiDAR nodes and autonomous vehicles across our campus. Our experimental results show that UniSense extends reliable perception range from local 80m to 140m, improving accuracy by 1.33× on average, up to 1.73×, over the state-of-the-art baselines, under communication constraints. The code and dataset are available at https://github.com/LetStarFly/UniSense. Haojie Ren, Wuyang Zhang, Shuyao Shi, Yanyong Zhang |
MobiSys | 2 |
| 2025 | MgHiSal: MLLM-guided hierarchical semantic alignment for multimodal knowledge graph completion
Wuyang Zhang, Yunxia Yin |
Knowl. Based Syst. | 2 |
| 2024 | Map++: Towards User-Participatory Visual SLAM Systems with Efficient Map Expansion and SharingabstractConstructing precise 3D maps is crucial for the development of future map-based systems such as self-driving and navigation. However, generating these maps in complex environments, such as multi-level parking garages or shopping malls, remains a formidable challenge. In this paper, we introduce a participatory sensing approach that delegates map-building tasks to map users, thereby enabling cost-effective and continuous data collection. The proposed method harnesses the collective efforts of users, facilitating the expansion and ongoing update of the maps as the environment evolves. Hanqi Zhu, Yifan Duan, Wuyang Zhang, Longfei Shangguan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang |
MobiCom | 4 |
| 2024 | Overcoming Noisy Labels and Non-IID Data in Edge Federated LearningabstractFederated learning (FL) enables edge devices to cooperatively train models without exposing their raw data. However, implementing a practical FL system at the network edge mainly faces three challenges: label noise, data non-IIDness, and device heterogeneity, which seriously harm model performance and slow down convergence speed. Unfortunately, none of the existing works tackle all three challenges simultaneously. To this end, we develop a novel FL system, called Aorta, which features adaptive dataset construction and aggregation weightassignment. On each client, Aorta first calibrates potentially noisy labels and then constructs a training dataset with low noise, balanced distribution, and proper size. To fully utilize limited data on clients, we propose a global model guided method to select clean data and progressively correct noisy labels. To achieve balanced class distribution and proper dataset size, we propose a distribution-and-capability-aware data augmentation method to generate local training data. On the server, Aorta assigns aggregation weights based on the quality of local models to ensure that high-quality models have a greater influence on the global model. The model quality is measured through its cosine similarity with a benchmark model, which is trained on a clean and balanced dataset. We conduct extensive experiments on four datasets with various settings, including different noise types/ratios and non-IID types/levels. Compared to the baselines, Aorta improves model accuracy up to 9.8% on the datasets with moderate noise and non-IIDness, while providing a speedup of 4.2× on average when achieving the same target accuracy. Yang Xu 0020, Yunming Liao, Lun Wang 0003, Hongli Xu 0001, Zhida Jiang, Wuyang Zhang |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | BOSE: Block-Wise Federated Learning in Heterogeneous Edge ComputingabstractAt the network edge, federated learning (FL) has gained attention as a promising approach for training deep learning (DL) models collaboratively across a large number of devices while preserving user privacy. However, FL still faces specific challenges related to the limited, heterogeneous and dynamic resources of devices. In most FL systems, all devices train the same model, while the devices with constrained resources, referred to as stragglers, will significantly slow down overall training process. It is intuitive to alleviate computation and communication load on the stragglers by training and transmitting a part of the model. Inspired by multi-exit models, we divide an original DL model into several non-overlapping blocks, which can be trained separately on the low-capability devices. Furthermore, we propose BOSE, a novel FL system that performs adaptiveblock-wisemodel training under resource constraints. Considering the diverse impacts of different blocks on model convergence and the varying training loads they incur, a naive block assignment strategy, e.g., uniformly random assignment, may not yield optimal model performance and fail to fully utilize available resources. To this end, we introduce two metrics, includinglearning speedanddevice-wise divergence, to measure the potential of blocks in promoting model convergence. Given resource budget, BOSE initially identifies a set of candidate blocks for each device and subsequently selects specific training blocks based on their potential for promoting model convergence. In general, blocks with higher potential are more likely to be chosen for training. Extensive experiments on a physical platform show that BOSE provides a 1.4$\times$$\sim$3.8$\times$speedup without sacrificing model accuracy, compared to the baselines. Lun Wang 0003, Yang Xu 0020, Hongli Xu 0001, Zhida Jiang, Min Chen 0033, Wuyang Zhang, Chen Qian 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2021 | Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingabstractAs mobile devices continuously generate streams of images and videos, a new class of mobile deep vision applications are rapidly emerging, which usually involve running deep neural networks on these multimedia data in real-time. To support such applications, having mobile devices offload the computation, especially the neural network inference, to edge clouds has proved effective. Existing solutions often assume there exists a dedicated and powerful server, to which the entire inference can be offloaded. In reality, however, we may not be able to find such a server but need to make do with less powerful ones. To address these more practical situations, we propose to partition the video frame and offload the partial inference tasks to multiple servers for parallel processing. This paper presents the design of Elf, a framework to accelerate the mobile deep vision applications with any server provisioning through the parallel offloading. Elf employs a recurrent region proposal prediction algorithm, a region proposal centric frame partitioning, and a resource-aware multi-offloading scheme. We implement and evaluate Elf upon Linux and Android platforms using four commercial mobile devices and three deep vision applications with ten state-of-the-art models. The comprehensive experiments show that Elf can speed up the applications by 4.85× with saving bandwidth usage by 52.6%, while with <1% application accuracy sacrifice. Wuyang Zhang, Zhezhi He, Zhenhua Jia, Yunxin Liu 0001, Marco Gruteser, Dipankar Raychaudhuri, Yanyong Zhang |
MobiCom | 1 |
| 2021 | Improving stochastic local search for uniform k-SAT by generating appropriate initial assignmentabstractAbstract Stochastic local search (SLS) algorithms are well known for their ability to efficiently find models of random instances of the SAT problem, especially for uniform random k‐SAT instances. Two processes affect most SLS solvers—the initial assignment of the variables and the heuristics that select which variable to flip. In the last few years, the work on generating the appropriate initial assignment has not been paid much attention or seen much progress, while most SLS solvers focused on the heuristic algorithm. The present work aims to improve SLS algorithms on uniform random k‐SAT instances by developing effective methods for generating the initial assignment of variables in a controlled way. First, the allocation strategy introduced recently for 3‐SAT instances is extended to initialize the initial assignment on random k‐SAT instances. Then a concept of an initial probability distribution of the clause‐to‐variable ratio of the instance is introduced to determine the parameters of the allocation strategy. This combined method is added to the beginning of six state‐of‐the‐art SLS algorithms in order to generate initial assignments of variables in a controlled way instead of generating them randomly, resulting in six extended SLS algorithms named WalkSATlm_E, DCCASat_E, Score2SAT_E, CSCCSat_E, Probsat_E, and Sparrow_E, respectively. They are then evaluated in terms of their capabilities and efficiency on uniform random k‐SAT instance from the random track of SAT Competitions in 2016, 2017, and 2018. Experimental results show that these improved SLS solvers outperform their original performance, especially WalkSAT_E, Score2SAT_E, and CSCCSat_E outperform the winner of the random track of SAT competition in 2017. In addition, based on the initial probability distribution method, the present work proposes a parameter tuning and analysis of random 3‐SAT instances and provides an additional comparative analysis with the state‐of‐the‐art random SLS solvers based on large‐scale experiments. Huimin Fu 0002, Wuyang Zhang, Guanfeng Wu, Yang Xu 0001, Jun Liu 0001 |
Comput. Intell. | 2 |
| 2019 | Hetero-Edge: Orchestration of Real-time Vision Applications on Heterogeneous Edge CloudsabstractRunning computer vision algorithms on images or videos collected by mobile devices represent a new class of latency-sensitive applications that expect to benefit from edge cloud computing. These applications often demand real-time responses (e.g., <;100 ms), which can not be satisfied by traditional cloud computing. However, the edge cloud architecture is inherently distributed and heterogeneous, requiring new approaches to resource allocation and orchestration. This paper presents the design and evaluation of a latency-aware edge computing platform, aiming to minimize the end-to-end latency for edge applications. The proposed platform is built on Apache Storm, and consists of multiple edge servers with heterogeneous computation (including both GPUs and CPUs) and networking resources. Central to our platform is an orchestration framework that breaks down an edge application into Storm tasks as defined by a directed acyclic graph (DAG) and then maps these tasks onto heterogeneous edge servers for efficient execution. An experimental proof-of-concept testbed is used to demonstrate that the proposed platform can indeed achieve low end-to-end latency: considering a real-time 3D scene reconstruction application, it is shown that the testbed can support up to 30 concurrent streams with an average perframe latency of 32ms, and can achieve 40% latency reduction relative to the baseline Storm scheduling approach. Wuyang Zhang, Sugang Li, Zhenhua Jia, Yanyong Zhang, Dipankar Raychaudhuri |
INFOCOM | 1 |
| 2019 | Updates with Multiple Service ClassesabstractA source submits status update jobs to a service facility for processing and delivery to a monitor. The status updates belong to service classes with different service requirements. We model the service requirements using a hyperexponential service time model. To avoid class-specific bias in the service process, the system implements an M/G/1/1 blocking queue; new arrivals are discarded if the server is busy. Using an age-of-information (AoI) metric to characterize timeliness of the updates, a stochastic hybrid system (SHS) approach is employed to derive the overall average AoI and the average AoI for each service class. We observe that both the overall AoI and class-specific AoI share a common penalty that is a function of the second moment of the average service time and they differ chiefly because of their different arrival rates. We show that each high-probability service class has an associated age-optimal update arrival rate while low-probability service classes incur an average age that is always decreasing in the update arrival rate. Roy D. Yates, Wuyang Zhang |
ISIT | 3 |
| 2018 | Cutting the Cord: Designing a High-quality Untethered VR System with Low Latency Remote RenderingabstractThis paper introduces an end-to-end untethered VR system design and open platform that can meet virtual reality latency and quality requirements at 4K resolution over a wireless link. High-quality VR systems generate graphics data at a data rate much higher than those supported by existing wireless-communication products such as Wi-Fi and 60GHz wireless communication. The necessary image encoding, makes it challenging to maintain the stringent VR latency requirements. To achieve the required latency, our system employs a Parallel Rendering and Streaming mechanism to reduce the add-on streaming latency, by pipelining the rendering, encoding, transmission and decoding procedures. Furthermore, we introduce a Remote VSync Driven Rendering technique to minimize display latency. To evaluate the system, we implement an end-to-end remote rendering platform on commodity hardware over a 60Ghz wireless network. Results show that the system can support current 2160x1200 VR resolution at 90Hz with less than 16ms end-to-end latency, and 4K resolution with 20ms latency, while keeping a visually lossless image quality to the user. Ruiguang Zhong, Wuyang Zhang, Yunxin Liu 0001, Jiansong Zhang 0001, Marco Gruteser |
MobiSys | 3 |
| 2018 | Continuous Low-Power Ammonia Monitoring Using Long Short-Term Memory Neural NetworksabstractAccurate and continuous ammonia monitoring is important for laboratory animal studies and many other applications. Existing solutions are often expensive, inaccurate, or unsuitable for long-term monitoring. In this work, we propose a new ammonia monitoring approach that is low-power, automatic, accurate, and wireless. Zhenhua Jia, Xinmeng Lyu, Wuyang Zhang, Richard P. Martin, Richard E. Howard, Yanyong Zhang |
SenSys | 3 |
| 2016 | SEGUE: Quality of Service Aware Edge Cloud Service MigrationabstractEdge cloud computing moves cloud services to the edge of the network, thereby allowing clients to access services with a significantly reduced network delay. This service migration is intended to enable a range of latency sensitive mobile applications. In this paper, we propose to manage user QoS by actively migrating services to different edge clouds in response to degraded server or network performance. Previous studies have proposed a distance-based Markov Decision Process (MDP) for optimizing migration decisions. These models provide the feasibility of applying MDP to edge cloud service migration decisions. However, these models fail to consider dynamic network and server states in migration decisions. In this work, we address these limitations by designing a comprehensive edge cloud migration decision system, which we call SEGUE. SEGUE achieves optimal migration decisions by providing a long-term optimal QoS to mobile users in the presence of link quality and server load variation. The basis of SEGUE is in its QoS-aware service migration and its state based MDP model which effectively incorporates the two dominant factors in making migration decisions: 1) network state, and 2) server state. An evaluation of SEGUE performance is given through an augmented reality application. Our results demonstrate that SEGUE reduces the response time of this application by 27.21% and 53.70% compared to the lowest load migration model and the least hop migration model, respectively. Wuyang Zhang, Yanyong Zhang, Dipankar Raychaudhuri |
CloudCom | 1 |