Xuedou Xiao

dblp:249/5388 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-8724-7254ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust mmWave Radar Sensing With Multisensor Temporal Calibration and Supervision
abstract
With the rapid development of Internet of Things (IoT) technologies, autonomous driving has become an integral part of the IoT ecosystem, where millimeter-wave radar plays a crucial role in ensuring robust perception under challenging weather and lighting conditions. However, its sparse and noisy data often require enhancement using high-end sensors such as LiDAR or RTK-GNSS, which are not common in commercial vehicles. This paper introduces mmEMP+, a self-supervised learning technique that leverages pervasive visual and inertial (VI) measurements to enhance radar sensing data. Using VI data to improve radar sensing introduces several challenges. First, moving objects in a scene are inaccurately reconstructed by VI structure-from-motion, which consequently fails to enhance radar sensing. Second, multipath effects generate spurious radar points that can distort the representation of the environment. Finally, the temporal misalignment between the camera, IMU, and mmWave radar results in mismatched data association, thereby degrading system performance. To address these issues, mmEMP+ first proposes a dynamic 3D reconstruction method to recover the positions of moving features accurately. Then, we develop a spatial-stability checking method to filter out spurious radar points. Finally, mmEMP+ devises a tightly coupled sensor fusion method to calibrate the multi-sensor temporal offset. Experiments on a real-world dataset show that mmEMP+ achieves performance comparable to high-channel LiDAR-supervised methods while using only low-cost sensors. We further validate its effectiveness in IoT-relevant applications such as object detection, localization, and mapping.
Kezhong Liu, Shengkai Zhang, Mozi Chen, Xuedou Xiao, Shuai Wang 0008, Zheng Yang 0002, Wei Wang 0050
IEEE Internet Things J.5
2026 Octopus: Optimizing Interactive Video QoE via Loosely Coupled Codec-Transport Adaptation
abstract
Enhancing the quality of experience (QoE) in interactive video streaming (IVS) remains a persistent challenge due to the need for ultra-low latency and rising bandwidth demands. Conventional algorithms, whether rule-based or learning-based, are obsessed with achieving tight coupling between encoding and sending bitrate adaptations for low-latency guarantee. However, our measurement studies reveal alarming harms of tight coupling in suppressing throughput, encoding bitrates and smoothness, as application- and transport-layer bitrate adaptations inherently have different mechanisms and goals. To tackle this problem, we propose Octopus, the first loosely coupled cross-layer bitrate adaptation algorithm for IVS to maximize QoE. Instead of blind synchronization, Octopus promotes mutual cooperation and independence between encoding and sending bitrate adaptations by integrating a multi-head network with shortcut connections and auto-regressive action modules. Additionally, based on meta-imitation reinforcement learning, we design a network condition-aware online adaptation scheme that enables the loosely coupled policy to swiftly adapt to diverse and dynamic wireless networks. We implement Octopus on a testbed, a microcosm of real-world deployment, with transceiver pairs running WebRTC on the WeChat for Business dataset. Results show that Octopus outperforms state-of-the-art algorithms, either improving bitrates by 37.1%, or optimizing stalling rate and smoothness by 54.1% and 9.2%, or achieving all-around improvements.
Xuedou Xiao, Mingxuan Yan, Yingying Zuo, Boxi Liu, Paul Ruan, Yang Cao 0002, Yue Cao 0002, Wei Wang 0050
IEEE Trans. Mob. Comput.1
2025 PDStream: Slashing Long- Tail Delay in Interactive Video Streaming via Pseudo-Dual Streaming
Xuedou Xiao, Yingying Zuo, Mingxuan Yan, Kezhong Liu, Wei Wang 0050
INFOCOM1
2025 Methodology and Benchmark for Automated Driving Theory Test of Large Language Models
abstract
Large Language Models (LLMs), with their strong generalization and inference capabilities, have been increasingly leveraged to address the challenges of handling corner cases in autonomous driving (AD). However, a critical unresolved issue remains: the lack of a comprehensive understanding and formal assessment of LLMs’ driving theory knowledge and practical skills. To address this issue, we propose the first dedicated driving theory test framework and benchmark for LLMs. That is a crucial yet unexplored area in the literature, particularly for safety-critical applications in autonomous driving and driver assistance. Our framework systematically evaluates LLMs’ competence in driving theory and hazard perception, akin to the official UK driving theory test, ensuring their qualification for critical driving-related tasks. To facilitate rigorous benchmarking, we construct a comprehensive dataset comprising over 700 multiple-choice questions (MCQs) and 54 hazard perception video tests sourced from the official UK driving theory examination. Additionally, we incorporate two standardized MCQ sets from the UK’s Driver and Vehicle Standards Agency (DVSA). For these two types of theoretical test items, we design tailored assessment methodologies and evaluation metrics, including accuracy, recall, precision, F1-score, real-time performance, and computational efficiency. The experimental results reveal that among all LLMs tested, only GPT-4o achieved an accuracy of 88. 21% in the MCQs test, successfully passing this component. However, in hazard perception testing, none of the evaluated models met the passing criteria under the given settings, highlighting the substantial improvements required before these models can be practically deployed for real-world driving applications. Our key insight is that the specific test questions LLMs fail to answer correctly directly reflect their deficiencies in understanding and flexibly applying traffic regulations, as well as in analyzing and responding to complex driving scenarios. This provides clear directions for future improvements.
Dashuai Pei, Jianhua He 0001, Kezhong Liu, Mozi Chen, Xuedou Xiao, Shengkai Zhang
IEEE Trans. Intell. Transp. Syst.6
2024 Task-Oriented Video Compressive Streaming for Real-Time Semantic Segmentation
abstract
Real-time semantic segmentation (SS) is a major task for various vision-based applications such as self-driving. Due to the limited computing resources and stringent performance requirements, streaming videos from camera-embedded mobile devices to edge servers for SS is a promising approach. While there are increasing efforts on task-oriented video compression, most SS-applicable algorithms apply more uniform compression, as the sensitive regions are less obvious and concentrated. Such processing results in low compression performance and significantly limits the capacity of edge servers supporting real-time SS. In this paper, we propose STAC, a novel task-oriented DNN-driven video compressive streaming algorithm tailed for SS, to strike accuracy-bitrate balance and adapt to time-varying bandwidth. It exploits DNN's gradients as sensitivity metrics for fine-grained spatial adaptive compression and includes a temporal adaptive scheme that integrates spatial adaptation with predictive coding. Furthermore, we design a new bandwidth-aware neural network, serving as a compatible configuration tuner to fit time-varying bandwidth and content. STAC is evaluated in a system with a commodity mobile device and an edge server with real-world network traces. Experiments show that STAC can save up to 63.7–75.2% of bandwidth or improve accuracy by 3.1–9.5% compared to state-of-the-art algorithms, while capable of adapting to time-varying bandwidth.
Xuedou Xiao, Yingying Zuo, Mingxuan Yan, Wei Wang 0050, Jianhua He 0001, Qian Zhang 0001
IEEE Trans. Mob. Comput.1
2023 From Ember to Blaze: Swift Interactive Video Adaptation via Meta-Reinforcement Learning
abstract
Maximizing quality of experience (QoE) for interactive video streaming has been a long-standing challenge, as its delay-sensitive nature makes it more vulnerable to bandwidth fluctuations. While reinforcement learning (RL) has demonstrated great potential, existing works are either limited by fixed models or require enormous data/time for online adaptation, which struggle to fit time-varying and diverse network states. Driven by these practical concerns, we perform large-scale measurements on WeChat for Business’s interactive video service to study real-world network fluctuations. Surprisingly, our analysis shows that, compared to time-varying network metrics, network sequences exhibit noticeable short-term continuity, sufficient for few-shot learning requirement. We thus propose Fiammetta, the first meta-RL-based bitrate adaptation algorithm for interactive video streaming. Building on the short-term continuity, Fiammetta accumulates learning experiences through offline meta-training and enables fast online adaptation to changing network states through few gradient updates. Moreover, Fiammetta innovatively incorporates a probing mechanism for real-time monitoring of network states, and proposes an adaptive meta-testing mechanism for seamless adaptation. We implement Fiammetta on a testbed whose end-to-end network follows the real-world WeChat for Business traces. The results show that Fiammetta outperforms prior algorithms significantly, improving video bitrate by 3.6%-16.2% without increasing stalling rate.
Xuedou Xiao, Mingxuan Yan, Yingying Zuo, Boxi Liu, Paul Ruan, Yang Cao 0002, Wei Wang 0050
INFOCOM1
2023 Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
abstract
Offloading computing to edge servers is a promising solution to support growing video understanding applications at resource-constrained IoT devices. Recent efforts have been made to enhance the scalability of such systems by reducing inference costs on edge servers. However, existing research is not directly applicable to pixel-level vision tasks such as video semantic segmentation (VSS), partly due to the fluctuating VSS accuracy and segment bitrate caused by the dynamic video content. In response, we present Penance, a new edge inference cost reduction framework. By exploiting softmax outputs of VSS models and the prediction mechanism of H.264/AVC codecs, Penance optimizes model selection and compression settings to minimize the inference cost while meeting the required accuracy within the available bandwidth constraints. We implement Penance in a commercial IoT device with only CPUs. Experimental results show that Penance consumes a negligible 6.8% more computation resources than the optimal strategy while satisfying accuracy and bandwidth constraints with a low failure rate.
Mingxuan Yan, Yi Wang 0118, Xuedou Xiao, Zhiqing Luo, Jianhua He 0001, Wei Wang 0050
ACM Multimedia3
2022 DNN-Driven Compressive Offloading for Edge-Assisted Semantic Video Segmentation
abstract
Deep learning has shown impressive performance in semantic segmentation, but it is still unaffordable for resource-constrained mobile devices. While offloading computation tasks is promising, the high traffic demands overwhelm the limited bandwidth. Existing compression algorithms are not fit for semantic segmentation, as the lack of obvious and concentrated regions of interest (RoIs) forces the adoption of uniform compression strategies, leading to low compression ratios or accuracy. This paper introduces STAC, a DNN-driven compression scheme tailored for edge-assisted semantic video segmentation. STAC is the first to exploit DNN’s gradients as spatial sensitivity metrics for spatial adaptive compression and achieves superior compression ratio and accuracy. Yet, it is challenging to adapt this content-customized compression to videos. Practical issues include varying spatial sensitivity and huge bandwidth consumption for compression strategy feedback and offloading. We tackle these issues through a spatiotemporal adaptive scheme, which (1) takes partial strategy generation operations offline to reduce communication load, and (2) propagates compression strategies and segmentation results across frames through dense optical flow, and adaptively offloads keyframes to accommodate video content. We implement STAC on a commodity mobile device. Experiments show that STAC can save up to 20.95% of bandwidth without losing accuracy, compared to the state-of-the-art algorithm.
Xuedou Xiao, Juecheng Zhang, Wei Wang 0050, Jianhua He 0001, Qian Zhang 0001
INFOCOM1
2022 Sensor-Assisted Rate Adaptation for UAV MU-MIMO Networks
abstract
Propelled by multi-user MIMO (MU-MIMO) technology, unmanned aerial vehicles (UAVs) as mobile hotspots have recently emerged as an attractive wireless communication paradigm. Rate adaptation (RA) becomes indispensable to enhance UAV communication robustness against UAV mobility-induced channel variances. However, existing MU-MIMO RA algorithms are mainly designed for ground communications with relatively stable channel coherence time, which incurs channel measurement staleness and sub-optimal rate selections when coping with highly dynamic air-to-ground links. In this paper, we propose SensRate, a new uplink MU-MIMO RA algorithm dedicated for low-altitude UAVs, which exploits inherent on-board sensors used for flight control with no extra cost. We propose a novel channel prediction algorithm that utilizes sensor-estimated flight states to assist channel direction prediction for each client and estimate inter-user interference for optimal rates. We provide an implementation of our design using a commercial UAV and show that it achieves an average throughput gain of$1.24\times $and$1.28\times $compared with the bestknown RA algorithm for 2- and 3-antenna APs, respectively.
Xuedou Xiao, Wei Wang 0050, Tao Jiang 0002
IEEE/ACM Trans. Netw.1
2020 Sensor-Augmented Neural Adaptive Bitrate Video Streaming on UAVs
abstract
Recent advances in unmanned aerial vehicle (UAV) technology have revolutionized a broad class of civil and military applications. However, the designs of wireless technologies that enable real-time streaming of high-definition video between UAVs and ground clients present a conundrum. Most existing adaptive bitrate (ABR) algorithms are not optimized for the air-to-ground links, which usually fluctuate dramatically due to the dynamic flight states of the UAV. In this paper, we present SA-ABR, a new sensor-augmented system that generates ABR video streaming algorithms with the assistance of various kinds of inherent sensor data that are used to pilot UAVs. By incorporating the inherent sensor data with network observations, SA-ABR trains a deep reinforcement learning (DRL) model to extract salient features from the flight state information and automatically learn an ABR algorithm to adapt to the varying UAV channel capacity through the training process. SA-ABR does not rely on any assumptions or models about UAV's flight states or the environment, but instead, it makes decisions by exploiting temporal properties of past throughput through the long short-term memory (LSTM) to adapt itself to a wide range of highly dynamic environments. We have implemented SA-ABR in a commercial UAV and evaluated it in the wild. We compare SA-ABR with a variety of existing state-of-the-art ABR algorithms, and the results show that our system outperforms the best known existing ABR algorithm by 21.4% in terms of the average quality of experience (QoE) reward.
Xuedou Xiao, Wei Wang 0050, Taobin Chen, Yang Cao 0002, Tao Jiang 0002, Qian Zhang 0001
IEEE Trans. Multim.1