EDBT 2026 Demo / reviewers in the wild / expert
Xiuxian Guan
dblp:286/8652
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0001-6133-8388ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AGRNav: Efficient and Energy-Saving Autonomous Navigation for Air-Ground Robots in Occlusion-Prone EnvironmentsabstractThe exceptional mobility and long endurance of air-ground robots are raising interest in their usage to navigate complex environments (e.g., forests and large buildings). However, such environments often contain occluded and unknown regions, and without accurate prediction of unobserved obstacles, the movement of the air-ground robot often suffers a sub-optimal trajectory under existing mapping-based and learning-based navigation methods. In this work, we present AGRNav, a novel framework designed to search for safe and energy-saving air-ground hybrid paths. AGRNav contains a lightweight semantic scene completion network (SCONet) with self-attention to enable accurate obstacle predictions by capturing contextual information and occlusion area features. The framework subsequently employs a query-based method for low-latency updates of prediction results to the grid map. Finally, based on the updated map, the hierarchical path planner efficiently searches for energy-saving paths for navigation. We validate AGRNav’s performance through benchmarks in both simulated and real-world environments, demonstrating its superiority over classical and state-of-the-art methods. The open-source code is available at https://github.com/jmwang0117/AGRNav. Junming Wang 0001, Zekai Sun, Xiuxian Guan, Tianxiang Shen, Zongyuan Zhang, Tianyang Duan, Dong Huang 0005, Shixiong Zhao, Heming Cui |
ICRA | 3 |
| 2023 | New Problems in Active Sampling for Mobile Robotic Online LearningabstractAI models deployed in real-world tasks (e.g., surveillance, implicit mapping, health care) typically need to be online trained for better modelling of the changing real-world environments and various online training methods (e.g., domain adaptation, few shot learning) are proposed for refining the AI models based on training input incrementally sampled from the real world. However, in the whole loop of AI model online training, there is a section rarely discussed: how to sample training input from the real world. In this paper, we show from the perspective of online training of AI models deployed on edge devices (e.g., robots) that several problems in sampling of training input on the device are affecting the time and energy consumption for the online training process to reach high performance. Notably, the online training relies on training input consecutively sampled from the real world and the consecutive samples from nearby states (e.g., position and orientation of a camera) are too similar and would limit the training accuracy gain per training iteration; on the other hand, while we can choose to sample more about the inaccurate samples to better final training accuracy, it is costly to obtain the accuracy statistics of samples via traditional ways such as validating, especially for AI models deployed on edge devices. These findings aim to raise research effort for practical online training of AI models, so that they can achieve resiliently and sustainably high performance in real-world tasks. Xiuxian Guan, Junming Wang 0001, Zekai Sun, Zongyuan Zhang, Tianyang Duan, Shengliang Deng, Fangming Liu, Heming Cui |
COMPSAC | 1 |
| 2023 | Coorp: Satisfying Low-Latency and High-Throughput Requirements of Wireless Network for Coordinated Robotic LearningabstractIn coordinated robotic learning, multiple robots share the same wireless channel for communication, and bring together latency-sensitive (LS) network flows for control and bandwidth-hungry (BH) flows for distributed learning. Unfortunately, existing wireless network supporting systems cannot coordinate these two network flows to meet their own requirements: 1) prioritized contention systems (e.g., EDCA) prevent LS messages from timely acquiring the wireless channel because multiple wireless network interface cards (WNICs) with BH messages are contending for the channel 2) global planning systems (e.g., SchedWiFi) have to reserve a notable time window in the shared channel for each LS flow, suffering from severe bandwidth degradation (up to 42%). We present the coordinated preemption method to meet both requirements for LS flows and BH flows. Globally (among multiple robots), coordinated preemption eliminates unnecessary contention of BH flows by making them transmit in a round-robin manner, such that LS flows have the highest chance to win the contention against BH flows, without sacrificing overall bandwidth from the perspective of coordinated robotic learning applications. Locally (within the same robot), coordinated preemption in real time predicts the periodic transmission of LS flows from the upper application and conservatively limits packets of BH flows buffered in the WNIC only before LS packets arriving, reducing the bandwidth devoted to preemption. COORP, our implementation of coordinated preemption, reduced the violation of latency requirements from 53.9% (EDCA) to 8.8% (comparable to SchedWiFi). Regarding learning quality, COORP achieved a comparable (at times the same) learning reward with EDCA, which grew up to 76% faster than SchedWiFi. Shengliang Deng, Xiuxian Guan, Zekai Sun, Shixiong Zhao, Tianxiang Shen, Xusheng Chen, Tianyang Duan, Jia Pan 0001, Libo Zhang 0001, Heming Cui |
IEEE Internet Things J. | 2 |
| 2023 | Fold3D: Rethinking and Parallelizing Computational and Communicational Tasks in the Training of Large DNN ModelsabstractTraining a large DNN (e.g., GPT3) efficiently on commodity clouds is challenging even with the latest 3D parallel training systems (e.g., Megatron v3.0). In particular, along the pipeline parallelism dimension, computational tasks that produce a whole DNN's gradients with multiple input batches should be concurrently activated; along the data parallelism dimension, a set of heavy-weight communications (for aggregating the accumulated outputs of computational tasks) isinevitably serializedafter the pipelined tasks, undermining the training performance (e.g., in Megatron, data parallelism caused all GPUs idle for over 44% of the training time) over commodity cloud networks. To deserialize these communicational and computational tasks, we propose the AIAO scheduling (for 3D parallelism) which slices a DNN into multiple segments, so that the computational tasks processing the same DNN segment can be scheduled together, and the communicational tasks that synchronize this segment can be launched and overlapped (deserialized) with other segments’ computational tasks. We realized this idea in ourFold3Dtraining system. Extensive evaluation showsFold3Deliminated most of the all-GPU 44% idle time in Megatron (caused by data parallelism), leading to 25.2%–42.1% training throughput improvement compared to four notable baselines over various settings;Fold3D's high performance scaled to many GPUs. Fanxin Li, Shixiong Zhao, Yuhao Qing, Xusheng Chen, Xiuxian Guan, Sen Wang 0004, Gong Zhang 0001, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | ROG: A High Performance and Robust Distributed Training System for Robotic IoTabstractCritical robotic tasks such as rescue and disaster response are more prevalently leveraging ML (Machine Learning) models deployed on a team of wireless robots, on which data parallel (DP) training over Internet of Things of these robots (robotic IoT) can harness the distributed hardware resources to adapt their models to changing environments as soon as possible. Unfortunately, due to the need for DP synchronization across all robots, the instability in wireless networks (i.e., fluctuating bandwidth due to occlusion and varying communication distance) often leads to severe stall of robots, which affects the training accuracy within a tight time budget and wastes energy stalling. Existing methods to cope with the instability of datacenter networks are incapable of handling such straggler effect. That is because they are conducting model-granulated transmission scheduling, which is much more coarse-grained than the granularity of transient network instability in real-world robotic IoT networks, making a previously reached schedule mismatch with the varying bandwidth during transmission. We present ROG, the first ROw-Granulated distributed training system optimized for ML training over unstable wireless networks. ROG confines the granularity of transmission and synchronization to each row of a layer’s parameters and schedules the transmission of each row adaptively to the fluctuating bandwidth. In this way the ML training process can update partial and the most important gradients of a stale robot to avoid triggering stalls, while provably guaranteeing convergence. The evaluation shows that, given the same training time, ROG achieved about 4.9%~6.5% training accuracy gain compared with the baselines and saved 20.4%~50.7% of the energy to achieve the same training accuracy. Xiuxian Guan, Zekai Sun, Shengliang Deng, Xusheng Chen, Shixiong Zhao, Zongyuan Zhang, Tianyang Duan, Chenshu Wu, Yong Cui 0001, Libo Zhang 0001, Rui Wang 0007, Heming Cui |
MICRO | 1 |
| 2022 | vPipe: A Virtualized Acceleration System for Achieving Efficient and Scalable Pipeline Parallel DNN TrainingabstractThe increasing computational complexity of DNNs achieved unprecedented successes in various areas such as machine vision and natural language processing (NLP), e.g., the recent advanced Transformer has billions of parameters. However, as large-scale DNNs significantly exceed GPU's physical memory limit, they cannot be trained by conventional methods such as data parallelism. Pipeline parallelism that partitions a large DNN into small subnets and trains them on different GPUs is a plausible solution. Unfortunately, the layer partitioning and memory management in existing pipeline parallel systems are fixed during training, making them easily impeded by out-of-memory errors and the GPU under-utilization. These drawbacks amplify when performing neural architecture search (NAS) such as the evolved Transformer, where different network architectures of Transformer needed to be trained repeatedly. vPipe is the first system that transparently provides dynamic layer partitioning and memory management for pipeline parallelism. vPipe has two unique contributions, including (1) an online algorithm for searching a near-optimal layer partitioning and memory management plan, and (2) a live layer migration protocol for re-balancing the layer distribution across a training pipeline. vPipe improved the training throughput of two notable baselines (Pipedream and GPipe) by 61.4-463.4 percent and 24.8-291.3 percent on various large DNNs and training settings. Shixiong Zhao, Fanxin Li, Xusheng Chen, Xiuxian Guan, Jianyu Jiang, Dong Huang 0005, Yuhao Qing, Sen Wang 0004, Peng Wang 0037, Gong Zhang 0001, Cheng Li 0001, Ping Luo 0002, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | MVSAS: Semantic-Aware Scheduling for Low Latency and High Precision in Wireless Multi-View ApplicationabstractMulti-view models for various multi-view applications (e.g., pose recognition, facial recognition) achieve higher accuracy when more sensing data (views) from different sensors are flexibly collected via wireless networks and combined into inference input. However, when the view number scales up, the application suffers a long latency to collect all the latest views before inference (vanilla workflow). We observed that collecting all the latest views before inference is unnecessary, because different views are often not equally important and important views have major contribution to the output. In this paper, we present a Multi-View Semantic-Aware Scheduling (MVSAS) system that automatically prioritizes views according to their importance and schedules the early transmission of the important views. We tackled the challenge to infer view importance by analyzing the inference intermediates and extracting the semantics (e.g., number of persons) of each view. Once important views are collected, needless to wait for other less important views, the important views are combined with stale version of less important views as inference input, so as to retain high accuracy while reducing the latency to collect views. Evaluation shows that MVSAS achieved at most 36.9% latency reduction while retaining at most 98.7% accuracy compared to the vanilla workflow. Xiuxian Guan, Zekai Sun, Shengliang Deng, Shixiong Zhao, Tianxiang Shen, Tsz On Li, Rui Wang 0007, Heming Cui |
ICPADS | 1 |
| 2020 | CTDMA: Color-aware TDMA Network System For Low latency and High Throughput in Dense D2D Wireless NetworkabstractThe usage of IoT devices is rapidly growing recently and it raises the challenge of supporting a dynamic network of dense IoT devices with low latency and high throughput. Among the network multiple access protocols, Time Division Multiple Access(TDMA) is most suitable for this scene for its time efficiency and easy implementation. However, TDMA protocols typically perform poorly when confronted with the merging collision problem incurred especially in dynamic networks. Inspired by the coloring mechanism and spatial reuse introduced by 802.11ax, we created a cluster-based TDMA network system that can predict incoming collisions and avoid the collisions by rescheduling or extending the frame length before they actually happen. It can achieve zero-collision TDMA networks with optimal frame length and thus enhances the overall throughput and latency. Xiuxian Guan, Zekai Sun, Shengliang Deng, Heming Cui |
MASS | 1 |