Guoxuan Chi

dblp:223/8076 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-6941-1612ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Generative AI for Wireless Communication and Sensing: Toward Unified Foundation Models
Zheng Yang 0002, Guoxuan Chi, Chenshu Wu, Yuchong Gao, Yunhao Liu 0001, Yonina C. Eldar, Jie Xu 0002, Tony Xiao Han
IEEE Trans. Commun.2
2025 Chameleon: Fast-Slow Neuro-Symbolic Lane Topology Extraction
abstract
Lane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving. This task requires complex reasoning, such as determining whether it is possible to turn left into a specific lane. To address this challenge, we introduce neuro-symbolic methods powered by vision-language foundation models (VLMs). Existing approaches have notable limitations: (1) Dense visual prompting with VLMs can achieve strong performance but is costly in terms of both financial resources and carbon footprint, making it impractical for robotics applications. (2) Neuro-symbolic reasoning methods for 3D scene understanding fail to integrate visual inputs when synthesizing programs, making them ineffective in handling complex corner cases. To this end, we propose a fast-slow neuro-symbolic lane topology extraction algorithm, named Chameleon, which alternates between a fast system that directly reasons over detected instances using synthesized programs and a slow system that utilizes a VLM with a chain-of-thought design to handle corner cases. Chameleon leverages the strengths of both approaches, providing an affordable solution while maintaining high performance. We evaluate the method on the OpenLane-V2 dataset, showing consistent improvements across various baseline detectors. Our code, data, and models are publicly available at https://github.com/XR-Lee/neural-symbolic
Zongzheng Zhang, Xinrun Li, Sizhe Zou, Guoxuan Chi, Siqi Li 0009, Xuchong Qiu, Guoliang Wang 0002, Guantian Zheng, Leichen Wang, Hang Zhao 0021, Hao Zhao 0002
ICRA4
2025 Delving into Mapping Uncertainty for Mapless Trajectory Prediction
abstract
Recent advances in autonomous driving are moving towards mapless approaches, where High-Definition (HD) maps are generated online directly from sensor data, reducing the need for expensive labeling and maintenance. However, the reliability of these online-generated maps remains uncertain. While incorporating map uncertainty into downstream trajectory prediction tasks has shown potential for performance improvements, current strategies provide limited insights into the specific scenarios where this uncertainty is beneficial. In this work, we first analyze the driving scenarios in which mapping uncertainty has the greatest positive impact on trajectory prediction and identify a critical, previously overlooked factor: the agent’s kinematic state. Building on these insights, we propose a novel Proprioceptive Scenario Gating that adaptively integrates map uncertainty into trajectory prediction based on forecasts of the ego vehicle’s future kinematics. This lightweight, self-supervised approach enhances the synergy between online mapping and trajectory prediction, providing interpretability around where uncertainty is advantageous and outperforming previous integration methods. Additionally, we introduce a Covariance-based Map Uncertainty approach that better aligns with map geometry, further improving trajectory prediction. Extensive ablation studies confirm the effectiveness of our approach, achieving up to 23.6% improvement in mapless trajectory prediction performance over the state-of-the-art method using the real-world nuScenes driving dataset. Our code, data, and models are publicly available at https://github.com/Ethan-Zheng136/Map-Uncertainty-for-Trajectory-Prediction.
Zongzheng Zhang, Xuchong Qiu, Boran Zhang, Guantian Zheng, Xunjiang Gu, Guoxuan Chi, Huan-ang Gao, Leichen Wang, Xinrun Li, Igor Gilitschenski, Hongyang Li 0001, Hang Zhao 0021, Hao Zhao 0002
IROS6
2025 RF-Prox: Radio-Based Proximity Estimation of Nondirectly Connected Devices
abstract
Recent years have witnessed an increasing number of mobile devices, posing a more diversified demand for device localization solutions. Existing methods can locate connected devices but fail to address the spatial proximity between devices lacking direct communication links. This limitation impedes numerous emerging applications, such as implicit control of IoT device and proximity-based autonomous aerial vehicles scheduling. In response to this technical challenge, we introduce RF-Prox, the pioneering system designed for the proximity estimation of nondirectly connected devices. RF-Prox determines the proximity between devices by extracting and analyzing the spatiotemporal correlation between two signals. RF-Prox introduces a multiresolution spatiotemporal encoder (MRSTE) that extracts multiscale features from complex-valued wireless signals, capturing both spatial and dynamic temporal characteristics. Additionally, the proximity metric adaptation network (PMAN) bridges the gap between high-dimensional signal characteristics and physical proximity. To enhance scalability, we leverage a transfer learning framework, significantly reducing the need for extensive data collection and retraining. Extensive experiments demonstrate RF-Prox’s outstanding performance across Wi-Fi and cellular networks, achieving fine-tuned accuracy rates of 98.6% indoors and 91.3% outdoors. Even without fine-tuning, the pretrained model achieves strong zero-shot performance, showcasing its exceptional performance in both proximity estimation accuracy and domain generalizability.
Yuchong Gao, Guoxuan Chi, Zheng Yang 0002, Shijie Cheng, Zhiqing Wei
IEEE Internet Things J.2
2024 RF-Diffusion: Radio Signal Generation via Time-Frequency Diffusion
abstract
Along with AIGC shines in CV and NLP, its potential in the wireless domain has also emerged in recent years. Yet, existing RF-oriented generative solutions are ill-suited for generating high-quality, time-series RF data due to limited representation capabilities. In this work, inspired by the stellar achievements of the diffusion model in CV and NLP, we adapt it to the RF domain and propose RF-Diffusion. To accommodate the unique characteristics of RF signals, we first introduce a novel Time-Frequency Diffusion theory to enhance the original diffusion model, enabling it to tap into the information within the time, frequency, and complex-valued domains of RF signals. On this basis, we propose a Hierarchical Diffusion Transformer to translate the theory into a practical generative DNN through elaborated design spanning network architecture, functional block, and complex-valued operator, making RF-Diffusion a versatile solution to generate diverse, high-quality, and time-series RF data. Performance comparison with three prevalent generative models demonstrates the RF-Diffusion's superior performance in synthesizing Wi-Fi and FMCW signals. We also showcase the versatility of RF-Diffusion in boosting Wi-Fi sensing systems and performing channel estimation in 5G networks.
Guoxuan Chi, Zheng Yang 0002, Chenshu Wu, Jingao Xu, Yuchong Gao, Yunhao Liu 0001, Tony Xiao Han
MobiCom1
2024 WiViD: Leveraging Wi-Fi and Vision for Depth Estimation via Multimodal Diffusion
abstract
Depth estimation is crucial for numerous applications, including autonomous driving, robotic navigation and aug-mented reality. Existing solutions based on LiDAR and mm Wave technologies are constrained by high deployment costs, while those utilizing monocular vision suffer from limited accuracy. To address these challenges, this paper proposes WiViD, a diffusion-based depth estimation system that leverages commercial Wi-Fi and vision. Diffusion models, with their ability to iteratively refine predictions, offer significant advantages in producing accurate and detailed estimations. We introduce a Multimodal Conditional Diffusion (MMCD) mechanism and design two encoding modules: the Complex-Valued CSI Encoder (CCE) and the Residual Image Encoder (RIE). These components fully exploit the spatio-temporal information inherent in Wi-Fi CSI and enable the effective fusion of Wi-Fi CSI and RGB image data, which results in high-precision and robust depth estimation. Experimental results in real-world scenarios demonstrate that WiViD out-performs state-of-the-art (SOTA) monocular methods, reducing the Absolute Relative Error (ARE) by 67.2 %, highlighting the advantages of WiViD in terms of accuracy and reliability.
Shijie Cheng, Yuchong Gao, Zheng Yang 0002, Guoxuan Chi, Tony Xiao Han
MSN4
2024 XFall: Domain Adaptive Wi-Fi-Based Fall Detection With Cross-Modal Supervision
abstract
Recent years have witnessed an increasing demand for human fall detection systems. Among all existing methods, Wi-Fi-based fall detection has become one of the most promising solutions due to its pervasiveness. However, when applied to a new domain, existing Wi-Fi-based solutions suffer from severe performance degradation caused by low generalizability. In this paper, we propose XFall, a domain-adaptive fall detection system based on Wi-Fi. XFall overcomes the generalization problem from three aspects. To advance cross-environment sensing, XFall exploits an environment-independent feature called speed distribution profile, which is irrelevant to indoor layout and device deployment. To ensure sensitivity across all fall types, an attention-based encoder is designed to extract the general fall representation by associating both the spatial and temporal dimensions of the input. To train a large model with limited amounts of Wi-Fi data, we design a cross-modal learning framework, adopting a pre-trained visual model for supervision during the training process. We implement and evaluate XFall on one of the latest commercial wireless products through a year-long deployment in real-world settings. The result shows XFall achieves an overall accuracy of 96.8%, with a miss alarm rate of 3.1% and a false alarm rate of 3.3%, outperforming the state-of-the-art solutions in both in-domain and cross-domain evaluation.
Guoxuan Chi, Guidong Zhang, Qiang Ma 0007, Zheng Yang 0002, Zhenguo Du, Houfei Xiao
IEEE J. Sel. Areas Commun.1
2023 Wi-Prox: Proximity Estimation of Non-Directly Connected Devices via Sim2Real Transfer Learning
abstract
Recent years have witnessed an increasing number of mobile devices, posing a more diversified demand for device localization solutions. While existing wireless localization solutions can obtain the relative locations of connected devices, they fall short in estimating the spatial relationships between devices that are not directly connected. To address this technical gap, we propose Wi-Prox, the first proximity estimation system for non-directly connected devices. Wi-Prox evaluates the spatial proximity of two devices by analyzing their received wireless signals. It integrates a novel multi-resolution spatial encoder that extracts multi-scale spatial features from complex-valued wireless signals, which are then analyzed and transformed into a domain-adaptive proximity metric. To enhance the general-izability of Wi-Prox, we adopt a simulation-to-reality transfer learning framework. Wi-Prox is pre-trained with a large amount of simulated data and then fine-tuned for real-world deployment, significantly reducing the need for real-world data collection. We implement Wi-Prox and evaluate its performance in both simulated and real environments. Our results indicate that a fine-tuned Wi-Prox achieves an average accuracy of 97.2% in selecting the most proximate device. Even without fine-tuning, a pre-trained Wi-Prox still manages an average accuracy of 93.8%, thereby demonstrating impressive performance in terms of both proximity estimation accuracy and domain generalizability.
Yuchong Gao, Guoxuan Chi, Guidong Zhang, Zheng Yang 0002
GLOBECOM2
2023 Locate, Tell, and Guide: Enabling Public Cameras to Navigate the Public
abstract
Indoor navigation is essential to a wide spectrum of applications in the era of mobile computing. Existing vision-based technologies suffer from both start-up costs and the absence of semantic information for navigation. We observe an opportunity to leverage pervasively deployed surveillance cameras to deal with the above drawbacks and revisit the problem of indoor navigation with a fresh perspective. In this paper, we proposeiSAT, a system that enables public surveillance cameras, as indoor navigating satellites, to locate users on the floorplan, tell users with semantic information about the surrounding environment, and guide users with navigation instructions. However, enabling public cameras to navigate is non-trivial due to 3 factors: absence of real scale, disparity of camera perspective, and lack of semantic information. To overcome these challenges,iSATleverages POI-assisted framework and adopts a novel coordinate transformation algorithm to associate public and mobile cameras, and further attaches semantic information to user location. Extensive experiments in 4 different scenarios show thatiSATachieves a localization accuracy of 0.48m and a navigation success rate of 90.5 percent, outperforming the state-of-th-art systems by$> 30\%$. Benefiting from our solution, all areas with public cameras can upgrade to smart spaces with visual navigation services.
Guoxuan Chi, Jingao Xu, Qian Zhang 0017, Qiang Ma 0007, Zheng Yang 0002
IEEE Trans. Mob. Comput.1
2023 Push the Limit of Millimeter-wave Radar Localization
abstract
Existing device-free localization systems have achieved centimeter-level accuracy and show their potential in a wide range of applications. However, today’s radio-based solutions fail to locate the target in millimeter-level due to their limited bandwidth and sampling rate, which constrains their applications in high-accuracy demand scenarios. We find an opportunity to break the bottleneck of existing radio-based localization systems by reconstructing the accurate signal spectral peak from the discrete samples, without changing either the bandwidth or the sampling rate of the radio hardware. This study proposes milliLoc , a millimeter-level radio-based localization system. We first derive a spectral peak reconstruction algorithm to reduce the ranging error from the previous centimeter-level to millimeter-level. Then, we improve the AoA measurement accuracy by leveraging the signal amplitude information. To ensure the practicality of milliLoc , we further extend our system to handle multi-target situations. We fully implement milliLoc on a commercial mmWave radar. Experiments show that milliLoc achieves a median ranging accuracy of 5.5 mm and decreases the AoA measurement error by 31.2% compared with the baseline. Our system fulfills the accuracy requirements of most application scenarios and can be easily integrated with other existing solutions, shedding light on high-accuracy location-based applications.
Guidong Zhang, Guoxuan Chi, Yi Zhang 0017, Zheng Yang 0002
ACM Trans. Sens. Networks2
2022 Wi-drone: wi-fi-based 6-DoF tracking for indoor drone flight control
abstract
After years of boom, drones and their applications are now entering indoors. Six-degree-of-freedom (6-DoF) pose tracking is the core of drone flight control, but existing solutions cannot be directly applied to indoor scenarios due to insufficient accuracy, low robustness to adverse texture and light conditions, and signal obstruction in indoor scenarios. To overcome the above limitations, we propose Wi-Drone, a Wi-Fi standalone 6-DoF tracking system for indoor drone flight control. Wi-Drone takes full advantage of both exte-roceptive and proprioceptive measurements of Wi-Fi to estimate the drone's absolute pose and relative motion, and fuse them in a tight-coupling manner to achieve their complementary benefits. We implement Wi-Drone and integrate it into a flight control system. The evaluation results show that Wi-Drone achieves a real-time performance with the average location accuracy of 26.1 cm and the rotation accuracy of 3.8°, which demonstrates its competency of flight control, compared to visual-inertial-based flight control. Such results also outperform existing Wi-Fi-based tracking solutions in terms of both dimensionality and accuracy.
Guoxuan Chi, Zheng Yang 0002, Jingao Xu, Chenshu Wu, Jianzhe Liang, Yunhao Liu 0001
MobiSys1
2021 FollowUpAR: enabling follow-up effects in mobile AR applications
abstract
Existing smartphone-based Augmented Reality (AR) systems are able to render virtual effects on static anchors. However, today's solutions lack the ability to render follow-up effects attached to moving anchors since they fail to track the 6 degrees of freedom (6-DoF) poses of them. We find an opportunity to accomplish the task by leveraging sensors capable of generating sparse point clouds on smartphones and fusing them with vision-based technologies. However, realizing this vision is non-trivial due to challenges in modeling radar error distributions and fusing heterogeneous sensor data. This study proposes FollowUpAR, a framework that integrates vision and sparse measurements to track object 6-DoF pose on smartphones. We derive a physical-level theoretical radar error distribution model based on an in-depth understanding of its hardware-level working principles and design a novel factor graph competent in fusing heterogeneous data. By doing so, FollowUpAR enables mobile devices to track anchor's pose accurately. We implement FollowUpAR on commodity smartphones and validate its performance with 800,000 frames in a total duration of 15 hours. The results show that FollowUpAR achieves a remarkable rotation tracking accuracy of 2.3° with a translation accuracy of 2.9mm, outperforming most existing tracking systems and comparable to state-of-the-art learning-based solutions. FollowUpAR can be integrated into ARCore and enable smartphones to render follow-up AR effects to moving objects.
Jingao Xu, Guoxuan Chi, Zheng Yang 0002, Danyang Li 0005, Qian Zhang 0017, Qiang Ma 0007
MobiSys2
2021 Enabling Surveillance Cameras to Navigate
abstract
Smartphone localization is essential to a wide spectrum of applications in the era of mobile computing. The ubiquity of smartphone mobile cameras and surveillance ambient cameras holds promise for offering sub-meter accuracy localization services thanks to the maturity of computer vision techniques. In general, ambient-camera-based solutions are able to localize pedestrians in video frames at fine-grained, but the tracking performance under dynamic environments remains unreliable. On the contrary, mobile-camera-based solutions are capable of continuously tracking pedestrians; however, they usually involve constructing a large volume of image database, a labor-intensive overhead for practical deployment. We observe an opportunity of integrating these two most promising approaches to overcome above limitations and revisit the problem of smartphone localization with a fresh perspective. However, fusing mobile-camera-based and ambient-camera-based systems is non-trivial due to disparity of camera in terms of perspectives, parameters and incorrespondence of localization results. In this article, we propose iMAC, an integrated mobile cameras and ambient cameras based localization system that achieves sub-meter accuracy and enhanced robustness with zero-human start-up effort. The key innovation of iMAC is a well-designed fusing frame to eliminate disparity of cameras including a construction of projection map function to automatically calibrate ambient cameras, an instant crowd fingerprints model to describe user motion patterns, and a confidence-aware matching algorithm to associate results from two sub-systems. We fully implement iMAC on commodity smartphones and validate its performance in five different scenarios. The results show that iMAC achieves a remarkable localization accuracy of 0.68 m, outperforming the state-of-the-art systems by >75%.
Jingao Xu, Guoxuan Chi, Danyang Li 0005, Xinglin Zhang 0001, Qiang Ma 0007, Zheng Yang 0002
ACM Trans. Sens. Networks3
2020 Enabling Surveillance Cameras to Navigate
abstract
Smartphone localization is essential to a wide spectrum of applications in the era of mobile computing. The ubiquity of smartphone mobile cameras and surveillance ambient cameras holds promise for offering sub-meter accuracy localization services thanks to the maturity of computer vision techniques. In general, ambient-camera-based solutions are able to localize pedestrians in video frames at fine-grained, but the tracking performance under dynamic environments remains unreliable. On the contrary, mobile-camera-based solutions are capable of continuously tracking pedestrians, however, they usually involve constructing a large volume of image database, a labor-intensive overhead for practical deployment. We observe an opportunity of integrating these two most promising approaches to overcome above limitations and revisit the problem of smartphone localization with a fresh perspective. However, fusing mobile-camera-based and ambient-camera-based systems is non-trivial due to disparity of camera in terms of perspectives, parameters and incorrespondence of localization results. In this paper, we propose iMAC, an integrated mobile cameras and ambient cameras based localization system that achieves sub-meter accuracy and enhanced robustness with zero-human start-up effort. The key innovation of iMAC is a well-designed fusing frame to eliminate disparity of cameras including a construction of projection map function to automatically calibrate ambient cameras, an instant crowd fingerprints model to describe user motion patterns, and a confidence-aware matching algorithm to associate results from two sub-systems. We fully implement iMAC on commodity smart-phones and validate its performance in five different scenarios. The results show that iMAC achieves a remarkable localization accuracy of 0.68m, outperforming the state-of-the-art systems by > 75%.
Jingao Xu, Guoxuan Chi, Danyang Li 0005, Xinglin Zhang 0001, Qiang Ma 0007, Zheng Yang 0002
ICCCN3
2018 Latency-Optimal Task Offloading for Mobile-Edge Computing System in 5G Heterogeneous Networks
abstract
Mobile edge computing (MEC) is an emerging technology to improve the quality of computation experience for mobile devices. As a promising paradigm to deal with latency-sensitive and computation-intensive tasks, it provides cloud computing capabilities in close proximity to mobile devices in the fifth-generation (5G) networks. As the radio and computational resources are both limited in 5G networks, reducing system latency by task scheduling and resource allocation has gained renewed interests. To minimize the weighted-sum latency of all users in multi-user MEC system, we formulate an optimization problem based on partial offloading strategy. Since the optimization problem is NP-hard, we transform it into a piece-wise convex problem and get the latency-optimal offloading strategy using the sub- gradient method. We further put forward a simplified algorithm which can achieve close-to- optimal performance in linear time. Our proposed strategies are verified by numerical results, which indicate that our algorithms significantly reduce the weighted-sum latency compared with other baseline strategies.
Guoxuan Chi, Yumei Wang, Xiang Liu 0017
VTC Spring1