Xuyang Bai

dblp:260/0505 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Energy-Efficient Simultaneous Wireless Information and Power Transfer System Design with Electromagnetic Modeling in Metal Cabinet
abstract
This paper proposes an energy-efficient simultaneous wireless information and power transfer (SWIPT) system for metal cabinet, aimed at supporting sustainable industrial wireless sensor networks (WSNs) as part of the envisioned 6G IoT connectivity. An innovative electromagnetic (EM) channel model is introduced, precisely characterizing electromagnetic propagation in this specific environment, including direct, scattered, and wall-reflected paths. The system combines hybrid power splitting (PS) and time switching (TS) with maximum ratio transmission (MRT) beamforming to optimize the trade-off between information rate and harvested power. For the two-stage task, the first stage maximizes the rate under a power constraint, while the second stage maximizes harvested power under a rate constraint. We propose a low-complexity adaptive algorithm that dynamically adjusts the PS and TS parameters to minimize energy wastage and fulfill task requirements. Simulations in a machine housing control cabinet demonstrate that the proposed design outperforms fixed-strategy counterparts, ensuring reliable communication and sustainable power delivery. This scalable solution aligns with green communication principles and is well-suited for industrial WSNs in constrained environments.
Youyang Xiang, Zhaoguo Ding, Chengjie Zhao, Xuyang Bai, Xianglu Li, Zhijiang Huang, Qilong Du, Jie Tian 0005
VTC2025-Fall5
2025 Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation of Indoor Scenes
abstract
In recent years, sparse voxel-based methods have become the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the euclidean and geodesic information. Intuitively, the euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% versus 72.5% and 73.6% in mIoU) with a simpler network structure (17M versus 30M and 38M parameters).
Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Characteristics of L-Band Microwave Scattering From Layered Rough Soil With a Full-Wave Volume Integral Equation Approach
abstract
Soil surfaces often exhibit moisture stratification as a result of natural processes such as precipitation and vertical infiltration, which is however commonly neglected in traditional L-band soil scattering modeling and inversion works. Understanding how different moisture stratifications affect microwave observations is therefore of critical importance. Existing soil scattering simulation algorithms face challenges in accurately modeling layered soils, especially when large lateral soil extents and internal roughness characteristics within the medium are involved. To address these challenges, this paper introduces a generalized and accurate soil scattering modeling method based on a numerical solution to Maxwell’s equations in the volume-integral form (NMM3D-VIE). To efficiently treat large lateral domains, the soil is modeled as laterally periodic, and an effective truncation in depth is implemented through half-space Green’s functions. Within each periodic unit, inhomogeneous layered moisture are constructed to represent realistic near-surface layered soil conditions. The VIE is solved using the discrete dipole approximation, and its volumetric discretization enables the handling of arbitrary layered configurations, including internal roughness. The proposed method is first compared against the surface integral equation approach (SIE) and the advanced integral equation method (AIEM) across various single-layer soil conditions, demonstrating its accuracy and broad applicability. Furthermore, we present a novel investigation of the scattering properties of various layered soil structures to emphasize the layering effects in soil scattering observations, including the bistatic scattering coefficient in the incidence plane and in the top hemisphere, as well as the emissivity. The feasibility of equivalent single-layer strategies for representing the scattering behavior of realistic layered soils is also investigated. Additionally, the phase characteristics for single-layer and multi-layer soil structures are also investigated for the first time, with the proposed full-wave approach. This advanced model provides theoretical support and guidance for future research in complex soil scattering, layered soil approximation models, and inversion methodologies.
Xuyang Bai, Shurun Tan
IEEE Trans. Geosci. Remote. Sens.1
2024 A Shared Computing Platform for Remote Sensing Community: The Framework Setup and user Interface
abstract
Retrieving key climate variables such as soil moisture and snow water equivalent from remote sensing data requires representative physical models. Up to date, there is no integrated remote sensing computing platform dedicated to modeling brightness temperature, backscatters, and relevant parameters based on microwave electromagnetic scattering mechanisms for complex soil, vegetation, and snow scenarios. In this paper, the Remote Sensing Hub (RSHub), a shared cloud computing platform is introduced to close the gap. The platform integrates multiple physical scattering models into a unified framework that supports soil/vegetation/snow scenarios, offering options for radiative transfer and full wave approaches. We demonstrate the use of the RSHub to predict brightness temperatures corresponding to vegetated land surface scenarios.
Yiwen Fang, Xuyang Bai, Yuanhao Cao, Shurun Tan
IGARSS3
2024 Vision-Centric BEV Perception: A Survey
abstract
In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being conducive to data fusion. The rapid advancements in deep learning have led to the proposal of numerous methods for addressing vision-centric BEV perception challenges. However, there has been no recent survey encompassing this novel and burgeoning research field. To catalyze future research, this paper presents a comprehensive survey of the latest developments in vision-centric BEV perception and its extensions. It compiles and organizes up-to-date knowledge, offering a systematic review and summary of prevalent algorithms. Additionally, the paper provides in-depth analyses and comparative results on various BEV perception tasks, facilitating the evaluation of future works and sparking new research directions. Furthermore, the paper discusses and shares valuable empirical implementation details to aid in the advancement of related algorithms.
Yuexin Ma, Xuyang Bai, Huitong Yang, Yuenan Hou, Yaming Wang, Yu Qiao 0001, Ruigang Yang, Xinge Zhu
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 One Training for Multiple Deployments: Polar-based Adaptive BEV Perception for Autonomous Driving
abstract
Current on-board chips usually have different computing power, which means multiple training processes are needed for adapting the same learning-based algorithm to different chips, costing huge computing resources. The situation becomes even worse for 3D perception methods with large models. Previous vision-centric 3D perception approaches are trained with regular grid-represented feature maps of fixed resolutions, which is not applicable to adapt to other grid scales, limiting wider deployment. In this paper, we leverage the Polar representation when constructing the BEV feature map from images in order to achieve the goal of training once for multiple deployments. Specifically, the feature along rays in Polar space can be easily adaptively sampled and projected to the feature in Cartesian space with arbitrary resolutions. To further improve the adaptation capability, we make multi-scale contextual information interact with each other to enhance the feature representation. Experiments on a large-scale autonomous driving dataset show that our method outperforms others as for the good property of one training for multiple deployments.
Huitong Yang, Xuyang Bai, Xinge Zhu, Yuexin Ma
ICRA2
2023 Layered Soil Remote Sensing With Multichannel Passive Microwave Observations Using a Physics-Embedded Artificial Intelligence Framework: A Theoretical Study
abstract
The vertical distribution of soil properties is crucial in accurately representing various hydrological and ecological processes such as freeze-thaw cycles and diurnal variations. In this paper, considering the complexity of the multi-parameter features of layered soil, we evaluate the potential to retrieve the vertical distribution of the moisture and temperature of soil using multi-channel passive microwave observations. To enhance the inversion efficiency and accuracy, a novel Physics-Embedded Artificial Neural Network (P-ANN) inversion algorithm combining multi-angle (30 to 50 degree), multi-frequency (L-, C-, and X-band), and multi-polarization (horizontal and vertical polarization) passive observations is proposed. In this approach, the multi-channel physical brightness temperature simulations corresponding to the predicted soil state parameters are integrated into the loss function of a standard fully connected feed-forward neural network, enabling efficient convergence with limited sampling data in the training process. Testing results exhibit that the inversion performance of P-ANN is superior to that of conventional neural network approaches which only adopts errors in soil states in the loss function to train the network. Test also shows the proposed P-ANN approach outperforms traditional optimization algorithms in dealing with layered soil retrieval. In order to further improve the retrieval accuracy, an advanced local optimization scheme is also proposed, where the output from P-ANN is further treated as the initial value to a local optimization algorithm, achieving even closer results to the ground truths without excessive computational costs. In addition, to estimate the reliability of the model predictions, this paper also establishes, in the testing process, a statistical relationship between the soil inversion error and the error of the corresponding brightness temperatures. When the trained neural network is in operation, the error of brightness temperature is calculated through the physical model, and the reliability of retrieval soil results is then acquired by putting the calculated brightness temperature errors into the pre-established statistical relationship. The proposed concepts and approaches have demonstrated the feasibility of using P-ANN model with multidimensional observations to invert the multi-variable layered soil structures. The proposed approach holds great potential for various remote sensing applications as well as solving a wide range of inverse problem challenges.
Xuyang Bai, Shurun Tan
IEEE Trans. Geosci. Remote. Sens.1
2022 TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers
abstract
LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are easily affected by such conditions, mainly due to a hard association of LiDAR points and image pixels, established by calibration matrices. We propose TransFusion, a robust solution to LiDAR-camera fusion with a soft-association mechanism to handle inferior image conditions. Specifically, our TransFusion consists of convolutional backbones and a detection head based on a transformer decoder. The first layer of the decoder predicts initial bounding boxes from a LiDAR point cloud using a sparse set of object queries, and its second decoder layer adaptively fuses the object queries with useful image features, leveraging both spatial and contextual relationships. The attention mechanism of the transformer enables our model to adaptively determine where and what information should be taken from the image, leading to a robust and effective fusion strategy. We additionally design an image-guided query initialization strategy to deal with objects that are difficult to detect in point clouds. TransFusion achieves state-of-the-art performance on large-scale datasets. We provide extensive experiments to demonstrate its robustness against degenerated image quality and calibration errors. We also extend the proposed method to the 3D tracking task and achieve the 1st place in the leader-board of nuScenes tracking, showing its effectiveness and generalization capability. [code release]
Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Hongbo Fu 0001, Chiew-Lan Tai
CVPR1
2022 LiDAL: Inter-frame Uncertainty Based Active Learning for 3D LiDAR Semantic Segmentation
Zeyu Hu, Xuyang Bai, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
ECCV (27)2
2021 PointDSC: Robust Point Cloud Registration Using Deep Spatial Consistency
abstract
Removing outlier correspondences is one of the critical steps for successful feature-based point cloud registration. Despite the increasing popularity of introducing deep learning techniques in this field, spatial consistency, which is essentially established by a Euclidean transformation between point clouds, has received almost no individual attention in existing learning frameworks. In this paper, we present PointDSC, a novel deep neural network that explicitly incorporates spatial consistency for pruning outlier correspondences. First, we propose a nonlocal feature aggregation module, weighted by both feature and spatial coherence, for feature embedding of the input correspondences. Second, we formulate a differentiable spectral matching module, supervised by pairwise spatial compatibility, to estimate the inlier confidence of each correspondence from the embedded features. With modest computation cost, our method outperforms the state-of-the-art hand- crafted and learning-based outlier rejection approaches on several real-world datasets by a significant margin. We also show its wide applicability by combining PointDSC with different 3D local descriptors. [code release]
Xuyang Bai, Zixin Luo, Lei Zhou 0011, Lei Li 0038, Zeyu Hu, Hongbo Fu 0001, Chiew-Lan Tai
CVPR1
2021 Learning to Match Features with Seeded Graph Matching Network
abstract
Matching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network consists of 1) Seeding Module, which initializes the matching by generating a small set of reliable matches as seeds. 2) Seeded Graph Neural Network, which utilizes seed matches to pass messages within/across images and predicts assignment costs. Three novel operations are proposed as basic elements for message passing: 1) Attentional Pooling, which aggregates keypoint features within the image to seed matches. 2) Seed Filtering, which enhances seed features and exchanges messages across images. 3) Attentional Unpooling, which propagates seed features back to original keypoints. Experiments show that our method reduces computational and memory complexity significantly compared with typical attention-based networks while competitive or higher performance is achieved.
Zixin Luo, Lei Zhou 0011, Xuyang Bai, Zeyu Hu, Chiew-Lan Tai, Long Quan
ICCV5
2021 VMNet: Voxel-Mesh Network for Geodesic-Aware 3D Semantic Segmentation
abstract
In recent years, sparse voxel-based methods have be-come the state-of-the-arts for 3D semantic segmentation of indoor scenes, thanks to the powerful 3D CNNs. Nevertheless, being oblivious to the underlying geometry, voxel-based methods suffer from ambiguous features on spatially close objects and struggle with handling complex and irregular geometries due to the lack of geodesic information. In view of this, we present Voxel-Mesh Network (VMNet), a novel 3D deep architecture that operates on the voxel and mesh representations leveraging both the Euclidean and geodesic information. Intuitively, the Euclidean information extracted from voxels can offer contextual cues representing interactions between nearby objects, while the geodesic information extracted from meshes can help separate objects that are spatially close but have disconnected surfaces. To incorporate such information from the two domains, we design an intra-domain attentive module for effective feature aggregation and an inter-domain attentive module for adaptive feature fusion. Experimental results validate the effectiveness of VMNet: specifically, on the challenging ScanNet dataset for large-scale segmentation of indoor scenes, it outperforms the state-of-the-art SparseConvNet and MinkowskiNet (74.6% vs 72.5% and 73.6% in mIoU) with a simpler network structure (17M vs 30M and 38M parameters). Code release: https://github.com/hzykent/VMNet
Zeyu Hu, Xuyang Bai, Jiaxiang Shang, Jiayu Dong, Xin Wang 0178, Guangyuan Sun, Hongbo Fu 0001, Chiew-Lan Tai
ICCV2
2021 Selection of product recycling channels based on extended TODIM method
Xianpei Hong, Xuyang Bai
Expert Syst. Appl.2
2020 D3Feat: Joint Learning of Dense Detection and Description of 3D Local Features
abstract
A successful point cloud registration often lies on robust establishment of sparse matches through discriminative 3D local features. Despite the fast evolution of learning-based 3D feature descriptors, little attention has been drawn to the learning of 3D feature detectors, even less for a joint learning of the two tasks. In this paper, we leverage a 3D fully convolutional network for 3D point clouds, and propose a novel and practical learning mechanism that densely predicts both a detection score and a description feature for each 3D point. In particular, we propose a keypoint selection strategy that overcomes the inherent density variations of 3D point clouds, and further propose a self-supervised detector loss guided by the on-the-fly feature matching results during training. Finally, our method achieves state-of-the-art results in both indoor and outdoor scenarios, evaluated on 3DMatch and KITTI datasets, and shows its strong generalization ability on the ETH dataset. Towards practical use, we show that by adopting a reliable feature detector, sampling a smaller number of features is sufficient to achieve accurate and fast point cloud alignment.
Xuyang Bai, Zixin Luo, Lei Zhou 0011, Hongbo Fu 0001, Long Quan, Chiew-Lan Tai
CVPR1
2020 ASLFeat: Learning Local Features of Accurate Shape and Localization
abstract
This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to acquire stronger geometric invariance. Second, the localization accuracy of detected keypoints is not sufficient to reliably recover camera geometry, which has become the bottleneck in tasks such as 3D reconstruction. In this paper, we present ASLFeat, with three light-weight yet effective modifications to mitigate above issues. First, we resort to deformable convolutional networks to densely estimate and apply local transformation. Second, we take advantage of the inherent feature hierarchy to restore spatial resolution and low-level details for accurate keypoint localization. Finally, we use a peakiness measurement to relate feature responses and derive more indicative detection scores. The effect of each modification is thoroughly studied, and the evaluation is extensively conducted across a variety of practical scenarios. State-of-the-art results are reported that demonstrate the superiority of our methods.
Zixin Luo, Lei Zhou 0011, Xuyang Bai, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan
CVPR3
2020 JSENet: Joint Semantic Segmentation and Edge Detection Network for 3D Point Clouds
Zeyu Hu, Mingmin Zhen, Xuyang Bai, Hongbo Fu 0001, Chiew-Lan Tai
ECCV (20)3