Xi Chen 0104

dblp:16/3283-104 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0003-2168-9057ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An accurate and resource-efficient network for surface anomaly detection via enhanced downsampling and activation representation
Xunkuai Zhou, Xi Chen 0104, Jie Chen 0003, Ben M. Chen
Adv. Eng. Informatics2
2025 SE-STDGNN: A Self-Evolving Spatial-Temporal Directed Graph Neural Network for Multi-Vehicle Trajectory Prediction
abstract
Vehicle trajectory prediction (VTP) is essential for microscopic traffic risk assessment, autonomous vehicle navigation, and traffic behavior analysis. Related research leveraging learning-based methodologies has yielded notable success on various benchmark trajectory datasets. However, these models often experience performance degradation when faced with dynamic changes in traffic conditions such as vehicle density, road types, and weather conditions, as they have not been exposed to these variations during the training process. To effectively address the need for real-time adaptation in dynamic traffic scenarios, we propose a novel framework titled self-evolving spatial-temporal directed graph neural network (SE-STDGNN). This model utilizes evolving graph convolution networks (EvolveGCNs) to aggregate spatial-temporal features of vehicles and their neighbors, which are then utilized by a trajectory prediction module to forecast future trajectories. Further, a self-evolving mechanism is introduced to adjust model parameters dynamically in the real-time operation. The efficacy of SE-STDGNN is validated using the public vehicle trajectory dataset AD4CHE.
Bingxin Han, Yijun Huang, Xi Chen 0104, Ben M. Chen
ICRA4
2025 Multi-View Stereo with Geometric Encoding for Dense Scene Reconstruction
abstract
Multi-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometric encoding for more accurate and complete depth estimation and point cloud reconstruction. First, the cross-view adaptive cost volume aggregation module is proposed to strengthen multi-view geometric cues encoding during cost volume construction. Then, the depth consistency optimization is performed in the 3D point space during learning by invoking ground-truth depth cues from adjacent views. Finally, the surface normal geometries are explicitly encoded to refine the sampled depth hypotheses to be consistent in the local neighbor regions. Extensive experiments on the standard MVS benchmarks including DTU, Tanks and Temples, and BlendedMVS demonstrate the state-of-the-art depth estimation and point cloud reconstruction performance of GE-MVS. The GE-MVS is further deployed in real-world experiments for UAV-based large-scale reconstruction, where our method outperforms the prevalent industrial reconstruction solutions concerning reconstruction efficiency and efficacy. Our project page is: https://cuhk-usr-group.github.io/GE-MVS/
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen
ICRA8
2025 End-to-End Underwater Multi-View Stereo for Dense Scene Reconstruction
abstract
Recent advancements in learning-based multi-view stereo (MVS) have demonstrated significant improvements over traditional counterpart, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and present the first large-scale UwMVS dataset for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on our dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, code and appendix are available at: https://cuhk-usr-group.github.io/UwMVS/
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen
ICRA7
2025 Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time Inspection
abstract
Defect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for unmanned aerial vehicles (UAVs). Given the significant challenges in automatically detecting defects of varying intensity and size within complex infrastructure, along with the tendency of lightweight models to lose detail and fail to fully capture features during the defect extraction process, we propose the CUPID_Block, a multi-level information fusion block to construct the backbone, featuring the CUPID_Conv module equipped with our proposed CCA (CrissCross Attention). Furthermore, CUPID features an auxiliary training branch that assimilates lower feature maps, helping to recover details lost in deeper convolutional layers. To verify the effectiveness of CUPID and to address the lack of a suitable dataset in the community, we establish a multi-scenario infrastructure defect dataset, CUBIT2024, to conduct extensive experiments. Finally, to assess the efficiency and adaptability of CUPID in UAV for online infrastructure inspection, we design a compact autonomous drone, CU-Astro, where the proposed CUPID is deployed on the Jetson Orin NX computer onboard to evaluate the speed and power consumption of the inference.
Benyun Zhao, Qigeng Duan, Guidong Yang, Jerry Tang, Zhenbo Song, Junjie Wen 0001, Xuchen Liu 0001, Qingxiang Li, Lei Lei 0010, Jihan Zhang, Xi Chen 0104, Mark W. Mueller, Ben M. Chen
ICRA11
2025 Towards interpretable and robust UAV-based foundation model for endangered species monitoring in complex ecosystems
Jihan Zhang, Mingqiao Han, K. H. Laurie, Benyun Zhao, Lei Lei 0010, Xi Chen 0104, Hon Chi Judy Wan, Siu Gin Cheung, Wenxing Hong, Ben M. Chen
Mach. Learn.6
2025 Multi-View Stereo With Geometric Encoding for Large-Scale Dense Scene Reconstruction
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Yun-Hui Liu 0001, Ben M. Chen
IEEE Trans Autom. Sci. Eng.6
2025 Sparse-to-Dense Prediction of Ocean Subsurface Temperature Using Multilevel Spatiotemporal Information Fusion
abstract
Accurately predicting ocean subsurface temperature is vital for advancing ocean and climate research, particularly given the sparse and costly nature of subsurface observations. This study introduces sparse-to-dense prediction of ocean subsurface temperature using multi-level spatiotemporal (ST) information fusion. The framework integrates interpretable ST decoupling, adaptive feature updating, and sparse-to-dense information fusion modules to address the challenge of sparse observations and ever-evolving dynamic environments. Comprehensive experiments focused on the Pacific demonstrate the superiority of the proposed methodology over peer methods. The proposed methodology achieves high-resolution predictions with a root mean square error of 0.2230, accuracy of 0.9846, and point-wise prediction errors below 0.5°C under 10% online random sparse observations (ORSO). Analyses of spatial and temporal temperature dynamics reveal long-term warming trends in the Pacific, including a temperature rise of up to 2.8°C at -100 m in low-latitude regions over the past 40 years, and identify the latitudinal slope of thermocline dynamics. This study advances the understanding of multi-scale thermal processes and variability in the Pacific, demonstrating the potential of application in climate studies, marine resource management, and environmental monitoring.
Lei Lei 0010, Guidong Yang, Zuoquan Zhao, Xi Chen 0104, Ben M. Chen
IEEE Trans. Geosci. Remote. Sens.4
2025 A Semi-Supervised Domain-Adaptive Framework for Real-World Underwater Image Enhancement
abstract
Underwater optical remote sensing is crucial for geoscience applications but often suffers from image degradation due to complex underwater environments. While learning-based methods have advanced underwater image enhancement (UIE), their efficacy in real-world UIE applications still faces challenges. This limitation arises from training predominantly on synthetic underwater images, resulting in a significantinter-domain gap when applied to real-world data. Additionally, diverse underwater conditions introduceintra-domain challenges, such as color casts and haze, further complicating the UIE process. To address these issues, we propose SSD-UIE, a semi-supervised domain-adaptive framework designed to mitigate bothinter- andintra-domain gaps. Our approach employs a systematic synthesis pipeline to reduce visualinter-domain discrepancies and introduces a Large Synthetic-Real Underwater Image Dataset (LSRUID) to facilitate the training of the framework. The Semantic-Blender is developed to handle semanticinter-domain differences, while the Intra-domain-aware Feature Extraction (IFE) branch and feature alignment strategy effectively addressintra-domain variability. Furthermore, the Dual-Trans Block is introduced to enhance the UIE performance while maintaining computational efficiency. Extensive experiments demonstrate that SSD-UIE outperforms state-of-the-art (SOTA) UIE methods in both qualitative and quantitative evaluations on real-world underwater images. Codes and dataset will be publicly available at https://github.com/RockWenJJ/SSD-UIE.git.
Junjie Wen 0001, Guidong Yang, Benyun Zhao, Dongyue Huang, Lei Lei 0010, Bo Zhang 0019, Zhi Gao 0005, Xi Chen 0104, Ben M. Chen
IEEE Trans. Geosci. Remote. Sens.8
2025 Toward End-to-End Underwater Multi-View Stereo for Real-World Dense Scene Reconstruction
abstract
Multi-view stereo (MVS) enables accurate and complete 3D reconstruction from multi-view imagery, serving as a core methodology in remote sensing applications across terrestrial and underwater domains. Recent advancements in learning-based MVS have demonstrated significant improvements over traditional counterparts, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and presenting the first large-scale synthetic UwMVS dataset preserving real-world underwater degradation properties for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on the dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, appendix, and supplementary video are available at https://yang-sober.github.io/UnderMVS/.
Guidong Yang, Junjie Wen 0001, Lei Lei 0010, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Zhi Gao 0005, Ben M. Chen
IEEE Trans. Geosci. Remote. Sens.6
2024 The optimization of Building Integrated Photovoltaics Systems in Urban Environments
abstract
The utilization of Building Integrated Photovoltaic (BIPV) systems in high-rise urban buildings is limited despite their potential benefits in reducing electricity costs and greenhouse gas emissions. This paper introduces an advanced methodology that comprehensively evaluates BIPV system performance in high-rise buildings. The approach considers facade orientations, rooftop areas, and PV panel types to assess energy generation potential under different scenarios. Microclimate data obtained using the Urban Weather Generator (UWG) is incorporated for a thorough analysis. A case study on a super high-rise building in Hong Kong demonstrates the effectiveness of the proposed method and highlights the BIPV system's potential. The paper emphasizes the importance of optimizing BIPV systems in urban settings by considering various factors to maximize their effectiveness.
Chi Chung Lee 0001, Elena Bian, Fanny Tang, Chi Ho Li, Chun Yin Li, Xi Chen 0104
IECON6
2024 Det-Recon-Reg: An Intelligent Framework Towards Automated Large-Scale Infrastructure Inspection
abstract
Visual inspection plays a predominant role in inspecting infrastructure surface. However, the generalization of existing visual inspection systems to large-scale real-world scenes remains challenging. In this paper, we introduce Det-Recon-Reg, an intelligent framework separating the complex inspection procedure into three stages: Detect, Reconstruct, and Register. (1) For defect detection (Detect), we present the first high-resolution defect dataset tailored for large-scale defect detection. Based on the dataset, we evaluate the most effective real-time object detection algorithms and push the boundary by proposing CUBIT-Net for real-world defect inspection. (2) For infrastructure reconstruction (Reconstruct), we propose a learning-based multi-view stereo (MVS) network to adapt to large-scale scenes, taking as input the multi-view images and outputting the point cloud reconstruction, where its performance has been validated on the standard MVS datasets, including BlendedMVS, DTU, and Tanks and Temples datasets. (3) For defect localization (Register), we propose an effective registration method based on the geographic information system that registers the detected defects onto the reconstructed infrastructure model to establish a global reference for maintenance measures. The real-world experiments further verify the effectiveness and efficiency of our proposed framework. More details about our proposed dataset, code, and appendix are available on our project page: https://cuhk-usr-group.github.io/large-scale-inspect-framework/.
Guidong Yang, Jihan Zhang, Benyun Zhao, Chuanxiang Gao, Yijun Huang, Junjie Wen 0001, Qingxiang Li, Jerry Tang, Xi Chen 0104, Ben M. Chen
IROS9
2023 Multi-View Stereo with Learnable Cost Metric
abstract
In this paper, we present LCM-MVSNet, a novel multi-view stereo (MVS) network with learnable cost metric (LCM) for more accurate and complete depth estimation and dense point cloud reconstruction. To adapt to the scene variation and improve the reconstruction quality in non-Lambertian low-textured scenes, we propose LCM to adaptively aggregate multi-view matching similarity into the 3D cost volume by leveraging sparse points hints. The proposed LCM benefits the MVS approaches in four folds, including depth estimation enhancement, reconstruction quality improvement, memory footprint reduction, and computational burden alleviation, allowing the depth inference for high-resolution images to achieve more accurate and complete reconstruction. Moreover, we improve the depth estimation by enhancing the propagation of shallow features via a bottom-up path and strengthen the end-to-end supervision by adapting the focal loss to reduce ambiguity caused by sample imbalance. Extensive experiments on two benchmark datasets show that our network achieves state-of-the-art performance on the DTU dataset and exhibits strong generalization ability with a competitive performance on the Tanks and Temples benchmark. Furthermore, we deploy our LCM-MVSNet into the real-world application for large-scale 3D reconstruction based on multi-view aerial images collected by self-developed UAV, demonstrating the robustness and scalability of our method. More detailed results are available in the Appendix11shorturl.at/rBG28
Guidong Yang, Xunkuai Zhou, Chuanxiang Gao, Benyun Zhao, Jihan Zhang, Xi Chen 0104, Ben M. Chen
IROS7