Benyun Zhao

dblp:325/6363 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
15since 2021 · last 2026
0009-0008-7298-2927ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AHD-Net: A damage detector and dataset for architectural heritage
Lingege Long, Benyun Zhao, Zhenkun Gan, Dayu Zhang, Qingxiang Li
Expert Syst. Appl.2
2025 Multi-View Stereo with Geometric Encoding for Dense Scene Reconstruction
abstract
Multi-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometric encoding for more accurate and complete depth estimation and point cloud reconstruction. First, the cross-view adaptive cost volume aggregation module is proposed to strengthen multi-view geometric cues encoding during cost volume construction. Then, the depth consistency optimization is performed in the 3D point space during learning by invoking ground-truth depth cues from adjacent views. Finally, the surface normal geometries are explicitly encoded to refine the sampled depth hypotheses to be consistent in the local neighbor regions. Extensive experiments on the standard MVS benchmarks including DTU, Tanks and Temples, and BlendedMVS demonstrate the state-of-the-art depth estimation and point cloud reconstruction performance of GE-MVS. The GE-MVS is further deployed in real-world experiments for UAV-based large-scale reconstruction, where our method outperforms the prevalent industrial reconstruction solutions concerning reconstruction efficiency and efficacy. Our project page is: https://cuhk-usr-group.github.io/GE-MVS/
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen
ICRA4
2025 End-to-End Underwater Multi-View Stereo for Dense Scene Reconstruction
abstract
Recent advancements in learning-based multi-view stereo (MVS) have demonstrated significant improvements over traditional counterpart, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and present the first large-scale UwMVS dataset for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on our dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, code and appendix are available at: https://cuhk-usr-group.github.io/UwMVS/
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen
ICRA3
2025 Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time Inspection
abstract
Defect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for unmanned aerial vehicles (UAVs). Given the significant challenges in automatically detecting defects of varying intensity and size within complex infrastructure, along with the tendency of lightweight models to lose detail and fail to fully capture features during the defect extraction process, we propose the CUPID_Block, a multi-level information fusion block to construct the backbone, featuring the CUPID_Conv module equipped with our proposed CCA (CrissCross Attention). Furthermore, CUPID features an auxiliary training branch that assimilates lower feature maps, helping to recover details lost in deeper convolutional layers. To verify the effectiveness of CUPID and to address the lack of a suitable dataset in the community, we establish a multi-scenario infrastructure defect dataset, CUBIT2024, to conduct extensive experiments. Finally, to assess the efficiency and adaptability of CUPID in UAV for online infrastructure inspection, we design a compact autonomous drone, CU-Astro, where the proposed CUPID is deployed on the Jetson Orin NX computer onboard to evaluate the speed and power consumption of the inference.
Benyun Zhao, Qigeng Duan, Guidong Yang, Jerry Tang, Zhenbo Song, Junjie Wen 0001, Xuchen Liu 0001, Qingxiang Li, Lei Lei 0010, Jihan Zhang, Xi Chen 0104, Mark W. Mueller, Ben M. Chen
ICRA1
2025 FHGS: Feature-Homogenized Gaussian Splatting
abstract
Scene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic color representation of gaussian primitives and the isotropic requirements of semantic features, leading to insufficient cross-view feature consistency. To overcome the limitation, we proposes FHGS (Feature-Homogenized Gaussian Splatting), a novel 3D feature distillation framework inspired by physical models, which freezes and distills 2D pre-trained features into 3D representations while preserving the real-time rendering efficiency of 3DGS. Specifically, our FHGS introduces the following innovations: Firstly, a universal feature fusion architecture is proposed, enabling robust embedding of large-scale pre-trained models' semantic features (e.g., SAM, CLIP) into sparse 3D structures. Secondly, a non-differentiable feature fusion mechanism is introduced, which enables semantic features to exhibit viewpoint independent isotropic distributions. This fundamentally balances the anisotropic rendering of gaussian primitives and the isotropic expression of features; Thirdly, a dual-driven optimization strategy inspired by electric potential fields is proposed, which combines external supervision from semantic feature fields with internal primitive clustering guidance. This mechanism enables synergistic optimization of global semantic alignment and local structural consistency. Extensive comparison experiments with other state-of-the-art methods on benchmark datasets demonstrate that our FHGS exhibits superior reconstruction performance in feature fusion, noise suppression, and geometric precision, while maintaining a significantly lower training time. This work establishes a novel Gaussian Splatting data structure, offering practical advancements for real-time semantic mapping, 3D stylization, and Vision-Language Navigation (VLN). Our code and additional results are available on our project page:https://fhgs.cuastro.org/.
Qigeng Duan, Benyun Zhao, Mingqiao Han, Yijun Huang, Ben M. Chen
NeurIPS2
2025 V2DGS:Visual Voxel Map-Based 2D Gaussian Splatting for Accurate Outdoor Reconstruction
Zhenbo Song, Qigeng Duan, Benyun Zhao, Jianfeng Lu 0003
PRCV (10)4
2025 Towards interpretable and robust UAV-based foundation model for endangered species monitoring in complex ecosystems
Jihan Zhang, Mingqiao Han, K. H. Laurie, Benyun Zhao, Lei Lei 0010, Xi Chen 0104, Hon Chi Judy Wan, Siu Gin Cheung, Wenxing Hong, Ben M. Chen
Mach. Learn.4
2025 Multi-View Stereo With Geometric Encoding for Large-Scale Dense Scene Reconstruction
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Yun-Hui Liu 0001, Ben M. Chen
IEEE Trans Autom. Sci. Eng.4
2025 A Semi-Supervised Domain-Adaptive Framework for Real-World Underwater Image Enhancement
abstract
Underwater optical remote sensing is crucial for geoscience applications but often suffers from image degradation due to complex underwater environments. While learning-based methods have advanced underwater image enhancement (UIE), their efficacy in real-world UIE applications still faces challenges. This limitation arises from training predominantly on synthetic underwater images, resulting in a significantinter-domain gap when applied to real-world data. Additionally, diverse underwater conditions introduceintra-domain challenges, such as color casts and haze, further complicating the UIE process. To address these issues, we propose SSD-UIE, a semi-supervised domain-adaptive framework designed to mitigate bothinter- andintra-domain gaps. Our approach employs a systematic synthesis pipeline to reduce visualinter-domain discrepancies and introduces a Large Synthetic-Real Underwater Image Dataset (LSRUID) to facilitate the training of the framework. The Semantic-Blender is developed to handle semanticinter-domain differences, while the Intra-domain-aware Feature Extraction (IFE) branch and feature alignment strategy effectively addressintra-domain variability. Furthermore, the Dual-Trans Block is introduced to enhance the UIE performance while maintaining computational efficiency. Extensive experiments demonstrate that SSD-UIE outperforms state-of-the-art (SOTA) UIE methods in both qualitative and quantitative evaluations on real-world underwater images. Codes and dataset will be publicly available at https://github.com/RockWenJJ/SSD-UIE.git.
Junjie Wen 0001, Guidong Yang, Benyun Zhao, Dongyue Huang, Lei Lei 0010, Bo Zhang 0019, Zhi Gao 0005, Xi Chen 0104, Ben M. Chen
IEEE Trans. Geosci. Remote. Sens.3
2025 Toward End-to-End Underwater Multi-View Stereo for Real-World Dense Scene Reconstruction
abstract
Multi-view stereo (MVS) enables accurate and complete 3D reconstruction from multi-view imagery, serving as a core methodology in remote sensing applications across terrestrial and underwater domains. Recent advancements in learning-based MVS have demonstrated significant improvements over traditional counterparts, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and presenting the first large-scale synthetic UwMVS dataset preserving real-world underwater degradation properties for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on the dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, appendix, and supplementary video are available at https://yang-sober.github.io/UnderMVS/.
Guidong Yang, Junjie Wen 0001, Lei Lei 0010, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Zhi Gao 0005, Ben M. Chen
IEEE Trans. Geosci. Remote. Sens.4
2024 EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement
abstract
In recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed to several factors: (1) Existing methods typically employ UIE as a pre-processing step, which inevitably introduces considerable computational overhead and latency. (2) The process of enhancing images prior to training object detectors may not necessarily yield performance improvements. (3) The complex underwater environments can induce significant domain shifts across different scenarios, seriously deteriorating the UOD performance. To address these challenges, we introduce EnYOLO, an integrated real-time framework designed for simultaneous UIE and UOD with domain-adaptation capability. Specifically, both the UIE and UOD task heads share the same network backbone and utilize a lightweight design. Furthermore, to ensure balanced training for both tasks, we present a multi-stage training strategy aimed at consistently enhancing their performance. Additionally, we propose a novel domain-adaptation strategy to align feature embeddings originating from diverse underwater environments. Comprehensive experiments demonstrate that our framework not only achieves state-of-the-art (SOTA) performance in both UIE and UOD tasks, but also shows superior adaptability when applied to different underwater scenarios. Our efficiency analysis further highlights the substantial potential of our framework for onboard deployment.
Junjie Wen 0001, Jinqiang Cui, Benyun Zhao, Bingxin Han, Xuchen Liu 0001, Zhi Gao 0005, Ben M. Chen
ICRA3
2024 SANet: Small but Accurate Detector for Aerial Flying Object
abstract
This paper proposes SANet, a small but accurate detector for aerial flying objects. The detector introduces an attention module into the feature extraction module (FEM) for enhancing the accuracy. This FEM with fewer convolutional kernel channels can reduce the parameters, speed up the inference time, and mitigate the computational burden. Furthermore, we optimize the Spatial Pyramid Pooling (SPP) module to enhance both the accuracy and speed. By analyzing the structure characteristic of the ResNet and RepVGG network that are usually utilized to extract features, a feature fusion module named RepNeck is designed to comprehensively fuse features extracted by the FEM, further enhancing the speed and accuracy. Eventually, we develop a neural network with an impressively small model size of only 4.5M. This network can achieve the state-of-the-art performance on three challenging datasets. Apart from its superior performance, our approach enjoys a real-time detection speed of 14.8 frames per second (fps) and power consumption of only 2.9W while the CPU and GPU temperatures are maintained below 50◦C even on an edge-computing device, highlighting the practicality of our approach for long-duration flying object detection and monitoring tasks.
Xunkuai Zhou, Benyun Zhao, Guidong Yang, Jihan Zhang, Li Li 0008, Ben M. Chen
ICRA2
2024 Det-Recon-Reg: An Intelligent Framework Towards Automated Large-Scale Infrastructure Inspection
abstract
Visual inspection plays a predominant role in inspecting infrastructure surface. However, the generalization of existing visual inspection systems to large-scale real-world scenes remains challenging. In this paper, we introduce Det-Recon-Reg, an intelligent framework separating the complex inspection procedure into three stages: Detect, Reconstruct, and Register. (1) For defect detection (Detect), we present the first high-resolution defect dataset tailored for large-scale defect detection. Based on the dataset, we evaluate the most effective real-time object detection algorithms and push the boundary by proposing CUBIT-Net for real-world defect inspection. (2) For infrastructure reconstruction (Reconstruct), we propose a learning-based multi-view stereo (MVS) network to adapt to large-scale scenes, taking as input the multi-view images and outputting the point cloud reconstruction, where its performance has been validated on the standard MVS datasets, including BlendedMVS, DTU, and Tanks and Temples datasets. (3) For defect localization (Register), we propose an effective registration method based on the geographic information system that registers the detected defects onto the reconstructed infrastructure model to establish a global reference for maintenance measures. The real-world experiments further verify the effectiveness and efficiency of our proposed framework. More details about our proposed dataset, code, and appendix are available on our project page: https://cuhk-usr-group.github.io/large-scale-inspect-framework/.
Guidong Yang, Jihan Zhang, Benyun Zhao, Chuanxiang Gao, Yijun Huang, Junjie Wen 0001, Qingxiang Li, Jerry Tang, Xi Chen 0104, Ben M. Chen
IROS3
2023 Multi-View Stereo with Learnable Cost Metric
abstract
In this paper, we present LCM-MVSNet, a novel multi-view stereo (MVS) network with learnable cost metric (LCM) for more accurate and complete depth estimation and dense point cloud reconstruction. To adapt to the scene variation and improve the reconstruction quality in non-Lambertian low-textured scenes, we propose LCM to adaptively aggregate multi-view matching similarity into the 3D cost volume by leveraging sparse points hints. The proposed LCM benefits the MVS approaches in four folds, including depth estimation enhancement, reconstruction quality improvement, memory footprint reduction, and computational burden alleviation, allowing the depth inference for high-resolution images to achieve more accurate and complete reconstruction. Moreover, we improve the depth estimation by enhancing the propagation of shallow features via a bottom-up path and strengthen the end-to-end supervision by adapting the focal loss to reduce ambiguity caused by sample imbalance. Extensive experiments on two benchmark datasets show that our network achieves state-of-the-art performance on the DTU dataset and exhibits strong generalization ability with a competitive performance on the Tanks and Temples benchmark. Furthermore, we deploy our LCM-MVSNet into the real-world application for large-scale 3D reconstruction based on multi-view aerial images collected by self-developed UAV, demonstrating the robustness and scalability of our method. More detailed results are available in the Appendix11shorturl.at/rBG28
Guidong Yang, Xunkuai Zhou, Chuanxiang Gao, Benyun Zhao, Jihan Zhang, Xi Chen 0104, Ben M. Chen
IROS4
2023 ADMNet: Anti-Drone Real-Time Detection and Monitoring
abstract
We propose a lightweight, effective, and efficient anti-drone network, namely ADMNet, for visually detecting and monitoring unfriendly drones with a constrained view field, flying against a complex environment. We merge an SPP module to the first head of YOLOv4 to improve accuracy and perform network compression to reduce inference latency and model size. To compensate for the accuracy loss caused by condensation, we propose an SPPS module and a ResNeck module for the neck of the network and implement an effective attention module for the backbone. Eventually, we present an accurate and compact ADMNet with barely 3.9 MB, ensuring low computational cost and real-time detection. Our method achieves state-of-the-art performance on three challenging real-world datasets (Average Precision @0.5IoU): Det-Fly 96.2%, NPS-Drones 92.0%, and TIBNet 89.7%. The throughput is higher than the prior work, in addition to its superior performance. The comparative testing in real-world scenarios proves that our method exhibits strong reliability and generalization ability. Deploying the network on drone onboard edge-computing devices enables real-time detection and monitoring of flying drones, highlighting the portability and viability of the ADMNet.
Xunkuai Zhou, Guidong Yang, Chuangxiang Gao, Benyun Zhao, Li Li 0008, Ben M. Chen
IROS5