VLDB 2026 Research / reviewers in the wild / expert
Ben M. Chen
dblp:55/4291
· DBLP profile ↗
77ranked-venue papers
0as first author
42since 2021 · last 2026
0000-0002-3839-5787ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 32 since 2021Systems, architecture and hardware · 37 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An accurate and resource-efficient network for surface anomaly detection via enhanced downsampling and activation representation
Xunkuai Zhou, Xi Chen 0104, Jie Chen 0003, Ben M. Chen |
Adv. Eng. Informatics | 4 |
| 2026 | An efficient and accurate network for gardenia fruit detection
Xunkuai Zhou, Yanni Wang, Jie Chen 0003, Ben M. Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Flying Vehicle Detection Under Complex Conditions With RGB-Infrared Imagery: A Large-Scale Open-Source Suite and Benchmark ApproachabstractWhile coordination among multiple flying vehicles improves aerial logistics efficiency, safe and orderly operation requires robust detection and collision-avoidance capabilities. These requirements apply to passenger aircraft as well as to urban traffic and maritime environments, including emerging underwater flying vehicles. However, existing detection methods often fail under challenging conditions such as low illumination or cluttered backgrounds. Their progress is further constrained by the lack of large-scale benchmarks and the high computational and memory costs required to achieve high accuracy, which limits their deployment in resource-constrained scenarios, such as air-to-air collision avoidance in aerial vehicles. To address this gap, we introduce FT55k, an open-source benchmark comprising over 55,000 annotated RGB and infrared images across diverse environments. We further provide baseline approaches tailored for platforms with different computational demands. Extensive experiments on FT55k and three public datasets demonstrate the superior accuracy and efficiency of our methods compared with state-of-the-art approaches. Notably, our approach is the first flying vehicle detection method with a computational cost below 0.5 BFLOPs, achieving real-time performance at 62.3 FPS on an edge-computing device. This work presents the first comprehensive benchmark for flying vehicle detection in complex environments, establishing a practical and scalable foundation for future research and deployment in intelligent transportation safety. Our datasets is publicly accessible athttps://github.com/chriszxk/Flying-Vehicle-Detection Xunkuai Zhou, Yijun Huang, Li Li 0008, Jie Chen 0003, Ben M. Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Sea-U-Whale: A Reconfigurable Marine Robot with Multi-Modal MotionabstractAs marine exploration becomes increasingly important, marine robots have been extensively studied in recent years. Despite some well-designed robots have already achieved to various successful missions, most existing robots struggle to adapt to diverse demands or tasks due to their fixed structure and complexity of the marine environment. To address these challenges, we present a novel reconfigurable marine robot named Sea-U-Whale. This system can dynamically adjust its actuator configuration in the marine environment, providing superior environmental adaptability, maneuverability, and ver-satile mobility. Considering the demands of unmanned ocean exploration, an active reconfiguration mechanism and three distinct vehicle modes are designed for optimal actuation in various marine scenarios. The multi-modal mobility of our system and its robust performance have been validated through extensive field tests and water tank experiments, demonstrating its potential in handling a wide range of mission profiles. Wendi Ding, Zuoquan Zhao, Ruixin Yan, Songqun Gao, Xuchen Liu 0001, Ben M. Chen |
ICRA | 7 |
| 2025 | SE-STDGNN: A Self-Evolving Spatial-Temporal Directed Graph Neural Network for Multi-Vehicle Trajectory PredictionabstractVehicle trajectory prediction (VTP) is essential for microscopic traffic risk assessment, autonomous vehicle navigation, and traffic behavior analysis. Related research leveraging learning-based methodologies has yielded notable success on various benchmark trajectory datasets. However, these models often experience performance degradation when faced with dynamic changes in traffic conditions such as vehicle density, road types, and weather conditions, as they have not been exposed to these variations during the training process. To effectively address the need for real-time adaptation in dynamic traffic scenarios, we propose a novel framework titled self-evolving spatial-temporal directed graph neural network (SE-STDGNN). This model utilizes evolving graph convolution networks (EvolveGCNs) to aggregate spatial-temporal features of vehicles and their neighbors, which are then utilized by a trajectory prediction module to forecast future trajectories. Further, a self-evolving mechanism is introduced to adjust model parameters dynamically in the real-time operation. The efficacy of SE-STDGNN is validated using the public vehicle trajectory dataset AD4CHE. Bingxin Han, Yijun Huang, Xi Chen 0104, Ben M. Chen |
ICRA | 5 |
| 2025 | Underwater Motions Analysis and Control of a Coupling-Tiltable Unmanned Aerial-Aquatic VehicleabstractCoupling-Tiltable Unmanned Aerial-Aquatic Vehicles (UAAVs) have gained increasing importance, yet lack comprehensive analysis and suitable controllers. This paper analyzes the underwater motion characteristics of a self-designed UAAV, Mirs-Alioth, and designs a controller for it. The effectiveness of the controller is validated through experiments. The singularities of Mirs-Alioth are derived as Singular Thrust Tilt Angle (STTA), which serve as an essential tool for an analysis of its underwater motion characteristics. The analysis reveals several key factors for designing the controller. These include the need for logic switching, using a Nussbaum function to compensate control direction uncertainty in the auxiliary channel, and employing an auxiliary controller to mitigate coupling effects. Based on these key points, a control scheme is designed. It consists of a controller that regulates the thrust tilt angle to the singular value, an auxiliary controller incorporating a Saturated Nussbaum function, and a logic switch. Eventually, two sets of experiments are conducted to validate the effectiveness of the controller and demonstrate the necessity of the Nussbaum function. Dongyue Huang, Minghao Dou, Xuchen Liu 0001, Xinlei Chen, Ben M. Chen |
ICRA | 8 |
| 2025 | Multi-View Stereo with Geometric Encoding for Dense Scene ReconstructionabstractMulti-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometric encoding for more accurate and complete depth estimation and point cloud reconstruction. First, the cross-view adaptive cost volume aggregation module is proposed to strengthen multi-view geometric cues encoding during cost volume construction. Then, the depth consistency optimization is performed in the 3D point space during learning by invoking ground-truth depth cues from adjacent views. Finally, the surface normal geometries are explicitly encoded to refine the sampled depth hypotheses to be consistent in the local neighbor regions. Extensive experiments on the standard MVS benchmarks including DTU, Tanks and Temples, and BlendedMVS demonstrate the state-of-the-art depth estimation and point cloud reconstruction performance of GE-MVS. The GE-MVS is further deployed in real-world experiments for UAV-based large-scale reconstruction, where our method outperforms the prevalent industrial reconstruction solutions concerning reconstruction efficiency and efficacy. Our project page is: https://cuhk-usr-group.github.io/GE-MVS/ Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen |
ICRA | 11 |
| 2025 | End-to-End Underwater Multi-View Stereo for Dense Scene ReconstructionabstractRecent advancements in learning-based multi-view stereo (MVS) have demonstrated significant improvements over traditional counterpart, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and present the first large-scale UwMVS dataset for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on our dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, code and appendix are available at: https://cuhk-usr-group.github.io/UwMVS/ Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen |
ICRA | 9 |
| 2025 | Lightweight Yet High-Performance Defect Detector for Uav-Based Large-Scale Infrastructure Real-Time InspectionabstractDefect diagnosis in urban infrastructure is crucial for public safety. Traditional manual inspections face significant challenges in terms of accuracy and cost-effectiveness. In this paper, we propose a lightweight and hardware-friendly large-scale infrastructure detector, CUPID, highly suitable for unmanned aerial vehicles (UAVs). Given the significant challenges in automatically detecting defects of varying intensity and size within complex infrastructure, along with the tendency of lightweight models to lose detail and fail to fully capture features during the defect extraction process, we propose the CUPID_Block, a multi-level information fusion block to construct the backbone, featuring the CUPID_Conv module equipped with our proposed CCA (CrissCross Attention). Furthermore, CUPID features an auxiliary training branch that assimilates lower feature maps, helping to recover details lost in deeper convolutional layers. To verify the effectiveness of CUPID and to address the lack of a suitable dataset in the community, we establish a multi-scenario infrastructure defect dataset, CUBIT2024, to conduct extensive experiments. Finally, to assess the efficiency and adaptability of CUPID in UAV for online infrastructure inspection, we design a compact autonomous drone, CU-Astro, where the proposed CUPID is deployed on the Jetson Orin NX computer onboard to evaluate the speed and power consumption of the inference. Benyun Zhao, Qigeng Duan, Guidong Yang, Jerry Tang, Zhenbo Song, Junjie Wen 0001, Xuchen Liu 0001, Qingxiang Li, Lei Lei 0010, Jihan Zhang, Xi Chen 0104, Mark W. Mueller, Ben M. Chen |
ICRA | 13 |
| 2025 | FHGS: Feature-Homogenized Gaussian SplattingabstractScene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic color representation of gaussian primitives and the isotropic requirements of semantic features, leading to insufficient cross-view feature consistency.
To overcome the limitation, we proposes FHGS (Feature-Homogenized Gaussian Splatting), a novel 3D feature distillation framework inspired by physical models, which freezes and distills 2D pre-trained features into 3D representations while preserving the real-time rendering efficiency of 3DGS.
Specifically, our FHGS introduces the following innovations: Firstly, a universal feature fusion architecture is proposed, enabling robust embedding of large-scale pre-trained models' semantic features (e.g., SAM, CLIP) into sparse 3D structures.
Secondly, a non-differentiable feature fusion mechanism is introduced, which enables semantic features to exhibit viewpoint independent isotropic distributions. This fundamentally balances the anisotropic rendering of gaussian primitives and the isotropic expression of features; Thirdly, a dual-driven optimization strategy inspired by electric potential fields is proposed, which combines external supervision from semantic feature fields with internal primitive clustering guidance. This mechanism enables synergistic optimization of global semantic alignment and local structural consistency.
Extensive comparison experiments with other state-of-the-art methods on benchmark datasets demonstrate that our FHGS exhibits superior reconstruction performance in feature fusion, noise suppression, and geometric precision, while maintaining a significantly lower training time.
This work establishes a novel Gaussian Splatting data structure, offering practical advancements for real-time semantic mapping, 3D stylization, and Vision-Language Navigation (VLN).
Our code and additional results are available on our project page:https://fhgs.cuastro.org/. Qigeng Duan, Benyun Zhao, Mingqiao Han, Yijun Huang, Ben M. Chen |
NeurIPS | 5 |
| 2025 | Towards interpretable and robust UAV-based foundation model for endangered species monitoring in complex ecosystems
Jihan Zhang, Mingqiao Han, K. H. Laurie, Benyun Zhao, Lei Lei 0010, Xi Chen 0104, Hon Chi Judy Wan, Siu Gin Cheung, Wenxing Hong, Ben M. Chen |
Mach. Learn. | 10 |
| 2025 | Multi-View Stereo With Geometric Encoding for Large-Scale Dense Scene Reconstruction
Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Yun-Hui Liu 0001, Ben M. Chen |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Sparse-to-Dense Prediction of Ocean Subsurface Temperature Using Multilevel Spatiotemporal Information FusionabstractAccurately predicting ocean subsurface temperature is vital for advancing ocean and climate research, particularly given the sparse and costly nature of subsurface observations. This study introduces sparse-to-dense prediction of ocean subsurface temperature using multi-level spatiotemporal (ST) information fusion. The framework integrates interpretable ST decoupling, adaptive feature updating, and sparse-to-dense information fusion modules to address the challenge of sparse observations and ever-evolving dynamic environments. Comprehensive experiments focused on the Pacific demonstrate the superiority of the proposed methodology over peer methods. The proposed methodology achieves high-resolution predictions with a root mean square error of 0.2230, accuracy of 0.9846, and point-wise prediction errors below 0.5°C under 10% online random sparse observations (ORSO). Analyses of spatial and temporal temperature dynamics reveal long-term warming trends in the Pacific, including a temperature rise of up to 2.8°C at -100 m in low-latitude regions over the past 40 years, and identify the latitudinal slope of thermocline dynamics. This study advances the understanding of multi-scale thermal processes and variability in the Pacific, demonstrating the potential of application in climate studies, marine resource management, and environmental monitoring. Lei Lei 0010, Guidong Yang, Zuoquan Zhao, Xi Chen 0104, Ben M. Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | A Semi-Supervised Domain-Adaptive Framework for Real-World Underwater Image EnhancementabstractUnderwater optical remote sensing is crucial for geoscience applications but often suffers from image degradation due to complex underwater environments. While learning-based methods have advanced underwater image enhancement (UIE), their efficacy in real-world UIE applications still faces challenges. This limitation arises from training predominantly on synthetic underwater images, resulting in a significantinter-domain gap when applied to real-world data. Additionally, diverse underwater conditions introduceintra-domain challenges, such as color casts and haze, further complicating the UIE process. To address these issues, we propose SSD-UIE, a semi-supervised domain-adaptive framework designed to mitigate bothinter- andintra-domain gaps. Our approach employs a systematic synthesis pipeline to reduce visualinter-domain discrepancies and introduces a Large Synthetic-Real Underwater Image Dataset (LSRUID) to facilitate the training of the framework. The Semantic-Blender is developed to handle semanticinter-domain differences, while the Intra-domain-aware Feature Extraction (IFE) branch and feature alignment strategy effectively addressintra-domain variability. Furthermore, the Dual-Trans Block is introduced to enhance the UIE performance while maintaining computational efficiency. Extensive experiments demonstrate that SSD-UIE outperforms state-of-the-art (SOTA) UIE methods in both qualitative and quantitative evaluations on real-world underwater images. Codes and dataset will be publicly available at https://github.com/RockWenJJ/SSD-UIE.git. Junjie Wen 0001, Guidong Yang, Benyun Zhao, Dongyue Huang, Lei Lei 0010, Bo Zhang 0019, Zhi Gao 0005, Xi Chen 0104, Ben M. Chen |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Toward End-to-End Underwater Multi-View Stereo for Real-World Dense Scene ReconstructionabstractMulti-view stereo (MVS) enables accurate and complete 3D reconstruction from multi-view imagery, serving as a core methodology in remote sensing applications across terrestrial and underwater domains. Recent advancements in learning-based MVS have demonstrated significant improvements over traditional counterparts, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and presenting the first large-scale synthetic UwMVS dataset preserving real-world underwater degradation properties for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on the dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, appendix, and supplementary video are available at https://yang-sober.github.io/UnderMVS/. Guidong Yang, Junjie Wen 0001, Lei Lei 0010, Benyun Zhao, Qingxiang Li, Xi Chen 0104, Zhi Gao 0005, Ben M. Chen |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Sensor-based Multi-Robot Coverage Control with Spatial Separation in Unstructured EnvironmentsabstractMulti-robot systems have increasingly become instrumental in tackling coverage problems. However, the challenge of optimizing task efficiency without compromising task success still persists, particularly in expansive, unstructured scenarios with dense obstacles. This paper presents an innovative, decentralized Voronoi-based coverage control approach to reactively navigate these complexities while guaranteeing safety. This approach leverages the active sensing capabilities of multi-robot systems to supplement GIS (Geographic Information System), offering a more comprehensive and real-time understanding of environments like post-disaster. Based on point cloud data, which is inherently non-convex and unstructured, this method efficiently generates collision-free Voronoi regions using only local sensing information through spatial decomposition and spherical mirroring techniques. Then, deadlock-aware guided map integrated with a gradient-optimized, centroid Voronoi-based coverage control policy, is constructed to improve efficiency by avoiding exhaustive searches and local sensing pitfalls. The effectiveness of our algorithm has been validated through extensive numerical simulations in high-fidelity environments, demonstrating significant improvements in task success rate, coverage ratio, and task execution time compared with others. Xinyi Wang 0007, Jiwen Xu, Chuanxiang Gao, Jihan Zhang, Ben M. Chen |
ICRA | 8 |
| 2024 | Air Bumper: A Collision Detection and Reaction Framework for Autonomous MAV NavigationabstractAutonomous navigation in unknown environments with obstacles remains challenging for micro aerial vehicles (MAVs) due to their limited onboard computing and sensing resources. Although various collision avoidance methods have been developed, it is still possible for drones to collide with unobserved obstacles due to unpredictable disturbances, sensor limitations, and control uncertainty. Instead of completely avoiding collisions, this article proposes Air Bumper, a collision detection and reaction framework, for fully autonomous flight in 3D environments to improve flight safety. Our framework only utilizes the onboard inertial measurement unit (IMU) to detect and estimate collisions. We further design a collision recovery control for rapid recovery and collision-aware mapping to integrate collision information into general LiDAR-based sensing and planning frameworks. Our simulation and experimental results show that the drone can rapidly detect, estimate, and recover from collisions with obstacles in 3D space and continue the flight smoothly with the help of the collision-aware map. In addition, we will open-source the implementation of Air Bumper on GitHub1. Ruoyu Wang 0032, Xinyi Wang 0007, Ben M. Chen |
ICRA | 5 |
| 2024 | SGCalib: A Two-stage Camera-LiDAR Calibration Method Using Semantic Information and Geometric FeaturesabstractExtrinsic calibration is an essential prerequisite for the applications of camera-LiDAR fusion. Existing methods either suffer from the complex offline setting of man-made targets or tend to produce suboptimal and unrobust results. In this paper, we propose an online two-stage calibration method that estimates robust and accurate extrinsic parameters between camera and LiDAR. This is a novel work to use semantic information and geometric features jointly in calibration to promote accuracy and robustness. In the first stage, we detect objects in the image and point cloud and build graphs on the objects using Delaunay triangulation. Then, we design a novel graph matching algorithm to associate the objects in the two data domains and extract pairs of 2D-3D points. Using the PnP solver, we get robust initial extrinsic parameters. Then, in the second stage, we design a new optimization formulation with semantic information and geometric features to generate accurate extrinsic parameters with the initial value from the first stage. Extensive experiments on solid-state LiDAR, conventional spinning LiDAR and KITTI datasets have verified the robustness and accuracy of our method which outperforms existing works. We will share the code publicly to benefit the community (after review stages). Zhi Gao 0005, Xinyi Liu 0002, Ben M. Chen |
ICRA | 6 |
| 2024 | EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image EnhancementabstractIn recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed to several factors: (1) Existing methods typically employ UIE as a pre-processing step, which inevitably introduces considerable computational overhead and latency. (2) The process of enhancing images prior to training object detectors may not necessarily yield performance improvements. (3) The complex underwater environments can induce significant domain shifts across different scenarios, seriously deteriorating the UOD performance. To address these challenges, we introduce EnYOLO, an integrated real-time framework designed for simultaneous UIE and UOD with domain-adaptation capability. Specifically, both the UIE and UOD task heads share the same network backbone and utilize a lightweight design. Furthermore, to ensure balanced training for both tasks, we present a multi-stage training strategy aimed at consistently enhancing their performance. Additionally, we propose a novel domain-adaptation strategy to align feature embeddings originating from diverse underwater environments. Comprehensive experiments demonstrate that our framework not only achieves state-of-the-art (SOTA) performance in both UIE and UOD tasks, but also shows superior adaptability when applied to different underwater scenarios. Our efficiency analysis further highlights the substantial potential of our framework for onboard deployment. Junjie Wen 0001, Jinqiang Cui, Benyun Zhao, Bingxin Han, Xuchen Liu 0001, Zhi Gao 0005, Ben M. Chen |
ICRA | 7 |
| 2024 | Sea-U-Foil: A Hydrofoil Marine Vehicle with Multi-Modal LocomotionabstractAutonomous Marine Vehicles (AMVs) have been widely used in many critical tasks such as surveillance, patrolling, marine environment monitoring, and hydrographic surveying. However, most typical AMVs cannot meet the diverse demands of different marine tasks. In this article, we design a new type of remote-controlled hydrofoil marine vehicle, named Sea-U-Foil, which is suitable for different marine scenarios. Sea-U-Foil features three distinct locomotion modes, displacement mode, foilborne mode, and submarine mode, which enable the platform flexible mobility, high-speed and high-load capacities, and superior concealment. Specifically, the submarine mode makes Sea-U-Foil unique among previous studies. In addition, the performance of Sea-U-Foil in foilborne mode outperforms those of most current unmanned surface vehicles (USVs) in terms of speed and payload. To the best of our knowledge, we are the first to introduce a new type of AMV that can work in displacement mode, foilborne mode, and submarine mode. We elaborate on the design principles and methodologies of Sea-U-Foil first, then validate the effectiveness of its tri-modal locomotion through extensive experiments. Zuoquan Zhao, Chuanxiang Gao, Wendi Ding, Ruixin Yan, Songqun Gao, Bingxin Han, Xuchen Liu 0001, Ben M. Chen |
ICRA | 10 |
| 2024 | SANet: Small but Accurate Detector for Aerial Flying ObjectabstractThis paper proposes SANet, a small but accurate detector for aerial flying objects. The detector introduces an attention module into the feature extraction module (FEM) for enhancing the accuracy. This FEM with fewer convolutional kernel channels can reduce the parameters, speed up the inference time, and mitigate the computational burden. Furthermore, we optimize the Spatial Pyramid Pooling (SPP) module to enhance both the accuracy and speed. By analyzing the structure characteristic of the ResNet and RepVGG network that are usually utilized to extract features, a feature fusion module named RepNeck is designed to comprehensively fuse features extracted by the FEM, further enhancing the speed and accuracy. Eventually, we develop a neural network with an impressively small model size of only 4.5M. This network can achieve the state-of-the-art performance on three challenging datasets. Apart from its superior performance, our approach enjoys a real-time detection speed of 14.8 frames per second (fps) and power consumption of only 2.9W while the CPU and GPU temperatures are maintained below 50◦C even on an edge-computing device, highlighting the practicality of our approach for long-duration flying object detection and monitoring tasks. Xunkuai Zhou, Benyun Zhao, Guidong Yang, Jihan Zhang, Li Li 0008, Ben M. Chen |
ICRA | 6 |
| 2024 | Accurate and Efficient Loop Closure Detection With Deep Binary Image Descriptor and Augmented Point Cloud RegistrationabstractLoop Closure Detection (LCD) is an essential component of Simultaneous Localization and Mapping (SLAM), helping to correct drift errors, facilitate map merging, or both by identifying previously observed scenes. Despite its importance, traditional LCD algorithms based on single sensor such as camera or LiDAR exhibit degraded performance in challenging scenarios due to their inherent limitations. To address this issue, we propose a novel LCD method based on camera-LiDAR fusion, exploiting the rich textural information from cameras and the accurate geometric data from LiDAR to ensure robustness and speed in challenging environments. Specifically, we first employ deep hashing learning to encode deep image features into binary image descriptors for extremely fast loop candidate (LC) retrieval. Then, LiDAR points are augmented with image color for accurate geometric verification. Finally, we incorporate a spatial-temporal consistency check that mandates an LC to have consistently matched neighbors to be accepted as true. Our method is extensively verified and compared with the state-of-the-art methods on various datasets encompassing both indoor and outdoor environments. Experimental results demonstrate that our method obtains the best performance, increasing the maximum recall rate at 100% precision by a significant margin of 20% while operating in real-time at an average speed of 30 fps. Zhi Gao 0005, Jianhua Cheng, Xinyi Liu 0002, Ben M. Chen |
IROS | 9 |
| 2024 | Det-Recon-Reg: An Intelligent Framework Towards Automated Large-Scale Infrastructure InspectionabstractVisual inspection plays a predominant role in inspecting infrastructure surface. However, the generalization of existing visual inspection systems to large-scale real-world scenes remains challenging. In this paper, we introduce Det-Recon-Reg, an intelligent framework separating the complex inspection procedure into three stages: Detect, Reconstruct, and Register. (1) For defect detection (Detect), we present the first high-resolution defect dataset tailored for large-scale defect detection. Based on the dataset, we evaluate the most effective real-time object detection algorithms and push the boundary by proposing CUBIT-Net for real-world defect inspection. (2) For infrastructure reconstruction (Reconstruct), we propose a learning-based multi-view stereo (MVS) network to adapt to large-scale scenes, taking as input the multi-view images and outputting the point cloud reconstruction, where its performance has been validated on the standard MVS datasets, including BlendedMVS, DTU, and Tanks and Temples datasets. (3) For defect localization (Register), we propose an effective registration method based on the geographic information system that registers the detected defects onto the reconstructed infrastructure model to establish a global reference for maintenance measures. The real-world experiments further verify the effectiveness and efficiency of our proposed framework. More details about our proposed dataset, code, and appendix are available on our project page: https://cuhk-usr-group.github.io/large-scale-inspect-framework/. Guidong Yang, Jihan Zhang, Benyun Zhao, Chuanxiang Gao, Yijun Huang, Junjie Wen 0001, Qingxiang Li, Jerry Tang, Xi Chen 0104, Ben M. Chen |
IROS | 10 |
| 2024 | VDTNet: A High-Performance Visual Network for Detecting and Tracking of Intruding DronesabstractThe misuse of drones can jeopardize public safety and privacy. The detection and catching of intruding drones are crucial and urgent issues to be investigated. This work proposes VDTNet, an accurate, lightweight, and fast network for visually detecting and tracking intruding drones. We first incorporate an SPP module into the first head of YOLOv4 to enhance detection accuracy. Model compression is utilized to shrink the model size and concurrently speed up inference. We then propose and insert an SPPS module and a ResNeck module into the neck, and introduce an effective attention module for the backbone to compensate for the accuracy drop brought on by compression. With the above strategies, we present the accurate and compact VDTNet with a model size of merely 3.9 MB, ensuring low computational cost and fast detection and tracking performance in real time. Extensive experiments on four challenging public datasets show that our proposed network outperforms state-of-the-art approaches. In real-world scenarios, the comparative ground-to-air detection testing proves the generalization ability of the VDTNet, and we further demonstrate the portability and practicability of the network by deploying it on drone onboard edge-computing devices for air-to-air real-time detection of the intruding drones. Xunkuai Zhou, Guidong Yang, Li Li 0008, Ben M. Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Sampling-based path planning under temporal logic constraints with real-time adaptationabstractReplanning in temporal logic tasks is extremely difficult during the online execution of robots. This study introduces an effective path planner that computes solutions for temporal logic goals and instantly adapts to non-static and partially unknown environments. Given prior knowledge and a task specification, the planner first identifies an initial feasible solution by growing a sampling-based search tree. While carrying out the computed plan, the robot maintains a solution library to continuously enhance the unfinished part of the plan and store backup plans. The planner updates existing plans when meeting unexpected obstacles or recognizing flaws in prior knowledge. Upon a high-level path is obtained, a trajectory generator tracks the path by dividing it into segments of motion primitives. Our planner is integrated into an autonomous mobile robot system, further deployed on a multicopter with limited onboard processing power. In simulation and real-world experiments, our planner is demonstrated to swiftly and effectively adjust to environmental uncertainties. Ruoyu Wang 0032, Xinyi Wang 0007, Ben M. Chen |
ICRA | 4 |
| 2023 | TJ-FlyingFish: Design and Implementation of an Aerial-Aquatic Quadrotor with Tiltable Propulsion UnitsabstractAerial-aquatic vehicles are capable to move in the two most dominant fluids, making them more promising for a wide range of applications. We propose a prototype with special designs for propulsion and thruster configuration to cope with the vast differences in the fluid properties of water and air. For propulsion, the operating range is switched for the different mediums by the dual-speed propulsion unit, providing sufficient thrust and also ensuring output efficiency. For thruster configuration, thrust vectoring is realized by the rotation of the propulsion unit around the mount arm, thus enhancing the underwater maneuverability. This paper presents a quadrotor prototype of this concept and the design details and realization in practice. Xuchen Liu 0001, Minghao Dou, Dongyue Huang, Songqun Gao, Ruixin Yan, Biao Wang 0004, Jinqiang Cui, Qinyuan Ren, LiHua Dou, Zhi Gao 0005, Jie Chen 0003, Ben M. Chen |
ICRA | 12 |
| 2023 | SyreaNet: A Physically Guided Underwater Image Enhancement Framework Integrating Synthetic and Real ImagesabstractUnderwater image enhancement (UIE) is vital for high-level vision-related underwater tasks. Although learning-based UIE methods have made remarkable achievements in recent years, it's still challenging for them to consistently deal with various underwater conditions, which could be caused by: 1) the use of the simplified atmospheric image formation model in UIE may result in severe errors; 2) the network trained solely with synthetic images might have difficulty in generalizing well to real underwater images. In this work, we, for the first time, propose a framework SyreaNet for UIE that integrates both synthetic and real data under the guidance of the revised underwater image formation model and novel domain adaptation (DA) strategies. First, an underwater image synthesis module based on the revised model is proposed. Then, a physically guided disentangled network is designed to predict the clear images by combining both synthetic and real underwater images. The intra- and inter-domain gaps are abridged by fully exchanging the domain knowledge. Extensive experiments demonstrate the superiority of our framework over other state-of-the-art (SOTA) learning-based UIE methods qualitatively and quantitatively. The code and dataset are publicly available at https://github.com/RockWenJJ/SyreaNet.git. Junjie Wen 0001, Jinqiang Cui, Zhenjun Zhao, Ruixin Yan, Zhi Gao 0005, LiHua Dou, Ben M. Chen |
ICRA | 7 |
| 2023 | An Interactive System for Multiple-Task Linear Temporal Logic Path PlanningabstractBeyond programming robots to accomplish a single high-level task at a time, people also hope robots follow instructions and complete a series of tasks while meeting their requirements. This paper presents an interactive software system that consists of a multiple-task linear temporal logic (LTL) path planner and a human-machine interface (HMI). The HMI transforms human oral instructions into task commands that can be understood by the machine. The planner grows a rapid random exploring tree to search for solutions for multiple tasks. When switching tasks, the search tree is re-initialized and reconnected to utilize the information gathered during the exploration of the workspace. The feasibility of the improved planner is theoretically guaranteed, and profiling in simulation shows an acceleration in planning. An experiment with a quadcopter is conducted to show that the combination of the multiple-task LTL planner and the HMI results in a synergistic effect in real-world applications. Xinyi Wang 0007, Ruoyu Wang 0032, Xunkuai Zhou, Guidong Yang, Shupeng Lai, Ben M. Chen |
IROS | 8 |
| 2023 | Multi-View Stereo with Learnable Cost MetricabstractIn this paper, we present LCM-MVSNet, a novel multi-view stereo (MVS) network with learnable cost metric (LCM) for more accurate and complete depth estimation and dense point cloud reconstruction. To adapt to the scene variation and improve the reconstruction quality in non-Lambertian low-textured scenes, we propose LCM to adaptively aggregate multi-view matching similarity into the 3D cost volume by leveraging sparse points hints. The proposed LCM benefits the MVS approaches in four folds, including depth estimation enhancement, reconstruction quality improvement, memory footprint reduction, and computational burden alleviation, allowing the depth inference for high-resolution images to achieve more accurate and complete reconstruction. Moreover, we improve the depth estimation by enhancing the propagation of shallow features via a bottom-up path and strengthen the end-to-end supervision by adapting the focal loss to reduce ambiguity caused by sample imbalance. Extensive experiments on two benchmark datasets show that our network achieves state-of-the-art performance on the DTU dataset and exhibits strong generalization ability with a competitive performance on the Tanks and Temples benchmark. Furthermore, we deploy our LCM-MVSNet into the real-world application for large-scale 3D reconstruction based on multi-view aerial images collected by self-developed UAV, demonstrating the robustness and scalability of our method. More detailed results are available in the Appendix11shorturl.at/rBG28 Guidong Yang, Xunkuai Zhou, Chuanxiang Gao, Benyun Zhao, Jihan Zhang, Xi Chen 0104, Ben M. Chen |
IROS | 8 |
| 2023 | ADMNet: Anti-Drone Real-Time Detection and MonitoringabstractWe propose a lightweight, effective, and efficient anti-drone network, namely ADMNet, for visually detecting and monitoring unfriendly drones with a constrained view field, flying against a complex environment. We merge an SPP module to the first head of YOLOv4 to improve accuracy and perform network compression to reduce inference latency and model size. To compensate for the accuracy loss caused by condensation, we propose an SPPS module and a ResNeck module for the neck of the network and implement an effective attention module for the backbone. Eventually, we present an accurate and compact ADMNet with barely 3.9 MB, ensuring low computational cost and real-time detection. Our method achieves state-of-the-art performance on three challenging real-world datasets (Average Precision @0.5IoU): Det-Fly 96.2%, NPS-Drones 92.0%, and TIBNet 89.7%. The throughput is higher than the prior work, in addition to its superior performance. The comparative testing in real-world scenarios proves that our method exhibits strong reliability and generalization ability. Deploying the network on drone onboard edge-computing devices enables real-time detection and monitoring of flying drones, highlighting the portability and viability of the ADMNet. Xunkuai Zhou, Guidong Yang, Chuangxiang Gao, Benyun Zhao, Li Li 0008, Ben M. Chen |
IROS | 7 |
| 2023 | FG-Net: A Fast and Accurate Framework for Large-Scale LiDAR Point Cloud UnderstandingabstractThis work presents FG-Net, a general deep learning framework for large-scale point cloud understanding without voxelizations, which achieves accurate and real-time performance with a single NVIDIA GTX 1080 8G GPU and an i7 CPU. First, a novel noise and outlier filtering method is designed to facilitate the subsequent high-level understanding tasks. For effective understanding purpose, we propose a novel plug-and-play module consisting of correlated feature mining and deformable convolution-based geometric-aware modeling, in which the local feature relationships and point cloud geometric structures can be fully extracted and exploited. For the efficiency issue, we put forward a new composite inverse density sampling (IDS)-based and learning-based operation and a feature pyramid-based residual learning strategy to save the computational cost and memory consumption, respectively. Compared with current methods which are only validated on limited datasets, we have done extensive experiments on eight real-world challenging benchmarks, which demonstrates that our approaches outperform state-of-the-art (SOTA) approaches in terms of accuracy, speed, and memory efficiency. Moreover, weakly supervised transfer learning is also conducted to demonstrate the generalization capacity of our method. Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
IEEE Trans. Cybern. | 4 |
| 2023 | Synergizing Low Rank Representation and Deep Learning for Automatic Pavement Crack DetectionabstractDue to the critical role of pavement crack detection for road maintenance and eventually ensuring safety, remarkable efforts have been devoted to this research area, and such a trend is further intensified for the coming unmanned vehicle era. However, such crack detection task still remains unexpectedly challenging in practice since the appearance of both cracks and the background are diverse and complex in real scenarios. In this work, we propose an automatic pavement crack detection method via synergizing low rank representation (LRR) and deep learning techniques. First, leveraging LRR which facilitates anomaly detection without making any specific assumption, we can easily discriminate most of the frames with cracks from the long sequence with a consistent pavement base, followed by a straightforward algorithm to localize the cracks. In order to achieve the intelligence of detecting cracks with different pavement basis under unconstrained imaging conditions, we resort to deep learning techniques and propose a deep convolutional neural network for crack detection leveraging on multi-level features and atrous spatial pyramid pooling (ASPP). We train this network based on the training data obtained in the previous stage in an end-to-end manner. Extensive experiments on a wide range of pavements demonstrate the high performance in terms of both accuracy and automaticity. Moreover, the dataset generated by us is much more extensive and challenging than public ones. We put it online athttps://gaozhinuswhu.comto benefit the community. Zhi Gao 0005, Min Cao 0001, Ziyao Li, Kangcheng Liu, Ben M. Chen |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Weakly Supervised 3D Scene Segmentation with Region-Level Boundary Awareness and Instance Discrimination
Kangcheng Liu, Yuzhi Zhao, Qiang Nie, Zhi Gao 0005, Ben M. Chen |
ECCV (28) | 5 |
| 2022 | Model-Based Reinforcement Learning with Self-attention Mechanism for Autonomous Driving in Dense Traffic
Junjie Wen 0001, Zuoquan Zhao, Jinqiang Cui, Ben M. Chen |
ICONIP (2) | 4 |
| 2022 | WeakLabel3D-Net: A Complete Framework for Real-Scene LiDAR Point Clouds Weakly Supervised Multi-Tasks UnderstandingabstractExisting state-of-the-art 3D point clouds understanding methods only perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework which simultaneously solves the downstream high-level understanding tasks, especially when labels are extremely limited. This work presents a general and simple framework to tackle point clouds understanding when labels are limited. We propose a novel unsupervised region expansion based clustering method for generating clusters. More importantly, we innovatively propose to learn to merge the over-divided clusters based on the local low-level geometric property similarities and the learned high-level feature similarities supervised by weak labels. Hence, the true weak labels guide pseudo labels merging taking both geometric and semantic feature correlations into consideration. Finally, the self-supervised data augmentation optimization module is proposed to guide the propagation of labels among semantically similar points within a scene. Experimental Results demonstrate that our framework has the best performance among the three most important weakly supervised point clouds understanding tasks including semantic segmentation, instance segmentation, and object detection even when limited points are labeled. Kangcheng Liu, Yuzhi Zhao, Zhi Gao 0005, Ben M. Chen |
ICRA | 4 |
| 2022 | A Memetic Algorithm for Curvature-Constrained Path Planning of Messenger UAV in Air-Ground CoordinationabstractThis paper addresses a UAV path planning problem for a team of cooperating heterogeneous vehicles composed of one unmanned aerial vehicle (UAV) and multiple unmanned ground vehicles (UGVs). The UGVs are used as mobile actuators and scattered in a large area. To achieve multi-UGV communication and collaboration, the UAV, modeled as a Dubins vehicle, serves as a messenger to fly over the effective communication range of all UGVs to relay information. The curvature-constrained path planning of the messenger UAV is formulated as a Dubins Traveling Salesman Problem with Dynamic Neighborhood (DTSPDN) which is a complex optimization problem involving coupled variables and contains dynamic constraints. We design an effective memetic algorithm to find the shortest route that enables the messenger UAV to visit all moving UGVs. This algorithm combines the genetic algorithm procedure, two kinds of local search operators based on gradient search and uniform sampling respectively, and a gradient-based repair operator to repair the solutions violating dynamic constraints. During the evolutionary process, a special phenomenon may occur that changing some decision variables (i.e., visiting sequence and location) may not affect the evaluation function value, but may alter the feasible region of another decision variable (i.e., visiting time) due to the encounter constraint between the UAV and UGV. To track and utilize the change of the feasible region, a transformation procedure is proposed to change one solution to another with less visiting time by analyzing the encounter pattern between UAV and UGV. The computational results on random instances with different scales demonstrate that the proposed approach can effectively generate better curvature-constrained tours to encounter all moving UGVs when compared to other four competitive algorithms in the literature. Note to Practitioners—This paper studies an emerging path planning problem for a UAV which is used to provide communication service for multiple moving UGVs. These UGVs are required to execute tasks (e.g., firefighting, search and rescue) within a large area. Due to their limited communication capabilities, they may be unable to obtain necessary information from other UGVs. The UAV serves as a messenger to fly over the effective communication range of all moving UGVs to relay information. We propose a novel memetic algorithm to efficiently search for the shortest tour that enables the messenger UAV to visit all moving UGVs. The memetic algorithm combines the parallel global search virtue of genetic algorithm with efficient local search procedure to improve the generated tour. A gradient-based repair procedure is also employed to make sure that the planned tour can guide the UAV to sequentially encounter each moving UGV. Simulations exhibit that the proposed approach can effectively generate high-quality tours for messenger UAV to rapidly visit all UGVs, which assists UGVs to achieve collaboration in large area. In future work, the proposed memetic algorithm will be extended to plan tours for multiple messenger UAVs. Bin Xin 0002, LiHua Dou, Jie Chen 0003, Ben M. Chen |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2022 | Formation-Containment Control of Euler-Lagrange Systems of Leaders With Bounded Unknown InputsabstractMotivated by the promising applications of multiple Euler-Lagrange (EL) systems, we study, in this article, the formation-containment (FC) control problem for multiple EL systems of leaders with bounded unknown control inputs and with communication among each other over directed topologies, which can cooperatively generate safe trajectories to avoid obstacles. Given the FC shapes, an algorithm is first proposed to obtain the stress matrix while satisfying certain conditions, based on which a novel adaptive distributed observer to the convex hull is proposed for every follower. An adaptive updating gain is applied to make the observer fully distributed without using the global information of the graph, and a continuous function is designed to restrain the influence of the inputs of the leaders. Then, a local control law using the adaptive distributed observer is presented to accomplish the FC control of EL systems. Based on the Lyapunov stability theory, it is proved that the FC error can be designed as small as possible by adjusting some parameters in the observer. Panpan Zhou, Ben M. Chen |
IEEE Trans. Cybern. | 2 |
| 2021 | FG-Conv: Large-Scale LiDAR Point Clouds Understanding Leveraging Feature Correlation Mining and Geometric-Aware ModelingabstractThis work presents a general deep learning framework for large-scale point clouds understanding without voxelizations, called FG-Conv, which achieves an accurate and real-time understanding of point clouds. Through our novel design combining feature level correlation mining and deformable convolutions based geometric aware modeling, the local feature relationships and geometric patterns can be captured. The attention mechanism is also adopted to enhance the global long-range feature correlations. Finally, the feature pyramid residual learning network is proposed to combine patterns at different resolutions in a memory-efficient way. Extensive experiments on real-world challenging datasets demonstrated that our approaches outperform state-of-the-art methods in terms of accuracy and efficiency. Weakly supervised transfer learning demonstrates the generalization capacity of our methods. Kangcheng Liu, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
ICRA | 4 |
| 2021 | Underwater Stability of a Morphable Aerial-Aquatic Quadrotor With Variable Thruster AnglesabstractThe design of aerial-aquatic multirotors can benefit from thruster rotation so that the thrusters can act directly in the lateral directions of surge and sway when submerged. This allows much more effective locomotion underwater as opposed to the aerial configuration where rotational acceleration is used to direct small components of thrust in the lateral directions. However, the introduction of lateral thruster components by rotating the thrusters about their respective arm axes creates additional coupled moment terms. Here, the dynamics of this design is analysed to show that critical angles of thruster rotation can achieve stable surge and sway movements while also decoupling lateral from rotational movements. Yu Herng Tan, Ben M. Chen |
ICRA | 2 |
| 2021 | Smooth quadrotor trajectory generation for tracking a moving target in cluttered environments
Lele Xi, Zhihong Peng, Lei Jiao 0005, Ben M. Chen |
Sci. China Inf. Sci. | 4 |
| 2021 | IPMGAN: Integrating physical model and generative adversarial network for underwater image enhancement
Xiaodong Liu 0008, Zhi Gao 0005, Ben M. Chen |
Neurocomputing | 3 |
| 2021 | Multivehicle Flocking With Collision Avoidance via Distributed Model Predictive ControlabstractFlocking control has been studied extensively along with the wide applications of multivehicle systems. In this article, the distributed flocking control strategy is studied for a network of autonomous vehicles with limited communication range. The main difference from the existing methods lies in that collision avoidance is considered a necessary condition while the vehicles are driven to follow a common desired trajectory under the proximity network. The sufficient conditions for system feasibility and stability are given by the proposed strategy. First, a centralized standard model predictive control (MPC) scheme is adopted to formulate the multivehicle flocking control problem by setting collision avoidance as an optimization constraint under the proximity network. Further, an equivalent distributed MPC (DMPC) is developed based on the consensus of local controllers under the existing framework of the alternating direction method of multiplier (ADMM). However, it may require infinite time to achieve consensus for all vehicles and, thus, the local controllers resulting in a limited number of ADMM iterations may not satisfy the given constraints. The constraints for each local controller are then modified so that the collision between vehicles is avoided all of the time. The feasibility and stability of the proposed method are analyzed under practical conditions. Simulation and experimental results show that the flocking of vehicles can track the common desired trajectory stably with no collisions by the proposed method. Yang Lyu, Jinwen Hu, Ben M. Chen, Chunhui Zhao 0002, Quan Pan 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | A Morphable Aerial-Aquatic Quadrotor with Coupled Symmetric Thrust VectoringabstractHybrid aerial-aquatic vehicles have the unique ability of travelling in both air and water and can benefit from both lower fluid resistance in air and energy efficient position holding in water. However, they have to address the differing requirements which make optimising a single design difficult. While existing examples have shown the possibility of such vehicles, they are mostly structurally identical to normal aerial vehicles with minor adjustments to work underwater. Instead of using rotational acceleration to direct a component of thrust in surge and sway, we propose a quadrotor based vehicle that tilts its rotors about the respective arm so that a larger component of thrust can be directed in the lateral plane or in the opposite direction without rotating the vehicle body. A small-scale prototype of this design is presented here, detailing the design considerations including mechanical actuation, static stability and waterproofing. Yu Herng Tan, Ben M. Chen |
ICRA | 2 |
| 2020 | A Target Tracking and Positioning Framework for Video Satellites Based on SLAMabstractWith the booming development in aerospace technology, the video satellite which observes the live phenomena on the ground by video shooting has gradually emerged as a new Earth observation method. And remote sensing comes into a "dynamic" era with the demand for new processing techniques, especially the near-real-time tracking and geo-positioning algorithm for ground moving targets. However, many researchers merely extract pixel-level trajectories in post-processed video products, resulting in fairly limited applications. We regard the video satellite as a robot flying in space and adopt the SLAM framework for the positioning of ground moving targets. The designed framework is based on the representative ORB-SLAM and we make improvements mainly in feature extraction, satellite pose estimation, moving target tracking and positioning. We coordinate a moving fishing boat with GPS-RTK (Real-time Kinematic) devices and a video satellite observing it simultaneously for verification and evaluation of our method. Experiments demonstrate that our framework provides reasonable geolocation of the moving target in satellite videos. Finally, some open problems and potential research directions are discussed. Zhi Gao 0005, Yongjun Zhang 0002, Ben M. Chen |
IROS | 4 |
| 2020 | MLFcGAN: Multilevel Feature Fusion-Based Conditional GAN for Underwater Image Color CorrectionabstractColor correction for underwater images has received increasing interest, due to its critical role in facilitating available mature vision algorithms for underwater scenarios. Inspired by the stunning success of deep convolutional neural network (DCNN) techniques in many vision tasks, especially the strength in extracting features in multiple scales, we propose a deep multiscale feature fusion net based on the conditional generative adversarial network (GAN) for underwater image color correction. In our network, multiscale features are extracted first, followed by augmenting local features in each scale with global features. This design was verified to facilitate more effective and faster network learning, resulting in better performance in both color correction and detail preservation. We conducted extensive experiments and compared the results with state-of-the-art approaches quantitatively and qualitatively, showing that our method achieves significant improvements. Xiaodong Liu 0008, Zhi Gao 0005, Ben M. Chen |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Motor-propeller Matching of Aerial Propulsion Systems for Direct Aerial-aquatic OperationabstractElectric aerial propulsion systems are commonly used for many small-scale unmanned aerial vehicles (UAVs), providing a light and powerful method of generating thrust. In the emerging area of aerial-aquatic vehicles, most existing prototypes rely on such systems to propel themselves in both air and water. As the density of water is three orders of magnitude larger than that of air, a spinning aerodynamic body in the medium will experience significantly higher torque at the same speed. This results in aerial propulsion systems to be heavily mismatched underwater, as the required torque is higher than the drive torque that a typical aerial motor can provide. Here, an in-depth investigation of such off-design operation is conducted. Based on numerical simulation, we identify the feasible operating range of such systems and present an evaluation framework that identifies a motor-propeller combination from a component database that maximises underwater performance while ensuring aerial thrust requirements are met. Yu Herng Tan, Ben M. Chen |
IROS | 2 |
| 2019 | Complex system and intelligent control: theories and applicationsabstractComplex systems are the systems that consist of a great many diverse and autonomous but interacting and interdependent components whose aggregate behaviors are nonlinear.As phased by Aristotle, "the whole is more than the sum of its parts;" properties of complex systems are not a simple summation of their individual parts.Complex systems are widespread.Typical examples of complex systems can be found in the human brain, flocking formation of migrating birds, power grid, transportation systems, autonomous vehicles, social networks, and communication networks.Complex systems have some distinct properties, such as highly nonlinear dynamics, emergence, adaptation, and self-organization, which are difficult to model precisely.Such properties lead to difficulties in understanding the behaviors of complex systems, to model them accurately, to control them, and to make them work in a specific way we desire.There is no generally agreed definition of intelligent control.Generally speaking, intelligent control is a class of control methods that use artificial intelligence techniques, such as fuzzy logic, neural networks, Jie Chen 0003, Ben M. Chen, Jian Sun 0003 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2019 | Safe navigation of quadrotors with jerk limited trajectoryabstractMany aerial applications require unmanned aerial systems operate in safe zones because of the presence of obstacles or security regulations. It is a non-trivial task to generate a smooth trajectory satisfying both dynamic constraints and motion limits of the unmanned vehicles while being inside the safe zones. Then the task becomes even more challenging for real-time applications, for which computational efficiency is crucial. In this study, we present a safe flying corridor navigation method, which combines jerk limited trajectories with an efficient testing method to update the position setpoints in real time. Trajectories are generated online and incrementally with a cycle time smaller than 10 μs, which is exceptionally suitable for vehicles with limited onboard computational capability. Safe zones are represented with multiple interconnected bounding boxes which can be arbitrarily oriented. The jerk limited trajectory generation algorithm has been extended to cover the cases with asymmetrical motion limits. The proposed method has been successfully tested and verified in flight simulations and actual experiments. Shupeng Lai, Menglu Lan, Ben M. Chen |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | SO-Net: Self-Organizing Network for Point Cloud AnalysisabstractThis paper presents SO-Net, a permutation invariant architecture for deep learning with orderless point clouds. The SO-Net models the spatial distribution of point cloud by building a Self-Organizing Map (SOM). Based on the SOM, SO-Net performs hierarchical feature extraction on individual points and SOM nodes, and ultimately represents the input point cloud by a single feature vector. The receptive field of the network can be systematically adjusted by conducting point-to-node k nearest neighbor search. In recognition tasks such as point cloud reconstruction, classification, object part segmentation and shape retrieval, our proposed network demonstrates performance that is similar with or better than state-of-the-art approaches. In addition, the training speed is significantly faster than existing point cloud recognition networks because of the parallelizability and simplicity of the proposed architecture. Our code is available at the project website1. Ben M. Chen, Gim Hee Lee |
CVPR | 2 |
| 2018 | Adaptive Weight Multi-Band Blending Based Fast Aerial Image Stitching and MappingabstractIn this paper, we implement a real-time method for incrementally stitching aerial images over a large area. To collect high-quality aerial images, the camera is mounted on a gimbal. Being isolated from the vibration of the drone, the camera can provide relatively stable and less blurry images. To achieve fast aerial mapping, instead of using traditional feature extraction and matching steps, this algorithm only relies on camera position and orientation for blending. The mosaic is fused and incrementally updated based on the adaptive weight multi-band blending algorithm. The weight matrix and Laplacian pyramid for blending are updated adaptively. The orthoimage can be reconstructed from the Laplacian pyramid incrementally. This method was applied in the aerial mapping task of the ninth International Micro Air Vehicle Conference and Flight Competition, achieving a high quality map that contributed to the team's championship in the outdoor competition. Xiaodong Liu 0008, Yu Herng Tan, Ben M. Chen |
ICARCV | 3 |
| 2018 | Development of an Autonomous Unmanned Surface Vehicle with Object Detection Using Deep LearningabstractA large number of research has been accomplished in the field of the Unmanned Surface Vehicle (USV) in recent years. As deep learning has the potential to raise the technology to the next level by teaching the algorithm to learn by itself, we aim at developing an autonomous USV which has the capabilities to acquire various types of data and information in offshore areas, process them, and then execute missions based on the situation with the aid of the deep convolutional neural network. This paper describes the implementation of such USV system outfitted with sensors for localization with autonomous navigation technologies and algorithms being adopted for the potential real-life applications such as identifying approaching vehicles to alert the ground station or exploring the surrounding environment of assigned locations. In this manuscript, Global Positioning System (GPS) and compass are equipped to provide the geolocation and the heading for autonomous navigation. Experimental results are provided to validate the proposed implementation. In the end, a summary of current progress is presented as well as the proposed future works. Junji Zhu, Feng Lin 0003, Ben M. Chen |
IECON | 5 |
| 2018 | Optimal Constrained Trajectory Generation for Quadrotors Through Smoothing SplinesabstractIn this paper, we present a trajectory generation method for quadrotors based on the optimal smoothing B-spline. Compared to existing methods which rely on polynomial splines or time optimal control techniques, our method systematically addresses the issue of axes-coupled and interval-wise constraints. These constraints can be used to construct safe flying zones and satisfy vehicle's physical limits. The proposed approach has also been extended to generate trajectories from the nominal plan which consists of not only points but also lines and planes, opening a door for new improvements and applications. Moreover, a closed-form solution can be obtained for cases without inequality constraints. Such a solution is numerically stable for the large-scale fitting problem, which allows us to directly fit the human sketching input from the touch device and capture all subtle details. Our approach is verified by various real flight experiments.. Shupeng Lai, Menglu Lan, Ben M. Chen |
IROS | 3 |
| 2017 | Deep learning for 2D scan matching and loop closureabstractAlthough 2D LiDAR based Simultaneous Localization and Mapping (SLAM) is a relatively mature topic nowadays, the loop closure problem remains challenging due to the lack of distinctive features in 2D LiDAR range scans. Existing research can be roughly divided into correlation based approaches e.g. scan-to-submap matching and feature based methods e.g. bag-of-words (BoW). In this paper, we solve loop closure detection and relative pose transformation using 2D LiDAR within an end-to-end Deep Learning framework. The algorithm is verified with simulation data and on an Unmanned Aerial Vehicle (UAV) flying in indoor environment. The loop detection ConvNet alone achieves an accuracy of 98.2% in loop closure detection. With a verification step using the scan matching ConvNet, the false positive rate drops to around 0.001%. The proposed approach processes 6000 pairs of raw LiDAR scans per second on a Nvidia GTX1080 GPU. Huangying Zhan, Ben M. Chen, Ian D. Reid 0001, Gim Hee Lee |
IROS | 3 |
| 2017 | Online schedule for autonomy of multiple unmanned aerial vehicles
Kemao Peng, Feng Lin 0003, Ben M. Chen |
Sci. China Inf. Sci. | 3 |
| 2017 | Autonomous reconfigurable hybrid tail-sitter UAV U-Lion
Kangli Wang, Yijie Ke, Ben M. Chen |
Sci. China Inf. Sci. | 3 |
| 2016 | Autonomous Mission Management for Forest Search with Multiple Unmanned Aerial VehiclesabstractAn autonomous mission management (AMM) system is designed with the enhanced hierarchical-distributed methodology (HDM) for multiple unmanned aerial vehicles (UAVs) to search a field of forest together. The main ideas of the enhanced HDM are hierarchical control and distributed implementation. The event control law is partitioned into the group and individual event control laws. The group event control law is to coordinate the group of UAVs to complete the designated mission and the individual event control laws are to complete the assigned submissions/ tasks accordingly. The group event control law is executed by the leader and any member can be designated or selected as the leader on the rules. The forest search is applied to verify the designed AMM system in simulation. The simulation results demonstrate that the designed AMM system is successful to complete the designated mission by collaborating the group of UAVs. Kemao Peng, Feng Lin 0003, Ben M. Chen |
ICINCO (1) | 3 |
| 2016 | BIT*-based path planning for micro aerial vehiclesabstractThis paper presents a 3D on-line path planning algorithm for micro sized aerial vehicles (MAVs). The proposed approach adopts a two-layered planning framework. The first layer of the algorithm utilizes a sampling-based planner named Batch Informed Trees (BIT*) to quickly find an geometric obstacle-free passage. The second layer takes into account the dynamic constrains of the vehicle. By adopting a two-point boundary value problems (TPBVPs) approach, dynamically feasible trajectories can be generated efficiently within the previously found passage for lower-level controller. The main contribution of this work is proposing a complete on-line 3D path planning algorithm which can be implemented on the MAV with limited computational power. Menglu Lan, Shupeng Lai, Yingcai Bi, Hailong Qin, Feng Lin 0003, Ben M. Chen |
IECON | 7 |
| 2016 | Semi-dense motion segmentation for moving cameras by discrete energy minimizationabstractWe present an approach for two-view motion segmentation for freely moving cameras, by formulating the epipolar constraint and spacial consistency into a discrete energy minimization problem, which can be efficiently solved using graph cut algorithms. With dense optical flow and proper sampling, a set of matched points is acquired for computing the fundamental matrix and the corresponding epipolar lines. The points distance to the epipolar lines, and their position on the image plane are used to construct a Markov Random Field (MRF) with discrete label-space. The 2-dimension label-space, i.e. motion area or static area, is computed using graph cut algorithms. We demonstrate the effectiveness of our method with the Johns Hopkins 155 motion dataset. Mo Shan, Menglu Lan, Yingcai Bi, Hailong Qin, Feng Lin 0003, Ben M. Chen |
IECON | 7 |
| 2016 | A brief survey of visual odometry for micro aerial vehiclesabstractRecently, visual odometry (VO) has experienced a rapid growth, which makes it viable for a range of applications. This survey paper attempts to provide a timely and comprehensive review of this field, focusing specifically on micro aerial vehicles (MAVs), with monocular, stereo or RGB-D cameras onboard. In this survey, the milestones in the development of VO will be reviewed, followed by an illustration of its general workflow, the commonly used datasets. The survey is concluded by an overall discussion. Mo Shan, Yingcai Bi, Hailong Qin, Zhi Gao 0005, Feng Lin 0003, Ben M. Chen |
IECON | 7 |
| 2016 | Survey of autopilot for multi-rotor unmanned aerial vehiclesabstractRecently, multi-rotor unmanned aerial vehicles (UAVs) have grown rapidly in popularity due to their simple structure, low cost and ease of operation. And they have been widely used in both academic research and commercial applications. As a core component of UAV, autopilot is normally used to realize autonomous control and navigation, including platform stabilization, flight trajectory generation, mission planning and so on. In this paper, we present a survey of existing autopilots for multi-rotor UAVs. First, we will introduce background and principle configuration of autopilots. Second, we will compare and analyze several popular autopilots in the market. Besides, several quality control methodologies will be introduced to improve the reliability and robustness of autopilots. Finally, we will discuss possible trend and technologies for autopilots in future. Zhaolin Yang, Feng Lin 0003, Ben M. Chen |
IECON | 3 |
| 2015 | A statistical approach for trajectory analysis and motion segmentation for freely moving camerasabstractThis paper examines the problem of motion segmentation by analyzing trajectories with statistical approach. We propose a statistical framework for motion segmentation, which makes no assumption on camera motion, camera model, number of moving objects and scene complexity. Long range trajectories are traced across frames and clustered by DTW metric. Various descriptors can be used to construct a weighted neighbor graph for the resulted clusters, following by spectral clustering to retrieve trajectories associated with motion. This framework is highly extensive because different descriptors can be combined into the bag-of-features, to build a more accurate neighbor graph to achieve better result. The algorithm is evaluated mainly with the Hopkins 155 database. Feng Lin 0003, Ben M. Chen |
IECON | 3 |
| 2015 | Wide area surveillance of urban environments using multiple Mini-VTOL UAVsabstractIn this paper, a system for the wide area surveillance of general urban environments using multiple Mini-VTOL UAVs is developed. Given the information of terrain and buildings in the target area, the problem of (robust) complete coverage of the urban environment is solved by a three-step approximation approach. Firstly, the target area and the observation area are discretized into two sets respectively. Secondly, the visibility between these two sets is checked. Finally, a set covering problem is solved based on the greedy approaches. Two case studies based on real-world data are carried out to demonstrate the effectiveness of our developed system. Mohammad Karimadini, Cheng Xiang 0001, Rodney Teo, Ben M. Chen, Tong Heng Lee |
IECON | 5 |
| 2015 | A high fidelity simulator for a quadrotor UAV using ROS and GazeboabstractFlight tests of prototype UAV systems can be restricted by spatial constraints and they may bring risks of damage due to failures. Motivated by these, we presented a simulation approach based on Robot Operating System (ROS) and Gazebo. Unlike other state-of-the-art quadrotor simulators, we implemented the dynamics model of the UAV in ROS to achieve high fidelity behavior of the UAV. A hierarchical navigation system is also presented in our paper. The system layers include simultaneous localization and mapping (SLAM), mapping framework in Cartesian and polar coordinates, A* global path planner, revised vector field histogram plus (VFH+) for optimal local path selection and online trajectory algorithm (OTA) with collision checking for obstacle avoidance. In order to cater for vision-based applications, quadrotor is equipped with a monocular camera in the simulation model. The implementation of circle and landing pad detection and tracking algorithm demonstrates the functionality of vision guidance. In our simulation, various aspects including complex indoor and outdoor environments and on-board sensors are capable of simultaneously interacting with our navigation system to achieve certain surveillance missions. In the end, we demonstrated the applicability of our complex quadrotor systems by performing an autonomous navigation task in simulated complex environments. In comparison with the experimental data, simulation results align with the ones in flight tests in terms of real flight behaviors during navigation tasks in general. Mengmi Zhang, Hailong Qin, Menglu Lan, Feng Lin 0003, Ben M. Chen |
IECON | 8 |
| 2014 | Vision-based detection and pose estimation for formation of micro aerial vehiclesabstractThis paper proposes an innovative method to detect micro aerial vehicles (MAVs) and estimate their relative pose in formation using a monocular on-board camera. Haar classifier is trained for autonomously detecting MAV in open scenes, like grasslands or obstruct-free playgrounds. In order to increase the robustness of the detection, a Kaiman filter has been employed to conduct image tracking. Contours of detected MAV have been extracted for shape matching. Point sets quantized from contours match with the given point sets using Hungarian algorithm and relaxation labeling based on shape contexts. Two techniques, affine and thin plate spline (TPS) transformation, are explored, while TPS is better in dealing with distorted shapes. In experiments, we develop and implement an innovative 2D shape-based pose estimation method by using only one monocular camera which results in fast and accurate performances. Mengmi Zhang, Feng Lin 0003, Ben M. Chen |
ICARCV | 3 |
| 2014 | Identification of stock market forces in the system adaptation framework
Xiaolian Zheng, Ben M. Chen |
Inf. Sci. | 2 |
| 2012 | Construction and Modeling of a Variable Collective Pitch Coaxial UAV
Jinqiang Cui, Fei Wang 0018, Zhengyin Qian, Ben M. Chen, Tong Heng Lee |
ICINCO (2) | 4 |
| 2012 | Minimum-time trajectory planning for helicopter UAVs using computational dynamic optimizationabstractIn this paper, we apply pseudospectral optimal control, a computational algorithm of dynamic optimization, to the problem of helicopter UAV for minimum-time trajectory planning in the presence of obstacles. The problem is formulated as a nonlinear optimal control subject to the dynamics and limitations of helicopter UAVs, in which the obstacles are formulated as inequality constraints by using p-norms. The dynamical system is defined by a set of fifteen states nonlinear differential equations developed for HeLion, a helicopter UAV constructed in National University of Singapore (NUS). The problem does not have an analytic solution. We numerically solve the problem using a pseudospectral method. Various terrain scenarios were tested, from a single obstacle to multiple obstacles. We found bifurcation points of minimum-time trajectories near obstacles. The bifurcation points and their relationship with the distance to the obstacle are analyzed. Nathan Xu, Wei Kang 0001, Ben M. Chen |
SMC | 4 |
| 2010 | Graphic interpretations of structural controllability for switched linear systemsabstractThis paper considers the controllability problem for switched linear systems. In particular, the structural controllability of switched linear systems is investigated. The structural controllability of switched linear systems is a generalization of the traditional controllability concept for dynamical systems, and purely based on the graphic topologies among state and input nodes. First, two kinds of graphic representations of switched linear systems are proposed. Second, several graph-theoretic characterizations of the structural controllability for switched linear systems are presented based on these two newly introduced graphs. Finally, the paper concludes with several illustrative examples and discussions of the results and future work. Hai Lin 0002, Ben M. Chen |
ICARCV | 3 |
| 2009 | Development of a vision-based ground target detection and tracking system for a small unmanned helicopter
Feng Lin 0003, Kai-Yew Lum, Ben M. Chen, Tong Heng Lee |
Sci. China Ser. F Inf. Sci. | 3 |
| 2006 | Discrete-time Robust Nonlinear Feedback Control for an HDD Servo System DesignabstractThis paper presents a discrete-time robust nonlinear control method to achieve fast and accurate set-point tracking for servo systems subject to actuator saturation and disturbances. The idea here is to use a combination of composite nonlinear feedback (CNF) control and disturbance estimation cum compensation. The CNF control is responsible for superior transient performance, i.e., to guarantee a fast response with low overshoot, while the disturbance estimator/compensator is used to remove the steady state bias that would otherwise be existent due to disturbances. Practical application in a micro hard disk drive servo system will be given to demonstrate the effectiveness of this control method Guoyang Cheng, Kemao Peng, Ben M. Chen, Tong Heng Lee |
ICARCV | 3 |
| 2006 | Explicit Constructions of Global Stabilization Control Laws for a Class of Nonminimum Phase Nonlinear SystemsabstractThis paper addresses a global stabilization problem for a class of nonminimum phase nonlinear systems. The nonlinearities of the system, which depend on the system output, can be unknown, but satisfy some linear growth conditions. The given system is first transformed into a special coordinate basis, in which the system zero dynamics is divided into a stable part and an unstable part. A sufficient solvability condition is then established for solving the global stabilization problem. Finally, the obtained result is utilized to solve a stabilization problem on a rotational/translational actuator (RTAC) system Weiyao Lan, Ben M. Chen |
ICARCV | 2 |
| 2006 | A partition approach for the restoration of camera images of planar and curled document
Shijian Lu, Ben M. Chen, Chi Chung Ko |
Image Vis. Comput. | 2 |
| 2005 | Perspective rectification of document images using fuzzy set and morphological operations
Shijian Lu, Ben M. Chen, Chi Chung Ko |
Image Vis. Comput. | 2 |
| 2004 | A MATLAB toolkit for composite nonlinear feedback controlabstractWe present in this article a MATLAB toolkit with a user-friendly graphical interface for composite nonlinear feedback control system design. The toolkit can be utilized to design a fast and smooth tracking controller for a class of linear systems with actuator and other nonlinearities as well as with external disturbances. The toolkit is capable of displaying both time-domain and frequency-domain responses on its main panel, and generating three different types of control laws, namely, the state feedback, the full order measurement feedback and the reduced order measurement feedback controllers. The usage and design procedure of the toolkit are illustrated by a practical example on the design of a hard disk drive servo system. The toolkit can be utilized to design servo systems that deal with point-and-shoot fast targeting. Guoyang Cheng, Ben M. Chen, Kemao Peng, Tong Heng Lee |
ICARCV | 2 |
| 2004 | Document image rectification using fuzzy sets and morphological operatorsabstractIn this paper, we deal with the problem of document image rectification from images captured by digital cameras. The improvement on the resolution of digital camera sensors has brought more and more applications for non-contact text capture. Unfortunately, perspective distortion coupled with resulting images makes it harder to properly identify the contents of captured texts using the traditional optical character recognition (OCR) system. We propose in this work a new technique, which is capable of removing distortion and recovering the fronto-parallel view of text with a single image. Different from reported approaches in the literature, the image rectification is carried out using character boundary and tip point, which are extracted from character strokes based on multiple fuzzy sets and morphological operators. The algorithm needs neither camera calibration nor high-contrast document boundary. Experimental results show our rectification process is fast and robust. Shijian Lu, Ben M. Chen, Chi Chung Ko |
ICIP | 2 |
| 2001 | A web-based virtual laboratory on a frequency modulation experimentabstractWith the rapid proliferation of Internet technologies, accessing and operating engineering instruments remotely anytime anywhere is fast becoming a reality. This paper presents a new Web-based virtual laboratory on a frequency modulation experiment for the teaching of an undergraduate course on communication principles in the National University of Singapore (NUS). The laboratory requires only a common Web browser to access and incorporate schemes for reducing data traffic and authenticating users. It enables students to have a natural hands-on experience of using an expensive spectrum analyzer on a one-to-one basis and provides a solution for distant engineering education. The system uses a double client-server structure where access to the experiment is via two rounds of client-server processing. The virtual laboratory can be accessed at the Web site http://vlab.ee.nus.edu.sg/vlab/freqmod/index.html. Chi Chung Ko, Ben M. Chen, Shaoyan Hu, V. Ramakrishnan, Chang Dong Cheng |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 1998 | Development of a Multi-Channel PC-Based Hard Disk Drive Bode-Plot GeneratorabstractA five-channel generator is developed to obtain the frequency response of a hard disk drive (HDD) by emulating the Bode-plot functions of a dynamic signal analyser (DSA). Written in LabVIEW, it runs on the Windows 95 platform using a standard Wintel PC, a commercial data acquisition (DAQ) board and an external hardware anti-alias filter board. This generator achieves a superior advantage at lower cost through its ability to evaluate the track-following servo loop frequency response of up to five HDDs simultaneously. Besides reducing test-time and providing a degree of accuracy comparable to a commercial DSA, it also includes an option to save all acquired data to disk for further analysis with other software engineering packages. Kin Wee Choo, Guoxiao Guo, Ben M. Chen |
Asian Test Symposium | 3 |