Huai Yu

dblp:197/2733 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-6043-3412ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 UniCalib: Targetless LiDAR-camera Calibration via Probabilistic Flow on Unified Depth Representations
abstract
Online targetless extrinsic LiDAR-camera calibration is essential for robust perception in computer vision applications such as autonomous driving. However, existing methods struggle with the significant modality gap between heterogeneous sensors and fail to handle unreliable correspondences arising from real-world challenges like occlusions and dynamic objects. To address these issues, we introduce UniCalib, a novel method that performs calibration by estimating a probabilistic flow on unified depth representations. UniCalib first bridges the modality gap by converting both the camera images and the sparse LiDAR points into unified, dense depth maps, enabling a unified encoder to learn consistent features. Subsequently, it learns a probabilistic flow field that captures the correspondence uncertainty to improve robustness. This probabilistic approach is reinforced by a reliability map and a perceptually weighted sparse flow loss, which guide the model to suppress the influence of unreliable regions. Experimental results on three datasets validate the accuracy and generalization of UniCalib. In particular, it achieves a mean translation error of 0.550cm and a rotation error of 0.044° on the KITTI dataset. The code is available at https://github.com/han-15/UniCalib.
Xubo Zhu, Ji Wu 0012, Ximeng Cai, Wen Yang 0001, Huai Yu, Gui-Song Xia
WACV6
2026 QuadricsReg: Large-Scale Point Cloud Registration Using Semantic Quadric Primitives
Ji Wu 0012, Huai Yu, Ximeng Cai, Mingfeng Wang, Wen Yang 0001, Gui-Song Xia
IEEE Trans. Robotics2
2025 LiDAR-enhanced 3D Gaussian Splatting Mapping
abstract
This paper introduces LiGSM, a novel LiDARenhanced 3D Gaussian Splatting (3DGS) mapping framework that improves the accuracy and robustness of 3D scene mapping by integrating LiDAR data. LiGSM constructs joint loss from images and LiDAR point clouds to estimate the poses and optimize their extrinsic parameters, enabling dynamic adaptation to variations in sensor alignment. Furthermore, it leverages LiDAR point clouds to initialize 3DGS, providing a denser and more reliable starting points compared to sparse SfM points. In scene rendering, the framework augments standard image-based supervision with depth maps generated from LiDAR projections, ensuring an accurate scene representation in both geometry and photometry. Experiments on public and self-collected datasets demonstrate that LiGSM outperforms comparative methods in pose tracking and scene rendering.
Huai Yu, Ji Wu 0012, Wen Yang 0001, Gui-Song Xia
ICRA2
2025 Unsupervised Multiview UAV Image Geolocalization via Iterative Rendering
abstract
Unmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) poses significant challenges due to the substantial view discrepancies between oblique UAV images and overhead satellite images. Existing methods heavily rely on supervised learning with labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these approaches are computationally expensive, prone to overfitting region-specific cues, and exhibit limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3D space from UAV observations for satellite image generation, providing a robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the perspective of the rendered image with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods.
Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Li Mi, Huai Yu, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.5
2024 FE-DeTr: Keypoint Detection and Tracking in Low-quality Image Frames with Events
abstract
Keypoint detection and tracking in traditional image frames are often compromised by image quality issues such as motion blur and extreme lighting conditions. Event cameras offer potential solutions to these challenges by virtue of their high temporal resolution and high dynamic range. However, they have limited performance in practical applications due to their inherent noise in event data. This paper advocates fusing the complementary information from image frames and event streams to achieve more robust keypoint detection and tracking. Specifically, we propose a novel keypoint detection network that fuses the textural and structural information from image frames with the high-temporal-resolution motion information from event streams, namely FE-DeTr. The network leverages a temporal response consistency for supervision, ensuring stable and efficient keypoint detection. Moreover, we use a spatio-temporal nearest-neighbor search strategy for robust keypoint tracking. Extensive experiments are conducted on a new dataset featuring both image frames and event data captured under extreme conditions. The experimental results confirm the superior performance of our method over both existing frame-based and event-based methods. Our code, pre-trained models, and dataset are available at https://github.com/yuyangpoi/FE-DeTr.
Xiangyuan Wang, Kuangyi Chen, Wen Yang 0001, Lei Yu 0006, Yannan Xing, Huai Yu
ICRA6
2024 QuadricsNet: Learning Concise Representation for Geometric Primitives in Point Clouds
abstract
This paper presents a novel framework to learn a concise geometric primitive representation for 3D point clouds. Different from representing each type of primitive individually, we focus on the challenging problem of how to achieve a concise and uniform representation robustly. We employ quadrics to represent diverse primitives with only 10 parameters and propose the first end-to-end learning-based framework, namely QuadricsNet, to parse quadrics in point clouds. The relationships between quadrics mathematical formulation and geometric attributes, including the type, scale and pose, are insightfully integrated for effective supervision of QuaidricsNet. Besides, a novel pattern-comprehensive dataset with quadrics segments and objects is collected for training and evaluation. Experiments demonstrate the effectiveness of our concise representation and the robustness of QuadricsNet. Our code is available at https://github.com/MichaelWu99-lab/QuadricsNet.
Ji Wu 0012, Huai Yu, Wen Yang 0001, Gui-Song Xia
ICRA2
2024 Detecting Line Segments in Motion-Blurred Images With Events
abstract
Making line segment detectors more reliable under motion blurs is one of the most important challenges for practical applications, such as visual SLAM and 3D line mapping. Existing line segment detection methods face severe performance degradation for accurately detecting and locating line segments when motion blur occurs. While event data shows strong complementary characteristics to images for minimal blur and edge awareness at high-temporal resolution, potentially beneficial for reliable line segment recognition. To robustly detect line segments over motion blurs, we propose to leverage the complementary information of images and events. Specifically, we first design a general frame-event feature fusion network to extract and fuse the detailed image textures and low-latency event edges, which consists of a channel-attention-based shallow fusion module and a self-attention-based dual hourglass module. We then utilize the state-of-the-art wireframe parsing networks to detect line segments on the fused feature map. Moreover, due to the lack of line segment detection datasets with pairwise motion-blurred images and events, we contribute two datasets, i.e., synthetic FE-Wireframe and realistic FE-Blurframe, for network training and evaluation. Extensive analyses on the component configurations demonstrate the design effectiveness of our fusion network. When compared to the state-of-the-arts, the proposed approach achieves the highest detection accuracy while maintaining comparable real-time performance. In addition to being robust to motion blur, our method also exhibits superior performance for line detection under high dynamic range scenes.
Huai Yu, Hao Li 0114, Wen Yang 0001, Lei Yu 0006, Gui-Song Xia
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning Cross-View Visual Geo-Localization Without Ground Truth
abstract
Cross-view geo-localization (CVGL) involves determining the geographical location of a query image by matching it with a corresponding GPS-tagged reference image. Current state-of-the-art methods predominantly rely on training models with labeled paired images, incurring substantial annotation costs and training burdens. In this study, we investigate the adaptation of frozen models for CVGL without requiring ground-truth pair labels. We observe that training on unlabeled cross-view images presents significant challenges, including establishing relationships within unlabeled data and reconciling view discrepancies between uncertain queries and references. To address these challenges, we propose a self-supervised learning framework to train a learnable adapter for a frozen foundation model (FM). This adapter is designed to map feature distributions from diverse views into a uniform space using unlabeled data exclusively. To establish relationships within unlabeled data, we introduce an expectation-maximization (EM)-based pseudolabeling module, which iteratively estimates matching between cross-view features and optimizes the adapter. To maintain the robustness of the FM’s representation, we incorporate an information consistency module with a reconstruction loss, ensuring that adapted features retain strong discriminative ability across views. Experimental results demonstrate that our proposed method achieves significant improvements over vanilla FMs and competitive accuracy compared to supervised methods while necessitating fewer training parameters and relying solely on unlabeled data. Evaluation of our adaptation for task-specific models further highlights its broad applicability. Particularly, on the University-1652 dataset, our method outperforms the FM baseline by a substantial margin, achieving about 39 points improvement in Recall@1 and more than 34 points increase in average precision (AP). The project is available athttps://collebt.github.io/EM-CVGL.
Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Huai Yu, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.4
2024 Extracting Building Footprints in SAR Images via Distilling Boundary Information From Optical Images
abstract
Buildings represent pivotal entities in remote sensing imagery for various applications like urban planning and land resource management. Predominantly, methods for building footprint extraction in the literature focus on optical imagery with visual attributes that faithfully mirror the physical world. Nevertheless, the acquisition of high-quality optical images presents formidable challenges due to the susceptibility to illumination conditions and scene visibility. In contrast, synthetic aperture radar (SAR) images can be acquired in all-weather and all-time situations, unburdened by the aforementioned constraints. However, the coherent imaging mechanism engenders intricate complexities for building footprint extraction SAR images. To address this issue, this paper introduces the Boundary Information Distillation Network (BIDNet) to improve the prediction accuracy in SAR images by distilling knowledge from optical images. The proposed approach adopts a teacher-student framework, featuring two customized components: the Explicit Distillation Module (EDM) and the Latent Distillation Module (LDM). Different from the conventional practice of directly aligning feature maps, BIDNet focuses on leveraging the more conspicuous boundary information in optical images. The EDM operates by simultaneously yielding a boundary map to emphasize the boundary area and assimilating the explicit low-level features of two modalities. The LDM represents the structural attributes within the high-level latent feature space and aligns the representations of the two modalities. Within this module, intrinsic self-correlations among features originating from boundary regions are encoded, and so are the cross-correlations established between features from boundary regions and alternative areas. The two modules also serve as the conduit for knowledge distillation from the teacher network to the student network, enabling the utilization of optical imagery for enhancing the building footprint extraction in SAR imagery. Extensive experiments demonstrate that our BIDNet achieves state-of-the-art performance on the Multi-Sensor All Weather Mapping (MSAW) dataset, outperforming the strong baseline by 4.3-7.2 points in f1-score and 4.9-8.0 points in IoU. The source code and trained models will be publicly available.
Lanxin Zeng, Wen Yang 0001, Jian Kang 0005, Huai Yu, Mihai Datcu, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.5
2023 PyPose: A Library for Robot Learning with Physics-based Optimization
abstract
Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-level semantic information and reliance on manual parametric tuning. To take advantage of these two complementary worlds, we present PyPose: a robotics-oriented, PyTorch-based library that combines deep perceptual models with physics-based optimization. PyPose's architecture is tidy and well-organized, it has an imperative style interface and is efficient and user-friendly, making it easy to integrate into real-world robotic applications. Besides, it supports parallel computing of any order gradients of Lie groups and Lie algebras and 2nd-order optimizers, such as trust region methods. Experiments show that PyPose achieves more than 10× speedup in computation compared to the state-of-the-art libraries. To boost future research, we provide concrete examples for several fields of robot learning, including SLAM, planning, control, and inertial navigation.
Chen Wang 0033, Dasong Gao, Junyi Geng, Yaoyu Hu, Yuheng Qiu, Bowen Li 0007, Fan Yang 0092, Brady G. Moon, Abhinav Pandey, Aryan, Jiahe Xu 0002, Daning Huang, Zhongqiang Ren, Shibo Zhao, Taimeng Fu, Pranay Reddy, Jingnan Shi, Rajat Talak, Kun Cao 0002, Yi Du 0001, Huai Yu, Shanzhao Wang, Siyu Chen 0036, Ananth Kashyap, Rohan Bandaru, Karthik Dantu, Jiajun Wu 0001, Lihua Xie 0001, Luca Carlone, Marco Hutter 0001, Sebastian A. Scherer
CVPR27
2023 Dynamic Coarse-to-Fine Learning for Oriented Tiny Object Detection
abstract
Detecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce severe mismatch and imbalance issues. Specifically, the position prior, positive sample feature, and instance are mismatched, and the learning of extreme-shaped objects is biased and unbalanced due to little proper feature supervision. To tackle these issues, we propose a dynamic prior along with the coarse-to-fine assigner, dubbed DCFL. For one thing, we model the prior, label assignment, and object representation all in a dynamic manner to alleviate the mismatch issue. For another, we leverage the coarse prior matching and finer posterior constraint to dynamically assign labels, providing appropriate and relatively balanced supervision for diverse instances. Extensive experiments on six datasets show substantial improvements to the baseline. Notably, we obtain the state-of-the-art performance for one-stage detectors on the DOTA-v1.5, DOTA-v2.0, and DIOR-R datasets under single-scale training and testing. Codes are available at https://github.com/Chasel-Tsui/mmrotate-dcfl.
Chang Xu 0027, Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia
CVPR5
2023 Self-Supervised Dense Depth Estimation with Panoramic Image and Sparse Lidar
abstract
The 360-depth estimation with spherical images and LiDAR data has recently become increasingly popular in autonomous driving and scene reconstruction. Compared with perspective images, spherical images have omnidirectional FoV, which exceedingly matches LiDAR data. However, the spherical distortion makes the 360-depth estimation a great challenge. To address this problem, we propose a self-supervised 360 depth estimation network in this paper. The network consists of a spherical convolution branch to extract panoramic image features and a ResNet branch to extract LiDAR features. Then an attention-based decoder is designed to estimate the depth. The reprojection error is used to self-supervise the network training. Experiments on the KITTI-360 dataset demonstrate the effectiveness of the proposed method.
Chenwei Lyu, Huai Yu, Zhipeng Zhao 0001, Pengliang Ji, Xiangli Yang, Wen Yang 0001
IGARSS2
2023 Cross-Modal 2D-3D Localization with Single-Modal Query
abstract
Global visual localization is an important task in geoscience with a plethora of applications such as SLAM and autonomous navigation. Current place recognition approaches restrict the modality of the query data which relies on the database data modality. However, real-world robots are equipped with different sensors in different application scenarios and it is difficult for data from a single fixed modality to accommodate all challenging environments. To overcome this limitation, we propose to build a generalized model that allows spherical images and point clouds to be retrieved under any single-modal query. Our 2D-3D dataset is created based on the KITTI360 dataset with spherical images and corresponding point clouds for training and evaluation. Extensive experimental results demonstrate the effectiveness of our proposed approach.
Zhipeng Zhao 0001, Huai Yu, Chenwei Lyu, Pengliang Ji, Xiangli Yang, Wen Yang 0001
IGARSS2
2022 RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection
Chang Xu 0027, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia
ECCV (9)4
2022 Unified Representation of Geometric Primitives for Graph-SLAM Optimization Using Decomposed Quadrics
abstract
In Simultaneous Localization And Mapping (SLAM) problems, high-level landmarks have the potential to build compact and informative maps compared to traditional point-based landmarks. In this work, we focus on the param-eterization of frequently used geometric primitives including points, lines, planes, ellipsoids, cylinders, and cones. We first present a unified representation based on quadrics, an algebraic representation of quadratic surfaces in 3D. Then we propose a decomposed model of quadrics that discloses the symmetry and degeneration properties of a primitive. Based on the decomposition, we develop geometrically meaningful quadrics factors for the graph-SLAM problem. Then in simulation, it is shown that the decomposed formulation has better efficiency and robustness to observation noises than baseline parame-terizations. Finally, in real-world experiments, the proposed back-end framework is demonstrated to be capable of building compact and regularized maps.
Weikun Zhen, Huai Yu, Yaoyu Hu, Sebastian A. Scherer
ICRA2
2022 Towards Robust Visual-Inertial Odometry with Multiple Non-Overlapping Monocular Cameras
abstract
We present a Visual-Inertial Odometry (VIO) algorithm with multiple non-overlapping monocular cameras aiming at improving the robustness of the VIO algorithm. An initialization scheme and tightly-coupled bundle adjustment for multiple non-overlapping monocular cameras are proposed. With more stable features captured by multiple cameras, VIO can maintain stable state estimation, especially when one of the cameras tracked unstable or limited features. We also address the high CPU usage rate brought by multiple cameras by proposing a GPU-accelerated frontend. Finally, we use our pedestrian carried system to evaluate the robustness of the VIO algorithm in several challenging environments. The results show that the multi-camera setup yields significantly higher estimation robustness than a monocular system while not increasing the CPU usage rate (reducing the CPU resource usage rate and computational latency by 40.4% and 50.6% on each camera). A demo video can be found at https://youtu.be/r7QvPth1m10.
Huai Yu, Wen Yang 0001, Sebastian A. Scherer
IROS2
2022 Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation
abstract
We present radiance field propagation (RFP), a novel approach to segmenting objects in 3D during reconstruction given only unlabeled multi-view images of a scene. RFP is derived from emerging neural radiance field-based techniques, which jointly encodes semantics with appearance and geometry. The core of our method is a novel propagation strategy for individual objects' radiance fields with a bidirectional photometric loss, enabling an unsupervised partitioning of a scene into salient or meaningful regions corresponding to different object instances. To better handle complex scenes with multiple objects and occlusions, we further propose an iterative expectation-maximization algorithm to refine object masks. To the best of our knowledge, RFP is the first unsupervised approach for tackling 3D scene object segmentation for neural radiance field (NeRF) without any supervision, annotations, or other cues such as 3D bounding boxes and prior knowledge of object class. Experiments demonstrate that RFP achieves feasible segmentation results that are more accurate than previous unsupervised image/scene segmentation approaches, and are comparable to existing supervised NeRF-based methods. The segmented object representations enable individual 3D object editing operations. Codes and datasets will be made publicly available.
Xinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai, Chi-Keung Tang
NeurIPS3
2022 Instance Switching-Based Contrastive Learning for Fine-Grained Airplane Detection
abstract
Detecting airplanes from high-resolution remote sensing images has a variety of applications. The characteristics of clear details, rich spatial and texture information of objects in high-resolution remote sensing images make it possible to identify different types of airplanes from backgrounds. However, airplanes usually exhibit slight inter-class discrepancy and unbalanced class distribution, which pose significant challenges to fine-grained detection of airplanes. In this paper, we propose the ISCL, an Instance Switching-based Contrastive Learning method for fine-grained airplane detection. Specifically, we introduce a Contrastive Learning-based Module (CLM) to widen the inter-class distance while narrowing the intra-class distance by optimizing feature space distribution with the InfoNCE+loss, which is built on a serial head in a cascaded way. Then, we design a Refined Instance Switching (ReIS) module to alleviate the class imbalance problem. To take full advantage of the CLM and ReIS, we further introduce an optimization strategy which is an organic combination of the two modules to widen the distances of different airplane categories that are easily confused. In addition, we contribute a fine-grained attribute-assisted dataset, dubbed GF-RarePlanes Dataset (GRD), to help the detectors better learn the subtle differences between the airplanes. Extensive experiments on two datasets (i.e., GF and FAIR1M) demonstrate that our proposed method can significantly improve the accuracy of fine-grained airplane detection under both HBB and OBB scenarios. Dataset and codes will be available at https://lanxin1011.github.io/ISCL/.
Lanxin Zeng, Haowen Guo, Wen Yang 0001, Huai Yu, Lei Yu 0006, Tongyuan Zou
IEEE Trans. Geosci. Remote. Sens.4
2022 Optical-Enhanced Oil Tank Detection in High-Resolution SAR Images
abstract
In recent years, object detection in high-resolution SAR images has made significant progress, especially after the introduction of deep learning. However, objects like dense oil tanks, which are compactly arranged in SAR images, are still challenging to recognize due to the unique imaging mechanism of SAR. Inspired by human learning from comparison, we propose a multi-stage framework for oil tank detection in SAR images using optical image enhancement. Specifically, in the training stage, we build a teacher-student network to align the semantic information between the two modalities, where the optical features are used to guide the corresponding SAR feature learning. While in the inference stage, the learned network detects oil tanks using only SAR images as input. Besides, a pre-training stage before training is applied to further improve the network’s ability for SAR feature extraction, which is realized by the proposed paired optical-SAR self-supervised learning. To verify the effectiveness of the proposed method, we perform experiments on our newly built SpaceNet6-OTD dataset. Extensive experiments demonstrate that the proposed method can effectively improve the accuracy of detecting oil tanks in SAR images. Datasets, codes, and more results will be released at: https://EIS-VIPG.github.io/SpaceNet6-OTD/.
Ruixiang Zhang, Haowen Guo, Wen Yang 0001, Huai Yu, Gui-Song Xia
IEEE Trans. Geosci. Remote. Sens.5
2021 ORStereo: Occlusion-Aware Recurrent Stereo Matching for 4K-Resolution Images
abstract
Stereo reconstruction models trained on small images do not generalize well to high-resolution data. Training a model on high-resolution image size faces difficulties of data availability and is often infeasible due to limited computing resources. In this work, we present the Occlusion-aware Recurrent binocular Stereo matching (ORStereo), which deals with these issues by only training on available low disparity range stereo images. ORStereo generalizes to unseen high-resolution images with large disparity ranges by formulating the task as residual updates and refinements of an initial prediction. ORStereo is trained on images with disparity ranges limited to 256 pixels, yet it can operate 4K-resolution input with over 1000 disparities using limited GPU memory. We test the model’s capability on both synthetic and real-world high-resolution images. Experimental results demonstrate that ORStereo achieves comparable performance on 4K-resolution images compared to state-of-the-art methods trained on large disparity ranges. Compared to the baseline methods that are only trained on low-resolution images, our method has 60% or less error on 4K-resolution images.
Yaoyu Hu, Huai Yu, Weikun Zhen, Sebastian A. Scherer
IROS3
2020 LiDAR-enhanced Structure-from-Motion
abstract
Although Structure-from-Motion (SfM) as a maturing technique has been widely used in many applications, state-of-the-art SfM algorithms are still not robust enough in certain situations. For example, images for inspection purposes are often taken in close distance to obtain detailed textures, which will result in less overlap between images and thus decrease the accuracy of estimated motion. In this paper, we propose a LiDAR-enhanced SfM pipeline that jointly processes data from a rotating LiDAR and a stereo camera pair to estimate sensor motions. We show that incorporating LiDAR helps to effectively reject falsely matched images and significantly improve the model consistency in large-scale environments. Experiments are conducted in different environments to test the performance of the proposed pipeline and comparison results with the state-of-the-art SfM algorithms are reported.
Weikun Zhen, Yaoyu Hu, Huai Yu, Sebastian A. Scherer
ICRA3
2020 Edge-Driven Object Matching for UAV Images and Satellite SAR Images
abstract
The task of matching between images acquired by terminal equipment and satellites is important and challenging due to the dramatic viewpoint changes and unknown orientations, especially with different imaging sensors. In this paper, we firstly present the task to match the optical/infrared images acquired by UAVs with satellite SAR images. Many previous works mainly focused on matching with the images of the same modality, and may not perform well to our task. To overcome the difficulties caused by the diversity among the three modalities of data, we mine the common features of them and propose a novel edge-driven matching framework to find the correspondence between the UAV images and SAR images. Experimental results demonstrate the effectiveness and superiority of our method.
Ruixiang Zhang, Huai Yu, Wen Yang 0001, Heng-Chao Li 0001
IGARSS3
2020 Monocular Camera Localization in Prior LiDAR Maps with 2D-3D Line Correspondences
abstract
Light-weight camera localization in existing maps is essential for vision-based navigation. Currently, visual and visual-inertial odometry (VO&VIO) techniques are well-developed for state estimation but with inevitable accumulated drifts and pose jumps upon loop closure. To overcome these problems, we propose an efficient monocular camera localization method in prior LiDAR maps using direct 2D-3D line correspondences. To handle the appearance differences and modality gaps between LiDAR point clouds and images, geometric 3D lines are extracted offline from LiDAR maps while robust 2D lines are extracted online from video sequences. With the pose prediction from VIO, we can efficiently obtain coarse 2D-3D line correspondences. Then the camera poses and 2D-3D correspondences are iteratively optimized by minimizing the projection error of correspondences and rejecting outliers. Experimental results on the EurocMav dataset and our collected dataset demonstrate that the proposed method can efficiently estimate camera poses without accumulated drifts or pose jumps in structured environments.
Huai Yu, Weikun Zhen, Wen Yang 0001, Ji Zhang 0003, Sebastian A. Scherer
IROS1
2019 Combined Convolutional and Structured Features for Power Line Detection in UAV Images
abstract
Power line detection plays an important role in automated UAV inspection system, which is crucial for real-time motion planning and navigation along power lines. Previous methods which adopt traditional filters and gradients may fail to capture complete power lines due to noisy background. To overcome this, we develop an accurate power line detection method using rich convolutional and structured features. The proposed method fully exploits multiscale and structured prior information to conduct both accurate and efficient detection. We evaluate the method on two well-annotated power line datasets and achieve state-of-the-art performance compared with previous methods.
Heng Zhang 0011, Wen Yang 0001, Huai Yu
IGARSS3
2018 Bottle Detection in the Wild Using Low-Altitude Unmanned Aerial Vehicles
abstract
In this paper, we propose a new dataset and benchmark for low altitude UAV object detection, aiming to find and localize waste plastic bottles in the wild, as well as to inspire the development of object detection models to be capable of detecting small and transparent objects. To this end, we collect 25, 407 UAV images of bottles with various kinds of backgrounds. Unlike traditional horizontal bounding box based annotation methods, we use the oriented bounding box to accurately and compactly annotate the bottles, which provides more detailed information for subsequent robotic grasping. The fully annotated images contain 34, 791 bottles, each of which is annotated by an arbitrary (5 d.o.f.) quadrilateral. To build a baseline for bottle detection, we evaluate several state-of-the-art object detection algorithms on our UAV-Bottle Dataset (UAV-BD), such as Faster R-CNN, SSD, YOLOv2 and RRPN. We also present an analysis of the dataset along with baseline approaches. Both the dataset and benchmark are made publicly available to the vision community on our website to advance research in the area of object detection from UAVs.
Jinwang Wang, Wei Guo 0006, Huai Yu, Lin Duan, Wen Yang 0001
FUSION4
2018 Accurate Registration of Multitemporal UAV Images Based on Detection of Major Changes
abstract
Accurate registration of multitemporal images captured by UAV usually involves affine transformation and complicated non-rigid transformation, which makes it very difficult to achieve satisfying results. Pixel-wise correspondence is effective for handling images with complicated non-rigidity. However, objects with changes in the scene deform severely because there should be no pixel-wise correspondence but the algorithm erroneously matches the pixels. In this paper, we propose a coarse-to-fine registration method for multitemporal UAV images. First, a projective model is used to eliminate large scale changes as well as perspective distortion. Then the major changes of different temporal UAV images are most detected, which is used to mask the dense matches in changed areas. Finally, the optical flow field method is used to handle complicated non-rigid changes by matching dense SIFT feature. Experimental results on a challenging set of multitemporal UAV images demonstrate the effectiveness of our approach.
Huai Yu, Jinwang Wang, Wen Yang 0001
FUSION2
2017 A low-rank fully convolutional network for classification based on a multi-dimensional description primitive of time series polarimetric sar images
abstract
Time series polarimetric SAR image classification relies on learned understanding of how the set of pixels in an image relate by relative position and how the information of different dates in a time series change as time goes on. In this paper, we firstly integrate the incoherent information in the spatial scale and the coherent information in the temporal scale to form the feature for time series polarimetric SAR images. Then we take advantage of the fully convolutional network (FCN) to make end-to-end, pixels-to-pixels 3-dimensional training, on which base we sparse the connection structure of deep learning network using low rank tensor decomposition to reduce the computational complexity of convolutional layers. Experiment results on a real polarimetric SAR data set preliminarily show the effectiveness of our presented approach.
Chu He, Gong Han, Huai Yu
IGARSS4
2017 Accurate object matching for UAV imagery using multi-scale best-buddies similarity
abstract
This paper presents a multi-scale Best-Buddies Similarity (BBS) method for matching objects in the UAV imagery. More precisely, we first extract proposed regions from the target image by Selective Search, then take a number of proposed regions with the highest confidence scores as templates and finally apply templates to the original BBS method and select the coordinate of the region with the highest overlap rate as the exact position of the query object in the target image. Experimental results on real data show the effectiveness of the proposed method.
Xu Lei 0002, Jinwang Wang, Kaimin Fu, Huai Yu, Wen Yang 0001
IGARSS4
2017 Stable feature point extraction for accurate multi-temporal SAR image registration
abstract
Feature extraction is an important issue for image interpretation, many valuable feature extraction methods have been proposed to address synthetic aperture radar (SAR) image registration. However, the current methods care little about changes over multi-temporal SAR images, which result in unstable output features. The unstable property is a latent factor to affect the accuracy of SAR image registration. To overcome this problem, we propose a new method to detect stable features by intersecting Coherent Scatters (CS) and SAR-FAST corners. Thus the stable features are not only corners but also located at the time-invariant area. Then, the stable features are described by SIFT descriptors and the coarse registration is achieved by matching these stable points. Finally, the Powell algorithm is used to search the optimal Mutual Information location for precise registration. Experimental results demonstrate the effectiveness of the proposed method.
Huai Yu, Wen Yang 0001, Mingsheng Liao
IGARSS1
2017 A UAV-based crack inspection system for concrete bridge monitoring
abstract
Crack inspection is one of the most important tasks for concrete bridge monitoring. Traditionally, human-based inspection systems using people's eyes are unsafe and costly. To overcome this problem, a novel unmanned aerial vehicle (UAV) based inspection system is developed to detect cracks on the lateral sides and underside of bridges. Applying obstacle avoidance modules, UAV can reach bridges' lateral sides and underside within a safe distance. Thus the onboard camera can capture images of bridge surface in a high resolution. To locate the positions of cracks and solve the problem of limited horizon of single image, a fast feature-based stitching algorithm is developed. Then an edge detection method using structured forests is utilized to detect cracks on the large panorama. With image distortion corrected, the positions of cracks can be located on the large panorama. Experimental results validate the applicability and efficiency of the proposed UAV-based crack inspection system.
Huai Yu, Wen Yang 0001, Heng Zhang 0011, Wanjun He
IGARSS1
2017 Accurate extraction of cracks on the underside of concrete bridges
abstract
Crack detection is crucial for maintaining the structural health and safety of concrete bridges. Previous studies mainly focus on the detection of thick and conspicuous cracks on bridge deck using high resolution images. However, it's difficult to obtain accurate crack detection results on the underside of bridges due to the fainter and thinner appearance and lower contrast of cracks to their background. To overcome this problem, we propose a modified Beamlet Tree-based crack detection method which searches for possible edges by hierarchically constructing difference filters. The proposed method is evaluated with real crack images on the underside of bridges collected by an unmanned aerial vehicle (UAV). Compared with competitors, our proposed method achieves better performance in terms of accuracy and completeness.
Heng Zhang 0011, Huai Yu, Wen Yang 0001
IGARSS2