Han Hu 0005

dblp:59/3754-5 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-1137-2208ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Higher Order Energy-Optimized Air-Ground Individual Tree Segmentation With Crown Morphology Fidelity
abstract
Accurate acquisition of carbon sequestration information at the individual tree scale is crucial for the refined assessment and management of carbon stocks in urban ecosystems. Traditional optical remote sensing methods, limited by two-dimensional spectral features, are more suitable for large-scale forest carbon stock estimation but struggle to achieve refined quantification at the individual tree level. LiDAR technology, which directly characterizes three-dimensional structural parameters of trees through point clouds, significantly outperforms optical methods in improving both segmentation quantity and morphological accuracy. To achieve a refined assessment and management of the urban carbon storage, on the one hand, we propose to integrate UAV-borne and ground-based point clouds, thereby overcoming the observational limitations of single-source data. On the other hand, to address the problems of under- and over-segmentation and morphological distortion in individual tree segmentation. We have respectively proposed two extreme value segmentation models and a high-order energy morphological optimization model. Specifically, the Canopy Skyline Extremum (CSE) model can solve the under-segmentation problem caused by the absence of canopy extreme points and the over-segmentation problem caused by pseudo-extreme points from the perspective of the canopy dimension. The Vertical Distribution Extremum (VDE) model of tree can solve the under-segmentation problem caused by the sparsity of trunk point clouds from the trunk dimension. For crown morphology fidelity, we construct, a voxel-based spatial correlation graph is constructed to characterize the distribution of branches and leaves. We use the Markov Random Field (MRF) energy optimization framework, integrate the prior knowledge of tree ecology, establish a high-order energy model that satisfies the objective cognitive morphology of trees, dynamically assign the membership relationship of branches and leaves, effectively solve the problem of interlacing and adhesion of branches and leaves between adjacent trees, overcome the problem of tree morphological distortion caused by the "one-size-fits-all" approach in traditional individual tree segmentation, improve the calculation accuracy of the tree crown diameter and canopy volume, and enhance the estimation accuracy of the carbon storage of individual trees to the decimeter level. Experiments were conducted on six datasets from two typical urban scenarios: street trees and landscaped gardens. Results validate the superior performance of our morphology-faithful segmentation. Quantitative comparisons with the state-of-the-arts show that our approach achieves optimal performance in both segmentation accuracy and stability. Further analysis confirms that the proposed morphological optimization preserves reasonable tree shapes, leading to more accurate individual tree attribute calculations. This study verifies that the proposed models significantly enhance the accuracy of urban individual tree segmentation (quantity and morphology), enabling decimeter-level urban carbon stock assessment, and provides a novel technical framework for precise ecosystem carbon sequestration monitoring.
Xuming Ge, Min Chen 0015, Han Hu 0005, Bo Xu 0003, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.3
2025 Asymmetric Mamba-CNN Collaborative Architecture for Large-Size Remote Sensing Image Semantic Segmentation
abstract
Large-size remote sensing images contain rich geographical information. Efficient and accurate semantic segmentation of these images is of significant importance in various fields. However, the massive memory requirements have hindered the development of semantic segmentation methods for large-size remote sensing images. Most existing methods struggle to balance memory usage, global modeling, and local representation accuracy. To address these issues, we propose a new semantic segmentation method for large-size remote sensing images, Mamba–CNN parallel network (MCPNet), which demonstrates impressive performance. The method is an asymmetric Mamba–convolutional neural network (CNN) hybrid architecture. Given the linear modeling complexity of Mamba, we construct the M-branch based on the visual state space (VSS) model, which processes downsampled images to reduce memory consumption while alleviating Mamba’s local forgetting problem. To further enhance the model’s capability in fine-grained detail extraction, we meticulously design a detail-preserving network (DPN) as the C-branch. This branch employs a split downsampling strategy and multiscale convolutional kernel groups to process large-size images, ensuring the preservation of spatial positional relationships while capturing fine-grained local details. Moreover, to effectively filter redundant information introduced by large-size images and bridge the semantic gap between the features extracted by CNN and Mamba, we propose a multigated feature fusion module (MG-FFM). This module progressively refines heterogeneous feature alignment through a bottom-up hierarchical refinement strategy, achieving a progressive fusion of semantics and details. Our method achieves state-of-the-art (SOTA) performance in terms of mean intersection over union (mIoU) and mF1 score on the self-constructed Yaan UAV dataset and two widely used public datasets (DeepGlobe and Inria Aerial) while consuming less GPU memory. The codes will be available athttps://github.com/fsqy-zhang/MCPNet
Min Chen 0015, Lianlei Shan, Caiyi Li, Han Hu 0005, Xuming Ge, Qing Zhu 0012, Bo Xu 0003
IEEE Trans. Geosci. Remote. Sens.6
2024 Landslide Extraction Using Fused Local and Nonlocal Attentional Features on Edge Device Toward Embedded UAV Emergency Response
abstract
Unmanned aerial vehicles (UAVs) have made significant contributions to landslide emergency response operations due to their precise and flexible imaging capabilities. However, the conventional workflow of generating orthophotos from UAV imagery and subsequent interpretation often exceeds the critical 72-hour rescue window. To address this challenge, this paper presents a onboard landslide extraction method for original UAV images utilizing a convolutional neural network (CNN). Given the abundance of overlapping images, the CNN is trained on labeled orthophotos. To minimize discrepancies between orthophotos and original UAV images, the proposed method integrates local and non-local features. Built upon the ResNet architecture, the method incorporates modules for extracting both shallow and deep features, enabling effective fusion through self-learning. This approach mitigates the issue of accuracy degradation caused by variations between training and testing data. Furthermore, considering the necessity of deploying the CNN-based landslide extraction model on low-power embedded platforms to achieve onboard landslide extraction, this paper introduces a quantitative model compression technique. Specifically, the model’s weight and activation value data precision are linearly mapped from 32-bit floating-point type to 8-bit integer type, guided by relative entropy minimization. This results in substantial reductions in memory access and computational complexity during model inference. Experimental results demonstrate that the proposed method yields outstanding extraction performance on both the Jiuzhaigou and Bijie landslide datasets. The time taken for extracting landslides from a single 6000x4000 pixel UAV image is reduced from 109.47 seconds to 4.75 seconds, which is less than the 5.13-second interval between camera shots, thereby achieving onboard landslide extraction.
Yulin Ding, Han Hu 0005, Qing Zhu 0012, Bo Xiang, Yunyong He
IEEE Trans. Geosci. Remote. Sens.3
2024 Semantic Image Translation for Repairing the Texture Defects of Building Models
abstract
The accurate representation of 3-D building models in urban environments is significantly hindered by challenges such as texture occlusion, blurring, and missing details, which are difficult to mitigate through standard photogrammetric texture mapping pipelines. Current image completion methods often struggle to produce structured results and effectively handle the intricate nature of highly structured façade textures with diverse architectural styles. Furthermore, existing image synthesis methods encounter difficulties in preserving high-frequency details and artificial regular structures, which are essential for achieving realistic façade texture synthesis. To address these challenges, we introduce a novel approach for synthesizing façade texture images that authentically reflect the architectural style from a structured label map, guided by a ground-truth façade image. In order to preserve fine details and regular structures, we propose a regularity-aware multidomain method that capitalizes on frequency information and corner maps. We also incorporate semantic region-adaptive normalization (SEAN) blocks into our generator to enable versatile style transfer. To generate plausible structured images without undesirable regions, we employ image completion techniques to remove occlusions according to semantics prior to image inference. Our proposed method is also capable of synthesizing texture images with specific styles for façades that lack preexisting textures, using manually annotated labels. Experimental results on publicly available façade image and 3-D model datasets demonstrate that our method yields superior results and effectively addresses issues associated with flawed textures.
Qisen Shang, Han Hu 0005, Haojia Yu, Bo Xu 0003, Libin Wang 0005, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.2
2024 StructuredMesh: 3-D Structured Optimization of Façade Components on Photogrammetric Mesh Models Using Binary Integer Programming
abstract
The absence of detailed façade elements in photogrammetric building mesh models falls short for complex applications. These models often present irregular surfaces, significant geometric noise, and subpar texture quality, complicating façade parsing. To mitigate these challenges, we introduce StructuredMesh, an innovative approach to augment façade details while adhering to architectural regularity. Our method captures multiview color and depth images with a virtual camera and implements an object detection pipeline to semiautomatically delineate façade elements—such as windows, doors, and balconies—from color images. These elements are mapped into 3-D space using depth information, forming a preliminary façade layout. By harnessing architectural intelligence, we engage binary integer programming (BIP) for 3-D layout optimization, targeting component positions, orientations, and sizes. The resulting refined layout underpins the façade reconstruction of the mesh model through instance replacement. Our method’s efficacy was confirmed on three distinct datasets. Employing the 3-D layout evaluation metrics we proposed, StructuredMesh surpasses conventional methods in numerical precision and regularity enhancement.
Libin Wang 0005, Han Hu 0005, Qisen Shang, Haowei Zeng, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.2
2023 3-D Line Segment Reconstruction With Depth Maps for Photogrammetric Mesh Refinement in Man-Made Environments
abstract
Three-dimensional (3D) line segments contain richer geometric and structural information than 3D point clouds in man-made environments, which is beneficial for providing constraints to refine point-cloud-based mesh models or build accurate wireframes. However, the efficient reconstruction of 3D line segments with high scene coverage from multi-view images is still challenging. In this study, the depth maps obtained from the point cloud generation procedure are exploited to decrease the search range of two-dimensional (2D) line segment correspondences to improve the efficiency, precision, and recall rate of 2D line segment matching, thereby improving the construction efficiency and scene coverage of 3D line segments. For a line segment on the reference image (called reference line segment) of an image pair, a reliable virtual line segment is produced by projecting several sampled points of the reference line segment onto the search image based on the corresponding depth information. Then, a purely geometrical similarity measurement under the constraints of the virtual line segment is designed to obtain 2D line segment matches. Using the 2D line segment correspondences of all image pairs, a multi-view clustering operation is performed to construct 3D line segments from the redundant 2D matches. Finally, a simple 3D-line-segment-based mesh model refinement method is designed and the reconstructed 3D line segments are employed to improve the quality of the point-cloud-based mesh model. In our experiments, five open-source datasets are adopted to qualitatively and quantitatively evaluate the performance of the proposed 3D line segment reconstruction method and the potential of 3D line segments on mesh model refinement. The experimental results show that the proposed 3D line segment reconstruction method performs better than the state-of-the-art methods. Specifically, on five open-source datasets, our method exhibits an average improvement of 47.28% in the number of reconstructed 3D line segments over the best one among the compared methods. Additionally, the experimental results of the mesh model refinement show that the addition of 3D line segments is beneficial for improving the quality of the point-cloud-based mesh model.
Tong Fang, Min Chen 0015, Han Hu 0005, Wen Li 0033, Xuming Ge, Qing Zhu 0012, Bo Xu 0003
IEEE Trans. Geosci. Remote. Sens.3
2022 Graph neural networks with constraints of environmental consistency for landslide susceptibility evaluation
abstract
In complex and heterogeneous geoenvironments, landslides exhibit varying features in different environments, and data in landslide inventories are imbalanced. Existing data-driven landslide susceptibility evaluation (LSE) methods overlook environmental heterogeneity and cannot reliably predict regions with few samples. Alternatively, global random negative sampling strategies may produce imbalanced positive and negative samples in some environments, contributing to inaccurate predictions. This article proposes a graph neural network (GNN) constrained by environmental consistency (GNN-EC) to overcome these problems. The GNN-EC consists of graphs with nodes, and edges. A graph represents the environmental relationships in the study area. Nodes are geographic units delineated from terrain polygon approximation. Edges capture the relationships between node-pairs. Additionally, the weights of edges reflect the similarity between two node environments. A GNN aggregates node information in the graph for LSE. Our experiment showed that the proposed method outperformed the common machine learning methods: increasing prediction accuracy by approximately 7, 5–6 and 3–4% compared to the artificial neural network (ANN), the support vector machine (SVM) and the random forest (RF), respectively. Moreover, our method can maintain high prediction accuracy, even with a small training set.
Haowei Zeng, Qing Zhu 0012, Yulin Ding, Han Hu 0005, Li Chen 0026, Xiao Xie, Min Chen 0015, Yanxia Yao
Int. J. Geogr. Inf. Sci.4
2021 Multientity Registration of Point Clouds for Dynamic Objects on Complex Floating Platform Using Object Silhouettes
abstract
This article is focused on a challenging topic emerging from the registration of point clouds, specifically the registration of dynamic objects with low overlapping ratio. This problem is especially difficult when the static scanner is installed on a floating platform, and the objects it scans are also floating. These issues make most of the automatic registration methods and software solutions invalid. To solve this problem, explicit exploration of the static region is necessary for both the coarse and fine registration steps. Fortunately, determining the corresponding regions can be eased by the intuitive realization that in urban environments, natural objects neither present straight boundaries nor stack vertically. This intuition has guided the authors to develop a robust approach for the detection of static regions using planar structures. Then, silhouettes of the objects are extracted from the planar structures, which assist in the determination of an SE(2) transformation in the horizontal direction by a novel line matching method. The silhouettes also enable identification of the correspondences of planes in the step of fine registration using a variant of the iterative closest point method. Experimental evaluations using point clouds of cargo ships with different sizes and shapes reveal the robustness and efficiency of the proposed method, which gives 100% success and reasonable accuracy in rapid time, suitable for an online system. In addition, the proposed method is evaluated systematically with regard to several practical situations caused by the floating platform, and it demonstrates good robustness to limited scanning time and noise.
Feng Wang 0044, Han Hu 0005, Xuming Ge, Bo Xu 0003, Ruofei Zhong, Yulin Ding, Xiao Xie, Qing Zhu 0012
IEEE Trans. Geosci. Remote. Sens.2
2021 MAP-Net: Multiple Attending Path Neural Network for Building Footprint Extraction From Remote Sensed Imagery
abstract
Building footprint extraction is a basic task in the fields of mapping, image understanding, computer vision, and so on. Accurately and efficiently extracting building footprints from a wide range of remote sensed imagery remains a challenge due to the complex structures, variety of scales, and diverse appearances of buildings. Existing convolutional neural network (CNN)-based building extraction methods are criticized for their inability to detect tiny buildings because the spatial information of CNN feature maps is lost during repeated pooling operations of the CNN. In addition, large buildings still have inaccurate segmentation edges. Moreover, features extracted by a CNN are always partially restricted by the size of the receptive field, and large-scale buildings with low texture are always discontinuous and holey when extracted. To alleviate these problems, multiscale strategies are introduced in the latest research works to extract buildings with different scales. The features with higher resolution generally extracted from shallow layers, which extracted insufficient semantic information for tiny buildings. This article proposes a novel multiple attending path neural network (MAP-Net) for accurately extracting multiscale building footprints and precise boundaries. Unlike existing multiscale feature extraction strategies, MAP-Net learns spatial localization-preserved multiscale features through a multiparallel path in which each stage is gradually generated to extract high-level semantic features with fixed resolution. Then, an attention module adaptively squeezes the channel-wise features extracted from each path for optimized multiscale fusion, and a pyramid spatial pooling module captures global dependence for refining discontinuous building footprints. Experimental results show that our method achieved 0.88%, 0.93%, and 0.45% F1-score and 1.53%, 1.50%, and 0.82% intersection over union (IoU) score improvements without increasing computational complexity compared with the latest HRNetv2 on the Urban 3-D, Deep Globe, and WHU data sets, respectively. Specifically, MAP-Net outperforms multiscale aggregation fully convolutional network (MA-FCN), which is the state-of-the-art (SOTA) algorithms with postprocessing and model voting strategies, on the WHU data set without pretraining and postprocessing. The TensorFlow implementation is available at https://github.com/lehaifeng/MAPNet.
Qing Zhu 0012, Han Hu 0005, Xiaoming Mei, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.3
2020 Interactive Correction of a Distorted Street-View Panorama for Efficient 3-D Façade Modeling
abstract
Façade features are important in large-scale level-of-detail 3 (LoD-3) reconstruction in urban environments, and street-view panoramas are arguably the best option for detailed 3-D façade modeling. However, despite the plethora of street-view panoramas available, few studies have explored the metric capabilities of panoramas. This is due in part to the complexities of system integration and in part to problems associated with projection (e.g., distortion at the tops of buildings), and deformation (e.g., the bending of straight structures). In an effort to solve these problems, this letter introduces a flexible and practical solution using only a single panorama. The key is to efficiently rectify panoramas using image-space line constrained deformation inspired by the as-rigid-as-possible deformation of surface meshes. The image is then re-projected using gnomonic projection on a properly selected tangent plane. The proposed approach requires a reasonable amount of user interaction to select and position the vertical line segments. The tangent point is also chosen empirically for each panorama. The rectified images can then be imported into off-the-shelf 3-D modeling solutions as reference images for interactive sketching. Experimental evaluations reveal the effectiveness of the image-space rectification: after proper scaling, the semantic-aware 3-D façade models achieve decimeter-level accuracy with respect to the reference surface mesh.
Qing Zhu 0012, Mier Zhang, Han Hu 0005, Feng Wang 0044
IEEE Geosci. Remote. Sens. Lett.3
2019 Image-Guided Registration of Unordered Terrestrial Laser Scanning Point Clouds for Urban Scenes
abstract
This paper presents an image-guided end-to-end registration approach for globally consistent 3-D registration of unordered terrestrial laser scanning (TLS) point clouds. The proposed method can handle arbitrary point clouds with reasonable pairwise overlap without knowledge about their initial position and orientation, without requiring artificial targets, and without needing to record the order of the scanning. One of the novel contributions of the proposed approach lies in the optimization of a scanning network. We retrieve the similarities of all scans based on a vocabulary tree using both the geometrically rectified panorama images and the corresponding 3-D point clouds. The approach also highlights the integral optimization in both the coarse and fine registration. A pose graph is introduced to realize global optimization at the end of the coarse step without primitives. After that, the results act as the inputs to start the pairwise fine registration, which is then followed by the minimum loop expansion (MLE) refinement. Comprehensive experiments demonstrated network optimization rates of over 60% using the image-guided strategy. Using the pose-graph optimization method, successful registration rates (SRRs) increased to 100% for all tested cases. The MLE not only accelerates the speed of the convergence but also improves registration accuracy, which reached 0.1 m and 0.1° in the translation and rotation angles, respectively.
Xuming Ge, Han Hu 0005, Bo Wu 0004
IEEE Trans. Geosci. Remote. Sens.2