EDBT 2026 Demo / reviewers in the wild / expert
Qingwu Hu
dblp:10/2415
· DBLP profile ↗
21ranked-venue papers
0as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A non-overlapping image stitching method for reconstruction of page in ancient Chinese books
Yizhou Lan, Daoyuan Zheng, Qingwu Hu, Shaohua Wang 0003, Shunli Wang 0003, Tong Yue, Jiayuan Li 0001 |
Comput. Vis. Image Underst. | 3 |
| 2025 | FTransDeepLab: Multimodal Fusion Transformer-Based DeepLabv3+ for Remote Sensing Semantic SegmentationabstractHigh-resolution remote sensing images contain rich color and texture information, but due to the inherent limitations of 2-D data, achieving high-quality semantic segmentation remains a challenge. Multimodal data fusion technology has emerged as an effective approach to overcome this issue. To accurately capture the semantic information in remote sensing images, this study designs a multimodal fusion Transformer-based DeepLabv3+ model for remote sensing semantic segmentation, named FTransDeepLab. Specifically, the network learns features from two modalities and is inspired by the DeepLab architecture. We extended the encoder by stacking the multiscale Segformer, encoding the input images into highly representative spatial features. Additionally, we introduced the multimodal feature rectification (MFR) module and the multimodal feature fusion (MFF) module. The MFR, composed of a channel attention module and a spatial attention module, enhances the model’s ability to capture essential features and improves performance by focusing on both global and local contexts. The MFF module utilizes a cross-attention mechanism to optimize the feature fusion process, which enhances representation learning by facilitating the interaction between diverse information and integrates features from different modalities. Finally, in the decoding path, the extracted high-level features are concatenated with low-level features to optimize the feature representation and upsampled to restore the size of input image. Extensive results on two datasets, the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam, have confirmed that the proposed FTransDeepLab can achieve superior performance compared to the state-of-the-art segmentation methods. Haixia Feng, Qingwu Hu, Shunli Wang 0003, Mingyao Ai, Daoyuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | CMINet: A Unified Cross-Modal Integration Framework for Crop Classification From Satellite Image Time SeriesabstractAccurate automatic interpretation of crops is crucial for agricultural monitoring and food security assessment. In recent years, there has been a rise in platforms generating multimodal and multitemporal remote sensing images at an unprecedented speed. These images provide rich temporal, spatial, and spectral information, enabling more comprehensive land cover classification. Therefore, single-modal crop classification can benefit from complementary modalities. However, given the distinct characteristics of different modality sensors, enhancing the performance of deep networks by integrating diverse modality data remains a significant challenge. Unlike previous methods, this article proposes a two-stage fusion network named CMINet for crop classification from satellite image time series (SITS). Specifically, we adopt a decoupled framework to encode the spatial and temporal features of SITS. At each level of the spatial encoder, we design a cross-modal transfer module (CMTM) to complement bimodality branch features by transferring knowledge from one modality to rectify the features of another modality. After rectified feature pairs are encoded by a temporal encoder, we develop a cross-attention fusion module (CAFM) to conduct adequate context exchange before merging. The seamless combination of these novel designs forms a robust multimodal representation, outperforming the state-of-the-art methods on two public multimodal crop classification datasets. Compared to existing methods, our CMINet improves at least 1.1% OA, 1.7% mF1, and 2.0% mIoU on the PASITS-R dataset and 2.6% OA, 1.1% mF1, and 2.8% mIoU on the South Sudan dataset. Yutong Hu 0010, Qingwu Hu, Jiayuan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | CSNet: Change Selection of Activations and Pseudomasks for Image-Level Weakly Supervised Change DetectionabstractWeakly supervised change detection (WSCD) of bi-temporal remote sensing (RS) images has gained attention for its ability to reduce reliance on labor-intensive pixel-level change masks. Recent methods leverage image-level weak supervision to generate change pseudomasks for detecting changed objects, typically using class activation map (CAM) technique combined with DenseCRF or the Segment Anything Model (SAM). However, these methods still face two main challenges: first, CAMs tend to produce weak or false activations for changed objects, and second, DenseCRF and SAM lead to unreliable pseudomask generation, particularly when complex variations occur within objects in bi-temporal images. To address these challenges, a change selection network (CSNet) is proposed to enhance the quality of change activation maps and pseudomasks, improving their ability to accurately extract changed regions in bi-temporal RS images. First, a change activation selection (CAS) module is designed to generate a weight mask that selects and aggregates change-representing features, effectively highlighting missed change activations and strengthening weak activations. Second, a bi-temporal image selection (BIS) strategy is developed, incorporating two selection rules to filter out image pairs with poor-quality mask derived from SAM, while retaining those with high-quality results. Finally, a change pseudomask generation (CPG) module integrated with an atrous-spatial pyramid pooling (ASPP) classifier is developed to predict accurate change pixels for final pseudomask generation. Experimental results demonstrate that the proposed CSNet outperforms existing WSCD methods, achieving 79.32% IoU in change pseudomasks for the WHU-CD dataset, 68.12% for the GZ-CD dataset and 75.44% for the GVLM dataset. This study proposes a novel method that enhances the performance of the weakly supervised paradigm in RS CD. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Shield Tunnel Dislocation Detection Method Based on Semantic Segmentation and Bolt Hole Positioning of MLS Point CloudabstractThe dislocation of ring segments in shield tunnels poses adverse effects on tunnel structural stability and waterproofing. Current methods can only detect limited types of dislocation from circular tunnels, posing challenges in meeting practical requirements. This paper introduces a dislocation detection method that applicable to shield tunnels of various shapes, enabling the detection of both inter-ring and intra-ring dislocations. Firstly, a point cloud lossless unfolding method is introduced, allowing for the unfolding of tunnel point clouds of any cross-section shape without compromising point cloud features. Subsequently, a subway tunnel point cloud segmentation workflow is proposed, enabling the direct application of image deep learning networks to tunnel point clouds. The bolts are then identified as markers for dislocation measurement, and a method for bolt area of interest positioning and pairing is introduced. This method calculates dislocation value through fitting reference surfaces from area of interest. A series of experiments conducted on circular and quasi-rectangular shield tunnels totaling 2.45 km and containing 20 billion points demonstrate that the proposed segmentation workflow, coupled with an image semantic segmentation model, achieves a maximum mIoU of 90.2%, which is 10% higher than directly segmenting tunnel point clouds. The proposed method went through repeated accuracy tests and compared with other methods, including total station measurement, RANSAC plane fitting and scanline method. The results of our proposed method demonstrate high precision and stability, with a measurement accuracy RMSE of 0.89 mm, MAE of 0.69 mm, and efficiency 20 times higher compared to total station measurement. Qingzhou Mao, Jian Li 0044, Qingwu Hu, Yiwen Tao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | AMIANet: Asymmetric Multimodal Interactive Augmentation Network for Semantic Segmentation of Remote Sensing ImageryabstractIn recent years, the inherent 2-D characteristics of optical images have led to a plateau in semantic segmentation performance. The complementary nature of light detection and ranging (LiDAR) point clouds and camera images can effectively enhance semantic segmentation capabilities, and thus, research into multimodal joint semantic segmentation is garnering increasing attention. However, the domain gaps between different dimensions present challenges for the fusion of multimodal data. In this article, we introduce a novel asymmetric multimodal interaction augmented network (AMIANet), which directly processes heterogeneous data from images and point clouds. The treatment of the disparities in modal data ensures consistency in the features of both modes. Through the newly developed synergistic multimodal interaction module (SMI Module), AMIANet is capable of combining the complementary characteristics of cross-modal data. This is achieved by interactively fusing and extracting precise and rich structural information from point cloud features to enhance image characteristics. The experimental results on the N3C-California, WHU-RRDSD, and ISPRS Vaihingen datasets demonstrate that AMIANet surpasses benchmark methods and current state-of-the-art (SOTA) approaches. The code will be available athttps://github.com/2012153946/AMIANet. Qingwu Hu, Wenlei Fan, Haixia Feng, Daoyuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Weakly Supervised Building Extraction From High-Resolution Remote Sensing Images Based on Building-Aware Clustering and Activation Refinement NetworkabstractWeakly supervised building extraction methods, utilizing image-level labels, offer a cost-effective solution by significantly reducing the need for pixel-level annotation in high-resolution (HR) remote sensing (RS) images. These methods often focus on class activation map (CAM) optimization based on features extracted from individual images, missing out on the benefits of associating building features from multiple RS images (i.e., n images) to improve CAMs. This limitation leaves room for improvement in both CAM optimization and pseudo-mask generation. To address this, we propose the building-aware clustering and activation refinement network (BAC-AR-Net), a novel weakly supervised network to enhance weakly supervised building extraction performance. The building-aware clustering (BAC) module aggregates and clusters feature maps from multiple building samples to obtain common features of buildings. The common features are subsequently used to extract regions with similar building semantics, thereby enhancing the accuracy and completeness of building coverage in CAMs. Additionally, the activation refinement module is designed to generate pseudo-masks with clear boundaries and an effective separation of buildings and background. Experiments were conducted on the ISPRS Potsdam and Vaihingen datasets as well as a self-built building dataset to verify the effectiveness of our proposed method. The results show the proposed method outperforms both the weakly supervised semantic segmentation and weakly supervised building extraction methods that use image-level labels, achieving IoU accuracies of 0.8556, 0.8163, and 0.7797 on the respective datasets. This study introduces a novel weakly supervised learning framework to the RS application, with a particular focus on building extraction and semantic segmentation tasks. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Jiayuan Li 0001, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Augmented Maximum Correntropy Criterion for Robust Geometric PerceptionabstractMaximum correntropy criterion (MCC) is a robust and powerful technique to handle heavy-tailed nonGaussian noise, which has many applications in the fields of vision, signal processing, machine learning, etc. In this article, we introduce several contributions to the MCC and propose an augmented MCC (AMCC), which raises the robustness of classic MCC variants for robust fitting to an unprecedented level. Our first contribution is to present an accurate bandwidth estimation algorithm based on the probability density function (PDF) matching, which solves the instability problem of the Silverman's rule. Our second contribution is to introduce the idea of graduated nonconvexity (GNC) and a worst-rejection strategy into MCC, which compensates for the sensitivity of MCC to high outlier ratios. Our third contribution is to provide a definition of local distribution measure to evaluate the quality of inliers, which makes the MCC no longer limited to random outliers but is generally suitable for both random and clustered outliers. Our fourth contribution is to show the generalizability of the proposed AMCC by providing eight application examples in geometry perception and performing comprehensive evaluations on five of them. Our experiments demonstrate that 1) AMCC is empirically robust to 80%$-$90% of random outliers across applications, which is much better than Cauchy M-estimation, MCC, and GNC-GM; 2) AMCC achieves excellent performance in clustered outliers, whose success rate is 60%$-$70% percentage points higher than the second-ranked method at 80% of outliers; 3) AMCC can run in real-time, which is 10$-$100 times faster than RANSAC-type methods in low-dimensional estimation problems with high outlier ratios. This gap will increase exponentially with the model dimension. Jiayuan Li 0001, Qingwu Hu, Xinyi Liu 0002, Yongjun Zhang 0002 |
IEEE Trans. Robotics | 2 |
| 2023 | QGORE: Quadratic-Time Guaranteed Outlier Removal for Point Cloud RegistrationabstractWith the development of 3D matching technology, correspondence-based point cloud registration gains more attention. Unfortunately, 3D keypoint techniques inevitably produce a large number of outliers, i.e., outlier rate is often larger than 95%. Guaranteed outlier removal (GORE) Bustos and Chin has shown very good robustness to extreme outliers. However, the high computational cost (exponential in the worst case) largely limits its usages in practice. In this paper, we propose the first$O(N^{2})$time GORE method, called quadratic-time GORE (QGORE), which preserves the globally optimal solution while largely increases the efficiency. QGORE leverages a simple but effective voting idea via geometric consistency for upper bound estimation, which achieves almost the same tightness as the one in GORE. We also present a one-point RANSAC by exploring “rotation correspondence” for lower bound estimation, which largely reduces the number of iterations of traditional 3-point RANSAC. Further, we propose al$_{p}$p-like adaptive estimator for optimization. Extensive experiments show that QGORE achieves the same robustness and optimality as GORE while being 1$\sim$2 orders faster. The source code will be made publicly available. Jiayuan Li 0001, Qingwu Hu, Yongjun Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | LNIFT: Locally Normalized Image for Rotation Invariant Multimodal Feature MatchingabstractSevere nonlinear radiation distortion (NRD) is the bottleneck problem of multimodal image matching. Although many efforts have been made in the past few years, such as the radiation-variation insensitive feature transform (RIFT) and the histogram of orientated phase congruency (HOPC), almost all these methods are based on frequency-domain information that suffers from high computational overhead and memory footprint. In this article, we propose a simple but very effective multimodal feature matching algorithm in the spatial domain, called locally normalized image feature transform (LNIFT). We first propose a local normalization filter to convert original images into normalized images for feature detection and description, which largely reduces the NRD between multimodal images. We demonstrate that normalized matching pairs have a much larger correlation coefficient than the original ones. We then detect oriented FAST and rotated brief (ORB) keypoints on the normalized images and use an adaptive nonmaximal suppression (ANMS) strategy to improve the distribution of keypoints. We also describe keypoints on the normalized images based on a histogram of oriented gradient (HOG), such as a descriptor. Our LNIFT achieves rotation invariance the same as ORB without any additional computational overhead. Thus, LNIFT can be performed in near real-time on images with 1024$\times 1024$pixels (only costs 0.32 s with 2500 keypoints). Four multimodal image datasets with a total of 4000 matching pairs are used for comprehensive evaluations, including synthetic aperture radar (SAR)–optical, infrared–optical, and depth–optical datasets. Experimental results show that LNIFT is far superior to RIFT in terms of efficiency (0.49 s versus 47.8 s on a$1024 \times 1024$image), success rate (99.9% versus 79.85%), and number of correct matches (309 versus 119). The source code and datasets will be publicly available athttps://ljy-rs.github.io/web. Jiayuan Li 0001, Wangyi Xu, Yongjun Zhang 0002, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Point Cloud Registration Based on One-Point RANSAC and Scale-Annealing Biweight EstimationabstractPoint cloud registration (PCR) is an important task in photogrammetry and remote sensing, whose goal is to seek a seven-parameter similarity transformation to register a pair of point clouds. Traditional iterative closest point (ICP) variants highly rely on the initial parameters, and most of them cannot deal with cross-source (multisource) point clouds with scale changes. In this article, we propose a flexible correspondence-based PCR method, which is initial-guess free, fast, and robust. We first decompose the full seven-parameter registration problem into three subproblems, i.e., scale, rotation, and translation estimations, based on line vectors. Then, we propose a one-point random sample consensus (RANSAC) algorithm to estimate the scale and translation parameters. For the rotation estimation, we introduce a graduated optimization strategy into Tukey’s biweight function and propose a scale-annealing biweight estimator. We evaluate the proposed method on both same-source and cross-source data. Results show that the proposed method is robust against over 99% outliers and is one to two orders of magnitude faster than its competitors. The source code of our method will be made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Learning From GPS Trajectories of Floating Car for CNN-Based Urban Road Extraction With High-Resolution Satellite ImageryabstractDeep learning has achieved great success in recent years, among which the convolutional neural network (CNN) method is outstanding in image segmentation and image recognition. It is also widely used in satellite imagery road extraction and, generally, can obtain accurate and extraction results. However, at present, the extraction of roads based on CNN still requires a lot of manual preparation work, and a large number of samples can be marked to achieve extraction, which has to take long drawing cycle and high production cost. In this article, a new CNN sample set production method is proposed, which uses the GPS trajectories of floating car as training set (GPSTasST), for the multilevel urban roads extraction from high-resolution remote sensing imagery. This method rasterizes the GPS trajectories of floating car into a raster map and uses the processed raster map to label the satellite image to obtain a road extraction sample set. CNN can extract roads from remote sensing imagery by learning the training set. The results show that the method achieves a harmonic mean of precision and recall higher than road extraction method from single data source while eliminating the manual labeling work, which shows the effectiveness of this work. Qingwu Hu, Jiayuan Li 0001, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Robust Geometric Model Estimation Based on Scaled Welsch q-NormabstractRobust estimation, which aims to recover the geometric transformation from outlier contaminated observations, is essential for many remote sensing and photogrammetry applications. This article presents a novel robust geometric model estimation method based on scaled Welsch q-norm (lq-norm, 0qLS) problem and a weighted least-squares (WLS) problem] by using alternating direction method of multipliers (ADMM) method. For the WLS problem, we introduce a coarse-to-fine strategy into the iterative reweighted least-squares (IRLS) method. We change the weight function by decreasing its scale parameter. This strategy can largely avoid that the solver gets stuck in local minimums. We adapt the proposed cost into classical remote sensing tasks and develop new robust feature matching (RFM), robust exterior orientation (REO), and robust absolute orientation (RAO) algorithms. Both synthetic and real experiments demonstrate that the proposed method significantly outperforms the other compared state-of-the-art methods. Our method is still robust even if the outlier rate is up to 90%. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | RIFT: Multi-Modal Image Matching Based on Radiation-Variation Insensitive Feature TransformabstractTraditional feature matching methods, such as scale-invariant feature transform (SIFT), usually use image intensity or gradient information to detect and describe feature points; however, both intensity and gradient are sensitive to nonlinear radiation distortions (NRD). To solve this problem, this paper proposes a novel feature matching algorithm that is robust to large NRD. The proposed method is called radiation-variation insensitive feature transform (RIFT). There are three main contributions in RIFT. First, RIFT uses phase congruency (PC) instead of image intensity for feature point detection. RIFT considers both the number and repeatability of feature points and detects both corner points and edge points on the PC map. Second, RIFT originally proposes a maximum index map (MIM) for feature description. The MIM is constructed from the log-Gabor convolution sequence and is much more robust to NRD than traditional gradient map. Thus, RIFT not only largely improves the stability of feature detection but also overcomes the limitation of gradient information for feature description. Third, RIFT analyses the inherent influence of rotations on the values of the MIM and realises rotation invariance. We use six different types of multi-modal image datasets to evaluate RIFT, including optical-optical, infrared-optical, synthetic aperture radar (SAR)-optical, depth-optical, map-optical, and day-night datasets. Experimental results show that RIFT is superior to SIFT and SAR-SIFT on multi-modal images. To the best of our knowledge, RIFT is the first feature matching algorithm that can achieve good performance on all the abovementioned types of multi-modal images. The source code of RIFT and the multi-modal image datasets are publicly available1. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Image Process. | 2 |
| 2019 | Haze and Thin Cloud Removal via Sphere Model Improved Dark Channel PriorabstractHaze and cloud seriously degrade the quality of optical remote sensing images, which largely decrease their interpretability and intelligibility. In this letter, we propose a two-stage haze and thin cloud removal method based on homomorphic filtering (HF) and sphere model improved dark channel prior (DCP). Compared with current dehazing methods, the most advantage of the proposed method is that our method can deal with uneven haze, thick haze, and thin cloud. We observe that haze and cloud are highly related to the illumination component and mainly located in the low frequency of an image. Thus, we adapt HF to enhance the haze image, which makes the distribution of haze more even. In the second stage, we analyze the drawback of DCP, i.e., the transmission estimated by DCP is very sensitive to noise. To draw this issue, we propose a novel sphere model to estimate a more accurate transmission map. The sphere model improved DCP is more suitable for thick haze images than the traditional DCP. Extensive experimental results show that the proposed method significantly outperforms the compared state-of-the-art methods. The source code and data sets used in the letter are made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Optimal Illumination and Color Consistency for Optical Remote-Sensing Image MosaickingabstractIllumination and color consistency are very important for optical remote-sensing image mosaicking. In this letter, we propose a simple but effective technique that simultaneously performs image illumination and color correction for multiview images. In this framework, we first present an uneven illumination removal algorithm based on bright channel prior, which guarantees the illumination consistency inside a single image. We then adapt a pairwise color-correction method to coarsely align the color tone between source and reference images. In this stage, we give a new single-image quality metric which combines brightness deviation, color cast, and entropy together for automatic reference-image selection. Finally, we perform a least-squares adjustment (LSA) procedure to obtain optimal illumination and color consistency among multiview images. In detail, we first perform a pairwise image matching by using SIFT algorithm; once sparse local patch correspondences obtained, the illumination and color relationship between images can be established based on a global gamma correction model; the illumination and color errors can then be minimized by LSA. Extensive experiments on both challenging synthetic and real optical remote-sensing image data sets show that it significantly outperforms the compared state-of-the-art approaches. All the source code and data sets used in this letter are made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Indoor Map Service System Based on Wechat Two-Dimensional Code
Qingwu Hu |
APWeb (2) | 2 |
| 2016 | A feature preserving algorithm for point cloud simplification based on hierarchical clusteringabstractThe efficiency and accuracy loss are the key issues for the point cloud simplification. In this paper, a feature preserving algorithm is proposed for point cloud simplification based on hierarchical clustering with the surface feature description. The surface variation is presented as the main criterion for the efficient hierarchical clustering method to simplify the mass and dense point cloud fast, meanwhile we retain the feature points to ensure a small accuracy loss. The experiment results show that the proposed method is efficient and has a good effect to maintain the features as the same degree of simplification. Qingwu Hu |
IGARSS | 3 |
| 2016 | Robust Feature Matching for Remote Sensing Image Registration Based on Lq-EstimatorabstractThis letter proposes a robust feature matching algorithm for remote sensing images based on lq-estimator. We start with a set of initial matches provided by a feature matching method such as scale-invariant feature transform and then focus on global transformation estimation from contaminated observations and outliers elimination as well. We use an affine model to describe the global transformation and minimize a new cost function based on lq-norm. We apply an augmented Lagrangian function and an alternating direction method of multipliers to solve such a nonconvex and nonsmooth optimization problem. Extensive experiments on real remote sensing data demonstrate that the proposed method is effective, efficient, and robust. Our method outperforms state-of-the-art methods and can easily handle situations with up to 90% outliers. In addition, the proposed method is much faster than RANSAC. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | A Study of Users' Movements Based on Check-In Data in Location-Based Social Networks
Jinzhou Cao, Qingwu Hu, Qingquan Li 0001 |
W2GIS | 2 |
| 2006 | An Ontology Definition Framework for Model Driven Development
Yucong Duan, Xiaolan Fu, Qingwu Hu, Yuqing Gu |
ICCSA (4) | 3 |