VLDB 2026 Research / reviewers in the wild / expert
Mingyao Ai
dblp:22/8502
· DBLP profile ↗
13ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Theory of computation · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Space-filling designs on Riemannian manifolds
Mingyao Ai, Yunfan Yang, Xiangshun Kong |
J. Complex. | 1 |
| 2025 | FTransDeepLab: Multimodal Fusion Transformer-Based DeepLabv3+ for Remote Sensing Semantic SegmentationabstractHigh-resolution remote sensing images contain rich color and texture information, but due to the inherent limitations of 2-D data, achieving high-quality semantic segmentation remains a challenge. Multimodal data fusion technology has emerged as an effective approach to overcome this issue. To accurately capture the semantic information in remote sensing images, this study designs a multimodal fusion Transformer-based DeepLabv3+ model for remote sensing semantic segmentation, named FTransDeepLab. Specifically, the network learns features from two modalities and is inspired by the DeepLab architecture. We extended the encoder by stacking the multiscale Segformer, encoding the input images into highly representative spatial features. Additionally, we introduced the multimodal feature rectification (MFR) module and the multimodal feature fusion (MFF) module. The MFR, composed of a channel attention module and a spatial attention module, enhances the model’s ability to capture essential features and improves performance by focusing on both global and local contexts. The MFF module utilizes a cross-attention mechanism to optimize the feature fusion process, which enhances representation learning by facilitating the interaction between diverse information and integrates features from different modalities. Finally, in the decoding path, the extracted high-level features are concatenated with low-level features to optimize the feature representation and upsampled to restore the size of input image. Extensive results on two datasets, the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam, have confirmed that the proposed FTransDeepLab can achieve superior performance compared to the state-of-the-art segmentation methods. Haixia Feng, Qingwu Hu, Shunli Wang 0003, Mingyao Ai, Daoyuan Zheng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | CSNet: Change Selection of Activations and Pseudomasks for Image-Level Weakly Supervised Change DetectionabstractWeakly supervised change detection (WSCD) of bi-temporal remote sensing (RS) images has gained attention for its ability to reduce reliance on labor-intensive pixel-level change masks. Recent methods leverage image-level weak supervision to generate change pseudomasks for detecting changed objects, typically using class activation map (CAM) technique combined with DenseCRF or the Segment Anything Model (SAM). However, these methods still face two main challenges: first, CAMs tend to produce weak or false activations for changed objects, and second, DenseCRF and SAM lead to unreliable pseudomask generation, particularly when complex variations occur within objects in bi-temporal images. To address these challenges, a change selection network (CSNet) is proposed to enhance the quality of change activation maps and pseudomasks, improving their ability to accurately extract changed regions in bi-temporal RS images. First, a change activation selection (CAS) module is designed to generate a weight mask that selects and aggregates change-representing features, effectively highlighting missed change activations and strengthening weak activations. Second, a bi-temporal image selection (BIS) strategy is developed, incorporating two selection rules to filter out image pairs with poor-quality mask derived from SAM, while retaining those with high-quality results. Finally, a change pseudomask generation (CPG) module integrated with an atrous-spatial pyramid pooling (ASPP) classifier is developed to predict accurate change pixels for final pseudomask generation. Experimental results demonstrate that the proposed CSNet outperforms existing WSCD methods, achieving 79.32% IoU in change pseudomasks for the WHU-CD dataset, 68.12% for the GZ-CD dataset and 75.44% for the GVLM dataset. This study proposes a novel method that enhances the performance of the weakly supervised paradigm in RS CD. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Weakly Supervised Building Extraction From High-Resolution Remote Sensing Images Based on Building-Aware Clustering and Activation Refinement NetworkabstractWeakly supervised building extraction methods, utilizing image-level labels, offer a cost-effective solution by significantly reducing the need for pixel-level annotation in high-resolution (HR) remote sensing (RS) images. These methods often focus on class activation map (CAM) optimization based on features extracted from individual images, missing out on the benefits of associating building features from multiple RS images (i.e., n images) to improve CAMs. This limitation leaves room for improvement in both CAM optimization and pseudo-mask generation. To address this, we propose the building-aware clustering and activation refinement network (BAC-AR-Net), a novel weakly supervised network to enhance weakly supervised building extraction performance. The building-aware clustering (BAC) module aggregates and clusters feature maps from multiple building samples to obtain common features of buildings. The common features are subsequently used to extract regions with similar building semantics, thereby enhancing the accuracy and completeness of building coverage in CAMs. Additionally, the activation refinement module is designed to generate pseudo-masks with clear boundaries and an effective separation of buildings and background. Experiments were conducted on the ISPRS Potsdam and Vaihingen datasets as well as a self-built building dataset to verify the effectiveness of our proposed method. The results show the proposed method outperforms both the weakly supervised semantic segmentation and weakly supervised building extraction methods that use image-level labels, achieving IoU accuracies of 0.8556, 0.8163, and 0.7797 on the respective datasets. This study introduces a novel weakly supervised learning framework to the RS application, with a particular focus on building extraction and semantic segmentation tasks. Daoyuan Zheng, Shaohua Wang 0003, Haixia Feng, Shunli Wang 0003, Mingyao Ai, Jiayuan Li 0001, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Optimal subsampling for large-scale quantile regression
Mingyao Ai |
J. Complex. | 1 |
| 2021 | Point Cloud Registration Based on One-Point RANSAC and Scale-Annealing Biweight EstimationabstractPoint cloud registration (PCR) is an important task in photogrammetry and remote sensing, whose goal is to seek a seven-parameter similarity transformation to register a pair of point clouds. Traditional iterative closest point (ICP) variants highly rely on the initial parameters, and most of them cannot deal with cross-source (multisource) point clouds with scale changes. In this article, we propose a flexible correspondence-based PCR method, which is initial-guess free, fast, and robust. We first decompose the full seven-parameter registration problem into three subproblems, i.e., scale, rotation, and translation estimations, based on line vectors. Then, we propose a one-point random sample consensus (RANSAC) algorithm to estimate the scale and translation parameters. For the rotation estimation, we introduce a graduated optimization strategy into Tukey’s biweight function and propose a scale-annealing biweight estimator. We evaluate the proposed method on both same-source and cross-source data. Results show that the proposed method is robust against over 99% outliers and is one to two orders of magnitude faster than its competitors. The source code of our method will be made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Learning From GPS Trajectories of Floating Car for CNN-Based Urban Road Extraction With High-Resolution Satellite ImageryabstractDeep learning has achieved great success in recent years, among which the convolutional neural network (CNN) method is outstanding in image segmentation and image recognition. It is also widely used in satellite imagery road extraction and, generally, can obtain accurate and extraction results. However, at present, the extraction of roads based on CNN still requires a lot of manual preparation work, and a large number of samples can be marked to achieve extraction, which has to take long drawing cycle and high production cost. In this article, a new CNN sample set production method is proposed, which uses the GPS trajectories of floating car as training set (GPSTasST), for the multilevel urban roads extraction from high-resolution remote sensing imagery. This method rasterizes the GPS trajectories of floating car into a raster map and uses the processed raster map to label the satellite image to obtain a road extraction sample set. CNN can extract roads from remote sensing imagery by learning the training set. The results show that the method achieves a harmonic mean of precision and recall higher than road extraction method from single data source while eliminating the manual labeling work, which shows the effectiveness of this work. Qingwu Hu, Jiayuan Li 0001, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Robust Geometric Model Estimation Based on Scaled Welsch q-NormabstractRobust estimation, which aims to recover the geometric transformation from outlier contaminated observations, is essential for many remote sensing and photogrammetry applications. This article presents a novel robust geometric model estimation method based on scaled Welsch q-norm (lq-norm, 0qLS) problem and a weighted least-squares (WLS) problem] by using alternating direction method of multipliers (ADMM) method. For the WLS problem, we introduce a coarse-to-fine strategy into the iterative reweighted least-squares (IRLS) method. We change the weight function by decreasing its scale parameter. This strategy can largely avoid that the solver gets stuck in local minimums. We adapt the proposed cost into classical remote sensing tasks and develop new robust feature matching (RFM), robust exterior orientation (REO), and robust absolute orientation (RAO) algorithms. Both synthetic and real experiments demonstrate that the proposed method significantly outperforms the other compared state-of-the-art methods. Our method is still robust even if the outlier rate is up to 90%. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | RIFT: Multi-Modal Image Matching Based on Radiation-Variation Insensitive Feature TransformabstractTraditional feature matching methods, such as scale-invariant feature transform (SIFT), usually use image intensity or gradient information to detect and describe feature points; however, both intensity and gradient are sensitive to nonlinear radiation distortions (NRD). To solve this problem, this paper proposes a novel feature matching algorithm that is robust to large NRD. The proposed method is called radiation-variation insensitive feature transform (RIFT). There are three main contributions in RIFT. First, RIFT uses phase congruency (PC) instead of image intensity for feature point detection. RIFT considers both the number and repeatability of feature points and detects both corner points and edge points on the PC map. Second, RIFT originally proposes a maximum index map (MIM) for feature description. The MIM is constructed from the log-Gabor convolution sequence and is much more robust to NRD than traditional gradient map. Thus, RIFT not only largely improves the stability of feature detection but also overcomes the limitation of gradient information for feature description. Third, RIFT analyses the inherent influence of rotations on the values of the MIM and realises rotation invariance. We use six different types of multi-modal image datasets to evaluate RIFT, including optical-optical, infrared-optical, synthetic aperture radar (SAR)-optical, depth-optical, map-optical, and day-night datasets. Experimental results show that RIFT is superior to SIFT and SAR-SIFT on multi-modal images. To the best of our knowledge, RIFT is the first feature matching algorithm that can achieve good performance on all the abovementioned types of multi-modal images. The source code of RIFT and the multi-modal image datasets are publicly available1. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Trans. Image Process. | 3 |
| 2019 | Haze and Thin Cloud Removal via Sphere Model Improved Dark Channel PriorabstractHaze and cloud seriously degrade the quality of optical remote sensing images, which largely decrease their interpretability and intelligibility. In this letter, we propose a two-stage haze and thin cloud removal method based on homomorphic filtering (HF) and sphere model improved dark channel prior (DCP). Compared with current dehazing methods, the most advantage of the proposed method is that our method can deal with uneven haze, thick haze, and thin cloud. We observe that haze and cloud are highly related to the illumination component and mainly located in the low frequency of an image. Thus, we adapt HF to enhance the haze image, which makes the distribution of haze more even. In the second stage, we analyze the drawback of DCP, i.e., the transmission estimated by DCP is very sensitive to noise. To draw this issue, we propose a novel sphere model to estimate a more accurate transmission map. The sphere model improved DCP is more suitable for thick haze images than the traditional DCP. Extensive experimental results show that the proposed method significantly outperforms the compared state-of-the-art methods. The source code and data sets used in the letter are made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2017 | Optimal Illumination and Color Consistency for Optical Remote-Sensing Image MosaickingabstractIllumination and color consistency are very important for optical remote-sensing image mosaicking. In this letter, we propose a simple but effective technique that simultaneously performs image illumination and color correction for multiview images. In this framework, we first present an uneven illumination removal algorithm based on bright channel prior, which guarantees the illumination consistency inside a single image. We then adapt a pairwise color-correction method to coarsely align the color tone between source and reference images. In this stage, we give a new single-image quality metric which combines brightness deviation, color cast, and entropy together for automatic reference-image selection. Finally, we perform a least-squares adjustment (LSA) procedure to obtain optimal illumination and color consistency among multiview images. In detail, we first perform a pairwise image matching by using SIFT algorithm; once sparse local patch correspondences obtained, the illumination and color relationship between images can be established based on a global gamma correction model; the illumination and color errors can then be minimized by LSA. Extensive experiments on both challenging synthetic and real optical remote-sensing image data sets show that it significantly outperforms the compared state-of-the-art approaches. All the source code and data sets used in this letter are made public. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2016 | Robust Feature Matching for Remote Sensing Image Registration Based on Lq-EstimatorabstractThis letter proposes a robust feature matching algorithm for remote sensing images based on lq-estimator. We start with a set of initial matches provided by a feature matching method such as scale-invariant feature transform and then focus on global transformation estimation from contaminated observations and outliers elimination as well. We use an affine model to describe the global transformation and minimize a new cost function based on lq-norm. We apply an augmented Lagrangian function and an alternating direction method of multipliers to solve such a nonconvex and nonsmooth optimization problem. Extensive experiments on real remote sensing data demonstrate that the proposed method is effective, efficient, and robust. Our method outperforms state-of-the-art methods and can easily handle situations with up to 90% outliers. In addition, the proposed method is much faster than RANSAC. Jiayuan Li 0001, Qingwu Hu, Mingyao Ai |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | Construction of uniform designs without replications
Bochuan Jiang, Mingyao Ai |
J. Complex. | 2 |