Jianqiao Luo

dblp:204/9738 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-7432-1531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Shape and Prototype-Guided Diffusion Model for Grape Amodal Completion in Vineyard Phenotyping Systems
abstract
Occlusions often lead to underestimated grape phenotypes, thereby impeding accurate vineyard management and yield prediction. In grape de-occlusion tasks, existing amodal completion methods perform poorly due to complex cluster structures and fine-grained local textures. To this end, we propose a shape and prototype guided diffusion model for high-fidelity amodal completion, by developing a variational shape embedding learning strategy and a dynamic prototype prior extraction module. We model the shape embedding as a mixture of von Mises–Fisher distributions, indicating the potential complete shape of occluded grape clusters. Subsequently, we formulate prototype priors as discrete embeddings that are consistent with patch features, which represent local texture characteristics of grape berries. The shape and prototype embeddings are integrated into the reverse diffusion process via cross-attention mechanisms, stabilizing structural predictions and mitigating common artifacts such as deformation and adhesion. We construct and release a grape amodal completion dataset collected from real-world vineyard environments. Experimental results on the grape dataset and public KINS dataset demonstrate the superiority of our method in terms of perceptual fidelity and phenotypic accuracy, highlighting its effectiveness for vision-based phenotyping in practical vineyard applications.
Yihan Wang 0006, Jianqiao Luo, Bailin Li, Mingtao Feng, Ajmal Mian
IEEE Trans Autom. Sci. Eng.3
2026 Uncertainty-Adaptive Volume for Unsupervised Homography Estimation
abstract
Estimating homography from an image pair is crucial for image alignment, and unsupervised methods that optimize feature reprojection error between target and warped source images have gained attention for their promising performance. In real-world scenes with multiple planes, such as moving objects, outlier rejection strategies are essential to mitigate the influence of non-dominant planes. Existing methods address this by learning a mask based on reprojection error, where high errors indicate non-dominant planes misaligned by homography. However, this error-fitting mask often overextends to the dominant plane, limiting the use of valid image regions for accurate estimation. This paper proposes a novel unsupervised method to compactly exclude non-dominant planes by introducing an uncertainty-adaptive cost volume for homography estimation. We first model uncertainty by assuming image features follow a Gaussian distribution derived from a prior Normal Inverse-Gamma distribution. The network-learned distribution parameters disentangle aleatoric uncertainty, distinguishing data-dependent errors within the total reprojection error. This uncertainty reflects inherent observation noise in image data, effectively indicating non-dominant planes. We then integrate this aleatoric uncertainty into the concatenation volume across image feature maps, creating an adaptive volume that filters out unreliable matching costs associated with non-dominant planes. This adaptive volume simplifies learning homography from the rich, redundant content in the concatenation volume, enabling more efficient and accurate estimation. Experiments demonstrate that our method outperforms existing approaches, achieving state-of-the-art performance both qualitatively and quantitatively.
Jianqiao Luo, Yaonan Wang 0001, Mingtao Feng, Zhen Zhou 0003, Xuebing Liu, Yang Mo
IEEE Trans. Circuits Syst. Video Technol.1
2026 Second-Order Robust Iterative Pose Optimization for Fine-Grained Cross-View Localization
abstract
Fine-grained cross-view localization seeks to estimate precise camera poses by matching ground images with GPS-tagged aerial imagery. Existing methods typically employ first-order iterative optimization to progressively update the camera pose based on cross-view feature correspondences. However, they rely on local features and neglect global and complementary contextual information, making them prone to local optima and slow convergence under large initial errors or strong disturbances. To overcome these limitations, we propose a second-order robust iterative pose estimation framework for fine-grained cross-view localization. Firstly, we devise a second-order deep iterative optimization module to capture complementary forward and backward motion cues, leading to a bidirectional correlation volume. A motion aggregator uses the volume to approximate the dynamics of second-order iterators, substantially facilitating convergence and robustness. In addition, a bidirectional motion-aware robust regularization module mitigates geometric distortions and outlier interference by leveraging bidirectional motion cues to generate fine-grained confidence maps, adaptively suppressing unreliable regions and enhancing the stability of iterative optimization and pose estimation accuracy. Extensive experiments demonstrate that the proposed framework achieves faster convergence and higher pose estimation accuracy than state-of-the-art methods, particularly under large initial errors and challenging conditions.
Mingtao Feng, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian
IEEE Trans. Image Process.3
2025 Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-Localization
abstract
Visual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing. This makes it challenging to automatically control the relative importance of the two tasks, potentially compromising the retrieval task in favor of the offset regression. Consequently, the model may encounter conflicting dominating gradients during joint training. To address this, we propose to model the semantic ambiguity during the offset regression process by integrating associated uncertainty scores, represented as 2D Gaussian distributions, to mitigate negative transfer effects within the joint tasks. We further introduce an uncertainty-aware similarity metric to enhance similarity assessment between query and reference images, accounting for their semantic ambiguities. This metric propagates uncertainty scores into the retrieval task, focusing on certain samples and learning discriminative feature embeddings, allowing the model to adaptively handle conflicting dominating gradients during joint training. Extensive experiments demonstrate that our method improves the overall performance of the joint tasks, achieving state-of-the-art results on the VIGOR and CVACT datasets.
Mingtao Feng, Fenghao Tian, Jianqiao Luo, Weisheng Dong, Yaonan Wang 0001, Ajmal Mian
AAAI3
2025 Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection
abstract
Tiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture to enhance the extinguished regions. We for the first time exploit the regions to be enhanced from the perspective of pixel-wise amount of information. Specifically, we model the entire image pixels feature information by minimizing Information Entropy loss, generating an information map to attentively highlight weak activated regions in an unsupervised way. To effectively assist the above phase with more attention to tiny objects, we next introduce the Position Gaussian Distribution Map, explicitly modeled using a Gaussian Mixture distribution, where each Gaussian component's parameters depend on the position and size of object instance labels, serving as supervision for further feature enhancement. Taking the information map as prior knowledge guidance, we construct a multi-scale position gaussian distribution map prediction module, simultaneously modulating the information map and distribution map to focus on tiny objects during training. Extensive experiments on three public tiny object datasets demonstrate the superiority of our method over current state-of-the-art competitors.
Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi
CVPR5
2025 Multi-range Adaptive Perception Transformer for Iterative Homography Estimation
abstract
Homography estimation is fundamental to various vision tasks. Iteration-based methods have recently achieved significant success in this field. However, errors introduced during iterations can lead to increased image deformation. Existing methods often focus on capturing local correspondences in the later stages of iteration while downplaying global ones, which may cause errors to persist and propagate into subsequent iterations, ultimately leading to error accumulation. To alleviate this issue, we propose Multi-range Adaptive Perception Transformer for Iterative Homography Estimation (MAPTHomo), which integrates Multi-range Attention (MRA) and Adaptive Perception Module (APM). Specifically, MRA captures both global and local correspondences, enabling the model to adapt to varying levels of deformation. The APM dynamically adjusts attention focus based on the current context. The combination of MRA and APM enhances the error-correction capability of the iterative process, effectively mitigating error accumulation. Extensive experiments demonstrate that MAPTHomo outperforms previous methods and exhibits strong generalization ability.
Tianming Li, Qing Zhu 0003, Zhen Zhou 0003, Jianqiao Luo, Yaonan Wang 0001
ICASSP4
2025 Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization
Mingtao Feng, Longlong Mei, Jianqiao Luo, Fenghao Tian, Jie Feng 0003, Weisheng Dong, Yaonan Wang 0001
ICCV4
2025 Generalizing to New Area: Self-Distillation Curriculum Learning for Fine-Grained Cross View Localization
abstract
Fine-grained cross-view localization seeks to predict ground-level camera positions within GPS-tagged aerial images by matching ground and aerial views. Existing methods often rely on large-scale ground truth annotations from specific regions, but performance degrades due to domain shifts when models trained in one area are applied to another. However, collecting region-specific annotations for each area is costly or infeasible. To address this, we propose a self-distillation curriculum learning framework that generalizes pretrained localization models to unseen new areas. Our approach introduces a Dirichlet-based quality assessment strategy to evaluate teacher-generated pseudo labels, where high uncertainty signals noisy predictions and low uncertainty indicates clean samples. This uncertainty is used to guide an easy-to-hard curriculum learning strategy, where easy samples are prioritized initially, and more challenging samples are progressively incorporated, enabling effective student training. Furthermore, we develop a joint optimization scheme that updates both the student model and pseudo labels, applying adaptive label smoothing to mitigate label noises and taking full advantage of new area data. Extensive experimental results on the VIGOR and KITTI benchmarks demonstrate that our method outperforms state-of-the-art approaches in new area localization, achieving superior accuracy without additional supervision.
Fenghao Tian, Mingtao Feng, Jianqiao Luo, Longlong Mei, Weisheng Dong, Yaonan Wang 0001
ACM Multimedia3
2025 DRLHomo: Disentangled Representation Learning for Cross-Modal Homography Estimation
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
PRCV (9)4
2025 Multiscale Spherical Feature Decoupling Network for Multimodal Image Registration
abstract
Multimodal image registration plays a crucial role in advancing Earth science. However, significant appearance variations and geometric deformations between multimodal images pose considerable challenges to this task. In this paper, we propose a novel multiscale spherical feature decoupling network (MSFDNet) for multimodal image registration by combining a multiscale iterative strategy with a multimodal decoupling strategy. MSFDNet adopts a multiscale architecture, with each scale incorporates a spherical feature decoupling (SFD) module with a carefully crafted three-stage decoupling strategy to bridge the modality gap. Specifically, we first introduce asymmetric shared and unique feature encoders to extract modality-shared and modality-unique features. Next, we design a spherical constraint learning (SCL) module to project the extracted features into spherical space, leveraging its regularized distance properties to enhance feature separability during the decoupling process. Finally, we propose a dual-path reconstruction mechanism that combines self-reconstruction with homography-guided cross-reconstruction to reconstruct the original multimodal images from the decoupled features, thereby simultaneously enhancing the learning of both feature decoupling and registration network. Based on the decoupled modality-shared features, we predict the registration function in a multiscale iterative manner, effectively bridging the geometric gap. Extensive experiments on multiple multimodal datasets validate the effectiveness of MSFDNet and demonstrate its state-of-the-art performance.
Tianming Li, Zhen Zhou 0003, Qing Zhu 0003, Jianqiao Luo, Yaonan Wang 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Uncertainty Guided Deep Lucas-Kanade Homography for Multimodal Image Alignment
abstract
Homography estimation for multimodal images poses a considerable challenge in computer vision because of content disparities and the diverse feature points captured by different sensors. Existing methods typically extract feature maps using neural networks and apply the Lucas-Kanade (LK) algorithm, which is based on the brightness constancy assumption, to solve the homography matrix. However, applying this assumption across all pixel features in multimodal images can lead to inaccuracies, as these images often contain noise, such as homogeneous regions or considerable appearance variations, which can corrupt the network’s training. To address this problem, we propose an uncertainty-guided deep LK (UG-DLK) framework that integrates uncertainty predictions to enhance the network’s iterative learning process. Specifically, we employ a probabilistic approach where the network predicts the distribution of the feature map rather than fixed values. By designing an uncertainty neighborhood estimator, we unfold the cost volume along the channels into 2-D slices, allowing the model to focus on neighborhood information at specific locations, effectively reducing the interference from spatial neighborhoods in the estimation of feature uncertainty. Through uncertainty modeling, the network can accurately identify scenes and objects that comply with the brightness constancy constraint, leading to more robust learning outcomes. Additionally, we introduce a novel loss function that incorporates feature uncertainty, leading to a smoother optimization landscape near the true homography parameters and reducing convergence oscillations. Our method, which is evaluated on benchmark datasets such as Google Maps, Google Earth, MSCOCO, and DPDN, demonstrates state-of-the-art performance, confirming the robustness and adaptability of our model across various scenarios.
Zhen Zhou 0003, Jianqiao Luo, Qing Zhu 0003, Yaonan Wang 0001, Hang Zhong, Mingtao Feng, Lin Chen 0034
IEEE Trans. Geosci. Remote. Sens.2
2025 Locally Aware Visual State Space for Small Defect Segmentation in Complex Component Images
abstract
Segmenting small defects within large imaging fields remains challenging in industrial scenarios due to the difficulty in distinguishing defects from complex component backgrounds and identifying defects comprising only a few pixels in high-resolution images. To address these issues, we propose a novel dual-branch feature extraction architecture, the locally aware visual state space block, which captures global contextual information while maintaining locally aware perception. In addition, we introduce the parallel quad-directional scanning fusion module to extract multiscale information, aggregating high-level features at different scales for enhanced global information fusion. To avoid losing small target details when upsampling the global segmentation mask to high-resolution input size, we develop progressive location refinement modules to incrementally refine small defect localization from the bottom up. Extensive experiments on our proposed small defect segmentation dataset and a public PCB dataset demonstrate that our method outperforms existing state-of-the-art methods in both performance and efficiency.
Jinghao Bian, Mingtao Feng, Weisheng Dong, Jianqiao Luo, Yaonan Wang 0001, Guangming Shi
IEEE Trans. Ind. Informatics5
2024 Unsupervised Homography Estimation With Pixel-Level SVDD
abstract
Homography estimation is a common image alignment method. Unsupervised learning, which uses unlabeled training and exhibits excellent performance, has attracted much attention in this field. When there are multiple planes in the scene, using features over the entire image for matching will lead to compromised results. However, existing methods for learning focused principal plane masks through deep neural networks lack explicit guidance. In this paper, we propose a novel unsupervised method to explicitly model anomaly descriptor removal and mask generation. Specifically, reliable feature descriptors are selected from a novel perspective, and regard the features that are not responsible for alignment as outliers. The pixel-level support vector data description (PL-SVDD) module is designed. This module learns the feature representation of image pixels and fits a hypersphere to exclude the feature redundancy information that is not responsible for alignment from the hypersphere, thereby optimizing the feature descriptor. Based on the optimized image features, a correlation learning (CL) module is designed. This module displays a generated mask through mathematical modeling to select reliable areas for homography estimation. Specifically, the feature descriptor of one unaligned images is modeled as a multivariate Gaussian distribution by Gaussian density estimation (GDE). Then, The Mahalanobis distance is combined with the multivariate Gaussian distribution of the model and the feature descriptor of another image to generate the mask. Experiments show that our method achieves good performance compared with previous methods.
Zhen Zhou 0003, Qing Zhu 0003, Mingtao Feng, Yaonan Wang 0001, Jianqiao Luo, Zhiqiang Miao, Lin Chen 0034, Yang Mo
IEEE Trans. Circuits Syst. Video Technol.5
2022 Single image denoising via multi-scale weighted group sparse coding
M. N. S. Swamy 0001, Jianqiao Luo, Bailin Li
Signal Process.3
2021 Topic-based label distribution learning to exploit label ambiguity for scene classification
Jianqiao Luo, Bailin Li
Neural Comput. Appl.1
2020 Gray-level image denoising with an improved weighted sparse coding
Jianqiao Luo, Bailin Li, M. N. S. Swamy 0001
J. Vis. Commun. Image Represent.2
2019 A classification model of railway fasteners based on computer vision
Jianqiao Luo, Bailin Li
Neural Comput. Appl.2
2017 Random Sampling Local Binary Pattern Encoding Based on Gaussian Distribution
abstract
The original local binary pattern (LBP) operator or LBP variants adopt the difference between the neighboring pixels and the center pixel to describe the pixel that does not consider the relationship between the neighboring pixels. The block region characteristics of an image are determined by the relationship between neighboring pixels, not just the neighboring pixels and the center pixel. In this letter, a new local neighborhood encoding method is proposed, which we call random sampling LBP (RSLBP). Based on the distribution of the image difference signal, point pairs are randomly selected in the local neighborhood, and LBP encoding is carried out after comparing the sums of pixels neighboring the random point. Image local difference is more obvious and noise resistance is better. By comparing the classification results of LBP and LBP variants with the proposed method, we show that the proposed method achieves better classification performance on the standard images library and the real fastener images, and the performance gain is significant when the noise level is high.
Bailin Li, Jianqiao Luo
IEEE Signal Process. Lett.4