Huabing Zhou

dblp:167/8080 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-5007-7303ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Task-aware dynamic routing network for cross-domain few-shot learning
Yanan Li 0006, Haoyang Ye, Huabing Zhou, Tao Lu 0001, Hao Lu 0003
Neurocomputing3
2026 CrossGlue: Cross-Modal Image matching via potential message investigation and visual-gradient message integration
Chaobo Yu, Zhonghui Pei, Huabing Zhou
J. Vis. Commun. Image Represent.4
2026 A DINO-based progressive semantic enhanced infrared and visible image fusion network
Shihan Yao, Zhonghui Pei, Huiqin Zhang, Haiyang Jiang 0021, Huabing Zhou
Neural Networks5
2026 Selecting and Pruning: A Differentiable Causal Sequentialized State-Space Model for Two-View Correspondence Learning
abstract
Two-view correspondence learning aims to discern true and false correspondences between image pairs by recognizing their underlying different information. Previous methods either treat the information equally or require the explicit storage of the entire context, tending to be laborious in real-world scenarios. Inspired by Mamba's inherent selectivity, we propose CorrMamba, a Correspondence filter leveraging Mamba's ability to selectively mine information from true correspondences while mitigating interference from false ones, thus achieving adaptive focus at a lower cost. To prevent Mamba from being potentially impacted by unordered keypoints that obscured its ability to mine spatial information, we customize a causal sequential learning approach based on the Gumbel-Softmax technique to establish causal dependencies between features in a fully autonomous and differentiable manner. Additionally, a local-context enhancement module is designed to capture critical contextual cues essential for correspondence pruning, complementing the core framework. Extensive experiments on relative pose estimation, visual localization, and analysis demonstrate that CorrMamba achieves state-of-the-art performance. Notably, in outdoor relative pose estimation, our method surpasses the previous SOTA by 2.58 absolute percentage points in AUC@20°, highlighting its practical superiority. Our code is publicly available at https://github.com/ShineFox/CorrMamba.
Hao Zhang 0073, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001
IEEE Trans. Image Process.5
2025 Multi-Shape Matching with Cycle Consistency Basis via Functional Maps
abstract
Multi-shape matching is a central problem in various applications of computer vision and graphics, where cycle consistency constraints play a pivotal role. For this issue, we propose a novel and efficient approach that models multi-shapes as directed graphs for two-stage optimization, i.e., optimizing pairwise correspondence accuracy using landmarks, and refining matching consistency through cycle consistency basis. Specifically, we utilize local mapping distortion to identify landmarks and extract the dimension of the functional space, which is then used to upsample in the spectral domain, thereby producing smoother results. Next, to optimize the consistency of correspondences, we introduce the cycle consistency basis, which succinctly describes all consistent cycles in the collection. We then propose cycle consistency refinement, which resolves inconsistencies in cycles efficiently via the alternating direction method of multipliers. Our approach simultaneously balances the accuracy and consistency of multi-shape matching, achieving lower correspondence errors. Extensive experiments on several public datasets demonstrate the superiority of our approach over current state-of-the-art methods.
Tianwei Ye, Huabing Zhou, Zhongyuan Wang 0001, Jiayi Ma 0001
AAAI3
2025 End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph Generation
Yanduo Zhang, Tao Lu 0001, Huiqin Zhang, Jiayi Ma 0001, Huabing Zhou
ICCV7
2025 DiffuseDoc: Document geometric rectification via diffusion model
Wenfei Xiong, Huabing Zhou, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001
Comput. Vis. Image Underst.2
2025 Exploring fusion domain: Advancing infrared and visible image fusion via IDFFN-GAN
abstract
Infrared (IR) and visible (VI) images are crucial in applications such as surveillance and night vision, where each modality provides complementary information—IR captures thermal details, while VI captures textures. Fusing these images is essential to combine their strengths , resulting in a more comprehensive and informative image. In this work, we introduce the “fusion domain” concept, a unique distribution that optimally blends IR and VI features. Our primary contribution is developing the Intermediate Domain Feature Fusion Network (IDFFN), which employs an MLP to learn the optimal domain factor for this fusion. Integrated with the IDFFN-GAN’s dual discriminators , our approach refines the fusion process to produce images that better preserve thermal and texture details from the proposed fusion and gradient domains. Experimental results demonstrate that our method achieves a 5% boost in Entropy (EN), a 50% rise in Spatial Frequency (SF), a 58% improvement in Average Gradient (AG), a 10% enhancement in Standard Deviation (SD), and a remarkable 198% increase in N A B / F compared to state-of-the-art techniques, ensuring superior retention of specific features from both modalities.
Xiaoqian Shi, Yanan Li 0006, Huabing Zhou
Neurocomputing4
2025 MFINet: a multi-scale feature interaction network for point cloud registration
Haiyuan Cao, Deng Chen, Yanduo Zhang, Huabing Zhou, Dawei Wen, Congcong Cao
Vis. Comput.4
2024 A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic Consistency
abstract
This work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Through the coupled optimization of the visual fusion loss and the lesion segmentation loss, visual-related and semantic-related features will be pulled into the same domain, effectively promoting accuracy improvement in a mutual-reinforcing manner. Further, we establish the robustness guarantees by constructing a two-level refinement constraint in the process of feature extraction and reconstruction. Benefiting from full consideration for common degradations in medical images, our framework can not only provide clear visual fusion results for doctor's observation, but also enhance the defense ability of lesion segmentation against these negatives. Extensive evaluations of visual fusion and lesion segmentation scenarios demonstrate the advantages of our method in terms of accuracy and robustness. Moreover, our proposed framework is generic, which can be well-compatible with existing lesion segmentation algorithms and improve their performance. The code is publicly available at https://github.com/HaoZhang1018/RMR-Fusion.
Hao Zhang 0073, Xuhui Zuo, Huabing Zhou, Tao Lu 0001, Jiayi Ma 0001
AAAI3
2024 PlantBiCNet: A new paradigm in plant science with bi-directional cascade neural network for detection and counting
Jianxiong Ye 0002, Zhenghong Yu, Yangxu Wang, Dunlu Lu, Huabing Zhou
Eng. Appl. Artif. Intell.5
2024 TasselLFANetV2: Exploring Vision Models Adaptation in Cross-Domain
abstract
The datasets collected by people are always just a sampling of the real world. In this letter, we explore the possibility of achieving high-quality domain adaptation (DA) without explicit adaptation. As a baseline, we implemented the significantly improved second-generation version of TasselLFANet, TasselLFANetV2. This model, with indicators reaching AP50of 0.981 and R2of 0.9684, demonstrates leading performance in two typical cross-domain settings of data distribution scenarios, agriculture and remote sensing (RS), exhibiting strong domain adaptation and generalization, surpassing advanced methods such as YOLOv8-UAV, PlantBiCNet, SLA, etc. We further studied the combination of regularization techniques and feature re-mapping modules can effectively alleviate the domain invariance of the model. What’s more, when the training set and validation set are set the same, the training performance of the model is better, but the premise is that there must be a proper data transformation strategy. This work provides a new perspective for understanding and solving the problem of domain difference in deep learning. The code, datasets can be accessed at https://github.com/Ye-Sk/TasselLFANetV2.
Zhenghong Yu, Jianxiong Ye 0002, Shengjie Liufu, Dunlu Lu, Huabing Zhou
IEEE Geosci. Remote. Sens. Lett.5
2024 CATNet: Convolutional attention and transformer for monocular depth estimation
Tongwei Lu, Xuanxuan Liu, Huabing Zhou, Yanduo Zhang
Pattern Recognit.4
2024 SegCLIP: Multimodal Visual-Language and Prompt Learning for High-Resolution Remote Sensing Semantic Segmentation
abstract
Remote sensing semantic segmentation is considered a key step in the intelligent interpretation of high-resolution remote sensing (HRRS) images, with widespread applications in fields such as hazard assessment, environmental monitoring, and urban planning. Recently, numerous deep learning-based semantic segmentation methods have emerged, achieving significant breakthroughs. However, the majority of current research still concentrates on representation learning in the visual feature space, with the potential of multimodal data sources yet to be fully explored. In recent years, the foundational visual language model, namely contrastive language-image pretraining (CLIP), has established a new paradigm in the visual field, demonstrating excellent generalization capabilities and deep semantic understanding across a variety of tasks. Inspired by prompt learning, we propose a prompting approach based on linguistic descriptions to enable CLIP to generate semantically distinct contextual information for remote sensing images. We introduce the SegCLIP network architecture, a novel framework specifically designed for semantic segmentation of HRRS images. Specifically, we have adapted CLIP to extract text information, thereby guiding the visual model in distinguishing among classes. Additionally, we have designed a cross-modal feature fusion (CFF) module that integrates linguistic and visual semantic features, ensuring semantic consistency across modalities. Finally, we have fully exploited the potential of text data and have used additional real text to refine ambiguous query features. Experimental evaluations confirm that the method exhibits superior performance on the LoveDA, iSAID, and UAVid public semantic segmentation datasets.
Bin Zhang 0033, Yuntao Wu, Huabing Zhou, Junjun Jiang, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Face Deformation Under Feature Transfer and Geometric Deformation
abstract
Image deformation refers to deforming objects in images into a target shape or posture. Although point-based image deformation algorithms have made breakthroughs in performance and visual effects, they are limited to warping the current structure information in the source image. For example, while opening a closed mouth in a face deformation, the point-based algorithm cannot generate teeth, which causes the mouth to twist weirdly. Deep learning-based face editing models can generate new parts but cannot achieve fine pixel-level manipulation. To overcome these challenges, we propose a two-step strategy for face deformation. We first generate a high-resolution intermediate image by blending the source image and the specific part of the target image via our face blending generative adversarial network (FB-GAN). Then, we employ a state-of-the-art point-based geometric deformation method to deform the intermediate image with target face guidance. Extensive experiments show that the proposed FB-GAN can generate realistic and high-resolution results and demonstrates that the two-step face deformation strategy can be better applied to human face deformation.
Dazhi Zhang, Huabing Zhou
IEEE Signal Process. Lett.4
2023 Semantic-Supervised Infrared and Visible Image Fusion Via a Dual-Discriminator Generative Adversarial Network
abstract
Image fusion synthesizes a new image from multiple images of the same scene. The synthesized image should be suitable for human visual perception and follow-up high-level image-processing tasks. However, existing methods focus on fusing low-level features, ignoring high-level semantic perception information. We propose a new end-to-end model to obtain a more semantically consistent image in infrared and visible image fusion, termedsemantic-supervised dual-discriminator generative adversarial network(SDDGAN). In particular, we design an information quantity discrimination (IQD) block to guide fusion progress. For each source image, the block determines the weight for preserving each semantic object’s feature. By this way, the generator learns to fuse various semantic objects via different weights to preserve their characteristics. Moreover, the dual discriminator is employed to identify the distribution of infrared and visible information in the fused image. Each discriminator acts on a certain modality (infrared/visible) of different semantic objects in the fused image to preserve and enhance their modality features. Thus, our fused image is more informative. Both the thermal radiation in the infrared image and the visible image texture details can be well preserved. Qualitative and quantitative experiments demonstrate the superiority of our SDDGAN over state-of-the-art methods in terms of visual effects, efficiency, and quantitative metrics.
Huabing Zhou, Yanduo Zhang, Jiayi Ma 0001, Haibin Ling
IEEE Trans. Multim.1
2022 Interpolation-based nonrigid deformation estimation under manifold regularization constraint
Huabing Zhou, Yulu Tian, Zhenghong Yu, Yanduo Zhang, Jiayi Ma 0001
Pattern Recognit.1
2022 Loop-Closure Detection Using Local Relative Orientation Matching
abstract
Loop-closure detection (LCD), which aims to recognize a previously visited location, is a crucial component of the simultaneous localization and mapping system. In this paper, a novel appearance-based LCD method is presented. In particular, we propose a simple yet surprisingly useful feature matching algorithm for real-time geometrical verification of candidate loop-closures, termed aslocal relative orientationmatching (LRO). It aims to efficiently establish reliable feature correspondences based on preserving local topological structures between the query image and candidate frame. To effectively retrieve candidate loop closures, we introduce the aggregated selective match kernel framework into the LCD task, which can effectively represent images and reduce the quantization noise of the traditional bag-of-words framework. In addition, the SuperPoint neural network is employed to extract reliable interest points and feature descriptors. Extensive experimental results demonstrate that our LRO can significantly improve the LCD performance, and the proposed overall LCD method can achieve much better performance over the current state-of-the-art on six publicly available datasets.
Jiayi Ma 0001, Xinyu Ye, Huabing Zhou, Xiaoguang Mei, Fan Fan 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Variation-Net: Interpretable Variation-Inspired Deep Network for Pansharpening
abstract
In this study, we propose Variation-net, an interpretable variation-inspired deep network for pansharpening, which aims to fuse panchromatic (PAN) and multispectral (MS) images for a high-resolution MS image. We first construct a novel variational pan-sharpening model with clear physical meanings. As the relationship between the PAN and MS images in the real situation is complex and nonlinear, we explore the similarity between PAN and MS images from the sparsity of nonlinear transforms in this variational pansharpening model. As a result, spatial details can be accurately transferred from PAN image to MS image. Furthermore, we build the Variation-net by unrolling the iterative shrinkage-thresholding algorithm to solve the proposed variational pansharpening model. Therefore, all modules in Variation-net have clear physical meanings and are easily observed, leading to good generalization capability. Meanwhile, nonlinear transforms and other parameters in the variational pansharpening model are learned end–to–end. The experiments demonstrate that Variation-net outperforms the state-of-the-art methods from the aspects of visual effect and objective quality analysis.
Kun Li 0025, Wei Zhang 0259, Xin Tian 0006, Jiayi Ma 0001, Huabing Zhou, Zhongyuan Wang 0001
ICME5
2021 Motion Field Consensus with Locality Preservation: A Geometric Confirmation Strategy for Loop Closure Detection
abstract
Loop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select semantically similar images based on the SuperPoint network and the aggregated selective match kernel in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. Based on the potential property of motion field in the LCD scene, a robust feature matching algorithm, termed as motion field consensus with locality preservation (MFC-LP), is proposed. In particular, we exploit the smoothness prior to guide the learning of the motion field for an image pair in a reproducing kernel Hilbert space (RKHS). Meanwhile, to enhance the local relevance of motion vectors, we design a locality preservation mechanism thus making the learned motion field more accurate. Extensive experiments on several publicly available datasets reveal that MFC-LP has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task.
Kaining Zhang, Xingyu Jiang 0005, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001
IROS4
2020 Cross-Weather Image Alignment via Latent Generative Model With Intensity Consistency
abstract
Image alignment/registration/correspondence is a critical prerequisite for many vision-based tasks, and it has been widely studied in computer vision. However, aligning images from different domains, such as cross-weather/season road scenes, remains a challenging problem. Inspired by the success of classic intensity-constancy-based image alignment methods and the modern generative adversarial network (GAN) technology, we propose a cross-weather road scene alignment method called latent generative model with intensity constancy. From a novel perspective, the alignment problem is formulated as a constrained 2D flow optimization problem with latent encoding, which can be decoded into an intensity-constancy image on the latent image manifold. The manifold is parameterized by a pre-trained GAN, which is able to capture statistic characteristics from large datasets. Moreover, we employ the learned manifold to constrain the warped latent image identical to the target image, thereby producing a realistic warping effect. Experimental results on several cross-weather/season road scene datasets demonstrate that our approach can significantly outperform the state-of-the-art methods.
Huabing Zhou, Jiayi Ma 0001, Chiu C. Tan 0001, Yanduo Zhang, Haibin Ling
IEEE Trans. Image Process.1
2019 Locality Preserving Matching
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Xiaojie Guo 0001
Int. J. Comput. Vis.4
2019 Color and depth image registration algorithm based on multi-vector-fields constraints
Daoqing Li, Li Peng 0003, Huabing Zhou, Deng Chen, Yanduo Zhang, Liang Xie 0001
Multim. Tools Appl.4
2019 Large scale image retrieval with DCNN and local geometrical constraint model
Huabing Zhou, Yiwei Tao, Jinshu Shi, Deng Chen, Yanduo Zhang, Liang Xie 0001
Multim. Tools Appl.1
2019 Nonrigid Point Set Registration With Robust Transformation Learning Under Manifold Regularization
abstract
This paper solves the problem of nonrigid point set registration by designing a robust transformation learning scheme. The principle is to iteratively establish point correspondences and learn the nonrigid transformation between two given sets of points. In particular, the local feature descriptors are used to search the correspondences and some unknown outliers will be inevitably introduced. To precisely learn the underlying transformation from noisy correspondences, we cast the point set registration into a semisupervised learning problem, where a set of indicator variables is adopted to help distinguish outliers in a mixture model. To exploit the intrinsic structure of a point set, we constrain the transformation with manifold regularization which plays a role of prior knowledge. Moreover, the transformation is modeled in the reproducing kernel Hilbert space, and a sparsity-induced approximation is utilized to boost efficiency. We apply the proposed method to learning motion flows between image pairs of similar scenes for visual homing, which is a specific type of mobile robot navigation. Extensive experiments on several publicly available data sets reveal the superiority of the proposed method over state-of-the-art competitors, particularly in the context of the degenerated data.
Jiayi Ma 0001, Jia Wu 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Quan Z. Sheng
IEEE Trans. Neural Networks Learn. Syst.5
2018 Facial Shape and Expression Transfer via Non-rigid Image Deformation
Huabing Zhou, Shiqiang Ren, Yuyu Kuang, Yanduo Zhang, Wei Zhang 0259, Tao Lu 0001, Hanwen Chen, Deng Chen
ICA3PP (3)1
2018 Face Hallucination Using Manifold-Regularized Group Locality-Constrained Representation
abstract
Sparsity and locality regularizations are successfully applied to face hallucination algorithms to ameliorate their ill-posed nature. However, most of patch-based face hallucination approaches only consider the manifold structure of single patch, thus resulting in unstable solution for image reconstruction. In this paper, we propose a novel face hallucination, termed manifold-regularized group locality-constrained representation (MGLR), in order to exploit the multiple manifold structures rooted in grouped self-similarly patches. Specifically, we first group similar patches to form a matrix which contains the recurrent non-local patches. Then graph regularization term is formulated to represent the group manifolds for better reconstruction quality. Taking advantages of grouped self-similar patches, MGLR can offer stable sparse solution to take advantage of the the accurate prior for super-resolution reconstruction. Experimental results on LFW database and CMU real-world images demonstrate the superiority of the proposed method over some state-of-the-art face methods both in terms of subjective and objective qualities.
Tao Lu 0001, Kangli Zeng, Junjun Jiang, Yanduo Zhang, Zhongyuan Wang 0001, Huabing Zhou
ICIP7
2018 Visual Homing via Guided Locality Preserving Matching
abstract
This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoramic images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from hundreds of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost true matches without sacrifice in accuracy. To apply our GLPM to the visual homing problem, we develop a method for dense motion flow estimation from sparse feature matches based on Tikhonov regularization. Moreover, the focus-of-contraction/focus-of-expansion is derived to determine homing directions. The effectiveness of our method is demonstrated on a panoramic database in both feature matching and visual homing.
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Yu Zhou 0016, Zheng Wang 0007, Xiaojie Guo 0001
ICRA4
2018 Non-rigid point set registration via global and local constraints
Changcai Yang, Meifang Zhang, Zejun Zhang 0001, Lifang Wei, Riqing Chen, Huabing Zhou
Multim. Tools Appl.6
2018 Guided Locality Preserving Feature Matching for Remote Sensing Image Registration
abstract
Feature matching, which refers to establishing reliable correspondences between two sets of feature points, is a critical prerequisite in feature-based image registration. This paper proposes a simple yet surprisingly effective approach, termed as guided locality preserving matching, for robust feature matching of remote sensing images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost the true matches without sacrifice in accuracy. Experiments on various real remote sensing image pairs demonstrate the generality of our method for handling both rigid and nonrigid image deformations, and it is more than two orders of magnitude faster than the state-of-the-art methods with better accuracy, making it practical for real-time applications.
Jiayi Ma 0001, Junjun Jiang, Huabing Zhou, Ji Zhao 0001, Xiaojie Guo 0001
IEEE Trans. Geosci. Remote. Sens.3
2017 Non-Rigid Point Set Registration with Robust Transformation Estimation under Manifold Regularization
abstract
In this paper, we propose a robust transformation estimation method based on manifold regularization for non-rigid point set registration. The method iteratively recovers the point correspondence and estimates the spatial transformation between two point sets. The correspondence is established based on existing local feature descriptors which typically results in a number of outliers. To achieve an accurate estimate of the transformation from such putative point correspondence, we formulate the registration problem by a mixture model with a set of latent variables introduced to identify outliers, and a prior involving manifold regularization is imposed on the transformation to capture the underlying intrinsic geometry of the input data. The non-rigid transformation is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both 2D and 3D data demonstrate that our method can yield superior results compared to other state-of-the-arts, especially in case of badly degraded data.
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou
AAAI4
2017 Book Page Identification Using Convolutional Neural Networks Trained by Task-Unrelated Dataset
Leyuan Liu 0001, Huabing Zhou, Jingying Chen 0001
ICIG (1)3
2017 Face hallucination using region-based deep convolutional networks
abstract
Most deep learning based face hallucinations exploit random patch prior from training samples, then to learn the mapping functions between low-resolution (LR) and high-resolution (HR) images, and achieve satisfactory reconstruction performance. However, most of them do not take into account the prior information on facial structure, which is pivotal for face hallucination. Different from random patch prior based deep learning approaches, in this paper, we utilize facial structural prior and develop a simple yet powerful face hallucination, named region-based deep convolutional networks (RDCN). Firstly, we divide facial image into several regions of interest, then to train multiple parallel subnetworks of these regions for exacting better structure priors, finally HR output is reconstructed by stitching facial parts. Experiments on the FEI database demonstrate that the proposed region-based convolution networks outperform other state-of-the-art, including recently proposed deep learning based approaches, both in subjective and objective reconstruction qualities.
Tao Lu 0001, Hao Wang 0237, Zixiang Xiong, Junjun Jiang, Yanduo Zhang, Huabing Zhou, Zhongyuan Wang 0001
ICIP6
2017 Non-rigid image deformation algorithm based on MRLS-TPS
abstract
In this paper, we propose a novel closed-form transformation estimation method based on moving regularized least squares optimization with thin-plate spline (MRLS-TPS) for non-rigid image deformation. The method takes the user-controlled point-offset-vectors as the input data, and estimates the spatial transformation about the two control point sets for each pixel. To achieve a realistic deformation, we formulates the transformation estimation as a vector-field interpolation problem by a moving regularized least squares method. Unlike MLS, the mapping function is modeled by a non-rigid function thin-plate spline with regularization technique, such that the deformation can satisfy both global linear affine motion and local non-rigid warping. We derive a closed-form solution of the transformation and achieve a fast implementation. In addition, the proposed method can give a wonderful user experience, fast and convenient manipulating. Extensive experiments on real images demonstrated the proposed method outperforms other state-of-the-art methods and the commercial software Adobe PhotoShop CS 6, especially in case of flexible object motion.
Huabing Zhou, Yuyu Kuang, Zhenghong Yu, Shiqiang Ren, Anna Dai, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001
ICIP1
2017 Non-rigid feature matching for image retrieval using global and local regularizations
abstract
In this paper, we propose a probabilistic method for feature matching of near-duplicate images undergoing non-rigid transformations. We start by creating a set of putative correspondences based on the feature similarity, and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose the non-parametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the Expectation Maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real near-duplicate images for both feature matching and image retrieval demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, especially in case of severe outliers.
Yong Ma 0001, Huabing Zhou, Jun Chen 0001, Jingshu Shi, Zhongyuan Wang 0001
ICME2
2017 A unified model for improving depth accuracy in kinect sensor
abstract
The Microsoft Kinect sensor has been widely used in many applications, but it suffers from the drawback of low depth accuracy. In this paper, we present a unified depth modification model to improve the Kinect depth accuracy by registering depth and color images in an iterative manner. Specifically, in each iteration, we first establish a coarse correspondence based on the feature descriptor of the canny edge. Then, we estimate the fine correspondence using a robust estimator called the L2E with the nonparametric model. Finally, we correct the depth data according to the correspondence results. In order to evaluate the effectiveness of our approach, we have performed extensive experiments and then analyzed the experimental results from the following respects: the accuracy of depth data, the accuracy of correspondence between color and depth images as well as the measurement error in the 3D reconstruction by our method. The experimental results show that our approach greatly improves the depth accuracy.
Li Peng 0003, Yanduo Zhang, Huabing Zhou, Deng Chen, Zhenghong Yu, Junjun Jiang, Jiayi Ma 0001
ICME3
2017 Feature guided non-rigid image/surface deformation via moving least squares with manifold regularization
abstract
In this paper, a novel closed-form transformation estimation method based on feature guided moving least squares together with manifold regularization is proposed for nonrigid image/surface deformation. The method takes the user-controlled point-offset-vectors and the feature points of the image/surface as input, and estimates the spatial transformation between the two control point sets for each pixel/voxel. To achieve a detail-preserving and realistic deformation, the transformation estimation is formulated as a vector-field interpolation problem using a feature guided moving least squares method, where a manifold regularization is imposed as a prior on the transformation to capture the underlying intrinsic geometry of the input image/surface. The non-rigid transformation is specified in a reproducing kernel Hilbert space. We derive a closed-form solution of the transformation and adopt a sparse approximation to achieve a fast implementation, which largely reduces the computation complexity without performance sacrifice. In addition, the proposed method can give a wonderful user experience, fast and convenient manipulating. Extensive experiments on both 2D and 3D data demonstrate that the proposed method can produce more natural deformations compared with other state-of-the-art methods.
Huabing Zhou, Jiayi Ma 0001, Yanduo Zhang, Zhenghong Yu, Shiqiang Ren, Deng Chen
ICME1
2017 Locality Preserving Matching
abstract
Seeking reliable correspondences between two feature sets is a fundamental and important task in computer vision. This paper attempts to remove mismatches from given putative image feature correspondences. To achieve the goal, an efficient approach, termed as locality preserving matching (LPM), is designed, the principle of which is to maintain the local neighborhood structures of those potential true matches. We formulate the problem into a mathematical model, and derive a closed-form solution with linearithmic time and linear space complexities. More specifically, our method can accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. Experiments on various real image pairs for general feature matching, as well as for visual homing and image retrieval demonstrate the generality of our method for handling different types of image deformations, and it is more than two orders of magnitude faster than state-of-the-art methods in the same range of or better accuracy.
Jiayi Ma 0001, Ji Zhao 0001, Hanqi Guo 0002, Junjun Jiang, Huabing Zhou, Yuan Gao 0015
IJCAI5
2016 Interpretation of Chinese Address Information Based on Multi-factor Inference
abstract
Interpretation of Chinese address information, which refers to translating the non-specification Chinese address into the concrete administration divisions, is a critical prerequisite in Internet geography semantic comprehension and Chinese address navigation. In this paper, We proposed Multi-factor(MF) Inference technology to solve the problem of vague address semantics. Firstly, we define the location and match factors of each geographical element in a series of address, which need to be included in a various address relationship. Secondly, the reliability of each possible address string is inferred. Finally, the highest reliability is choosed. Expriments show that our method outperforms current state-of-the-art methds.
Yanhui Duan, Huabing Zhou
ISPDC3
2016 Animation Generating Based on MRLS Image Deformation
abstract
Animation generating, which refers to generating a movie or animation according to one or more images, has a number of useful application ranging from movie effects to virtual reality. This paper proposes a novel animation generating method via MRLS image deformation and designs an animation system, which can generate a vividly animation from a static image by choosing set of point handles and deformation paths. Firstly, users need to choose a set of fixed control points and a distorted image deformation path. Then, construct deformation mapping functions according to the control points and the deformation path, and get a series of deformed images with the mapping function. Finally, users get the animation by packaging the deformed images into an AVI video files frame-by-frame. Experiment results show that the non-rigid deformation animations produced with our method are deformed smoothly, naturally and realistically.
Shiqiang Ren, Huabing Zhou, Zhengjun Li
ISPDC2
2016 An Image-Based Approach to Automatic Crop Organ Extraction via Low-Rank Matrix Recovery
abstract
Automatic extraction of crop organ from images is a crucial step for quantitatively acquiring crop growth information in precision agriculture. There has been some attempt on this task, but the performance is not satisfactory. In this paper, we proposed an image-based method based on low-rank matrix recovery to extract organ accurately. In our method, a crop image is considered to be compose of two factors: background and organ. In a certain feature space, the image is represented as a low-rank matrix plus sparse noises. The organ is then extracted by identifying the sparse noises when using low-rank matrix recovery algorithm. In order to ensure the rank of background is low, a linear transform for the feature space is introduced and needs to be learned from historical data. Dynamic threshold segmentation followed by vegetation removing techniques are ultimately adopted in the final step. The experimental results on the benchmark farmland dataset show that our method achieve competitive performance, compared with the other well-established methods, yielding the highest performance of 93.9% with the lowest standard deviation of 2.86%, which means our method is more robust and not sensitive to the complex environmental elements and different cultivars.
Zhenghong Yu, Haichang Yin, Haijie Feng, Minfang Chen, Huabing Zhou, Tongwei Lu, Feng Min
ISPDC5
2016 Non-rigid Point Set Registration via Coherent Spatial Mapping and Local Structures Preserving
abstract
Non-rigid point set registration is a fundamental problem for many computer vision technologies. In this paper, we proposed a new non-rigid point set registration method based on coherent spatial mapping (CSM) and local geometrical constraint. Our central idea is to express each point as a weighted sum of several nearest neighbors and the same relation holds after the transformation. The registration problem is solved by minimizing an error function, which combines the the global model and local geometrical constraint. The registration experiments are undertaken on various synthetic and real data. The results demonstrate that the proposed approach is robust and is superior to the state-of-the-art methods.
Meifang Zhang, Changcai Yang, Lifang Wei, Zejun Zhang 0001, Riqing Chen, Huabing Zhou
ISPDC6
2016 Hardware-Efficient Architecture of Photo Core Transform in JPEG XR for Low-Cost Applications
abstract
This paper proposes a novel two-input, two-output architecture of photo core transform (PCT) in JPEG XR. First, the lifting operations of PCT are optimized such that some operations can be reused. Then the time multiplexing technique is used by combining lifting steps to design a hardware-efficient architecture. Experimental results based on field programmable gate array (FPGA) demonstrate that our architecture has good performance in reducing hardware resources and power consumption, which could be an efficient alternative in low-cost applications.
Shuiping Zhang, Huabing Zhou, Haihui Wang
ISPDC2
2016 Nonrigid Feature Matching for Remote Sensing Images via Probabilistic Inference With Global and Local Regularizations
abstract
In this letter, we propose a probabilistic method for the feature matching of remote sensing images which undergo nonrigid transformations. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose nonparametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the expectation-maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real remote sensing images demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, particularly in case of severe outliers.
Huabing Zhou, Jiayi Ma 0001, Changcai Yang, Renfeng Liu, Ji Zhao 0001
IEEE Geosci. Remote. Sens. Lett.1
2015 Robust Feature Matching for Remote Sensing Image Registration via Locally Linear Transforming
abstract
Feature matching, which refers to establishing reliable correspondence between two sets of features (particularly point features), is a critical prerequisite in feature-based registration. In this paper, we propose a flexible and general algorithm, which is called locally linear transforming (LLT), for both rigid and nonrigid feature matching of remote sensing images. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. We formulate this as a maximum-likelihood estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. To ensure the well-posedness of the problem, we develop a local geometrical constraint that can preserve local structures among neighboring feature points, and it is also robust to a large number of outliers. The problem is solved by using the expectation-maximization algorithm (EM), and the closed-form solutions of both rigid and nonrigid transformations are derived in the maximization step. In the nonrigid case, we model the transformation between images in a reproducing kernel Hilbert space (RKHS), and a sparse approximation is applied to the transformation that reduces the method computation complexity to linearithmic. Extensive experiments on real remote sensing images demonstrate accurate results of LLT, which outperforms current state-of-the-art methods, particularly in the case of severe outliers (even up to 80%).
Jiayi Ma 0001, Huabing Zhou, Ji Zhao 0001, Yuan Gao 0015, Junjun Jiang, Jinwen Tian
IEEE Trans. Geosci. Remote. Sens.2