VLDB 2026 Research / reviewers in the wild / expert
Yang Yang 0032
dblp:48/450-32
· DBLP profile ↗
43ranked-venue papers
3as first author
30since 2021 · last 2026
0000-0002-4607-0501ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A generalizable rapid multi-target capture framework based on a single conventional pan-tilt-zoom camera
Fangxu Jiao, Deyun Ren, Yang Yang 0032, Shan Zhao 0008, Zidu Yin |
Expert Syst. Appl. | 6 |
| 2026 | Geo-Mamba: Geometry-informed state-space learning of functional brain organization
Yuwei Cao, Tingting Dan, Yang Yang 0032, Guorong Wu 0001 |
Medical Image Anal. | 3 |
| 2026 | A Real-Time Automation Framework for Multi-Target Capture in Large Scenes With Predictive Scheduling and Zoom ControlabstractThis paper presents a Multi-Target Dynamic Capture Framework for intelligent and real-time visual monitoring in large-scale dynamic environments using a single pan-tilt-zoom (PTZ) camera. The proposed framework integrates target detection, trajectory tracking, motion forecasting, capture scheduling, and adaptive zoom control into a unified automation pipeline. A computationally efficient trajectory prediction module estimates future target locations based on historical motion states and normalized velocities. To address control latency and spatial contention, a scheduling strategy is developed to assign capture timestamps by jointly modeling PTZ adjustment time and global target distribution. The system achieves an average throughput of 12 FPS on 2560×1440 RTSP streams, ensuring real-time responsiveness under constrained camera resources. Experimental evaluations show that MTDCF significantly enhances capture success rate and image resolution, outperforming baseline methods in both prediction accuracy and scheduling effectiveness. This work offers a scalable and control-aware solution for automated multi-target acquisition in resource-limited visual surveillance systems. Fangxu Jiao, Shan Zhao 0008, Yang Yang 0032, Zidu Yin |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | An end-to-end semantic-guided infrared and visible registration-fusion network for advanced visual tasks
Meng Sang, Housheng Xie, Jingrui Meng, Yukuan Zhang, Junhui Qiu, Shan Zhao 0008, Yang Yang 0032 |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Ensemble fractional fuzzy dispersion entropy: A low bias approach for data analysis
Chengjiang Zhou, Longkun He, Xuanyu Liao, Jie Li 0023, Yang Yang 0032 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Advancing multi-object tracking through occlusion-awareness and trajectory optimization
Yukuan Zhang, Yunhua Jia, Yang Yang 0032 |
Knowl. Based Syst. | 4 |
| 2025 | ISGLNet: Infrared Small Target Detection With Intrinsic Sensitivity and Guided LearningabstractInfrared small target detection (IRSTD) plays a critical role in both civilian and military applications, yet it still faces inherent challenges stemming from faint targets, complex noise interference, and difficulties in preserving shape integrity. Despite significant progress in detecting general small targets, existing methods often struggle to balance detection accuracy and false alarms due to limited sensitivity to low-intensity signals, inaccurate perception of confusing noise, and inadequate edge refinement. To break this dilemma, we propose ISGLNet, which centers on a U-shaped architecture specifically tailored to preserve salient target responses, along with a guided learning strategy that progressively enhances target–noise distinction while refining boundary details. Specifically, we introduce the Context-aware Local-Global Module (CLGM) as the cornerstone of the model, which incorporates multi-branch large receptive fields and multi-dimensional adaptive attention mechanisms, effectively capturing rich contexts while preserving critical target information. This ensures reliable feature modeling throughout the extraction and fusion process. Furthermore, the Multi-frequency Perception Module (MFPM) and the Edge Refinement Module (ERM) replace conventional skip connections to refine semantic patterns through guidance. Among these, the MFPM operates in the deeper layers, primarily identifying discriminative clues by evaluating and dynamically selecting multi-frequency information to amplify the distinction between targets and complex noise. The ERM further works in the shallower layers with a progressive strategy to refine uncertain target boundaries, enabling precise segmentation of fine-grained target shapes. Extensive experiments on multiple public datasets demonstrate that ISGLNet achieves superior performance in both detection and segmentation accuracy. The code is available at https://github.com/fuqingzhang/ISGLNet. Fuqing Zhang, Anning Pan, Jing Yang 0060, Shen Deng, Shan Zhao 0008, Chengjiang Zhou, Yang Yang 0032 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Revealing Cortical Spreading Pathway of Neuropathological Events by Neural Optimal Mass TransportabstractPositron Emission Tomography (PET) is essential for understanding the pathophysiological mechanisms underlying neurodegenerative diseases like Alzheimer's disease (AD). However, existing approaches primarily focus on stereotypical patterns of pathology burden, lacking the ability to elucidate the underlying propagation mechanisms by which pathologies spread throughout the brain over time. Given that many neurodegenerative diseases exhibit prion-like pathology spread, it is essential to uncover the spot-to-spot flow field between consecutive PET snapshots. To address this, we reformulate the problem of identifying latent cortical propagation pathways of neuropathological burden within the well-established framework of optimal mass transport (OMT). In this formulation, the dynamic spreading of pathology across longitudinal PET scans is inherently constrained by the geometry of the brain cortex. To solve this problem, we introduce a variational framework that characterizes the dynamical system of pathology propagation in the brain, ultimately reducing to a Wasserstein geodesic between two density distributions of pathology accumulation. Furthermore, we hypothesize that a well-characterized mechanism of pathology propagation will enable the prediction of future pathology accumulation at the individual level, paving the way for personalized disease progression modeling. Building on the principles of physics-informed deep models, we derive the governing equation of the underlying OMT model and introduce an explainable, generative adversarial network-inspired framework. Our approach (1) parameterizes population-level OMT dynamics through a flow adjuster and (2) predicts the spreading flow in unseen subjects using a trained flow driver. We validate the accuracy of our model on publicly available datasets, demonstrating its effectiveness in forecasting future pathology accumulation. Since our deep model adheres to the second law of thermodynamics, we further explore the propagation dynamics of tau aggregates throughout the progression of AD. In contrast to traditional methods, our physics-informed approach enhances both accuracy and interpretability, demonstrating its potential to reveal novel neurobiological mechanisms driving disease progression. Tingting Dan, Yanquan Huang, Yang Yang 0032, Guorong Wu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | VRFF: Video Registration and Fusion FrameworkabstractThrough the process of infrared and visible image registration and fusion, we can generate a composite image that encapsulates the features of both infrared and visible images, thereby enhancing decision-making capabilities in subsequent advanced visual tasks. Despite significant strides made in image registration and fusion research, a noticeable gap remains in the field of video registration and fusion. The direct application of image registration and fusion techniques to videos often results in a flickering effect. To address this issue, we redefine the workflow of the image registration and fusion framework (IRFF), which includes stages of matching key points (MKPs) extraction, image alignment, and image fusion. Based on this workflow, we propose a new video registration and fusion framework (VRFF). In the image alignment stage, we propose the integrated previous frames (IPF) strategy and employ the Moment algorithm, both of which are based on temporal relationships. Additionally, we adopt a new strategy to retrain the MKPs extraction network and redesign the image fusion network for enhanced performance. Experimental results demonstrate that the VRFF exhibits superior performance on video streams. Furthermore, we also explore the effectiveness of fused images in advanced visual applications. Meng Sang, Housheng Xie, Yang Yang 0032 |
IJCNN | 3 |
| 2024 | Data-driven hierarchical learning approach for multi-point servo control of Pan-Tilt-Zoom cameras
Xiangshuai Zhai, Zidu Yin, Yang Yang 0032 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | OFACD: An end-to-end change detection network for small UAVs remote sensing with viewpoint differences
Yaxin Dong, Kai Yan 0002, Shen Deng, Yang Yang 0032 |
Image Vis. Comput. | 6 |
| 2024 | AIPT: Adaptive information perception for online multi-object tracking
Yukuan Zhang, Housheng Xie, Yunhua Jia, Jingrui Meng, Meng Sang, Junhui Qiu, Shan Zhao 0008, Yang Yang 0032 |
Knowl. Based Syst. | 8 |
| 2024 | RCVS: A Unified Registration and Fusion Framework for Video StreamsabstractThe infrared and visible cross-modal registration and fusion can generate more comprehensive representations of object and scene information. Previous frameworks primarily focus on addressing the modality disparities and the impact of preserving diverse modality information on the performance of registration and fusion tasks among different static image pairs. However, these frameworks overlook the practical deployment on real-world devices, particularly in the context of video streams. Consequently, the resulting video streams often suffer from instability in registration and fusion, characterized by fusion artifacts and inter-frame jitter. In light of these considerations, this paper proposes a unified registration and fusion scheme for video streams, termed RCVS. It utilizes a robust matcher and spatial-temporal calibration module to achieve stable registration of video sequences. Subsequently, RCVS combines a fast lightweight fusion network to provide stable fusion video streams for infrared and visible imaging. Additionally, we collect a infrared and visible video dataset HDO, which comprises high-quality infrared and visible video data captured across diverse scenes. Our RCVS exhibits superior performance in video stream registration and fusion tasks, adapting well to real-world demands. Overall, our proposed framework and HDO dataset offer the first effective and comprehensive benchmark in this field, solving stability and real-time challenges in infrared and visible video stream fusion while assessing different solution performances to foster development in this area. Housheng Xie, Meng Sang, Yukuan Zhang, Yang Yang 0032, Shan Zhao 0008, Jianbo Zhong |
IEEE Trans. Multim. | 4 |
| 2023 | Recent progress in image denoising: A training strategy perspectiveabstractAbstract Image denoising is one of the hottest topics in image restoration area, it has achieved great progress both in terms of quantity and quality in recent years, especially after the wide and intensive application of deep neural networks. In many deep learning based image denoising models, the performance can greatly benefit from the prepared clean/noisy image pairs used for model training, however, it also limits the application of these models in real denoising scenes. Therefore, more and more researchers tend to develop models that can be learned without image pairs, namely the denoising models that can be well generalised in real‐world denoising tasks. This motivates to make a survey on the recent development of image denoising methods. In this paper, the typical denoising methods from the perspective of model training are reviewed, the reviewed methods are categorised into four classes: the models need clean/noisy image pairs to train, the models trained on multiple noisy images, the models can be learned from a single noisy image, and the visual transformer based models. The denoising results of different denoisers were compared on some public datasets to discover the performance and advantages. The challenges and future directions in image denoising area are also discussed. Wencong Wu, Mingfei Chen, Yungang Zhang, Yang Yang 0032 |
IET Image Process. | 5 |
| 2023 | StateNet: Deep State Learning for Robust Feature Matching of Remote Sensing ImagesabstractSeeking good correspondences between two images is a fundamental and challenging problem in the remote sensing (RS) community, and it is a critical prerequisite in a wide range of feature-based visual tasks. In this article, we propose a flexible and general deep state learning network for both rigid and nonrigid feature matching, which provides a mechanism to change the state of matches into latent canonical forms, thereby weakening the degree of randomness in matching patterns. Different from the current conventional strategies (i.e., imposing a global geometric constraint or designing additional handcrafted descriptor), the proposed StateNet is designed to perform alternating two steps: 1) recalibrates matchwise feature responses in the spatial domain and 2) leverages the spatially local correlation across two sets of feature points for transformation update. For this purpose, our network contains two novel operations: adaptive dual-aggregation convolution (ADAConv) and point rendering layer (PRL). These two operations are differentiable, so our network can be inserted into the existing classification architecture to reduce the cost of establishing reliable correspondences. To demonstrate the robustness and universality of our approach, extensive experiments on various real image pairs for feature matching are conducted. Experiments reveal the superiority of our StateNet significantly over the state-of-the-art alternatives. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Yujing Rao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Robust probability model based on variational Bayes for point set registration
Hualong Cao, Yang Yang 0032, Ziyun Zhou |
Knowl. Based Syst. | 4 |
| 2022 | VLSG-SANet: A feature matching algorithm for remote sensing image registration
Linjie Xing, Jiaxuan Chen 0002, Shuang Chen 0008, Haicheng Bai, Lin Xing, Chengjiang Zhou, Yang Yang 0032 |
Knowl. Based Syst. | 8 |
| 2022 | Robust Feature Matching via Hierarchical Local Structure VisualizationabstractFeature matching, which refers to seeking good correspondence between two feature point sets, is a critical prerequisite in many applications of remote sensing and photogrammetry. This work can be viewed as an extension of LSV-ANet. Traditional local structure visualization (LSV) descriptor is very sensitive to outliers existing in the small region around a feature point, which limits the ability of LSV-ANet to recognize fine-grained patterns and its generalizability for complex matching scenes. Thus, in this letter, we propose a hierarchical local structure visualization (HLSV) to solve this issue. Specifically, local structures of feature points are first decoupled by hierarchical tensor and then reconstructed in a learning fashion, which explicitly permits us to learn a nonmutually exclusive relationship for enhancement structure manipulation. In addition, in order to further improve the network generalization ability, we design a permutation-invariant network layer to ensure that the output of model is invariant to the input order. To demonstrate the robustness of the HLSV, extensive experiments on various real remote sensing image pairs for feature matching are conducted. The experiment results reveal that HLSV is superior to the current eight state-of-the-art alternatives. Jiaxuan Chen 0002, Shuang Chen 0008, Yang Yang 0032, Haicheng Bai |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Real-Time Garbage Object Detection With Data Augmentation and Feature Fusion Using SUAV Low-Altitude Remote Sensing ImagesabstractRecently, a number of nature reserves have been shut down because of serious pollution from tourist garbage. Garbage monitoring in high-altitude natural reserves using small unmanned aerial vehicle (SUAV) remote sensing is an important and urgent need for environmental protection. In order to help cleaners to eliminate garbage more conveniently and quickly, a novel approach is proposed to detect scattered garbage regions in real time using low-altitude remote sensing videos captured by SUAVs. First, the high-resolution, low-altitude, multitemporal remote sensing images and videos containing scattered garbage were collected through SUAV and then proposed a data augmentation method to expand the training samples. Second, the Yolov4 detection network was used to classify the scattered garbage regions. Finally, the location of the object was roughly calculated according to the altitude, flight direction, global positioning system, and digital elevation model (DEM). Then, the garbage object was marked on the video, while the object location was marked on the map. Experimental results show that the proposed method achieves a mean accuracy of 91.34% and provides better performances on the real data set compared with state-of-the-art methods. Hao Li 0093, Quanjing Li, Yang Yang 0032, Kun Yang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Dual-Features Student-t Distribution Mixture Model Based Remote Sensing Image RegistrationabstractIn the work, we present a novel multiviewpoint and multitemporal remote sensing image registration method based on a dual-features Student-$t$distribution mixture model (DSMM) under a variational Bayesian (VB) framework. The main contributions of the work are: 1) guided image filter (GIF) is adopted to smooth edges and strengthen ridges of images for heightening characteristics of feature point; 2) a Student-$t$distribution mixture model (SMM) based DSMM designs a global and local descriptor to estimate correspondences from local to global scale; and 3) local structure constraints are designed to preserve relationships of neighbors of points and the scale of neighborhood structure of points to constrain transformation. The experimental results demonstrate the better performance of our DSMM against five state-of-the-art methods. Li Liang 0006, Qiqi He, Hualong Cao, Yang Yang 0032, Mina Han |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Multi-Source Remote Sensing Intelligent Characterization Technique-Based Disaster Regions Detection in High-Altitude Mountain Forest AreasabstractNatural disasters frequently have caused huge impact on life and property losses, in Southwest China. To provide assistance for disaster relief, areas damaged in natural disasters is quickly located by utilizing satellite remote sensing images based-deep learning object detection technology. However, the current detection technology, for the detection of damaged objects discretely in the disaster area, has some challenges, such as partial missing of multi-source images and extremely sparse targets with weak features or occlusion at large scales. Furthermore, we propose an object detection network based on dynamic extraction of multi-source images features to solve above problems. To train our proposed network, we collect multi-source remote sensing images before and after the disaster. Finally, it is verified that when the detection error rate is less than 5%, the accuracy of the detection model reaches more than 85%. Hualong Cao, Kai Yan 0002, Haicheng Bai, Yang Yang 0032, Lin Xing, Chengjiang Zhou |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | An Estimation of Distribution Algorithm Based on Variational Bayesian for Point-Set RegistrationabstractPoint-set registration is widely used in computer vision and pattern recognition. However, it has become a challenging problem since the current registration algorithms suffer from the complexities of the point-set distributions. To solve this problem, we propose a robust registration algorithm based on the estimation of distribution algorithm (EDA) to optimize the complex distributions from a global search mechanism. We propose an EDA probability model based on the asymmetric generalized Gaussian mixture model, which describes the area in the solution space as comprehensively as possible and constructs a probability model of complex distribution points, especially for missing and outliers. We propose a transformation and a Gaussian evolution strategy in the selection mechanism of EDA to process the deformation, rotation, and denoising of selected dominant individuals. Considering the complexity of the model, we choose to optimize from the perspective of variational Bayesian, and introduce a prior probability distribution through local variation to reinforce the convergence of the algorithm in dealing with complex point sets. In addition, a local search mechanism based on the simulated annealing algorithm is added to realize the coarse-to-fine registration. Experimental results show that our method has the best robustness compared with the state-of-the-art registration algorithms. Hualong Cao, Qiqi He, Zenghui Xiong, Yang Yang 0032 |
IEEE Trans. Evol. Comput. | 6 |
| 2022 | LSV-ANet: Deep Learning on Local Structure Visualization for Feature MatchingabstractFeature matching is a fundamental and important task in many applications of remote sensing and photogrammetry. Remote sensing images often involve complex spatial relationships due to the ground relief variations and imaging viewpoint changes. Therefore, using a pre-defined geometrical model will probably lead to inferior matching accuracy. In order to find good correspondences, we propose a simple yet efficient deep learning network, which we term the “local structure visualization-attention” network (LSV-ANet). Our main aim is to transform outlier detection into a dynamic visual similarity evaluation. Specifically, we first map the local spatial distribution into a regular grid as descriptor LSV, and then customized a spatial SCale Attention (SCA) module and a spatial STructure Attention (STA) module, which explicitly allows structure manipulation and scale selection of LSV within the network. Finally, the embedded SCA and STA deduce optimal LSV for solving feature matching task by training the LSV-ANet end-to-end. In order to demonstrate the robustness and universality of our LSV-ANet, extensive experiments on various real image pairs for general feature matching are conducted and compared against eight state-of-the-art methods. The experiment results demonstrate the superiority of our method over state of the art. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yang Yang 0032, Linjie Xing, Yujing Rao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | IGS-Net: Seeking Good Correspondences via Interactive Generative Structure LearningabstractFeature matching, which aims to seek good correspondences from an image pair of the same or similar scene, is one of the important studies on digital remote sensing (RS) image processing. However, alongside the common degradation problems, such as geometric distortion, RS images also often face nonlinear radiation distortions, thereby posing more complex matching patterns. To reduce the cost of establishing reliable correspondences, this article proposes a simple, but effective end-to-end hierarchical learning framework, termed interactive generative structure learning network (IGS-Net). The key thinking of our approach is to offer a structure self-generate learning mechanism, called interactive generative structure learning (IGSL) block, for modeling the local context information of potential correspondences. Specifically, IGSL contains two novel operations: adaptive structure-aware representation (ASR) and physical constraint embedding. Besides, we introduce a coarse-to-fine geometry estimation pipeline aligning two sets of feature points to weaken the degree of randomness in matching patterns, thus improving the generalization ability of representation learning. Overall, this differentiable representation learning architecture can be inserted into existing classification models easily for robust outlier detection and removal. In order to demonstrate that our IGS-Net can boost the baselines, we intensively experiment on both single modality and multimodal RS image datasets. The large amounts of experiment results reveal that the matching performances of IGS-Net are significantly improved over eight state-of-the-art competitors. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yujing Rao, Chengjiang Zhou, Yang Yang 0032 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | A Hierarchical Consensus Attention Network for Feature Matching of Remote Sensing ImagesabstractFeature matching, referring to establishing high reliable correspondences between two or more scenes with overlapping regions, is of extremely significance to various remote sensing (RS) tasks, such as panorama mosaic and change detection. In this work, we propose an end-to-end deep network for mismatch removal, named hierarchical consensus attention network (HCA-Net), which is one of the critical steps in matching pipeline. Unlike existing practices, our HCA-Net does not rely on global geometric constraints and handcrafted structural representations. The key principle of the proposed HCA-Net is to adaptively enhance neighborhood consensus before evaluating correspondence. To this end, we design a consensus attention mechanism to regularize sparse matches directly. More specifically, consensus attention consists of two novel operations: an encoder–decoder module for calculating compatibility scores and a context-based density representation module. Such attention mechanism can be easily plugged into the existing inlier/outlier classification model in a stacked way to reject outliers. We also propose a hierarchical global-aware network for further improving the accuracy of outlier detection. We compare the proposed HCA-Net with seven state-of-the-art algorithms on several datasets (including various RS images), and the results reveal that our method significantly outperforms the other competitors. Shuang Chen 0008, Jiaxuan Chen 0002, Yujing Rao, Xiaoxian Chen, Haicheng Bai, Lin Xing, Chengjiang Zhou, Yang Yang 0032 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2022 | Learning Relaxed Neighborhood Consistency for Feature MatchingabstractFeature matching is a critical prerequisite in many applications of remote sensing, and its aim is to establish reliable correspondences between two sets of features. Existing attempts typically involve estimating the underlying image transformations to remove false matches in putative matches. However, the image transformation could vary with different application scenarios, which means that using a predefined geometrical model may lead to inferior matching accuracy, especially if the image transformation is nonrigid. This article casts the mismatch removal into a neighborhood consistency evaluation problem under a customized learning framework. With only seven training image pairs involving approximately 8000 putative matches, our method can handle different types of images or transformation models (affine, homography, piecewise-linear transformation, and others). Extensive experiments on feature matching and image registration are conducted to demonstrate the superiority of our method over the eight state-of-the-art competitors. Shuang Chen 0008, Jiaxuan Chen 0002, Zenghui Xiong, Linjie Xing, Yang Yang 0032, Kai Yan 0002, Hao Li 0093 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | CSR-Net: Learning Adaptive Context Structure Representation for Robust Feature CorrespondenceabstractFeature matching, which refers to identifying and then corresponding the same or similar visual pattern from two or more images, is a key technique in any image processing task that requires establishing good correspondences between images. Given potential correspondences (matches) in two scenes, a novel whole-part deep learning framework, termed as Context Structure Representation Network (CSR-Net), is designed to infer the probabilities of arbitrary correspondences being inliers. Traditional approaches commonly build the local relation between correspondences by manually engineered criteria. Different from existing attempts, the main idea of our work is to learn explicitly neighborhood structure of each correspondence, allowing us to formulate the matching problem into a dynamic local structure consensus evaluation in an end-to-end fashion. For this purpose, we propose a permutation-invariant STructure Representation (STR) learning module, which can easily merge different types of networks into a unified architecture to deal with sparse matches directly. By the collaborative use of STR, we introduce a Context-Aware Attention (CAA) mechanism to adaptively re-calibrate structure features via a rotation-invariant context aware encoding and simple feature gating, thus arising the ability of fine-grained patterns recognition. Moreover, to further weaken the cost of establishing reliable correspondences, the CSR-Net is formulated as whole-part consensus learning, where the aim of whole level is compensating rigid transformations. In order to demonstrate our CSR-Net can effectively boost the baselines, we intensively experiment on image matching and other visual tasks. The results of the experiment confirm that the matching performances of CSR-Net have significantly improved over nine state-of-the-art competitors. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yuan Dai, Yang Yang 0032 |
IEEE Trans. Image Process. | 5 |
| 2021 | Progressive structure network-based multiscale feature fusion for object detection in real-time application
Lvjiyuan Jiang, Hao Li 0093, Kai Yan 0002, Yang Yang 0032, Yungang Zhang, Lianliu Qiao, Cuilian Fu |
Eng. Appl. Artif. Intell. | 6 |
| 2021 | Robust Stepwise Correspondence Refinement for Low-Altitude Remote Sensing Image RegistrationabstractLow-altitude aerial photography using small unmanned aerial vehicles (UAVs) often exists image viewpoint changes as human, nature, and equipment factors. The various viewpoint changes bring a series of problems to the subsequent applications of low-altitude remote sensing images. In this letter, a robust neighborhood structure-invariant descriptor is proposed to solve the scaling, rotation, deformation, and their mixture problems existing in the images. To estimate a robust correspondence under severe mismatches, a dynamic strategy is designed. Using the designed descriptor, a classifier is trained to distinguish between true matches and mismatches. Our method compares with six state-of-the-art methods on 55 small UAV images and performs a feature-invariant capability in horizontal/vertical rotation, scaling, mixture, and extreme situations, and gives the best performance in feature matching and image registration. Xiaoying Gong, Yang Yang 0032 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | A Multilayer Fusion Network With Rotation- Invariant and Dynamic Feature Representation for Multiview Low-Altitude Image RegistrationabstractDue to human and natural factors, when the small unmanned aerial vehicles (UAVs) are monitoring the ground, multiview transformation problems such as image distortion and low overlap will occur, which will inhibit the accuracy of low-altitude image registration and limit the subsequent application. In this letter, we propose a mismatch removal method based on the Siamese architecture to solve the issues of multiview images. A dynamic neighbor-guided patch representation is designed to enhance the representation of each feature point. Meanwhile, a multilayer fusion is used to obtain more comprehensive information on feature points, and whether a pair of points correspond depends on the similarity of its descriptors. The network is trained by adding a rotation-invariant layer to solve the inevitable rotation and image distortion in multiview scenarios. The experimental results prove that our method can deal with the scenarios of the horizontal rotation, vertical rotation, mixture, scaling, and extreme, and is better than the other five state-of-the-art methods in most scenarios. Xiaoying Gong, Yang Yang 0032 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Image registration using two-layer cascade reciprocal pipeline and context-aware dissimilarity measure
Li Liang 0006, Xuying Hao, Yang Yang 0032, Kun Yang 0007, Lijia Liang, Qinglu Yang |
Neurocomputing | 4 |
| 2020 | Remote Sensing Image Registration With Adjustable Threshold and Variational Mixture TransformationabstractThis letter proposes a remote sensing (RS) image registration method based on adjustable threshold and variational mixture transformation (VMT). The main contributions of our method are: 1) an adjustable threshold strategy proposed to guarantee sufficient inlier pairs for image spatial transformation and 2) a VMT, which achieves a coarse-to-fine process, consisting of three main steps as follows: a) a rigid transformation is employed to achieve approximate correspondence; b) a guided Gaussian mixture model is proposed to better distinguish outliers; and c) a nonrigid transformation is applied to achieve an accurate registration. We test the performance of our algorithm in feature matching and RS image registration and compare it with the six state-of-the-art methods. Our method shows the best performances in most scenarios. Xinke Ma, Xuying Hao, Yang Yang 0032, Kun Yang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Point set registration with mixture framework and variational inference
Xinke Ma, Shijin Xu, Qinglu Yang, Yang Yang 0032, Kun Yang 0007, Sim Heng Ong |
Pattern Recognit. | 5 |
| 2020 | Adaptive Hierarchical Probabilistic Model Using Structured Variational Inference for Point Set RegistrationabstractPoint set registration plays an important role in computer vision and pattern recognition. In this article, we propose an adaptive hierarchical probabilistic model (HPM) under a variational Bayesian (VB) framework for point set registration problem. The main contributions of this article are given as follows. First, a dynamic putative inlier estimation strategy is proposed through the hesitant fuzzy Einstein weighted averaging based membership calculation and component estimation using symmetric cross entropy. Second, a student-t mixture model based HPM is designed to solve outlier and occlusion problems during registration. Third, a VB-based transformation updating is proposed to construct a robust and adjustable transformation for effectively fitting target point set while further resisting outliers. The performances of the proposed method in point set and image registrations against 11 state-of-the-art methods are evaluated, in which our method gives the best performance in most scenarios. Qiqi He, Shijin Xu, Yang Yang 0032, Rui Yu 0004, Yuhe Liu |
IEEE Trans. Fuzzy Syst. | 4 |
| 2020 | A Context-Aware Locality Measure for Inlier Pool Enrichment in Stepwise Image RegistrationabstractWe present a feature-based image registration method, the stepwise image registration (SIR), with a closed-form solution. Our SIR creates an inlier pool and a candidate pool as the initialization, and then gradually enriches the inlier pool and refines the transformation. In each step, the enriched correspondence exclusively tunes the transformation coefficient within the confirmed inlier pairs, instead of updating the mapping using the complete putative set. In turn, the refined transformation prunes inconsistent mismatches to alleviate the incoming matching ambiguity. The context-aware locality measure (CALM) is designed for dissimilarity measure. The capability of the CALM can be enhanced by the progressive inlier pool enrichment. Finally, a retrieval process is performed based on the finest CALM and alignment, by which the inlier pool is maximized. Extensive experiments of enrichment evaluation, feature matching, image registration, and image retrieval demonstrate the favorable performance of our SIR against state-of-the-art methods. The code and datasets are available at https://github.com/sucv/SIR. Su Zhang 0004, Xuying Hao, Yang Yang 0032, Cuntai Guan |
IEEE Trans. Image Process. | 4 |
| 2019 | Robust Variational Bayesian Point Set RegistrationabstractIn this work, we propose a hierarchical Bayesian network based point set registration method to solve missing correspondences and various massive outliers. We construct this network first using the finite Student s t latent mixture model (TLMM), in which distributions of latent variables are estimated by a tree-structured variational inference (VI) so that to obtain a tighter lower bound under the Bayesian framework. We then divide the TLMM into two different mixtures with isotropic and anisotropic covariances for correspondences recovering and outliers identification, respectively. Finally, the parameters of mixing proportion and covariances are both taken as latent variables, which benefits explaining of missing correspondences and heteroscedastic outliers. In addition, a cooling schedule is adopted to anneal prior on covariances and scale variables within designed two phases of transformation, it anneal priors on global and local variables to perform a coarse-to- fine registration. In experiments, our method outperforms five state-of-the-art methods in synthetic point set and realistic imaging registrations. Xinke Ma, Li Liang 0006, Liu Yuhe, Shijin Xu, Sim Heng Ong, Yang Yang 0032 |
ICCV | 7 |
| 2019 | Non-Rigid Image Registration With Dynamic Gaussian Component Density and Space Curvature PreservationabstractImage registration plays an important role in military and civilian applications, such as natural disaster damage assessment, environmental monitoring, ground change detection and military damage assessment, etc. This work presents a new feature-based non-rigid image registration method. The main contributions of this work are: (i) a dynamic Gaussian component density is designed to better exploit available potential image information and provide sufficient inlier pairs for image transformation; (ii) a spatial structure preservation, which consists of an image transformation space curvature preservation and a local spatial structure constrain, is proposed to constrain the image transforming cost as well as the local structure of feature points during feature point set registration. The performances of the proposed method in multi-spectral natural images, lowaltitude aerial images and medical images against four types of nine state-of-the-art methods are tested where our method shows the best performances in most scenarios. Zhuoqian Yang, Yang Yang 0032, Kun Yang 0007, Ziquan Wei |
IEEE Trans. Image Process. | 2 |
| 2018 | Nonrigid Image Registration for Low-Altitude SUAV Images With Large Viewpoint ChangesabstractLow-altitude aerial photography using small unmanned aerial vehicles (SUAVs) with large viewpoint changes causes nonrigid distortions and low overlap ratios. We present a nonrigid feature-based low-altitude SUAV image-registration method. The key idea of our method is to maintain a high matching ratio on inliers while taking advantage of outliers for varying the warping grids. Thus, accurate image transformation over the overlapping areas as well as a good approximation of the real transformation over the nonoverlapping areas can be obtained. Experiments on feature matching and image registration are performed using 42 pairs of SUAV images. Our method exhibited a favorable performance as compared with four state-of-the-art methods, even with up to 80% outliers. Su Zhang 0004, Kun Yang 0007, Yang Yang 0032 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Non-rigid point set registration using dual-feature finite mixture model and global-local structural preservation
Su Zhang 0004, Kun Yang 0007, Yang Yang 0032, Ziquan Wei |
Pattern Recognit. | 3 |
| 2017 | Point Set Registration with Global-Local Correspondence and Transformation EstimationabstractWe present a new point set registration method with global-local correspondence and transformation estimation (GL-CATE). The geometric structures of point sets are exploited by combining the global feature, the point-to-point Euclidean distance, with the local feature, the shape distance (SD) which is based on the histograms generated by an elliptical Gaussian soft count strategy. By using a bidirectional deterministic annealing scheme to directly control the searching ranges of the two features, the mixture-feature Gaussian mixture model (MGMM) is constructed to recover the correspondences of point sets. A new vector based structure constraint term is formulated to regularize the transformation. The accuracy of transformation updating is improved by constraining spatial structure at both global and local scales. An annealing scheme is applied to progressively decrease the strength of the regularization and to achieve the maximum overlap. Both of the aforementioned processes are incorporated in the EM algorithm, a unified optimization framework. We test the performances of our GL-CATE in contour registration, sequence images, real images, medical images, fingerprint images and remote sensing images, and compare with eight state-of-the-art methods where our GL-CATE shows favorable performances in most scenarios. Su Zhang 0004, Yang Yang 0032, Kun Yang 0007, Sim Heng Ong |
ICCV | 2 |
| 2015 | A robust global and local mixture distance based non-rigid point set registration
Yang Yang 0032, Sim Heng Ong, Kelvin Weng Chiong Foong |
Pattern Recognit. | 1 |
| 2012 | Image-based estimation of biomechanical relationship between masticatory muscle activities and mandibular movementabstractCurrent techniques for investigating the functional roles of masticatory muscles are not suited to explaining subject-specific biomechanical relationship between mandibular movements and masticatory muscle activities. The aims of this study were to estimate the muscle tensions of subject-specific masticatory muscles in dental occlusion through the three-dimensional morphologic changes (3DMCs) of the muscles, and to explain the subject-specific biomechanical relationship between the mandibular movement and the muscle tensions. One healthy adult subject underwent magnetic resonance (MR) scans of the head at mandibular rest position (M0) and maximum intercuspation (M1). Based on the two sets of MR images, the mandibular movement was measured by the position changes of the mental protuberance of the mandible and the muscle tension from M0 to M1 for each masticatory muscle was estimated by its 3DMCs. The results showed the subject-specific biomechanical relationship between the mandibular movement and the muscle tensions, and the mandibular movement could be explained by these related muscle tensions anatomically and functionally. Yang Yang 0032, Kelvin Weng Chiong Foong, Sim Heng Ong, Masakazu Yagi, Kenji Takada |
BIBE | 1 |
| 2012 | An image-based method for quantification of lateral pterygoid muscle deformationabstractThe aims of this study were to present a method quantifying and visualizing the deformation of subject-specific lateral pterygoid muscles (LPM) during a simulated jaw-opening movement. A normal adult male subject underwent magnetic resonance (MR) scans of the head at three mandibular positions: mandibular rest (M0), medium jaw-opened (M1), and maximum jaw-opened (M2) positions. The 3D models of the LPM were reconstructed from the three sets of MR images. The deformations of each muscle in the two cases (M0->M1 and M1->M2) were quantified in terms of the displacements of region correspondences between the muscle models before and after the mandibular position changed. The 3D models of the subject-specific LPM were reconstructed, and the directions and magnitudes of deformations of each muscle in the two cases were accurately quantified and visualized in the three anatomic planes. The functional activities along the entire body and at specific compartments of subject-specific LPM were quantified and visualised using the quantified 3D deformations of the LPM as a new descriptor. The presented method defined the deformations of the subject-specific LPM, and revealed the anatomic architectural and biomechanical characteristics of the subject-specific LPM appropriately and meaningfully in the simulated jaw-opening movement. Yang Yang 0032, Kelvin Weng Chiong Foong, Sim Heng Ong, Masakazu Yagi, Kenji Takada |
BIBE | 1 |