VLDB 2026 Research / reviewers in the wild / expert
Fan Fan 0001
dblp:20/4226-1
· DBLP profile ↗
45ranked-venue papers
2as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 14 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LACT-Fusion: Linear attention-Guided cross-Modal learning for infrared and visible image fusion
Zhao Cai, Yong Ma 0001, Qi Peng 0001, Jun Huang 0008, Fan Fan 0001 |
Knowl. Based Syst. | 7 |
| 2026 | Unleashing the potential of Mamba: A novel approach for low-light image enhancement
Yong Ma 0001, Jun Huang 0008, You Du, Fan Fan 0001, Zhiqing Zhao |
Knowl. Based Syst. | 5 |
| 2026 | A general outlier filtering method for feature matching via local motion consistency based Markov network
Fan Fan 0001, Songchu Deng, Yong Ma 0001, Jun Huang 0008 |
Pattern Recognit. | 1 |
| 2026 | Light field image blind super-resolution via degradation representation learning
Kailing Yong, Fan Fan 0001, Jun Huang 0008, You Du, Haonan Tian, Yong Ma 0001 |
Pattern Recognit. | 2 |
| 2025 | Multimodal Image Matching Based on Cross-Modality Completion Pre-trainingabstractThe differences in imaging devices cause multimodal images to have modal differences and geometric distortions, complicating the matching task. Deep learning-based matching methods struggle with multimodal images due to the lack of large annotated multimodal datasets. To address these challenges, we propose XCP-Match based on cross-modality completion pre-training. XCP-Match has two phases. (1) Self-supervised cross-modality completion pre-training based on real multimodal image dataset. We develop a novel pre-training model to learn cross-modal semantic features. The pre-training uses masked image modeling method for cross-modality completion, and introduces an attention-weighted contrastive loss to emphasize matching in overlapping areas. (2) Supervised fine-tuning for multimodal image matching based on the augmented MegaDepth dataset. XCP-Match constructs a complete matching framework to overcome geometric distortions and achieve precise matching. Two-phase training encourages the model to learn deep cross-modal semantic information, improving adaptation to modal differences without needing large annotated datasets. Experiments demonstrate that XCP-Match outperforms existing algorithms on public datasets. Meng Yang 0031, Fan Fan 0001, Jun Huang 0008, Yong Ma 0001, Xiaoguang Mei, Zhanchuan Cai, Jiayi Ma 0001 |
IJCAI | 2 |
| 2025 | CorrNeXt: Making the ConvNet-Style Correspondence Pruner Stronger for Two-View GeometryabstractThe uproar over two-view correspondence pruning stems from the advent of the ConvNet-style paradigm, which showcases intrinsic proficiency in local context aggregation, tackling the context-agnostic deficiency of MLP-based methods fundamentally and delivering impressive pruning capability. To further unlock the potential of such a paradigm, this perspective study revisits its design decisions and introduces CorrNeXt, a cutting-edge ConvNet-style pruner that incorporates multiple simple but effective improvements. Firstly, we explicitly integrate 2D relative spatial knowledge into motion field modeling, arming the interconversion between unordered sparse motion vectors and ordered image-structured ones with positional awareness. Secondly, considering that existing methods struggle with perceiving global context due to limited receptive field of small-kernel convolution, we devise a context-orthogonal aggregation module that decomposes computationally expensive large-kernel depthwise convolution along channel dimension into a small square kernel, two orthogonal band kernels, and an identity mapping, enjoying large receptive field while maintaining efficiency. Thirdly, we deploy a motion field pyramid architecture that obtains and fuses multi-level motion fields, thereby facilitating the handling of the motion field's discontinuities in case of large scene disparity. Ultimately, we propose an elastic inference strategy that allows the model to introspect the confidence of its predictions at each layer, through which CorrNeXt is endowed with the flexibility of adaptively determining inference termination according to the difficulty of each image pair. Thorough experimentation affirms CorrNeXt's remarkable capabilities. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
ACM Multimedia | 3 |
| 2025 | Infrared Small Target Detection via Local-Global Feature FusionabstractDue to the high-luminance (HL) background clutter in infrared (IR) images, the existing IR small target detection methods struggle to achieve a good balance between efficiency and performance. Addressing the issue of HL clutter, which is difficult to suppress, leading to a high false alarm rate, this letter proposes an IR small target detection method based on local-global feature fusion (LGFF). We develop a fast and efficient local feature extraction operator and utilize global rarity to characterize the global feature of small targets, effectively suppressing a significant amount of HL clutter. By integrating local and global features, we achieve further enhancement of the targets and robust suppression of the clutter. Experimental results demonstrate that the proposed method outperforms existing methods in terms of target enhancement, clutter removal, and real-time performance. Yong Ma 0001, Fan Fan 0001, Jun Huang 0008 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Shifting Neighbors Within Temporal Contexts for Slow-Moving Infrared Small Target Detection
Yong Ma 0001, Fan Fan 0001, Jun Huang 0008 |
IEEE Signal Process. Lett. | 3 |
| 2025 | An End-to-End Network for Rotary Motion Deblurring in the Polar Coordinate SystemabstractNon-blind rotary motion deblurring (RMD) aims to restore a latent image from its blurred image. Since the integration path of rotary motion blurring (RMB) is a circle, RMD is modelled as a typical motion deblurring in the polar coordinate system (PCS). However, existing PCS-based methods use hand-designed image priors and are limited by transformation errors, including Cartesian-to-polar transformation (CPT) error and polar-to-Cartesian transformation (PCT) error. In this paper, we analyze the impact of transformation errors on the restored image and propose a novel end-to-end network which introduces a convolutional neural network (CNN) to learn image priors. Specifically, considering the CPT error, we construct a degradation model and solve it in an unrolling way, effectively reducing the ringing artifacts. For the PCT error, we develop a PCT error correction module (PCM) to reconstruct the lost details and textures. Experiments show our method performs against state-of-the-art (SOTA) approaches on synthetic and real-world rotary motion blur datasets by a large margin. The code and model are available athttps://github.com/Jinhui-Qin/RMD_PCS. Jinhui Qin, Yong Ma 0001, Jun Huang 0008, Zhanchuan Cai, Fan Fan 0001, You Du |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | DSTransNet: Dynamic Feature Selection Network With Feature Enhancement and Multiattention for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) has significantly benefited from UNet-based neural models in recent years. However, current methodologies face challenges in achieving optimal compromise between missed detections and false alarms. To overcome this limitation, we rethink the role of each structural component within UNet-based architectures applied for IRSTD. Accordingly, we conceptualize the UNet’s encoder as specializing in feature extraction, the skip connections in feature selection, and the decoder in fusion-based reconstruction. Building upon these conceptualizations, we propose the DSTransNet. Within the feature extraction stage, the edge shape receptive field (ESR) module enhances edge and shape feature extraction and expands the receptive field via multiple convolutional branches, thereby reducing missed detections. At the feature selection stage, the reliable dynamic selection filtering (RDSF) module employs dynamic feature selection, leveraging encoder-based self-attention and decoder-based cross-attention of the Transformer to suppress background features resembling small targets and mitigate false alarms. During the feature fusion-based reconstruction stage, the cross-attention of spaces and channels (CSCE) module emphasizes small target features via spatial and channel cross-attention, reconstructing more accurate multi-scale detection masks. Extensive experiments on the SIRST, NUDT-SIRST, and SIRST-Aug datasets demonstrate that the proposed DSTransNet method outperforms state-of-the-art IRSTD approaches. The code is available at https://github.com/RuiminHuang/DSTransNet. Ruimin Huang, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Toward Robust Infrared Small Target Detection via Frequency and Spatial Feature FusionabstractInfrared small target detection (IRSTD) faces significant challenges due to the small scale and low intensity of targets, which are characterized by extremely sparse features. Most existing methods primarily concentrate on spatial features while neglecting the significant cluttered interference inherent in complex backgrounds. Such an oversight poses substantial challenges in distinguishing targets from background noise, thereby limiting detection performance. Drawing inspiration from the frequency characteristics that differentiate targets from backgrounds in infrared images, we introduce an innovative detection network that leverages high- and low-frequency partitioning and interaction. Specifically, we introduce a patch-wise fast Fourier transform (PFFT), which divides the input image into patches and applies the Fourier transform to each patch. Subsequently, we employ convolutional neural networks (CNNs) for learnable high- and low-frequency partitioning and propose a learnable frequency augmentation module (FAM) to enhance the interfrequency and intrafrequency feature. This methodology effectively harnesses the spatial information inherent in both high and low frequencies to suppress background clutter and accurately extract sparse target features. Furthermore, to further integrate frequency information with spatial information, we propose a frequency spatial fusion module (FSFM) to merge features from frequency and spatial domains. Experimental results show that our method surpasses state-of-the-art techniques on four publicly available datasets. Yong Ma 0001, Fan Fan 0001, Jun Huang 0008, Ruimin Huang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Mutually Reinforcing Learning of Decoupled Degradation and Diffusion Enhancement for Unpaired Low-Light Image LighteningabstractDenoising Diffusion Probabilistic Model (DDPM) has demonstrated exceptional performance in low-light enhancement task. However, the dependency on paired training datas has left the generality of DDPM in low-light enhancement largely untapped. Therefore, this paper proposes a mutually reinforcing learning framework of decoupled degradation and diffusion enhancement, named MRLIE, which leverages style guidance from unpaired low-light images to generate pseudo-image pairs that are consistent with the target domain, thereby optimizing the latter diffusion enhancement network in a supervised manner. During the degradation process, the diffusion loss of fixed enhancement network serves as a evaluation metric for structure consistency and is combined with adversarial style loss to form the optimization objective for degradation network. Such loss design ensures that scene structure information is retained during the degradation process. During the enhancement process, the degradation network with frozen parameters continuously generates pseudo-paired low-/normal-light image pairs as training datas, thus the diffusion enhancement network could be progressively optimized. On the whole, the two processes are interdependent and could achieve cooperative improvement in terms of degradation realism and enhancement quality through iterative optimization. Additionally, we propose the Retinex-based decoupled degradation strategy for simulating the complex degradation in real low-light imaging, which ensures the color correction and noise suppression capabilities of latter diffusion enhancement network. Extensive experiments show that MRLIE can achieve promising results and better generality across various datasets. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | HMFENet: Hierarchical Matching Guided Feature Enhancement Network for Few-Shot RGB-Thermal Urban Scene SegmentationabstractRGB-Thermal semantic segmentation provides reliable support for intelligent traffic perception systems, such as road safety monitoring and autonomous driving perception, by fusing visible and thermal imaging modalities under adverse weather conditions and low-light environments at night. However, the scarcity of multimodal data and the high cost of annotations severely limit the generalization capability of traditional models. To address the core demands of urban scene segmentation, we propose a Hierarchical Matching Guided Feature Enhancement Network (HMFENet) tailored for few-shot learning. It tackles two major challenges: 1) scale diversity of traffic objects (e.g., vehicles and pedestrians) under limited labeled data, which significantly degrades segmentation accuracy; 2) information redundancy across multimodal features, which undermines the enhancement effect of the thermal modality on traffic object segmentation. HMFENet employs a hierarchical dense matching mechanism to establish multi-scale and multi-level feature alignment between query images and support samples. Additionally, it incorporates a mutual information minimization constraint to optimize cross-modal complementarity, thereby enhancing segmentation robustness in complex urban scenes. Experiments on the urban scene dataset, Tokyo Multi-Spectral-$4^{i}$demonstrate that the proposed method achieves state-of-the-art results: an improvement of 5.9% and 9.4% in mean mIoU for critical traffic objects under 1-shot and 5-shot settings, respectively, compared to baseline models. Furthermore, the complementary effect of the thermal modality contributes to a 2.5% improvement under the 1-shot setting. The proposed method provides a feasible solution for deploying multimodal traffic perception systems with low annotation costs. The source code is available athttps://github.com/Zhou-xy99/HMFENet. Yong Ma 0001, Jun Huang 0008, Zhanchuan Cai, Fan Fan 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Universal Infrared Image Nonuniformity Correction via Stripe-Aware Attention NetworkabstractInfrared image nonuniformity correction aims to remove the column-wise stripe noise. Most existing methods just consider stripe noise whereas failing to handle real captured nonuniformity, as directional characteristic of stripe is severely disrupted by random Gaussian noise. Moreover, deep learning-based methods proposed in recent years are blocked by limited receptive field thus cannot accurately distinguish vertical structure and vertical stripes. To address these issues, we propose a universal infrared image nonuniformity correction method based on stripe-aware attention network. We seek to improve the performance of our algorithm by first restoring the damaged stripe directional characteristics, then maximizing the utilization of the prior characteristics. On the one hand, we construct the two-stage framework, in which denoising network is firstly applied to eliminate Gaussian noise and preserve stripes as scene information. As a result, the prior directional characteristics are restored, thereby enhancing the ability of subsequent sub-network to perceive stripe noise. On the other hand, due to the distinct long-range pixel correlations of vertical structures and vertical textures, we introduce a column-wise stripe attention mechanism (CSA) that can capture long-range dependencies of target pixels in the vertical direction. This significantly improves the discriminative ability of algorithm towards vertical structures and stripes, with minimal computational cost. Extensive experiments show that the proposed method can achieve promising results and has better universality for different infrared scenarios. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | PTET: A progressive token exchanging transformer for infrared and visible image fusion
Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Linfeng Tang, Xinyu Xiang |
Image Vis. Comput. | 4 |
| 2024 | PSD-ELGAN: A pseudo self-distillation based CycleGAN with enhanced local adversarial interaction for single image dehazing
Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
Neural Networks | 4 |
| 2024 | MC-Net: Integrating Multi-Level Geometric Context for Two-View Correspondence LearningabstractIn two-view correspondence learning, prevalent multi-layer perceptron (MLP)-based methods struggle with context capturing. To remedy this issue, recent advances innovatively stack convolutional neural network (CNN)-based Resblocks sequentially, showing an inherent proficiency in local context extraction. Yet, such non-issue-specific designs inherit the drawback of CNN’s difficulty in aggregating global context, leading to performance bottlenecks. To address this problem, this prospective study further explores the potential of the CNN-based framework and proposes MC-Net, a top-performing network that integrates both local and global context elegantly and seamlessly. Specifically, considering that sparse motion vectors and a dense motion field can be converted into each other through interpolation and sampling, we first transform unordered matches into image-structured data by estimating the dense motion field implicitly. Then, we design a hierarchical rectifying module to rectify the error of each ordered motion vector with CNN at multiple levels, enabling MC-Net to perceive global context from coarse-level features and local context from fine-level features simultaneously, which facilitates to tackle the discontinuities of the motion field in case of large scene disparity. Finally, we reconstruct comprehensive context-embedded features from rectified motion fields at all levels. Also, instead of using the residuals between rectified and pre-rectified motion vectors at the same layer to reject outliers as in previous studies, which seriously affects the inlier prediction accuracy, we rethink this operation meticulously and modify it to the difference between motion vectors obtained from each layer’s reconstruction and ones from the first layer before transformation, ensuring purer residuals and enhancing the matching performance without extra computational burden. Extensive experiments show that MC-Net outperforms state-of-the-arts on multiple domains and datasets. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Robust Feature Matching via Graph Neighborhood Motion ConsensusabstractIn this paper, we propose an effective method for mismatch removal, termed as graph neighborhood motion consensus, to address the feature matching problem which plays a pivotal role in various computer vision tasks. In our method, we convert each feature correspondence into a motion field sample and model it with the probabilistic graphical model (PGM). To differentiate mismatches from true matches, we firstly design a metric based on neighborhood topology consensus and neighborhood interaction to evaluate the correctness of each match. We also design a variance-based similarity search module to make the information used more reliable for better matching performance. To derive the solution of PGM, we build a model to transform the problem into an integer quadratic programming problem and obtain its closed-form solution with linear time complexity. Extensive experiments on general feature matching, fundamental matrix estimation and image registration tasks demonstrate that our proposed method can achieve superior performance over several state-of-the-art approaches. Jun Huang 0008, Yijia Gong, Fan Fan 0001, Yong Ma 0001, Qinglei Du, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Cycle-Retinex: Unpaired Low-Light Image Enhancement via Retinex-Inline CycleGANabstractLow-light image enhancement aims to recover normal-light images from the images captured under dim environments. Most existing methods could just improve the light appearance globally whereas failing to handle other degradation such as dense noise, color offset and extremely low-light. Moreover, unsupervised methods proposed in recent years lack reliable physical model as the basis, thus universality is greatly limited. To address these problems, we propose a novel low-light image enhancement method via Retinex-inline cycle-consistent generative adversarial network named Cycle-Retinex, whose training is totally dependent on unpaired datasets. Specifically, we organically combine Retinex theory with CycleGAN, by which we decouple low-light image enhancement task into two sub-tasks, i.e. illumination map enhancement and reflectance map restoration. Retinex theory helps CycleGAN simplify low-light image enhancement problem and CycleGAN provides synthetic paired images to guide the training of Retinex decomposition network. We further introduce a self-augmented method to address the color distortion and noise problem, thus making the network learn to enhance low-light images adaptively. Extensive experiments show that the proposed method can achieve promising results. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Learning-based correspondence classifier with self-attention hierarchical network
Mingfan Chu, Yong Ma 0001, Xiaoguang Mei, Jun Huang 0008, Fan Fan 0001 |
Appl. Intell. | 5 |
| 2023 | Improved DBSCAN for Infrared Cluster Small Target DetectionabstractWith the development of modern weapons such as UAV swarms and multi-warhead missiles, infrared (IR) cluster small target detection technology has become increasingly important. However, the difficulty in characterizing cluster multi-targets leads to poor detection performance of existing methods. On the one hand, this paper proposes improved DBSCAN (IDBSCAN) to accurately extract the features of cluster multi-targets with unknown number and distribution. On the other hand, IDBSCAN-based difference measure (IDBSCAN-DM) is proposed, which fuses saliency and distribution features to further enhance cluster multi-targets. Specifically, we first design the multiscale sliding window to quickly extract candidate targets. Then, the IDBSCAN-based local window is constructed and IDBSCAN-DM is computed for better target enhancement and background suppression. Finally, adaptive threshold segmentation is performed on the IDBSCAN-DM map to detect real targets. Extensive comparative experiments demonstrate that the proposed method achieves better target enhancement and higher probability of detection. Zhaobing Qiu, Yong Ma 0001, Fan Fan 0001, Jun Huang 0008, You Du |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Hyperspectral image denoising via spectral noise distribution bootstrap
Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2023 | Progressive Hyperspectral Image Destriping With an Adaptive Frequencial FocusabstractLimited by the imaging paradigm, stripes is pervasive in remote sensing scenes, and its intensity, density, and periodicity differ dramatically among different imaging systems. Worse, it always co-exists with random noises caused by unstable imaging condition. However, current destriping methods are victim to undue ideal assumptions and fail to accurately eliminate stripes against diverse practical degradation, yielding excessive or inadequate destriping results. This study proposes a progressive hyperspectral destriping method with an adaptive frequency focus for accurate destriping and delicate restoration. Specifically, a hierarchical decomposition and reconstruction framework based on progressive wavelet learning encodes the degraded input to the frequency domain with smaller scales, easing the difficulty of restoration. Then, to avoid excessive or insufficient destriping, we devote specific efforts to finely separating noise and preserving details in the high-frequency domain. First, we devise a gradient-aware frequency attention block based on the prominent unidirectional pattern of stripes, empowering to adaptively assign weights according to their sensitivity to the spatial gradient. Second, we design a focal high-frequency loss item that is dynamically scaled according to feature distance in the high-frequency domain, profiting in identifying and preserving details. Extensive experiments conducted on data with synthetic stripes and realistic satellite scenes validate the superiority of the proposed method over the current state-of-the-art methods. The code is available at https://github.com/EtPan/PHID. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Seamless UAV Hyperspectral Image Stitching Using Optimal Seamline Detection via Graph Cuts
Zongyi Peng, Yong Ma 0001, Hao Li 0034, Fan Fan 0001, Xiaoguang Mei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Adversarial Autoencoder Network for Hyperspectral UnmixingabstractSpectral unmixing (SU), which refers to extracting basic features (i.e., endmembers) at the subpixel level and calculating the corresponding proportion (i.e., abundances), has become a major preprocessing technique for the hyperspectral image analysis. Since the unmixing procedure can be explained as finding a set of low-dimensional representations that reconstruct the data with their corresponding bases, autoencoders (AEs) have been effectively designed to address unsupervised SU problems. However, their ability to exploit the prior properties remains limited, and noise and initialization conditions will greatly affect the performance of unmixing. In this article, we propose a novel technique network for unsupervised unmixing which is based on the adversarial AE, termed as adversarial autoencoder network (AAENet), to address the above problems. First, the image to be unmixed is assumed to be partitioned into homogeneous regions. Then, considering the spatial correlation between local pixels, the pixels in the same region are assumed to share the same statistical properties (means and covariances) and abundance can be modeled to follow an appropriate prior distribution. Then the adversarial training procedure is adapted to transfer the spatial information into the network. By matching the aggregated posterior of the abundance with a certain prior distribution to correct the weight of unmixing, the proposed AAENet exhibits a more accurate and interpretable unmixing performance. Compared with the traditional AE method, our approach can greatly enhance the performance and robustness of the model by using the adversarial procedure and adding the abundance prior to the framework. The experiments on both the simulated and real hyperspectral data demonstrate that the proposed algorithm can outperform the other state-of-the-art methods. Qiwen Jin, Yong Ma 0001, Fan Fan 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Hyperspectral Image Stitching via Optimal Seamline DetectionabstractHyperspectral images (HSIs) with both spatial and spectral information have found broad applications. Since most cameras have narrow viewing angle, generating panoramic images is essential to show a large-range view of the environment. So far, there are few studies on HSI stitching, and the stitching result still suffers from some problems, such as blurring and ghosting, geometric misalignment, visible seam, and spectral distortion. Hence, to address the above disadvantages, we propose a novel HSI stitching strategy using optimal seamline detection approach in this letter. First, we use a fast and robust seam estimation method to determine the seamline in each single band of HSI. This method works in RGB images, and we have modified it to be used in a single-band gray-scale image of HSI. Then, to guarantee the integrity of spatial and spectral information of hundreds of bands of HSI, we propose to apply the structural similarity (SSIM) index to select the optimal one among all band candidate seamlines and use the selected optimal seamline to stitch all the remaining bands. The experimental results demonstrate that our proposed approach outperforms traditional HSI stitching approach in both spatial and spectral performances. Zongyi Peng, Yong Ma 0001, Xiaoguang Mei, Jun Huang 0008, Fan Fan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Adaptive Scale Patch-Based Contrast Measure for Dim and Small Infrared Target DetectionabstractEffective detection of small infrared (IR) targets buried in complex backgrounds and heavy noise plays an important role in IR search and track (IRST) systems. In this letter, an IR small target detection method called adaptive scale patch-based contrast measure (ASPCM) is proposed. Compared with existing detection methods based on the human visual system (HVS), our method can estimate the size of the potential target at each location. According to the estimated size, the local contrast between the target and the background is greatly enhanced, and the background clutter can be further suppressed. Experimental results on four real sequences of different complex background demonstrate that the proposed method can detect targets effectively and efficiently, even in the case of complex backgrounds and heavy noises. Zhaobing Qiu, Yong Ma 0001, Fan Fan 0001, Jun Huang 0008, Minghui Wu 0007 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Global Sparsity-Weighted Local Contrast Measure for Infrared Small Target DetectionabstractLocal contrast measure (LCM) proves effective in infrared (IR) small target detection. Existing LCM-based methods focus on mining local features of small targets to improve detection performance. As a result, they struggle to reduce false alarms while maintaining detection rates, especially with high-contrast background interference. To address this issue, this letter proposes global sparsity-weighted local contrast measure (GSWLCM), which fuses both global and local features of small targets. First, robust local contrast measure (RLCM) is proposed to remove low-contrast backgrounds and extract candidate targets. Then, to suppress high-contrast backgrounds, we customize the random walker (RW) to extract candidate target pixels, construct the global histogram and calculate global sparsity. Finally, GSWLCM fusing global and local features is calculated and the target is detected by adaptive threshold segmentation. Extensive experimental results show that the proposed method is effective in suppressing high-contrast backgrounds and has better detection performance than several state-of-the-art methods. Zhaobing Qiu, Yong Ma 0001, Fan Fan 0001, Jun Huang 0008 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Hyperspectral Anomaly Detection With Robust Graph AutoencodersabstractAnomaly detection of hyperspectral data has been gaining particular attention for its ability in detecting targets in an unsupervised manner. Autoencoder (AE), together with its variants can not only extract intrinsic features automatically but also detect anomalies that differ dramatically from others. Many AE-driven algorithms are, thus, proposed for anomaly detection in hyperspectral imagery (HSI), but they suffer from two problems: 1) when there exist anomalies in the training set, AE can generalize so well that it can also learn the abnormal patterns well, thereby reducing the ability to distinguish anomalies from the background and 2) geometric structure among samples are lost in latent space of AE, which is vital in hyperspectral anomaly detection. To tackle these problems, we propose a robust anomaly detector based on the AE framework, named robust graph AE (RGAE) detector, in this article. To be specific, we propose a robust AE framework with$\ell _{2,1}$-norm that is robust to noise and anomalies during training. Meanwhile, we embed a superpixel segmentation-based graph regularization term (SuperGraph) into AE. This strategy can preserve the geometric structure and the local spatial consistency of HSI simultaneously and also effectively reduce the searching space and execution time for each pixel. Extensive experiments are conducted on five datasets, and the results demonstrate that our method has a better detection performance, after comparing with other state-of-the-art hyperspectral anomaly detectors. Ganghui Fan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SQAD: Spatial-Spectral Quasi-Attention Recurrent Network for Hyperspectral Image DenoisingabstractThis article presents a novel end-to-end model based on encoder–decoder architecture for hyperspectral image (HSI) denoising, named spatial-spectral quasi-attention recurrent network, denoted as SQAD. The central goal of this work is to incorporate the intrinsic properties of HSI noise to construct a practical feature extraction module while maintaining high-quality spatial and spectral information. Accordingly, we first design a spatial-spectral quasi-recurrent attention unit (QARU) to address that issue. QARU is the basic building block in our model, consisting of spatial component and spectral component, and each of them involves a two-step calculation. Remarkably, the quasi-recurrent pooling function in the spectral component could explore the relevance of spatial features in the spectral domain. The spectral attention calculation could strengthen the correlation between adjacent spectra and provide the intrinsic properties of HSI noise distribution in the spectral dimension. Apart from this, we also design a unique skip connection consisting of channelwise concatenation and transition block in our model to convey the detailed information and promote the fusion of the low-level features with the high-level ones. Such a design helps maintain better structural characteristics, and spatial and spectral fidelities when reconstructing the clean HSI. Qualitative and quantitative experiments are performed on publicly available datasets. The results demonstrate that SQAD outperforms the state-of-the-art methods of visual effect and objective evaluation metrics. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Learning Spatial-Parallax Prior Based on Array Thermal Camera for Infrared Image EnhancementabstractIn this article, an array thermal camera equipment is developed to capture multiple infrared images with spatial and parallax information. Based on the captured images, an end-to-end method called spatial–parallax prior network (SPPN) is proposed. Specifically, we design a spatial–parallax prior block with two symmetric branches to extract spatial and parallax features in an interactive guidance manner. Then, to effectively integrate spatial and parallax features, we introduce a channel attention mechanism to enable the network to focus on and fuse the most useful information adaptively. In this way, spatial and parallax information can be fully utilized without any explicit alignment operation. Finally, considering the scarcity and poor quality of infrared training data, we leverage transfer learning to better train the network. Extensive experimental results demonstrate that the proposed SPPN consistently outperforms the current state-of-the-art methods, providing a highly effective and scalable solution for the improvement of infrared image quality. Jiayi Ma 0001, Wenjing Gao, Yong Ma 0001, Jun Huang 0008, Fan Fan 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Loop-Closure Detection Using Local Relative Orientation MatchingabstractLoop-closure detection (LCD), which aims to recognize a previously visited location, is a crucial component of the simultaneous localization and mapping system. In this paper, a novel appearance-based LCD method is presented. In particular, we propose a simple yet surprisingly useful feature matching algorithm for real-time geometrical verification of candidate loop-closures, termed aslocal relative orientationmatching (LRO). It aims to efficiently establish reliable feature correspondences based on preserving local topological structures between the query image and candidate frame. To effectively retrieve candidate loop closures, we introduce the aggregated selective match kernel framework into the LCD task, which can effectively represent images and reduce the quantization noise of the traditional bag-of-words framework. In addition, the SuperPoint neural network is employed to extract reliable interest points and feature descriptors. Extensive experimental results demonstrate that our LRO can significantly improve the LCD performance, and the proposed overall LCD method can achieve much better performance over the current state-of-the-art on six publicly available datasets. Jiayi Ma 0001, Xinyu Ye, Huabing Zhou, Xiaoguang Mei, Fan Fan 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image ClassificationabstractSince CapsNet [1] shattered all previous records of algorithms for image recognition, the capsule's conception has attracted bright attention. It interprets an object by the geometrical arrangement of parts. We think it can be transferred to hyperspectral images. In a hyperspectral data cube, each pixel spectrum can be regarded as a continuous curve representing its inherent properties. In the spatial domain, there are various spatial distributions in different positionsand there is usually a specific structural relationship between adjacently distributed categories. Based on HSI data's aforementioned structural characteristics, combined with the stacked capsule autoencoder, we propose our model to achieve an unsupervised HSI classification. In our model, the ConvLSTM is employed to discover part capsules of HSI, and we utilize Set Transformer to encode relations among all parts and indicate object capsules. The decoders of both phases use Gaussian mixture models to reconstruct specific information. Experimental results of the Pavia Center dataset show the exceptional of our model. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jiayi Ma 0001 |
ICASSP | 4 |
| 2021 | A Double-Neighborhood Gradient Method for Infrared Small Target DetectionabstractEffective and efficient infrared (IR) small target detection is essential for IR search and tracking (IRST) systems. The current methods have some limitations in background suppression or detection of targets close to each other. In this letter, a double-neighborhood gradient method (DNGM) is proposed. First, a new technology of the tri-layer sliding window is designed to measure the double-neighborhood gradient. Then, the DNGM is obtained by multiplying the double-neighborhood gradient. In this way, even the sizes of the targets may vary, ranging from 2 ×1 to 9 ×9 pixels, the target can be better highlighted under a fixed scale, and background interference can be suppressed. Finally, the target is segmented from the DNGM salience map by an adaptive threshold. Experiments illustrate that the proposed method can avoid the “expansion effect” of the traditional multiscale human vision system (HVS) method and can accurately detect multiple targets close to each other. Besides, the proposed method is more robust and real-time than the existing methods. Yong Ma 0001, Fan Fan 0001, Minghui Wu 0007, Jun Huang 0008 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | A Generative Adversarial Network For Medical Image FusionabstractIn this paper, a novel end-to-end model for fusing medical images characterizing structural information, i.e., IS, and images characterizing functional information, i.e., IF, of different resolutions is proposed, which is achieved by using a conditional generative adversarial network with multiple generators and multiple discriminators (MGMDcGAN). In the first cGAN, a real-like fused image is generated by a generator, simultaneously fooling two discriminators. While the discriminators are to distinguish the fused image from source images. Besides, to prevent the functional information from being weakened in the final fused image when enhancing the dense structure information, we employ the second cGAN with a mask calculated. Meanwhile, the structural information in ISand the functional information in IFthe final fused image can be concurrently kept. Furthermore, our MGMDcGAN is a unified method, which is applicable to different kinds of medical image fusion, including MRI-PET, MRISPECT, and CT-SPECT. Extensive experiments on publicly available datasets substantiate the superiority of our MGMDcGAN over the current state-of-the-art. Zhuliang Le, Jun Huang 0008, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICIP | 3 |
| 2020 | Learning to find reliable correspondences with local neighborhood consensus
Xiaoguang Mei, Yong Ma 0001, Jun Huang 0008, Fan Fan 0001, Jiayi Ma 0001 |
Neurocomputing | 5 |
| 2020 | A generative adversarial network with adaptive constraints for multi-focus image fusion
Jun Huang 0008, Zhuliang Le, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001 |
Neural Comput. Appl. | 5 |
| 2019 | Gaussian Mixture Model for Hyperspectral Unmixing with Low-Rank RepresentationabstractGaussian mixture model (GMM) can estimate not only the abundances and distribution parameters but also distinct end-member set for each pixel. However, the traditional GMM unmixing model only has proper smoothness and sparsity prior constraints on the abundances and thus cannot excavate the local spatial information in hyperspectral image (HSI). Thus, we propose a new unmixing method with superpixel segmentation (SS) and low-rank representation (LRR) based on GMM called GMM-SS-LRR, which can consider the local spatial correlation of HSI. First, we adopt the principal component analysis (PCA) to obtain the first principal component of HSI, which contains the most information for the entire HSI. Then, we adopt the SS in the first principal component of HSI to obtain the homogeneous regions, and the abundances in each homogeneous region have the underlying low-rank property. Finally, we unmix the pixels in each homogeneous region of HSI depending on the low-rank property of abundances. Experiments on synthetic datasets and real H-SIs demonstrate that the proposed GMM-SS-LRR is efficient compared with other current popular methods. Qiwen Jin, Yong Ma 0001, Xiaoguang Mei, Xiaobing Dai, Hao Li 0034, Fan Fan 0001, Jun Huang 0008 |
IGARSS | 6 |
| 2019 | GRU with Spatial Prior for Hyperspectral Image ClassificationabstractNeural networks have been successfully used to extract deep features for many hyperspectral tasks. In this study, we propose a tiny effective model based on gate recurrent unit (GRU) with spectral-spatial information for hyperspectral image classification. In our method, the core GRU cell can learn interspectral correlations within an entirely continuous spectrum input, and spatial information is the initial state of this GRU cell as a prior. Experimental results demonstrate that our method can fully utilize spectral and spatial information to obtain competitive performance. Erting Pan, Yong Ma 0001, Xiaobing Dai, Fan Fan 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IGARSS | 4 |
| 2019 | Spectral-Spatial Classification of Hyperspectral Image based on a Joint Attention NetworkabstractDeep neural networks have been successfully applied to extracting deep features for many hyperspectral tasks. Attention mechanism has been widely used in computer vision, inspired by this, we have designed a joint attention network for spectral-spatial classification of hyperspectral image. In our method, recurrent neural network (RNN) with attention can learn inner spectral correlations within a continuous spectrum, convolutional neural network (CNN) with attention is designed to focus on saliency features and spatial dependency in the neighbor regions. Experimental results demonstrate that our method can fully utilize spectral and spatial information to obtain competitive performance. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Xiaobing Dai, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001 |
IGARSS | 5 |
| 2018 | Wasserstein Introspective Neural NetworksabstractWe present Wasserstein introspective neural networks (WINN) that are both a generator and a discriminator within a single model. WINN provides a significant improvement over the recent introspective neural networks (INN) method by enhancing INN's generative modeling capability. WINN has three interesting properties: (1) A mathematical connection between the formulation of the INN algorithm and that of Wasserstein generative adversarial networks (WGAN) is made. (2) The explicit adoption of the Wasserstein distance into INN results in a large enhancement to INN, achieving compelling results even with a single classifier - e.g., providing nearly a 20 times reduction in model size over INN for unsupervised generative modeling. (3) When applied to supervised classification, WINN also gives rise to improved robustness against adversarial examples in terms of the error reduction. In the experiments, we report encouraging results on unsupervised learning problems including texture, face, and object modeling, as well as a supervised classification task against adversarial attacks. Our code is available online1. Kwonjoon Lee, Weijian Xu, Fan Fan 0001, Zhuowen Tu |
CVPR | 3 |
| 2018 | Robust GBM hyperspectral image unmixing with superpixel segmentation based low rank and sparse representation
Xiaoguang Mei, Yong Ma 0001, Chang Li 0001, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
Neurocomputing | 4 |
| 2017 | Hyperspectral image denoising with superpixel segmentation and low-rank representation
Fan Fan 0001, Yong Ma 0001, Chang Li 0001, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
Inf. Sci. | 1 |
| 2016 | Infrared and visible image fusion using total variation model
Yong Ma 0001, Jun Chen 0019, Chen Chen 0003, Fan Fan 0001, Jiayi Ma 0001 |
Neurocomputing | 4 |
| 2014 | A Robust Infrared Small Target Detection Algorithm Based on Human Visual SystemabstractRobust human visual system (HVS) properties can effectively improve the infrared (IR) small target detection capabilities, such as detection rate, false alarm rate, speed, etc. However, current algorithms based on HVS usually improve one or two of the aforementioned detection capabilities while sacrificing the others. In this letter, a robust IR small target detection algorithm based on HVS is proposed to pursue good performance in detection rate, false alarm rate, and speed simultaneously. First, an HVS size-adaptation process is used, and the IR image after preprocessing is divided into subblocks to improve detection speed. Then, based on HVS contrast mechanism, the improved local contrast measure, which can improve detection rate and reduce false alarm rate, is proposed to calculate the saliency map, and a threshold operation along with a rapid traversal mechanism based on HVS attention shift mechanism is used to get the target subblocks quickly. Experimental results show the proposed algorithm has good robustness and efficiency for real IR small target detection applications. Jinhui Han, Yong Ma 0001, Bo Zhou 0006, Fan Fan 0001, Kun Liang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |