VLDB 2026 Research / reviewers in the wild / expert
Jiayi Ma 0001
dblp:96/9989
· DBLP profile ↗
335ranked-venue papers
32as first author
234since 2021 · last 2026
0000-0003-3264-3265ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 168 · 12 first-author · 132 since 2021Graphics, computer vision, multimedia, augmented reality and games · 167 · 16 first-author · 114 since 2021Applied, interdisciplinary, general and emerging computing · 55 · 6 first-author · 36 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingabstractMulti-modal image matching is a fundamental task in multi-view and multi-modal image processing. Its key challenge lies in extracting features that remain consistent despite drastic appearance variations across modalities. However, the learning of the feature is hindered by the scarcity and the inaccurate alignment of existing multi-modal datasets. To address this, we propose a knowledge distillation framework termed SGPFeat that transfers rich prior knowledge from large-scale unimodal tasks to enhance multi-modal representation learning. Specifically, semantic priors from a vision foundation model guide the feature extractor to identify shared semantic structures across modalities, enabling better generalization under large appearance gaps. In parallel, geometric priors derived from accurately aligned visible-light datasets improve detection precision on noisy aligned multi-modal pairs. Furthermore, we introduce a Heterogeneous Feature Aggregation (HFA) module to facilitate effective distillation and feature representation. Extensive experiments demonstrate that semantic and geometric priors bring significant improvement for our SGPFeat across diverse multi-modal image matching benchmarks. Yuxin Deng 0002, Botian Wang, Kaining Zhang, Hao Zhang 0073, Jiayi Ma 0001 |
AAAI | 5 |
| 2026 | GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View GeometryabstractRecent progress in two-view geometry increasingly emphasizes enforcing smoothness and global consistency priors when estimating motion fields between pairs of images. However, in complex real-world scenes, characterized by extreme viewpoint and scale changes as well as pronounced depth discontinuities, the motion field often exhibits diverse and heterogeneous motion patterns. Most existing methods lack targeted modeling strategies and fail to explicitly account for this variability, resulting in estimated motion fields that diverge from their true underlying structure and distribution. We observe that Mixture-of-Experts (MoE) can assign dedicated experts to motion sub-fields, enabling a divide-and-conquer strategy for heterogeneous motion patterns. Building on this insight, we re-architect motion field modeling in two-view geometry with GeoMoE, a streamlined framework. Specifically, we first devise a Probabilistic Prior-Guided Decomposition strategy that exploits inlier probability signals to perform a structure-aware decomposition of the motion field into heterogeneous sub-fields, sharply curbing outlier-induced bias. Next, we introduce an MoE-Enhanced Bi-Path Rectifier that enhances each sub-field along spatial-context and channel-semantic paths and routes it to a customized expert for targeted modeling, thereby decoupling heterogeneous motion regimes, suppressing cross-sub-field interference and representational entanglement, and yielding fine-grained motion-field rectification. With this minimalist design, GeoMoE outperforms prior state-of-the-art methods in relative pose and homography estimation and shows strong generalization. Jiajun Le, Jiayi Ma 0001 |
AAAI | 2 |
| 2026 | Probabilistic Deformation Consistency for Unsupervised Shape MatchingabstractIn this paper, we propose a novel unsupervised shape matching framework based on probabilistic deformation consistency in the spectral domain, termed as PDCMatch. Axiomatic optimization methods suffer from expensive geodesic distance calculations and vulnerability to local optima, and learning-based methods typically lack geometric consistency in pointwise correspondences. To overcome both limitations, we develop a non-Euclidean probabilistic deformation model that jointly estimates the underlying deformation and the correspondence probability via a linear Expectation-Maximization procedure. Building on this formulation, we further design a task-specific deformation loss that explicitly encourages geometric smoothness and structural consistency in an unsupervised manner. This tailored loss function plays a central role in improving the matching performance across challenging scenarios. Extensive experiments on public benchmarks involving near-isometric shapes, anisotropic meshing, cross-dataset generalization, topological noise, and non-isometric shapes demonstrate that our method consistently outperforms state-of-the-art methods, highlighting both its effectiveness and generalizability. Tianwei Ye, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
AAAI | 5 |
| 2026 | Diff-NAT: Better Naturalistic and Aggressive Adversarial Attacks via Class-Optimized Diffusion for Object DetectionabstractRecent advances in naturalistic physical adversarial patch generation show great promise in protecting personal privacy against detector-based malicious surveillance while remaining inconspicuous to human observers. In this work, we present the first systematic categorization and in-depth re-examination of existing methods into three representative paradigms, revealing a pervasive imbalance: enforcing naturalness constraints inherently restricts the adversarial search space, thus limiting attack performance. To address this challenge, we propose a novel paradigm based on class-optimized diffusion, termed Diff-NAT. Diff-NAT leverages pretrained diffusion models as powerful natural image priors and introduces a unified iterative framework that jointly optimizes two complementary components: semantic-level textual prompts and instance-level latent codes. Specifically, prompt optimization enables broad traversal across inter-class semantic regions, while latent refinement allows for fine-grained manipulation within class objectives. This dual-level optimization facilitates progressive navigation toward adversarial distributions embedded within the natural semantic manifold. Extensive experiments in both digital and physical settings demonstrate that Diff-NAT outperforms existing SOTA approaches in terms of both visual realism and aggressiveness. Qinglong Yan, Tong Zou, Xunpeng Yi, Xinyu Xiang, Xuying Wu, Hao Zhang 0073, Jiayi Ma 0001 |
AAAI | 7 |
| 2026 | Robust Fusion Controller: Degradation-Aware Image Fusion with Fine-Grained Language InstructionsabstractCurrent image fusion methods struggle to adapt to real-world environments encompassing diverse degradations with spatially varying characteristics. To address this challenge, we propose a robust fusion controller (RFC) capable of achieving degradation-aware image fusion through fine-grained language instructions, ensuring its reliable application in adverse environments. Specifically, RFC first parses language instructions to innovatively derive the functional condition and the spatial condition, where the former specifies the degradation type to remove, while the latter defines its spatial coverage. Then, a composite control priori is generated through a multi-condition coupling network, achieving a seamless transition from abstract language instructions to latent control variables. Subsequently, we design a hybrid attention-based fusion network to aggregate multi-modal information, in which the obtained composite control priori is deeply embedded to linearly modulate the intermediate fused features. To ensure the alignment between language instructions and control outcomes, we introduce a novel language-feature alignment loss, which constrains the consistency between feature-level gains and the composite control priori. Extensive experiments on publicly available datasets demonstrate that our RFC is robust against various composite degradations, particularly in highly challenging flare scenarios. Hao Zhang 0073, Yanping Zha, Qingwei Zhuang, Jiayi Ma 0001 |
AAAI | 5 |
| 2026 | MAP-MIL: Dual-branch collaborative learning of mask enhancement and pseudo-bag generation for whole slide image classification
Zequn Liu, Liangkuan Zhu, Yining Xie, Jiayi Ma 0001 |
Expert Syst. Appl. | 7 |
| 2026 | Locality Optimization Refinement with Deformation for Shape Matching via Functional Maps
Jiayi Ma 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Multi-view parallel convolutional network for organ segmentation in mediastinal region on CT imagesabstract• This paper is the first to propose a method for organ segmentation in mediastinal region on CT images. • This paper proposes multi-view parallel convolution module to capture the organ’s unique and complex morphological characteristics. • This paper introduces a region fusion small-kernel deformable attention mechanism to address the issue of spatial information and detailed deformation feature loss in medical image segmentation. • This paper design efficient dual-channel bottleneck structures to shared parameters and optimize the computation path. In lung CT images, mediastinal organ segmentation is crucial for localizing different mediastinal regions. However, existing medical image segmentation methods exhibit significant limitations in modeling the diverse topological structures of organs, sensitivity to intra-class morphological variations, and inter-class feature differentiation. To address these limitations, we propose a novel multi-view parallel convolutional network (MVPCNet), built on an efficient U-shaped encoder-decoder framework. The shallow and deep information encoders are respectively composed of alternating multi-view parallel convolution module (MVPM) and the dual-path backbone structure (DPBS) at different scales. MVPM is designed as a parallel convolutional structure to enhance the model’s ability to capture complex structural features, enabling complementary extraction of morphological and detailed features. DPBS comprises the efficient dual-channel bottleneck structures (EDC-BS) and the region fusion small-kernel deformable attention mechanism (RF-SKDA). EDC-BS employs a branched convolutional architecture, effectively reducing computational complexity while ensuring accurate recognition of the same organ across varying morphologies. RF-SKDA captures the spatial structural information of different organs by combining regional and global average pooling, and further extracts organ-specific morphological features through the deformable convolutions. The decoder utilizes lightweight parameterization through depthwise separable convolutions and integrates multi-scale features during the decoding process. Experimental results demonstrate that MVPCNet achieves an average Dice Coefficient of 90.59 % and an mIoU of 82.80 % on mediastinal organ dataset. With a parameter size of only 8.21 MB, it outperforms advanced medical segmentation algorithms and classical lightweight semantic segmentation models. Yining Xie, Jiayi Ma 0001, Fengjiao Wang |
Neural Networks | 3 |
| 2026 | MT-IDS: A multi-task information decoupling strategy for identifying lymph node metastasis in the mediastinal region
Yining Xie, Fengjiao Wang, Jiayi Ma 0001 |
Neural Networks | 5 |
| 2026 | Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer ApproachabstractImage fusion aims to blend complementary information from multiple sensing modalities, yet existing approaches remain limited in robustness, adaptability, and controllability. Most current fusion networks are tailored to specific tasks and lack the ability to flexibly incorporate user intent, especially in complex scenarios involving low-light degradation, color shifts, or exposure imbalance. Moreover, the absence of ground-truth fused images and the small scale of existing datasets make it difficult to train an end-to-end model that simultaneously understands high-level semantics and performs fine-grained multimodal alignment. We therefore present DiTFuse, an instruction-driven Diffusion Transformer (DiT) framework that performs end-to-end, semantics-aware fusion within a single model. By jointly encoding two images and natural-language instructions in a shared latent space, DiTFuse enables hierarchical and fine-grained control over fusion dynamics, overcoming the limitations of pre-fusion and post-fusion pipelines that struggle to inject high-level semantics. The training phase employs a multi-degradation masked-image modeling strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ground truth images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion-as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multi-image fusion scenarios, including instruction-conditioned segmentation. Jiayang Li 0004, Chengjie Jiang, Junjun Jiang, Pengwei Liang, Jiayi Ma 0001, Liqiang Nie |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Mask-DiFuser: A Masked Diffusion Model for Unified Unsupervised Image FusionabstractThe absence of ground truth (GT) in most fusion tasks poses significant challenges for model optimization, evaluation, and generalization. Existing fusion methods achieving complementary context aggregation predominantly rely on hand-crafted fusion rules and sophisticated loss functions, which introduce subjectivity and often fail to adapt to complex real-world scenarios. To address this challenge, we propose Mask-DiFuser, a novel fusion paradigm that ingeniously transforms the unsupervised image fusion task into a dual masked image reconstruction task by incorporating masked image modeling with a diffusion model, overcoming various issues arising from the absence of GT. In particular, we devise a dual masking scheme to simulate complementary information and employ a diffusion model to restore source images from two masked inputs, thereby aggregating complementary contexts. A content encoder with an attention parallel feature mixer is deployed to extract and integrate complementary features, offering local content guidance. Moreover, a semantic encoder is developed to supply global context which is integrated into the diffusion model via a cross-attention mechanism. During inference, Mask-DiFuser begins with a Gaussian distribution and iteratively denoises it conditioned on multi-source images to directly generate fused images. The masked diffusion model, learning priors from high-quality natural images, ensures that fusion results align more closely with human visual perception. Extensive experiments on several fusion tasks, including infrared-visible, medical, multi-exposure, and multi-focus image fusion, demonstrate that Mask-DiFuser significantly outshines SOTA fusion alternatives. Linfeng Tang, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | TOFusion: Text-guided and object-aware infrared and visible image fusion
Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2026 | MSTDNet: Multi-scale traffic object detection network with smooth information perception
Jie Hua 0005, Zhongyuan Wang 0001, Hua Zou 0002, Gang Wu 0010, Jiayi Ma 0001 |
Pattern Recognit. | 6 |
| 2026 | No-reference dehazed image quality assessment via perception-driven interactive feature representation learning
Hangyu Nie, Ziqiang Huang, Miao Qi, Junjun Jiang, Jiayi Ma 0001, Wei Liu 0123 |
Pattern Recognit. | 5 |
| 2026 | Land-cover prior diffusion probabilistic model for remote sensing image super resolution
Zhizheng Zhang 0009, Jiayi Ma 0001, Jindou Zhang, Yu Wang 0140, Zhenghao Liao, Gui Cheng, Mingqiang Guo, Liang Wu 0005 |
Pattern Recognit. | 3 |
| 2026 | SPEN: Sub-Pixel Position Error Estimation Network for Multi-Modal Image Matching
Maoqing Hu, Bin Sun 0001, Shutao Li 0001, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Self-Iteration Image Haze Removal Using a Deep Curve-Dehazing ModelabstractThis paper proposes a novel dehazing method termed Haze-Restoration Curve Model (HRCM), which transforms the single-image dehazing task into a specific curve estimation problem, achieving haze removal through an intuitive and simple nonlinear curve mapping. Unlike methods based on Atmospheric Scattering Model (ASM), HRCM does not require the computation of complex physical parameters. Instead, it estimates two intuitive curvature adjustment coefficients. Moreover, compared to recent end-to-end dehazing methods, HRCM circumvents the challenging modeling of static mapping functions, thereby improving the generalization ability and dehazing performance of the model. All of these are attributed to a meticulously designed dehazing curve, which first reversing the hazy image to highlight obscured regions, and then specifies a set of high-order functions to remap hazy pixels for image restoration. Moreover, to estimate the curve parameters, we designed a dual-branch Deep Dehaze Curve Estimation Network(DDCEN), which consists of the Residual Swin Transformer Block(RTSB) and the Large kernel convolutional Attention Block(LAB). Specifically, RTSB captures the global fog density distribution features of foggy images by introducing window self-attention and shifted window mechanisms, providing support for global semantic information for subsequent parameter estimation. LAB captures local multi-scale features by constructing a large receptive field, and uses the attention mechanism of feature pooling in horizontal and vertical directions to focus on detail regions, refining the local details of the parameter map. Extensive experiments on synthetic and real-world hazy image datasets demonstrate that the proposed approach achieves superior performance in terms of quantitative accuracy and subjective visual quality compared to the current state-of-the-art methods. The source code of our HRCM is available at https://github.com/larrylanrui/HRCM. Wei Liu 0123, Rui Nan, Jiayi Ma 0001, Xin Chen 0003, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Think Twice Before Determining: Toward Scene-Aware Visual Reasoning for Mirror DetectionabstractMirror detection (MD) aims to overcome interference caused by reflections and locate mirror regions. Existing methods focus on designing components to explicitly establish the associations between physical entities and corresponding imagings, or utilizing rotation to construct symmetric consistency. We observe that: a) incomplete and incorrect correspondence between entities and imagings; b) other physical materials (e.g., glass) exhibit characteristics partially similar to mirrors, causing confusion when they co-occur; c) complex interfering factors (e.g., occlusion) and reflection mechanisms may expand vector space several times over. To address these issues in a unified manner, we formulate the scene-aware visual reasoning network (SVRNet) based on visual prompts. Specifically, we construct the prototype-guided prompt chain reasoning (PPCR) that generates a mixed chain of thought reasoning based on maximal difference heterogeneous prototypes to construct comprehensive spatial location and semantic perception. Noise may accumulate gradually through the chain, and crucial clues may also disappear. Therefore, we design the prompt evolution (PE) to filter out noise and enhance the coupling between prompts. We further develop the mixture of prompt injection expert (MPIE) to dynamically select the optimal injection strategy in the low-rank space based on specific scene. Due to reflection interference and random parameter space introducing potential ambiguity, we formulate the three-way evidence-aware (TEA) loss to quantify the uncertainty, thereby providing reliable predictions. To leverage historical knowledge and further disentangle representations, we propose the frequency prototype contrastive (FPC) loss for learning more generalizable features across images. Finally, we relabel 25,828 images and formulate the first point-supervised MD framework. Extensive experiments conducted on four mirror benchmarks under three settings demonstrate that our method surpasses state-of-the-art approaches. Promising results are also achieved on six related benchmarks, showing its generality. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Jiayi Ma 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Breaking Low-Light Fusion Barrier: Unsupervised Darkness and Noise-Aware Visible and Infrared Image Fusion NetworkabstractInfrared and visible image fusion aims to integrate complementary information but suffers from severe residual degradations under low-light conditions. Existing methods face two main limitations: darkness-aware fusion focuses on illumination enhancement while neglecting noise suppression, and degradation-aware supervised fusion relies on paired data and struggles with enhancement-denoising imbalance due to multi-task optimization conflicts. To address these issues, we propose BLFusion, an unsupervised darkness- and noise-aware fusion framework that performs illumination enhancement and noise suppression in two dedicated stages while fusing complementary information without high-quality references. First, a Retinex-guided state space model-based decomposition network models illumination degradation to brighten dark visible images. Then, an unsupervised denoising fusion network jointly performs fusion and denoising, where noise correlation is disrupted by shuffling and a blind-spot network with dilated convolutions estimates clean representations from surrounding pixels. Finally, noise-free features from both modalities are fused to generate the final image. Moreover, we construct the MRLL dataset with 500 well-aligned infrared-visible image pairs, filling the gap for real-world noise-degraded nighttime scenarios. Experiments demonstrate that BLFusion outperforms state-of-the-art methods and generalizes robustly across diverse low-light and noisy conditions. The MRLL dataset and code are publicly available at https://github.com/ChenDoubleJ/BLFusion-MRLL. Han Xu 0001, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | MDbFusion++: A Visible and Infrared Image Fusion Framework Capable for Motion DeblurringabstractExisting image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++. Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | SigMa: Semantic Similarity-Guided Semi-Dense Feature MatchingabstractRecent advancements have led the image matching community to increasingly focus on obtaining subpixel-level correspondences in a detector-free manner, i.e., semi-dense feature matching. Existing methods tend to overfocus on low-level local features while ignoring equally important high-level semantic information. To tackle these shortcomings, we propose SigMa, a semantic similarity-guided semi-dense feature matching method, which leverages the strengths of both local features and high-level semantic features. First, we design a dual-branch feature extractor, comprising a convolutional network and a vision foundation model, to extract low-level local features and high-level semantic features, respectively. To fully retain the advantages of these two features and effectively integrate them, we also introduce a cross-domain feature adapter, which could overcome their spatial resolution mismatches, channel dimensionality variations, and inter-domain gaps. Furthermore, we observe that performing the transformer on the whole feature map is unnecessary because of the similarity of local representations. We design a guided pooling method based on semantic similarity. This strategy performs attention computation by selecting highly semantically similar regions, aiming to minimize information loss while maintaining computational efficiency. Extensive experiments on multiple datasets demonstrate that our method achieves a competitive accuracy-efficiency trade-off across various tasks and exhibits strong generalization capabilities across different datasets. Additionally, we conduct a series of ablation studies and analysis experiments to validate the effectiveness and rationality of our method's design. Our code is publicly available at https://github.com/ShineFox/SigMa. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Selecting and Pruning: A Differentiable Causal Sequentialized State-Space Model for Two-View Correspondence LearningabstractTwo-view correspondence learning aims to discern true and false correspondences between image pairs by recognizing their underlying different information. Previous methods either treat the information equally or require the explicit storage of the entire context, tending to be laborious in real-world scenarios. Inspired by Mamba's inherent selectivity, we propose CorrMamba, a Correspondence filter leveraging Mamba's ability to selectively mine information from true correspondences while mitigating interference from false ones, thus achieving adaptive focus at a lower cost. To prevent Mamba from being potentially impacted by unordered keypoints that obscured its ability to mine spatial information, we customize a causal sequential learning approach based on the Gumbel-Softmax technique to establish causal dependencies between features in a fully autonomous and differentiable manner. Additionally, a local-context enhancement module is designed to capture critical contextual cues essential for correspondence pruning, complementing the core framework. Extensive experiments on relative pose estimation, visual localization, and analysis demonstrate that CorrMamba achieves state-of-the-art performance. Notably, in outdoor relative pose estimation, our method surpasses the previous SOTA by 2.58 absolute percentage points in AUC@20°, highlighting its practical superiority. Our code is publicly available at https://github.com/ShineFox/CorrMamba. Hao Zhang 0073, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light ImagingabstractDenoising images captured under extreme low-light conditions remains a persistent challenge in computational photography, primarily due to low signal-to-noise ratios and sensor-specific noise characteristics. These variations often require per-sensor noise calibration to achieve effective denoising. Although recent calibration-free methods aim to reduce this dependency through synthetic noise modeling or few-shot fine-tuning, their performance often degrades in extreme low-light scenarios across different sensors due to mismatches between synthetic and real-world noise. To address this gap, we introduce CLIP-Guided Denoising (CLD), the first framework to leverage large-scale vision models pretrained on sRGB images for cross-domain feature fusion, effectively guiding RAW image denoising across diverse sensors. Although not trained on RAW data, CLIP embeddings offer semantically robust and noise-invariant features that help guide the denoising network to focus on the underlying image content rather than fitting to specific noise distributions. Extensive experiments on the SID and ELD datasets demonstrate that CLD achieves state-of-the-art performance in calibration-free settings, significantly outperforming prior methods under extreme low-light conditions and achieving robust generalization across unseen sensor domains. Mingde Qiao, Junjun Jiang, Zhanghong Zhao, Junhui Hou, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | DSPFusion: Image Fusion via Degradation and Semantic Dual-Prior GuidanceabstractExisting infrared-visible image fusion methods are mainly tailored for high-quality source images. Although recent studies have begun to explore degradation-aware fusion, most existing methods still focus on specific degradation types, while unified frameworks that aim to handle diverse degradations often depend on auxiliary textual prompts, which limits their practicality in automatic fusion scenarios. This work presents a Degradation and Semantic Prior dual-guided framework for degraded image Fusion (DSPFusion), which jointly performs degradation-aware restoration and complementary information aggregation in a unified architecture without relying on auxiliary prompts. Specifically, it first extracts modality-specific degradation priors from degraded infrared and visible images, while capturing compact semantic embeddings from paired source images as low-quality semantic priors to encode global scene context. Then, a semantic prior diffusion model is devised to restore high-quality scene semantic priors in a compact latent space, providing global scene guidance with low computational overhead and enabling over $30\times $ inference speedup compared with mainstream diffusion model-based image fusion schemes, such as DDFM. Guided by the restored semantic priors and degradation priors, the enhancement and fusion network adaptively suppresses degradations and aggregates complementary information. Extensive experiments under both degraded and normal scenarios demonstrate that DSPFusion effectively handles representative degradations, preserves complementary information, and achieves competitive performance with low computational cost, thereby broadening the practical application scope of image fusion. The source code is publicly available at https://github.com/Linfeng-Tang/DSPFusion. Linfeng Tang, Yeda Wang, Guoqing Wang 0001, Yixuan Yuan, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Diff-MEF: Cross-Modal Diffusion Framework With Text Prompts and Semantic Perception for Multi-Exposure Image FusionabstractThe absence of real-world ground truth (GT) remains a challenge in multi-exposure image fusion (MEF). Benchmarks synthesizing pseudo GT through algorithm ensembles. Existing methods, hampered by inherent imperfections of pseudo GT and fixed mapping relationships, show limited performance and robustness. To address the limitations, we propose a novel cross-modal diffusion framework that synergizes text prompts and semantic perception for MEF, termed as Diff-MEF. First, it reformulates MEF as a probabilistic estimation task with conditional diffusion model for progressive transition and fusion. Then, we explicitly infer semantic and exposure priors as text prompts and semantic perception to improve performance and robustness. The priors are synergized through multi-modal prior embedding and optimization guidance. On the one hand, regarding cross-modal interaction, multi-modal priors, including segmentation masks, and exposure- and content-aware text prompts, are embedded into diffusion process by dedicated encoders and refine visual features through a text-segmentation refinement module. On the other hand, a semantic-level contrastive loss builds a regularization between cross-modal features in the semantic space of CLIP to mitigate degradations introduced by pseudo GT and fusion distortions. Experiments demonstrate that Diff-MEF outperforms SOTA methods and pseudo GT with superior fusion performance and robustness across diverse exposure scenarios. Code is available at https://github.com/hanna-xu/Diff-MEF. Han Xu 0001, Yunfei Huang, Linfeng Tang, Jiayi Ma 0001, Guangcan Liu |
IEEE Trans. Image Process. | 4 |
| 2026 | M2PL-GAN: Multi-View Multi-Level Pathology Semantic Perception Learning for H&E-to-IHC Virtual StainingabstractImmunohistochemistry (IHC) staining is crucial for determining tumor subtypes, obtaining protein expression information, and developing personalized treatment plans. But compared with hematoxylin and eosin (H&E) staining, IHC staining is more complex and expensive. With the advancement of deep learning, converting H&E stained images into IHC stained images has gradually emerged as a solution for obtaining IHC staining. However, current virtual staining processes suffer from difficulties in aligning pathological semantic features, posing significant challenges for network training, which poses significant challenges for network training. To solve these issues, we propose a multi-view multi-level pathology semantic perception learning method for H&E-to-IHC virtual staining (M2PL-GAN). Unlike prior approaches, M2PL-GAN introduces a comprehensive semantic learning paradigm from three views: structural contextual relations, feature distribution, and topology-aware fine-grained semantics. These correspond to the Context-aware Correlation Mechanism (CACM), the Local-aware Distribution Alignment Mechanism (LDAM), and the Graph- aware Bidirectional Contrastive Learning Mechanism (GBCLM) respectively. Among them, CACM enhances contextual consistency by establishing semantic correlations between virtual and real IHC images at local scales. LDAM ensures alignment of semantic feature distributions between virtual and real IHC images, mitigating semantic shifts caused by HE-IHC staining. GBCLM leverages graph neural network to capture topology-aware semantic representations and optimizes semantic feature alignment through bidirectional contrastive learning. Extensive experiments on both public and private datasets demonstrate that our method outperforms state-of-the-art approaches in both quantitative metrics and qualitative evaluations. Our code is available in https://github.com/Pikachu-one/M2PL-GAN. Zequn Liu, Liangkuan Zhu, Yining Xie, Xiaoqing Hu, Haochen Qi, Jiayi Ma 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2026 | DDFNet: Dual-Neighborhoods Dynamic Fusion Network for Image Feature MatchingabstractEstablishing reliable correspondences is a fundamental task in computer vision. Constructing neighbor graphs in feature space with position information to mine correspondence consistency has become a common strategy for recognizing correct correspondences (inliers). However, these neighbors may include a high ratio of incorrect correspondences (outliers), only using the correspondence consistency from feature space will probably be difficult to guarantee the matching accuracy. To address this issue, we propose a novel motion consistent space to find consistent neighbors that are independent of the correspondence's position and have a larger search range. On top of that, we build two neighbor graphs according to the feature space and motion consistent space separately, and expand a shift annular convolution to retain rich neighbor graph structure information and fully exploit the neighborhood context. Then, we design a dynamic feature fusion block to dynamically fuse these dual-neighbor graphs to flexibly cope with various complex scenarios. Finally, we develop a Dual-Neighborhoods Dynamic Fusion Network (DDFNet) for accurately identifying inliers and retrieving camera poses. Experimental results demonstrate that our proposed DDFNet outperforms the state-of-the-art methods. Source code:https://github.com/1211193023/DDFNet. Changcai Yang, Fengyuan Zhuang, Lifang Wei, Jiayi Ma 0001, Riqing Chen |
IEEE Trans. Multim. | 5 |
| 2025 | DeMo: Deep Motion Field Consensus with Learnable Kernels for Two-view Correspondence LearningabstractAs a long-range prior, motion consensus essentially forces the overall spatial transformation between a pair of images to be smooth and consistent, which is naturally well-suited for two-view correspondence learning. However, such precious property remains under-explored by most existing studies due to the modeling challenges posed by the sparsity and uneven distributions of putative correspondences. In this paper, we propose DeMo, a novel and cutting-edge network for outlier rejection, which possesses the capacity to fully capture global motion consensus clues by way of consensus interpolation over the entire high-dimensional motion field generated by putative correspondences. Specifically, through incorporating regularization techniques into a Reproducing Kernel Hilbert Space (RKHS), a concise interpolation formula can be derived for the high-dimensional motion field, which inherently allows a closed-form solution. Subsequently, learnable deep kernels are collaboratively used to flexibly and efficiently capture the relationships between global inputs, thus maintaining the entire motion field consensus. In addition, to remedy the cubic computational overhead of explicit interpolation, a scene-adaptive sampling strategy is introduced, which implicitly selects the more scene-representative motions, reducing the computational complexity of motion consensus interpolation to be approximately linear while maintaining the accuracy. Moreover, to deal with underlying depth discontinuities caused by complicated scene variations, a local consensus complementation block is designed, which maintains local bilateral consensus across both feature and spatial channels. Without bells and whistles, DeMo achieves superior performance in various geometric tasks, including relative pose estimation, homography estimation, and visual localization. Jiajun Le, Zizhuo Li, Yixuan Yuan, Jiayi Ma 0001 |
AAAI | 5 |
| 2025 | Multi-Shape Matching with Cycle Consistency Basis via Functional MapsabstractMulti-shape matching is a central problem in various applications of computer vision and graphics, where cycle consistency constraints play a pivotal role. For this issue, we propose a novel and efficient approach that models multi-shapes as directed graphs for two-stage optimization, i.e., optimizing pairwise correspondence accuracy using landmarks, and refining matching consistency through cycle consistency basis. Specifically, we utilize local mapping distortion to identify landmarks and extract the dimension of the functional space, which is then used to upsample in the spectral domain, thereby producing smoother results. Next, to optimize the consistency of correspondences, we introduce the cycle consistency basis, which succinctly describes all consistent cycles in the collection. We then propose cycle consistency refinement, which resolves inconsistencies in cycles efficiently via the alternating direction method of multipliers. Our approach simultaneously balances the accuracy and consistency of multi-shape matching, achieving lower correspondence errors. Extensive experiments on several public datasets demonstrate the superiority of our approach over current state-of-the-art methods. Tianwei Ye, Huabing Zhou, Zhongyuan Wang 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2025 | Cross-Modal Stealth: A Coarse-to-Fine Attack Framework for RGB-T TrackerabstractCurrent research on adversarial attacks mainly focuses on RGB trackers, with no existing methods for attacking RGB-T cross-modal trackers. To fill this gap and overcome its challenges, we propose a progressive adversarial patch generation framework and achieve cross-modal stealth. On the one hand, we design a coarse-to-fine architecture grounded in the latent space to progressively and precisely uncover the vulnerabilities of RGB-T trackers. On the other hand, we introduce a correlation-breaking loss that disrupts the modal coupling within trackers, spanning from the pixel to the semantic level. These two design elements ensure that the proposed method can overcome the obstacles posed by cross-modal information complementarity in implementing attacks. Furthermore, to enhance the reliable application of the adversarial patches in real world, we develop a point tracking-based reprojection strategy that effectively mitigates performance degradation caused by multi-angle distortion during imaging. Extensive experiments demonstrate the superiority of our method. Xinyu Xiang, Qinglong Yan, Hao Zhang 0073, Jianfeng Ding, Han Xu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
AAAI | 7 |
| 2025 | Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic AmalgamationabstractImage feature matching is a cardinal problem in computer vision, aiming to establish accurate correspondences between two-view images. Existing methods are constrained by the performance of feature extractors and struggle to capture local information affected by sparse texture or occlusions. Recognizing that human eyes consider not only similar local geometric features but also high-level semantic information of scene objects when matching images, this paper introduces SemaGlue. This novel algorithm perceives and incorporates semantic information into the matching process. In contrast to recent approaches that leverage semantic consistency to narrow the scope of matching areas, SemaGlue achieves semantic amalgamation with the designed Semantic-Aware Fusion (SAF) Block by injecting abundant semantic features from the pre-trained segmentation model. Moreover, the Cross-Domain Alignment (CDA) Block is proposed to address domain alignment issues, bridging the gaps between semantic and geometric domains to ensure applicable semantic amalgamation. Extensive experiments demonstrate that SemaGlue outperforms state-of-the-art methods across various applications such as homography estimation, relative pose estimation, and visual localization. Zhenjie Zhu, Zizhuo Li, Tao Lu 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2025 | ACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response DecouplingabstractThe research on adversarial attacks against trackers primarily concentrates on the RGB modality, whereas the methodology for attacking RGB-T multi-modal trackers has seldom been explored so far. This work represents an innovative attempt to develop an adaptive cross attack framework via multi-modal response decoupling, generating multi-modal adversarial patches to evade RGB-T trackers. Specifically, a modal-aware adaptive attack strategy is introduced to weaken the modality with high common information contribution alternately and iteratively, achieving the modal decoupling attack. In order to perturb the judgment of the modal balance mechanism in the tracker, we design a modal disturbance loss to increase the distance of the response map of the single-modal adversarial samples in the tracker. Besides, we also propose a novel spatio-temporal joint attack loss to progressively deteriorate the tracker’s perception of the target. Moreover, the design of the shared adversarial shape enables the generated multi-modal adversarial patches to be readily deployed in real-world scenarios, effectively reducing the interference of the patch posting process on the shape attack of the infrared adversarial layer. Extensive digital and physical domain experiments demonstrate the effectiveness of our multi-modal adversarial patch attack. Our code is available at https://github.com/Xinyu-Xiang/ACAttack. Xinyu Xiang, Qinglong Yan, Hao Zhang 0073, Jiayi Ma 0001 |
CVPR | 4 |
| 2025 | Adapting Dense Matching for Homography Estimation with Grid-based AccelerationabstractCurrent deep homography estimation methods are typically constrained to processing low-resolution image pairs due to network architecture and computational limitations. For high-resolution images, downsampling is often required, which can greatly degrade estimation accuracy. In contrast, image matching methods, which match pixels and compute homography from correspondences, provide greater resolution flexibility. So in this work, we revisit the traditional image matching paradigm for homography estimation and propose GFNet, a Grid Flow regression Network that adapts the high-accuracy dense matching framework for homography estimation while enhancing efficiency through a grid-based strategy—estimating flow only over a coarse grid by leveraging homography’s global smoothness. We demonstrate the effectiveness of GFNet on a wide range of experiments on multiple datasets, including the common scene MSCOCO, multimodal datasets VIS-IR and GoogleMap, and the dynamic scene VIRAT. Notably, on 448×448 GoogleMap, GFNet achieves an improvement of +13.5% in auc@3 while reducing MACs by ~47% compared to the SOTA dense matching method. Additionally, it shows a 1.8× improvement in auc@3 over the SOTA deep homography method. Code is available at https://github.com/KN-Zhang/GFNet. Kaining Zhang, Yuxin Deng 0002, Jiayi Ma 0001, Paolo Favaro |
CVPR | 3 |
| 2025 | LaMamba: Layer-Aware Mamba with Frequency-Spatial Fusion for Remote Sensing Image Super-ResolutionabstractRecent advances in Vision Mamba architectures have shown great promise for remote sensing image super-resolution (RSISR), owing to their efficient state space modeling. However, existing Mamba-based methods face two critical challenges: inadequate cross-layer feature transmission and limited frequency-spatial modeling, both of which hinder the reconstruction of fine textures and complex scene structures. To address these issues, we propose LaMamba, a novel RSISR framework that jointly enhances hierarchical feature transmission and frequency-spatial modeling through two key modules. The Layer-aware Feature Integration (LaFI) module adaptively promotes inter-layer feature interaction, effectively preserving shallow spatial cues in deep networks. Meanwhile, the Spatial-Frequency Nonlinear Mapping (SFNM) module adopts a dual-domain fusion strategy by integrating a Gated Dual Feature Interaction (GDFI) and a Magnitude-Phase Adaptive Aggregation (MPAA). GDFI alleviates spatial-domain channel redundancy, while MPAA enhances frequency-domain representation by decoupling magnitude and phase components. The SFNM module then fuses spatial-frequency information via nonlinear mappings. Extensive experiments on six public RSISR benchmarks demonstrate that LaMamba achieves state-of-the-art performance while maintaining computational efficiency, highlighting its practicality for remote sensing applications. The source code is available at https://github.com/Lmy-0914/LaMamba. Chengyi Xiong, Zhirong Gao, Jiayi Ma 0001 |
ECAI | 4 |
| 2025 | ArgMatch: Adaptive Refinement Gathering for Efficient Dense Matching
Yuxin Deng 0002, Kaining Zhang, Linfeng Tang, Jiaqi Yang 0002, Jiayi Ma 0001 |
ICCV | 5 |
| 2025 | TemCoCo: Temporally Consistent Multi-Modal Video Fusion with Visual-Semantic Collaboration
Meiqi Gong, Hao Zhang 0073, Xunpeng Yi, Linfeng Tang, Jiayi Ma 0001 |
ICCV | 5 |
| 2025 | Balancing Task-Invariant Interaction and Task-Specific Adaptation for Unified Image FusionabstractUnified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem facilitates task-invariant knowledge sharing, it often overlooks task-specific characteristics, thereby limiting the overall performance. Existing general image fusion methods incorporate explicit task identification to enable adaptation to different fusion tasks. However, this dependence during inference restricts the model's generalization to unseen fusion tasks. To address these issues, we propose a novel unified image fusion framework named "TITA", which dynamically balances both Task-invariant Interaction and Task-specific Adaptation. For task-invariant interaction, we introduce the Interaction-enhanced Pixel Attention (IPA) module to enhance pixel-wise interactions for better multi-source complementary information extraction. For task-specific adaptation, the Operation-based Adaptive Fusion (OAF) module dynamically adjusts operation weights based on task properties. Additionally, we incorporate the Fast Adaptive Multitask Optimization (FAMO) strategy to mitigate the impact of gradient conflicts across tasks during joint training. Extensive experiments demonstrate that TITA not only achieves competitive performance compared to specialized methods across three image fusion scenarios but also exhibits strong generalization to unseen fusion tasks. The source codes are released at https://github.com/huxingyuabc/TITA. Junjun Jiang, Chenyang Wang 0002, Kui Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICCV | 6 |
| 2025 | CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image MatchingabstractThis prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling context interaction over the entire coarse feature map elicits highly redundant computation due to the neighboring representation similarity of tokens, a covisibility-guided token condenser is introduced to adaptively aggregate tokens in light of their covisibility scores that are dynamically estimated, thereby ensuring computational efficiency while improving the representational capacity of aggregated tokens simultaneously. Secondly, considering that feature interaction with massive non-covisible areas is distracting, which may degrade feature distinctiveness, a covisibility-assisted attention mechanism is deployed to selectively suppress irrelevant message broadcast from non-covisible reduced tokens, resulting in robust and compact attention to relevant rather than all ones. Thirdly, we find that at the fine-level stage, current methods adjust only the target view's keypoints to subpixel level, while those in the source view remain restricted at the coarse level and thus not informative enough, detrimental to keypoint location-sensitive usages. A simple yet potent fine correlation module is developed to refine the matching candidates in both source and target views to subpixel level, attaining attractive performance improvement. Thorough experimentation across an array of public benchmarks affirms CoMatch's promising accuracy, efficiency, and generalizability. Zizhuo Li, Linfeng Tang, Jiayi Ma 0001 |
ICCV | 5 |
| 2025 | Robust Test-Time Adaptation for Single Image Denoising Using Deep Gaussian Prior
Pengwei Liang, Jiayi Ma 0001, Junjun Jiang, Zhe Peng |
ICCV | 4 |
| 2025 | Deep Adaptive Unfolded Network via Spatial Morphology Stripping and Spectral Filtration for Pan-Sharpening
Hebaixu Wang, Jiayi Ma 0001 |
ICCV | 2 |
| 2025 | End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph Generation
Yanduo Zhang, Tao Lu 0001, Huiqin Zhang, Jiayi Ma 0001, Huabing Zhou |
ICCV | 6 |
| 2025 | LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesabstractCurrent advanced research on infrared and visible image fusion primarily focuses on improving fusion performance, often neglecting the applicability on real-time fusion devices. In this paper, we propose a novel approach that towards extremely fast fusion via distillation to learnable lookup tables specifically designed for image fusion, termed as LUT-Fuse. Firstly, we develop a look-up table structure that utilizing low-order approximation encoding and high-level joint contextual scene encoding, which is well-suited for multi-modal fusion. Moreover, given the lack of ground truth in multi-modal image fusion, we naturally proposed the efficient LUT distillation strategy instead of traditional quantization LUT methods. By integrating the performance of the multi-modal fusion network (MM-Net) into the MM-LUT model, our method achieves significant breakthroughs in efficiency and performance. It typically requires less than one-tenth of the time compared to the current lightweight SOTA fusion algorithms, ensuring high operational speed across various scenarios, even in low-power mobile devices. Extensive experiments validate the superiority, reliability, and stability of our fusion approach. The code is available at https://github.com/zyb5/LUT-Fuse. Xunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan, Han Xu 0001, Jiayi Ma 0001 |
ICCV | 6 |
| 2025 | HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D RegistrationabstractGeometric constraints between feature matches are critical in 3D point cloud registration problems. Existing approaches typically model unordered matches as a consistency graph and sample consistent matches to generate hypotheses. However, explicit graph construction introduces noise, posing great challenges for handcrafted geometric constraints to render consistency. To overcome this, we propose HyperGCT, a flexible dynamic Hyper-GNN-learned geometric ConstrainT that leverages high-order consistency among 3D correspondences. To our knowledge, HyperGCT is the first method that mines robust geometric constraints from dynamic hypergraphs for 3D registration. By dynamically optimizing the hypergraph through vertex and edge feature aggregation, HyperGCT effectively captures the correlations among correspondences, leading to accurate hypothesis generation. Extensive experiments on 3DMatch, 3DLoMatch, KITTI-LC, and ETH show that HyperGCT achieves state-of-the-art performance. Furthermore, HyperGCT is robust to graph noise, demonstrating a significant advantage in terms of generalization. Xiyu Zhang 0001, Jiayi Ma 0001, Zhaoshuai Qi, Fei Hui, Jiaqi Yang 0002, Yanning Zhang 0001 |
ICCV | 2 |
| 2025 | Multimodal Image Matching Based on Cross-Modality Completion Pre-trainingabstractThe differences in imaging devices cause multimodal images to have modal differences and geometric distortions, complicating the matching task. Deep learning-based matching methods struggle with multimodal images due to the lack of large annotated multimodal datasets. To address these challenges, we propose XCP-Match based on cross-modality completion pre-training. XCP-Match has two phases. (1) Self-supervised cross-modality completion pre-training based on real multimodal image dataset. We develop a novel pre-training model to learn cross-modal semantic features. The pre-training uses masked image modeling method for cross-modality completion, and introduces an attention-weighted contrastive loss to emphasize matching in overlapping areas. (2) Supervised fine-tuning for multimodal image matching based on the augmented MegaDepth dataset. XCP-Match constructs a complete matching framework to overcome geometric distortions and achieve precise matching. Two-phase training encourages the model to learn deep cross-modal semantic information, improving adaptation to modal differences without needing large annotated datasets. Experiments demonstrate that XCP-Match outperforms existing algorithms on public datasets. Meng Yang 0031, Fan Fan 0001, Jun Huang 0008, Yong Ma 0001, Xiaoguang Mei, Zhanchuan Cai, Jiayi Ma 0001 |
IJCAI | 7 |
| 2025 | Towards Perfection: Building Inter-component Mutual Correction for Retinex-based Low-light Image EnhancementabstractIn low-light image enhancement, Retinex-based deep learning methods have garnered significant attention due to their exceptional interpretability. These methods decompose images into mutually independent illumination and reflectance components, allows each component to be enhanced separately. In fact, achieving perfect decomposition of illumination and reflectance components proves to be quite challenging, with some residuals still existing after decomposition. In this paper, we formally name these residuals as inter-component residuals (ICR), which has been largely underestimated by previous methods. In our investigation, ICR not only affects the accuracy of the decomposition but also causes enhanced components to deviate from the ideal outcome, ultimately reducing the final synthesized image quality. To address this issue, we propose a novel Inter-correction Retinex model (IRetinex) to alleviate ICR during the decomposition and enhancement stage. In the decomposition stage, we leverage inter-component residual reduction module to reduce the feature similarity between illumination and reflectance components. In the enhancement stage, we utilize the feature similarity between the two components to detect and mitigate the impact of ICR within each enhancement unit. Extensive experiments on three low-light benchmark datasets demonstrated that by reducing ICR, our method outperforms state-of-the-art approaches both qualitatively and quantitatively. Our code is available at: https://github.com/caoluyang0830/IRetinex.git. Luyang Cao, Han Xu 0001, Jian Zhang 0090, Lei Qi 0001, Jiayi Ma 0001, Yinghuan Shi, Yang Gao 0001 |
ACM Multimedia | 5 |
| 2025 | CorrNeXt: Making the ConvNet-Style Correspondence Pruner Stronger for Two-View GeometryabstractThe uproar over two-view correspondence pruning stems from the advent of the ConvNet-style paradigm, which showcases intrinsic proficiency in local context aggregation, tackling the context-agnostic deficiency of MLP-based methods fundamentally and delivering impressive pruning capability. To further unlock the potential of such a paradigm, this perspective study revisits its design decisions and introduces CorrNeXt, a cutting-edge ConvNet-style pruner that incorporates multiple simple but effective improvements. Firstly, we explicitly integrate 2D relative spatial knowledge into motion field modeling, arming the interconversion between unordered sparse motion vectors and ordered image-structured ones with positional awareness. Secondly, considering that existing methods struggle with perceiving global context due to limited receptive field of small-kernel convolution, we devise a context-orthogonal aggregation module that decomposes computationally expensive large-kernel depthwise convolution along channel dimension into a small square kernel, two orthogonal band kernels, and an identity mapping, enjoying large receptive field while maintaining efficiency. Thirdly, we deploy a motion field pyramid architecture that obtains and fuses multi-level motion fields, thereby facilitating the handling of the motion field's discontinuities in case of large scene disparity. Ultimately, we propose an elastic inference strategy that allows the model to introspect the confidence of its predictions at each layer, through which CorrNeXt is endowed with the flexibility of adaptively determining inference termination according to the difficulty of each image pair. Thorough experimentation affirms CorrNeXt's remarkable capabilities. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
ACM Multimedia | 5 |
| 2025 | Projection-Manifold Regularized Latent Diffusion for Robust General Image FusionabstractThis study proposes PDFuse, a robust, general training-free image fusion framework built on pre-trained latent diffusion models with projection–manifold regularization. By redefining fusion as a diffusion inference process constrained by multiple source images, PDFuse can adapt to varied image modalities and produce high-fidelity outputs utilizing the diffusion prior. To ensure both source consistency and full utilization of generative priors, we develop novel projection–manifold regularization, which consists of two core mechanisms. On the one hand, the Multi-source Information Consistency Projection (MICP) establishes a projection system between diffusion latent representations and source images, solved efficiently via conjugate gradients to inject multi-source information into the inference. On the other hand, the Latent Manifold-preservation Guidance (LMG) aligns the latent distribution of diffusion variables with that of the sources, guiding generation to respect the model’s manifold prior. By alternating these mechanisms, PDFuse strikes an optimal balance between fidelity and generative quality, achieving superior fusion performance across diverse tasks. Moreover, PDFuse constructs a canonical interference operator set. It synergistically incorporates it into the aforementioned dual mechanisms, effectively leveraging generative priors to address various degradation issues during the fusion process without requiring clean data for supervising training. Extensive experimental evidence substantiates that PDFuse achieves highly competitive performance across diverse image fusion tasks. The code is publicly available at https://github.com/Leiii-Cao/PDFuse. Hao Zhang 0073, Jiayi Ma 0001 |
NeurIPS | 4 |
| 2025 | ControlFusion: A Controllable Image Fusion Network with Language-Vision Degradation PromptsabstractCurrent image fusion methods struggle with real-world composite degradations and lack the flexibility to accommodate user-specific needs. To address this, we propose ControlFusion, a controllable fusion network guided by language-vision prompts that adaptively mitigates composite degradations. On the one hand, we construct a degraded imaging model based on physical mechanisms, such as the Retinex theory and atmospheric scattering principle, to simulate composite degradations and provide a data foundation for addressing realistic degradations. On the other hand, we devise a prompt-modulated restoration and fusion network that dynamically enhances features according to degradation prompts, enabling adaptability to varying degradation levels. To support user-specific preferences in visual quality, a text encoder is incorporated to embed user-defined degradation types and levels as degradation prompts. Moreover, a spatial-frequency collaborative visual adapter is designed to autonomously perceive degradations from source images, thereby reducing complete reliance on user instructions. Extensive experiments demonstrate that ControlFusion outperforms SOTA fusion methods in fusion quality and degradation handling, particularly under real-world and compound degradations. Linfeng Tang, Yeda Wang, Zhanchuan Cai, Junjun Jiang, Jiayi Ma 0001 |
NeurIPS | 5 |
| 2025 | DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image RestorationabstractDiffusion models have achieved remarkable progress in universal image restoration. However, existing methods perform naive inference in the reverse process, which leads to cumulative errors under limited sampling steps and large step intervals. Moreover, they struggle to balance the commonality of degradation representations with restoration quality, often depending on complex compensation mechanisms that enhance fidelity at the expense of efficiency. To address these challenges, we introduce \textbf{DGSolver}, a diffusion generalist solver with universal posterior sampling. We first derive the exact ordinary differential equations for generalist diffusion models to unify degradation representations and design tailored high-order solvers with a queue-based accelerated sampling strategy to improve both accuracy and efficiency. We then integrate universal posterior sampling to better approximate manifold-constrained gradients, yielding a more accurate noise estimation and correcting errors in inverse inference. Extensive experiments demonstrate that DGSolver outperforms state-of-the-art methods in restoration accuracy, stability, and scalability, both qualitatively and quantitatively. Code and models are publicly available at https://github.com/MiliLab/DGSolver. Hebaixu Wang, Jing Zhang 0037, Di Wang 0023, Jiayi Ma 0001, Bo Du 0001 |
NeurIPS | 5 |
| 2025 | Deno-IF: Unsupervised Noisy Visible and Infrared Image Fusion MethodabstractMost image fusion methods are designed for ideal scenarios and struggle to handle noise. Existing noise-aware fusion methods are supervised and heavily rely on constructed paired data, limiting performance and generalization. This paper proposes a novel unsupervised noisy visible and infrared image fusion method, comprising two key modules. First, when only noisy source images are available, a convolutional low-rank optimization module decomposes clean components based on convolutional low-rank priors, guiding subsequent optimization. The unsupervised approach eliminates data dependency and enhances generalization across various and variable noise. Second, a unified network jointly realizes denoising and fusion. It consists of both intra-modal recovery and inter-modal recovery and fusion, also with a convolutional low-rankness loss for regularization. By exploiting the commonalities of denoising and fusion, the joint framework significantly reduces network complexity while expanding functionality. Extensive experiments validate the effectiveness and generalization of the proposed method for image fusion under various and variable noise conditions. The code is publicly available at https://github.com/hanna-xu/Deno-IF. Han Xu 0001, Yuyang Li 0005, Yunfei Deng, Jiayi Ma 0001, Guangcan Liu |
NeurIPS | 4 |
| 2025 | Multi-receptive field interaction network for shape from polarization
Yini Peng, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
Sci. China Inf. Sci. | 5 |
| 2025 | DiffuseDoc: Document geometric rectification via diffusion model
Wenfei Xiong, Huabing Zhou, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 5 |
| 2025 | SD-MIL: Multiple instance learning with dual perception of scale and distance information fusion for whole slide image classification
Yining Xie, Zequn Liu, Wei Zhang 0259, Jiayi Ma 0001 |
Expert Syst. Appl. | 6 |
| 2025 | Feature Matching via Graph Clustering with Local Affine Consensus
Jiayi Ma 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | C2RF: Bridging Multi-modal Image Registration and Fusion via Commonality Mining and Contrastive Learning
Linfeng Tang, Qinglong Yan, Xinyu Xiang, Leyuan Fang, Jiayi Ma 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Pixel2Pixel: A Pixelwise Approach for Zero-Shot Single Image DenoisingabstractWe propose Pixel2Pixel, a novel zero-shot image denoising framework that leverages the non-local self-similarity of images to generate a large number of training samples using only the input noisy image. This framework employs a compact convolutional neural network architecture to achieve high-quality image denoising. Given a single observed noisy image, we first aim to obtain multiple images with different noise versions. We ensure that the content remains as consistent as possible with the true signal of the noisy image while keeping the noise independent. Specifically, we construct a pixel bank tensor, where each pixel consists of the most similar pixels from the non-local region of the noisy image. Then, multiple training samples, also known as pseudo instances, can be derived from the pixel bank by randomly pixel sampling. By harnessing pixel-wise random sampling, Pixel2Pixel generates a large number of training pseudo instances, thus avoiding reliance on specific training data. In addition, this non-local pixel selection and random sampling strategy helps to break down the spatial correlation of real-world noise as well. Since the proposed method does not require accurate priors on the noise distribution and clean training images, it is suitable for a wide range of noise types and different noise levels, exhibiting strong generalization ability, especially in real noisy scenes. Extensive experiments across various noise types show that Pixel2Pixel outperforms existing methods. Junjun Jiang, Pengwei Liang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Modeling the Label Distributions for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models by weak labels, which is receiving significant attention due to its low annotation cost. Existing approaches focus on generating pseudo labels for supervision while largely ignoring to leverage the inherent semantic correlation among different pseudo labels. We observe that pseudo-labeled pixels that are close to each other in the feature space are more likely to share the same class, and those closer to the distribution centers tend to have higher confidence. Motivated by this, we propose to model the underlying label distributions and employ cross-label constraints to generate more accurate pseudo labels. In this paper, we develop a unified WSSS framework named Adaptive Gaussian Mixtures Model, which leverages a GMM to model the label distributions. Specifically, we calculate the feature distribution centers of pseudo-labeled pixels and build the GMM by measuring the distance between the centers and each pseudo-labeled pixel. Then, we introduce an Online Expectation-Maximization (OEM) algorithm and a novel maximization loss to optimize the GMM adaptively, aiming to learn more discriminative decision boundaries between different class-wise Gaussian mixtures. Based on the label distributions, we leverage the GMM to generate high-quality pseudo labels for more reliable supervision. Our framework is capable of solving different forms of weak labels: image-level labels, points, scribbles, blocks, and bounding-boxes. Extensive experiments on PASCAL, COCO, Cityscapes, and ADE20 K datasets demonstrate that our framework can effectively provide more reliable supervision and outperform the state-of-the-art methods under all settings. Linshan Wu, Zhun Zhong, Jiayi Ma 0001, Yunchao Wei, Hao Chen 0011, Leyuan Fang, Shutao Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Diff-Retinex++: Retinex-Driven Reinforced Diffusion Model for Low-Light Image EnhancementabstractThis paper proposes a Retinex-driven reinforced diffusion model for low-light image enhancement, termed Diff-Retinex++, to address various degradations caused by low light. Our main approach integrates the diffusion model with Retinex-driven restoration to achieve physically-inspired generative enhancement, making it a pioneering effort. To be detailed, Diff-Retinex++ consists of two-stage view modules, including the Denoising Diffusion Model (DDM), and the Retinex-Driven Mixture of Experts Model (RMoE). First, DDM treats low-light image enhancement as one type of image generation task, benefiting from the powerful generation ability of diffusion model to handle the enhancement. Second, we design the Retinex theory into the plug-and-play supervision attention module. It leverages the latent features in the backbone and knowledge distillation to learn Retinex rules, and further regulates these latent features through the attention mechanism. In this way, it couples the relationship between Retinex decomposition and image enhancement in a new view, achieving dual improvement. In addition, the Low-Light Mixture of Experts preserves the vividness of the diffusion model and fidelity of the Retinex-driven restoration to the greatest extent. Ultimately, the iteration of DDM and RMoE achieves the goal of Retinex-driven reinforced diffusion model. Extensive experiments conducted on real-world low-light datasets qualitatively and quantitatively demonstrate the effectiveness, superiority, and generalization of the proposed method. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | OmniFuse: Composite Degradation-Robust Image Fusion With Language-Driven SemanticsabstractExisting image fusion methods struggle to accommodate composite degradation and do not support users flexibly modulating the semantic objects of interest. To address these challenges, this study proposes a composite degradation-robust image fusion framework with language-driven semantics, called OmniFuse. Firstly, OmniFuse establishes a novel multi-modal information fusion paradigm based on the latent diffusion model (LDM). By projecting the information fusion function into the latent space of the LDM, the information fusion process is seamlessly integrated with the diffusion process. Thus, OmniFuse fully leverages the powerful generative capabilities of LDM to eliminate composite degradation, thereby achieving highly robust image fusion. Secondly, OmniFuse develops a language-driven controllable fusion strategy to strengthen fusion flexibility. It employs a language-driven feature fusion module (LFFM) to receive the specified localization priori, dynamically aggregating multi-modal features. Within LFFM, a visual enhancement regularization is introduced to highlight objects of interest for capturing perceptual attention, while reverse semantic driving is established to strengthen their semantic attributes. Together, the visual and semantic constraints can implicitly correct the imperfect localization priori, further refining the accuracy of language-driven control. Extensive experiments demonstrate the omnipotent performance of OmniFuse, with significant advantages in robustness and flexibility compared to state-of-the-art methods. Hao Zhang 0073, Xuhui Zuo, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | DeMatch++: Two-View Correspondence Learning via Deep Motion Field Decomposition and Respective Local-Context AggregationabstractTwo-view correspondence learning has increasingly focused on the coherence and smoothness of motion fields between image pairs. Conventional methods either regularize the complexity of the field function at substantial computational expense, or apply local filters that prove ineffective for large scene disparities. In this paper, we present DeMatch++, a novel network drawing inspiration from Fourier decomposition principles that decomposes the motion field to retain its primary "low-frequency" and smooth components. This approach achieves implicit regularization with lower computational overhead while exhibiting inherent piecewise smoothness. Specifically, our method decomposes the noise-contaminated motion field into multiple linearly independent basis vectors, generating smooth sub-fields that preserve the main energy of the original field. These sub-fields facilitate the recovery of a cleaner motion field for precise vector derivation. Within this framework, we aggregate local context within each sub-field while enhancing global information across all sub-fields. We also employ a masked decomposition strategy that mitigates the influence of false matches, and construct a compact representation to suppress redundant sub-fields. The complete pipeline is formulated as a discrete learnable architecture, circumventing the need for dense field computation. Extensive experiments demonstrate that DeMatch++ outperforms state-of-the-art methods while maintaining computational efficiency and piecewise smoothness. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Deep blind super-resolution for hyperspectral images
Yong Ma 0001, Xiaoguang Mei, Qihai Chen, Minghui Wu 0007, Jiayi Ma 0001 |
Pattern Recognit. | 6 |
| 2025 | Hierarchical diffusion models for generating various pattern vehicles in infrared aerial images
Nan Zhang 0030, Youmeng Liu, Hao Liu 0064, Tian Tian 0006, Jiayi Ma 0001, Jinwen Tian |
Pattern Recognit. | 5 |
| 2025 | An Infrared and Visible Image Fusion Method Based on Semantic-Sensitive Mask Selection and Bidirectional-Collaboration Region FusionabstractMask is considered as an important prior for fusion, which could selectively enhance specific regions to generate ideal fused images. However, masks used in the existing methods exhibit limitations in the precise representation of targets, and more importantly, these masks are generated from a single modality, which restricts the effective integration of multi-modal information. To address this issue, we propose a competitive mask-guidance fusion method for infrared and visible images. A multi-modal semantic-sensitive mask selection network is proposed to generate complementary-mask maps, which organically integrate advantageous target regions of different modalities by competitively comparing the qualities of masks. In this network, a pseudosiamese architecture is designed to obtain respective target masks, and specifically, a spatial-aligned-based feature aggregation module is devised to produce high-quality pseudo-labels which are served as references for the generation of the complementary-mask maps. Furthermore, we propose a bidirectional-collaboration region fusion strategy, which enhances the expression of advantageous target regions from each modality inforeground while suppressing the contribution of corresponding regions from the other modality in background. Compared to methods on public datasets, the results show that our method significantly enhances the description of semantic-sensitive targets in fused images, including the saliency and the integrity of structural information. Code are available athttps://github.com/xbsj-cool/MSCRFusion. Guomin Zhang, Yining Xie, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | SDSFusion: A Semantic-Aware Infrared and Visible Image Fusion Network for Degraded ScenesabstractA single-modal infrared or visible image offers limited representation in scenes with lighting degradation or extreme weather. We propose a multi-modal fusion framework, named SDSFusion, for all-day and all-weather infrared and visible image fusion. SDSFusion exploits the commonality in image processing to achieve enhancement, fusion, and semantic task interaction in a unified framework guided by semantic awareness and multi-scale features and losses. To address the disparity between infrared and visible images in degraded scenes, we differentiate modal features in a unified fusion model. Unlike existing joint fusion methods, we propose an adversarial generative network that refines the reconstruction of low-light images by embedding fused features. It provides feature-level brightness supplementation and image reconstruction to refine brightness and contrast. Extensive experiments in degraded scenes confirm that our approach is superior to state-of-the-art approaches in visual quality and performance, demonstrating the effectiveness of interaction improvement. The code will be posted at: https://github.com/Liling-yang/SDSFusion. Jun Chen 0019, Liling Yang, Wei Yu 0018, Wenping Gong, Zhanchuan Cai, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | MaeFuse: Transferring Omni Features With Pretrained Masked Autoencoders for Infrared and Visible Image Fusion via Guided TrainingabstractIn this paper, we introduce MaeFuse, a novel autoencoder model designed for Infrared and Visible Image Fusion (IVIF). The existing approaches for image fusion often rely on training combined with downstream tasks to obtain high-level visual information, which is effective in emphasizing target objects and delivering impressive results in visual quality and task-specific applications. Instead of being driven by downstream tasks, our model called MaeFuse utilizes a pretrained encoder from Masked Autoencoders (MAE), which facilities the omni features extraction for low-level reconstruction and high-level vision tasks, to obtain perception friendly features with a low cost. In order to eliminate the domain gap of different modal features and the block effect caused by the MAE encoder, we further develop a guided training strategy. This strategy is meticulously crafted to ensure that the fusion layer seamlessly adjusts to the feature space of the encoder, gradually enhancing the fusion performance. The proposed method can facilitate the comprehensive integration of feature vectors from both infrared and visible modalities, thus preserving the rich details inherent in each modal. MaeFuse not only introduces a novel perspective in the realm of fusion techniques but also stands out with impressive performance across various public datasets. The code is available at https://github.com/Henry-Lee-real/MaeFuse. Jiayang Li 0004, Junjun Jiang, Pengwei Liang, Jiayi Ma 0001, Liqiang Nie |
IEEE Trans. Image Process. | 4 |
| 2025 | Learning Feature Matching via Matchable Keypoint-Assisted Graph Neural NetworkabstractAccurately matching local features between a pair of images corresponding to the same 3D scene is a challenging computer vision task. Previous studies typically utilize attention-based graph neural networks (GNNs) with fully-connected graphs over keypoints within/across images for visual and geometric information reasoning. However, in the background of local feature matching, a significant number of keypoints are non-repeatable due to factors like occlusion and failure of the detector, and thus irrelevant for message passing. The connectivity with non-repeatable keypoints not only introduces redundancy, resulting in limited efficiency (quadratic computational complexity w.r.t. the keypoint number), but also interferes with the representation aggregation process, leading to limited accuracy. Aiming at the best of both worlds on accuracy and efficiency, we propose MaKeGNN, a sparse attention-based GNN architecture which bypasses non-repeatable keypoints and leverages matchable ones to guide compact and meaningful message passing. More specifically, our Bilateral Context-Aware Sampling (BCAS) Module first dynamically samples two small sets of well-distributed keypoints with high matchability scores from the image pair. Then, our Matchable Keypoint-Assisted Context Aggregation (MKACA) Module regards sampled informative keypoints as message bottlenecks and thus constrains each keypoint only to retrieve favorable contextual information from intra- and inter-matchable keypoints, evading the interference of irrelevant and redundant connectivity with non-repeatable ones. Furthermore, considering the potential noise in initial keypoints and sampled matchable ones, the MKACA module adopts a matchability-guided attentional aggregation operation for purer data-dependent context propagation. By these means, MaKeGNN outperforms the state-of-the-arts on multiple highly challenging benchmarks, while significantly reducing computational and memory complexity compared to typical attentional GNNs. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | FusionINV: A Diffusion-Based Approach for Multimodal Image FusionabstractInfrared images exhibit a significantly different appearance compared to visible counterparts. Existing infrared and visible image fusion (IVF) methods fuse features from both infrared and visible images, producing a new "image" appearance not inherently captured by any existing device. From an appearance perspective, infrared, visible, and fused images belong to different data domains. This difference makes it challenging to apply fused images because their domain-specific appearance may be difficult for downstream systems, e.g., pre-trained segmentation models. Therefore, accurately assessing the quality of the fused image is challenging. To address those problem, we propose a novel IVF method, FusionINV, which produces fused images with an appearance similar to visible images. FusionINV employs the pre-trained Stable Diffusion (SD) model to invert infrared images into the noise feature space. To inject visible-style appearance information into the infrared features, we leverage the inverted features from visible images to guide this inversion process. In this way, we can embed all the information of infrared and visible images in the noise feature space, and then use the prior of the pre-trained SD model to generate visually friendly images that align more closely with the RGB distribution. Specially, to generate the fused image, we design a tailored fusion rule within the denoising process that iteratively fuses visible-style infrared and visible features. In this way, the fused image falls into the visible domain and can be directly applied to existing downstream machine systems. Thanks to advancements in image inversion, FusionINV can directly produce fused images in a training-free manner. Extensive experiments demonstrate that FusionINV achieves outstanding performance in both human visual evaluation and machine perception tasks. The code is available at https://github.com/erfect2020/FusionINV. Pengwei Liang, Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Mutually Reinforcing Learning of Decoupled Degradation and Diffusion Enhancement for Unpaired Low-Light Image LighteningabstractDenoising Diffusion Probabilistic Model (DDPM) has demonstrated exceptional performance in low-light enhancement task. However, the dependency on paired training datas has left the generality of DDPM in low-light enhancement largely untapped. Therefore, this paper proposes a mutually reinforcing learning framework of decoupled degradation and diffusion enhancement, named MRLIE, which leverages style guidance from unpaired low-light images to generate pseudo-image pairs that are consistent with the target domain, thereby optimizing the latter diffusion enhancement network in a supervised manner. During the degradation process, the diffusion loss of fixed enhancement network serves as a evaluation metric for structure consistency and is combined with adversarial style loss to form the optimization objective for degradation network. Such loss design ensures that scene structure information is retained during the degradation process. During the enhancement process, the degradation network with frozen parameters continuously generates pseudo-paired low-/normal-light image pairs as training datas, thus the diffusion enhancement network could be progressively optimized. On the whole, the two processes are interdependent and could achieve cooperative improvement in terms of degradation realism and enhancement quality through iterative optimization. Additionally, we propose the Retinex-based decoupled degradation strategy for simulating the complex degradation in real low-light imaging, which ensures the color correction and noise suppression capabilities of latter diffusion enhancement network. Extensive experiments show that MRLIE can achieve promising results and better generality across various datasets. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Grid-Guided Sparse Laplacian Consensus for Robust Feature MatchingabstractFeature matching is a fundamental concern widely employed in computer vision applications. This paper introduces a novel and efficacious method named Grid-guided Sparse Laplacian Consensus, rooted in the concept of smooth constraints. To address challenging scenes such as severe deformation and independent motions, we devise grid-based adaptive matching guidance to construct multiple transformations based on motion coherence. Specifically, we obtain a set of precise yet sparse seed correspondences through motion statistics, facilitating the generation of an adaptive number of candidate correspondence sets. In addition, we propose an innovative formulation grounded in graph Laplacian for correspondence pruning, wherein mapping function estimation is formulated as a Bayesian model. We solve this utilizing EM algorithm with seed correspondences as initialization for optimal convergence. Sparse approximation is leveraged to reduce the time-space burden. A comprehensive set of experiments are conducted to demonstrate the superiority of our method over other state-of-the-art methods in both robustness to serious deformations and generalizability for various descriptors, as well as generalizability to multi motions. Additionally, experiments in geometric estimation, image registration, loop closure detection, and visual localization highlight the significance of our method across diverse scenes for high-level tasks. Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | URFusion: Unsupervised Unified Degradation-Robust Image Fusion NetworkabstractWhen dealing with low-quality source images, existing image fusion methods either fail to handle degradations or are restricted to specific degradations. This study proposes an unsupervised unified degradation-robust image fusion network, termed as URFusion, in which various types of degradations can be uniformly eliminated during the fusion process, leading to high-quality fused images. URFusion is composed of three core modules: intrinsic content extraction, intrinsic content fusion, and appearance representation learning and assignment. It first extracts degradation-free intrinsic content features from images affected by various degradations. These content features then provide feature-level rather than image-level fusion constraints for optimizing the fusion network, effectively eliminating degradation residues and reliance on ground truth. Finally, URFusion learns the appearance representation of images and assigns the statistical appearance representation of high-quality images to the content-fused result, producing the final high-quality fused image. Extensive experiments on multi-exposure image fusion and multi-modal image fusion tasks demonstrate the advantages of URFusion in fusion performance and suppression of multiple types of degradations. The code is available at https://github.com/hanna-xu/URFusion. Han Xu 0001, Xunpeng Yi, Chen Lu 0004, Guangcan Liu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Universal Infrared Image Nonuniformity Correction via Stripe-Aware Attention NetworkabstractInfrared image nonuniformity correction aims to remove the column-wise stripe noise. Most existing methods just consider stripe noise whereas failing to handle real captured nonuniformity, as directional characteristic of stripe is severely disrupted by random Gaussian noise. Moreover, deep learning-based methods proposed in recent years are blocked by limited receptive field thus cannot accurately distinguish vertical structure and vertical stripes. To address these issues, we propose a universal infrared image nonuniformity correction method based on stripe-aware attention network. We seek to improve the performance of our algorithm by first restoring the damaged stripe directional characteristics, then maximizing the utilization of the prior characteristics. On the one hand, we construct the two-stage framework, in which denoising network is firstly applied to eliminate Gaussian noise and preserve stripes as scene information. As a result, the prior directional characteristics are restored, thereby enhancing the ability of subsequent sub-network to perceive stripe noise. On the other hand, due to the distinct long-range pixel correlations of vertical structures and vertical textures, we introduce a column-wise stripe attention mechanism (CSA) that can capture long-range dependencies of target pixels in the vertical direction. This significantly improves the discriminative ability of algorithm towards vertical structures and stripes, with minimal computational cost. Extensive experiments show that the proposed method can achieve promising results and has better universality for different infrared scenarios. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | VRTNet: Vector Rectifier Transformer for Two-View Correspondence LearningabstractFinding reliable correspondences in two-view image and recovering the camera poses are key problems in photogrammetry and image signal processing. Multilayer perceptron (MLP) has a wide application in two-view correspondence learning for which is good at learning disordered sparse correspondences, but it is susceptible to the dominant outliers and requires additional functional blocks to capture context information. CNN can naturally extract local context information, but it cannot handle disordered data and extract global context and channel information. In order to overcome the shortcomings of MLP and CNN, we design a correspondence learning network based on Transformer, named Vector Rectifier Transformer (VRTNet). Transformer is an encoder-decoder structure which can handle disordered sparse correspondences and output sequences of arbitrary length. Therefore, we design two sub-Transformers in VRTNet to achieve the mutual conversion between disordered and ordered correspondences. The self-attention and cross-attention mechanisms in them allow VRTNet to focus on the global context relations of all correspondences. To capture local context and channel information, we propose rectifier network (including CNN and channel attention block) as the backbone of VRTNet, which avoids the complex design of additional blocks. Rectifier network can correct the errors of ordered correspondences to obtain rectified correspondences. Finally, outliers are removed by comparing original and rectified correspondences. VRTNet performs better than the state-of-the-art methods in the tasks of relative pose estimation, outlier removal and image registration. Meng Yang 0031, Jun Chen 0019, Xin Tian 0006, Longsheng Wei, Jiayi Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | General Hyperspectral Image Super-Resolution via Meta-Transfer LearningabstractRecent advances in deep learning-based methods have led to significant progress in the hyperspectral super-resolution (SR). However, the scarcity and the high dimension of data have hindered further development since deep models require sufficient data to learn stable patterns. Moreover, the huge domain differences between hyperspectral image (HSI) datasets pose a significant challenge in generalizability. To address these problems, we present a general hyperspectral SR framework via meta-transfer learning (MTL). We randomly sample various spectral ranges for SR tasks during MTL, allowing the model to accumulate diverse task experiences. Additionally, we implement a task schedule to gradually expand the number of bands, bridging the significant domain differences between datasets. By leveraging multiple datasets, we are able to achieve better performance and greater generalizability, making it applicable under various circumstances. Meanwhile, as a general framework, our scheme can be applied to existing methods to obtain performance improvements. In addition, we design an advanced network architecture based on the multifusion features to further improve the performance. Experiments demonstrate that our method not only achieves superior performance in both qualitative and quantitative terms but also can adapt robustly to a new and difficult sample, where few epochs can yield quite considerable results. Yingsong Cheng, Xinya Wang, Yong Ma 0001, Xiaoguang Mei, Minghui Wu 0007, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Seed to Prune: A Seeded Graph Neural Network for Two-View Correspondence LearningabstractWe present a simple yet tough-to-beat method dubbed SGNNet, for correspondence learning. Instead of focusing on devising sophisticated geometric extractors to explore the global or local contextual information involving all sparse correspondences as most existing studies have done, which may be biased by heavy outliers, we propose to first delve into elaborate contextual information encoded in several specific reliable correspondences, and later leverage it to achieve per-correspondence representation updating. To this end, the proposed network contains three pivotal modules: 1) dynamic seeding module, which aims to dynamically sample a set of reliable matches from the putative set as seeds to guide the network learning; 2) intraseed attention module (ISAM), which intends to capture the geometrical relations among seed matches and further leverage them to enhance seed features; and 3) dynamic unseeding module, which is designed to sufficiently aggregate favorable contextual information from seed matches and broadcast it back to features of original matches. With all the aforementioned components, the proposed SGNNet is capable of rejecting outliers from putative correspondences effectively. Extensive experiments indicate that our method beats current solid baselines and sets new SOTA scores across multiple domains and datasets. Notably, SGNNet attains an AUC@5° of 56.43% on YFCC100M without RANSAC, surpassing the most cutting-edge model by 4.51 absolute percentage points and exceeding the 55% AUC@5° bar for the first time. Project page: https://github.com/ZizhuoLi/SGNNet. Zizhuo Li, Jie Jiang 0015, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Image Deblurring by Exploring In-Depth Properties of TransformerabstractImage deblurring continues to achieve impressive performance with the development of generative models. Nonetheless, there still remains a displeasing problem if one wants to improve perceptual quality and quantitative scores of recovered image at the same time. In this study, drawing inspiration from the research of transformer properties, we introduce the pretrained transformers to address this problem. In particular, we leverage deep features extracted from a pretrained vision transformer (ViT) to encourage recovered images to be sharp without sacrificing the performance measured by the quantitative metrics. The pretrained transformer can capture the global topological relations (i.e., self-similarity) of image, and we observe that the captured topological relationships about the sharp image will change when blur occurs. By comparing the transformer features between recovered image and target one, the pretrained transformer provides high-resolution blur-sensitive semantic information, which is critical in measuring the sharpness of the deblurred image. On the basis of the advantages, we present two types of novel perceptual losses to guide image deblurring. One regards the features as vectors and computes the discrepancy between representations extracted from recovered image and target one in Euclidean space. The other type considers the features extracted from an image as a distribution and compares the distribution discrepancy between recovered image and target one. We demonstrate the effectiveness of transformer properties in improving the perceptual quality while not sacrificing the quantitative scores peak signal-to-noise ratio (PSNR) over the most competitive models, such as Uformer, Restormer, and NAFNet, on defocus deblurring and motion deblurring tasks. The code is available at https://github. com/erfect2020/TransformerPerceptualLoss. Pengwei Liang, Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | SDGMNet: Statistic-Based Dynamic Gradient Modulation for Local Descriptor LearningabstractRescaling the backpropagated gradient of contrastive loss has made significant progress in descriptor learning. However, current gradient modulation strategies have no regard for the varying distribution of global gradients, so they would suffer from changes in training phases or datasets. In this paper, we propose a dynamic gradient modulation, named SDGMNet, for contrastive local descriptor learning. The core of our method is formulating modulation functions with dynamically estimated statistical characteristics. Firstly, we introduce angle for distance measure after deep analysis on backpropagation of pair-wise loss. On this basis, auto-focus modulation is employed to moderate the impact of statistically uncommon individual pairs in stochastic gradient descent optimization; probabilistic margin cuts off the gradients of proportional triplets that have achieved enough optimization; power adjustment balances the total weights of negative pairs and positive pairs. Extensive experiments demonstrate that our novel descriptor surpasses previous state-of-the-art methods in several tasks including patch verification, retrieval, pose estimation, and 3D reconstruction. Yuxin Deng 0002, Jiayi Ma 0001 |
AAAI | 2 |
| 2024 | ResMatch: Residual Attention Learning for Feature MatchingabstractAttention-based graph neural networks have made great progress in feature matching. However, the literature lacks a comprehensive understanding of how the attention mechanism operates for feature matching. In this paper, we rethink cross- and self-attention from the viewpoint of traditional feature matching and filtering. To facilitate the learning of matching and filtering, we incorporate the similarity of descriptors into cross-attention and relative positions into self-attention. In this way, the attention can concentrate on learning residual matching and filtering functions with reference to the basic functions of measuring visual and spatial correlation. Moreover, we leverage descriptor similarity and relative positions to extract inter- and intra-neighbors. Then sparse attention for each point can be performed only within its neighborhoods to acquire higher computation efficiency. Extensive experiments, including feature matching, pose estimation and visual localization, confirm the superiority of the proposed method. Our codes are available at https://github.com/ACuOoOoO/ResMatch. Yuxin Deng 0002, Kaining Zhang, Yansheng Li 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2024 | Deep Unfolded Network with Intrinsic Supervision for Pan-SharpeningabstractExisting deep pan-sharpening methods lack the learning of complementary information between PAN and MS modalities in the intermediate layers, and exhibit low interpretability due to their black-box designs. To this end, an interpretable deep unfolded network with intrinsic supervision for pan-sharpening is proposed. Building upon the observation degradation process, it formulates the pan-sharpening task as a variational model minimization with spatial consistency prior and spectral projection prior. The former prior requires a joint component decomposition of PAN and MS images to extract intrinsic features. By being supervised in the intermediate layers, it can selectively provide high-frequency information for spatial enhancement. The latter prior constrains the intensity correlation between MS and PAN images derived from physical observations, so as to improve spectral fidelity. To further enhance the transparency of network design, we develop an iterative solution algorithm following the half-quadratic splitting to unfold the deep model. It rigorously adheres to the variational model, significantly enhancing the interpretability behind network design and efficiently alternating the optimization of the network. Extensive experiments demonstrate the advantages of our method compared to state-of-the-arts, showcasing its remarkable generalization capability to real-world scenes. Our code is publicly available at https://github.com/Baixuzx7/DISPNet. Hebaixu Wang, Meiqi Gong, Xiaoguang Mei, Hao Zhang 0073, Jiayi Ma 0001 |
AAAI | 5 |
| 2024 | Locality Preserving Refinement for Shape Matching with Functional MapsabstractIn this paper, we address the nonrigid shape matching with outliers by a novel and effective pointwise map refinement method, termed Locality Preserving Refinement. For accurate pointwise conversion from a given functional map, our method formulates a two-step procedure. Firstly, starting with noisy point-to-point correspondences, we identify inliers by leveraging the neighborhood support, which yields a closed-form solution with linear time complexity. After obtained the reliable correspondences of inliers, we refine the pointwise correspondences for outliers using local linear embedding, which operates in an adaptive spectral similarity space to further eliminate the ambiguities that are difficult to handle in the functional space. By refining pointwise correspondences with local consistency thus embedding geometric constraints into functional spaces, our method achieves considerable improvement in accuracy with linearithmic time and space cost. Extensive experiments on public benchmarks demonstrate the superiority of our method over the state-of-the-art methods. Our code is publicly available at https://github.com/XiaYifan1999/LOPR. Yuan Gao 0015, Jiayi Ma 0001 |
AAAI | 4 |
| 2024 | A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic ConsistencyabstractThis work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Through the coupled optimization of the visual fusion loss and the lesion segmentation loss, visual-related and semantic-related features will be pulled into the same domain, effectively promoting accuracy improvement in a mutual-reinforcing manner. Further, we establish the robustness guarantees by constructing a two-level refinement constraint in the process of feature extraction and reconstruction. Benefiting from full consideration for common degradations in medical images, our framework can not only provide clear visual fusion results for doctor's observation, but also enhance the defense ability of lesion segmentation against these negatives. Extensive evaluations of visual fusion and lesion segmentation scenarios demonstrate the advantages of our method in terms of accuracy and robustness. Moreover, our proposed framework is generic, which can be well-compatible with existing lesion segmentation algorithms and improve their performance. The code is publicly available at https://github.com/HaoZhang1018/RMR-Fusion. Hao Zhang 0073, Xuhui Zuo, Huabing Zhou, Tao Lu 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2024 | Unmixing Before Fusion: A Generalized Paradigm for Multi-Source-Based Hyperspectral Image SynthesisabstractIn the realm of AI, data serves as a pivotal resource. Real-world hyperspectral images (HSIs), bearing wide spectral characteristics, are particularly valuable. However, the acquisition of HSIs is always costly and time-intensive, resulting in a severe data-thirsty issue in HSI research and applications. Current solutions have not been able to generate a sufficient volume of diverse and reliable synthetic HSIs. To this end, our study formulates a novel, generalized paradigm for HSI synthesis, i.e., unmixing before fusion, that initiates with unmixing across multi-source data and follows by fusion-based synthesis. By integrating unmixing, this work maps unpaired HSI and RGB data to a low-dimensional abundance space, greatly alleviating the difficulty of generating high-dimensional samples. Moreover, incorporating abundances inferred from unpaired RGB images into generative models allows for cost-effective supplementation of various realistic spatial distributions in abundance synthesis. Our proposed paradigm can be instrumental with a series of deep generative models, filling a significant gap in the field and enabling the generation of vast high-quality HSI samples for large-scale downstream tasks. Extension experiments on downstream tasks demonstrate the effectiveness of synthesized HSIs. The code is available at HSI-Synthesis.github.io. Yang Yu 0045, Erting Pan, Xinya Wang, Xiaoguang Mei, Jiayi Ma 0001 |
CVPR | 6 |
| 2024 | MRFS: Mutually Reinforcing Image Fusion and SegmentationabstractThis paper proposes a coupled learning framework to break the performance bottleneck of infrared-visible image fusion and segmentation, called MRFS. By leveraging the intrinsic consistency between vision and semantics, it emphasizes mutual reinforcement rather than treating these tasks as separate issues. First, we embed weakened information recovery and salient information integration into the image fusion task, employing the CNN-based interactive gated mixed attention (IGM-Att) module to extract high-quality visual features. This aims to satisfy human visual perception, producing fused images with rich textures, high contrast, and vivid colors. Second, a transformer-based progressive cycle attention (PC-Att) module is developed to enhance semantic segmentation. It establishes single-modal self-reinforcement and cross-modal mutual complementarity, enabling more accurate decisions in machine semantic perception. Then, the cascade of IGM-Att and PC-Att couples image fusion and semantic segmentation tasks, implicitly bringing vision-related and semantics-related features into closer alignment. Therefore, they mutually provide learning priors to each other, resulting in visually satisfying fused images and more accurate segmentation decisions. Extensive experiments on public datasets showcase the advantages of our method in terms of visual satisfaction and decision accuracy. The code is publicly available at https://github.com/HaoZhang1018/MRFS. Hao Zhang 0073, Xuhui Zuo, Jie Jiang 0015, Chunchao Guo, Jiayi Ma 0001 |
CVPR | 5 |
| 2024 | Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image FusionabstractImage fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
CVPR | 5 |
| 2024 | DeMatch: Deep Decomposition of Motion Field for Two-View Correspondence LearningabstractTwo-view correspondence learning has recently focused on considering the coherence and smoothness of the motion field between an image pair. Dominant schemes include controlling the complexity of the field function with regularization or smoothing the field with local filters, but the former suffers from heavy computational burden, and the latter fails to accommodate discontinuities in the case of large scene disparities. In this paper, inspired by Fourier expansion, we propose a novel network called DeMatch, which decomposes the motion field to retain its main “low-frequency” and smooth part. This achieves implicit regularization with lower computational cost and generates piece-wise smoothness naturally. Specifically, we first decompose the rough motion field that is contaminated by false matches into several different sub-fields, which are highly smooth and contain the main energy of the original field. Then, with these smooth sub-fields, we recover a cleaner motion field from which correct motion vectors are subsequently derived. We also design a special masked decomposition strategy to further mitigate the negative influence of false matches. All the mentioned processes are finally implemented in a discrete and learnable manner, avoiding the difficulty of calculating real dense fields. Extensive experiments reveal that DeMatch outperforms state-of-the-art methods in multiple tasks and shows promising low computational usage and piecewise smoothness property. The code and trained models are publicly available at https://github.com/SuhZhang/DeMatch. Zizhuo Li, Yuan Gao 0015, Jiayi Ma 0001 |
CVPR | 4 |
| 2024 | Dispel Darkness for Better Fusion: A Controllable Visual Enhancer Based on Cross-Modal Conditional Adversarial LearningabstractWe propose a controllable visual enhancer, named DDBF, which is based on cross-modal conditional adversarial learning and aims to dispel darkness and achieve better visible and infrared modalities fusion. Specifically, a guided restoration module (GRM) is firstly designed to enhance weakened information in the low-light visible modality. The GRM utilizes the light-invariant high-contrast characteristics of the infrared modality as the central target distribution, and constructs a multilevel conditional adversarial sample set to enable continuous controlled brightness enhancement of visible images. Then, we develop an information fusion module (IFM) to integrate the advantageous features of the enhanced visible image and the infrared image. Thanks to customized explicit information preservation and hue fidelity constraints, the IFM produces visually pleasing results with rich textures, significant contrast, and vivid colors. The brightened visible image and the final fused image compose the dual output of our DDBF to meet the diverse visual preferences of users. We evaluate DDBF on the public datasets, achieving state-of-the-art performances of low-light enhancement and information integration that is available for both day and night scenarios. The experiments also demonstrate that our DDBF is effective in improving decision accuracy for object detection and semantic segmentation. Moreover, we offer a user-friendly interface for the convenient application of our model. The code is publicly available at https://github.com/HaoZhang1018/DDBF. Hao Zhang 0073, Linfeng Tang, Xinyu Xiang, Xuhui Zuo, Jiayi Ma 0001 |
CVPR | 5 |
| 2024 | CLIFF: Continual Latent Diffusion for Open-Vocabulary Object Detection
Wuyang Li, Xinyu Liu 0001, Jiayi Ma 0001, Yixuan Yuan |
ECCV (55) | 3 |
| 2024 | Mdbfusion: A Visible And Infrared Image Fusion Framework Capable For Motion DeblurringabstractExisting image fusion methods focus on containing more complementary information, but source images always suffer from motion blur owing to object motion, which results in distorted details in fused images and further deteriorates performance on high-level tasks. This paper proposes a novel visible and infrared image fusion framework capable for motion deblurring (MDbFusion++), which can simultaneously perform image fusion and deblurring within a mutually reinforcing framework. MDbFusion++ employs a coarse-to-fine image restoration strategy and comprises two key components: a coarse deblurring part (CDP) and a fine deblurring and fusion part (FDFP). Firstly, CDP transfers multi-modal images into features corresponding to spatial locations and creatively leverages infrared features to coarsely compensate motion blurred visible ones through adaptive weights module (AWM). Subsequently, FDFP further restores fine visible features and achieves multi-modal images fusion in spatial and frequency domains with the help of multi-domain enhancement module (MEM). The deblurred visible features provide clear information to improve fusion results, and the improved fused images, in turn, provide gradient feedback to further improve deblurring effects. We evaluate our network in terms of both image deblurring and fusion, and extensive comparative experiments demonstrate the superior performance and distinct advantages of MDbFusion++. Jun Chen 0019, Wei Yu 0018, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001 |
ICIP | 5 |
| 2024 | Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference CostabstractWe aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weights/gradients manipulation, our method is architecture-based with a flexible asymmetric structure for the primary and auxiliary tasks, which produces different networks for training and inference. Specifically, starting from two single task networks/branches (each representing a task), we propose a novel method with evolving networks where only primary-to-auxiliary links exist as the cross-task connections after convergence. These connections can be removed during the primary task inference, resulting in a single-task inference cost. We achieve this by formulating a Neural Architecture Search (NAS) problem, where we initialize bi-directional connections in the search space and guide the NAS optimization converging to an architecture with only the single-side primary-to-auxiliary connections. Moreover, our method can be incorporated with optimization-based auxiliary learning approaches. Extensive experiments with six tasks on NYU v2, CityScapes, and Taskonomy datasets using VGG, ResNet, and ViT backbones validate the promising performance. The codes are available at https://github.com/ethanygao/Aux-NAS. Yuan Gao 0015, Wenhan Luo, Lin Ma 0002, Jin-Gang Yu, Gui-Song Xia, Jiayi Ma 0001 |
ICLR | 7 |
| 2024 | Sparse-to-dense Multimodal Image Registration via Multi-Task LearningabstractAligning image pairs captured by different sensors or those undergoing significant appearance changes is crucial for various computer vision and robotics applications. Existing approaches cope with this problem via either Sparse feature Matching (SM) or Dense direct Alignment (DA) paradigms. Sparse methods are efficient but lack accuracy in textureless scenes, while dense ones are more accurate in all scenes but demand for good initialization. In this paper, we propose SDME, a Sparse-to-Dense Multimodal feature Extractor based on a novel multi-task network that simultaneously predicts SM and DA features for robust multimodal image registration. We propose the sparse-to-dense registration paradigm: we first perform initial registration via SM and then refine the result via DA. By using the well-designed SDME, the sparse-to-dense approach combines the merits from both SM and DA. Extensive experiments on MSCOCO, GoogleEarth, VIS-NIR and VIS-IR-drone datasets demonstrate that our method achieves remarkable performance on multimodal cases. Furthermore, our approach exhibits robust generalization capabilities, enabling the fine-tuning of models initially trained on single-modal datasets for use with smaller multimodal datasets. Our code is available at https://github.com/KN-Zhang/SDME. Kaining Zhang, Jiayi Ma 0001 |
ICML | 2 |
| 2024 | Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening
Meiqi Gong, Hao Zhang 0073, Hebaixu Wang, Jun Chen 0019, Jun Huang 0008, Xin Tian 0006, Jiayi Ma 0001 |
IJCAI | 7 |
| 2024 | DRMF: Degradation-Robust Multi-Modal Image Fusion via Composable Diffusion PriorabstractExisting multi-modal image fusion algorithms are typically designed for high-quality images and fail to tackle degradation (e.g., low light, low resolution, and noise), which restricts image fusion from unleashing the potential in practice. In this work, we present Degradation-Robust Multi-modality image Fusion (DRMF), leveraging the powerful generative properties of diffusion models to counteract various degradations during image fusion. Our critical insight is that generative diffusion models driven by different modalities and degradation are inherently complementary during the denoising process. Specifically, we pre-train multiple degradation-robust conditional diffusion models for different modalities to handle degradations. Subsequently, the diffusion priori combination module is devised to integrate generative priors from pre-trained uni-modal models, enabling effective multi-modal image fusion. Extensive experiments demonstrate that DRMF excels in infrared-visible and medical image fusion, even under complex degradations. Our code is available at https://github.com/Linfeng-Tang/DRMF. Linfeng Tang, Yuxin Deng 0002, Xunpeng Yi, Qinglong Yan, Yixuan Yuan, Jiayi Ma 0001 |
ACM Multimedia | 6 |
| 2024 | TeRF: Text-driven and Region-aware Flexible Visible and Infrared Image FusionabstractThe fusion of visible and infrared images aims to produce high-quality fusion images with rich textures and salient target information. Existing methods lack interactivity and flexibility in the execution of fusion. It is unfeasible to express the requirements to modify the fusion effect, and the different regions in the source images are treated equally across the identical fusion model, which causes fusion homogenization and low distinction. Besides, their pre-defined fusion strategies invariably lead to monotonous effects, which are insufficiently comprehensive. They fail to adequately consider data credibility, scene illumination, and noise degradation inherent in the source information. To address these issues, we propose the Te xt-driven and Region-aware Flexible visible and infrared image fusion, termed as TeRF. On the one hand, we propose a flexible image fusion framework with multiple large language and vision models, which facilitates the visual-text interaction. On the other hand, we aggregate comprehensive fine-tuning paradigms for the different fusion requirements to build a unified fine-tuning pipeline. It allows the linguistic selection of the regions and effects, yielding visually appealing fusion outcomes. Extensive experiments demonstrate the competitiveness of our method both qualitatively and quantitatively compared to existing state-of-the-art methods. Our code is publicly available at https://github.com/Baixuzx7/TeRF. Hebaixu Wang, Hao Zhang 0073, Xunpeng Yi, Xinyu Xiang, Leyuan Fang, Jiayi Ma 0001 |
ACM Multimedia | 6 |
| 2024 | DiffGlue: Diffusion-Aided Image Feature MatchingabstractAs one of the most fundamental computer vision problems, image feature matching aims to establish correct correspondences between two-view images. Existing studies enhance the descriptions of feature points with graph neural network (GNN), identifying correspondences with the predicted assignment matrix. However, this pipeline easily falls into a suboptimal result during training for the solution space is extremely complex, and is inaccessible to the prior that can guide the information propagation and network convergence. In this paper, we propose a novel method called DiffGlue that introduces the Diffusion Model into the sparse image feature matching framework. Concretely, based on the incrementally iterative diffusion and denoising processes, DiffGlue can be guided by the prior from the Diffusion Model and trained step by step on the optimization path, approaching the optimal solution progressively. Besides, it contains a special Assignment-Guided Attention as a bridge to merge the Diffusion Model and sparse image feature matching, which injects the inherent prior into GNN thereby ameliorating the message delivery. Extensive experiments reveal that DiffGlue converges faster and better, outperforming state-of-the-arts on several applications such as homography estimation, relative pose estimation, and visual localization. The code is available at https://github.com/SuhZhang/DiffGlue. Jiayi Ma 0001 |
ACM Multimedia | 2 |
| 2024 | Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion ModelabstractExisting multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, etc. Additionally, these methods often overlook the specificity of foreground objects, weakening the salience of the objects of interest within the fused images. To address these challenges, this study proposes a novel interactive multi-modal image fusion framework based on the text-modulated diffusion model, called Text-DiFuse. First, this framework integrates feature-level information integration into the diffusion process, allowing adaptive degradation removal and multi-modal information fusion. This is the first attempt to deeply and explicitly embed information fusion within the diffusion process, effectively addressing compound degradation in image fusion. Second, by embedding the combination of the text and zero-shot location model into the diffusion fusion process, a text-controlled fusion re-modulation strategy is developed. This enables user-customized text control to improve fusion performance and highlight foreground objects in the fused images. Extensive experiments on diverse public datasets show that our Text-DiFuse achieves state-of-the-art fusion performance across various scenarios with complex degradation. Moreover, the semantic segmentation experiment validates the significant enhancement in semantic performance achieved by our text-controlled fusion re-modulation strategy. The code is publicly available at https://github.com/Leiii-Cao/Text-DiFuse. Hao Zhang 0073, Jiayi Ma 0001 |
NeurIPS | 3 |
| 2024 | CRetinex: A Progressive Color-Shift Aware Retinex Model for Low-Light Image Enhancement
Han Xu 0001, Hao Zhang 0073, Xunpeng Yi, Jiayi Ma 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Open Set Recognition in Real World
Zhen Yang 0026, Jun Yue 0004, Pedram Ghamisi, Shiliang Zhang, Jiayi Ma 0001, Leyuan Fang |
Int. J. Comput. Vis. | 5 |
| 2024 | Multi-image super-resolution based low complexity deep network for image compressive sensing reconstructionabstractDeep learning (DL) has been widely utilized in image compressive sensing (CS) to enhance the quality and speed of reconstruction. The typical deep network for CS reconstruction comprises an initial reconstruction subnetwork, followed by a cascaded deep refinement reconstruction subnetworks. This paper introduces a new low-complexity image CS deep reconstruction framework, GSRCS, which leverages multi-image based deep super-resolution technology to better address the cost constraints of practical applications. The proposed initial reconstruction module generates multiple low-resolution images in parallel by grouping the input measurements, while a high-quality, high-resolution reconstructed image is produced through a multi-image deep super-resolution network. The theoretical derivation and experimental results demonstrate that this method significantly reduces system complexity in terms of parameters and floating-point arithmetic operations, while achieving competitive reconstruction performance compared to the state-of-the-arts. Specifically, the average number of parameters is reduced by over 63%, and the computational complexity is decreased by more than 88%. Qiming Xiong, Zhirong Gao, Jiayi Ma 0001, Yong Ma 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | PSD-ELGAN: A pseudo self-distillation based CycleGAN with enhanced local adversarial interaction for single image dehazing
Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
Neural Networks | 5 |
| 2024 | U-Match: Exploring Hierarchy-Aware Local Context for Two-View Correspondence LearningabstractRejecting outlier correspondences is one of the critical steps for successful feature-based two-view geometry estimation, and contingent heavily upon local context exploration. Recent advances focus on devising elaborate local context extractors whereas typically adopting explicit neighborhood relationship modeling at a specific scale, which is intrinsically flawed and inflexible, because 1) severe outliers often populated in putative correspondences and 2) the uncertainty in the distribution of inliers and outliers make the network incapable of capturing adequate and reliable local context from such neighborhoods, therefore resulting in the failure of pose estimation. This prospective study proposes a novel network called U-Match that has the flexibility to enable implicit local context awareness at multiple levels, naturally circumventing the aforementioned issues that plague most existing studies. Specifically, to aggregate multi-level local context implicitly, a hierarchy-aware graph representation module is designed to flexibly encode and decode hierarchical features. Moreover, considering that global context always works collaboratively with local context, an orthogonal local-and-global information fusion module is presented to integrate complementary local and global context in a redundancy-free manner, thus yielding compact feature representations to facilitate correspondence learning. Thorough experimentation across relative pose estimation, homography estimation, visual localization, and point cloud registration affirms U-Match's remarkable capabilities. Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | T-Net++: Effective Permutation-Equivariance Network for Two-View Correspondence PruningabstractWe propose a conceptually novel, flexible, and effective framework (named T-Net++) for the task of two-view correspondence pruning. T-Net++ comprises two unique structures: the "-'' structure and the "|'' structure. The "-'' structure utilizes an iterative learning strategy to process correspondences, while the "|'' structure integrates all feature information of the "-'' structure and produces inlier weights. Moreover, within the "|'' structure, we design a new Local-Global Attention Fusion module to fully exploit valuable information obtained from concatenating features through channel-wise and spatial-wise relationships. Furthermore, we develop a Channel-Spatial Squeeze-and-Excitation module, a modified network backbone that enhances the representation ability of important channels and correspondences through the squeeze-and-excitation operation. T-Net++ not only preserves the permutation-equivariance manner for correspondence pruning, but also gathers rich contextual information, thereby enhancing the effectiveness of the network. Experimental results demonstrate that T-Net++ outperforms other state-of-the-art correspondence pruning methods on various benchmarks and excels in two extended tasks. Guobao Xiao, Xin Liu 0091, Xiaoqin Zhang 0002, Jiayi Ma 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Latent Semantic Consensus for Deterministic Geometric Model FittingabstractEstimating reliable geometric model parameters from the data with severe outliers is a fundamental and important task in computer vision. This paper attempts to sample high-quality subsets and select model instances to estimate parameters in the multi-structural data. To address this, we propose an effective method called Latent Semantic Consensus (LSC). The principle of LSC is to preserve the latent semantic consensus in both data points and model hypotheses. Specifically, LSC formulates the model fitting problem into two latent semantic spaces based on data points and model hypotheses, respectively. Then, LSC explores the distributions of points in the two latent semantic spaces, to remove outliers, generate high-quality model hypotheses, and effectively estimate model instances. Finally, LSC is able to provide consistent and reliable solutions within only a few milliseconds for general multi-structural model fitting, due to its deterministic fitting nature and efficiency. Compared with several state-of-the-art model fitting methods, our LSC achieves significant superiority for the performance of both accuracy and speed on synthetic data and real images. Guobao Xiao, Jun Yu 0002, Jiayi Ma 0001, Deng-Ping Fan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | ConvMatch: Rethinking Network Design for Two-View Correspondence LearningabstractMultilayer perceptron (MLP) has become the de facto backbone in two-view correspondence learning, for it can extract effective deep features from unordered correspondences individually. However, the problem of natively lacking context information limits its performance although many context-capturing modules are appended in the follow-up studies. In this paper, from a novel perspective, we design a correspondence learning network called ConvMatch that for the first time can leverage a convolutional neural network (CNN) as the backbone, inherently capable of context aggregation. Specifically, with the observation that sparse motion vectors and a dense motion field can be converted into each other with interpolating and sampling, we regularize the putative motion vectors by estimating the dense motion field implicitly, then rectify the errors caused by outliers in local areas with CNN, and finally obtain correct motion vectors from the rectified motion field. Moreover, we propose global information injection and bilateral convolution, to fit the overall spatial transformation better and accommodate the discontinuities of the motion field in case of large scene disparity. Extensive experiments reveal that ConvMatch consistently outperforms state-of-the-arts for relative pose estimation, homography estimation, and visual localization. Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | SSL-Net: Sparse semantic learning for identifying reliable correspondences
Shunxing Chen, Guobao Xiao, Ziwei Shi, Junwen Guo, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2024 | Semantic attention-based heterogeneous feature aggregation network for image fusion
Zhiqiang Ruan, Guobao Xiao, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2024 | MC-Net: Integrating Multi-Level Geometric Context for Two-View Correspondence LearningabstractIn two-view correspondence learning, prevalent multi-layer perceptron (MLP)-based methods struggle with context capturing. To remedy this issue, recent advances innovatively stack convolutional neural network (CNN)-based Resblocks sequentially, showing an inherent proficiency in local context extraction. Yet, such non-issue-specific designs inherit the drawback of CNN’s difficulty in aggregating global context, leading to performance bottlenecks. To address this problem, this prospective study further explores the potential of the CNN-based framework and proposes MC-Net, a top-performing network that integrates both local and global context elegantly and seamlessly. Specifically, considering that sparse motion vectors and a dense motion field can be converted into each other through interpolation and sampling, we first transform unordered matches into image-structured data by estimating the dense motion field implicitly. Then, we design a hierarchical rectifying module to rectify the error of each ordered motion vector with CNN at multiple levels, enabling MC-Net to perceive global context from coarse-level features and local context from fine-level features simultaneously, which facilitates to tackle the discontinuities of the motion field in case of large scene disparity. Finally, we reconstruct comprehensive context-embedded features from rectified motion fields at all levels. Also, instead of using the residuals between rectified and pre-rectified motion vectors at the same layer to reject outliers as in previous studies, which seriously affects the inlier prediction accuracy, we rethink this operation meticulously and modify it to the difference between motion vectors obtained from each layer’s reconstruction and ones from the first layer before transformation, ensuring purer residuals and enhancing the matching performance without extra computational burden. Extensive experiments show that MC-Net outperforms state-of-the-arts on multiple domains and datasets. Zizhuo Li, Chunbao Su, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | CGR-Net: Consistency Guided ResFormer for Two-View Correspondence LearningabstractAccurately identifying correct correspondences (inliers) in two-view images is a fundamental task in computer vision. Recent studies usually adopt Graph Neural Networks or stack local graphs into global ones to establish neighborhood relations. However, the smoothing properties of Graph Convolutional Neural network (GCN) cause the model to fall into local extreme, which leads to the issue of indistinguishability between inliers and outliers. Especially when the initial correspondences contain a large number of incorrect correspondences (outliers), these studies suffer from severe performance degradation. To address the above issues and refocus perspective information on distinct features, we design a Consistency Guided ResFormer Network (CGR-Net) that uses consistent correspondences to guide model perspective focusing, thereby avoiding the negative impact of outliers. Specifically, we design an efficient Graph Score Calculation module, which aims to compute global graph scores by enhancing the representation of important features and comprehensively capturing the contextual relationships between correspondences. Then, we propose a Consistency Guided Correspondences Selection module to dynamically fuse global graph scores and consistency graphs and construct a novel consistency matrix to accurately recognize inliers. Extensive experiments on various challenging tasks demonstrate that our CGR-Net outperforms state-of-the-art methods. Our code is released athttps://github.com/XiaojieLi11/CGR-Net. Changcai Yang, Jiayi Ma 0001, Fengyuan Zhuang, Lifang Wei, Riqing Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | UnmixDiff: Unmixing-Based Diffusion Model for Hyperspectral Image SynthesisabstractThe scarcity of hyperspectral images (HSIs) hinders the development of processing methods and downstream applications. HSI synthesis, which aims to generate realistic samples from the existing datasets, is undoubtedly a prospective and economical solution for the HSI data shortage problem. Inspired by the impressive performance of the diffusion model (DM) in image synthesis tasks, this article initiatively proposes an unmixing diffusion (UnmixDiff) model for high-quality HSI generation. The method starts with training an unmixing network to learn the distribution characteristic of objects (abundance). By incorporating the unmixing autoencoder into the DM, the UnmixDiff transforms the HSI generation into the abundance domain, which maintains the consistency of the generated spectral profile, reduces the computational complexity, and introduces a clear physical interpretation into the hyperspectral image synthesis tasks. After that, we construct a diffusion generation model in abundance space to generate realistic abundance maps. Instead of synthesizing original hyperspectral images, the proposed UnmixDiff synthesizes abundance maps to simulate objects’ distribution rather than superficial textures. In this way, realistic HSI samples are generated by mixing the synthesized abundance with the scene end-members. With the comparative experiments, the proposed method achieves state-of-the-art performance in HSI synthesis tasks, effectively alleviating the HSI data scarcity and supporting widespread HSI applications. The code is available athttps://github.com/yuyang95/UnmixingDM. Yang Yu 0045, Erting Pan, Yong Ma 0001, Xiaoguang Mei, Qihai Chen, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | DTNet: A Specialized Dual-Tuning Network for Infrared Vehicle Detection in Aerial ImagesabstractVehicle detection in infrared aerial images is vital for both military and civilian applications, as infrared imaging remains effective under low-light conditions and various adverse weather scenarios. However, the longer wavelengths of long-wave infrared, compared to visible light, make diffraction more noticeable, leading to low-frequency degradation of vehicle information. Thermal radiation from the environment and optical system leads to higher noise in infrared images. Additionally, atmospheric transport models for various weather conditions can degrade infrared images to different extents. These factors lead to a reduced signal-to-noise ratio, which complicates the extraction of clear features. To overcome these challenges, we propose the Dual-Tuning Network (DTNet), an advanced framework for vehicle detection in infrared aerial images, developed based on the mechanisms of infrared imaging. Specifically, the core component of DTNet is the Dual-Tuning Block (DTBlock), which works alongside the Dynamic Guided Filtering Module (DGFM) and the Point Spread Recovery Module (PSRM) for feature extraction. DTBlock decomposes feature maps into low- and high-frequency components with learnable low-pass filters. DGFM eliminates disturbance from the optical system and background thermal radiation in the high-frequency component of feature maps, while preserving the details and texture of vehicles. PSRM aggregates vehicle features in the low-frequency component of feature maps, which is proposed with reference to diffraction and atmospheric models. The concept of Dual-Tuning refers to enhancing the signal and suppressing interference in the low and high frequency parts, respectively. Experimental results on the DroneVehicle public dataset for infrared vehicle detection indicate that our proposed approach achieves state-of-the-art (SOTA) performance. Moreover, extensive ablation studies confirm the superior capability of our DTNet in robust feature extraction from infrared images. Nan Zhang 0030, Youmeng Liu, Hao Liu 0064, Tian Tian 0006, Jiayi Ma 0001, Jinwen Tian |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | SegCLIP: Multimodal Visual-Language and Prompt Learning for High-Resolution Remote Sensing Semantic SegmentationabstractRemote sensing semantic segmentation is considered a key step in the intelligent interpretation of high-resolution remote sensing (HRRS) images, with widespread applications in fields such as hazard assessment, environmental monitoring, and urban planning. Recently, numerous deep learning-based semantic segmentation methods have emerged, achieving significant breakthroughs. However, the majority of current research still concentrates on representation learning in the visual feature space, with the potential of multimodal data sources yet to be fully explored. In recent years, the foundational visual language model, namely contrastive language-image pretraining (CLIP), has established a new paradigm in the visual field, demonstrating excellent generalization capabilities and deep semantic understanding across a variety of tasks. Inspired by prompt learning, we propose a prompting approach based on linguistic descriptions to enable CLIP to generate semantically distinct contextual information for remote sensing images. We introduce the SegCLIP network architecture, a novel framework specifically designed for semantic segmentation of HRRS images. Specifically, we have adapted CLIP to extract text information, thereby guiding the visual model in distinguishing among classes. Additionally, we have designed a cross-modal feature fusion (CFF) module that integrates linguistic and visual semantic features, ensuring semantic consistency across modalities. Finally, we have fully exploited the potential of text data and have used additional real text to refine ambiguous query features. Experimental evaluations confirm that the method exhibits superior performance on the LoveDA, iSAID, and UAVid public semantic segmentation datasets. Bin Zhang 0033, Yuntao Wu, Huabing Zhou, Junjun Jiang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | DHM-Net: Deep Hypergraph Modeling for Robust Feature MatchingabstractWe present a novel deep hypergraph modeling architecture (called DHM-Net) for feature matching in this paper. Our network focuses on learning reliable correspondences between two sets of initial feature points by establishing a dynamic hypergraph structure that models group-wise relationships and assigns weights to each node. Compared to existing feature matching methods that only consider pair-wise relationships via a simple graph, our dynamic hypergraph is capable of modeling nonlinear higher-order group-wise relationships among correspondences in an interaction capturing and attention representation learning fashion. Specifically, we propose a novel Deep Hypergraph Modeling block, which initializes an overall hypergraph by utilizing neighbor information, and then adopts node-to-hyperedge and hyperedge-to-node strategies to propagate interaction information among correspondences while assigning weights based on hypergraph attention. In addition, we propose a Differentiation Correspondence-Aware Attention mechanism to optimize the hypergraph for promoting representation learning. The proposed mechanism is able to effectively locate the exact position of the object of importance via the correspondence aware encoding and simple feature gating mechanism to distinguish candidates of inliers. In short, we learn such a dynamic hypergraph format that embeds deep group-wise interactions to explicitly infer categories of correspondences. To demonstrate the effectiveness of DHM-Net, we perform extensive experiments on both real-world outdoor and indoor datasets. Particularly, experimental results show that DHM-Net surpasses the state-of-the-art method by a sizable margin. Our approach obtains an 11.65% improvement under error threshold of 5° for relative pose estimation task on YFCC100M dataset. Code will be released at https://github.com/CSX777/DHM-Net. Shunxing Chen, Guobao Xiao, Junwen Guo, Qiangqiang Wu, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Progressive Learning With Cross-Window Consistency for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation focuses on the exploration of a small amount of labeled data and a large amount of unlabeled data, which is more in line with the demands of real-world image understanding applications. However, it is still hindered by the inability to fully and effectively leverage unlabeled images. In this paper, we reveal that cross-window consistency (CWC) is helpful in comprehensively extracting auxiliary supervision from unlabeled data. Additionally, we propose a novel CWC-driven progressive learning framework to optimize the deep network by mining weak-to-strong constraints from massive unlabeled data. More specifically, this paper presents a biased cross-window consistency (BCC) loss with an importance factor, which helps the deep network explicitly constrain confidence maps from overlapping regions in different windows to maintain semantic consistency with larger contexts. In addition, we propose a dynamic pseudo-label memory bank (DPM) to provide high-consistency and high-reliability pseudo-labels to further optimize the network. Extensive experiments on three representative datasets of urban views, medical scenarios, and satellite scenes with consistent performance gain demonstrate the superiority of our framework. Our code is released at https://jack-bo1220.github.io/project/CWC.html. Bo Dang 0002, Yansheng Li 0001, Yongjun Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | MERF: A Practical HDR-Like Image Generator via Mutual-Guided Learning Between Multi-Exposure Registration and FusionabstractIn this paper, we present a novel high dynamic range (HDR)-like image generator that utilizes mutual-guided learning between multi-exposure registration and fusion, leading to promising dynamic multi-exposure image fusion. The method consists of three main components: the registration network, the fusion network, and the dual attention network which seamlessly integrates registration and fusion processes. Initially, within the registration network, the estimation of deformation fields among multi-exposure image sequences is conducted following an exposure-invariant feature extraction phase. This leads to enhanced accuracy by mitigating discrepancies across domains. Subsequently, the fusion network utilizes a progressive frequency fusion module in two distinct stages, addressing color correction and detail preservation within low and high-frequency domains, respectively. To facilitate the mutual enhancement of the registration and fusion networks, we undertake a mutual-guided learning strategy encompassing their physical connection and constraint paradigm. Firstly, a dual attention network bridges the registration and fusion networks, addressing ghosting, which is beyond the scope of registration and facilitates information exchange between input images. Secondly, a meticulously designed generative adversarial network-like iterative training schema guides the overall network framework, thereby yielding high-quality HDR-like images through mutual enhancement. Comprehensive experiments on publicly available datasets validate the superiority of our method over existing state-of-the-art approaches. Wenhui Hong, Hao Zhang 0073, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Incrementally Adapting Pretrained Model Using Network Prior for Multi-Focus Image FusionabstractMulti-focus image fusion can fuse the clear parts of two or more source images captured at the same scene with different focal lengths into an all-in-focus image. On the one hand, previous supervised learning-based multi-focus image fusion methods relying on synthetic datasets have a clear distribution shift with real scenarios. On the other hand, unsupervised learning-based multi-focus image fusion methods can well adapt to the observed images but lack the general knowledge of defocus blur that can be learned from paired data. To avoid the problems of existing methods, this paper presents a novel multi-focus image fusion model by considering both the general knowledge brought by the supervised pretrained backbone and the extrinsic priors optimized on specific testing sample to improve the performance of image fusion. To be specific, the Incremental Network Prior Adaptation (INPA) framework is proposed to incrementally integrate features extracted from the pretrained strong baselines into a tiny prior network (6.9% parameters of the backbone network) to boost the performance for test samples. We evaluate our method on both synthetic and real-world public datasets (Lytro, MFI-WHU, and Real-MFF) and show that our method outperforms existing supervised learning-based methods and unsupervised learning based methods. Junjun Jiang, Chenyang Wang 0002, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Exploring the Spectral Prior for Hyperspectral Image Super-ResolutionabstractIn recent years, many single hyperspectral image super-resolution methods have emerged to enhance the spatial resolution of hyperspectral images without hardware modification. However, existing methods typically face two significant challenges. First, they struggle to handle the high-dimensional nature of hyperspectral data, which often results in high computational complexity and inefficient information utilization. Second, they have not fully leveraged the abundant spectral information in hyperspectral images. To address these challenges, we propose a novel hyperspectral super-resolution network named SNLSR, which transfers the super-resolution problem into the abundance domain. Our SNLSR leverages a spatial preserve decomposition network to estimate the abundance representations of the input hyperspectral image. Notably, the network acknowledges and utilizes the commonly overlooked spatial correlations of hyperspectral images, leading to better reconstruction performance. Then, the estimated low-resolution abundance is super-resolved through a spatial spectral attention network, where the informative features from both spatial and spectral domains are fully exploited. Considering that the hyperspectral image is highly spectrally correlated, we customize a spectral-wise non-local attention module to mine similar pixels along spectral dimension for high-frequency detail recovery. Extensive experiments demonstrate the superiority of our method over other state-of-the-art methods both visually and metrically. Our code is publicly available at https://github.com/HuQ1an/SNLSR. Xinya Wang, Junjun Jiang, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Learning a Non-Locally Regularized Convolutional Sparse Representation for Joint Chromatic and Polarimetric DemosaickingabstractDivision of focal plane color polarization camera becomes the mainstream in polarimetric imaging for it directly captures color polarization mosaic image by one snapshot, so image demosaicking is an essential task. Current color polarization demosaicking (CPDM) methods are prone to unsatisfied results since it's difficult to recover missed 15 or 14 pixels out of 16 pixels in color polarization mosaic images. To address this problem, a non-locally regularized convolutional sparse regularization model, which is advantaged in denoising and edge maintaining, is proposed to recall more information for CPDM task, and the CPDM task is transformed into an energy function to be solved by ADMM optimization. Finally, the optimal model generates informative and clear results. The experimental results, including reconstructed synthetic and real-world scenes, demonstrate that our proposed method outperforms the current state-of-the-art methods in terms of quantitative measurements and visual quality. The source code is available at https://github.com/roydon-luo/NLCSR-CPDM. Yidong Luo, Junchao Zhang 0001, Jianbo Shao, Jiandong Tian, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | HitFusion: Infrared and Visible Image Fusion for High-Level Vision Tasks Using TransformerabstractThis study proposes an innovative network to fuse infrared and visible images, called HitFusion, which uses the cross-feature transformer module and is compatible with high-level vision tasks. Firstly, existing image fusion approaches primarily concentrate on optimizing human visual perception and image metrics. To enhance the performance of the fusion network in subsequent high-level vision tasks, a segmentation network and a corresponding loss are introduced into the fusion network training process. Specifically, we devise a three-stage training strategy to render the fusion network more suitable for high-level vision tasks, guided by the segmentation network and broadening the fusion network's training set to boost its generalization capability. Secondly, current transformer-based image fusion methods neglect the interaction between visible texture features and infrared contrast features. To tackle this, the cross-feature transformer module is proposed, allowing the fusion network to learn the cross-feature correlation and long-range dependencies between source images, thus achieving fusion results with good complementarity. Finally, a dual-branch fusion network is proposed, based on the distinct characteristics of different images, that targets the extraction of deep features from source images utilizing contrast residual and texture enhancement modules to achieve improved fusion results. Extensive experimental results reveal that our HitFusion method excels in both qualitative and quantitative assessments, while also demonstrating superior performance in addressing high-level vision tasks. Jun Chen 0019, Jianfeng Ding, Jiayi Ma 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | OFPF-MEF: An Optical Flow Guided Dynamic Multi-Exposure Image Fusion Network With Progressive Frequencies LearningabstractIn this paper, we propose a novel optical flow guided network with progressive frequencies learning, achieving promising dynamic multi-exposure image fusion. Specifically, the proposed method consists of the optical flow alignment block and the progressive frequencies fusion block, where the former is to alleviate the ghost caused by the camera and object motions, and the latter is dedicated to synthesizing the desired color and details. First, in the optical flow alignment block, we estimate the optical flow between two source images and utilize the deformable convolutional network to achieve spatial alignment guided by the estimated optical flow. Second, in the progressive frequencies fusion block, color correction and details preservation are implemented in two gradual phases,i.e., low and high frequencies. For the low frequency fusion phase, we combine the convolutional neural network and swin transformer to capture local and global features, so as to consider the color correction from a complete perspective. For the high frequency fusion phase, an attention gate is designed to evaluate the important details from source images, bringing fewer artifacts and halos. Finally, the low and high frequency fusion phases are connected through a residual mapping strategy to generate a desired image with reasonable colors and rich details. Extensive experiments on publicly available datasets reveal that our method outperforms the state-of-the-art for both static and dynamic scenarios. Moreover, our method is superior in running efficiency over most of the state-of-the-art methods. Wenhui Hong, Hao Zhang 0073, Jiayi Ma 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Robust Feature Matching via Graph Neighborhood Motion ConsensusabstractIn this paper, we propose an effective method for mismatch removal, termed as graph neighborhood motion consensus, to address the feature matching problem which plays a pivotal role in various computer vision tasks. In our method, we convert each feature correspondence into a motion field sample and model it with the probabilistic graphical model (PGM). To differentiate mismatches from true matches, we firstly design a metric based on neighborhood topology consensus and neighborhood interaction to evaluate the correctness of each match. We also design a variance-based similarity search module to make the information used more reliable for better matching performance. To derive the solution of PGM, we build a model to transform the problem into an integer quadratic programming problem and obtain its closed-form solution with linear time complexity. Extensive experiments on general feature matching, fundamental matrix estimation and image registration tasks demonstrate that our proposed method can achieve superior performance over several state-of-the-art approaches. Jun Huang 0008, Yijia Gong, Fan Fan 0001, Yong Ma 0001, Qinglei Du, Jiayi Ma 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | CAMF: An Interpretable Infrared and Visible Image Fusion Network Based on Class Activation MappingabstractImage fusion aims to integrate the complementary information of source images and synthesize a single fused image. Existing image fusion algorithms apply hand-crafted fusion rules to merge deep features which cause information loss and limit the fusion performance of methods since the uninterpretability of deep learning. To overcome the above shortcomings, we propose a learnable fusion rule for infrared and visible image fusion based on class activation mapping. Our proposed fusion rule can selectively preserve meaningful information and reduce distortion. More specifically, we first train an encoder-decoder network and an auxiliary classifier based on the shared encoder. Then, the class activation weights are taken out from the auxiliary classifier, which indicates the importance of each channel. Finally, the deep features extracted by the encoder are adaptively fused according to the class activation weights and the fused image is reconstructed from the fused features via the pre-trained decoder. Note that our learnable fusion rule can automatically measure the importance of each deep feature without human participation. Moreover, it fully preserves the significant features of source images such as salient targets and texture details. Extensive experiments manifest our superiority over state-of-the-art algorithms. Visualization of feature maps and their corresponding weights reveals the high interpretability of our method. Linfeng Tang, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Cycle-Retinex: Unpaired Low-Light Image Enhancement via Retinex-Inline CycleGANabstractLow-light image enhancement aims to recover normal-light images from the images captured under dim environments. Most existing methods could just improve the light appearance globally whereas failing to handle other degradation such as dense noise, color offset and extremely low-light. Moreover, unsupervised methods proposed in recent years lack reliable physical model as the basis, thus universality is greatly limited. To address these problems, we propose a novel low-light image enhancement method via Retinex-inline cycle-consistent generative adversarial network named Cycle-Retinex, whose training is totally dependent on unpaired datasets. Specifically, we organically combine Retinex theory with CycleGAN, by which we decouple low-light image enhancement task into two sub-tasks, i.e. illumination map enhancement and reflectance map restoration. Retinex theory helps CycleGAN simplify low-light image enhancement problem and CycleGAN provides synthetic paired images to guide the training of Retinex decomposition network. We further introduce a self-augmented method to address the color distortion and noise problem, thus making the network learn to enhance low-light images adaptively. Extensive experiments show that the proposed method can achieve promising results. Kangle Wu, Jun Huang 0008, Yong Ma 0001, Fan Fan 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | DRLIE: Flexible Low-Light Image Enhancement via Disentangled RepresentationsabstractLow-light image enhancement (LIME) aims to convert images with unsatisfied lighting into desired ones. Different from existing methods that manipulate illumination in uncontrollable manners, we propose a flexible framework to take user-specified guide images as references to improve the practicability. To achieve the goal, this article models an image as the combination of two components, that is, content and exposure attribute, from an information decoupling perspective. Specifically, we first adopt a content encoder and an attribute encoder to disentangle the two components. Then, we combine the scene content information of the low-light image with the exposure attribute of the guide image to reconstruct the enhanced image through a generator. Extensive experiments on public datasets demonstrate the superiority of our approach over state-of-the-art alternatives. Particularly, the proposed method allows users to enhance images according to their preferences, by providing specific guide images. Our source code and the pretrained model are available at https://github.com/Linfeng-Tang/DRLIE. Linfeng Tang, Jiayi Ma 0001, Hao Zhang 0073, Xiaojie Guo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Interpretable Model-Driven Deep Network for Hyperspectral, Multispectral, and Panchromatic Image FusionabstractSimultaneously fusing hyperspectral (HS), multispectral (MS), and panchromatic (PAN) images brings a new paradigm to generate a high-resolution HS (HRHS) image. In this study, we propose an interpretable model-driven deep network for HS, MS, and PAN image fusion, called HMPNet. We first propose a new fusion model that utilizes a deep before describing the complicated relationship between the HRHS and PAN images owing to their large resolution difference. Consequently, the difficulty of traditional model-based approaches in designing suitable hand-crafted priors can be alleviated because this deep prior is learned from data. We further solve the optimization problem of this fusion model based on the proximal gradient descent (PGD) algorithm, achieved by a series of iterative steps. By unrolling these iterative steps into several network modules, we finally obtain the HMPNet. Therefore, all parameters besides the deep prior are learned in the deep network, simplifying the selection of optimal parameters in the fusion and achieving a favorable equilibrium between the spatial and spectral qualities. Meanwhile, all modules contained in the HMPNet have explainable physical meanings, which can improve its generalization capability. In the experiment, we exhibit the advantages of the HMPNet over other state-of-the-art methods from the aspects of visual comparison and quantitative analysis, where a series of simulated as well as real datasets are utilized for validation. Xin Tian 0006, Kun Li 0025, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Omniscient Video Super-Resolution with Explicit-Implicit AlignmentabstractWhen considering the temporal relationships, most previous video super-resolution (VSR) methods follow the iterative or recurrent framework. The iterative framework adopts neighboring low-resolution (LR) frames from a sliding window, while the recurrent framework utilizes the output generated in the previous SR procedure. The hybrid framework combines them but still cannot fully leverage the temporal relationships. Meanwhile, the existing methods are limited in the receptive field of the optical flow or lack semantic constrains on motion information. In this work, we propose an omniscient framework to fully explore the temporal relationships in the video, which encompasses both LR frames and SR outputs from the past, present, and future. The omniscient framework is more generic because the iterative, recurrent, and hybrid frameworks can be regarded as its special cases. Besides, when addressing the motion information, most previous VSR methods adopt the explicit motion estimation and compensation, while many recent methods turn to implicit alignment. In implicit alignment methods, because basic non-local means suffers from heavy computational costs, we improve it by capturing the non-local correlations in a relatively local manner to reduce the complexity. Moreover, we integrate the explicit and implicit methods into an explicit-implicit alignment module to better utilize motion information. We have conducted extensive experiments on public datasets, which show that our method is superior over the state-of-the-art methods in objective metrics, subjective visual quality, and complexity. In particular, on datasets of Vid4 and UDM10, our method improves PSNR by 0.19 dB, 0.49 dB against the most advanced method BasicVSR++, respectively. Peng Yi 0002, Zhongyuan Wang 0002, Laigan Luo, Kui Jiang, Zheng He 0001, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | From Noise Addition to Denoising: A Self-Variation Capture Network for Point Cloud OptimizationabstractPoint clouds obtained from 3D scanners are often noisy and cannot be directly used for subsequent high-level tasks. In this article, we propose a novel point cloud optimization method capable of denoising and homogenizing point clouds. Our idea is based on the assumption that the noise is generally much smaller than the effective signal. We perform noise perturbation on the noisy point cloud to get a new noisy point cloud, called self-variation point cloud. The noisy point cloud and self-variation point cloud have different noise distribution, but the same point cloud distribution. We compute the potential commonality between two noisy point clouds to obtain a clean point cloud. To implement our idea, we propose a Self-Variation Capture Network (SVCNet). We perturb the point cloud features in the latent space to obtain self-variation feature vectors, and capture the commonality between two noisy feature vectors through the feature aggregation and averaging. In addition, an edge constraint module is introduced to suppress low-pass effects during denoising. Our denoising method does not take into account the noise characteristics, and can filter the drift noise located on the underlying surface, resulting in a uniform distribution of the generated point cloud. The experimental results show that our algorithm outperforms the current state-of-the-art algorithms, especially in generating more uniform point clouds. In addition, extended experiments demonstrate the potential of our algorithm for point clouds upsampling. Tianming Zhao 0003, Peng Gao 0012, Tian Tian 0006, Jiayi Ma 0001, Jinwen Tian |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Unsupervised Multi-Exposure Image Fusion Breaking Exposure Limits via Contrastive LearningabstractThis paper proposes an unsupervised multi-exposure image fusion (MEF) method via contrastive learning, termed as MEF-CL. It breaks exposure limits and performance bottleneck faced by existing methods. MEF-CL firstly designs similarity constraints to preserve contents in source images. It eliminates the need for ground truth (actually not exist and created artificially) and thus avoids negative impacts of inappropriate ground truth on performance and generalization. Moreover, we explore a latent feature space and apply contrastive learning in this space to guide fused image to approximate normal-light samples and stay away from inappropriately exposed ones. In this way, characteristics of fused images (e.g., illumination, colors) can be further improved without being subject to source images. Therefore, MEF-CL is applicable to image pairs of any multiple exposures rather than a pair of under-exposed and over-exposed images mandated by existing methods. By alleviating dependence on source images, MEF-CL shows better generalization for various scenes. Consequently, our results exhibit appropriate illumination, detailed textures, and saturated colors. Qualitative, quantitative, and ablation experiments validate the superiority and generalization of MEF-CL. Our code is publicly available at https://github.com/hanna-xu/MEF-CL. Han Xu 0001, Liang Haochen, Jiayi Ma 0001 |
AAAI | 3 |
| 2023 | ConvMatch: Rethinking Network Design for Two-View Correspondence LearningabstractMultilayer perceptron (MLP) has been widely used in two-view correspondence learning for only unordered correspondences provided, and it extracts deep features from individual correspondence effectively. However, the problem of lacking context information limits its performance and hence, many extra complex blocks are designed to capture such information in the follow-up studies. In this paper, from a novel perspective, we design a correspondence learning network called ConvMatch that for the first time can leverage convolutional neural network (CNN) as the backbone to capture better context, thus avoiding the complex design of extra blocks. Specifically, with the observation that sparse motion vectors and dense motion field can be converted into each other with interpolating and sampling, we regularize the putative motion vectors by estimating dense motion field implicitly, then rectify the errors caused by outliers in local areas with CNN, and finally obtain correct motion vectors from the rectified motion field. Extensive experiments reveal that ConvMatch with a simple CNN backbone consistently outperforms state-of-the-arts including MLP-based methods for relative pose estimation and homography estimation, and shows promising generalization ability to different datasets and descriptors. Our code is publicly available at https://github.com/SuhZhang/ConvMatch. Jiayi Ma 0001 |
AAAI | 2 |
| 2023 | Robust and Scalable Gaussian Process Regression and Its ApplicationsabstractThis paper introduces a robust and scalable Gaussian process regression (GPR) model via variational learning. This enables the application of Gaussian processes to a wide range of real data, which are often large-scale and contaminated by outliers. Towards this end, we employ a mixture likelihood model where outliers are assumed to be sampled from a uniform distribution. We next derive a variational formulation that jointly infers the mode of data, i.e., inlier or outlier, as well as hyperparameters by maximizing a lower bound of the true log marginal likelihood. Compared to previous robust GPR, our formulation approximates the exact posterior distribution. The inducing variable approximation and stochastic variational inference are further introduced to our variational framework, extending our model to large-scale data. We apply our model to two challenging real-world applications, namely feature matching and dense gene expression imputation. Extensive experiments demonstrate the superiority of our model in terms of robustness and speed. Notably, when matching 4k feature points, its inference is completed in milliseconds with almost no false matches. The code is at github.com/YifanLu2000/Robust-Scalable-GPR. Jiayi Ma 0001, Leyuan Fang, Xin Tian 0006, Junjun Jiang |
CVPR | 2 |
| 2023 | Sparsely Annotated Semantic Segmentation with Adaptive Gaussian MixturesabstractSparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent relation between labeled and unlabeled pixels. In this paper, we observe that pixels that are close to each other in the feature space are more likely to share the same class. Inspired by this, we propose a novel SASS framework, which is equipped with an Adaptive Gaussian Mixture Model (AGMM). Our AGMM can effectively endow reliable supervision for unlabeled pixels based on the distributions of labeled and unlabeled pixels. Specifically, we first build Gaussian mixtures using labeled pixels and their relatively similar unlabeled pixels, where the labeled pixels act as centroids, for modeling the feature distribution of each class. Then, we leverage the reliable information from labeled pixels and adaptively generated GMM predictions to supervise the training of unlabeled pixels, achieving online, dynamic, and robust selfsupervision. In addition, by capturing category-wise Gaussian mixtures, AGMM encourages the model to learn discriminative class decision boundaries in an end-to-end contrastive learning manner. Experimental results conducted on the PASCAL VOC 2012 and Cityscapes datasets demonstrate that our AGMM can establish new state-of-the-art SASS performance. Code is available at https://github.com/Luffy03/AGMM-SASS Linshan Wu, Zhun Zhong, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Hao Chen 0011 |
CVPR | 6 |
| 2023 | LSTFE-Net: Long Short-Term Feature Enhancement Network for Video Small Object DetectionabstractVideo small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropriate features. Since information from frames at different positions contributes differently to small objects, it is not ideal to assume that using one universal method will extract proper features. We find that context information from the long-term frame and temporal information from the short-term frame are two useful cues for video small object detection. To fully utilize these two cues, we propose a long short-term feature enhancement network (LSTFE-Net) for video small object detection. First, we develop a plug-and-play spatiotemporal feature alignment module to create temporal correspondences between the short-term and current frames. Then, we propose a frame selection module to select the long-term frame that can provide the most additional context information. Finally, we propose a long short-term feature aggregation module to fuse long short-term features. Compared to other state-of-the-art methods, our LSTFE-Net achieves 4.4% absolute boosts in AP on the FL-Drones dataset. More details can be found at https://github.com/xiaojs18/LSTFE-Net. Jinsheng Xiao, Yuanxu Wu, Yunhua Chen, Zhongyuan Wang 0001, Jiayi Ma 0001 |
CVPR | 6 |
| 2023 | Diff-Retinex: Rethinking Low-light Image Enhancement with A Generative Diffusion ModelabstractIn this paper, we rethink the low-light image enhancement task and propose a physically explainable and generative diffusion model for low-light image enhancement, termed as Diff-Retinex. We aim to integrate the advantages of the physical model and the generative network. Furthermore, we hope to supplement and even deduce the information missing in the low-light image through the generative network. Therefore, Diff-Retinex formulates the lowlight image enhancement problem into Retinex decomposition and conditional image generation. In the Retinex decomposition, we integrate the superiority of attention in Transformer and meticulously design a Retinex Transformer decomposition network (TDN) to decompose the image into illumination and reflectance maps. Then, we design multi-path generative diffusion networks to reconstruct the normal-light Retinex probability distribution and solve the various degradations in these components respectively, including dark illumination, noise, color deviation, loss of scene contents, etc. Owing to generative diffusion model, Diff-Retinex puts the restoration of low-light subtle detail into practice. Extensive experiments conducted on real-world low-light datasets qualitatively and quantitatively demonstrate the effectiveness, superiority, and generalization of the proposed method. Xunpeng Yi, Han Xu 0001, Hao Zhang 0073, Linfeng Tang, Jiayi Ma 0001 |
ICCV | 5 |
| 2023 | U-Match: Two-view Correspondence Learning with Hierarchy-aware Local Context AggregationabstractLocal context capturing has become the core factor for achieving leading performance in two-view correspondence learning. Recent advances have devised various local context extractors whereas typically adopting explicit neighborhood relation modeling that is restricted and inflexible. To address this issue, we introduce U-Match, an attentional graph neural network that has the flexibility to enable implicit local context awareness at multiple levels. Specifically, a hierarchy-aware graph representation (HAGR) module is designed and fleshed out by local context pooling and unpooling operations. The former encodes local context by adaptively sampling a set of nodes to form a coarse-grained graph, while the latter decodes local context by recovering the coarsened graph back to its original size. Moreover, an orthogonal fusion module is proposed for the collaborative use of HAGR module, which integrates complementary local and global information into compact feature representations without redundancy. Extensive experiments on different visual tasks prove that our method significantly surpasses the state-of-the-arts. In particular, U-Match attains an AUC at 5 degree threshold of 60.53% on the challenging YFCC100M dataset without RANSAC, outperforming the strongest prior model by 8.61 absolute percentage points. Our code is publicly available at https://github.com/ZizhuoLi/U-Match. Zizhuo Li, Jiayi Ma 0001 |
IJCAI | 3 |
| 2023 | Robust Model Reasoning and Fitting via Dual Sparsity PursuitabstractIn this paper, we contribute to solving a threefold problem: outlier rejection, true model reasoning and parameter estimation with a unified optimization modeling. To this end, we first pose this task as a sparse subspace recovering problem, to search a maximum of independent bases under an over-embedded data space. Then we convert the objective into a continuous optimization paradigm that estimates sparse solutions for both bases and errors. Wherein a fast and robust solver is proposed to accurately estimate the sparse subspace parameters and error entries, which is implemented by a proximal approximation method under the alternating optimization framework with the ``optimal'' sub-gradient descent. Extensive experiments regarding known and unknown model fitting on synthetic and challenging real datasets have demonstrated the superiority of our method against the state-of-the-art. We also apply our method to multi-class multi-model fitting and loop closure detection, and achieve promising results both in accuracy and efficiency. Code is released at: https://github.com/StaRainJ/DSP. Xingyu Jiang 0005, Jiayi Ma 0001 |
NeurIPS | 2 |
| 2023 | Hierarchical image peeling: A flexible scale-space filtering framework
Yuanbin Fu, Jiayi Ma 0001, Xiaojie Guo 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Improving sparse graph attention for feature matching by informative keypoints exploration
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 4 |
| 2023 | STP-SOM: Scale-Transfer Learning for Pansharpening via Estimating Spectral Observation Model
Hao Zhang 0073, Jiayi Ma 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Querying Labeled for Unlabeled: Cross-Image Semantic Consistency Guided Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation aims to learn a semantic segmentation model via limited labeled images and adequate unlabeled images. The key to this task is generating reliable pseudo labels for unlabeled images. Existing methods mainly focus on producing reliable pseudo labels based on the confidence scores of unlabeled images while largely ignoring the use of labeled images with accurate annotations. In this paper, we propose a Cross-Image Semantic Consistency guided Rectifying (CISC-R) approach for semi-supervised semantic segmentation, which explicitly leverages the labeled images to rectify the generated pseudo labels. Our CISC-R is inspired by the fact that images belonging to the same class have a high pixel-level correspondence. Specifically, given an unlabeled image and its initial pseudo labels, we first query a guiding labeled image that shares the same semantic information with the unlabeled image. Then, we estimate the pixel-level similarity between the unlabeled image and the queried labeled image to form a CISC map, which guides us to achieve a reliable pixel-level rectification for the pseudo labels. Extensive experiments on the PASCAL VOC 2012, Cityscapes, and COCO datasets demonstrate that the proposed CISC-R can significantly improve the quality of the pseudo labels and outperform the state-of-the-art methods. Code is available at https://github.com/Luffy03/CISC-R. Linshan Wu, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Zhun Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | MURF: Mutually Reinforcing Multi-Modal Image Registration and FusionabstractExisting image fusion methods are typically limited to aligned source images and have to "tolerate" parallaxes when images are unaligned. Simultaneously, the large variances between different modalities pose a significant challenge for multi-modal image registration. This study proposes a novel method called MURF, where for the first time, image registration and fusion are mutually reinforced rather than being treated as separate issues. MURF leverages three modules: shared information extraction module (SIEM), multi-scale coarse registration module (MCRM), and fine registration and fusion module (F2M). The registration is carried out in a coarse-to-fine manner. During coarse registration, SIEM first transforms multi-modal images into mono-modal shared information to eliminate the modal variances. Then, MCRM progressively corrects the global rigid parallaxes. Subsequently, fine registration to repair local non-rigid offsets and image fusion are uniformly implemented in F2M. The fused image provides feedback to improve registration accuracy, and the improved registration result further improves the fusion result. For image fusion, rather than solely preserving the original source information in existing methods, we attempt to incorporate texture enhancement into image fusion. We test on four types of multi-modal data (RGB-IR, RGB-NIR, PET-MRI, and CT-MRI). Extensive registration and fusion results validate the superiority and universality of MURF. Han Xu 0001, Jiteng Yuan, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Variational Bayesian deep network for blind Poisson denoising
Hao Liang 0008, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
Pattern Recognit. | 4 |
| 2023 | Hyperspectral image denoising via spectral noise distribution bootstrap
Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2023 | Hyperspectral image destriping and denoising from a task decomposition view
Erting Pan, Yong Ma 0001, Xiaoguang Mei, Jun Huang 0008, Qihai Chen, Jiayi Ma 0001 |
Pattern Recognit. | 6 |
| 2023 | MSINet: Mining scale information from digital surface models for semantic segmentation of aerial images
Chengli Peng, Haifeng Li 0007, Chao Tao 0001, Yansheng Li 0001, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2023 | JRA-Net: Joint representation attention network for correspondence learning
Ziwei Shi, Guobao Xiao, Linxin Zheng, Jiayi Ma 0001, Riqing Chen |
Pattern Recognit. | 4 |
| 2023 | APUNet: Attention-guided upsampling network for sparse and non-uniform point cloud
Tianming Zhao 0003, Linfeng Li 0002, Tian Tian 0006, Jiayi Ma 0001, Jinwen Tian |
Pattern Recognit. | 4 |
| 2023 | Patch-guided point matching for point cloud registration with low overlap
Tianming Zhao 0003, Linfeng Li 0002, Tian Tian 0006, Jiayi Ma 0001, Jinwen Tian |
Pattern Recognit. | 4 |
| 2023 | Multipatch Progressive Pansharpening With Knowledge DistillationabstractIn this paper, we propose a novel multi-patch and multi-stage pansharpening method with knowledge distillation, termed as PSDNet. Different from existing pansharpening methods that typically input single-size patches to the network and implement pansharpening in an overall stage, we design multi-patch inputs and a multi-stage network for more accurate and finer learning. First, multi-patch inputs allow the network to learn more accurate spatial and spectral information by reducing the number of object types. We employ small patches in the early part to learn accurate local information, as small patches contain fewer object types. Then, the later part exploits large patches to fine-tune it for the overall information. Second, the multi-stage network is designed to reduce the difficulty of the previous single-step pansharpening and progressively generate elaborate results. In addition, instead of the traditional perceptual loss, which hardly relates to the specific task or the designed network, we introduce distillation loss to reinforce the guidance of the ground truth. Extensive experiments are conducted to demonstrate the superior performance of our proposed PSDNet to existing state-of-the-art methods. Our code is available at https://github.com/Meiqi-Gong/PSDNet. Meiqi Gong, Hao Zhang 0073, Han Xu 0001, Xin Tian 0006, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Learning for Feature Matching via Graph Context AttentionabstractEstablishing reliable correspondences via a deep learning network is an important task in remote sensing, photogrammetry and other computer vision fields. It usually requires mining the relationship among correspondences to aggregate both local and global context. However, current methods are insufficient to effectively acquire context information with high reliability. In this paper, we propose a graph context attention based network (called GCA-Net) to capture and leverage abundant contextual information for feature matching. Specifically, we design a graph context attention block which generates multi-path graph contexts and softly fuses them to combine respective advantages. In addition, for building the graph context containing stronger representation ability and outlier resistance ability, we further design a local-global channel mining block to gather context information by focusing on the significant part as well as to mine dependencies among channels of correspondences in both local and global aspects. The proposed GCA-Net is able to effectively infer the probability of correspondences being inliers or outliers and estimate the essential matrix meanwhile. Extensive experimental results for outlier removal and relative pose estimation demonstrate that GCA-Net outperforms the state-of-art methods on both outdoor and indoor datasets (i.e.,YFCC100M and SUN3D). In addition, experiments extended to remote sensing and point cloud scenes also demonstrate the powerful generalization capability of our network. Junwen Guo, Guobao Xiao, Shunxing Chen, Shiping Wang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Coarse-to-Fine Cross-Domain Learning Fusion Network for PansharpeningabstractDeep learning (DL) based pansharpening methods have shown great advantages in fusing multispectral (MS) and panchromatic (PAN) images to obtain a high-resolution MS image in remote sensing applications. However, most DL methods have low generalization capability that will cause severe spatial or spectral distortions, especially when a large distribution gap exists between training data from a source domain and testing data from another target domain. To overcome this problem, we propose a coarse-to-fine adaption learning fusion network for pansharpening. We first learn the priori mapping relationships between MS and PAN images in the source domain through the coarse-fusion network, which combines the advantages of UNet and Transformer architectures that helps to explore texture information of different characteristics. To generate a clear fusion result with good preservation of spatial and spectral information in the target domain, the fine-fusion network is further proposed to adjust the spatial and spectral information of the coarse-fusion image in an unsupervised learning manner based on the target-specific knowledge. Therefore, the generalization capability can be effectively improved because the general mapping relationship from the source domain and the specific target knowledge from the target domain are both considered. Experiments on simulated and real datasets are conducted to demonstrate the superiority of our proposed method over other state-of-the-art DL methods in terms of visual quality and quantitative analysis. Chengjie Ke, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | SPGAN-DA: Semantic-Preserved Generative Adversarial Network for Domain Adaptive Remote Sensing Image Semantic SegmentationabstractUnsupervised domain adaptation for remote sensing semantic segmentation seeks to adapt a model trained on the labeled source domain to the unlabeled target domain. One of the most promising ways is to translate images from the source domain to the target domain to align the spectral information or imaging mode by the generative adversarial network (GAN). However, source-to-target translation often brings bias in the translated images causing limited performance, as semantic information is not well considered in the translation procedure. To overcome this limitation, we present an innovative semantic-preserved generative adversarial network (SPGAN), designed to mitigate the image translation bias and then leverage the translated images as well as unlabeled target images by class distribution alignment (CDA) module to train a domain adaptive semantic segmentation model. The above two stages are coupled together to form a unified framework called SPGAN-DA. Specifically, we first conduct semantic invariant translation from source to target domain, which is achieved by introducing representation-invariant and semantic-preserved constraints to the GAN model. To further narrow the landscape layout gap between the translated and target images, CDA semantic segmentation is proposed. CDA semantic segmentation consists of two aspects. At the model input level, object discrepancy is eliminated by introducing the ClassMix operation. At the model output level, boundary enhancement is proposed to refine the performance of object boundaries. Extensive experiments on three typical remote sensing cross-domain semantic segmentation benchmarks demonstrate the effectiveness and generality of our proposed method, which competes favorably against existing state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Yongjun Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Progressive Hyperspectral Image Destriping With an Adaptive Frequencial FocusabstractLimited by the imaging paradigm, stripes is pervasive in remote sensing scenes, and its intensity, density, and periodicity differ dramatically among different imaging systems. Worse, it always co-exists with random noises caused by unstable imaging condition. However, current destriping methods are victim to undue ideal assumptions and fail to accurately eliminate stripes against diverse practical degradation, yielding excessive or inadequate destriping results. This study proposes a progressive hyperspectral destriping method with an adaptive frequency focus for accurate destriping and delicate restoration. Specifically, a hierarchical decomposition and reconstruction framework based on progressive wavelet learning encodes the degraded input to the frequency domain with smaller scales, easing the difficulty of restoration. Then, to avoid excessive or insufficient destriping, we devote specific efforts to finely separating noise and preserving details in the high-frequency domain. First, we devise a gradient-aware frequency attention block based on the prominent unidirectional pattern of stripes, empowering to adaptively assign weights according to their sensitivity to the spatial gradient. Second, we design a focal high-frequency loss item that is dynamically scaled according to feature distance in the high-frequency domain, profiting in identifying and preserving details. Extensive experiments conducted on data with synthetic stripes and realistic satellite scenes validate the superiority of the proposed method over the current state-of-the-art methods. The code is available at https://github.com/EtPan/PHID. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Dual Spatial-Spectral Pyramid Network With Transformer for Hyperspectral Image FusionabstractMultispectral image (MSI) and hyperspectral image (HSI) fusion can combine the best of both worlds to produce images with both high spatial and spectral resolution. In this paper, we have designed a network for fusing MSIs and HSIs, called DSPNet. On the one hand, in order to ensure the accuracy of the spectral dimension, i.e. spectral fidelity, we designed the spectral pyramid (SpePy) module and the multiscale spectral information fusion (MLSIF) module. The former extracts the multiscale local spectral information that captures the subtle spectral details and variations between different spectra. The latter establishes long-range dependency in the spectral dimension through the spectral-wise multi-head hybrid-attention (S-MHA) mechanism, thus enabling the network to focus on the local spectral information needed to recover the spectral details. On the other hand, to address the spatial information of MSIs, we designed the spatial pyramid (SpaPy) module. The SpaPy module can extract the non-local spatial information of MSIs at different scales, which enables the network to adapt to different remote-sensing scenes. Experiments performed on simulated and real data demonstrate the superiority of our method over the state-of-the-art methods both qualitatively and quantitatively. Han Xu 0001, Yong Ma 0001, Minghui Wu 0007, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | PG-Net: Progressive Guidance Network via Robust Contextual Embedding for Efficient Point Cloud RegistrationabstractBuilding high-quality correspondences is critical in the feature-based point cloud registration pipelines. However, existing single-sequence learning frameworks are difficult to accurately and adequately capture contextual information, leaving a large proportion of outliers between two low-overlap scenes. In this paper, we present a progressive guidance network (PG-Net) to gather rich contextual information and exclude outliers. Specifically, we design a novel iterative structure that exploits the inlier probabilities of correspondences to guide the classification of initial correspondences progressively. This structure can mitigate outlier effects with robust contextual information to obtain more accurate model estimation. In addition, to sufficiently capture contextual information, we propose a grouped dense fusion attention feature embedding module to enhance the representation of inliers and significant channel-spatial. Meanwhile, we propose a two-stage neural spectral matching module to compute the inlier probability of each correspondence and estimate a 3D transformation model in a coarse-to-fine manner. Experiments results on indoor and outdoor datasets using distinct 3D local descriptors demonstrate that our PG-Net surpasses state-of-the-art outlier removal methods. Especially compared to the recent outlier removal network PointDSC, our PG-Net improves the registration recall by 4.06% on the indoor dataset with the FPFH descriptor. Source code: https://github.com/changcaiyang/PG-Net. Xin Liu 0091, Luanyuan Dai, Jiayi Ma 0001, Lifang Wei, Changcai Yang, Riqing Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Neighborhood Manifold Preserving Matching for Visual Place RecognitionabstractThis article proposes an effective and efficient visual place recognition (VPR) approach, which can make full use of semantic, sequential, and spatial geometric information in VPR tasks. Rather than previous methods focusing on extracting discriminative and compact features to represent images, we improve VPR performance from candidate selection and geometric verification. To this end, we propose a fast feature matching algorithm for real-time geometrical verification of candidate places, termed neighborhood manifold preserving matching (NMP). To generate high-quality candidates, we design a dynamic sequence partitioning strategy based on NMP, which is able to utilize the inherently sequential nature of spatial data to cluster images into places. By sequence-to-sequence matching, the ambiguity of single frame matching can be reduced. Extensive experiments demonstrate that our VPR method outperforms the current state-of-the-art methods, and our geometric verification and candidate selection strategies are easily plugged into other VPR pipelines to significantly improve the VPR performance. Xinyu Ye, Jiayi Ma 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | ReDFeat: Recoupling Detection and Description for Multimodal Feature LearningabstractDeep-learning-based local feature extraction algorithms that combine detection and description have made significant progress in visible image matching. However, the end-to-end training of such frameworks is notoriously unstable due to the lack of strong supervision of detection and the inappropriate coupling between detection and description. The problem is magnified in cross-modal scenarios, in which most methods heavily rely on the pre-training. In this paper, we recouple independent constraints of detection and description of multimodal feature learning with a mutual weighting strategy, in which the detected probabilities of robust features are forced to peak and repeat, while features with high detection scores are emphasized during optimization. Different from previous works, those weights are detached from back propagation so that the detected probability of indistinct features would not be directly suppressed and the training would be more stable. Moreover, we propose the Super Detector, a detector that possesses a large receptive field and is equipped with learnable non-maximum suppression layers, to fulfill the harsh terms of detection. Finally, we build a benchmark that contains cross visible, infrared, near-infrared and synthetic aperture radar image pairs for evaluating the performance of features in feature matching and image registration tasks. Extensive experiments demonstrate that features trained with the recoulped detection and description, named ReDFeat, surpass previous state-of-the-arts in the benchmark, while the model can be readily trained from scratch. The code is released at https://github.com/ACuOoOoO/ReDFeat. Yuxin Deng 0002, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Density-Guided Incremental Dominant Instance Exploration for Two-View Geometric Model FittingabstractExisting two-view multi-model fitting methods typically follow a two-step manner, i.e., model generation and selection, without considering their interaction. Therefore, in the first step, these methods have to generate a considerable number of instances in order to cover all desired ones, which not only offers no guarantees, but also introduces unnecessary expensive calculations. To address this challenge, this study presents a new algorithm, termed as D2Fitting, that incrementally explores dominant instances. Particularly, rather than viewing model generation and selection as two disjoint parts, D2Fitting fully considers their interaction, and thus performs these two subroutines alternatively under a simple yet effective optimization framework. This design can avoid generating too many redundant instances, thus reducing computational overhead and allowing the proposed D2Fitting being real-time. Meanwhile, we further design a novel density-guided sampler to sample high-quality minimal subsets during the model generation process, so as to fully exploit the spatial distribution of the input data. Also, to mitigate the influence of noise on the subsets sampled by the proposed sampler, a global-residual optimization strategy is investigated for the minimal subset refinement. With all the ingredients mentioned above, the proposed D2Fitting can accurately estimate the number and parameters of geometric models and efficiently segment the input data simultaneously. Extensive experiments on several public datasets demonstrate the significant superiority of D2Fitting over several state-of-the-arts. Zizhuo Li, Jiayi Ma 0001, Guobao Xiao |
IEEE Trans. Image Process. | 2 |
| 2023 | PGFNet: Preference-Guided Filtering Network for Two-View Correspondence LearningabstractAccurate correspondence selection between two images is of great importance for numerous feature matching based vision tasks. The initial correspondences established by off-the-shelf feature extraction methods usually contain a large number of outliers, and this often leads to the difficulty in accurately and sufficiently capturing contextual information for the correspondence learning task. In this paper, we propose a Preference-Guided Filtering Network (PGFNet) to address this problem. The proposed PGFNet is able to effectively select correct correspondences and simultaneously recover the accurate camera pose of matching images. Specifically, we first design a novel iterative filtering structure to learn the preference scores of correspondences for guiding the correspondence filtering strategy. This structure explicitly alleviates the negative effects of outliers so that our network is able to capture more reliable contextual information encoded by the inliers for network learning. Then, to enhance the reliability of preference scores, we present a simple yet effective Grouped Residual Attention block as our network backbone, by designing a feature grouping strategy, a feature grouping manner, a hierarchical residual-like manner and two grouped attention operations. We evaluate PGFNet by extensive ablation studies and comparative experiments on the tasks of outlier removal and camera pose estimation. The results demonstrate outstanding performance gains over the existing state-of-the-art methods on different challenging scenes. The code is available at https://github.com/guobaoxiao/PGFNet. Xin Liu 0091, Guobao Xiao, Riqing Chen, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion ModelsabstractColor plays an important role in human visual perception, reflecting the spectrum of objects. However, the existing infrared and visible image fusion methods rarely explore how to handle multi-spectral/channel data directly and achieve high color fidelity. This paper addresses the above issue by proposing a novel method with diffusion models, termed as Dif-Fusion, to generate the distribution of the multi-channel input data, which increases the ability of multi-source information aggregation and the fidelity of colors. In specific, instead of converting multi-channel images into single-channel data in existing fusion methods, we create the multi-channel data distribution with a denoising network in a latent space with forward and reverse diffusion process. Then, we use the the denoising network to extract the multi-channel diffusion features with both visible and infrared information. Finally, we feed the multi-channel diffusion features to the multi-channel fusion module to directly generate the three-channel fused image. To retain the texture and intensity information, we propose multi-channel gradient loss and intensity loss. Along with the current evaluation metrics for measuring texture and intensity fidelity, we introduce Delta E as a new evaluation metric to quantify color fidelity. Extensive experiments indicate that our method is more effective than other state-of-the-art image fusion methods, especially in color fidelity. The source code is available at https://github.com/GeoVectorMatrix/Dif-Fusion. Jun Yue 0004, Leyuan Fang, Shaobo Xia, Yue Deng 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Loop Closure Detection With Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD) is an indispensable module in simultaneous localization and mapping. It is responsible to recognize pre-visited areas during the navigation of a robot, providing auxiliary information to revise pose estimation. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. Specifically, to avoid unnecessary memory costs, consecutive images are segmented into sequences as per the similarity of their global features. Then the sequence descriptor is incrementally inserted into hierarchical navigable small world for the construction of reference database, from which the most similar image for the query one is searched parallelly. To further identify whether the candidate pair is geometry-consistent, a feature matching method termed as bidirectional manifold representation consensus (BMRC) is proposed. It constructs local neighborhood structures of feature points via manifold representation, and formulates the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Meanwhile, an accelerated version of it is introduced (BMRC*), which performs about 63% faster than BMRC in an image pair with 352 initial correspondences. Extensive experiments on nine publicly available datasets demonstrate that BMRC and BMRC* perform well in feature matching and the proposed pipeline has remarkable performance in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Correspondence Attention Transformer: A Context-Sensitive Network for Two-View Correspondence LearningabstractSeeking reliable correspondences then recovering camera poses from a set of putative correspondences extracted from two images of the same scene is a fundamental problem in computer vision. Recent advances have demonstrated that this problem can be effectively solved by using a deep architecture based on the multi-layer perceptron, where the context normalization is designed to make the network permutation-equivariant and embed global information in the sparse point data. However, the context normalization simply normalizes the feature maps according to their distribution and treats each correspondence equally, leading to difficulties in adequately capturing scene geometry encoded by the inliers, especially in case of severe outliers. To address this issue, this paper designs a context-sensitive network based on the self-attention mechanism, termed as correspondence attention transformer (CAT), to enhance the consistent geometry information of inliers and simultaneously suppress outliers during embedding global information. In particular, we design an attention-style structure to aggregate features from all correspondences, i.e., a spatial attention namely CAT-S, which provides each correspondence with information exchange from others in the putative set. To capture the contextual information in a more comprehensive and robust way, we also introduce a multi-head mechanism in our structure to exploit the geometrical context from different aspects. Moreover, considering the high memory request in spatial attention, we propose a covariance normalized channel attention CAT-C in our framework, which can largely reduce the memory consumption and parameter scale, but it asks for eigenvalue decomposition in each attention block thus resulting in more runtime. Anyway, these two attention mechanisms can realize information exchange from the spatial or channel aspect, which both contribute to constructing the geometrical context between inliers and encourage the network to pay more attention to the feature subset about potential inliers. Extensive experiments have been conducted over both indoor and outdoor datasets on the tasks of camera pose estimation, outlier removal, and image registration, which demonstrate the superiority of our method that realizes a large performance improvement compared with the current state-of-the-art approaches. Jiayi Ma 0001, Aoxiang Fan, Guobao Xiao, Riqing Chen |
IEEE Trans. Multim. | 1 |
| 2023 | DMEF: Multi-Exposure Image Fusion Based on a Novel Deep Decomposition MethodabstractIn this paper, we propose a novel deep decomposition approach based on Retinex theory for multi-exposure image fusion, termed as DMEF. According to the assumption of Retinex theory, we firstly decompose the source images into illumination and reflection maps by the data-driven decomposition network, among which we introduce the pathwise interaction block that reactivates the deep features lost in one path and embeds them into another path. Therefore, loss of illumination and reflection features during decomposition can be effectively suppressed. And then the high dynamic range illumination map could be obtained by fusing the separated illumination maps in the fusion network. Thus, the reconstructed details in under-exposed and over-exposed regions will be clearer with the help of the fused reflection map which contains complete high-frequency scene information. Finally, the fused illumination and reflection maps are multiplied pixel-by-pixel to obtain the final fused image. Moreover, to retain the discontinuity in the illumination map where gradient of reflection map changes steeply, we introduce the structure-preservation smoothness loss function to retain the structure information and eliminate visual artifacts in these regions. The superiority of our proposed network is demonstrated by applying extensive experiments compared with other state-of-the-art fusion methods subjectively and objectively. Kangle Wu, Jun Chen 0019, Jiayi Ma 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | ACE-MEF: Adaptive Clarity Evaluation-Guided Network With Illumination Correction for Multi-Exposure Image FusionabstractFor a natural scene with nonuniform environment light, the captured visible images are always under- or over-exposed because of the limited dynamic range of digital imaging devices. Multi-exposure image fusion (MEF) is a mainstream and effective solution. For a local region that has friendly visual effect in one exposure setting but extremely bad-exposed in another, most existing MEF methods have the ability to transfer the scene detail information to the fused images. However, they will be affected by the over-high or -low light inevitably thus resulting in local visibility reduction. To address this issue, we propose an adaptive clarity evaluation-guided network with illumination correction for MEF in a coarse-to-fine manner, which is termed as ACE-MEF. To be specific, our ACE-MEF is mainly composed of two modules: clarity preservation network (CPN) and illumination adjustment network (IAN). Based on the adaptive clarity evaluation, CPN could be trained to coarsely preserve the environment light and texture details of the clearer regions in source images. Therefore, the need for labeled reference images that are time-consuming to obtain could be mitigated. By measuring the parameter maps of gamma function, IAN is able to refine and correct the local bad-exposed regions so that more details could be further revealed. Extensive experiments demonstrate that our method outperforms multiple state-of-the-art algorithms qualitatively and quantitatively. Kangle Wu, Jun Chen 0019, Yang Yu 0045, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Semantic-Supervised Infrared and Visible Image Fusion Via a Dual-Discriminator Generative Adversarial NetworkabstractImage fusion synthesizes a new image from multiple images of the same scene. The synthesized image should be suitable for human visual perception and follow-up high-level image-processing tasks. However, existing methods focus on fusing low-level features, ignoring high-level semantic perception information. We propose a new end-to-end model to obtain a more semantically consistent image in infrared and visible image fusion, termedsemantic-supervised dual-discriminator generative adversarial network(SDDGAN). In particular, we design an information quantity discrimination (IQD) block to guide fusion progress. For each source image, the block determines the weight for preserving each semantic object’s feature. By this way, the generator learns to fuse various semantic objects via different weights to preserve their characteristics. Moreover, the dual discriminator is employed to identify the distribution of infrared and visible information in the fused image. Each discriminator acts on a certain modality (infrared/visible) of different semantic objects in the fused image to preserve and enhance their modality features. Thus, our fused image is more informative. Both the thermal radiation in the infrared image and the visible image texture details can be well preserved. Qualitative and quantitative experiments demonstrate the superiority of our SDDGAN over state-of-the-art methods in terms of visual effects, efficiency, and quantitative metrics. Huabing Zhou, Yanduo Zhang, Jiayi Ma 0001, Haibin Ling |
IEEE Trans. Multim. | 4 |
| 2023 | Smoothness-Driven Consensus Based on Compact Representation for Robust Feature MatchingabstractFor robust feature matching, a popular and particularly effective method is to recover smooth functions from the data to differentiate the true correspondences (inliers) from false correspondences (outliers). In the existing works, the well-established regularization theory has been extensively studied and exploited to estimate the functions while controlling its complexity to enforce the smoothness constraint, which has shown prominent advantages in this task. However, despite the theoretical optimality properties, the high complexities in both time and space are induced and become the main obstacle of their application. In this article, we propose a novel method for multivariate regression and point matching, which exploits the sparsity structure of smooth functions. Specifically, we use compact Fourier bases for constructing the function, which inherently allows a coarse-to-fine representation. The smoothness constraint can be explicitly imposed by adopting a few low-frequency bases for representation, resulting in reduced computational complexities of the induced multivariate regression algorithm. To cope with potential gross outliers, we formulate the learning problem into a Bayesian framework with latent variables indicating the inliers and outliers and a mixture model accounting for the distribution of data, where a fast expectation-maximization solution can be derived. Extensive experiments are conducted on synthetic data and real-world image matching, and point set registration datasets, which demonstrates the advantages of our method against the current state-of-the-art methods in terms of both scalability and robustness. Aoxiang Fan, Xingyu Jiang 0005, Yong Ma 0001, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Adversarial Autoencoder Network for Hyperspectral UnmixingabstractSpectral unmixing (SU), which refers to extracting basic features (i.e., endmembers) at the subpixel level and calculating the corresponding proportion (i.e., abundances), has become a major preprocessing technique for the hyperspectral image analysis. Since the unmixing procedure can be explained as finding a set of low-dimensional representations that reconstruct the data with their corresponding bases, autoencoders (AEs) have been effectively designed to address unsupervised SU problems. However, their ability to exploit the prior properties remains limited, and noise and initialization conditions will greatly affect the performance of unmixing. In this article, we propose a novel technique network for unsupervised unmixing which is based on the adversarial AE, termed as adversarial autoencoder network (AAENet), to address the above problems. First, the image to be unmixed is assumed to be partitioned into homogeneous regions. Then, considering the spatial correlation between local pixels, the pixels in the same region are assumed to share the same statistical properties (means and covariances) and abundance can be modeled to follow an appropriate prior distribution. Then the adversarial training procedure is adapted to transfer the spatial information into the network. By matching the aggregated posterior of the abundance with a certain prior distribution to correct the weight of unmixing, the proposed AAENet exhibits a more accurate and interpretable unmixing performance. Compared with the traditional AE method, our approach can greatly enhance the performance and robustness of the model by using the adversarial procedure and adding the abundance prior to the framework. The experiments on both the simulated and real hyperspectral data demonstrate that the proposed algorithm can outperform the other state-of-the-art methods. Qiwen Jin, Yong Ma 0001, Fan Fan 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | MS2DG-Net: Progressive Correspondence Learning via Multiple Sparse Semantics Dynamic GraphabstractEstablishing superior-quality correspondences in an image pair is pivotal to many subsequent computer vision tasks. Using Euclidean distance between correspondences to find neighbors and extract local information is a common strategy in previous works. However, most such works ignore similar sparse semantics information between two given images and cannot capture local topology among correspondences well. Therefore, to deal with the above problems, Multiple Sparse Semantics Dynamic Graph Network (MS2DG-Net) is proposed, in this paper, to predict probabilities of correspondences as inliers and recover camera poses. MS2 DG-Net dynamically builds sparse semantics graphs based on sparse semantics similarity between two given images, to capture local topology among correspondences, while maintaining permutation-equivariant. Extensive experiments prove that MS2 DG-Net outperforms state-of-the-art methods in outlier removal and camera pose estimation tasks on the public datasets with heavy outliers. Source code:https://github.com/changcaiyang/MS2DG-Net Luanyuan Dai, Yizhang Liu, Jiayi Ma 0001, Lifang Wei, Taotao Lai, Changcai Yang, Riqing Chen |
CVPR | 3 |
| 2022 | Coherent Point Drift Revisited for Non-rigid Shape Matching and RegistrationabstractIn this paper, we explore a new type of extrinsic method to directly align two geometric shapes with point-to-point correspondences in ambient space by recovering a deformation, which allows more continuous and smooth maps to be obtained. Specifically, the classic coherent point drift is revisited and generalizations have been proposed. First, by observing that the deformation model is essentially defined with respect to Euclidean space, we generalize the kernel method to non-Euclidean domains. This generally leads to better results for processing shapes, which are known as two-dimensional manifolds. Second, a generalized probabilistic model is proposed to address the sensibility of coherent point drift method to local optima. Instead of directly optimizing over the objective of coherent point drift, the new model allows to focus on a group of most confident ones, thus improves the robustness of the registration system. Experiments are conducted on multiple public datasets with comparison to state-of-the-art competitors, demonstrating the superiority of our method which is both flexible and efficient to improve the matching accuracy due to our extrinsic alignment objective in ambient space. Aoxiang Fan, Jiayi Ma 0001, Xin Tian 0006, Xiaoguang Mei |
CVPR | 2 |
| 2022 | RFNet: Unsupervised Network for Mutually Reinforcing Multi-modal Image Registration and FusionabstractIn this paper, we propose a novel method to realize multimodal image registration and fusion in a mutually reinforcing framework, termed as RFNet. We handle the registration in a coarse-to-fine fashion. For the first time, we exploit the feedback of image fusion to promote the registration accuracy rather than treating them as two separate issues. The fine-registered results also improve the fusion performance. Specifically, for image registration, we solve the bottlenecks of defining registration metrics applicable for multi-modal images and facilitating the network convergence. The metrics are defined based on image translation and image fusion respectively in the coarse and fine stages. The convergence is facilitated by the designed metrics and a deformable convolution-based network. For image fusion, we focus on texture preservation, which not only increases the information amount and quality of fusion results but also improves the feedback of fusion results. The proposed method is evaluated on multi-modal images with large global parallaxes, images with local misalignments and aligned images to validate the performances of registration and fusion. The results in these cases demonstrate the effectiveness of our method. Han Xu 0001, Jiayi Ma 0001, Jiteng Yuan, Zhuliang Le, Wei Liu 0005 |
CVPR | 2 |
| 2022 | Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001 |
ECCV (27) | 7 |
| 2022 | Fusion from Decomposition: A Self-Supervised Decomposition Approach for Image Fusion
Pengwei Liang, Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ECCV (18) | 4 |
| 2022 | CUFD: An encoder-decoder network for visible and infrared image fusion based on common and unique feature decomposition
Han Xu 0001, Meiqi Gong, Xin Tian 0006, Jun Huang 0008, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 5 |
| 2022 | Feature Matching via Motion-Consistency Driven Probabilistic Graphical Model
Jiayi Ma 0001, Aoxiang Fan, Xingyu Jiang 0005, Guobao Xiao |
Int. J. Comput. Vis. | 1 |
| 2022 | RANet: A relation-aware network for two-view correspondence learning
Guorong Lin, Xin Liu 0091, Fangfang Lin, Guobao Xiao, Jiayi Ma 0001 |
Neurocomputing | 5 |
| 2022 | Learning Two-View Correspondences and Geometry Using Neighbor-Aware NetworkabstractFinding reliable correspondences between two images is a fundamental problem in remote-sensing image registration. In the face of sparse and unordered correspondences, the previous learning-based works often focus on global information and ignore valuable local information. To gather rich local information, we propose a simple and effective approach, the principle of which is to establish a local neighborhood structure for putative correspondences and then extract and aggregate neighborhood context. Specifically, we design a residual block with an innovative normalization operation and then we construct a sub-net, by introducing a competition mechanism in the local neighborhood to enhance the expression ability of features. Finally, we propose a learning-based network, to improve the performance of outlier rejection by extracting neighborhood context and global context. Extensive experiments on relative pose estimation demonstrate that the proposed network surpasses current state-of-the-art approaches on both challenging indoor and outdoor datasets. Ziwei Shi, Guobao Xiao, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Dilated projection correction network based on autoencoder for hyperspectral image super-resolution
Xinya Wang, Jiayi Ma 0001, Junjun Jiang, Xiao-Ping Zhang 0002 |
Neural Networks | 2 |
| 2022 | Efficient Deterministic Search With Robust Loss Functions for Geometric Model FittingabstractGeometric model fitting is a fundamental task in computer vision, which serves as the pre-requisite of many downstream applications. While the problem has a simple intrinsic structure where the solution can be parameterized within a few degrees of freedom, the ubiquitously existing outliers are the main challenge. In previous studies, random sampling techniques have been established as the practical choice, since optimization-based methods are usually too time-demanding. This prospective study is intended to design efficient algorithms that benefit from a general optimization-based view. In particular, two important types of loss functions are discussed, \emph{i.e.} truncated and$l_1$losses, and efficient solvers have been derived for both upon specific approximations. Based on this philosophy, a class of algorithms are introduced to perform deterministic search for the inliers or geometric model. Recommendations are made based on theoretical and experimental analyses. Compared with the existing solutions, the proposed methods are both simple in computation and robust to outliers. Extensive experiments are conducted on publicly available datasets for geometric estimation, which demonstrate the superiority of our methods compared with the state-of-the-art ones. Additionally, we apply our method to the recent benchmark for wide-baseline stereo evaluation, leading to a significant improvement of performance. Aoxiang Fan, Jiayi Ma 0001, Xingyu Jiang 0005, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | U2Fusion: A Unified Unsupervised Image Fusion NetworkabstractThis study proposes a novel unified and unsupervised end-to-end image fusion network, termed as U2Fusion, which is capable of solving different fusion problems, including multi-modal, multi-exposure, and multi-focus cases. Using feature extraction and information measurement, U2Fusion automatically estimates the importance of corresponding source images and comes up with adaptive information preservation degrees. Hence, different fusion tasks are unified in the same framework. Based on the adaptive degrees, a network is trained to preserve the adaptive similarity between the fusion result and source images. Therefore, the stumbling blocks in applying deep learning for image fusion, e.g., the requirement of ground-truth and specifically designed metrics, are greatly mitigated. By avoiding the loss of previous fusion capabilities when training a single model for different tasks sequentially, we obtain a unified model that is applicable to multiple fusion tasks. Moreover, a new aligned infrared and visible image dataset, RoadScene (available at https://github.com/hanna-xu/RoadScene), is released to provide a new option for benchmark evaluation. Qualitative and quantitative experimental results on three typical image fusion tasks validate the effectiveness and universality of U2Fusion. Our code is publicly available at https://github.com/hanna-xu/U2Fusion. Han Xu 0001, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | A Progressive Fusion Generative Adversarial Network for Realistic and Consistent Video Super-ResolutionabstractHow to effectively fuse temporal information from consecutive frames remains to be a non-trivial problem in video super-resolution (SR), since most existing fusion strategies (direct fusion, slow fusion, or 3D convolution) either fail to make full use of temporal information or cost too much calculation. To this end, we propose a novel progressive fusion network for video SR, in which frames are processed in a way of progressive separation and fusion for the thorough utilization of spatio-temporal information. We particularly incorporate multi-scale structure and hybrid convolutions into the network to capture a wide range of dependencies. We further propose a non-local operation to extract long-range spatio-temporal correlations directly, taking place of traditional motion estimation and motion compensation (ME&MC). This design relieves the complicated ME&MC algorithms, but enjoys better performance than various ME&MC schemes. Finally, we improve generative adversarial training for video SR to avoid temporal artifacts such as flickering and ghosting. In particular, we propose a frame variation loss with a single-sequence training method to generate more realistic and temporally consistent videos. Extensive experiments on public datasets show the superiority of our method over state-of-the-art methods in terms of performance and complexity. Our code is available at https://github.com/psychopa4/MSHPFNL. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | CSDA-Net: Seeking reliable correspondences by channel-Spatial difference augment network
Shunxing Chen, Linxin Zheng, Guobao Xiao, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2022 | Robust image matching via local graph structure consensus
Xingyu Jiang 0005, Xiao-Ping Zhang 0002, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2022 | Guided neighborhood affine subspace embedding for feature matching
Zizhuo Li, Yong Ma 0001, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
Pattern Recognit. | 5 |
| 2022 | BSCA-Net: Bit Slicing Context Attention network for polyp segmentation
Jichun Wu, Guobao Xiao, Junwen Guo, Geng Chen 0001, Jiayi Ma 0001 |
Pattern Recognit. | 6 |
| 2022 | Infrared and visible image fusion via parallel scene and texture learning
Meilong Xu, Linfeng Tang, Hao Zhang 0073, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2022 | Interpolation-based nonrigid deformation estimation under manifold regularization constraint
Huabing Zhou, Yulu Tian, Zhenghong Yu, Yanduo Zhang, Jiayi Ma 0001 |
Pattern Recognit. | 6 |
| 2022 | Hyperspectral Anomaly Detection With Robust Graph AutoencodersabstractAnomaly detection of hyperspectral data has been gaining particular attention for its ability in detecting targets in an unsupervised manner. Autoencoder (AE), together with its variants can not only extract intrinsic features automatically but also detect anomalies that differ dramatically from others. Many AE-driven algorithms are, thus, proposed for anomaly detection in hyperspectral imagery (HSI), but they suffer from two problems: 1) when there exist anomalies in the training set, AE can generalize so well that it can also learn the abnormal patterns well, thereby reducing the ability to distinguish anomalies from the background and 2) geometric structure among samples are lost in latent space of AE, which is vital in hyperspectral anomaly detection. To tackle these problems, we propose a robust anomaly detector based on the AE framework, named robust graph AE (RGAE) detector, in this article. To be specific, we propose a robust AE framework with$\ell _{2,1}$-norm that is robust to noise and anomalies during training. Meanwhile, we embed a superpixel segmentation-based graph regularization term (SuperGraph) into AE. This strategy can preserve the geometric structure and the local spatial consistency of HSI simultaneously and also effectively reduce the searching space and execution time for each pixel. Extensive experiments are conducted on five datasets, and the results demonstrate that our method has a better detection performance, after comparing with other state-of-the-art hyperspectral anomaly detectors. Ganghui Fan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | D2TNet: A ConvLSTM Network With Dual-Direction Transfer for Pan-SharpeningabstractIn this article, we propose an efficient convolutional long short-term memory (ConvLSTM) network with dual-direction transfer for pan-sharpening, termed D2TNet. We design a specially structured ConvLSTM network that allows for dual-directional communication, including multiscale information and multilevel information. On the one hand, due to the sensitivity of spatial information to scales and the sensitivity of spectral information to levels, multiscale and multilevel information is extracted to facilitate the fuller use of source images. On the other hand, ConvLSTM is employed to capture the strong dependencies between multiscale information and multilevel information. Besides, we introduce a multiscale loss to enable different scales contributing to each other to generate high-resolution multispectral images that are closer to the ground truth. Extensive experiments, including qualitative evaluation, quantitative evaluation, and efficiency comparison, are implemented to verify that our D2TNet outperforms state-of-the-art methods indeed. Meiqi Gong, Jiayi Ma 0001, Han Xu 0001, Xin Tian 0006, Xiao-Ping Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | TANet: An Unsupervised Two-Stream Autoencoder Network for Hyperspectral UnmixingabstractSpectral unmixing is a major technique for the further development of hyperspectral analysis. It aims to determine the corresponding proportion (fractional abundance) of the basic spectral signatures (endmembers) blindly at the subpixel level. Recently, the learning-based method has received much attention in hyperspectral unmixing, and autoencoders have been effectively designed to solve the unsupervised scenarios of unmixing. However, their ability to extract physically meaningful endmembers remains limited, and the performance has not been satisfactory. In this article, we propose a novel two-stream network, termed TANet, to address the above problems. The network consists of a two-stream architecture. First, superpixel segmentation is adopted as preprocessing to extract the endmember bundles from the image. Then, the first stream learns a mapping from the pseudopure pixels to their corresponding abundances. The second stream is conducting the same untied-weighted autoencoder to minimize reconstruction errors from the original pixel data. By learning from the pure or nearly pure candidate pixels to correct the weights of unmixing, the proposed TANet exhibits a more accurate and interpretable unmixing performance. Extensive experiments on both synthetic and real hyperspectral data demonstrate that the proposed TANet can outperform the other state-of-the-art approaches. Qiwen Jin, Yong Ma 0001, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SQAD: Spatial-Spectral Quasi-Attention Recurrent Network for Hyperspectral Image DenoisingabstractThis article presents a novel end-to-end model based on encoder–decoder architecture for hyperspectral image (HSI) denoising, named spatial-spectral quasi-attention recurrent network, denoted as SQAD. The central goal of this work is to incorporate the intrinsic properties of HSI noise to construct a practical feature extraction module while maintaining high-quality spatial and spectral information. Accordingly, we first design a spatial-spectral quasi-recurrent attention unit (QARU) to address that issue. QARU is the basic building block in our model, consisting of spatial component and spectral component, and each of them involves a two-step calculation. Remarkably, the quasi-recurrent pooling function in the spectral component could explore the relevance of spatial features in the spectral domain. The spectral attention calculation could strengthen the correlation between adjacent spectra and provide the intrinsic properties of HSI noise distribution in the spectral dimension. Apart from this, we also design a unique skip connection consisting of channelwise concatenation and transition block in our model to convey the detailed information and promote the fusion of the low-level features with the high-level ones. Such a design helps maintain better structural characteristics, and spatial and spectral fidelities when reconstructing the clean HSI. Qualitative and quantitative experiments are performed on publicly available datasets. The results demonstrate that SQAD outperforms the state-of-the-art methods of visual effect and objective evaluation metrics. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Cross Fusion Net: A Fast Semantic Segmentation Network for Small-Scale Semantic Information Capturing in Aerial ScenesabstractCapturing accurate multiscale semantic information from the images is of great importance for high-quality semantic segmentation. Over the past years, a large number of methods attempt to improve the multiscale information capturing ability of the networks via various means. However, these methods always suffer unsatisfactory efficiency (e.g., speed or accuracy) on the images that include a large number of small-scale objects, for example, aerial images. In this article, we propose a new network named cross fusion net (CF-Net) for fast and effective extraction of the multiscale semantic information, especially for small-scale semantic information. In particular, the proposed CF-Net can capture more accurate small-scale semantic information from two aspects. On the one hand, we develop a channel attention refinement block to select the informative features. On the other hand, we propose a cross fusion block to enlarge the receptive field of the low-level feature maps. As a result, the network can encode more accurate semantic information from the small-scale objects, and the segmentation accuracy of the small-scale objects is improved accordingly. We have compared the proposed CF-Net with several state-of-the-art semantic segmentation methods on two popular aerial image segmentation data sets. Experimental results reveal that the average$F_{1}$score gain brought by our CF-Net is about 0.43% and the$F_{1}$score gain of the small-scale objects (e.g., cars) is about 2.61%. In addition, our CF-Net has the fastest inference speed, which proves its superiority in the aerial scenes. Our code will be released at:https://github.com/pcl111/CF-Net. Chengli Peng, Kaining Zhang, Yong Ma 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Variational Pansharpening by Exploiting Cartoon-Texture SimilaritiesabstractPansharpening aims to fuse a multispectral (MS) image with low spatial resolution and a panchromatic (PAN) image with a high-spatial resolution to produce an image with both high spectral and high spatial resolution. In this study, we propose a variational pansharpening method by exploiting cartoon-texture similarities. After decomposition of the PAN image, the cartoon component always contains the global structure information, while the texture component includes the locally patterned information. This enables that the fused high-spatial resolution MS image can preserve the global and local spatial details (e.g., high-order information) well after leveraging the similarities of cartoon and texture components from PAN and MS images. To explore such cartoon-texture similarities, we describe cartoon similarity as gradient sparsity, formulated as a reweighted total variation term. Meanwhile, we use group low-rank constraint for texture similarity that is presented as repetitive texture patterns. By incorporating a data fidelity term for preserving the spectral information on the basis that the down-sampled fused MS image is consistent with the MS image, we further formulate pansharpening as an optimization problem and solve it efficiently using the alternative direction multiplier method. Extensive experiments have been conducted on a series of satellite data sets, and we also carry out a simulated vegetation coverage change experiment to verify the efficiency of the proposed method in remote sensing. The qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art pansharpening methods in terms of both visual effect and objective metrics. Xin Tian 0006, Yuerong Chen, Changcai Yang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | VP-Net: An Interpretable Deep Network for Variational PansharpeningabstractIn this study, we propose an interpretable deep network for variational pansharpening (VP), named VP-Net. Different from traditional priors using linear operators, such as the gradient, we construct a prior based on the similarity between panchromatic (PAN) and high-resolution multispectral (HRMS) images by a nonlinear operator that can be learned through a deep network. Considering the spectral difference of various satellite multispectral (MS) imaging platforms, we specifically seek the aforementioned similarity from the PAN image and the intensity of the HRMS image to reduce the spectral distortion. Based on this prior, we propose a novel VP model by further incorporating a data fidelity term from the low-resolution MS image. Specifically, we build the VP-Net by unrolling the variable splitting method for an optimal solution to this model. Consequently, all modules in VP-Net have clear physical meanings and strong generalization capabilities. Meanwhile, all parameters and the aforementioned nonlinear operator are learned in VP-Net, avoiding the difficulty of selecting optimal handcrafted parameters in traditional methods. Therefore, VP-Net not only achieves an optimal balance between spatial and spectral qualities but also has a strong generalization capability across different types of training and testing data. In the experiment, we first demonstrate the superiority of the proposed method over the current state of the arts in terms of both visual effect and quantitative analysis on different satellite datasets. Moreover, we carry out an normalized difference vegetation index (NDVI) experiment to demonstrate its potential in remote sensing. Xin Tian 0006, Kun Li 0025, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | HyperFusion: A Computational Approach for Hyperspectral, Multispectral, and Panchromatic Image FusionabstractFusing hyperspectral image (HSI) and multispectral image (MSI) of high spatial resolution is typically utilized to obtain HSIs of high spatial resolution. However, the spatial quality of most existing methods is unsatisfactory due to the limited spatial resolution of an MSI. To further improve the spatial resolution of the fused HSI while keeping the spectral information well, we propose a new computational paradigm, named HyperFusion, which simultaneously fuses HSI, MSI, and panchromatic (PAN) image. To achieve this goal, we first establish two data fidelity terms based on a physical observation that HSI and MSI can be treated as degraded versions of the fused HSI. Consequently, the spatial and spectral information from HSI and MSI can be well preserved. To efficiently transfer the spatial details of PAN into the fused HSI while keeping the spectral information well, we further construct a prior constraint from PAN based on the structural similarity. Meanwhile, we impose another low-rank prior constraint on the coefficient matrix to accurately describe the latent characteristics of the HSI with high spatial resolution. By incorporating the aforementioned data fidelity terms and prior constraints, we finally formulate the objective as an optimization problem and utilize the alternative direction multiplier method to solve it efficiently. Comprehensive experiments on simulated and real datasets are carried out to demonstrate the superiority of HyperFusion over other state of the arts in terms of visual quality and quantitative analysis. We also adopt a simulated experiment of vegetation coverage index analysis to verify the effectiveness of HyperFusion in remote sensing applications. Xin Tian 0006, Wei Zhang 0259, Yuerong Chen, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Group-Based Embedding Learning and Integration Network for Hyperspectral Image Super-ResolutionabstractAlthough natural image super-resolution methods have achieved impressive performance, single hyperspectral image super-resolution still remains a challenge due to the high dimensionality. In recent years, many single hyperspectral image super-resolution methods adopted the group-convolution strategy to design the network for reducing the computational burden. However, these methods still process all spectral bands at once during the deep feature extraction and reconstruction, which increases the difficulty of fully exploring the inherent data characteristic of hyperspectral images. Moreover, the advanced group-based methods make insufficient exploitation of complementary information contained in different bands, resulting in limited reconstruction performance. In this paper, we propose a novel group-based single hyperspectral image super-resolution method termed GELIN to reconstruct high-resolution images in a group-by-group manner, which alleviates the difficulty of feature extraction and reconstruction for hyperspectral images. Specifically, a spatial-spectral embedding learning module is designed to extract rewarding spatial details and explore the correlations among spectra simultaneously. Considering the high similarity among different bands, a neighboring group integration module is proposed to fully exploit the complementary information contained in neighboring image groups to recover missing details in the target image group. Experimental results on both natural and remote sensing hyperspectral datasets demonstrate that the proposed method is superior to other state-of-the-art methods both visually and metrically. Xinya Wang, Junjun Jiang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hyperspectral Image Super-Resolution via Recurrent Feedback Embedding and Spatial-Spectral Consistency RegularizationabstractHyperspectral images with tens to hundreds of spectral bands usually suffer from low spatial resolution due to the limitation of the amount of incident energy. Without auxiliary images, the single hyperspectral image super-resolution (SR) method is still a challenging problem because of the high-dimensionality characteristic and special spectral patterns of hyperspectral images. Failing to thoroughly explore the coherence among hyperspectral bands and preserve the spatial–spectral structure of the scene, the performance of existing methods is still limited. In this article, we propose a novel single hyperspectral image SR method termed RFSR, which models the spectrum correlations from a sequence perspective. Specifically, we introduce a recurrent feedback network to fully exploit the complementary and consecutive information among the spectra of the hyperspectral data. With the group strategy, each grouping band is first super-resolved by exploring the consecutive information among groups via feedback embedding. For better preservation of the spatial–spectral structure among hyperspectral data, a regularization network is subsequently appended to enforce spatial–spectral correlations over the intermediate estimation. Experimental results on both natural and remote sensing hyperspectral images demonstrate the advantage of our approach over the state-of-the-art methods. Xinya Wang, Jiayi Ma 0001, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Robust Feature Matching for Remote Sensing Image Registration via Guided Hyperplane FittingabstractFeature matching is a fundamental problem in feature-based remote sensing image registration. Due to the ground relief variations and imaging viewpoint changes, remote sensing images often involve local distortions, leading to difficulties in high-accuracy image registration. To address this issue, in this article, we propose a robust feature matching method called First Neighbor Relation Guided (FNRG) for remote sensing image registration via guided hyperplane fitting. The key idea of FNRG is to exploit the first neighbor relation of feature points between two images for seeking consistent seeds in a parameter-free manner. To boost more consistent matches based on the consistent seeds, we formulate the feature matching problem into an affine hyperplane fitting problem by imposing the motion consistency, and then we design a hyperplane updating strategy to refine the fitting model. We also introduce a locality preserving structure-based cost function to promote the matching performance of the hyperplane updating strategy. Our method can mine consistent matches from thousands of putative ones within only a few milliseconds, and it also can handle the data with a large-scale change, rotation, or severe nonrigid deformation. Extensive experiments on the remote sensing image data sets with different types of image transformations show that the proposed method achieves significant superiority over several state-of-the-art methods. Guobao Xiao, Huan Luo 0001, Leyi Wei, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Local Affine Preservation With Motion Consistency for Feature Matching of Remote Sensing ImagesabstractAs a fundamental and essential task in the field of remote sensing and photogrammetry, feature matching endeavors to establish reliable correspondences between two sets of feature points extracted from an image pair of the same scene. In this article, we propose an efficient and general algorithm, which is called local affine preservation (LAP) matching, for robust feature matching of remote sensing images. We start by constructing the putative point correspondences according to the similarity of well-designed feature descriptors and then focus on removing false matches from the putative set. The key idea of LAP is to search motion-consistent neighborhoods and maintain the local neighborhood topological structures of the true putative matches. To this end, we present a local geometric constraint, which exploits the property of affine invariance to measure the preservation degree of neighborhood topology, since the property is still held under both rigid and complex nonrigid transformations for a minimum topological unit. Moreover, in order to avoid the random distribution of outliers to destroy the neighborhood structure preservation of inliers, a neighbor mining strategy is introduced to search motion-consistent neighbors for each correspondence. We formulate the problem into an optimization model and derive a closed-form solution with linearithmic time complexity. Extensive experimental results on remote sensing images demonstrate that our LAP is able to achieve better performance over the current state-of-the-art approaches. Xinyu Ye, Jiayi Ma 0001, Huilin Xiong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Learning Spatial-Parallax Prior Based on Array Thermal Camera for Infrared Image EnhancementabstractIn this article, an array thermal camera equipment is developed to capture multiple infrared images with spatial and parallax information. Based on the captured images, an end-to-end method called spatial–parallax prior network (SPPN) is proposed. Specifically, we design a spatial–parallax prior block with two symmetric branches to extract spatial and parallax features in an interactive guidance manner. Then, to effectively integrate spatial and parallax features, we introduce a channel attention mechanism to enable the network to focus on and fuse the most useful information adaptively. In this way, spatial and parallax information can be fully utilized without any explicit alignment operation. Finally, considering the scarcity and poor quality of infrared training data, we leverage transfer learning to better train the network. Extensive experimental results demonstrate that the proposed SPPN consistently outperforms the current state-of-the-art methods, providing a highly effective and scalable solution for the improvement of infrared image quality. Jiayi Ma 0001, Wenjing Gao, Yong Ma 0001, Jun Huang 0008, Fan Fan 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Fast and Robust Loop-Closure Detection via Convolutional Auto-Encoder and Motion ConsensusabstractLoop-closure detection is an indispensable module in the visual simultaneous localization and mapping (vSLAM) system. It typically consists of three main steps: image representation, loop-closure candidate selection, and loop-closure event verification. This article proposes a novel approach for loop-closure detection. In particular, we first introduce a lightweight convolutional auto-encoder network trained by the deep perceptual similarity loss for image representation. We then propose an image-to-sequence selection approach based on place sequence division and distance-weighted voting for loop-closure candidate selection. Furthermore, we propose a motion vector consensus constraint to improve locality preserving matching, which can be used for efficient loop-closure event verification that is robust for various complex environments. Extensive experiments have been conducted on four publicly available datasets. The results demonstrate that our method is able to achieve better recall performance than the state-of-the-art and meet the real-time requirement of vSLAM systems. Jiayi Ma 0001, Shenyue Wang, Kaining Zhang, Zheng He 0001, Jun Huang 0008, Xiaoguang Mei |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Multi-Task Interaction Learning for Spatiospectral Image Super-ResolutionabstractHigh spatial resolution and high spectral resolution images (HR-HSIs) are widely applied in geosciences, medical diagnosis, and beyond. However, how to get images with both high spatial resolution and high spectral resolution is still a problem to be solved. In this paper, we present a deep spatial-spectral feature interaction network (SSFIN) for reconstructing an HR-HSI from a low-resolution multispectral image (LR-MSI), e.g., RGB image. In particular, we introduce two auxiliary tasks, i.e., spatial super-resolution (SR) and spectral SR to help the network recover the HR-HSI better. Since higher spatial resolution can provide more detailed information about image texture and structure, and richer spectrum can provide more attribute information, we propose a spatial-spectral feature interaction block (SSFIB) to make the spatial SR task and the spectral SR task benefit each other. Therefore, we can make full use of the rich spatial and spectral information extracted from the spatial SR task and spectral SR task, respectively. Moreover, we use a weight decay strategy (for the spatial and spectral SR tasks) to train the SSFIN, so that the model can gradually shift attention from the auxiliary tasks to the primary task. Both quantitative and visual results on three widely used HSI datasets demonstrate that the proposed method achieves a considerable gain compared to other state-of-the-art methods. Source code is available at https://github.com/junjun-jiang/SSFIN. Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Locality-Guided Global-Preserving Optimization for Robust Feature MatchingabstractFeature matching is a fundamental problem in many computer vision tasks. This paper proposes a novel effective framework for mismatch removal, named LOcality-guided Global-preserving Optimization (LOGO). To identify inliers from a putative matching set generated by feature descriptor similarity, we introduce a fixed-point progressive approach to optimize a graph-based objective, which represents a two-class assignment problem regarding an affinity matrix containing global structures. We introduce a strategy that a small initial set with a high inlier ratio exploits the topology of the affinity matrix to elicit other inliers based on their reliable geometry, which enhances the robustness to outliers. Geometrically, we provide a locality-guided matching strategy, i.e., using local topology consensus as a criterion to determine the initial set, thus expanding to yield the final feature matching set. In addition, we apply local affine transformations based on reference points to determine the local consensus and similarity scores of nodes and edges, ensuring the validity and generality for various scenarios including complex nonrigid transformations. Extensive experiments demonstrate the effectiveness and robustness of the proposed LOGO, which is competitive with the current state-of-the-art methods. It also exhibits favorable potential for high-level vision tasks, such as essential and fundamental matrix estimation, image registration and loop closure detection. Jiayi Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | MSA-Net: Establishing Reliable Correspondences by Multiscale Attention NetworkabstractIn this paper, we propose a novel multi-scale attention based network (called MSA-Net) for feature matching problems. Current deep networks based feature matching methods suffer from limited effectiveness and robustness when applied to different scenarios, due to random distributions of outliers and insufficient information learning. To address this issue, we propose a multi-scale attention block to enhance the robustness to outliers, for improving the representational ability of the feature map. In addition, we also design a novel context channel refine block and a context spatial refine block to mine the information context with less parameters along channel and spatial dimensions, respectively. The proposed MSA-Net is able to effectively infer the probability of correspondences being inliers with less parameters. Extensive experiments on outlier removal and relative pose estimation have shown the performance improvements of our network over current state-of-the-art methods with less parameters on both outdoor and indoor datasets. Notably, our proposed network achieves an 11.7% improvement at error threshold 5° without RANSAC than the state-of-the-art method on relative pose estimation task when trained on YFCC100M dataset. Linxin Zheng, Guobao Xiao, Ziwei Shi, Shiping Wang, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Loop-Closure Detection Using Local Relative Orientation MatchingabstractLoop-closure detection (LCD), which aims to recognize a previously visited location, is a crucial component of the simultaneous localization and mapping system. In this paper, a novel appearance-based LCD method is presented. In particular, we propose a simple yet surprisingly useful feature matching algorithm for real-time geometrical verification of candidate loop-closures, termed aslocal relative orientationmatching (LRO). It aims to efficiently establish reliable feature correspondences based on preserving local topological structures between the query image and candidate frame. To effectively retrieve candidate loop closures, we introduce the aggregated selective match kernel framework into the LCD task, which can effectively represent images and reduce the quantization noise of the traditional bag-of-words framework. In addition, the SuperPoint neural network is employed to extract reliable interest points and feature descriptors. Extensive experimental results demonstrate that our LRO can significantly improve the LCD performance, and the proposed overall LCD method can achieve much better performance over the current state-of-the-art on six publicly available datasets. Jiayi Ma 0001, Xinyu Ye, Huabing Zhou, Xiaoguang Mei, Fan Fan 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Appearance-Based Loop Closure Detection via Locality-Driven Accurate Motion Field LearningabstractLoop closure detection (LCD) is of significant importance in simultaneous localization and mapping. It represents the robot’s ability to recognize whether the current surrounding corresponds to a previously observed one. In this paper, we conduct this task in a two-step strategy: candidate frame selection and loop closure verification. The first step aims to search semantically similar images for the query one using features obtained by Key.Net with HardNet. Instead of adopting the traditional Bag-of-Words strategy, we utilize the aggregated selective match kernel to calculate the similarity between images. Subsequently, based on the potential property of motion field in the LCD scene, we propose a novel feature matching method,i.e., exploiting the smoothness prior and learning the motion field for an image pair in a reproducing kernel Hilbert space (RKHS), to implement loop closure verification. Concretely, we formulate the learning problem into a Bayesian framework with latent variables indicating the true/false correspondences and a mixture model accounting for the distribution of data. Furthermore, we propose a locality-driven mechanism to enhance the local relevance of motion vectors and term the algorithm as locality-driven accurate motion field learning (LAL). To satisfy the requirement of efficiency in the LCD task, we use a sparse approximation and search a suboptimal solution for the motion field in the RKHS, termed as LAL*. Extensive experiments are conducted on public datasets for feature matching and LCD tasks. The quantitative results demonstrate the effectiveness of our method over the current state-of-the-art, meanwhile showing its potential for long-term visual localization. The codes of LAL and LAL* are publicly available athttps://github.com/KN-Zhang/LAL. Kaining Zhang, Xingyu Jiang 0005, Jiayi Ma 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Multi-Focus Image Fusion Based on Multi-Scale Gradients and Image MattingabstractMulti-focus image fusion technology is to extract different focused regions of the same scene among partially focused images and merge them together to generate a composite image where all objects are clear. Two crucial points to multi-focus image fusion are the effective focus measurement method to evaluate the sharpness of the source images and the accurate segmentation method to extract the focused regions. In conventional multi-focus image fusion methods, the decision map obtained according to the focus measurement is sensitive to mis-registration, or produces an uneven boundary lines. In this paper, the maximum value in the top-hat transform and the bottom-hat transform is used as the gradient measurement value, and the complementary features between multiple scales are used to achieve accurate focus measurement for initial segmentation. In order to obtain a better fusion decision map, a robust image matting algorithm is used to refine the trimap generated by the initial segmentation. Then, make full use of the strong correlation between the source images to optimize the edge regions of the decision map to improve the image fusion quality. Finally, a fusion image is constructed based on the fusion decision map and the source images. We perform qualitative and quantitative experiments on publicly available databases to verify the effectiveness of the method. The results show that compared with several state-of-the-art algorithms, the proposed fusion method can obtain accurate decision maps and achieve better performance in visual perception and quantitative analysis. Jun Chen 0019, Linbo Luo 0002, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | DBDnet: A Deep Boosting Strategy for Image DenoisingabstractIn this paper, we propose a new deep network architecture named deep boosting denoising net (DBDnet) for image denoising. It is a residual learning network that can generate a noise map from a noisy observation. In detail, it first generates a coarse noise map via a simple structure, and then updates the noise map gradually via a boosting function. The motivation of our DBDnet stems from the observation that the noise map recovered by any algorithm cannot ideally equal the ground-truth noise map, which typically contains noise. We call this noise NoN,i.e., noise of noise map. Based on this observation, we formulate the denoising as a process of reducing NoN, and the role of DBDnet is to eliminate the NoN from the coarse noise map. In particular, we analyze the process of reducing NoN theoretically, and propose an NoN eliminating module to simulate it accordingly. We evaluate the proposed DBDnet on images polluted by different levels of additive white Gaussian noise and real noise. Experiment results demonstrate that our DBDnet can attain better denoising performance compared with state-of-the-art methods on several kinds of image denoising tasks. In particular, for the Gaussian denoising and real image denoising tasks, the average improvements of the PSNR values brought by our DBDnet are about 0.25 dB and 1.01 dB, respectively. In addition, we find and verify that the deep boosting insight can be easily introduced into the state-of-the-art image denoising network, and promotes its denoising performance. Our code is publicly available athttps://github.com/jiayi-ma/DBDNet. Jiayi Ma 0001, Chengli Peng, Xin Tian 0006, Junjun Jiang |
IEEE Trans. Multim. | 1 |
| 2022 | Real-Time and Accurate UAV Pedestrian Detection for Social Distancing Monitoring in COVID-19 PandemicabstractCoronavirus Disease 2019 (COVID-19) is a highly infectious virus that has created a health crisis for people all over the world. Social distancing has proved to be an effective non-pharmaceutical measure to slow down the spread of COVID-19. As unmanned aerial vehicle (UAV) is a flexible mobile platform, it is a promising option to use UAV for social distance monitoring. Therefore, we propose a lightweight pedestrian detection network to accurately detect pedestrians by human head detection in real-time and then calculate the social distancing between pedestrians on UAV images. In particular, our network follows the PeleeNet as backbone and further incorporates the multi-scale features and spatial attention to enhance the features of small objects, like human heads. The experimental results on Merge-Head dataset show that our method achieves 92.22% AP (average precision) and 76 FPS (frames per second), outperforming YOLOv3 models and SSD models and enabling real-time detection in actual applications. The ablation experiments also indicate that multi-scale feature and spatial attention significantly contribute the performance of pedestrian detection. The test results on UAV-Head dataset show that our method can also achieve high precision pedestrian detection on UAV images with 88.5% AP and 75 FPS. In addition, we have conducted a precision calibration test to obtain the transformation matrix from images (vertical images and tilted images) to real-world coordinate. Based on the accurate pedestrian detection and the transformation matrix, the social distancing monitoring between individuals is reliably achieved. Gui Cheng, Jiayi Ma 0001, Zhongyuan Wang 0001, Jiaming Wang 0001, DeRen Li |
IEEE Trans. Multim. | 3 |
| 2022 | Multilayer Spectral-Spatial Graphs for Label Noisy Robust Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) analysis, label information is a scarce resource and it is unavoidably affected by human and nonhuman factors, resulting in a large amount of label noise. Although most of the recent supervised HSI classification methods have achieved good classification results, their performance drastically decreases when the training samples contain label noise. To address this issue, we propose a label noise cleansing method based on spectral-spatial graphs (SSGs). In particular, an affinity graph is constructed based on spectral and spatial similarity, in which pixels in a superpixel segmentation-based homogeneous region are connected, and their similarities are measured by spectral feature vectors. Then, we use the constructed affinity graph to regularize the process of label noise cleansing. In this manner, we transform label noise cleansing to an optimization problem with a graph constraint. To fully utilize spatial information, we further develop multiscale segmentation-based multilayer SSGs (MSSGs). It can efficiently merge the complementary information of multilayer graphs and thus provides richer spatial information compared with any single-layer graph obtained from isolation segmentation. Experimental results show that MSSG reduces the level of label noise. Compared with the state of the art, the proposed MSSG method exhibits significantly enhanced classification accuracy toward the training data with noisy labels. The significant advantages of the proposed method over four major classifiers are also demonstrated. The source code is available at https://github.com/junjun-jiang/MSSG. Junjun Jiang, Jiayi Ma 0001, Xianming Liu 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | To Choose or to Fuse? Scale Selection for Crowd CountingabstractIn this paper, we address the large scale variation problem in crowd counting by taking full advantage of the multi-scale feature representations in a multi-level network. We implement such an idea by keeping the counting error of a patch as small as possible with a proper feature level selection strategy, since a specific feature level tends to perform better for a certain range of scales. However, without scale annotations, it is sub-optimal and error-prone to manually assign the predictions for heads of different scales to specific feature levels. Therefore, we propose a Scale-Adaptive Selection Network (SASNet), which automatically learns the internal correspondence between the scales and the feature levels. Instead of directly using the predictions from the most appropriate feature level as the final estimation, our SASNet also considers the predictions from other feature levels via weighted average, which helps to mitigate the gap between discrete feature levels and continuous scale variation. Since the heads in a local patch share roughly a same scale, we conduct the adaptive selection strategy in a patch-wise style. However, pixels within a patch contribute different counting errors due to the various difficulty degrees of learning. Thus, we further propose a Pyramid Region Awareness Loss (PRA Loss) to recursively select the most hard sub-regions within a patch until reaching the pixel level. With awareness of whether the parent patch is over-estimated or under-estimated, the fine-grained optimization with the PRA Loss for these region-aware hard pixels helps to alleviate the inconsistency problem between training target and evaluation metric. The state-of-the-art results on four datasets demonstrate the superiority of our approach. The code will be available at: https://github.com/TencentYoutuResearch/CrowdCounting-SASNet. Qingyu Song 0001, Changan Wang, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Jian Wu 0001, Jiayi Ma 0001 |
AAAI | 8 |
| 2021 | Robust Graph Autoencoder for Hyperspectral Anomaly DetectionabstractAutoencoder can not only extract features in an unsupervised manner, but also selects samples out that differs significantly from others. However, autoencoder is sensitive to noise and anomalies during training, and the relationships between pixels are discarded. In order to tackle these problems, we propose a robust graph autoencoder (RGAE) for hyperspectral anomaly detection. To be specific, we first redesign the objective function to encourage the network more robust to noise and anomalies. Meanwhile, a superpixel segmentation-based graph regularization term (SuperGraph) is incorporated into AE to preserve the geometric structure and spatial information simultaneously. Experiments with three real data sets are conducted to evaluate the performance, and the detection results demonstrate that our method outperforms other state-of-the-art hyperspectral anomaly detectors. Ganghui Fan, Yong Ma 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
ICASSP | 5 |
| 2021 | UTDN: An Unsupervised Two-Stream Dirichlet-Net for Hyperspectral UnmixingabstractRecently, the learning-based method has received much attention in the unsupervised hyperspectral unmixing, yet their ability to extract physically meaningful endmembers remains limited and the performance has not been satisfactory. In this paper, we propose a novel two-stream Dirichlet-net, termed as uTDN, to address the above problems. The weight-sharing architecture makes it possible to transfer the intrinsic properties of the endmembers during the process of unmixing, which can help to correct the network converging towards a more accurate and interpretable unmixing solution. Besides, the stick-breaking process is adopted to encourage the latent representation to follow a Dirichlet distribution, where the physical property of the estimated abundance can be naturally incorporated. Extensive experiments on both synthetic and real hyperspectral data demonstrate that the proposed uTDN can outperform the other state-of-the-art approaches. Qiwen Jin, Yong Ma 0001, Xiaoguang Mei, Hao Li 0034, Jiayi Ma 0001 |
ICASSP | 5 |
| 2021 | Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image ClassificationabstractSince CapsNet [1] shattered all previous records of algorithms for image recognition, the capsule's conception has attracted bright attention. It interprets an object by the geometrical arrangement of parts. We think it can be transferred to hyperspectral images. In a hyperspectral data cube, each pixel spectrum can be regarded as a continuous curve representing its inherent properties. In the spatial domain, there are various spatial distributions in different positionsand there is usually a specific structural relationship between adjacently distributed categories. Based on HSI data's aforementioned structural characteristics, combined with the stacked capsule autoencoder, we propose our model to achieve an unsupervised HSI classification. In our model, the ConvLSTM is employed to discover part capsules of HSI, and we utilize Set Transformer to encode relations among all parts and indicate object capsules. The decoders of both phases use Gaussian mixture models to reconstruct specific information. Experimental results of the Pavia Center dataset show the exceptional of our model. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Fan Fan 0001, Jiayi Ma 0001 |
ICASSP | 5 |
| 2021 | Omniscient Video Super-ResolutionabstractMost recent video super-resolution (SR) methods either adopt an iterative manner to deal with low-resolution (LR) frames from a temporally sliding window, or leverage the previously estimated SR output to help reconstruct the current frame recurrently. A few studies try to combine these two structures to form a hybrid framework but have failed to give full play to it. In this paper, we propose an omniscient framework to not only utilize the preceding SR output, but also leverage the SR outputs from the present and future. The omniscient framework is more generic because the iterative, recurrent and hybrid frameworks can be regarded as its special cases. The proposed omniscient framework enables a generator to behave better than its counterparts under other frameworks. Abundant experiments on public datasets show that our method is superior to the state-of-the-art methods in objective metrics, subjective visual effects and complexity. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICCV | 7 |
| 2021 | Uniformity in Heterogeneity: Diving Deep into Count Interval Partition for Crowd CountingabstractRecently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of counts instead of the count values themselves. However, an inappropriate interval setting might make the count error contributions from different intervals extremely imbalanced, leading to inferior counting performance. Therefore, we propose a novel count interval partition criterion called Uniform Error Partition (UEP), which always keeps the expected counting error contributions equal for all intervals to minimize the prediction risk. Then to mitigate the inevitably introduced discretization errors in the count quantization process, we propose another criterion called Mean Count Proxies (MCP). The MCP criterion selects the best count proxy for each interval to represent its count value during inference, making the overall expected discretization error of an image nearly negligible. As far as we are aware, this work is the first to delve into such a classification task and ends up with a promising solution for count interval partition. Following the above two theoretically demonstrated criterions, we propose a simple yet effective model termed Uniform Error Partition Network (UEPNet), which achieves state-of-the-art performance on several challenging datasets. The codes will be available at: TencentYoutuResearch/CrowdCounting-UEPNet. Changan Wang, Qingyu Song 0001, Boshen Zhang, Yabiao Wang, Ying Tai, Xuyi Hu, Chengjie Wang 0001, Jiayi Ma 0001, Yang Wu 0001 |
ICCV | 9 |
| 2021 | T-Net: Effective Permutation-Equivariant Network for Two-View Correspondence LearningabstractWe develop a conceptually simple, flexible, and effective framework (named T-Net) for two-view correspondence learning. Given a set of putative correspondences, we reject outliers and regress the relative pose encoded by the essential matrix, by an end-to-end framework, which is consisted of two novel structures: "−" structure and "|" structure. " − " structure adopts an iterative strategy to learn correspondence features. "|" structure integrates all the features of the iterations and outputs the correspondence weight. In addition, we introduce Permutation-Equivariant Context Squeeze-and-Excitation module, an adapted version of SE module, to process sparse correspondences in a permutation-equivariant way and capture both global and channel-wise contextual information. Extensive experiments on outdoor and indoor scenes show that the proposed T-Net achieves state-of-the-art performance. On outdoor scenes (YFCC100M dataset), T-Net achieves an mAP of 52.28%, a 34.22% precision increase from the best-published result (38.95%). On indoor scenes (SUN3D dataset), T-Net (19.71%) obtains a 21.82% precision increase from the best-published result (16.18%). Source code: https://github.com/x-gb/T-Net. Guobao Xiao, Linxin Zheng, Jiayi Ma 0001 |
ICCV | 5 |
| 2021 | Pan-Sharpening Via High-Pass Modification Convolutional Neural NetworkabstractMost existing deep learning-based pan-sharpening methods have several widely recognized issues, such as spectral distortion and insufficient spatial texture enhancement, we propose a novel pan-sharpening convolutional neural network based on a high-pass modification b lock. Different from existing methods, the proposed block is designed to learn the high-pass information, leading to enhance spatial information in each band of the multi-spectral-resolution images. To facilitate the generation of visually appealing pan-sharpened images, we propose a perceptual loss function and further optimize the model based on high-level features in the near-infrared space. Experiments demonstrate the superior performance of the proposed method compared to the state-of the-art pan-sharpening methods, both quantitatively and qualitatively. The proposed model is open-sourced at https://github.com/jiaming-wang/HMB. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
ICIP | 6 |
| 2021 | Zero-Shot Multi-Focus Image FusionabstractMulti-focus image fusion (MFIF) is an effective way to eliminate the out-of-focus blur generated in the imaging process. The difficulties in focus level estimation and the lack of real training set for supervised learning make MFIF remain a challenging task after decades of research. According to DIP [1], a neural network can capture the low-level statistics of a single image and can be used as a prior for solving many low-level problems. Based on this idea, we propose a novel architecture named IM-Net comprised of I-Net to model the deep prior of the fused image and M-Net to model the deep prior of the focus map. Without any large scale training set, our method achieves zero-shot learning through the extracted prior information. Experiments on extensively used dataset demonstrate the effectiveness of our approach. Junjun Jiang, Xianming Liu 0005, Jiayi Ma 0001 |
ICME | 4 |
| 2021 | Variation-Net: Interpretable Variation-Inspired Deep Network for PansharpeningabstractIn this study, we propose Variation-net, an interpretable variation-inspired deep network for pansharpening, which aims to fuse panchromatic (PAN) and multispectral (MS) images for a high-resolution MS image. We first construct a novel variational pan-sharpening model with clear physical meanings. As the relationship between the PAN and MS images in the real situation is complex and nonlinear, we explore the similarity between PAN and MS images from the sparsity of nonlinear transforms in this variational pansharpening model. As a result, spatial details can be accurately transferred from PAN image to MS image. Furthermore, we build the Variation-net by unrolling the iterative shrinkage-thresholding algorithm to solve the proposed variational pansharpening model. Therefore, all modules in Variation-net have clear physical meanings and are easily observed, leading to good generalization capability. Meanwhile, nonlinear transforms and other parameters in the variational pansharpening model are learned end–to–end. The experiments demonstrate that Variation-net outperforms the state-of-the-art methods from the aspects of visual effect and objective quality analysis. Kun Li 0025, Wei Zhang 0259, Xin Tian 0006, Jiayi Ma 0001, Huabing Zhou, Zhongyuan Wang 0001 |
ICME | 4 |
| 2021 | Visual Place Recognition via Local Affine Preserving MatchingabstractVisual Place Recognition (VPR) is a crucial component for long-term mobile robot autonomy. In this paper, we exploit a coarse-to-fine paradigm to recognize places. In particular, we first select candidate frames for each query image, and then check the spatial geometric relationship between the query and its candidate frames to determine the final place match. In the coarse match stage, we employ the deep learning network to extract global features that encode semantic information of images, then by comparing the similarity between features to obtain a candidate list of the query place. In the fine match stage, we propose an effective and efficient feature matching algorithm for real-time geometrical verification of candidate places, termed as local affine preserving matching (LAP). Extensive experimental results demonstrate that our LAP can significantly promote the VPR performance, and the proposed overall VPR method can achieve much better performance over the current state-of-the-art approaches. Xinyu Ye, Jiayi Ma 0001 |
ICRA | 2 |
| 2021 | Appearance-based Loop Closure Detection via Bidirectional Manifold Representation ConsensusabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select candidates on-line according to the similarity of global semantic features in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. To this end, a robust feature matching algorithm, termed as bidirectional manifold representation consensus (BMRC), is proposed. In particular, we utilize manifold representation to construct local neighborhood structures of feature points and formulate the matching problem into an optimization model, enabling linearithmic time complexity via a closed-form solution. Furthermore, we propose a dynamic place partition strategy based on BMRC to segment image streams with similar content into a place, which can mine more valid candidate frames, improving the recall rate of the whole system. Extensive experiments on several publicly available datasets reveal that BMRC has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Zizhuo Li, Jiayi Ma 0001 |
ICRA | 3 |
| 2021 | Motion Field Consensus with Locality Preservation: A Geometric Confirmation Strategy for Loop Closure DetectionabstractLoop closure detection (LCD), which aims to deal with the drift emerging when robots travel around the route, plays a key role in a simultaneous localization and mapping system. Unlike most current methods which focus on seeking an appropriate representation of images, we propose a novel two-stage pipeline dominated by the estimation of spatial geometric relationship. When a query image occurs, we select semantically similar images based on the SuperPoint network and the aggregated selective match kernel in the first stage, and then conduct robust geometric confirmation to verify true loop-closing pairs in the second stage. Based on the potential property of motion field in the LCD scene, a robust feature matching algorithm, termed as motion field consensus with locality preservation (MFC-LP), is proposed. In particular, we exploit the smoothness prior to guide the learning of the motion field for an image pair in a reproducing kernel Hilbert space (RKHS). Meanwhile, to enhance the local relevance of motion vectors, we design a locality preservation mechanism thus making the learned motion field more accurate. Extensive experiments on several publicly available datasets reveal that MFC-LP has a good performance in the general feature matching task and the proposed pipeline outperforms the current state-of-the-art approaches in the LCD task. Kaining Zhang, Xingyu Jiang 0005, Xiaoguang Mei, Huabing Zhou, Jiayi Ma 0001 |
IROS | 5 |
| 2021 | Image Matching from Handcrafted to Deep Features: A SurveyabstractAbstract As a fundamental and critical task in various visual applications, image matching can identify then correspond the same or similar structure/content from two or more images. Over the past decades, growing amount and diversity of methods have been proposed for image matching, particularly with the development of deep learning techniques over the recent years. However, it may leave several open questions about which method would be a suitable choice for specific applications with respect to different scenarios and task requirements and how to design better image matching methods with superior performance in accuracy, robustness and efficiency. This encourages us to conduct a comprehensive and systematic review and analysis for those classical and latest techniques. Following the feature-based image matching pipeline, we first introduce feature detection, description, and matching techniques from handcrafted methods to trainable ones and provide an analysis of the development of these methods in theory and practice. Secondly, we briefly introduce several typical image matching-based applications for a comprehensive understanding of the significance of image matching. In addition, we also provide a comprehensive and objective comparison of these classical and latest techniques through extensive experiments on representative datasets. Finally, we conclude with the current status of image matching technologies and deliver insightful discussions and prospects for future works. This survey can serve as a reference for (but not limited to) researchers and engineers in image matching and related fields. Jiayi Ma 0001, Xingyu Jiang 0005, Aoxiang Fan, Junjun Jiang, Junchi Yan |
Int. J. Comput. Vis. | 1 |
| 2021 | Segmentation by Continuous Latent Semantic Analysis for Multi-structure Model Fitting
Guobao Xiao, Hanzi Wang, Jiayi Ma 0001, David Suter |
Int. J. Comput. Vis. | 3 |
| 2021 | Beyond Brightening Low-light Images
Xiaojie Guo 0001, Jiayi Ma 0001, Wei Liu 0005, Jiawan Zhang |
Int. J. Comput. Vis. | 3 |
| 2021 | SDNet: A Versatile Squeeze-and-Decomposition Network for Real-Time Image Fusion
Hao Zhang 0073, Jiayi Ma 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Locality-constrained sparse representation for hyperspectral image classification
Yuanshu Zhang, Yong Ma 0001, Xiaobing Dai, Hao Li 0034, Xiaoguang Mei, Jiayi Ma 0001 |
Inf. Sci. | 6 |
| 2021 | MPIN: a macro-pixel integration network for light field super-resolutionabstractMost existing light field (LF) super-resolution (SR) methods either fail to fully use angular information or have an unbalanced performance distribution because they use parts of views. To address these issues, we propose a novel integration network based on macro-pixel representation for the LF SR task, named MPIN. Restoring the entire LF image simultaneously, we couple the spatial and angular information by rearranging the four-dimensional LF image into a two-dimensional macro-pixel image. Then, two special convolutions are deployed to extract spatial and angular information, separately. To fully exploit spatial-angular correlations, the integration resblock is designed to merge the two kinds of information for mutual guidance, allowing our method to be angular-coherent. Under the macro-pixel representation, an angular shuffle layer is tailored to improve the spatial resolution of the macro-pixel image, which can effectively avoid aliasing. Extensive experiments on both synthetic and real-world LF datasets demonstrate that our method can achieve better performance than the state-of-the-art methods qualitatively and quantitatively. Moreover, the proposed method has an advantage in preserving the inherent epipolar structures of LF images with a balanced distribution of performance. Xinya Wang, Jiayi Ma 0001, Wenjing Gao, Junjun Jiang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Hyperspectral Anomaly Detection via Integration of Feature Extraction and Background PurificationabstractAnomaly detection (AD) has become a hotspot in hyperspectral imagery (HSI) processing due to its advantage in detecting potential targets without prior knowledge, and a variety of algorithms are proposed for a better performance. However, they usually either fail to extract intrinsic features underlying HSIs, or suffer from the contamination of noise and anomalies. To address these problems, we propose a new anomaly detector by integrating fractional Fourier transform (FrFT) with low rank and sparse matrix decomposition (LRaSMD). First, distinctive features of HSI data are extracted via FrFT. Then, row-constrained LRaSMD (RC-LRaSMD), which is more practical and stable than the traditional LRaSMD, is employed to separate background from noise and anomalies. Finally, we implement an atom-selection strategy to construct the background covariance matrix for detection. The experimental results with several HSI data sets demonstrate satisfying detection performance compared with other state-of-the-art detectors. Yong Ma 0001, Ganghui Fan, Qiwen Jin, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | Bilateral attention decoder: A lightweight decoder for real-time semantic segmentation
Chengli Peng, Tian Tian 0006, Chen Chen 0001, Xiaojie Guo 0001, Jiayi Ma 0001 |
Neural Networks | 5 |
| 2021 | Enhanced image prior for unsupervised remoting sensing super-resolution
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
Neural Networks | 6 |
| 2021 | Ranking list preservation for feature matching
Junjun Jiang, Xingyu Jiang 0005, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2021 | Mining consistent correspondences using co-occurrence statistics
Guobao Xiao, Shiping Wang, Han Wang 0001, Jiayi Ma 0001 |
Pattern Recognit. | 4 |
| 2021 | Robust Feature Matching for Remote Sensing Image Registration via Linear Adaptive FilteringabstractAs a fundamental and critical task in feature-based remote sensing image registration, feature matching refers to establishing reliable point correspondences from two images of the same scene. In this article, we propose a simple yet efficient method termed linear adaptive filtering (LAF) for both rigid and nonrigid feature matching of remote sensing images and apply it to the image registration task. Our algorithm starts with establishing putative feature correspondences based on local descriptors and then focuses on removing outliers using geometrical consistency priori together with filtering and denoising theory. Specifically, we first grid the correspondence space into several nonoverlapping cells and calculate a typical motion vector for each one. Subsequently, we remove false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a Gaussian kernel convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. In addition, an adaptive parameter setting strategy and posterior probability estimation based on the expectation-maximization algorithm enhance the robustness of our method to different data. Most importantly, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Consequently, some sparse point-based tasks may inspire from our method when they are achieved by deep learning techniques. Extensive feature matching and image registration experiments on several remote sensing data sets demonstrate the superiority of our approach over the state of the art. Xingyu Jiang 0005, Jiayi Ma 0001, Aoxiang Fan, Haiping Xu, Geng Lin, Tao Lu 0001, Xin Tian 0006 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | FusionNDVI: A Computational Fusion Approach for High-Resolution Normalized Difference Vegetation IndexabstractNormalized difference vegetation index (NDVI), derived from the near-infrared and red bands of a multispectral (MS) image, has been widely used in remote sensing. To obtain a high-resolution (HR) NDVI, existing attempts typically first generate an HR-MS image using pansharpening and then calculate the HR NDVI accordingly. However, some inaccurate spatial information will be simultaneously introduced into NDVIs, influencing their spatial quality seriously. To overcome this challenge, we investigate a computational fusion approach from a novel perspective for HR NDVI in this study. Rather than pansharpening an HR-MS image, we define an HR vegetation index calculated based on an available HR panchromatic image and an estimated HR red band (VIPR) and fuse the low-resolution (LR) NDVI and HR VIPR directly to acquire an HR NDVI. In particular, we adopt a nonlocal gradient sparsity constraint to force a similar nonlocal spatial structure in the fused NDVI and VIPR, where the VIPR is dynamically updated by adding a constraint to reconstruct the HR red band. We further integrate a data fidelity term to constrain the relationship between the fused NDVI and its LR version, and an efficient strategy based on the alternative direction multiplier method is developed to solve the nonconvex optimization problem. The extensive experimental results demonstrate that the proposed method achieves superior fusion performance over the state of the art, exhibiting its wide application aspect in remote sensing. Xin Tian 0006, Mengliang Zhang, Changcai Yang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | SDPNet: A Deep Network for Pan-Sharpening With Enhanced Information RepresentationabstractIn this article, we propose a surface- and deep-level constraint-based pan-sharpening network, termed SDPNet, to address the pan-sharpening problem. Focusing on the two primary goals of pan-sharpening, i.e., spatial and spectral information preservations, we first design two encoder-decoder networks to extract deep-level features from two types of source images, in addition to surface-level characteristics, as the enhanced information representation. The unique feature maps that characterize the unique information in source images can be obtained through the deep-level feature extraction. We further design a pan-sharpening network with densely connected blocks to strengthen feature propagation and reduce parameter number, where the unique feature maps are utilized to efficiently constrain the similarity between the pan-sharpened result and the ground truth, thus avoiding information distortion. Both qualitative and quantitative comparisons on the reduced-resolution and full-resolution source images demonstrate the advantages of our method over state-of-the-art methods. Our code is publicly available at https://github.com/hanna-xu/SDPNet. Han Xu 0001, Jiayi Ma 0001, Hao Zhang 0073, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Robust ENF Estimation Based on Harmonic Enhancement and Maximum Weight CliqueabstractThe electric network frequency (ENF) is an important and extensively researched forensic criterion to authenticate digital recordings, but currently it is still challenging to extract reliable ENF traces from recordings in uncontrollable environments. In this paper, we present a framework for robust ENF extraction from real-world audio recordings, featuring multi-tone harmonic ENF enhancement and graph-based harmonic selection. We first extend the recently developed single-tone robust filtering algorithm (RFA) to the multi-tone scenario and propose a harmonic robust filtering algorithm (HRFA). It can enhance each harmonic component without cross-component interference, thus alleviating the effects of unwanted noise and audio content. In addition, considering the fact that some harmonic components could still be severely corrupted after the HRFA, interfering rather than facilitating ENF estimation, we propose a graph-based harmonic selection algorithm (GHSA), which finds a subset of harmonic components having the overall highest mutual cross-correlation. Noticeably, the harmonic selection problem is found to be equivalent to the maximum weight clique problem in graph theory, and the Bron-Kerbosch algorithm is adopted in the GHSA. With the enhanced and carefully selected harmonic components, both the existing maximum likelihood estimator (MLE) and weighted MLE are incorporated to yield the final ENF estimation results. The proposed framework is evaluated using both synthetic signals and the ENF-WHU dataset consisting of 130 real-world audio recordings, demonstrating its advantages over both the existing single- and multi-tone competitors. This work further improves the applicability of the ENF as a forensic criterion in real-world situations. Guang Hua 0001, Han Liao, Dengpan Ye, Jiayi Ma 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Correlation Discrepancy Insight Network for Video Re-identificationabstractVideo-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods. Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2020 | FusionDN: A Unified Densely Connected Network for Image FusionabstractIn this paper, we present a new unsupervised and unified densely connected network for different types of image fusion tasks, termed as FusionDN. In our method, the densely connected network is trained to generate the fused image conditioned on source images. Meanwhile, a weight block is applied to obtain two data-driven weights as the retention degrees of features in different source images, which are the measurement of the quality and the amount of information in them. Losses of similarities based on these weights are applied for unsupervised learning. In addition, we obtain a single model applicable to multiple fusion tasks by applying elastic weight consolidation to avoid forgetting what has been learned from previous tasks when training multiple tasks sequentially, rather than train individual models for every fusion task or jointly train tasks roughly. Qualitative and quantitative results demonstrate the advantages of FusionDN compared with state-of-the-art methods in different fusion tasks. Han Xu 0001, Jiayi Ma 0001, Zhuliang Le, Junjun Jiang, Xiaojie Guo 0001 |
AAAI | 2 |
| 2020 | Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and IntensityabstractIn this paper, we propose a fast unified image fusion network based on proportional maintenance of gradient and intensity (PMGI), which can end-to-end realize a variety of image fusion tasks, including infrared and visible image fusion, multi-exposure image fusion, medical image fusion, multi-focus image fusion and pan-sharpening. We unify the image fusion problem into the texture and intensity proportional maintenance problem of the source images. On the one hand, the network is divided into gradient path and intensity path for information extraction. We perform feature reuse in the same path to avoid loss of information due to convolution. At the same time, we introduce the pathwise transfer block to exchange information between different paths, which can not only pre-fuse the gradient information and intensity information, but also enhance the information to be processed later. On the other hand, we define a uniform form of loss function based on these two kinds of information, which can adapt to different fusion tasks. Experiments on publicly available datasets demonstrate the superiority of our PMGI over the state-of-the-art in terms of both visual effect and quantitative metric in a variety of fusion tasks. In addition, our method is faster compared with the state-of-the-art. Hao Zhang 0073, Han Xu 0001, Xiaojie Guo 0001, Jiayi Ma 0001 |
AAAI | 5 |
| 2020 | MTL-NAS: Task-Agnostic Neural Architecture Search Towards General-Purpose Multi-Task LearningabstractWe propose to incorporate neural architecture search (NAS) into general-purpose multi-task learning (GP-MTL). Existing NAS methods typically define different search spaces according to different tasks. In order to adapt to different task combinations (i.e., task sets), we disentangle the GP-MTL networks into single-task backbones (optionally encode the task priors), and a hierarchical and layerwise features sharing/fusing scheme across them. This enables us to design a novel and general task-agnostic search space, which inserts cross-task edges (i.e., feature fusion connections) into fixed single-task network backbones. Moreover, we also propose a novel single-shot gradient-based search algorithm that closes the performance gap between the searched architectures and the final evaluation architecture. This is realized with a minimum entropy regularization on the architecture weights during the search phase, which makes the architecture weights converge to near-discrete values and therefore achieves a single model. As a result, our searched model can be directly used for evaluation without (re-)training from scratch. We perform extensive experiments using different single-task backbones on various task sets, demonstrating the promising performance obtained by exploiting the hierarchical and layerwise features, as well as the desirable generalizability to different i) task sets and ii) single-task backbones. The code of our paper is available at https://github.com/bhpfelix/MTLNAS. Yuan Gao 0015, Haoping Bai, Zequn Jie, Jiayi Ma 0001, Kui Jia, Wei Liu 0005 |
CVPR | 4 |
| 2020 | Multi-Scale Progressive Fusion Network for Single Image DerainingabstractRain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary information for rain streak representation. In this work, we explore the multi-scale collaborative representation for rain streaks from the perspective of input image scales and hierarchical deep features in a unified framework, termed multi-scale progressive fusion network (MSPFN) for single image rain streak removal. For the similar rain streaks at different positions, we employ recurrent calculation to capture the global texture, thus allowing to explore the complementary and redundant information at the spatial dimension to characterize target rain streaks. Besides, we construct multi-scale pyramid structure, and further introduce the attention mechanism to guide the fine fusion of these correlated information from different scales. This multi-scale progressive fusion strategy not only promotes the cooperative representation, but also boosts the end-to-end training. Our proposed method is extensively evaluated on several benchmark datasets and achieves the state-of-the-art results. Moreover, we conduct experiments on joint deraining, detection, and segmentation tasks, and inspire a new research direction of vision task driven image deraining. The source code is available at https://github.com/kuihua/MSPFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Baojin Huang, Yimin Luo, Jiayi Ma 0001, Junjun Jiang |
CVPR | 7 |
| 2020 | Geometric Estimation via Robust Subspace Recovery
Aoxiang Fan, Xingyu Jiang 0005, Junjun Jiang, Jiayi Ma 0001 |
ECCV (22) | 5 |
| 2020 | A Generative Adversarial Network For Medical Image FusionabstractIn this paper, a novel end-to-end model for fusing medical images characterizing structural information, i.e., IS, and images characterizing functional information, i.e., IF, of different resolutions is proposed, which is achieved by using a conditional generative adversarial network with multiple generators and multiple discriminators (MGMDcGAN). In the first cGAN, a real-like fused image is generated by a generator, simultaneously fooling two discriminators. While the discriminators are to distinguish the fused image from source images. Besides, to prevent the functional information from being weakened in the final fused image when enhancing the dense structure information, we employ the second cGAN with a mask calculated. Meanwhile, the structural information in ISand the functional information in IFthe final fused image can be concurrently kept. Furthermore, our MGMDcGAN is a unified method, which is applicable to different kinds of medical image fusion, including MRI-PET, MRISPECT, and CT-SPECT. Extensive experiments on publicly available datasets substantiate the superiority of our MGMDcGAN over the current state-of-the-art. Zhuliang Le, Jun Huang 0008, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICIP | 5 |
| 2020 | Infrared and visible image fusion via gradientlet filter
Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 1 |
| 2020 | Spectral-spatial classification for hyperspectral image based on a single GRU
Erting Pan, Xiaoguang Mei, Quande Wang, Yong Ma 0001, Jiayi Ma 0001 |
Neurocomputing | 5 |
| 2020 | Learning to find reliable correspondences with local neighborhood consensus
Xiaoguang Mei, Yong Ma 0001, Jun Huang 0008, Fan Fan 0001, Jiayi Ma 0001 |
Neurocomputing | 6 |
| 2020 | Infrared and visible image fusion based on target-enhanced multiscale transform decomposition
Jun Chen 0019, Linbo Luo 0002, Xiaoguang Mei, Jiayi Ma 0001 |
Inf. Sci. | 5 |
| 2020 | Sparse unmixing of hyperspectral data with bandwise model
Chang Li 0001, Yu Liu 0023, Juan Cheng 0004, Rencheng Song, Jiayi Ma 0001, Chenhong Sui, Xun Chen 0001 |
Inf. Sci. | 5 |
| 2020 | Mutually Guided Image FilteringabstractFiltering images is required by numerous multimedia, computer vision and graphics tasks. Despite diverse goals of different tasks, making effective rules is key to the filtering performance. Linear translation-invariant filters with manually designed kernels have been widely used. However, their performance suffers from content-blindness. To mitigate the content-blindness, a family of filters, called joint/guided filters, have attracted a great amount of attention from the community. The main drawback of most joint/guided filters comes from the ignorance of structural inconsistency between the reference and target signals like color, infrared, and depth images captured under different conditions. Simply adopting such guidelines very likely leads to unsatisfactory results. To address the above issues, this paper designs a simple yet effective filter, named mutually guided image filter (muGIF), which jointly preserves mutual structures, avoids misleading from inconsistent structures and smooths flat regions. The proposed muGIF is very flexible, which can work in various modes including dynamic only (self-guided), static/dynamic (reference-guided) and dynamic/dynamic (mutually guided) modes. Although the objective of muGIF is in nature non-convex, by subtly decomposing the objective, we can solve it effectively and efficiently. The advantages of muGIF in effectiveness and flexibility are demonstrated over other state-of-the-art alternatives on a variety of applications. Our code is publicly available at https://sites.google.com/view/xjguo/mugif. Xiaojie Guo 0001, Yu Li 0003, Jiayi Ma 0001, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Minimal Case Relative Pose Computation Using Ray-Point-Ray FeaturesabstractCorners are popular features for relative pose computation with 2D-2D point correspondences. Stable corners may be formed by two 3D rays sharing a common starting point. We call such elements ray-point-ray (RPR) structures. Besides a local invariant keypoint given by the lines' intersection, their reprojection also defines a corner orientation and an inscribed angle in the image plane. The present paper investigates such RPR features, and aims at answering the fundamental question of what additional constraints can be formed from correspondences between RPR features in two views. In particular, we show that knowing the value of the inscribed angle between the two 3D rays poses additional constraints on the relative orientation. Using the latter enables the solution of the relative pose problem with as few as 3 correspondences across the two images. We provide a detailed analysis of all minimal cases distinguishing between 90-degree RPR-structures and structures with an arbitrary, known inscribed angle. We furthermore investigate the special cases of a known directional correspondence and planar motion, the latter being solvable with only a single RPR correspondence. We complete the exposition by outlining an image processing technique for robust RPR-feature extraction. Our results suggest high practicality in man-made environments, where 90-degree RPR-structures naturally occur. Ji Zhao 0001, Laurent Kneip, Yijia He, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Semantic segmentation using stride spatial pyramid pooling and dual attention decoder
Chengli Peng, Jiayi Ma 0001 |
Pattern Recognit. | 2 |
| 2020 | A Variational Pansharpening Method Based on Gradient Sparse RepresentationabstractBy exploiting the gradient similarity between multispectral (MS) and panchromatic (PAN) images, a variational pansharpening method based on gradient sparse representation is proposed, based on the observation that the gradients of corresponding MS and PAN images with different resolutions have the similar sparse coefficients under certain specific dictionaries. By adding a data fidelity term to preserve the spectral information, an optimization model is constructed as a minimization problem of an energy function. The problem can be solved by the gradient descent method efficiently. Experiments on different satellite data reveal that the proposed method outperforms the state-of-the-art methods in terms of visual effect and objective quality analysis. Xin Tian 0006, Yuerong Chen, Changcai Yang, Jiayi Ma 0001 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Multi-Temporal Ultra Dense Memory Network for Video Super-ResolutionabstractVideo super-resolution (SR) aims to reconstruct the corresponding high-resolution (HR) frames from consecutive low-resolution (LR) frames. It is crucial for video SR to harness both inter-frame temporal correlations and intra-frame spatial correlations among frames. Previous video SR methods based on convolutional neural network (CNN) mostly adopt a single-channel structure and a single memory module, so they are unable to fully exploit inter-frame temporal correlations specific for video. To this end, this paper proposes a multi-temporal ultra-dense memory (MTUDM) network for video super-resolution. Particularly, we embed convolutional long-short-term memory (ConvLSTM) into ultra-dense residual block (UDRB) to construct an ultra-dense memory block (UDMB) for extracting and retaining spatio-temporal correlations. This design also reduces the layer depth by expanding the width, thus avoiding training difficulties, such as gradient exploding and vanishing under a large model. We further adopt multi-temporal information fusion (MTIF) strategy to merge the extracted temporal feature maps in consecutive frames, improving the accuracy without requiring much extra computational cost. The experimental results on extensive public datasets demonstrate that our method outperforms the state-of-the-art methods by a large margin. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Context-Patch Face Hallucination Based on Thresholding Locality-Constrained Representation and Reproducing LearningabstractFace hallucination is a technique that reconstructs high-resolution (HR) faces from low-resolution (LR) faces, by using the prior knowledge learned from HR/LR face pairs. Most state-of-the-arts leverage position-patch prior knowledge of the human face to estimate the optimal representation coefficients for each image patch. However, they focus only the position information and usually ignore the context information of the image patch. In addition, when they are confronted with misalignment or the small sample size (SSS) problem, the hallucination performance is very poor. To this end, this paper incorporates the contextual information of the image patch and proposes a powerful and efficient context-patch-based face hallucination approach, namely, thresholding locality-constrained representation and reproducing learning (TLcR-RL). Under the context-patch-based framework, we advance a thresholding-based representation method to enhance the reconstruction accuracy and reduce the computational complexity. To further improve the performance of the proposed algorithm, we propose a promotion strategy called reproducing learning. By adding the estimated HR face to the training set, which can simulate the case that the HR version of the input LR face is present in the training set, it thus iteratively enhances the final hallucination result. Experiments demonstrate that the proposed TLcR-RL method achieves a substantial increase in the hallucinated results, both subjectively and objectively. In addition, the proposed framework is more robust to face misalignment and the SSS problem, and its hallucinated HR face is still very good when the LR test face is from the real world. The MATLAB source code is available at https://github.com/junjun-jiang/TLcR-RL. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Akiko Aizawa, Kiyoharu Aizawa |
IEEE Trans. Cybern. | 4 |
| 2020 | Ensemble Super-Resolution With a Reference DatasetabstractBy developing sophisticated image priors or designing deep(er) architectures, a variety of image super-resolution (SR) approaches have been proposed recently and achieved very promising performance. A natural question that arises is whether these methods can be reformulated into a unifying framework and whether this framework assists in SR reconstruction? In this paper, we present a simple but effective single image SR method based on ensemble learning, which can produce a better performance than that could be obtained from any of SR methods to be ensembled (or called component super-resolvers). Based on the assumption that better component super-resolver should have larger ensemble weight when performing SR reconstruction, we present a maximum a posteriori (MAP) estimation framework for the inference of optimal ensemble weights. Especially, we introduce a reference dataset, which is composed of high-resolution (HR) and low-resolution (LR) image pairs, to measure the SR abilities (prior knowledge) of different component super-resolvers. To obtain the optimal ensemble weights, we propose to incorporate the reconstruction constraint, which states that the degenerated HR estimation should be equal to the LR observation one, as well as the prior knowledge of ensemble weights into the MAP estimation framework. Moreover, the proposed optimization problem can be solved by an analytical solution. We study the performance of the proposed method by comparing with different competitive approaches, including four state-of-the-art nondeep learning-based methods, four latest deep learning-based methods, and one ensemble learning-based method, and prove its effectiveness and superiority on some general image datasets and face image datasets. Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Suhua Tang, Ruimin Hu, Jiayi Ma 0001 |
IEEE Trans. Cybern. | 6 |
| 2020 | Robust Feature Matching Using Spatial Clustering With Heavy OutliersabstractThis paper focuses on removing mismatches from given putative feature matches created typically based on descriptor similarity. To achieve this goal, existing attempts usually involve estimating the image transformation under a geometrical constraint, where a pre-defined transformation model is demanded. This severely limits the applicability, as the transformation could vary with different data and is complex and hard to model in many real-world tasks. From a novel perspective, this paper casts the feature matching into a spatial clustering problem with outliers. The main idea is to adaptively cluster the putative matches into several motion consistent clusters together with an outlier/mismatch cluster. To implement the spatial clustering, we customize the classic density based spatial clustering method of applications with noise (DBSCAN) in the context of feature matching, which enables our approach to achieve quasi-linear time complexity. We also design an iterative clustering strategy to promote the matching performance in case of severely degraded data. Extensive experiments on several datasets involving different types of image transformations demonstrate the superiority of our approach over state-of-the-art alternatives. Our approach is also applied to near-duplicate image retrieval and co-segmentation and achieves promising performance. Xingyu Jiang 0005, Jiayi Ma 0001, Junjun Jiang, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | DDcGAN: A Dual-Discriminator Conditional Generative Adversarial Network for Multi-Resolution Image FusionabstractIn this paper, we proposed a new end-to-end model, termed as dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Our method establishes an adversarial game between a generator and two discriminators. The generator aims to generate a real-like fused image based on a specifically designed content loss to fool the two discriminators, while the two discriminators aim to distinguish the structure differences between the fused image and two source images, respectively, in addition to the content loss. Consequently, the fused image is forced to simultaneously keep the thermal radiation in the infrared image and the texture details in the visible image. Moreover, to fuse source images of different resolutions, e.g., a low-resolution infrared image and a high-resolution visible image, our DDcGAN constrains the downsampled fused image to have similar property with the infrared image. This can avoid causing thermal radiation information blurring or visible texture detail loss, which typically happens in traditional methods. In addition, we also apply our DDcGAN to fusing multi-modality medical images of different resolutions, e.g., a low-resolution positron emission tomography image and a high-resolution magnetic resonance image. The qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our DDcGAN over the state-of-the-art, in terms of both visual effect and quantitative metrics. Jiayi Ma 0001, Han Xu 0001, Junjun Jiang, Xiaoguang Mei, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | Deterministic Model Fitting by Local-Neighbor Preservation and Global-Residual OptimizationabstractGeometric model fitting has been widely used in many computer vision tasks. However, it remains as a challenging task when handing multiple-structural data contaminated by noises and outliers. Most previous work on model fitting cannot guarantee the consistency of their solutions due to their randomness, precluding them from many real-world applications. In this research, we propose a fast two-view approximately deterministic model fitting scheme (called LGF), to provide consistent solutions for multiple-structural data. The proposed LGF scheme starts from defining preference function by preserving local neighborhood relationship, and then adopts the min-hash technique to roughly sample subsets. By this way, it is able to cover all model instances in data in the parameter space with a high probability. After that, LGF refines the previous sampled subsets by globalresidual optimization. Furthermore, we propose a simple yet effective model selection framework to estimate the number and the parameters of model instances in data. Extensive experiments on real images show that the proposed LGF scheme is able to observe superior or very competitive performance on both accuracy and speed over several state-of-the-art model fitting methods. Guobao Xiao, Jiayi Ma 0001, Shiping Wang, Chang Wen Chen |
IEEE Trans. Image Process. | 2 |
| 2020 | MEF-GAN: Multi-Exposure Image Fusion via Generative Adversarial NetworksabstractIn this paper, we present an end-to-end architecture for multi-exposure image fusion based on generative adversarial networks, termed as MEF-GAN. In our architecture, a generator network and a discriminator network are trained simultaneously to form an adversarial relationship. The generator is trained to generate a real-like fused image based on the given source images which is expected to fool the discriminator. Correspondingly, the discriminator is trained to distinguish the generated fused images from the ground truth. The adversarial relationship makes the fused image not limited to the restriction of the content loss. Therefore, the fused images are closer to the ground truth in terms of probability distribution, which can compensate for the insufficiency of single content loss. Moreover, aiming at the problem that the luminance of multi-exposure images varies greatly with spatial location, the self-attention mechanism is employed in our architecture to allow for attention-driven and long-range dependency. Thus, local distortion, confusing results, or inappropriate representation can be corrected in the fused image. Qualitative and quantitative experiments are performed on publicly available datasets, where the results demonstrate that MEF-GAN outperforms the state-of-the-art, in terms of both visual effect and objective evaluation metrics. Our code is publicly available at https://github.com/jiayi-ma/MEF-GAN. Han Xu 0001, Jiayi Ma 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Image Process. | 2 |
| 2020 | Cross-Weather Image Alignment via Latent Generative Model With Intensity ConsistencyabstractImage alignment/registration/correspondence is a critical prerequisite for many vision-based tasks, and it has been widely studied in computer vision. However, aligning images from different domains, such as cross-weather/season road scenes, remains a challenging problem. Inspired by the success of classic intensity-constancy-based image alignment methods and the modern generative adversarial network (GAN) technology, we propose a cross-weather road scene alignment method called latent generative model with intensity constancy. From a novel perspective, the alignment problem is formulated as a constrained 2D flow optimization problem with latent encoding, which can be decoded into an intensity-constancy image on the latent image manifold. The manifold is parameterized by a pre-trained GAN, which is able to capture statistic characteristics from large datasets. Moreover, we employ the learned manifold to constrain the warped latent image identical to the target image, thereby producing a realistic warping effect. Experimental results on several cross-weather/season road scene datasets demonstrate that our approach can significantly outperform the state-of-the-art methods. Huabing Zhou, Jiayi Ma 0001, Chiu C. Tan 0001, Yanduo Zhang, Haibin Ling |
IEEE Trans. Image Process. | 2 |
| 2019 | NDDR-CNN: Layerwise Feature Fusing in Multi-Task CNNs by Neural Discriminative Dimensionality ReductionabstractIn this paper, we propose a novel Convolutional Neural Network (CNN) structure for general-purpose multi-task learning (MTL), which enables automatic feature fusing at every layer from different tasks. This is in contrast with the most widely used MTL CNN structures which empirically or heuristically share features on some specific layers (e.g., share all the features except the last convolutional layer). The proposed layerwise feature fusing scheme is formulated by combining existing CNN components in a novel way, with clear mathematical interpretability as discriminative dimensionality reduction, which is referred to as Neural Discriminative Dimensionality Reduction (NDDR). Specifically, we first concatenate features with the same spatial resolution from different tasks according to their channel dimension. Then, we show that the discriminative dimensionality reduction can be fulfilled by 1×1 Convolution, Batch Normalization, and Weight Decay in one CNN. The use of existing CNN components ensures the end-to-end training and the extensibility of the proposed NDDR layer to various state-of-the-art CNN architectures in a "plug-and-play" manner. The detailed ablation analysis shows that the proposed NDDR layer is easy to train and also robust to different hyperparameters. Experiments on different task sets with various base network architectures demonstrate the promising performance and desirable generalizability of our proposed method. The code of our paper is available at https://github.com/ethanygao/NDDR-CNN. Yuan Gao 0015, Jiayi Ma 0001, Ming-Bo Zhao, Wei Liu 0005, Alan L. Yuille |
CVPR | 2 |
| 2019 | Progressive Filtering for Feature MatchingabstractIn this paper, we propose a simple yet efficient method termed as Progressive Filtering for Feature Matching, which is able to establish accurate correspondences between two images of common or similar scenes. Our algorithm first grids the correspondence space and calculates a typical motion vector for each cell, and then removes false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive matching strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. The density estimation is utilized to address the island samples and accelerate the convergency of the mismatch removal procedure. In addition, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Extensive experiments on several representative real images involving different types of geometric transformations demonstrate the superiority of our approach over the state-of-the-art. Xingyu Jiang 0005, Jiayi Ma 0001, Jun Chen 0019 |
ICASSP | 2 |
| 2019 | Image Super-resolution via Deep Aggregation NetworkabstractDeep convolutional neural networks (CNNs) have recently made a considerable achievement in the single-image super-resolution (SISR) problem. Most CNN architectures for SIS-R incorporate skip connections to integrate features, and treat them equally. However, this neglects the discrimination of features, and consequently, achieving relatively poor performance. To address this problem, we introduce a deep aggregation network that merging extraction and aggregation nodes in a tree structure, which can aggregate features progressively. In particular, we rescale the information in the aggregation node by modelling the interaction between channels, which shares the same insight on the attention mechanism for improving the discriminative ability of network. In the extraction node, we introduce an mlpconv layer into a dense unit that is parallel to the convolutional layer and can improve the nonlinear mapping capability, where the residual learning is utilized to accelerate the training process. Extensive experiments conducted on several publicly available datasets have demonstrated the superiority of our model over state-of-the-art in objective metrics and visual impressions. Xinya Wang, Jiayi Ma 0001, Junjun Jiang |
ICASSP | 2 |
| 2019 | Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsabstractMost previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for video SR, which is designed to make better use of spatio-temporal information and is proved to be more efficient and effective than the existing direct fusion, slow fusion or 3D convolution strategies. Under this progressive fusion framework, we further introduce an improved non-local operation to avoid the complex motion estimation and motion compensation (ME&MC) procedures as in previous video SR approaches. Extensive experiments on public datasets demonstrate that our method surpasses state-of-the-art with 0.96 dB in average, and runs about 3 times faster, while requires only about half of the parameters. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Jiayi Ma 0001 |
ICCV | 5 |
| 2019 | GRU with Spatial Prior for Hyperspectral Image ClassificationabstractNeural networks have been successfully used to extract deep features for many hyperspectral tasks. In this study, we propose a tiny effective model based on gate recurrent unit (GRU) with spectral-spatial information for hyperspectral image classification. In our method, the core GRU cell can learn interspectral correlations within an entirely continuous spectrum input, and spatial information is the initial state of this GRU cell as a prior. Experimental results demonstrate that our method can fully utilize spectral and spatial information to obtain competitive performance. Erting Pan, Yong Ma 0001, Xiaobing Dai, Fan Fan 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IGARSS | 7 |
| 2019 | Spectral-Spatial Classification of Hyperspectral Image based on a Joint Attention NetworkabstractDeep neural networks have been successfully applied to extracting deep features for many hyperspectral tasks. Attention mechanism has been widely used in computer vision, inspired by this, we have designed a joint attention network for spectral-spatial classification of hyperspectral image. In our method, recurrent neural network (RNN) with attention can learn inner spectral correlations within a continuous spectrum, convolutional neural network (CNN) with attention is designed to focus on saliency features and spatial dependency in the neighbor regions. Experimental results demonstrate that our method can fully utilize spectral and spatial information to obtain competitive performance. Erting Pan, Yong Ma 0001, Xiaoguang Mei, Xiaobing Dai, Fan Fan 0001, Xin Tian 0006, Jiayi Ma 0001 |
IGARSS | 7 |
| 2019 | Learning a Generative Model for Fusing Infrared and Visible Images via Conditional Generative Adversarial Network with Dual DiscriminatorsabstractIn this paper, we propose a new end-to-end model, called dual-discriminator conditional generative adversarial network (DDcGAN), for fusing infrared and visible images of different resolutions. Unlike the pixel-level methods and existing deep learning-based methods, the fusion task is accomplished through the adversarial process between a generator and two discriminators, in addition to the specially designed content loss. The generator is trained to generate real-like fused images to fool discriminators. The two discriminators are trained to calculate the JS divergence between the probability distribution of downsampled fused images and infrared images, and the JS divergence between the probability distribution of gradients of fused images and gradients of visible images, respectively. Thus, the fused images can compensate for the features that are not constrained by the single content loss. Consequently, the prominence of thermal targets in the infrared image and the texture details in the visible image can be preserved or even enhanced in the fused image simultaneously. Moreover, by constraining and distinguishing between the downsampled fused image and the low-resolution infrared image, DDcGAN can be preferably applied to the fusion of different resolution images. Qualitative and quantitative experiments on publicly available datasets demonstrate the superiority of our method over the state-of-the-art. Han Xu 0001, Pengwei Liang, Wei Yu 0018, Junjun Jiang, Jiayi Ma 0001 |
IJCAI | 5 |
| 2019 | Locality Preserving Matching
Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Xiaojie Guo 0001 |
Int. J. Comput. Vis. | 1 |
| 2019 | Face hallucination through differential evolution parameter map learning with facial structure prior
Junjun Jiang, Jiayi Ma 0001, Suhua Tang, Yi Yu 0001, Kiyoharu Aizawa |
Inf. Sci. | 2 |
| 2019 | Deep transfer learning for military object recognition under small training set condition
Wei Yu 0018, Pengwei Liang, Hanqi Guo 0002, Likun Xia, Yong Ma 0001, Jiayi Ma 0001 |
Neural Comput. Appl. | 8 |
| 2019 | Feature-guided Gaussian mixture model for image matching
Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Yuan Gao 0015 |
Pattern Recognit. | 1 |
| 2019 | Gaussian field estimator with manifold regularization for retinal image registration
Jiahao Wang 0001, Jun Chen 0019, Shuaibin Zhang, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
Signal Process. | 7 |
| 2019 | Hyperspectral Image Classification in the Presence of Noisy LabelsabstractLabel information plays an important role in a supervised hyperspectral image classification problem. However, current classification methods all ignore an important and inevitable problem-labels may be corrupted and collecting clean labels for training samples is difficult and often impractical. Therefore, how to learn from the database with noisy labels is a problem of great practical importance. In this paper, we study the influence of label noise on hyperspectral image classification and develop a random label propagation algorithm (RLPA) to cleanse the label noise. The key idea of RLPA is to exploit knowledge (e.g., the superpixel-based spectral-spatial constraints) from the observed hyperspectral images and apply it to the process of label propagation. Specifically, the RLPA first constructs a spectral-spatial probability transform matrix (SSPTM) that simultaneously considers the spectral similarity and superpixel-based spatial information. It then randomly chooses some training samples as “clean” samples and sets the rest as unlabeled samples, and propagates the label information from the “clean” samples to the rest unlabeled samples with the SSPTM. By repeating the random assignment (of “clean” labeled samples and unlabeled samples) and propagation, we can obtain multiple labels for each training sample. Therefore, the final propagated label can be calculated by a majority vote algorithm. Experimental studies show that the RLPA can reduce the level of noisy label and demonstrates the advantages of our proposed method over four major classifiers with a significant margin-the gains in terms of the average overall accuracy, average accuracy, and kappa are impressive, e.g., 9.18%, 9.58%, and 0.1043. The MATLAB source code is available at https://github.com/junjun-jiang/RLPA. Junjun Jiang, Jiayi Ma 0001, Zheng Wang 0007, Chen Chen 0001, Xianming Liu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Multiscale Locality and Rank Preservation for Robust Feature Matching of Remote Sensing ImagesabstractAs a fundamental and important task in many applications of remote sensing and photogrammetry, feature matching tries to seek correspondences between the two feature sets extracted from an image pair of the same object or scene. This paper focuses on eliminating mismatches from a set of putative feature correspondences constructed according to the similarity of existing well-designed feature descriptors. Considering the stable local topological relationship of the potential true correspondences, we propose a simple yet efficient method named multiscale Top K Rank Preservation (mTopKRP) for robust feature matching. To this end, we first search the K-nearest neighbors of each feature point and generate a ranking list accordingly. Then we design a metric based on the weighted Spearman's footrule distance to describe the similarity of two ranking lists specifically for the matching problem. We build a mathematical optimization model and derive its closed-form solution, enabling our method to establish reliable correspondences in linearithmic time complexity, which requires only tens of milliseconds to handle over 1000 putative matches. We also introduce a multiscale strategy for neighborhood construction, which increases the robustness of our method and can deal with different types of degradation, even when the image pair suffers from a large scale change, rotation, nonrigid deformation, or a large number of mismatches. Extensive experiments on several representative remote sensing image data sets demonstrate the superiority of our method over state of the art. Xingyu Jiang 0005, Junjun Jiang, Aoxiang Fan, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Graph-Regularized Locality-Constrained Joint Dictionary and Residual Learning for Face Sketch SynthesisabstractFace sketch synthesis is a crucial issue in digital entertainment and law enforcement. It can bridge the considerable texture discrepancy between face photos and sketches. Most of the current face sketch synthesis approaches directly to learn the relationship between the photos and sketches, and it is very difficult for them to generate the individual specific features, which we call rare characteristics. In this paper, we propose a novel face sketch synthesis approach through residual learning. In contrast to traditional approaches, which aim to reconstruct a sketch image directly (i.e., learn the mapping relationship between the photo and sketch), we aim to predict the residual image by learning the mapping relationship between the photo and residual, i.e., the difference between the photo and sketch, given an observed photo. This technique will render optimizing the residual mapping easier than optimizing the original mapping and deriving rare characteristic information. We also introduce a joint dictionary learning algorithm by preserving the local geometry structure of a data space. Through the learned joint dictionary, we transform the face sketch synthesis from an image space to a new and compact space; the new and compact space is spanned by learned dictionary atoms, where the manifold assumption can be further guaranteed. Results show that the proposed method demonstrates an impressive performance in the face sketch synthesis task on three public face sketch datasets and various real-world photos. These results are derived by comparing the proposed method with several state-of-the-art techniques, including certain recently proposed deep learning-based approaches. Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Xianming Liu 0005, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | LMR: Learning a Two-Class Classifier for Mismatch RemovalabstractFeature matching, which refers to establishing reliable correspondence between two sets of features, is a critical prerequisite in a wide spectrum of vision-based tasks. Existing attempts typically involve the mismatch removal from a set of putative matches based on estimating the underlying image transformation. However, the transformation could vary with different data. Thus, a pre-defined transformation model is often demanded, which severely limits the applicability. From a novel perspective, this paper casts the mismatch removal into a two-class classification problem, learning a general classifier to determine the correctness of an arbitrary putative match, termed as Learning for Mismatch Removal (LMR). The classifier is trained based on a general match representation associated with each putative match through exploiting the consensus of local neighborhood structures based on a multiple K -nearest neighbors strategy. With only ten training image pairs involving about 8000 putative matches, the learned classifier can generate promising matching results in linearithmic time complexity on arbitrary testing data. The generality and robustness of our approach are verified under several representative supervised learning techniques as well as on different training and testing data. Extensive experiments on feature matching, visual homing, and near-duplicate image retrieval are conducted to reveal the superiority of our LMR over the state-of-the-art competitors. Jiayi Ma 0001, Xingyu Jiang 0005, Junjun Jiang, Ji Zhao 0001, Xiaojie Guo 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Multi-Memory Convolutional Neural Network for Video Super-ResolutionabstractVideo super-resolution (SR) is focused on reconstructing high-resolution (HR) frames from consecutive lowresolution (LR) frames. Most previous video SR methods based on convolutional neural network (CNN) use a direct connection and single-memory module within the network, and they thus fail to make full use of spatio-temporal complementary information from LR observed frames. To fully exploit spatio-temporal correlations between adjacent LR frames and reveal more realistic details, this paper proposes a multi-memory convolutional neural network (MMCNN) for video SR, cascading an optical flow network and an image-reconstruction network. A serial of residual blocks engaged in utilizing intra-frame spatial correlations are proposed for feature extraction and reconstruction. Particularly, instead of using single-memory module, we embed convolutional long short-term memory (ConvLSTM) into the residual block, thus form a multi-memory residual block to progressively extract and retain inter-frame temporal correlations between consecutive LR frames. We conduct extensive experiments on numerous testing datasets with respect to different scaling factors. Our proposed MMCNN shows superiority over the state-of-the-art methods in terms of PSNR and visual quality and surpasses the best counterpart method 1 dB at most. The code and datasets are available at https://github.com/psychopa4/MMCNN. Zhongyuan Wang 0001, Peng Yi 0002, Kui Jiang, Junjun Jiang, Zhen Han 0002, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 7 |
| 2019 | Nonrigid Point Set Registration With Robust Transformation Learning Under Manifold RegularizationabstractThis paper solves the problem of nonrigid point set registration by designing a robust transformation learning scheme. The principle is to iteratively establish point correspondences and learn the nonrigid transformation between two given sets of points. In particular, the local feature descriptors are used to search the correspondences and some unknown outliers will be inevitably introduced. To precisely learn the underlying transformation from noisy correspondences, we cast the point set registration into a semisupervised learning problem, where a set of indicator variables is adopted to help distinguish outliers in a mixture model. To exploit the intrinsic structure of a point set, we constrain the transformation with manifold regularization which plays a role of prior knowledge. Moreover, the transformation is modeled in the reproducing kernel Hilbert space, and a sparsity-induced approximation is utilized to boost efficiency. We apply the proposed method to learning motion flows between image pairs of similar scenes for visual homing, which is a specific type of mobile robot navigation. Extensive experiments on several publicly available data sets reveal the superiority of the proposed method over state-of-the-art competitors, particularly in the context of the degenerated data. Jiayi Ma 0001, Jia Wu 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Quan Z. Sheng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Residual Learning for Face Sketch SynthesisabstractFace sketch synthesis plays an important role in both digital entertainment and law enforcement. It can bridge the great texture discrepancy between face photos and sketches. Most of the current face sketch synthesis approaches directly learn the relationship between the photos and sketches, and it is very difficult for them to generate the individual specific details, which we call rare features. To address this problem, in this paper we propose a novel face sketch synthesis through residual learning. In contrast the traditional approaches, which try to construct the sketch image directly, we aim at predicting the residual image (between the photo and sketch), given the photo observation. In addition, we also introduce a couple dictionary learning algorithm through preserving the local geometry structure of data space, which is usually ignored by existing methods. Our proposed method shows impressive results on the face sketch synthesis task, when compared with some state-of-the-arts including some recent proposed deep learning based approaches. Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Jiayi Ma 0001 |
ICASSP | 4 |
| 2018 | Feature Matching Based on Top K Rank SimilarityabstractFeature matching plays a key component in many computer vision and pattern recognition tasks. Observing that the spatial neighborhood relationship (representing the topological structures of an image scene) is generally well preserved between two feature points of an image pair, some mismatch removing methods based on maintaining the local neighborhood structures of the potential true matches have been proposed. How to define the local neighborhood structure is an issue of vital importance. In this paper, we propose a robust and efficient method, called Top$K$Rank Preservation (Top-KRP), for mismatch removal from given putative point set matching correspondences. Instead of preserving the intersection of neighbors, TopKRP aims at preserving the top$K$rank of two feature points. The developed approach is validated on numerous challenging real image pairs for general feature matching, and the experimental results demonstrate that it outperforms several state-of-the-art feature matching methods, especially in case of a large number of mismatches. Junjun Jiang, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
ICASSP | 5 |
| 2018 | Visual Homing via Guided Locality Preserving MatchingabstractThis study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoramic images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from hundreds of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost true matches without sacrifice in accuracy. To apply our GLPM to the visual homing problem, we develop a method for dense motion flow estimation from sparse feature matches based on Tikhonov regularization. Moreover, the focus-of-contraction/focus-of-expansion is derived to determine homing directions. The effectiveness of our method is demonstrated on a panoramic database in both feature matching and visual homing. Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou, Yu Zhou 0016, Zheng Wang 0007, Xiaojie Guo 0001 |
ICRA | 1 |
| 2018 | Deep CNN Denoiser and Multi-layer Neighbor Component Embedding for Face HallucinationabstractMost of the current face hallucination methods, whether they are shallow learning-based or deep learning-based, all try to learn a relationship model between Low-Resolution (LR) and High-Resolution (HR) spaces with the help of a training set. They mainly focus on modeling image prior through either model-based optimization or discriminative inference learning. However, when the input LR face is tiny, the learned prior knowledge is no longer effective and their performance will drop sharply. To solve this problem, in this paper we propose a general face hallucination method that can integrate model-based optimization and discriminative inference. In particular, to exploit the model based prior, the Deep Convolutional Neural Networks (CNN) denoiser prior is plugged into the super-resolution optimization model with the aid of image-adaptive Laplacian regularization. Additionally, we further develop a high-frequency details compensation method by dividing the face image to facial components and performing face hallucination in a multi-layer neighbor embedding manner. Experiments demonstrate that the proposed method can achieve promising super-resolution results for tiny input LR faces. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001 |
IJCAI | 5 |
| 2018 | Robust GBM hyperspectral image unmixing with superpixel segmentation based low rank and sparse representation
Xiaoguang Mei, Yong Ma 0001, Chang Li 0001, Fan Fan 0001, Jun Huang 0008, Jiayi Ma 0001 |
Neurocomputing | 6 |
| 2018 | Hyperspectral Image Classification With Discriminative Kernel Collaborative Representation and Tikhonov RegularizationabstractRecently, collaborative representation has received much attention in the hyperspectral image (HSI) classification due to its simplicity and effectiveness. However, the existing collaborative representation-based HSI classification methods ignore the correlation among different classes. To overcome this problem, we propose a discriminative kernel collaborative representation and Tikhonov regularization method (DKCRT) for HSI classification, which can make the kernel collaborative representation of different classes to be more discriminative. Specifically, the kernel trick is adopted to map the original HSI into a high space to improve the class separability. Besides, distance-weighted kernel Tikhonov regularization is adopted to enforce these training samples to have large representation coefficients, which are similar to the test sample in the high-dimensional feature space. Moreover, we add a discriminative regularization term to further enhance the separability of different classes, which can take the correlation among different classes into consideration. Furthermore, to take the spatial information of HSI into consideration, we extend the DKCRT to a joint version named JDKCRT. Experiments on real HSIs demonstrate the efficiency of the proposed DKCRT and JDKCRT. Yong Ma 0001, Chang Li 0001, Hao Li 0034, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Spatial-Spectral Total Variation Regularized Low-Rank Tensor Decomposition for Hyperspectral Image DenoisingabstractSeveral bandwise total variation (TV) regularized low-rank (LR)-based models have been proposed to remove mixed noise in hyperspectral images (HSIs). These methods convert high-dimensional HSI data into 2-D data based on LR matrix factorization. This strategy introduces the loss of useful multiway structure information. Moreover, these bandwise TV-based methods exploit the spatial information in a separate manner. To cope with these problems, we propose a spatial–spectral TV regularized LR tensor factorization (SSTV-LRTF) method to remove mixed noise in HSIs. From one aspect, the hyperspectral data are assumed to lie in an LR tensor, which can exploit the inherent tensorial structure of hyperspectral data. The LRTF-based method can effectively separate the LR clean image from sparse noise. From another aspect, HSIs are assumed to be piecewisely smooth in the spatial domain. The TV regularization is effective in preserving the spatial piecewise smoothness and removing Gaussian noise. These facts inspire the integration of the LRTF with TV regularization. To address the limitations of bandwise TV, we use the SSTV regularization to simultaneously consider local spatial structure and spectral correlation of neighboring bands. Both simulated and real data experiments demonstrate that the proposed SSTV-LRTF method achieves superior performance for HSI mixed-noise removal, as compared to the state-of-the-art TV regularized and LR-based methods. Haiyan Fan, Chang Li 0001, Yulan Guo, Gangyao Kuang, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | SuperPCA: A Superpixelwise PCA Approach for Unsupervised Feature Extraction of Hyperspectral ImageryabstractAs an unsupervised dimensionality reduction method, the principal component analysis (PCA) has been widely considered as an efficient and effective preprocessing step for hyperspectral image (HSI) processing and analysis tasks. It takes each band as a whole and globally extracts the most representative bands. However, different homogeneous regions correspond to different objects, whose spectral features are diverse. Therefore, it is inappropriate to carry out dimensionality reduction through a unified projection for an entire HSI. In this paper, a simple but very effective superpixelwise PCA (SuperPCA) approach is proposed to learn the intrinsic low-dimensional features of HSIs. In contrast to classical PCA models, the SuperPCA has four main properties: 1) unlike the traditional PCA method based on a whole image, the SuperPCA takes into account the diversity in different homogeneous regions, that is, different regions should have different projections; 2) most of the conventional feature extraction models cannot directly use the spatial information of HSIs, while the SuperPCA is able to incorporate the spatial context information into the unsupervised dimensionality reduction by superpixel segmentation; 3) since the regions obtained by superpixel segmentation have homogeneity, the SuperPCA can extract potential low-dimensional features even under noise; and 4) although the SuperPCA is an unsupervised method, it can achieve a competitive performance when compared with supervised approaches. The resulting features are discriminative, compact, and noise-resistant, leading to an improved HSI classification performance. Experiments on three public data sets demonstrate that the SuperPCA model significantly outperforms the conventional PCA-based dimensionality reduction baselines for HSI classification, and some state-of-the-art feature extraction approaches. The MATLAB source code is available at https://github.com/junjun-jiang/SuperPCA. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Zhongyuan Wang 0001, Zhihua Cai, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Learning Source-Invariant Deep Hashing Convolutional Neural Networks for Cross-Source Remote Sensing Image RetrievalabstractDue to the urgent demand for remote sensing big data analysis, large-scale remote sensing image retrieval (LSRSIR) attracts increasing attention from researchers. Generally, LSRSIR can be divided into two categories as follows: uni-source LSRSIR (US-LSRSIR) and cross-source LSRSIR (CS-LSRSIR). More specifically, US-LSRSIR means the inquiry remote sensing image and images in the searching data set come from the same remote sensing data source, whereas CS-LSRSIR is designed to retrieve remote sensing images with a similar content to the inquiry remote sensing image that are from a different remote sensing data source. In the literature, US-LSRSIR has been widely exploited, but CS-LSRSIR is rarely discussed. In practical situations, remote sensing images from different kinds of remote sensing data sources are continually increasing, so there is a great motivation to exploit CS-LSRSIR. Therefore, this paper focuses on CS-LSRSIR. To cope with CS-LSRSIR, this paper proposes source-invariant deep hashing convolutional neural networks (SIDHCNNs), which can be optimized in an end-to-end manner using a series of well-designed optimization constraints. To quantitatively evaluate the proposed SIDHCNNs, we construct a dual-source remote sensing image data set that contains eight typical land-cover categories and 10 000 dual samples in each category. Extensive experiments show that the proposed SIDHCNNs can yield substantial improvements over several baselines involving the most recent techniques. Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Large-Scale Remote Sensing Image Retrieval by Deep Hashing Neural NetworksabstractAs one of the most challenging tasks of remote sensing big data mining, large-scale remote sensing image retrieval has attracted increasing attention from researchers. Existing large-scale remote sensing image retrieval approaches are generally implemented by using hashing learning methods, which take handcrafted features as inputs and map the high-dimensional feature vector to the low-dimensional binary feature vector to reduce feature-searching complexity levels. As a means of applying the merits of deep learning, this paper proposes a novel large-scale remote sensing image retrieval approach based on deep hashing neural networks (DHNNs). More specifically, DHNNs are composed of deep feature learning neural networks and hashing learning neural networks and can be optimized in an end-to-end manner. Rather than requiring to dedicate expertise and effort to the design of feature descriptors, we can automatically learn good feature extraction operations and feature hashing mapping under the supervision of labeled samples. To broaden the application field, DHNNs are evaluated under two representative remote sensing cases: scarce and sufficient labeled samples. To make up for a lack of labeled samples, DHNNs can be trained via transfer learning for the former case. For the latter case, DHNNs can be trained via supervised learning from scratch with the aid of a vast number of labeled samples. Extensive experiments on one public remote sensing image data set with a limited number of labeled samples and on another public data set with plenty of labeled samples show that the proposed remote sensing image retrieval approach based on DHNNs can remarkably outperform state-of-the-art methods under both of the examined conditions. Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Hu Zhu, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2018 | Guided Locality Preserving Feature Matching for Remote Sensing Image RegistrationabstractFeature matching, which refers to establishing reliable correspondences between two sets of feature points, is a critical prerequisite in feature-based image registration. This paper proposes a simple yet surprisingly effective approach, termed as guided locality preserving matching, for robust feature matching of remote sensing images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two images. We formulate it into a mathematical model, and derive a simple closed-form solution with linearithmic time and linear space complexities. This enables our method to accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. To handle extremely large proportions of outliers, we further design a guided matching strategy based on the proposed method, using the matching result on a small putative set with a high inlier ratio to guide the matching on a large putative set. This strategy can also significantly boost the true matches without sacrifice in accuracy. Experiments on various real remote sensing image pairs demonstrate the generality of our method for handling both rigid and nonrigid image deformations, and it is more than two orders of magnitude faster than the state-of-the-art methods with better accuracy, making it practical for real-time applications. Jiayi Ma 0001, Junjun Jiang, Huabing Zhou, Ji Zhao 0001, Xiaojie Guo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Non-Rigid Point Set Registration with Robust Transformation Estimation under Manifold RegularizationabstractIn this paper, we propose a robust transformation estimation method based on manifold regularization for non-rigid point set registration. The method iteratively recovers the point correspondence and estimates the spatial transformation between two point sets. The correspondence is established based on existing local feature descriptors which typically results in a number of outliers. To achieve an accurate estimate of the transformation from such putative point correspondence, we formulate the registration problem by a mixture model with a set of latent variables introduced to identify outliers, and a prior involving manifold regularization is imposed on the transformation to capture the underlying intrinsic geometry of the input data. The non-rigid transformation is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both 2D and 3D data demonstrate that our method can yield superior results compared to other state-of-the-arts, especially in case of badly degraded data. Jiayi Ma 0001, Ji Zhao 0001, Junjun Jiang, Huabing Zhou |
AAAI | 1 |
| 2017 | Non-rigid image deformation algorithm based on MRLS-TPSabstractIn this paper, we propose a novel closed-form transformation estimation method based on moving regularized least squares optimization with thin-plate spline (MRLS-TPS) for non-rigid image deformation. The method takes the user-controlled point-offset-vectors as the input data, and estimates the spatial transformation about the two control point sets for each pixel. To achieve a realistic deformation, we formulates the transformation estimation as a vector-field interpolation problem by a moving regularized least squares method. Unlike MLS, the mapping function is modeled by a non-rigid function thin-plate spline with regularization technique, such that the deformation can satisfy both global linear affine motion and local non-rigid warping. We derive a closed-form solution of the transformation and achieve a fast implementation. In addition, the proposed method can give a wonderful user experience, fast and convenient manipulating. Extensive experiments on real images demonstrated the proposed method outperforms other state-of-the-art methods and the commercial software Adobe PhotoShop CS 6, especially in case of flexible object motion. Huabing Zhou, Yuyu Kuang, Zhenghong Yu, Shiqiang Ren, Anna Dai, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001 |
ICIP | 8 |
| 2017 | Context-patch based face hallucination via thresholding locality-constrained representation and reproducing learningabstractFace hallucination, which refers to predicting a HighResolution (HR) face image from an observed Low-Resolution (LR) one, is a challenging problem. Most state-of-the-arts employ local face structure prior to estimate the optimal representations for each patch by the training patches of the same position, and achieve good reconstruction performance. However, they do not take into account the contextual information of image patch, which is very useful for the expression of human face. Different from position-patch based methods, in this paper we leverage the contextual information and develop a robust and efficient context-patch face hallucination algorithm, called Thresholding Locality-constrained Representation with Reproducing learning (TLcR-RL). In TLcR-RL, we use a thresholding strategy to enhance the stability of patch representation and the reconstruction accuracy. Additionally, we develop a reproducing learning to iteratively enhance the estimated result by adding the estimated HR face to the training set. Experiments demonstrate that the performance of our proposed framework has a substantial increase when compared to state-of-the-arts, including recently proposed deep learning based method. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Guo-Jun Qi, Akiko Aizawa |
ICME | 4 |
| 2017 | A unified model for improving depth accuracy in kinect sensorabstractThe Microsoft Kinect sensor has been widely used in many applications, but it suffers from the drawback of low depth accuracy. In this paper, we present a unified depth modification model to improve the Kinect depth accuracy by registering depth and color images in an iterative manner. Specifically, in each iteration, we first establish a coarse correspondence based on the feature descriptor of the canny edge. Then, we estimate the fine correspondence using a robust estimator called the L2E with the nonparametric model. Finally, we correct the depth data according to the correspondence results. In order to evaluate the effectiveness of our approach, we have performed extensive experiments and then analyzed the experimental results from the following respects: the accuracy of depth data, the accuracy of correspondence between color and depth images as well as the measurement error in the 3D reconstruction by our method. The experimental results show that our approach greatly improves the depth accuracy. Li Peng 0003, Yanduo Zhang, Huabing Zhou, Deng Chen, Zhenghong Yu, Junjun Jiang, Jiayi Ma 0001 |
ICME | 7 |
| 2017 | Feature guided non-rigid image/surface deformation via moving least squares with manifold regularizationabstractIn this paper, a novel closed-form transformation estimation method based on feature guided moving least squares together with manifold regularization is proposed for nonrigid image/surface deformation. The method takes the user-controlled point-offset-vectors and the feature points of the image/surface as input, and estimates the spatial transformation between the two control point sets for each pixel/voxel. To achieve a detail-preserving and realistic deformation, the transformation estimation is formulated as a vector-field interpolation problem using a feature guided moving least squares method, where a manifold regularization is imposed as a prior on the transformation to capture the underlying intrinsic geometry of the input image/surface. The non-rigid transformation is specified in a reproducing kernel Hilbert space. We derive a closed-form solution of the transformation and adopt a sparse approximation to achieve a fast implementation, which largely reduces the computation complexity without performance sacrifice. In addition, the proposed method can give a wonderful user experience, fast and convenient manipulating. Extensive experiments on both 2D and 3D data demonstrate that the proposed method can produce more natural deformations compared with other state-of-the-art methods. Huabing Zhou, Jiayi Ma 0001, Yanduo Zhang, Zhenghong Yu, Shiqiang Ren, Deng Chen |
ICME | 2 |
| 2017 | Locality Preserving MatchingabstractSeeking reliable correspondences between two feature sets is a fundamental and important task in computer vision. This paper attempts to remove mismatches from given putative image feature correspondences. To achieve the goal, an efficient approach, termed as locality preserving matching (LPM), is designed, the principle of which is to maintain the local neighborhood structures of those potential true matches. We formulate the problem into a mathematical model, and derive a closed-form solution with linearithmic time and linear space complexities. More specifically, our method can accomplish the mismatch removal from thousands of putative correspondences in only a few milliseconds. Experiments on various real image pairs for general feature matching, as well as for visual homing and image retrieval demonstrate the generality of our method for handling different types of image deformations, and it is more than two orders of magnitude faster than state-of-the-art methods in the same range of or better accuracy. Jiayi Ma 0001, Ji Zhao 0001, Hanqi Guo 0002, Junjun Jiang, Huabing Zhou, Yuan Gao 0015 |
IJCAI | 1 |
| 2017 | Visual homing by robust interpolation for sparse motion flowabstractIn this paper, we propose a visual homing method by mismatch removal and robust interpolation of sparse motion flows. First, a smoothness prior is proposed and verified by using synthetic and real panoramic images. Then a mismatch removal method is developed for panoramic images based on such smoothness prior. Finally, an interpolation function for the motion flow is obtained simultaneously as a byproduct of mismatch removal, and the focus-of-expansion or focus-of-contraction is used to determine homing directions. The proposed visual homing can be used alone. Also the mismatch removal method can be used as a pre-processing method, and integrated with any other visual homing method which depends on precious keypoints matching. The effectiveness of our method is demonstrated by a panoramic dataset for visual homing. Ji Zhao 0001, Jiayi Ma 0001 |
IROS | 2 |
| 2017 | Mutually Guided Image FilteringabstractImage filtering is helpful to numerous multimedia, computer vision and graphics tasks. Linear translation-invariant filters with manually designed kernels have been widely used. However, their performance suffers from the content-blindness, say identically treating noises, textures and structures. To mitigate the content-blindness, a family of filters, called joint/guided filters, has attracted much attention from the community, the principle of which is transferring the structure in the reference image to the target one. The main drawback of most joint/guided filters comes from the ignorance of structural inconsistency between the reference and target signals that can be like color, infrared and depth images captured under different conditions. Simply adopting such guidances very likely leads to unsatisfactory results. To address the above issues, this paper designs a simple yet effective filter, named as mutually guided image filter (muGIF), which jointly preserves mutual structures, avoids misleading from inconsistent structures and smooths flat regions. The proposed muGIF is very flexible, which can perform in one of dynamic only (self-guided), static/dynamic and dynamic/dynamic modes. Although the objective of muGIF is in nature non-convex, by subtly decomposing the objective, we can solve it effectively and efficiently. The advantages of muGIF in terms of effectiveness and flexibility are demonstrated over other state-of-the-art alternatives on a variety of applications. Xiaojie Guo 0001, Yu Li 0003, Jiayi Ma 0001 |
ACM Multimedia | 3 |
| 2017 | Statistical Inference of Gaussian-Laplace Distribution for Person VerificationabstractMetric learning is an important issue in the person verification problem, which is to identify whether a pair of face or human body images is about the same person. Due to low running cost, the non-iterative statistical inference methods for metric learning show their efficiency and effectiveness to large scale datasets and on-line updating person verification applications. The KISSME method is a typical one that constructs the metric based on two assumptions that both of the discrepancy spaces of negative pairs and positive pairs should be Gaussian structures. However, we find that, in fact, the distribution of discrepancies of positive pairs might tend to the Laplace distribution rather than the Gaussian distribution. Based on this finding, we propose a metric learning method by exploiting Gaussian-Laplace distribution statistical inference, where the Gaussian distribution of negative discrepancies and the Laplace distribution of positive discrepancies are considered together. Experiments conducted on two human body datasets (VIPeR and Market-1501) and one face dataset (LFW) show its superiority in terms of effectiveness and efficiency as compared with the state-of-the-art approaches, no matter the appearance description is handcrafted or deep learned. Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Jiayi Ma 0001, Shin'ichi Satoh 0001 |
ACM Multimedia | 5 |
| 2017 | Hyperspectral image denoising with superpixel segmentation and low-rank representation
Fan Fan 0001, Yong Ma 0001, Chang Li 0001, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001 |
Inf. Sci. | 6 |
| 2017 | Feature guided Gaussian mixture model with semi-supervised EM and local geometric constraint for retinal image registration
Jiayi Ma 0001, Junjun Jiang, Chengyin Liu, Yansheng Li 0001 |
Inf. Sci. | 1 |
| 2017 | Spatial-Aware Collaborative Representation for Hyperspectral Remote Sensing Image ClassificationabstractRepresentation-residual-based classifiers have attracted much attention in recent years in hyperspectral image (HSI) classification. How to obtain the optimal representa-tion coefficients for the classification task is the key problem of these methods. In this letter, spatial-aware collaborative representation (CR) is proposed for HSI classification. In order to make full use of the spatial-spectral information, we propose a closed-form solution, in which the spatial and spectral features are both utilized to induce the distance-weighted regularization terms. Different from traditional CR-based HSI classification algorithms, which model the spatial feature in a preprocessing or postprocessing stage, we directly incorporate the spatial information by adding a spatial regularization term to the representation objective function. The experimental results on three HSI data sets verify that our proposed approach outperforms the state-of-the-art classifiers. Junjun Jiang, Chen Chen 0001, Yi Yu 0001, Xinwei Jiang, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Noise Robust Face Image Super-Resolution Through Smooth Sparse RepresentationabstractFace image super-resolution has attracted much attention in recent years. Many algorithms have been proposed. Among them, sparse representation (SR)-based face image super-resolution approaches are able to achieve competitive performance. However, these SR-based approaches only perform well under the condition that the input is noiseless or has small noise. When the input is corrupted by large noise, the reconstruction weights (or coefficients) of the input low-resolution (LR) patches using SR-based approaches will be seriously unstable, thus leading to poor reconstruction results. To this end, in this paper, we propose a novel SR-based face image super-resolution approach that incorporates smooth priors to enforce similar training patches having similar sparse coding coefficients. Specifically, we introduce the fused least absolute shrinkage and selection operator-based smooth constraint and locality-based smooth constraint to the least squares representation-based patch representation in order to obtain stable reconstruction weights, especially when the noise level of the input LR image is high. Experiments are carried out on the benchmark FEI face database and CMU+MIT face database. Visual and quantitative comparisons show that the proposed face image super-resolution method yields superior reconstruction results when the input LR face image is contaminated by strong noise. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Xinwei Jiang, Zheng Wang 0007 |
IEEE Trans. Cybern. | 2 |
| 2017 | Robust Sparse Hyperspectral Unmixing With ell2, 1 NormabstractSparse unmixing (SU) of hyperspectral data have recently received particular attention for analyzing remote sensing images, which aims at finding the optimal subset of signatures to best model the mixed pixel in the scene. However, most SU methods are based on the commonly admitted linear mixing model, which ignores the possible nonlinear effects (i.e., nonlinearity), and the nonlinearity is merely treated as outlier. Besides, the traditional SU algorithms often adopt the$\ell _{2}$norm loss function, which makes them sensitive to noises and outliers. In this paper, we propose a robust SU (RSU) method with$\ell _{2,1}$norm loss function, which is robust for noises and outliers. Then, the RSU can be solved by the alternative direction method of multipliers. Finally, the experiments on both synthetic data sets and real hyperspectral images demonstrate that the proposed RSU is efficient for solving the hyperspectral SU problem compared with the state-of-the-art algorithms. Yong Ma 0001, Chang Li 0001, Xiaoguang Mei, Chengyin Liu, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Semi-Supervised Sparse Representation Based Classification for Face Recognition With Insufficient Labeled SamplesabstractThis paper addresses the problem of face recognition when there is only few, or even only a single, labeled examples of the face that we wish to recognize. Moreover, these examples are typically corrupted by nuisance variables, both linear (i.e., additive nuisance variables, such as bad lighting and wearing of glasses) and non-linear (i.e., non-additive pixel-wise nuisance variables, such as expression changes). The small number of labeled examples means that it is hard to remove these nuisance variables between the training and testing faces to obtain good recognition performance. To address the problem, we propose a method called semi-supervised sparse representation-based classification. This is based on recent work on sparsity, where faces are represented in terms of two dictionaries: a gallery dictionary consisting of one or more examples of each person, and a variation dictionary representing linear nuisance variables (e.g., different lighting conditions and different glasses). The main idea is that: 1) we use the variation dictionary to characterize the linear nuisance variables via the sparsity framework and 2) prototype face images are estimated as a gallery dictionary via a Gaussian mixture model, with mixed labeled and unlabeled samples in a semi-supervised manner, to deal with the non-linear nuisance variations between labeled and unlabeled samples. We have done experiments with insufficient labeled samples, even when there is only a single labeled sample per person. Our results on the AR, Multi-PIE, CAS-PEAL, and LFW databases demonstrate that the proposed method is able to deliver significantly improved performance over existing methods. Yuan Gao 0015, Jiayi Ma 0001, Alan L. Yuille |
IEEE Trans. Image Process. | 2 |
| 2017 | SRLSP: A Face Image Super-Resolution Algorithm Using Smooth Regression With Local Structure PriorabstractThe performance of traditional face recognition systems is sharply reduced when encountered with a low-resolution (LR) probe face image. To obtain much more detailed facial features, some face super-resolution (SR) methods have been proposed in the past decade. The basic idea of a face image SR is to generate a high-resolution (HR) face image from an LR one with the help of a set of training examples. It aims at transcending the limitations of optical imaging systems. In this paper, we regard face image SR as an image interpolation problem for domain-specific images. A missing intensity interpolation method based on smooth regression with a local structure prior (LSP), named SRLSP for short, is presented. In order to interpolate the missing intensities in a target HR image, we assume that face image patches at the same position share similar local structures, and use smooth regression to learn the relationship between LR pixels and missing HR pixels of one position patch. Performance comparison with the state-of-the-art SR algorithms on two public face databases and some real-world images shows the effectiveness of the proposed method for a face image SR in general. In addition, we conduct a face recognition experiment on the extended Yale-B face database based on the super-resolved HR faces. Experimental results clearly validate the advantages of our proposed SR method over the state-of-the-art SR methods in face recognition application. Junjun Jiang, Chen Chen 0001, Jiayi Ma 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu |
IEEE Trans. Multim. | 3 |
| 2017 | Single Image Super-Resolution via Locally Regularized Anchored Neighborhood Regression and Nonlocal MeansabstractThe goal of learning-based image super resolution (SR) is to generate a plausible and visually pleasing high-resolution (HR) image from a given low-resolution (LR) input. The SR problem is severely underconstrained, and it has to rely on examples or some strong image priors to reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e., projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. Encouraged by recent developments in image prior modeling, where the state-of-the-art algorithms are formed with nonlocal self-similarity and local geometry priors, we seek an SR algorithm of similar nature that will incorporate these two priors into the learning from LR space to HR space. The nonlocal self-similarity prior takes advantage of the redundancy of similar patches in natural images, while the local geometry prior of the data space can be used to regularize the modeling of the nonlinear relationship between LR and HR spaces. Based on the above two considerations, we first apply the local geometry prior to regularize the patch representation, and then utilize the nonlocal means filter to improve the super-resolved outcome. Experimental results verify the effectiveness of the proposed algorithm compared with the state-of-the-art SR methods. Junjun Jiang, Chen Chen 0001, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 6 |
| 2016 | Hyperspectral image denoising based on low-rank representation and superpixel segmentationabstractRecently, low-rank representation (LRR) based methods have been used for hyperspectral image (HSI) denoising, which can simultaneously remove different types of noise: Gaussian noise, impulse noise, dead lines, and so on. However, the LRR based method does not make full use of the spatial information in HSI. In this paper, we integrate the superpixel segmentation (SS) into the LRR, and propose a novel denoising method named SS-LRR. We first use the principle component analysis (PCA) to obtain the first principle component of HSI. Then the superpixel segmentation is adopted to the first principle component of HSI to get homogeneous regions. Finally, we employ the LRR to each homogeneous region of HSI, which enable us to simultaneously remove all the above mentioned mixed noise. Extensive experiments on both simulated and real hyperspectral images demonstrate that the proposed SS-LRR is efficient for HSI denoising. Jiayi Ma 0001, Chang Li 0001, Yong Ma 0001, Zhongyuan Wang 0001 |
ICIP | 1 |
| 2016 | Robust image matching via feature guided Gaussian mixture modelabstractIn this paper, we propose a novel feature guided Gaussian mixture model (FG-GMM) for image matching, which typically requires matching two sets of feature points extracted from the given images. We formulate the problem as estimation of a feature guided mixture of densities: a GMM is fitted to one point set, such that both the centers and local features of the Gaussian densities are constrained to coincide with the other point set. The problem is solved under a unified maximum-likelihood framework together with an iterative semi-supervised Expectation-Maximization (EM) algorithm initialized by the confident feature correspondences. The image transformation is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on various real images show the robustness of our approach, which consistently outperforms other state-of-the-art methods. Jiayi Ma 0001, Junjun Jiang, Yuan Gao 0015, Jun Chen 0019, Chengyin Liu |
ICME | 1 |
| 2016 | Registration of remote sensing images with non-rigid distortionsabstractIn this paper, we propose a novel formulation for building accurate pixel-wise alignments between remote sensing images under non-rigid distortions. Our formulation involves two variables: the first is a discrete displacement flow field similar to optical flow which controls the pixel-wise correspondence and allows piecewise smoothness, while the second is a continuous spatial transformation which fits for a few confidential sparse feature correspondences. An additional term is introduced to ensure the coherence between the two variables, and the continuous spatial transformation plays a role of anchor for optimizing the discrete displacement flow field. Experiments on real remote sensing images demonstrate that our approach greatly outperforms state-of-the-art methods. Jiayi Ma 0001, Jun Chen 0019, Yong Ma 0001 |
IGARSS | 1 |
| 2016 | Robust sparse unmixing of hyperspectral dataabstractSparse unmixing (SU) of hyperspectral data has recently received particular attention for analyzing remote sensing images, which aims at finding the optimal subset of signatures to best model the mixed pixel in the scene. However, most SU methods are based on the commonly admitted linear mixing model (LMM), which ignores the possible nonlinear effects (i.e. nonlinearity), and the nonlinearity is merely treated as outlier. Besides, the traditional SU algorithms often adopt the ℒ2norm loss function, which makes them sensitive to noises and outliers. In this paper, we propose a robust sparse unmixing (RSU) method with ℒ2,1norm loss function, which is robust for noises and outliers. Then, the RSU can be solved by the alternative direction method of multipliers (ADMM). Finally, experiments on synthetic datasets demonstrate that the proposed RSU is efficient for solving the hyperspectral SU problem compared with state-of-the-art algorithms. Yong Ma 0001, Chang Li 0001, Jiayi Ma 0001 |
IGARSS | 3 |
| 2016 | Smooth sparse representation for noise robust face super-resolutionabstractFace super-resolution has attracted much attention in recent years. Many algorithms have been proposed. Among them, sparse representation based face super-resolution approaches are able to achieve competitive performance. However, these sparse representation based approaches only perform well under the condition that the input is noiseless or has small noise. When the input is corrupted by large noise, the reconstruction weights of the input LR patches using sparse representation based approaches will be seriously unstable, thus leading to poor reconstruction results. To this end, in this paper, we propose a novel sparse representation based face super-resolution approach that incorporates a smooth prior to enforce similar training patches having similar sparse coding coefficients. Specifically, we introduce the fused Lasso to the least squares representation of the input LR image in order to obtain a stable sparse representation, especially when the noise level of the input LR image is high. Experiments are carried out on the benchmark FEI face dataset. Visual and quantitative comparisons show that the proposed face super-resolution method achieves comparable performance to the state-of-the-art methods under noiseless condition, and yields superior super-resolution results when the input LR face image is contaminated by strong noise. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Zhongyuan Wang 0001, Tao Lu 0001 |
VCIP | 2 |
| 2016 | Sparse unmixing of hyperspectral data based on robust linear mixing modelabstractRecently, sparse unmixing (SU) of hyperspectral data has received particular attention for analyzing remote sensing images. However, most of SU methods are based on the commonly admitted linear mixing model (LMM), which ignores the possible nonlinear effects (i.e. nonlinearity). In this paper, we proposed a new method named robust collaborative sparse regression (RCSR) for hyperspectral unmixing, which is based on the robust LMM (rLMM). The rLMM takes the nonlinearity into consideration, and the nonlinearity is merely treated as outlier, which has the underlying sparse property. The RCSR takes the collaborative sparse property of the abundance and sparsely distributed additive property of the outlier into consideration, which can be formed as a robust joint sparse regression problem. Experiments on synthetic datasets demonstrate that the proposed RCSR is efficient for solving the hyperspectral SU problem compared with other five state-of-the-art algorithms. Chang Li 0001, Yong Ma 0001, Yuan Gao 0015, Zhongyuan Wang 0001, Jiayi Ma 0001 |
VCIP | 5 |
| 2016 | Object localization by density-based spatial clusteringabstractRegion search is widely used for object localization in computer vision area. After projecting the score of an image classifier into an image plane, region search aims to find regions that precisely localize desired objects. The popular region search methods, such as efficient subwindow search and efficient region search, usually find regions with maximal score. In this paper, we observe that there is a large score density around a desired object usually. Based on this observation, we proposed a region search method by density-based spatial clustering. The resulted regions of this method can guarantee that their density is above a threshold. Besides, this method has linear time complexity for regularly sampling feature points, which is useful for popular bag-of-words feature representation. We demonstrate its superiority for synthetic data and real image dataset for weakly-supervised localization task. Ya Lu, Ji Zhao 0001, Jiayi Ma 0001 |
VCIP | 3 |
| 2016 | Multimodal retinal image registration using edge map and feature guided Gaussian mixture modelabstractIn this paper, we propose a method for multimodal retinal image registration based on feature guided Gaussian mixture model (GMM) and edge map. We extract two sets of feature points from the edge maps of two images, and formulate image registration as the estimation of a feature guided mixture of densities: a GMM is fitted to one point set, such that both the centers and local features of the Gaussian densities are constrained to coincide with the other point set. The problem is solved under a maximum-likelihood framework together with an iterative EM algorithm initialized by confident feature matches, where the image transformation is modeled by an affine function. Extensive experiments on various retinal images show the robustness of our method, which consistently outperforms other state-of-the-arts, especially when the data is badly degraded. Jiayi Ma 0001, Junjun Jiang, Jun Chen 0019, Chengyin Liu, Chang Li 0001 |
VCIP | 1 |
| 2016 | Hyperspectral image denoising with segmentation-based low rank representationabstractRecently, low-rank representation (LRR) based hyperspectral image (HSI) denoising method has been proven to be a powerful tool for removing different kinds of noise simultaneously, such as Gaussian, dead pixels and impulse noise. However, the LRR based method cannot make full use of the spatial information in HSI. In this paper, we integrate the graph based segmentation (GS) into the LRR, and propose a novel denoising method named GS-LRR. We first use the principle component analysis (PCA) to obtain the first principle component of HSI. Then the graph based segmentation is adopted to the first principle component of HSI to get homogeneous regions. Finally, we employ the LRR to each homogeneous region of HSI, which enable us to simultaneously remove all the above mentioned mixed noise. Extensive experiments on both simulated and real HSIs demonstrate the efficiency of the proposed GS-LRR. Jiayi Ma 0001, Junjun Jiang, Chang Li 0001 |
VCIP | 1 |
| 2016 | Infrared and visible image fusion using total variation model
Yong Ma 0001, Jun Chen 0019, Chen Chen 0003, Fan Fan 0001, Jiayi Ma 0001 |
Neurocomputing | 5 |
| 2016 | A novel spatio-temporal saliency approach for robust dim moving target detection from airborne infrared image sequences
Yansheng Li 0001, Yongjun Zhang 0002, Jin-Gang Yu, Yihua Tan, Jinwen Tian, Jiayi Ma 0001 |
Inf. Sci. | 6 |
| 2016 | An Infrared Small Target Detecting Algorithm Based on Human Visual SystemabstractInfrared (IR) small target detection with high detection rate, low false alarm rate, and multiscale detection ability is a challenging task since raw IR images usually have low contrast and complex background. In recent years, robust human visual system (HVS) properties have been introduced into the IR small target detection field. However, existing algorithms based on HVS, such as difference of Gaussians (DoG) filters, are sensitive to not only real small targets but also background edges, which results in a high false alarm rate. In this letter, the difference of Gabor (DoGb) filters is proposed and improved (IDoGb), which is an extension of DoG but is sensitive to orientations and can better suppress the complex background edges, then achieves a lower false alarm rate. In addition, multiscale detection can be also achieved. Experimental results show that the IDoGb filter produces less false alarms at the same detection rate, while consuming only about 0.1 s for a single frame. Jinhui Han, Yong Ma 0001, Jun Huang 0008, Xiaoguang Mei, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | GBM-Based Unmixing of Hyperspectral Data Using Bound Projected Optimal Gradient MethodabstractThe generalized bilinear model (GBM) has been widely used for the nonlinear unmixing of hyperspectral images, and traditional GBM solvers include the Bayesian algorithm, the gradient descent algorithm, the semi-nonnegative-matrix-factorization algorithm, etc. However, they suffer from one of the following problems: high computational cost, sensitive to initialization, and the pixelwise algorithm hinders us from applying to large hyperspectral images. In this letter, we apply Nesterov's optimal gradient method to solve the least-square problem under the bound constraint, which is named as the bound projected optimal gradient method (BPOGM). The BPOGM can achieve the optimal convergence rate of$O(1/k^{2})$, with$k$denoting the number of iterations in BPOGM. We further apply the BPOGM to solve the GBM-based unmixing problem. Experiments on both synthetic data sets and real hyperspectral images demonstrate that the BPOGM is efficient for solving the GBM-based unmixing problem. Chang Li 0001, Yong Ma 0001, Jun Huang 0008, Xiaoguang Mei, Chengyin Liu, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2016 | Hyperspectral Image Classification With Robust Sparse RepresentationabstractRecently, the sparse representation-based classification (SRC) methods have been successfully used for the classification of hyperspectral imagery, which relies on the underlying assumption that a hyperspectral pixel can be sparsely represented by a linear combination of a few training samples among the whole training dictionary. However, the SRC-based methods ignore the sparse representation residuals (i.e., outliers), which may make the SRC not robust for outliers in practice. To overcome this problem, we propose a robust SRC (RSRC) method which can handle outliers. Moreover, we extend the RSRC to the joint robust sparsity model named JRSRC, where pixels in a small neighborhood around the test pixel are simultaneously represented by linear combinations of a few training samples and outliers. The JRSRC can also deal with outliers in hyperspectral classification. Experiments on real hyperspectral images demonstrate that the proposed RSC and JRSRC have better performances than the orthogonal matching pursuit (OMP) and simultaneous OMP, respectively. Moreover, the JRSRC outperforms some other popular classifiers. Chang Li 0001, Yong Ma 0001, Xiaoguang Mei, Chengyin Liu, Jiayi Ma 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2016 | Nonrigid Feature Matching for Remote Sensing Images via Probabilistic Inference With Global and Local RegularizationsabstractIn this letter, we propose a probabilistic method for the feature matching of remote sensing images which undergo nonrigid transformations. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose nonparametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the expectation-maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real remote sensing images demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, particularly in case of severe outliers. Huabing Zhou, Jiayi Ma 0001, Changcai Yang, Renfeng Liu, Ji Zhao 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Image retrieval based on image-to-class similarity
Jun Chen 0019, Yong Wang 0036, Linbo Luo 0002, Jin-Gang Yu, Jiayi Ma 0001 |
Pattern Recognit. Lett. | 5 |
| 2016 | Facial Image Hallucination Through Coupled-Layer Neighbor EmbeddingabstractAs the facial image captured by a low-cost camera is typically very low resolution (LR), blurring, and noisy, traditional neighbor-embedding-based facial image hallucination methods from one single manifold (i.e., the LR image manifold) fail to reliably estimate the intention geometrical structure, consequently leading to a bias to the image reconstruction result. In this paper, we introduce the notion of neighbor embedding (NE) from the LR and the high-resolution (HR) image manifolds simultaneously and propose a novel NE model, termed the coupled-layer NE (CLNE), for facial image hallucination. CLNE differs substantially from other NE models in that it has two layers: the LR and the HR layers. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights of the LR patches; the HR layer is the intrinsic geometry that can geometrically constrain the reconstruction weights. With this coupled-constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust NE through iteratively updating the LR patch reconstruction weights and the estimated HR patch. The experimental results in simulation and real conditions confirm that the proposed method outperforms the related state-of-the-art methods in both quantitative and visual comparisons. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Non-Rigid Point Set Registration by Preserving Global and Local StructuresabstractIn previous work on point registration, the input point sets are often represented using Gaussian mixture models and the registration is then addressed through a probabilistic approach, which aims to exploit global relationships on the point sets. For non-rigid shapes, however, the local structures among neighboring points are also strong and stable and thus helpful in recovering the point correspondence. In this paper, we formulate point registration as the estimation of a mixture of densities, where local features, such as shape context, are used to assign the membership probabilities of the mixture model. This enables us to preserve both global and local structures during matching. The transformation between the two point sets is specified in a reproducing kernel Hilbert space and a sparse approximation is adopted to achieve a fast implementation. Extensive experiments on both synthesized and real data show the robustness of our approach under various types of distortions, such as deformation, noise, outliers, rotation, and occlusion. It greatly outperforms the state-of-the-art methods, especially when the data is badly degraded. Jiayi Ma 0001, Ji Zhao 0001, Alan L. Yuille |
IEEE Trans. Image Process. | 1 |
| 2015 | Density-based region search with arbitrary shape for object localisationabstractRegion search is widely used for object localisation in computer vision. After projecting the score of an image classifier into an image plane, region search aims to find regions that precisely localise desired objects. The recently proposed region search methods, such as efficient subwindow search and efficient region search, usually find regions with maximal score. For some classifiers and scenarios, the projected scores are nearly all positive or very noisy, then maximising the score of a region results in localising nearly the entire images as objects, or causes localisation results unstable. In this study, the authors observe that the projected scores with large magnitudes are mainly concentrated on or around objects. On the basis of this observation, they propose a region search method for object localisation, named level set maximum‐weight connected subgraph (LS‐MWCS). It localises objects by searching regions by graph mode‐seeking rather than the maximal score. The score density by localised region can be controlled by a parameter flexibly. They also prove an interesting property of the proposed LS‐MWCS, which guarantees that the region with desired density can be found. Moreover, the LS‐MWCS can be efficiently solved by the belief propagation scheme. The effectiveness of the author's method is validated on the problem of weakly‐supervised object localisation. Quantitative results on synthetic and real data demonstrate the superiorities of their method compared to other state‐of‐the‐art methods. Ji Zhao 0001, Deyu Meng, Jiayi Ma 0001 |
IET Comput. Vis. | 3 |
| 2015 | Non-rigid visible and infrared face registration via regularized Gaussian fields criterion
Jiayi Ma 0001, Ji Zhao 0001, Yong Ma 0001, Jinwen Tian |
Pattern Recognit. | 1 |
| 2015 | Non-rigid point set registration via coherent spatial mapping
Jun Chen 0019, Jiayi Ma 0001, Changcai Yang |
Signal Process. | 2 |
| 2015 | Image Feature Matching via Progressive Vector Field ConsensusabstractIn this letter, we propose a simple yet effective approach, named Progressive Vector Field Consensus (PVFC), for addressing the problem of finding more true feature correspondences between images. The key idea is to progressively perform feature matching based on Vector Field Consensus, and hence greatly boost the number of true matches as well as avoid false matches. More specifically, it uses matching results on a small putative correspondence set with high inlier ratio to guide the matching on a large putative correspondence set which probably covers the whole true correspondences. We model the transformation between images in a reproducing kernel Hilbert space, and a sparse approximation is applied to the transformation to avoid high computational complexity. Our results quantitatively show that our PVFC outperforms state-of-the-art methods, both in accuracy and in efficiency. Moreover, the progressive framework is general and can be applied to other cases for robust estimation. Jiayi Ma 0001, Yong Ma 0001, Ji Zhao 0001, Jinwen Tian |
IEEE Signal Process. Lett. | 1 |
| 2015 | Robust Feature Matching for Remote Sensing Image Registration via Locally Linear TransformingabstractFeature matching, which refers to establishing reliable correspondence between two sets of features (particularly point features), is a critical prerequisite in feature-based registration. In this paper, we propose a flexible and general algorithm, which is called locally linear transforming (LLT), for both rigid and nonrigid feature matching of remote sensing images. We start by creating a set of putative correspondences based on the feature similarity and then focus on removing outliers from the putative set and estimating the transformation as well. We formulate this as a maximum-likelihood estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. To ensure the well-posedness of the problem, we develop a local geometrical constraint that can preserve local structures among neighboring feature points, and it is also robust to a large number of outliers. The problem is solved by using the expectation-maximization algorithm (EM), and the closed-form solutions of both rigid and nonrigid transformations are derived in the maximization step. In the nonrigid case, we model the transformation between images in a reproducing kernel Hilbert space (RKHS), and a sparse approximation is applied to the transformation that reduces the method computation complexity to linearithmic. Extensive experiments on real remote sensing images demonstrate accurate results of LLT, which outperforms current state-of-the-art methods, particularly in the case of severe outliers (even up to 80%). Jiayi Ma 0001, Huabing Zhou, Ji Zhao 0001, Yuan Gao 0015, Junjun Jiang, Jinwen Tian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | A robust and outlier-adaptive method for non-rigid point registration
Yuan Gao 0015, Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Dazhi Zhang |
Pattern Anal. Appl. | 2 |
| 2014 | MsLRR: A Unified Multiscale Low-Rank Representation for Image SegmentationabstractIn this paper, we present an efficient multiscale low-rank representation for image segmentation. Our method begins with partitioning the input images into a set of superpixels, followed by seeking the optimal superpixel-pair affinity matrix, both of which are performed at multiple scales of the input images. Since low-level superpixel features are usually corrupted by image noise, we propose to infer the low-rank refined affinity matrix. The inference is guided by two observations on natural images. First, looking into a single image, local small-size image patterns tend to recur frequently within the same semantic region, but may not appear in semantically different regions. The internal image statistics are referred to as replication prior, and we quantitatively justified it on real image databases. Second, the affinity matrices at different scales should be consistently solved, which leads to the cross-scale consistency constraint. We formulate these two purposes with one unified formulation and develop an efficient optimization procedure. The proposed representation can be used for both unsupervised or supervised image segmentation tasks. Our experiments on public data sets demonstrate the presented method can substantially improve segmentation accuracy. Xiaobai Liu, Jiayi Ma 0001, Hai Jin 0001, Yanduo Zhang |
IEEE Trans. Image Process. | 3 |
| 2014 | Robust Point Matching via Vector Field ConsensusabstractIn this paper, we propose an efficient algorithm, called vector field consensus, for establishing robust point correspondences between two sets of points. Our algorithm starts by creating a set of putative correspondences which can contain a very large number of false correspondences, or outliers, in addition to a limited number of true correspondences (inliers). Next, we solve for correspondence by interpolating a vector field between the two point sets, which involves estimating a consensus of inlier points whose matching follows a nonparametric geometrical constraint. We formulate this a maximum a posteriori (MAP) estimation of a Bayesian model with hidden/latent variables indicating whether matches in the putative set are outliers or inliers. We impose nonparametric geometrical constraints on the correspondence, as a prior distribution, using Tikhonov regularizers in a reproducing kernel Hilbert space. MAP estimation is performed by the EM algorithm which by also estimating the variance of the prior model (initialized to a large value) is able to obtain good estimates very quickly (e.g., avoiding many of the local minima inherent in this formulation). We illustrate this method on data sets in 2D and 3D and demonstrate that it is robust to a very large number of outliers (even up to 90%). We also show that in the special case where there is an underlying parametric geometrical model (e.g., the epipolar line constraint) that we obtain better results than standard alternatives like RANSAC if a large number of outliers are present. This suggests a two-stage strategy, where we use our nonparametric model to reduce the size of the putative set and then apply a parametric variant of our approach to estimate the geometric parameters. Our algorithm is computationally efficient and we provide code for others to use it. In addition, our approach is general and can be applied to other problems, such as learning with a badly corrupted training data set. Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Alan L. Yuille, Zhuowen Tu |
IEEE Trans. Image Process. | 1 |
| 2013 | Robust Estimation of Nonrigid Transformation for Point Set RegistrationabstractWe present a new point matching algorithm for robust nonrigid registration. The method iteratively recovers the point correspondence and estimates the transformation between two point sets. In the first step of the iteration, feature descriptors such as shape context are used to establish rough correspondence. In the second step, we estimate the transformation using a robust estimator called L_2E. This is the main novelty of our approach and it enables us to deal with the noise and outliers which arise in the correspondence step. The transformation is specified in a functional space, more specifically a reproducing kernel Hilbert space. We apply our method to nonrigid sparse image feature correspondence on 2D images and 3D surfaces. Our results quantitatively show that our approach outperforms state-of-the-art methods, particularly when there are a large number of outliers. Moreover, our method of robustly estimating transformations from correspondences is general and has many other applications. Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Zhuowen Tu, Alan L. Yuille |
CVPR | 1 |
| 2013 | LIPID: Local Image Permutation Interval DescriptorabstractImage representation through local descriptors is the basis of numerous computer vision applications. In the past decade, many local image descriptors such as SIFT and SURF have been proposed, yet algorithms requiring low memory and computation complexity are still preferred. Binary descriptors such as BRIEF have been suggested to satisfy this demand, showing a comparable performance but much faster computation speed. In this paper, we propose a novel local image descriptor, LIPID, which employs intensity permutation and interval division to yield an effective performance in terms of speed and recognition. Our method is inspired by LUCID, proposed by Ziegler and Christiansen [8]. An extensive evaluation on the well-known benchmark datasets reveals the robustness and effectiveness of LIPID as well as its capability to handle illumination changes and texture images. Tian Tian 0006, Ishwar K. Sethi, Delie Ming, Jiayi Ma 0001 |
ICMLA (2) | 5 |
| 2013 | Regularized vector field learning with sparse approximation for mismatch removal
Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian, Xiang Bai, Zhuowen Tu |
Pattern Recognit. | 1 |
| 2013 | Nonrigid Image Deformation Using Moving Regularized Least SquaresabstractThis letter presents an image deformation method based on Moving Regularized Least Squares optimization. The user controls the deformation by simply choosing a set of point handles in the input image, and also the target positions that the source point handles should be deformed to. The deformation function in our method is nonrigid and specified in a functional space, more specifically a reproducing kernel Hilbert space. The proposed method possesses three characteristics: 1) it is able to create detail-preserving and intuitive deformations; 2) the solution of the deformation function has a simple closed-form; 3) it is extremely computationally efficient which can be performed in real-time (less than 0.1 milliseconds per frame for an image of size 500 ×500). We compare our method to a state-of-the-art method which is modeled by rigid transformations; the qualitative and quantitative results demonstrate the benefits of using the nonrigid formulation in aspects of both accuracy and efficiency. Moreover, the proposed method is general and it can be applied to other applications for interpolation. Jiayi Ma 0001, Ji Zhao 0001, Jinwen Tian |
IEEE Signal Process. Lett. | 1 |
| 2012 | Mismatch removal via coherent spatial mappingabstractWe propose a method for removing mismatches from given putative point correspondences in image pairs. Our algorithm aims to recover the underlying coherent spatial mapping which related to inliers. The thin-plate spline (TPS) is chosen to parameterize the coherent spatial mapping, and we formulate the solution of it as a maximum likelihood problem. The mismatches could be successfully removed after the EM algorithm, which we used for solving the problem, converges. The quantitative results on various experimental data demonstrate that our method outperforms many state-of-the-art methods. Moreover, the proposed method is also able to handle the case that image pairs contain non-rigid motions. Jiayi Ma 0001, Ji Zhao 0001, Yu Zhou 0016, Jinwen Tian |
ICIP | 1 |
| 2011 | A robust method for vector field learning with application to mismatch removingabstractWe propose a method for vector field learning with outliers, called vector field consensus (VFC). It could distinguish inliers from outliers and learn a vector field fitting for the inliers simultaneously. A prior is taken to force the smoothness of the field, which is based on the Tiknonov regularization in vector-valued reproducing kernel Hilbert space. Under a Bayesian framework, we associate each sample with a latent variable which indicates whether it is an inlier, and then formulate the problem as maximum a posteriori problem and use Expectation Maximization algorithm to solve it. The proposed method possesses two characteristics: 1) robust to outliers, and being able to tolerate 90% outliers and even more, 2) computationally efficient. As an application, we apply VFC to solve the problem of mismatch removing. The results demonstrate that our method outperforms many state-of-the-art methods, and it is very robust. Ji Zhao 0001, Jiayi Ma 0001, Jinwen Tian, Jie Ma 0003, Dazhi Zhang |
CVPR | 2 |