VLDB 2026 Research / reviewers in the wild / expert
Jinyang Liu 0004
dblp:36/1731-4
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-0429-7422ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Selective Re-learning Mechanism for Hyperspectral Fusion ImagingabstractHyperspectral fusion imaging is challenged by high computational cost due to the abundant spectral information. We find that pixels in regions with smooth spatial-spectral structure can be reconstructed well using a shallow network, while only those in regions with complex spatial-spectral structure require a deeper network. However, existing methods process all pixels uniformly, which ignores this property. To leverage this property, we propose a Selective Re-Learning Fusion Network (SRLF) that initially extracts features from all pixels uniformly and then selectively refines distorted feature points. Specifically, SRLF first employs a Preliminary Fusion Module with robust global modeling capability to generate a preliminary fusion feature. Afterward, it applies a Selective Re-Learning Module to focus on improving distorted feature points in the preliminary fusion feature. To achieve targeted learning, we present a novel Spatial-Spectral Structure-Guided Selective Re-Learning Mechanism (SSG-SRL) that integrates the observation model to identify the feature points with spatial or spectral distortions. Only these distorted points are sent to the corresponding re-learning blocks, reducing both computational cost and the risk of overfitting. Finally, we develop an SRLF-Net, composed of multiple cascaded SRLFs, which surpasses multiple state-of-the-art methods on several datasets with minimal computational cost. Yuanye Liu, Jinyang Liu 0004, Renwei Dian, Shutao Li 0001 |
CVPR | 2 |
| 2025 | Asymptotic Spectral Mapping for Hyperspectral Image FusionabstractThe fusion of low-resolution hyperspectral images (LR HSI) and high-resolution multispectral images (HR MSI) is a crucial approach for generating hyperspectral images (HSI). However, existing hyperspectral image fusion methods often rely on a single feature mapping process, which makes it difficult to accommodate the significant differences in features between the source images and the real images. Consequently, the generated images frequently exhibit varying degrees of information loss across different spectral bands and limit the overall performance of the fusion. To address this issue, we propose a novel hyperspectral image fusion network. Specifically, an asymptotic spectral mapping module is designed to enhance the detail information fitting capabilities of the fusion network. This module transforms the fitting process of missing information into multiple sets of fitting processes with varying degrees, which can map features with different spectral fidelity to various scales and gradually fit spectral information, thereby reducing spectral distortion. Additionally, we introduce an adaptive defect optimization loss that guides the network to focus on reconstructing regions with substantial spectral differences between LR HSI and HSI, optimizing the network’s constraints regarding the similarity between predicted and real images. Experimental results demonstrate that the proposed fusion network outperforms existing state-of-the-art methods across diverse datasets. Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Lishan Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Heterospectral Structure Compensation Sampling for Hyperspectral Fusion Computational ImagingabstractExisting hyperspectral fusion computational imaging methods primarily rely on using high-resolution multispectral images (HRMSI) to provide spatial details for low-resolution hyperspectral images (LRHSI), thereby enabling the reconstruction of hyperspectral images. However, these methods are often limited by the low spectral resolution of the HRMSI, making the sampled tensors unable to provide effective information for the LRHSI in a finer spectral range. To achieve more accurate computational imaging results, we propose a Heterospectral Structure Compensation Sampling (HSC-sampling) mechanism. Unlike traditional spatial sampling methods, which directly calculate the interpolation between adjacent pixels, this mechanism analyzes the structural complementarity among different bands in LRHSI. It utilizes the information from other bands to compensate for the missing details in the current band. Additionally, a novel Multi-phase Mixed Modeling (M2M) approach is designed, expanding the model's analytical capabilities into multiple phases to accommodate the high-dimensional nature of HSI data. Specifically, it extracts fusion features from three phases and organizes the generated features along with the input features into a multi-variate mixed cube based on phase relationships, thereby capturing feature correlations across different phases. Based on the HSC-sampling mechanism and the M2M approach, we construct a Merging Residual Concatenation (MRC) hyperspectral fusion computational imaging network. Compared to other state-of-the-art methods, this network achieves significant improvements in fusion performance across multiple datasets. Moreover, the effectiveness of the HSC-sampling mechanism has been demonstrated in various hyperspectral imaging tasks. Code is available at: https://github.com/1318133/HSC-Sampling. Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Yuanye Liu |
IEEE Trans. Image Process. | 1 |
| 2025 | Continuous Feature Representation for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to discover objects that are seamlessly embedded in the environment. Existing COD methods have made significant progress by typically representing features in a discrete way with arrays of pixels. However, limited by discrete representation, these methods need to align features of different scales during decoding, which causes some subtle discriminative clues to become blurred. This is a huge blow to the task of identifying camouflaged objects from clear subtle clues. To address this issue, we propose a novel continuous feature representation network (CFRN), which aims to represent features of different scales as a continuous function for COD. Specifically, a Swin transformer encoder is first exploited to explore the global context between camouflaged objects and the background. Then, an object-focusing module (OFM) deployed layer by layer is designed to deeply mine subtle discriminative clues, thereby highlighting the body of camouflaged objects and suppressing other distracting objects at different scales. Finally, a novel frequency-based implicit feature decoder (FIFD) is proposed, which directly decodes the predictions at arbitrary coordinates in the continuous function with implicit neural representations, thus propagating clearer discriminative clues. Extensive experiments on four challenging COD benchmarks demonstrate that our method significantly outperforms state-of-the-art methods. The source code will be available at https://github.com/SongZeHNU/CFRN. Xudong Kang, Xiaohui Wei 0001, Jinyang Liu 0004, Zheng Lin 0005, Shutao Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Mosaic Pattern Excavation Transformer for Spectral ImagingabstractSingle spectral image demosaicing for multispectral filter array (MSFA) is an essential task in spectral imaging, aiming to recover a mosaic-free spectral image from its mosaic raw counterpart. Existing deep learning-based methods typically improve the reconstruction performance by indiscriminately stacking CNN-based blocks, failing to effectively handle the intertwined spatio-spectral correlations caused by spatial sub-sampling and spectral aliasing. In this paper, we propose Mosaic Pattern Excavation Transformer (MPEFormer) to achieve better reconstruction by effectively modelling the intertwined spatio-spectral correlations. Specifically, the proposed three-branch model integrates low-frequency information, edge information, and fine high-frequency details essential for spectral image reconstruction, with the third branch serving as the core component. In this branch, we design the Dual Fusion Self-attention Block (DFSAB) and the Mosaic Pattern-guided Spectral Modulation Module (MPSM). DFSAB incorporates the Mosaic Pattern Excavation Self-attention (MPESA) mechanism, which effectively captures non-local spatio-spectral correlations induced by the MSFA pattern distributed across the whole image, thereby enhancing the expressive capability of the model. By dynamically integrating various MSFA pattern-related dependencies, MPSM enables adaptive recalibration of spectral information. Extensive experimental results demonstrate the effectiveness of our MPEFormer, highlighting its greater potential over the state-of-the-art MSFA demosaicing methods. The code will be uploaded at https://github.com/Matsuri247/MPEFormer. Yaohang Wu, Jinyang Liu 0004, Renwei Dian, Shutao Li 0001, Yining Yang |
IEEE Trans. Image Process. | 2 |
| 2025 | Multi-Granularity Context Perception Network for Open Set Recognition of Camouflaged ObjectsabstractOpen set recognition (OSR) aims to identify whether a test sample belongs to a semantic class in the classifier training set. Existing OSR methods exhibit prominent performance on various image datasets. However, they are primarily designed for general object recognition rather than more complex camouflaged object recognition. When an object is camouflaged, i.e., it exhibits a similar pattern to the background, it is difficult to finely identify it and differentiate between known and unknown categories. To address this problem, we propose a novel multi-granularity context perception network (MCPNet) for OSR of camouflaged objects, which can accurately identify camouflaged objects by fusing coarse-grained and fine-grained context features. In MCPNet, the vision transformer is first utilized to extract coarse-grained context features to locate the approximate location of camouflaged objects. Then, an adaptive local focus module (ALFM) is proposed to pick out the most discriminative regions and learn the fine-grained context of these regions. Finally, multi-granular context features are fused to obtain recognition results. During the training, a contrastive clustering module (CCM) is introduced to guide the network to effectively utilize multi-granularity context to generate high-confidence decision boundaries. We also built two camouflaged object classification datasets named ACOC and NCOC which mainly consist of artificial camouflage and natural camouflage respectively to facilitate research in OSR of camouflaged objects. Experimental results on two datasets show that MCPNet outperforms state-of-the art methods. Xudong Kang, Xiaohui Wei 0001, Renwei Dian, Jinyang Liu 0004, Shutao Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Denoiser Learning for Infrared and Visible Image FusionabstractInfrared image (IR) and visible image (VI) fusion creates fusion images that contain richer information and gain improved visual effects. Existing methods generally use the operators of manual design, such as intensity and gradient operators, to mine the image information. However, it is hard for them to achieve a complete and accurate description of information, which limits the image fusion performance. To this end, a novel information measurement method is proposed to achieve IR and VI fusion. Its core idea is to guide a generator in achieving image fusion by learning the denoisers. Specifically, by using denoisers to restore fusion images with different noise interference to source images, a mutual competition relationship is formed between denoisers, which helps the generator thoroughly explore the data specificity of the source images and guide it to achieve more accurate feature representation. In addition, a semantic adaptive measurement loss function is proposed to constrain the generator, which fuses semantic information adaptively by considering the semantic information density of different source images. The results of quantitative and qualitative experiments have shown that the proposed method can achieve a higher quality information fusion and has a faster fusion speed on three public datasets when compared with advanced methods. Jinyang Liu 0004, Shutao Li 0001, Lishan Tan, Renwei Dian |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MDENet: Multidomain Differential Excavating Network for Remote Sensing Image Change DetectionabstractRemote sensing image change detection can analyze alterations on the Earth’s surface within a specific region. However, the accuracy of change detection has consistently been hindered by the style differences in captured images caused by seasonal or lighting variations, as well as the challenge of distinguishing similar features between the background and foreground in the scene. To this end, a multidomain differential excavating network (MDENet) for change detection is introduced. Using the novel multidomain differential collaboration module (MDCM) to precisely capture object features on the frequency and spatial domains across diverse temporal domains, it enables simultaneous querying of global and local change information. Moreover, the multineighborhood frequency gate attention (MFGatt) is devised to eliminate the impact of image style relevance information and consolidate attention toward object localization, thereby enhancing the adaptability of the network to variations in image style. Extensive experiments have illustrated that our proposed network achieves better detection accuracy compared with current state-of-the-art (SOTA) methods on various datasets. Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Xudong Kang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Focus Relationship Perception for Unsupervised Multi-Focus Image FusionabstractMulti-focus image fusion can extract the focus regions from different source images and combine them into a fully clear image. Existing unsupervised methods typically use gradient information to measure the focus regions in images and generate a fusion weight map, but ordinary gradient operators are difficult to measure information accurately in regions with weaker textures. In addition, using only gradient information as a constraint cannot make the model fully distinguish all the focus regions in the image, which seriously restricts the clarity of the fusion image. To address these issues, a novel unsupervised multi-focus image fusion method is proposed in this paper. Specifically, a neighborhood information fusion network is designed to generate an initial fusion weight map. It can capture features within different neighborhood ranges at once, which enhances the information association between different regions. In addition, to further improve the feature extraction ability of the model in the regions with low texture information, a local difference evaluation loss function is proposed. It is combined with the gradient measure loss function to constrain the network. Finally, a fusion weight optimization module is proposed to improve the clarity of the fusion image in the repeated defocusing regions and overexposed regions of different source images, which redistributes the weights of different source images. The proposed fusion method is compared with advanced methods on three public multi-focus datasets. Experimental results indicate that the proposed method has achieved better performance in qualitative and quantitative aspects. Jinyang Liu 0004, Shutao Li 0001, Renwei Dian |
IEEE Trans. Multim. | 1 |
| 2024 | A Lightweight Pixel-Level Unified Image Fusion NetworkabstractIn recent years, deep-learning-based pixel-level unified image fusion methods have received more and more attention due to their practicality and robustness. However, they usually require a complex network to achieve more effective fusion, leading to high computational cost. To achieve more efficient and accurate image fusion, a lightweight pixel-level unified image fusion (L-PUIF) network is proposed. Specifically, the information refinement and measurement process are used to extract the gradient and intensity information and enhance the feature extraction capability of the network. In addition, these information are converted into weights to guide the loss function adaptively. Thus, more effective image fusion can be achieved while ensuring the lightweight of the network. Extensive experiments have been conducted on four public image fusion datasets across multimodal fusion, multifocus fusion, and multiexposure fusion. Experimental results show that L-PUIF can achieve better fusion efficiency and has a greater visual effect compared with state-of-the-art methods. In addition, the practicability of L-PUIF in high-level computer vision tasks, i.e., object detection and image segmentation, has been verified. Jinyang Liu 0004, Shutao Li 0001, Renwei Dian, Xiaohui Wei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Toward Efficient Remote Sensing Image Change Detection via Cross-Temporal Context LearningabstractChange detection (CD) aims to find areas of specific changes in multi-temporal remote sensing images. The existing methods fail to adequately explore the cross-temporal global context, making the establishment of spatial-temporal deep global associations insufficient and inefficient. As a result, their performance is vulnerable to complex and various objects in changing scenes. Hence, we propose a cross-temporal context learning network, termed as CCLNet, where the intra- and inter-temporal long-range dependency are mined and interactively fused, to fully exploit the cross-temporal context information. Specifically, a lightweight convolutional neural network is first used to extract deep semantic features. Then, a well-designed cross-temporal fusion transformer (CFT) is proposed to locate the changing objects in the scene by establishing the long-range dependency across bitemporal images. Thanks to this, the temporal-specific information extraction and cross-temporal information integration are seamlessly integrated into the same network, thereby significantly improving the discriminative features of changing objects. Furthermore, this allows us using naive backbones with low computational cost to achieve reliable CD performance. Experiments on mainstream benchmarks show that our proposed method can handle CD task faster than state-of-the-art methods while maintaining better or comparable matching accuracy on a single RTX3090. Xiaohui Wei 0001, Xudong Kang, Shutao Li 0001, Jinyang Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |