VLDB 2026 Research / reviewers in the wild / expert
Shangqi Deng
dblp:311/4372 · also Shang-Qi Deng
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-6971-5114ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive 3D Convolution for Remote Sensing Image FusionabstractRemote sensing image fusion aims to create a high-resolution multi/hyper-spectral image from a high-resolution image with limited spectral information and a low-resolution image with abundant spectral data. Recently, deep learning (DL) techniques have shown significant effectiveness in this area. Most DL-based methods approach image fusion as a 2D problem by encoding spectral information into feature map channels. However, our research suggests that this strategy introduces notable spectral distortions. In contrast, some methods consider spectral data as an additional dimension, utilizing standard 3D convolutions to preserve spectral information. Nevertheless, in a standard 3D convolutional layer, the same set of kernels is applied across all input regions, which we have found to be sub-optimal for image fusion. Furthermore, standard 3D convolutions necessitate substantial computational resources. To address these challenges, we propose a novel convolutional paradigm called Adaptive 3D Convolution (Ada3D) for remote sensing image fusion. Ada3D applies a unique set of 3D kernels to each input voxel, enabling the capture of fine-grained details. These adaptive kernels are generated through a two-step process: 1) spatial and spectral kernels are derived from their respective image sources and 2) these two types of kernels are then combined to form content-aware 3D kernels that effectively integrate spatial and spectral information. Additionally, adaptive biases are introduced to enhance the convolutional outcome at the voxel level. Furthermore, we incorporate the group convolution technique to reduce computational complexity. As a result, Ada3D offers full adaptivity in an efficient manner. Evaluation results across five datasets demonstrate that our method achieves state-of-the-art (SOTA) performance, underscoring the superiority of Ada3D. The code is available at https://github.com/PSRben/Ada3D. Siran Peng, Xiangyu Zhu 0001, Shangqi Deng, Liang-Jian Deng, Zhen Lei 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image FusionabstractImplicit Neural Representation (INR) methods have demonstrated great potential in arbitrary-scale super-resolution tasks. This success is primarily due to their ability to continuously represent images using coordinates. In the task of remote sensing image fusion, INR methods have also shown promising applications. However, the previous INR methods neglect channel-wise modeling, while sharing a single kernel across all channels at each position, resulting in a lack of sensitivity to data specificity. To address these issues, we propose the OcTree Implicit Adaptive Sampling (OTIAS) method, which innovatively applies the octree structure to restore data from both horizontal and vertical directions, effectively incorporating spatial and spectral information from hyperspectral data. Additionally, we introduce a novel method to adaptively generate interpolation kernels based on coordinates. This approach efficiently produces customized interpolation kernel parameters for octree nodes, tailored to different spectral information. Overall, our method achieves state-of-the-art performance on the CAVE and Harvard datasets with 4× and 8× scaling factors, outperforming existing approaches. Shangqi Deng, Liang-Jian Deng, Ping Wei 0001 |
AAAI | 1 |
| 2025 | PanAdapter: Two-Stage Fine-Tuning with Spatial-Spectral Priors Injecting for PansharpeningabstractPansharpening is a challenging image fusion task that involves restoring images using two different modalities: low-resolution multispectral images (LRMS) and high-resolution panchromatic (PAN). Many end-to-end specialized models based on deep learning (DL) have been proposed, yet the scale and performance of these models are limited by the size of dataset. Given the superior parameter scales and feature representations of pre-trained models, they exhibit outstanding performance when transferred to downstream tasks with small datasets. Therefore, we propose an efficient fine-tuning method, namely PanAdapter, which utilizes additional advanced semantic information from pre-trained models to alleviate the issue of small-scale datasets in pansharpening tasks. Specifically, targeting the large domain discrepancy between image restoration and pansharpening tasks, the PanAdapter adopts a two-stage training strategy for progressively adapting to the downstream task. In the first stage, we fine-tune the pre-trained CNN model and extract task-specific priors at two scales by proposed Local Prior Extraction (LPE) module. In the second stage, we feed the extracted two-scale priors into two branches of cascaded adapters respectively. At each adapter, we design two parameter-efficient modules for allowing the two branches to interact and be injected into the frozen pre-trained VisionTransformer (ViT) blocks. We demonstrate that by only training the proposed LPE modules and adapters with a small number of parameters, our approach can benefit from pre-trained image restoration models and achieve state-of-the-art performance in several benchmark pansharpening datasets. RuoCheng Wu, Zien Zhang, Shangqi Deng, Yule Duan 0001, Liang-Jian Deng |
AAAI | 3 |
| 2025 | TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent Diffusion
Ziyang Ren, Ping Wei 0001, Shangqi Deng, Haowen Tang, Jiapeng Li 0003 |
ICCV | 3 |
| 2025 | Physics-informed Neural Operator for PansharpeningabstractOver the past decades, pansharpening has contributed greatly to numerous remote sensing applications, with methods evolving from theoretically grounded models to deep learning approaches and their hybrids. Though promising, existing methods rarely address pansharpening through the lens of underlying physical imaging processes. In this work, we revisit the spectral imaging mechanism and propose a novel physics‐informed neural operator framework for pansharpening, termed PINO, which faithfully models the end‐to‐end electro‐optical sensor process. Specifically, PINO operates as: (1) First, a spatial-spectral encoder pair is introduced to aggregate multi-granularity high-resolution panchromatic (PAN) and low-resolution multispectral (LRMS) features.
(2) Subsequently, an iterative neural integral process utilizes these fused spatial-spectral characteristics to learn a continuous radiance field $L_i(x, y, \lambda)$ over spatial coordinates and wavelength, effectively emulating band-wise spectral integration. (3) Finally, the learned radiance field is modulated by the sensor’s spectral responsivity $R_b(\lambda)$ to produce physically consistent spatial–spectral fusion products. This physics-grounded fusion paradigm offers a principled solution for reconstructing high-resolution multispectral and hyperspectral images in accordance with sensor imaging physics, effectively harnessing the unique advantages of spectral data to better uncover real-world characteristics. Experiments on multiple benchmark datasets show that our method surpasses state-of-the-art fusion algorithms, achieving reduced spectral aberrations and finer spatial textures. Furthermore, extension to hyperspectral (HS) data demonstrates its generalizability and universality. The code will be available upon potential acceptance. Junming Hou, Chenxu Wu, Xiaofeng Cong, Shangqi Deng, Junling Li, Liang-Jian Deng |
NeurIPS | 6 |
| 2025 | NeRI: Implicit Neural Representation for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) remains challenging due to the weak spatial features of targets and their susceptibility to background clutter. Recent studies have improved detection performance through the embedding of additional spatial representations. However, these feature prompting methods rely on discretely sampled feature spaces, which weaken high-frequency information and consequently limit their representational efficiency. To overcome this, we propose NeRI, a network that leverages the potential of implicit neural representations (INRs) through a continuous formulation to learn mappings from spatial coordinates to the high-frequency structural representations of targets. Specifically, these mappings are realized through INR Blocks (INRBs) integrated into different encoder layers, providing continuous spatial guidance from multi-scale inputs and enabling more accurate localization and distinction. In addition, to better model the distinction between foreground and background, we construct a hybrid U-shaped block (HUB) that combines a U-shaped Transformer block (UTB) and multi-scale convolution block (MCB). The UTB component effectively increases network depth and facilitates long-range dependency modeling across different scales, while the MCB employs convolutions with varying receptive fields to capture fine-grained local information, thereby enabling the two components to fully exploit their complementary strengths. Finally, we propose a simple yet effective spatial–semantic fusion (SSF) module that reweights and integrates spatial information from diverse layers to enhance the expressive power of the features. The proposed NeRI offers a robust solution for the accurate separation of targets from backgrounds. Experimental validation, conducted on three public datasets (i.e., NUDT-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrates the superior performance of NeRI compared to other methods. Open-source implementations will be available at https://github.com/Shangwei-Deng/NeRI. Shangwei Deng, Qianwen Ma, Shangqi Deng, Ziqian Chen, Ruoqi Lian, Bincheng Li, Kepeng Xu, Xiaobo Li 0004, Haofeng Hu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | IOVarNet: Inner-Outer Variation Synergy Network for Infrared Small Target DetectionabstractSparsity and weak characteristics of targets pose significant challenges in infrared small target detection (IRSTD). For convolutional neural network-based methods, the increase in semantic information during propagation is often accompanied by the degradation of spatial features, which hampers the performance of IRSTD. In this paper, we proposed the Inner-outer Variation Synergy Network (IOVarNet) for IRSTD, which explicitly enhances the spatial response of targets by reinforcing their structural representations across different layers of the network. Specifically, IOVarNet leverages the delicate Total Variation-inspired Module, which takes the form of a partial differential equation, and incorporates it through the Inner-outer Variational Synergy architecture to supplement the target’s structural information at both the inner and outer layers of the encoder and decoder. Besides, the Dual Space Attention mechanism was introduced to enhance the semantic distinction between the target and background, while fusing spatial features from different layers. Experimental validation was conducted on three public datasets (i.e., NUDT-SIRST, NUAA-SIRST, and IRSTD-1K), demonstrating the performance superiority of IOVarNet over other methods. Open-source implementations will be available at https://github.com/Shangwei-Deng/IOVarNet. Shangwei Deng, Qianwen Ma, Bincheng Li, Liaoran Jin, Kepeng Xu, Shangqi Deng, Xiaobo Li 0004, Haofeng Hu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Mamba Collaborative Implicit Neural Representation for Hyperspectral and Multispectral Remote Sensing Image FusionabstractHyperspectral remote sensing images (HSIs) capture detailed spectral characteristics of features, while multispectral remote sensing images (MSIs) provide clear spatial distribution. Fusing these two types of images can enhance feature identification and classification accuracy. Current deep learning algorithms achieve high fusion quality but struggle with balancing global effective perception and lightweight computation. Moreover, these algorithms typically discretely handle data mapping, which contrasts with the continuous nature of the world. Recently, the Mamba has shown significant potential for complex long-range modeling, addressing the computational complexity of global perception. Concurrently, implicit neural representation (INR) offers high-quality solutions for continuous domain modeling. To this end, this study introduces a novel network architecture that combines Mamba and INR, termed the Mamba cooperative INR fusion network (MCIFNet). MCIFNet effectively captures global image information and generates fused images in a continuous domain through point-to-point processing. The network comprises two main units: potential space projection and semantic extraction and fusion. The potential space projection unit performs shallow encoding of hyperspectral and MSIs, mapping them to a latent feature space. The semantic extraction and fusion unit (SEFU) uses scale adaptive residual state spatial and implicit spatial-spectral fusion (ISSF) modules to extract deep features from the bimodal images, generating fused images point-by-point. A series of fusion experiments with$4\times $,$8\times $, and$16\times $scale factors demonstrate that MCIFNet surpasses popular algorithms in both spatial detail and spectral information reconstruction, while also providing more lightweight performance. The code for MCIFNet will be shared onhttps://github.com/chunyuzhu/MCIFNet. Chunyu Zhu, Shangqi Deng, Xuan Song 0002, Yachao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Fourier-enhanced Implicit Neural Fusion Network for Multispectral and Hyperspectral Image FusionabstractRecently, implicit neural representations (INR) have made significant strides in various vision-related domains, providing a novel solution for Multispectral and Hyperspectral Image Fusion (MHIF) tasks. However, INR is prone to losing high-frequency information and is confined to the lack of global perceptual capabilities. To address these issues, this paper introduces a Fourier-enhanced Implicit Neural Fusion Network (FeINFN) specifically designed for MHIF task, targeting the following phenomena: The Fourier amplitudes of the HR-HSI latent code and LR-HSI are remarkably similar; however, their phases exhibit different patterns. In FeINFN, we innovatively propose a spatial and frequency implicit fusion function (Spa-Fre IFF), helping INR capture high-frequency information and expanding the receptive field. Besides, a new decoder employing a complex Gabor wavelet activation function, called Spatial-Frequency Interactive Decoder (SFID), is invented to enhance the interaction of INR features. Especially, we further theoretically prove that the Gabor wavelet activation possesses a time-frequency tightness property that favors learning the optimal bandwidths in the decoder. Experiments on two benchmark MHIF datasets verify the state-of-the-art (SOTA) performance of the proposed method, both visually and quantitatively. Also, ablation studies demonstrate the mentioned contributions. The code can be available at https://github.com/294coder/Efficient-MIF. Yu-Jie Liang, Zihan Cao, Shangqi Deng, Hong-Xia Dou, Liang-Jian Deng |
NeurIPS | 3 |
| 2023 | Bidirectional Dilation Transformer for Multispectral and Hyperspectral Image FusionabstractTransformer-based methods have proven to be effective in achieving long-distance modeling, capturing the spatial and spectral information, and exhibiting strong inductive bias in various computer vision tasks. Generally, the Transformer model includes two common modes of multi-head self-attention (MSA): spatial MSA (Spa-MSA) and spectral MSA (Spe-MSA). However, Spa-MSA is computationally efficient but limits the global spatial response within a local window. On the other hand, Spe-MSA can calculate channel self-attention to accommodate high-resolution images, but it disregards the crucial local information that is essential for low-level vision tasks. In this study, we propose a bidirectional dilation Transformer (BDT) for multispectral and hyperspectral image fusion (MHIF), which aims to leverage the advantages of both MSA and the latent multiscale information specific to MHIF tasks. The BDT consists of two designed modules: the dilation Spa-MSA (D-Spa), which dynamically expands the spatial receptive field through a given hollow strategy, and the grouped Spe-MSA (G-Spe), which extracts latent features within the feature map and learns local data behavior. Additionally, to fully exploit the multiscale information from both inputs with different spatial resolutions, we employ a bidirectional hierarchy strategy in the BDT, resulting in improved performance. Finally, extensive experiments on two commonly used datasets, CAVE and Harvard, demonstrate the superiority of BDT both visually and quantitatively. Furthermore, the related code will be available at the GitHub page of the authors. Shangqi Deng, Liang-Jian Deng, Ran Ran 0001 |
IJCAI | 1 |
| 2023 | PSRT: Pyramid Shuffle-and-Reshuffle Transformer for Multispectral and Hyperspectral Image FusionabstractA Transformer has received a lot of attention in computer vision. Because of global self-attention, the computational complexity of Transformer is quadratic with the number of tokens, leading to limitations for practical applications. Hence, the computational complexity issue can be efficiently resolved by computing the self-attention in groups of smaller fixed-size windows. In this article, we propose a novel pyramid Shuffle-and-Reshuffle Transformer (PSRT) for the task of multispectral and hyperspectral image fusion (MHIF). Considering the strong correlation among different patches in remote sensing images and complementary information among patches with high similarity, we design Shuffle-and-Reshuffle (SaR) modules to consider the information interaction among global patches in an efficient manner. Besides, using pyramid structures based on window self-attention, the detail extraction is supported. Extensive experiments on four widely used benchmark datasets demonstrate the superiority of the proposed PSRT with a few parameters compared with several state-of-the-art approaches. The related code is available athttps://github.com/Deng-shangqi/PSRT. Shangqi Deng, Liang-Jian Deng, Ran Ran 0001, Danfeng Hong, Gemine Vivone |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | QIS-GAN: A Lightweight Adversarial Network With Quadtree Implicit Sampling for Multispectral and Hyperspectral Image FusionabstractMultispectral and Hyperspectral Image Fusion (MHIF) involves the fusion of high spatial resolution multispectral images (HR-MSI) and low spatial resolution hyperspectral images (LR-HSI) to generate high spatial resolution hyperspectral images (HR-HSI), has gained significant attention in the field of remote sensing imaging. While CNN and Transformer models have shown effectiveness in MHIF, existing CNN or Transformer-based algorithms are overburdened with model size, making it difficult to achieve an effective trade-off between fusion accuracy and degree of lightweight. Recently, Implicit Neural Representation (INR) has been proven good interpretability and the ability to exploit coordinate information in 2D tasks. Nonetheless, INR-based fusion networks have certain limitations, such as the need for deeper super-resolution networks as shallow encoders, and insufficient representation capability on high upsampling ratios. To address these challenges, we present the Quadtree Implicit Sampling (QIS), which employs a hierarchical sampling from the perspective of the quadtree, to enhance the capacity of the overall network. Furthermore, the remarkable design of QIS allows us to adopt a lightweight structure as the shallow encoder, greatly alleviating the network burden and achieving lightweight. Inspired by generative adversarial models, we incorporate QIS as a lightweight generator into the GAN framework named QIS-GAN and leverage a discriminator to increase the fidelity of fused images. The results showcase the superior performance of QIS-GAN on the MHIF tasks with upsampling ratios of ×4, ×8, and ×16, surpassing the state-of-the-art in several datasets. The code for our approach will be available at https://github.com/chunyuzhu/QIS-GAN. Chunyu Zhu, Shangqi Deng, Yingjie Zhou 0001, Liang-Jian Deng |
IEEE Trans. Geosci. Remote. Sens. | 2 |