Jie Guo 0009

dblp:77/2751-9 · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0002-6223-5492ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Artificial intelligence and machine learning · 15 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 TF-SNN: Temporal focus-based dynamic neuron regulation framework for spiking neural networks
Jie Guo 0009, Junxiang Wu, Mingjin Zhang, Yunsong Li 0001
Expert Syst. Appl.2
2026 Diff-Transformer: Heterogeneous Feature Fusion Network for Multisource Remote Sensing Classification
abstract
Multimodal remote sensing image classification has emerged as a key research area in remote sensing, with extensive applications in real-world scenarios. However, these images are collected by different sensors and contain multiple features such as spectrum, space, height and texture. Due to the differences in the characteristics of these data, existing methods have poor results in extracting and fusing heterogeneous features, which limits the improvement of classification performance. To address this problem, we propose a new heterogeneous feature extraction and fusion framework DTFNet, which utilizes the diffusion model and Transformer architecture. In the feature extraction stage, different networks are constructed to extract heterogeneous features while reducing redundancy. The dual-branch diffusion feature extraction (DBDFE) network based on the diffusion model is introduced to process data from different sensors, avoiding the limitation of extracting all features with a single network. In the feature fusion stage, the extracted diffusion features are fused with the original features to preserve the integrity of the original data. The cross-fusion transformer (CFT) module uses a convolutional neural network (CNN) to complete the local feature transformation and integration and models the long-range dependencies between heterogeneous features through cross-transformer encoders. Experimental results show that the classification accuracy of DTFNet on the three datasets reaches 92.38%, 80.08% and 95.02% respectively, which is significantly better than the existing state-of-the-art methods, demonstrating its effectiveness and superiority.
Zhihao Ying, Jie Guo 0009, Yunsong Li 0001, Yu'e Gao
IEEE Trans. Circuits Syst. Video Technol.2
2026 Snapshot Compressive Imaging via Degradation Cue and Spectral Latent Diffusion
abstract
The goal of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image from a 2D measurement. However, current reconstruction methods still face significant challenges in fully leveraging degradation and image prior. Many methods estimate degradation solely from a single measurement rather than learning from the real imaging process, resulting in inaccurate prior modeling. Moreover, the high compression of the CASSI measurement leads to the loss of spectral-spatial context, and the existing priors fail to fully capture it - for instance, in complex scenarios (such as S5, S9 in Table I), the performance gap can be as high as 3 dB. To address these issues, this paper introduces a novel reconstruction method with Degradation Cue Learning and Spectral Latent Diffusion (DCL-SLD), which comprises two key components: the Degradation Cue Learning (DCL) module and the Spectral Latent Diffusion (SLD) module. In the spatial domain, the DCL module employs a pre-trained image encoder and a feature distribution transmission strategy to extract degraded information and integrate it into the feature, enabling reconstruction through learned visual context. In the spectral domain, the SLD module leverages a latent diffusion model based on spectral correlations to generate a low-rank vector representation, effectively preserving contextual relationships within the high-dimensional structure. By enhancing priors in both dimensions, the model significantly improves its ability to exploit contextual information for more accurate recovery. Extensive experimental results on both simulation and real datasets demonstrate the superior performance of DCL-SLD over state-of-the-art methods.
Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001
IEEE Trans. Image Process.3
2026 Degradation-Adaptive Denoising: Aligning Diffusion Models With Physics of Video Snapshot Compressive Imaging
abstract
Video Snapshot Compressive Imaging (SCI) captures multiple video frames in a single exposure, enabling efficient reconstruction of high-speed scenes for motion analysis and event detection. Existing SCI in coded aperture compressive temporal imaging (CACTI) methods predominantly rely on feedforward deep networks with fixed denoising strategies. However, they lack alignment with the SCI physical inverse model and struggle to balance motion detail recovery and static background smoothing. In this paper, we propose PCD-Diffusion for Video SCI, the first diffusion-based reconstruction framework for Video SCI, which reformulates the inverse problem as a progressive denoising process. Specifically, we design a Physically-Constrained Dynamic Diffusion (PCD-Diffusion) model, introducing a region-adaptive diffusion schedule and spatiotemporal residual estimation. This method explicitly aligns the denoising process with SCI's spatially non-uniform and temporally evolving residual distribution. Additionally, a motion prior-guided diffusion schedule and a Gauss-guided spatiotemporal adaptive residual estimation dynamically steer the denoising trajectory, ensuring accurate motion detail restoration and physically consistent reconstructions. Extensive results on simulated and real datasets verify the superior reconstruction fidelity and temporal coherence of the proposed PCD-Diffusion framework over existing approaches. Code will be released upon publication.
Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001
IEEE Trans. Image Process.3
2025 IRMamba: Pixel Difference Mamba with Layer Restoration for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) focuses on identifying small targets in infrared images. Despite advancements with deep learning, challenges persist due to the IR long-range imaging mechanism, where targets are small, dim, and easily lost in noise and background clutter. Current deep learning methods struggle to suppress noise and background interference while preserving fine details, leading to missed detections and false alarms. To address these issues, we propose IRMamba, an encoder-decoder architecture featuring Pixel Difference Mamba (PDMamba) and a Layer Restoration Module (LRM). Specifically, PDMamba integrates the intensity and directional information of pixel differences between scanning positions and their central neighborhoods into the state equation of the state space model (SSM). This enhances target detail representation and suppresses background interference by capturing local 2D dependencies from a global perspective. In addition, LRM incorporates the double-depth image prior into the iterative convergence algorithm, and utilizes the inter-layer interrelationships to gradually reverse the separation of the target layer, achieving noise suppression and refined reconstruction of the image mask. Experiments conducted on multiple public datasets, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1K, demonstrate the significant advantages of IRMamba over SOTA methods.
Mingjin Zhang, Fei Gao 0006, Jie Guo 0009
AAAI4
2025 MOCID: Motion Context and Displacement Information Learning for Moving Infrared Small Target Detection
abstract
In the field of Moving Infrared Small Target Detection (MIRSTD), current methods typically use sequential modeling with two individual modules for spatial and temporal processing. However, such a modeling strategy lacks clear guidance on the motion and displacement difference between moving targets and background noise, thereby limiting the feature discriminability and resulting in error-prone target localization. This paper addresses this issue from clip and frame levels and proposes a novel architecture MOCID for MIRSTD. For clip-level feature fusion, we design a spatio-temporal backbone consisting of several proposed Fourier-inspired Spatio-temporal Attention (FISTA) layers. Each FISTA layer sequentially processes the features from spatial and temporal views to capture clip-level temporal motion context, where Fourier Transformation and Inverse Fourier Transformation are employed for each view. This context is then embedded into dynamic convolutional kernels for subsequent spatial feature extraction, thereby enabling clear motion difference guidance and generating comprehensive features. For frame-level feature fusion, we design a Displacement-aware Mamba Module (DAM) to capture detailed frame-to-frame displacement information. DAM utilizes an innovative Temporal Interpolation and Displacement-aware Scan technique to perform spatio-temporal difference-aware displacement modeling, introducing elaborate temporal indicators into feature extraction. Combining the above improvements, our model captures comprehensive motion and displacement contexts, significantly improving the detection of the small target. Extensive experiments demonstrate that MOCID achieves state-of-the-art detection accuracy on popular IRDST and DAUB datasets. Furthermore, MOCID offers a superior balance between throughput and performance compared to other methods. The code for this work will be made publicly available.
Mingjin Zhang, Yuanjun Ouyang, Fei Gao 0006, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037
AAAI4
2025 SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining
abstract
Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully capture contextual information, leading to limited detection accuracy and adaptability in complex environments. Inspired by vision-language models, this paper proposes a novel framework, SAIST, which integrates textual information with image modalities to enhance IRSTD performance. The framework consists of two main components: Scene Recognition Contrastive Language-Image Pretraining (SR-CLIP) and CLIP-guided Segment Anything Model (CG-SAM). SR-CLIP generates a set of visual descriptions through object-object similarity and object-scene relevance, embedding them into learnable prompts to refine the textual description set. This reduces the domain gap between vision and language, generating precise textual and visual prompts. CG-SAM utilizes the prompts generated by SR-CLIP to accurately guide the Mask Decoder in learning prior knowledge of background features, while incorporating infrared imaging equations to improve small target recognition in complex backgrounds and significantly reduce the false alarm rate. Additionally, this paper introduces the first multimodal IRSTD dataset, MIRSTD, which contains abundant image-text pairs. Experimental results demonstrate that the proposed SAIST method outperforms existing state-of-the-art approaches.
Mingjin Zhang, Fei Gao 0006, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037
CVPR4
2025 Multimodal Prior Learning with Double Constraint Alignment for Snapshot Spectral Compressive Imaging
abstract
The objective of snapshot spectral compressive imaging reconstruction is to recover the 3D hyperspectral image (HSI) from a 2D measurement. Existing methods either focus on network architecture design or simply introduce image-level prior to the model. However, these methods lack guiding information for accurate reconstruction. Recognizing that textual description contain rich semantic information that can significantly enhance details, this paper introduces a novel framework, CAMM, which integrates text information into the model to improve the performance. The framework comprises two key components: Fine-grained Alignment Module (FAM) and Multimodal Fusion Mamba (MFM). Specifically, FAM is used to reduce the knowledge gap between the RGB domain obtained by the pre-trained vision-language model and the HSI domain. Through the double constraints of distribution similarity and entropy, the adaptive alignment of different complexity features is realized, which makes the encoded features more accurate. MFM aims to identify the guiding effect of RGB features and text features on HSI in space and channel dimensions. Instead of fusing features directly, it integrates prior at image-level and text-level prior into Mamba's state-space equation, so that each scanning step can be accurately guided. This kind of positive feedback adjustment ensures the authenticity of the guiding information. To our knowledge, this is the first text-guided model for compressive spectral imaging. Extensive experimental results the public datasets demonstrate the superior performance of CAMM, validating the effectiveness of our proposed method.
Mingjin Zhang, Longyi Li, Fei Gao 0006, Qiming Zhang 0001, Jie Guo 0009
IJCAI5
2025 WMRNet: Wavelet Mamba With Reversible Structure for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) is of great practical significance in many real-world applications, such as maritime rescue and early warning systems, benefiting from the unique and excellent infrared imaging ability in adverse weather and low-light conditions. Nevertheless, segmenting small targets from the background remains a challenge. When the subsampling frequency during image processing does not satisfy the Nyquist criterion, the aliasing effect occurs, which makes it extremely difficult to identify small targets. To address this challenge, we propose a novel Wavelet Mamba with Reversible Structure Network (WMRNet) for infrared small target detection in this paper. Specifically, WMRNet consists of a Discrete Wavelet Mamba (DW-Mamba) module and a Third-order Difference Equation guided Reversible (TDE-Rev) structure. DW-Mamba employs the Discrete Wavelet Transform to decompose images into multiple subbands, integrating this information into the state equations of a state space model. This method minimizes frequency interference while preserving a global perspective, thereby effectively reducing background aliasing. The TDE-Rev aims to suppress edge aliasing effects by refining the target edges, which first processes features with an explicit neural structure derived from the second-order difference equations and then promotes feature interactions through a reversible structure. Extensive experiments on the public IRSTD-1k and SIRST datasets demonstrate that the proposed WMRNet outperforms the state-of-the-art methods.
Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 MDEformer: Mixed Difference Equation Inspired Transformer for Compressed Video Quality Enhancement
abstract
Deep learning methods have achieved impressive performance in compressed video quality enhancement tasks. However, these methods rely excessively on practical experience by manually designing the network structure and do not fully exploit the potential of the feature information contained in the video sequences, i.e., not taking full advantage of the multiscale similarity of the compressed artifact information and not seriously considering the impact of the partition boundaries in the compressed video on the overall video quality. In this article, we propose a novel Mixed Difference Equation inspired Transformer (MDEformer) for compressed video quality enhancement, which provides a relatively reliable principle to guide the network design and yields a new insight into the interpretable transformer. Specifically, drawing on the graphical concept of the mixed difference equation (MDE), we utilize multiple cross-layer cross-attention aggregation (CCA) modules to establish long-range dependencies between encoders and decoders of the transformer, where partition boundary smoothing (PBS) modules are inserted as feedforward networks. The CCA module can make full use of the multiscale similarity of compression artifacts to effectively remove compression artifacts, and recover the texture and detail information of the frame. The PBS module leverages the sensitivity of smoothing convolution to partition boundaries to eliminate the impact of partition boundaries on the quality of compressed video and improve its overall quality, while not having too much impacts on non-boundary pixels. Extensive experiments on the MFQE 2.0 dataset demonstrate that the proposed MDEformer can eliminate compression artifacts for improving the quality of the compressed video, and surpasses the state-of-the-arts (SOTAs) in terms of both objective metrics and visual quality.
Mingjin Zhang, Haichen Bai, Wenteng Shang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 IRPruneDeXt: Efficient Infrared Small Target Detection via Musical Wavelet-Regularized Channel Pruning
abstract
Infrared small target detection (IRSTD) refers to detecting faint targets in infrared (IR) images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and computation inefficiencies. In this pioneering study, we introduce the concept of utilizing network pruning to enhance the efficiency of IRSTD. Due to the challenge posed by low signal-to-noise ratios (SNRs) and the absence of detailed semantic information in IR images, directly applying existing pruning techniques yields suboptimal performance. To address this, we propose a novel wavelet structure-regularized multidimensional musical scale soft channel pruning (SCP) method, giving rise to the efficient IRPruneDeXt model. Our approach involves representing the weight matrix in the wavelet domain and formulating a wavelet channel pruning (WCP) strategy. We incorporate wavelet regularization to induce structural sparsity without incurring extra memory usage. Additionally, we design a multidimensional musical scale soft channel reconstruction (MMSCR) method that adapts the strategy across temporal and spatial dimensions to preserve key target information and prevent premature pruning. By leveraging interactions between criteria, it balances pruning and reconstruction through a musical scale feedback effect, achieving an optimal sparse structure while maintaining overall sparsity. Through extensive experiments on many widely used benchmarks, our IRPruneDeXt method surpasses established techniques in both model complexity and accuracy. Specifically, when employing U-net as the baseline network, IRPruneDeXt achieves a 65.68% reduction in parameters and a 51.77% decrease in floating-point operations (FLOPs) while improving intersection over union (IoU) from 73.31% to 76.17% and normalized IoU (nIoU) from 70.92% to 75.08%. The code is available at github.com/hd0013/IRPruneDet.
Mingjin Zhang, Jin Feng, Handi Yang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Computational Fluid Dynamic Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) aims to identify and locate small targets amidst background noise. It is highly valuable in various practical application domains, such as maritime rescue and early warning systems deployed in challenging conditions such as harsh weather, low illumination, and long imaging distances. Different from existing works that either adopt well-designed backbone networks or devise specific modules to improve them from different aspects, in this article, we formulate the learning process of IRSTD from a novel perspective, i.e., the mechanism of pixel movement. Considering that the movement of pixels passing through the layers of the network for IRSTD can be analogized to the flow of particles in a fluid dynamic system, we propose a computational fluid dynamic network (CFD-Net) derived from computational fluid dynamics. Technically, we leverage the superiority of the unilateral difference equation with third-order accuracy and devise a unilateral differential residual structure as the backbone of CFD-Net. This design ensures that the pixel stream only flows in the forward direction. In addition, a switch-controlled multidirectional treatment tank (SMTT) is introduced to CFD-Net to dynamically guide the pixel stream to the appropriate path for different targets with varying shapes and orientations, facilitating learning robust target representation and improving detection performance. The proposed CFD-Net is evaluated on the IRSTD-1k and SIRST datasets and is found to outperform existing state-of-the-art (SOTA) methods.
Mingjin Zhang, Ke Yue, Jie Guo 0009, Qiming Zhang 0001, Jing Zhang 0037, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel Pruning
abstract
Infrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and computation inefficiencies. In this pioneering study, we introduce the concept of utilizing network pruning to enhance the efficiency of IRSTD. Due to the challenge posed by low signal-to-noise ratios and the absence of detailed semantic information in infrared images, directly applying existing pruning techniques yields suboptimal performance. To address this, we propose a novel wavelet structure-regularized soft channel pruning method, giving rise to the efficient IRPruneDet model. Our approach involves representing the weight matrix in the wavelet domain and formulating a wavelet channel pruning strategy. We incorporate wavelet regularization to induce structural sparsity without incurring extra memory usage. Moreover, we design a soft channel reconstruction method that preserves important target information against premature pruning, thereby ensuring an optimal sparse structure while maintaining overall sparsity. Through extensive experiments on two widely-used benchmarks, our IRPruneDet method surpasses established techniques in both model complexity and accuracy. Specifically, when employing U-net as the baseline network, IRPruneDet achieves a 64.13% reduction in parameters and a 51.19% decrease in FLOPS, while improving IoU from 73.31% to 75.12% and nIoU from 70.92% to 74.30%. The code is available at https://github.com/hd0013/IRPruneDet.
Mingjin Zhang, Handi Yang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037
AAAI3
2024 IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection
Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001, Jing Zhang 0037
ECCV (67)3
2024 Explore Hybrid Modeling for Moving Infrared Small Target Detection
Mingjin Zhang, Shilong Liu 0005, Yuanjun Ouyang, Jie Guo 0009, Zhihong Tang, Yunsong Li 0001
ACM Multimedia4
2024 VmambaSCI: Dynamic Deep Unfolding Network with Mamba for Compressive Spectral Imaging
abstract
Snapshot spectral compressive imaging can capture spectral information across multiple wavelengths in one imaging. The coded aperture snapshot spectral imaging (CASSI) method, aims to recover 3D spectral cubes from 2D measurements. Most existing approaches employ a deep unfolding framework based on Transformer, which alternately address a data subproblem and a prior subproblem. However, these frameworks lack flexibility regarding the sensing matrix and inter-stage interactions. In addition, the quadratic computational complexity of global Transformer and the restricted receptive field of local Transformer impact reconstruction efficiency and accuracy. In this paper, we propose a dynamic deep unfolding network with mamba for compressive spectral imaging, called VmambaSCI. We integrate spatial-spectral information from the sensing matrix into the data module and utilizes spatial adaptive operations in the stage interaction of the prior module. Furthermore, recognizing that the imaging process causes aliasing of spatial and spectral information, we develop a dual-domain scanning mamba (DSMamba), featuring a novel spatial-channel scanning method for enhanced efficiency and accuracy. To our knowledge, VmambaSCI is the first Mamba-based model for compressive spectral imaging. Experimental results on the public databases, CAVE and KAIST, demonstrate the superiority of the proposed VmambaSCI over the state-of-the-art approaches.
Mingjin Zhang, Longyi Li, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
ACM Multimedia4
2024 Single-Frame Infrared Small Target Detection via Gaussian Curvature Inspired Network
abstract
Single-frame infrared small target detection (SIRSTD) is in urgent demand for many practical tasks, such as fire rescue and urban management systems, benefiting from the excellent performance of infrared (IR) imaging in harsh climates and low-light environments. SIRSTD strives to segment small targets from the background as accurately as possible. However, in a real-world application, complex background environments with high brightness and strong edges have similar physical characteristics to small IR targets, which makes it extremely difficult to separate small targets. To address this challenge, we propose a novel Gaussian Curvature Inspired Network (GCI-Net). Inspired by the well-known Gaussian curvature, we develop a Gaussian curvature-based branch (GCB) to eliminate the smoothing noise and preserve the target structure texture information. In addition, we design a complementary patch-group attention (PGA) module that relies on the complementary relationship between low-level and high-level features to provide accurate guidance for GCB. The curvature information generated by the GCB is continuously optimized under the constraint of the curvature information of the ground truth. The proposed GCI-Net provides a reliable guarantee for accurate separation of small targets from the background. We conduct extensive experiments on the public IRSTD-1k and SIRST datasets. The experimental results demonstrate that the proposed GCI-Net outperforms the state-of-the-art (SOTA) methods.
Mingjin Zhang, Ke Yue, Boyang Li 0007, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Heat Transfer-Inspired Network for Image Super-Resolution Reconstruction
abstract
Image super-resolution (SR) is a critical image preprocessing task for many applications. How to recover features as accurately as possible is the focus of SR algorithms. Most existing SR methods tend to guide the image reconstruction process with gradient maps, frequency perception modules, etc. and improve the quality of recovered images from the perspective of enhancing edges, but rarely optimize the neural network structure from the system level. In this article, we conduct an in- depth exploration for the inner nature of the SR network structure. In light of the consistency between thermal particles in the thermal field and pixels in the image domain, we propose a novel heat-transfer-inspired network (HTI-Net) for image SR reconstruction based on the theoretical basis of heat transfer. With the finite difference theory, we use a second-order mixed-difference equation to redesign the residual network (ResNet), which can fully integrate multiple information to achieve better feature reuse. In addition, according to the thermal conduction differential equation (TCDE) in the thermal field, the pixel value flow equation (PVFE) in the image domain is derived to mine deep potential feature information. The experimental results on multiple standard databases demonstrate that the proposed HTI-Net has superior edge detail reconstruction effect and parameter performance compared with the existing SR methods. The experimental results on the microscope chip image (MCI) database consisting of realistic low-resolution (LR) and high-resolution (HR) images show that the proposed HTI-Net for image SR reconstruction can improve the effectiveness of the hardware Trojan detection system.
Mingjin Zhang, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution
abstract
Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between spectral features. This results in inadequate utilization of spectral information and artifacts after upsampling. To address this issue, we propose ES-SAformer, an ESSA attention-embedded Transformer network for single-HSI-SR with an iterative refining structure. Specifically, we first introduce a robust and spectral-friendly similarity metric, i.e., the spectral correlation coefficient of the spectrum (SCC), to replace the original attention matrix and incorporates inductive biases into the model to facilitate training. Built upon it, we further utilize the kernelizable attention technique with theoretical support to form a novel efficient SCC-kernel-based self-attention (ESSA) and reduce attention computation to linear complexity. ESSA enlarges the receptive field for features after upsampling without bringing much computation and allows the model to effectively utilize spatial-spectral information from different scales, resulting in the generation of more natural high-resolution images. Without the need for pretraining on large-scale datasets, our experiments demonstrate ESSA’s effectiveness in both visual quality and quantitative results. The code will be released at ESSAformer.
Mingjin Zhang, Chi Zhang 0080, Qiming Zhang 0001, Jie Guo 0009, Xinbo Gao 0001, Jing Zhang 0037
ICCV4
2023 Fluid Micelle Network for Image Super-Resolution Reconstruction
abstract
Most existing convolutional neural-network-based super-resolution (SR) methods focus on designing effective neural blocks but rarely describe the image SR mechanism from the perspective of image evolution in the SR process. In this study, we explore a new research routine by abstracting the movement of pixels in the reconstruction process as the flow of fluid in the field of fluid dynamics (FD), where explicit motion laws of particles have been discovered. Specifically, a novel fluid micelle network is devised for image SR based on the theory of FD that follows the residual learning scheme but learns the residual structure by solving the finite difference equation in FD. The pixel motion equation in the SR process is derived from the Navier-Stokes (N-S) FD equation, establishing a guided branch that is aware of edge information. Thus, the second-order residual drives the network for feature extraction, and the guided branch corrects the direction of the pixel stream to supplement the details. Experiments on popular benchmarks and a real-world microscope chip image dataset demonstrate that the proposed method outperforms other modern methods in terms of both objective metrics and visual quality. The proposed method can also reconstruct clear geometric structures, offering the potential for real-world applications.
Mingjin Zhang, Jing Zhang 0037, Xinbo Gao 0001, Jie Guo 0009, Dacheng Tao
IEEE Trans. Cybern.5
2023 Dim2Clear Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) is important for many practical applications such as hazardous aircraft warning, especially when the target is not visible in visible light image due to atmospheric conditions such as fog and cloud. However, IRSTD is challenging due to noises, small and dim targets. To address this challenge, we propose a novel Dim2Clear Network (Dim2Clear) for IRSTD in this paper. Specifically, the Dim2Clear consists of a U-Net backbone encoder, a context mixer decoder (CMD) based on spatial and frequency attention (SFA), and an eyeball-shaped enhancement module (EEM). The CMD is composed of cascaded regular residual blocks where two SFA modules are inserted. Each SFA module receives features from different residual blocks and generates spatial attention map from them to modulate the low-level features, which are then decomposed into low and high frequencies using the discrete cosine transformation. Accordingly, features are further modulated according to the generated frequency attention maps. In this way, SFA can extract both spatial context and frequency context to improve the feature representation capacity. In addition, we design an EEM to suppress the noise and enhance the signal-to-noise ratio in the segmentation results from the perspective of image super-resolution. Experiments on the SIRST dataset and our newly constructed IRSTD-1k dataset show that the proposed Dim2Clear outperforms state-of-the-art methods.
Mingjin Zhang, Rui Zhang 0124, Jing Zhang 0037, Jie Guo 0009, Yunsong Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 ISNet: Shape Matters for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) refers to extracting small and dim targets from blurred backgrounds, which has a wide range of applications such as traffic management and marine rescue. Due to the low signal-to-noise ratio and low contrast, infrared targets are easily submerged in the background of heavy noise and clutter. How to detect the precise shape information of infrared targets remains challenging. In this paper, we propose a novel infrared shape network (ISNet), where Taylor finite difference (TFD) -inspired edge block and two-orientation attention aggregation (TOAA) block are devised to address this problem. Specifically, TFD-inspired edge block aggregates and enhances the comprehensive edge information from different levels, in order to improve the contrast between target and background and also lay a foundation for extracting shape information with mathematical interpretation. TOAA block calculates the lowlevel information with attention mechanism in both row and column directions and fuses it with the high-level information to capture the shape characteristic of targets and suppress noises. In addition, we construct a new benchmark consisting of 1, 000 realistic images in various target shapes, different target sizes, and rich clutter backgrounds with accurate pixel-level annotations, called IRSTD-1k. Experiments on public datasets and IRSTD-1 k demonstrate the superiority of our approach over representative state-of-the-art IRSTD methods. The dataset and code are available at github.com/RuiZhang97/ISNet.
Mingjin Zhang, Rui Zhang 0124, Yuxiang Yang 0001, Haichen Bai, Jing Zhang 0037, Jie Guo 0009
CVPR6
2022 SAR-to-Optical Image Translation via Neural Partial Differential Equations
abstract
Synthetic Aperture Radar (SAR) becomes prevailing in remote sensing while SAR images are challenging to interpret by human visual perception due to the active imaging mechanism and speckle noise. Recent researches on SAR-to-optical image translation provide a promising solution and have attracted increasing attentions, though still suffering from low optical image quality with geometric distortion due to the large domain gap. In this paper, we mitigate this issue from a novel perspective, i.e., neural partial differential equations (PDE). First, based on the efficient numerical scheme for solving PDE, i.e., Taylor Central Difference (TCD), we devise a basic TCD residual block to build the backbone network, which promotes the extraction of useful information in SAR images by aggregating and enhancing features from different levels. Furthermore, inspired by the Perona-Malik Diffusion (PMD), we devise a PMD neural module to implement feature diffusion through layers, aiming at removing the noises in smooth regions while preserving the geometric structures. Assembling them together, we propose a novel SAR-to-Optical image translation network named S2O-NPDE, which delivers optical images with finer structures and less noise while enjoying an explainability advantage from explicit mathematical derivation. Experiments on the popular SEN1-2 dataset show that our model outperforms state-of-the-art methods in terms of both objective metrics and visual quality.
Mingjin Zhang, Chengyu He, Jing Zhang 0037, Yuxiang Yang 0001, Xiaoqi Peng, Jie Guo 0009
IJCAI6
2022 RKformer: Runge-Kutta Transformer with Random-Connection Attention for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) refers to segmenting the small targets from infrared images, which is of great significance in practical applications. However, due to the small scale of targets as well as noise and clutter in the background, current deep neural network-based methods struggle in extracting features with discriminative semantics while preserving fine details. In this paper, we address this problem by proposing a novel RKformer model with an encoder-decoder structure, where four specifically designed Runge-Kutta transformer (RKT) blocks are stacked sequentially in the encoder. Technically, it has three key designs. First, we adopt a parallel encoder block (PEB) of the transformer and convolution to take their advantages in long-range dependency modeling and locality modeling for extracting semantics and preserving details. Second, we propose a novel random-connection attention (RCA) block, which has a reservoir structure to learn sparse attention via random connections during training. RCA encourages the target to attend to sparse relevant positions instead of all the large-area background pixels, resulting in more informative attention scores. It has fewer parameters and computations than the original self-attention in the transformer while performing better. Third, inspired by neural ordinary differential equations (ODE), we stack two PEBs with several residual connections as the basic encoder block to implement the Runge-Kutta method for solving ODE, which can effectively enhance the feature and suppress noise. Experiments on the public NUAA-SIRST dataset and IRSTD-1k dataset demonstrate the superiority of the RKformer over state-of-the-art methods.
Mingjin Zhang, Haichen Bai, Jing Zhang 0037, Rui Zhang 0124, Jie Guo 0009, Xinbo Gao 0001
ACM Multimedia6
2022 Edge-Conditioned Feature Transform Network for Hyperspectral and Multispectral Image Fusion
abstract
Despite recent advances achieved by deep learning techniques in the fusion of low-spatial-resolution hyperspectral image (LR-HSI) and high-spatial-resolution multispectral image (HR-MSI), it remains a challenge to reconstruct the high-spatial-resolution HSI (HR-HSI) with more accurate spatial details and less spectral distortions, since the low-level structure information such as sharp edges tends to be weakened or lost as the network depth grows. To tackle this issue, we creatively propose an edge-conditioned feature transform network (EC-FTN) in this article, which is mainly composed of three parts, namely, feature extraction network (FEN), feature fusion and transformation network (FFTN), and image reconstruction network (IRN). First, two computationally efficient FENs with 3-D convolutions and reshaping layers are employed to extract the joint spectral-spatial features of input images. Then, the FFTN conditioned on the edge map prior can fuse and transform the features adaptively, in which a fusion node and several cascaded feature modulation modules (FMMs) equipped with feature-wise modulation layers are constructed. Specifically, the edge map is generated via transfer learning, i.e., by applying the Sobel operator to feature maps of the red-green-blue (RGB) version of HR-MSI resulting from the pretrained VGG16 model without extra training. Finally, the desired HR-HSI is recovered from the transformed features through IRN. Furthermore, we elaborately design a weighted combinatorial loss function consisting of mean absolute error, image gradient difference, and spectral angle terms to guide the training. Experiments on both ground-based and remotely sensed datasets demonstrate that our EC-FTN outperforms state-of-the-art methods in visual and quantitive evaluations, as well as in fine details reconstruction.
Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Jie Guo 0009, Xianyun Wu, Yanzi Shi, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2020 Hyperspectral Pansharpening Using Deep Prior and Dual Attention Residual Network
abstract
Convolutional neural networks (CNNs) have recently achieved impressive improvements on hyperspectral (HS) pansharpening. However, most of the CNN-based HS pansharpening approaches would have to first upsample the low-resolution hyperspectral image (LR-HSI) using bicubic interpolation or data-driven training strategy, which inevitably lose some details or greatly rely on the learning process. In addition, most previous methods regard the pansharpening as a black-box problem and treat diverse features equally, thus hindering the discriminative ability of CNNs. To conquer these issues, a novel HS pansharpening method using deep hyperspectral prior (DHP) and dual-attention residual network (DARN) is proposed in this article. Specifically, we first upsample the LR-HSI to the scale of the panchromatic (PAN) image through the DHP algorithm, which can better preserve spatial and spectral information without learning from large data sets. The upsampled result is then concatenated with the PAN image to form the input of the DARN, where several channel-spatial attention residual blocks (CSA ResBlocks) are stacked to map the residual HSI between the reference HSI and the upsampled HSI. In each CSA ResBlock, two complementary attention modules, i.e., channel attention and spatial attention modules, are designed to adaptively learn more informative features of spectral channels and spatial locations simultaneously, which can effectively boost the fusion accuracy. Finally, the fused HSI is obtained by the summation of the upsampled HSI and the reconstructed residual HSI. The experimental results of both simulated and real HS data sets demonstrate that the performance of our DHP-DARN method is superior over the state-of-the-art HS pansharpening approaches.
Yuxuan Zheng, Jiaojiao Li 0001, Yunsong Li 0001, Jie Guo 0009, Xianyun Wu, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2012 VLSI Architecture of Arithmetic Coder Used in SPIHT
abstract
A high-throughput memory-efficient arithmetic coder architecture for the set partitioning in hierarchical trees (SPIHT) image compression is proposed based on a simple context model in this paper. The architecture benefits from various optimizations performed at different levels of arithmetic coding from higher algorithm abstraction to lower circuits' implementations. First, the complex context model used by software is mitigated by designing a simple context model, which just uses the brother nodes' states in the coding zerotree of SPIHT to form context symbols for the arithmetic coding. The simple context model results in a regular access pattern during reading the wavelet transform coefficients, which is convenient to the hardware implementation, but at a cost of slight performance loss. Second, in order to avoid rescanning the wavelet transform coefficients, a breadth first search SPIHT without lists algorithm is used instead of SPIHT with lists algorithm. Especially, the coding bit-planes of each zero tree are processed in parallel. Third, an out-of-order execution mechanism for different types of context is proposed that can allocate the context symbol to the idle arithmetic coding core with a different order that of the input. For the balance of the input rate of the wavelet coefficients, eight arithmetic coders are replicated in the compression system. And in one arithmetic coder, there exists four cores to process different contexts. Fourth, several dedicated circuits are designed to further improve the throughput of the architecture. The common bit detection (CBD) circuit is used for unrolling the renormalization stage of the arithmetic coding. The carry look-ahead adder (CLA) and fast multiplier-divider are also employed to shorten the critical path in the architecture. Moreover, an adaptive clock switch mechanism can stop some invalid bit-planes' clock for the power saving purpose according to the input images. Experimental results demonstrate that the proposed architecture attains a throughput of 902.464 Mb/s at its maximum and achieves savings of 20.08% in power consumption over full bit-planes coding scheme based on field-programmable gate arrays (FPGAs).
Kai Liu 0021, Eugeniy Belyaev, Jie Guo 0009
IEEE Trans. Very Large Scale Integr. Syst.3