Xinying Wang 0005

dblp:06/3244-5 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-3704-4747ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 GMSR: Gradient-integrated mamba for spectral reconstruction from RGB images
Xinying Wang 0005, Zhixiong Huang, Jiawen Zhu 0003, Paolo Gamba, Lin Feng 0001
Neural Networks1
2025 Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual Tracking
abstract
Efficient tracking has garnered attention for its ability to operate on resource-constrained platforms for real-world deployment beyond desktop GPUs. Current efficient trackers mainly follow precision-oriented trackers, adopting a one-stream framework with lightweight modules. However, blindly adhering to the one-stream paradigm may not be optimal, as incorporating template computation in every frame leads to redundancy, and pervasive semantic interaction between template and search region places stress on edge devices. In this work, we propose a novel asymmetric Siamese tracker named AsymTrack for efficient tracking. AsymTrack disentangles template and search streams into separate branches, with template computing only once during initialization to generate modulation signals. Building on this architecture, we devise an efficient template modulation mechanism to unidirectional inject crucial cues into the search features, and design an object perception enhancement module that integrates abstract semantics and local details to overcome the limited representation in lightweight tracker. Extensive experiments demonstrate that AsymTrack offers superior speed-precision trade-offs across different platforms compared to the current state-of-the-arts. For instance, AsymTrack-T achieves 60.8% AUC on LaSOT and 224/81/84 FPS on GPU/CPU/AGX, surpassing HiT-Tiny by 6.0% AUC with higher speeds.
Jiawen Zhu 0003, Huayi Tang, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Huchuan Lu
AAAI4
2025 Underwater variable zoom: Depth-guided perception network for underwater image enhancement
Zhixiong Huang, Xinying Wang 0005, Chengpei Xu, Jinjiang Li 0001, Lin Feng 0001
Expert Syst. Appl.2
2025 S3-Net: Learning spectral-spatio self-similarity for hyperspectral image super-resolution
Xinying Wang 0005, Zhixiong Huang, Jiawen Zhu 0003, Xiang-Hai Wang 0001, Lin Feng 0001
Neural Networks1
2025 Lightweight Edge-Guided Super-Resolution Network for Remote Sensing Images
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) techniques have achieved significant progress, but challenges remain in preserving critical edge details essential for high-quality image reconstruction, These details are crucial for tasks like object recognition, change detection, and accurate analysis in remote sensing imagery. Furthermore, existing RSISR methods typically require substantial computational resources, making them unsuitable for resource-constrained edge devices. To address these challenges, we propose a novel Edge-Guided Super-Resolution Network (EGSRN). The network employs an Edge Extraction Module (Edge Net) to explicitly extract edge information from low-resolution images, combined with multi-layer Feature Extraction Modules (FEM) and an Edge Information Fusion (EIF) mechanism to progressively integrate edge and image features. This design enables precise recovery of edge details, significantly enhancing the overall visual quality of the reconstructed images. Edge-aware processing enhances visual fidelity while also improving the accuracy of downstream tasks, such as classification, object detection, and change analysis. Furthermore, the network incorporates lightweight designs such as depthwise separable convolutions and channel shuffling to effectively reduce computational demands. Comprehensive experiments were conducted on two remote sensing datasets, and the model’s parameter count and floating-point operations (FLOPs) were evaluated. Results demonstrate that the proposed method achieves an excellent balance between performance and model complexity, delivering superior super-resolution reconstruction quality while maintaining low computational costs, making it well-suited for resource-limited real-world applications.
Zhixiong Huang, Xinying Wang 0005, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 WFA-SRNet: A Wavelet-Guided and Feature-Aware Network for Remote Sensing Image Super-Resolution
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) methods have achieved remarkable progress. However, effectively preserving high-frequency details remains a significant challenge, as these features are critical for downstream tasks such as object detection, change analysis, and scene classification. Moreover, relying solely on the information contained in low-resolution images often results in the loss of structural details, thereby degrading reconstruction quality. To address these issues, we propose a novel Wavelet-guided and Feature-Aware Super-Resolution Network (WFA-SRNet). The proposed network adopts a dual-branch architecture, consisting of a Feature Extraction Block (FEB) and a High-Frequency Extraction Block (HFE), to collaboratively model semantic structures and fine-grained textures. Specifically, FEB integrates a Shift-Window Cross Attention (SWCA) mechanism and a dictionary-based similarity matching strategy to capture non-local self-similarities, while the HFE branch incorporates a wavelet-domain high-frequency modeling module (WD-HFE), which explicitly decomposes and reconstructs frequency components via Discrete Wavelet Transform (DWT) and Inverse DWT (IDWT) to enhance edge and texture recovery. Furthermore, a Fusion Attention (FA) module is designed to guide the integration of multi-source features from both semantic and high-frequency pathways. Extensive experiments on multiple benchmark remote sensing datasets demonstrate that WFA-SRNet achieves superior reconstruction performance, particularly in restoring structural and textural details. Additionally, the proposed method significantly improves the accuracy of downstream classification tasks, showing strong potential for practical RSISR applications.
Xinying Wang 0005, Zhixiong Huang, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 DCR-SRNet: A Degradation-Contrastive and Wavelet-Guided Network for Blind Remote Sensing Image Super-Resolution
abstract
Recently, deep learning-based remote sensing image super-resolution (RSISR) has achieved remarkable progress. However, conventional super-resolution methods usually assume a fixed and known degradation process (e.g., bicubic downsampling), which often leads to significant performance degradation when applied to real-world data with diverse and unknown degradations. To overcome this limitation, we propose DCR-SRNet, a novel Degradation-Contrastive and Wavelet-Guided Network for blind RSISR. The proposed network incorporates three key innovations: First, we design a contrastive degradation representation learning strategy that disentangles degradation priors from scene semantics by pulling together representations of identical degradations across different scenes while pushing apart those of different degradations within the same scene. Second, we introduce a wavelet-guided patch-wise weighted loss module, which employs wavelet decomposition and patch-level discrimination scores to adaptively reweight the pixel-wise loss, thereby enhancing the recovery of edge and texture details. Third, we design an adaptive modulation block (AMB) that injects degradation priors into the reconstruction process through feature- and channel-wise modulation, enabling robust adaptation to diverse degradations. Extensive experiments on three benchmark remote sensing datasets demonstrate that DCR-SRNet significantly outperforms state-of-the-art methods, particularly in preserving structural and textural details.
Zhixiong Huang, Xinying Wang 0005, Shenglan Liu 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 DISD-Net: A Dynamic Interactive Network With Self-Distillation for Cross-Subject Multi-Modal Emotion Recognition
abstract
Multi-modal Emotion Recognition (MER) has demonstrated competitive performance in affective computing, owing to synthesizing information from diverse modalities. However, many existing approaches still face unresolved challenges, such as: (i) how to learn compact yet representative features from multi-modal data simultaneously and (ii) how to address differences among subjects and enhance the generalization of the emotion recognition model, given the diverse nature of individual biological signals. To this end, we propose a Dynamic Interactive Network with Self-Distillation (DISD-Net) for cross-subject MER. The DISD-Net incorporates a dynamin interactive module to capture the intra- and inter-modal interactions from multi-modal data. Additionally, to enhance compactness in modal representations, we leverage the soft labels generated by the DISD-Net model as supplemental training guidance. This involves incorporating self-distillation, aiming to transfer the knowledge that the DISD-Net model contains hard and soft labels to each modality. Finally, domain adaptation (DA) is seamlessly integrated into the dynamic interactive and self-distillation components, forming a unified framework to extract subject-invariant multi-modal emotional features. Experimental results indicate that the proposed model achieves a mean accuracy of 75.00% with a standard deviation of 7.68% for the DEAP dataset and a mean accuracy of 65.65% with a standard deviation of 5.08% for the SEED-IV dataset.
Cheng Cheng 0013, Xinying Wang 0005, Lin Feng 0001, Ziyu Jia
IEEE Trans. Multim.3
2024 SRRT: Exploring Search Region Regulation for Visual Object Tracking
abstract
The dominant trackers generate a fixed-size rectangular region based on the previous prediction or initial bounding box as the model input, i.e., search region. While this manner obtains promising tracking efficiency, a fixed-size search region lacks flexibility and is likely to fail in some cases, e.g., fast motion and distractor interference. Trackers tend to lose the target object due to the limited search region or experience interference from distractors due to the excessive search region. Drawing inspiration from the pattern humans track an object, we propose a novel tracking paradigm, called Search Region Regulation Tracking (SRRT) that applies a small eyereach when the target is captured and zooms out the search field when the target is about to be lost. SRRT applies a proposed search region regulator to estimate an optimal search region dynamically for each frame, by which the tracker can flexibly respond to transient changes in the location of object occurrences. To adapt the object’s appearance variation during online tracking, we further propose a locking-state determined updating strategy for reference frame updating. The proposed SRRT is concise without bells and whistles, yet achieves evident improvements and competitive results with other state-of-the-art trackers on eight benchmarks. On the large-scale LaSOT benchmark, SRRT improves SiamRPN++ and TransT with absolute gains of 4.6% and 3.1% in terms of AUC. The code and models will be released.
Jiawen Zhu 0003, Xin Chen 0032, Xinying Wang 0005, Dong Wang 0004, Wenda Zhao 0003, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.4
2023 MCT-Net: Multi-hierarchical cross transformer for hyperspectral and multispectral image fusion
Xiang-Hai Wang 0001, Xinying Wang 0005, Ruoxi Song, Xiao-Yang Zhao 0003, Keyun Zhao
Knowl. Based Syst.2
2023 SS-INR: Spatial-Spectral Implicit Neural Representation Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to the limitation of imaging equipment, it is difficult to acquire hyperspectral images with high spatial resolution directly. Existing approaches improve the resolution of HSIs by fusing multispectral image (MSI) and hyperspectral image (HSI). However, most of them are only feed-forward. They only learn low- to high-resolution feature mappings without considering the ill-posedness of super-resolution tasks, leading to a large solution space of mapping functions and making it difficult to learn a complete mapping function. Moreover, there is a large resolution difference between HSI and MSI, and some up-sampling operations are inevitably employed in the network. Nevertheless, traditional upsampling methods only represent pixel points in a discrete way, failing to adequately restore the continuous spatial and spectral information. To this end, this paper proposes a spatial-spectral implicit neural representation network for hyperspectral and multispectral image fusion (SS-INR). Inspired by the success of implicit neural representation(INR) in continuum reconstruction, we design spatial-INR and spectral-INR for spatial and spectral resolution reconstruction, respectively. SS-INR contains two processes: forward fusion (FF) and back-projection fusion(BPF). In the FF process, the input HSI is first spatially upsampled with Spatial-INR to overcome spatial resolution differences while performing initial fusion with MSI. In the BPF process, we explore the spatial and spectral degradation processes and use them as prior knowledge for error correction. Extensive experiments on five public hyperspectral datasets demonstrate the effectiveness of SS-INR, and SS-INR achieves competitive results compared with existing state-of-the-art fusion methods. The source code for SS-INR will be released at https://github.com/wxy11-27/SS-INR.
Xinying Wang 0005, Cheng Cheng 0013, Shenglan Liu 0001, Ruoxi Song, Xiang-Hai Wang 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 FSL-Unet: Full-Scale Linked Unet With Spatial-Spectral Joint Perceptual Attention for Hyperspectral and Multispectral Image Fusion
abstract
The application of hyperspectral image (HSI) is more and more extensive, but the lower spatial resolution seriously affects its application effect. Using low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (MSI) fusion technology to achieve super-resolution reconstruction of HSI has become a mainstream method. However, most of the existing fusion methods do not make full use of the large-scale range of remote sensing images, and neglect the preservation of spatial-spectral information in the fusion process. Considering that the spectral information in fused high-resolution hyperspectral image (HR-HSI) mainly depends on HSI, and the spatial information mainly depends on MSI, this paper proposes a full-scale linked Unet with spatial-spectral joint perceptual attention for hyperspectral and multispectral image fusion (FSL-Unet). The FSL-Unet consists of two modules, the first is spatial-spectral attention extraction module (SSAE), which is used to calculate the spectral attention of LR-HSI and the spatial attention of HR-MSI at different scales. The second is the full-scale link U-shaped fusion module (FLUF), which adopts a multi-level feature extraction strategy, using denser full-scale skip connections to explore feature information in a finer-grained range, enabling flexible combination of multi-scale and multi-path features. At the same time, we propose spatial-spectral joint peceptual attention (SSJPA) on the encoder side of FLUF. SSJPA can make full use of the attention maps computed by the SSAE, and then effectively embed spatial and spectral information into the fused image, enabling uninterrupted information transfer and aggregation. To demonstrate the effectiveness of FSL-Unet, we selected five public hyperspectral datasets for experiments. Compared with other eight state-of-the-art fusion methods, the experimental results show that the FSL-Unet achieves competitive results. The source code for FSL-Unet can be downloaded from https://github.com/wxy11-27/FSL-Unet.
Xiang-Hai Wang 0001, Xinying Wang 0005, Keyun Zhao, Xiao-Yang Zhao 0003, Chuanming Song 0001
IEEE Trans. Geosci. Remote. Sens.2