EDBT 2026 Demo / reviewers in the wild / expert
Yuzeng Chen
dblp:325/1321
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-6757-9051ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STAR: A Unified Spatiotemporal Fusion Framework for Satellite Video Object TrackingabstractSatellite video object tracking (SVOT) delivers comprehensive spatiotemporal insights for Earth surface observation, yet existing SVOT methods confront several critical challenges including data scarcity, modality restrictions, paradigm gaps, and underutilization of multidimensional features, sealing the performance ceiling. This study proposes STAR, a unified spatiotemporal fusion framework for satellite video object tracking, mitigating these issues. To optimize satellite video scenes, STAR first introduces a scene enhancement module for generating enhanced multi-modal representations. Then, the extraction-correlation-adaptation module is designed, incorporating a multi-modal hierarchical Transformer architecture with local and unified relation modeling, which jointly achieves feature extraction, relation learning, and domain adaptation. Additionally, the temporal decoding structure is introduced to integrate deep temporal features via attention propagation. Finally, the inertial navigation module models physical temporal features, including an awareness selector to assess the tracking confidence-uncertainty and an inertial navigation scheme to manage anomalous interferences and continuous trajectory. Inspired by the prompt learning pattern, STAR introduces a minimal number of tunable parameters yet achieves competitive performance across various SVOT benchmarks. Implementation details and evaluation results will be available at: https://github.com/YZCU/STAR. Yuzeng Chen, Qiangqiang Yuan, Yi Xiao 0003, Te Han |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Hyperspectral Video Tracking With Spectral-Spatial Fusion and Memory EnhancementabstractHyperspectral video (HSV) provides rich spectral-spatial-temporal information, enabling the capture of complex object dynamics beyond the limitations of conventional single- and multi-modal tracking. However, current HSV tracking methods face challenges such as data scarcity, band gaps, spectral fragmentation, temporal underutilization, and high computational load, which constrain performance. In this article, we present SpectralTrack, a novel HSV tracking framework with spectral-spatial fusion and memory enhancement. SpectralTrack incorporates an explicit visual prompting module to mitigate band gaps and spectral fragmentation. We further introduce an extraction-matching-interaction module, which leverages a template-bridging search adapter and a multi-layer perceptron adapter within a multi-modal Transformer architecture for efficient cross-modal feature extraction-matching-interaction. Additionally, a memory perception module enhances state reasoning by injecting temporal prompts to refine spectral and spatial cues. SpectralTrack follows parameter-efficient fine-tuning and feature-level fusion to alleviate data scarcity and reduce computational overhead. We instantiate two variants, SpectralTrack and SpectralTrack+, across nine HSV tracking datasets, demonstrating superior effectiveness over extensive trackers. Implementations and results will be available at https://github.com/YZCU/SpectralTrack. Yuzeng Chen, Qiangqiang Yuan, Hong Xie 0002, Yi Xiao 0003, Renxiang Guan, Xinwang Liu 0002, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multi-Axis Feature Diversity Enhancement for Remote Sensing Video Super-ResolutionabstractHow to aggregate spatial-temporal information plays an essential role in video super-resolution (VSR) tasks. Despite the remarkable success, existing methods adopt static convolution to encode spatial-temporal information, which lacks flexibility in aggregating information in large-scale remote sensing scenes, as they often contain heterogeneous features (e.g., diverse textures). In this paper, we propose a spatial feature diversity enhancement module (SDE) and channel diversity enhancement module (CDE), which explore the diverse representation of different local patterns while aggregating the global response with compactly channel-wise embedding representation. Specifically, SDE introduces multiple learnable filters to extract representative spatial variants and encodes them to generate a dynamic kernel for enriched spatial representation. To explore the diversity in the channel dimension, CDE exploits the discrete cosine transform to transform the feature into the frequency domain. This enriches the channel representation while mitigating massive frequency loss caused by pooling operation. Based on SDE and CDE, we further devise a multi-axis feature diversity enhancement (MADE) module to harmonize the spatial, channel, and pixel-wise features for diverse feature fusion. These elaborate strategies form a novel network for satellite VSR, termed MADNet, which achieves favorable performance against state-of-the-art method BasicVSR++ in terms of average PSNR by 0.14 dB on various video satellites, including JiLin-1, Carbonite-2, SkySat-1, and UrtheCast. Code will be available at https://github.com/XY-boy/MADNet. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Shiqi Wang 0001, Chia-Wen Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | Frequency-Assisted Mamba for Remote Sensing Image Super-ResolutionabstractRecent progress in remote sensing image (RSI) super-resolution (SR) has exhibited remarkable performance using deep neural networks, e.g., Convolutional Neural Networks and Transformers. However, existing SR methods often suffer from either a limited receptive field or quadratic computational overhead, resulting in sub-optimal global representation and unacceptable computational costs in large-scale RSI. To alleviate these issues, we develop the first attempt to integrate the Vision State Space Model (Mamba) for RSI-SR, which specializes in processing large-scale RSI by capturing long-range dependency with linear complexity. To achieve better SR reconstruction, building upon Mamba, we devise a Frequency-assisted Mamba framework, dubbed FMSR, to explore the spatial and frequent correlations. In particular, our FMSR features a multi-level fusion architecture equipped with the Frequency Selection Module (FSM), Vision State Space Module (VSSM), and Hybrid Gate Module (HGM) to grasp their merits for effective spatial-frequency fusion. Considering that global and local dependencies are complementary and both beneficial for SR, we further recalibrate these multi-level features for accurate feature fusion via learnable scaling adaptors. Extensive experiments on AID, DOTA, and DIOR benchmarks demonstrate that our FMSR outperforms state-of-the-art Transformer-based methods HAT-L in terms of PSNR by 0.11 dB on average, while consuming only 28.05% and 19.08% of its memory consumption and complexity, respectively. Yi Xiao 0003, Qiangqiang Yuan, Kui Jiang, Yuzeng Chen, Qiang Zhang 0011, Chia-Wen Lin |
IEEE Trans. Multim. | 4 |
| 2024 | PHTrack: Prompting for Hyperspectral Video TrackingabstractHyperspectral (HS) video captures continuous spectral information of objects, enhancing material identification in tracking tasks. It is expected to overcome the inherent limitations of red-green–blue (RGB) and multimodal tracking, such as finite spectral cues and cumbersome modality alignment. However, HS tracking faces challenges such as data anxiety, bandgaps, and huge volumes. In this study, inspired by prompt learning in language models, we propose the prompting for hyperspectral video tracking (PHTrack) framework. PHTrack learns prompts to adapt foundation models, mitigating data anxiety and enhancing performance and efficiency. First, the modality prompter (MOP) is proposed to capture rich spectral cues and bridge bandgaps for improved model adaptation and knowledge enhancement. In addition, the distillation prompter (DIP) is developed to refine cross-modal features. PHTrack follows feature-level fusion, effectively managing huge volumes compared to traditional decision-level fusion fashions. Extensive experiments validate the proposed framework, offering valuable insights for future research. The code and data will be available athttps://github.com/YZCU/PHTrack Yuzeng Chen, Xin Su 0003, Jie Li 0022, Yi Xiao 0003, Qiangqiang Yuan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | SPIRIT: Spectral Awareness Interaction Network With Dynamic Template for Hyperspectral Object TrackingabstractHyperspectral (HS) video is able to capture abundant spectral, spatial, and temporal information about objects, which overcomes the limitations of common red-green-blue (RGB) video in complex scenarios such as similar appearances and background clutters (BCs). However, most trackers apply hand-crafted features extracted from manually selected bands instead of deep features for object representations due to limited HS data and the band gap problem. Each HS image consists of many bands, and it is challenging to fully interact with the band information while maintaining tracking speed. To this end, this article proposes a novel end-to-end spectral awareness interaction network with a dynamic template (SPIRIT) for HS video object tracking. First, a spectral awareness module (SAM) is proposed to learn band contributions with consideration of nonlinear and global interactions between HS bands. It can also cooperate with the feature extraction module pretrained with RGB data to attenuate the band gap and data-hungry. Second, an interaction module (IM) is proposed to achieve inter and intraband feature interactions to enhance tracking performance while improving efficiency. Furthermore, the proposed method contains a novel update module (UM) that evaluates the tracking confidence of the current state to adapt to object changes and attenuate tracking drifts. Extensive experiments demonstrate the superiority of our approach compared to state-of-the-arts (SOTAs) while meeting real-time demands. Yuzeng Chen, Qiangqiang Yuan, Yi Xiao 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Spatial-Spectral Graph Contrastive Clustering With Hard Sample Mining for Hyperspectral ImagesabstractHyperspectral image (HSI) clustering is a fundamental yet challenging task that groups image pixels with similar features into distinct clusters. Among various approaches, contrastive learning methods, which employ the concept of encouraging semantically similar samples to move closer together while pushing semantically inconsistent samples apart, have garnered significant attention due to their promising performance. However, the most prevalent approaches face two major limitations: 1) treating all samples indiscriminately during optimization, where the abundance of well-categorized samples overwhelms the feature learning process and 2) tending to introduce noise when constructing positive sample pairs through view augmentation or searching the nearest neighbors, which would cause semantic drift of sample features. To solve these issues, we propose a graph autoencoder-based deep clustering framework named spatial–spectral graph contrastive clustering with hard sample mining (SSGCC) that constructs spatial–spectral dual views without data augmentation and focuses more on hard samples rather than treating all samples equally with the aid of spatial–spectral features. Concretely, we extract the spectral features and the neighborhood spatial features of the samples as dual branches to avoid the noise caused by data augmentation and develop the cluster-oriented consistency learning to facilitate the exchange of knowledge between the two spectral–spatial perspectives. In addition, we propose a hard sample mining-based contrastive learning scheme with the aid of spatial–spectral features. To better measure the importance of the samples, we combine spatial features and spectral features to calculate the similarity between sample pairs. The weights of hard sample pairs are dynamically up-weight while the easy ones are down-weighting to improve the discriminative capability. Extensive experiments on four benchmark HSI datasets demonstrate the effectiveness and superiority of the proposed methods against state-of-the-art ones. Renxiang Guan, Wenxuan Tu, Hao Yu 0017, Dayu Hu, Yuzeng Chen, Chang Tang, Qiangqiang Yuan, Xinwang Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | SDC-GAE: Structural Difference Compensation Graph Autoencoder for Unsupervised Multimodal Change DetectionabstractMultimodal change detection (MCD) is a crucial technology for applications in natural resource monitoring, disaster assessment, and urban planning. To address the reliance on labeled data and enhance the robustness of structural features in the existing methods, we propose a structure difference compensation graph autoencoder (SDC-GAE) for unsupervised MCD. It is recognized that the registered multimodal images exhibit consistency in structural features in unchanged areas, while the structural features in changed areas are distinct. SDC-GAE utilizes a graph convolutional network (GCN) to extract deep structural features from multimodal images. It uses the structural features of one time-phase image to reconstruct its spectral features in the spectral feature space of the target image. Through structural difference compensation, SDC-GAE learns the structural disparities between different images, with the compensation value directly reflecting the intensity of the changes. The SDC-GAE loss function consists of three components: image reconstruction loss, which evaluates the spectral feature discrepancy between the reconstructed and target images, guiding the model to reduce these differences via structural difference compensation; sparse constraint loss, which accounts for the fact that changes are typically confined to a few areas, ensuring the sparsity of the detected changes; and structural consistency loss, which aligns the structural features of the reconstructed image closely with those of the target image. The efficacy of our method is validated through experiments on eight multimodal datasets, where it is compared with the state-of-the-art methods. Te Han, Yuzeng Chen, Yuqiang Guo, Shujing Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Heterogeneous Image Change Detection Based on Two-Stage Joint Feature LearningabstractHeterogeneous image change detection, in contrast to homogeneous image change detection, has been a research hotspot due to the information complementary of different imaging mechanisms. However, the imaging difference leads to challenges on change detection by image comparison. To address the incomparability among heterogeneous images and improve the efficiency of heterogeneous image change detection, this paper proposes a novel heterogeneous image change detection method based two-stage joint feature learning. Assuming that the change is few and the image differences in unchanged areas between heterogeneous images are related to the imaging and environmental differences, it maps heterogeneous images into a similar feature space for comparison. Firstly, the bi-temporal similar feature maps with high similarity are extracted after joint feature learning of heterogeneous image. And the similar feature maps are used for joint feature learning optimized by a similarity measure in order to map them to an approximate feature space for comparison. Then the change map is obtained by segmenting the difference between the optimal feature maps. The experiments prove its superiority over existing methods on two heterogeneous image datasets (optical and synthetic aperture radar (SAR) images). Te Han, Yuzeng Chen |
IGARSS | 3 |