Wen Guo 0003

dblp:49/2045-3 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-3691-4942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Occlusion-Aware Multi-Object Tracking via Joint Diffusion Motion Prediction and Appearance Purification
abstract
Occlusion, a pervasive challenge in Multi-Object Tracking (MOT) within complex scenes, severely degrades tracking performance. Current methods still face numerous problems when handling occlusions. Motion prediction struggles to accommodate diverse motion patterns, resulting in failure during short-term occlusions. Concurrently, appearance features possess insufficient discriminative power under occluded conditions, which frequently leads to identity switches following long-term occlusion. To enhance the performance of MOT under such challenging conditions, we propose an innovative Occlusion-Aware Multi-Object Tracking via Joint Diffusion Motion Prediction and Appearance Purification (OAMOT). For short-term occlusions, a Diffusion Motion Recovery Model (DMRM) is developed, which integrates residual and noise diffusion branches to recover trajectories precisely under varied motion patterns. For long-term occlusions, the SAM Appearance Purification for ReID (SAPR) module is proposed; the module leverages the mask generation mechanism of SAM as foreground attention to enhance feature discriminability. Furthermore, a Lightweight Attention Predictor (LAP) is integrated into the ReID network to achieve SAM-quality attention during inference without significant computational overhead. Experimental results on the MOT17, MOT20, and DanceTrack datasets demonstrate that the proposed OAMOT method outperforms current state-of-the-art multi-object tracking techniques across multiple evaluation metrics. The code can be available on https://github.com/wangtuo111/OAMOT.
Wen Guo 0003, Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.1
2026 History-Guided Prompt Generation for Vision-and-Language Navigation
abstract
Vision-and-language navigation (VLN) has garnered extensive attention in the field of embodied artificial intelligence. VLN involves time series information, where historical observations contain rich contextual knowledge and play a crucial role in navigation. However, current methods do not explicitly excavate the connection between rich contextual information in history and the current environment, and ignore adaptive learning of clues related to the current environment. Therefore, we explore a Prompt Learning-based strategy which adaptively mines information in history that is highly relevant to the current environment to enhance the agent's perception of the current environment and propose a history-guided prompt generation (HGPG) framework. Specifically, HGPG includes two parts, one is an entropy-based history acquisition module that assesses the uncertainty of the action probability distribution from the preceding step to determine whether historical information should be used at the current time step. The other part is the prompt generation module that transforms historical context into prompt vectors by sampling from an end-to-end learned token library. These prompt tokens serve as discrete, knowledge-rich representations that encode semantic cues from historical observations in a compact form, making them easier for the decision network to understand and utilize. In addition, we share the token library across various navigation tasks, mining common features between different tasks to improve generalization to unknown environments. Extensive experimental results on four mainstream VLN benchmarks (R2R, REVERIE, SOON, R2R-CE) demonstrate the effectiveness of our proposed method. Code is available at https://github.com/Wzmshdong/HGPG.
Wen Guo 0003, Zongmeng Wang, Yufan Hu, Junyu Gao 0002
IEEE Trans. Cybern.1
2025 What Is in the Frequency: Wavelet-Guided Semantic Understanding for Infrared Small Target Detection
abstract
The task of infrared small target detection holds significant application value in military surveillance and civilian security. Small targets typically exhibit characteristics such as weak texture, low contrast, and small scale, which make them susceptible to interference from complex backgrounds. Although existing deep learning-based models have incorporated background modeling to enhance contextual awareness, high-frequency and low-frequency information are often entangled during feature propagation. The entanglement makes it difficult to distinguish targets from the background, consequently leading to false alarms and missed detections. To address the issue, we propose a dual-band detection network based on spectral decoupling, termed the Frequency-Guided Semantic Understanding Network (FSUNet) for infrared small target detection. The encoder and decoder of our network are composed of Wavelet-Guided Semantic Disentangling Blocks (W-SD) and Dual-Band Spectral Cooperative Blocks (D-SC), respectively. The W-SD explicitly separates high-frequency edge and low-frequency structural semantic features using wavelet frequency-domain priors. The D-SC is implemented by combining a Dual-Band Refinement (DBR) and a Semantic Re-Sampling (SRS) mechanism. The DBR filters and optimizes effective information for the high-frequency and low-frequency branches separately, suppressing noise interference. The SRS fuses features from the two branches, by dynamically correcting their statistical differences and reconstructing channel weights to highlight discriminative features, thereby enhancing target distinguishability and noise-resistant robustness. Experimental results on public datasets SIRST, NUDT-SIRST, and IRSTD-1K demonstrate that our FSUNet method outperforms existing methods from the perspective of detection accuracy and robustness, especially in complex backgrounds with low contrast. The code can be available on https://github.com/fulongcai/FSUNet-main.
Wen Guo 0003, Fulong Cai, Wuzhou Quan
IEEE Trans. Geosci. Remote. Sens.1
2024 Noise-aware progressive multi-scale deepfake detection
Xinmiao Ding, Shuai Pang, Wen Guo 0003
Multim. Tools Appl.3
2024 Feature Disentanglement Network: Multi-Object Tracking Needs More Differentiated Features
abstract
To reduce computational redundancies, a common approach is to integrate detection and re-identification (Re-ID) into a single network in multi-object tracking (MOT), referred to as “tracking by detection.” Most of the previous research has focused on resolving the conflict between the detection and Re-ID branches, considering it a simple coupling. In our work, we uncover that the entangled state between the detection and Re-ID tasks is much more complex than previous idea, resulting in a form of competition that degrades performance. To address the preceding issue, we propose a feature disentanglement network that deeply disentangles the intricately interwoven latent space of features and provides differentiated feature maps for each individual task. Furthermore, considering the demand for shallow semantic features in the feature re-ID branch, we also introduce a feature re-globalization module to enrich the shallow semantics. By integrating two distinct networks into a one-shot online MOT method, we develop a robust MOT tracker (named HDGTrack ). We conduct extensive experiments on a number of benchmarks, and our experimental results demonstrate that our method significantly outperforms state-of-the-art MOT methods. Besides, HDGTrack is efficient and can run at 13.9 (MOT17) and 8.7 (MOT20) frames per second.
Wen Guo 0003, Wuzhou Quan, Junyu Gao 0002, Tianzhu Zhang 0001, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Multi-view region proposal network predictive learning for tracking
Wen Guo 0003, Bin Shan
Multim. Syst.1
2023 DTEMPan: Dual Texture-Edge Maintaining Transformer for Pansharpening
abstract
Pansharpening plays a crucial role in the domain of remote sensing image processing, as it allows for the generation of high-resolution multispectral images. The main objective of pansharpening is to fuse low-resolution multispectral and high-resolution panchromatic images, resulting in high-resolution multispectral images that exhibit uniform spectra and enhanced spatial details. Consequently, the primary focus of related research is to preserve accurate features from both input images and achieve superior image reconstruction. In this paper, we introduce a novel pansharpening framework called Dual Texture-Edge Maintaining Transformer (DTEMPan). Our framework achieves exceptional fusion results by leveraging a novel, more interpretable, and powerful architecture that considers pansharpening as dual, distinct deep sub-semantic branches. It independently reconstructs sub-semantic layer information, leading to improved performance. The DTEMPan architecture incorporates a dual transformer design comprising shared perception encoders and two parallel, effective semantic-level decoders. The hybrid multi-scale texture maintaining decoder and the precise edge maintaining decoder are responsible for reconstructing the general low-frequency and rare high-frequency signals, respectively. Through the integration of complementary information from both decoders, DTEMPan is capable of reconstructing accurate edges and high spatial information while preserving rich spectral details. Extensive experimental evaluations have demonstrated it significantly outperforms state-of-the-art methods on a number of benchmarks. Our code is available at https://github.com/D-Walter/DTEMPan.
Wuzhou Quan, Wen Guo 0003
IEEE Trans. Geosci. Remote. Sens.2
2022 Multi-cue multi-hypothesis tracking with re-identification for multi-object tracking
Wen Guo 0003, Yuelong Jin, Bin Shan, Xinmiao Ding
Multim. Syst.1
2015 Max-Confidence Boosting With Uncertainty for Visual Tracking
abstract
The challenges in visual tracking call for a method which can reliably recognize the subject of interests in an environment, where the appearance of both the background and the foreground change with time. Many existing studies model this problem as tracking by classification with online updating of the classification models, however, most of them overlook the ambiguity in visual modeling and do not consider the prior information in the tracking process. In this paper, we present a novel visual tracking method called max-confidence boosting (MCB), which explores a new way of online updating ambiguous visual phenomenon. The MCB framework models uncertainty in prior knowledge utilizing the indeterministic labels, which are used in updating models from previous frames and the new frame. Our proposed MCB tracker allows ambiguity in the tracking process and can effectively alleviate the drift problem. Many experimental results in challenging video sequences verify the success of our method, and our MCB tracker outperforms a number of the state-of-the-art tracking by classification methods.
Wen Guo 0003, Liangliang Cao, Tony X. Han, Shuicheng Yan, Changsheng Xu
IEEE Trans. Image Process.1
2010 Visual attention based small object segmentation in natual images
abstract
Small object segmentation is a challenging task in image processing and computer vision. In this paper we propose a visual attention based segmentation approach to segment interesting objects with small size in natural images. Different from traditional methods which use the single feature vectors, visual attention analysis is used on local and global features to extract the region of interesting objects. Within the region selected by visual attention analysis, Gaussian Mixture Model (GMM) is applied to further locate the object region. By incorporation of visual attention analysis into object segmentation, the proposed approach is able to narrow the searching region for object segmentation so as to increase the segmentation accuracy and reduce the computational complex. Experimental results demonstrate that the proposed approach is efficient for object segmentation in natural images, especially for small objects. The proposed method outperforms traditional GMM based segmentation significantly.
Wen Guo 0003, Changsheng Xu, Songde Ma, Min Xu 0001
ICIP1