Hongzhou Wu

dblp:349/8844 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DFS-PINN: A Dynamic Feature Separation Physics-Informed Neural Network
Zhuo Zhang 0020, Wei Wang 0130, Hongzhou Wu, Xi Yang 0020, Canqun Yang
Comput. Aided Des.5
2025 TFS: Revisiting Temporal Language Grounding from Frequency Spiking Perspective
abstract
Temporal Language Grounding (TLG) aims to localize moments in untrimmed videos that are most relevant to natural language queries. While existing weakly-supervised methods have achieved significant success in exploring cross-modal relationships, they still face a critical bottleneck: the interference of task-irrelevant information in query embeddings. To address this issue, we propose TLG Frequency Spiking (TFS), a dimensional mask derived from the frequency domain that models the varying importance specific to different queries. By enhancing the understanding of queries, TFS effectively optimizes the cross-modal alignment of visual and textual modalities. Experimental results show that TFS significantly outperforms state-of-the-art baselines on both the Charades-STA and ActivityNet-Captions datasets.
Hongzhou Wu, Chuxiong Sun
ICASSP2
2025 RBDN: A Robust Background Denoising Network for Weakly Supervised Temporal Language Grounding
abstract
Temporal Language Grounding (TLG) with weak supervision aims to retrieve events in untrimmed videos corresponding to text queries using only video-text pairs as annotations. However, noisy background semantics in videos lead to inconsistencies in cross-modal alignment, which further hinder the grounding of query events. To address this, we introduce a Robust Background Denoising Network (RBDN), which refines backgrounds and mitigates their negative impact on events. RBDN leverages robust PCA and Frequency Augmentation techniques to filter out spatially static and temporally regular background information, thereby denoising the undesired gradient influence. Extensive experiments on well-known datasets, such as Charades-Sta and ActivityNet-Captions, demonstrate that our approach significantly outperforms current benchmarks.
Zehua Zang, Hongzhou Wu, Jiangmeng Li
ICME3
2024 Foreground Enhanced Network for Weakly Supervised Temporal Language Grounding
Hongzhou Wu, Xuechen Zhao, Xiang Zhang 0008
CogSci1
2024 MSFR: Stance Detection Based on Multi-Aspect Semantic Feature Representation via Hierarchical Contrastive Learning
abstract
Zero-shot stance detection aims to determine the stance of previously unseen targets during the inference phase. Achieving effective feature alignment from seen targets to unseen targets is crucial for zero-shot stance detection. In this paper, we propose MSFR, a hierarchical contrastive learning framework, which consists of two core components: inter-aspect contrastive learning for distinguishing aspect-level features and intra-aspect contrastive learning for capturing attribute-level features. Specifically, inter-aspect contrastive learning first maps the global features of an utterance to multiple aspects that influence semantic expression (referred to as aspect-level feature differentiation). This process facilitates the alignment of semantic features across different factors of seen and unseen targets. Intra-aspect contrastive learning enhances the distinguishability of features within the same aspect (referred to as attribute-level feature differentiation) and improves the model’s fine-grained generalization capability. Experimental results demonstrate the superior performance of our model compared to competing baseline models.
Xuechen Zhao, Feng Xie 0003, Bin Zhou 0004, Hongzhou Wu, Liqun Gao
ICASSP6
2023 A Unified Framework for Unseen Target Stance Detection based on Feature Enhancement via Graph Contrastive Learning
Xuechen Zhao, Jiaying Zou, Feng Xie 0003, Hongzhou Wu, Bin Zhou 0004
CogSci6
2023 Atomic-action-based Contrastive Network for Weakly Supervised Temporal Language Grounding
abstract
As one knows, an event often consists of several actions while each action is atomic. Inspired by this insight, we propose a novel framework named Atomic-action-based Contrastive Network model (ACN) for weakly supervised temporal language grounding task to localize the query-related event moment in an untrimmed video, without access to any temporal annotations. Specifically, ACN first determines the accurate moment boundary of each action in a query-agnostic way. This can adequately exploit homogeneous visual cues while impeding the heterogeneity of the query from hurting the atomicity of visual action, i.e., action boundary. To effectively localize the query-related event, we seek the discriminative words in the given query, and explore a composite-grained contrastive module to retrieve those corresponding atomic actions in the common latent space across modalities. This boosts feature discrimination of visual event segment to remove irrelevant action video segments. Experiments on two popular datasets show the efficacy of our model.
Hongzhou Wu, Xuechen Zhao, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo
ICME1
2023 Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to search for a video segment that matches the search intent in a query sentence, which has received increasing attention in recent years, due to its practical values in various fields. Existing efforts devoted to this interesting yet challenging task typically encode the query sentence and video segments into unstructured global representations for cross-modal interaction and fusion, which may fail to accurately capture the search intent in complex queries with multi-granularity semantics.
Xiang Zhang 0008, Xun Yang 0001, Yibing Zhan, Long Lan, Jianfeng Dong, Hongzhou Wu
ACM Multimedia7