VLDB 2026 Research / reviewers in the wild / expert
Hongzhou Wu
dblp:349/8844
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFS-PINN: A Dynamic Feature Separation Physics-Informed Neural Network
Zhuo Zhang 0020, Wei Wang 0130, Hongzhou Wu, Xi Yang 0020, Canqun Yang |
Comput. Aided Des. | 5 |
| 2025 | TFS: Revisiting Temporal Language Grounding from Frequency Spiking PerspectiveabstractTemporal Language Grounding (TLG) aims to localize moments in untrimmed videos that are most relevant to natural language queries. While existing weakly-supervised methods have achieved significant success in exploring cross-modal relationships, they still face a critical bottleneck: the interference of task-irrelevant information in query embeddings. To address this issue, we propose TLG Frequency Spiking (TFS), a dimensional mask derived from the frequency domain that models the varying importance specific to different queries. By enhancing the understanding of queries, TFS effectively optimizes the cross-modal alignment of visual and textual modalities. Experimental results show that TFS significantly outperforms state-of-the-art baselines on both the Charades-STA and ActivityNet-Captions datasets. Hongzhou Wu, Chuxiong Sun |
ICASSP | 2 |
| 2025 | RBDN: A Robust Background Denoising Network for Weakly Supervised Temporal Language GroundingabstractTemporal Language Grounding (TLG) with weak supervision aims to retrieve events in untrimmed videos corresponding to text queries using only video-text pairs as annotations. However, noisy background semantics in videos lead to inconsistencies in cross-modal alignment, which further hinder the grounding of query events. To address this, we introduce a Robust Background Denoising Network (RBDN), which refines backgrounds and mitigates their negative impact on events. RBDN leverages robust PCA and Frequency Augmentation techniques to filter out spatially static and temporally regular background information, thereby denoising the undesired gradient influence. Extensive experiments on well-known datasets, such as Charades-Sta and ActivityNet-Captions, demonstrate that our approach significantly outperforms current benchmarks. Zehua Zang, Hongzhou Wu, Jiangmeng Li |
ICME | 3 |
| 2024 | Foreground Enhanced Network for Weakly Supervised Temporal Language Grounding
Hongzhou Wu, Xuechen Zhao, Xiang Zhang 0008 |
CogSci | 1 |
| 2024 | MSFR: Stance Detection Based on Multi-Aspect Semantic Feature Representation via Hierarchical Contrastive LearningabstractZero-shot stance detection aims to determine the stance of previously unseen targets during the inference phase. Achieving effective feature alignment from seen targets to unseen targets is crucial for zero-shot stance detection. In this paper, we propose MSFR, a hierarchical contrastive learning framework, which consists of two core components: inter-aspect contrastive learning for distinguishing aspect-level features and intra-aspect contrastive learning for capturing attribute-level features. Specifically, inter-aspect contrastive learning first maps the global features of an utterance to multiple aspects that influence semantic expression (referred to as aspect-level feature differentiation). This process facilitates the alignment of semantic features across different factors of seen and unseen targets. Intra-aspect contrastive learning enhances the distinguishability of features within the same aspect (referred to as attribute-level feature differentiation) and improves the model’s fine-grained generalization capability. Experimental results demonstrate the superior performance of our model compared to competing baseline models. Xuechen Zhao, Feng Xie 0003, Bin Zhou 0004, Hongzhou Wu, Liqun Gao |
ICASSP | 6 |
| 2023 | A Unified Framework for Unseen Target Stance Detection based on Feature Enhancement via Graph Contrastive Learning
Xuechen Zhao, Jiaying Zou, Feng Xie 0003, Hongzhou Wu, Bin Zhou 0004 |
CogSci | 6 |
| 2023 | Atomic-action-based Contrastive Network for Weakly Supervised Temporal Language GroundingabstractAs one knows, an event often consists of several actions while each action is atomic. Inspired by this insight, we propose a novel framework named Atomic-action-based Contrastive Network model (ACN) for weakly supervised temporal language grounding task to localize the query-related event moment in an untrimmed video, without access to any temporal annotations. Specifically, ACN first determines the accurate moment boundary of each action in a query-agnostic way. This can adequately exploit homogeneous visual cues while impeding the heterogeneity of the query from hurting the atomicity of visual action, i.e., action boundary. To effectively localize the query-related event, we seek the discriminative words in the given query, and explore a composite-grained contrastive module to retrieve those corresponding atomic actions in the common latent space across modalities. This boosts feature discrimination of visual event segment to remove irrelevant action video segments. Experiments on two popular datasets show the efficacy of our model. Hongzhou Wu, Xuechen Zhao, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo |
ICME | 1 |
| 2023 | Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to search for a video segment that matches the search intent in a query sentence, which has received increasing attention in recent years, due to its practical values in various fields. Existing efforts devoted to this interesting yet challenging task typically encode the query sentence and video segments into unstructured global representations for cross-modal interaction and fusion, which may fail to accurately capture the search intent in complex queries with multi-granularity semantics. Xiang Zhang 0008, Xun Yang 0001, Yibing Zhan, Long Lan, Jianfeng Dong, Hongzhou Wu |
ACM Multimedia | 7 |