Wenwu Wang 0001

dblp:61/5537-1 · DBLP profile ↗
← Back
9ranked-venue papers in the field
0as first author
5since 2021 · last 2026
0000-0002-8393-5703ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 7Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 STGFMamba: Spatio-temporal graph Fourier-enhanced Mamba for traffic prediction
Xinyuan Zhou, Ruiyi Lu, Zhiang Hou, Yao Ren, Wenwu Wang 0001, Shiyong Lan
Inf. Sci.6
2025 TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
abstract
Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on implicitly modeling audio and visual features through weak labels, without mining semantic relationships for different modalities and explicit modeling of event temporal dependencies. This makes it difficult for the model to accurately parse event information for each segment under weak supervision, especially when high similarity between segmental modal features leads to ambiguous event boundaries. Hence, we propose a multimodal optimization framework, TeMTG, that combines text enhancement and multi-hop temporal graph modeling. Specifically, we leverage pre-trained multimodal models to generate modality-specific text embeddings, and fuse them with audio-visual features to enhance the semantic representation of these features. In addition, we introduce a multi-hop temporal graph neural network, which explicitly models the local temporal relationships between segments, capturing the temporal continuity of both short-term and long-range events. Experimental results demonstrate that our proposed method achieves state-of-the-art (SOTA) performance in multiple key indicators in the LLP dataset.
Yaru Chen 0003, Peiliang Zhang, Fei Li 0022, Faegheh Sardari, Ruohao Guo, Wenwu Wang 0001
ICMR7
2024 Regime Learning for Differentiable Particle Filters
abstract
Differentiable particle filters are an emerging class of models that combine sequential Monte Carlo techniques with the flexibility of neural networks to perform state space inference. This paper concerns the case where the system may switch between a finite set of state-space models, i.e. regimes. No prior approaches effectively learn both the individual regimes and the switching process simultaneously. In this paper, we propose the neural network based regime learning differentiable particle filter (RLPF) to address this problem. We further design a training procedure for the RLPF and other related algorithms. We demonstrate competitive performance compared to the previous state-of-the-art algorithms on a pair of numerical experiments.
John-Joseph Brady, Yuhui Luo, Wenwu Wang 0001, Victor Elvira, Yunpeng Li 0001
FUSION3
2022 UAV-enabled Edge Computing for Optimal Task Distribution in Target Tracking
Shidrokh Goudarzi, Wenwu Wang 0001, Pei Xiao 0001, Lyudmila Mihaylova, Simon J. Godsill
FUSION2
2022 Deep Learning for Audio Visual Emotion Recognition
Tassadaq Hussain, Wenwu Wang 0001, Nidhal Bouaynaya, Hassan M. Fathallah-Shaykh, Lyudmila Mihaylova
FUSION2
2014 Audio-visual tracking of a variable number of speakers with a random finite set approach
Volkan Kilic, Xionghu Zhong, Mark Barnard, Wenwu Wang 0001, Josef Kittler
FUSION4
2014 A Bayesian performance bound for time-delay of arrival based acoustic source tracking in a reverberant environment
Xionghu Zhong, Wenwu Wang 0001, Syed M. Naqvi, Chng Eng Siong
FUSION2
2013 Audio-visual face detection for tracking in a meeting room environment
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
FUSION2
2013 Acoustic source tracking in a reverberant environment using a pairwise synchronous microphone network
Xionghu Zhong, Arash Mohammadi 0001, Wenwu Wang 0001, A. Benjamin Premkumar, Amir Asif
FUSION3