VLDB 2026 Research / reviewers in the wild / expert
Wenwen Pan 0003
dblp:96/8644-3
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-4300-7694ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene ReconstructionabstractDynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible. In this paper, we propose Sparse4DGS, the first method for sparse-frame dynamic scene reconstruction. We observe that dynamic reconstruction methods fail in both canonical and deformed spaces under sparse-frame settings, especially in areas with high texture richness. Sparse4DGS tackles this challenge by focusing on texture-rich areas. For the deformation network, we propose Texture-Aware Deformation Regularization, which introduces a texture-based depth alignment loss to regulate Gaussian deformation. For the canonical Gaussian field, we introduce Texture-Aware Canonical Optimization, which incorporates texture-based noise into the gradient descent process of canonical Gaussians. Extensive experiments show that when taking sparse frames as inputs, our method outperforms existing dynamic or few-shot techniques on NeRF-Synthetic, HyperNeRF, NeRF-DS, and our iPhone-4D datasets. Changyue Shi, Chuxiao Yang, Wenwen Pan 0003, Jiajun Ding, Zhou Yu 0001, Jun Yu 0002 |
AAAI | 5 |
| 2025 | Growing a Twig to Accelerate Large Vision-Language Models
Zhenwei Shao, Zhou Yu 0001, Wenwen Pan 0003, Hongyuan Zhang 0001, Wei Chen 0001, Jun Yu 0002 |
ICCV | 4 |
| 2025 | Benchmarking and Enhancing Geospatial Visual Reasoning Over Street MapsabstractRecent advances in large multimodal models (LMMs) have enabled substantial progress in various visual question answering (VQA) benchmarks, including the challenging text-centric ones that require a simultaneous understanding of both the visual and textual contents in the images. Despite the prominence of existing text-centric VQA benchmarks, they either have limited textual information or have a limited number of questions requiring complex reasoning skills beyond the basic OCR. To this end, we present SMVQA—a novel text-centric VQA benchmark based on street map images. SMVQA contains more than 10K real-world street map images from the open geospatial database OpenStreetMap. Each image in SMVQA is also associated with detailed geospatial annotations, enabling it to automatically generate up to 57.5K distinctive QA pairs of five representative question types. In addition to the standard test split, SMVQA introduces an extra test split to verify the generalization abilities over out-of-domain images and novel reasoning skills. The evaluation of the state-of-the-art open-source and commercial LMMs reflects the great challenge posed by SMVQA. The latest LMMs, such as GPT-4o, only achieve accuracies of 49.9%, showing plenty of room for improvement. To further improve the latest LMMs’ performance on SMVQA, we introduce a LMM-based agentic framework LHR, which consists of the localizing, highlighting, and reasoning stages. Specifically, LHR first prompts the LMM to localize region-of-interest (RoI) to the question and then highlight the RoI and perform chain-of-thought reasoning for answer prediction. By integrating LHR with GPT-4o, we observe a significant improvement over the vanilla counterpart, showing the effectiveness of our framework. Wenwen Pan 0003, Haiting Zhou, Zhenwei Shao, Shuai Shao 0012, Suguo Zhu, Min Tan 0005, Jun Yu 0002, Zhou Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multi-Granularity Relational Attention Network for Audio-Visual Question AnsweringabstractRecent methods for video question answering (VideoQA), aiming to generate answers based on given questions and video content, have made significant progress in cross-modal interaction. From the perspective of video understating, these existing frameworks concentrate on the various levels of visual content, partially assisted by subtitles. However, audio information is also instrumental in helping get correct answers, especially in videos with real-life scenarios. Indeed, in some cases, both audio and visual contents are required and complement each other to answer questions, which is defined as audio-visual question answering (AVQA). In this paper, we focus on importing raw audio for AVQA and contribute in three ways. Firstly, due to no dataset annotating QA pairs for raw audio, we introduce E-AVQA, a manually annotated and large-scale dataset involving multiple modalities. E-AVQA consists of 34,033 QA pairs on 33,340 clips of 18,786 videos from the e-commerce scenarios. Secondly, we propose a multi-granularity relational attention method with contrastive constraints between audio and visual features after the interaction, named MGN, which captures local sequential representation by leveraging the pairwise potential attention mechanism and obtains global multi-modal representation via designing the novel ternary potential attention mechanism. Thirdly, our proposed MGN outperforms the baseline on dataset E-AVQA, achieving 20.73% on [email protected] and 19.81% on BLEU@1, demonstrating its superiority with at least 1.02 improvement on [email protected] and about 10% on timing complexity over the baseline. Linjun Li, Tao Jin 0004, Hao Jiang 0062, Wenwen Pan 0003, Jian Wang 0119, Shuwen Xiao, Yan Xia 0006, Zhou Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Video Moment Retrieval With Noisy LabelsabstractVideo moment retrieval (VMR) aims to localize the target moment in an untrimmed video according to the given nature language query. The existing algorithms typically rely on clean annotations to train their models. However, making annotations by human labors may introduce much noise. Thus, the video moment retrieval models will not be well trained in practice. In this article, we present a simple yet effective video moment retrieval framework via bottom-up schema, which is in end-to-end manners and robust to noisy label training. Specifically, we extract the multimodal features by syntactic graph convolutional networks and multihead attention layers, which are fused by the cross gates and the bilinear approach. Then, the feature pyramid networks are constructed to encode plentiful scene relationships and capture high semantics. Furthermore, to mitigate the effects of noisy annotations, we devise the multilevel losses characterized by two levels: a frame-level loss that improves noise tolerance and an instance-level loss that reduces adverse effects of negative instances. For the frame level, we adopt the Gaussian smoothing to regard noisy labels as soft labels through the partial fitting. For the instance level, we exploit a pair of structurally identical models to let them teach each other during iterations. This leads to our proposed robust video moment retrieval model, which experimentally and significantly outperforms the state-of-the-art approaches on standard public datasets ActivityCaption and textually annotated cooking scene (TACoS). We also evaluate the proposed approach on the different manual annotation noises to further demonstrate the effectiveness of our model. Wenwen Pan 0003, Zhou Zhao 0001, Wencan Huang, Liyong Fu, Jun Yu 0002, Fei Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | TAVT: Towards Transferable Audio-Visual Text GenerationabstractAudio-visual text generation aims to understand multi-modality contents and translate them into texts.Although various transfer learning techniques of text generation have been proposed, they focused on uni-modal analysis (e.g., text-to-text, visual-to-text) and lack consideration of multi-modal content and cross-modal relation.Motivated by the fact that humans can recognize the timbre of the same low-level concepts (e.g., footstep, rainfall, and laughing), even in different visual conditions, we aim to mitigate the domain discrepancies by audiovisual correlation.In this paper, we propose a novel Transferable Audio-Visual Text Generation framework, named TAVT, which consists of two key components: Audio-Visual Meta-Mapper (AVMM) and Dual Counterfactual Contrastive Learning (DCCL).(1) AVMM first introduces a universal auditory semantic space and drifts the domain-invariant low-level concepts into visual prefixes.Then the reconstructbased learning encourages the AVMM to learn "which pixels belong to the same sound" and achieve audio-enhanced visual prefix.The welltrained AVMM can be further applied to unimodal setting.(2) Furthermore, DCCL leverages the destructive counterfactual transformations to provide cross-modal constraints for AVMM from the perspective of feature distribution and text generation.(3) The experimental results show that TAVT outperforms the stateof-the-art methods across multiple domains (cross-datasets, cross-categories) and various modal settings (uni-modal, multi-modal). Tao Jin 0004, Wenwen Pan 0003, Linjun Li, Xize Cheng, Ye Wang 0018, Zhou Zhao 0001 |
ACL (1) | 3 |
| 2023 | Exploring Group Video Captioning with Efficient Relational ApproximationabstractCurrent video captioning efforts most focus on describing a single video while the need for captioning videos in groups has increased considerably. In this study, we propose a new task, group video captioning, which aims to infer the desired content among a group of target videos and describe it with another group of related reference videos. This task requires the model to effectively summarize the target videos and accurately describe the distinguishing content compared to the reference videos, and it becomes more difficult as the video length increases. To solve this problem, 1) First, we propose an efficient relational approximation (ERA) to identify the shared content among videos while the complexity is linearly related to the number of videos. 2) Then, we introduce a contextual feature refinery with intra-group self-supervision to capture the contextual information and further refine the common properties. 3) In addition, we construct two group video captioning datasets derived from the YouCook2 and the ActivityNet Captions. The experimental results demonstrate the effectiveness of our method on this new task. Tao Jin 0004, Ye Wang 0018, Wenwen Pan 0003, Linjun Li, Xize Cheng, Zhou Zhao 0001 |
ICCV | 4 |
| 2023 | Unsupervised Domain Adaptation for Referring Semantic SegmentationabstractIn this paper, we study the task of referring semantic segmentation in a highly practical setting, in which labeled visual data with corresponding text descriptions are available in the source, but only unlabeled visual data (without text descriptions) are available in the target. It is a challenging task that has many difficulties: (1) how to obtain proper queries for the target domain; (2) how to adapt visual-text joint distribution shifts; (3) how to maintain the original segmentation performance. Thus, we propose a cycle-consistent vision-language matching network to narrow down the domain gap and ease adaptation difficulty. Our model has significant practical applications since they are capable generalising to new data sources without requiring corresponding text annotations. First, a pseudo-text selector is devised to handle the missing modality, through the pre-trained clip model to measure the gap between query features of the source and visual features of the target. Next, a cross-domain segmentation predictor is adopted, which prompts the joint representations to be domain invariant and minimize the discrepancy between two domains. Then, we present a cycle-consistent query matcher to learn discriminative features via reconstructing visual features from masks. Instead of doing the textual comparison, we match the visual features to the pseudo queries. Extensive experiments show the effectiveness of our method. Wenwen Pan 0003, Zhou Zhao 0001, Mingmin Zhang 0001, Fei Wu 0001 |
ACM Multimedia | 2 |
| 2022 | Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross- Modal Denoising NetworksabstractAudio-Guided video object segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from the background in a video sequence according to the referring audio expressions. However, existing referring video object segmentation works mainly focus on the guidance of text-based referring expressions, due to the lack of modeling the semantic representations of audio-video interaction contents. In this paper, we consider the problem of audio-guided video semantic segmentation from the viewpoint of end-to-end denoising encoder-decoder network learning. We propose the wavelet-based encoder network to learn the cross-modal representations of the video contents with audio-form queries. Specifically, we adopt the multi-head cross-modal attention layers to explore the potential relations of video and query contents. A 2-dimension discrete wavelet trans-form is merged into the transformer encoder to decompose the audio-video features. Next, we maximize mutual information between the encoded features and multi-modal features after cross-modal attention layers to enhance the au-dio guidance. Then, a self attention-free decoder network is developed to generate the target masks with frequency-domain transforms. In addition, we construct the first large-scale audio-guided video semantic segmentation dataset. The extensive experiments show the effectiveness of our method11Code is available at: https://github.com/asudahkzj/Wnet.git. Wenwen Pan 0003, Zhou Zhao 0001, Jieming Zhu, Xiuqiang He 0001, Lianli Gao, Jun Yu 0002, Fei Wu 0001, Qi Tian 0001 |
CVPR | 1 |
| 2022 | Seek-and-Hide: Adversarial Steganography via Deep Reinforcement LearningabstractThe goal of image steganography is to hide a full-sized image, termed secret, into another, termed cover. Prior image steganography algorithms can conceal only one secret within one cover. In this paper, we propose an adaptive local image steganography (AdaSteg) system that allows for scale- and location-adaptive image steganography. By adaptively hiding the secret on a local scale, the proposed system makes the steganography more secured, and further enables multi-secret steganography within one single cover. Specifically, this is achieved via two stages, namely the adaptive patch selection stage and secret encryption stage. Given a pair of secret and cover, first, the optimal local patch for concealment is determined adaptively by exploiting deep reinforcement learning with the proposed steganography quality function and policy network. The secret image is then converted into a patch of encrypted noises, resembling the process of generating adversarial examples, which are further encoded to a local region of the cover to realize a more secured steganography. Furthermore, we propose a novel criterion for the assessment of local steganography, and also collect a challenging dataset that is specialized for the task of image steganography, thus contributing to a standardized benchmark for the area. Experimental results demonstrate that the proposed model yields results superior to the state of the art in both security and capacity. Wenwen Pan 0003, Yanling Yin, Xinchao Wang, Yongcheng Jing, Mingli Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Towards Fast and High-Quality Sign Language ProductionabstractSign Language Production (SLP) aims to automatically translate a spoken language description to its corresponding sign language video. The core procedure of SLP is to transform sign gloss intermediaries into sign pose sequences (G2P). Most existing methods for G2P are based on sequential autoregression or sequence-to-sequence encoder-decoder learning. However, by generating target pose frames conditioned on the previously generated ones, these models are prone to bringing issues such as error accumulation and high inference latency. In this paper, we argue that such issues are mainly caused by adopting autoregressive manner. Hence, we propose a novel Non-AuToregressive (NAT) model with a parallel decoding scheme, as well as an External Aligner for sequence alignment learning. Specifically, we extract alignments from the external aligner by monotonic alignment search for gloss duration prediction, which is used by a length regulator to expand the source gloss sequence to match the length of the target sign pose sequence for parallel sign pose generation. Furthermore, we devise a spatial-temporal graph convolutional pose generator in the NAT model to generate smoother and more natural sign pose sequences. Extensive experiments conducted on PHOENIX14T dataset show that our proposed model outperforms state-of-the-art autoregressive models in terms of speed and quality. Wencan Huang, Wenwen Pan 0003, Zhou Zhao 0001, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2020 | Collaboration by Competition: Self-coordinated Knowledge Amalgamation for Multi-talent Student Learning
Sihui Luo 0001, Wenwen Pan 0003, Xinchao Wang, Dazhou Wang, Haihong Tang, Mingli Song |
ECCV (6) | 2 |