VLDB 2026 Research / reviewers in the wild / expert
Wansen Wu
dblp:275/1821
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-0467-3830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Query Rephrasing for Context Independence in Scene Knowledge-guided Visual GroundingabstractScene Knowledge-guided Visual Grounding (SK-VG) aims to locate the specific object in an image that is referred to by an open-ended query, utilizing textual scene knowledge for guidance. Besides the grounding ability, SK-VG models are expected to infer more information of the target object based on the expatiatory contextual knowledge, given that the queries in SK-VG are often insufficient for the models to identify the object. Existing models perform well when the queries contain obvious visual clues but falter in hard situations where visual attributes are absent from the query. In this paper, we propose a Large Language Models (LLMs) based Inference framework for Scene Knowledge-guided Visual Grounding (LISK-VG) to address such a problem by rephrasing abstract queries with visually distinguishable object descriptions. Specifically, we explicitly decompose the SK-VG task into two steps, i.e., inference and grounding. The inference step takes advantage of the semantic understanding capabilities of LLMs to fully recognize the useful contextual information in the scene knowledge, which is then injected to the query. LISK-VG, in this way, transforms the originally context-dependent query into a more self-informative form that characterizes the target object mainly by the visual features on the given image. This process significantly facilitates the target localization by avoiding joint understanding of the veiled query and the long-winded knowledge text, which might be largely distractive for the grounding task. Then in the second step, the model performs precise grounding based on the converted query with an explicit object description. The experimental results conducted on the SK-VG dataset demonstrate that our LISK-VG significantly outperforms other approaches in terms of accuracy, particularly exhibiting substantial advantages on hard cases. Xilong Qin, Haixiang Zhu, Wansen Wu, Yue Hu 0016 |
IJCNN | 4 |
| 2024 | DAP: Domain-Aware Prompt Learning for Vision-and-Language NavigationabstractFollowing language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents. With strong representation capabilities, pretrained vision-and-language models are widely used in VLN. However, most of them are trained on web-crawled generalpurpose datasets, which incurs a considerable domain gap when used for VLN tasks. To address the problem, we propose a novel and model-agnostic Domain-Aware Prompt learning (DAP) framework. For equipping the pretrained models with specific object-level and scene-level cross-modal alignment in VLN tasks, DAP applies a low-cost prompt tuning paradigm to learn soft visual prompts for extracting in-domain image semantics. Specifically, we first generate a set of in-domain image-text pairs with the help of the CLIP model. Then we introduce soft visual prompts in the input space of the visual encoder in a pretrained model. DAP injects in-domain visual knowledge into the visual encoder of the pretrained model in an efficient way. Experimental results on both R2R and REVERIE show the superiority of DAP compared to existing state-of-the-art methods. Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin |
ICASSP | 3 |
| 2024 | Enhancing Multimodal Sentiment Analysis via Learning from Large Language ModelabstractMultimodal sentiment analysis (MSA) detects human sentiments by understanding data from multiple modalities, such as text and images. Existing research primarily strives for an effective multimodal fusion framework to derive informative representations. However, these methods neglect the necessity of exploiting external knowledge to aid in analyzing sentiments. As a result, the lack of external commonsense embarrasses these models when the opinion cues come in an implicit and obscure manner. To address the limitation, in this paper, we propose an Auxiliary Rationale Knowledge enhanced framework, namely ARK, which improves MSA models via learning from a multimodal large language model (MLLM). Specifically, based on text-image pairs, we employ Chain-of-Thought prompting to generate image descriptions and rationales from the MLLM as auxiliary knowledge, thus enriching the original samples with commonsense knowledge encoded within the MLLM. By combining the source text with image descriptions, we are able to effectively handle MSA through a Text+Text paradigm. In this paradigm, smaller pre-trained language models (LMs) can be tasked for sentiment classification via prompt-tuning. Besides, rationales are leveraged as additional supervision to facilitate the learning of reasoning abilities by LMs. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches across four datasets. Our data and code are available at https://github.com/ningpang/ArkMSA. Ning Pang, Wansen Wu, Yue Hu 0016, Kai Xu 0014, Quanjun Yin, Long Qin 0004 |
ICME | 2 |
| 2024 | PANDA: Prompt-Based Context- and Indoor-Aware Pretraining for Vision and Language Navigation
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin |
MMM (1) | 3 |
| 2024 | ACT: Action-assoCiated and Target-Related Representations for Object Navigation
Youkai Wang, Yue Hu 0016, Wansen Wu, Ting Liu 0018, Yong Peng 0006 |
MMM (1) | 3 |
| 2024 | Vision-language navigation: a survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, Yue Hu 0016 |
Neural Comput. Appl. | 1 |
| 2024 | Visual Grounding With Dual Knowledge DistillationabstractVisual grounding is a task that seeks to predict the specific location of an object or region described by a linguistic expression within an image. Despite the recent success, existing methods still suffer from two problems. First, most methods use independently pre-trained unimodal feature encoders for extracting expressive feature embeddings, thus resulting in a significant semantic gap between unimodal embeddings and limiting the effective interaction of visual-linguistic contexts. Second, existing attention-based approaches equipped with the global receptive field have a tendency to neglect the local information present in the images. This limitation restricts the semantic understanding required to distinguish between referred objects and the background, consequently leading to inadequate localization performance. Inspired by the recent advance in knowledge distillation, in this paper, we propose a DUal knowlEdge disTillation (DUET) method for visual grounding models to bridge the cross-modal semantic gap and improve localization performance simultaneously. Specifically, we utilize the CLIP model as the teacher model to transfer the semantic knowledge to a student model, in which the vision and language modalities are linked into a unified embedding space. Besides, we design a self-distillation method for the student model to acquire localization knowledge by performing the region-level contrastive learning to make the predicted region close to the positive samples. To this end, this work further proposes a Semantics-Location Aware sampling mechanism to generate high-quality self-distillation samples. Extensive experiments on five datasets and ablation studies demonstrate the state-of-the-art performance of DUET and its orthogonality with different student models, thereby making DUET adaptable to a wide range of visual grounding architectures. Our code are available on DUET. Wansen Wu, Meng Cao 0002, Yue Hu 0016, Yong Peng 0006, Long Qin 0004, Quanjun Yin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Dynamic Multi-modal Prompting for Efficient Visual Grounding
Wansen Wu, Ting Liu 0018, Youkai Wang, Kai Xu 0014, Quanjun Yin, Yue Hu 0016 |
PRCV (7) | 1 |
| 2022 | The emergence of collective obstacle avoidance based on a visual perception mechanism
Jingtao Qi, Yandong Xiao, Yingmei Wei, Wansen Wu |
Inf. Sci. | 5 |
| 2021 | Generation and Extraction Combined Dialogue State Tracking with Hierarchical Ontology IntegrationabstractRecently, the focus of dialogue state tracking has expanded from single domain to multiple domains.The task is characterized by the shared slots between domains.As the scenario gets more complex, the out-of-vocabulary problem also becomes more severe.Current models are not satisfactory for addressing the challenges of ontology integration between domains and out-of-vocabulary problems.To address the problem, we explore the hierarchical semantics of the ontology and enhance the interrelation between slots with masked hierarchical attention.In state value decoding stage, we address the out-of-vocabulary problem by combining generation method and extraction method together.We evaluate the performance of our model on two representative datasets, MultiWOZ in English and CrossWOZ in Chinese.The results show that our model yields a significant performance gain over current state-of-the-art state tracking model and it is more robust to out-of-vocabulary problem compared with other methods. Xinmeng Li, Wansen Wu, Quanjun Yin |
EMNLP (1) | 3 |