Qunbo Wang

dblp:228/4336 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories
abstract
Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to simulated or off-street environments, and often rely on precise goal formats, such as specific coordinates or images. This limits their effectiveness for autonomous agents like last-mile delivery robots navigating unfamiliar cities. To address these limitations, we introduce UrbanNav, a scalable framework that trains embodied agents to follow free-form language instructions in diverse urban settings. Leveraging web-scale city walking videos, we develop an scalable annotation pipeline that aligns human navigation trajectories with language instructions grounded in real-world landmarks. UrbanNav encompasses over 1,500 hours of navigation data and 3 million instruction-trajectory-landmark triplets, capturing a wide range of urban scenarios. Our model learns robust navigation policies to tackle complex urban scenarios, demonstrating superior spatial reasoning, robustness to noisy instructions, and generalization to unseen urban settings. Experimental results show that UrbanNav significantly outperforms existing methods, highlighting the potential of large-scale web video data to enable language-guided, real-world urban navigation for embodied agents.
Yanghong Mei, Yirong Yang, Longteng Guo, Qunbo Wang, Ming-Ming Yu, Xingjian He, Wenjun Wu 0001, Jing Liu 0001
AAAI4
2026 A real-time multi-Automated Guided Vehicles scheduling approach with long-term planning under persistent resource contention
Qunbo Wang, Runmei Li, Junsheng Wu
Eng. Appl. Artif. Intell.2
2026 Agents Trainer: Automatically Training Multi-Agent Reinforcement Learning Models for Drone Swarm Using Language Model-Based Agents
Jiabin Lou, Rongye Shi, Ming-Ming Yu, Yuanshuai Wang, Qunbo Wang, Wenjun Wu 0001
IEEE Trans Autom. Sci. Eng.6
2026 TAS-DAQ: Task-Adaptive Sparse Prediction With Dense Query Auxiliary Supervisory for Efficient 3D Object Detection
abstract
Detecting 3D objects from surround-view images focuses on capturing the spatio-temporal positions of the surrounding environment, serving as a pivotal capability for vision-centric autonomous driving and robotics. While existing approaches primarily employ either dense BEV queries or sparse 3D queries, both paradigms have inherent limitations: dense queries suffer from redundant feature interactions and optimization conflicts, while sparse queries rely on high-quality initialization and struggle with error propagation in complex scenarios. To address these challenges, we proposeTAS-DAQ, a novel two-stage framework that synergizes dense and sparse query strategies. In Stage I, we generate geometry-aware coarse queries through the BEV feature providing robust initialization, thereby ensuring robust query initialization with explicit 3D priors. Stage II introduces a learnable Query Bank with temporal fusion to iteratively refine sparse queries by capturing discriminative instance features across views and frames. Moreover, considering the optimization conflicts caused by redundant query interactions in dense paradigms, we introduce adaptive query aggregation in the query bank that dynamically prioritizes high-confidence queries from BEV features, effectively addressing query error propagation while enhancing instance-level representation consistency. Extensive experiments on the nuScenes R50 benchmark demonstrate state-of-the-art performance, achieving56.9 % NDSand46.1% mAP.
Yirong Yang, Qunbo Wang, Longteng Guo, Ruyi Ji, Ming-Ming Yu, Wenjun Wu 0001, Jing Liu 0001
IEEE Trans. Multim.3
2025 COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional components such as external knowledge bases or map information to enhance performance. These additions, while boosting performance, also lead to larger models and increased computational costs. In this paper, to achieve both high performance and low computational costs, we propose a novel architecture with the COmbination of Selective MemOrization (COSMO). Specifically, COSMO integrates state-space modules and transformer modules, and incorporates two VLN-customized selective state space modules: the Round Selective Scan (RSS) and the Cross-modal Selective State Space Module (CS3). RSS facilitates comprehensive inter-modal interactions within a single scan, while the CS3 module adapts the selective state space module into a dual-stream architecture, thereby enhancing the acquisition of cross-modal interactions. Experimental validations on three mainstream VLN benchmarks, REVERIE, R2R, and R2R-CE, not only demonstrate competitive navigation performance of our model but also show a significant reduction in computational costs.
Yanyuan Qiao, Qunbo Wang, Zike Yan, Qi Wu 0001, Zhihua Wei 0001, Jing Liu 0001
ICCV3
2025 C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
abstract
Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the real-world requirement for continual adaptation to evolving scenarios. To facilitate related studies, we introduce the continual object navigation benchmark, which requires agents to acquire navigation skills for new object categories while avoiding catastrophic forgetting of previously learned knowledge. To tackle this challenge, we propose C-Nav, a continual visual navigation framework that integrates two key innovations: (1) A dual-path anti-forgetting mechanism, which comprises feature distillation that aligns multi-modal inputs into a consistent representation space to ensure representation consistency, and feature replay that retains temporal features within the action decoder to ensure policy consistency. (2) An adaptive sampling strategy that selects diverse and informative experiences, thereby reducing redundancy and minimizing memory overhead. Extensive experiments across multiple model architectures demonstrate that C-Nav consistently outperforms existing approaches, achieving superior performance even compared to baselines with full trajectory retention, while significantly lowering memory requirements. The code will be publicly available at \url{https://bigtree765.github.io/C-Nav-project}.
Mingming Yu, Fei Zhu 0004, Wenzhuo Liu, Yirong Yang, Qunbo Wang, Wenjun Wu 0001, Jing Liu 0001
NeurIPS5
2025 GroundingMate: Aiding Object Grounding for Goal-Oriented Vision-and-Language Navigation
abstract
Goal-Oriented Vision-and-Language Navigation (VLN) aims to enable agents to navigate to specified locations and identify designated target objects following natural language instruction. This approach has gained popularity due to its close alignment with real-world scenarios. However, existing studies have predominantly focused on enhancing navigation performance, neglecting the ability to locate objects at the navigation endpoint. This oversight has resulted in a significant discrepancy between the success rates of navigation and object grounding. The challenge is compounded by the complex reasoning required by the instructions and the necessity to synthesize multiperspective images of objects, which overwhelms traditional object grounding methods. We leverage the Multi-Modal Large Language Model (MLLM) to bridge this gap, allowing agents to seek assistance from these models when struggling to locate the target object. The agent conducts a multi-stage evaluation to discern the cause of its confusion and promptly extracts and updates the most relevant information for MLLM to assess. Our method is plug-and-play and model-agnostic, facilitating integration with numerous existing VLN strategies without the need for retraining. Implementing our approach across four distinct methods has improved performance on the REVERIE and SOON datasets, demonstrating the effectiveness and generalizability of our technique.
Qianyi Liu, Yanyuan Qiao, Junyou Zhu, Longteng Guo, Qunbo Wang, Xingjian He, Qi Wu 0001, Jing Liu 0001
WACV7
2025 Spatial-temporal context-aware network for 3D-Craft generation
Ruyi Ji, Qunbo Wang, Boying Wang, Hangu Zhang, Yanni Wang
Appl. Intell.2
2025 TagRec: Temporal-Aware Graph Contrastive Learning With Theoretical Augmentation for Sequential Recommendation
abstract
Sequential recommendation systems aim to predict the future behaviors of users based on their historical interactions. Despite the success of neural architectures like Transformer and Graph Neural Networks, these models often struggle with the inherent challenge of sparse data in accurately predicting future user behaviors. To alleviate the data sparsity problem, some methods leverage the contrastive learning to generate contrastive views, assuming the items appear discretely at the same time intervals and focusing on the sequence order. However, these approaches neglect the crucial temporal-aware collaborative patterns hidden within the user-item interactions, leading to a limited variety of contrastive pairs and less informative embeddings. The proposed framework,Temporal-awaregraph contrastive learning with theoretical guarantees for sequentialRecommendation (TagRec), integrates temporal-aware collaborative patterns with adaptive data augmentation to generate more informative user and item representations. TagRec employs a temporal-aware graph neural network to embed the original graph, then generates augmented graphs through the addition of interactions via latent user interest mining, the dropping of redundant interaction edges, and the perturbation of temporal information. Theoretical guarantees are provided that these augmentations enhance the graph’s utility. Extensive experiments on real-world datasets demonstrate the superiority of the proposed approach over the state-of-the-art recommendation methods.
Tianhao Peng 0002, Haitao Yuan 0002, Yuchen Li 0006, Peihong Dai, Qunbo Wang, Senzhang Wang, Wenjun Wu 0001
IEEE Trans. Knowl. Data Eng.6
2025 FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
Yanyuan Qiao, Qunbo Wang, Longteng Guo, Zhihua Wei 0001, Jing Liu 0001
IEEE Trans. Multim.3
2024 Soft Knowledge Prompt: Help External Knowledge Become a Better Teacher to Instruct LLM in Knowledge-based VQA
abstract
LLM has achieved impressive performance on multi-modal tasks, which have received everincreasing research attention.Recent research focuses on improving prediction performance and reliability (e.g., addressing the hallucination problem).They often prepend relevant external knowledge to the input text as an extra prompt.However, these methods would be affected by the noise in the knowledge and the context length limitation of LLM.In our work, we focus on making better use of external knowledge and propose a method to actively extract valuable information in the knowledge to produce the latent vector as a soft prompt, which is then fused with the image embedding to form a knowledge-enhanced context to instruct LLM.The experimental results on knowledge-based VQA benchmarks show that the proposed method enjoys better utilization of external knowledge and helps the model achieve better performance.
Qunbo Wang, Ruyi Ji, Tianhao Peng 0002, Wenjun Wu 0001, Zechao Li, Jing Liu 0001
ACL (1)1
2024 Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
abstract
While large visual-language models (LVLM) have shown promising results on traditional visual question answering benchmarks, it is still challenging for them to answer complex VQA problems which requires diverse world knowledge.Motivated by the research of retrievalaugmented generation in the field of natural language processing, we use Dense Passage Retrieval (DPR) to retrieve related knowledge to help the model answer questions.However, DPR conduct retrieving in natural language space, which may not ensure comprehensive acquisition of image information.Thus, the retrieved knowledge is not truly conducive to helping answer the question, affecting the performance of the overall system.To address this issue, we propose a novel framework that leverages the visual-language model to select the key knowledge retrieved by DPR and answer questions.The framework consists of two modules: Selector and Answerer, where both are initialized by the LVLM and parameterefficiently finetuned by self-bootstrapping: find key knowledge in the retrieved knowledge documents using the Selector, and then use them to finetune the Answerer to predict answers; obtain the pseudo-labels of key knowledge documents based on the predictions of the Answerer and weak supervision labels, and then finetune the Selector to select key knowledge; repeat.Our framework significantly enhances the performance of the baseline on the challenging open-domain Knowledge-based VQA benchmark, OK-VQA, achieving a state-ofthe-art accuracy of 62.83%.Our code is publicly available at https://github.com/ haodongze/Self-KSel-QAns.
Dongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang 0016, Jing Liu 0001
EMNLP2
2024 Semantic-Visual Graph Reasoning for Visual Dialog
abstract
Visual dialog (VisDial) requires models to answer questions based on both the dialog history and image contents. Traditional approaches simply extract relevant visual and textual information for answering questions, ignoring the relationships between the entities in the dialog and the relationships between the objects in the image. These fine-grained information is crucial to help correctly answer the questions in VisDial. In this work, we propose a Semantic-Visual Graph Reasoning framework (SVG) for VisDial. Specifically, we first construct a semantic graph to capture the semantic relationships between different entities in the current question and the dialog history. Secondly, we construct a semantics-aware visual graph to capture high-level visual semantics including key objects of the image and their visual relationships. Extensive experimental results on the VisDial v0.9 and v1.0 show that our method has shown superior performance compared to the state-of-the-art models across most evaluation metrics.
Dongze Hao, Qunbo Wang, Jing Liu 0001
ICME2
2024 Coordinating explicit and implicit knowledge for knowledge-based VQA
Qunbo Wang, Jing Liu 0001, Wenjun Wu 0001
Pattern Recognit.1
2024 HCCL: Hierarchical Counterfactual Contrastive Learning for Robust Visual Question Answering
abstract
Despite most state-of-the-art models having achieved amazing performance in Visual Question Answering (VQA) , they usually utilize biases to answer the question. Recently, some studies synthesize counterfactual training samples to help the model to mitigate the biases. However, these synthetic samples need extra annotations and often contain noises. Moreover, these methods simply add synthetic samples to the training data to train the model with the cross-entropy loss, which cannot make the best use of synthetic samples to mitigate the biases. In this article, to mitigate the biases in VQA more effectively, we propose a Hierarchical Counterfactual Contrastive Learning (HCCL) method. Firstly, to avoid introducing noises and extra annotations, our method automatically masks the unimportant features in original pairs to obtain positive samples and create mismatched question-image pairs as negative samples. Then our method uses feature-level and answer-level contrastive learning to make the original sample close to positive samples in the feature space, while away from negative samples in both feature and answer spaces. In this way, the VQA model can learn the robust multimodal features and focus on both visual and language information to produce the answer. Our HCCL method can be adopted in different baselines, and the experimental results on VQA v2, VQA-CP, and GQA-OOD datasets show that our method is effective in mitigating the biases in VQA, which improves the robustness of the VQA model.
Dongze Hao, Qunbo Wang, Jing Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Deep Bayesian Active Learning for Learning to Rank: A Case Study in Answer Selection (Extended Abstract)
abstract
Active learning can select informative data for model training to reduce the amount of labelling efforts required. Because traditional active learning methods cannot be directly used for deep learning, researchers have proposed multiple deep active learning methods. However, none of the previous research efforts on deep active learning algorithms presents a specific framework for learning-to-rank tasks. In this work, we introduce a novel deep active learning framework based on Deep Expected Loss Optimization (DELO) for the answer selection task.
Qunbo Wang, Wenjun Wu 0001, Yuxing Qi, Yongchi Zhao
ICDE1
2023 VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
abstract
Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audio, and Subtitle, and Text by exploring an automatically generated large-scale omni-modality video caption dataset called VAST-27M. Specifically, we first collect 27 million open-domain video clips and separately train a vision and an audio captioner to generate vision and audio captions. Then, we employ an off-the-shelf Large Language Model (LLM) to integrate the generated captions, together with subtitles and instructional prompts into omni-modality captions. Based on the proposed VAST-27M dataset, we train an omni-modality video-text foundational model named VAST, which can perceive and process vision, audio, and subtitle modalities from video, and better support various tasks including vision-text, audio-text, and multi-modal video-text tasks (retrieval, captioning and QA). Extensive experiments have been conducted to demonstrate the effectiveness of our proposed VAST-27M corpus and VAST foundation model. VAST achieves 22 new state-of-the-art results on various cross-modality benchmarks.
Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Jing Liu 0001
NeurIPS3
2022 Deep Bayesian Active Learning for Learning to Rank: A Case Study in Answer Selection
abstract
Given a question and a set of candidate answers, answer selection is the task of identifying the best answer, which can be viewed as a kind of learning-to-rank tasks. Learning to rank arises in many information retrieval applications, where deep learning models can achieve inspiring results. Training a deep learning model often requires large scale annotated data that are expensive and time-consuming to obtain. Active learning presents a promising approach to this problem by selecting more informative training data to reduce the amount of labelling efforts required. Because traditional active learning methods cannot be directly used for deep learning, researchers have proposed multiple deep active learning methods. However, none of the previous research efforts on deep active learning algorithms presents a specific framework for learning-to-rank tasks. In this work, we introduce a novel deep active learning framework based onDeepExpectedLossOptimization (DELO) for the answer selection task. It adopts a data acquisition function based on model uncertainty with Bayesian deep learning and the expected loss optimization. Moreover, a two-step batch-mode procedure, combining DELO and other data acquisition strategies is proposed to further improve the performance of active learning. Experimental results verify the effectiveness of the proposed framework.
Qunbo Wang, Wenjun Wu 0001, Yuxing Qi, Yongchi Zhao
IEEE Trans. Knowl. Data Eng.1
2021 Combining Label-wise Attention and Adversarial Training for Tag Prediction of Web Services
abstract
Tagging is well regarded as one of the best ways of managing web services, in which keywords are assigned by users to describe the published services. As users are required to select multiple tags from a large set of candidate tags based on their own understanding, such user-attached tags are not always reliable and may affect the efficiency of service discovery. To alleviate the issue, tag prediction can suggest users appropriate tags for web services based on the textual descriptions of their functionality. Therefore, it is necessary to design tag prediction methods to support service search and recommendation. In this work, we propose a tag prediction model that adopts BERT-based label-wise attention mechanism, and use adversarial training to further improve the model performance. Experimental results on the service datasets collected from ProgrammableWeb show that the proposed method can achieve better prediction performance than other state-of-art methods.
Qunbo Wang, Wenjun Wu 0001, Yongchi Zhao, Yuzhang Zhuang, Yanni Wang
ICWS1
2021 Graph active learning for GCN-based zero-shot classification
Qunbo Wang, Wenjun Wu 0001, Yongchi Zhao, Yuzhang Zhuang
Neurocomputing1
2020 Combination of Active Learning and Self-Paced Learning for Deep Answer Selection with Bayesian Neural Network
abstract
Answer Selection is an important subtask of Question Answering tasks. For this learning-to-rank problem, deep learning methods have outperformed traditional methods. To train a high-quality deep answer selection model, it often requires large amounts of labeled data, which is a costly and noise-prone process. Active learning and semi-supervised learning are usually applied in the modelling training procedure to achieve optimal accuracy with fewer labeled training samples. However, traditional active learning methods rely on good uncertainty estimates that are hard to obtain with standard neural networks. And the performance of semi-supervised learning methods are always affected adversely by the quality of the pseudo-labeled data. In this work, we propose a new framework integrating active learning and self-paced learning in training deep answer selection models. This framework proposes an uncertainty quantification method based on Bayesian neural network, which can guide active learning and self-paced learning in the same iterative process of model training. Experiments were conducted on two kinds of deep answer selection models with real-world datasets including YahooCQA and SemiEvalCQA. The results reveal that the proposed method can significantly reduce the labeled samples for model training.
Qunbo Wang, Wenjun Wu 0001, Yuxing Qi, Zhimin Xin
ECAI1