VLDB 2026 Research / reviewers in the wild / expert
Yuxuan Wang 0004
dblp:94/3940-4
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-3889-8560ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | v-HUB: A Benchmark for Video Humor Understanding from Vision and SoundabstractZhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang 0004, Qinrong Cui, Wei Bi, Song-Chun Zhu, Zilong Zheng |
ACL (1) | 4 |
| 2025 | Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingabstractMulti-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal conversations, MMC requires stronger character-centered understanding abilities as there are many interlocutors appearing in both the visual and textual context. To facilitate the study of this problem, we present Friends-MMC in this paper, an MMC dataset that contains 24,000+ unique utterances paired with video context. To explore the character-centered understanding of the dialogue, we also annotate the speaker of each utterance, the names and bounding bboxes of faces that appear in the video. Based on this Friends-MMC dataset, we further study two fundamental MMC tasks: conversation speaker identification and conversation response prediction, both of which have the multi-party nature with the video or image as visual context. For conversation speaker identification, we demonstrate the inefficiencies of existing methods such as pre-trained models, and propose a simple yet effective baseline method that leverages an optimization solver to utilize the context of two modalities to achieve better performance. For conversation response prediction, we fine-tune generative dialogue models on Friend-MMC, and analyze the benefits of speaker information. The code and dataset will be publicly available, and thus we call for more attention on modelling speaker information when understanding conversations. Yueqian Wang, Xiaojun Meng, Yuxuan Wang 0004, Jianxin Liang, Qun Liu 0001, Dongyan Zhao 0001 |
AAAI | 3 |
| 2025 | Probing and Inducing Combinational Creativity in Vision-Language Models
Yongqian Peng, Yuxuan Wang 0004, Yizhou Wang 0001, Chi Zhang 0017, Yixin Zhu 0001, Zilong Zheng |
CogSci | 4 |
| 2025 | OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video ContextsabstractThe rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential, evaluating their real-world interactive capabilities in streaming video contexts remains a formidable challenge. In this work, we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks. Moreover, we propose a novel framework, Multi-modal Multiplexing Modeling (M4), designed to enable an inference-efficient streaming model that can see, listen while generating. Extensive experimental results reveal that the existing MLLMs fall short in interactive streaming understanding, particularly struggling with proactive tasks and multi-turn queries. Our proposed M4, though lightweight, demonstrates a significant improvement in handling proactive tasks and real-time interactions. Yuxuan Wang 0004, Yueqian Wang, Dongyan Zhao 0001, Zilong Zheng |
CVPR | 1 |
| 2025 | VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Yuxuan Wang 0004, Yiqi Song, Cihang Xie, Yang Liu 0003, Zilong Zheng |
ICCV | 1 |
| 2025 | TokenSwift: Lossless Acceleration of Ultra Long Sequence GenerationabstractGenerating ultra-long sequences with large language models (LLMs) has become increasingly crucial but remains a highly time-intensive task, particularly for sequences up to 100K tokens. While traditional speculative decoding methods exist, simply extending their generation limits fails to accelerate the process and can be detrimental. Through an in-depth analysis, we identify three major challenges hindering efficient generation: frequent model reloading, dynamic key-value (KV) management and repetitive generation. To address these issues, we introduce TokenSwift, a novel framework designed to substantially accelerate the generation process of ultra-long sequences while maintaining the target model’s inherent quality. Experimental results demonstrate that TokenSwift achieves over $3 \times$ speedup across models of varying scales (1.5B, 7B, 8B, 14B) and architectures (MHA, GQA). This acceleration translates to hours of time savings for ultra-long sequence generation, establishing TokenSwift as a scalable and effective solution at unprecedented lengths. Junzhe Shen, Zixia Jia, Yuxuan Wang 0004, Zilong Zheng |
ICML | 4 |
| 2025 | Multi-scale Temporal Prediction via Incremental Generation and Multi-agent CollaborationabstractAccurate temporal prediction is the bridge between comprehensive scene understanding and embodied artificial intelligence. However, predicting multiple fine-grained states of scene at multiple temporal scales is difficult for vision-language models.
We formalize the Multi‐Scale Temporal Prediction (MSTP) task in general and surgical scene by decomposing multi‐scale into two orthogonal dimensions: the temporal scale, forecasting states of human and surgery at varying look‐ahead intervals, and the state scale, modeling a hierarchy of states in general and surgical scene. For instance in general scene, states of contacting relationship are finer-grained than states of spatial relationship. For instance in surgical scene, medium‐level steps are finer‐grained than high‐level phases yet remain constrained by their encompassing phase.
To support this unified task, we introduce the first MSTP Benchmark, featuring synchronized annotations across multiple state scales and temporal scales. We further propose a novel method, Incremental Generation and Multi‐agent Collaboration (IG-MC), which integrates two key innovations. Firstly, we propose an plug-and-play incremental generation to keep high-quality temporal prediction that continuously synthesizes up-to-date visual previews at expanding temporal scales to inform multiple decision-making agents, ensuring decision content and generated visuals remain synchronized and preventing performance degradation as look‐ahead intervals lengthen.
Secondly, we propose a decision‐driven multi‐agent collaboration framework for multiple states prediction, comprising generation, initiation, and multi‐state assessment agents that dynamically triggers and evaluates prediction cycles to balance global coherence and local fidelity. Extensive experiments on the MSTP Benchmark in general and surgical scene show that IG‐MC is a generalizable plug-and-play method for MSTP, demonstrating the effectiveness of incremental generation and the stability of decision‐driven multi‐agent collaboration. Zhitao Zeng, Guojian Yuan, Junyuan Mao, Yuxuan Wang 0004, Xiaoshuang Jia, Yueming Jin |
NeurIPS | 4 |
| 2024 | STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question AnsweringabstractRecently we have witnessed the rapid development of video question answering models. However, most models can only handle simple videos in terms of temporal reasoning, and their performance tends to drop when answering temporal-reasoning questions on long and informative videos. To tackle this problem we propose STAIR, a Spatial-Temporal Reasoning model with Auditable Intermediate Results for video question answering. STAIR is a neural module network, which contains a program generator to decompose a given question into a hierarchical combination of several sub-tasks, and a set of lightweight neural modules to complete each of these sub-tasks. Though neural module networks are already widely studied on image-text tasks, applying them to videos is a non-trivial task, as reasoning on videos requires different abilities. In this paper, we define a set of basic video-text sub-tasks for video question answering and design a set of lightweight modules to complete them. Different from most prior works, modules of STAIR return intermediate outputs specific to their intentions instead of always returning attention maps, which makes it easier to interpret and collaborate with pre-trained models. We also introduce intermediate supervision to make these intermediate outputs more accurate. We conduct extensive experiments on several video question answering datasets under various settings to show STAIR's performance, explainability, compatibility with pre-trained models, and applicability when program annotations are not available. Code: https://github.com/yellow-binary-tree/STAIR Yueqian Wang, Yuxuan Wang 0004, Dongyan Zhao 0001 |
AAAI | 2 |
| 2024 | Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding BridgeabstractDespite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the inefficiency in temporal grounding and limited pre-trained context window size. In this work, we introduce Temporal Grounding Bridge (TGB), a novel framework that bootstraps MLLMs with advanced temporal grounding capabilities and broadens their contextual scope. Our framework significantly enhances the temporal capabilities of current MLLMs through three key innovations: an efficient multi-span temporal grounding algorithm applied to low-dimension temporal features projected from flow; a multimodal length extrapolation training paradigm that utilizes low-dimension temporal features to extend the training context window size; and a bootstrapping framework that bridges our model with pluggable MLLMs without requiring annotation. We validate TGB across seven video benchmarks and demonstrate substantial performance improvements compared with prior MLLMs. Notably, our model, initially trained on sequences of four frames, effectively handles sequences up to 16 longer without sacrificing performance, highlighting its scalability and effectiveness in real-world applications. Our code is publicly available. Yuxuan Wang 0004, Yueqian Wang, Pengfei Wu 0003, Jianxin Liang, Dongyan Zhao 0001, Yang Liu 0003, Zilong Zheng |
EMNLP | 1 |
| 2023 | VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsabstractYeah, we're pretty lousy with pens around here, so knock yourself out.Really? Thanks.Here are a few notes on Yuxuan Wang 0004, Zilong Zheng, Xueliang Zhao, Jinpeng Li 0003, Yueqian Wang, Dongyan Zhao 0001 |
ACL (1) | 1 |
| 2023 | Overview of the NLPCC 2023 Shared Task 10: Learn to Watch TV: Multimodal Dialogue Understanding and Response Generation
Yueqian Wang, Yuxuan Wang 0004, Dongyan Zhao 0001 |
NLPCC (3) | 2 |
| 2022 | Overview of the NLPCC 2022 Shared Task: Multi-modal Dialogue Understanding and Generation
Yuxuan Wang 0004, Xueliang Zhao, Dongyan Zhao 0001 |
NLPCC (2) | 1 |