VLDB 2026 Research / reviewers in the wild / expert
Junbin Xiao
dblp:242/4518
· DBLP profile ↗
35ranked-venue papers
9as first author
32since 2021 · last 2026
0000-0001-5573-6195ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 22 · 8 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sycophancy in vision-language models: A systematic analysis and an inference-time mitigation framework
Yunpu Zhao, Rui Zhang 0040, Junbin Xiao, Changxin Ke, Ruibo Hou, Yifan Hao 0001, Ling Li 0001 |
Neurocomputing | 3 |
| 2026 | ADVersa: Abductive Driving Accident Video UnderstandingabstractUnderstanding traffic accident scenes is a long-standing research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video-based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model, which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis, and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding. Lei-Lei Li, Jianwu Fang, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Scene-Text Grounding for Text-Based Video Question AnsweringabstractExisting efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decision-making and heavy reliance on scene-text recognition. In this paper, we propose to studyGrounded TextVideoQAby forcing models to answer questions and spatio-temporally localize the relevant scene-text regions, thus decoupling QA from scene-text recognition and promoting research towards interpretable QA. The task has three-fold significance. First, it encourages scene-text evidence versus other short-cuts for answer predictions. Second, it directly accepts scene-text regions as visual answers, thus circumventing the problem of ineffective answer evaluation by stringent string matching. Third, it isolates the challenges inherited in VideoQA and scene-text recognition. This enables the diagnosis of the root causes for failure predictions,e.g., wrong QA or wrong scene-text recognition? To achieve Grounded TextVideoQA, we propose the T2S-QA model that highlights a disentangled temporal-to-spatial contrastive learning strategy for weakly-supervised scene-text grounding and grounded TextVideoQA. To facilitate evaluation, we construct a new datasetViTXT-GQAwhich features 52Kscene-text bounding boxes within 2.2Ktemporal segments related to 2Kquestions and 729 videos. With ViTXT-GQA, we perform extensive experiments and demonstrate the severe limitations of existing techniques in Grounded TextVideoQA. While T2S-QA achieves superior results, the large performance gap with human leaves ample space for improvement. Our further analysis of oracle scene-text inputs posits that the major challenge is scene-text recognition. To advance the research of Grounded TextVideoQA, our dataset and code are athttps://github.com/zhousheng97/ViTXT-GQA.git Junbin Xiao, Xun Yang 0001, Peipei Song, Dan Guo 0001, Angela Yao, Meng Wang 0001, Tat-Seng Chua |
IEEE Trans. Multim. | 2 |
| 2025 | On the Consistency of Video Large Language Models in Temporal ComprehensionabstractVideo large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency –a key indicator for robustness and trustworthiness of temporal grounding. After the model identifies an initial moment within the video content, we apply a series of probes to check if the model’s responses align with this initial grounding as an indicator of reliable comprehension. Our results reveal that current Video-LLMs are sensitive to variations in video contents, language queries, and task settings, unveiling severe deficiencies in maintaining consistency. We further explore common prompting and instruction-tuning methods as potential solutions, but find that their improvements are often unstable. To that end, we propose event temporal verification tuning that explicitly accounts for consistency, and demonstrate significant improvements for both grounding and consistency. Our data and code are open-sourced at https://github.com/minjoong507/Consistency-of-Video-LLM. Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang, Angela Yao |
CVPR | 2 |
| 2025 | EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question AnsweringabstractWe introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are designed to elicit identification and reasoning on scene text in an egocentric and dynamic environment. With EgoTextVQA, we comprehensively evaluate 10 prominent multimodal large language models. Currently, all models struggle, and the best results (Gemini 1.5 Pro) are around 33% accuracy, highlighting the severe deficiency of these techniques in egocentric QA assistance. Our further investigations suggest that precise temporal grounding and multi-frame reasoning, along with high resolution and auxiliary scene-text inputs, are key for better performance. With thorough analyses and heuristic suggestions, we hope EgoTextVQA can serve as a solid testbed for research in egocentric scene-text QA assistance. Our dataset is released at: https://github.com/zhousheng97/EgoTextVQA. Junbin Xiao, Qingyun Li, Yicong Li 0004, Xun Yang 0001, Dan Guo 0001, Meng Wang 0001, Tat-Seng Chua, Angela Yao |
CVPR | 2 |
| 2025 | Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object Categories
Yicong Li 0004, Zhenyuan Ma, Junbin Xiao, Xiang Wang 0010, Angela Yao |
ICCV | 4 |
| 2025 | Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
ICCV | 3 |
| 2025 | Visual Intention Grounding for Egocentric Assistants
Pengzhan Sun 0001, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 0004, Arjun R. Akula, Angela Yao |
ICCV | 2 |
| 2025 | Unleashing the Power of LLMs for Medical Video Answer Localization
Junbin Xiao, Qingyun Li, Yusen Yang, Liang Qiu 0002, Angela Yao |
MICCAI (7) | 1 |
| 2025 | Bottom-Up and Top-Down Thoughts for Visual Intention GroundingabstractRIO (Reasoning Intention-oriented Object) is a visual grounding task aimed at locating object within an image that best matches a given intention. Due to the implicit referential nature of the intention descriptions and the inherent complexity of the scenes depicted in images, existing methods struggle to effectively accomplish this task.In this paper, we propose a novel training-free approach that decouples visual processing and reasoning to effectively solve this task. Our approach comprises two complementary pipelines: a top-down pipeline that first performs textual reasoning before proceeding to visual processing, and a bottom-up pipeline that initially conducts visual processing followed by textual reasoning. We then employ a meticulously designed strategy to merge the outputs of these parallel pipelines, thereby achieving complementarity. Experimental results demonstrate that our approach attains performance levels comparable to those of fine-tuned visual grounding models on the RIO dataset. Additional experiments conducted on the SKVG dataset further attest to the generalizability of our method. Kangcheng Liu, Junbin Xiao, Rui Zhang 0040, Hanqi Lv, Zidong Du |
ICMR | 2 |
| 2025 | Video Question Answering and BeyondabstractVideo Question Answering (Video QA) has emerged as a central task in multimodal learning. This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers. We begin with an introduction to VideoQA preliminaries, tracing how methods have adapted from third-person view short videos to capture egocentric and long-ranged spatial-temporal dynamics. We then focus on the impact of large multimodal models. Next, we expand the scope to spatial understanding beyond videos. Each topic is discussed through the lens of tasks, datasets, methods, and evaluation protocols. Finally, we conclude with future directions, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration. This tutorial aims to equip participants with both a historical perspective and a forward-looking roadmap for advancing Video QA in the LLM era. Yicong Li 0004, Junbin Xiao, Angela Yao, Tat-Seng Chua |
ACM Multimedia | 2 |
| 2025 | Can Multimodal Large Language Models Understand Human Values in Videos?abstractHuman values are core principles to determine what is right, desirable, and important for individuals and societies. The deep integration of large language models (LLMs) into human life has facilitated their remarkable performance in value understanding. Multimodal content is rich in value-laden information. The development of multimodal large language models (MLLMs) provides new perspectives for multimodal value understanding. However, MLLMs’ capacity for understanding basic human values in the video domain remains underexplored. To bridge this gap, our work focuses on evaluating their ability to understand specific human values embedded in videos. We assess 11 advanced MLLMs based on the VVALUES video dataset, using a four-module strategy designed to answer two questions: Do MLLMs understand human values, and Do their design and training elements impact the performance? Based on the evaluation results, we derive valuable findings across eight aspects that provide deeper insights into the value understanding of MLLMs and offer useful guides for future research in this field. Junbin Xiao, Zhulin Tao, Libiao Jin |
MMAsia | 4 |
| 2025 | EgoBlind: Towards Egocentric Visual Assistance for the BlindabstractWe present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It also features 5,311 questions directly posed or verified by the blind to reflect their in-situation needs for visual assistance. Each question has an average of 3 manually annotated reference answers to reduce subjectiveness.Using EgoBlind, we comprehensively evaluate 16 advanced MLLMs and find that all models struggle. The best performers achieve an accuracy near 60\%, which is far behind human performance of 87.4\%. To guide future advancements, we identify and summarize major limitations of existing MLLMs in egocentric visual assistance for the blind and explore heuristic solutions for improvement. With these efforts, we hope that EgoBlind will serve as a foundation for developing effective AI assistants to enhance the independence of the blind and visually impaired. Data and code are available at \url{https://github.com/doc-doc/EgoBlind}. Junbin Xiao, Nanxin Huang, Zhulin Tao, Xun Yang 0001, Richang Hong, Meng Wang 0001, Angela Yao |
NeurIPS | 1 |
| 2025 | Question-Answering Dense Video Events
Hangyu Qin, Junbin Xiao, Angela Yao |
SIGIR | 2 |
| 2025 | VideoQA in the Era of LLMs: An Empirical Study
Junbin Xiao, Nanxin Huang, Hangyu Qin, Yicong Li 0004, Fengbin Zhu, Zhulin Tao, Jianxing Yu, Tat-Seng Chua, Angela Yao |
Int. J. Comput. Vis. | 1 |
| 2025 | Transformer-Empowered Invariant Grounding for Video Question AnsweringabstractVideo Question Answering (VideoQA) is the task of answering questions about a video. At its core is the understanding of the alignments between video scenes and question semantics to yield the answer. In leading VideoQA models, the typical learning objective, empirical risk minimization (ERM), tends to over-exploit the spurious correlations between question-irrelevant scenes and answers, instead of inspecting the causal effect of question-critical scenes, which undermines the prediction with unreliable reasoning. In this work, we take a causal look at VideoQA and propose a modal-agnostic learning framework, named Invariant Grounding for VideoQA (IGV), to ground the question-critical scene, whose causal relations with answers are invariant across different interventions on the complement. With IGV, leading VideoQA models are forced to shield the answering from the negative influence of spurious correlations, which significantly improves their reasoning ability. To unleash the potential of this framework, we further provide a Transformer-Empowered Invariant Grounding for VideoQA (TIGV), a substantial instantiation of IGV framework that naturally integrates the idea of invariant grounding into a transformer-style backbone. Experiments on four benchmark datasets validate our design in terms of accuracy, visual explainability, and generalization ability over the leading baselines. Our code is available at https://github.com/yl3800/TIGV. Yicong Li 0004, Xiang Wang 0010, Junbin Xiao, Wei Ji 0008, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | LASO: Language-Guided Affordance Segmentation on 3D ObjectabstractSegmenting affordance in 3D data is key for bridging perception and action in robots. Existing efforts mostly focus on the visual side and overlook the affordance knowledge from a semantic aspect. This oversight not only limits their generalization to unseen objects, but more importantly, hinders their synergy with large language models (LLMs) which are excellent task planners that can decompose an overarching command into agent-actionable instructions. With this regard, we propose a novel task, Language-guided Affordance Segmentation on 3D Object (LASO), which challenges a model to segment a 3D object's part relevant to a given affordance question. To facilitate the task, we contribute a dataset comprising 19,751 point-question pairs, covering 8434 object shapes and 870 expert-crafted questions. As a pioneer solution, we further propose PointRefer, which highlights an adaptive fusion module to identify target affordance regions at different scales. To ensure a text-aware segmentation, we adopt a set of affordance queries conditioned on linguistic cues to generate dynamic kernels. These kernels are further used to convolute with point features and generate a segmentation mask. Comprehensive experiments and analyses validate PointRefer's effectiveness. With these efforts, We hope that LASO can steer the direction of 3D affordance, guiding it towards enhanced integration with the evolving capabilities of LLMs. Code and data are available at https://github.com/yl3800/LASO. Yicong Li 0004, Na Zhao 0004, Junbin Xiao, Xiang Wang 0010, Tat-Seng Chua |
CVPR | 3 |
| 2024 | Abductive Ego-View Accident Video Understanding for Safe Driving PerceptionabstractWe present MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild ego-view accident videos, each with temporally aligned text descriptions. We annotate over 2.23 mil-lion object boxes and 58,650 pairs of video-based accident reasons, covering 58 accident categories. MM-AU supports various accident understanding tasks, particularly multimodal video diffusion to understand accident cause-effect chains for safe driving. With MM-AU, we present an Abductive accident Video unders tanding framework for Safe Driving perception (AdVersa-SD). AdVersa-SD performs video diffusion via an Object-Centric Video Diffusion (OAVD) method which is driven by an abductive CLIP model. This model involves a contrastive interaction loss to learn the pair co-occurrence of normal, near-accident, accident frames with the corresponding text descriptions, such as accident reasons, prevention advice, and accident categories. OAVD enforces the object region learning while fixing the content of the original frame background in video generation, to find the dominant objects for certain accidents. Extensive experiments verify the abductive ability of AdVersa-SD and the superiority of OAVD against the state-of-the-art diffusion models. Additionally, we provide care-ful benchmark evaluations for object detection and accident reason answering since AdVersa-SD relies on precise object and accident reason information. Jianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
CVPR | 4 |
| 2024 | Can I Trust Your Answer? Visually Grounded Video Question AnsweringabstractWe study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding. Specifically, by forcing vision-language models (VLMs) to answer questions and simultaneously provide visual evidence, we seek to ascertain the extent to which the predictions of such techniques are genuinely anchored in relevant video content, versus spurious correlations from language or irrelevant visual context. Towards this, we construct NExT-GQA - an extension of NExT-QA with 10.5K temporal grounding (or location) labels tied to the original QA pairs. With NExT-GQA, we scrutinize a series of state-of-the-art VLMs. Through post-hoc attention analysis, we find that these models are extremely weak in substantiating the answers despite their strong QA performance. This exposes the limitation of current VLMs in making reliable predictions. As a remedy, we further explore and propose a grounded-QA method via Gaussian mask optimization and cross-modal learning. Experiments with different backbones demonstrate that this grounding mechanism improves both grounding and QA. With these efforts, we aim to push towards trustworthy VLMs in VQA systems. Our dataset and code are available at https://github.com/doc-doc/NExT-GQA. Junbin Xiao, Angela Yao, Yicong Li 0004, Tat-Seng Chua |
CVPR | 1 |
| 2024 | Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question AnsweringabstractIn Large Language Models (LLMs), text generation that involves knowledge representation is often fraught with the risk of "hallucinations'', where models confidently produce erroneous or fabricated content. These inaccuracies often stem from intrinsic biases in the pre-training stage or from the incorporation of human preference biases during the fine-tuning process. To mitigate these issues, we take inspiration from Goldman's causal theory of knowledge, which asserts that knowledge is not merely about having a true belief but also involves a causal connection between the belief and the truth of the proposition. We instantiate this theory within the context of Knowledge Question Answering (KQA) by constructing a causal graph that delineates the pathways between the candidate knowledge and belief. Through the application of the do-calculus rules from structural causal models, we devise an unbiased estimation framework based on this causal graph, thereby establishing a methodology for knowledge modeling grounded in causal inference. The resulting CORE framework (short for "Causal knOwledge REasoning'') is comprised of four essential components: question answering, causal reasoning, belief scoring, and refinement. Together, they synergistically improve the KQA system by fostering faithful reasoning and introspection. Extensive experiments are conducted on ScienceQA and HotpotQA datasets, which demonstrate the effectiveness and rationality of the CORE framework. Jiawei Wang 0025, Da Cao, Shaofei Lu, Zhanchang Ma, Junbin Xiao, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2023 | FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video PlatformsabstractShort video platforms have become an important channel for news sharing, but also a new breeding ground for fake news. To mitigate this problem, research of fake news video detection has recently received a lot of attention. Existing works face two roadblocks: the scarcity of comprehensive and largescale datasets and insufficient utilization of multimodal information. Therefore, in this paper, we construct the largest Chinese short video dataset about fake news named FakeSV, which includes news content, user comments, and publisher profiles simultaneously. To understand the characteristics of fake news videos, we conduct exploratory analysis of FakeSV from different perspectives. Moreover, we provide a new multimodal detection model named SV-FEND, which exploits the cross-modal correlations to select the most informative features and utilizes the social context information for detection. Extensive experiments evaluate the superiority of the proposed method and provide detailed comparisons of different methods and modalities for future works. Our dataset and codes are available in https://github.com/ICTMCG/FakeSV. Peng Qi 0005, Yuyan Bu, Juan Cao 0001, Wei Ji 0008, Ruihao Shui, Junbin Xiao, Danding Wang, Tat-Seng Chua |
AAAI | 6 |
| 2023 | Discovering Spatio-Temporal Rationales for Video Question AnsweringabstractThis paper strives to solve complex video question answering (VideoQA) which features long video containing multiple objects and events at different time. To tackle the challenge, we highlight the importance of identifying question-critical temporal moments and spatial objects from the vast amount of video content. Towards this, we propose a Spatio-Temporal Rationalization (STR), a differentiable selection module that adaptively collects question-critical moments and objects using cross-modal interaction. The discovered video moments and objects are then served as grounded rationales to support answer reasoning. Based on STR, we further propose TranSTR, a Transformerstyle neural network architecture that takes STR as the core and additionally underscores a novel answer interaction mechanism to coordinate STR for answer decoding. Experiments on four datasets show that TranSTR achieves new state-of-the-art (SoTA). Especially, on NExT-QA and Causal-VidQA which feature complex VideoQA, it significantly surpasses the previous SoTA by 5.8% and 6.8%, respectively. We then conduct extensive studies to verify the importance of STR as well as the proposed answer interaction mechanism. With the success of TranSTR and our comprehensive analysis, we hope this work can spark more future efforts in complex VideoQA. Code will be released at https://github.com/yl3800/TranSTR. Yicong Li 0004, Junbin Xiao, Xiang Wang 0010, Tat-Seng Chua |
ICCV | 2 |
| 2023 | Deconfounded Multimodal Learning for Spatio-temporal Video GroundingabstractThe task of spatio-temporal video grounding involves identifying the spatial and temporal regions in a video that correspond to the objects or actions described in a given textual description. However, current models used for spatio-temporal video grounding often rely heavily on spatio-temporal priors to make the predictions. As a result, they may suffer from spurious correlations and lack the ability to generalize well to new or diverse scenarios. To overcome this limitation, we introduce a deconfounded multimodal learning framework, which utilizes a structural causal model to treat dataset biases as a confounder and subsequently remove their confounding effect. Through this framework, we can perform causal intervention on the multimodal input and derive an unbiased estimation formula through the do-calculus technique. In order to tackle the challenge of diverse and often unobservable confounders, we further propose a novel retrieval-based approach with a causal mask mechanism. The proposed method leverages analogical reasoning to facilitate deconfounded learning and mitigate dataset biases, enabling unbiased spatio-temporal prediction without explicitly modeling the confounding factors. Extensive experiments on two challenging benchmarks have well verified the effectiveness and rationality of our proposed solution. Jiawei Wang 0025, Zhanchang Ma, Da Cao, Yuquan Le, Junbin Xiao, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2023 | Contrastive Video Question Answering via Video Graph TransformerabstractWe propose to perform video question answering (VideoQA) in a Contrastive manner via a Video Graph Transformer model (CoVGT). CoVGT's uniqueness and superiority are three-fold: 1) It proposes a dynamic graph transformer module which encodes video by explicitly capturing the visual objects, their relations and dynamics, for complex spatio-temporal reasoning. 2) It designs separate video and text transformers for contrastive learning between the video and text to perform QA, instead of multi-modal transformer for answer classification. Fine-grained video-text communication is done by additional cross-modal interaction modules. 3) It is optimized by the joint fully- and self-supervised contrastive objectives between the correct and incorrect answers, as well as the relevant and irrelevant questions respectively. With superior video encoding and QA solution, we show that CoVGT can achieve much better performances than previous arts on video reasoning tasks. Its performances even surpass those models that are pretrained with millions of external data. We further show that CoVGT can also benefit from cross-modal pretraining, yet with orders of magnitude smaller data. The results demonstrate the effectiveness and superiority of CoVGT, and additionally reveal its potential for more data-efficient pretraining. Junbin Xiao, Pan Zhou 0002, Angela Yao, Yicong Li 0004, Richang Hong, Shuicheng Yan, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringabstractVideo question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been focused on designing sophisticated cross-modal interactions to fuse the information from two modalities, while encoding the video and question holistically as frame and word sequences. Despite their success, these methods are essentially revolving around the sequential nature of video- and question-contents, providing little insight to the problem of question-answering and lacking interpretability as well. In this work, we argue that while video is presented in frame sequence, the visual elements (e.g., objects, actions, activities and events) are not sequential but rather hierarchical in semantic space. To align with the multi-granular essence of linguistic concepts in language queries, we propose to model video as a conditional graph hierarchy which weaves together visual facts of different granularity in a level-wise manner, with the guidance of corresponding textual cues. Despite the simplicity, our extensive experiments demonstrate the superiority of such conditional hierarchical graph architecture, with clear performance improvements over prior methods and also better generalization across different type of questions. Further analyses also demonstrate the model's reliability as it shows meaningful visual-textual evidences for the predicted answers. Junbin Xiao, Angela Yao, Zhiyuan Liu 0001, Yicong Li 0004, Wei Ji 0008, Tat-Seng Chua |
AAAI | 1 |
| 2022 | Invariant Grounding for Video Question AnsweringabstractVideo Question Answering (VideoQA) is the task of an-swering questions about a video. At its core is understanding the alignments between visual scenes in video and linguistic semantics in question to yield the answer. In leading VideoQA models, the typical learning objective, empirical risk minimization (ERM), latches on superficial correlations between video-question pairs and answers as the alignments. However, ERM can be problematic, because it tends to over-exploit the spurious correlations between question-irrelevant scenes and answers, instead of inspecting the causal effect of question-critical scenes. As a result, the VideoQA models suffer from unreliable reasoning. In this work, we first take a causal look at VideoQA and argue that invariant grounding is the key to ruling out the spurious correlations. Towards this end, we propose a new learning framework, Invariant Grounding for VideoQA (IGV), to ground the question-critical scene, whose causal relations with answers are invariant across different interventions on the complement. With IGV, the VideoQA mod-els are forced to shield the answering process from the negative influence of spurious correlations, which significantly improves the reasoning ability. Experiments on three benchmark datasets validate the superiority of IGV in terms of accuracy, visual explainability, and generalization ability over the leading baselines. Our code is available at https://github.com/y13800/IGV. Yicong Li 0004, Xiang Wang 0010, Junbin Xiao, Wei Ji 0008, Tat-Seng Chua |
CVPR | 3 |
| 2022 | Video Graph Transformer for Video Question Answering
Junbin Xiao, Pan Zhou 0002, Tat-Seng Chua, Shuicheng Yan |
ECCV (36) | 1 |
| 2022 | Video Question Answering: Datasets, Algorithms and ChallengesabstractThis survey aims to organize the recent advances in video question answering (VideoQA) and point towards future directions.We firstly categorize the datasets into: 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, and 2) factoid VideoQA and inference VideoQA, according to the technical challenges in comprehending the questions and deriving the correct answers.We then summarize the VideoQA techniques, including those mainly designed for Factoid QA (such as the early spatio-temporal attention-based methods and the recent Transformer-based ones) and those targeted at explicit relation and logic inference (such as neural modular networks, neural symbolic methods, and graph-structured methods).Aside from the backbone techniques, we also delve into specific models and derive some common and useful insights either for video modeling, question answering, or for cross-modal correspondence learning.Finally, we present the research trends of studying beyond factoid VideoQA to inference VideoQA, as well as towards the robustness and interpretability.Additionally, we maintain a repository, https://github.com/VRU-NExT/ VideoQA, to keep trace of the latest VideoQA papers, datasets, and their open-source implementations if available.With these efforts, we strongly hope this survey could shed light on the follow-up VideoQA research. Yaoyao Zhong, Wei Ji 0008, Junbin Xiao, Yicong Li 0004, Weihong Deng, Tat-Seng Chua |
EMNLP | 3 |
| 2022 | Equivariant and Invariant Grounding for Video Question AnsweringabstractVideo Question Answering (VideoQA) is the task of answering the natural language questions about a video. Producing an answer requires understanding the interplay across visual scenes in video and linguistic semantics in question. However, most leading VideoQA models work as black boxes, which make the visual-linguistic alignment behind the answering process obscure. Such black-box nature calls for visual explainability that reveals "What part of the video should the model look at to answer the question?". Only a few works present the visual explanations in a post-hoc fashion, which emulates the target model's answering process via an additional method. Yicong Li 0004, Xiang Wang 0010, Junbin Xiao, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2021 | NExT-QA: Next Phase of Question-Answering to Explaining Temporal ActionsabstractWe introduce NExT-QA, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions. Based on the dataset, we set up multi-choice and open-ended QA tasks targeting causal action reasoning, temporal action reasoning, and common scene comprehension. Through extensive analysis of baselines and established VideoQA techniques, we find that top-performing methods excel at shallow scene descriptions but are weak in causal and temporal action reasoning. Furthermore, the models that are effective on multi-choice QA, when adapted to open-ended QA, still struggle in generalizing the answers. This raises doubt on the ability of these models to reason and highlights possibilities for improvement. With detailed results for different question types and heuristic observations for future works, we hope NExT-QA will guide the next generation of VQA research to go beyond superficial description towards a deeper understanding of videos. (The dataset and related resources are available at https://github.com/doc-doc/NExT-QA.git). Junbin Xiao, Xindi Shang, Angela Yao, Tat-Seng Chua |
CVPR | 1 |
| 2021 | VidVRD 2021: The Third Grand Challenge on Video Relation DetectionabstractACM Multimedia 2021 Video Relation Understanding Challenge is the third grand challenge which aims at exploring the relationship of subjects and objects appearing in videos for fine-grained and high-level video understanding. Given a video, the video relation detection model should output a serious of relation triplet subject, predicate, object and the corresponding trajectories of subject and object. The goal of this task is to promote research on developing video semantic understanding model, so as to perform complex inferences and mining of visual knowledge in videos. In this paper, we make a comprehensive and detailed introduction of this task, conclude the proposed algorithms in the last few years, and propose future direction for research in this task. Wei Ji 0008, Yicong Li 0004, Xindi Shang, Junbin Xiao, Tongwei Ren, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2021 | Video Visual Relation Detection via Iterative InferenceabstractThe core problem of video visual relation detection (VidVRD) lies in accurately classifying the relation triplets, which comprise of the classes of subject and object entities, and the predicate classes of various relationships between them. Existing VidVRD approaches classify these three relation components in either independent or cascaded manner, thus fail to fully exploit the inter-dependency among them. In order to utilize this inter-dependency in tackling the challenges of visual relation recognition in videos, we propose a novel iterative relation inference approach for VidVRD. We derive our model from the viewpoint of joint relation classification which is light-weight yet effective, and propose a training approach to better learn the dependency knowledge from the likely correct triplet combinations. As such, the proposed inference approach is able to gradually refine each component based on its learnt dependency and the other two's predictions. Our ablation studies show that this iterative relation inference can empirically converge in a few steps and consistently boost the performance over baselines. Further, we incorporate it into a newly designed VidVRD architecture, named VidVRD-II (Iterative Inference), which generalizes well across different datasets. Experiments show that VidVRD-II achieves the start-of-the-art performance on both of ImageNet-VidVRD and VidOR benchmark datasets. Xindi Shang, Yicong Li 0004, Junbin Xiao, Wei Ji 0008, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2020 | Visual Relation Grounding in Videos
Junbin Xiao, Xindi Shang, Xun Yang 0001, Sheng Tang, Tat-Seng Chua |
ECCV (6) | 1 |
| 2019 | Annotating Objects and Relations in User-Generated VideosabstractUnderstanding the objects and relations between them is indispensable to fine-grained video content analysis, which is widely studied in recent research works in multimedia and computer vision. However, existing works are limited to evaluating with either small datasets or indirect metrics, such as the performance over images. The underlying reason is that the construction of a large-scale video dataset with dense annotation is tricky and costly. In this paper, we address several main issues in annotating objects and relations in user-generated videos, and propose an annotation pipeline that can be executed at a modest cost. As a result, we present a new dataset, named VidOR, consisting of 10k videos (84 hours) together with dense annotations that localize 80 categories of objects and 50 categories of predicates in each video. We have made the training and validation set public and extendable for more tasks to facilitate future research on video object and relation recognition. Xindi Shang, Donglin Di, Junbin Xiao, Xun Yang 0001, Tat-Seng Chua |
ICMR | 3 |
| 2019 | Relation Understanding in Videos: A Grand Challenge OverviewabstractACM Multimedia 2019 Video Relation Understanding Challenge is the first grand challenge aiming at pushing video content analysis at the relational and structural level. This year, the challenge asks the participants to explore and develop innovative algorithms to detect object entities and their relations based on a large-scale user-generated video dataset. The tasks will advance the foundation of future visual systems that are able to perform complex inferences. This paper presents an overview of the grand challenge, including background, detailed descriptions of the three proposed tasks, the corresponding datasets for training, validation and testing, and the evaluation process. Xindi Shang, Junbin Xiao, Donglin Di, Tat-Seng Chua |
ACM Multimedia | 2 |