VLDB 2026 Research / reviewers in the wild / expert
Difei Gao
dblp:172/9932
· DBLP profile ↗
32ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0001-8494-3492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 20 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EmoAgent: A Multi-Agent Framework for Diverse Affective Image ManipulationabstractAffective Image Manipulation (AIM) aims to alter visual elements within an image to evoke specific emotional responses from viewers. However, existing AIM approaches rely on rigidone-to-onemappings between emotions and visual cues, making them ill-suited for the inherently subjective and diverse ways in which humans perceive and express emotion. To address this, we introduce a novel task setting termedDiverse AIM (D-AIM), aiming to generate multiple visually distinct yet emotionally consistent image edits from a single source image and target emotion. We proposeEmoAgent, the first multi-agent framework tailored specifically for D-AIM. EmoAgent explicitly decomposes the manipulation process into three specialized phases executed by collaborative agents: a Planning Agent that generates diverse emotional editing strategies, an Editing Agent that precisely executes these strategies, and a Critic Agent that iteratively refines the results to ensure emotional accuracy. This collaborative design empowers EmoAgent to modelone-to-manyemotion-to-visual mappings, enabling semantically diverse and emotionally faithful edits. Extensive quantitative and qualitative evaluations demonstrate that EmoAgent substantially outperforms state-of-the-art approaches in both emotional fidelity and semantic diversity, effectively generating multiple distinct visual edits that convey the same target emotion. Qi Mao 0002, Haobo Hu, Yujie She, Difei Gao, Libiao Jin |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | ShowUI: One Vision-Language-Action Model for GUI Visual AgentabstractBuilding Graphical User Interface (GUI) assistants holds significant promise for enhancing human workflow productivity. While most agents are language-based, relying on closed-source API with text-rich meta-information (e.g., HTML or accessibility tree), they show limitations in perceiving UI visuals as humans do, highlighting the need for GUI visual agents. In this work, we develop a vision-language-action model in digital world, namely ShowUI, which features the following innovations: (i) UI-Guided Visual Token Selection to reduce computational costs by formulating screenshots as an UI connected graph, adaptively identifying their redundant relationship and serve as the criteria for token selection during self-attention blocks; (ii) Interleaved Vision-Language-Action Streaming that flexibly unifies diverse needs within GUI tasks, enabling effective management of visual-action history in navigation or pairing multi-turn query-action sequences per screenshot to enhance training efficiency; (iii) Small-scale High-quality GUI Instruction-following Datasets by careful data curation and employing a resampling strategy to address significant data type imbalances. With above components, ShowUI, a lightweight 2B model using 256K data, achieves a strong 75.1% accuracy in zero-shot screenshot grounding. Its UI-guided token selection further reduces 33% of redundant visual tokens during training and speeds up the performance by 1.4×. Navigation experiments across web [12], mobile [35], and online [39] environments further underscore the effectiveness and potential of our model in advancing GUI visual agents. The models are available at https://github.com/showlab/ShowUI. Qinghong Lin, Difei Gao, Zhengyuan Yang, Zechen Bai, Stan Weixian Lei, Zheng Shou 0001 |
CVPR | 3 |
| 2025 | Factorized Learning for Temporally Grounded Video-Language Models
Wenzheng Zeng, Difei Gao, Zheng Shou 0001, Hwee Tou Ng |
ICCV | 2 |
| 2025 | Grounding Multimodal Large Language Model in GUI WorldabstractRecent advancements in Multimodal Large Language Models (MLLMs) have accelerated the development of Graphical User Interface (GUI) agents capable of automating complex tasks across digital platforms. However, precise GUI element grounding remains a key challenge for accurate interaction and generalization. In this work, we present an effective GUI grounding framework, which includes an automated data collection engine that gathers extensive GUI screenshots and annotations to ensure broad generalization. We also propose a lightweight and flexible GUI grounding module designed to efficiently localize UI elements by pre-training on the collected data, and introduce a novel method to integrate this module with MLLMs for the effective execution of GUI tasks. Our approach demonstrates superior performance in task accuracy and adaptability, as validated by benchmarks such as ScreenSpot, MiniWob, AITW, and Mind2Web. Weixian Lei, Difei Gao, Zheng Shou 0001 |
ICLR | 2 |
| 2025 | Can I Trust You? Advancing GUI Task Automation with Action Trust Score
Haiyang Mei, Difei Gao, Xiaopeng Wei, Xin Yang 0011, Zheng Shou 0001 |
ACM Multimedia | 2 |
| 2025 | GUI-Narrator: Detecting and Captioning Computer GUI Actions
Qinchen Wu, Difei Gao, Qinghong Lin, Zhuoyu Wu, Zheng Shou 0001 |
ACM Multimedia | 2 |
| 2025 | Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Junhao Zhang 0001, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao 0001, Lingmin Ran, Yuchao Gu, Difei Gao, Zheng Shou 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | HOVER: Hyperbolic Video-Text RetrievalabstractVideo-text retrieval is a crucial task in numerous computer vision applications. In this paper, we focus on video-text retrieval involving complex action compositions, where a single video encompasses multiple primitive actions such as "sitting up", "opening door", "cooking food", and "eating." Despite the common occurrences in real-world scenarios, such action-compositional videos have received limited research attention, often leading to significant performance degradations in existing retrieval methods. To address this challenge, we present Hyperbolic Video-tExt Retrieval (HOVER), which models the hierarchical semantic relationships between videos and texts by embedding them in a low-dimensional hyperbolic space. Since hyperbolic space provides a geometric prior that naturally aligns with hierarchical data, it allows for more efficient and generalizable representations of video-text semantic hierarchies. HOVER first longitudinally decomposes each video into a hierarchical action tree, where primitive mono-actions are represented as leaf nodes and increasingly complex action compositions as parent nodes. The semantic structures and temporal dependencies of videos/texts are then encoded in hyperbolic space by exploiting hyperbolic distance, norm, and relative cosine similarity. Experimental results show that HOVER significantly outperforms traditional Euclidean-based methods, particularly in scenarios with limited training labels, achieving a notable performance improvement of 28.83%. Additionally, the hyperbolic video-text embeddings learned by HOVER demonstrate strong generalization across new datasets containing videos with varying levels of action complexity. The source code is available at https://github.com/shi-rq/HOVER. Jun Wen 0001, Ruiqi Shi, Wei Ji 0008, Menglin Yang 0001, Difei Gao, Junsong Yuan 0001, Roger Zimmermann |
IEEE Trans. Image Process. | 6 |
| 2024 | VideoLLM-online: Online Video Large Language Model for Streaming VideoabstractRecent Large Language Models (LLMs) have been en-hanced with vision capabilities, enabling them to compre-hend images, videos, and interleaved vision-language con-tent. However, the learning methods of these large multi-modal models (LMMs) typically treat videos as predeter-mined clips, rendering them less effective and efficient at handling streaming video inputs. In this paper, we pro-pose a novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training ob-jective designed to perform language modeling for contin-uous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming di-alogue format, and (3) an optimized inference pipeline to speed up interactive chat in real-world video streams. With our LIVE framework, we develop a simplified model called VideoLLM-online and demonstrate its significant advan-tages in processing streaming videos. For instance, our VideoLLM-online-7B model can operate at over 10 FPS on an A100 GPU for a 5-minute video clip from Ego4D narration. Moreover, VideoLLM-online also showcases state-of-the-art performance on public offline video bench-marks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at showlab.github. iolvideollm-online. Joya Chen, Zhaoyang Lv, Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, Zheng Shou 0001 |
CVPR | 6 |
| 2024 | AssistGUI: Task-Oriented PC Graphical User Interface AutomationabstractGraphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, Assistgui, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose a multi-agent collaboration framework, which incorporates four agents to perform task decomposition, GUI parsing, action generation, and reflection. Our experimental results reveal that our multi-agent collaboration mechanism outshines existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain. Difei Gao, Lei Ji 0001, Zechen Bai, Mingyu Ouyang, Dongxing Mao, Qinchen Wu, Peiyi Wang, Xiangwu Guo, Hengxu Wang, Luowei Zhou, Zheng Shou 0001 |
CVPR | 1 |
| 2024 | VIT-LENS: Towards Omni-modal RepresentationsabstractAiming to advance AI agents, large foundation models significantly improve reasoning and instruction execution, yet the current focus on vision and language neglects the potential of perceiving diverse modalities in open-world environments. However, the success of data-driven vision and language models is costly or even infeasible to be reproduced for rare modalities. In this paper, we present Vit-lens that facilitates efficient omni-modal representation learning by perceiving novel modalities with a pretrained- ViT and aligning them to a pre-defined space. Specifically, the modality-specific lens is tuned to project any-modal signals to an intermediate embedding space, which are then processed by a strong ViT with pre-trained visual knowledge. The encoded representations are optimized toward aligning with the modal-independent space, pre-defined by off-the-shelf foundation models. Vit-lensprovides a unified solution for representation learning of increasing modalities with two appealing advantages: (i) Unlocking the great potential of pretrained- ViTs to novel modalities effectively with efficient parameters and data regime; (ii) Enabling emergent down- stream capabilities through modality alignment and shared ViT parameters. We tailor Vit-lensto learn representations for 3D point cloud, depth, audio, tactile and EEG, and set new state-of-the-art results across various understanding tasks, such as zero-shot classification. By seamlessly integrating Vit-lensinto Multimodal Foundation Models, we enable Any-modality to Text and Image Generation in a zero-shot manner. Code and models are available at https://github.com/TencentARC/ViT-Lens. Weixian Lei, Yixiao Ge, Difei Gao, Dylan Sun 0001, Yuying Ge, Ying Shan, Zheng Shou 0001 |
CVPR | 5 |
| 2024 | Learning Video Context as Interleaved Multimodal Sequences
Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia, Joya Chen, Ziteng Gao, Jinheng Xie, Xuhong Xiao, Zheng Shou 0001 |
ECCV (49) | 3 |
| 2024 | Delocate: Detection and Localization for Deepfake Videos with Randomly-Located Tampered Traces
Xin Liao 0001, Difei Gao, Satoshi Tsutsui, Zheng Qin 0001, Zheng Shou 0001 |
IJCAI | 3 |
| 2024 | AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video CreationabstractGraphical User Interface (GUI) Automation has shown significant potential recently. Previous works built GUI Agent systems to handle short-procedure tasks such as element grounding or functional assistance. In this paper, we propose a novel PC-Copilot, AssistEditor, that focuses on automating the video editing workflow. Unlike previous approaches, our system does not require users to input specific commands to control the computer. Instead, users simply describe their requirements, such as the content and style of the video, and upload the necessary materials. The system then autonomously translates these requirements into detailed actions for controlling video understanding models and professional video editing software, e.g., Premiere Pro to produce the final video. This functionality is enabled by a collaborative AI agent framework of multiple GUI agents, each capable of dialogue, knowledge retrieval, and software usage. These agents have distinct roles, including interacting with users to gather requirements, generating storyboards, and performing editing tasks. This approach significantly streamlines the video editing process, making advanced editing accessible to users with varying levels of expertise. Difei Gao, Zechen Bai, Qinghong Lin, Zheng Shou 0001 |
ACM Multimedia | 1 |
| 2024 | VideoGUI: A Benchmark for GUI Automation from Instructional VideosabstractGraphical User Interface (GUI) automation holds significant promise for enhancing human productivity by assisting with computer tasks. Existing task formulations primarily focus on simple tasks that can be specified by a single, language-only instruction, such as “Insert a new slide.” In this work, we introduce VideoGUI, a novel multi-modal benchmark designed to evaluate GUI assistants on visual-centric GUI tasks. Sourced from high-quality web instructional videos, our benchmark focuses on tasks involving professional and novel software (e.g., Adobe Pho- toshop or Stable Diffusion WebUI) and complex activities (e.g., video editing). VideoGUI evaluates GUI assistants through a hierarchical process, allowing for identification of the specific levels at which they may fail: (i) high-level planning: reconstruct procedural subtasks from visual conditions without language descrip- tions; (ii) middle-level planning: generate sequences of precise action narrations based on visual state (i.e., screenshot) and goals; (iii) atomic action execution: perform specific actions such as accurately clicking designated elements. For each level, we design evaluation metrics across individual dimensions to provide clear signals, such as individual performance in clicking, dragging, typing, and scrolling for atomic action execution. Our evaluation on VideoGUI reveals that even the SoTA large multimodal model GPT4o performs poorly on visual-centric GUI tasks, especially for high-level planning. The data and code are available at https://github.com/showlab/videogui. Qinghong Lin, Difei Gao, Qinchen Wu, Mingyi Yan, Zhengyuan Yang, Zheng Shou 0001 |
NeurIPS | 3 |
| 2024 | LOVA3: Learning to Visual Question Answering, Asking and AssessmentabstractQuestion answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. However, current Multimodal Large Language Models (MLLMs) primarily focus on question answering, often neglecting the full potential of questioning and assessment skills. In this study, we introduce LOVA3, an innovative framework named ``Learning tO Visual Question Answering, Asking and Assessment,'' designed to equip MLLMs with these additional capabilities. Our approach involves the creation of two supplementary training tasks GenQA and EvalQA, aiming at fostering the skills of asking and assessing questions in the context of images. To develop the questioning ability, we compile a comprehensive set of multimodal foundational tasks. For assessment, we introduce a new benchmark called EvalQABench, comprising 64,000 training samples (split evenly between positive and negative samples) and 5,000 testing samples. We posit that enhancing MLLMs with the capabilities to answer, ask, and assess questions
will enhance their multimodal comprehension, ultimately improving overall performance. To validate this hypothesis, we train MLLMs using the LOVA3 framework and evaluate them on a range of multimodal datasets and benchmarks. Our results demonstrate consistent performance gains, underscoring the critical role of these additional tasks in fostering comprehensive intelligence in MLLMs. Hengyuan Zhao, Pan Zhou 0002, Difei Gao, Zechen Bai, Zheng Shou 0001 |
NeurIPS | 3 |
| 2024 | Event Graph Guided Compositional Spatial-Temporal Reasoning for Video Question AnsweringabstractVideo question answering (VideoQA) is challenging since it requires the model to extract and combine multi-level visual concepts from local objects to global actions from complex events for compositional reasoning. Existing works represent the video with fixed-duration clip features that make the model struggle in capturing the crucial concepts in multiple granularities. To overcome this shortcoming, we propose to represent the video with an Event Graph in a hierarchical structure whose nodes correspond to visual concepts of different levels (object, relation, scene and action) and edges indicate their spatial-temporal relationships. We further propose a H ierarchical S patial- T emporal T ransformer (HSTT) which takes nodes from the graph as visual input to realize compositional reasoning guided by the event graph. To fully exploit the spatial-temporal context delivered from the graph structure, on the one hand, we encode the nodes in the order of their semantic hierarchy (depth) and occurrence time (breadth) with our improved graph search algorithm; On the other hand, we introduce edge-guided attention to combine the spatial-temporal context among nodes according to their edge connections. HSTT then performs QA by cross-modal interactions guaranteed by the hierarchical correspondence between the multi-level event graph and the cross-level question. Experiments on the recent challenging AGQA and STAR datasets show that the proposed method clearly outperforms the existing VideoQA models by a large margin, including those pre-trained with large-scale external data. Our code is available at https://github.com/ByZ0e/HSTT. Ziyi Bai, Ruiping Wang 0001, Difei Gao, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA TaskabstractVQA is an ambitious task aiming to answer any image-related question. However, in reality, it is hard to build such a system once for all since the needs of users are continuously updated, and the system has to implement new functions. Thus, Continual Learning (CL) ability is a must in developing advanced VQA systems. Recently, a pioneer work split a VQA dataset into disjoint answer sets to study this topic. However, CL on VQA involves not only the expansion of label sets (new Answer sets). It is crucial to study how to answer questions when deploying VQA systems to new environments (new Visual scenes) and how to answer questions requiring new functions (new Question types). Thus, we propose CLOVE, a benchmark for Continual Learning On Visual quEstion answering, which contains scene- and function-incremental settings for the two aforementioned CL scenarios. In terms of methodology, the main difference between CL on VQA and classification is that the former additionally involves expanding and preventing forgetting of reasoning mechanisms, while the latter focusing on class representation. Thus, we propose a real-data-free replay-based method tailored for CL on VQA, named Scene Graph as Prompt for Symbolic Replay. Using a piece of scene graph as a prompt, it replays pseudo scene graphs to represent the past images, along with correlated QA pairs. A unified VQA model is also proposed to utilize the current and replayed data to enhance its QA ability. Finally, experimental results reveal challenges in CLOVE and demonstrate the effectiveness of our method. Code and data are available at https://github.com/showlab/CLVQA. Stan Weixian Lei, Difei Gao, Jay Zhangjie Wu, Wei Liu 0005, Mengmi Zhang, Zheng Shou 0001 |
AAAI | 2 |
| 2023 | CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingabstractZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, W.k. Chan, Chong-Wah Ngo, Mike Zheng Shou, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhijian Hou, Wanjun Zhong, Lei Ji 0001, Difei Gao, Kun Yan 0004, Wing Kwong Chan, Chong-Wah Ngo, Zheng Shou 0001, Nan Duan 0001 |
ACL (1) | 4 |
| 2023 | Affordance Grounding from Demonstration Video to Target ImageabstractHumans excel at learning from expert demonstrations and solving their own problems. To equip intelligent robots and assistants, such as AR glasses, with this ability, it is essential to ground human hand interactions (i.e., affordances) from demonstration videos and apply them to a target image like a user's AR glass view. This video-to-image affordance grounding task is challenging due to (1) the need to predict fine-grained affordances, and (2) the limited training data, which inadequately covers video-image discrepancies and negatively impacts grounding. To tackle them, we propose Affordance Transformer (Afformer), which has a fine-grained transformer-based decoder that gradually refines affordance grounding. Moreover, we introduce Mask Affordance Hand (MaskAHand), a self-supervised pre-training technique for synthesizing video-image data and simulating context changes, enhancing affordance grounding across video-image discrepancies. Afformer with MaskAHand pre-training achieves state-of-the-art performance on multiple benchmarks, including a sub-stantial 37% improvement on the OPRA dataset. Code is made available at https://github.com/showlab/afformer. Joya Chen, Difei Gao, Qinghong Lin, Zheng Shou 0001 |
CVPR | 2 |
| 2023 | MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringabstractTo build Video Question Answering (VideoQA) systems capable of assisting humans in daily activities, seeking answers from long-form videos with diverse and complex events is a must. Existing multi-modal VQA models achieve promising performance on images or short video clips, especially with the recent success of large-scale multi-modal pre-training. However, when extending these methods to long-form videos, new challenges arise. On the one hand, using a dense video sampling strategy is computationally prohibitive. On the other hand, methods relying on sparse sampling struggle in scenarios where multi-event and multi-granularity visual reasoning are required. In this work, we introduce a new model named$\mathcal{M}ulti{-}$· modal Iterative$\mathcal{S}$.patial-temporal Transformer$(\mathcal{MIST})$) to better adapt pre-trained models for long-form VideoQA. Specifically,$\mathcal{MIST}$decomposes traditional dense spatial-temporal self-attention into cascaded segment and region selection modules that adaptively select frames and image regions that are closely relevant to the question itself. Visual concepts at different granularities are then processed efficiently through an attention module. In addition,$\mathcal{MIST}$iteratively conducts selection and attention over multiple layers to support reasoning over multiple events. The experimental results on four VideoQA datasets, including AGQA, NExT-QA, STAR, and Env-QA, show that$\mathcal{MIST}$achieves state-of-the-art performance and is superior at efficiency. The code is available at github.com/showlab/mist. Difei Gao, Luowei Zhou, Lei Ji 0001, Linchao Zhu, Yi Yang 0001, Zheng Shou 0001 |
CVPR | 1 |
| 2023 | GazeVQA: A Video Question Answering Dataset for Multiview Eye-Gaze Task-Oriented CollaborationsabstractThe usage of exocentric and egocentric videos in Video Question Answering (VQA) is a new endeavor in human-robot interaction and collaboration studies.Particularly for egocentric videos, one may leverage eye-gaze information to understand human intentions during the task.In this paper, we build a novel task-oriented VQA dataset, called GazeVQA, for collaborative tasks where gaze information is captured during the task process.GazeVQA is designed with a novel QA format that covers thirteen different reasoning types to capture multiple aspects of task information and user intent.For each participant, GazeVQA consists of more than 1,100 textual questions and more than 500 labeled images that were annotated with the assistance of the Segment Anything Model.In total, 2,967 video clips, 12,491 labeled images, and 25,040 questions from 22 participants were included in the dataset.Additionally, inspired by the assisting models and common ground theory for industrial task collaboration, we propose a new AI model called AssistGaze that is designed to answer the questions with three different answer types, namely textual, image, and video.AssistGaze can effectively ground the perceptual input into semantic information while reducing ambiguities.We conduct comprehensive experiments to demonstrate the challenges of GazeVQA 1 and the effectiveness of AssistGaze 2 . Muhammet Furkan Ilaslan, Chenan Song, Joya Chen, Difei Gao, Weixian Lei, Qianli Xu, Joo Lim, Zheng Shou 0001 |
EMNLP | 4 |
| 2023 | UniVTG: Towards Unified Video-Language Temporal GroundingabstractVideo Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing on social media. Most methods in this direction develop task-specific models that are trained with type-specific labels, such as moment retrieval (time interval) and highlight detection (worthiness curve), which limits their abilities to generalize to various VTG tasks and labels. In this paper, we propose to Unify the diverse VTG labels and tasks, dubbed UniVTG, along three directions: Firstly, we revisit a wide range of VTG labels and tasks and define a unified formulation. Based on this, we develop data annotation schemes to create scalable pseudo supervision. Secondly, we develop an effective and flexible grounding model capable of addressing each task and making full use of each label. Lastly, thanks to the unified framework, we are able to unlock temporal grounding pretraining from large-scale diverse labels and develop stronger grounding abilities e.g., zero-shot grounding. Extensive experiments on three tasks (moment retrieval, highlight detection and video summarization) across seven datasets (QVHighlights, Charades-STA, TACoS, Ego4D, YouTube Highlights, TVSum, and QFVS) demonstrate the effectiveness and flexibility of our proposed framework. The codes are available at https://github.com/showlab/UniVTG. Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan 0001, Zheng Shou 0001 |
ICCV | 5 |
| 2023 | Learning to Learn: How to Continuously Teach Humans and MachinesabstractCurriculum design is a fundamental component of education. For example, when we learn mathematics at school, we build upon our knowledge of addition to learn multiplication. These and other concepts must be mastered before our first algebra lesson, which also reinforces our addition and multiplication skills. Designing a curriculum for teaching either a human or a machine shares the underlying goal of maximizing knowledge transfer from earlier to later tasks, while also minimizing forgetting of learned tasks. Prior research on curriculum design for image classification focuses on the ordering of training examples during a single offline task. Here, we investigate the effect of the order in which multiple distinct tasks are learned in a sequence. We focus on the online class-incremental continual learning setting, where algorithms or humans must learn image classes one at a time during a single pass through a dataset. We find that curriculum consistently influences learning outcomes for humans and for multiple continual machine learning algorithms across several benchmark datasets. We introduce a novel-object recognition dataset for human curriculum learning experiments and observe that curricula that are effective for humans are highly correlated with those that are effective for machines. As an initial step towards automated curriculum design for online class-incremental learning, we propose a novel algorithm, dubbed Curriculum Designer (CD), that designs and ranks curricula based on inter-class feature similarities. We find significant overlap between curricula that are empirically highly effective and those that are highly ranked by our CD. Our study establishes a framework for further research on teaching humans and machines to learn continuously using optimized curricula. Our code and data are available through this link. Parantak Singh, Ankur Sikarwar, Weixian Lei, Difei Gao, Morgan B. Talbot, Ying Sun 0001, Zheng Shou 0001, Gabriel Kreiman, Mengmi Zhang |
ICCV | 5 |
| 2023 | CRIC: A VQA Dataset for Compositional Reasoning on Vision and CommonsenseabstractAlternatively inferring on the visual facts and commonsense is fundamental for an advanced visual question answering (VQA) system. This ability requires models to go beyond the literal understanding of commonsense. The system should not just treat objects as the entrance to query background knowledge, but fully ground commonsense to the visual world and imagine the possible relationships between objects, e.g., "fork, can lift, food". To comprehensively evaluate such abilities, we propose a VQA benchmark, Compositional Reasoning on vIsion and Commonsense(CRIC), which introduces new types of questions about CRIC, and an evaluation metric integrating the correctness of answering and commonsense grounding. To collect such questions and rich additional annotations to support the metric, we also propose an automatic algorithm to generate question samples from the scene graph associated with the images and the relevant knowledge graph. We further analyze several representative types of VQA models on the CRIC dataset. Experimental results show that grounding the commonsense to the image region and joint reasoning on vision and commonsense are still challenging for current approaches. The dataset is available at https://cricvqa.github.io. Difei Gao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
Difei Gao, Licheng Yu, Weixian Lei, Matt Feiszli, Zheng Shou 0001 |
ECCV (35) | 2 |
| 2022 | AssistQ: Affordance-Centric Question-Driven Task Completion for Egocentric Assistant
Benita Wong, Joya Chen, Stan Weixian Lei, Dongxing Mao, Difei Gao, Zheng Shou 0001 |
ECCV (36) | 6 |
| 2022 | Egocentric Video-Language PretrainingabstractVideo-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP. Qinghong Lin, Jinpeng Wang 0001, Mattia Soldan, Michael Wray, Rui Yan 0001, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu 0005, Zheng Shou 0001 |
NeurIPS | 7 |
| 2021 | Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic EnvironmentsabstractVisual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects’ state changes, which are still challenging for current AI systems. To equip AI system with the ability to understand dynamic ENVironments, we build a video Question Answering dataset named Env-QA. Env-QA contains 23K egocentric videos, where each video is composed of a series of events about exploring and interacting in the environment. It also provides 85K questions to evaluate the ability of understanding the composition, layout, and state changes of the environment presented by the events in videos. Moreover, we propose a video QA model, Temporal Segmentation and Event Attention network (TSEA), which introduces event-level video representation and corresponding attention mechanisms to better extract environment information and answer questions. Comprehensive experiments demonstrate the effectiveness of our framework and show the formidable challenges of Env-QA in terms of long-term state tracking, multi-event temporal reasoning and event counting, etc. Difei Gao, Ruiping Wang 0001, Ziyi Bai, Xilin Chen 0001 |
ICCV | 1 |
| 2020 | Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene TextabstractAnswering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g., names of places, products, and sports teams. To overcome this difficulty, only resorting to pre-trained word embedding models is far from enough. A desired model should utilize the rich information in multiple modalities of the image to help understand the meaning of scene texts, e.g., the prominent text on a bottle is most likely to be the brand. Following this idea, we propose a novel VQA approach, Multi-Modal Graph Neural Network (MM-GNN). It first represents an image as a graph consisting of three sub-graphs, depicting visual, semantic, and numeric modalities respectively. Then, we introduce three aggregators which guide the message passing from one graph to another to utilize the contexts in various modalities, so as to refine the features of nodes. The updated nodes have better features for the downstream question answering module. Experimental evaluations show that our MM-GNN represents the scene texts better and obviously facilitates the performances on two VQA tasks that require reading scene texts. Difei Gao, Kenneth Li 0002, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2017 | Visual Textbook Network: Watch Carefully before Answering Visual Questions
Difei Gao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 1 |
| 2015 | Correlated warped Gaussian processes for gender-specific age estimationabstractFacial age estimation is a challenging problem in computer vision. Existing methods can be classified into two categories: global and person-specific. In practice, the person-specific methods have shown better performance, however it still has some inherit problems such as over learning and mis-assignment of age estimators for unseen facial images. To fix these problems, this paper proposes correlated warped Gaussian processes (CWGP) regression for gender-specific age estimation. It uses two correlated regressors to accurately approximate the gender-specific mapping from facial features to age. Extensive experiments demonstrate the superiority of our method over state-of-the-art methods. Difei Gao, Lili Pan 0001, Risheng Liu, Mei Xie |
ICIP | 1 |