EDBT 2026 Demo / reviewers in the wild / expert
Fabian Caba Heilbron
dblp:144/5671 · also Fabian Caba, Fabian David Caba Heilbron
· DBLP profile ↗
38ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-3129-1985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 24 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 19 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CineVerse: Consistent Keyframe Synthesis for Cinematic Scene CompositionabstractMulti-shot generation requires preserving the identity of characters and settings across frames. Cinematic scene composition goes beyond standard multi-shot generation, introducing additional challenges such as expressing complex interactions among multiple characters and visual effects to convey creative narratives—challenges existing datasets cannot fully address. We present CineVerse, a large-scale dataset of diverse movie scenes labeled with shot-level annotations tailored for filmmaking. CineVerse includes refined scene descriptions, shot-type information, and newly extracted shot, character, setting descriptions. We validate our dataset by developing a baseline framework that first generates a scene plan containing detailed information for the overall scene and each individual shot, then produces a set of coherent keyframes. Our results show significant improvements in controlling and synthesizing cinematic content through the added context provided by CineVerse. Quynh Phung, Long Mai, Fabian Caba Heilbron, Feng Liu 0015, Jia-Bin Huang 0001, Cusuh Ham |
WACV | 3 |
| 2026 | VRAgent: Self-Refining Agent for Zero-Shot Multimodal Video RetrievalabstractRecent advances in Vision-Language Models (VLMs) and Large Language Models (LLMs) have demonstrated remarkable zero-shot capabilities while becoming increasingly accessible. They have enabled zero-shot text-to-video search, yet they still struggle on real-world queries that demand temporally aligned reasoning across vision and speech. We present VRAgent, an agentic retrieval framework that leverages a central LLM as a planner which (i) decomposes a free-form user query into a tool-instruction set spanning visual, dialogue and other modality-specific retrievers, and (ii) iteratively self-refines this plan by scoring the outputs and rewriting the next tool-instruction set. The resulting closed-loop optimization acts at test time and requires no additional training data or gradient updates.VRAgent is modular by design—adding a new modality is as simple as adding a corresponding foundation model into the toolbox. On our newly proposed MM-MSRVTT and TVR-1200 multimodal benchmarks, VRAgent improves average recall by +8.3% and +1.7% over the best zero-shot baselines, while on single-modality MSR-VTT and DiDeMo it obtains consistent gains of +3.7% and +4.1%. An interactive variant that asks the user up to two multiple-choice questions pushes average recall to 79.7% on MSR-VTT, underscoring the value of on-the-fly human feedback. Ketul Shah, Pankaj Nathani, Rama Chellappa, Fabian Caba Heilbron |
WACV | 4 |
| 2025 | Improving Personalized Search with Regularized Low-Rank Parameter UpdatesabstractPersonalized vision-language retrieval seeks to recognize new concepts (e.g., “my dog Fido”) from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the concept in different contexts. In this paper, we show how to effectively adapt the internal representation of a vision-language dual encoder model for personalized vision-language retrieval. We find that regularized low-rank adaption of a small set of parameters in the language encoder’s final layer serves as a highly effective alternative to textual inversion for recognizing the personal concept while preserving general knowledge. Additionally, we explore strategies for combining parameters of multiple learned personal concepts, finding that parameter addition is effective. To evaluate how well general knowledge is preserved in a finetuned representation, we introduce a metric that measures image retrieval accuracy based on captions generated by a vision language model (VLM). Our approach achieves state-of-the-art accuracy on two benchmarks for personalized image retrieval with natural language queries – DeepFashion2 and ConCon-Chi – outperforming the prior art by 4% − 22% on personal retrievals. Fiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman, James M. Rehg, Bryan C. Russell |
CVPR | 3 |
| 2025 | Discovering Divergent Representations Between Text-To-Image Models
Lisa Dunlap, Joseph Gonzalez 0001, Trevor Darrell, Fabian Caba Heilbron, Josef Sivic, Bryan C. Russell |
ICCV | 4 |
| 2025 | ResidualViT for Efficient Temporally Dense Video EncodingabstractSeveral video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" reasoning over frames sampled at high temporal resolution. However, computing frame-level features for these tasks is computationally expensive given the temporal resolution requirements. In this paper, we make three contributions to reduce the cost of computing features for temporally dense tasks. First, we introduce a vision transformer (ViT) architecture, dubbed ResidualViT, that leverages the large temporal redundancy in videos to efficiently compute temporally dense frame-level features. Our architecture incorporates (i) learnable residual connections that ensure temporal consistency across consecutive frames and (ii) a token reduction module that enhances processing speed by selectively discarding temporally redundant information while reusing weights of a pretrained foundation model. Second, we propose a lightweight distillation strategy to approximate the frame-level features of the original foundation model. Finally, we evaluate our approach across four tasks and five datasets, in both zero-shot and fully supervised settings, demonstrating significant reductions in computational cost (up to 60%) and improvements in inference speed (up to 2.5x faster), all while closely approximating the accuracy of the original foundation model. Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic, Bryan C. Russell |
ICCV | 2 |
| 2024 | VideoMap: Supporting Video Exploration, Brainstorming, and Prototyping in the Latent SpaceabstractVideo editing is a creative and complex endeavor and we believe that there is potential for reimagining a new video editing interface to better support the creative and exploratory nature of video editing. We take inspiration from latent space exploration tools that help users find patterns and connections within complex datasets. We present VideoMap, a proof-of-concept video editing interface that operates on video frames projected onto a latent space. We support intuitive navigation through map-inspired navigational elements and facilitate transitioning between different latent spaces through swappable lenses. We built three VideoMap components to support editors in three common video tasks. In a user study with both professionals and non-professionals, editors found that VideoMap helps reduce grunt work, offers a user-friendly experience, provides an inspirational way of editing, and effectively supports the exploratory nature of video editing. We further demonstrate the versatility of VideoMap by implementing three extended applications. For interactive examples, we invite you to visit our project page: https://chuanenlin.com/videomap. David Chuan-En Lin, Fabian Caba Heilbron, Joon-Young Lee, Oliver Wang, Nikolas Martelaro |
Creativity & Cognition | 2 |
| 2024 | Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a PriorabstractThis paper investigates the challenge of extracting highlight moments from videos. To perform this task, we need to understand what constitutes a highlight for arbitrary video domains while at the same time being able to scale across different domains. Our key insight is that photographs taken by photographers tend to capture the most remarkable or photogenic moments of an activity. Drawing on this insight, we present Videogenic, a technique capable of creating domain-specific highlight videos for a diverse range of domains. In a human evaluation study (N=50), we show that a high-quality photograph collection combined with CLIP-based retrieval (which uses a neural network with semantic knowledge of images) can serve as an excellent prior for finding video highlights. In a within-subjects expert study (N=12), we demonstrate the usefulness of Videogenic in helping video editors create highlight videos with lighter workload, shorter task completion time, and better usability. David Chuan-En Lin, Fabian Caba Heilbron, Joon-Young Lee, Oliver Wang, Nikolas Martelaro |
Creativity & Cognition | 2 |
| 2024 | Scaling Up Video Summarization Pretraining with Large Language ModelsabstractLong-form video content constitutes a significant portion of internet traffic, making automated video summarization an essential research problem. However, existing video summarization datasets are notably limited in their size, constraining the effectiveness of state-of-the-art methods for generalization. Our work aims to overcome this limitation by capitalizing on the abundance of long-form videos with dense speech-to-video alignment and the remarkable capabilities of recent large language models (LLMs) in summarizing long text. We introduce an automated and scalable pipeline for generating a large-scale video summarization dataset using LLMs as Oracle summarizers. By leveraging the generated dataset, we analyze the limitations of existing approaches and propose a new video summarization model that effectively addresses them. To facilitate further research in the field, our work also presents a new benchmark dataset that contains 1200 long videos each with high-quality summaries annotated by professionals. Extensive experiments clearly indicate that our proposed approach sets a new state-of-the-art in video summarization across several benchmarks. Dawit Mureja Argaw, Seunghyun Yoon 0002, Fabian Caba Heilbron, Hanieh Deilamsalehy, Trung Bui, Franck Dernoncourt, Joon Son Chung |
CVPR | 3 |
| 2024 | Towards Automated Movie Trailer GenerationabstractMovie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To stream-line this process, we propose an automatic trailer generation framework that generates plausible trailers from a full movie by automating shot selection and composition. Our approach draws inspiration from machine translation techniques and models the movies and trailers as sequences of shots, thus formulating the trailer generation problem as a sequence-to-sequence task. We introduce Trailer Generation Transformer (TGT), a deep-learning framework utilizing an encoder-decoder architecture. TGT movie encoder is tasked with contextualizing each movie shot representation via self-attention, while the autoregressive trailer de-coder predicts the feature representation of the next trailer shot, accounting for the relevance of shots' temporal order in trailers. Our TGT significantly outperforms previous methods on a comprehensive suite of metrics. Dawit Mureja Argaw, Mattia Soldan, Alejandro Pardo, Chen Zhao 0002, Fabian Caba Heilbron, Joon Son Chung, Bernard Ghanem |
CVPR | 5 |
| 2024 | Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsabstractWhile there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this work, we introduce Concept Weaver, a method for composing customized text-to-image diffusion models at inference time. Specifically, the method breaks the process into two steps: creating a template image aligned with the semantics of input prompts, and then personalizing the template using a concept fusion strategy. The fusion strategy incorporates the appearance of the target concepts into the template image while retaining its structural details. The results indicate that our method can generate multiple custom concepts with higher identity fidelity compared to alternative approaches. Furthermore, the method is shown to seamlessly handle more than two concepts and closely follow the semantic meaning of the input prompt without blending appearances across different subjects. Gihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee, Jong Chul Ye, Fabian Caba Heilbron |
CVPR | 6 |
| 2024 | Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
Ishan Rajendrakumar Dave, Fabian Caba Heilbron, Mubarak Shah, Simon Jenni |
ECCV (8) | 2 |
| 2024 | Multi-modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation
Linzi Xing, Quan Hung Tran, Fabian Caba Heilbron, Franck Dernoncourt, Seunghyun Yoon 0002, Trung Bui, Giuseppe Carenini |
MMM (3) | 3 |
| 2023 | PIVOT: Prompting for Video Continual LearningabstractModern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on such dynamic annotated sets. Continual learning directly approaches this problem, with the ultimate goal of devising methods where a deep neural network effectively learns relevant patterns for new (unseen) classes, without significantly altering its performance on previously learned ones. In this paper, we address the problem of continual learning for video data. We introduce PIVOT, a novel method that leverages extensive knowledge in pre-trained models from the image domain, thereby reducing the number of trainable parameters and the associated forgetting. Unlike previous methods, ours is the first approach that effectively uses prompting mechanisms for continual learning without any in-domain pre-training. Our experiments show that PIVOT improves state-of-the-art methods by a significant 27% on the 20-task ActivityNet setup. Andrés Villa, Juan Leon Alcazar, Motasem Alfarra, Kumail Alhamoud, Julio Hurtado, Fabian Caba Heilbron, Alvaro Soto, Bernard Ghanem |
CVPR | 6 |
| 2023 | Meta-Personalizing Vision-Language Models to Find Named Instances in VideoabstractLarge-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for moments in a video where a specific object instance such as “My dog Biscuit” appears. We present the following three contributions to address this problem. First, we describe a method to meta-personalize a pre-trained VLM, i.e., learning how to learn to personalize a VLM at test time to search in video. Our method extends the VLM's token vocabulary by learning novel word embeddings specific to each instance. To capture only instance-specific features, we represent each instance embedding as a combination of shared and learned global category features. Second, we propose to learn such personalization without explicit human supervision. Our approach automatically identifies moments of named visual instances in video using transcripts and vision-language similarity in the VLM's embedding space. Finally, we introduce This-Is-My, a personal video instance retrieval benchmark. We evaluate our approach on This-Is-My and Deep-Fashion2 and show that we obtain a 15% relative improvement over the state of the art on the latter dataset. Chun-Hsiao Yeh, Bryan C. Russell, Josef Sivic, Fabian Caba Heilbron, Simon Jenni |
CVPR | 4 |
| 2023 | Long-range Multimodal Pretraining for Movie UnderstandingabstractLearning computer vision models from (and for) movies has a long-standing history. While great progress has been attained, there is still a need for a pretrained multimodal model that can perform well in the ever-growing set of movie understanding tasks the community has been establishing. In this work, we introduce Long-range Multimodal Pretraining, a strategy, and a model that leverages movie data to train transferable multimodal and cross-modal encoders. Our key idea is to learn from all modalities in a movie by observing and extracting relationships over a long-range. After pretraining, we run ablation studies on the LVU benchmark and validate our modeling choices and the importance of learning from long-range time spans. Our model achieves state-of-the-art on several LVU tasks while being much more data efficient than previous works. Finally, we evaluate our model’s transferability by setting a new state-of-the-art in five different benchmarks. Dawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In-So Kweon, Fabian Caba Heilbron |
ICCV | 5 |
| 2023 | Localizing Moments in Long Video Via Multimodal GuidanceabstractThe recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with interesting findings: current grounding methods alone fail at tackling this challenging task and setup due to their inability to process long video sequences. In this paper, we propose a method for improving the performance of natural language grounding in long videos by identifying and pruning out non-describable windows. We design a guided grounding framework consisting of a Guidance Model and a base grounding model. The Guidance Model emphasizes describable windows, while the base grounding model ana-lyzes short temporal windows to determine which segments accurately match a given language query. We offer two designs for the Guidance Model: Query-Agnostic and Query-Dependent, which balance efficiency and accuracy. Experiments demonstrate that our proposed method outperforms state-of-the-art models by 4.1% in MAD and 4.52% in Ego4D (NLQ), respectively. Code, data and MAD’s audio features necessary to reproduce our experiments are available at: https://github.com/waybarrios/guidance-based-video-grounding. Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo, Fabian Caba Heilbron, Bernard Ghanem |
ICCV | 4 |
| 2023 | Efficient Adaptive Human-Object Interaction Detection with Concept-guided MemoryabstractHuman Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to handle long-tailed distributions of HOIs in complex HOI scenes in realistic settings. This observation motivates us to design an HOI detector that can be trained even with long-tailed labeled data and can leverage existing knowledge from pre-trained models. Inspired by the powerful generalization ability of the large Vision-Language Models (VLM) on classification and retrieval tasks, we propose an efficient Adaptive HOI Detector with Concept-guided Memory (ADA-CM). ADA-CM has two operating modes. The first mode makes it tunable without learning new parameters in a training-free paradigm. Its second mode incorporates an instance-aware adapter mechanism that can further efficiently boost performance if updating a lightweight set of parameters can be afforded. Our proposed method achieves competitive results with state-of-the-art on the HICO-DET and V-COCO datasets with much less training time. Code can be found at https://github.com/ltttpku/ADA-CM. Ting Lei 0001, Fabian Caba Heilbron, Qingchao Chen, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 2 |
| 2022 | FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks
Santiago Castro, Fabian Caba Heilbron |
BMVC | 2 |
| 2022 | MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio DescriptionsabstractThe recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recent works have begun to discover significant limitations in these datasets, suggesting that state-of-the-art techniques commonly overfit to hidden dataset biases. In this work, we present MAD (Movie Audio Descriptions), a novel benchmark that departs from the paradigm of augmenting existing video datasets with text annotations and focuses on crawling and aligning available audio descriptions of mainstream movies. MAD contains over 384, 000 natural language sentences grounded in over 1, 200 hours of videos and exhibits a significant reduction in the currently diagnosed biases for video-language grounding datasets. MAD's collection strategy enables a novel and more challenging version of video-language grounding, where short temporal moments (typically seconds long) must be accurately grounded in diverse long-form videos that can last up to three hours. We have released MAD's data and baselines code at https://github.com/Soldelli/MAD. Mattia Soldan, Alejandro Pardo, Juan Leon Alcazar, Fabian Caba Heilbron, Chen Zhao 0002, Silvio Giancola, Bernard Ghanem |
CVPR | 4 |
| 2022 | vCLIMB: A Novel Video Class Incremental Learning BenchmarkabstractContinual learning (CL) is under-explored in the video domain. The few existing works contain splits with imbalanced class distributions over the tasks, or study the problem in unsuitable datasets. We introduce vCLIMB, a novel video continual learning benchmark. vCLIMB is a standardized test-bed to analyze catastrophic forgetting of deep models in video continual learning. In contrast to previous work, we focus on class incremental continual learning with models trained on a sequence of disjoint tasks, and distribute the number of classes uniformly across the tasks. We perform in-depth evaluations of existing CL methods in vCLIMB, and observe two unique challenges in video data. The selection of instances to store in episodic memory is performed at the frame level. Second, untrimmed training data influences the effectiveness of frame sampling strategies. We address these two challenges by proposing a temporal consistency regularization that can be applied on top of memory-based continual learning methods. Our approach significantly improves the baseline, by up to 24% on the untrimmed continual learning task. The code of our benchmark can be found at: https://vclimb.netlify.app/. Andrés Villa, Kumail Alhamoud, Victor Escorcia, Fabian Caba Heilbron, Juan Leon Alcazar, Bernard Ghanem |
CVPR | 4 |
| 2022 | The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing
Dawit Mureja Argaw, Fabian Caba Heilbron, Joon-Young Lee, Markus Woodson, In-So Kweon |
ECCV (8) | 2 |
| 2022 | MovieCuts: A New Dataset and Benchmark for Cut Type Recognition
Alejandro Pardo, Fabian Caba Heilbron, Juan Leon Alcazar, Ali K. Thabet, Bernard Ghanem |
ECCV (7) | 2 |
| 2022 | Transcript to Video: Efficient Clip Sequencing from TextsabstractAmong numerous videos shared on the web, well-edited ones always attract more attention. However, it is difficult for inexperienced users to make well-edited videos because it requires professional expertise and immense manual labor. To meet the demands for non-experts, we present Transcript-to-Video -- a weakly-supervised framework that uses texts as input to automatically create video sequences from an extensive collection of shots. Specifically, we propose a Content Retrieval Module and a Temporal Coherent Module to learn visual-language representations and model shot sequencing styles, respectively. For fast inference, we introduce an efficient search strategy for real-time video clip sequencing. Quantitative results and user studies demonstrate empirically that the proposed learning framework can retrieve content-relevant shots while creating plausible video sequences in terms of style. Besides, the run-time performance analysis shows that our framework can support real-world applications. Project page: http://www.xiongyu.me/projects/transcript2video/ Fabian Caba Heilbron, Dahua Lin |
ACM Multimedia | 2 |
| 2021 | MAAS: Multi-modal Assignation for Active Speaker DetectionabstractActive speaker detection requires a mindful integration of multi-modal cues. Current methods focus on modeling and fusing short-term audiovisual features for individual speakers, often at frame level. We present a novel approach to active speaker detection that directly addresses the multi-modal nature of the problem and provides a straightforward strategy, where independent visual features (speakers) in the scene are assigned to a previously detected speech event. Our experiments show that a small graph data structure built from local information can approximate an instantaneous audio-visual assignment problem. More-over, the temporal extension of this initial graph achieves a new state-of-the-art performance on the AVA-ActiveSpeaker dataset with a mAP of 88.8%. Juan Leon Alcazar, Fabian Caba Heilbron, Ali K. Thabet, Bernard Ghanem |
ICCV | 2 |
| 2021 | Learning to Cut by Watching MoviesabstractVideo content creation keeps growing at an incredible pace; yet, creating engaging stories remains challenging and requires non-trivial video editing expertise. Many video editing components are astonishingly hard to automate primarily due to the lack of raw video materials. This paper focuses on a new task for computational video editing, namely the task of raking cut plausibility. Our key idea is to leverage content that has already been edited to learn fine-grained audiovisual patterns that trigger cuts. To do this, we first collected a data source of more than 10K videos, from which we extract more than 255K cuts. We devise a model that learns to discriminate between real and artificial cuts via contrastive learning. We set up a new task and a set of baselines to benchmark video cut generation. We observe that our proposed model outperforms the baselines by large margins. To demonstrate our model in real-world applications, we conduct human studies in a collection of unedited videos. The results show that our model does a better job at cutting than random and alternative baselines. Alejandro Pardo, Fabian Caba Heilbron, Juan Leon Alcazar, Ali K. Thabet, Bernard Ghanem |
ICCV | 2 |
| 2021 | RefineLoc: Iterative Refinement for Weakly-Supervised Action LocalizationabstractVideo action detectors are usually trained using datasets with fully-supervised temporal annotations. Building such datasets is an expensive task. To alleviate this problem, recent methods have tried to leverage weak labeling, where videos are untrimmed and only a video-level label is available. In this paper, we propose RefineLoc, a novel weaklysupervised temporal action localization method. RefineLoc uses an iterative refinement approach by estimating and training on snippet-level pseudo ground truth at every iteration. We show the benefit of this iterative approach and present an extensive analysis of five different pseudo ground truth generators. We show the effectiveness of our model on two standard action datasets, ActivityNet v1.2 and THUMOS14. RefineLoc shows competitive results with the stateof-the-art in weakly-supervised temporal localization. Additionally, our iterative refinement process is able to significantly improve the performance of two state-of-the-art methods, setting a new state-of-the-art on THUMOS14. Alejandro Pardo, Humam Alwassel, Fabian Caba Heilbron, Ali K. Thabet, Bernard Ghanem |
WACV | 3 |
| 2020 | Active Speakers in ContextabstractCurrent methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper introduces the Active Speaker Context, a novel representation that models relationships between multiple speakers over long time horizons. Our new model learns pairwise and temporal relations from a structured ensemble of audiovisual observations. Our experiments show that a structured feature ensemble already benefits active speaker detection performance. We also find that the proposed Active Speaker Context improves the state-of-the-art on the AVA-ActiveSpeaker dataset achieving an mAP of 87.1%. Moreover, ablation studies verify that this result is a direct consequence of our long-term multi-speaker analysis. Juan Leon Alcazar, Fabian Caba Heilbron, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Andrés Arbeláez, Bernard Ghanem |
CVPR | 2 |
| 2020 | Temporally Distributed Networks for Fast Video Semantic SegmentationabstractWe present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inherent temporal continuity in videos, we distribute these sub-networks over sequential frames. Therefore, at each time step, we only need to perform a lightweight computation to extract a sub-features group from a single sub-network. The full features used for segmentation are then recomposed by application of a novel attention propagation module that compensates for geometry deformation between frames. A grouped knowledge distillation loss is also introduced to further improve the representation power at both full and sub-feature levels. Experiments on Cityscapes, CamVid, and NYUD-v2 demonstrate that our method achieves state-of-the-art accuracy with significantly faster speed and lower latency. Ping Hu 0001, Fabian Caba Heilbron, Oliver Wang, Zhe Lin 0001, Stan Sclaroff, Federico Perazzi |
CVPR | 2 |
| 2018 | Diagnosing Error in Temporal Action Detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem |
ECCV (3) | 2 |
| 2018 | Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization
Humam Alwassel, Fabian Caba Heilbron, Bernard Ghanem |
ECCV (9) | 2 |
| 2018 | What Do I Annotate Next? An Empirical Study of Active Learning for Action Localization
Fabian Caba Heilbron, Joon-Young Lee, Hailin Jin, Bernard Ghanem |
ECCV (11) | 1 |
| 2017 | SCC: Semantic Context Cascade for Efficient Action DetectionabstractDespite the recent advances in large-scale video analysis, action detection remains as one of the most challenging unsolved problems in computer vision. This snag is in part due to the large volume of data that needs to be analyzed to detect actions in videos. Existing approaches have mitigated the computational cost, but still, these methods lack rich high-level semantics that helps them to localize the actions quickly. In this paper, we introduce a Semantic Cascade Context (SCC) model that aims to detect action in long video sequences. By embracing semantic priors associated with human activities, SCC produces high-quality class-specific action proposals and prune unrelated activities in a cascade fashion. Experimental results in ActivityNet unveils that SCC achieves state-of-the-art performance for action detection while operating at real time. Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, Bernard Ghanem |
CVPR | 1 |
| 2016 | Fast Temporal Activity Proposals for Efficient Detection of Human Actions in Untrimmed VideosabstractIn many large-scale video analysis scenarios, one is interested in localizing and recognizing human activities that occur in short temporal intervals within long untrimmed videos. Current approaches for activity detection still struggle to handle large-scale video collections and the task remains relatively unexplored. This is in part due to the computational complexity of current action recognition approaches and the lack of a method that proposes fewer intervals in the video, where activity processing can be focused. In this paper, we introduce a proposal method that aims to recover temporal segments containing actions in untrimmed videos. Building on techniques for learning sparse dictionaries, we introduce a learning framework to represent and retrieve activity proposals. We demonstrate the capabilities of our method in not only producing high quality proposals but also in its efficiency. Finally, we show the positive impact our method has on recognition performance when it is used for action detection, while running at 10FPS. Fabian Caba Heilbron, Juan Carlos Niebles, Bernard Ghanem |
CVPR | 1 |
| 2016 | DAPs: Deep Action Proposals for Action Understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, Bernard Ghanem |
ECCV (3) | 2 |
| 2015 | Robust Manhattan Frame estimation from a single RGB-D imageabstractThis paper proposes a new framework for estimating the Manhattan Frame (MF) of an indoor scene from a single RGB-D image. Our technique formulates this problem as the estimation of a rotation matrix that best aligns the normals of the captured scene to a canonical world axes. By introducing sparsity constraints, our method can simultaneously estimate the scene MF, the surfaces in the scene that are best aligned to one of three coordinate axes, and the outlier surfaces that do not align with any of the axes. To test our approach, we contribute a new set of annotations to determine ground truth MFs in each image of the popular NYUv2 dataset. We use this new benchmark to experimentally demonstrate that our method is more accurate, faster, more reliable and more robust than the methods used in the literature. We further motivate our technique by showing how it can be used to address the RGB-D SLAM problem in indoor scenes by incorporating it into and improving the performance of a popular RGB-D SLAM method. Bernard Ghanem, Ali K. Thabet, Juan Carlos Niebles, Fabian Caba Heilbron |
CVPR | 4 |
| 2015 | ActivityNet: A large-scale video benchmark for human activity understandingabstractIn spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on simple actions and movements occurring on manually trimmed videos. In this paper we introduce ActivityNet, a new large-scale video benchmark for human activity understanding. Our benchmark aims at covering a wide range of complex human activities that are of interest to people in their daily living. In its current version, ActivityNet provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours. We illustrate three scenarios in which ActivityNet can be used to compare algorithms for human activity understanding: untrimmed video classification, trimmed activity classification and activity detection. Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, Juan Carlos Niebles |
CVPR | 1 |
| 2014 | Camera Motion and Surrounding Scene Appearance as Context for Action Recognition
Fabian Caba Heilbron, Ali K. Thabet, Juan Carlos Niebles, Bernard Ghanem |
ACCV (4) | 1 |
| 2014 | Collecting and Annotating Human Activities in Web VideosabstractRecent efforts in computer vision tackle the problem of human activity understanding in video sequences. Traditionally, these algorithms require annotated video data to learn models. In this paper, we introduce a novel data collection framework, to take advantage of the large amount of video data available on the web. We use this new framework to retrieve videos of human activities in order to build datasets for training and evaluating computer vision algorithms. We rely on Amazon Mechanical Turk workers to obtain high accuracy annotations. An agglomerative clustering technique brings the possibility to achieve reliable and consistent annotations for temporal localization of human activities in videos. Using two different datasets, Olympics Sports and our novel Daily Human Activities dataset, we show that our collection/annotation framework achieves robust annotations for human activities in large amount of video data. Fabian Caba Heilbron, Juan Carlos Niebles |
ICMR | 1 |