Victor Escorcia

dblp:169/4850 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-0070-1566ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning
abstract
There has been immense progress recently in the visual quality of Stable Diffusion-based Super Resolution (SD-SR). However, deploying large diffusion models on computationally restricted devices such as mobile phones remains impractical due to the large model size and high latency. This is compounded for SR as it often operates at high res (e.g. 4K×3K). In this work, we introduce Edge-SD-SR, the first parameter efficient and low latency diffusion model for image super-resolution. Edge-SD-SR consists of ~169M parameters, including UNet, encoder and decoder, and has a complexity of only ∼ 142 GFLOPs. To maintain a high visual quality on such low compute budget, we introduce a number of training strategies: (i) A novel conditioning mechanism on the low-resolution input, coined bidirectional conditioning, which tailors the SD model for the SR task. (ii) Joint training of the UNet and encoder, while decoupling the encodings of the HR and LR images and using a dedicated schedule. (iii) Finetuning the decoder using the UNet’s output to directly tailor the decoder to the latents obtained at inference time. Edge-SD-SR runs efficiently on device, e.g. it can upscale a 128 × 128 patch to 512 × 512 in 38 msec while running on a Samsung S24 DSP, and of a 512 × 512 to 2, 048 × 2, 048 (requiring 25 model evaluations) in just ∼ 1.1 sec. Furthermore, we show that Edge-SD-SR matches or even outperforms state-of-the-art SR approaches on the most established SR benchmarks.
Isma Hadji, Mehdi Noroozi, Victor Escorcia, Anestis Zaganidis, Brais Martínez, Georgios Tzimiropoulos
CVPR3
2022 vCLIMB: A Novel Video Class Incremental Learning Benchmark
abstract
Continual learning (CL) is under-explored in the video domain. The few existing works contain splits with imbalanced class distributions over the tasks, or study the problem in unsuitable datasets. We introduce vCLIMB, a novel video continual learning benchmark. vCLIMB is a standardized test-bed to analyze catastrophic forgetting of deep models in video continual learning. In contrast to previous work, we focus on class incremental continual learning with models trained on a sequence of disjoint tasks, and distribute the number of classes uniformly across the tasks. We perform in-depth evaluations of existing CL methods in vCLIMB, and observe two unique challenges in video data. The selection of instances to store in episodic memory is performed at the frame level. Second, untrimmed training data influences the effectiveness of frame sampling strategies. We address these two challenges by proposing a temporal consistency regularization that can be applied on top of memory-based continual learning methods. Our approach significantly improves the baseline, by up to 24% on the untrimmed continual learning task. The code of our benchmark can be found at: https://vclimb.netlify.app/.
Andrés Villa, Kumail Alhamoud, Victor Escorcia, Fabian Caba Heilbron, Juan Leon Alcazar, Bernard Ghanem
CVPR3
2022 SOS! Self-supervised Learning over Sets of Handled Objects in Egocentric Action Recognition
Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martínez
ECCV (13)1
2021 Knowing What, Where and When to Look: Video Action modelling with Attention
Juan-Manuel Pérez-Rúa, Brais Martínez, Xiatian Zhu, Antoine Toisoul, Victor Escorcia, Tao Xiang 0002
BMVC5
2021 TNT: Text-Conditioned Network with Transductive Inference for Few-Shot Video Classification
Andrés Villa, Juan-Manuel Pérez-Rúa, Vladimir Araujo, Juan Carlos Niebles, Victor Escorcia, Alvaro Soto
BMVC5
2021 Boundary-sensitive Pre-training for Temporal Localization in Videos
abstract
Many video analysis tasks require temporal localization for the detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is due to large scale annotation of temporal boundaries in untrimmed videos being expensive. Therefore, no suitable datasets exist that enable pre-training in a manner sensitive to temporal boundaries. In this paper for the first time, we investigate model pre-training for temporal localization by introducing a novel boundary-sensitive pretext (BSP) task. Instead of relying on costly manual annotations of temporal boundaries, we propose to synthesize temporal boundaries in existing video action classification datasets. By defining different ways of synthesizing boundaries, BSP can then be simply conducted in a self-supervised manner via the classification of the boundary types. This enables the learning of video representations that are much more transferable to downstream temporal localization tasks. Extensive experiments show that the proposed BSP is superior and complementary to the existing action classification-based pre-training counterpart, and achieves new state-of-the-art performance on several temporal localization tasks. Please visit our website for more details https://frostinassiky.github.io/bsp.
Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martínez, Xiatian Zhu, Li Zhang 0040, Bernard Ghanem, Tao Xiang 0002
ICCV3
2020 Guess where? Actor-supervision for spatiotemporal action localization
Victor Escorcia, Cuong Duc Dao, Mihir Jain, Bernard Ghanem, Cees Snoek
Comput. Vis. Image Underst.1
2018 Diagnosing Error in Temporal Action Detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem
ECCV (3)3
2017 End-to-End, Single-Stream Temporal Action Detection in Untrimmed Videos
Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei 0001, Juan Carlos Niebles
BMVC2
2017 SST: Single-Stream Temporal Action Proposals
abstract
Our paper presents a new approach for temporal detection of human actions in long, untrimmed video sequences. We introduce Single-Stream Temporal Action Proposals (SST), a new effective and efficient deep architecture for the generation of temporal action proposals. Our network can run continuously in a single stream over very long input video sequences, without the need to divide input into short overlapping clips or temporal windows for batch processing. We demonstrate empirically that our model outperforms the state-of-the-art on the task of temporal action proposal generation, while achieving some of the fastest processing speeds in the literature. Finally, we demonstrate that using SST proposals in conjunction with existing action classifiers results in improved state-of-the-art temporal action detection performance.
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, Juan Carlos Niebles
CVPR2
2017 SCC: Semantic Context Cascade for Efficient Action Detection
abstract
Despite the recent advances in large-scale video analysis, action detection remains as one of the most challenging unsolved problems in computer vision. This snag is in part due to the large volume of data that needs to be analyzed to detect actions in videos. Existing approaches have mitigated the computational cost, but still, these methods lack rich high-level semantics that helps them to localize the actions quickly. In this paper, we introduce a Semantic Cascade Context (SCC) model that aims to detect action in long video sequences. By embracing semantic priors associated with human activities, SCC produces high-quality class-specific action proposals and prune unrelated activities in a cascade fashion. Experimental results in ActivityNet unveils that SCC achieves state-of-the-art performance for action detection while operating at real time.
Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, Bernard Ghanem
CVPR3
2016 DAPs: Deep Action Proposals for Action Understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, Bernard Ghanem
ECCV (3)1
2015 On the relationship between visual attributes and convolutional networks
abstract
One of the cornerstone principles of deep models is their abstraction capacity, i.e. their ability to learn abstract concepts from `simpler' ones. Through extensive experiments, we characterize the nature of the relationship between abstract concepts (specifically objects in images) learned by popular and high performing convolutional networks (conv-nets) and established mid-level representations used in computer vision (specifically semantic visual attributes). We focus on attributes due to their impact on several applications, such as object description, retrieval and mining, and active (and zero-shot) learning. Among the findings we uncover, we show empirical evidence of the existence of Attribute Centric Nodes (ACNs) within a conv-net, which is trained to recognize objects (not attributes) in images. These special conv-net nodes (1) collectively encode information pertinent to visual attribute representation and discrimination, (2) are unevenly and sparsely distribution across all layers of the conv-net, and (3) play an important role in conv-net based object recognition.
Victor Escorcia, Juan Carlos Niebles, Bernard Ghanem
CVPR1
2015 ActivityNet: A large-scale video benchmark for human activity understanding
abstract
In spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on simple actions and movements occurring on manually trimmed videos. In this paper we introduce ActivityNet, a new large-scale video benchmark for human activity understanding. Our benchmark aims at covering a wide range of complex human activities that are of interest to people in their daily living. In its current version, ActivityNet provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours. We illustrate three scenarios in which ActivityNet can be used to compare algorithms for human activity understanding: untrimmed video classification, trimmed activity classification and activity detection.
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, Juan Carlos Niebles
CVPR2