EDBT 2026 Demo / reviewers in the wild / expert
Juan-Manuel Pérez-Rúa
dblp:172/9703
· DBLP profile ↗
21ranked-venue papers
10as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Flow Fields in Attention for Controllable Person Image GenerationabstractControllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person’s appearance or pose. However, prior methods often distort fine-grained details from the reference image, despite achieving high overall image quality. We attribute these distortions to inadequate attention to corresponding regions in the reference image. To address this, we thereby propose learning flow fields in attention (Leffa), which explicitly guides the target query to attend to the correct reference key in the attention layer during training. Specifically, it is realized via a regularization loss on top of the attention map within a diffusionbased baseline. Our extensive experiments show that Leffa achieves state-of-the-art performance in controlling appearance and pose, significantly reducing fine-grained detail distortion while maintaining high image quality. Additionally, we show that our loss is model-agnostic and can be used to improve the performance of other diffusion models. Zijian Zhou 0002, Shikun Liu, Kam Woh Ng, Tian Xie 0003, Yuren Cong, Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Aditya Patel, Tao Xiang 0002, Miaojing Shi, Sen He 0001 |
CVPR | 10 |
| 2024 | GenTron: Diffusion Transformers for Image and Video GenerationabstractIn this study, we explore Transformer-based diffusion models for image and video generation. Despite the dominance of Transformer architectures in various fields due to their flexibility and scalability, the visual generative domain primarily utilizes CNN-based U-Net architectures, particularly in diffusion-based models. We introduce GenTron, a family of Generative models employing Transformer-based diffusion, to address this gap. Our initial step was to adapt Diffusion Transformers (DiTs) from class to text conditioning, a process involving thorough empirical exploration of the conditioning mechanism. We then scale GenTron from approximately 900M to over 3B parameters, observing improvements in visual quality. Furthermore, we extend GenTron to text-to-video generation, incorporating novel motion-free guidance to enhance video quality. In human evaluations against SDXL, GenTron achieves a 51.1% win rate in visual quality (with a 19.8% draw rate), and a 42.3% win rate in text alignment (with a 42.9% draw rate). GenTron notably performs well in T2I-CompBench, highlighting its compositional generation ability. We hope GenTron could provide meaningful insights and serve as a valuable reference for future research. Please refer to the website11https://www.shoufachen.com/gentron_website/ and the arXiv version for the most up-to-date results: https://arxiv.org/abs/2312.04557. Shoufa Chen, Mengmeng Xu 0006, Jiawei Ren 0001, Yuren Cong, Sen He 0001, Yanping Xie, Animesh Sinha, Ping Luo 0002, Tao Xiang 0002, Juan-Manuel Pérez-Rúa |
CVPR | 10 |
| 2024 | FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingabstractText-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts.
A major challenge in this task is to ensure that all frames in the edited video are visually consistent.
Most recent works apply advanced text-to-image diffusion models to this task by inflating 2D spatial attention in the U-Net into spatio-temporal attention.
Although temporal context can be added through spatio-temporal attention, it may introduce some irrelevant information for each patch and therefore cause inconsistency in the edited video.
In this paper, for the first time, we introduce optical flow into the attention module in diffusion model's U-Net to address the inconsistency issue for text-to-video editing.
Our method, FLATTEN, enforces the patches on the same flow path across different frames to attend to each other in the attention module, thus improving the visual consistency in the edited videos.
Additionally, our method is training-free and can be seamlessly integrated into any diffusion based text-to-video editing methods and improve their visual consistency.
Experiment results on existing text-to-video editing benchmarks show that our proposed method achieves the new state-of-the-art performance. In particular, our method excels in maintaining the visual consistency in the edited videos. Yuren Cong, Mengmeng Xu 0006, Christian Simon, Shoufa Chen, Jiawei Ren 0001, Yanping Xie, Juan-Manuel Pérez-Rúa, Bodo Rosenhahn, Tao Xiang 0002, Sen He 0001 |
ICLR | 7 |
| 2024 | Boundary Denoising for Video Activity LocalizationabstractVideo activity localization aims at understanding the semantic content in long, untrimmed videos and retrieving actions of interest. The retrieved action with its start and end locations can be used for highlight generation, temporal action detection, etc. Unfortunately, learning the exact boundary location of activities is highly challenging because temporal activities are continuous in time, and there are often no clear-cut transitions between actions. Moreover, the definition of the start and end of events is subjective, which may confuse the model. To alleviate the boundary ambiguity, we propose to study the video activity localization problem from a denoising perspective. Specifically, we propose an encoder-decoder model named DenosieLoc. During training, a set of temporal spans is randomly generated from the ground truth with a controlled noise scale. Then, we attempt to reverse this process by boundary denoising, allowing the localizer to predict activities with precise boundaries and resulting in faster convergence speed. Experiments show that DenosieLoc advances
several video activity understanding tasks. For example, we observe a gain of +12.36% average mAP on the QV-Highlights dataset.
Moreover, DenosieLoc achieves state-of-the-art performance on the MAD dataset but with much fewer predictions than others. Mengmeng Xu 0006, Mattia Soldan, Jialin Gao, Shuming Liu 0001, Juan-Manuel Pérez-Rúa, Bernard Ghanem |
ICLR | 5 |
| 2023 | Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query LocalizationabstractThis paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual query localization. We first identify grave implicit biases in current query-conditioned model design and visual query datasets. Then, we directly tackle such biases at both frame and object set levels. Concretely, our method solves these issues by expanding limited annotations and dynamically dropping object proposals during training. Additionally, we propose a novel transformer-based module that allows for object-proposal set context to be considered while incorporating query information. We name our module Conditioned Contextual Transformer or CocoFormer. Our experiments show the proposed adaptations improve egocentric query detection, leading to a better visual query localization system in both 2D and 3D configurations. Thus, we can improve frame-level detection performance from 26.28% to 31.26% in AP, which correspondingly improves the VQ2D and VQ3D localization scores by significant margins. Our improved context-aware query object detector ranked first and second respectively in the VQ2D and VQ3D tasks in the 2nd Ego4D challenge. In addition to this, we showcase the relevance of our proposed model in the Few-Shot Detection (FSD) task, where we also achieve SOTA results. Our code is available at https://github.com/facebookresearch/vq2d_cvpr. Mengmeng Xu 0006, Yanghao Li, Cheng-Yang Fu, Bernard Ghanem, Tao Xiang 0002, Juan-Manuel Pérez-Rúa |
CVPR | 6 |
| 2021 | Knowing What, Where and When to Look: Video Action modelling with Attention
Juan-Manuel Pérez-Rúa, Brais Martínez, Xiatian Zhu, Antoine Toisoul, Victor Escorcia, Tao Xiang 0002 |
BMVC | 1 |
| 2021 | TNT: Text-Conditioned Network with Transductive Inference for Few-Shot Video Classification
Andrés Villa, Juan-Manuel Pérez-Rúa, Vladimir Araujo, Juan Carlos Niebles, Victor Escorcia, Alvaro Soto |
BMVC | 2 |
| 2021 | Few-shot Action Recognition with Prototype-centered Attentive Learning
Xiatian Zhu, Antoine Toisoul, Juan-Manuel Pérez-Rúa, Li Zhang 0040, Brais Martínez, Tao Xiang 0002 |
BMVC | 3 |
| 2021 | Boundary-sensitive Pre-training for Temporal Localization in VideosabstractMany video analysis tasks require temporal localization for the detection of content changes. However, most existing models developed for these tasks are pre-trained on general video action classification tasks. This is due to large scale annotation of temporal boundaries in untrimmed videos being expensive. Therefore, no suitable datasets exist that enable pre-training in a manner sensitive to temporal boundaries. In this paper for the first time, we investigate model pre-training for temporal localization by introducing a novel boundary-sensitive pretext (BSP) task. Instead of relying on costly manual annotations of temporal boundaries, we propose to synthesize temporal boundaries in existing video action classification datasets. By defining different ways of synthesizing boundaries, BSP can then be simply conducted in a self-supervised manner via the classification of the boundary types. This enables the learning of video representations that are much more transferable to downstream temporal localization tasks. Extensive experiments show that the proposed BSP is superior and complementary to the existing action classification-based pre-training counterpart, and achieves new state-of-the-art performance on several temporal localization tasks. Please visit our website for more details https://frostinassiky.github.io/bsp. Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martínez, Xiatian Zhu, Li Zhang 0040, Bernard Ghanem, Tao Xiang 0002 |
ICCV | 2 |
| 2021 | Space-time Mixing Attention for Video TransformerabstractThis paper is on video recognition using Transformers. Very recent attempts in this area have demonstrated promising results in terms of recognition accuracy, yet they have been also shown to induce, in many cases, significant computational overheads due to the additional modelling of the temporal information. In this work, we propose a Video Transformer model the complexity of which scales linearly with the number of frames in the video sequence and hence induces no overhead compared to an image-based Transformer model. To achieve this, our model makes two approximations to the full space-time attention used in Video Transformers: (a) It restricts time attention to a local temporal window and capitalizes on the Transformer's depth to obtain full temporal coverage of the video sequence. (b) It uses efficient space-time mixing to attend jointly spatial and temporal locations without inducing any additional cost on top of a spatial-only attention model. We also show how to integrate 2 very lightweight mechanisms for global temporal-only attention which provide additional accuracy improvements at minimal computational cost. We demonstrate that our model produces very high recognition accuracy on the most popular video recognition datasets while at the same time being significantly more efficient than other Video Transformer models. Adrian Bulat, Juan-Manuel Pérez-Rúa, Swathikiran Sudhakaran, Brais Martínez, Georgios Tzimiropoulos |
NeurIPS | 2 |
| 2021 | Low-Fidelity Video Encoder Optimization for Temporal Action LocalizationabstractMost existing temporal action localization (TAL) methods rely on a transfer learning pipeline: by first optimizing a video encoder on a large action classification dataset (i.e., source domain), followed by freezing the encoder and training a TAL head on the action localization dataset (i.e., target domain). This results in a task discrepancy problem for the video encoder – trained for action classification, but used for TAL. Intuitively, joint optimization with both the video encoder and TAL head is a strong baseline solution to this discrepancy. However, this is not operable for TAL subject to the GPU memory constraints, due to the prohibitive computational cost in processing long untrimmed videos. In this paper, we resolve this challenge by introducing a novel low-fidelity (LoFi) video encoder optimization method. Instead of always using the full training configurations in TAL learning, we propose to reduce the mini-batch composition in terms of temporal, spatial, or spatio-temporal resolution so that jointly optimizing the video encoder and TAL head becomes operable under the same memory conditions of a mid-range hardware budget. Crucially, this enables the gradients to flow backwards through the video encoder conditioned on a TAL supervision loss, favourably solving the task discrepancy problem and providing more effective feature representations. Extensive experiments show that the proposed LoFi optimization approach can significantly enhance the performance of existing TAL methods. Encouragingly, even with a lightweight ResNet18 based video encoder in a single RGB stream, our method surpasses two-stream (RGB + optical-flow) ResNet50 based alternatives, often by a good margin. Our code is publicly available at https://github.com/saic-fi/lofiactionlocalization. Mengmeng Xu 0006, Juan-Manuel Pérez-Rúa, Xiatian Zhu, Bernard Ghanem, Brais Martínez |
NeurIPS | 2 |
| 2020 | Incremental Few-Shot Object DetectionabstractExisting object detection methods typically rely on the availability of abundant labelled training samples per class and offline model training in a batch mode. These requirements substantially limit their scalability to open-ended accommodation of novel classes with limited labelled training data, both in terms of model accuracy and training efficiency during deployment. We present the first study aiming to go beyond these limitations by considering the Incremental Few-Shot Detection (iFSD) problem setting, where new classes must be registered incrementally (without revisiting base classes) and with few examples. To this end we propose OpeN-ended Centre nEt (ONCE), a detector designed for incrementally learning to detect novel class objects with few examples. This is achieved by an elegant adaptation of the efficient CentreNet detector to the few-shot learning scenario, and meta-learning a class-wise code generator model for registering novel classes. ONCE fully respects the incremental learning paradigm, with novel class registration requiring only a single forward pass of few-shot training samples, and no access to base classes - thus making it suitable for deployment on embedded devices, etc. Extensive experiments conducted on both the standard object detection (COCO, PASCAL VOC) and fashion landmark detection (DeepFashion2) tasks show the feasibility of iFSD for the first time, opening an interesting and very important line of research. Juan-Manuel Pérez-Rúa, Xiatian Zhu, Timothy M. Hospedales, Tao Xiang 0002 |
CVPR | 1 |
| 2020 | ROAM: A Rich Object Appearance Model with Application to RotoscopingabstractRotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given an initial closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We further extend this model by so-called trimaps which serve as an input to alpha-matting algorithms to allow truly seamless compositing. To this end, we leverage local classifiers attached to the roto-curves to define a confidence measure that is well-suited to define trimaps with adaptive band-widths. The resulting trimaps are parametric, temporally consistent and remain fully editable by the artist. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling. Juan-Manuel Pérez-Rúa, Ondrej Miksik, Tomás Crivelli, Patrick Bouthemy, Philip Torr 0001, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | MFAS: Multimodal Fusion Architecture SearchabstractWe tackle the problem of finding good architectures for multimodal classification problems. We propose a novel and generic search space that spans a large number of possible fusion architectures. In order to find an optimal architecture for a given dataset in the proposed search space, we leverage an efficient sequential model-based exploration approach that is tailored for the problem. We demonstrate the value of posing multimodal fusion as a neural architecture search problem by extensive experimentation on a toy dataset and two other real multimodal datasets. We discover fusion architectures that exhibit state-of-the-art performance for problems with different domain and dataset size, including the \ntu~dataset, the largest multimodal action recognition dataset available. Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, Frédéric Jurie |
CVPR | 1 |
| 2018 | Efficient Progressive Neural Architecture Search
Juan-Manuel Pérez-Rúa, Moez Baccouche, Stéphane Pateux |
BMVC | 1 |
| 2017 | ROAM: A Rich Object Appearance Model with Application to RotoscopingabstractRotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given a first closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling. Ondrej Miksik, Juan-Manuel Pérez-Rúa, Philip Torr 0001, Patrick Pérez |
CVPR | 2 |
| 2016 | Discovering motion hierarchies via tree-structured coding of trajectories
Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez, Patrick Bouthemy |
BMVC | 1 |
| 2016 | Determining Occlusions from Space and Time Image ReconstructionsabstractThe problem of localizing occlusions between consecutive frames of a video is important but rarely tackled on its own. In most works, it is tightly interleaved with the computation of accurate optical flows, which leads to a delicate chicken-and-egg problem. With this in mind, we propose a novel approach to occlusion detection where visibility or not of a point in next frame is formulated in terms of visual reconstruction. The key issue is now to determine how well a pixel in the first image can be "reconstructed" from co-located colors in the next image. We first exploit this reasoning at the pixel level with a new detection criterion. Contrary to the ubiquitous displaced-framedifference and forward-backward flow vector matching, the proposed alternative does not critically depend on a precomputed, dense displacement field, while being shown to be more effective. We then leverage this local modeling within an energy-minimization framework that delivers occlusion maps. An easy-to-obtain collection of parametric motion models is exploited within the energy to provide the required level of motion information. Our approach outperforms state-of-the-art detection methods on the challenging MPI Sintel dataset. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Bouthemy, Patrick Pérez |
CVPR | 1 |
| 2016 | Hierarchical motion decomposition for dynamic scene parsingabstractA number of applications in video analysis rely on a per-frame motion segmentation of the scene as key preprocessing step. Moreover, different settings in video production require extracting segmentation masks of multiple moving objects and object parts in a hierarchical fashion. In order to tackle this problem, we propose to analyze and exploit the compositional structure of scene motion to provide a segmentation which is not purely driven by local image information. Specifically, we leverage a hierarchical motion-based partition of the scene to capture a mid-level understanding of the dynamic video content. We present experimental results showing the strengths of this approach in comparison to current video segmentation approaches. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez, Patrick Bouthemy |
ICIP | 1 |
| 2016 | Object-guided motion estimation
Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez |
Comput. Vis. Image Underst. | 1 |
| 2015 | Background-foreground tracking for video object segmentationabstractWe present a method to segment objects of interest in video sequences by combining robust background and foreground point tracking with joint color and motion-based segmentation. Our approach is sequential in time, avoiding a global processing of the video, while being simple and generic. This makes the method attractive for online applications, including video editing or augmented reality, as it can be adapted for both automated and interactive work-flows. We present visual and quantitative experiments to compare with existing algorithms, showing promising results. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez |
ICIP | 1 |