Ana Cristina Murillo

dblp:58/4977 · also A. C. Murillo, Ana C. Murillo · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-7580-9037ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 3 first-author · 13 since 2021Systems, architecture and hardware · 17 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 CineMPC: A Fully Autonomous Drone Cinematography System Incorporating Zoom, Focus, Pose, and Scene Composition (Abstract Reprint)
abstract
We present CineMPC, a complete cinematographic system that autonomously controls a drone to film multiple targets recording user-specified aesthetic objectives. Existing solutions in autonomous cinematography control only the camera extrinsics, namely, its position and orientation. In contrast, CineMPC is the first solution that includes the camera intrinsic parameters in the control loop, which are essential tools for controlling cinematographic effects such as focus, zoom, and depth of field. The system is validated in real-world experiments.
Pablo Pueyo, Juan Dendarieta, Eduardo Montijano, Ana Cristina Murillo, Mac Schwager
AAAI4
2026 FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
abstract
Finding information in hour-long videos is a challenging task even for top-performing Vision Language Models (VLMs), as encoding visual content quickly exceeds available context windows. To tackle this challenge, we present FALCONEye1, a novel video agent based on a training-free, model-agnostic meta-architecture composed of a VLM and a Large Language Model (LLM). FALCONEye answers open-ended questions using an exploration-based search algorithm guided by calibrated confidence from the VLM’s answers. We also introduce the FALCON-Bench benchmark, extending Question Answering problem to Video Answer Search—requiring models to return both the answer and its supporting temporal window for open-ended questions in hour-long videos. With just a 7B VLM and a lightweight LLM, FALCONEye outscores all open-source 7B VLMs and comparable agents in FALCON-Bench. It further demonstrates its generalization capability in MLVU benchmark with shorter videos and different tasks, surpassing GPT-4o on single-detail tasks while slashing inference cost by roughly an order of magnitude.
Carlos Plou, Cesar Borja, Ruben Martinez-Cantin, Ana Cristina Murillo
WACV4
2026 EventSleep2: Sleep activity recognition on complete night sleep recordings with an event camera
abstract
Sleep is fundamental to health, and society is more and more aware of the impact and relevance of sleep disorders. Traditional diagnostic methods, like polysomnography, are intrusive and resource-intensive. Instead, research is focusing on developing novel, less intrusive or portable methods that combine intelligent sensors with activity recognition for diagnosis support and scoring. Event cameras offer a promising alternative for automated, in-home sleep activity recognition due to their excellent low-light performance and low power consumption. This work introduces EventSleep2-data , a significant extension to the EventSleep dataset, featuring 10 complete night recordings (around 7 h each) of volunteers sleeping in their homes. Unlike the original short and controlled recordings, this new dataset captures natural, full-night sleep sessions under realistic conditions. This new data incorporates challenging real-world scene variations, an efficient movement-triggered sparse data recording pipeline, and synchronized 2-channel EEG data for a subset of recordings. We also present EventSleep2-net , a novel event-based sleep activity recognition approach with a dual-head architecture to simultaneously analyze motion classes and static poses. The model is specifically designed to handle the motion-triggered, sparse nature of complete night recordings. Unlike the original EventSleep architecture, EventSleep2-net can predict both movement and static poses even during long periods with no events. We demonstrate state-of-the-art performance on both EventSleep1-data, the original dataset, and EventSleep2-data, with comprehensive ablation studies validating our design decisions. Together, EventSleep2-data and EventSleep2-net overcome the limitations of the previous setup and enable continuous, full-night analysis for real-world sleep monitoring, significantly advancing the potential of event-based vision for sleep disorder studies.
Nerea Gallego, Carlos Plou, Miguel Marcos, Pablo Urcola, Luis Montesano, Eduardo Montijano, Ruben Martinez-Cantin, Ana Cristina Murillo
Comput. Vis. Image Underst.8
2026 Temporal video segmentation with natural language using text-video cross attention and Bayesian order-priors
abstract
Video is a crucial perception component in both robotics and wearable devices, two key technologies to enable innovative assistive applications, such as navigation and procedure execution assistance tools. Video understanding tasks are essential to enable these systems to interpret and execute complex instructions in real-world environments. One such task is step grounding, which involves identifying the temporal boundaries of activities based on natural language descriptions in long, untrimmed videos. This paper introduces Bayesian-VSLNet, a probabilistic formulation of step grounding that predicts a likelihood distribution over segments and refines it through Bayesian inference with temporal-order priors. These priors disambiguate cyclic and repeated actions that frequently appear in procedural tasks, enabling precise step localization in long videos. Our evaluations demonstrate superior performance over existing methods, achieving state-of-the-art results in the Ego4D Goal-Step dataset, winning the Goal Step challenge at the EgoVis 2024 CVPR. Furthermore, experiments on additional benchmarks confirm the generality of our approach beyond Ego4D. In addition, we present qualitative results in a real-world robotics scenario, illustrating the potential of this task to improve human–robot interaction in practical applications. Code is released at https://github.com/cplou99/BayesianVSLNet .
Carlos Plou, Lorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-Cantin, Ana Cristina Murillo
Comput. Vis. Image Underst.5
2024 SpectralWaste Dataset: Multimodal Data for Waste Sorting Automation
abstract
The increase in non-biodegradable waste is a worldwide concern. Recycling facilities play a crucial role, but their automation is hindered by the complex characteristics of waste recycling lines like clutter or object deformation. In addition, the lack of publicly available labeled data for these environments makes developing robust perception systems challenging. Our work explores the benefits of multimodal perception for object segmentation in real waste management scenarios. First, we present SpectralWaste, the first dataset collected from an operational plastic waste sorting facility that provides synchronized hyperspectral and conventional RGB images. This dataset contains labels for several categories of objects that commonly appear in sorting plants and need to be detected and separated from the main trash flow for several reasons, such as security in the management line or reuse. Additionally, we propose a pipeline employing different object segmentation architectures and evaluate the alternatives on our dataset, conducting an extensive analysis for both multimodal and unimodal alternatives. Our evaluation pays special attention to efficiency and suitability for real-time processing and demonstrates how hyperspectral imaging can bring a boost to RGB-only perception in these realistic industrial settings without much computational overhead.
Sara Casao, Fernando Peña 0002, Alberto Sabater, Rosa Castillón, Darío Suárez Gracia, Eduardo Montijano, Ana Cristina Murillo
IROS7
2024 CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
abstract
This paper introduces CLIPSwarm, a new algorithm designed to automate the modeling of swarm drone formations based on natural language. The algorithm begins by enriching a provided word, to compose a text prompt that serves as input to an iterative approach to find the formation that best matches the provided word. The algorithm iteratively refines formations of robots to align with the textual description, employing different steps for "exploration" and "exploitation". Our framework is currently evaluated on simple formation targets, limited to contour shapes. A formation is visually represented through alpha-shape contours and the most representative color is automatically found for the input word. To measure the similarity between the description and the visual representation of the formation, we use CLIP [1], encoding text and images into vectors and assessing their similarity. Sub-sequently, the algorithm rearranges the formation to visually represent the word more effectively, within the given constraints of available drones. Control actions are then assigned to the drones, ensuring robotic behavior and collision-free movement. Experimental results demonstrate the system’s efficacy in accurately modeling robot formations from natural language descriptions. The algorithm’s versatility is showcased through the execution of drone shows in photorealistic simulation with varying shapes. We refer the reader to the supplementary video for a visual reference of the results.
Pablo Pueyo, Eduardo Montijano, Ana Cristina Murillo, Mac Schwager
IROS3
2024 Distributed multi-target tracking and active perception with mobile camera networks
abstract
Smart cameras are an essential component in surveillance and monitoring applications, and they have been typically deployed in networks of fixed camera locations. The addition of mobile cameras, mounted on robots, can overcome some of the limitations of static networks such as blind spots or back-lightning, allowing the system to gather the best information at each time by active positioning. This work presents a hybrid camera system, with static and mobile cameras, where all the cameras collaborate to observe people moving freely in the environment and efficiently visualize certain attributes from each person. Our solution combines a multi-camera distributed tracking system, to localize with precision all the people, with a control scheme that moves the mobile cameras to the best viewpoints for a specific classification task. The main contribution of this paper is a novel framework that exploits the synergies that result from the cooperation of the tracking and the control modules, obtaining a system closer to the real-world application and capable of high-level scene understanding. The static camera network provides global awareness of the control scheme to move the robots. In exchange, the mobile cameras onboard the robots provide enhanced information about the people on the scene. We perform a thorough analysis of the people monitoring application performance under different conditions thanks to the use of a photo-realistic simulation environment. Our experiments demonstrate the benefits of collaborative mobile cameras with respect to static or individual camera setups.
Sara Casao, Álvaro Serra-Gómez, Ana Cristina Murillo, Wendelin Böhmer, Javier Alonso-Mora, Eduardo Montijano
Comput. Vis. Image Underst.3
2024 CineMPC: A Fully Autonomous Drone Cinematography System Incorporating Zoom, Focus, Pose, and Scene Composition
abstract
We present CineMPC, a complete cinematographic system that autonomously controls a drone to film multiple targets recording user-specified aesthetic objectives. Existing solutions in autonomous cinematography control only the camera extrinsics, namely its position, and orientation. In contrast, CineMPC is the first solution that includes the camera intrinsic parameters in the control loop, which are essential tools for controlling cinematographic effects like focus, depth-of-field, and zoom. The system estimates the relative poses between the targets and the camera from an RGB-D image and optimizes a trajectory for the extrinsic and intrinsic camera parameters to film the artistic and technical requirements specified by the user. The drone and the camera are controlled in a nonlinear Model Predicted Control (MPC) loop by re-optimizing the trajectory at each time step in response to current conditions in the scene. The perception system of CineMPC can track the targets' position and orientation despite the camera effects. Experiments in a photo-realistic simulation and with a real platform demonstrate the capabilities of the system to achieve a full array of cinematographic effects that are not possible without the control of the intrinsics of the camera. Code for CineMPC is implemented following a modular architecture in ROS and released to the community
Pablo Pueyo, Juan Dendarieta, Eduardo Montijano, Ana Cristina Murillo, Mac Schwager
IEEE Trans. Robotics4
2023 A Framework for Fast Prototyping of Photo-realistic Environments with Multiple Pedestrians
abstract
Robotic applications involving people often require advanced perception systems to better understand complex real-world scenarios. To address this challenge, photo-realistic and physics simulators are gaining popularity as a means of generating accurate data labeling and designing scenarios for evaluating generalization capabilities, e.g., lighting changes, camera movements or different weather conditions. We develop a photo-realistic framework built on Unreal Engine and AirSim to generate easily scenarios with pedestrians and mobile robots. The framework is capable to generate random and customized trajectories for each person and provides up to 50 ready-to-use people models along with an API for their metadata retrieval. We demonstrate the usefulness of the proposed framework with a use case of multi-target tracking, a popular problem in real pedestrian scenarios. The notable feature variability in the obtained perception data is presented and evaluated.
Sara Casao, Andrés Otero, Álvaro Serra-Gómez, Ana Cristina Murillo, Javier Alonso-Mora, Eduardo Montijano
ICRA4
2023 CineTransfer: Controlling a Robot to Imitate Cinematographic Style from a Single Example
abstract
This work presents CineTransfer, an algorithmic framework that drives a robot to record a video sequence that mimics the cinematographic style of an input video. We propose features that abstract the aesthetic style of the input video, so the robot can transfer this style to a scene with visual details that are significantly different from the input video. The framework builds upon CineMPC, a tool that allows users to control cinematographic features, like subjects' position on the image and the depth of field, by manipulating the intrinsics and extrinsics of a cinematographic camera. However, CineMPC requires a human expert to specify the desired style of the shot (composition, camera motion, zoom, focus, etc). CineTransfer bridges this gap, aiming a fully autonomous cinematographic platform. The user chooses a single input video as a style guide. CineTransfer extracts and optimizes two important style features, the composition of the subject in the image and the scene depth of field, and provides instructions for CineMPC to control the robot to record an output sequence that matches these features as closely as possible. In contrast with other style transfer methods, our approach is a lightweight and portable framework which does not require deep network training or extensive datasets. Experiments with real and simulated videos demonstrate the system's ability to analyze and transfer style between recordings, and are available in the supplementary video11https://youtu.be/_QzNz5WUtpk
Pablo Pueyo, Eduardo Montijano, Ana Cristina Murillo, Mac Schwager
IROS3
2023 Tracking Adaptation to Improve SuperPoint for 3D Reconstruction in Endoscopy
Oscar León Barbed, J. M. M. Montiel, Pascal Fua, Ana Cristina Murillo
MICCAI (1)4
2023 CycleSTTN: A Learning-Based Temporal Model for Specular Augmentation in Endoscopy
abstract
Feature detection and matching is a computer vision problem that underpins different computer assisted techniques in endoscopy, including anatomy and lesion recognition, camera motion estimation, and 3D reconstruction. This problem is made extremely challenging due to the abundant presence of specular reflections. Most of the solutions proposed in the literature are based on filtering or masking out these regions as an additional processing step. There has been little investigation into explicitly learning robustness to such artefacts with single-step end-to-end training. In this paper, we propose an augmentation technique (CycleSTTN) that adds temporally consistent and realistic specularities to endoscopic videos. Such videos can act as ground truth data with known texture occluded behind the added specularities. We demonstrate that our image generation technique produces better results than a standard CycleGAN model. Additionally, we leverage this data augmentation to re-train a deep-learning based feature extractor (SuperPoint) and show that it improves. CycleSTTN code is made available here .
Rema Daher, Oscar León Barbed, Ana Cristina Murillo, Francisco Vasconcelos 0001, Danail Stoyanov
MICCAI (10)3
2023 Event Transformer$^+$. A Multi-Purpose Solution for Efficient Event Data Processing
abstract
Event cameras record sparse illumination changes with high temporal resolution and high dynamic range. Thanks to their sparse recording and low consumption, they are increasingly used in applications such as AR/VR and autonomous driving. Current top-performing methods often ignore specific event-data properties, leading to the development of generic but computationally expensive algorithms, while event-aware methods do not perform as well. We proposeEvent Transformer$^+$+, that improves our seminal workEvTwith a refined patch-based event representation and a more robust backbone to achieve more accurate results, while still benefiting from event-data sparsity to increase its efficiency. Additionally, we show how our system can work with different data modalities and propose specific output heads, for event-stream classification (i.e. action recognition) and per-pixel predictions (dense depth estimation). Evaluation results show better performance to the state-of-the-art while requiring minimal computation resources, both on GPU and CPU.
Alberto Sabater, Luis Montesano, Ana Cristina Murillo
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 CineMPC: Controlling Camera Intrinsics and Extrinsics for Autonomous Cinematography
abstract
We present CineMPC, an algorithm to autonomously control a UAV-borne video camera in a nonlinear Model Predicted Control (MPC) loop. CineMPC controls both the position and orientation of the camera-the camera extrinsics-as well as the lens focal length, focal distance, and aperture-the camera intrinsics. While some existing solutions autonomously control the position and orientation of the camera, no existing solutions also control the intrinsic parameters, which are essential tools for rich cinematographic expression. The intrinsic parameters control the parts of the scene that are focused or blurred, the viewers' perception of depth in the scene and the position of the targets in the image. CineMPC closes the loop from camera images to UAV trajectory and lens parameters in order to follow the desired relative trajectory and image composition as the targets move through the scene. Experiments using a photo-realistic environment demon-strate the capabilities of the proposed control framework to successfully achieve a full array of cinematographic effects not possible without full camera control.
Pablo Pueyo, Eduardo Montijano, Ana Cristina Murillo, Mac Schwager
ICRA3
2021 Semi-Supervised Semantic Segmentation with Pixel-Level Contrastive Learning from a Class-wise Memory Bank
abstract
This work presents a novel approach for semi-supervised semantic segmentation. The key element of this approach is our contrastive learning module that enforces the segmentation network to yield similar pixel-level feature representations for same-class samples across the whole dataset. To achieve this, we maintain a memory bank which is continuously updated with relevant and high-quality feature vectors from labeled data. In an end-to-end training, the features from both labeled and unlabeled data are optimized to be similar to same-class samples from the memory bank. Our approach not only outperforms the current state-of-the-art for semi-supervised semantic segmentation but also for semi-supervised domain adaptation on well-known public benchmarks, with larger improvements on the most challenging scenarios, i.e., less available labeled data. Code is available at https://github.com/Shathe/SemiSeg-Contrastive
Iñigo Alonso 0002, Alberto Sabater, David Ferstl, Luis Montesano, Ana Cristina Murillo
ICCV5
2021 Domain Adaptation in LiDAR Semantic Segmentation by Aligning Class Distributions
abstract
LiDAR semantic segmentation provides 3D semantic information about the environment, an essential cue for intelligent systems, such as autonomous vehicles, during their decision making processes. Unfortunately, the annotation process for this task is very expensive. To overcome this, it is key to find models that generalize well or adapt to additional domains where labeled data is limited. This work addresses the problem of unsupervised domain adaptation for LiDAR semantic segmentation models. We propose simple but effective strategies to reduce the domain shift by aligning the data distribution on the input space. Besides, we present a learning-based module to align the distribution of the semantic classes of the target domain to the source domain. Our approach achieves new state-of-the-art results on three different public datasets, which showcase adaptation to three different domains.
Iñigo Alonso 0002, Luis Riazuelo, Luis Montesano, Ana Cristina Murillo
ICINCO4
2021 Distributed Multi-Target Tracking in Camera Networks
abstract
Most recent works on multi-target tracking with multiple cameras focus on centralized systems. In contrast, this paper presents a multi-target tracking approach implemented in a distributed camera network. The advantages of distributed systems lie in lighter communication management, greater robustness to failures and local decision making. On the other hand, data association and information fusion are more challenging than in a centralized setup, mostly due to the lack of global and complete information. The proposed algorithm boosts the benefits of the Distributed-Consensus Kalman Filter with the support of a re-identification network and a distributed tracker manager module to facilitate consistent information. These techniques complement each other and facilitate the cross-camera data association in a simple and effective manner. We evaluate the whole system with known public data sets under different conditions demonstrating the advantages of combining all the modules. In addition, we compare our algorithm to some existing centralized tracking methods, outperforming their behavior in terms of accuracy and bandwidth usage.
Sara Casao, Abel Naya, Ana Cristina Murillo, Eduardo Montijano
ICRA3
2020 Robust and efficient post-processing for video object detection
abstract
Object recognition in video is an important task for plenty of applications, including autonomous driving perception, surveillance tasks, wearable devices or IoT networks. Object recognition using video data is more challenging than using still images due to blur, occlusions or rare object poses. Specific video detectors with high computational cost or standard image detectors together with a fast post-processing algorithm achieve the current state-of-the-art. This work introduces a novel post-processing pipeline that overcomes some of the limitations of previous post-processing methods by introducing a learning-based similarity evaluation between detections across frames. Our method improves the results of stat-of-the-art specific video detectors, specially regarding fast moving objects, and presents low resource requirements. And applied to efficient still image detectors, such as YOLO, provides comparable results to much more computationally intensive detectors.
Alberto Sabater, Luis Montesano, Ana Cristina Murillo
IROS3
2020 Incremental Learning of Object Models From Natural Human-Robot Interactions
abstract
In order to perform complex tasks in realistic human environments, robots need to be able to learn new concepts in the wild, incrementally, and through their interactions with humans. This article presents an end-to-end pipeline to learn object models incrementally during the human-robot interaction (HRI). The pipeline we propose consists of three parts: 1) recognizing the interaction type; 2) detecting the object that the interaction is targeting; and 3) learning incrementally the models from data recorded by the robot sensors. Our main contributions lie in the target object detection, guided by the recognized interaction, and in the incremental object learning. The novelty of our approach is the focus on natural, heterogeneous, and multimodal HRIs to incrementally learn new object models. Throughout the article, we highlight the main challenges associated with this problem, such as high degree of occlusion and clutter, domain change, low-resolution data, and interaction ambiguity. This article shows the benefits of using multiview approaches and combining visual and language features, and our experimental results outperform standard baselines.
Pablo Azagra, Javier Civera 0001, Ana Cristina Murillo
IEEE Trans Autom. Sci. Eng.3
2020 MiniNet: An Efficient Semantic Segmentation ConvNet for Real-Time Robotic Applications
abstract
Efficient models for semantic segmentation, in terms of memory, speed, and computation, could boost many robotic applications with strong computational and temporal restrictions. This article presents a detailed analysis of different techniques for efficient semantic segmentation. Following this analysis, we have developed a novel architecture, MiniNet-v2, an enhanced version of MiniNet. MiniNet-v2 is built considering the best option depending on CPU or GPU availability. It reaches comparable accuracy to the state-of-the-art models but uses less memory and computational resources. We validate and analyze the details of our architecture through a comprehensive set of experiments on public benchmarks (Cityscapes, Camvid, and COCO-Text datasets), showing its benefits over relevant prior work. Our experiments include a sample application where these models can boost existing robotic applications. Alonso, Íñigo; Riazuelo, Luis; Murillo, Ana C.
Iñigo Alonso 0002, Luis Riazuelo, Ana Cristina Murillo
IEEE Trans. Robotics3
2019 Integral Actions Towards Women in Engineering Recognition
abstract
This work presents integral actions towards women in engineering recognition organized according to educational stages they are directed: Stage 1, from early childhood education, primary education and secondary education; Stage 2, during university and Stage 3 after university. At stage 1 actions are devised to increase girl's interest on Science, Technology, Engineering and Mathematics (STEM). Emphasis is put on showing women in engineering as role models, illustrating engineers work and stressing the importance of diversity in working groups. Stage 2 is focused on making male and female students aware of the gender gap in engineering and the importance of diversity for innovation and training female students on known female narrow circumstances. Finally, at stage 3, the objective is to retain and promote women in the engineering profession. The specific actions developed at the three stages are presented. Their impact is discussed in order to accomplish effective actions for achieving gender balance towards excellence in Engineering.
Natalia Ayuso-Escuer, Sandra Baldassarri, Raquel Trillo Lado, Rosario Aragues, Belén Masiá, Pilar Molina-Gaudó, Ana Cristina Murillo, Eva Cerezo Bagdasari, María Villarroya-Gaudó
ETFA7
2019 Performance of object recognition in wearable videos
abstract
Wearable technologies are enabling plenty of new applications of computer vision, from life logging to health assistance. Many of them are required to recognize the elements of interest in the scene captured by the camera This work studies the problem of object detection and localization on videos captured by this type of camera. Wearable videos are a much more challenging scenario for object detection than standard images or even another type of videos, due to lower quality images (e.g. poor focus) or high clutter and occlusion common in wearable recordings. Existing work typically focuses on detecting the objects of focus or those being manipulated by the user wearing the camera. We perform a more general evaluation of the task of object detection in this type of video, because numerous applications, such as marketing studies, also need detecting objects which are not in focus by the user. This work presents a thorough study of the well known YOLO architecture, that offers an excellent trade-off between accuracy and speed, for the particular case of object detection in wearable video. We focus our study on the public ADL Dataset, but we also use additional public data for complementary evaluations. We run an exhaustive set of experiments with different variations of the original architecture and its training strategy. Our experiments drive to several conclusions about the most promising directions for our goal and point us to further research steps to improve detection in wearable videos.
Alberto Sabater, Luis Montesano, Ana Cristina Murillo
ETFA3
2019 Enhancing V-SLAM Keyframe Selection with an Efficient ConvNet for Semantic Analysis
abstract
Selecting relevant visual information from a video is a challenging task on its own and even more in robotics, due to strong computational restrictions. This work proposes a novel keyframe selection strategy based on image quality and semantic information, which boosts strategies currently used in Visual-SLAM (V-SLAM). Commonly used V-SLAM methods select keyframes based only on relative displacements and amount of tracked feature points. Our strategy to select more carefully these keyframes allows the robotic systems to make better use of them. With minimal computational cost, we show that our selection includes more relevant keyframes, which are useful for additional posterior recognition tasks, without penalizing the existing ones, mainly place recognition. A key ingredient is our novel CNN architecture to run a quick semantic image analysis at the onboard CPU of the robot. It provides sufficient accuracy significantly faster than related works. We demonstrate our hypothesis with several public datasets with challenging robotic data.
Iñigo Alonso 0002, Luis Riazuelo, Ana Cristina Murillo
ICRA3
2018 Semantic Segmentation from Sparse Labeling Using Multi-Level Superpixels
abstract
Semantic segmentation is a challenging problem that can benefit numerous robotics applications, since it provides information about the content at every image pixel. Solutions to this problem have recently witnessed a boost on performance and results thanks to deep learning approaches. Unfortunately, common deep learning models for semantic segmentation present several challenges which hinder real life applicability in many domains. A significant challenge is the need of pixel level labeling on large amounts of training images to be able to train those models, which implies a very high cost. This work proposes and validates a simple but effective approach to train dense semantic segmentation models from sparsely labeled data. Labeling only a few pixels per image reduces the human interaction required. We find many available datasets, e.g., environment monitoring data, that provide this kind of sparse labeling. Our approach is based on augmenting the sparse annotation to a dense one with the proposed adaptive superpixel segmentation propagation. We show that this label augmentation enables effective learning of state-of-the-art segmentation models, getting similar results to those models trained with dense ground-truth. We demonstrate the applicability of the presented approach to different image modalities in real domains (underwater, aerial and urban scenarios) with publicly available datasets.
Iñigo Alonso 0002, Ana Cristina Murillo
IROS2
2018 A generic tool for interactive complex image editing
Ana B. Cambra, Ana Cristina Murillo, Adolfo Muñoz 0001
Vis. Comput.2
2017 A multimodal dataset for object model learning from natural human-robot interaction
abstract
Learning object models in the wild from natural human interactions is an essential ability for robots to perform general tasks. In this paper we present a robocentric multimodal dataset addressing this key challenge. Our dataset focuses on interactions where the user teaches new objects to the robot in various ways. It contains synchronized recordings of visual (3 cameras) and audio data which provide a challenging evaluation framework for different tasks. Additionally, we present an end-to-end system that learns object models using object patches extracted from the recorded natural interactions. Our proposed pipeline follows these steps: (a) recognizing the interaction type, (b) detecting the object that the interaction is focusing on, and (c) learning the models from the extracted data. Our main contribution lies in the steps towards identifying the target object patches of the images. We demonstrate the advantages of combining language and visual features for the interaction recognition and use multiple views to improve the object modelling. Our experimental results show that our dataset is challenging due to occlusions and domain change with respect to typical object learning frameworks. The performance of common out-of-the-box classifiers trained on our data is low. We demonstrate that our algorithm outperforms such baselines.
Pablo Azagra, Florian Golemo, Yoan Mollard, Manuel Lopes 0001, Javier Civera 0001, Ana Cristina Murillo
IROS6
2016 Dense Labeling with User Interaction: an Example for Depth-Of-Field Simulation
Ana B. Cambra, Adolfo Muñoz 0001, Josechu J. Guerrero, Ana Cristina Murillo
BMVC4
2014 Line-based global descriptor for omnidirectional vision
abstract
Scene understanding is a widely studied problem in computer vision. Many works approach this problem in indoor environments assuming constraints about the scene, such as the typical Manhattan World assumption. The goal of this work is to design and evaluate a global descriptor for indoor panoramic images that encloses information about the 3D structure. This descriptor is based on the detection of representative lines of the scene, which encode the scene structure. Our work focuses on omnidirectional imagery, where observed lines are longer than in conventional images and the whole scene is captured in a single image. Experiments using two public datasets analyze the performance of the descriptor for scene categorization. We also analyze the influence of different parameters and show sample results for a navigation assistance application.
Alejandro Rituerto, Ana Cristina Murillo, Josechu J. Guerrero
ICIP2
2013 From Bikers to Surfers: Visual Recognition of Urban Tribes
abstract
Iljung S. Kwak1 [email protected] Ana C. Murillo2 [email protected] Peter N. Belhumeur3 [email protected] David Kriegman1 [email protected] Serge Belongie1 [email protected] 1 Dept. of Computer Science and Engineering University of California, San Diego, USA. 2 Dpt. Informatica e Ing. Sistemas Inst. Investigacion en Ingenieria de Aragon. University of Zaragoza, Spain. 3 Department of Computer Science Columbia University, USA.
Iljung S. Kwak, Ana Cristina Murillo, Peter N. Belhumeur, David J. Kriegman, Serge J. Belongie
BMVC2
2013 Localization in Urban Environments Using a Panoramic Gist Descriptor
abstract
Vision-based topological localization and mapping for autonomous robotic systems have received increased research interest in recent years. The need to map larger environments requires models at different levels of abstraction and additional abilities to deal with large amounts of data efficiently. Most successful approaches for appearance-based localization and mapping with large datasets typically represent locations using local image features. We study the feasibility of performing these tasks in urban environments using global descriptors instead and taking advantage of the increasingly common panoramic datasets. This paper describes how to represent a panorama using the global gist descriptor, while maintaining desirable invariance properties for location recognition and loop detection. We propose different gist similarity measures and algorithms for appearance-based localization and an online loop-closure detection method, where the probability of loop closure is determined in a Bayesian filtering framework using the proposed image representation. The extensive experimental validation in this paper shows that their performance in urban environments is comparable with local-feature-based approaches when using wide field-of-view images.
Ana Cristina Murillo, Gautam Singh, Jana Kosecka, Josechu J. Guerrero
IEEE Trans. Robotics1
2011 Label propagation in videos indoors with an incremental non-parametric model update
abstract
Semantic interpretation of the environment can significantly improve the capabilities of our autonomous robots. This work is focused on automatic semantic label propagation in video of indoor environments acquired by a mobile robot. Using a small number of training examples, we propose a new approach to recognize and label dominant background regions of interest, such as floor, wall and doors, and separate them from the remaining of foreground/object image categories. Our approach performs the labeling at the level of image superpixels. A simple non-parametric model is initialized from a few hand labeled examples in the first frame, and then it is propagated and updated along the sequence. We demonstrate the promising results obtained with our proposal in five different indoor sequences from different environments. The obtained semantic labeling can be used both for autonomous navigation and to provide better context for subsequent object detection.
J. Rituerto, Ana Cristina Murillo, Jana Kosecka
IROS2
2009 Improving topological maps for safer and robust navigation
abstract
Nowadays we frequently find big amounts of data to work with, what facilitates many robotic tasks and helps to solve perception problems. At the same time, this fact origins an interesting ongoing research problem: how to organize and arrange big sets of information to be useful in later uses. Topological mapping is a very useful tool to arrange and deal with big amounts of reference images for robotic tasks. There are many previous works on topological mapping and many others use this kind of maps for topological localization, planning and navigation. This work is focused on the problem of carefully design topological map building processes that facilitate the posterior robot tasks that use them and make them safer. We propose a new hierarchy of topological maps focused on this aspect. The experiments included in this paper were run outdoors using omnidirectional images and GPS information, and show the good topological maps obtained and how they allow robust and safer localization and navigation tasks.
Ana Cristina Murillo, Pablo Abad, Josechu J. Guerrero, Carlos Sagüés
IROS1
2008 Localization and Matching Using the Planar Trifocal Tensor With Bearing-Only Data
abstract
This paper addresses the robot and landmark localization problem from bearing-only data in three views, simultaneously to the robust association of this data. The localization algorithm is based on the 1-D trifocal tensor, which relates linearly the observed data and the robot localization parameters. The aim of this work is to bring this useful geometric construction from computer vision closer to robotic applications. One contribution is the evaluation of two linear approaches of estimating the 1-D tensor: the commonly used approach that needs seven bearing-only correspondences and another one that uses only five correspondences plus two calibration constraints. The results in this paper show that the inclusion of these constraints provides a simpler and faster solution and better estimation of robot and landmark locations in the presence of noise. Moreover, a new method that makes use of scene planes and requires only four correspondences is presented. This proposal improves the performance of the two previously mentioned methods in typical man-made scenarios with dominant planes, while it gives similar results in other cases. The three methods are evaluated with simulation tests as well as with experiments that perform automatic real data matching in conventional and omnidirectional images. The results show sufficient accuracy and stability to be used in robotic tasks such as navigation, global localization or initialization of simultaneous localization and mapping (SLAM) algorithms.
Josechu J. Guerrero, Ana Cristina Murillo, Carlos Sagüés
IEEE Trans. Robotics2
2007 SURF features for efficient robot localization with omnidirectional images
abstract
Many robotic applications work with visual reference maps, which usually consist of sets of more or less organized images. In these applications, there is a compromise between the density of reference data stored and the capacity to identify later the robot localization, when it is not exactly in the same position as one of the reference views. Here we propose the use of a recently developed feature, SURF, to improve the performance of appearance-based localization methods that perform image retrieval in large data sets. This feature is integrated with a vision-based algorithm that allows both topological and metric localization using omnidirectional images in a hierarchical approach. It uses pyramidal kernels for the topological localization and three-view geometric constraints for the metric one. Experiments with several omnidirectional images sets are shown, including comparisons with other typically used features (radial lines and SIFT). The advantages of this approach are proved, showing the use of SURF as the best compromise between efficiency and accuracy in the results.
Ana Cristina Murillo, Josechu J. Guerrero, Carlos Sagüés
ICRA1
2006 Localization with Omnidirectional Images using the Radial Trifocal Tensor
abstract
In this paper we present a technique to linearly recover 2D structure and motion in man made environments from three uncalibrated omnidirectional views. We use vertical lines from the scene which are projected as radial lines in the images and are automatically matched. The algorithm is based on a 1D radial trifocal tensor which encodes the relations of the three views and the projected lines. We include experiments with real images, which demonstrate the good performance of the method and its application to robotic tasks, such as robot localization based in a database of reference images or to obtain the initial values of robot and landmarks localization in SLAM algorithms
Carlos Sagüés, Ana Cristina Murillo, Josechu J. Guerrero, Toon Goedemé, Tinne Tuytelaars, Luc Van Gool
ICRA2
2006 Robot and Landmark Localization using Scene Planes and the 1D Trifocal Tensor
abstract
This paper presents a method for robot and landmarks 2D localization, in man made environments, taking profit of scene planes. The method uses bearing-only measurements that are robustly matched in three views. In our experiments we obtain them from vertical lines corresponding to natural landmarks. With these three view line-matches a trifocal tensor can be computed. This tensor contains the three views geometry and is used to estimate the aforementioned localization. As it is very usual to find a planar surface, we use the homography corresponding to that plane to obtain the tensor with one match less than the general case method. This implies lower computational complexity, mainly when trying a robust estimation, where we see a reduction in the number of iterations needed. Another advantage of obtaining an homography during the process is that it can help to automatically detect singular situations, such us totally planar scenes. It is shown that our proposal performs similarly to the general case method in a general scenario and better in case that we have some dominant plane in the scene. This paper includes simulated results proving this, as well as examples with real images, both with conventional and omnidirectional cameras
Ana Cristina Murillo, Josechu J. Guerrero, Carlos Sagüés
IROS1
2006 From lines to epipoles through planes in two views
Carlos Sagüés, Ana Cristina Murillo, F. Escudero, Josechu J. Guerrero
Pattern Recognit.2