Luis Miguel Bergasa

dblp:59/3675 · DBLP profile ↗
← Back
79ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-0087-3077ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 5 first-author · 13 since 2021Systems, architecture and hardware · 16Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 PO-GUISE+: Pose and Object Guided Transformer Token Selection for Efficient Driver Action Recognition
abstract
We address the task of identifying distracted driving by analyzing in-car videos using efficient transformers. Although transformer models have achieved outstanding performance in human action recognition tasks, their high computational costs limit their application onboard a vehicle. We introduce PO-GUISE+, a multi-task video transformer that, given an input clip, predicts the distracted driving action, the driver’s pose, and the interacting object. Our enhanced features for token selection are specifically adapted to driver actions by leveraging information about object interaction and the driver’s pose. With PO-GUISE+, we significantly reduce the model’s computational demands while maintaining or improving baseline accuracy across various computational budgets. Additionally, to evaluate our model’s performance in real-world scenarios, we have developed benchmarks on a Jetson computing platform, demonstrating its effectiveness across different configurations and computational budgets. Our model outperforms current state-of-the-art results on the Drive&Act, 100-Driver, and 3MDAD datasets, while having superior efficiency compared to existing video transformer-based methods.
Ricardo Pizarro 0002, Roberto Valle, Rafael Barea, José Miguel Buenaposada, Luis Baumela, Luis Miguel Bergasa
IEEE Trans. Intell. Transp. Syst.6
2025 Design and Development of a Digital Twin for Monitoring Railway Infrastructure
abstract
The aim of this work is to develop a digital twin application to ensure an optimal level of reliability when launching a larger project based on the identification of trains and detection of defects that allow for safe freight transport on spanish trains. The digital twin (DT) framework consists of three parts: the “physical product” which consists of a scanning camera placed on a track gantry, the “virtual product” which includes a model based on real-time data representing the freight car detected by the perception system, and the data flow connections. The camera images will be post-processed through an artificial intelligence detection model (YOLOv8), trained to detect all the elements necessary for the safety of the vehicle and the cargo. Field studies have demonstrated the effectiveness of the proposed digital twin framework and its potential to identify railcars and detect defects in freight wagons.
Javier Fuentes, Franck Fierro, Rafael Barea, María Elena López Guillén, Luis Miguel Bergasa
IV5
2025 Pose-guided token selection for the recognition of activities of daily living
abstract
Large pre-trained video transformers are becoming the standard architecture for video processing due to their exceptional accuracy. However, their computational complexity has been a major obstacle to their practical application in problems that require the recognition of precise motion patterns in video, such as in the recognition of Activities of Daily Living (ADL). Techniques like token pruning help mitigate their computational cost, but overlook some specific aspects of this task such as the actor movement. To address this we propose an improved token selection method that integrates semantic information from the ADL recognition task with that of human motion. Our model relies on a multi-task architecture that infers human pose and activity classification from RGB videos. We show that guiding token pruning with motion information significantly improves the trade-off between higher efficiency, obtained by reducing the number of tokens, and accuracy of the classification task. We evaluate our model on three popular ADL recognition benchmarks with their respective cross-subject and cross-view setups. In our experiments, a video transformer modified with our proposed modules sets a new state-of-the-art on the ADL recognition task whilst achieving significant reductions in computational cost. • PO-GUISE is a human motion and ADL-guided token selection for video transformers. • The resulting model improves the accuracy-GFLOPs trade-off during inference. • Our model integrates heatmap tokens for temporal and multi-actor prediction. • Sets new state-of-the-art results on ADL benchmarks at a reduced computational cost.
Ricardo Pizarro 0002, Roberto Valle, José Miguel Buenaposada, Luis Miguel Bergasa, Luis Baumela
Image Vis. Comput.4
2024 Decision Making for Autonomous Driving Stack: Shortening the Gap from Simulation to Real-World Implementations
abstract
This paper introduces a novel methodology for implementing a practical Decision Making module within an Autonomous Driving Stack, focusing on merge scenarios in urban environments. Our approach leverages Deep Reinforcement Learning and Curriculum Learning, structured into three stages: initial training in a lightweight simulator (SUMO), refinement in a high-fidelity simulation (CARLA) through a Digital Twin, and final validation in real-world scenarios with Parallel Execution. We propose a Partially Observable Markov Decision Process framework and employ the Trust Region Policy Optimization algorithm to train our agent. Our method significantly narrows the gap between simulated training and real-world application, offering a cost-effective and flexible solution for Autonomous Driving development. The paper details the experimental setup and outcomes in each stage, demonstrating the effectiveness of the proposed methodology.
Rodrigo Gutiérrez-Moreno, Rafael Barea, María Elena López Guillén, Juan Felipe Arango, Pedro A. Revenga, Luis Miguel Bergasa
IV6
2024 Do You Act Like You Talk? Exploring Pose-based Driver Action Classification with Speech Recognition Networks
abstract
Recognizing distractions on the road is crucial to reduce traffic accidents. Video-based networks are typically used, but are limited by their computational cost and are vulnerable to viewpoint changes. In this paper, we propose a novel approach for pose-based driver action classification using speech recognition networks, which is lighter and more viewpoint invariant that video-based one. We leverage the similarity in the encoding of information between audio and pose data, representing poses as key points over time. Our architecture is based on Squeezeformer, an efficient attentionbased speech recognition network. We introduce a selection of data augmentation techniques to enhance generalization. Experiments on the Drive&Act dataset demonstrate superior performance compared to state-of-the-art methods. Additionally, we explore the integration of object information and the impact of viewpoint changes. Our results highlight the effectiveness and robustness of speech recognition networks in pose-based action classification.
Pablo Pardo-Decimavilla, Luis Miguel Bergasa, Santiago Montiel-Marín, Miguel Antunes, Angel Llamazares
IV2
2024 DRVMon-VM: Distracted driver recognition using large pre-trained video transformers
abstract
Recent advancements in video transformers have significantly impacted the field of human action recognition. Leveraging these models for distracted driver action recognition could potentially revolutionize road safety measures and enhance Human-Machine Interaction (HMI) technologies. A factor that limits their potential use is the need for extensive data for model training. In this paper, we propose DRVMon-VM, a novel approach for the recognition of distracted driver actions. This is based on a large pre-trained video transformer called VideoMaeV2 (backbone) and a classification head as decoder, which are fine-tuned using a dual learning rate strategy and a medium-sized driver actions database complemented by various data augmentation techniques. Our proposed model exhibits a substantial improvement, exceeding previous results by 7.34% on the challenging Drive&Act dataset, thereby setting a new benchmark in this field.
Ricardo Pizarro 0002, Luis Miguel Bergasa, Luis Baumela, José Miguel Buenaposada, Rafael Barea
IV2
2024 Leveraging Driver Attention for an End-to-End Explainable Decision-Making From Frontal Images
abstract
Explaining the decision made by end-to-end autonomous driving is a difficult task. These approaches take raw sensor data and compute the decision as a black box with large deep learning models. Understanding the output of deep learning is a complex challenge due to the complicated nature of explainability; as data passes through the network, it becomes untraceable, making it difficult to understand. Explainability increases confidence in the decision by making the black box that drives the vehicle transparent to the user inside. Achieving a Level 5 autonomous vehicle necessitates the resolution of that challenging task. In this work, we propose a model that leverages the driver’s attention to obtain explainable decisions based on an attention map and the scene context. Our novel architecture addresses the task of obtaining a decision and its explanation from a single RGB sequence of the driving scene ahead. We base this architecture on the Transformer architecture with some efficiency tricks in order to use it at a reasonable frame rate. Moreover, we integrate in this proposal our previous ARAGAN model, which obtains SOTA attention maps, to improve the performance of the model thanks to understand the sequence as a human does. We train and validate our proposal on the BDD-OIA dataset, achieving on-pair results or even better than other state-of-the-art methods. Additionally, we present a simulation-based proof of concept demonstrating the model’s performance as a copilot in a close-loop vehicle to driver interaction.
Javier Araluce, Luis Miguel Bergasa, Manuel Ocaña, Angel Llamazares, María Elena López Guillén
IEEE Trans. Intell. Transp. Syst.2
2024 Efficient Baselines for Motion Prediction in Autonomous Driving
abstract
Motion Prediction (MP) of multiple surroundings agents is a crucial task in arbitrarily complex environments, from simple robots to Autonomous Driving Stacks (ADS). Current techniques tackle this problem using end-to-end pipelines, where the input data is usually a rendered top-view of the physical information and the past trajectories of the most relevant agents; leveraging this information is a must to obtain optimal performance. In that sense, a reliable ADS must produce reasonable predictions on time. However, despite many approaches use simple ConvNets and LSTMs to obtain the social latent features, State-of-the-Art (SOTA) models might be too complex for real-time applications when using both sources of information (map and past trajectories) as well as little interpretable, specifically considering the physical information. Moreover, the performance of such models highly depends on the number of available inputs for each particular traffic scenario, which are expensive to obtain, particularly, annotated High-Definition (HD) maps. In this work, we propose several efficient baselines for the well-known Argoverse 1 Motion Forecasting Benchmark. We aim to develop compact models using SOTA techniques for MP, including attention mechanisms and GNNs. Our lightweight models use standard social information and interpretable map information such as points from the driveable area and plausible centerlines by means of a novel physics-based heuristic step based on kinematic constraints, in opposition to black-box CNN-based or too-complex graphs methods for map encoding, to generate plausible multi-modal trajectories achieving up-to-pair accuracy with less operations and parameters than other SOTA methods. Our code is publicly available at https://github.com/Cram3r95/mapfe4mp.
Carlos Gómez Huélamo, Marcos V. Conde, Rafael Barea, Manuel Ocaña, Luis Miguel Bergasa
IEEE Trans. Intell. Transp. Syst.5
2023 Hybrid Decision Making for Autonomous Driving in Complex Urban Scenarios
abstract
Autonomous driving presents significant challenges due to the variability of behaviours exhibited by surrounding vehicles and the diversity of scenarios encountered. To address these challenges, we propose a hybrid architecture that combines traditional and deep learning techniques. Our architecture includes strategy, tactical and execution modules. Specifically, the strategy module defines the trajectory to be followed. Then, the tactical decision module employs a proximal policy optimization algorithm and deep reinforcement learning. Finally, the maneuver execution module uses a linear-quadratic regulator controller for trajectory tracking and a predictive model controller for lane change execution. This hybrid architecture and the comparison with other classical approaches are the main contributions of this research. Experimental results demonstrate that the proposed framework solves concatenated complex urban scenarios optimally.
Rodrigo Gutiérrez-Moreno, Rafael Barea, María Elena López Guillén, Juan Felipe Arango, Navil Abdeselam, Luis Miguel Bergasa
IV6
2022 ARAGAN: A dRiver Attention estimation model based on conditional Generative Adversarial Network
abstract
Predicting driver’s attention in complex driving scenarios is becoming a hot topic due to it helps the design of some autonomous driving tasks, optimizing visual scene understanding and contributing knowledge to the decision making. We introduce ARAGAN, a driver attention estimation model based on a conditional Generative Adversarial Network (cGAN). This architecture uses some of the most challenging and novel deep learning techniques to develop this task. It fuses adversarial learning with Multi-Head Attention mechanisms. To the best of our knowledge, this combination has never been applied to predict driver’s attention. Adversarial mechanism learns to map an attention image from an RGB traffic image while mapping the loss function. Attention mechanism contributes to the deep learning paradigm finding the most interesting feature maps inside the tensors of the net. In this work, we have adapted this concept to find the saliency areas in a driving scene. An ablation study with different architectures has been carried out, obtained the results in terms of some saliency metrics. Besides, a comparison with other state-of-the-art models has been driven, outperforming results in accuracy and performance, and showing that our proposal is adequate to be used on real-time applications. ARAGAN has been trained in BDDA and tested in BDDA and DADA2000, which are two of the most complex driver attention datasets available for research.
Javier Araluce, Luis Miguel Bergasa, Manuel Ocaña, Rafael Barea, María Elena López Guillén, Pedro A. Revenga
IV2
2022 HD maps: Exploiting OpenDRIVE potential for Path Planning and Map Monitoring
abstract
Autonomous vehicle (AV) is one of the most challenging engineering tasks of our era. High-Definition (HD) maps are a fundamental tool in the development of AVs, being considered as pseudo sensors that provide a trusted baseline that other sensors cannot. Our approach is focused on the use of OpenDRIVE standard based HD maps in order to conduct the different mapping and planning tasks involved in Autonomous Driving (AD). In this paper we present a method for exploiting the HD map potential for two specific purposes: i) Global Path Planning and ii) Monitoring the relevant lanes and regulatory elements around the ego-vehicle to support the perception module. Mapping and planning modules are connected to the other modules of the AV stack by using ROS (Robot Operating System). Our AD architecture has been validated both in local and CARLA Autonomous Driving Leaderboard cloud, where we can appreciate a considerable improvement in the metrics by incorporating information from the HD map, not only used to conduct the Global Path Planning task but also providing prior information to the Perception module. Code is available in https://github.com/AlejandroDiazD/opendrive-mapping-planning.
Alejandro Diaz-Diaz, Manuel Ocaña, Angel Llamazares, Carlos Gómez Huélamo, Pedro A. Revenga, Luis Miguel Bergasa
IV6
2022 How to build and validate a safe and reliable Autonomous Driving stack? A ROS based software modular architecture baseline
abstract
The implementation of Autonomous Driving stacks (ADS) is one of the most challenging engineering tasks of our era. Autonomous Vehicles (AVs) are expected to be driven in highly dynamic environments with a reliability greater than human beings and full autonomy. Furthermore, one of the most important topics is the way to democratize and accelerate the development and research of holistic validation to ensure the robustness of the vehicle. In this paper we present a powerful ROS (Robot Operating System) based modular ADS that achieves state-of-the-art results in challenging scenarios based on the CARLA (Car Learning to Act) simulator, outperforming several strong baselines in a novel evaluation setting which involves non-trivial traffic scenarios and adverse environmental conditions (Qualitative results). Our proposal ranks in second position in the CARLA Autonomous Driving Leaderboard (Map Track) and gets the best score considering modular pipelines, as a preliminary stage before implementing it in our real-world autonomous electric car. To encourage the use research in holistic development and testing, our code is publicly available at https://github.com/RobeSafe-UAH/CARLA_Leaderboard.
Carlos Gómez Huélamo, Alejandro Díaz, Javier Araluce, Miguel E. Ortiz, Rodrigo Gutiérrez, Juan Felipe Arango, Angel Llamazares, Luis Miguel Bergasa
IV8
2022 Beyond Supervised Deep Learning for Autonomous Driving
Luis Miguel Bergasa
VEHITS1
2022 Fine-tuning your answers: a bag of tricks for improving VQA models
Roberto Arroyo, Sergio Álvarez, Aitor Aller, Luis Miguel Bergasa, Miguel E. Ortiz
Multim. Tools Appl.4
2022 $360^{\circ }$ real-time and power-efficient 3D DAMOT for autonomous driving applications
abstract
Abstract Autonomous Driving (AD) promises an efficient, comfortable and safe driving experience. Nevertheless, fatalities involving vehicles equipped with Automated Driving Systems (ADSs) are on the rise, especially those related to the perception module of the vehicle. This paper presents a real-time and power-efficient 3D Multi-Object Detection and Tracking (DAMOT) method proposed for Intelligent Vehicles (IV) applications, allowing the vehicle to track $$360^{\circ }$$ 360 ∘ surrounding objects as a preliminary stage to perform trajectory forecasting to prevent collisions and anticipate the ego-vehicle to future traffic scenarios. First, we present our DAMOT pipeline based on Fast Encoders for object detection and a combination of a 3D Kalman Filter and Hungarian Algorithm, used for state estimation and data association respectively. We extend our previous work ellaborating a preliminary version of sensor fusion based DAMOT, merging the extracted features by a Convolutional Neural Network (CNN) using camera information for long-term re-identification and obstacles retrieved by the 3D object detector. Both pipelines exploit the concepts of lightweight Linux containers using the Docker approach to provide the system with isolation, flexibility and portability, and standard communication in robotics using the Robot Operating System (ROS). Second, both pipelines are validated using the recently proposed KITTI-3DMOT evaluation tool that demonstrates the full strength of 3D localization and tracking of a MOT system. Finally, the most efficient architecture is validated in some interesting traffic scenarios implemented in the CARLA (Car Learning to Act) open-source driving simulator and in our real-world autonomous electric car using the NVIDIA AGX Xavier, an AI embedded system for autonomous machines, studying its performance in a controlled but realistic urban environment with real-time execution (results).
Carlos Gómez Huélamo, Javier del Egido, Luis Miguel Bergasa, Rafael Barea, María Elena López Guillén, Javier Araluce, Miguel Antunes
Multim. Tools Appl.3
2022 Train here, drive there: ROS based end-to-end autonomous-driving pipeline validation in CARLA simulator using the NHTSA typology
abstract
Abstract Urban complex scenarios are the most challenging situations in the field of Autonomous Driving (AD). In that sense, an AD pipeline should be tested in countless environments and scenarios, escalating the cost and development time exponentially with a physical approach. In this paper we present a validation of our fully-autonomous driving architecture using the NHTSA (National Highway Traffic Safety Administration) protocol in the CARLA simulator, focusing on the analysis of our decision-making module, based on Hierarchical Interpreted Binary Petri Nets (HIBPN). First, the paper states the importance of using hyper-realistic simulators, as a preliminary help to real test, as well as an appropriate design of the traffic scenarios as the two current keys to build safe and robust AD technology. Second, our pipeline is introduced, which exploits the concepts of standard communication in robotics using the Robot Operating System (ROS) and the Docker approach to provide the system with isolation, flexibility and portability, describing the main modules and approaches to perform the navigation. Third, the CARLA simulator is described, outlining the steps carried out to merge our architecture with the simulator and the advantages to create ad-hoc driving scenarios for use cases validation instead of just modular evaluation. Finally, the architecture is validated using some challenging driving scenarios such as Pedestrian Crossing, Stop, Adaptive Cruise Control (ACC) and Unexpected Pedestrian. Some qualitative (video files: Simulation Use Cases ) and quantitative (linear velocity and trajectory splitted in the corresponding HIBPN states) results are presented for each use case, as well as an analysis of the temporal graphs associated to the Vulnerable Road Users (VRU) cases, validating our architecture in simulation as a preliminary stage before implementing it in our real autonomous electric car.
Carlos Gómez Huélamo, Javier del Egido, Luis Miguel Bergasa, Rafael Barea, María Elena López Guillén, Juan Felipe Arango, Javier Araluce, Joaquín López
Multim. Tools Appl.3
2022 Deep reinforcement learning based control for Autonomous Vehicles in CARLA
abstract
Abstract Nowadays, Artificial Intelligence (AI) is growing by leaps and bounds in almost all fields of technology, and Autonomous Vehicles (AV) research is one more of them. This paper proposes the using of algorithms based on Deep Learning (DL) in the control layer of an autonomous vehicle. More specifically, Deep Reinforcement Learning (DRL) algorithms such as Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG) are implemented in order to compare results between them. The aim of this work is to obtain a trained model, applying a DRL algorithm, able of sending control commands to the vehicle to navigate properly and efficiently following a determined route. In addition, for each of the algorithms, several agents are presented as a solution, so that each of these agents uses different data sources to achieve the vehicle control commands. For this purpose, an open-source simulator such as CARLA is used, providing to the system with the ability to perform a multitude of tests without any risk into an hyper-realistic urban simulation environment, something that is unthinkable in the real world. The results obtained show that both DQN and DDPG reach the goal, but DDPG obtains a better performance. DDPG perfoms trajectories very similar to classic controller as LQR. In both cases RMSE is lower than 0.1m following trajectories with a range 180-700m. To conclude, some conclusions and future works are commented.
Óscar Pérez-Gil, Rafael Barea, María Elena López Guillén, Luis Miguel Bergasa, Carlos Gómez Huélamo, Rodrigo Gutiérrez, Alejandro Díaz
Multim. Tools Appl.4
2021 Validation Method of a Self-Driving Architecture for Unexpected Pedestrian Scenario in CARLA Simulator
abstract
This paper introduces a method to validate autonomous navigation frameworks, in simulation using CARLA Simulator, fulfilling the requirements of the Euro-NCAP evaluation. We propose the protocol for evaluating an unexpected pedestrian scenario, where a walker suddenly invades the road and the vehicle has to react in a safe way. Standard validation metrics are created for this use case, which are generalizables for other use cases. To support the proposal, we describe our ROS (Robot Operating System) based Self-Driving architecture, open source and implemented in an electric vehicle. Then, we explain the procedures and requirements needed for the validation protocol that we propose. Finally, we show the metrics and results obtained in simulation for different ego-vehicle velocities and weather conditions. The scenarios implemented in Carla are publicly available22https://github.com/RobeSafe-UAH/scenarios.
Rodrigo Gutiérrez, Juan Felipe Arango, Carlos Gómez Huélamo, Luis Miguel Bergasa, Rafael Barea, Javier Araluce
IV4
2021 SmartMOT: Exploiting the fusion of HDMaps and Multi-Object Tracking for Real-Time scene understanding in Intelligent Vehicles applications
abstract
Behaviour prediction in multi-agent and dynamic environments is crucial in the context of intelligent vehicles, due to the complex interactions and representations of road participants (such as vehicles, cyclists or pedestrians) and road context information (e.g. traffic lights, lanes and regulatory elements). This paper presents SmartMOT, a simple yet powerful pipeline that fuses the concepts of tracking-by-detection and semantic information of HD maps, in particular using the OpenDrive format specification to describe the road network's logic, to design a real-time and power-efficient Multi-Object Tracking (MOT) pipeline which is then used to predict the future trajectories of the obstacles assuming a CTRV (Constant Turn Rate and Velocity) model. The system pipeline is fed by the monitorized lanes around the ego-vehicle, which are calculated by the planning layer, the ego-vehicle status, that contains its odometry and velocity and the corresponding Bird's Eye View (BEV) detections. Based on some well-defined traffic rules, HD map geometric and semantic information are used in the initial stage of the tracking module, made up by a BEV Kalman Filter and Hungarian algorithm are used for state estimation and data association respectively, to track only the most relevant detections around the ego-vehicle, as well as in the subsequent steps to predict new relevant traffic participants or delete trackers that go outside the monitorized area, helping the perception layer to understand the scene in terms of behavioural use cases to feed the executive layer of the vehicle. First, our system pipeline is described, exploiting the concepts of lightweight Linux containers using Docker to provide the system with isolation, flexibility and portability, and standard communication in robotics using the Robot Operating System (ROS). Second, the system is validated (Qualitative results11SmartMOT: https://cutt.ly/uk9ziaq) in the CARLA simulator fulfilling the requirements of the Euro-NCAP evaluation for Unexpected Vulnerable Road Users (VRU), where a pedestrian suddenly jumps into the road and the vehicle has to avoid collision or reduce the impact velocity as much as possible. Finally, a comparison between our HD map based perception strategy and our previous work with rectangular based approach is carried out, demonstrating how incorporating enriched topological map information increases the reliability of the Autonomous Driving (AD) stack. Code is publicly available https://github.com/Cram3r95/map-filtered-mot as a ROS package.
Carlos Gómez Huélamo, Luis Miguel Bergasa, Rodrigo Gutiérrez, Juan Felipe Arango, Alejandro Díaz
IV2
2021 Deep Reinforcement Learning based control algorithms: Training and validation using the ROS Framework in CARLA Simulator for Self-Driving applications
abstract
This paper presents a Deep Reinforcement Learning (DRL) framework adapted and trained for Autonomous Vehicles (AVs) purposes. To do that, we propose a novel software architecture for training and validating DRL based control algorithms that exploits the concepts of standard communication in robotics using the Robot Operating System (ROS), the Docker approach to provide the system with portability, isolation and flexibility, and CARLA (CAR Learning to Act) as our hyper-realistic open-source simulation platform. First, the algorithm is introduced in the context of Self-Driving and DRL tasks. Second, we highlight the steps to merge the proposed algorithm with ROS, Docker and the CARLA simulator, as well as how the training stage is carried out to generate our own model, specifically designed for the AV paradigm. Finally, regarding our proposed validation architecture, the paper compares the trained model with other state-of-the-art traditional control approaches, demonstrating the full strength of our DL based control algorithm, as a preliminary stage before implementing it in our real-world autonomous electric car.
Óscar Pérez-Gil, Rafael Barea, María Elena López Guillén, Luis Miguel Bergasa, Carlos Gómez Huélamo, Rodrigo Gutiérrez, Alejandro Díaz
IV4
2020 PASS: Panoramic Annular Semantic Segmentation
abstract
Pixel-wise semantic segmentation is capable of unifying most of driving scene perception tasks, and has enabled striking progress in the context of navigation assistance, where an entire surrounding sensing is vital. However, current mainstream semantic segmenters are predominantly benchmarked against datasets featuring narrow Field of View (FoV), and a large part of vision-based intelligent vehicles use only a forward-facing camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the whole surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the efforts entailed to create fully dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments, and fuse them to fulfill panoramic scene parsing. The innovation lies in the network adaptation to enable smooth and seamless segmentation, combined with an extended set of heterogeneous data augmentations to attain robustness in panoramic imagery. A comprehensive variety of experiments demonstrates the effectiveness for real-world surrounding perception in a single PASS, while the adaptation proposal is exceptionally positive for state-of-the-art efficient networks.
Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Kaiwei Wang
IEEE Trans. Intell. Transp. Syst.3
2019 Bridging the Day and Night Domain Gap for Semantic Segmentation
abstract
Perception in autonomous vehicles has progressed exponentially in the last years thanks to the advances of vision-based methods such as Convolutional Neural Networks (CNNs). Current deep networks are both efficient and reliable, at least in standard conditions, standing as a suitable solution for the perception tasks of autonomous vehicles. However, there is a large accuracy downgrade when these methods are taken to adverse conditions such as nighttime. In this paper, we study methods to alleviate this accuracy gap by using recent techniques such as Generative Adversarial Networks (GANs). We explore diverse options such as enlarging the dataset to cover these domains in unsupervised training or adapting the images on-the-fly during inference to a comfortable domain such as sunny daylight in a pre-processing step. The results show some interesting insights and demonstrate that both proposed approaches considerably reduce the domain gap, allowing IV perception systems to work reliably also at night.
Eduardo Romera, Luis Miguel Bergasa, Kailun Yang 0001, José M. Álvarez 0004, Rafael Barea
IV2
2019 Can we PASS beyond the Field of View? Panoramic Annular Semantic Segmentation for Real-World Surrounding Perception
abstract
Pixel-wise semantic segmentation unifies distinct scene perception tasks in a coherent way, and has catalyzed notable progress in autonomous and assisted navigation, where a whole surrounding perception is vital. However, current mainstream semantic segmenters are normally benchmarked against datasets with narrow Field of View (FoV), and most vision-based navigation systems use only a forward-view camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the entire surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the effort entailed to create dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments and fuse them to fulfill smooth and seamless panoramic scene parsing. Beyond the enlarged FoV,we extend focal length-related and style transfer-based data augmentations, to robustify the semantic segmenter against distortions and blurs in panoramic imagery. A comprehensive variety of experiments demonstrates the qualified robustness of our proposal for realworld surrounding understanding.
Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Dongming Sun, Kaiwei Wang
IV3
2018 Train Here, Deploy There: Robust Segmentation in Unseen Domains
abstract
Semantic Segmentation methods play a key role in today’s Autonomous Driving research, since they provide a global understanding of the traffic scene for upper-level tasks like navigation. However, main research efforts are being put on enlarging deep architectures to achieve marginal accuracy boosts in existing datasets, forgetting that these algorithms must be deployed in a real vehicle with images that were not seen during training. On the other hand, achieving robustness in any domain is not an easy task, since deep networks are prone to overfitting even with thousands of training images. In this paper, we study in a systematic way what is the gap between the concepts of “accuracy” and “robustness”. A comprehensive set of experiments demonstrates the relevance of using data augmentation to yield models that can produce robust semantic segmentation outputs in any domain. Our results suggest that the existing domain gap can be significantly reduced when appropriate augmentation techniques regarding geometry (position and shape) and texture (color and illumination) are applied. In addition, the proposed training process results in better calibrated models, which is of special relevance to assess the robustness of current systems.
Eduardo Romera, Luis Miguel Bergasa, José M. Álvarez 0004, Mohan M. Trivedi
Intelligent Vehicles Symposium2
2018 CNN-based Fisheye Image Real-Time Semantic Segmentation
abstract
Semantic segmentation based on Convolutional Neural Networks (CNNs) has been proven as an efficient way of facing scene understanding for autonomous driving applications. Traditionally, environment information is acquired using narrow-angle pin-hole cameras, but autonomous vehicles need wider field of view to perceive the complex surrounding, especially in urban traffic scenes. Fisheye cameras have begun to play an increasingly role to cover this need. This paper presents a real-time CNN-based semantic segmentation solution for urban traffic images using fisheye cameras. We adapt our Efficient Residual Factorized CNN (ERFNet) architecture to handle distorted fish-eye images. A new fisheye image dataset for semantic segmentation from the existing CityScapes dataset is generated to train and evaluate our CNN. We also test a data augmentation suggestion for fisheye image proposed in [1]. Experiments show outstanding results of our proposal regarding other methods of the state of the art.
Álvaro Sáez, Luis Miguel Bergasa, Eduardo Romera, María Elena López Guillén, Rafael Barea, Rafael Sanz
Intelligent Vehicles Symposium2
2018 Unifying terrain awareness through real-time semantic segmentation
abstract
Active research on computer vision accelerates the progress in autonomous driving. Following this trend, we aim to leverage the recently emerged methods for Intelligent Vehicles (IV), and transfer them to develop navigation assistive technologies for the Visually Impaired (VI). This topic grows notoriously challenging as it requires to detect a variety of scenes towards higher level of assistance. Computer vision based techniques with monocular detectors or depth sensors sprung up within years of research. These separate approaches achieved remarkable results with relatively low processing time, and improved the mobility of visually impaired people to a large extent. However, running all detectors jointly increases the latency and burdens the computational resources. In this paper, we put forward to seize pixel-wise semantic segmentation to cover the perception needs of navigational assistance in a unified way. This is critical not only for the terrain awareness regarding traversable areas, sidewalks, stairs and water hazards, but also for the avoidance of short-range obstacles, fast-approaching pedestrians and vehicles. At the heart of our proposal is a combination of efficient residual factorized network (ERFNet), pyramid scene parsing network (PSPNet) and 3D point cloud based segmentation. This approach proves to be with qualified accuracy and speed for real-world applications by a comprehensive set of experiments on a wearable navigation system.
Kailun Yang 0001, Luis Miguel Bergasa, Eduardo Romera, Ruiqi Cheng, Tianxue Chen, Kaiwei Wang
Intelligent Vehicles Symposium2
2018 ERFNet: Efficient Residual Factorized ConvNet for Real-Time Semantic Segmentation
abstract
Semantic segmentation is a challenging task that addresses most of the perception needs of intelligent vehicles (IVs) in an unified way. Deep neural networks excel at this task, as they can be trained end-to-end to accurately classify multiple object categories in an image at pixel level. However, a good tradeoff between high quality and computational resources is yet not present in the state-of-the-art semantic segmentation approaches, limiting their application in real vehicles. In this paper, we propose a deep architecture that is able to run in real time while providing accurate semantic segmentation. The core of our architecture is a novel layer that uses residual connections and factorized convolutions in order to remain efficient while retaining remarkable accuracy. Our approach is able to run at over 83 FPS in a single Titan X, and 7 FPS in a Jetson TX1 (embedded device). A comprehensive set of experiments on the publicly available Cityscapes data set demonstrates that our system achieves an accuracy that is similar to the state of the art, while being orders of magnitude faster to compute than other architectures that achieve top precision. The resulting tradeoff makes our model an ideal approach for scene understanding in IV applications. The code is publicly available at: https://github.com/Eromera/erfnet.
Eduardo Romera, José M. Álvarez 0004, Luis Miguel Bergasa, Roberto Arroyo
IEEE Trans. Intell. Transp. Syst.3
2018 Guest Editorial Introduction to the Special Issue on Robust and Efficient Vision Techniques for Intelligent Vehicles
abstract
In recent years, intelligent vehicles have been a hot topic for both research and industry communities. Since the whole system is a comprehensive integration of many advanced techniques, their respective development and improvement become fundamentally important.
Qi Wang 0009, Luis Miguel Bergasa, José M. Álvarez 0004
IEEE Trans. Intell. Transp. Syst.2
2017 Efficient ConvNet for real-time semantic segmentation
abstract
Semantic segmentation is a task that covers most of the perception needs of intelligent vehicles in an unified way. ConvNets excel at this task, as they can be trained end-to-end to accurately classify multiple object categories in an image at the pixel level. However, current approaches normally involve complex architectures that are expensive in terms of computational resources and are not feasible for ITS applications. In this paper, we propose a deep architecture that is able to run in real-time while providing accurate semantic segmentation. The core of our ConvNet is a novel layer that uses residual connections and factorized convolutions in order to remain highly efficient while still retaining remarkable performance. Our network is able to run at 83 FPS in a single Titan X, and at more than 7 FPS in a Jetson TX1 (embedded GPU). A comprehensive set of experiments demonstrates that our system, trained from scratch on the challenging Cityscapes dataset, achieves a classification performance that is among the state of the art, while being orders of magnitude faster to compute than other architectures that achieve top precision. This makes our model an ideal approach for scene understanding in intelligent vehicles applications.
Eduardo Romera, José M. Álvarez 0004, Luis Miguel Bergasa, Roberto Arroyo
Intelligent Vehicles Symposium3
2016 Fusion and binarization of CNN features for robust topological localization across seasons
abstract
The extreme variability in the appearance of a place across the four seasons of the year is one of the most challenging problems in life-long visual topological localization for mobile robotic systems and intelligent vehicles. Traditional solutions to this problem are based on the description of images using hand-crafted features, which have been shown to offer moderate invariance against seasonal changes. In this paper, we present a new proposal focused on automatically learned descriptors, which are processed by means of a technique recently popularized in the computer vision community: Convolutional Neural Networks (CNNs). The novelty of our approach relies on fusing the image information from multiple convolutional layers at several levels and granularities. In addition, we compress the redundant data of CNN features into a tractable number of bits for efficient and robust place recognition. The final descriptor is reduced by applying simple compression and binarization techniques for fast matching using the Hamming distance. An exhaustive experimental evaluation confirms the improved performance of our proposal (CNN-VTL) with respect to state-of-the-art methods over varied long-term datasets recorded across seasons.
Roberto Arroyo, Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Eduardo Romera
IROS3
2015 Towards life-long visual localization using an efficient matching of binary sequences from images
abstract
Life-long visual localization is one of the most challenging topics in robotics over the last few years. The difficulty of this task is in the strong appearance changes that a place suffers due to dynamic elements, illumination, weather or seasons. In this paper, we propose a novel method (ABLE-M) to cope with the main problems of carrying out a robust visual topological localization along time. The novelty of our approach resides in the description of sequences of monocular images as binary codes, which are extracted from a global LDB descriptor and efficiently matched using FLANN for fast nearest neighbor search. Besides, an illumination invariant technique is applied. The usage of the proposed binary description and matching method provides a reduction of memory and computational costs, which is necessary for long-term performance. Our proposal is evaluated in different life-long navigation scenarios, where ABLE-M outperforms some of the main state-of-the-art algorithms, such as WI-SURF, BRIEF-Gist, FAB-MAP or SeqSLAM. Tests are presented for four public datasets where a same route is traversed at different times of day or night, along the months or across all four seasons.
Roberto Arroyo, Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Eduardo Romera
ICRA3
2015 Fast pixelwise road inference based on Uniformly Reweighted Belief Propagation
abstract
The future of autonomous vehicles and driver assistance systems is underpinned by the need of fast and efficient approaches for road scene understanding. Despite the large explored paths for road detection, there is still a research gap for incorporating image understanding capabilities in intelligent vehicles. This paper presents a pixelwise segmentation of roads from monocular images. The proposal is based on a probabilistic graphical model and a set of algorithms and configurations chosen to speed up the inference of the road pixels. In brief, the proposed method employs Conditional Random Fields and Uniformly Reweighted Belief Propagation. Besides, the approach is ranked on the KITTI ROAD dataset yielding state-of-the-art results with the lowest runtime per image using a standard PC.
Mario Passani, José Javier Yebes Torres, Luis Miguel Bergasa
Intelligent Vehicles Symposium3
2015 Expert video-surveillance system for real-time detection of suspicious behaviors in shopping malls
Roberto Arroyo, José Javier Yebes Torres, Luis Miguel Bergasa, Iván García 0001, Javier Almazán
Expert Syst. Appl.3
2014 Fast and effective visual place recognition using binary codes and disparity information
abstract
We present a novel approach for place recognition and loop closure detection based on binary codes and disparity information using stereo images. Our method (ABLE-S) applies the Local Difference Binary (LDB) descriptor in a global framework to obtain a robust global image description, which is initially based on intensity and gradient pairwise comparisons. LDB has a higher descriptiveness power than other popular alternatives such as BRIEF, which only relies on intensity. In addition, we integrate disparity information into the binary descriptor (D-LDB). Disparity provides valuable information which decreases the effect of some typical problems in place recognition such as perceptual aliasing. The KITTI Odometry dataset is mainly used to test our approach due to its varied environments, challenging situations and length. Additionally, a loop closure ground-truth is introduced in this work for the KITTI Odometry benchmark with the aim of standardizing a robust evaluation methodology for comparing different previous algorithms against our method and for future benchmarking of new proposals. Attending to the presented results, our method allows a fast and more effective visual loop closure detection compared to state-of-the-art algorithms such as FAB-MAP, WI-SURF and BRIEF-Gist.
Roberto Arroyo, Pablo Fernández Alcantarilla, Luis Miguel Bergasa, José Javier Yebes Torres, Sebastián Bronte
IROS3
2014 Real-time sequential model-based non-rigid SFM
abstract
Tracking non-rigid objects from video is useful in robotic systems such as HMIs or robotic manipulator arms which interact with deformable objects. This paper proposes a method for sequential model-based 3D reconstruction of deformable objects and camera localization in real time. Non-rigid SFM methods commonly process a video sequence offline in a batch way. While there are real-time methods for rigid models, reconstruction of deformable 3D shapes for real-time applications is still unsolved. Dense approaches offer promising results, but processing all frames in batch, offline. We propose a real-time non-rigid reconstruction method based on a known deformable model. Object shape and pose is tracked by real-time estimation of camera pose and deformation coefficients. An extensive evaluation of the algorithm on several data sets, and comparison with state-of-the-art techniques is performed. The tests include different outlier rates, noise levels and occlusions handling.
Sebastián Bronte, Marco Paladini, Luis Miguel Bergasa, Lourdes Agapito, Roberto Arroyo
IROS3
2014 Bidirectional loop closure detection on panoramas for visual navigation
abstract
Visual loop closure detection plays a key role in navigation systems for intelligent vehicles. Nowadays, state-of-the-art algorithms are focused on unidirectional loop closures, but there are situations where they are not sufficient for identifying previously visited places. Therefore, the detection of bidirectional loop closures when a place is revisited in a different direction provides a more robust visual navigation. We propose a novel approach for identifying bidirectional loop closures on panoramic image sequences. Our proposal combines global binary descriptors and a matching strategy based on cross-correlation of sub-panoramas, which are defined as the different parts of a panorama. A set of experiments considering several binary descriptors (ORB, BRISK, FREAK, LDB) is provided, where LDB excels as the most suitable. The proposed matching proffers a reliable bidirectional loop closure detection, which is not efficiently solved in any other previous research. Our method is successfully validated and compared against FAB-MAP and BRIEF-Gist. The Ford Campus and the Oxford New College datasets are considered for evaluation.
Roberto Arroyo, Pablo Fernández Alcantarilla, Luis Miguel Bergasa, José Javier Yebes Torres, Sergio Gamez
Intelligent Vehicles Symposium3
2014 DriveSafe: An app for alerting inattentive drivers and scoring driving behaviors
abstract
This paper presents DriveSafe, a new driver safety app for iPhones that detects inattentive driving behaviors and gives corresponding feedback to drivers, scoring their driving and alerting them in case their behaviors are unsafe. It uses computer vision and pattern recognition techniques on the iPhone to assess whether the driver is drowsy or distracted using the rear-camera, the microphone, the inertial sensors and the GPS. We present the general architecture of DriveSafe and evaluate its performance using data from 12 drivers in two different studies. The first one evaluates the detection of some inattentive driving behaviors obtaining an overall precision of 82% at 92% of recall. The second one compares the scores between DriveSafe vs the commercial AXA Drive app obtaining a better valuation to its operation. DriveSafe is the first app for smartphones based on inbuilt sensors able to detect inattentive behaviors evaluating the quality of the driving at the same time. It represents a new disruptive technology because, on the one hand, it provides similar ADAS features that found in luxury cars, and on the other hand, it presents a viable alternative for the “blackboxes” installed in vehicles by the insurance companies.
Luis Miguel Bergasa, Daniel Almeria, Javier Almazán, José Javier Yebes Torres, Roberto Arroyo
Intelligent Vehicles Symposium1
2014 Supervised learning and evaluation of KITTI's cars detector with DPM
abstract
This paper carries out a discussion on the supervised learning of a car detector built as a Discriminative Part-based Model (DPM) from images in the recently published KITTI benchmark suite as part of the object detection and orientation estimation challenge. We present a wide set of experiments and many hints on the different ways to supervise and enhance the well-known DPM on a challenging and naturalistic urban dataset as KITTI. The evaluation algorithm and metrics, the selection of a clean but representative subset of training samples and the DPM tuning are key factors to learn an object detector in a supervised fashion. We provide evidence of subtle differences in performance depending on these aspects. Besides, the generalization of the trained models to an independent dataset is validated by 5-fold cross-validation.
José Javier Yebes Torres, Luis Miguel Bergasa, Roberto Arroyo, Alberto Lazaro
Intelligent Vehicles Symposium2
2014 Text Detection and Recognition on Traffic Panels From Street-Level Imagery Using Visual Appearance
abstract
Traffic sign detection and recognition has been thoroughly studied for a long time. However, traffic panel detection and recognition still remains a challenge in computer vision due to its different types and the huge variability of the information depicted in them. This paper presents a method to detect traffic panels in street-level images and to recognize the information contained on them, as an application to intelligent transportation systems (ITS). The main purpose can be to make an automatic inventory of the traffic panels located in a road to support road maintenance and to assist drivers. Our proposal extracts local descriptors at some interest keypoints after applying blue and white color segmentation. Then, images are represented as a “bag of visual words” and classified using Naïve Bayes or support vector machines. This visual appearance categorization method is a new approach for traffic panel detection in the state of the art. Finally, our own text detection and recognition method is applied on those images where a traffic panel has been detected, in order to automatically read and save the information depicted in the panels. We propose a language model partly based on a dynamic dictionary for a limited geographical area using a reverse geocoding service. Experimental results on real images from Google Street View prove the efficiency of the proposed method and give way to using street-level images for different applications on ITS.
Álvaro Gonzalez, Luis Miguel Bergasa, José Javier Yebes Torres
IEEE Trans. Intell. Transp. Syst.2
2013 Full auto-calibration of a smartphone on board a vehicle using IMU and GPS embedded sensors
abstract
Nowadays, smartphones are widely used in the world, and generally, they are equipped with many sensors. In this paper we study how powerful the low-cost embedded IMU and GPS could become for Intelligent Vehicles. The information given by accelerometer and gyroscope is useful if the relations between the smartphone reference system, the vehicle reference system and the world reference system are known. Commonly, the magnetometer sensor is used to determine the orientation of the smartphone, but its main drawback is the high influence of electromagnetic interference. In view of this, we propose a novel automatic method to calibrate a smartphone on board a vehicle using its embedded IMU and GPS, based on longitudinal vehicle acceleration. To the best of our knowledge, this is the first attempt to estimate the yaw angle of a smartphone relative to a vehicle in every case, even on non-zero slope roads. Furthermore, in order to decrease the impact of IMU noise, an algorithm based on Kalman Filter and fitting a mixture of Gaussians is introduced. The results show that the system achieves high accuracy, the typical error is 1%, and is immune to electromagnetic interference.
Javier Almazán, Luis Miguel Bergasa, José Javier Yebes Torres, Rafael Barea, Roberto Arroyo
Intelligent Vehicles Symposium2
2013 Traffic panels detection using visual appearance
abstract
Traffic signs detection has been thoroughly studied for a long time. However, road panels detection still remains a challenge in computer vision due to the huge variability of types of traffic panels, as the information depicted in them is not restricted. This paper presents a method to detect traffic panels in street-level images as an application to Intelligent Transportation Systems (ITS), since the main purpose can be to make an automatic inventory of the traffic panels located in a road to support maintenance and to assist drivers in order to improve human quality of life. The proposed method extracts local descriptors at some interest points after applying a color detection method for blue and white pixels. Then, the images are modeled using a Bag of Visual Words technique and classified using Naïve Bayes theory and SVM. Experimental results on real images from Google Street View prove the efficiency of the proposed method and give way to using street-level images for different applications on robotics and ITS.
Álvaro González, Luis Miguel Bergasa, José Javier Yebes Torres, Javier Almazán
Intelligent Vehicles Symposium2
2013 Gauge-SURF descriptors
Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Andrew J. Davison
Image Vis. Comput.2
2013 A text reading algorithm for natural images
Álvaro Gonzalez, Luis Miguel Bergasa
Image Vis. Comput.2
2012 A character recognition method in natural scene images
Álvaro Gonzalez, Luis Miguel Bergasa, José Javier Yebes Torres, Sebastián Bronte
ICPR2
2012 Text location in complex images
Álvaro Gonzalez, Luis Miguel Bergasa, José Javier Yebes Torres, Sebastián Bronte
ICPR2
2012 On combining visual SLAM and dense scene flow to increase the robustness of localization and mapping in dynamic environments
abstract
In this paper, we introduce the concept of dense scene flow for visual SLAM applications. Traditional visual SLAM methods assume static features in the environment and that a dominant part of the scene changes only due to camera egomotion. These assumptions make traditional visual SLAM methods prone to failure in crowded real-world dynamic environments with many independently moving objects, such as the typical environments for the visually impaired. By means of a dense scene flow representation, moving objects can be detected. In this way, the visual SLAM process can be improved considerably, by not adding erroneous measurements into the estimation, yielding more consistent and improved localization and mapping results. We show large-scale visual SLAM results in challenging indoor and outdoor crowded environments with real visually impaired users. In particular, we performed experiments inside the Atocha railway station and in the city-center of Alcalá de Henares, both in Madrid, Spain. Our results show that the combination of visual SLAM and dense scene flow allows to obtain an accurate localization, improving considerably the results of traditional visual SLAM methods and GPS-based approaches.
Pablo Fernández Alcantarilla, José Javier Yebes Torres, Javier Almazán, Luis Miguel Bergasa
ICRA4
2012 Vision-based drowsiness detector for real driving conditions
abstract
This paper presents a non-intrusive approach for drowsiness detection, based on computer vision. It is installed in a car and it is able to work under real operation conditions. An IR camera is placed in front of the driver, in the dashboard, in order to detect his face and obtain drowsiness clues from their eyes closure. It works in a robust and automatic way, without prior calibration. The presented system is composed of 3 stages. The first one is preprocessing, which includes face and eye detection and normalization. The second stage performs pupil position detection and characterization, combining it with an adaptive lighting filtering to make the system capable of dealing with outdoor illumination conditions. The final stage computes PERCLOS from eyes closure information. In order to evaluate this system, an outdoor database was generated, consisting of several experiments carried out during more than 25 driving hours. A study about the performance of this proposal, showing results from this testbench, is presented.
Iván García 0001, Sebastián Bronte, Luis Miguel Bergasa, Javier Almazán, José Javier Yebes Torres
Intelligent Vehicles Symposium3
2012 Text recognition on traffic panels from street-level imagery
abstract
Text detection and recognition in images taken in uncontrolled environments still remains a challenge in computer vision. This paper presents a method to extract the text depicted in road panels in street view images as an application to Intelligent Transportation Systems (ITS). It applies a text detection algorithm to the whole image together with a panel detection method to strengthen the detection of text in road panels. Word recognition is based on Hidden Markov Models, and a Web Map Service is used to increase the effectiveness of the recognition. In order to compute the distance from the vehicle to the panels, a function that estimates the distance in meters from the text height in pixels has been obtained. After computing the direction vector of the vehicle, world coordinates are computed for each panel. Experimental results on real images from Google Street View prove the efficiency of our proposal and give way to using street-level images for different applications on ITS such as traffic signs inventory or driver assistance.
Álvaro Gonzalez, Luis Miguel Bergasa, José Javier Yebes Torres, Javier Almazán
Intelligent Vehicles Symposium2
2012 Face pose estimation with automatic 3D model creation in challenging scenarios
Pedro Jiménez, Luis Miguel Bergasa, Jesús Nuevo, Pablo Fernández Alcantarilla
Image Vis. Comput.2
2012 Gaze Fixation System for the Evaluation of Driver Distractions Induced by IVIS
abstract
We present a method to monitor driver distraction based on a stereo camera to estimate the face pose and gaze of a driver in real time. A coarse eye direction is composed of face pose estimation to obtain the gaze and driver's fixation area in the scene, which is a parameter that gives much information about the distraction pattern of the driver. The system does not require any subject-specific calibration; it is robust to fast and wide head rotations and works under low-lighting conditions. The system provides some consistent statistics, which help psychologists to assess the driver distraction patterns under influence of different in-vehicle information systems (IVISs). These statistics are objective, as the drivers are not required to report their own distraction states. The proposed gaze fixation system has been tested on a set of challenging driving experiments directed by a team of psychologists in a naturalistic driving simulator. This simulator mimics conditions present in real driving, including weather changes, maneuvering, and distractions due to IVISs. Professional drivers participated in the tests.
Pedro Jiménez, Luis Miguel Bergasa, Jesús Nuevo, Noelia Hernández, Iván García 0001
IEEE Trans. Intell. Transp. Syst.2
2011 Visibility learning in large-scale urban environment
abstract
A crucial step in many vision based applications, such as localization and structure from motion, is the data association between a large map of known 3D points and 2D features perceived by a new camera. In this paper, we propose a novel approach to predict the visibility of known 3D points with respect to a query camera in large-scale environments. In our approach, we model the visibility of each 3D point with respect to a camera pose using a memory-based learning algorithm, in which a distance metric between cameras is learned in an entirely non-parametric way. We show that by fully exploiting the geometric relationships between the 3D map and the camera poses, as well as the related appearance information, the resulting prediction is much more robust and efficient than conventional approaches. We demonstrate the performance of our algorithm on a large urban 3D model in terms of both speed and accuracy.
Pablo Fernández Alcantarilla, Kai Ni 0001, Luis Miguel Bergasa, Frank Dellaert
ICRA3
2011 Occupant Monitoring System for Traffic Control Based on Visual Categorization
abstract
This paper presents the basics of Bag of visual words method, which will be used for an occupant monitoring system that integrates a small onboard camera inside vehicles. It is intended to detect passengers' faces because it is the most appealing characteristic of occupants in a vehicle. This work proposes the implementation of visual categorization by means of two classification methods (Naïve Bayes and Multi-class SVM) that build multi-category image models using the invariant descriptors (SIFT and SURF) extracted from the images under analysis. Bag of visual words approach requires training in order to cluster invariant descriptors and learn the data distribution depending on the classification algorithm. Once the model is created, the category of every test image can be determined by querying a visual dictionary like searching a word in a text dictionary. The performance of the classifiers will be evaluated doing several comparative tests and using standard multi-category image databases. Experimental results and the conclusions are presented.
José Javier Yebes Torres, Pablo Fernández Alcantarilla, Luis Miguel Bergasa
Intelligent Vehicles Symposium3
2011 Face tracking with automatic model construction
Jesús Nuevo, Luis Miguel Bergasa, David Fernández Llorca, Manuel Ocaña
Image Vis. Comput.2
2011 Automatic LightBeam Controller for driver assistance
Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Pedro Jiménez, Ignacio Parra, David Fernández Llorca, Miguel Ángel Sotelo, S. S. Mayoral
Mach. Vis. Appl.2
2011 Automatic Traffic Signs and Panels Inspection System Using Computer Vision
abstract
Computer vision techniques applied to systems used on road maintenance, which are related either to traffic signs or to the road itself, are playing a major role in many countries because of the higher investment on public works of this kind. These systems are able to collect a wide range of information automatically and quickly, with the aim of improving road safety. In this context, the correct visibility of traffic signs and panels is vital for the safety of drivers. This paper describes an approach to the VISUAL Inspection of Signs and panEls (“VISUALISE”), which is an automatic inspection system, mounted onboard a vehicle, which performs inspection tasks at conventional driving speeds. VISUALISE allows for an improvement in the awareness of the road signaling state, supporting planning and decision making on the administration's and infrastructure operators' side. A description of the main computer vision techniques and some experimental results obtained from thousands of kilometers are presented. Finally, the conclusions of the system are described.
Álvaro Gonzalez, Miguel Ángel García Garrido, David Fernández Llorca, Miguel Gavilán, J. Pablo Fernandez, Pablo Fernández Alcantarilla, Ignacio Parra, Fernando Herranz, Luis Miguel Bergasa, Miguel Ángel Sotelo, Pedro A. Revenga
IEEE Trans. Intell. Transp. Syst.9
2010 Visual odometry priors for robust EKF-SLAM
abstract
One of the main drawbacks of standard visual EKF-SLAM techniques is the assumption of a general camera motion model. Usually this motion model has been implemented in the literature as a constant linear and angular velocity model. Because of this, most approaches cannot deal with sudden camera movements, causing them to lose accurate camera pose and leading to a corrupted 3D scene map. In this work we propose increasing the robustness of EKF-SLAM techniques by replacing this general motion model with a visual odometry prior, which provides a real-time relative pose prior by tracking many hundreds of features from frame to frame. We perform fast pose estimation using the two-stage RANSAC-based approach from [1]: a two-point algorithm for rotation followed by a one-point algorithm for translation. Then we integrate the estimated relative pose into the prediction step of the EKF. In the measurement update step, we only incorporate a much smaller number of landmarks into the 3D map to maintain real-time operation. Incorporating the visual odometry prior in the EKF process yields better and more robust localization and mapping results when compared to the constant linear and angular velocity model case. Our experimental results, using a handheld stereo camera as the only sensor, clearly show the benefits of our method against the standard constant velocity model.
Pablo Fernández Alcantarilla, Luis Miguel Bergasa, Frank Dellaert
ICRA2
2010 Learning visibility of landmarks for vision-based localization
abstract
We aim to perform robust and fast vision-based localization using a pre-existing large map of the scene. A key step in localization is associating the features extracted from the image with the map elements at the current location. Although the problem of data association has greatly benefited from recent advances in appearance-based matching methods, less attention has been paid to the effective use of the geometric relations between the 3D map and the camera in the matching process. In this paper we propose to exploit the geometric relationship between the 3D map and the camera pose to determine the visibility of the features. In our approach, we model the visibility of every map feature with respect to the camera pose using a non-parametric distribution model. We learn these non-parametric distributions during the 3D reconstruction process, and develop efficient algorithms to predict the visibility of features during localization. With this approach, the matching process only uses those map features with the highest visibility score, yielding a much faster algorithm and superior localization results. We demonstrate an integrated system based on the proposed idea and highlight its potential benefits for the localization in large and cluttered environments.
Pablo Fernández Alcantarilla, Sang Min Oh, Gian Luca Mariottini, Luis Miguel Bergasa, Frank Dellaert
ICRA4
2010 RSMAT: Robust simultaneous modeling and tracking
Jesús Nuevo, Luis Miguel Bergasa, Pedro Jiménez
Pattern Recognit. Lett.2
2010 Low-cost GPS sensor improvement using stereovision fusion
David Schleicher, Luis Miguel Bergasa, Manuel Ocaña, Rafael Barea, María Elena López Guillén
Signal Process.2
2009 WiFi Localization System based on Fuzzy Logic to Deal with Signal Variations
abstract
The goal of this paper is to study some of the most important WiFi signal variations, large and small scale variations and how they affect to WiFi localization systems. Moreover, the paper shows how to use soft computing techniques to deal with these uncertainties in WiFi localization systems. This work describes how to reduce uncertainty produced by small scale variations in indoor environments using fuzzy techniques. Some experimental results and conclusions are presented.
Noelia Hernández, Fernando Herranz, Manuel Ocaña, Luis Miguel Bergasa, Jose Maria Alonso-Moral, Luis Magdalena
ETFA4
2009 Real-time hierarchical GPS aided visual SLAM on urban environments
abstract
In this paper we present a new real-time hierarchical (topological/metric) Visual SLAM system focusing on the localization of a vehicle in large-scale outdoor urban environments. It is exclusively based on the visual information provided by both a low-cost wide-angle stereo camera and a low-cost GPS. Our approach divides the whole map into local sub-maps identified by the so-called fingerprint (reference poses). At the sub-map level (low level SLAM), 3D sequential mapping of natural landmarks and the vehicle location/orientation are obtained using a top-down Bayesian method to model the dynamic behavior. A higher topological level (high level SLAM) based on references poses has been added to reduce the global accumulated drift, keeping real-time constraints. Using this hierarchical strategy, we keep local consistency of the metric sub-maps, by mean of the EKF, and global consistency by using the topological map and the MultiLevel Relaxation (MLR) algorithm. GPS measurements are integrated at both levels, improving global estimation. Some experimental results for different large-scale urban environments are presented, showing an almost constant processing time.
David Schleicher, Luis Miguel Bergasa, Manuel Ocaña, Rafael Barea, María Elena López Guillén
ICRA2
2009 Real-Time Hierarchical Outdoor SLAM Based on Stereovision and GPS Fusion
abstract
This paper presents a new real-time hierarchical (topological/metric) simultaneous localization and mapping (SLAM) system. It can be applied to the robust localization of a vehicle in large-scale outdoor urban environments, improving the current vehicle navigation systems, most of which are only based on Global Positioning System (GPS). Then, it can be used on autonomous vehicle guidance with recurrent trajectories (bus journeys, theme park internal journeys, etc.). It is exclusively based on the information provided by both a low-cost, wide-angle stereo camera and a low-cost GPS. Our approach divides the whole map into local submaps identified by the so-called fingerprints (vehicle poses). In this submap level (low-level SLAM), a metric approach is carried out. There, a 3-D sequential mapping of visual natural landmarks and the vehicle location/orientation are obtained using a top-down Bayesian method to model the dynamic behavior. GPS measurements are integrated within this low-level improving vehicle positioning. A higher topological level (high-level SLAM) based on fingerprints and the multilevel relaxation (MLR) algorithm has been added to reduce the global error within the map, keeping real-time constraints. This level provides nearly consistent estimation, keeping a small degradation with GPS unavailability. Some experimental results for large-scale outdoor urban environments are presented, showing an almost constant processing time.
David Schleicher, Luis Miguel Bergasa, Manuel Ocaña, Rafael Barea, María Elena López Guillén
IEEE Trans. Intell. Transp. Syst.2
2007 Model-based load localisation for an autonomous hot metal carrier
abstract
Hot metal carriers (HMCs) are large forklift-type vehicles used to move molten metal in aluminium smelters. The molten metal is contained in bucket-like crucibles, that the HMC picks up. In this paper we explore the feasibility of using active appearance models to recognise and localise the handle of a crucible, from a camera on board an autonomous HMC. A two-dimensional model is built that mimics the apparent perspective deformations of the three- dimensional handle. The model is fitted to the handle using efficient algorithms, that include M-estimators for improved robustness. The fitting algorithm also provides and estimate of the actual distance to the crucible. We evaluate the accuracy and robustness of the approach in different lighting conditions.
Jesús Nuevo, Cédric Pradalier, Luis Miguel Bergasa
IROS3
2007 Real-time wide-angle stereo visual SLAM on large environments using SIFT features correction
abstract
This paper presents a new method for real-time SLAM calculation applied to autonomous robot navigation in large environments without restrictions. It is exclusively based on the information provided by a cheap wide-angle stereo camera. Our approach divide the global map into local sub- maps identified by the so-called SIFT fingerprint. At the sub- map level (low level SLAM), 3D sequential mapping of natural land-marks and the robot location/orientation are obtained using a top-down Bayesian method to model the dynamic behavior. A high abstraction level to reduce the global accumulated drift, keeping real-time constraints, has been added (high level SLAM). This uses a SIFT correction method based on the sub-maps' fingerprints. A comparison of the low SLAM level using our method and SIFT features has been carried out. Some experimental results using a real large environment are presented.
David Schleicher, Luis Miguel Bergasa, Rafael Barea, María Elena López Guillén, Manuel Ocaña, Jesús Nuevo
IROS2
2007 Combination of Feature Extraction Methods for SVM Pedestrian Detection
abstract
This paper describes a comprehensive combination of feature extraction methods for vision-based pedestrian detection in Intelligent Transportation Systems. The basic components of pedestrians are first located in the image and then combined with a support-vector-machine-based classifier. This poses the problem of pedestrian detection in real cluttered road images. Candidate pedestrians are located using a subtractive clustering attention mechanism based on stereo vision. A components-based learning approach is proposed in order to better deal with pedestrian variability, illumination conditions, partial occlusions, and rotations. Extensive comparisons have been carried out using different feature extraction methods as a key to image understanding in real traffic conditions. A database containing thousands of pedestrian samples extracted from real traffic images has been created for learning purposes at either daytime or nighttime. The results achieved to date show interesting conclusions that suggest a combination of feature extraction methods as an essential clue for enhanced detection performance
Ignacio Parra, David Fernández Llorca, Miguel Ángel Sotelo, Luis Miguel Bergasa, Pedro A. Revenga, Jesús Nuevo, Manuel Ocaña, Miguel Ángel García Garrido
IEEE Trans. Intell. Transp. Syst.4
2006 Training Method Improvements of a WiFi Navigation System Based on POMDP
abstract
The framework of this paper is the robotics navigation inside buildings using WiFi signal strength measure. This navigation is achieved using a partially observable Markov decision process (POMDP). In the localization phase we used WiFi signal strength and ultrasound measures as observations. The localization system works in two stages: map construction and localization stage. The map construction stage usually requires a great effort, therefore in this paper we address the problem of minimizing this calibration effort using an automatic training method. We describe the method based on simultaneous localization and mapping (SLAM) techniques and in a robust local navigation task. This automatic method is compared with a manual method to obtain a deterministic map. Also we demonstrate that using this one in a on-line training stage the system is able to adapt the WiFi map to the variations of the WiFi signal measure. Additionally, we analyze the optimal parameters for this automatic training system. The system has been tested in a real environment using two commercial robotic platforms. Some experimental results and the conclusions are presented
Manuel Ocaña, Luis Miguel Bergasa, Miguel Ángel Sotelo, Ramón Flores, María Elena López Guillén, Rafael Barea
IROS2
2006 Real-Time Simultaneous Localization and Mapping using a Wide-Angle Stereo Camera and Adaptive Patches
abstract
This paper presents a new method for real-time ego-motion calculation applied to the location/orientation of a cheap wide-angle stereo camera in a 3D environment. The objective is to apply it to a mobile robot navigation system. To achieve that, the goal is to solve the simultaneous localization and mapping (SLAM) problem. Our approach consists in the 3D sequential mapping of natural landmarks by means of a stereo camera, which also provides means to obtain the camera location/orientation. The dynamic behavior is modeled using a top-down Bayesian method. The results show a comparison between our system and a monocular visual SLAM system using a hand-waved camera. Several improvements related to no priori environment knowledge requirements, lower processing time (real-time constrained) and higher robustness is presented
David Schleicher, Luis Miguel Bergasa, Rafael Barea, María Elena López Guillén, Manuel Ocaña
IROS2
2006 Real-time system for monitoring driver vigilance
abstract
This paper presents a nonintrusive prototype computer vision system for monitoring a driver's vigilance in real time. It is based on a hardware system for the real-time acquisition of a driver's images using an active IR illuminator and the software implementation for monitoring some visual behaviors that characterize a driver's level of vigilance. Six parameters are calculated: Percent eye closure (PERCLOS), eye closure duration, blink frequency, nodding frequency, face position, and fixed gaze. These parameters are combined using a fuzzy classifier to infer the level of inattentiveness of the driver. The use of multiple visual parameters and the fusion of these parameters yield a more robust and accurate inattention characterization than by using a single parameter. The system has been tested with different sequences recorded in night and day driving conditions in a motorway and with different users. Some experimental results and conclusions about the performance of the system are presented
Luis Miguel Bergasa, Jesús Nuevo, Miguel Ángel Sotelo, Rafael Barea, María Elena López Guillén
IEEE Trans. Intell. Transp. Syst.1
2005 Pedestrian recognition for intelligent transportation systems
David Fernández Llorca, Ignacio Parra, Miguel Ángel Sotelo, Luis Miguel Bergasa, Pedro A. Revenga, Jesús Nuevo, Manuel Ocaña
ICINCO4
2005 SVM-based Obstacles Recognition for Road Vehicle Applications
Miguel Ángel Sotelo, Jesús Nuevo, David Fernández Llorca, Ignacio Parra, Luis Miguel Bergasa, Manuel Ocaña, Ramón Flores
IJCAI5
2005 Indoor robot navigation using a POMDP based on WiFi and ultrasound observations
abstract
This paper presents a robot navigation system for indoor environments using a partially observable Markov decision process (POMDP) based on WiFi signal strength and ultrasound observations. The paper represents the first one in using WiFi sensor readings as an observation in a POMDP. We present an algorithm based on an EM-SLAM that we called WSLAM (Wifi simultaneous localization and mapping) that is able to learn the observation and transition matrix in autonomous mode. With this algorithm we obtain a minimum calibration effort. We demonstrate that this system is useful to navigate in indoor environments with a real robot. Some experimental results are shown. Finally, the conclusions and future works are presented.
Manuel Ocaña, Luis Miguel Bergasa, Miguel Ángel Sotelo, Ramón Flores
IROS2
2003 A planning architecture for topological robot navigation in uncertain domains
abstract
This paper presents a new navigation architecture for autonomous mobile robots working in uncertain domains. Partially Observable Markov Decision Processes (POMDPs) are suitable mathematical models for solving localization, planning and learning problems in uncertain navigation systems based on a topological representation of the environment. This paper focuses on the planning module, consisting of a two-level layered architecture (a local policy and a global policy) that simplifies the problem of finding optimal policies in POMDPs. The proposed system naturally integrates several planning objectives, such as guiding to a goal room, reducing location uncertainty, and exploring. Some experimental results are shown, carried out with an assistant robot developed in the Electronics Department of the University of Alcala.
María Elena López Guillén, Luis Miguel Bergasa, Rafael Barea, Marisol Escudero
ETFA (1)2
2001 Using a New Model of Recurrent Neural Network for Control
Luciano Boquete, Luis Miguel Bergasa, Rafael Barea, Ricardo García, Manuel Mazo 0001
Neural Process. Lett.2
2000 E.O.G. guidance of a weelchair using spiking neural networks
Rafael Barea, Luciano Boquete, Manuel Mazo 0001, María Elena López Guillén, Luis Miguel Bergasa
ESANN5
2000 Neurocontrol of a binary distillation column
M. A. Torres, M. E. Pardo, J. M. Pupo, Luciano Boquete, Rafael Barea, Luis Miguel Bergasa
ESANN6
2000 E.O.G. Guidance of a Wheelchair Using Neural Networks
abstract
Presents a method to control and guide mobile robots. In this case, to send different commands we have used electrooculography (EOG) techniques, so that, control is made by means of the ocular position (eye displacement into its orbit). A neural network is used to identify the inverse eye model, therefore the saccadic eye movements can be detected and where the user is looking can be determined. This control technique can be useful in multiple applications, but in this work it is used to guide an autonomous robot (wheelchair) as a system to help to people with severe disabilities. The system consists of a standard electric wheelchair with an on-board computer, sensors and graphical user interface running on a computer.
Rafael Barea, Luciano Boquete, Manuel Mazo 0001, María Elena López Guillén, Luis Miguel Bergasa
ICPR5
2000 Commands Generation by Face Movements Applied to the Guidance of a Wheelchair for Handicapped People
abstract
Describes a vision-based commands generation system, by face movements, applied to the guidance of an electric wheelchair for handicapped people with severe disabilities. Using a 2D color face tracker and a fuzzy detector the system computes face movements of the user and, depending on them, some commands are generated to drive the wheelchair. The system is non-intrusive and it allows visibility and freedom of head movements. It is able to learn the face movements of the user in an automatic initial setup, working even for people of different races. It is adaptive and, therefore, robust to light and background changes in inside environments. We report on some experimental results of this kind of guidance and some conclusions about its performance.
Luis Miguel Bergasa, Manuel Mazo 0001, Alfredo Gardel Vicente, Rafael Barea, Luciano Boquete
ICPR1
2000 Industrial inspection using Gaussian functions in a colour space
Luis Miguel Bergasa, Nicola Duffy, Gerard Lacey, Manuel Mazo 0001
Image Vis. Comput.1
2000 Unsupervised and adaptive Gaussian skin-color model
Luis Miguel Bergasa, Manuel Mazo 0001, Alfredo Gardel Vicente, Miguel Ángel Sotelo, Luciano Boquete
Image Vis. Comput.1