EDBT 2026 Demo / reviewers in the wild / expert
Marcos Nieto Doncel
dblp:98/4510 · also Marcos Nieto
· DBLP profile ↗
49ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-9879-0992ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collective Perception Fusion at the Edge: Extending the Edge Dynamic Map Architecture
Mikel García, Marcos Nieto Doncel, Naiara Aginako |
INFOCOM | 2 |
| 2026 | Occlusion-Aware Driver Monitoring using VLM-Enhanced Situational Understanding
Paola Natalia Cañas, Alexander Diez, Marcos Nieto Doncel, Igor Rodríguez, Oumayma Sghairi, Martí Sánchez |
IV | 3 |
| 2026 | Automotive Scenario Mining Using Semantic Graph Databases
Mikel García, Ainara Escribano, Itziar Urbieta, Ioannis Kontopodis, Marcos Nieto Doncel, Naiara Aginako |
IV | 5 |
| 2026 | Exploring visual language models for driver gaze estimation: A task-based approach to debugging AIabstractVisual Language Models (VLMs) have demonstrated superior context understanding and generalization across various tasks compared to models tailored for specific tasks. However, due to their complexity and limited information on their training processes, estimating their performance on specific tasks often requires exhaustive testing, which can be costly and may not account for edge cases. To leverage the zero-shot capabilities of VLMs in safety-critical applications like Driver Monitoring Systems, it is crucial to characterize their knowledge and abilities to ensure consistent performance. This research proposes a methodology to explore and gain a deeper understanding of the functioning of these models in driver’s gaze estimation. It involves detailed task decomposition, identification of necessary data knowledge and abilities (e.g., understanding gaze concepts), and exploration through targeted prompting strategies. Applying this methodology to several VLMs (Idefics2, Qwen2-VL, Moondream, GPT-4o) revealed significant limitations, including sensitivity to prompt phrasing, vocabulary mismatches, reliance on image-relative spatial frames, and difficulties inferring non-visible elements. The findings from this evaluation have highlighted specific areas for improvement and guided the development of more effective prompting and fine-tuning strategies, resulting in enhanced performance comparable with traditional CNN-based approaches. This research is also useful for initial model filtering, for selecting the best model among alternatives and for understanding the model’s limitations and expected behaviors, thereby increasing reliability. Paola Natalia Cañas, Alejandro H. Artiles, Marcos Nieto Doncel, Igor Rodríguez |
Comput. Vis. Image Underst. | 3 |
| 2025 | Design and Implementation of a Data Model for AI Trustworthiness Assessment in CCAM
Ruben Naranjo, Nerea Aranjuelo, Marcos Nieto Doncel, Itziar Urbieta, Javier Fernández 0004, Itsaso Rodríguez-Moreno |
ICAART (3) | 3 |
| 2025 | Interpretable Railway Object Classification Using Part-Prototype Networks
Ruben Naranjo, Iker Sancho, Nerea Aranjuelo, Itsaso Rodríguez-Moreno, Marcos Nieto Doncel |
IJCCI (3) | 5 |
| 2025 | Synthetic Dataset Generation Using Logical Scenario Files for Automotive Perception TestingabstractConducting extensive recording campaigns to asses the safety of newly developed Automated Driving Systems (ADS) or perception algorithms has been proved to be a costly and time consuming process. This is one of the reasons why the automotive industry is adopting the scenario-based testing methodology, to verify and validate the safety of the developed ADS in their expected operating domain. The exterior perception system is the first component in the sense-plan-act process of Connected Cooperative and Automated Vehicles (CCAVs). In this context, high-fidelity simulation engines are used to replicate sensor setups at reduced cost and higher scalability than driving and capturing data from real sensors. The use of logical automotive scenario descriptions allows defining certain parameter ranges, contexts and actions to execute simulations that fulfill the desired conditions. This work proposes a methodology for generating synthetic labelled datasets to test and validate automotive perception systems using logical scenario files. Decoupling the desired sensor setup from the simulation allows reproduction and testing the same situation under different sensor setups and conditions. We implement the methodology to validate three 3D LiDAR-based object detectors in three different sensor setups. The generated sample dataset will be made public here11https://github.com/Vicomtech/Synthetic-OpenLABEL-Scenarios. Mikel García, Aitor Iglesias, Martí Sánchez, Ruben Naranjo, Jon Ander Iñiguez de Gordoa, Marcos Nieto Doncel, Naiara Aginako |
IV | 6 |
| 2025 | DiverSim: A Customizable Simulation Tool to Generate Diverse Vulnerable Road User Datasets
Jon Ander Iñiguez de Gordoa, Martín Hormaetxea, Marcos Nieto Doncel, Gorka Vélez, Andoni Mujika |
VEHITS | 3 |
| 2024 | Real-Time Multi-Camera System for on-Board Train Path EstimationabstractAs the need for more efficient, safe and environmentally friendly transportation increases, the implementation of ADAS systems on trains is becoming essential. In this work, we propose a flexible pipeline for accurately determining the position of the ego-track in railway scenarios. This pipeline is capable of detecting and classifying switches in the track, and use that information to determine which of both branches is part of the ego-track. We evaluate this pipeline on RGB cameras using the OSDaR23 dataset and achieve a Train Centerline Offset Metric (TCOM) of under 15cm until a distance of 75m. We also demonstrate the flexibility of the pipeline by combining the information of wide-angle and zoom camera images, which provides an improvement in TCOM of around 15cm in distances longer than 100m. Maider Larrazabal, Aitor Iglesias, Nerea Aranjuelo, Mikel Labayen, Marcos Nieto Doncel |
IS | 6 |
| 2024 | Annotation Pipeline for Railway Track SegmentationabstractAs ADAS (Advanced Driver Assistance Systems) are becoming more common in the railway environment, the need for data to be able to develop those systems keeps increasing. We are reaching higher Grades of automation (GoA) in trains and subways, but there are not that many publicly available datasets to train the models that help us reach that autonomy; specially when compared to the number of datasets with automotive data. In this paper, we propose a pipeline for annotating rail tracks by using ASAM's OpenLABEL format and its labelling tool WebLabel. With the proposed pipeline we have demonstrated WebLabel's fitness for track annotation and we hope it will help in the release of future datasets for the railway environment. The pipeline helps reduce time in annotation in all the proposed steps. we see a significant change specially by projecting annotations between cameras in sequences that were recorded with more than one camera, where the total time of annotation was reduced by 43.41 %. We also propose a semiautomatic annotation step, where we segment the tracks and load the preannotations into WebLabel before the manual annotation. By doing this we managed to reduce the time spent annotating by 8.25%. Itziar Sagastiberri, Maider Larrazabal, Martí Sánchez, Nerea Aranjuelo, Marcos Nieto Doncel, Daniel Ochoa de Eribe |
IS | 5 |
| 2024 | Efficient class-agnostic obstacle detection for UAV-assisted waterway inspection systemsabstractAbstract Ensuring the safety of water airport runways is essential for the correct operation of seaplane flights. Among other tasks, airport operators must identify and remove various objects that may have drifted into the runway area. In this paper, the authors propose a complete and embedded‐friendly waterway obstacle detection pipeline that runs on a camera‐equipped drone. This system uses a class‐agnostic version of the YOLOv7 detector, which is capable of detecting objects regardless of its class. Additionally, through the usage of the GPS data of the drone and camera parameters, the location of the objects are pinpointed with 0.58 m Distance Root Mean Square. In our own annotated dataset, the system is capable of generating alerts for detected objects with a recall of 0.833 and a precision of 1. Jon Ander Iñiguez de Gordoa, Juan Diego Ortega, Marcos Nieto Doncel |
IET Comput. Vis. | 4 |
| 2024 | WebLabel: OpenLABEL-compliant multi-sensor labellingabstractAbstract Annotated datasets have become crucial for training Machine Learning (ML) models for developing Autonomous Vehicles (AVs) and their functions. Generating these datasets usually involves a complex coordination of automation and manual effort. Moreover, most available labelling tools focus on specific media types (e.g., images or video). Consequently, they cannot perform complex labelling tasks for multi-sensor setups. Recently, ASAM published OpenLABEL, a standard designed to specify an annotation format flexible enough to support the development of automated driving features and to guarantee interoperability among different systems and providers. In this work, we present WebLabel, the first multipurpose web application tool for labelling complex multi-sensor data that is fully compliant with OpenLABEL 1.0. The proposed work analyses several labelling use cases demonstrating the standard's benefits and the application's flexibility to cover various heterogeneous requirements: image labelling, multi-view video object annotation, point-cloud view-based labelling for 3D geometries and action recognition. Itziar Urbieta, Andoni Mujika, Gonzalo Piérola, Eider Irigoyen, Marcos Nieto Doncel, Estíbaliz Loyo, Naiara Aginako |
Multim. Tools Appl. | 5 |
| 2024 | Correction to: WebLabel: OpenLABEL-compliant multi-sensor labelling
Itziar Urbieta, Andoni Mujika, Gonzalo Piérola, Eider Irigoyen, Marcos Nieto Doncel, Estíbaliz Loyo, Naiara Aginako |
Multim. Tools Appl. | 5 |
| 2023 | Multi-camera BEV video-surveillance system for efficient monitoring of social distancingabstractThe current sanitary emergency situation caused by COVID-19 has increased the interest in controlling the flow of people in indoor infrastructures, to ensure compliance with the established security measures. Top view camera-based solutions have proven to be an effective and non-invasive approach to accomplish this task. Nevertheless, current solutions suffer from scalability problems: they cover limited range areas to avoid dealing with occlusions and only work with single camera scenarios. To overcome these problems, we present an efficient and scalable people flow monitoring system that relies on three main pillars: an optimized top view human detection neural network based on YOLO-V4, capable of working with data from cameras at different heights; a multi-camera 3D detection projection and fusion procedure, which uses the camera calibration parameters for an accurate real-world positioning; and a tracking algorithm which jointly processes the 3D detections coming from all the cameras, allowing the traceability of individuals across the entire infrastructure. The conducted experiments show that the proposed system generates robust performance indicators and that it is suitable for real-time applications to control sanitary measures in large infrastructures. Furthermore, the proposed projection approach achieves an average positioning error below 0.2 meters, with an improvement of more than 4 times compared to other methods. David Montero 0002, Nerea Aranjuelo, Peter Leskovský, Estíbaliz Loyo, Marcos Nieto Doncel, Naiara Aginako |
Multim. Tools Appl. | 5 |
| 2022 | Automatic UAV-based airport pavement inspection using mixed real and virtual scenariosabstractRunway and taxiway pavements are exposed to high stress during their projected lifetime, which inevitably leads to a decrease in their condition over time. To make sure airport pavement condition ensure uninterrupted and resilient operations, it is of utmost importance to monitor their condition and conduct regular inspections. UAV-based inspection is recently gaining importance due to its wide range monitoring capabilities and reduced cost. In this work, we propose a vision-based approach to automatically identify pavement distress using images captured by UAVs. The proposed method is based on Deep Learning (DL) to segment defects in the image. The DL architecture leverages the low computational capacities of embedded systems in UAVs by using an optimised implementation of EfficientNet feature extraction and Feature Pyramid Network segmentation. To deal with the lack of annotated data for training we have developed a synthetic dataset generation methodology to extend available distress datasets. We demonstrate that the use of a mixed dataset composed of synthetic and real training images yields better results when testing the training models in real application scenarios. Jon Ander Iñiguez de Gordoa, Juan Diego Ortega, Sara García, Francisco Javier Iriarte, Marcos Nieto Doncel |
ICMV | 6 |
| 2022 | Method for the automatic measurement of camera-calibration quality in a surround-view systemabstractOver the last decade, the automotive industry has introduced advanced driving assistance systems (ADAS) and automated driving (AD) features into roads to reduce fatality rates. One of these ADAS is the surround-view system, which provides an orthographic view of the vehicle by using at least four fish-eye lens cameras embedded in it. Small bumps or temperature changes may modify these cameras' relative poses leading to some geometrical mismatches between views in the top-view projection plane. In addition, terrain irregularities may misalign the orthographic view with the ground plane surface. Both problems can be solved by reestimating the relative poses of the cameras with respect to a single common point in the vehicle. This procedure, also known as recalibration, is offline performed in technical garages, or by online calibration mechanisms on engine start. However, it is a slow and cumbersome process. Research to date studies how to optimally recalibrate these cameras in an online manner, neglecting the practical aspects of when this procedure should be undertaken. Therefore, depending on the functionalities for which the embedded cameras are required, a compromise between using out-of-calibration cameras and the consequences derived from the recalibration process must be considered. This would prevent reestimating the cameras’ relative poses in situations where misalignment between adjacent cameras may not be noticeable. For this reason, a novel approach that measures the degree of calibration between cameras embedded in a vehicle is proposed. This method extracts relevant features from the predefined regions of interest of each camera by using the histogram of oriented gradients (HOG) descriptor. Then, features that belong to adjacent cameras are compared by employing the cosine similarity metric. The proposed method is evaluated on the open-source AD research simulator CARLA providing detailed analysis to objectively highlight the usefulness of this method in studying the degree of calibration of a camera array in a surround-view system. Martí Sánchez, Jon Ander Iñiguez de Gordoa, Marcos Nieto Doncel, Pablo Carballeira |
ICMV | 3 |
| 2022 | Exploiting AirSim as a Cross-dataset Benchmark for Safe UAV Landing and Monocular Depth Estimation Models
Jon Ander Iñiguez de Gordoa, Javier Barandiarán, Marcos Nieto Doncel |
IJCCI | 3 |
| 2022 | Efficient large-scale face clustering using an online Mixture of Gaussians
David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Automated Annotation of Lane Markings Using LIDAR and OdometryabstractLane markings are mymargin a key element for Autonomous Driving. The generation of high definition maps and ground-truth data require extensive manual labor. In this paper, we present an efficient and robust method for the offline annotation of lane markings, using low-density LIDAR point clouds and odometry information. The odometry is used to accumulate the scans and to process them using blocks following the trajectory of the vehicle. At each block, candidate lane marking points are detected by generating virtual scan-lines and applying a dynamically optimized filter function to the LIDAR intensity values. The lane markings are tracked block wise, and their width is estimated and classified as either solid or dashed. The results are lists of connected 3D points that represent the different lane markings. The accuracy of the proposed method was tested against manually labeled recordings. A novel evaluation methodology focused on the lateral precision of detections is presented. Moreover, a web user interface was used to load the produced annotations, achieving a reduction of 60% in the annotation time, as compared to a fully manual baseline. Javier Barandiarán, Marcos Nieto Doncel, Andoni Cortés Vidal, Oihana Otaegui Madurga, Julián Flórez 0001, Manuel Graña |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Analysis of Classifier Training on Synthetic Data for Cross-Domain DatasetsabstractA major challenges of deep learning (DL) is the necessity to collect huge amounts of training data. Often, the lack of a sufficiently large dataset discourages the use of DL in certain applications. Typically, acquiring the required amounts of data costs considerable time, material and effort. To mitigate this problem, the use of synthetic images combined with real data is a popular approach, widely adopted in the scientific community to effectively train various detectors. In this study, we examined the potential of synthetic data-based training in the field of intelligent transportation systems. Our focus is on camera-based traffic sign recognition applications for advanced driver assistance systems and autonomous driving. The proposed augmentation pipeline of synthetic datasets includes novel augmentation processes such as structured shadows and gaussian specular highlights. A well-known DL model was trained with different datasets to compare the performance of synthetic and real image-based trained models. Additionally, a new, detailed method to objectively compare these models is proposed. Synthetic images are generated using a semi-supervised errors-guide method which is also described. Our experiments showed that a synthetic image-based approach outperforms in most cases real image-based training when applied to cross-domain test datasets (+10% precision for GTSRB dataset) and consequently, the generalization of the model is improved decreasing the cost of acquiring images. Andoni Cortés, Clemente Rodríguez Lafuente, Gorka Vélez, Javier Barandiarán, Marcos Nieto Doncel |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | MFR 2021: Masked Face Recognition CompetitionabstractThis paper presents a summary of the Masked Face Recognition Competitions (MFR) held within the 2021 International Joint Conference on Biometrics (IJCB 2021). The competition attracted a total of 10 participating teams with valid submissions. The affiliations of these teams are diverse and associated with academia and industry in nine different countries. These teams successfully submitted 18 valid solutions. The competition is designed to motivate solutions aiming at enhancing the face recognition accuracy of masked faces. Moreover, the competition considered the deployability of the proposed solutions by taking the compactness of the face recognition models into account. A private dataset representing a collaborative, multisession, real masked, capture scenario is used to evaluate the submitted solutions. In comparison to one of the topperforming academic face recognition solutions, 10 out of the 18 submitted solutions did score higher masked face verification accuracy. Fadi Boutros, Naser Damer, Jan Niklas Kolf, Kiran B. Raja, Florian Kirchbuchner, Ramachandra Raghavendra, Arjan Kuijper, Pengcheng Fang, Fei Wang 0032, David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Asaki Kataoka, Kohei Ichikawa, Shizuma Kubo, Jie Zhang 0071, Shiguang Shan, Klemen Grm, Vitomir Struc, Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Pedro C. Neto, Ana Filipa Sequeira, João Ribeiro Pinto, Mohsen Saffari, Jaime S. Cardoso 0001 |
IJCB | 14 |
| 2021 | Accurate 3D Object Detection from Point Cloud Data using Bird's Eye View Representations
Nerea Aranjuelo, Guus Engels, David Montero 0002, Marcos Nieto Doncel, Ignacio Arganda-Carreras, Luis Unzueta, Oihana Otaegui Madurga |
IJCCI | 4 |
| 2020 | 3D Object Detection from LiDAR Data using Distance Dependent Feature ExtractionabstractThis paper presents a new approach to 3D object detection that leverages the properties of the data obtained by a LiDAR sensor. State-of-the-art detectors use neural network architectures based on assumptions valid for camera images. However, point clouds obtained from LiDAR are fundamentally different. Most detectors use shared filter kernels to extract features which do not take into account the range dependent nature of the point cloud features. To show this, different detectors are trained on two splits of the KITTI dataset: close range (objects up to 25 meters from LiDAR) and long-range. Top view images are generated from point clouds as input for the networks. Combined results outperform the baseline network trained on the full dataset with a single backbone. Additional research compares the effect of using different input features when converting the point cloud to image. The results indicate that the network focuses on the shape and structure of the objects, rather than exact values of the input. This work proposes an improvement for 3D object detectors by taking into account the properties of LiDAR point clouds over distance. Results show that training separate networks for close-range and long-range objects boosts performance for all KITTI benchmark difficulties. Guus Engels, Nerea Aranjuelo, Ignacio Arganda-Carreras, Marcos Nieto Doncel, Oihana Otaegui Madurga |
VEHITS | 4 |
| 2020 | User-adaptive Eyelid Aperture Estimation for Blink Detection in Driver Monitoring Systems
Juan Diego Ortega, Marcos Nieto Doncel, Luis Salgado, Oihana Otaegui Madurga |
VEHITS | 2 |
| 2019 | A Cloud-Based AI Framework for Machine Learning Orchestration: A "Driving or Not-Driving" Case-Study for Self-Driving CarsabstractSelf-driving cars rely on a plethora of algorithms in order to perform safe driving manoeuvres. Training those models is expensive (e.g. hardware cost, storage, energy) and requires continuous updates. This paper proposes a cloud-based framework for continuous training of self-driving AI models. In addition to training standalone models, the framework is capable of leveraging pre-trained models in expediting the training on environment changes (e.g. new driver or new car model). As use-case, this paper focuses on a driver's behaviour while the vehicle's control is being transferred between the driver and the self-driving AI. A human driver can hand over the control of a vehicle's driving tasks to an automated system, when that system's confidence level is high enough. Reciprocally, there are situations where that control has to be handed back to the human driver. This paper proposes a novel real-time system for Driving Not-Driving (DND) detection, which is able to capture the ability of the driver to re-take control of a vehicle when the automated driving system transitions from a higher to a lower level of automation (e.g. L3 to L2 vehicle automation). We are using a computer vision-based Driver Monitoring System (DMS) that captures in real-time head and eye movements. These are captured in the car and transferred to the cloud where a DND model is trained for a specific driver. The DND classification model is deployed in the vehicle and predicts if the driver is ready or not to resume control at a given time. The cloud-based framework proposed in this paper shows an end-to-end cycle of collecting, training and deploying self-driving AI technology, with the additional features of continuous and transfer learning. Cristian Olariu, Haytham Assem, Juan Diego Ortega, Marcos Nieto Doncel |
IV | 4 |
| 2018 | Constant-time monocular object detection using scene geometry
Marcos Nieto Doncel, Juan Diego Ortega, Peter Leskovský, Orti Senderos |
Pattern Anal. Appl. | 1 |
| 2016 | Guest EditorialabstractAs digital technologies advance, video has become ubiquitous and hence a rich source of information. Video analytics (or video content analysis) is an important area of computer vision which is concerned with the process of making sense of video content in order to ultimately understand video. Video analytics appears in different forms, such as activity recognition, motion detection, object detection and recognition, person detection and recognition, event and scenario recognition, anomaly detection and identity recognition and verification. Video analytics can be applied in a wide range of domains including healthcare, retail, transport, smart homes, safety and security. The aim of this Special Issue is to raise the awareness of the importance of video analytics. The specific objectives are: (1) to report the latest developments; (2) to identify major research challenges and; (3) to provide visions of future development. A total of 21 papers were submitted to this Special Issue including invited papers and, following a rigorous peer-review process, a total of 11 papers were accepted for inclusion in this Special Issue. The invited paper “Video Analytics Revisited” by Ayesha Choudhary and Santanu Chaudhury presents a concise yet in-depth survey of video analytics. Research problems are discussed and important current applications of video analytics are reviewed. The paper “Human Action Recognition Using Histogram of Motion Intensity and Direction from Multiple Views” by SungYong Chun et al. presents an approach to human activity recognition from multiple views based on estimation of local motion from multiple camera views. A new motion descriptor, histogram of motion intensity and direction, is proposed to capture local motion characteristics of human activity. Classification is done using a support vector machine. Experimental evaluation has demonstrated superior performance of their approach, outperforming 3D optical flow-based approaches with lower computational requirements. The paper “Video Anomaly Detection Using Deep Incremental Slow Feature Analysis Network” by Xing Hu et al. presents an approach to anomaly detection using automatically learned features instead of hand-crafted features. A Deep Incremental Slow Feature Analysis (D-IncSFA) network is proposed, which learns progressively abstract and global high-level features from raw data. The D-IncSFA network has the functionalities of both feature extractor and anomaly detector so anomaly detection can be completed in one step. Their approach can detect global anomaly such as crowd panic and local anomaly and is intended to be universal in order to work in different scenarios, with little human intervention and low memory and computational requirements. The paper “Multiple Deep Features Learning for Object Retrieval in Surveillance Videos” by Haiyun Guo et al. aims to address the challenge of efficiently indexing and retrieving objects of interest from large-scale surveillance videos. A multiple deep features learning approach to object retrieval in surveillance videos is proposed, which is based on the discriminative convolutional neural network (CNN). The CNN model is pre-trained on ImageNet ILSVRC12 and then fine-tuned on their dataset. To improve the retrieval performance, the deep features are encoded into short binary codes by Locality-Sensitive Hash and fused to retrieve the object of interest. Experiments on a dataset of 100k objects extracted from multi-camera surveillance videos have demonstrated good performance of the proposed approach, compared with other common visual features. The paper “A Two-layer Discriminative Model for Human Activity Recognition” by Mouna Selmi et al. studies human activity recognition with a focus on the role of local interest point features like spatio-temporal interest points. This paper presents a new approach that explicitly models the sequential aspect of activities. A support vector machine provides a vector of conditional class probabilities for each window that summarises all discriminant information that is relevant for sequence recognition. The sequence of these stochastic vectors is then fed to a hidden conditional random field for inference at the sequence level. Experiments on various human activity datasets have demonstrated that the proposed approach compares favourably with current state-of-the-art. The paper “A New Fusional Framework Combining Sparse Selection and Clustering for Key Frame Extraction” by Mengjuan Fei et al. studies key frame extraction, a type of video summarisation, which facilitates rapid browsing and efficient video indexing. This paper proposes a syncretic key frame extraction framework (SS-MIAHC) that combines sparse selection and mutual information-based agglomerative hierarchical clustering (MIAHC) to generate effective video summaries. The proposed framework overcomes issues such as information redundancy and computational complexity. The experiments conducted on two benchmark datasets demonstrate that the proposed SS-MIAHC framework is superior to conventional methods. The paper “Multi-Object Tracking using Dominant Sets” by Yonatan T. Tesfaye et al. studies multi-object tracking and addresses the challenges of identity switches and difficulties in handling long-term occlusions by formulating the tracking task as a problem of finding dominant sets in an auxiliary edge weighted graph. This is a novel approach to multi-object tracking, which has been demonstrated to have superior performance compared with several state of the art methods in experiments on three different challenging datasets. The paper “Contextualized Learning-free 3D Body Pose Estimation from 2D Body Features in Monocular Images” by Luis Unzueta et al. presents a method for 3D human body pose estimation from a monocular camera based on a learning-free hierarchical optimisation procedure and contextual information. This approach explicitly considers and preserves the relations between the 3D subject's overall scale; its depth with respect to the camera; and its configuration related to the reference floor. Thus, it can obtain more coherent reconstructions with respect to the shared 3D world, compared to other state-of-the-art approaches, efficiently and without the need for learning 2D/3D mapping models from training data. Therefore, it is not affected by data characteristic differences between training and deployment stages. The paper “‘Owl’ and ‘Lizard’: Patterns of Head Pose and Eye Pose in Driver Gaze Classification” by Lex Fridman et al. studies gaze tracking in the car through estimating head pose and eye pose from monocular video. New research questions are asked, which are answered by evaluating data drawn from an on-road study of 40 drivers. The main insight of the paper is conveyed through the analogy of an “owl” and “lizard” which describes the degree to which the eyes and the head move when shifting gaze. When the head moves a lot (“owl”), not much classification improvement is attained by estimating eye pose on top of head pose. On the other hand, when the head stays still and only the eyes move (“lizard”), classification accuracy increases significantly from adding in eye pose. The paper “Forensic Video Solution Using Facial Feature Based Synoptic Video Footage Record” by B.Yogameena et al. proposes a solution to identify a specific person quickly which is valuable in analysing incidents/crimes. The main idea of this paper is to reduce the enormous volume of video data by using an object based video synopsis. SVM is used to classify the weak and strong features. These strong features are used to recognise the person. The algorithm works well even in complicated situations such as expression changes, pose, illumination variations and even if the face is partially or fully occluded in few frames. The advantage of synoptic video helps to recognize the person who is not occluded in some other frames. Experimental results on benchmark and real time datasets demonstrate the effectiveness of the proposed algorithm. The paper “Facial Video based Detection of Physical Fatigue for Maximal Muscle Activity” by Mohammad A. Haque et al. studies video based detection of physical fatigue. This paper presents an efficient noncontact system for detecting non-localised physical fatigue from maximal muscle activity using facial videos acquired in a realistic environment with natural lighting where subjects were allowed to voluntarily move their head, change their facial expression and vary their pose. Experimental results show that the proposed system outperforms video based existing system for physical fatigue detection. Hui Wang 0001, Marcos Nieto Doncel, Zhen Lei 0001, Suzanne Little |
IET Comput. Vis. | 2 |
| 2014 | Perspective Multiscale Detection and Tracking of Persons
Marcos Nieto Doncel, Juan Diego Ortega, Andoni Cortés, Seán Gaines |
MMM (2) | 1 |
| 2013 | Perspective Multiscale Detection of Vehicles for Real-Time Forward Collision Avoidance Systems
Juan Diego Ortega, Marcos Nieto Doncel, Andoni Cortés, Julián Flórez 0001 |
ACIVS | 2 |
| 2013 | The Impact of Video Transcoding Parameters on Event Detection for Surveillance SystemsabstractThe process of transcoding videos apart from being computationally intensive, can also be a rather complex procedure. The complexity refers to the choice of appropriate parameters for the transcoding engine, with the aim of decreasing video sizes, transcoding times and network bandwidth without degrading video quality beyond some threshold that event detectors lose their accuracy. This paper explains the need for transcoding, and then studies different video quality metrics. Commonly used algorithms for motion and person detection are briefly described, with emphasis in investigating the optimum transcoding configuration parameters. The analysis of the experimental results reveals that the existing video quality metrics are not suitable for automated systems and that the detection of persons is affected by the reduction of bit rate and resolution, while motion detection is more sensitive to frame rate. Emmanouil Kafetzakis, Christos Xilouris, Michail-Alexandros Kourtis, Marcos Nieto Doncel, Iveel Jargalsaikhan, Suzanne Little |
ISM | 4 |
| 2013 | Interactive surveillance event detection at TRECVid2012abstractThis demonstration shows the integration of video analysis and search tools to facilitate the interactive retrieval of video segments depicting specific activities from surveillance footage. The implementation was developed by members of the SAVASA project for participation in the interactive surveillance event detection (SED) task of TRECVid 2012. This year, for the first time, the purpose of the interactive SED task was to evaluate systems' ability to support users in identifying video segments that depict a specific activity (event) in a large collection of surveillance video footage. Project partners worked together to analyse video and provide a query interface enabling users to search and identify matching video segments. The collaborative integration of components from multiple partners and the participation of end user partners in evaluating the system are the novel aspects of this work. Suzanne Little, Iveel Jargalsaikhan, Kathy M. Clawson, Marcos Nieto Doncel, Cem Direkoglu, Noel E. O'Connor, Alan F. Smeaton, Jun Liu 0001, Bryan W. Scotney, Hui Wang 0001, Seán Gaines, Aitor Rodriguez, Pedro J. Sánchez, Ana Martínez Llorens, Karina Villarroel Paniza, Roberto Gimenez, Raúl Santos de la Cámara, Anna Mereu, Celso Prados, Emmanouil Kafetzakis |
ICMR | 4 |
| 2013 | An information retrieval approach to identifying infrequent events in surveillance videoabstractThis paper presents work on integrating multiple computer vision-based approaches to surveillance video analysis to support user retrieval of video segments showing human activities. Applied computer vision using real-world surveillance video data is an extremely challenging research problem, independently of any information retrieval (IR) issues. Here we describe the issues faced in developing both generic and specific analysis tools and how they were integrated for use in the new TRECVid interactive surveillance event detection task. We present an interaction paradigm and discuss the outcomes from face-to-face end user trials and the resulting feedback on the system from both professionals, who manage surveillance video, and computer vision or machine learning experts. We propose an information retrieval approach to finding events in surveillance video rather than solely relying on traditional annotation using specifically trained classifiers. Suzanne Little, Iveel Jargalsaikhan, Kathy M. Clawson, Marcos Nieto Doncel, Cem Direkoglu, Noel E. O'Connor, Alan F. Smeaton, Bryan W. Scotney, Hui Wang 0001, Jun Liu 0001 |
ICMR | 4 |
| 2013 | An Empirical Evaluation of Interest Point DetectorsabstractImage interest point extraction and matching across images is a commonplace task in computer vision–based applications, across widely diverse domains, such as 3D reconstruction, augmented reality, or tracking. We present an empirical evaluation of state-of-the-art interest point detection algorithms measuring several parameters, such as efficiency, robustness to image domain geometric transformations—that is, similarity—affine or projective transformations, as well as invariance to photometric transformations such as light intensity or image noise. Iñigo Barandiaran, Manuel Graña, Marcos Nieto Doncel |
Cybern. Syst. | 3 |
| 2012 | Adaptive Multicue Background Subtraction for Robust Vehicle Counting and ClassificationabstractIn this paper, we present a robust vision-based system for vehicle tracking and classification devised for traffic flow surveillance. The system performs in real time, achieving good results, even in challenging situations, such as with moving casted shadows on sunny days, headlight reflections on the road, rainy days, and traffic jams, using only a single standard camera. We propose a robust adaptive multicue segmentation strategy that detects foreground pixels corresponding to moving and stopped vehicles, even with noisy images due to compression. First, the approach adaptively thresholds a combination of luminance and chromaticity disparity maps between the learned background and the current frame. It then adds extra features derived from gradient differences to improve the segmentation of dark vehicles with casted shadows and removes headlight reflections on the road. The segmentation is further used by a two-step tracking approach, which combines the simplicity of a linear 2-D Kalman filter and the complexity of a 3-D volume estimation using Markov chain Monte Carlo (MCMC) methods. Experimental results show that our method can count and classify vehicles in real time with a high level of performance under different environmental situations comparable with those of inductive loop detectors. Luis Unzueta, Marcos Nieto Doncel, Andoni Cortés, Javier Barandiarán, Oihana Otaegui Madurga, Pedro J. Sánchez |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2011 | Road environment modeling using robust perspective analysis and recursive Bayesian segmentation
Marcos Nieto Doncel, Jon Arróspide Laborda, Luis Salgado |
Mach. Vis. Appl. | 1 |
| 2011 | Line segment detection using weighted mean shift procedures on a 2D slice sampling strategy
Marcos Nieto Doncel, Carlos Cuevas, Luis Salgado, Narciso García |
Pattern Anal. Appl. | 1 |
| 2011 | Simultaneous estimation of vanishing points and their converging lines using the EM algorithm
Marcos Nieto Doncel, Luis Salgado |
Pattern Recognit. Lett. | 1 |
| 2010 | Multiple object tracking using an automatic variable-dimension particle filterabstractObject tracking through particle filtering has been widely addressed in recent years. However, most works assume a constant number of objects or utilize an external detector that monitors the entry or exit of objects in the scene. In this work, a novel tracking method based on particle filtering that is able to automatically track a variable number of objects is presented. As opposed to classical prior data assignment approaches, adaptation of tracks to the measurements is managed globally. Additionally, the designed particle filter is able to generate hypotheses on the presence of new objects in the scene, and to confirm or dismiss them by gradually adapting to the global observation. The method is especially suited for environments where traditional object detectors render noisy measurements and frequent artifacts, such as that given by a camera mounted on a vehicle, where it is proven to yield excellent results. Jon Arróspide Laborda, Luis Salgado, Marcos Nieto Doncel |
ICIP | 3 |
| 2010 | Non-linear optimization for robust estimation of vanishing pointsabstractA new method for robust estimation of vanishing points is introduced in this paper. It is based on the MSAC (M-estimator Sample and Consensus) algorithm and on the definition of a new distance function between a vanishing point and a given orientation. Apart from the robustness, our method represents a flexible and efficient solution, since it allows to work with different type of image data, and its iterative nature makes better use of the available information to obtain more accurate estimates. The key issue of the work is the proposed distance function, that makes the error to be independent from the position of an hypothesized vanishing point, which allows to work with points at the infinity. Besides, the estimation process is guided by a non-linear optimization process that enhances the accuracy of the system. The robustness of our proposal, compared with other methods in the literature is shown with a set of tests carried out for both synthetic data and real images. The results show that our approach obtain excellent levels of accuracy and that is definitely robust against the presence of large amounts of outliers, outperforming other state of the art approaches. Marcos Nieto Doncel, Luis Salgado |
ICIP | 1 |
| 2010 | Plane rectification through robust vanishing point tracking using the Expectation-Maximization algorithmabstractThis paper introduces a new strategy for plane rectification in sequences of images, based on the Expectation-Maximization (EM) algorithm. Our approach is able to compute simultaneously the parameters of the dominant vanishing point in the image plane and the most significant lines passing through it. It is based on a novel definition of the likelihood distribution of the gradient image considering both the position and the orientation of the gradient pixels. Besides, the mixture model in which the EM algorithm operates is extended, compared to other works, to consider an additional component to control the presence of outliers. Some synthetic data tests are described to show the robustness and efficiency of the proposed method. The plane rectification results show that the method is able to remove the perspective and affine distortion of real traffic sequences without the need to compute two vanishing points. Marcos Nieto Doncel, Luis Salgado |
ICIP | 1 |
| 2010 | Vehicle detection and tracking using homography-based plane rectification and particle filteringabstractThis paper presents a full system for vehicle detection and tracking in non-stationary settings based on computer vision. The method proposed for vehicle detection exploits the geometrical relations between the elements in the scene so that moving objects (i.e., vehicles) can be detected by analyzing motion parallax. Namely, the homography of the road plane between successive images is computed. Most remarkably, a novel probabilistic framework based on Kalman filtering is presented for reliable and accurate homography estimation. The estimated homography is used for image alignment, which in turn allows to detect the moving vehicles in the image. Tracking of vehicles is performed on the basis of a multidimensional particle filter, which also manages the exit and entries of objects. The filter involves a mixture likelihood model that allows a better adaptation of the particles to the observed measurements. The system is specially designed for highway environments, where it has been proven to yield excellent results. Jon Arróspide Laborda, Luis Salgado, Marcos Nieto Doncel |
Intelligent Vehicles Symposium | 3 |
| 2009 | Measurement-based reclustering for multiple object tracking with particle filtersabstractMultiple object tracking is a main research area in the computer vision field. Particle filters have shown their performance as a powerful tool allowing to track visual objects giving temporal coherence to incoming observations, as well as offering an excellent framework for this task due to its inherent multimodality. However, traditional algorithms for particle filters do not cope directly with multiple objects and several considerations have to be addressed. In this work, an efficient reclustering strategy is proposed, which takes into account new measurements according to a novelty function, and provides a criterium to determine the minimum required number of particles to be drawn for each tracked object. To show its performance, this strategy has been used as a multiple 2D object tracking for video-surveillance applications. Excellent results are obtained, in terms of efficiency and accuracy. Marcos Nieto Doncel, Carlos Cuevas, Luis Salgado |
ICIP | 1 |
| 2008 | On-board robust vehicle detection and tracking using adaptive quality evaluationabstractThis paper presents a robust method for real-time vehicle detection and tracking in dynamic traffic environments. The proposed strategy aims to find a trade-off between the robustness shown by time-uncorrelated detection techniques and the speed-up obtained with tracking algorithms. It combines both advantages by continuously evaluating the quality of the tracking results along time and triggering new detections to restart the tracking process when quality falls behind a certain quality requirement. Robustness is also ensured within the tracking algorithm with an outlier rejection stage and the use of stochastic filtering. Several sequences from real traffic situations have been tested, obtaining highly accurate multiple vehicle detections. Jon Arróspide Laborda, Luis Salgado, Marcos Nieto Doncel, Fernando Jaureguizar |
ICIP | 3 |
| 2008 | Robust multiple lane road modeling based on perspective analysisabstractRoad modeling is the first step towards environment perception within driver assistance video-based systems. Typically, lane modeling allows applications such as lane departure warning or lane invasion by other vehicles. In this paper, a new monocular image processing strategy that achieves a robust multiple lane model is proposed. The identification of multiple lanes is done by firstly detecting the own lane and estimating its geometry under perspective distortion. The perspective analysis and curve fitting allows to hypothesize adjacent lanes assuming some a priori knowledge about the road. The verification of these hypotheses is carried out by a confidence level analysis. Several types of sequences have been tested, with different illumination conditions, presence of shadows and significant curvature, all performing in realtime. Results show the robustness of the system, delivering accurate multiple lane road models in most situations. Marcos Nieto Doncel, Luis Salgado, Fernando Jaureguizar, Jon Arróspide Laborda |
ICIP | 1 |
| 2007 | Real-Time Vanishing Point Estimation in Road Sequences Using Adaptive Steerable Filter Banks
Marcos Nieto Doncel, Luis Salgado |
ACIVS | 1 |
| 2007 | Robust Vehicle Detection Through Multidimensional Classification for on Board Video Based SystemsabstractThis paper presents a new in-vehicle real-time vehicle detection strategy which hypothesizes the presence of vehicles in rectangular sub-regions based on the robust classification of features vectors result of a combination of multiple morphological vehicle features. One vector is extracted for each region of the image likely containing vehicles as a multidimensional likelihood measure with respect to a simplified vehicle model. A supervised training phase set the representative vectors of the classes vehicle and non-vehicle, so that the hypothesis is verified or not according to the Mahalanobis distance between the feature vector and the representative vectors. Excellent results have been obtained in several video sequences accurately detecting vehicles with very different aspect-ratio, color, size, etc, while minimizing the number of missing detections and false alarms. Daniel Alonso, Luis Salgado, Marcos Nieto Doncel |
ICIP (4) | 3 |
| 2006 | Fast Mode Decision on H.264/AVC Main Profile Encoding Based on PSNR PredictionsabstractIn this paper we propose a new fast mode decision (FMD) algorithm to reduce the computational load of the motion estimation (ME) process of the new video coding standard H.264/AVC main profile. The algorithm firstly computes the skip/direct mode (respectively for P and B-frames), for all macroblocks (MBs), and then decides, for each MB if no mores modes are needed. This decision is based on the generation of predictions of the expected PSNR of the frame to be encoded. These predictions are performed with the distortion cost obtained for skip/direct mode jointly with the average distortion cost obtained for the previous encoded frames. Results show that the computational load have been dramatically reduced, with average encoding time saving of-36.38%, while nearly negligible loss of coding efficiency. Marcos Nieto Doncel, Luis Salgado, Julián Cabrera |
ICIP | 1 |
| 2006 | Sequence Independent very Fast Mode Decision Algorithm on H.264/AVC Baseline ProfileabstractIn this paper we propose a new fast mode decision (FMD) algorithm for H.264/AVC to reduce the computational load of the motion estimation (ME) process. It is oriented to dramatically reduce the encoding time regardless of the level of motion present at the sequences (high or low-motion). The algorithm decides as the best mode the skip mode or mode 1 without need to compute the rest of coding alternatives if some conditions are satisfied. These conditions ensure these modes to be used appropriately for low motion and high motion sequences respectively to achieve an appropriate rate distortion (RD) cost based on the results of the previous encoded frames. Tests have shown reductions of the encoding time around-81% for all kinds of sequences, while moderate loss of coding efficiency. Luis Salgado, Marcos Nieto Doncel |
ICIP | 2 |
| 2005 | Fast Mode Decision and Motion Estimation with Object Segmentation in H.264/AVC Encoding
Marcos Nieto Doncel, Luis Salgado, Julián Cabrera |
ACIVS | 1 |