José María Martínez Sanchez

dblp:47/3503 · also José M. Martínez 0001 · DBLP profile ↗
← Back
70ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-2236-1769ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2026 Soft-Labelling for Budget-Constrained Semantic Segmentation: Bringing Coherence to label Down-Sampling
abstract
In semantic segmentation, training data down-sampling is commonly performed due to resource limitations, the need to adapt image size to the model input, or to improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled colour and ground-truth label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labelling that better conserves label information after down-sampling, thereby, fully aligning soft-labels with image data to keep the distribution of the sampled pixels for down-sampling. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that our proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly fewer computational resources than foremost methods. This proposal enables competitive research for semantic segmentation under resource constraints.
Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Media4
2025 An image-processing toolkit for remote photoplethysmography
abstract
Abstract Objective. Image-processing-based remote photoplethysmography algorithms are usually composed of steps where different methods are used, and often, researchers perform these steps using methods that are not necessarily the best for their application. With our toolkit, we want to provide easy and fast access to different state-of-the-art methods for the most common image-processing steps in remote photoplethysmography algorithms. Methods. Our toolkit was programmed in Python and was developed with sequential workflow in mind, making it versatile and easy to use in interactive environments. It also includes tools so the users can modify or extend it if they want to, and will be updated as new methods for the different steps are published. Results. Our use case examples and validation show an effective approach and how the toolkit can be used for exhaustive evaluation and ablation studies in a simple way. We also show how choosing different methods can affect the final heart rate estimation accuracy at the cost of computation time. Conclusion. With this toolkit we are providing researchers with a versatile, easy-to-use tool, with access to different methods for the most common steps in remote photoplethysmography algorithms. Significance. Our toolkit is a relevant tool for researchers in the remote photoplethysmography field due to their versatility, ease of use, and adaptability. (It will be available on https://github.com/Montyro/rppgtk github upon acceptance).
Javier Montalvo, Álvaro García-Martín, José María Martínez Sanchez
Multim. Tools Appl.3
2024 Long-Term Geo-Positioned Re-Identification Dataset of Urban Elements
abstract
This paper introduces UrbAM-ReID, a new long-term geo-positioned urban ReID dataset. It is composed by four subdatasets recording the same trajectory at the UAM Campus, each one recorded in different seasons and including an inverse direction recording. While most of the current datasets in the state-of-the-art focus on person re-identification, with vehicles as the second most explored object, our work specifically addresses urban objects re-identification, currently, waste containers, rubbish bins, and crosswalks. The dataset provides different attributes of the annotated objects, like their classes, their foreground or background status and the geo-position. Several evaluation configurations can be defined to simulate realistic scenarios that may arise in actual situations within the management of urban elements, considering the utilization of just visual data, or incorporating additional attributes, providing different complexity levels. Finally, the dataset is used for defining a benchmark where two state-of-the-art systems are evaluated. The dataset and supplementary material is available in https://github.com/vpulab/UrbAMReID
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez
ICIP3
2023 Vehicle Re-Identification Based on Unsupervised Domain Adaptation by Incremental Generation of Pseudo-Labels
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez
CIARP3
2023 Enhancing vehicle re-identification via synthetic training datasets and re-ranking based on video-clips information
abstract
Abstract Vehicle re-identification (ReID) aims to find a specific vehicle identity across multiple non-overlapping cameras. The main challenge of this task is the large intra-class and small inter-class variability of vehicles appearance, sometimes related with large viewpoint variations, illumination changes or different camera resolutions. To tackle these problems, we proposed a vehicle ReID system based on ensembling deep learning features and adding different post-processing techniques. In this paper, we improve that proposal by: incorporating large-scale synthetic datasets in the training step; performing an exhaustive ablation study showing and analyzing the influence of synthetic content in ReID datasets, in particular CityFlow-ReID and VeRi-776; and extending post-processing by including different approaches to the use of gallery video-clips of the target vehicles in the re-ranking step. Additionally, we present an evaluation framework in order to evaluate CityFlow-ReID: as this dataset has not public ground truth annotations, AI City Challenge provided an on-line evaluation service which is no more available; our evaluation framework allows researchers to keep on evaluating the performance of their systems in the CityFlow-ReID dataset.
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós
Multim. Tools Appl.3
2023 Graph Neural Networks for Cross-Camera Data Association
abstract
Cross-camera image data association is essential for many multi-camera computer vision tasks, such as multi-camera pedestrian detection, multi-camera multi-target tracking, 3D pose estimation, etc. This association task is typically modeled as a bipartite graph matching problem and often solved by applying minimum-cost flow techniques, which may be computationally demanding for large data. Furthermore, cameras are usually treated by pairs, obtaining local solutions, rather than finding a global solution at once for all multiple cameras. Other key issue is that of the affinity function: the widespread usage of non-learnable pre-defined distances, such as the Euclidean and Cosine ones. This paper proposes an effective approach for cross-camera data-association focused on a global solution, instead of processing cameras by pairs. To avoid the usage of fixed distances and thresholds, we leverage the connectivity of Graph Neural Networks, previously unused in this scope, using a Message Passing Network to jointly learn features and similarity functions. We validate the proposal for pedestrian cross-camera association, showing results over the EPFL multi-camera pedestrian dataset. Our approach considerably outperforms the literature data association techniques, without requiring to be trained in the same scenario in which it is tested. Our code is available athttps://www-vpu.eps.uam.es/publications/gnn
Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Pablo Carballeira
IEEE Trans. Circuits Syst. Video Technol.3
2022 Online clustering-based multi-camera vehicle tracking in scenarios with overlapping FOVs
abstract
Abstract Multi-Target Multi-Camera (MTMC) vehicle tracking is an essential task of visual traffic monitoring, one of the main research fields of Intelligent Transportation Systems. Several offline approaches have been proposed to address this task; however, they are not compatible with real-world applications due to their high latency and post-processing requirements. This lack of suitable approaches motivates our proposal: A new low-latency online approach for MTMC tracking in scenarios with partially overlapping fields of view (FOVs), such as road intersections. Firstly, the proposed approach detects vehicles at each camera. Then, the detections are merged between cameras by applying cross-camera clustering based on appearance and location. Lastly, the clusters containing different detections of the same vehicle are temporally associated to compute the tracks on a frame-by-frame basis. The experiments show promising low-latency results while addressing real-world challenges such as the a priori unknown and time-varying number of targets and the continuous state estimation of them without performing any post-processing of the trajectories. Our code is available at http://www-vpu.eps.uam.es/publications/Online-MTMC-Tracking .
Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Marcos Escudero-Viñolo
Multim. Tools Appl.3
2019 On guiding video object segmentation
abstract
This paper presents a novel approach for segmenting moving objects in unconstrained environments using guided convolutional neural networks. This guiding process relies on foreground masks from independent algorithms (i.e. state-of-the-art algorithms) to implement an attention mechanism that incorporates the spatial location of foreground and background to compute their separated representations. Our approach initially extracts two kinds of features for each frame using colour and optical flow information. Such features are combined following a multiplicative scheme to benefit from their complementarity. These unified colour and motion features are later processed to obtain the separated foreground and background representations. Then, both independent representations are concatenated and decoded to perform foreground segmentation. Experiments conducted on the challenging DAVIS 2016 dataset demonstrate that our guided representations not only outperform non-guided, but also recent and top-performing video object segmentation algorithms.
Diego Ortego, Kevin McGuinness, Juan C. SanMiguel, Eric Arazo Sanchez, José María Martínez Sanchez, Noel E. O'Connor
CBMI5
2019 Incorporating wheelchair users in people detection
Rafael Martin Nieto, Álvaro García-Martín, José María Martínez Sanchez
Multim. Tools Appl.3
2019 Hierarchical Improvement of Foreground Segmentation Masks in Background Subtraction
abstract
A plethora of algorithms have been defined for foreground segmentation, a fundamental stage for many computer vision applications. In this paper, we propose a post-processing framework to improve the foreground segmentation performance of background subtraction algorithms. We define a hierarchical framework for extending segmented foreground pixels to undetected foreground object areas and for removing erroneously segmented foreground. First, we create a motion-aware hierarchical image segmentation of each frame that prevents merging foreground and background image regions. Then, we estimate the quality of the foreground mask through the fitness of the binary regions in the mask and the hierarchy of segmented regions. Finally, the improved foreground mask is obtained as an optimal labeling by jointly exploiting foreground quality and spatial color relations in a pixel-wise fully connected conditional random field. Experiments are conducted over four large and heterogeneous data sets with varied challenges (CDNET2014, LASIESTA, SABS, and BMC) demonstrating the capability of the proposed framework to improve background subtraction results.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
IEEE Trans. Circuits Syst. Video Technol.3
2019 Automatic Vacant Parking Places Management System Using Multicamera Vehicle Detection
abstract
This paper presents a multicamera system for vehicles detection and their corresponding mapping into the parking spots of a parking lot. Approaches from the state-of-the-art system, which work properly in controlled scenarios, have been validated using small amount of sequences and without more challenging realistic conditions (illumination changes and different weather conditions). On the other hand, most of them are not complete systems, but provide only parts of them, usually detectors. The proposed system has been designed for realistic scenarios considering different cases of occlusion, illumination changes, and different climatic conditions; a real scenario (the International Pittsburgh Airport parking lot) has been targeted with the condition that existing parking security cameras can be used, avoiding the deployment of new cameras or other sensors infrastructures. For design and validation, a new multicamera data set has been recorded. The system is based on existing object detectors (the results of two of them are shown) and different proposed postprocessing stages. The results clearly show that the proposed system works correctly in challenging scenarios including almost total occlusions, illumination changes, and different weather conditions.
Rafael Martin Nieto, Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez
IEEE Trans. Intell. Transp. Syst.4
2017 Stand-alone quality estimation of background subtraction algorithms
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.3
2016 Rejection based multipath reconstruction for background estimation in SBMnet 2016 dataset
abstract
Background Estimation in video consists in extracting a foreground-free image from a set of training frames. In this paper, we overview a temporal-spatial block-level approach for background estimation in video and present their results in the SBMnet dataset. First, the employed approach uses a Temporal Analysis module to obtain a compact representation of the training data that is later clustered by a threshold-free technique to generate background candidates at each block location. Then, a Spatial Analysis module iteratively reconstructs the background using a multipath reconstruction guided by background smoothness constraints. The experimental results in the SBMnet dataset demonstrates the utility of the employed approach against stationary objects and its weaknesses when motion information is involved.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
ICPR3
2016 Rejection based multipath reconstruction for background estimation in video sequences with stationary objects
abstract
Background estimation in video consists in extracting a foreground-free image from a set of training frames. Moving and stationary objects may affect the background visibility, thus invalidating the assumption of many related literature where background is the temporal dominant data. In this paper, we present a temporal-spatial block-level approach for background estimation in video to cope with moving and stationary objects. First, a Temporal Analysis module obtains a compact representation of the training data by motion filtering and dimensionality reduction. Then, a threshold-free hierarchical clustering determines a set of candidates to represent the background for each spatial location (block). Second, a Spatial Analysis module iteratively reconstructs the background using these candidates. For each spatial location , multiple reconstruction hypotheses (paths) are explored to obtain its neighboring locations by enforcing inter-block similarities and intra-block homogeneity constraints in terms of color discontinuity, color dissimilarity and variability. The experimental results show that the proposed approach outperforms the related state-of-the-art over challenging video sequences in presence of moving and stationary objects.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.3
2016 Visual attention based on a joint perceptual space of color and brightness for improved video tracking
Víctor Fernández-Carbajales Canete, Miguel Ángel García, José María Martínez Sanchez
Pattern Recognit.3
2015 Post-processing approaches for improving people detection performance
Álvaro García-Martín, José María Martínez Sanchez
Comput. Vis. Image Underst.2
2015 People detection in surveillance: classification and evaluation
abstract
Nowadays, people detection in video surveillance environments is a task that has been generating great interest. There are many approaches trying to solve the problem either in controlled scenarios or in very specific surveillance applications. The main objective of this study is to give a comprehensive and extensive evaluation of the state of the art of people detection regardless of the final surveillance application. For this reason, first, the different processing tasks involved in the automatic people detection in video sequences have been defined, then a proper classification of the state of the art of people detection has been made according to the two most critical tasks, object detection and person model, that are needed in every detection approach. Finally, experiments have been performed on an extensive dataset with different approaches that completely cover the proposed classification and support the conclusions drawn from the state of the art.
Álvaro García-Martín, José María Martínez Sanchez
IET Comput. Vis.2
2015 Long-Term Stationary Object Detection Based on Spatio-Temporal Change Detection
abstract
We present a block-wise approach to detect stationary objects based on spatio-temporal change detection. First, block candidates are extracted by filtering out consecutive blocks containing moving objects. Then, an online clustering approach groups similar blocks at each spatial location over time via statistical variation of pixel ratios. The stability changes are identified by analyzing the relationships between the most repeated clusters at regular sampling instants. Finally, stationary objects are detected as those stability changes that exceed an alarm time and have not been visualized before. Unlike previous approaches making use of Background Subtraction, the proposed approach does not require foreground segmentation and provides robustness to illumination changes, crowds and intermittent object motion. The experiments over an heterogeneous dataset demonstrate the ability of the proposed approach for short- and long-term operation while overcoming challenging issues.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
IEEE Signal Process. Lett.3
2014 A semi-supervised system for players detection and tracking in multi-camera soccer videos
Rafael Martín, José María Martínez Sanchez
Multim. Tools Appl.2
2014 A synthetic training framework for providing gesture scalability to 2.5D pose-based hand gesture recognition systems
Javier Molina, José María Martínez Sanchez
Mach. Vis. Appl.2
2014 A natural and synthetic corpus for benchmarking of hand gesture recognition systems
Javier Molina, José A. Pajuelo, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez
Mach. Vis. Appl.5
2013 An automatic system for sports analytics in multi-camera tennis videos
abstract
This paper presents an automatic system which after a simple previous configuration is able to detect and track each one of the players on the court or field in single player sports. After that, the system is able to extract statistics and performance of the players. This system is complete, general and modular, to be improved and modified by future work. The system is based on a monocamera detection and tracking system, originally designed for video surveillance, which has been adapted for its use in the individual sports domain. Target sports of the developed system are individual sports (e.g., tennis, paddle tennis) where the players have its own side of the field.
Rafael Martin Nieto, José María Martínez Sanchez
AVSS2
2013 Combining MPEG Tools to Generate Video Summaries Adapted to the Terminal and Network
abstract
Moving picture experts group (MPEG) standards provide tools for a broad range of purposes, covering from coding to metadata description tools. In this paper, the combined use of tools from different MPEG standards is described in the context of a video summarization application. The main objective of the framework is the efficient generation of summaries, integrated with their adaptation to the user's terminal and network. The MPEG-4 Scalable Video Coding specification is used for fast adaptation and summary bitstream generation. MPEG-21 digital item adaptation tools are used to describe metadata related to the user's terminal and network. MPEG-7 tools are used to describe the summary. Finally, the framework is compared with alternative approaches (variations and transcoding), in terms of efficiency, rate-distortion performance and other aspects.
Luis Herranz, José María Martínez Sanchez
Comput. J.2
2013 Preface for the special issue of MTAP following CBMI 2011
José María Martínez Sanchez, Bernard Mérialdo, Jenny Benois-Pineau, Joemon M. Jose
Multim. Tools Appl.1
2013 Real-time user independent hand gesture recognition from time-of-flight camera video using static and dynamic models
Javier Molina, Marcos Escudero-Viñolo, Alessandro Signoriello, Montse Pardàs, Christian Ferran Bennström, Jesús Bescós, Ferran Marqués, José María Martínez Sanchez
Mach. Vis. Appl.8
2013 A semantic-guided and self-configurable framework for video analysis
Juan C. SanMiguel, José María Martínez Sanchez
Mach. Vis. Appl.2
2012 People-background segmentation with unequal error cost
abstract
We address the problem of segmenting a video in two classes of different semantic value, namely background and people, with the goal of guaranteeing that no people (or body parts) are classified as background. Body parts classified as background are given a higher classification error cost (segmentation with bias on background), as opposed to traditional approaches focused on people detection. To generate the people-background segmentation mask, the proposed approach first combines detection confidence maps of body parts and then extends them in order to derive a background mask, which is finally post-processed using morphological operators. Experiments validate the performance of our algorithm in different complex indoor and outdoor scenes with both static and moving cameras.
Álvaro García-Martín, Andrea Cavallaro, José María Martínez Sanchez
ICIP3
2012 Standalone evaluation of deterministic video tracking
abstract
We present an approach for performance evaluation of deterministic video trackers without ground-truth data. The proposed approach detects if a tracker is correctly operating over time using two main steps. First, it transforms the output of the localization step into a distribution of the target state, which emulates a multi-hypothesis tracker. Then, the uncertainty of such distribution is estimated to determine the time instants when the tracker is stable. A time-reversed analysis is used to identify tracker recovery after unsuccessful operation. The proposed approach is demonstrated on the well-known MeanShift tracker. The results over a heterogeneous dataset show that the proposed approach outperforms the related state-of-the-art methods in presence of tracking challenges such as occlusions, illumination and scale changes, and clutter.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
ICIP3
2012 Bounded non-deterministic planning for multimedia adaptation
Fernando López Hernández 0001, Dietmar Jannach, José María Martínez Sanchez, Christian Timmerer, Narciso García, Hermann Hellwagner
Appl. Intell.3
2012 A semantic-based probabilistic approach for real-time video event recognition
Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.2
2012 On collaborative people detection and tracking in complex scenarios
Álvaro García-Martín, José María Martínez Sanchez
Image Vis. Comput.2
2012 On-line video abstract generation of multimedia news
Víctor Valdés, José María Martínez Sanchez
Multim. Tools Appl.2
2012 A corpus for benchmarking of people detection algorithms
Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós
Pattern Recognit. Lett.2
2012 Adaptive Online Performance Evaluation of Video Trackers
abstract
We propose an adaptive framework to estimate the quality of video tracking algorithms without ground-truth data. The framework is divided into two main stages, namely, the estimation of the tracker condition to identify temporal segments during which a target is lost and the measurement of the quality of the estimated track when the tracker is successful. A key novelty of the proposed framework is the capability of evaluating video trackers with multiple failures and recoveries over long sequences. Successful tracking is identified by analyzing the uncertainty of the tracker, whereas track recovery from errors is determined based on the time-reversibility constraint. The proposed approach is demonstrated on a particle filter tracker over a heterogeneous data set. Experimental results show the effectiveness and robustness of the proposed framework that improves state-of-the-art approaches in the presence of tracking challenges such as occlusions, illumination changes, and clutter and on sequences containing multiple tracking errors and recoveries.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
IEEE Trans. Image Process.3
2012 Scalable Comic-Like Video Summaries and Layout Disturbance
abstract
This paper describes an efficient system for scalable video summarization that exploits comic-like summaries and multi-scale representations to facilitate interactivity and balance between content coverage and compactness. Due to the layout disturbance induced by the transitions between scales, a new heuristic algorithm is proposed to restrict changes to bounded summary segments. Conducted user evaluations show that the proposed methodology improves usability while keeping the summaries compact and informative.
Luis Herranz, Janko Calic, José María Martínez Sanchez, Marta Mrak
IEEE Trans. Multim.3
2012 Automatic evaluation of video summaries
abstract
This article describes a method for the automatic evaluation of video summaries based on the training of individual predictors for different quality measures from the TRECVid 2008 BBC Rushes Summarization Task. The obtained results demonstrate that, with a large set of evaluation data, it is possible to train fully automatic evaluation systems based on visual features automatically extracted from the summaries. The proposed approach will enable faster and easier estimation of the results of newly developed abstraction algorithms and the study of which summary characteristics influence their perceived quality.
Víctor Valdés, José María Martínez Sanchez
ACM Trans. Multim. Comput. Commun. Appl.2
2011 Discrimination of abandoned and stolen object based on active contours
abstract
In this paper we propose an approach based on active contours to discriminate previously detected static foreground regions between abandoned and stolen. Firstly, the static foreground object contour is extracted. Then, an active contour adjustment is performed on the current and the background frames. Finally, similarities between the initial contour and the two adjustments are studied to decide whether the object is abandoned or stolen. Three different methods have been tested for this adjustment. Experimental results over a heterogeneous dataset show that the proposed method outperforms state-of-art approaches and provides a robust solution against non-accurate data (i.e., foreground static objects wrongly segmented) that is common in complex scenarios.
Luis Caro Campos, Juan C. SanMiguel, José María Martínez Sanchez
AVSS3
2011 Improving the efficiency and accuracy of visual attention
abstract
Visual attention is the cognitive process of selectively focusing on certain areas of a visual scene while ignoring the others. It is a desirable capability for intelligent video surveillance systems, as it allows them to control the aim of mobile cameras or to selectively process the most relevant parts of the captured images. This paper proposes an adaptation of a well-known biologically-inspired visual attention model in order to increase its computational efficiency without sacrificing its accuracy, and shows that the latter can be further improved through a supervised training stage that fine-tunes the model to the particular application scope in which the system is being utilized. Experimental results and comparisons with previous visual attention techniques are shown and discussed.
Víctor Fernández-Carbajales Canete, Miguel Ángel García, José María Martínez Sanchez
AVSS3
2011 People detection based on appearance and motion models
abstract
The main contribution of this paper is a new people detection algorithm based on motion information. The algorithm builds a people motion model based on the Implicit Shape Model (ISM) Framework and the MoSIFT descriptor. We also propose a detection system that integrates appearance, motion and tracking information. Experimental results over sequences extracted from the TRECVID dataset show that our new people motion detector produces results comparable to the state of the art and that the proposed multimodal fusion system improves the obtained results combining the three information sources.
Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez
AVSS3
2011 A model for preference-driven multimedia adaptation decision-making in the MPEG-21 framework
Fernando López Hernández 0001, José María Martínez Sanchez, Narciso García
Multim. Tools Appl.2
2010 Robust Real Time Moving People Detection in Surveillance Scenarios
abstract
In this paper an improved real time algorithm for detecting pedestrians in surveillance video is proposed. The algorithm is based on people appearance and defines a person model as the union of four models of body parts. Firstly, motion segmentation is performed to detect moving pixels. Then, moving regions are extracted and tracked. Finally, the detected moving objects are classified as human or nonhuman objects. In order to test and validate the algorithm, we have developed a dataset containing annotated surveillance sequences of different complexity levels focused on the pedestrians detection. Experimental results over this dataset show that our approach performs considerably well at real time and even better than other real and non-real time approaches from the state of art.
Álvaro García-Martín, José María Martínez Sanchez
AVSS2
2010 On the Evaluation of Background Subtraction Algorithms without Ground-Truth
abstract
In video-surveillance systems, the moving object segmentation stage (commonly based on background subtraction) has to deal with several issues like noise, shadows and multimodal backgrounds. Hence, its failure is inevitable and its automatic evaluation is a desirable requirement for online analysis. In this paper, we propose a hierarchy of existing performance measures not-based on ground-truth for video object segmentation. Then, four measures based on color and motion are selected and examined in detail with different segmentation algorithms and standard test sequences for video object segmentation. Experimental results show that color-based measures perform better than motion-based measures and background multimodality heavily reduces the accuracy of all obtained evaluation results.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS2
2010 Stationary foreground detection using background subtraction and temporal difference in video surveillance
abstract
In this paper we describe a new algorithm focused on obtaining stationary foreground regions, which is useful for applications like the detection of abandoned/stolen objects and parked vehicles. Firstly, a sub-sampling scheme based on background subtraction techniques is implemented to obtain stationary foreground regions. Secondly, some modifications are introduced on this base algorithm with the purpose of reducing the amount of stationary foreground detected. Finally, we evaluate the proposed algorithm and compare results with the base algorithm using video surveillance sequences from PETS 2006, PETS 2007 and I-LIDS for AVSS 2007 datasets. Experimental results show that the proposed algorithm increases the detection of stationary foreground regions as compared to the base algorithm.
Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez
ICIP3
2010 Evaluation of on-line quality estimators for object tracking
abstract
Failure of tracking algorithms is inevitable in real and on-line tracking systems. The online estimation of the track quality is therefore desirable for detecting tracking failures while the algorithm is operating. In this paper, we propose a taxonomy and present a comparative evaluation of online quality estimators for video object tracking. The measures are compared over a heterogeneous video dataset with standard sequences. Among other results, the experiments show, that the Observation Likelihood (OL) measure is an appropriate quality measure for overall tracking performance evaluation, while the Template Inverse Matching (TIM) measure is appropriate to detect the start and the end instants of tracking failures.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
ICIP3
2010 On the Advantages of the Use of Bitstream Extraction for Video Summary Generation
Luis Herranz, José María Martínez Sanchez
MMM2
2010 A framework for video abstraction systems analysis and modelling from an operational point of view
Víctor Valdés, José María Martínez Sanchez
Multim. Tools Appl.2
2010 A Framework for Scalable Summarization of Video
abstract
Video summaries provide compact representations of video sequences, with the length of the summary playing an important role, trading off the amount of information conveyed and how fast it can be visualized. This letter proposes scalable summarization as a method to easily adapt the summary to a suitable length, according to the requirements in each case, along with a suitable framework. The analysis algorithm uses a novel iterative ranking procedure in which each summary is the result of the extension of the previous one, balancing information coverage and visual pleasantness. The result of the algorithm is a ranked list, a scalable representation of the sequence useful for summarization. The summary is then efficiently generated from the bitstream of the sequence using bitstream extraction.
Luis Herranz, José María Martínez Sanchez
IEEE Trans. Circuits Syst. Video Technol.2
2009 Comparative Evaluation of Stationary Foreground Object Detection Algorithms Based on Background Subtraction Techniques
abstract
In several video surveillance applications, such as the detection of abandoned/stolen objects or parked vehicles,the detection of stationary foreground objects is a critical task. In the literature, many algorithms have been proposed that deal with the detection of stationary foreground objects, the majority of them based on background subtraction techniques. In this paper we discuss various stationary object detection approaches comparing them in typical surveillance scenarios (extracted from standard datasets). Firstly, the existing approaches based on background-subtraction are organized into categories. Then, a representative technique of each category is selected and described. Finally, a comparative evaluation using objective and subjective criteria is performed on video surveillance sequences selected from the PETS 2006 and i-LIDS for AVSS 2007 datasets, analyzing the advantages and drawbacks of each selected approach.
Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez
AVSS3
2009 An Ontology for Event Detection and its Application in Surveillance Video
abstract
In this paper, we propose an ontology for representing the prior knowledge related to video event analysis. It is composed of two types of knowledge related to the application domain and the analysis system. Domain knowledge involves all the high level semantic concepts in the context of each examined domain (objects, events, context...) whilst system knowledge involves the capabilities of the analysis system (algorithms, reactions to events...). The proposed ontology has been structured in two parts: the basic ontology (composed of the basic concepts and their specializations) and the domain-specific extensions. Additionally, a video analysis framework based on the proposed ontology is defined for the analysis of different application domains showing the potential use of the proposed ontology. In order to show the real applicability of the proposed ontology, it is specialized for the underground video-surveillance domain showing some results that demonstrate the usability and effectiveness of the proposed ontology.
Juan C. SanMiguel, José María Martínez Sanchez, Álvaro García-Martín
AVSS2
2009 Shadow detection in video surveillance by maximizing agreement between independent detectors
abstract
This paper starts from the idea of automatically choosing the appropriate thresholds for a shadow detection algorithm. It is based on the maximization of the agreement between two independent shadow detectors without training data. Firstly, this shadow detection algorithm is described and then, it is adapted to analyze video surveillance sequences. Some modifications are introduced to increase its robustness in generic surveillance scenarios and to reduce its overall computational cost (critical in some video surveillance applications). Experimental results show that the proposed modifications increase the detection reliability as compared to some previous shadow detection algorithms and performs considerably well across a variety of multiple surveillance scenarios.
Juan C. SanMiguel, José María Martínez Sanchez
ICIP2
2009 An efficient summarization algorithm based on clustering and bitstream extraction
abstract
Visualizing video is a very time consuming task. Video summaries are compact representations and very useful in video management systems. However, in some applications, in which the summary must be generated on demand, it usually implies a high delay for the user. Video summarization is a very demanding task in terms of processing resources, and it is very difficult to obtain a summary of a large video with low delay. In this paper a flexible summarization framework is presented which can create storyboards and video skims using bitstream extraction to generate the summary after an appropriate analysis algorithm. This technique is very efficient and achieves very low delay while keeping a reasonable quality of video summaries.
Luis Herranz, José María Martínez Sanchez
ICME2
2009 An integrated approach to summarization and adaptation using H.264/MPEG-4 SVC
Luis Herranz, José María Martínez Sanchez
Signal Process. Image Commun.2
2009 On the use of hierarchical prediction structures for efficient summary generation of H.264/AVC bitstreams
Luis Herranz, José María Martínez Sanchez
Signal Process. Image Commun.2
2009 Visual Tools for ROI Montage in an Image2Video Application
abstract
This letter presents an image to video adaptation system that transmodes images in order to be viewed on small displays without a significant loss of information. The followed approach is to automate the process of manual browsing and zooming through an image by simulating the movement of a virtual camera. The presented system proposes two modules: one for analysis (ROI Extraction: for defining the regions of interest (ROIs) in the image) and the second for transmoding (ROIs2Video: for creating a video from any predefined list of relevance annotated ROIs). The separation of the image analysis and transmoding steps allows to use the ROIs2Video module in different application domains, as the transmoding will only be driven by the available relevance values associated to the ROIs. Results are presented for automatic and manual ROI generation modules. User evaluation results show that the developed system offers a visually attractive solution.
Fernando Barreiro-Megino, José María Martínez Sanchez, Víctor Valdés
IEEE Trans. Circuits Syst. Video Technol.2
2008 Robust Unattended and Stolen Object Detection by Fusing Simple Algorithms
abstract
In this paper a new approach for detecting unattended or stolen objects in surveillance video is proposed. It is based on the fusion of evidence provided by three simple detectors. As a first step, the moving regions in the scene are detected and tracked. Then, these regions are classified as static or dynamic objects and human or nonhuman objects. Finally, objects detected as static and nonhuman are analyzed with each detector. Data from these detectors are fused together to select the best detection hypotheses. Experimental results show that the fusion-based approach increases the detection reliability as compared to the detectors and performs considerably well across a variety of multiple scenarios operating at realtime.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS2
2008 Generation of scalable summaries based on iterative GoP ranking
abstract
Video skims and image storyboards are two widely used abstractions for representing the essence of a video sequence, crucial for effective browsing and retrieval applications. In this paper we propose a flexible approach to find a reasonable balance between the semantic coverage and naturalness of the generated summaries, targeting a wide range of summarization ratios. The result of the algorithm is a scalable representation of the information required for summarization (a ranked list of GoPs) with a number of advantages in terms of efficient generation and potential applications.
Luis Herranz, José María Martínez Sanchez
ICIP2
2008 A ground truth for motion-based video-object segmentation
abstract
This paper describes the design procedure followed to generate a ground truth for the evaluation of motion-based algorithms for video-object segmentation. A thorough review and classification of the critical factors that affect the behavior of segmentation algorithms results in a set of video scripts which have then been filmed. Foreground objects have been recorded in a chroma studio, in order to automatically obtain pixel-level high quality segmentation masks for each generated sequence. The resulting corpus (segmentation ground-.truth plus filmed sequences mounted over different backgrounds) is available for research purposes under a license agreement.
Fabricio Tiburzi, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez
ICIP4
2007 On the effect of motion segmentation techniques in description based adaptive video transmission
abstract
This paper presents the results of analysing the effect of different motion segmentation techniques in a system that transmits the information captured by a static surveillance camera in an adaptative way based on the on-line generation of descriptions and their descriptions at different levels of detail. The video sequences are analyzed to detect the regions of activity (motion analysis) and to differentiate them from the background, and the corresponding descriptions (mainly MPEG-7 moving regions) are generated together with the textures of the moving regions and the associated background image. Depending on the available bandwidth, different levels of transmission are specified, ranging from just sending the descriptions generated to a transmission with all the associated images corresponding to the moving objects and background. We study the effect of three motion segmentation algorithms in several aspects such as accurate segmentation, size of the descriptions generated, computational efficiency and reconstructed data quality.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS2
2007 TV-Anytime Phase 1 and MPEG-7
abstract
Abstract Personal video recorders have the capability to change the media delivery industry fundamentally, and in this context, many believe the real international age of personal digital recorders (PDRs) will arrive with the use of “open” systems. The world reached an important milestone with the publication of the TV‐Anytime Phase 1 specifications for unidirectional broadcast and metadata services over bidirectional networks. TV‐Anytime is a worldwide prestandardization body; this article gives an overview of the main features of TV‐Anytime's metadata specification and its relationship to MPEG‐7 and provides insight into ways two organizations concerned with standards work together. Phase 2 has since been completed and TV‐Anytime has been adopted by various international standards organizations dealing with telecommunications and is now in the implementation phase.
Jean-Pierre Evain, José María Martínez Sanchez
J. Assoc. Inf. Sci. Technol.2
2007 MPEG-7 tools for Universal Multimedia Access
abstract
Abstract Universal Multimedia Access (UMA) deals with seamless access to once‐only‐created content via any kind of terminal and any kind of network connectivity, which implies that the content should be adapted in order to fit a variety of terminal and network characteristics, as well as user preferences. The MPEG‐7 standard offers some support for UMA within its section on Multimedia Description Schemes (MDS). Within the standard, several groups of tools serve this purpose. For instance, the Navigation and Access Tools provide some Description Schemes that allow the description of adapted content variations and summaries and allow for preprocessed content versions. Some support is also found in the Content Metadata Tools (Media and Usage Tools), for real‐time ease in creation of online content versions and in limited support for session description, which is completed in MPEG‐21.
José María Martínez Sanchez
J. Assoc. Inf. Sci. Technol.1
2007 Content-driven adaptation of on-line video
Jesús Bescós, José María Martínez Sanchez, Luis Herranz, Fabricio Tiburzi
Signal Process. Image Commun.2
2006 Video summaries generation and access via personalized delivery of multimedia presentations adapted to service and terminal
abstract
This article is centered on describing the provision of universal multimedia access services for video summaries via multimedia presentations that allow the integration of multimedia messaging service (MMS)-enabled terminals in the framework of the deferred time environment (DTE) of the DYMAS system. The system uses the framework of MPEG-7 and MPEG-21 to provide description metadata of the multimedia content and the usage context (including terminal, network capabilities, and user preferences), respectively. These descriptions are the base for the main functionalities of the complete system that provides personalized access to content (filtering by user preferences or via querying) that is first adapted to a multimedia presentation (generating a video summary that is represented via keyframes with synchronized audio clips) by the Presentation Builder and afterward adapted to the current service and terminal by the Adaptation Engine. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 785–800, 2006.
Marta Padilla, José María Martínez Sanchez, Luis Herranz
Int. J. Intell. Syst.2
2005 A unified model for techniques on video-shot transition detection
abstract
A first step required to allow video indexing and retrieval of visual data is to perform a temporal segmentation, that is, to find the location of camera-shot transitions, which can be either abrupt (i.e., cuts) or gradual (e.g., fades, dissolves, wipes). After a critical review of most approaches seeking to solve this problem, we propose a unified detection model (both for abrupt and all types of gradual transitions), as well as an implementation whose results improve upon those of all the inspected reports. The innovation of the approach presented here is centered on mapping the space of inter-frame distances onto a new space of decision better suited to achieving a sequence-independent thresholding. This mapping aims to consider frame ordering information within the thresholding process; it is based on the parametric modeling of the patterns that transitions generate on the distances' output. As opposed to most reviewed works, our results are detailed over a large and representative sample of more than 1500 cuts and 250 gradual transitions, which make up a significant part (200 min) of the MPEG-7 testing material; this ensures a high degree of confidence in the validity of our approach.
Jesús Bescós, Guillermo Cisneros, José María Martínez Sanchez, José Manuel Menéndez, Julián Cabrera
IEEE Trans. Multim.3
2003 MISS: A Generic Model for MetaInformation SubSystems
José María Martínez Sanchez, Julián Cabrera, Jesús Bescós, Guillermo Cisneros, José Manuel Menéndez
Multim. Tools Appl.1
2002 Towards universal access to content using MPEG-7
abstract
This paper presents a system providing functionalities for cataloging multimedia content using MPEG-7 and accessing to content and descriptions. The cataloging application indexes content using MPEG-7 and creates annotated variations in order to have the capability of offering media content to a large amount of different terminals and through different access networks. The created multimedia database, both descriptions and content (original sources and variations), is used by, currently, two applications: a searching application, which allows a user selecting specific tags to find the desired media; and a filtering application for transparent access to the content database using profiles. These profiles can be selected from a profiles database or created by the user specifying content preferences, and network and terminal parameters.
José María Martínez Sanchez, César González, Oscar Fernández, Clara García, Jaime de Ramón
ACM Multimedia1
2002 Authoring 744: first results
abstract
This paper presents the first results of the Authoring744 research initiative, which uses MPEG-7 to synthesize MPEG-4 content. The objective is to use MPEG-7 content descriptions to synthesize content, instead of creating descriptions by analyzing existing content. The output uses MPEG-4 XMT as the representation format, which is further used to create an MPEG-4 binary format, which can in turn be played.
José María Martínez Sanchez, Luis F. Rubio, Francisco Morán
ACM Multimedia1
2001 Robust digital image watermarking using DWT, DFT and quality based average
abstract
This paper presents a digital image watermarking system that complements a wavelet based insertion module, with a resynchronization module, and a method for selecting the watermark using an estimated-quality-based average. The proposed system has been tested with attacks performed by Stirmark, obtaining results of robustness over 90%.
Eduardo Fullea, José María Martínez Sanchez
ACM Multimedia2
2000 A Unified Approach to Gradual Shot Transition Detection
abstract
Nowadays, content-based retrieval of video material is based on the availability of meta-data linked to it. Current approaches to automatically extract these data start from a temporal segmentation of the audio-visual material, that is, a location of the camera shot transitions. Although abrupt transition detection is a problem almost solved, success on gradual transition detection is still very low. We present a detector for all types of gradual transitions (chromatic, geometric, and mixed), based on in depth modelling of the patterns that these transitions generate on a specific frame distance. Results are presented over a representative sample of more than 250 gradual transitions, which belong to a significant part (200') of the MPEG-7 testing material, achieving a recall (percentage of effects correctly detected) and a precision (percentage of non-false positives in the set of detection) notably higher than the ones reported so far.
Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros, Julián Cabrera, José María Martínez Sanchez
ICIP5
1999 Generic meta data browsing system for multimedia document retrieval
abstract
This paper presents the architecture of a multimedia information-indexing subsystem, the Meta-information subsystem. This subsystem makes use of meta data, that is, content-based information extracted from multimedia materials and other associated information. It also provides new mechanisms to store, query and browse meta data, and manages their relationship with the multimedia materials stored in the system.
Julián Cabrera, José María Martínez Sanchez, Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros
MMSP2
1998 Developing Multimedia Applications: System Modeling and Implementation
José María Martínez Sanchez, Jesús Bescós, Guillermo Cisneros
Multim. Tools Appl.1