VLDB 2026 Research / reviewers in the wild / expert
José María Martínez Sanchez
dblp:47/3503 · also José M. Martínez 0001
· DBLP profile ↗
70ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0002-2236-1769ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Soft-Labelling for Budget-Constrained Semantic Segmentation: Bringing Coherence to label Down-SamplingabstractIn semantic segmentation, training data down-sampling is commonly performed due to resource limitations, the need to adapt image size to the model input, or to improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled colour and ground-truth label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labelling that better conserves label information after down-sampling, thereby, fully aligning soft-labels with image data to keep the distribution of the sampled pixels for down-sampling. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that our proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly fewer computational resources than foremost methods. This proposal enables competitive research for semantic segmentation under resource constraints. Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Media | 4 |
| 2025 | An image-processing toolkit for remote photoplethysmographyabstractAbstract Objective. Image-processing-based remote photoplethysmography algorithms are usually composed of steps where different methods are used, and often, researchers perform these steps using methods that are not necessarily the best for their application. With our toolkit, we want to provide easy and fast access to different state-of-the-art methods for the most common image-processing steps in remote photoplethysmography algorithms. Methods. Our toolkit was programmed in Python and was developed with sequential workflow in mind, making it versatile and easy to use in interactive environments. It also includes tools so the users can modify or extend it if they want to, and will be updated as new methods for the different steps are published. Results. Our use case examples and validation show an effective approach and how the toolkit can be used for exhaustive evaluation and ablation studies in a simple way. We also show how choosing different methods can affect the final heart rate estimation accuracy at the cost of computation time. Conclusion. With this toolkit we are providing researchers with a versatile, easy-to-use tool, with access to different methods for the most common steps in remote photoplethysmography algorithms. Significance. Our toolkit is a relevant tool for researchers in the remote photoplethysmography field due to their versatility, ease of use, and adaptability. (It will be available on https://github.com/Montyro/rppgtk github upon acceptance). Javier Montalvo, Álvaro García-Martín, José María Martínez Sanchez |
Multim. Tools Appl. | 3 |
| 2024 | Long-Term Geo-Positioned Re-Identification Dataset of Urban ElementsabstractThis paper introduces UrbAM-ReID, a new long-term geo-positioned urban ReID dataset. It is composed by four subdatasets recording the same trajectory at the UAM Campus, each one recorded in different seasons and including an inverse direction recording. While most of the current datasets in the state-of-the-art focus on person re-identification, with vehicles as the second most explored object, our work specifically addresses urban objects re-identification, currently, waste containers, rubbish bins, and crosswalks. The dataset provides different attributes of the annotated objects, like their classes, their foreground or background status and the geo-position. Several evaluation configurations can be defined to simulate realistic scenarios that may arise in actual situations within the management of urban elements, considering the utilization of just visual data, or incorporating additional attributes, providing different complexity levels. Finally, the dataset is used for defining a benchmark where two state-of-the-art systems are evaluated. The dataset and supplementary material is available in https://github.com/vpulab/UrbAMReID Paula Moral, Álvaro García-Martín, José María Martínez Sanchez |
ICIP | 3 |
| 2023 | Vehicle Re-Identification Based on Unsupervised Domain Adaptation by Incremental Generation of Pseudo-Labels
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez |
CIARP | 3 |
| 2023 | Enhancing vehicle re-identification via synthetic training datasets and re-ranking based on video-clips informationabstractAbstract Vehicle re-identification (ReID) aims to find a specific vehicle identity across multiple non-overlapping cameras. The main challenge of this task is the large intra-class and small inter-class variability of vehicles appearance, sometimes related with large viewpoint variations, illumination changes or different camera resolutions. To tackle these problems, we proposed a vehicle ReID system based on ensembling deep learning features and adding different post-processing techniques. In this paper, we improve that proposal by: incorporating large-scale synthetic datasets in the training step; performing an exhaustive ablation study showing and analyzing the influence of synthetic content in ReID datasets, in particular CityFlow-ReID and VeRi-776; and extending post-processing by including different approaches to the use of gallery video-clips of the target vehicles in the re-ranking step. Additionally, we present an evaluation framework in order to evaluate CityFlow-ReID: as this dataset has not public ground truth annotations, AI City Challenge provided an on-line evaluation service which is no more available; our evaluation framework allows researchers to keep on evaluating the performance of their systems in the CityFlow-ReID dataset. Paula Moral, Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós |
Multim. Tools Appl. | 3 |
| 2023 | Graph Neural Networks for Cross-Camera Data AssociationabstractCross-camera image data association is essential for many multi-camera computer vision tasks, such as multi-camera pedestrian detection, multi-camera multi-target tracking, 3D pose estimation, etc. This association task is typically modeled as a bipartite graph matching problem and often solved by applying minimum-cost flow techniques, which may be computationally demanding for large data. Furthermore, cameras are usually treated by pairs, obtaining local solutions, rather than finding a global solution at once for all multiple cameras. Other key issue is that of the affinity function: the widespread usage of non-learnable pre-defined distances, such as the Euclidean and Cosine ones. This paper proposes an effective approach for cross-camera data-association focused on a global solution, instead of processing cameras by pairs. To avoid the usage of fixed distances and thresholds, we leverage the connectivity of Graph Neural Networks, previously unused in this scope, using a Message Passing Network to jointly learn features and similarity functions. We validate the proposal for pedestrian cross-camera association, showing results over the EPFL multi-camera pedestrian dataset. Our approach considerably outperforms the literature data association techniques, without requiring to be trained in the same scenario in which it is tested. Our code is available athttps://www-vpu.eps.uam.es/publications/gnn Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Pablo Carballeira |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Online clustering-based multi-camera vehicle tracking in scenarios with overlapping FOVsabstractAbstract Multi-Target Multi-Camera (MTMC) vehicle tracking is an essential task of visual traffic monitoring, one of the main research fields of Intelligent Transportation Systems. Several offline approaches have been proposed to address this task; however, they are not compatible with real-world applications due to their high latency and post-processing requirements. This lack of suitable approaches motivates our proposal: A new low-latency online approach for MTMC tracking in scenarios with partially overlapping fields of view (FOVs), such as road intersections. Firstly, the proposed approach detects vehicles at each camera. Then, the detections are merged between cameras by applying cross-camera clustering based on appearance and location. Lastly, the clusters containing different detections of the same vehicle are temporally associated to compute the tracks on a frame-by-frame basis. The experiments show promising low-latency results while addressing real-world challenges such as the a priori unknown and time-varying number of targets and the continuous state estimation of them without performing any post-processing of the trajectories. Our code is available at http://www-vpu.eps.uam.es/publications/Online-MTMC-Tracking . Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Marcos Escudero-Viñolo |
Multim. Tools Appl. | 3 |
| 2019 | On guiding video object segmentationabstractThis paper presents a novel approach for segmenting moving objects in unconstrained environments using guided convolutional neural networks. This guiding process relies on foreground masks from independent algorithms (i.e. state-of-the-art algorithms) to implement an attention mechanism that incorporates the spatial location of foreground and background to compute their separated representations. Our approach initially extracts two kinds of features for each frame using colour and optical flow information. Such features are combined following a multiplicative scheme to benefit from their complementarity. These unified colour and motion features are later processed to obtain the separated foreground and background representations. Then, both independent representations are concatenated and decoded to perform foreground segmentation. Experiments conducted on the challenging DAVIS 2016 dataset demonstrate that our guided representations not only outperform non-guided, but also recent and top-performing video object segmentation algorithms. Diego Ortego, Kevin McGuinness, Juan C. SanMiguel, Eric Arazo Sanchez, José María Martínez Sanchez, Noel E. O'Connor |
CBMI | 5 |
| 2019 | Incorporating wheelchair users in people detection
Rafael Martin Nieto, Álvaro García-Martín, José María Martínez Sanchez |
Multim. Tools Appl. | 3 |
| 2019 | Hierarchical Improvement of Foreground Segmentation Masks in Background SubtractionabstractA plethora of algorithms have been defined for foreground segmentation, a fundamental stage for many computer vision applications. In this paper, we propose a post-processing framework to improve the foreground segmentation performance of background subtraction algorithms. We define a hierarchical framework for extending segmented foreground pixels to undetected foreground object areas and for removing erroneously segmented foreground. First, we create a motion-aware hierarchical image segmentation of each frame that prevents merging foreground and background image regions. Then, we estimate the quality of the foreground mask through the fitness of the binary regions in the mask and the hierarchy of segmented regions. Finally, the improved foreground mask is obtained as an optimal labeling by jointly exploiting foreground quality and spatial color relations in a pixel-wise fully connected conditional random field. Experiments are conducted over four large and heterogeneous data sets with varied challenges (CDNET2014, LASIESTA, SABS, and BMC) demonstrating the capability of the proposed framework to improve background subtraction results. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Automatic Vacant Parking Places Management System Using Multicamera Vehicle DetectionabstractThis paper presents a multicamera system for vehicles detection and their corresponding mapping into the parking spots of a parking lot. Approaches from the state-of-the-art system, which work properly in controlled scenarios, have been validated using small amount of sequences and without more challenging realistic conditions (illumination changes and different weather conditions). On the other hand, most of them are not complete systems, but provide only parts of them, usually detectors. The proposed system has been designed for realistic scenarios considering different cases of occlusion, illumination changes, and different climatic conditions; a real scenario (the International Pittsburgh Airport parking lot) has been targeted with the condition that existing parking security cameras can be used, avoiding the deployment of new cameras or other sensors infrastructures. For design and validation, a new multicamera data set has been recorded. The system is based on existing object detectors (the results of two of them are shown) and different proposed postprocessing stages. The results clearly show that the proposed system works correctly in challenging scenarios including almost total occlusions, illumination changes, and different weather conditions. Rafael Martin Nieto, Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Stand-alone quality estimation of background subtraction algorithms
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 3 |
| 2016 | Rejection based multipath reconstruction for background estimation in SBMnet 2016 datasetabstractBackground Estimation in video consists in extracting a foreground-free image from a set of training frames. In this paper, we overview a temporal-spatial block-level approach for background estimation in video and present their results in the SBMnet dataset. First, the employed approach uses a Temporal Analysis module to obtain a compact representation of the training data that is later clustered by a threshold-free technique to generate background candidates at each block location. Then, a Spatial Analysis module iteratively reconstructs the background using a multipath reconstruction guided by background smoothness constraints. The experimental results in the SBMnet dataset demonstrates the utility of the employed approach against stationary objects and its weaknesses when motion information is involved. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
ICPR | 3 |
| 2016 | Rejection based multipath reconstruction for background estimation in video sequences with stationary objectsabstractBackground estimation in video consists in extracting a foreground-free image from a set of training frames. Moving and stationary objects may affect the background visibility, thus invalidating the assumption of many related literature where background is the temporal dominant data. In this paper, we present a temporal-spatial block-level approach for background estimation in video to cope with moving and stationary objects. First, a Temporal Analysis module obtains a compact representation of the training data by motion filtering and dimensionality reduction. Then, a threshold-free hierarchical clustering determines a set of candidates to represent the background for each spatial location (block). Second, a Spatial Analysis module iteratively reconstructs the background using these candidates. For each spatial location , multiple reconstruction hypotheses (paths) are explored to obtain its neighboring locations by enforcing inter-block similarities and intra-block homogeneity constraints in terms of color discontinuity, color dissimilarity and variability. The experimental results show that the proposed approach outperforms the related state-of-the-art over challenging video sequences in presence of moving and stationary objects. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 3 |
| 2016 | Visual attention based on a joint perceptual space of color and brightness for improved video tracking
Víctor Fernández-Carbajales Canete, Miguel Ángel García, José María Martínez Sanchez |
Pattern Recognit. | 3 |
| 2015 | Post-processing approaches for improving people detection performance
Álvaro García-Martín, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 2 |
| 2015 | People detection in surveillance: classification and evaluationabstractNowadays, people detection in video surveillance environments is a task that has been generating great interest. There are many approaches trying to solve the problem either in controlled scenarios or in very specific surveillance applications. The main objective of this study is to give a comprehensive and extensive evaluation of the state of the art of people detection regardless of the final surveillance application. For this reason, first, the different processing tasks involved in the automatic people detection in video sequences have been defined, then a proper classification of the state of the art of people detection has been made according to the two most critical tasks, object detection and person model, that are needed in every detection approach. Finally, experiments have been performed on an extensive dataset with different approaches that completely cover the proposed classification and support the conclusions drawn from the state of the art. Álvaro García-Martín, José María Martínez Sanchez |
IET Comput. Vis. | 2 |
| 2015 | Long-Term Stationary Object Detection Based on Spatio-Temporal Change DetectionabstractWe present a block-wise approach to detect stationary objects based on spatio-temporal change detection. First, block candidates are extracted by filtering out consecutive blocks containing moving objects. Then, an online clustering approach groups similar blocks at each spatial location over time via statistical variation of pixel ratios. The stability changes are identified by analyzing the relationships between the most repeated clusters at regular sampling instants. Finally, stationary objects are detected as those stability changes that exceed an alarm time and have not been visualized before. Unlike previous approaches making use of Background Subtraction, the proposed approach does not require foreground segmentation and provides robustness to illumination changes, crowds and intermittent object motion. The experiments over an heterogeneous dataset demonstrate the ability of the proposed approach for short- and long-term operation while overcoming challenging issues. Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez |
IEEE Signal Process. Lett. | 3 |
| 2014 | A semi-supervised system for players detection and tracking in multi-camera soccer videos
Rafael Martín, José María Martínez Sanchez |
Multim. Tools Appl. | 2 |
| 2014 | A synthetic training framework for providing gesture scalability to 2.5D pose-based hand gesture recognition systems
Javier Molina, José María Martínez Sanchez |
Mach. Vis. Appl. | 2 |
| 2014 | A natural and synthetic corpus for benchmarking of hand gesture recognition systems
Javier Molina, José A. Pajuelo, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez |
Mach. Vis. Appl. | 5 |
| 2013 | An automatic system for sports analytics in multi-camera tennis videosabstractThis paper presents an automatic system which after a simple previous configuration is able to detect and track each one of the players on the court or field in single player sports. After that, the system is able to extract statistics and performance of the players. This system is complete, general and modular, to be improved and modified by future work. The system is based on a monocamera detection and tracking system, originally designed for video surveillance, which has been adapted for its use in the individual sports domain. Target sports of the developed system are individual sports (e.g., tennis, paddle tennis) where the players have its own side of the field. Rafael Martin Nieto, José María Martínez Sanchez |
AVSS | 2 |
| 2013 | Combining MPEG Tools to Generate Video Summaries Adapted to the Terminal and NetworkabstractMoving picture experts group (MPEG) standards provide tools for a broad range of purposes, covering from coding to metadata description tools. In this paper, the combined use of tools from different MPEG standards is described in the context of a video summarization application. The main objective of the framework is the efficient generation of summaries, integrated with their adaptation to the user's terminal and network. The MPEG-4 Scalable Video Coding specification is used for fast adaptation and summary bitstream generation. MPEG-21 digital item adaptation tools are used to describe metadata related to the user's terminal and network. MPEG-7 tools are used to describe the summary. Finally, the framework is compared with alternative approaches (variations and transcoding), in terms of efficiency, rate-distortion performance and other aspects. Luis Herranz, José María Martínez Sanchez |
Comput. J. | 2 |
| 2013 | Preface for the special issue of MTAP following CBMI 2011
José María Martínez Sanchez, Bernard Mérialdo, Jenny Benois-Pineau, Joemon M. Jose |
Multim. Tools Appl. | 1 |
| 2013 | Real-time user independent hand gesture recognition from time-of-flight camera video using static and dynamic models
Javier Molina, Marcos Escudero-Viñolo, Alessandro Signoriello, Montse Pardàs, Christian Ferran Bennström, Jesús Bescós, Ferran Marqués, José María Martínez Sanchez |
Mach. Vis. Appl. | 8 |
| 2013 | A semantic-guided and self-configurable framework for video analysis
Juan C. SanMiguel, José María Martínez Sanchez |
Mach. Vis. Appl. | 2 |
| 2012 | People-background segmentation with unequal error costabstractWe address the problem of segmenting a video in two classes of different semantic value, namely background and people, with the goal of guaranteeing that no people (or body parts) are classified as background. Body parts classified as background are given a higher classification error cost (segmentation with bias on background), as opposed to traditional approaches focused on people detection. To generate the people-background segmentation mask, the proposed approach first combines detection confidence maps of body parts and then extends them in order to derive a background mask, which is finally post-processed using morphological operators. Experiments validate the performance of our algorithm in different complex indoor and outdoor scenes with both static and moving cameras. Álvaro García-Martín, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 3 |
| 2012 | Standalone evaluation of deterministic video trackingabstractWe present an approach for performance evaluation of deterministic video trackers without ground-truth data. The proposed approach detects if a tracker is correctly operating over time using two main steps. First, it transforms the output of the localization step into a distribution of the target state, which emulates a multi-hypothesis tracker. Then, the uncertainty of such distribution is estimated to determine the time instants when the tracker is stable. A time-reversed analysis is used to identify tracker recovery after unsuccessful operation. The proposed approach is demonstrated on the well-known MeanShift tracker. The results over a heterogeneous dataset show that the proposed approach outperforms the related state-of-the-art methods in presence of tracking challenges such as occlusions, illumination and scale changes, and clutter. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 3 |
| 2012 | Bounded non-deterministic planning for multimedia adaptation
Fernando López Hernández 0001, Dietmar Jannach, José María Martínez Sanchez, Christian Timmerer, Narciso García, Hermann Hellwagner |
Appl. Intell. | 3 |
| 2012 | A semantic-based probabilistic approach for real-time video event recognition
Juan C. SanMiguel, José María Martínez Sanchez |
Comput. Vis. Image Underst. | 2 |
| 2012 | On collaborative people detection and tracking in complex scenarios
Álvaro García-Martín, José María Martínez Sanchez |
Image Vis. Comput. | 2 |
| 2012 | On-line video abstract generation of multimedia news
Víctor Valdés, José María Martínez Sanchez |
Multim. Tools Appl. | 2 |
| 2012 | A corpus for benchmarking of people detection algorithms
Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós |
Pattern Recognit. Lett. | 2 |
| 2012 | Adaptive Online Performance Evaluation of Video TrackersabstractWe propose an adaptive framework to estimate the quality of video tracking algorithms without ground-truth data. The framework is divided into two main stages, namely, the estimation of the tracker condition to identify temporal segments during which a target is lost and the measurement of the quality of the estimated track when the tracker is successful. A key novelty of the proposed framework is the capability of evaluating video trackers with multiple failures and recoveries over long sequences. Successful tracking is identified by analyzing the uncertainty of the tracker, whereas track recovery from errors is determined based on the time-reversibility constraint. The proposed approach is demonstrated on a particle filter tracker over a heterogeneous data set. Experimental results show the effectiveness and robustness of the proposed framework that improves state-of-the-art approaches in the presence of tracking challenges such as occlusions, illumination changes, and clutter and on sequences containing multiple tracking errors and recoveries. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
IEEE Trans. Image Process. | 3 |
| 2012 | Scalable Comic-Like Video Summaries and Layout DisturbanceabstractThis paper describes an efficient system for scalable video summarization that exploits comic-like summaries and multi-scale representations to facilitate interactivity and balance between content coverage and compactness. Due to the layout disturbance induced by the transitions between scales, a new heuristic algorithm is proposed to restrict changes to bounded summary segments. Conducted user evaluations show that the proposed methodology improves usability while keeping the summaries compact and informative. Luis Herranz, Janko Calic, José María Martínez Sanchez, Marta Mrak |
IEEE Trans. Multim. | 3 |
| 2012 | Automatic evaluation of video summariesabstractThis article describes a method for the automatic evaluation of video summaries based on the training of individual predictors for different quality measures from the TRECVid 2008 BBC Rushes Summarization Task. The obtained results demonstrate that, with a large set of evaluation data, it is possible to train fully automatic evaluation systems based on visual features automatically extracted from the summaries. The proposed approach will enable faster and easier estimation of the results of newly developed abstraction algorithms and the study of which summary characteristics influence their perceived quality. Víctor Valdés, José María Martínez Sanchez |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | Discrimination of abandoned and stolen object based on active contoursabstractIn this paper we propose an approach based on active contours to discriminate previously detected static foreground regions between abandoned and stolen. Firstly, the static foreground object contour is extracted. Then, an active contour adjustment is performed on the current and the background frames. Finally, similarities between the initial contour and the two adjustments are studied to decide whether the object is abandoned or stolen. Three different methods have been tested for this adjustment. Experimental results over a heterogeneous dataset show that the proposed method outperforms state-of-art approaches and provides a robust solution against non-accurate data (i.e., foreground static objects wrongly segmented) that is common in complex scenarios. Luis Caro Campos, Juan C. SanMiguel, José María Martínez Sanchez |
AVSS | 3 |
| 2011 | Improving the efficiency and accuracy of visual attentionabstractVisual attention is the cognitive process of selectively focusing on certain areas of a visual scene while ignoring the others. It is a desirable capability for intelligent video surveillance systems, as it allows them to control the aim of mobile cameras or to selectively process the most relevant parts of the captured images. This paper proposes an adaptation of a well-known biologically-inspired visual attention model in order to increase its computational efficiency without sacrificing its accuracy, and shows that the latter can be further improved through a supervised training stage that fine-tunes the model to the particular application scope in which the system is being utilized. Experimental results and comparisons with previous visual attention techniques are shown and discussed. Víctor Fernández-Carbajales Canete, Miguel Ángel García, José María Martínez Sanchez |
AVSS | 3 |
| 2011 | People detection based on appearance and motion modelsabstractThe main contribution of this paper is a new people detection algorithm based on motion information. The algorithm builds a people motion model based on the Implicit Shape Model (ISM) Framework and the MoSIFT descriptor. We also propose a detection system that integrates appearance, motion and tracking information. Experimental results over sequences extracted from the TRECVID dataset show that our new people motion detector produces results comparable to the state of the art and that the proposed multimodal fusion system improves the obtained results combining the three information sources. Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez |
AVSS | 3 |
| 2011 | A model for preference-driven multimedia adaptation decision-making in the MPEG-21 framework
Fernando López Hernández 0001, José María Martínez Sanchez, Narciso García |
Multim. Tools Appl. | 2 |
| 2010 | Robust Real Time Moving People Detection in Surveillance ScenariosabstractIn this paper an improved real time algorithm for detecting pedestrians in surveillance video is proposed. The algorithm is based on people appearance and defines a person model as the union of four models of body parts. Firstly, motion segmentation is performed to detect moving pixels. Then, moving regions are extracted and tracked. Finally, the detected moving objects are classified as human or nonhuman objects. In order to test and validate the algorithm, we have developed a dataset containing annotated surveillance sequences of different complexity levels focused on the pedestrians detection. Experimental results over this dataset show that our approach performs considerably well at real time and even better than other real and non-real time approaches from the state of art. Álvaro García-Martín, José María Martínez Sanchez |
AVSS | 2 |
| 2010 | On the Evaluation of Background Subtraction Algorithms without Ground-TruthabstractIn video-surveillance systems, the moving object segmentation stage (commonly based on background subtraction) has to deal with several issues like noise, shadows and multimodal backgrounds. Hence, its failure is inevitable and its automatic evaluation is a desirable requirement for online analysis. In this paper, we propose a hierarchy of existing performance measures not-based on ground-truth for video object segmentation. Then, four measures based on color and motion are selected and examined in detail with different segmentation algorithms and standard test sequences for video object segmentation. Experimental results show that color-based measures perform better than motion-based measures and background multimodality heavily reduces the accuracy of all obtained evaluation results. Juan C. SanMiguel, José María Martínez Sanchez |
AVSS | 2 |
| 2010 | Stationary foreground detection using background subtraction and temporal difference in video surveillanceabstractIn this paper we describe a new algorithm focused on obtaining stationary foreground regions, which is useful for applications like the detection of abandoned/stolen objects and parked vehicles. Firstly, a sub-sampling scheme based on background subtraction techniques is implemented to obtain stationary foreground regions. Secondly, some modifications are introduced on this base algorithm with the purpose of reducing the amount of stationary foreground detected. Finally, we evaluate the proposed algorithm and compare results with the base algorithm using video surveillance sequences from PETS 2006, PETS 2007 and I-LIDS for AVSS 2007 datasets. Experimental results show that the proposed algorithm increases the detection of stationary foreground regions as compared to the base algorithm. Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez |
ICIP | 3 |
| 2010 | Evaluation of on-line quality estimators for object trackingabstractFailure of tracking algorithms is inevitable in real and on-line tracking systems. The online estimation of the track quality is therefore desirable for detecting tracking failures while the algorithm is operating. In this paper, we propose a taxonomy and present a comparative evaluation of online quality estimators for video object tracking. The measures are compared over a heterogeneous video dataset with standard sequences. Among other results, the experiments show, that the Observation Likelihood (OL) measure is an appropriate quality measure for overall tracking performance evaluation, while the Template Inverse Matching (TIM) measure is appropriate to detect the start and the end instants of tracking failures. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 3 |
| 2010 | On the Advantages of the Use of Bitstream Extraction for Video Summary Generation
Luis Herranz, José María Martínez Sanchez |
MMM | 2 |
| 2010 | A framework for video abstraction systems analysis and modelling from an operational point of view
Víctor Valdés, José María Martínez Sanchez |
Multim. Tools Appl. | 2 |
| 2010 | A Framework for Scalable Summarization of VideoabstractVideo summaries provide compact representations of video sequences, with the length of the summary playing an important role, trading off the amount of information conveyed and how fast it can be visualized. This letter proposes scalable summarization as a method to easily adapt the summary to a suitable length, according to the requirements in each case, along with a suitable framework. The analysis algorithm uses a novel iterative ranking procedure in which each summary is the result of the extension of the previous one, balancing information coverage and visual pleasantness. The result of the algorithm is a ranked list, a scalable representation of the sequence useful for summarization. The summary is then efficiently generated from the bitstream of the sequence using bitstream extraction. Luis Herranz, José María Martínez Sanchez |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Comparative Evaluation of Stationary Foreground Object Detection Algorithms Based on Background Subtraction TechniquesabstractIn several video surveillance applications, such as the detection of abandoned/stolen objects or parked vehicles,the detection of stationary foreground objects is a critical task. In the literature, many algorithms have been proposed that deal with the detection of stationary foreground objects, the majority of them based on background subtraction techniques. In this paper we discuss various stationary object detection approaches comparing them in typical surveillance scenarios (extracted from standard datasets). Firstly, the existing approaches based on background-subtraction are organized into categories. Then, a representative technique of each category is selected and described. Finally, a comparative evaluation using objective and subjective criteria is performed on video surveillance sequences selected from the PETS 2006 and i-LIDS for AVSS 2007 datasets, analyzing the advantages and drawbacks of each selected approach. Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez |
AVSS | 3 |
| 2009 | An Ontology for Event Detection and its Application in Surveillance VideoabstractIn this paper, we propose an ontology for representing the prior knowledge related to video event analysis. It is composed of two types of knowledge related to the application domain and the analysis system. Domain knowledge involves all the high level semantic concepts in the context of each examined domain (objects, events, context...) whilst system knowledge involves the capabilities of the analysis system (algorithms, reactions to events...). The proposed ontology has been structured in two parts: the basic ontology (composed of the basic concepts and their specializations) and the domain-specific extensions. Additionally, a video analysis framework based on the proposed ontology is defined for the analysis of different application domains showing the potential use of the proposed ontology. In order to show the real applicability of the proposed ontology, it is specialized for the underground video-surveillance domain showing some results that demonstrate the usability and effectiveness of the proposed ontology. Juan C. SanMiguel, José María Martínez Sanchez, Álvaro García-Martín |
AVSS | 2 |
| 2009 | Shadow detection in video surveillance by maximizing agreement between independent detectorsabstractThis paper starts from the idea of automatically choosing the appropriate thresholds for a shadow detection algorithm. It is based on the maximization of the agreement between two independent shadow detectors without training data. Firstly, this shadow detection algorithm is described and then, it is adapted to analyze video surveillance sequences. Some modifications are introduced to increase its robustness in generic surveillance scenarios and to reduce its overall computational cost (critical in some video surveillance applications). Experimental results show that the proposed modifications increase the detection reliability as compared to some previous shadow detection algorithms and performs considerably well across a variety of multiple surveillance scenarios. Juan C. SanMiguel, José María Martínez Sanchez |
ICIP | 2 |
| 2009 | An efficient summarization algorithm based on clustering and bitstream extractionabstractVisualizing video is a very time consuming task. Video summaries are compact representations and very useful in video management systems. However, in some applications, in which the summary must be generated on demand, it usually implies a high delay for the user. Video summarization is a very demanding task in terms of processing resources, and it is very difficult to obtain a summary of a large video with low delay. In this paper a flexible summarization framework is presented which can create storyboards and video skims using bitstream extraction to generate the summary after an appropriate analysis algorithm. This technique is very efficient and achieves very low delay while keeping a reasonable quality of video summaries. Luis Herranz, José María Martínez Sanchez |
ICME | 2 |
| 2009 | An integrated approach to summarization and adaptation using H.264/MPEG-4 SVC
Luis Herranz, José María Martínez Sanchez |
Signal Process. Image Commun. | 2 |
| 2009 | On the use of hierarchical prediction structures for efficient summary generation of H.264/AVC bitstreams
Luis Herranz, José María Martínez Sanchez |
Signal Process. Image Commun. | 2 |
| 2009 | Visual Tools for ROI Montage in an Image2Video ApplicationabstractThis letter presents an image to video adaptation system that transmodes images in order to be viewed on small displays without a significant loss of information. The followed approach is to automate the process of manual browsing and zooming through an image by simulating the movement of a virtual camera. The presented system proposes two modules: one for analysis (ROI Extraction: for defining the regions of interest (ROIs) in the image) and the second for transmoding (ROIs2Video: for creating a video from any predefined list of relevance annotated ROIs). The separation of the image analysis and transmoding steps allows to use the ROIs2Video module in different application domains, as the transmoding will only be driven by the available relevance values associated to the ROIs. Results are presented for automatic and manual ROI generation modules. User evaluation results show that the developed system offers a visually attractive solution. Fernando Barreiro-Megino, José María Martínez Sanchez, Víctor Valdés |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Robust Unattended and Stolen Object Detection by Fusing Simple AlgorithmsabstractIn this paper a new approach for detecting unattended or stolen objects in surveillance video is proposed. It is based on the fusion of evidence provided by three simple detectors. As a first step, the moving regions in the scene are detected and tracked. Then, these regions are classified as static or dynamic objects and human or nonhuman objects. Finally, objects detected as static and nonhuman are analyzed with each detector. Data from these detectors are fused together to select the best detection hypotheses. Experimental results show that the fusion-based approach increases the detection reliability as compared to the detectors and performs considerably well across a variety of multiple scenarios operating at realtime. Juan C. SanMiguel, José María Martínez Sanchez |
AVSS | 2 |
| 2008 | Generation of scalable summaries based on iterative GoP rankingabstractVideo skims and image storyboards are two widely used abstractions for representing the essence of a video sequence, crucial for effective browsing and retrieval applications. In this paper we propose a flexible approach to find a reasonable balance between the semantic coverage and naturalness of the generated summaries, targeting a wide range of summarization ratios. The result of the algorithm is a scalable representation of the information required for summarization (a ranked list of GoPs) with a number of advantages in terms of efficient generation and potential applications. Luis Herranz, José María Martínez Sanchez |
ICIP | 2 |
| 2008 | A ground truth for motion-based video-object segmentationabstractThis paper describes the design procedure followed to generate a ground truth for the evaluation of motion-based algorithms for video-object segmentation. A thorough review and classification of the critical factors that affect the behavior of segmentation algorithms results in a set of video scripts which have then been filmed. Foreground objects have been recorded in a chroma studio, in order to automatically obtain pixel-level high quality segmentation masks for each generated sequence. The resulting corpus (segmentation ground-.truth plus filmed sequences mounted over different backgrounds) is available for research purposes under a license agreement. Fabricio Tiburzi, Marcos Escudero-Viñolo, Jesús Bescós, José María Martínez Sanchez |
ICIP | 4 |
| 2007 | On the effect of motion segmentation techniques in description based adaptive video transmissionabstractThis paper presents the results of analysing the effect of different motion segmentation techniques in a system that transmits the information captured by a static surveillance camera in an adaptative way based on the on-line generation of descriptions and their descriptions at different levels of detail. The video sequences are analyzed to detect the regions of activity (motion analysis) and to differentiate them from the background, and the corresponding descriptions (mainly MPEG-7 moving regions) are generated together with the textures of the moving regions and the associated background image. Depending on the available bandwidth, different levels of transmission are specified, ranging from just sending the descriptions generated to a transmission with all the associated images corresponding to the moving objects and background. We study the effect of three motion segmentation algorithms in several aspects such as accurate segmentation, size of the descriptions generated, computational efficiency and reconstructed data quality. Juan C. SanMiguel, José María Martínez Sanchez |
AVSS | 2 |
| 2007 | TV-Anytime Phase 1 and MPEG-7abstractAbstract Personal video recorders have the capability to change the media delivery industry fundamentally, and in this context, many believe the real international age of personal digital recorders (PDRs) will arrive with the use of “open” systems. The world reached an important milestone with the publication of the TV‐Anytime Phase 1 specifications for unidirectional broadcast and metadata services over bidirectional networks. TV‐Anytime is a worldwide prestandardization body; this article gives an overview of the main features of TV‐Anytime's metadata specification and its relationship to MPEG‐7 and provides insight into ways two organizations concerned with standards work together. Phase 2 has since been completed and TV‐Anytime has been adopted by various international standards organizations dealing with telecommunications and is now in the implementation phase. Jean-Pierre Evain, José María Martínez Sanchez |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | MPEG-7 tools for Universal Multimedia AccessabstractAbstract Universal Multimedia Access (UMA) deals with seamless access to once‐only‐created content via any kind of terminal and any kind of network connectivity, which implies that the content should be adapted in order to fit a variety of terminal and network characteristics, as well as user preferences. The MPEG‐7 standard offers some support for UMA within its section on Multimedia Description Schemes (MDS). Within the standard, several groups of tools serve this purpose. For instance, the Navigation and Access Tools provide some Description Schemes that allow the description of adapted content variations and summaries and allow for preprocessed content versions. Some support is also found in the Content Metadata Tools (Media and Usage Tools), for real‐time ease in creation of online content versions and in limited support for session description, which is completed in MPEG‐21. José María Martínez Sanchez |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | Content-driven adaptation of on-line video
Jesús Bescós, José María Martínez Sanchez, Luis Herranz, Fabricio Tiburzi |
Signal Process. Image Commun. | 2 |
| 2006 | Video summaries generation and access via personalized delivery of multimedia presentations adapted to service and terminalabstractThis article is centered on describing the provision of universal multimedia access services for video summaries via multimedia presentations that allow the integration of multimedia messaging service (MMS)-enabled terminals in the framework of the deferred time environment (DTE) of the DYMAS system. The system uses the framework of MPEG-7 and MPEG-21 to provide description metadata of the multimedia content and the usage context (including terminal, network capabilities, and user preferences), respectively. These descriptions are the base for the main functionalities of the complete system that provides personalized access to content (filtering by user preferences or via querying) that is first adapted to a multimedia presentation (generating a video summary that is represented via keyframes with synchronized audio clips) by the Presentation Builder and afterward adapted to the current service and terminal by the Adaptation Engine. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 785–800, 2006. Marta Padilla, José María Martínez Sanchez, Luis Herranz |
Int. J. Intell. Syst. | 2 |
| 2005 | A unified model for techniques on video-shot transition detectionabstractA first step required to allow video indexing and retrieval of visual data is to perform a temporal segmentation, that is, to find the location of camera-shot transitions, which can be either abrupt (i.e., cuts) or gradual (e.g., fades, dissolves, wipes). After a critical review of most approaches seeking to solve this problem, we propose a unified detection model (both for abrupt and all types of gradual transitions), as well as an implementation whose results improve upon those of all the inspected reports. The innovation of the approach presented here is centered on mapping the space of inter-frame distances onto a new space of decision better suited to achieving a sequence-independent thresholding. This mapping aims to consider frame ordering information within the thresholding process; it is based on the parametric modeling of the patterns that transitions generate on the distances' output. As opposed to most reviewed works, our results are detailed over a large and representative sample of more than 1500 cuts and 250 gradual transitions, which make up a significant part (200 min) of the MPEG-7 testing material; this ensures a high degree of confidence in the validity of our approach. Jesús Bescós, Guillermo Cisneros, José María Martínez Sanchez, José Manuel Menéndez, Julián Cabrera |
IEEE Trans. Multim. | 3 |
| 2003 | MISS: A Generic Model for MetaInformation SubSystems
José María Martínez Sanchez, Julián Cabrera, Jesús Bescós, Guillermo Cisneros, José Manuel Menéndez |
Multim. Tools Appl. | 1 |
| 2002 | Towards universal access to content using MPEG-7abstractThis paper presents a system providing functionalities for cataloging multimedia content using MPEG-7 and accessing to content and descriptions. The cataloging application indexes content using MPEG-7 and creates annotated variations in order to have the capability of offering media content to a large amount of different terminals and through different access networks. The created multimedia database, both descriptions and content (original sources and variations), is used by, currently, two applications: a searching application, which allows a user selecting specific tags to find the desired media; and a filtering application for transparent access to the content database using profiles. These profiles can be selected from a profiles database or created by the user specifying content preferences, and network and terminal parameters. José María Martínez Sanchez, César González, Oscar Fernández, Clara García, Jaime de Ramón |
ACM Multimedia | 1 |
| 2002 | Authoring 744: first resultsabstractThis paper presents the first results of the Authoring744 research initiative, which uses MPEG-7 to synthesize MPEG-4 content. The objective is to use MPEG-7 content descriptions to synthesize content, instead of creating descriptions by analyzing existing content. The output uses MPEG-4 XMT as the representation format, which is further used to create an MPEG-4 binary format, which can in turn be played. José María Martínez Sanchez, Luis F. Rubio, Francisco Morán |
ACM Multimedia | 1 |
| 2001 | Robust digital image watermarking using DWT, DFT and quality based averageabstractThis paper presents a digital image watermarking system that complements a wavelet based insertion module, with a resynchronization module, and a method for selecting the watermark using an estimated-quality-based average. The proposed system has been tested with attacks performed by Stirmark, obtaining results of robustness over 90%. Eduardo Fullea, José María Martínez Sanchez |
ACM Multimedia | 2 |
| 2000 | A Unified Approach to Gradual Shot Transition DetectionabstractNowadays, content-based retrieval of video material is based on the availability of meta-data linked to it. Current approaches to automatically extract these data start from a temporal segmentation of the audio-visual material, that is, a location of the camera shot transitions. Although abrupt transition detection is a problem almost solved, success on gradual transition detection is still very low. We present a detector for all types of gradual transitions (chromatic, geometric, and mixed), based on in depth modelling of the patterns that these transitions generate on a specific frame distance. Results are presented over a representative sample of more than 250 gradual transitions, which belong to a significant part (200') of the MPEG-7 testing material, achieving a recall (percentage of effects correctly detected) and a precision (percentage of non-false positives in the set of detection) notably higher than the ones reported so far. Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros, Julián Cabrera, José María Martínez Sanchez |
ICIP | 5 |
| 1999 | Generic meta data browsing system for multimedia document retrievalabstractThis paper presents the architecture of a multimedia information-indexing subsystem, the Meta-information subsystem. This subsystem makes use of meta data, that is, content-based information extracted from multimedia materials and other associated information. It also provides new mechanisms to store, query and browse meta data, and manages their relationship with the multimedia materials stored in the system. Julián Cabrera, José María Martínez Sanchez, Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros |
MMSP | 2 |
| 1998 | Developing Multimedia Applications: System Modeling and Implementation
José María Martínez Sanchez, Jesús Bescós, Guillermo Cisneros |
Multim. Tools Appl. | 1 |