Juan C. SanMiguel

dblp:72/8109 · also Juan Carlos San Miguel, Juan Carlos SanMiguel Avedillo · DBLP profile ↗
← Back
43ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0002-4999-2851ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 11 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
Diego Ortego, Marlon Rodríguez, Mario Almagro-Cádiz, Kunal Dahiya, David Jiménez, Juan C. SanMiguel
AAAI6
2026 Soft-Labelling for Budget-Constrained Semantic Segmentation: Bringing Coherence to label Down-Sampling
abstract
In semantic segmentation, training data down-sampling is commonly performed due to resource limitations, the need to adapt image size to the model input, or to improve data augmentation. This down-sampling typically employs different strategies for the image data and the annotated labels. Such discrepancy leads to mismatches between the down-sampled colour and ground-truth label images. Hence, the training performance significantly decreases as the down-sampling factor increases. In this paper, we bring together the down-sampling strategies for the image data and the training labels. To that aim, we propose a novel framework for label down-sampling via soft-labelling that better conserves label information after down-sampling, thereby, fully aligning soft-labels with image data to keep the distribution of the sampled pixels for down-sampling. This proposal also produces reliable annotations for under-represented semantic classes. Altogether, it allows training competitive models at lower resolutions. Experiments show that our proposal outperforms other down-sampling strategies. Moreover, state-of-the-art performance is achieved for reference benchmarks, but employing significantly fewer computational resources than foremost methods. This proposal enables competitive research for semantic segmentation under resource constraints.
Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Media3
2025 A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-Driven Diffusion Models
abstract
Pedestrian Attribute Recognition (PAR) is a challenging task as models are required to generalize across numerous attributes in real-world data. Traditional approaches focus on complex methods, yet recognition performance is often constrained by training dataset limitations, particularly the under-representation of certain attributes. In this paper, we propose a data-centric approach to improve PAR by synthetic data augmentation guided by textual descriptions. First, we define a protocol to identify weakly recognized attributes across multiple datasets. Second, we propose a prompt-driven pipeline that leverages diffusion models to generate synthetic pedestrian images while preserving the consistency of PAR datasets. Finally, we derive a strategy to seamlessly incorporate synthetic samples into training data, which considers prompt-based annotation rules and modifies the loss function. Results on popular PAR datasets demonstrate that our approach not only boosts recognition of underrepresented attributes but also improves overall model performance beyond the targeted attributes. Notably, this approach strengthens zero-shot generalization without requiring architectural changes of the model, presenting an efficient and scalable solution to improve the recognition of attributes of pedestrians in the real world.
Sawaiz A. Chaudhry, Juan C. SanMiguel, Álvaro García-Martín, Pablo Ayuso-Albizu, Pablo Carballeira
AVSS3
2025 Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion Models
abstract
Pedestrian Attribute Recognition (PAR) involves identifying various human attributes from images with applications in intelligent monitoring systems. The scarcity of large-scale annotated datasets hinders the generalization of PAR models, specially in complex scenarios involving occlusions, varying poses, and diverse environments. Recent advances in diffusion models have shown promise for generating diverse and realistic synthetic images, allowing to expand the size and variability of training data. However, the potential of diffusion-based data expansion for generating PAR-like images remains underexplored. Such expansion may enhance the robustness and adaptability of PAR models in real-world scenarios. This paper investigates the effectiveness of diffusion models in generating synthetic pedestrian images tailored to PAR tasks. We identify key parameters of img2img diffusion-based data expansion —including text prompts, image properties, and the latest enhancements in diffusion-based data augmentation—and examine their impact on the quality of generated images for PAR. Furthermore, we employ the best-performing expansion approach to generate synthetic images for training PAR models, by enriching the zero-shot datasets. Experimental results show that prompt alignment and image properties are critical factors in image generation, with optimal selection leading to a 4.5% improvement in PAR recognition performance.
Pablo Ayuso-Albizu, Juan C. SanMiguel, Pablo Carballeira
AVSS2
2025 Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
abstract
Abstract Unsupervised domain adaptation (UDA) offers a compelling solution to bridge the gap between labelled synthetic data and unlabelled real‐world data for training semantic segmentation models, given the high costs associated with manual annotation. However, the visual differences between the synthetic and real images pose significant challenges to their practical applications. This work addresses these challenges through synthetic‐to‐real style transfer leveraging diffusion models. The authors’ proposal incorporates semantic controllers to guide the diffusion process and low‐rank adaptations (LoRAs) to ensure that style‐transferred images align with real‐world aesthetics while preserving semantic layout. Moreover, the authors introduce quality metrics to rank the utility of generated images, enabling the selective use of high‐quality images for training. To further enhance reliability, the authors propose a novel loss function that mitigates artefacts from the style transfer process by incorporating only pixels aligned with the original semantic labels. Experimental results demonstrate that the authors’ proposal outperforms selected state‐of‐the‐art methods for image generation and UDA training, achieving optimal performance even with a smaller set of high‐quality generated images. The authors’ code and models are available at http://www‐vpu.eps.uam.es/ControllingSem4UDA/ .
Henrietta Ridley, Roberto Alcover-Couso, Juan C. SanMiguel
IET Comput. Vis.3
2025 Gradient-based class weighting for unsupervised domain adaptation in dense prediction visual tasks
abstract
In unsupervised domain adaptation (UDA), where models are trained on source data (e.g., synthetic) and adapted to target data (e.g., real-world) without target annotations, addressing the challenge of significant class imbalance remains an open issue. Despite progress in bridging the domain gap, existing methods often experience performance degradation when confronted with highly imbalanced dense prediction visual tasks like semantic segmentation. This discrepancy becomes especially pronounced due to the lack of equivalent priors between the source and target domains, turning class imbalanced techniques used for other areas (e.g., image classification) ineffective in UDA scenarios. This paper proposes a class-imbalance mitigation strategy that incorporates class-weights into the UDA learning losses, with the novelty of estimating these weights dynamically through the gradients of the per-class losses, defining a Gradient-based class weighting (GBW) approach. The proposed GBW naturally increases the contribution of classes whose learning is hindered by highly-represented classes, and has the advantage of automatically adapting to training outcomes, avoiding explicit curricular learning patterns common in loss-weighing strategies. Extensive experimentation validates the effectiveness of GBW across architectures (Convolutional and Transformer), UDA strategies (adversarial, self-training and entropy minimization), tasks (semantic and panoptic segmentation), and datasets. Analysis shows that GBW consistently increases the recall of under-represented classes. • A novel class-imbalance algorithm that uses per-class gradients to assign effective class weights (GBW). • Class weights are computed through an optimization process that maximizes the decrease in the loss function. • The proposed method demonstrates significant and consistent performance improvements across various UDA methods. • The class weights provide a per-class complexity measure throughout training, establishing an automatic and adaptable curriculum.
Roberto Alcover-Couso, Marcos Escudero-Viñolo, Juan C. SanMiguel, Jesús Bescós
Pattern Recognit.3
2025 Per-class curriculum for Unsupervised Domain Adaptation in semantic segmentation
abstract
Abstract Accurate training of deep neural networks for semantic segmentation requires a large number of pixel-level annotations of real images, which are expensive to generate or not even available. In this context, Unsupervised Domain Adaptation (UDA) can transfer knowledge from unlimited synthetic annotations to unlabeled real images of a given domain. UDA methods are composed of an initial training stage with labeled synthetic data followed by a second stage for feature alignment between labeled synthetic and unlabeled real data. In this paper, we propose a novel approach for UDA focusing the initial training stage, which leads to increased performance after adaptation. We introduce a curriculum strategy where each semantic class is learned progressively. Thereby, better features are obtained for the second stage. This curriculum is based on: (1) a class-scoring function to determine the difficulty of each semantic class, (2) a strategy for incremental learning based on scoring and pacing functions that limits the required training time unlike standard curriculum-based training and (3) a training loss to operate at class level. We extensively evaluate our approach as the first stage of several state-of-the-art UDA methods for semantic segmentation. Our results demonstrate significant performance enhancements across all methods: improvements of up to 10% for entropy-based techniques and 8% for adversarial methods. These findings underscore the dependency of UDA on the accuracy of the initial training. The implementation is available at https://github.com/vpulab/PCCL .
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Pablo Carballeira
Vis. Comput.2
2025 Layer-wise model merging for unsupervised domain adaptation in segmentation tasks
abstract
Abstract Merging parameters of multiple models has resurfaced as an effective strategy to enhance task performance and robustness, but prior work is limited by the high costs of ensemble creation and inference. In this paper, we leverage the abundance of freely accessible trained models to introduce a cost-free approach to model merging. It focuses on a layer-wise integration of merged models, aiming to maintain the distinctiveness of the task-specific final layers while unifying the initial layers, which are primarily associated with feature extraction. This approach ensures parameter consistency across all layers, essential for boosting performance. Moreover, it facilitates seamless integration of knowledge, enabling effective merging of models from different datasets and tasks. Specifically, we investigate its applicability in unsupervised domain adaptation (UDA), an unexplored area for model merging, for semantic and panoptic segmentation. Experimental results demonstrate substantial UDA improvements without additional costs for merging same-architecture models from distinct datasets ( $$\uparrow 2.6\%$$ ↑ 2.6 % mIoU) and different-architecture models with a shared backbone ( $$\uparrow 6.8\%$$ ↑ 6.8 % mIoU). Furthermore, merging semantic and panoptic segmentation models increases mPQ by 7%. These findings are validated across a wide variety of UDA strategies, architectures and datasets. The code will be publicly available upon acceptance in the LWMM repository: http://www-vpu.eps.uam.es/LWMM/ .
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Jose M. Martínez
Vis. Comput.2
2024 Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion Models
abstract
Diffusion models represent a new paradigm in text-to-image generation. Beyond generating high-quality images from text prompts, models such as Stable Diffusion have been successfully extended to the joint generation of se-mantic segmentation pseudo-masks. However, current ex-tensions primarily rely on extracting attentions linked to prompt words used for image synthesis. This approach limits the generation of segmentation masks derived from word tokens not contained in the text prompt. In this work, we introduce Open- Vocabulary Attention Maps (OVAM)-a training-free method for text-to-image diffusion models that enables the generation of attention maps for any word. In addition, we propose a lightweight optimization process based on OVAM for finding tokens that generate accurate attention maps for an object class with a single annotation. We evaluate these tokens within existing state-of-the-art Stable Diffusion extensions. The best-performing model im-proves its mIoU from 52.1 to 86.6 for the synthetic images' pseudo-masks, demonstrating that our optimized tokens are an efficient way to improve the performance of existing methods without architectural changes or retraining. The implementation is available at github.com/vpulablovam.
Pablo Marcos-Manchón, Roberto Alcover-Couso, Juan C. SanMiguel, Jose M. Martínez
CVPR3
2023 Human skeletons and change detection for efficient violence detection in surveillance videos
abstract
In our constantly monitored world, surveillance cameras play a crucial role in curbing crime and violence in public spaces by serving as a deterrent. To enhance their effectiveness, there is a growing need for automated tools that can detect crimes in real time. In this paper, we propose a novel deep learning architecture that accurately and efficiently detects violent crimes in surveillance videos. We rely on what we believe are the most essential pieces of information to detect violence, namely: human bodies and their interaction. To this end, we employ human pose extractors and change detectors as the input of our proposal. Subsequently, we combine them using a novel method, which relies on additions instead of multiplications to guarantee the transmission of information even when one of the inputs provides a zero-valued signal; outperforming other combination alternatives of the literature. Finally, to account for both spatial and temporal information, we use a convolutional alternative of the standard LSTM, the ConvLSTM. The experiments performed on several benchmark datasets demonstrate the efficacy and efficiency of our proposal, achieving state-of-the-art results with much fewer trainable parameters. We release the code to replicate the proposed architecture at https://github.com/atmguille/Violence-Detection-With-Human-Skeletons.
Guillermo Garcia-Cobo, Juan C. SanMiguel
Comput. Vis. Image Underst.2
2023 On exploring weakly supervised domain adaptation strategies for semantic segmentation using synthetic data
abstract
Abstract Pixel-wise image segmentation is key for many Computer Vision applications. The training of deep neural networks for this task has expensive pixel-level annotation requirements, thus, motivating a growing interest on synthetic data to provide unlimited data and its annotations. In this paper, we focus on the generation and application of synthetic data as representative training corpuses for semantic segmentation of urban scenes. First, we propose a synthetic data generation protocol, which identifies key features affecting performance and provides datasets with variable complexity. Second, we adapt two popular weakly supervised domain adaptation approaches (combined training, fine-tuning) to employ synthetic and real data. Moreover, we analyze several backbone models, real/synthetic datasets and their proportions when combined. Third, we propose a new curriculum learning strategy to employ several synthetic and real datasets. Our major findings suggest the high performance impact of pace and order of synthetic and real data presentation, achieving state of the art results for well-known models. The results by training with the proposed dataset outperform popular alternatives, thus demonstrating the effectiveness of the proposed protocol. Our code and dataset are available at http://www-vpu.eps.uam.es/publications/WSDA_semantic/
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Álvaro García-Martín
Multim. Tools Appl.2
2023 Attention-Based Knowledge Distillation in Scene Recognition: The Impact of a DCT-Driven Loss
abstract
Knowledge Distillation (KD) is a strategy for the definition of a set of transferability gangways to improve the efficiency of Convolutional Neural Networks. Feature-based Knowledge Distillation is a subfield of KD that relies on intermediate network representations, either unaltered or depth-reduced via maximum activation maps, as the source knowledge. In this paper, we propose and analyze the use of a 2D frequency transform of the activation maps before transferring them. We pose that—by using global image cues rather than pixel estimates, this strategy enhances knowledge transferability in tasks such as scene recognition, defined by strong spatial and contextual relationships between multiple and varied concepts. To validate the proposed method, an extensive evaluation of the state of the art in scene recognition is presented. Experimental results provide strong evidence that the proposed strategy enables the student network to better focus on the relevant image areas learnt by the teacher network, hence leading to better descriptive features and higher transferred performance than every other state-of-the-art alternative. We publicly release the training and evaluation framework used in this paper athttps://www-vpu.eps.uam.es/publications/DCTBasedKDForSceneRecognition.
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Juan C. SanMiguel
IEEE Trans. Circuits Syst. Video Technol.4
2023 Graph Neural Networks for Cross-Camera Data Association
abstract
Cross-camera image data association is essential for many multi-camera computer vision tasks, such as multi-camera pedestrian detection, multi-camera multi-target tracking, 3D pose estimation, etc. This association task is typically modeled as a bipartite graph matching problem and often solved by applying minimum-cost flow techniques, which may be computationally demanding for large data. Furthermore, cameras are usually treated by pairs, obtaining local solutions, rather than finding a global solution at once for all multiple cameras. Other key issue is that of the affinity function: the widespread usage of non-learnable pre-defined distances, such as the Euclidean and Cosine ones. This paper proposes an effective approach for cross-camera data-association focused on a global solution, instead of processing cameras by pairs. To avoid the usage of fixed distances and thresholds, we leverage the connectivity of Graph Neural Networks, previously unused in this scope, using a Message Passing Network to jointly learn features and similarity functions. We validate the proposal for pedestrian cross-camera association, showing results over the EPFL multi-camera pedestrian dataset. Our approach considerably outperforms the literature data association techniques, without requiring to be trained in the same scenario in which it is tested. Our code is available athttps://www-vpu.eps.uam.es/publications/gnn
Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Pablo Carballeira
IEEE Trans. Circuits Syst. Video Technol.2
2022 Detection-aware multi-object tracking evaluation
abstract
How would you fairly evaluate two multi-object tracking algorithms (i.e. trackers), each one employing a different object detector? Detectors keep improving, thus trackers can make less effort to estimate object states over time. Is it then fair to compare a new tracker employing a new detector with another tracker using an old detector? In this paper, we propose a novel performance measure, named Tracking Effort Measure (TEM), to evaluate trackers that use different detectors. TEM estimates the improvement that the tracker does with respect to its input data (i.e. detections) at frame level (intra-frame complexity) and sequence level (inter-frame complexity). We evaluate TEM over well-known datasets, four trackers and eight detection sets. Results show that, unlike conventional tracking evaluation measures, TEM can quantify the effort done by the tracker with a reduced correlation on the input detections. Its implementation will be made publicly available online.1.1https://github.com/vpulab/MOT-evaluation
Juan C. SanMiguel, Jorge Muñoz, Fabio Poiesi
AVSS1
2022 Online clustering-based multi-camera vehicle tracking in scenarios with overlapping FOVs
abstract
Abstract Multi-Target Multi-Camera (MTMC) vehicle tracking is an essential task of visual traffic monitoring, one of the main research fields of Intelligent Transportation Systems. Several offline approaches have been proposed to address this task; however, they are not compatible with real-world applications due to their high latency and post-processing requirements. This lack of suitable approaches motivates our proposal: A new low-latency online approach for MTMC tracking in scenarios with partially overlapping fields of view (FOVs), such as road intersections. Firstly, the proposed approach detects vehicles at each camera. Then, the detections are merged between cameras by applying cross-camera clustering based on appearance and location. Lastly, the clusters containing different detections of the same vehicle are temporally associated to compute the tracks on a frame-by-frame basis. The experiments show promising low-latency results while addressing real-world challenges such as the a priori unknown and time-varying number of targets and the continuous state estimation of them without performing any post-processing of the trajectories. Our code is available at http://www-vpu.eps.uam.es/publications/Online-MTMC-Tracking .
Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Marcos Escudero-Viñolo
Multim. Tools Appl.2
2019 On guiding video object segmentation
abstract
This paper presents a novel approach for segmenting moving objects in unconstrained environments using guided convolutional neural networks. This guiding process relies on foreground masks from independent algorithms (i.e. state-of-the-art algorithms) to implement an attention mechanism that incorporates the spatial location of foreground and background to compute their separated representations. Our approach initially extracts two kinds of features for each frame using colour and optical flow information. Such features are combined following a multiplicative scheme to benefit from their complementarity. These unified colour and motion features are later processed to obtain the separated foreground and background representations. Then, both independent representations are concatenated and decoded to perform foreground segmentation. Experiments conducted on the challenging DAVIS 2016 dataset demonstrate that our guided representations not only outperform non-guided, but also recent and top-performing video object segmentation algorithms.
Diego Ortego, Kevin McGuinness, Juan C. SanMiguel, Eric Arazo Sanchez, José María Martínez Sanchez, Noel E. O'Connor
CBMI3
2019 Hierarchical Improvement of Foreground Segmentation Masks in Background Subtraction
abstract
A plethora of algorithms have been defined for foreground segmentation, a fundamental stage for many computer vision applications. In this paper, we propose a post-processing framework to improve the foreground segmentation performance of background subtraction algorithms. We define a hierarchical framework for extending segmented foreground pixels to undetected foreground object areas and for removing erroneously segmented foreground. First, we create a motion-aware hierarchical image segmentation of each frame that prevents merging foreground and background image regions. Then, we estimate the quality of the foreground mask through the fitness of the binary regions in the mask and the hierarchy of segmented regions. Finally, the improved foreground mask is obtained as an optimal labeling by jointly exploiting foreground quality and spatial color relations in a pixel-wise fully connected conditional random field. Experiments are conducted over four large and heterogeneous data sets with varied challenges (CDNET2014, LASIESTA, SABS, and BMC) demonstrating the capability of the proposed framework to improve background subtraction results.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
IEEE Trans. Circuits Syst. Video Technol.2
2017 Adaptive people detection based on cross-correlation maximization
abstract
Applying people detectors to unseen data is challenging since patterns distributions may significantly differ from the ones of the training dataset. In this paper, we propose a framework to adapt people detectors during runtime classification. Such adaptation takes advantage of multiple detectors to identify their best configurations (i.e. detection thresholds) without requiring manually labeled ground truth. We maximize the mutual information of detectors by pair-wise correlating their outputs to obtain a set of hypotheses for the detection thresholds. These hypotheses are later combined by weighted voting to obtain a final decision for the detection threshold of each detector. The proposed approach does not require re-training detectors and uses standard people detector outputs, i.e., bounding boxes, therefore it can employ various types of detectors. The experimental results demonstrate that the proposed approach outperforms state-of-the-art detectors whose optimal configuration is learned from training data.
Álvaro García-Martín, Juan C. SanMiguel
ICIP2
2017 Efficient estimation of target detection quality
abstract
The capability of determining the quality of target detections is important for applications using smart cameras, such as autonomous robotics and surveillance. We propose to estimate the quality of target detections by integrating the target location uncertainty over polygonal domains, which represent the fields of view of the cameras. We define a framework based on numerical integration that easily accommodates multiple models for uncertainty and fields of view. We perform quadrature-based integration combined with importance sampling to provide accurate quality estimations while reducing the computational cost. The proposed method outperforms alternative approaches in terms of estimation accuracy and execution time. We validate the proposed approach with a recent distributed multi-camera multi-target tracker and improved it by considering realistic fields of view. Results demonstrate the effectiveness of the proposed method in decreasing the state estimation error.
Juan C. SanMiguel, Andrea Cavallaro
ICIP1
2017 Stand-alone quality estimation of background subtraction algorithms
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.2
2017 Multi-Tracker Partition Fusion
abstract
We propose a decision-level approach to fuse the output of multiple trackers based on their estimated individual performance. The proposed approach is composed of three main steps. First, we group trackers into clusters based on the spatiotemporal pair-wise correlation of their short-term trajectories. Then, we evaluate performance based on reverse-time analysis with an adaptive reference frame and define the cluster with trackers that appear to be successfully following the target as the on-target cluster. Finally, the state estimations produced by trackers in the on-target cluster are fused to obtain the target state. The proposed fusion approach uses standard tracker outputs and can therefore combine various types of trackers. We tested the proposed approach with several combinations of state-of-the-art trackers and also compared it with individual trackers and other fusion approaches. The results show that the proposed approach improves the state estimation accuracy under multiple tracking challenges.
ObaidUllah Khalid, Juan C. SanMiguel, Andrea Cavallaro
IEEE Trans. Circuits Syst. Video Technol.2
2017 Energy Consumption Models for Smart Camera Networks
abstract
Camera networks require heavy visual data processing and high-bandwidth communication. In this paper, we identify key factors underpinning the development of resource-aware algorithms and we propose a comprehensive energy consumption model for the resources employed by smart camera networks, which are composed of cameras that process data locally and collaborate with their neighbors. We account for the main parameters that influence consumption when sensing (frame size and frame rate), processing (dynamic frequency scaling and task load), and communication (output power and bandwidth) are considered. Next, we define an abstraction based on clock frequency and duty cycle that accounts for active, idle, and sleep operational states. We demonstrate the importance of the proposed model for a multicamera tracking task and show how one may significantly reduce consumption with only minor performance degradation when choosing to operate with an appropriately reduced hardware capacity. Moreover, we quantify the dependency on local computation resources and bandwidth availability. The proposed consumption model can be easily adjusted to account for new platforms, thus providing a valuable tool for the design of resource-aware algorithms and further research in resource-aware camera networks.
Juan C. SanMiguel, Andrea Cavallaro
IEEE Trans. Circuits Syst. Video Technol.1
2016 Rejection based multipath reconstruction for background estimation in SBMnet 2016 dataset
abstract
Background Estimation in video consists in extracting a foreground-free image from a set of training frames. In this paper, we overview a temporal-spatial block-level approach for background estimation in video and present their results in the SBMnet dataset. First, the employed approach uses a Temporal Analysis module to obtain a compact representation of the training data that is later clustered by a threshold-free technique to generate background candidates at each block location. Then, a Spatial Analysis module iteratively reconstructs the background using a multipath reconstruction guided by background smoothness constraints. The experimental results in the SBMnet dataset demonstrates the utility of the employed approach against stationary objects and its weaknesses when motion information is involved.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
ICPR2
2016 Rejection based multipath reconstruction for background estimation in video sequences with stationary objects
abstract
Background estimation in video consists in extracting a foreground-free image from a set of training frames. Moving and stationary objects may affect the background visibility, thus invalidating the assumption of many related literature where background is the temporal dominant data. In this paper, we present a temporal-spatial block-level approach for background estimation in video to cope with moving and stationary objects. First, a Temporal Analysis module obtains a compact representation of the training data by motion filtering and dimensionality reduction. Then, a threshold-free hierarchical clustering determines a set of candidates to represent the background for each spatial location (block). Second, a Spatial Analysis module iteratively reconstructs the background using these candidates. For each spatial location , multiple reconstruction hypotheses (paths) are explored to obtain its neighboring locations by enforcing inter-block similarities and intra-block homogeneity constraints in terms of color discontinuity, color dissimilarity and variability. The experimental results show that the proposed approach outperforms the related state-of-the-art over challenging video sequences in presence of moving and stationary objects.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.2
2015 Temporal validation of Particle Filters for video tracking
Juan C. SanMiguel, Andrea Cavallaro
Comput. Vis. Image Underst.1
2015 Long-Term Stationary Object Detection Based on Spatio-Temporal Change Detection
abstract
We present a block-wise approach to detect stationary objects based on spatio-temporal change detection. First, block candidates are extracted by filtering out consecutive blocks containing moving objects. Then, an online clustering approach groups similar blocks at each spatial location over time via statistical variation of pixel ratios. The stability changes are identified by analyzing the relationships between the most repeated clusters at regular sampling instants. Finally, stationary objects are detected as those stability changes that exceed an alarm time and have not been visualized before. Unlike previous approaches making use of Background Subtraction, the proposed approach does not require foreground segmentation and provides robustness to illumination changes, crowds and intermittent object motion. The experiments over an heterogeneous dataset demonstrate the ability of the proposed approach for short- and long-term operation while overcoming challenging issues.
Diego Ortego, Juan C. SanMiguel, José María Martínez Sanchez
IEEE Signal Process. Lett.2
2014 Consensus protocols for distributed tracking in wireless camera networks
Sandeep Katragadda, Juan C. SanMiguel, Andrea Cavallaro
FUSION2
2014 Multi-feature stationary foreground detection for crowded video-surveillance
abstract
We propose a novel approach for stationary foreground detection in crowds based on the spatio-temporal evolution of multiple features. A generic framework is presented to detect stationarity where history images model the spatio-temporal feature patterns. A feature is proposed based on structural information over each pixel neighborhood for dealing with shadows and illumination changes. A multifeature detector is composed by combining the history images of three features (namely, foreground, motion and structural information) to estimate the foreground stationarity over time, which is later thresholded to detect stationary regions. Experimental results over challenging video-surveillance sequences show the improvement of the proposed approach against related work as structural information reduces false detections, which are common in crowded places.
Diego Ortego, Juan C. SanMiguel
ICIP2
2013 Stationary foreground detection for video-surveillance based on foreground and motion history images
abstract
Stationary foreground detection is a common stage in many video-surveillance applications. In this paper, we propose an approach for stationary foreground detection in video based on the spatio-temporal variation of foreground and motion data. Foreground data are obtained by Background Subtraction to detect regions of interest. Motion data allows to filter out the moving regions and it is estimated using median filters over sliding windows. Spatiotemporal patterns of both data are computed through history images and the final detection is obtained using a two-threshold scheme that considers motion activity. Partial visibility of stationary foreground for short-time intervals is handled to increase robustness. The results over challenging video-surveillance sequences show an improvement of the proposed approach against the related work.
Diego Ortego, Juan C. SanMiguel
AVSS2
2013 A semantic-guided and self-configurable framework for video analysis
Juan C. SanMiguel, José María Martínez Sanchez
Mach. Vis. Appl.1
2013 Skin detection by dual maximization of detectors agreement for video monitoring
Juan C. SanMiguel, Sergio Suja
Pattern Recognit. Lett.1
2012 Standalone evaluation of deterministic video tracking
abstract
We present an approach for performance evaluation of deterministic video trackers without ground-truth data. The proposed approach detects if a tracker is correctly operating over time using two main steps. First, it transforms the output of the localization step into a distribution of the target state, which emulates a multi-hypothesis tracker. Then, the uncertainty of such distribution is estimated to determine the time instants when the tracker is stable. A time-reversed analysis is used to identify tracker recovery after unsuccessful operation. The proposed approach is demonstrated on the well-known MeanShift tracker. The results over a heterogeneous dataset show that the proposed approach outperforms the related state-of-the-art methods in presence of tracking challenges such as occlusions, illumination and scale changes, and clutter.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
ICIP1
2012 A semantic-based probabilistic approach for real-time video event recognition
Juan C. SanMiguel, José María Martínez Sanchez
Comput. Vis. Image Underst.1
2012 Adaptive Online Performance Evaluation of Video Trackers
abstract
We propose an adaptive framework to estimate the quality of video tracking algorithms without ground-truth data. The framework is divided into two main stages, namely, the estimation of the tracker condition to identify temporal segments during which a target is lost and the measurement of the quality of the estimated track when the tracker is successful. A key novelty of the proposed framework is the capability of evaluating video trackers with multiple failures and recoveries over long sequences. Successful tracking is identified by analyzing the uncertainty of the tracker, whereas track recovery from errors is determined based on the time-reversibility constraint. The proposed approach is demonstrated on a particle filter tracker over a heterogeneous data set. Experimental results show the effectiveness and robustness of the proposed framework that improves state-of-the-art approaches in the presence of tracking challenges such as occlusions, illumination changes, and clutter and on sequences containing multiple tracking errors and recoveries.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
IEEE Trans. Image Process.1
2011 Discrimination of abandoned and stolen object based on active contours
abstract
In this paper we propose an approach based on active contours to discriminate previously detected static foreground regions between abandoned and stolen. Firstly, the static foreground object contour is extracted. Then, an active contour adjustment is performed on the current and the background frames. Finally, similarities between the initial contour and the two adjustments are studied to decide whether the object is abandoned or stolen. Three different methods have been tested for this adjustment. Experimental results over a heterogeneous dataset show that the proposed method outperforms state-of-art approaches and provides a robust solution against non-accurate data (i.e., foreground static objects wrongly segmented) that is common in complex scenarios.
Luis Caro Campos, Juan C. SanMiguel, José María Martínez Sanchez
AVSS2
2010 On the Evaluation of Background Subtraction Algorithms without Ground-Truth
abstract
In video-surveillance systems, the moving object segmentation stage (commonly based on background subtraction) has to deal with several issues like noise, shadows and multimodal backgrounds. Hence, its failure is inevitable and its automatic evaluation is a desirable requirement for online analysis. In this paper, we propose a hierarchy of existing performance measures not-based on ground-truth for video object segmentation. Then, four measures based on color and motion are selected and examined in detail with different segmentation algorithms and standard test sequences for video object segmentation. Experimental results show that color-based measures perform better than motion-based measures and background multimodality heavily reduces the accuracy of all obtained evaluation results.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS1
2010 Stationary foreground detection using background subtraction and temporal difference in video surveillance
abstract
In this paper we describe a new algorithm focused on obtaining stationary foreground regions, which is useful for applications like the detection of abandoned/stolen objects and parked vehicles. Firstly, a sub-sampling scheme based on background subtraction techniques is implemented to obtain stationary foreground regions. Secondly, some modifications are introduced on this base algorithm with the purpose of reducing the amount of stationary foreground detected. Finally, we evaluate the proposed algorithm and compare results with the base algorithm using video surveillance sequences from PETS 2006, PETS 2007 and I-LIDS for AVSS 2007 datasets. Experimental results show that the proposed algorithm increases the detection of stationary foreground regions as compared to the base algorithm.
Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez
ICIP2
2010 Evaluation of on-line quality estimators for object tracking
abstract
Failure of tracking algorithms is inevitable in real and on-line tracking systems. The online estimation of the track quality is therefore desirable for detecting tracking failures while the algorithm is operating. In this paper, we propose a taxonomy and present a comparative evaluation of online quality estimators for video object tracking. The measures are compared over a heterogeneous video dataset with standard sequences. Among other results, the experiments show, that the Observation Likelihood (OL) measure is an appropriate quality measure for overall tracking performance evaluation, while the Template Inverse Matching (TIM) measure is appropriate to detect the start and the end instants of tracking failures.
Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez
ICIP1
2009 Comparative Evaluation of Stationary Foreground Object Detection Algorithms Based on Background Subtraction Techniques
abstract
In several video surveillance applications, such as the detection of abandoned/stolen objects or parked vehicles,the detection of stationary foreground objects is a critical task. In the literature, many algorithms have been proposed that deal with the detection of stationary foreground objects, the majority of them based on background subtraction techniques. In this paper we discuss various stationary object detection approaches comparing them in typical surveillance scenarios (extracted from standard datasets). Firstly, the existing approaches based on background-subtraction are organized into categories. Then, a representative technique of each category is selected and described. Finally, a comparative evaluation using objective and subjective criteria is performed on video surveillance sequences selected from the PETS 2006 and i-LIDS for AVSS 2007 datasets, analyzing the advantages and drawbacks of each selected approach.
Álvaro Bayona, Juan C. SanMiguel, José María Martínez Sanchez
AVSS2
2009 An Ontology for Event Detection and its Application in Surveillance Video
abstract
In this paper, we propose an ontology for representing the prior knowledge related to video event analysis. It is composed of two types of knowledge related to the application domain and the analysis system. Domain knowledge involves all the high level semantic concepts in the context of each examined domain (objects, events, context...) whilst system knowledge involves the capabilities of the analysis system (algorithms, reactions to events...). The proposed ontology has been structured in two parts: the basic ontology (composed of the basic concepts and their specializations) and the domain-specific extensions. Additionally, a video analysis framework based on the proposed ontology is defined for the analysis of different application domains showing the potential use of the proposed ontology. In order to show the real applicability of the proposed ontology, it is specialized for the underground video-surveillance domain showing some results that demonstrate the usability and effectiveness of the proposed ontology.
Juan C. SanMiguel, José María Martínez Sanchez, Álvaro García-Martín
AVSS1
2009 Shadow detection in video surveillance by maximizing agreement between independent detectors
abstract
This paper starts from the idea of automatically choosing the appropriate thresholds for a shadow detection algorithm. It is based on the maximization of the agreement between two independent shadow detectors without training data. Firstly, this shadow detection algorithm is described and then, it is adapted to analyze video surveillance sequences. Some modifications are introduced to increase its robustness in generic surveillance scenarios and to reduce its overall computational cost (critical in some video surveillance applications). Experimental results show that the proposed modifications increase the detection reliability as compared to some previous shadow detection algorithms and performs considerably well across a variety of multiple surveillance scenarios.
Juan C. SanMiguel, José María Martínez Sanchez
ICIP1
2008 Robust Unattended and Stolen Object Detection by Fusing Simple Algorithms
abstract
In this paper a new approach for detecting unattended or stolen objects in surveillance video is proposed. It is based on the fusion of evidence provided by three simple detectors. As a first step, the moving regions in the scene are detected and tracked. Then, these regions are classified as static or dynamic objects and human or nonhuman objects. Finally, objects detected as static and nonhuman are analyzed with each detector. Data from these detectors are fused together to select the best detection hypotheses. Experimental results show that the fusion-based approach increases the detection reliability as compared to the detectors and performs considerably well across a variety of multiple scenarios operating at realtime.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS1
2007 On the effect of motion segmentation techniques in description based adaptive video transmission
abstract
This paper presents the results of analysing the effect of different motion segmentation techniques in a system that transmits the information captured by a static surveillance camera in an adaptative way based on the on-line generation of descriptions and their descriptions at different levels of detail. The video sequences are analyzed to detect the regions of activity (motion analysis) and to differentiate them from the background, and the corresponding descriptions (mainly MPEG-7 moving regions) are generated together with the textures of the moving regions and the associated background image. Depending on the available bandwidth, different levels of transmission are specified, ranging from just sending the descriptions generated to a transmission with all the associated images corresponding to the moving objects and background. We study the effect of three motion segmentation algorithms in several aspects such as accurate segmentation, size of the descriptions generated, computational efficiency and reconstructed data quality.
Juan C. SanMiguel, José María Martínez Sanchez
AVSS1