EDBT 2026 Demo / reviewers in the wild / expert
Pablo Carballeira
dblp:59/3210
· DBLP profile ↗
18ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-7199-698XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-Driven Diffusion ModelsabstractPedestrian Attribute Recognition (PAR) is a challenging task as models are required to generalize across numerous attributes in real-world data. Traditional approaches focus on complex methods, yet recognition performance is often constrained by training dataset limitations, particularly the under-representation of certain attributes. In this paper, we propose a data-centric approach to improve PAR by synthetic data augmentation guided by textual descriptions. First, we define a protocol to identify weakly recognized attributes across multiple datasets. Second, we propose a prompt-driven pipeline that leverages diffusion models to generate synthetic pedestrian images while preserving the consistency of PAR datasets. Finally, we derive a strategy to seamlessly incorporate synthetic samples into training data, which considers prompt-based annotation rules and modifies the loss function. Results on popular PAR datasets demonstrate that our approach not only boosts recognition of underrepresented attributes but also improves overall model performance beyond the targeted attributes. Notably, this approach strengthens zero-shot generalization without requiring architectural changes of the model, presenting an efficient and scalable solution to improve the recognition of attributes of pedestrians in the real world. Sawaiz A. Chaudhry, Juan C. SanMiguel, Álvaro García-Martín, Pablo Ayuso-Albizu, Pablo Carballeira |
AVSS | 6 |
| 2025 | Enhancing Zero-Shot Pedestrian Attribute Recognition with Synthetic Data Generation: A Comparative Study with Image-To-Image Diffusion ModelsabstractPedestrian Attribute Recognition (PAR) involves identifying various human attributes from images with applications in intelligent monitoring systems. The scarcity of large-scale annotated datasets hinders the generalization of PAR models, specially in complex scenarios involving occlusions, varying poses, and diverse environments. Recent advances in diffusion models have shown promise for generating diverse and realistic synthetic images, allowing to expand the size and variability of training data. However, the potential of diffusion-based data expansion for generating PAR-like images remains underexplored. Such expansion may enhance the robustness and adaptability of PAR models in real-world scenarios. This paper investigates the effectiveness of diffusion models in generating synthetic pedestrian images tailored to PAR tasks. We identify key parameters of img2img diffusion-based data expansion —including text prompts, image properties, and the latest enhancements in diffusion-based data augmentation—and examine their impact on the quality of generated images for PAR. Furthermore, we employ the best-performing expansion approach to generate synthetic images for training PAR models, by enriching the zero-shot datasets. Experimental results show that prompt alignment and image properties are critical factors in image generation, with optimal selection leading to a 4.5% improvement in PAR recognition performance. Pablo Ayuso-Albizu, Juan C. SanMiguel, Pablo Carballeira |
AVSS | 3 |
| 2025 | Per-class curriculum for Unsupervised Domain Adaptation in semantic segmentationabstractAbstract Accurate training of deep neural networks for semantic segmentation requires a large number of pixel-level annotations of real images, which are expensive to generate or not even available. In this context, Unsupervised Domain Adaptation (UDA) can transfer knowledge from unlimited synthetic annotations to unlabeled real images of a given domain. UDA methods are composed of an initial training stage with labeled synthetic data followed by a second stage for feature alignment between labeled synthetic and unlabeled real data. In this paper, we propose a novel approach for UDA focusing the initial training stage, which leads to increased performance after adaptation. We introduce a curriculum strategy where each semantic class is learned progressively. Thereby, better features are obtained for the second stage. This curriculum is based on: (1) a class-scoring function to determine the difficulty of each semantic class, (2) a strategy for incremental learning based on scoring and pacing functions that limits the required training time unlike standard curriculum-based training and (3) a training loss to operate at class level. We extensively evaluate our approach as the first stage of several state-of-the-art UDA methods for semantic segmentation. Our results demonstrate significant performance enhancements across all methods: improvements of up to 10% for entropy-based techniques and 8% for adversarial methods. These findings underscore the dependency of UDA on the accuracy of the initial training. The implementation is available at https://github.com/vpulab/PCCL . Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Pablo Carballeira |
Vis. Comput. | 4 |
| 2024 | Synthmanticlidar: A Synthetic Dataset For Semantic Segmentation On Lidar ImagingabstractSemantic segmentation on LiDAR imaging is increasingly gaining attention, as it can provide useful knowledge for perception systems and potential for autonomous driving. However, collecting and labeling real LiDAR data is an expensive and time-consuming task. While datasets such as SemanticKITTI [1] have been manually collected and labeled, the introduction of simulation tools such as CARLA [2], has enabled the creation of synthetic datasets on demand. In this work, we present a modified CARLA simulator designed with LiDAR semantic segmentation in mind, with new classes, more consistent object labeling with their counter-parts from real datasets such as SemanticKITTI, and the possibility to adjust the object class distribution. Using this tool, we have generated SynthmanticLiDAR, a synthetic dataset for semantic segmentation on LiDAR imaging, designed to be similar to SemanticKITTI, and we evaluate its contribution to the training process of different semantic segmentation algorithms by using a naive transfer learning approach. Our results show that incorporating SynthmanticLiDAR into the training process improves the overall performance of tested algorithms, proving the usefulness of our dataset, and therefore, our adapted CARLA simulator. The dataset and simulator are available in https:// github.com/vpulab/SynthmanticLiDAR. Javier Montalvo, Pablo Carballeira, Álvaro García-Martín |
ICIP | 2 |
| 2024 | Improved transferability of self-supervised learning models through batch normalization finetuning
Kirill Sirotkin, Marcos Escudero-Viñolo, Pablo Carballeira, Álvaro García-Martín |
Appl. Intell. | 3 |
| 2023 | Graph Neural Networks for Cross-Camera Data AssociationabstractCross-camera image data association is essential for many multi-camera computer vision tasks, such as multi-camera pedestrian detection, multi-camera multi-target tracking, 3D pose estimation, etc. This association task is typically modeled as a bipartite graph matching problem and often solved by applying minimum-cost flow techniques, which may be computationally demanding for large data. Furthermore, cameras are usually treated by pairs, obtaining local solutions, rather than finding a global solution at once for all multiple cameras. Other key issue is that of the affinity function: the widespread usage of non-learnable pre-defined distances, such as the Euclidean and Cosine ones. This paper proposes an effective approach for cross-camera data-association focused on a global solution, instead of processing cameras by pairs. To avoid the usage of fixed distances and thresholds, we leverage the connectivity of Graph Neural Networks, previously unused in this scope, using a Message Passing Network to jointly learn features and similarity functions. We validate the proposal for pedestrian cross-camera association, showing results over the EPFL multi-camera pedestrian dataset. Our approach considerably outperforms the literature data association techniques, without requiring to be trained in the same scenario in which it is tested. Our code is available athttps://www-vpu.eps.uam.es/publications/gnn Elena Luna, Juan C. SanMiguel, José María Martínez Sanchez, Pablo Carballeira |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | A study on the distribution of social biases in self-supervised learning visual modelsabstractDeep neural networks are efficient at learning the data distribution if it is sufficiently sampled. However, they can be strongly biased by non-relevant factors implicitly incorporated in the training data. These include operational biases, such as ineffective or uneven data sampling, but also ethical concerns, as the social biases are implicitly present—even inadvertently, in the training data or explicitly defined in unfair training schedules. In tasks having impact on human processes, the learning of social biases may produce discriminatory, unethical and untrustworthy consequences. It is often assumed that social biases stem from supervised learning on labelled data, and thus, Self-Supervised Learning (SSL) wrongly appears as an efficient and bias-free solution, as it does not require labelled data. However, it was recently proven that a popular SSL method also incorporates biases. In this paper, we study the biases of a varied set of SSL visual models, trained using ImageNet data, using a method and dataset designed by psychological experts to measure social biases. We show that there is a correlation between the type of the SSL model and the number of biases that it incorporates. Furthermore, the results also suggest that this number does not strictly depend on the model's accuracy and changes throughout the network. Finally, we conclude that a careful SSL model selection process can reduce the number of social biases in the deployed model, whilst keeping high performance. The code is available at https://github.com/vpulab/SB-SSL. Kirill Sirotkin, Pablo Carballeira, Marcos Escudero-Viñolo |
CVPR | 2 |
| 2022 | Method for the automatic measurement of camera-calibration quality in a surround-view systemabstractOver the last decade, the automotive industry has introduced advanced driving assistance systems (ADAS) and automated driving (AD) features into roads to reduce fatality rates. One of these ADAS is the surround-view system, which provides an orthographic view of the vehicle by using at least four fish-eye lens cameras embedded in it. Small bumps or temperature changes may modify these cameras' relative poses leading to some geometrical mismatches between views in the top-view projection plane. In addition, terrain irregularities may misalign the orthographic view with the ground plane surface. Both problems can be solved by reestimating the relative poses of the cameras with respect to a single common point in the vehicle. This procedure, also known as recalibration, is offline performed in technical garages, or by online calibration mechanisms on engine start. However, it is a slow and cumbersome process. Research to date studies how to optimally recalibrate these cameras in an online manner, neglecting the practical aspects of when this procedure should be undertaken. Therefore, depending on the functionalities for which the embedded cameras are required, a compromise between using out-of-calibration cameras and the consequences derived from the recalibration process must be considered. This would prevent reestimating the cameras’ relative poses in situations where misalignment between adjacent cameras may not be noticeable. For this reason, a novel approach that measures the degree of calibration between cameras embedded in a vehicle is proposed. This method extracts relevant features from the predefined regions of interest of each camera by using the histogram of oriented gradients (HOG) descriptor. Then, features that belong to adjacent cameras are compared by employing the cosine similarity metric. The proposed method is evaluated on the open-source AD research simulator CARLA providing detailed analysis to objectively highlight the usefulness of this method in studying the degree of calibration of a camera array in a surround-view system. Martí Sánchez, Jon Ander Iñiguez de Gordoa, Marcos Nieto Doncel, Pablo Carballeira |
ICMV | 4 |
| 2022 | Semantic-driven multi-camera pedestrian detectionabstractAbstract In the current worldwide situation, pedestrian detection has reemerged as a pivotal tool for intelligent video-based systems aiming to solve tasks such as pedestrian tracking, social distancing monitoring or pedestrian mass counting. Pedestrian detection methods, even the top performing ones, are highly sensitive to occlusions among pedestrians, which dramatically degrades their performance in crowded scenarios. The generalization of multi-camera setups permits to better confront occlusions by combining information from different viewpoints. In this paper, we present a multi-camera approach to globally combine pedestrian detections leveraging automatically extracted scene context. Contrarily to the majority of the methods of the state-of-the-art, the proposed approach is scene-agnostic, not requiring a tailored adaptation to the target scenario–e.g., via fine-tuning. This noteworthy attribute does not require ad hoc training with labeled data, expediting the deployment of the proposed method in real-world situations. Context information, obtained via semantic segmentation, is used (1) to automatically generate a common area of interest for the scene and all the cameras, avoiding the usual need of manually defining it, and (2) to obtain detections for each camera by solving a global optimization problem that maximizes coherence of detections both in each 2D image and in the 3D scene. This process yields tightly fitted bounding boxes that circumvent occlusions or miss detections. The experimental results on five publicly available datasets show that the proposed approach outperforms state-of-the-art multi-camera pedestrian detectors, even some specifically trained on the target scenario, signifying the versatility and robustness of the proposed method without requiring ad hoc annotations nor human-guided configuration. Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Pablo Carballeira |
Knowl. Inf. Syst. | 4 |
| 2022 | FVV Live: A Real-Time Free-Viewpoint Video System With Consumer Electronics HardwareabstractFVV Live is a novel end-to-end free-viewpoint video system, designed for real-time operation, using consumer-grade cameras and hardware, which enables low deployment costs and easy installation for immersive event-broadcasting or videoconferencing. All the blocks of the system have been designed to maximize perceptual video quality, overcoming the limitations imposed by hardware and network, which impact directly the accuracy of depth data and thus the quality of virtual view synthesis. Therefore, it does not sacrifice perceptual video quality with respect to high-end counterparts. The results presented in this paper correspond to an implementation with nine stereo-based depth cameras. However, the design of the acquisition block of FVV Live allows scalability for an arbitrary number of cameras. In addition, FVV Live presents low motion-to-photon and end-to-end delays, which enables a responsive free-viewpoint navigation and bilateral immersive communications. Moreover, the visual quality of FVV Live has been assessed through subjective assessment with satisfactory results, and additional comparative tests show that it is preferred over state-of-the-art DIBR alternatives. Pablo Carballeira, Carlos Carmona, César Díaz, Daniel Berjón, Daniel Corregidor, Julián Cabrera, Francisco Morán, Carmen Doblado, Sergio Arnaldo, María del Mar Martín, Narciso García |
IEEE Trans. Multim. | 1 |
| 2021 | Robust people indoor localization with omnidirectional cameras using a Grid of Spatial-Aware Classifiers
Carlos R. del-Blanco, Pablo Carballeira, Fernando Jaureguizar, Narciso García |
Signal Process. Image Commun. | 2 |
| 2020 | QoE Analysis of Dense Multiview Video With Head-Mounted DevicesabstractThis paper presents a system and methodology for the analysis of quality of experience factors for dense multiview (MV) video using a head-mounted device (HMD). An MV-HMD player has been designed and implemented to immerse the users in a virtual environment, where they are placed in front of a virtual lightfield display that shows a different viewpoint depending on the position of their head. This paper describes a methodology for the analysis of the subjective perception of the transition among views (motion parallax), which is specific to the visualization of MV content. While previous works simulated the user movement by predefined view paths or used complex devices to track them, this system allows the observer to move freely, varying the perspective of the scene while easily tracking the observer's position. This paper is, up to our knowledge, the first providing a complete framework for the assessment of this subjective factor using an HMD. The subjective results obtained using this framework are used to 1) assess the influence of the user movement, display settings, and content characteristics in the perception of smoothness in the view transition, and 2) analyze the performance and limitations of a prediction model for subjective smoothness scores. Javier Cubelos, Pablo Carballeira, Jesús Gutiérrez 0001, Narciso García |
IEEE Trans. Multim. | 2 |
| 2017 | Toward the realization of six degrees-of-freedom with compressed light fieldsabstract360° video, supporting the ability to present views consistent with the rotation of the viewer's head along three axes (roll, pitch, yaw) is the current approach for creation of immersive video experiences. Nevertheless, a more fully natural, photorealistic experience - with support of visual cues that facilitate coherent psycho-visual sensory fusion without the side-effect of cyber-sickness - is desired. 360° video applications that additionally enable the user to translate in x, y, and z directions are clearly a subsequent frontier to be realized toward the goal of sensory fusion without cyber-sickness. Such support of full Six Degrees-of-Freedom (6 DoF) for next generation immersive video is a natural application for light fields. However, a significant obstacle to the adoption of light field technologies is the large data necessary to ensure that the light rays corresponding to the viewer's position relative to 6-DoF are properly delivered, either from captured light information or synthesized from available views. Experiments to improve known methods for view synthesis and depth estimation are therefore a fundamental next step to establish a reference framework within which compression technologies can be evaluated. This paper describes a testbed and experiments to enable smooth and artefact-free view transitions that can later be used in a framework to study how best to compress the data. Arianne T. Hinds, Didier Doyen, Pablo Carballeira |
ICME | 3 |
| 2017 | Subjective Assessment of Super Multiview Video with Coding ArtifactsabstractThe subjective assessment of super multiview (SMV) video considers two main perceptual factors: image quality and visual comfort at the viewpoint transition. While previous works only covered raw content with high levels of visual comfort, this work supersedes them by targeting the subjective assessment of SMV content with coding artifacts. The outcome of this analysis yields important conclusions regarding the relationship between these two factors, indicating that the perceived image quality is independent from the view point change speed, and the perceived visual comfort at the view point transition is independent from the image quality. These conclusions facilitate the extension of the scope of existing subjective perception models, designed for raw SMV content, to coded content. Rocio Recio, Pablo Carballeira, Jesús Gutiérrez 0001, Narciso García |
IEEE Signal Process. Lett. | 2 |
| 2016 | Analysis of the depth-shift distortion as an estimator for view synthesis distortion
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
Signal Process. Image Commun. | 1 |
| 2015 | Novel multi-feature Bag-of-Words descriptor via subspace random projection for efficient human-action recognitionabstractHuman-action recognition through local spatio-temporal features have been widely applied because of their simplicity and its reasonable computational complexity. The most common method to represent such features is the well-known Bag-of-Words approach, which turns a Multiple-Instance Learning problem into a supervised learning one, which can be addressed by a standard classifier. In this paper, a learning framework for human-action recognition that follows the previous strategy is presented. First, spatio-temporal features are detected. Second, they are described by HOG-HOF descriptors, and then represented by a Bag of Words approach to create a feature vector representation. The resulting high dimensional features are reduced by means of a subspace-random-projection technique that is able to retain almost all the original information. Lastly, the reduced feature vectors are delivered to a classifier called Citation K-Nearest Neighborhood, especially adapted to Multiple-Instance Learning frameworks. Excellent results have been obtained, outperforming other state-of-the art approaches in a public database. Ana I. Maqueda, Arturo Ruano, Carlos R. del-Blanco, Pablo Carballeira, Fernando Jaureguizar, Narciso García |
AVSS | 4 |
| 2014 | Systematic analysis of the decoding delay in multiview video
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
J. Vis. Commun. Image Represent. | 1 |
| 2008 | 3D reconstruction with uncalibrated cameras using the six-line conic varietyabstractWe present new algorithms for the recovery of the Euclidean structure from a projective calibration of a set of cameras with square pixels but otherwise arbitrarily varying intrinsic and extrinsic parameters. Our results, based on a novel geometric approach, include a closed-form solution for the case of three cameras and two known vanishing points and an efficient one-dimensional search algorithm for the case of four cameras and one known vanishing point. In addition, an algorithm for a reliable automatic detection of vanishing points on the images is presented. These techniques fit in a 3D reconstruction scheme oriented to urban scenes reconstruction. The satisfactory performance of the techniques is demonstrated with tests on synthetic and real data. Pablo Carballeira, José Ignacio Ronda, Antonio Valdés |
ICIP | 1 |