EDBT 2026 Demo / reviewers in the wild / expert
Narciso García
dblp:05/4319 · also Narciso García Santos
· DBLP profile ↗
137ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-0397-894XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 110 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 17 · 3 since 2021Human-computer interaction and ubiquitous computing · 13 · 11 since 2021Computer networks · 7 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Immersive remote learning for individuals with intellectual disabilitiesabstractExtended reality (XR) technologies are transforming domains like gaming, healthcare, and training, offering benefits such as enhanced safety, cost reduction, and immersive experiences. However, current XR systems often overlook accessibility for individuals with diverse abilities and remain unexplored in inclusive learning contexts. The Incluverso-5G project seeks to address this gap by leveraging XR to improve the lives of individuals with intellectual disabilities, including employment inclusion. This paper presents an immersive tool to support remote training for certified cooking courses, enabling users unable to attend in-person classes to access training and certification.The system includes an advanced 360-degree content viewer and a real-time immersive supervision tool, utilizing features like six degrees of freedom (6-DoF), color passthrough, and hand tracking for interactive learning. A four-step remote learning process—sensitization, instructional content, supervised practice, and certification—ensures accessibility and usability. The tool empowers individuals to acquire valuable skills and achieve professional certifications, demonstrating the potential of XR for inclusive training. Carlos Cortés 0001, Marta Goyena, Marta Orduna, Ainhoa Fernández-Alcaide, María Nava-Ruiz, Jesús Gutiérrez 0001, Pablo Pérez 0001, Narciso García |
IMX | 8 |
| 2026 | Viewpoint-invariant soccer pitch registration using geometric and learned featuresabstractAutomatic registration of broadcast soccer images to a standardized field model enables advanced analytics, augmented reality overlays, and precise player tracking. We propose a fully automatic, viewpoint-independent homography estimation pipeline fusing three complementary geometric cues: white field markings (lines and elliptical arcs), grass-band delimitations, and a binary playing-field mask. Detected primitives are first richly labeled — classifying lines as longitudinal or transversal, characterizing grass-tone transitions, and encoding four-quadrant intersection patterns — to reduce correspondence ambiguity. We then generate and prune candidate subsets of primitives, establish plausible matches to model elements via intersection-pattern rules and projective cross-ratio invariants, and systematically evaluate homography hypotheses using bidirectional mask-projection accuracies and mean reprojection error. An experimental evaluation on the LaSoDa benchmark demonstrates that the proposed method achieves highly accurate registrations with ground-truth primitives and robust performance in the fully automatic end-to-end pipeline. Furthermore, comparative experiments with recent state-of-the-art approaches confirm improved precision and robustness across diverse broadcast scenarios. Carlos Cuevas, Daniel Berjón, Narciso García |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Quality assessment of 3D reconstructed meshes: Bridging objective metrics, subjective perception, and behavioral cuesabstractAssessing the quality of 3D reconstructed models remains a key challenge in multimedia applications, especially in the context of cultural heritage, where visual fidelity and perceptual realism are equally crucial. This study investigates how reconstruction parameters, as well as existing objective quality metrics, align with human perception. In addition, we analyze how perceived quality and user interaction are related. A dataset of 3D models was generated by varying the number of input images, mesh complexity, and texture resolution. Results from a subjective study show that texture resolution significantly affects perceived quality, whereas variations in number of images and mesh complexity have a limited impact. Furthermore, interaction behavior was found to vary with perceived quality, with participants spending more time and exploring larger viewing angles for models receiving higher scores. These findings highlight the need for perceptually grounded, interaction-aware evaluation methodologies and provide guidelines for future perceptual optimization of 3D reconstruction pipelines. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, David Barbero García, Daniel Berjón, Francisco Morán, Narciso García, Federica Battisti, Jesús Gutiérrez 0001, Marco Carli, Julián Cabrera |
Signal Process. Image Commun. | 8 |
| 2025 | NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation ProblemabstractPath planning is usually solved by addressing either the (high-level) route planning problem (waypoint sequencing to achieve the final goal) or the (low-level) path planning problem (trajectory prediction between two waypoints avoiding collisions). However, real-world problems usually require simultaneous solutions to the route and path planning subproblems with a holistic and efficient approach. In this paper, we introduce NaviFormer, a deep reinforcement learning model based on a Transformer architecture that solves the global navigation problem by predicting both high-level routes and low-level trajectories. To evaluate NaviFormer, several experiments have been conducted, including comparisons with other algorithms. Results show competitive accuracy from NaviFormer since it can understand the constraints and difficulties of each subproblem and act consequently to improve performance. Moreover, its superior computation speed proves its suitability for real-time missions. Daniel Fuertes, Andrea Cavallaro, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IROS | 5 |
| 2025 | Embodied design for inclusive XR learning systemabstractEmbodied design can be an effective approach for developing extended reality (XR) experiences by enabling design teams to ideate and prototype through hands-on physical interaction. This body-centric approach is particularly valuable for addressing the complexities and paradigm shifts inherent in XR systems, where users engage significantly with their bodies in virtual environments, i.e. through head movements, body orientation, and hand tracking among others. In this paper, we present an embodied design method, which we used to design an XR-based cooking learning tool aimed at fostering autonomous learning for individuals with intellectual disabilities. The learning tool features 360-degree video recipes with a configurable combination of immersive video and passthrough views to adapt to different user needs and contexts. Our method includes functional prototypes to iterate current design concepts and explore usability across different possible interaction modalities. These tools were used to implement the Wizard of Oz technique and simulate different interaction techniques-such as voice commands and simple button-based inputs. To assess the usability of these tools, we employed the subjective Quality of Experience (QoE) assessment. The inclusion of this form of evaluation directly within the embodied design process is a novel contribution, providing valuable information about the usability of the prototypes as embodied design tools. Finally, we present preliminary results of the workshop, which illustrate the usefulness of embodied design for the development of XR interactive experiences. Carlos Cortés 0001, José Manuel Vega-Cebrián, Pablo Pérez 0001, María Nava-Ruiz, Narciso García, Elena Márquez Segura |
QoMEX | 5 |
| 2025 | Analysis of Objective 3D Mesh Quality Metrics for Cultural HeritageabstractExtended reality technologies are increasingly used in cultural heritage for preserving and accessing sites and artworks, where 3D model acquisition and rendering are key. Despite progress in reconstruction methodologies, a standardized approach to quality assessment is still missing. This study aims to evaluate objective quality metrics —both image-based and model-based, Full Reference and No Reference— applied to 3D models generated using the Structure from Motion algorithm. By varying parameters such as the number of images, number of triangles, and texture resolution, we examine the impact of these factors on metric outcomes, aiming to assess their reliability in cultural heritage applications. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, Jesús Gutiérrez 0001, Daniel Berjón, Francisco Morán, Federica Battisti, Narciso García, Marco Carli, Julián Cabrera |
QoMEX | 9 |
| 2025 | The Impact of Segmentation Methods for Avatar representation on User Experience in a Task-Based XR experienceabstractFigure 1: Three different hand representations based on different pass through techniques while users are performing three different tasks. Carlos Cortés 0001, David Barbero García, Jesús Gutiérrez 0001, Narciso García |
IMX | 4 |
| 2025 | A Study on Immersive Behavioral Therapy for Individuals with Intellectual Disabilities with Fear of StairsabstractThe increasing deployment of immersive technologies is opening up opportunities in the field of behavioral therapies.In particular, the use of eXtended Reality technologies allows the development of therapies that require bringing users to a remote location without Marta Goyena, Carlos Cortés 0001, Marta Orduna, Matteo Dal Magro, Ainhoa Fernández-Alcaide, María Nava-Ruiz, Jesús Gutiérrez 0001, Pablo Pérez 0001, Narciso García |
IMX | 9 |
| 2025 | Immersive Cognitive Training for Job Integration with people with Intellectual DisabilitiesabstractThe combination of eXtended Reality (XR) technologies with physiological signals offers innovative approaches to cognitive training for individuals with intellectual disabilities.This study presents a tool designed for cognitive training aimed at job integration through interactive and immersive environments.The tool was developed in collaboration with therapists from the Juan XXIII Foundation, an occupational center, and focuses on enhancing cognitive skills through serious games for job training.These games simulate realistic environments, such as a cafeteria and a supermarket, where users must complete various tasks.These tasks are designed to progressively increase in complexity as users advance, allowing them to improve their skills.Therapists can monitor and control the training sessions in real-time through an application, while users engage with the immersive scenarios.Additionally, the system incorporates non-invasive biosensors to track physiological data, such as heart rate, eye movement, etc. providing valuable insights into the Marta Goyena, Matteo Dal Magro, Martina Merolli, David Barbero García, Carlos Cortés 0001, Marta Orduna, Ainhoa Fernández-Alcaide, María Nava-Ruiz, Jesús Gutiérrez 0001, Pablo Pérez 0001, Narciso García |
IMX | 11 |
| 2025 | TOP-Former: A Multi-Agent Transformer Approach for the Team Orienteering ProblemabstractRoute planning for a fleet of vehicles is an important task in applications such as package delivery, surveillance, or transportation, often integrated within larger Intelligent Transportation Systems (ITS). This problem is commonly formulated as a Vehicle Routing Problem (VRP) known as the Team Orienteering Problem (TOP). Existing solvers for this problem primarily rely on either linear programming, which provides accurate solutions but requires computation times that grow with the size of the problem, or heuristic methods, which typically find suboptimal solutions in a shorter time. In this paper, we introduce TOP-Former, a multi-agent route planning neural network designed to efficiently and accurately solve the Team Orienteering Problem. The proposed algorithm is based on a centralized Transformer neural network capable of learning to encode the scenario (modeled as a graph) and analyze the complete context of all agents to deliver fast, precise, and collaborative solutions. Unlike other neural network-based approaches that adopt a more local perspective, TOP-Former is trained to understand the global situation of the vehicle fleet and generate solutions that maximize long-term expected returns. Extensive experiments demonstrate that the presented system outperforms most state-of-the-art methods in terms of both accuracy and computation speed. Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Engaging yet ineffective? A video coding tool's impact on learningabstractThis study investigates the use of an innovative educational tool, a video coding app, to improve students’ understanding of video coding concepts. The app was implemented in a university-level computer science course, and data were collected on students’ perceptions of the tool and their performance on knowledge tests. In this study, 7 participants conducted two practical sessions using the app. In each session, they performed a pre- and post-test evaluation. After each session, they also completed two subjective questionnaires to measure the perception of learning and usability of the tool. Results showed that the app was perceived as engaging and easy to use by students. However, analysis of test performance did not show a significant impact on learning outcomes. This suggests that the quality of experience may not always be correlated with objective results when evaluating new learning tools. As a lesson learned, we always recommend using pilot studies when evaluating the performance of new teaching tools. In addition, such studies should subjectively and objectively measure the influence of the tool on student performance. Carlos Cortés 0001, Carlos Cuevas, Narciso García |
QoMEX | 3 |
| 2024 | Real-Time Free Viewpoint Video for Immersive VideoconferencingabstractIn this work, we propose a demo of an immersive videoconference system using Free Viewpoint Video (FVV) technology. It makes use of the FVV Live system, which covers the entire FVV pipeline (capture, view rendering, and visualization) while working in real-time. The FVV Live system consists of nine cameras that capture an environment and a view renderer that uses the information from the cameras to generate a synthetic view at an arbitrary point.It is designed as a hybrid demo. While the capture and rendering processes take place at our premises, FVV Live can be visualized through devices connected to the Internet.The system allows immersive navigation of a virtual scene with 6 degrees of freedom, and interaction with live-captured avatars integrated in such scene. For this purpose, it uses WebRTC connections to update the position of the virtual camera and to receive the FVV Live view encoded as a video.Additionally, the user will be recorded by a simple camera and microphone setup, and the generated streams will be transmitted to our premises through the same WebRTC server. This way, people being recorded by FVV Live will be able to see and hear the user, enabling bidirectional communication. Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Daniel Berjón, Francisco Morán, César Díaz, Jesús Gutiérrez 0001, Fernando Jaureguizar, Narciso García, Julián Cabrera |
QoMEX | 9 |
| 2024 | Automatic highlight detection in videos of martial arts trickingabstractAbstract We propose a novel strategy for the automatic detection of highlight events in user-generated tricking videos, to the best of our knowledge, the first one specifically tailored for this complex sport. Most current methods for related sports leverage high-level semantics such as predefined camera angles or common editing practices, or rely on depth cameras to achieve automatic detection. However, our approach only relies on the contents (themselves) in the frames of a given video, and consists in a four stage pipeline. The first stage identifies foreground key points of interest along with an estimation of their motion in the video frames. In the second stage, these points are grouped into regions of interest based on their proximity and motion. Their behavior over time is evaluated in the third stage to generate an attention map indicating the regions participating in the most relevant events. The fourth and final stage provides the extracted video sequences where highlights have been identified. Experimental results attest to the effectiveness of our approach, which shows high recall and precision values at frame level, with detections that fit well the ground truth events. Marcos Rodrigo, Carlos Cuevas, Daniel Berjón, Narciso García |
Multim. Tools Appl. | 4 |
| 2024 | Delay Threshold for Social Interaction in Volumetric eXtended Reality CommunicationabstractImmersive technologies like eXtended Reality (XR) are the next step in videoconferencing. In this context, understanding the effect of delay on communication is crucial. This article presents the first study on the impact of delay on collaborative tasks using a realistic Social XR system. Specifically, we design an experiment and evaluate the impact of end-to-end delays of 300, 600, 900, 1,200, and 1,500 ms on the execution of a standardized task involving the collaboration of two remote users that meet in a virtual space and construct block-based shapes. To measure the impact of the delay in this communication scenario, objective and subjective data were collected. As objective data, we measured the time required to execute the tasks and computed conversational characteristics by analyzing the recorded audio signals. As subjective data, a questionnaire was prepared and completed by every user to evaluate different factors such as overall quality, perception of delay, annoyance using the system, level of presence, cybersickness, and other subjective factors associated with social interaction. The results show a clear influence of the delay on the perceived quality and a significant negative effect as the delay increases. Specifically, the results indicate that the acceptable threshold for end-to-end delay should not exceed 900 ms. This article additionally provides guidelines for developing standardized XR tasks for assessing interaction in Social XR environments. Carlos Cortés 0001, Irene Viola 0001, Jesús Gutiérrez 0001, Jack Jansen 0001, Shishir Subramanyam, Evangelos Alexiou, Pablo Pérez 0001, Narciso García, Pablo César |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2023 | Content-immersive Subjective Quality Assessment in Long Duration 360-degree VideosabstractThis paper presents a comparison between three content-immersive methodologies to evaluate video quality: Absolute Category Rating (ACR) with unrepeated scenes, Single Stimulus Continuous Quality Evaluation (SSCQE), and Single Stimulus Discrete Quality Evaluation (SSDQE). Three tests were conducted with different participants assigned to each methodology to assess 360-degree video quality in an immersive communications context. Additionally, socioemotional aspects such as presence, attitude, or attention were measured for each condition. All three methodologies offer similar results in terms of video quality evaluation, and they all keep similar levels of social and spatial presence. Both SSCQE and SSDQE allow narrative and context while evaluating video quality, but SSDQE has less impact in participants' attention to the video content. These results suggest that it is possible to extend ITU-T P.919 with the support of long sequences with in-sequence SSDQE video quality evaluation, without significant impact on the socioemotional properties of the experience. This would result on a standard content-immersive method, methodology suitable for 360-degree video content, favoring the ecological validity of the subjective assessment tests. Marta Orduna, Pablo Pérez 0001, Jesús Gutiérrez 0001, Narciso García |
QoMEX | 4 |
| 2023 | Solving routing problems for multiple cooperative Unmanned Aerial Vehicles using Transformer networksabstractMissions involving Unmanned Aerial Vehicle usually consist of reaching a set of regions, performing some actions in each region, and returning to a determined depot after all the regions have been successfully visited or before the fuel/battery is totally consumed. Hence, planning a route becomes an important task for many applications, especially if a team of Unmanned Aerial Vehicles is considered. From this team, coordination and cooperation are expected to optimize results of the mission. In this paper, a system for managing multiple cooperative Unmanned Aerial Vehicles is presented. This system divides the routing problem into two stages: initial planning and routing solving. Initial planning is a first step where the regions to be visited are grouped in multiple clusters according to a distance criterion, with each cluster being assigned to each of the Unmanned Aerial Vehicles. Routing solving computes the best route for every agent considering the clusters of the initial planning and a variant of the Orienteering Problem. This variant introduces the concept of shared regions, allowing an Unmanned Aerial Vehicle to visit regions from other clusters and compensating for the suboptimal region clustering of the previous stage. The Orienteering Problem with shared regions is solved using the deep learning architecture Transformer along with a deep reinforcement learning framework. This architecture is able to obtain high-quality solutions much faster than conventional optimization approaches. Extensive results and comparison with other Combinatorial Optimization algorithms, including cooperative and non-cooperative scenarios, have been performed to show the benefits of the proposed solution. Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar, Juan José Navarro, Narciso García |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Soccer line mark segmentation and classification with stochastic watershed transformabstractAugmented reality applications are beginning to change the way sports are broadcast, providing richer experiences and valuable insights to fans. The first step of augmented reality systems is camera calibration, possibly based on detecting the line markings of the playing field. Most existing proposals for line detection rely on edge detection and Hough transform, but radial distortion and extraneous edges cause inaccurate or spurious detections of line markings. We propose a novel strategy to automatically and accurately segment and classify line markings. First, line points are segmented thanks to a stochastic watershed transform that is robust to radial distortions, since it makes no assumptions about line straightness, and is unaffected by the presence of players or the ball. The line points are then linked to primitive structures (straight lines and ellipses) thanks to a very efficient procedure that makes no assumptions about the number of primitives that appear in each image. The strategy has been tested on a new and public database composed by 60 annotated images from matches in five stadiums. The results obtained have proven that the proposed strategy is more robust and accurate than existing approaches, achieving successful line mark detection even under challenging conditions. Daniel Berjón, Carlos Cuevas, Narciso García |
Signal Process. Image Commun. | 3 |
| 2023 | Methodology to Assess Quality, Presence, Empathy, Attitude, and Attention in 360-degree Videos for Immersive CommunicationsabstractThis paper proposes a methodology to assess video quality, spatial and social presence, empathy, attitude, and attention in 360-degree videos for immersive communications. The methodology is validated in an experiment which simulates an immersive communication environment where participants attend three conversations of different genre (everyday conversation, educational, and discussion) and from actor and observer acquisition perspectives. We consider three experimental conditions: (A) visualizing and rating the perceptual quality of contents in a Head-Mounted Display (HMD), (B) visualizing the contents in an HMD, and (C) visualizing the contents in an HMD where participants can see their hands and take notes. In all conditions participants visualize the same 360-degree videos, designed and acquired in the context of international experiences. Fifty-four participants were evenly distributed among A, B, and C conditions taking into account their international experience backgrounds (working or studying in a foreign country), obtaining a balanced and diverse sample of participants. Finally, the annotated dataset, Student Experiences Around the World dataset (SEAW-dataset), obtained from the experiment is made publicly available for the research community. Marta Orduna, Pablo Pérez 0001, Jesús Gutiérrez 0001, Narciso García |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | UPM-GTI-Face: A dataset for the evaluation of the impact of distance and masks in face detection and recognition systemsabstractWe present a novel dataset for the evaluation of face detection and recognition algorithms in challenging surveillance scenarios. The dataset consists in 4K images of different subjects captured at annotated distances ranging from 1 to 30 meters, both in indoor and outdoor environments, and under two face mask conditions (with and without). To the best of our knowledge, this is the only existing dataset that addresses the joint impact of masks and distances in a rigorous manner. We also propose an end-to-end fully automatic face detection and recognition system to provide baseline results on this dataset. Face detection is performed using Tiny Faces network, while face recognition is performed using VGG Face network. Experimental results show very high detection and recognition rates up to a distance of 20 meters, where the impact of distance is clear (especially for the latter). The use of face masks degrades the detection range and produces less consistent recognition results. Marcos Rodrigo, Ester Gonzalez-Sosa, Carlos Cuevas, Narciso García |
AVSS | 4 |
| 2022 | Impact of Self-View Latency on Quality of Experience: Analysis of Natural Interaction in XR EnvironmentsabstractThe rise of eXtended Reality (XR) has led to multiple ways of including the user’s body in interactive experiences. However, the delay limits of self-view rendering in interactive XR remain unexplored. This article presents a minimum self-view latency system and an interactive task-based experiment to study the influence of different levels of self-view delay on Quality of Experience (QoE) and task performance. During the experiment, 23 users tested 8 delay conditions (from 190 to 597 ms) while building block-based models. The results show a hard threshold in terms of involvement and overall quality around 450ms. However, the impact on adaptation and execution time was less pronounced. This indicates that although users adapted to the task in a certain way, their immersion was severely affected above a certain self-view delay value. Carlos Cortés 0001, Jesús Gutiérrez 0001, Pablo Pérez 0001, Irene Viola 0001, Pablo César, Narciso García |
ICIP | 6 |
| 2022 | Evaluation of the Performance of an Immersive System for Tele-educationabstractTele-education was already a solution for people who cannot attend lessons in person (such as inaccessibility in rural areas or illness issues). However, COVID has revealed problems in tele-education with current technology, causing adolescents and children to slow down their learning curves and experience problems of social distancing with their classmates. This paper presents a user study to validate an immersive communication system for tele-education purposes. This system streams in real time a class using 360-degree cameras, allowing remote students to explore the whole scene and improving the feeling of being in the classroom with their colleagues. Additionally, the prototype provides notifications to the remote students about events (such as a changes in the teacher’s presentation or classmates raising their hands) that occur outside their viewport to indicate in which direction they should move their heads to visualize them. Marta Orduna, Jesús Gutiérrez 0001, Alejandro Sánchez, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García |
IMX | 7 |
| 2022 | Grass band detection in soccer images for improved image registrationabstractThe registration of images of soccer matches is a key stage in many computer vision applications. Until now, this task has been typically carried out from key points obtained from the white line marks drawn on the field of play, but in many cases this does not yield enough keypoints for a robust registration. This article proposes a strategy to detect the borders between the grass bands of the field of play and therefore makes it possible to locate many more key points that will allow to carry out a subsequent registration of the images. First, a preprocessing is applied to obtain a grayscale image in which the grass bands are easily distinguishable, and also to obtain a binary mask of the entire field of play that determines the area of interest. Then, a local analysis is carried out to detect most of the borders between grass bands. Finally, a global analysis based on the intersections between lines is applied to group the detected borders and rule out false detections. The strategy has been evaluated on two databases composed of hundreds of annotated images from matches in several stadiums with different characteristics and light conditions. The results obtained have shown that most of the lines delimiting the grass bands are found successfully, while the number of false detections is very small. Carlos Cuevas, Daniel Berjón, Narciso García |
Signal Process. Image Commun. | 3 |
| 2022 | A Novel System for Nighttime Vehicle Detection Based on Foveal Classifiers With Real-Time PerformanceabstractVehicle monitoring using camera networks is an important task for traffic applications. Moreover, it becomes critical in nighttime, when the probability of an accident considerably increases as visibility conditions worsen. Typical approaches are mostly based on the assumption that regions delimiting vehicle lights are well defined, so that they are segmented and then associated to vehicle entities. However, this assumption fails in images acquired by existing traffic camera networks, where vehicle lights are revealed as flashes and other complex light patterns, occupying large and even disconnected image regions. In this work, a real-time vehicle detection algorithm for nighttime situations has been presented, which is able to locate vehicles in the image by analyzing the previous complex light patterns. For this purpose, a novel machine learning framework based on a grid of foveal classifiers has been designed. Every classifier in the grid processes the same global image descriptor (only one descriptor is computed per image). However, every one of them is trained to predict a different output depending on the classifier position in the grid and the vehicle ground-truth location. Additionally, only point-based annotations are required to train the grid of foveal classifiers, speeding up the cost of creating the required databases. Experimental results prove the effectiveness of the proposed method in a new created nighttime database with point-based annotations. Andrés Bell, Tomás Mantecón, César Díaz, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | FVV Live: A Real-Time Free-Viewpoint Video System With Consumer Electronics HardwareabstractFVV Live is a novel end-to-end free-viewpoint video system, designed for real-time operation, using consumer-grade cameras and hardware, which enables low deployment costs and easy installation for immersive event-broadcasting or videoconferencing. All the blocks of the system have been designed to maximize perceptual video quality, overcoming the limitations imposed by hardware and network, which impact directly the accuracy of depth data and thus the quality of virtual view synthesis. Therefore, it does not sacrifice perceptual video quality with respect to high-end counterparts. The results presented in this paper correspond to an implementation with nine stereo-based depth cameras. However, the design of the acquisition block of FVV Live allows scalability for an arbitrary number of cameras. In addition, FVV Live presents low motion-to-photon and end-to-end delays, which enables a responsive free-viewpoint navigation and bilateral immersive communications. Moreover, the visual quality of FVV Live has been assessed through subjective assessment with satisfactory results, and additional comparative tests show that it is preferred over state-of-the-art DIBR alternatives. Pablo Carballeira, Carlos Carmona, César Díaz, Daniel Berjón, Daniel Corregidor, Julián Cabrera, Francisco Morán, Carmen Doblado, Sergio Arnaldo, María del Mar Martín, Narciso García |
IEEE Trans. Multim. | 11 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 27 |
| 2022 | Subjective Assessment Experiments That Recruit Few Observers With Repetitions (FOWR)abstractRecent studies have shown that it is possible to characterize subject bias and variance in subjective assessment tests. Apparent differences among subjects can, for the most part, be explained by random factors. Building on that theory, we propose a subjective test design where four to six team members each rate the stimuli multiple times. The results are comparable to a high performing objective metric. This provides a quick and simple way to analyze new technologies and perform pre-tests for subjective assessment. Pablo Pérez 0001, Lucjan Janowski, Narciso García, Margaret H. Pinson |
IEEE Trans. Multim. | 3 |
| 2021 | EVENT-CLASS: Dataset of events in the classroomabstractThis work-in-progress presents a dataset of 360degree videos, called EVENT-CLASS, with associated characteristics in the context of tele-education. The sequences (video and audio) have been captured considering several environments, lighting conditions, acquisition perspectives, and cameras, enriching the dataset. EVENT-CLASS will be helpful for numerous applications related to tele-education, including quality assessment tests, and with the aim of improving the immersive experience of remote users thanks to the detection of relevant events that happen in the class. In this sense, this paper presents preliminary results of using transfer learning for person detection in 360degree scenes, based on Detectron2, and provides insights on the influence of applying it to equirectangular projection and over the viewport. Ongoing works will allow to include more videos and ground-truth annotations to the dataset. Marta Orduna, Jesús Gutiérrez 0001, Carlos Manzano, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García |
QoMEX | 8 |
| 2021 | Robust people indoor localization with omnidirectional cameras using a Grid of Spatial-Aware Classifiers
Carlos R. del-Blanco, Pablo Carballeira, Fernando Jaureguizar, Narciso García |
Signal Process. Image Commun. | 4 |
| 2020 | Robust Nighttime Vehicle Detection Based on Foveal ClassifiersabstractVisual vehicle surveillance has become an important research field due to its wide range of traffic applications. This task becomes more relevant in nighttime because accidents considerably increase. Typically, this problem is addressed by segmenting the bright image regions produced by vehicle lights, assuming they are well defined. But, often there are only flashes that occupy large image regions, invalidating the previous strategy. Thus, a real-time vehicle detection algorithm for nighttime that addresses the previous challenge is presented. First, the whole image is characterized by only one descriptor. Then, a grid of foveal classifiers that share the same previous image descriptor (unlike the traditional sliding window scheme) estimates the vehicle positions. Every classifier is trained to detect vehicles in specific image regions by analyzing the complex light patterns in the night. Furthermore, a new nighttime database has been also created to assess the effectiveness of the proposed method. Andrés Bell, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
ISCAS | 4 |
| 2020 | Influence of Video Delay on Quality, Presence, and Sickness in Viewport Adaptive Immersive StreamingabstractCurrent extended reality applications use omnidirectional video to boost users' immersion and sense of presence. As contents from distant video sources cannot be instantaneously delivered, the end-to-end delay becomes a key problem when user actions cannot be simultaneously matched by system reactions. Thus, we have designed and executed an experiment to assess its influence on the quality of experience, the sense of presence, and the sickness caused. To do it, we have developed a viewport adaptive simulator to render simultaneously two layers of immersive video to allow different adaptation schemes and delay values. Twenty observers have assessed 180 test videos from 9 sources. Our analysis shows a clear influence of the delay condition and the adaptation scheme on the perceived quality. Moreover, it also shows that the adaptation schemes and delay conditions have a small influence on the sense of presence and little effect on observers sickness. Carlos Cortés 0001, Pablo Pérez 0001, Jesús Gutiérrez 0001, Narciso García |
QoMEX | 4 |
| 2020 | Techniques and applications for soccer video analysis: A survey
Carlos Cuevas, Daniel Quilon, Narciso García |
Multim. Tools Appl. | 3 |
| 2020 | Automatic soccer field of play registration
Carlos Cuevas, Daniel Quilon, Narciso García |
Pattern Recognit. | 3 |
| 2020 | QoE Analysis of Dense Multiview Video With Head-Mounted DevicesabstractThis paper presents a system and methodology for the analysis of quality of experience factors for dense multiview (MV) video using a head-mounted device (HMD). An MV-HMD player has been designed and implemented to immerse the users in a virtual environment, where they are placed in front of a virtual lightfield display that shows a different viewpoint depending on the position of their head. This paper describes a methodology for the analysis of the subjective perception of the transition among views (motion parallax), which is specific to the visualization of MV content. While previous works simulated the user movement by predefined view paths or used complex devices to track them, this system allows the observer to move freely, varying the perspective of the scene while easily tracking the observer's position. This paper is, up to our knowledge, the first providing a complete framework for the assessment of this subjective factor using an HMD. The subjective results obtained using this framework are used to 1) assess the influence of the user movement, display settings, and content characteristics in the perception of smoothness in the view transition, and 2) analyze the performance and limitations of a prediction model for subjective smoothness scores. Javier Cubelos, Pablo Carballeira, Jesús Gutiérrez 0001, Narciso García |
IEEE Trans. Multim. | 4 |
| 2020 | XLR (piXel Loss Rate): A Lightweight Indicator to Measure Video QoE in IP NetworksabstractA novel Key Quality Indicator for video delivery applications, XLR (piXel Loss Rate), is defined, characterized, and evaluated. The proposed indicator is an objective measure that captures the effects of transmission errors in the received video, has a good correlation with subjective Mean Opinion Scores, and provides comparable results with state-of-the-art Full-Reference metrics. Moreover, XLR can be estimated using only a lightweight analysis on the compressed bitstream, thus allowing a No-Reference operational method. Therefore, XLR can be used for measuring the quality of experience without latency at any network location. Thus, it is a relevant tool for network planning, specially in new high-demanding scenarios. The experiments carried out show the outstanding performance of its linear-dimension score and the reliability of the bitstream-based estimation. César Díaz, Pablo Pérez 0001, Julián Cabrera, Jaime J. Ruiz, Narciso García |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2019 | ImageCLEF 2019: Multimedia Retrieval in Lifelogging, Medical, Nature, and Security Applications
Bogdan Ionescu, Henning Müller, Renaud Péteri, Duc-Tien Dang-Nguyen, Luca Piras 0001, Michael Riegler 0001, Minh-Triet Tran, Mathias Lux, Cathal Gurrin, Yashin Dicente Cid, Vitali Liauchuk, Vassili Kovalev, Asma Ben Abacha, Sadid A. Hasan, Vivek V. Datla, Joey Liu, Dina Demner-Fushman, Obioma Pelka, Christoph M. Friedrich, Jon Chamberlain, Adrian F. Clark, Alba Garcia Seco de Herrera, Narciso García, Ergina Kavallieratou, Carlos R. del-Blanco, Carlos Cuevas, Nikos Vasilopoulos, Konstantinos Karampidis |
ECIR (2) | 23 |
| 2019 | Subjective Assessment of Adaptive Media Playout for Video StreamingabstractAdaptive Media Playout (AMP), the adaptive modification of media playback speed, has been previously proposed as a technique to modify the playout delay of a user with respect to other users or with real time, thus having a crucial role in social TV and live video streaming. In this paper, we study the subjective impact of the playout rate change, which had received very limited attention in the past. In particular, we analyze the effect of the subject and the source content in the subjective opinion.We have found that most of the previous works underestimated the impact of AMP in perceived quality, and that the speed change rate should not be modified beyond 10% from the original, as a general rule. We have also developed a user scoring model which takes into account the variability between users and sources. Our results can help developing cost models for playout control systems based on AMP, as well as provide some insights to the analysis of other types of subjective assessment tests. Pablo Pérez 0001, Narciso García, Álvaro Villegas |
QoMEX | 2 |
| 2019 | Perceptually Equivalent Resolution in Handheld Devices for Streaming Bandwidth SavingabstractWe present the description, results, and analysis of the experiments conducted to find the equivalent resolution associated with handheld devices. That is, the resolution from which users stop perceiving quality improvements if better resolutions are presented to them in such devices. Thus, it is the maximum resolution that it is worth considering for generating and delivering video, as long as sequences are not too intensively compressed. Therefore, the detection of the equivalent resolutions allows for notable savings in bandwidth consumption. Subjective assessments have been carried out on fifty subjects using a set of video sequences of very different nature and four handheld devices with a broad range of screen dimensions. The results prove that the equivalent resolution in current handheld devices is 720p as higher resolutions are not valued by users. Mateo Camara, César Díaz, Juan Casal, Jorge Ruano, Narciso García |
IEEE Signal Process. Lett. | 5 |
| 2018 | Event-Based Vision Meets Deep Learning on Steering Prediction for Self-Driving CarsabstractEvent cameras are bio-inspired vision sensors that naturally capture the dynamics of a scene, filtering out redundant information. This paper presents a deep neural network approach that unlocks the potential of event cameras on a challenging motion-estimation task: prediction of a vehicle's steering angle. To make the best out of this sensor-algorithm combination, we adapt state-of-the-art convolutional architectures to the output of event sensors and extensively evaluate the performance of our approach on a publicly available large scale event-camera dataset (~1000 km). We present qualitative and quantitative explanations of why event cameras allow robust steering prediction even in cases where traditional cameras fail, e.g. challenging illumination conditions and fast motion. Finally, we demonstrate the advantages of leveraging transfer learning from traditional to event-based vision, and show that our approach outperforms state-of-the-art algorithms based on standard cameras. Ana I. Maqueda, Antonio Loquercio, Guillermo Gallego 0002, Narciso García, Davide Scaramuzza 0001 |
CVPR | 4 |
| 2018 | Real-time nonparametric background subtraction with tracking-based foreground update
Daniel Berjón, Carlos Cuevas, Francisco Morán, Narciso García |
Pattern Recognit. | 4 |
| 2018 | Automatic Depth Extraction from 2D Images Using a Cluster-Based Learning FrameworkabstractThere has been a significant increase in the availability of 3D players and displays in the last years. Nonetheless, the amount of 3D content has not experimented an increment of such magnitude. To alleviate this problem, many algorithms for converting images and videos from 2D to 3D have been proposed. Here, we present an automatic learning-based 2D-3D image conversion approach, based on the key hypothesis that color images with similar structure likely present a similar depth structure. The presented algorithm estimates the depth of a color query image using the prior knowledge provided by a repository of color + depth images. The algorithm clusters this database attending to their structural similarity, and then creates a representative of each color-depth image cluster that will be used as prior depth map. The selection of the appropriate prior depth map corresponding to one given color query image is accomplished by comparing the structural similarity in the color domain between the query image and the database. The comparison is based on a K-Nearest Neighbor framework that uses a learning procedure to build an adaptive combination of image feature descriptors. The best correspondences determine the cluster, and in turn the associated prior depth map. Finally, this prior estimation is enhanced through a segmentation-guided filtering that obtains the final depth map estimation. This approach has been tested using two publicly available databases, and compared with several state-of-the-art algorithms in order to prove its efficiency. José L. Herrera, Carlos R. del-Blanco, Narciso García |
IEEE Trans. Image Process. | 3 |
| 2017 | Subjective Assessment of Super Multiview Video with Coding ArtifactsabstractThe subjective assessment of super multiview (SMV) video considers two main perceptual factors: image quality and visual comfort at the viewpoint transition. While previous works only covered raw content with high levels of visual comfort, this work supersedes them by targeting the subjective assessment of SMV content with coding artifacts. The outcome of this analysis yields important conclusions regarding the relationship between these two factors, indicating that the perceived image quality is independent from the view point change speed, and the perceived visual comfort at the view point transition is independent from the image quality. These conclusions facilitate the extension of the scope of existing subjective perception models, designed for raw SMV content, to coded content. Rocio Recio, Pablo Carballeira, Jesús Gutiérrez 0001, Narciso García |
IEEE Signal Process. Lett. | 4 |
| 2017 | Detection of Stationary Foreground Objects Using Multiple Nonparametric Background-Foreground Models on a Finite State MachineabstractThere is a huge proliferation of surveillance systems that require strategies for detecting different kinds of stationary foreground objects (e.g., unattended packages or illegally parked vehicles). As these strategies must be able to detect foreground objects remaining static in crowd scenarios, regardless of how long they have not been moving, several algorithms for detecting different kinds of such foreground objects have been developed over the last decades. This paper presents an efficient and high-quality strategy to detect stationary foreground objects, which is able to detect not only completely static objects but also partially static ones. Three parallel nonparametric detectors with different absorption rates are used to detect currently moving foreground objects, short-term stationary foreground objects, and long-term stationary foreground objects. The results of the detectors are fed into a novel finite state machine that classifies the pixels among background, moving foreground objects, stationary foreground objects, occluded stationary foreground objects, and uncovered background. Results show that the proposed detection strategy is not only able to achieve high quality in several challenging situations but it also improves upon previous strategies. Carlos Cuevas, Daniel Berjón, Narciso García |
IEEE Trans. Image Process. | 4 |
| 2016 | Hand Gesture Recognition Using Infrared Imagery Provided by Leap Motion Controller
Tomás Mantecón, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
ACIVS | 4 |
| 2016 | Detection of stationary foreground objects: A survey
Carlos Cuevas, Narciso García |
Comput. Vis. Image Underst. | 3 |
| 2016 | Labeled dataset for integral evaluation of moving object detection algorithms: LASIESTA
Carlos Cuevas, Eva María Yáñez, Narciso García |
Comput. Vis. Image Underst. | 3 |
| 2016 | Perceptual Quality of HTTP Adaptive Streaming Strategies: Cross-Experimental Analysis of Multi-Laboratory and Crowdsourced Subjective StudiesabstractToday's packet-switched networks are subject to bandwidth fluctuations that cause degradation of the user experience of multimedia services. In order to cope with this problem, HTTP adaptive streaming (HAS) has been proposed in recent years as a video delivery solution for the future Internet and being adopted by an increasing number of streaming services, such as Netflix and Youtube. HAS enables service providers to improve users' quality of experience (QoE) and network resource utilization by adapting the quality of the video stream to the current network conditions. However, the resulting time-varying video quality caused by adaptation introduces a new type of impairment and thus novel QoE research challenges. Despite various recent attempts to investigate these challenges, many fundamental questions regarding HAS perceptual performance are still open. In this paper, the QoE impact of different technical adaptation parameters, including chunk length, switching amplitude, switching frequency, and temporal recency, are investigated. In addition, the influence of content on perceptual quality of these parameters is analyzed. To this end, a large number of adaptation scenarios have been subjectively evaluated in four laboratory experiments and one crowdsourcing study. A statistical analysis of the combined data set reveals results that partly contradict widely held assumptions and provide novel insights in perceptual quality of adapted video sequences, e.g., interaction effects between quality switching direction (up/down) and switching strategy (smooth/abrupt). The large variety of experimental configurations across different studies ensures the consistency and external validity of the presented results that can be utilized for enhancing the perceptual performance of adaptive streaming services. Samira Tavakoli, Sebastian Egger-Lampl, Michael Seufert, Raimund Schatz, Kjell Brunnström, Narciso García |
IEEE J. Sel. Areas Commun. | 6 |
| 2016 | Analysis of the depth-shift distortion as an estimator for view synthesis distortion
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
Signal Process. Image Commun. | 4 |
| 2016 | Visual Face Recognition Using Bag of Dense Derivative Depth PatternsabstractA novel biometric face recognition algorithm using depth cameras is proposed. The key contribution is the design of a novel and highly discriminative face image descriptor called bag of dense derivative depth patterns (Bag-D3P). This descriptor is composed of four different stages that fully exploit the characteristics of depth information: 1) dense spatial derivatives to encode the 3-D local structure; 2) face-adaptive quantization of the previous derivatives; 3) multibag of words that creates a compact vector description from the quantized derivatives; and 4) spatial block division to add global spatial information. The proposed system can recognize people faces from a wide range of poses, not only frontal ones, increasing its applicability to real situations. Last, a new face database of high-resolution depth images has been created and made it public for evaluation purposes. Tomás Mantecón, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IEEE Signal Process. Lett. | 4 |
| 2016 | Temporal Pyramid Matching of Local Binary Subpatterns for Hand-Gesture RecognitionabstractHuman-computer Interaction systems based on hand-gesture recognition are nowadays of great interest to establish a natural communication between humans and machines. However, the visual recognition of gestures and other human poses remains a challenging problem. In this paper, the original volumetric spatiograms of local binary patterns descriptor has been extended to efficiently and robustly encode the spatial and temporal information of hand gestures. This enhancement mitigates the dimensionality problems of the previous approach, and considers more temporal information to achieve a higher recognition rate. Excellent results have been obtained, outperforming other existing approaches of the state of the art. Ana I. Maqueda, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IEEE Signal Process. Lett. | 4 |
| 2016 | Optimal Piecewise Linear Function Approximation for GPU-Based ApplicationsabstractMany computer vision and human-computer interaction applications developed in recent years need evaluating complex and continuous mathematical functions as an essential step toward proper operation. However, rigorous evaluation of these kind of functions often implies a very high computational cost, unacceptable in real-time applications. To alleviate this problem, functions are commonly approximated by simpler piecewise-polynomial representations. Following this idea, we propose a novel, efficient, and practical technique to evaluate complex and continuous functions using a nearly optimal design of two types of piecewise linear approximations in the case of a large budget of evaluation subintervals. To this end, we develop a thorough error analysis that yields asymptotically tight bounds to accurately quantify the approximation performance of both representations. It provides an improvement upon previous error estimates and allows the user to control the tradeoff between the approximation error and the number of evaluation subintervals. To guarantee real-time operation, the method is suitable for, but not limited to, an efficient implementation in modern graphics processing units, where it outperforms previous alternative approaches by exploiting the fixed-function interpolation routines present in their texture units. The proposed technique is a perfect match for any application requiring the evaluation of continuous functions; we have measured in detail its quality and efficiency on several functions, and, in particular, the Gaussian function because it is extensively used in many areas of computer vision and cybernetics, and it is expensive to evaluate. Daniel Berjón, Guillermo Gallego 0002, Carlos Cuevas, Francisco Morán, Narciso García |
IEEE Trans. Cybern. | 5 |
| 2015 | Enhanced gesture-based human-computer interaction through a Compressive Sensing reduction scheme of very large and efficient depth feature descriptorsabstractIn this paper, a hand gesture-based recognition system is presented with the aim of recognizing finger-spelling using the American Sign Language. The solution makes use of the depth imagery acquired by the new Kinect 2 sensor that provides more depth resolution. The main novelty is the introduction of a Compressive Sensing step to reduce the dimension of a depth-based feature descriptor, called Depth Spatiograms of Quantized Patterns, which is very discriminative, but also too large for its practical application. The system is composed by three steps: 1) depth-based feature descriptor computation that robustly characterizes the hand gesture; 2) Compressive Sensing based dimensionality reduction that shortens the previous highly discriminative but also large feature vector with almost no information lost; and 3) Support Vector Machine based classification that recognizes the performed hand gestures. Promising recognition results have been obtained in an American Sign Language based database. Tomás Mantecón, Ana Mantecon, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
AVSS | 5 |
| 2015 | Novel multi-feature Bag-of-Words descriptor via subspace random projection for efficient human-action recognitionabstractHuman-action recognition through local spatio-temporal features have been widely applied because of their simplicity and its reasonable computational complexity. The most common method to represent such features is the well-known Bag-of-Words approach, which turns a Multiple-Instance Learning problem into a supervised learning one, which can be addressed by a standard classifier. In this paper, a learning framework for human-action recognition that follows the previous strategy is presented. First, spatio-temporal features are detected. Second, they are described by HOG-HOF descriptors, and then represented by a Bag of Words approach to create a feature vector representation. The resulting high dimensional features are reduced by means of a subspace-random-projection technique that is able to retain almost all the original information. Lastly, the reduced feature vectors are delivered to a classifier called Citation K-Nearest Neighborhood, especially adapted to Multiple-Instance Learning frameworks. Excellent results have been obtained, outperforming other state-of-the art approaches in a public database. Ana I. Maqueda, Arturo Ruano, Carlos R. del-Blanco, Pablo Carballeira, Fernando Jaureguizar, Narciso García |
AVSS | 6 |
| 2015 | An extension to the PRO-MPEG COP3 codes for unequal error protection of real-time video transmissionabstractWe propose and evaluate an extension to the Application-Layer FEC (AL-FEC) codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes), consisting in allowing the use of a number of matrices of dissimilar size per FEC block. So, unequal protection of the data packets in the video stream is enabled, since dissimilar code rates can be applied to different groups of data packets. This boosts the efficiency of the protection scheme, increasing the video quality of the sequence presented to final users, even if the resulting packet loss rate (PLR) after channel decoding remains the same. Evaluation results show a significantly better performance of the Pro-MPEG COP3 codes when the proposed protection extension is incorporated. César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ICIP | 4 |
| 2015 | Learning-based depth estimation from 2D images using GIST and saliencyabstractAlthough there has been a significant proliferation of 3D displays in the last decade, the availability of 3D content is still scant compared to the volume of 2D data. To fill this gap, automatic 2D to 3D conversion algorithms are needed. In this paper, we present an automatic approach, inspired by machine learning principles, for estimating the depth of a 2D image. The depth of a query image is inferred from a dataset of color and depth images by searching this repository for images that are photometrically similar to the query. We measure the photometric similarity between two images by comparing their GIST descriptors. Since not all regions in the query image require the same visual attention, we give more weight in the GIST-descriptor comparison to regions with high saliency. Subsequently, we fuse the depths of the most similar images and adaptively filter the result to obtain a depth estimate. Our experimental results indicate that the proposed algorithm outperforms other state-of-the-art approaches on the commonly-used Kinect-NYU dataset. José L. Herrera, Janusz Konrad, Carlos R. del-Blanco, Narciso García |
ICIP | 4 |
| 2015 | Q-learning based control algorithm for HTTP adaptive streamingabstractWe present a control algorithm based on Q-Learning for an HTTP Adaptive Streaming (HAS) Client in order to optimize the Quality of Experience (QoE) of the user. First, we propose a model with a suitable number of variables in an attempt to find a reasonable tradeoff between the complexity of the model and its capacity to capture appropriately the dynamics of the system. Second, we define a novel reward function that takes into consideration factors related to the user's QoE. Results will show, that our Q-learning algorithm is able to learn and efficiently control the selection of the segment qualities. In addition, we will show that our proposed approach outperforms another Q-learning approach. Virginia Martin, Julián Cabrera, Narciso García |
VCIP | 3 |
| 2015 | Seamless, Static Multi-Texturing of 3D MeshesabstractAbstract In the context of 3D reconstruction, we present a static multi‐texturing system yielding a seamless texture atlas calculated by combining the colour information from several photos from the same subject covering most of its surface. These pictures can be provided by shooting just one camera several times when reconstructing a static object, or a set of synchronized cameras, when dealing with a human or any other moving object. We suppress the colour seams due to image misalignments and irregular lighting conditions that multi‐texturing approaches typically suffer from, while minimizing the blurring effect introduced by colour blending techniques. Our system is robust enough to compensate for the almost inevitable inaccuracies of 3D meshes obtained with visual hull–based techniques: errors in silhouette segmentation, inherently bad handling of concavities, etc. Rafael Pagés, Daniel Berjón, Francisco Morán, Narciso García |
Comput. Graph. Forum | 4 |
| 2015 | Human-computer interaction based on visual hand-gesture recognition using volumetric spatiograms of local binary patterns
Ana I. Maqueda, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
Comput. Vis. Image Underst. | 4 |
| 2015 | Quality of Experience of adaptive video streaming: Investigation in service parameters and subjective quality assessment methodology
Samira Tavakoli, Kjell Brunnström, Jesús Gutiérrez 0001, Narciso García |
Signal Process. Image Commun. | 4 |
| 2014 | Learning 3D structure from 2D images using LBP featuresabstractAn automatic machine learning strategy for computing the 3D structure of monocular images from a single image query using Local Binary Patterns is presented. The 3D structure is inferred through a training set composed by a repository of color and depth images, assuming that images with similar structure present similar depth maps. Local Binary Patterns are used to characterize the structure of the color images. The depth maps of those color images with a similar structure to the query image are adaptively combined and filtered to estimate the final depth map. Using public databases, promising results have been obtained outperforming other state-of-the-art algorithms and with a computational cost similar to the most efficient 2D-to-3D algorithms. José L. Herrera, Carlos R. del-Blanco, Narciso García |
ICIP | 3 |
| 2014 | Depth-based face recognition using local quantized patterns adapted for range dataabstractA depth-based face recognition algorithm specially adapted to high range resolution data acquired by the new Microsoft Kinect 2 sensor is presented. A novel descriptor called Depth Local Quantized Pattern descriptor has been designed to make use of the extended range resolution of the new sensor. This descriptor is a substantial modification of the popular Local Binary Pattern algorithm. One of the main contributions is the introduction of a quantification step, increasing its capacity to distinguish different depth patterns. The proposed descriptor has been used to train and test a Support Vector Machine classifier, which has proven to be able to accurately recognize different people faces from a wide range of poses. In addition, a new depth-based face database acquired by the new Kinect 2 sensor have been created and made public to evaluate the proposed face recognition system. Tomás Mantecón, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
ICIP | 4 |
| 2014 | Subjective Quality Study of Adaptive Streaming of Monoscopic and Stereoscopic VideoabstractNowadays, HTTP adaptive streaming (HAS) has become a reliable distribution technology offering significant advantages in terms of both user perceived Quality of Experience (QoE) and resource utilization for content and network service providers. By trading-off the video quality, HAS is able to adapt to the available bandwidth and display requirements so that it can deliver the video content to a variety of devices over the Internet. However, until now there is not enough knowledge of how the adaptation techniques affect the end user's visual experience. Therefore, this paper presents a comparative analysis of different bitrate adaptation strategies in adaptive streaming of monoscopic and stereoscopic video. This has been done through a subjective experiment of testing the end-user response to the video quality variations, considering the visual comfort issue. The experimental outcomes have made a good insight into the factors that can influence on the QoE of different adaptation strategies. Samira Tavakoli, Jesús Gutiérrez 0001, Narciso García |
IEEE J. Sel. Areas Commun. | 3 |
| 2014 | Systematic analysis of the decoding delay in multiview video
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Advanced background modeling with RGB-D sensors through classifiers combination and inter-frame foreground prediction
Massimo Camplani, Carlos R. del-Blanco, Luis Salgado, Fernando Jaureguizar, Narciso García |
Mach. Vis. Appl. | 5 |
| 2014 | Camera Localization UsingTrajectories and MapsabstractWe propose a new Bayesian framework for automatically determining the position (location and orientation) of an uncalibrated camera using the observations of moving objects and a schematic map of the passable areas of the environment. Our approach takes advantage of static and dynamic information on the scene structures through prior probability distributions for object dynamics. The proposed approach restricts plausible positions where the sensor can be located while taking into account the inherent ambiguity of the given setting. The proposed framework samples from the posterior probability distribution for the camera position via data driven MCMC, guided by an initial geometric analysis that restricts the search space. A Kullback-Leibler divergence analysis is then used that yields the final camera position estimate, while explicitly isolating ambiguous settings. The proposed approach is evaluated in synthetic and real environments, showing its satisfactory performance in both ambiguous and unambiguous settings. Raúl Mohedano, Andrea Cavallaro, Narciso García |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Multi-sensor background subtraction by fusing multiple region-based probabilistic classifiers
Massimo Camplani, Carlos R. del-Blanco, Luis Salgado, Fernando Jaureguizar, Narciso García |
Pattern Recognit. Lett. | 5 |
| 2013 | Depth video coding for free viewpoint video oriented to the synthetic view perceptual qualityabstractIn this paper we propose a novel depth encoding algorithm based on the synthetic view perceptual quality optimization. In Free Viewpoint Video, depth sequences are never shown to the observer but they are used for the view synthesis. Due to the compression quality losses, the depth distortion generates errors in the synthetic view pixel position. We propose to encode the depth sequences by a novel approach, divided into two parts. The first one is a rate distortion optimization method based on the minimization of the synthetic view pixel position error. The second one is an extension of the first one and improves the performance taking into account the perceptual distortion, evaluated by exploiting the Human Visual System characteristics described by the Just Noticeable Distortion. However, the Just Noticeable Distortion is a model designed for the traditional video, hence we also propose a Just Noticeable Pixel Displacement evaluation method which, considering the synthetic view perceptual distortion, is able to estimate when an error in the synthetic view pixel position is not noticed. The results show how the proposed algorithm not only improves the synthetic view quality measured by the Video Quality Metric (highly correlated to the subjective quality) but also the objective quality measured by the PSNR, with an achieved improvement of 0.3 dB with a corresponding bit saving of 13%. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 4 |
| 2013 | A combined active contours method for segmentation using localization and multiresolutionabstractImage segmentation is a fundamental step in many image processing applications. To achieve high-quality segmentations active contours are commonly used. However, state of art strategies are not able to provide successful results in all the conditions. Additionally, the strategies that get the best overall results are computationally expensive and need to manually set some parameters, which decreases their usability. Here, we propose a novel active contours-based segmentation method that, through the combination of boundary-based and region-based energies and a multiresolution analysis, provides very high-quality results while significantly increasing both the computational efficiency and the usability of previous approaches. Eva María Yáñez, Carlos Cuevas, Narciso García |
ICIP | 3 |
| 2013 | Stochastic modelling of peer-assisted VoD streaming in managed networks
Sasho Gramatikov, Fernando Jaureguizar, Julián Cabrera, Narciso García |
Comput. Networks | 4 |
| 2013 | Improved background modeling for real-time spatio-temporal non-parametric moving object detection strategies
Carlos Cuevas, Narciso García |
Image Vis. Comput. | 2 |
| 2013 | Low Complexity Mode Decision and Motion Estimation for H.264/AVC Based Depth Maps Encoding in Free Viewpoint VideoabstractWithin free viewpoint video, the 3-D reconstruction of the scene is created from a high number of viewpoints. Every viewpoint is represented by a traditional sequence, called texture, and its associated depth information. This is known as a View plus Depth environment. In this paper, a novel low complexity mode decision and motion estimation algorithm for the H.264/AVC based encoding of depth sequences is proposed. Given that a texture sequence and its associated depth represent the same scene from the same point of view, they should have similar motion characteristics. The complexity reduction of the depth encoding, in the proposed algorithm, is obtained by taking advantage of the texture motion information that has been previously processed by a traditional H.264/AVC encoder. The characteristics of depth and texture sequences are analyzed, focusing on similarities and differences that are properly managed to design an algorithm able to detect when the motion information of the texture might be usefully exploited in the encoding of the corresponding depth sequence. The proposed method is able to achieve the same objective quality, measured by means of the PSNR and the VQM, than a traditional H.264/AVC encoder with a reduction of up to 58% of the computational burden. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Efficient Moving Object Detection for Lightweight Applications on Smart CamerasabstractRecently, the number of electronic devices with smart cameras has grown enormously. These devices require new, fast, and efficient computer vision applications that include moving object detection strategies. In this paper, a novel and high-quality strategy for real-time moving object detection by nonparametric modeling is presented. It is suitable for its application to smart cameras operating in real time in a large variety of scenarios. While the background is modeled using an innovative combination of chromaticity and gradients, reducing the influence of shadows and reflected light in the detections, the foreground model combines this information and spatial information. The application of a particle filter allows to update the spatial information and provides a priori knowledge about the areas to analyze in the following images, enabling an important reduction in the computational requirements and improving the segmentation results. The quality of the results and the achieved computational efficiency show the suitability of the proposed strategy to enable new applications and opportunities in last generation of electronic devices. Carlos Cuevas, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Inference of Complex Trajectories by Means of a Multibehavior and Multiobject Tracking AlgorithmabstractVisual tracking of multiple objects is a fundamental aspect of many video-based systems. Today, there are reliable algorithms that can track a small number of objects in restricted situations. However, the tracking of a large number of objects in uncontrolled situations involving interacting objects with complex dynamics is still a challenge. In this situation, the typical assumptions of linearity and independence of object motions are not fulfilled, causing a low tracking performance. This paper proposes a novel Bayesian tracking algorithm for interacting objects that are able to reliably simulate several object behaviors with an uncalibrated camera, which can be positioned in an arbitrary perspective. Three different models of object behavior are used to simulate and predict the object dynamics, where the proportion of hypotheses of each possible behavior of an object depends on the dynamics (position, velocity, etc.) of the other objects in the scene. Experimental results on public databases prove the reliability and robustness of the proposed tracking algorithm in the presence of object interactions. Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Versatile Bayesian classifier for moving object detection by non-parametric background-foreground modelingabstractAlong the recent years, several moving object detection strategies by non-parametric background-foreground modeling have been proposed. To combine both models and to obtain the probability of a pixel to belong to the foreground, these strategies make use of Bayesian classifiers. However, these classifiers do not allow to take advantage of additional prior information at different pixels. So, we propose a novel and efficient alternative Bayesian classifier that is suitable for this kind of strategies and that allows the use of whatever prior information. Additionally, we present an effective method to dynamically estimate prior probability from the result of a particle filter-based tracking strategy. Carlos Cuevas, Raúl Mohedano, Narciso García |
ICIP | 3 |
| 2012 | Subjective study of adaptive streaming strategies for 3DTVabstractAlthough the delivery of 3D video services to households is nowadays a reality thanks to frame-compatible formats, many efforts are being made to obtain efficient methods to transmit 3D content offering a high quality of experience to the end users. In this paper, a stereoscopic video streaming scenario is considered, and the perceptual impact of various strategies applicable to adaptive streaming situations are compared. Specifically, the mechanisms are based on switching between copies of the content with different coding qualities, on discarding frames of the sequence, on switching from 3D to 2D, and on using asymmetric coding of the stereo views. In addition, when video freezes happen, the possibility of keeping the end-to-end latency or maintaining the continuity of the video are considered. These aspects were evaluated carrying out a subjective assessment test, considering also visual discomfort issues, using a methodology designed to keep as far as possible domestic viewing conditions. Jesús Gutiérrez 0001, Pablo Pérez 0001, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 5 |
| 2012 | Depth perceptual video coding for free viewpoint video based on H.264/AVCabstractA novel scheme for depth sequences compression, based on a perceptual coding algorithm, is proposed. A depth sequence describes the object position in the 3D scene, and is used, in Free Viewpoint Video, for the generation of synthetic video sequences. In perceptual video coding the human visual system characteristics are exploited to improve the compression efficiency. As depth sequences are never shown, the perceptual video coding, assessed over them, is not effective. The proposed algorithm is based on a novel perceptual rate distortion optimization process, assessed over the perceptual distortion of the rendered views generated through the encoded depth sequences. The experimental results show the effectiveness of the proposed method, able to obtain a very considerable improvement of the rendered view perceptual quality. Gianluca Cernigliaro, Matteo Naccari, Fernando Jaureguizar, Julián Cabrera, Narciso García |
PCS | 5 |
| 2012 | Adaptive protection scheme for MVC-encoded stereoscopic video streaming in IP-based networksabstractWe present an adaptive unequal error protection (UEP) strategy built on the 1-D interleaved parity Application Layer Forward Error Correction (AL-FEC) code for protecting the transmission of stereoscopic 3D video content encoded with Multiview Video Coding (MVC) through IP-based networks. Our scheme targets the minimization of quality degradation produced by packet losses during video transmission in time-sensitive application scenarios. To that end, based on a novel packet-level distortion model, it selects in real time the most suitable packets within each Group of Pictures (GOP) to be protected and the most convenient FEC technique parameters, i.e., the size of the FEC generator matrix. In order to make these decisions, it considers the relevance of the packet, the behavior of the channel, and the available bitrate for protection purposes. Simulation results validate both the distortion model introduced to estimate the importance of packets and the optimization of the FEC technique parameter values. César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García |
VCIP | 4 |
| 2012 | Bounded non-deterministic planning for multimedia adaptation
Fernando López Hernández 0001, Dietmar Jannach, José María Martínez Sanchez, Christian Timmerer, Narciso García, Hermann Hellwagner |
Appl. Intell. | 5 |
| 2011 | A new fast motion estimation and mode decision algorithm for H.264 depth maps encoding in free viewpoint TVabstractIn this paper, we consider a scenario where 3D scenes are modeled through a View+Depth representation. This representation is to be used at the rendering side to generate synthetic views for free viewpoint video. The encoding of both type of data (view and depth) is carried out using two H.264/AVC encoders. In this scenario we address the reduction of the encoding complexity of depth data. Firstly, an analysis of the Mode Decision and Motion Estimation processes has been conducted for both view and depth sequences, in order to capture the correlation between them. Taking advantage of this correlation, we propose a fast mode decision and motion estimation algorithm for the depth encoding. Results show that the proposed algorithm reduces the computational burden with a negligible loss in terms of quality of the rendered synthetic views. Quality measurements have been conducted using the Video Quality Metric. Gianluca Cernigliaro, Matteo Naccari, Fernando Jaureguizar, Julián Cabrera, Fernando Pereira 0001, Narciso García |
ICIP | 6 |
| 2011 | Automatic bandwidth estimation strategy for high-quality non-parametric modeling based moving object detectionabstractHere, a novel and efficient moving object detection strategy by non-parametric modeling is presented. Whereas the foreground is modeled by combining color and spatial information, the background model is constructed exclusively with color information, thus resulting in a great reduction of the computational and memory requirements. The estimation of the background and foreground covariance matrices, allows us to obtain compact moving regions while the number of false detections is reduced. Additionally, the application of a tracking strategy provides a priori knowledge about the spatial position of the moving objects, which improves the performance of the Bayesian classifier. Carlos Cuevas, Narciso García |
ICIP | 2 |
| 2011 | Simultaneous 3D object tracking and camera parameter estimation by Bayesian methods and transdimensional MCMC samplingabstractMulti-camera 3D tracking systems with overlapping cameras represent a powerful mean for scene analysis, as they potentially allow greater robustness than monocular systems and provide useful 3D information about object location and movement. However, their performance relies on accurately calibrated camera networks, which is not a realistic assumption in real surveillance environments. Here, we introduce a multi-camera system for tracking the 3D position of a varying number of objects and simultaneously refining the calibration of the network of overlapping cameras. Therefore, we introduce a Bayesian framework that combines Particle Filtering for tracking with recursive Bayesian estimation methods by means of adapted transdimensional MCMC sampling. Additionally, the system has been designed to work on simple motion detection masks, making it suitable for camera networks with low transmission capabilities. Tests show that our approach allows a successful performance even when starting from clearly inaccurate camera calibrations, which would ruin conventional approaches. Raúl Mohedano, Narciso García |
ICIP | 2 |
| 2011 | Bayesian visual surveillance: A model for detecting and tracking a variable number of moving objectsabstractAn automatic detection and tracking framework for visual surveillance is proposed, which is able to handle a variable number of moving objects. Video object detectors generate an unordered set of noisy, false, missing, split, and merged measurements that make extremely complex the tracking task. Especially challenging are split detections (one object is split into several measurements) and merged detections (several objects are merged into one detection). Few approaches address this problem directly, and the existing ones use heuristics methods, or assume a known number of objects, or are not suitable for on-line applications. In this paper, a Bayesian Visual Surveillance Model is proposed that is able to manage undesirable measurements. Particularly, split and merged measurements are explicitly modeled by stochastic processes. Inference is accurately performed through a particle filtering approach that combines ancestral and MCMC sampling. Experimental results have shown a high performance of the proposed approach in real situations. Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
ICIP | 3 |
| 2011 | Inter-packet symbol approach to Reed-Solomon FEC codes for RTP-multimedia stream protectionabstractThis paper presents an alternative Forward Error Correction scheme, based on Reed-Solomon codes, with the aim of protecting the transmission of RTP-multimedia streams: the inter-packet symbol approach. This scheme is based on an alternative bit structure that allocates each symbol of the Reed-Solomon code in several RTP-media packets. This characteristic permits to exploit better the recovery capability of Reed-Solomon codes against bursty packet losses. The performance of our approach has been studied in terms of encoding/decoding time versus recovery capability, and compared with other proposed schemes in the literature. The theoretical analysis has shown that our approach allows the use of a lower size of the Galois Fields compared to other solutions. This lower size results in a decrease of the required encoding/decoding time while keeping a comparable recovery capability. Finally, experimental results have been carried out to assess the performance of our approach compared to other schemes in a simulated environment, where models for wireless and wireline channels have been considered. Filippo Casu, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ISCC | 4 |
| 2011 | Qualitative Monitoring of Video Quality of ExperienceabstractReal-time monitoring of multimedia Quality of Experience is a critical task for the providers of multimedia delivery services: from television broadcasters to IP content delivery networks or IPTV. For such scenarios, meaningful metrics are required which can generate useful information to the service providers that overcome the limitations of pure Quality of Service monitoring probes. However, most of objective multimedia quality estimators, aimed at modeling the Mean Opinion Score, are difficult to apply to massive quality monitoring. Thus we propose a lightweight and scalable monitoring architecture called Qualitative Experience Monitoring (QuEM), based on detecting identifiable impairment events such as the ones reported by the customers of those services. We also carried out a subjective assessment test to validate the approach and calibrate the metrics. Preliminary results of this test set support our approach. Pablo Pérez 0001, Jesús Gutiérrez 0001, Jaime J. Ruiz, Narciso García |
ISM | 4 |
| 2011 | Subjective evaluation of transmission errors in IPTV and 3DTVabstractThe increase of multimedia services delivered over packet-based networks has entailed greater quality expectations of the end-users. This has led to an intensive research on techniques for evaluating the quality of experience perceived by the viewers of audiovisual content, considering the different degradations that it could suffer along the broadcasting system. In this paper, a comprehensive study of the impact of transmission errors affecting video and audio in IPTV is presented. With this aim, subjective assessment tests were carried out proposing a novel methodology trying to keep as close as possible home environment viewing conditions. Also 3DTV content in side-by-side format has been used in the experiments to compare the impact of the degradations. The results provide a better understanding of the effects of transmission errors, and show that the QoE related to the first approach of 3DTV is acceptable, but the visual discomfort that it causes should be reduced. Jesús Gutiérrez 0001, Pablo Pérez 0001, Fernando Jaureguizar, Julián Cabrera, Narciso García |
VCIP | 5 |
| 2011 | A model for preference-driven multimedia adaptation decision-making in the MPEG-21 framework
Fernando López Hernández 0001, José María Martínez Sanchez, Narciso García |
Multim. Tools Appl. | 3 |
| 2011 | Line segment detection using weighted mean shift procedures on a 2D slice sampling strategy
Marcos Nieto Doncel, Carlos Cuevas, Luis Salgado, Narciso García |
Pattern Anal. Appl. | 4 |
| 2010 | Fast mode decision for multiview video coding based on scene geometryabstractA new fast mode decision (FMD) algorithm for multi-view video coding (MVC) is presented. The codification of the views is based on the analysis of the homogeneity of the depth map and corrected with the motion analysis of a reference view, which is encoded based on traditional methods and on the use of the disparity differences between the views. This approach reduces the burden of the rate-distortion motion analysis using the availability of a depth map and the presence of the disparity vectors, which are assumed to be provided by the acquisition process. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 4 |
| 2010 | Tracking-based non-parametric background-foreground classification in a chromaticity-gradient spaceabstractThis work presents a novel background-foreground classification technique based on adaptive non-parametric kernel estimation in a color-gradient space of components. By combining normalized color components with their gradients, shadows are efficiently suppressed from the results, while the luminance information in the moving objects is preserved. Moreover, a fast multi-region iterative tracking strategy applied over previously detected foreground regions allows to construct a robust foreground modeling, which combined with the background model increases noticeably the quality in the detections. The proposed strategy has been applied to different kind of sequences, obtaining satisfactory results in complex situations such as those given by dynamic backgrounds, illumination changes, shadows and multiple moving objects. Carlos Cuevas, Narciso García |
ICIP | 2 |
| 2010 | A wireless video transmission control approach through Stochastic Dynamic ProgrammingabstractThis paper presents an intelligent, rate-limited multicast video transmission optimization scheme for video distribution over 802.11 wireless networks based on packet retransmissions. We propose a problem formulation which involves the characteristics of the encoded stream together with the behaviour of the wireless channel. Using standard Stochastic Dynamic Programming techniques, optimal control policies are obtained off-line. These policies are optimal in the sense of minimizing the expected distortion at the terminal. In addition, the on-line complexity or our approach is very low since the optimization problem is solved off-line. The performance of our scheme has been evaluated in a real scenario and compared with that of a limited rate ARQ algorithm. The results for our proposed system show a higher packet recovery rate and a better protection of information with a higher priority. Victor Miguel, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ICIP | 4 |
| 2010 | Robust multi-camera tracking from schematic descriptionsabstractAlthough monocular 2D tracking has been largely studied in the literature, it suffers from some inherent problems, mainly when handling persistent occlusions, that limit its performance in practical situations. Tracking methods combining observations from multiple cameras seem to solve these problems. However, most multi-camera systems require detailed information from each view, making it impossible their use in real networks with low transmission rate. In this paper, we present a robust multi-camera 3D tracking method which works on schematic descriptions of the observations performed by each camera of the system, allowing thus its performance in real surveillance networks. It is based on unspecific 2D detection systems working independently in each camera, whose results are smartly combined by means of a Bayesian association method based on geometry and color, allowing the 3D tracking of the objects of the scene with a Particle Filter. The tests performed show the excellent performance of the system, even correcting possible failures of the 2D processing modules. Raúl Mohedano, Narciso García |
ICIP | 2 |
| 2010 | Capabilities and limitations of mono-camera pedestrian-based autocalibrationabstractMany environments lack enough architectural information to allow an autocalibration based on features extracted from the scene structure. Nevertheless, the observation over time of walking people can generally be used to estimate the vertical vanishing point and the horizon line in the acquired image. However, this information is not enough to allow the calibration of a general camera without presuming excessive simplifications. This paper presents a study on the capabilities and limitations of the mono-camera calibration methods based solely on the knowledge of the vertical vanishing point and the horizon line in the image. The mathematical analysis sets the conditions to assure the feasibility of the mono-camera pedestrian-based autocalibration. In addition, examples of applications are presented and discussed. Raúl Mohedano, Narciso García |
ICIP | 2 |
| 2010 | Visual tracking of multiple interacting objects through Rao-Blackwellized Data Association Particle FilteringabstractA multiple object visual tracking framework is presented, which is able to manage complex object interactions, missing detections and clutter. The main contribution is the ability to deal with complex situations in which the interacting objects can change their dynamics while they are occluded. This is achieved by explicitly estimating putative locations of the occluded objects. The tracking is modeled by a Rao-Blackwellized Data Association Particle Filter (RBDAPF), which has a tractable substructure that allows to analytically compute the object positions, while the object-measurement associations are approximated by Particle Filtering. Besides improving the accuracy, this filter decomposition reduces the computational cost, since the complexity with the number of objects becomes linear instead of exponential. The Particle Filter efficiently manages the measurements from visible and occluded objects, the clutter, and missing measurements to estimate the correct data associations that lead to a robust tracking. Experimental results on surveillance videos show that the proposed RBDAPF framework is able to track multiple interacting objects in complex situations. Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
ICIP | 3 |
| 2010 | An adaptive, real-time, traffic monitoring system
Tomás Rodríguez, Narciso García |
Mach. Vis. Appl. | 2 |
| 2009 | Object Tracking from Unstabilized Platforms by Particle Filtering with Embedded Camera Ego MotionabstractVisual tracking with moving cameras is a challenging task. The global motion induced by the moving camera moves the target object outside the expected search area, according to the object dynamics. The typical approach is to use a registration algorithm to compensate the camera motion. However, in situations involving several moving objects, and backgrounds highly affected by the aperture problem, image registration quality may be very low, decreasing dramatically the performance of the tracking. In this work, a novel approach is proposed to successfully tackle the tracking with moving cameras in complex situations, which involve several independent moving objects. The key idea is to compute several hypotheses for the camera motion, instead of estimating deterministically only one. These hypotheses are combined with the object dynamics in a particle filter framework to predict the most probable object locations. Then, each hypothetical object location is evaluated by the measurement model using a spatiogram, which is a region descriptor based on color and spatial distributions. Experimental results show that the proposed strategy allows to accurately track an object in complex situations affected by strong ego motion. Carlos R. del-Blanco, Narciso García, Luis Salgado, Fernando Jaureguizar |
AVSS | 2 |
| 2008 | 3D Tracking Using Multi-view Based Particle Filters
Raúl Mohedano, Narciso García, Luis Salgado, Fernando Jaureguizar |
ACIVS | 2 |
| 2008 | Automatic Feature-Based Stabilization of Video with Intentional Motion through a Particle Filter
Carlos R. del-Blanco, Fernando Jaureguizar, Luis Salgado, Narciso García |
ACIVS | 4 |
| 2008 | Robust 3D people tracking and positioning system in a semi-overlapped multi-camera environmentabstractPeople positioning and tracking in 3D indoor environments are challenging tasks due to background clutter and occlusions. Current works are focused on solving people occlusions in low-cluttered backgrounds, but fail in high-cluttered scenarios, specially when foreground objects occlude people. In this paper, a novel 3D people positioning and tracking system is presented, which shows itself robust to both possible occlusion sources: static scene objects and other people. The system holds on a set of multiple cameras with partially overlapped fields of view. Moving regions are segmented independently in each camera stream by means of a new background modeling strategy based on Gabor filters. People detection is carried out on these segmentations through a template-based correlation strategy. Detected people are tracked independently in each camera view by means of a graph-based matching strategy, which estimates the best correspondences between consecutive people segmentations. Finally, 3D tracking and positioning of people is achieved by geometrical consistency analysis over the tracked 2D candidates, using head position (instead of object centroids) to increase robustness to foreground occlusions. Raúl Mohedano, Carlos R. del-Blanco, Fernando Jaureguizar, Luis Salgado, Narciso García |
ICIP | 5 |
| 2007 | Aerial Moving Target Detection Based on Motion Vector Field Analysis
Carlos R. del-Blanco, Fernando Jaureguizar, Luis Salgado, Narciso García |
ACIVS | 4 |
| 2007 | Recursive Camera Autocalibration with the Kalman FilterabstractGiven a projective reconstruction of a 3D scene, we address the problem of recovering the Euclidean structure of the scene in a recursive way. This leads to the application of Kalman filtering to the problem of camera autocalibration and to new algorithms for the autocalibration of cameras with varying parameters. This has benefits in saving memory and computational effort, and obtaining faster updates of the 3D Euclidean structure of the scene under consideration. Guillermo Gallego 0002, José Ignacio Ronda, Antonio Valdés, Narciso García |
ICIP (5) | 4 |
| 2007 | Target Detection Through Robust Motion Segmentation and Tracking Restrictions in Aerial Flir ImagesabstractAn efficient automatic moving target detection and tracking system in airborne forward looking infrared (FLIR) imagery is presented in this paper. Due to camera ego-motion, these detection and tracking tasks are challenging problems. Besides, previously proposed techniques are not suitable for aerial images, as the predominant regions are non-textured. The proposed system efficiently estimates not only the camera motion but also the target motion, by means of an accurate motion vector field computation and robust motion parameters estimation technique. This information allows accurately to segment each target, and tracking them with ego-motion compensation. Verification of tracking restrictions helps detecting true targets while reducing very significantly the false alarm rate. Excellent results have been obtained over real FLIR sequences. Carlos R. del-Blanco, Fernando Jaureguizar, Luis Salgado, Narciso García |
ICIP (5) | 4 |
| 2007 | A New Class Partitioned Discrete Model for the Characterization of RTP Multicast Transmission through IEEE 802.11 ChannelsabstractMulticast of real time protocol (RTP) flows (i.e. multicast video streaming) within 802.11 wireless local area networks (WLANs) has to deal with the inherent unreliability of wireless channels and the lack of link-level error protection of multicast transmissions. This problem can be solved in part by means of sophisticated rate control strategies, but no efficient discrete channel models for this scenario have been developed yet. In this paper a novel packet error rate class partitioned (PERCP) model for the characterization of wireless channels is proposed. Its characteristics make it suitable for its integration in rate-control schemes, and it is proved to perform better than simplified Gilbert models in all considered scenarios. Juan C. Plaza, Julián Cabrera, Fernando Jaureguizar, Narciso García |
VTC Fall | 4 |
| 2005 | Moving Objects Segmentation Based on Automatic Foreground / Background Identification of Static Elements
Laurent Isenegger, Luis Salgado, Narciso García |
ACIVS | 3 |
| 2005 | Robust and accurate registration of images with unknown relative orientation and exposureabstractNowadays, low-cost digital cameras embedded in mobile phones are becoming the main source of visual information. Nevertheless, its use is restricted to simple store and forward applications due to the limitations in the spatial resolution and dynamic range of those cameras. A registration algorithm is proposed to allow image mosaicing and therefore to increase camera applications. It is able to find the geometric relation between any image in the set, lacking any prior knowledge of the relative position and exposure of any image in the set, and working on low-overlapping pairs. It also deals with partial occlusions. The registration algorithm is based on the use of Zernike moments and RANSAC robust fitting to guarantee stability, and KLT tracker to provide accuracy. Besides, a radiometric compensation stage allows the development of a completely automatic and seamless mosaicing system. Pablo Pérez 0001, Narciso García |
ICIP (3) | 2 |
| 2004 | Fast face segmentation in component color spaceabstractA fast face segmentation color-based method is presented. The characterization of the decision thresholds along the skin color distribution, constituting a non-rectangular decision region for skin detection in the YCgCr color space, is analyzed. Besides, a transformation based on the skin representation in the chrominance plane is proposed for skin detection improvement. Basically, it is based on the rotation of the Cg and Cr axes in the chrominance plane, so that a boundary box can be used for face segmentation. Segmentation tests have been performed on a set of portrait-like images with different lighting conditions. Two types of decision thresholds have been obtained for detecting the whole head or just the skin region. The face detection results achieved in the segmentation process have also been compared to those using the YCgCr decision thresholds. Juan J. De Dios, Narciso García |
ICIP | 2 |
| 2004 | Unsupervised segmentation algorithm of hrtem images
Ainhoa Mendizabal, Julián Cabrera, Luis Salgado, Narciso García, Juan C. Gonzalez |
ICIP | 4 |
| 2004 | Comparison of wavelet-based three-dimensional model coding techniquesabstractTruly hierarchical three-dimensional model coding (3DMC) techniques based on subdivision surfaces have, over merely progressive ones, the advantage of yielding pyramidally nested meshes that inherently approximate the surface of a three-dimensional (3-D) object at increasing levels of detail (LODs). Such a nesting makes it natural to edit (hence animate) the model at different LODs. From a compression efficiency viewpoint, it is also the most advantageous, because it permits us to group into hierarchical sets the 3-D details needed to reposition the new vertices predicted along the subdivision process. We first review extensively previous work on both progressive and hierarchical 3DMC. We then describe our hierarchical 3DMC technique based on subdivision surfaces, which is inspired by the set partitioning in hierarchical trees (SPIHT) coding scheme and has in turn inspired the wavelet subdivision surfaces tool of MPEG-4's Animation Framework eXtension (AFX). We finally compare our technique with several others. Francisco Morán, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | DCT based segmentation applied to a scalable zenithal people counterabstractThis paper faces the problem of detecting the number of customers crossing an uncontrolled access to a big store, from the information retrieved by a zenithal camera. The solution operates in real time and is scalable in two ways: allows the coverage of wide accesses (spatial scalability), and works with a level of detail enough to allow detection of carried objects (functional scalability). In order to cope with the main problems of these systems - shadows, sudden changes in global lighting, and sporadic camera motion due to vibration -, a novel DCT based segmentation is presented. Jesús Bescós, José Manuel Menéndez, Narciso García |
ICIP (3) | 3 |
| 2003 | Stochastic rate-control of interframe video coders for VBR channelsabstractWe propose a new algorithm for the real-time control of an inter-frame video coder operating with a variable rate channel such as wireless channels or the Internet. Using techniques of stochastic dynamic programming we obtain off-line optimal policies from stochastic models of the channel and coder which minimize the average expected distortion. The on-line complexity of our approach is only that required to identify the state of the system (source and channel). The state of the channel is obtained based on the ARQ error-control mechanism, and the source state is computed as complexity measurements on each incoming frame. Simulation results based on this new approach are provided and compared to other proposed rate-control strategies. They show how our model-based optimal policies outperform the other considered approaches keeping a negligible on-line computational cost. This result is very interesting when considering an alternative to traditional costly solutions based on deterministic dynamic programming. Julián Cabrera, José Ignacio Ronda, Antonio Ortega, Narciso García |
ICIP (3) | 4 |
| 2003 | Face detection based on a new color space YCgCrabstractA new color space, YCgCr, is described and applied for face detection. Although similar to the YCbCr color space, it differs in the use of the Cg color component instead of the Cb one. Here, the fundamentals of this new color space YCgCr are presented and its capabilities and advantages over YCbCr are analyzed. After describing the face extraction technique and the required decision values, the parameters which determine the decision thresholds are modeled and represented in the Cg-Cr plane. Based on the representation of the experimental results, the decision thresholds are modified. Finally, the segmentations results achieved with this new color space are compared with those obtained in YCbCr. Juan J. De Dios, Narciso García |
ICIP (3) | 2 |
| 2002 | Subdivision surfaces in MPEG-4abstractSS (subdivision surfaces) are a powerful modelling paradigm for truly hierarchical (instead of merely progressive) 3D surface (as opposed to mesh) coding. Two SS-based tools of the MPEG-4 animation framework extension allow us to derive a piecewise smooth surface from an initial control mesh: if the subdivision process is run in its predefined form, the initial mesh is simply smoothed; if 3D details are added to the positions of the new vertices appearing after each subdivision step, a particular target surface may be approximated with an increasing accuracy. In both cases, multiresolution editing/animation is possible. Francisco Morán, Patrick Gioia, Michael Steliaros, Narciso García, Mikaël Bourges-Sévenier |
ICIP (3) | 4 |
| 2002 | Dynamic object segmentation for outdoor analysis
Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón |
VCIP | 2 |
| 2001 | SADWT for efficient mesh based video codingabstractThe application of SADWT (shape adaptive discrete wavelet transform) has been widely proposed for object-oriented video coding, making it possible to transform and code the arbitrarily-shaped regions obtained by a segmentation of the scene content (which are usually objects in the sense of being a visually observable part of the image context). We present an application which is not restricted to objects but is meant to deal with any kind of region that is desired to be treated in a special way according to the coding objectives. Hereby, the main idea is to achieve a more efficient use of the available bits on areas with high prediction error (important for very low bit-rate coding) or skin-coloured areas (important for video telephony or conferencing). The coding of these areas can be independent, i.e. without taking into account the rest of the image, or could be controlled by multi-object bit distribution algorithms. Promising results for subjective quality tests have already been obtained at fixed bit rates with fixed and variable allocation strategies and our 2D irregular mesh based video coder, as well as for very-low bit-rate coding. José Ignacio Ronda, Martina Eckert, Narciso García |
ICIP (3) | 3 |
| 2001 | Automatic antibiograms inhibition halo determination through texture and directional filtering analysisabstractA new segmentation and analysis strategy to automatically measure the inhibition halos (IHs) in antibiograms is presented. It is based on the application of a combined texture and directional filtering analysis technique on the result of a segmentation process which incorporates colour analysis and Hough transform applications. The computational efficiency is highly improved by applying the different processing stages to selected areas of the acquired images and exploiting a priori information about the IHs shape and distribution. Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón, Damián Ruiz-Coll |
ICIP (2) | 2 |
| 2001 | Region-based control points determination for multivector motion descriptionabstractMotion description is a key clement in region-based video-coding approaches, becoming essential for low and very low bit-rate transmissions. Among the different advanced motion description approaches, multivector motion description (MMD) models the region motion through a variable number of motion vectors (MVs) applied on specific region points called control points (CPs). However, its application for region-based video coding would require the transmission of both MVs and CPs in order to adequately compensate each region motion. Here, a region-based control point determination strategy for MMD is proposed, which is strictly based on region contour information and the application of motion-interpolation accuracy constraints to drive the CP's location procedure. This approach offers two main advantages while keeping a reasonable computational cost: (1) it is able to operate on arbitrary-shaped regions of the image and (2) it allows the complete elimination of the transmission of any information related to the number and position of the CPs. Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | A Monocular Vision System for Autonomous Vehicle GuidanceabstractThe improvement of vehicle security is a major priority of the car industry. Both active and passive security-systems have experienced a great development during the last decade, but the main research is still focused on minimising the errors committed by the driver rather than trying to avoid them. In this paper, a basic monocular vision system, focused on road location solely, is described. Working with only one video camera hinders the exact 3D reconstruction of the scene: no information about distances and dimensions is available, unless a priori artificial constraints are taken. The problem increases because of the movement of the camera with respect to the scene. In the present system, the road searching and following algorithms operate on the two-dimensional image plane, and 2D to 3D conversion is not accomplished. The system obtains excellent final results, succeeding in more than 75% of the analysed images from the tested sequences. Alberto L. Maganto, José Manuel Menéndez, Luis Salgado, Enrique Rendón, Narciso García |
ICIP | 5 |
| 2000 | Hierarchical Coding of 3D Models with Subdivision SurfacesabstractWe build on state of the art methods for multiresolution embedded coding of images, such as Said and Pearlman's (1996) set partitioning in hierarchical trees, and combine them with ideas for 3D objects modelling with subdivision surfaces, to obtain a new technique for hierarchical 3D model coding. The compression ratios we obtain are better than or similar to previously reported ones but, perhaps more importantly, the truly hierarchical coding of 3D objects we propose allows their efficient multiresolution animation. This kind of technique could have a major impact on VRML and MPEG-4, the two ISO standards that now deal with the coding of 3D objects, which are in both cases static and linearly approximated by polygonal meshes. In future versions of those standards, that will have to address the coding of dynamic 3D objects, these will most likely be modelled with higher order primitives such as subdivision surfaces. Francisco Morán, Narciso García |
ICIP | 2 |
| 2000 | Efficient Prediction Error Regions Determination for Region-Based Video Coding Through Shape Adaptive DCTabstractAn efficient strategy to determine the prediction error regions to be coded within a region-based prediction error coding scheme is presented. Prediction error coding is based on the segmentation of the displaced field difference (DFD) and coding the resulting arbitrary shaped DFD regions using shape adaptive DCT. Efficiency in the determination of the DFD regions to be coded is achieved by eliminating from the selection process the direct computation of the cost of region contours and textures coding. With this scheme, perceptual distortion of the decoded images is reduced while quality is locally improved on relevant image areas. Comparative results with the complete H.263 coder are shown. Luis Salgado, José Manuel Menéndez, Enrique Rendón, Narciso García, Raúl Larrosa |
ICIP | 4 |
| 2000 | Evaluation of DWT and DCT for irregular mesh-based motion compensation in predictive video coding
Martina Eckert, Damián Ruiz-Coll, José Ignacio Ronda, Fernando Jaureguizar, Narciso García |
VCIP | 5 |
| 2000 | Efficient image segmentation for region-based motion estimation and compensationabstractAn intra-frame segmentation strategy to assist region-based motion estimation and compensation is presented. It is based on the multiresolution application of a histogram clustering and a probabilistic relaxation-labeling algorithm, followed by a local gradient-based bottom-up merging procedure. Specially suited for region-based video coding, it strongly differs from other proposals in that it generates arbitrary shaped image regions with pixel accuracy at a low computational cost, while allowing full reconstruction of the segmentation at the decoder without the transmission of any region description information. José Manuel Menéndez, Luis Salgado, Enrique Rendón, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1999 | Multiresolution Image Segmentation for Region-Based Motion Estimation and CompensationabstractAn intra-frame segmentation strategy to assist region-based motion estimation and compensation is presented. It is based on the multiresolution application of a histogram clustering and a probabilistic relaxation labelling algorithm, followed by a focal gradient-based bottom-up merging procedure. Specially suited for region-based video coding, it strongly differs from other proposals in that it generates arbitrary shaped image regions with pixel accuracy at a low computational cost, while allowing full reconstruction of the segmentation at the decoder without the transmission of any region description information. Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón |
ICIP (2) | 2 |
| 1999 | Determination of Control Point Sets for Motion Description Based on Motion Interpolation Accuracy Constraints
Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón |
ICIP (3) | 2 |
| 1999 | Model-based analytical FOE determination
José Manuel Menéndez, Narciso García, Luis Salgado, Enrique Rendón |
Signal Process. Image Commun. | 2 |
| 1999 | Rate control and bit allocation for MPEG-4abstractIn previous years, an interest has developed in the coded representations of video signals allowing independent manipulation of semantically independent elements (objects). Along these lines, the ISO standard MPEG-4 enhances the traditional concept of the video sequence to convert it into a synchronized set of visual objects organized in a flexible way. The real-time generation of a bitstream according to this new paradigm, and suitable for its transmission through either fixed- or variable-rate channels, results in a challenging new bit allocation and rate control problem, which has to satisfy complex application requirements. This paper formalizes this new issue by focusing on the design of rate control systems for real-time applications. The proposed approach relies on the modelization of the source and the optimization of a cost criterion based on signal quality parameters. Different cost criteria are provided, corresponding to a set of relevant definitions of the object priority concept. Algorithms are introduced to minimize the average distortion of the objects, to guarantee desired qualities to the most relevant ones, and to keep constant ratios among the object qualities. The techniques have been applied to a coder implementing the MPEG-4 video verification model, showing good properties in terms of achievement of the control objectives. José Ignacio Ronda, Martina Eckert, Fernando Jaureguizar, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1998 | Contourless Region-Based Video Coding for Very Low Bit-RatesabstractIn spite of their advantages, region-based approaches present a heavy burden in the required amount of information to describe region contours. To avoid these limitations, a new region-based motion estimation and compensation strategy is proposed, which allows the operation on arbitrary shaped regions and the ability to reconstruct them without any contour information. The strategy is enhanced by the use of multivector motion estimation and compensation. Results showing the quality and advantages of the strategy are provided. Luis Salgado, Narciso García, José Manuel Menéndez, Enrique Rendón |
ICIP (1) | 2 |
| 1997 | Adaptive Palette Determination for Color Images Based on Kohonen NetworksabstractA new method for the determination of adaptive palettes of color images with a variation of Kohonen (1984) networks, for the purposes of image storage and display, is presented. New neighborhood relationships in one and three dimensions are defined among neurons and the scope of these binds is also extended. The presented method is compared with popular palette selection algorithms, the resulting quality, provided network convergence, is competitive with the other methods, both for natural color and graphic images. Additionally, the resulting palettes show interesting topological relationships among colors. Enrique Rendón, Luis Salgado, José Manuel Menéndez, Narciso García |
ICIP (1) | 4 |
| 1996 | An algorithm for FOE localizationabstractAn analytical procedure for the estimation of the focus of expansion (FOE) location is introduced. The method is applied to a vector field that has been obtained from the analysis of the relative translational movement of a rigid body with respect to the acquiring camera (monocular system). Two basic properties of complex differentiation theory are applied, assuming that each vector of the optic flow is a real-valued complex function of the form v(x,y)=v/sub x/(x,y)+iv/sub y/,(x,y). These properties provide the set of points for which both components of the vector field v/sub x/(x,y) and v/sub y/(x,y) are harmonic and, as a subset, the set of points for which both components are null. The use of perspective projection states that the FOE is obtained as the locus where v/sup x/(x,y)=0=v/sub y/(x,y) and, therefore, where the conditions stated by both properties meet. José Manuel Menéndez, Narciso García, Luis Salgado, Enrique Rendón |
ICIP (3) | 2 |
| 1996 | Buffer-constrained coding of video sequences with quasi-constant qualityabstractThe problem of controlling a video coder in order to adapt its variable output bit-rate to the input requirements of the communication channel has received a lot of attention. While in the reported work the performance of the control system is mainly measured in terms of the resulting average coding distortion and its degree of variability, only the minimization of the first has received enough attention as a case for systematic optimization, while the joint minimization of the two terms has only received a heuristic treatment. This paper investigates the compromise between average distortion and distortion variability within the frame of the model-based optimal control for overflow-safe operation of the coder. For this purpose, the systematic obtaining of control policies which optimize a general measure of the dynamical performance of the system is addressed as an extension of the classical formulation of the buffer control problem. By obtaining optimal solutions for different parts of the general performance measure, the compromise between the average value of the distortion and its average variability is made explicit and its qualitative aspects analyzed. Results are provided for an ITU-standard hybrid video coder. José Ignacio Ronda, Fernando Jaureguizar, Narciso García |
ICIP (3) | 3 |
| 1994 | Image communication open architecture
Rüdiger Strack, Christof Blum, David A. Duce, Dale C. Sutcliffe, Narciso García, María José Pérez-Luque, Eckhard Moeller, Hauke Peyn |
Comput. Graph. | 5 |
| 1993 | Stochastic optimal control of variable bit-rate video codersabstractTransmission of variable bit-rate encoded video over either constant or variable bit-rate communication channels imposes restrictions on the shape of the bit generation process in the encoder. In order that these restrictions can be met, a coder-network interface system has be to included consisting of a buffer and a coder controller device. In this paper we show how the availability of appropriate statistical models of the behavior of the source-coder system allows the formalization of the problem of the controller design as a stochastic optimal control problem, which can be solved by direct application of dynamic programming algorithms. Focusing the problem in the adaptation of a hybrid DCT coder to a fixed bit-rate channel, examples of the three stages of the process (source-coder modelization, problem definition, and optimal policy computation) are provided. José Ignacio Ronda, Fernando Jaureguizar, Narciso García |
VCIP | 3 |
| 1990 | Hybrid DCT encoding of TV and HDTV: a comparative studyabstractThe different nature of the TV and HDTV signals, which hold dissimilar features from both the statistical and perceptual points of view, implies that the same behavior cannot be expected from the encoding procedure. Here, it is shown a comparison of the performances of a Hybrid DCT bit rate compression system when applied to information of each type. This study is carried out taking as a basis the statistics of the output data which result from the processing of corresponding HDTV and TV sequences, and it is described by several sets of data, relevant at different stages of the encoding process, such as DCT-domain energy distributions, quality versus bit rate curves, and final symbol statistics. Narciso García, José Ignacio Ronda, Fernando Jaureguizar |
VCIP | 1 |
| 1990 | Nonlinear spatial filtering of FLIR imagesabstractFLIR images contain a high level of noise, mainly caused by the acquisition system. Different classical non linear filters have been proposed to reduce this noise. Here, those having the best performance have been selected, studied in depth, and evaluated using statistical and subjective criteria. The best filters were AM and DWMTM, on which a detailed comparison study was carried out, finally choosing the AM one. María José Pérez-Luque, Narciso García |
VCIP | 3 |
| 1989 | Analysis of predictive schemes in pyramidal image codingabstractA model of the pyramid encoding generation algorithm is introduced, providing an approximation to the pyramid generation algorithm from which a theoretical expression for the expected prediction error can be derived. An expression for the improvement of the prediction error over standard predictive techniques is also obtained. Experimental results are provided, both to check the derived expressions and to test the method on real images. Quantization errors propagate throughout the pyramid, forcing a layer-dependent quantization mechanism, which is demonstrated. Overall results show good reconstructed images for a bit rate around 0.5 bit/pixel.> Alberto Sanz, Narciso García |
ICASSP | 3 |
| 1985 | Hierarchical predictive approach to image codingabstractProperties related with statistical redundancy of hierarchical image codes have been unveiled recently, allowing for a potential high bit rate reduction in storage or transmission. Appropriate schemes can be developed to take advantage of these properties, either for lossless or lossy applications. Here, a predictive approach for hierarchical line encoding is presented, giving rise to a two step procedure (hierarchical plus predictive). Both predictor and quantiser have been adapted to the different pixel weight within the hierarchical code; the selective quantiser has so been developed as a function of code outgrowth layer. Zero bit quantisation is also used to reduce bit rate in the large uniform areas of last code layers. Performance of the proposed scheme has been tested on a significative set of images, behaving well even with high entropy pictures. Results offer encoding at a bit rate around 0.5 bit/pixel, while subjective quality is still kept high. Alberto Sanz, Narciso García |
ICASSP | 3 |
| 1984 | Faster phase only image reconstructionabstractA novel starting approach for the phase only image reconstruction, based on the use of standard starting amplitude functions, is presented. These functions (hyperbolic, exponential) have a superior performance over other starting approaches, achieving the same reconstruction quality with considerably less number of iterations. The relative importance of phase and amplitude spectra is discussed for several image situations. Based on this discussion, the partial reconstruction approach is introduced, being it enough normally. Narciso García, Alberto R. Calero |
ICASSP | 1 |
| 1984 | Universal compression lossless code statistically builtabstractA universal compression statistical code based on a hierarchical transform of an image is presented. The properties of the transform that make it valuable for lossless image compression are studied. Based on them, a unique Huffman code, valid for every hierarchical transform, is constructed. Performance of the proposed coding scheme has been tested on a complete collection of images including small objects, faces, groups, remote sensing, ... obtained in different conditions. Results show that, on the average, a nearly 2:1 lossless compression is achieved, still keeping low the necessary computing requirements of the approach. Narciso García, Alberto Sanz |
ICASSP | 1 |
| 1984 | Approximation Quality Improvement Techniques in Progressive Image TransmissionabstractProgressive transmission of images has proved to be an efficient way of reaching effective bit rate reductions in highly interactive contexts such as telebrowsing systems, or low bandwidth communication channels. However, the lack of quality of the intermediate approximations has so far limited the obtained effective compression to a small factor. Three different and compatible techniques are presented to improve this quality: use of smoother interpolation schemes in its generation, a nonuniform decomposition procedure to favor the early refinement of relevant areas of the image, and an extension of the transmission hierarchy to pixel quantization producing successive refinement of the grey level resolution of the approximations. Combined use of these approaches yields effective bit rate reductions of up to 12.5:1, considerably extending the range and power of progressive transmission methods. Alberto Sanz, Narciso García |
IEEE J. Sel. Areas Commun. | 3 |
| 1982 | On the use of splines in hierarchical image transmissionabstractA quick identification of the received image is a desirable feature in low bandwith transmission systems. A gross quality approximation of the image is obtained at early stages when it is hierarchically encoded. This allows faster recognition than raster scanning methods, without transmission overhead. Spline interpolation facilitates this process, greatly improving the efficiency of hierarchical approaches. The same subjective impression is obtained with half of the information. Cubic convolution interpolation has all the advantages of the former, and can be more easily calculated. Recognition with less amount of received bits can be achieved with a non uniform development of the hierarchy. In order to automatize this procedure, two algorithms are presented with their advantages and drawbacks. In all the considered cases, transmission can be ended when the receiver decides. So, compression can be achieved at the cost of an approximation errors. Alberto Sanz, Narciso García |
ICASSP | 3 |