César Díaz

dblp:169/3881 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0003-2030-9390ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enabling Real-Time Collaborative Cultural Experiences with Free Viewpoint Video
abstract
In this demo, we showcase a system that applies extended reality (XR) and volumetric video to enhance cultural experiences. The platform enables real-time interaction between users and a presenter within a virtual environment enriched with 3D assets. Users can visualize a volumetric representation of the presenter captured in real-time using a Free Viewpoint Video (FVV) system. Additionally, the presenter can control the elements in the scene using hand gestures, recognized by an artificial intelligence (AI) model. We propose a methodology to assess the subjective quality of this experience through user studies. The complete system will be demonstrated, including the real-time volumetric capture and the immersive application to participate in the cultural experience.
Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Isabel Rodríguez, César Díaz, Jesús Gutiérrez 0001, Julián Cabrera
QoMEX5
2024 Real-Time Free Viewpoint Video for Immersive Videoconferencing
abstract
In this work, we propose a demo of an immersive videoconference system using Free Viewpoint Video (FVV) technology. It makes use of the FVV Live system, which covers the entire FVV pipeline (capture, view rendering, and visualization) while working in real-time. The FVV Live system consists of nine cameras that capture an environment and a view renderer that uses the information from the cameras to generate a synthetic view at an arbitrary point.It is designed as a hybrid demo. While the capture and rendering processes take place at our premises, FVV Live can be visualized through devices connected to the Internet.The system allows immersive navigation of a virtual scene with 6 degrees of freedom, and interaction with live-captured avatars integrated in such scene. For this purpose, it uses WebRTC connections to update the position of the virtual camera and to receive the FVV Live view encoded as a video.Additionally, the user will be recorded by a simple camera and microphone setup, and the generated streams will be transmitted to our premises through the same WebRTC server. This way, people being recorded by FVV Live will be able to see and hear the user, enabling bidirectional communication.
Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Daniel Berjón, Francisco Morán, César Díaz, Jesús Gutiérrez 0001, Fernando Jaureguizar, Narciso García, Julián Cabrera
QoMEX6
2022 Evaluation of the Performance of an Immersive System for Tele-education
abstract
Tele-education was already a solution for people who cannot attend lessons in person (such as inaccessibility in rural areas or illness issues). However, COVID has revealed problems in tele-education with current technology, causing adolescents and children to slow down their learning curves and experience problems of social distancing with their classmates. This paper presents a user study to validate an immersive communication system for tele-education purposes. This system streams in real time a class using 360-degree cameras, allowing remote students to explore the whole scene and improving the feeling of being in the classroom with their colleagues. Additionally, the prototype provides notifications to the remote students about events (such as a changes in the teacher’s presentation or classmates raising their hands) that occur outside their viewport to indicate in which direction they should move their heads to visualize them.
Marta Orduna, Jesús Gutiérrez 0001, Alejandro Sánchez, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García
IMX5
2022 A Novel System for Nighttime Vehicle Detection Based on Foveal Classifiers With Real-Time Performance
abstract
Vehicle monitoring using camera networks is an important task for traffic applications. Moreover, it becomes critical in nighttime, when the probability of an accident considerably increases as visibility conditions worsen. Typical approaches are mostly based on the assumption that regions delimiting vehicle lights are well defined, so that they are segmented and then associated to vehicle entities. However, this assumption fails in images acquired by existing traffic camera networks, where vehicle lights are revealed as flashes and other complex light patterns, occupying large and even disconnected image regions. In this work, a real-time vehicle detection algorithm for nighttime situations has been presented, which is able to locate vehicles in the image by analyzing the previous complex light patterns. For this purpose, a novel machine learning framework based on a grid of foveal classifiers has been designed. Every classifier in the grid processes the same global image descriptor (only one descriptor is computed per image). However, every one of them is trained to predict a different output depending on the classifier position in the grid and the vehicle ground-truth location. Additionally, only point-based annotations are required to train the grid of foveal classifiers, speeding up the cost of creating the required databases. Experimental results prove the effectiveness of the proposed method in a new created nighttime database with point-based annotations.
Andrés Bell, Tomás Mantecón, César Díaz, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García
IEEE Trans. Intell. Transp. Syst.3
2022 FVV Live: A Real-Time Free-Viewpoint Video System With Consumer Electronics Hardware
abstract
FVV Live is a novel end-to-end free-viewpoint video system, designed for real-time operation, using consumer-grade cameras and hardware, which enables low deployment costs and easy installation for immersive event-broadcasting or videoconferencing. All the blocks of the system have been designed to maximize perceptual video quality, overcoming the limitations imposed by hardware and network, which impact directly the accuracy of depth data and thus the quality of virtual view synthesis. Therefore, it does not sacrifice perceptual video quality with respect to high-end counterparts. The results presented in this paper correspond to an implementation with nine stereo-based depth cameras. However, the design of the acquisition block of FVV Live allows scalability for an arbitrary number of cameras. In addition, FVV Live presents low motion-to-photon and end-to-end delays, which enables a responsive free-viewpoint navigation and bilateral immersive communications. Moreover, the visual quality of FVV Live has been assessed through subjective assessment with satisfactory results, and additional comparative tests show that it is preferred over state-of-the-art DIBR alternatives.
Pablo Carballeira, Carlos Carmona, César Díaz, Daniel Berjón, Daniel Corregidor, Julián Cabrera, Francisco Morán, Carmen Doblado, Sergio Arnaldo, María del Mar Martín, Narciso García
IEEE Trans. Multim.3
2022 Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919
abstract
Recently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community.
Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García
IEEE Trans. Multim.19
2021 EVENT-CLASS: Dataset of events in the classroom
abstract
This work-in-progress presents a dataset of 360degree videos, called EVENT-CLASS, with associated characteristics in the context of tele-education. The sequences (video and audio) have been captured considering several environments, lighting conditions, acquisition perspectives, and cameras, enriching the dataset. EVENT-CLASS will be helpful for numerous applications related to tele-education, including quality assessment tests, and with the aim of improving the immersive experience of remote users thanks to the detection of relevant events that happen in the class. In this sense, this paper presents preliminary results of using transfer learning for person detection in 360degree scenes, based on Detectron2, and provides insights on the influence of applying it to equirectangular projection and over the viewport. Ongoing works will allow to include more videos and ground-truth annotations to the dataset.
Marta Orduna, Jesús Gutiérrez 0001, Carlos Manzano, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García
QoMEX6
2020 XLR (piXel Loss Rate): A Lightweight Indicator to Measure Video QoE in IP Networks
abstract
A novel Key Quality Indicator for video delivery applications, XLR (piXel Loss Rate), is defined, characterized, and evaluated. The proposed indicator is an objective measure that captures the effects of transmission errors in the received video, has a good correlation with subjective Mean Opinion Scores, and provides comparable results with state-of-the-art Full-Reference metrics. Moreover, XLR can be estimated using only a lightweight analysis on the compressed bitstream, thus allowing a No-Reference operational method. Therefore, XLR can be used for measuring the quality of experience without latency at any network location. Thus, it is a relevant tool for network planning, specially in new high-demanding scenarios. The experiments carried out show the outstanding performance of its linear-dimension score and the reliability of the bitstream-based estimation.
César Díaz, Pablo Pérez 0001, Julián Cabrera, Jaime J. Ruiz, Narciso García
IEEE Trans. Netw. Serv. Manag.1
2019 Perceptually Equivalent Resolution in Handheld Devices for Streaming Bandwidth Saving
abstract
We present the description, results, and analysis of the experiments conducted to find the equivalent resolution associated with handheld devices. That is, the resolution from which users stop perceiving quality improvements if better resolutions are presented to them in such devices. Thus, it is the maximum resolution that it is worth considering for generating and delivering video, as long as sequences are not too intensively compressed. Therefore, the detection of the equivalent resolutions allows for notable savings in bandwidth consumption. Subjective assessments have been carried out on fifty subjects using a set of video sequences of very different nature and four handheld devices with a broad range of screen dimensions. The results prove that the equivalent resolution in current handheld devices is 720p as higher resolutions are not valued by users.
Mateo Camara, César Díaz, Juan Casal, Jorge Ruano, Narciso García
IEEE Signal Process. Lett.2
2015 An extension to the PRO-MPEG COP3 codes for unequal error protection of real-time video transmission
abstract
We propose and evaluate an extension to the Application-Layer FEC (AL-FEC) codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes), consisting in allowing the use of a number of matrices of dissimilar size per FEC block. So, unequal protection of the data packets in the video stream is enabled, since dissimilar code rates can be applied to different groups of data packets. This boosts the efficiency of the protection scheme, increasing the video quality of the sequence presented to final users, even if the resulting packet loss rate (PLR) after channel decoding remains the same. Evaluation results show a significantly better performance of the Pro-MPEG COP3 codes when the proposed protection extension is incorporated.
César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García
ICIP1
2013 Enhancement of Pro-MPEG COP3 codes and application to layer-aware FEC protection of two-layered video transmission
abstract
In this paper we propose an enhancement of the Application-Layer FEC codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes) through allowing the introduction of a third dimension. The potential addition of an extra set of protection packets augments the number of possible combinations of data packets within a FEC block for parity packet computation. This enables a finer optimization process of the parameters of the FEC codes for a better adaptation to the specific conditions of the communication channel, increasing their capability. Additionally, we propose a Layer-Aware FEC scheme in which the enhanced Pro-MPEG COP3 codes are used to protect two-layered video streams. Experiment results reveal a gain in the introduction of this protection mechanism, when compared to the standard codes.
César Díaz, Cornelius Hellge, Julián Cabrera, Fernando Jaureguizar, Thomas Schierl
ICIP1
2012 Adaptive protection scheme for MVC-encoded stereoscopic video streaming in IP-based networks
abstract
We present an adaptive unequal error protection (UEP) strategy built on the 1-D interleaved parity Application Layer Forward Error Correction (AL-FEC) code for protecting the transmission of stereoscopic 3D video content encoded with Multiview Video Coding (MVC) through IP-based networks. Our scheme targets the minimization of quality degradation produced by packet losses during video transmission in time-sensitive application scenarios. To that end, based on a novel packet-level distortion model, it selects in real time the most suitable packets within each Group of Pictures (GOP) to be protected and the most convenient FEC technique parameters, i.e., the size of the FEC generator matrix. In order to make these decisions, it considers the relevance of the packet, the behavior of the channel, and the available bitrate for protection purposes. Simulation results validate both the distortion model introduced to estimate the importance of packets and the optimization of the FEC technique parameter values.
César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García
VCIP1