VLDB 2026 Research / reviewers in the wild / expert
Julián Cabrera
dblp:70/469
· DBLP profile ↗
41ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-7154-2451ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Computer networks · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quality assessment of 3D reconstructed meshes: Bridging objective metrics, subjective perception, and behavioral cuesabstractAssessing the quality of 3D reconstructed models remains a key challenge in multimedia applications, especially in the context of cultural heritage, where visual fidelity and perceptual realism are equally crucial. This study investigates how reconstruction parameters, as well as existing objective quality metrics, align with human perception. In addition, we analyze how perceived quality and user interaction are related. A dataset of 3D models was generated by varying the number of input images, mesh complexity, and texture resolution. Results from a subjective study show that texture resolution significantly affects perceived quality, whereas variations in number of images and mesh complexity have a limited impact. Furthermore, interaction behavior was found to vary with perceived quality, with participants spending more time and exploring larger viewing angles for models receiving higher scores. These findings highlight the need for perceptually grounded, interaction-aware evaluation methodologies and provide guidelines for future perceptual optimization of 3D reconstruction pipelines. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, David Barbero García, Daniel Berjón, Francisco Morán, Narciso García, Federica Battisti, Jesús Gutiérrez 0001, Marco Carli, Julián Cabrera |
Signal Process. Image Commun. | 12 |
| 2025 | Analysis of Objective 3D Mesh Quality Metrics for Cultural HeritageabstractExtended reality technologies are increasingly used in cultural heritage for preserving and accessing sites and artworks, where 3D model acquisition and rendering are key. Despite progress in reconstruction methodologies, a standardized approach to quality assessment is still missing. This study aims to evaluate objective quality metrics —both image-based and model-based, Full Reference and No Reference— applied to 3D models generated using the Structure from Motion algorithm. By varying parameters such as the number of images, number of triangles, and texture resolution, we examine the impact of these factors on metric outcomes, aiming to assess their reliability in cultural heritage applications. Anna Ferrarotti, Isabel Rodríguez, Javier Usón, Sara Baldoni, Jesús Gutiérrez 0001, Daniel Berjón, Francisco Morán, Federica Battisti, Narciso García, Marco Carli, Julián Cabrera |
QoMEX | 11 |
| 2025 | Enabling Real-Time Collaborative Cultural Experiences with Free Viewpoint VideoabstractIn this demo, we showcase a system that applies extended reality (XR) and volumetric video to enhance cultural experiences. The platform enables real-time interaction between users and a presenter within a virtual environment enriched with 3D assets. Users can visualize a volumetric representation of the presenter captured in real-time using a Free Viewpoint Video (FVV) system. Additionally, the presenter can control the elements in the scene using hand gestures, recognized by an artificial intelligence (AI) model. We propose a methodology to assess the subjective quality of this experience through user studies. The complete system will be demonstrated, including the real-time volumetric capture and the immersive application to participate in the cultural experience. Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Isabel Rodríguez, César Díaz, Jesús Gutiérrez 0001, Julián Cabrera |
QoMEX | 7 |
| 2024 | Analysis and Development of Deep Learning Depth Estimation Techniques for Volumetric Capture and Free Viewpoint VideoabstractVolumetric capture is an important topic in eXtended Reality (XR) as it enables the integration of realistic three-dimensional content into virtual scenarios and immersive applications. Certain systems are even capable of delivering these volumetric captures live and in real-time, opening the door to interactive use cases such as immersive videoconferencing. One example of such systems is FVV Live, a Free Viewpoint Video (FVV) application capable of working in real-time with low delay Javier Usón, Julián Cabrera |
MMSys | 2 |
| 2024 | Real-Time Free Viewpoint Video for Immersive VideoconferencingabstractIn this work, we propose a demo of an immersive videoconference system using Free Viewpoint Video (FVV) technology. It makes use of the FVV Live system, which covers the entire FVV pipeline (capture, view rendering, and visualization) while working in real-time. The FVV Live system consists of nine cameras that capture an environment and a view renderer that uses the information from the cameras to generate a synthetic view at an arbitrary point.It is designed as a hybrid demo. While the capture and rendering processes take place at our premises, FVV Live can be visualized through devices connected to the Internet.The system allows immersive navigation of a virtual scene with 6 degrees of freedom, and interaction with live-captured avatars integrated in such scene. For this purpose, it uses WebRTC connections to update the position of the virtual camera and to receive the FVV Live view encoded as a video.Additionally, the user will be recorded by a simple camera and microphone setup, and the generated streams will be transmitted to our premises through the same WebRTC server. This way, people being recorded by FVV Live will be able to see and hear the user, enabling bidirectional communication. Javier Usón, Victoria Muñoz, Carlos Cortés 0001, Daniel Berjón, Francisco Morán, César Díaz, Jesús Gutiérrez 0001, Fernando Jaureguizar, Narciso García, Julián Cabrera |
QoMEX | 10 |
| 2024 | Analysis and design framework for the development of indoor scene understanding assistive solutions for the person with visual impairment/blindnessabstractAbstract This paper discusses the challenges of the current state of computer vision-based indoor scene understanding assistive solutions for the person with visual impairment (P-VI)/blindness. It focuses on two main issues: the lack of user-centered approach in the development process and the lack of guidelines for the selection of appropriate technologies. First, it discusses the needs of users of an assistive solution through state-of-the-art analysis based on a previous systematic review of literature and commercial products and on semi-structured user interviews. Then it proposes an analysis and design framework to address these needs. Our paper presents a set of structured use cases that help to visualize and categorize the diverse real-world challenges faced by the P-VI/blindness in indoor settings, including scene description, object finding, color detection, obstacle avoidance and text reading across different contexts. Next, it details the functional and non-functional requirements to be fulfilled by indoor scene understanding assistive solutions and provides a reference architecture that helps to map the needs into solutions, identifying the components that are necessary to cover the different use cases and respond to the requirements. To further guide the development of the architecture components, the paper offers insights into various available technologies like depth cameras, object detection, segmentation algorithms and optical character recognition (OCR), to enable an informed selection of the most suitable technologies for the development of specific assistive solutions, based on aspects like effectiveness, price and computational cost. In conclusion, by systematically analyzing user needs and providing guidelines for technology selection, this research contributes to the development of more personalized and practical assistive solutions tailored to the unique challenges faced by the P-VI/blindness. Mohammad Moeen Valipoor, Angélica de Antonio Jiménez, Julián Cabrera |
Multim. Syst. | 3 |
| 2022 | Evaluation of the Performance of an Immersive System for Tele-educationabstractTele-education was already a solution for people who cannot attend lessons in person (such as inaccessibility in rural areas or illness issues). However, COVID has revealed problems in tele-education with current technology, causing adolescents and children to slow down their learning curves and experience problems of social distancing with their classmates. This paper presents a user study to validate an immersive communication system for tele-education purposes. This system streams in real time a class using 360-degree cameras, allowing remote students to explore the whole scene and improving the feeling of being in the classroom with their colleagues. Additionally, the prototype provides notifications to the remote students about events (such as a changes in the teacher’s presentation or classmates raising their hands) that occur outside their viewport to indicate in which direction they should move their heads to visualize them. Marta Orduna, Jesús Gutiérrez 0001, Alejandro Sánchez, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García |
IMX | 4 |
| 2022 | FVV Live: A Real-Time Free-Viewpoint Video System With Consumer Electronics HardwareabstractFVV Live is a novel end-to-end free-viewpoint video system, designed for real-time operation, using consumer-grade cameras and hardware, which enables low deployment costs and easy installation for immersive event-broadcasting or videoconferencing. All the blocks of the system have been designed to maximize perceptual video quality, overcoming the limitations imposed by hardware and network, which impact directly the accuracy of depth data and thus the quality of virtual view synthesis. Therefore, it does not sacrifice perceptual video quality with respect to high-end counterparts. The results presented in this paper correspond to an implementation with nine stereo-based depth cameras. However, the design of the acquisition block of FVV Live allows scalability for an arbitrary number of cameras. In addition, FVV Live presents low motion-to-photon and end-to-end delays, which enables a responsive free-viewpoint navigation and bilateral immersive communications. Moreover, the visual quality of FVV Live has been assessed through subjective assessment with satisfactory results, and additional comparative tests show that it is preferred over state-of-the-art DIBR alternatives. Pablo Carballeira, Carlos Carmona, César Díaz, Daniel Berjón, Daniel Corregidor, Julián Cabrera, Francisco Morán, Carmen Doblado, Sergio Arnaldo, María del Mar Martín, Narciso García |
IEEE Trans. Multim. | 6 |
| 2021 | EVENT-CLASS: Dataset of events in the classroomabstractThis work-in-progress presents a dataset of 360degree videos, called EVENT-CLASS, with associated characteristics in the context of tele-education. The sequences (video and audio) have been captured considering several environments, lighting conditions, acquisition perspectives, and cameras, enriching the dataset. EVENT-CLASS will be helpful for numerous applications related to tele-education, including quality assessment tests, and with the aim of improving the immersive experience of remote users thanks to the detection of relevant events that happen in the class. In this sense, this paper presents preliminary results of using transfer learning for person detection in 360degree scenes, based on Detectron2, and provides insights on the influence of applying it to equirectangular projection and over the viewport. Ongoing works will allow to include more videos and ground-truth annotations to the dataset. Marta Orduna, Jesús Gutiérrez 0001, Carlos Manzano, Julián Cabrera, César Díaz, Pablo Pérez 0001, Narciso García |
QoMEX | 5 |
| 2020 | XLR (piXel Loss Rate): A Lightweight Indicator to Measure Video QoE in IP NetworksabstractA novel Key Quality Indicator for video delivery applications, XLR (piXel Loss Rate), is defined, characterized, and evaluated. The proposed indicator is an objective measure that captures the effects of transmission errors in the received video, has a good correlation with subjective Mean Opinion Scores, and provides comparable results with state-of-the-art Full-Reference metrics. Moreover, XLR can be estimated using only a lightweight analysis on the compressed bitstream, thus allowing a No-Reference operational method. Therefore, XLR can be used for measuring the quality of experience without latency at any network location. Thus, it is a relevant tool for network planning, specially in new high-demanding scenarios. The experiments carried out show the outstanding performance of its linear-dimension score and the reliability of the bitstream-based estimation. César Díaz, Pablo Pérez 0001, Julián Cabrera, Jaime J. Ruiz, Narciso García |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Voronoi-based Objective Quality Metrics for Omnidirectional VideoabstractOmnidirectional video (ODV) represents one of the latest and most promising trends in immersive media. The success of ODV depends on the ability to deliver high-quality ODV to the viewers. For this reason, new methods are needed to measure ODV quality that takes into account the interactive look around nature and the spherical representation of ODV. In this paper, we study full-reference objective quality metrics for ODV based on typical encoding distortions in adaptive streaming systems, namely, scaling and compression. The contribution of this paper is three-fold. First, we propose new objective metrics that take into account the unique aspects of ODV. The proposed metrics are based on the subdivision of a given ODV into multiple patches using the spherical Voronoi diagram. Second, we introduce a new dataset of 75 impaired ODVs with different resolutions and compression levels, together with the subjective quality scores gathered during an experiment with 21 participants. Third, we evaluate the proposed Voronoi-based objective metrics using our dataset. The evaluation of the proposed objective metrics and the comparison with existing metrics show that the proposed metrics achieve a better correlation with the subjective scores. The ODV dataset together with the subjective quality scores and the code of the proposed quality metrics are available with this paper. Simone Croci, Cagri Ozcinar, Emin Zerman, Julián Cabrera, Aljoscha Smolic |
QoMEX | 4 |
| 2018 | Omnidirectional Video Streaming Using Visual Attention-Driven Dynamic Tiling for VRabstractThis paper proposes a new adaptive omnidirectional video (ODV) streaming system that uses visual attention (VA) maps. The proposed method benefits from a novel approach to VA-based bitrate allocation algorithm and dynamic tiling, providing enhanced virtual reality (VR) video experiences. The main contribution of this paper is the use of VA maps: (i) to distribute a given bitrate budget among a set of tiles of a given ODV and, (ii) to decide an optimal tiling structure (i.e., tile scheme) per chunk. For this, a novel objective metric is proposed: the visual attention spherical weighted (VASW) PSNR. This metric operates in the spherical domain and by means of a VA probabilistic model aims at capturing the quality of the actual areas observed by the users when navigating through the ODV content. We evaluate the proposed system performance with varying bandwidth conditions and the tracked head orientations from disjoint user experiments. Results show that the proposed system significantly outperforms the existing tiled-based streaming method. Cagri Ozcinar, Julián Cabrera, Aljoscha Smolic |
VCIP | 2 |
| 2016 | Analysis of the depth-shift distortion as an estimator for view synthesis distortion
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
Signal Process. Image Commun. | 2 |
| 2015 | An extension to the PRO-MPEG COP3 codes for unequal error protection of real-time video transmissionabstractWe propose and evaluate an extension to the Application-Layer FEC (AL-FEC) codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes), consisting in allowing the use of a number of matrices of dissimilar size per FEC block. So, unequal protection of the data packets in the video stream is enabled, since dissimilar code rates can be applied to different groups of data packets. This boosts the efficiency of the protection scheme, increasing the video quality of the sequence presented to final users, even if the resulting packet loss rate (PLR) after channel decoding remains the same. Evaluation results show a significantly better performance of the Pro-MPEG COP3 codes when the proposed protection extension is incorporated. César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ICIP | 2 |
| 2015 | Q-learning based control algorithm for HTTP adaptive streamingabstractWe present a control algorithm based on Q-Learning for an HTTP Adaptive Streaming (HAS) Client in order to optimize the Quality of Experience (QoE) of the user. First, we propose a model with a suitable number of variables in an attempt to find a reasonable tradeoff between the complexity of the model and its capacity to capture appropriately the dynamics of the system. Second, we define a novel reward function that takes into consideration factors related to the user's QoE. Results will show, that our Q-learning algorithm is able to learn and efficiently control the selection of the segment qualities. In addition, we will show that our proposed approach outperforms another Q-learning approach. Virginia Martin, Julián Cabrera, Narciso García |
VCIP | 2 |
| 2014 | Systematic analysis of the decoding delay in multiview video
Pablo Carballeira, Julián Cabrera, Fernando Jaureguizar, Narciso García |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Depth video coding for free viewpoint video oriented to the synthetic view perceptual qualityabstractIn this paper we propose a novel depth encoding algorithm based on the synthetic view perceptual quality optimization. In Free Viewpoint Video, depth sequences are never shown to the observer but they are used for the view synthesis. Due to the compression quality losses, the depth distortion generates errors in the synthetic view pixel position. We propose to encode the depth sequences by a novel approach, divided into two parts. The first one is a rate distortion optimization method based on the minimization of the synthetic view pixel position error. The second one is an extension of the first one and improves the performance taking into account the perceptual distortion, evaluated by exploiting the Human Visual System characteristics described by the Just Noticeable Distortion. However, the Just Noticeable Distortion is a model designed for the traditional video, hence we also propose a Just Noticeable Pixel Displacement evaluation method which, considering the synthetic view perceptual distortion, is able to estimate when an error in the synthetic view pixel position is not noticed. The results show how the proposed algorithm not only improves the synthetic view quality measured by the Video Quality Metric (highly correlated to the subjective quality) but also the objective quality measured by the PSNR, with an achieved improvement of 0.3 dB with a corresponding bit saving of 13%. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 3 |
| 2013 | Enhancement of Pro-MPEG COP3 codes and application to layer-aware FEC protection of two-layered video transmissionabstractIn this paper we propose an enhancement of the Application-Layer FEC codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes) through allowing the introduction of a third dimension. The potential addition of an extra set of protection packets augments the number of possible combinations of data packets within a FEC block for parity packet computation. This enables a finer optimization process of the parameters of the FEC codes for a better adaptation to the specific conditions of the communication channel, increasing their capability. Additionally, we propose a Layer-Aware FEC scheme in which the enhanced Pro-MPEG COP3 codes are used to protect two-layered video streams. Experiment results reveal a gain in the introduction of this protection mechanism, when compared to the standard codes. César Díaz, Cornelius Hellge, Julián Cabrera, Fernando Jaureguizar, Thomas Schierl |
ICIP | 3 |
| 2013 | Stochastic modelling of peer-assisted VoD streaming in managed networks
Sasho Gramatikov, Fernando Jaureguizar, Julián Cabrera, Narciso García |
Comput. Networks | 3 |
| 2013 | Low Complexity Mode Decision and Motion Estimation for H.264/AVC Based Depth Maps Encoding in Free Viewpoint VideoabstractWithin free viewpoint video, the 3-D reconstruction of the scene is created from a high number of viewpoints. Every viewpoint is represented by a traditional sequence, called texture, and its associated depth information. This is known as a View plus Depth environment. In this paper, a novel low complexity mode decision and motion estimation algorithm for the H.264/AVC based encoding of depth sequences is proposed. Given that a texture sequence and its associated depth represent the same scene from the same point of view, they should have similar motion characteristics. The complexity reduction of the depth encoding, in the proposed algorithm, is obtained by taking advantage of the texture motion information that has been previously processed by a traditional H.264/AVC encoder. The characteristics of depth and texture sequences are analyzed, focusing on similarities and differences that are properly managed to design an algorithm able to detect when the motion information of the texture might be usefully exploited in the encoding of the corresponding depth sequence. The proposed method is able to achieve the same objective quality, measured by means of the PSNR and the VQM, than a traditional H.264/AVC encoder with a reduction of up to 58% of the computational burden. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Subjective study of adaptive streaming strategies for 3DTVabstractAlthough the delivery of 3D video services to households is nowadays a reality thanks to frame-compatible formats, many efforts are being made to obtain efficient methods to transmit 3D content offering a high quality of experience to the end users. In this paper, a stereoscopic video streaming scenario is considered, and the perceptual impact of various strategies applicable to adaptive streaming situations are compared. Specifically, the mechanisms are based on switching between copies of the content with different coding qualities, on discarding frames of the sequence, on switching from 3D to 2D, and on using asymmetric coding of the stereo views. In addition, when video freezes happen, the possibility of keeping the end-to-end latency or maintaining the continuity of the video are considered. These aspects were evaluated carrying out a subjective assessment test, considering also visual discomfort issues, using a methodology designed to keep as far as possible domestic viewing conditions. Jesús Gutiérrez 0001, Pablo Pérez 0001, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 4 |
| 2012 | Depth perceptual video coding for free viewpoint video based on H.264/AVCabstractA novel scheme for depth sequences compression, based on a perceptual coding algorithm, is proposed. A depth sequence describes the object position in the 3D scene, and is used, in Free Viewpoint Video, for the generation of synthetic video sequences. In perceptual video coding the human visual system characteristics are exploited to improve the compression efficiency. As depth sequences are never shown, the perceptual video coding, assessed over them, is not effective. The proposed algorithm is based on a novel perceptual rate distortion optimization process, assessed over the perceptual distortion of the rendered views generated through the encoded depth sequences. The experimental results show the effectiveness of the proposed method, able to obtain a very considerable improvement of the rendered view perceptual quality. Gianluca Cernigliaro, Matteo Naccari, Fernando Jaureguizar, Julián Cabrera, Narciso García |
PCS | 4 |
| 2012 | Adaptive protection scheme for MVC-encoded stereoscopic video streaming in IP-based networksabstractWe present an adaptive unequal error protection (UEP) strategy built on the 1-D interleaved parity Application Layer Forward Error Correction (AL-FEC) code for protecting the transmission of stereoscopic 3D video content encoded with Multiview Video Coding (MVC) through IP-based networks. Our scheme targets the minimization of quality degradation produced by packet losses during video transmission in time-sensitive application scenarios. To that end, based on a novel packet-level distortion model, it selects in real time the most suitable packets within each Group of Pictures (GOP) to be protected and the most convenient FEC technique parameters, i.e., the size of the FEC generator matrix. In order to make these decisions, it considers the relevance of the packet, the behavior of the channel, and the available bitrate for protection purposes. Simulation results validate both the distortion model introduced to estimate the importance of packets and the optimization of the FEC technique parameter values. César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García |
VCIP | 2 |
| 2011 | A new fast motion estimation and mode decision algorithm for H.264 depth maps encoding in free viewpoint TVabstractIn this paper, we consider a scenario where 3D scenes are modeled through a View+Depth representation. This representation is to be used at the rendering side to generate synthetic views for free viewpoint video. The encoding of both type of data (view and depth) is carried out using two H.264/AVC encoders. In this scenario we address the reduction of the encoding complexity of depth data. Firstly, an analysis of the Mode Decision and Motion Estimation processes has been conducted for both view and depth sequences, in order to capture the correlation between them. Taking advantage of this correlation, we propose a fast mode decision and motion estimation algorithm for the depth encoding. Results show that the proposed algorithm reduces the computational burden with a negligible loss in terms of quality of the rendered synthetic views. Quality measurements have been conducted using the Video Quality Metric. Gianluca Cernigliaro, Matteo Naccari, Fernando Jaureguizar, Julián Cabrera, Fernando Pereira 0001, Narciso García |
ICIP | 4 |
| 2011 | Inter-packet symbol approach to Reed-Solomon FEC codes for RTP-multimedia stream protectionabstractThis paper presents an alternative Forward Error Correction scheme, based on Reed-Solomon codes, with the aim of protecting the transmission of RTP-multimedia streams: the inter-packet symbol approach. This scheme is based on an alternative bit structure that allocates each symbol of the Reed-Solomon code in several RTP-media packets. This characteristic permits to exploit better the recovery capability of Reed-Solomon codes against bursty packet losses. The performance of our approach has been studied in terms of encoding/decoding time versus recovery capability, and compared with other proposed schemes in the literature. The theoretical analysis has shown that our approach allows the use of a lower size of the Galois Fields compared to other solutions. This lower size results in a decrease of the required encoding/decoding time while keeping a comparable recovery capability. Finally, experimental results have been carried out to assess the performance of our approach compared to other schemes in a simulated environment, where models for wireless and wireline channels have been considered. Filippo Casu, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ISCC | 2 |
| 2011 | Subjective evaluation of transmission errors in IPTV and 3DTVabstractThe increase of multimedia services delivered over packet-based networks has entailed greater quality expectations of the end-users. This has led to an intensive research on techniques for evaluating the quality of experience perceived by the viewers of audiovisual content, considering the different degradations that it could suffer along the broadcasting system. In this paper, a comprehensive study of the impact of transmission errors affecting video and audio in IPTV is presented. With this aim, subjective assessment tests were carried out proposing a novel methodology trying to keep as close as possible home environment viewing conditions. Also 3DTV content in side-by-side format has been used in the experiments to compare the impact of the degradations. The results provide a better understanding of the effects of transmission errors, and show that the QoE related to the first approach of 3DTV is acceptable, but the visual discomfort that it causes should be reduced. Jesús Gutiérrez 0001, Pablo Pérez 0001, Fernando Jaureguizar, Julián Cabrera, Narciso García |
VCIP | 4 |
| 2010 | Fast mode decision for multiview video coding based on scene geometryabstractA new fast mode decision (FMD) algorithm for multi-view video coding (MVC) is presented. The codification of the views is based on the analysis of the homogeneity of the depth map and corrected with the motion analysis of a reference view, which is encoded based on traditional methods and on the use of the disparity differences between the views. This approach reduces the burden of the rate-distortion motion analysis using the availability of a depth map and the presence of the disparity vectors, which are assumed to be provided by the acquisition process. Gianluca Cernigliaro, Fernando Jaureguizar, Julián Cabrera, Narciso García |
ICIP | 3 |
| 2010 | A wireless video transmission control approach through Stochastic Dynamic ProgrammingabstractThis paper presents an intelligent, rate-limited multicast video transmission optimization scheme for video distribution over 802.11 wireless networks based on packet retransmissions. We propose a problem formulation which involves the characteristics of the encoded stream together with the behaviour of the wireless channel. Using standard Stochastic Dynamic Programming techniques, optimal control policies are obtained off-line. These policies are optimal in the sense of minimizing the expected distortion at the terminal. In addition, the on-line complexity or our approach is very low since the optimization problem is solved off-line. The performance of our scheme has been evaluated in a real scenario and compared with that of a limited rate ARQ algorithm. The results for our proposed system show a higher packet recovery rate and a better protection of information with a higher priority. Victor Miguel, Julián Cabrera, Fernando Jaureguizar, Narciso García |
ICIP | 2 |
| 2007 | Combined Markov Model for Packet Loss Characterization in UMTS ChannelsabstractMobile communication channels such as UMTS channels are affected by errors which show certain statistical dependencies. In this paper we tackle the characterization of error occurrences from a packet level perspective. We present a novel approach which consists of a generative model based on the interconnection of a Markov gap model, aimed to characterize error-free bursts, with several hidden Markov models (HMMs), which characterize the distribution of bursts with erroneous packets. We demonstrate that this model provides accurate predictions on UMTS radio links and compare its accuracy with that of traditional models based on Markov chains. Victor Gomis, Julián Cabrera, Juan C. Plaza, Fernando Jaureguizar |
VTC Fall | 2 |
| 2007 | A New Class Partitioned Discrete Model for the Characterization of RTP Multicast Transmission through IEEE 802.11 ChannelsabstractMulticast of real time protocol (RTP) flows (i.e. multicast video streaming) within 802.11 wireless local area networks (WLANs) has to deal with the inherent unreliability of wireless channels and the lack of link-level error protection of multicast transmissions. This problem can be solved in part by means of sophisticated rate control strategies, but no efficient discrete channel models for this scenario have been developed yet. In this paper a novel packet error rate class partitioned (PERCP) model for the characterization of wireless channels is proposed. Its characteristics make it suitable for its integration in rate-control schemes, and it is proved to perform better than simplified Gilbert models in all considered scenarios. Juan C. Plaza, Julián Cabrera, Fernando Jaureguizar, Narciso García |
VTC Fall | 2 |
| 2006 | Fast Mode Decision on H.264/AVC Main Profile Encoding Based on PSNR PredictionsabstractIn this paper we propose a new fast mode decision (FMD) algorithm to reduce the computational load of the motion estimation (ME) process of the new video coding standard H.264/AVC main profile. The algorithm firstly computes the skip/direct mode (respectively for P and B-frames), for all macroblocks (MBs), and then decides, for each MB if no mores modes are needed. This decision is based on the generation of predictions of the expected PSNR of the frame to be encoded. These predictions are performed with the distortion cost obtained for skip/direct mode jointly with the average distortion cost obtained for the previous encoded frames. Results show that the computational load have been dramatically reduced, with average encoding time saving of-36.38%, while nearly negligible loss of coding efficiency. Marcos Nieto Doncel, Luis Salgado, Julián Cabrera |
ICIP | 3 |
| 2005 | Fast Mode Decision and Motion Estimation with Object Segmentation in H.264/AVC Encoding
Marcos Nieto Doncel, Luis Salgado, Julián Cabrera |
ACIVS | 3 |
| 2005 | A unified model for techniques on video-shot transition detectionabstractA first step required to allow video indexing and retrieval of visual data is to perform a temporal segmentation, that is, to find the location of camera-shot transitions, which can be either abrupt (i.e., cuts) or gradual (e.g., fades, dissolves, wipes). After a critical review of most approaches seeking to solve this problem, we propose a unified detection model (both for abrupt and all types of gradual transitions), as well as an implementation whose results improve upon those of all the inspected reports. The innovation of the approach presented here is centered on mapping the space of inter-frame distances onto a new space of decision better suited to achieving a sequence-independent thresholding. This mapping aims to consider frame ordering information within the thresholding process; it is based on the parametric modeling of the patterns that transitions generate on the distances' output. As opposed to most reviewed works, our results are detailed over a large and representative sample of more than 1500 cuts and 250 gradual transitions, which make up a significant part (200 min) of the MPEG-7 testing material; this ensures a high degree of confidence in the validity of our approach. Jesús Bescós, Guillermo Cisneros, José María Martínez Sanchez, José Manuel Menéndez, Julián Cabrera |
IEEE Trans. Multim. | 5 |
| 2004 | Adaptive segmentation for gymnastic exercises based on change detection over multiresolution combined differencesabstractA new adaptive segmentation strategy is proposed to segment gymnasts in sport sequences accurately. It is based on a Markov random fields (MRF) change detection analysis operating on a multiresolution combination of static and dynamic image differences. After a morphological analysis of the segmented masks, estimated motion information in the area of interest is incorporated to improve the efficiency of the segmentation process. Although presented in the particular context of gymnastic exercises, the new segmentation strategy could be applied to other applications where moving objects on a quasi-static background need to be segmented. María José Aramburu Cabo, Luis Salgado, Julián Cabrera |
ICIP | 3 |
| 2004 | Unsupervised segmentation algorithm of hrtem images
Ainhoa Mendizabal, Julián Cabrera, Luis Salgado, Narciso García, Juan C. Gonzalez |
ICIP | 2 |
| 2003 | Stochastic rate-control of interframe video coders for VBR channelsabstractWe propose a new algorithm for the real-time control of an inter-frame video coder operating with a variable rate channel such as wireless channels or the Internet. Using techniques of stochastic dynamic programming we obtain off-line optimal policies from stochastic models of the channel and coder which minimize the average expected distortion. The on-line complexity of our approach is only that required to identify the state of the system (source and channel). The state of the channel is obtained based on the ARQ error-control mechanism, and the source state is computed as complexity measurements on each incoming frame. Simulation results based on this new approach are provided and compared to other proposed rate-control strategies. They show how our model-based optimal policies outperform the other considered approaches keeping a negligible on-line computational cost. This result is very interesting when considering an alternative to traditional costly solutions based on deterministic dynamic programming. Julián Cabrera, José Ignacio Ronda, Antonio Ortega, Narciso García |
ICIP (3) | 1 |
| 2003 | MISS: A Generic Model for MetaInformation SubSystems
José María Martínez Sanchez, Julián Cabrera, Jesús Bescós, Guillermo Cisneros, José Manuel Menéndez |
Multim. Tools Appl. | 2 |
| 2002 | Stochastic rate-control of video coders for wireless channelsabstractWe introduce a new approach to deal with the transmission of real-time video over wireless channels, based on a priori stochastic models for both source and channel. This new problem formulation captures in a natural way the stochastic nature of the channel as well as the uncertainty regarding the properties of the video sequence. Our formulation leads to an optimal control problem that can be solved off-line, employing standard stochastic dynamic programming techniques. The outcome of this optimization is an off-line control policy that is optimal in the sense of minimizing the average coding distortion. The on-line computational cost of the new approach is thus very low: all that is required during run-time is to identify the state of the system (source and channel). Unlike other optimization-based rate control techniques, which require a search for the optimal operating point, the operating points here for each allowable state of the system have been precalculated. We consider wireless packet-based transmission with Automatic Repeat reQuest (ARQ) error control. While a standard model has been adopted to characterize the channel behavior, a new model based on the concept of coding complexity has been devised in order to characterize the video source. Simulation results based on this new approach are provided and compared to other proposed rate-control strategies. They show how the use of model-based optimal policies has negligible on-line computational cost while providing a transmission quality comparable to that achieved with more costly deterministic dynamic programming techniques, and significantly better than for simpler algorithms that do not explicitly take into account the channel state. Julián Cabrera, José Ignacio Ronda, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Stochastic rate-control of video coders for wireless channelsabstractWe study the transmission of real-time video over wireless channels, proposing a formulation of the problem that includes a priori stochastic models for both source and channel. Using techniques of stochastic dynamic programming, we obtain offline optimal policies for each system state that minimize the average expected frame distortion. The online complexity of our approach is only that required to identify the state of the system (source and channel). The state of the channel is obtained based on the ARQ error-control mechanism, and the source state is computed as a complexity measurement on each incoming frame. Simulation results based on this new approach are provided and compared to other proposed rate-control strategies. They show how our model-based optimal policies require negligible on-line computational cost while providing a transmission quality comparable to that achieved with more complex deterministic dynamic programming techniques, and better than for simpler algorithms such as TMN. Julián Cabrera, José Ignacio Ronda, Antonio Ortega |
ICIP (1) | 1 |
| 2000 | A Unified Approach to Gradual Shot Transition DetectionabstractNowadays, content-based retrieval of video material is based on the availability of meta-data linked to it. Current approaches to automatically extract these data start from a temporal segmentation of the audio-visual material, that is, a location of the camera shot transitions. Although abrupt transition detection is a problem almost solved, success on gradual transition detection is still very low. We present a detector for all types of gradual transitions (chromatic, geometric, and mixed), based on in depth modelling of the patterns that these transitions generate on a specific frame distance. Results are presented over a representative sample of more than 250 gradual transitions, which belong to a significant part (200') of the MPEG-7 testing material, achieving a recall (percentage of effects correctly detected) and a precision (percentage of non-false positives in the set of detection) notably higher than the ones reported so far. Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros, Julián Cabrera, José María Martínez Sanchez |
ICIP | 4 |
| 1999 | Generic meta data browsing system for multimedia document retrievalabstractThis paper presents the architecture of a multimedia information-indexing subsystem, the Meta-information subsystem. This subsystem makes use of meta data, that is, content-based information extracted from multimedia materials and other associated information. It also provides new mechanisms to store, query and browse meta data, and manages their relationship with the multimedia materials stored in the system. Julián Cabrera, José María Martínez Sanchez, Jesús Bescós, José Manuel Menéndez, Guillermo Cisneros |
MMSP | 1 |