EDBT 2026 Demo / reviewers in the wild / expert
Yago Sánchez de la Fuente
dblp:17/8235 · also Yago Sanchez
· DBLP profile ↗
34ranked-venue papers
13as first author
8since 2021 · last 2026
0000-0003-2950-8900ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 8 since 2021Computer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CodecGS-IR: Implicit Representation for Decoder-friendly Video-based Gaussian Splat CompressionabstractCoding of Gaussian splats has drawn the attention of academia and standardization bodies lately. A commonly used approach involves projecting explicit 3D attributes onto 2D planes of a video and compressing it with existing video coding standards. However, such an approach produces excessively high sample rates, resulting in a significant bottleneck that might make it unsuitable for existing hardware decoders if the number of splats is high. This paper presents a framework that utilizes an implicit representation, where Gaussian splat attributes are represented by compact feature planes. By reducing the dependency of video resolution on the number of splats, our approach significantly reduces the required sample rate, making it suitable for deployed devices. In addition, the proposed framework achieves a superior rate-distortion trade-off, providing high-fidelity reconstruction at low bitrates without the excessive sample rates associated with the conventional video-based anchor. Soonbin Lee, Simon Sasse, Yago Sánchez de la Fuente, Robert Skupin, Tomas M. Borges, Cornelius Hellge, Thomas Schierl |
DCC | 3 |
| 2025 | Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video Codecsabstract3D Gaussian Splatting is a recognized method for 3D scene representation, known for its high rendering quality and speed. However, its substantial data requirements present challenges for practical applications. In this paper, we introduce an efficient compression technique that significantly reduces storage overhead by using compact representation. We propose a unified architecture that combines point cloud data and feature planes through a progressive tri-plane structure. Our method utilizes 2D feature planes, enabling continuous spatial representation. To further optimize these representations, we incorporate entropy modeling in the frequency domain, specifically designed for standard video codecs. We also propose channel-wise bit allocation to achieve a better trade-off between bitrate consumption and feature plane representation. Consequently, our model effectively leverages spatial correlations within the feature planes to enhance rate-distortion performance using standard, non-differentiable video codecs. Experimental results demonstrate that our method outperforms existing methods in data compactness while maintaining high rendering quality. Our project page is available at https://fraunhoferhhi.github.io/CodecGS Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge |
ICCV | 3 |
| 2025 | Digital signatures for trustworthy authentication of elementary video streamsabstractThis paper introduces a method for digitally signing and verifying elementary video bitstreams. The method verifies temporal consistency of the video while allowing random access into the bitstream and adaptation to temporal and spatial scalability including sub-bitstream extraction. It was adopted by the Joint Video Experts Team (JVET) into the Versatile supplemental enhancement information messages for coded video bitstreams (VSEI) specification version 4 by introducing three new Digitally Signed Content SEI messages. Karsten Sühring, Tobias Hinz, Jonathan Pfaff, Yago Sánchez de la Fuente, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
VCIP | 4 |
| 2024 | ECRF: Entropy-Constrained Neural Radiance Fields Compression with Frequency Domain OptimizationabstractExplicit feature-grid based NeRF models have shown promising results in terms of rendering quality and significant speed-up in training. However, these methods often require a significant amount of data to represent a single scene or object. In this work, we present a compression model that aims to minimize the entropy in the frequency domain in order to effectively reduce the data size. First, we propose using the discrete cosine transform (DCT) on the tensorial radiance fields to compress the feature-grid. This feature-grid is transformed into coefficients, which are then quantized and entropy encoded, following a similar approach to the traditional video coding pipeline. Furthermore, to achieve a higher level of sparsity, we propose using an entropy parameterization technique for the frequency domain, specifically for DCT coefficients of the feature-grid. Since the transformed coefficients are optimized during the training phase, the proposed model does not require any fine-tuning or additional information. Our model only requires a lightweight compression pipeline for encoding and decoding, making it easier to apply volumetric radiance field methods for real-world applications. Experimental results demonstrate that our proposed frequency domain entropy model can achieve superior compression performance across various datasets. Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge |
MMSP | 3 |
| 2023 | L4S Congestion Control Algorithm for Interactive Low Latency Applications over 5GabstractIn recent years, applications such as cloud gaming and virtual video conferencing have gained increasing popularity and new applications, such as immersive applications, have emerged that require very low latency in order to guarantee quality of service. Such applications can benefit from using the Low Latency, Low Loss, Scalable Throughput (L4S) specification that is currently being defined by the IETF. This paper presents a congestion control algorithm aiming at achieving a target latency over a 5G connection, which together with L4S can be used in delay-critical applications. The described algorithm has been implemented into WebRTC and uses ECN marking to adapt its sending rate. The developed algorithm has been compared to the Google Congestion Control (GCC) as a baseline on various data rate patterns and shown to be more responsive to throughput variations more successfully avoiding latency spikes that surpass the acceptable latency. Jangwoo Son, Yago Sánchez de la Fuente, Christian Hampe, Dominik Schnieders, Thomas Schierl, Cornelius Hellge |
ICME | 2 |
| 2022 | Region-Of-Interest Coding Schemes For Http Adaptive Streaming With VVCabstractIndependently coding areas in a video have potential for improved coding efficiency in Region-of-Interest (RoI) streaming applications where clients can freely switch between RoI and full view streams. This paper investigates three schemes for RoI coding in streaming scenarios based on VVC to enable harnessing open GOP coding efficiency gains. Two schemes are based on Motion-Constrained-Tile-Sets while a third scheme relies on the new VVC feature referred to as independent subpictures. The reported results show substantial bitrate gains for the RoI streams compared to naïve closed-GOP coding of up to -8.68% YUV BD-rate while for the full stream minor coding efficiency losses are reported. Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl, Moncef Gabbouj |
ICIP | 2 |
| 2021 | Open GOP Resolution Switching in HTTP Adaptive Streaming with VVCabstractThe user experience in adaptive HTTP streaming relies on offering bitrate ladders with suitable operation points for all users and typically involves multiple resolutions. While open GOP coding structures are generally known to provide substantial coding efficiency benefit, their use in HTTP streaming has been precluded through lacking support of reference picture resampling (RPR) in AVC and HEVC. The newly emerging Versatile Video Coding (VVC) standard supports RPR, but only conversational scenarios were primarily investigated during the design of VVC. This paper aims at enabling usage of RPR in HTTP streaming scenarios through analysing the drift potential of VVC coding tools and presenting a constrained encoding method that avoids severe drift artefacts in resolution switching with open GOP coding in VVC. In typical live streaming configurations, the presented method achieves -8.7% BD-rate reduction compared to closed GOP coding while in a typical Video on Demand configuration, -1.89% BD-rate reduction is reported. The constraints penalty compared to regular open GOP coding is 0.65% BD-rate in the worst case. The presented method was integrated into the publicly available open source VVC encoder VVenC v0.3. Robert Skupin, Christian Bartnik, Adam Wieckowski, Yago Sánchez de la Fuente, Benjamin Bross, Cornelius Hellge, Thomas Schierl |
PCS | 4 |
| 2021 | The High-Level Syntax of the Versatile Video Coding (VVC) StandardabstractVersatile Video Coding (VVC), a.k.a. ITU-T H.266 | ISO/IEC 23090-3, is the new generation video coding standard that has just been finalized by the Joint Video Experts Team (JVET) of ITU-T VCEG and ISO/IEC MPEG at its$19^{\mathrm {th}}$meeting ending on July 1, 2020. This paper gives an overview of the VVC high-level syntax (HLS), which forms its system and transport interface. Comparisons to the HLS designs in High Efficiency Video Coding (HEVC) and Advanced Video Coding (AVC), the previous major video coding standards, are included. When discussing new HLS features introduced into VVC or differences relative to HEVC and AVC, the reasoning behind the design differences and the benefits they bring are described. The HLS of VVC enables newer and more versatile use cases such as video region extraction, composition and merging of content from multiple coded video bitstreams, and viewport-adaptive 360° immersive media. Ye-Kui Wang, Robert Skupin, Miska M. Hannuksela, Sachin Deshpande, Hendry, Virginie Drugeon, Rickard Sjöberg, Byeongdoo Choi, Vadim Seregin, Yago Sánchez de la Fuente, Jill M. Boyce, Wade Wan, Gary J. Sullivan |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2020 | Rate Assignment in 360-Degree Video Tiled Streaming Using Random Forest RegressionabstractStreaming of high-resolution 360-degree video is typically done in a viewport-dependent fashion such as in the tile-based viewport-dependent profile of MPEG OMAF wherein clients continuously adapt their tile selection according to the user viewport. From the perspective of a streaming service operator, tile rate assignment is crucial to ensure that a given target bitrate is obeyed while quality distribution among tiles leads to a favorable user experience. This paper addresses rate assignment in a distributed tile encoding system for such multi-resolution tiled streaming services based on the emerging Versatile Video Coding Standard. A model for rate assignment is derived based on random forest regression using spatio-temporal activity and encodings of a learning dataset. Robert Skupin, Kai Bitterschulte, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl |
ICASSP | 3 |
| 2019 | Encoding Configurations for Tile-Based 360° Videoabstract360° video streaming has drawn a lot of attention in the last years. In order to provide a good quality of experience with the existing device capability and available bandwidths several viewport dependent techniques have been investigated and developed. One of these techniques, as specified in MPEG OMAF, is to use tile-based streaming, with the 360° video being offered in a tiled manner at various resolutions. Thus, each user can retrieve the tiles at different resolutions, thereby selecting the tiles that match its viewport at high-resolution. This paper provides an analysis on different configurations for tile-based 360° video streaming, comparing the quality that can be achieved by using them for different values of end-to-end delays. Yago Sánchez de la Fuente, Gurdeep Singh Bhullar, Robert Skupin, Cornelius Hellge, Thomas Schierl |
ISM | 1 |
| 2019 | HTML5 MSE playback of MPEG 360 VR tiled streaming: JavaScript implementation of MPEG-OMAF viewport-dependent video profile with HEVC tilesabstractVirtual Reality (VR) and 360-degree video streaming have gained significant attention in recent years. First standards have been published in order to avoid market fragmentation. For instance, 3GPP released its first VR specification to enable 360-degree video streaming over 5G networks which relies on several technologies specified in ISO/IEC 23090-2, also known as MPEG-OMAF. While some implementations of OMAF-compatible players have already been demonstrated at several trade shows, so far, no web browser-based implementations have been presented. In this demo paper we describe a browser-based JavaScript player implementation of the most advanced media profile of OMAF: HEVC-based viewport-dependent OMAF video profile, also known as tile-based streaming, with multi-resolution HEVC tiles. We also describe the applied workarounds for the implementation challenges we encountered with state-of-the-art HTML5 browsers. The presented implementation was tested in the Safari browser with support of HEVC video through the HTML5 Media Source Extensions API. In addition, the WebGL API was used for rendering, using region-wise packing metadata as defined in OMAF. Dimitri Podborski, Jangwoo Son, Gurdeep Singh Bhullar, Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl |
MMSys | 5 |
| 2018 | Tile-Based Rate Assignment for 360-Degree Video Based on Spatio-Temporal Activity MetricsabstractTile-based video systems have recently emerged as a viable solution to overcome the challenges of 360-degree video. For instance, the HEVC based viewport-dependent profile of MPEG OMAF allows serving clients independently coded tiles of the 360-degree video at varying resolution to enhance fidelity within the actual user viewport. During streaming, the client constantly adapts its tile selection and feeds a single merged bitstream to the video decoder. This paper addresses the open issue of rate assignment in a distributed encoding system in such a multi-resolution tiled streaming scenario. A model for tile rate assignment based on the spatio-temporal activity of the video is presented to reduce variance of the quality distribution and experimental results are reported. Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl |
ISM | 2 |
| 2017 | HEVC tile based streaming to head mounted displaysabstract360° video streaming to clients using Virtual Reality head mounted displays is a challenge for traditional video delivery. As transmission of the complete content in a desirable quality sacrifices a large fraction of available client and network resources, adaptivity to the user viewport promises substantial benefits. An efficient way to achieve viewport adaptive streaming without per-user or per-orientation encoding, i.e. essentially transcoding, is to make use of motion-constrained HEVC tiles. DASH can be used for tiled streaming, where tiled content resides on the server at multiple resolutions. The DASH client selects the resolutions of each tile according to the current viewport. This demonstration paper presents an agile and responsive streaming prototype system for 360° video content. In order to achieve acceptable responsiveness, the DASH client relies on small buffer sizes and shifted random access points across tiles when suitable. Robert Skupin, Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl |
CCNC | 2 |
| 2017 | Random access point period optimization for viewport adaptive tile based streaming of 360° videoabstract360° video streaming introduces stricter requirements to the established transmission chain than in traditional streaming services. Transmission of the complete 360° video in desirable quality can lead to multiple times UHD resolution which would waste a large fraction of network resources as the larger fraction of the video is not presented on the end device. Adaptivity to the user viewport promises substantial benefits and contrary to per-user or per-orientation encoding, tiled based streaming provides a scalable solution. We consider tiled streaming of 360° video using the cubic projection, where tiled content resides on the server at 2 different resolutions. User traces have been collected for a range of content and based on consumption patterns, tiles at the server have been optimized. More concretely, the paper optimizes the Random Access Point period that leads to the minimum transmitted bitrate, while ensuring that users watch most of the time high resolution content. Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl |
ICIP | 1 |
| 2017 | Viewport-dependent 360 degree video streaming based on the emerging Omnidirectional Media Format (OMAF) standardabstract360 degree video streaming has gained much interest recently. The Omnidirectional MediA Format (OMAF) standard, currently in development by the Moving Picture Experts Group (MPEG), standardizes means for storage and delivery of 360 degree coded video based on a well-established standards ecosystem. This demonstration system uses the viewport-dependent OMAF media profile in which visual content within the current viewport is transmitted and displayed in higher fidelity than content outside the viewport to exceed the visual fidelity of viewport-independent solutions. This is achieved by dividing the video frame into tiles and using HEVC motion-constrained tile set (MCTS) encoding at multiple resolutions. Robert Skupin, Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl |
ICIP | 2 |
| 2017 | Tile based panoramic streaming using shifted IDR representationsabstractPanoramic streaming enables users to interactively navigate through high-spatial resolution videos and create an immersive and personalized user experience. Since transmission of high-resolution videos in desirable quality is not feasible given the limited throughput of access and home network links, our work is based on tile-based streaming, where only a spatial subset of the video is transmitted. In this paper, we propose a highly responsive DASH client algorithm that allows users to rapidly change the set of downloaded tiles. The high responsiveness is achieved by using small buffers, which are usually very sensitive to variations of throughput and instantaneous media bitrate and are therefore prone to playback interruptions. The proposed rate adaptation algorithm for DASH, in combination with a peak bitrate reducing RAP configuration referred to as shifted IDRs, outperforms the state of the art rate adaptation algorithms. Experiments report a decreased number of playback interruptions and quality changes while maintaining the average video quality. Dimitri Podborski, Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl |
ICME | 2 |
| 2017 | Spatio-Temporal Activity based Tiling for Panorama StreamingabstractIn panorama streaming an arbitrary Region-of-Interest (RoI) of a high-resolution video is transmitted, allowing users to navigate interactively within the videos. Transmitting the whole video becomes unfeasible due to the required high bitrates and sending a single video per user, which is encoded for that specific user (i.e. its RoI), has scalability issues. Tile based panoramic streaming overcomes the mentioned drawbacks by allowing users to receive a set of tiles that match their RoI instead of the whole set of tiles. However, optimal tiling -- so that the transmitted bitrate of the RoI content is minimized - is content dependent. In this paper, we propose a model based on a spatio-temporal activity metric so that optimization of the tiling process can be performed in a low complexity manner. Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl |
NOSSDAV | 1 |
| 2017 | Standardization status of 360 degree video coding and deliveryabstractThe emergence of consumer level capturing and display devices for 360 degree video creates new and promising segments in entertainment, education, professional training, and other markets. In order to avoid market fragmentation and ensure interoperability of 360 degree video ecosystems, industry and academia cooperate in standardization efforts in this field. In the video coding domain, 360 degree video invalidates many established procedures, e.g., concerning evaluation of the visual quality, while the specific content characteristics offer potential for higher compression efficiency beyond the current standards. Likewise, 360 degree video puts stricter demands on the system level aspects of transmission but may also offer the potential to enhance existing transport schemes. The Joint Collaborative Team on Video Coding (JCT-VC) as well as the Joint Video Exploration Team (JVET) already started investigations into 360 degree video coding while numerous activities in the Systems subgroup of the Moving Picture Experts Group (MPEG) started to investigate application requirements and delivery aspects of 360 degree video. This paper reports on the current status of the outlined standardization efforts. Robert Skupin, Yago Sánchez de la Fuente, Ye-Kui Wang, Miska M. Hannuksela, Jill M. Boyce, Mathias Wien |
VCIP | 2 |
| 2017 | Video processing for panoramic streaming using HEVC and its scalable extensionsabstractPanoramic streaming is a particular way of video streaming where an arbitrary Region-of-Interest (RoI) is transmitted from a high-spatial resolution video, i.e. a video covering a very “wide-angle” (much larger than the human field-of-view – e.g. 360°). Some transport schemes for panoramic video delivery have been proposed and demonstrated within the past decade, which allow users to navigate interactively within the high-resolution videos. With the recent advances of head mounted displays, consumers may soon have immersive and sufficiently convenient end devices at reach, which could lead to an increasing demand for panoramic video experiences. The solution proposed within this paper is built upon tile-based panoramic streaming, where users receive a set of tiles that match their RoI, and consists in a low-complexity compressed domain video processing technique for using H.265/HEVC and its scalable extensions (H.265/SHVC and H.265/MV-HEVC). The proposed technique generates a single video bitstream out of the selected tiles so that a single hardware decoder can be used. It overcomes the scalability issue of previous solutions not using tiles and the battery consumption issue inherent of tile-based panorama streaming, where multiple parallel software decoders are used. In addition, the described technique is capable of reducing peak streaming bitrate during changes of the RoI, which is crucial for allowing a truly immersive and low latency video experience. Besides, it makes it possible to use Open GOP structures without incurring any playback interruption at switching events, which provides a better compression efficiency compared to closed GOP structures. Yago Sánchez de la Fuente, Robert Skupin, Thomas Schierl |
Multim. Tools Appl. | 1 |
| 2016 | Peak bitrate reduction for multi-party video conferencing using SHVCabstractReal-time video applications, such as multi party video conferencing, involve the simultaneous transport of multiple and potentially multi-layered video sources to participating or interested parties. It is desirable to mix these multiple source videos into a single video stream at intermediary nodes in the network, e.g. at Multipoint Control Units (MCU). This has the advantage of reduced application and transport complexity on the client device while allowing single hardware decoder devices to consume the content. This paper proposes a solution, which uses the scalable extension of H.265/HEVC (SHVC), and that generates a single bitstream out of several sources by a low-complexity operation. In addition, the presented technique drastically reduces the peak bitrate at layout change events in comparison to state-of-the-art solutions. Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl |
ICIP | 1 |
| 2016 | Shifted IDR Representations for Low Delay Live DASH Streaming Using HEVC TilesabstractDynamic Adaptive Streaming over HTTP (DASH) has gained a lot of popularity in the last years and has become a widespread solution for media delivery. It is mainly used for VoD services but has being also standardized and can be used for live scenarios. Although there has been a lot of work on low delay DASH services, it is still very challenging to cope with network variations. In fact, low delay configurations prevent DASH users from having a large media buffer to cope with network variations and they can easily lead to deadline misses, which results in a service of low Quality of Experience (QoE). In this paper we propose a solution based on HEVC Tiles and an advanced DASH adaptation algorithm that improves drastically the performance by heavily reducing the deadline misses. Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl |
ISM | 1 |
| 2016 | Tile Based HEVC Video for Head Mounted Displaysabstract360° Video services with resolutions of UHD and beyond for Virtual Reality head mounted displays are a challenging task due to limits of video decoders in constrained end devices. Adaptivity to the current user viewport is a promising approach but incurs significant encoding overhead when encoding per user or set of viewports. A more efficient way to achieve viewport adaptive streaming is to facilitate motion-constrained HEVC tiles. Original content resolution within the user viewport is preserved while content currently not presented to the user is delivered in lower resolution. A lightweight aggregation of varying resolution tiles into a single HEVC bitstream can be carried out on-the-fly and allows usage of a single decoder instance on the end device. Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl |
ISM | 2 |
| 2016 | Step response metric for video encoder rate control characterisationabstractAdaptive streaming is a widely-used solution to avoid playback interruptions and compensate for insufficient bandwidth or momentary congestion. A temporary reduction in bitrate results in less or no video break-up and rebuffering. The rate control mechanism of a codec implementation allows the selection of a target bitrate and tries to meet this target within certain constraints e.g. number of frames, with a certain percentage overshoot. However the target bitrate is not always instantaneously achievable, and the number of bits required to encode a frame are highly content-dependent. In this paper the Codec Step Response metric is proposed to measure the effectiveness of a codec rate control at responding to instantaneous rate variations. This metric allows the measurements of the suitability of a rate control implementation for an application or access network. Ralf Globisch, Keith L. Ferguson, Yago Sánchez de la Fuente, Thomas Schierl |
NOSSDAV | 3 |
| 2015 | Compressed domain video processing for tile based panoramic streaming using HEVCabstractPanoramic streaming is a particular way of video streaming where an arbitrary Region-of-Interest (RoI) of high-spatial resolution videos is transmitted. It allows users to navigate interactively around the video and thus select anytime the portion of it they are interested in. The most basic approach consists of each user interacting with the system indicating the desired RoI. Then an encoder associated with each user encodes the desired RoI. However, such a system does not scale well. Instead, we consider tile based panoramic streaming, where users receive a set of tiles that match their RoI, and propose a low-complexity compressed domain video processing technique for tile based panoramic streaming using H.265/HEVC that generates a single video bitstream out of the selected tiles so that single hardware decoders can be used to decode the RoI video stream. Yago Sánchez de la Fuente, Robert Skupin, Thomas Schierl |
ICIP | 1 |
| 2015 | Compressed domain video compositing with HEVCabstractVideo compositing such as blending of user interfaces or advertisements on top of video content is used in many applications. Compositing is usually carried out in the pixel domain either after decoding on the end device or based on transcoding before or during transport, e.g. on cloud resources. This paper proposes a novel method to create a composition of several coded input videos in the compressed domain, i.e. without performing entropy coding at runtime. The method entails merging of input video bitstreams into a single output video bitstream and insertion of pre-encoded inter-predicted composition pictures. Such a lightweight approach is computationally much less demanding than transcoding-based compositing and can be beneficial for service scalability. This paper explores the coding performance of the proposed method though experiments with the composition of a transparent ticker overlay on top of video sequences. The proposed method is reported to outperform transcoding-based pixel domain compositing in rate distortion and quality. Robert Skupin, Yago Sánchez de la Fuente, Thomas Schierl |
PCS | 2 |
| 2014 | Low complexity cloud-video-mixing using HEVCabstractReal-time video applications, such as video conferencing and video surveillance systems, typically involve the simultaneous transport of multiple video sources to interested parties that consume the content. It may be desirable to mix these multiple source videos into a single video stream at intermediary nodes in the network. This has the advantage of reduced application and transport complexity on the client device and also makes it possible for devices with a single hardware decoder to consume the content. A typical approach is to apply transcoding operations to the original videos, i.e. the videos are decoded, merged and encoded into a single video stream. This paper proposes an alternative solution to video transcoding, which uses the new video coding standard HEVC and has a much lower processing complexity. We consider how our approach can be realized in real-world applications such as a cloud video mixer. Such systems typically require some degree of dynamics and personalization and we provide some insight into how transport signaling complexities can be addressed. Yago Sánchez de la Fuente, Ralf Globisch, Thomas Schierl, Thomas Wiegand 0001 |
CCNC | 1 |
| 2014 | Low latency DASH based streaming over LTEabstractDynamic Adaptive Streaming over HTTP (DASH) is becoming the de facto technique for video delivery, especially for VoD services. Although 3GPP has specified carriage of DASH over eMBMS for Live Streaming, eMBMS is not available everywhere (operators are starting service rollout) and it is only worthwhile for reasonably large number of users due to the static SFN resource allocation for those services. Thus, Live Streaming using DASH over unicast connections is still necessary, which may suffer from playback interruptions when the throughput varies since the low end-to-end latency for live streaming requires small buffers. In order to cope with network throughput variations in mobile networks, we propose the usage of scalable video coding, combining it with parallel TCP connections and prioritizing the most important data of the scalable video. We show that using LTE non-GBR bearers for prioritization playback interruptions can be avoided. Yago Sánchez de la Fuente, Edward Grinshpun, David W. Faucher, Thomas Schierl, Sameer Sharma |
VCIP | 1 |
| 2012 | Advanced downlink LTE radio resource management for HTTP-streamingabstractVideo traffic contributes to the majority of data packets transported over cellular wireless. Future broadband wireless access networks based on 3GPP's Long Term Evolution offer mechanisms for optimized transmission with high data rates and low delay. However, especially when packets are transmitted in the LTE downlink and if services are run over-the-top (OTT), optimization of radio resources in a multi-user environment for video services becomes infeasible. The current market trend is moving to OTT solutions, also for video transmission, where an emerging standard based on HTTP streaming - DASH - is expected to have a huge success in the upcoming years. The solution presented in this paper consists of a novel technique, which combines LTE features with knowledge on DASH sessions for optimization of the wireless resources. The combined optimization yields an improved transmission of videos over cellular wireless systems which are based on LTE and LTE-Advanced. Thomas Wirth, Yago Sánchez de la Fuente, Bernd Holfeld, Thomas Schierl |
ACM Multimedia | 2 |
| 2012 | Efficient HTTP-based streaming using Scalable Video Coding
Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec |
Signal Process. Image Commun. | 1 |
| 2011 | Improved caching for HTTP-based Video on Demand using Scalable Video CodingabstractHTTP-based delivery for Video on Demand (VoD) has been gaining popularity within recent years. Progressive Download over HTTP, typically used in VoD, takes advantage of the widely deployed network caches to release video servers from sending the same content to a high number of users in the same VoD service. However, due to the inherent heterogeneity of user demands, which may result in requesting the same video content in different resolutions or qualities, the caching efficiency is expected to decrease due to a higher variety in requested media files. The use of Scalable Video Coding allows different representations of the same content to be combined in a single file, whose parts, aka layers, are requested sequentially by a user up to the maximum desired quality. In this paper we show the benefits of using Scalable Video Coding to maintain the same set of possible video content representations, while at the same time maximizing the caching efficiency. Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec |
CCNC | 1 |
| 2011 | iDASH: improved dynamic adaptive streaming over HTTP using scalable video codingabstractHTTP-based delivery for Video on Demand (VoD) has been gaining popularity within recent years. Progressive Download over HTTP, typically used in VoD, takes advantage of the widely deployed network caches to relieve video servers from sending the same content to a high number of users in the same access network. However, due to a sharp increase in the requests at peak hours or due to cross-traffic within the network, congestion may arise in the cache feeder link or access link respectively. Since the connection characteristics may vary over the time, with Dynamic Adaptive Streaming over HTTP (DASH), a technique that has been recently proposed, video clients may dynamically adapt the requested video quality for ongoing video flows, to match their current download rate as good as possible. In this work we show the benefits of using the Scalable Video Coding (SVC) for such a DASH environment. Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec |
MMSys | 1 |
| 2011 | Priority-based Media Delivery using SVC with RTP and HTTP streaming
Thomas Schierl, Yago Sánchez de la Fuente, Ralf Globisch, Cornelius Hellge, Thomas Wiegand 0001 |
Multim. Tools Appl. | 2 |
| 2010 | P2P group communication using Scalable Video CodingabstractP2P-streaming has become of high interest in the last years, since it reduces the load on expensive servers, due to the participation of receivers in the media transmission. In this paper, P2P content delivery is shown to be a promising technique for video group communication, for which the main requirement is low-delay. Combining low-delay encoding and low-delay P2P Application Layer Multicast makes it possible to fulfill the delay constraints for interactive group communication applications. For such an application, congestion is a considerable problem, since it causes packet loss or late arrival of the packets, degrading the quality of the service. The results presented in this paper show how rate adaptation in combination with the Scalable Video Coding (SVC) helps to overcome problems in the network, providing a better solution than when non-adaptive single layer coding is transmitted. Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001 |
ICIP | 1 |
| 2010 | Improving P2P live-content delivery using SVCabstractP2P content delivery techniques for video transmission have become of high interest in the last years. With the involvement of client into the delivery process, P2P approaches can significantly reduce the load and cost on servers, especially for popular services. However, previous studies have already pointed out the unreliability of P2P-based live streaming approaches due to peer churn, where peers may ungracefully leave the P2P infrastructure, typically an overlay networks. Peers ungracefully leaving the system cause connection losses in the overlay, which require repair operations. During such repair operations, which typically take a few roundtrip times, no data is received from the lost connection. While taking low delay for fast-channel tune-in into account as a key feature for broadcast-like streaming applications, the P2P live streaming approach can only rely on a certain media pre-buffer during such repair operations. In this paper, multi-tree based Application Layer Multicast as a P2P overlay technique for live streaming is considered. The use of Flow Forwarding (FF), a.k.a. Retransmission, or Forward Error Correction (FEC) in combination with Scalable video Coding (SVC) for concealment during overlay repair operations is shown. Furthermore the benefits of using SVC over the use of AVC single layer transmission are presented. Thomas Schierl, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Wiegand 0001 |
VCIP | 2 |