Cornelius Hellge

dblp:58/2455 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-4465-9136ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 6 first-author · 9 since 2021Computer networks · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 CodecGS-IR: Implicit Representation for Decoder-friendly Video-based Gaussian Splat Compression
abstract
Coding of Gaussian splats has drawn the attention of academia and standardization bodies lately. A commonly used approach involves projecting explicit 3D attributes onto 2D planes of a video and compressing it with existing video coding standards. However, such an approach produces excessively high sample rates, resulting in a significant bottleneck that might make it unsuitable for existing hardware decoders if the number of splats is high. This paper presents a framework that utilizes an implicit representation, where Gaussian splat attributes are represented by compact feature planes. By reducing the dependency of video resolution on the number of splats, our approach significantly reduces the required sample rate, making it suitable for deployed devices. In addition, the proposed framework achieves a superior rate-distortion trade-off, providing high-fidelity reconstruction at low bitrates without the excessive sample rates associated with the conventional video-based anchor.
Soonbin Lee, Simon Sasse, Yago Sánchez de la Fuente, Robert Skupin, Tomas M. Borges, Cornelius Hellge, Thomas Schierl
DCC6
2025 Compression of 3D Gaussian Splatting with Optimized Feature Planes and Standard Video Codecs
abstract
3D Gaussian Splatting is a recognized method for 3D scene representation, known for its high rendering quality and speed. However, its substantial data requirements present challenges for practical applications. In this paper, we introduce an efficient compression technique that significantly reduces storage overhead by using compact representation. We propose a unified architecture that combines point cloud data and feature planes through a progressive tri-plane structure. Our method utilizes 2D feature planes, enabling continuous spatial representation. To further optimize these representations, we incorporate entropy modeling in the frequency domain, specifically designed for standard video codecs. We also propose channel-wise bit allocation to achieve a better trade-off between bitrate consumption and feature plane representation. Consequently, our model effectively leverages spatial correlations within the feature planes to enhance rate-distortion performance using standard, non-differentiable video codecs. Experimental results demonstrate that our method outperforms existing methods in data compactness while maintaining high rendering quality. Our project page is available at https://fraunhoferhhi.github.io/CodecGS
Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge
ICCV5
2024 ECRF: Entropy-Constrained Neural Radiance Fields Compression with Frequency Domain Optimization
abstract
Explicit feature-grid based NeRF models have shown promising results in terms of rendering quality and significant speed-up in training. However, these methods often require a significant amount of data to represent a single scene or object. In this work, we present a compression model that aims to minimize the entropy in the frequency domain in order to effectively reduce the data size. First, we propose using the discrete cosine transform (DCT) on the tensorial radiance fields to compress the feature-grid. This feature-grid is transformed into coefficients, which are then quantized and entropy encoded, following a similar approach to the traditional video coding pipeline. Furthermore, to achieve a higher level of sparsity, we propose using an entropy parameterization technique for the frequency domain, specifically for DCT coefficients of the feature-grid. Since the transformed coefficients are optimized during the training phase, the proposed model does not require any fine-tuning or additional information. Our model only requires a lightweight compression pipeline for encoding and decoding, making it easier to apply volumetric radiance field methods for real-world applications. Experimental results demonstrate that our proposed frequency domain entropy model can achieve superior compression performance across various datasets.
Soonbin Lee, Fangwen Shu, Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge
MMSP5
2024 Improving QoE-Privacy Tradeoff in XR Streaming
abstract
Viewpoint and position (VP)-adaptive mixed reality (XR) streaming requires uploading the trajectory of user behavior, thereby causing privacy leakage. Preserving privacy leads to the performance loss of VP prediction of users, subsequently degrading the quality of experience (QoE). This work is the first to improve the QoE-privacy tradeoff for XR streaming. We find the key difference of sample importance for privacy attack and VP prediction, based on which we propose a framework to improve the tradeoff. By designing the noisy entropy function and remapping function, it can achieve a better tradeoff than simply adding the noise to the actual trajectory. The performance of the framework is evaluated with the state-of-the-art VP predictors and practical XR streaming platforms. The results show that the loss of QoE can be mitigated by 55%$\sim$100% to achieve the same privacy level as simply adding noise. When the privacy level achieves the maximum, the QoE is degraded by 0$\sim$3%, compared to VP-adaptive streaming without any privacy-preserving.
Xing Wei 0003, Cornelius Hellge, Chenyang Yang 0001, Jangwoo Son
IEEE Signal Process. Lett.2
2024 Distributed Machine-Learning for Early HARQ Feedback Prediction in Cloud RANs
abstract
In this work, we propose novel HARQ prediction schemes for Cloud RANs (C-RANs) that use feedback over a rate-limited feedback channel (2 - 6 bits) from the Remote Radio Heads (RRHs) to predict at the User Equipment (UE) the decoding outcome at the BaseBand Unit (BBU) ahead of actual decoding. In particular, we propose a Dual Autoencoding 2-Stage Gaussian Mixture Model (DA2SGMM) that is trained in an end-to-end fashion over the whole C-RAN setup. Using realistic link-level simulations in the sub-THz band at 100 GHz, we show that the novel DA2SGMM HARQ prediction scheme clearly outperforms all other adapted and state-of-the-art schemes. The DA2SGMM shows a superior performance in terms of blockage detection as well as HARQ prediction in the no-blockage and single-blockage cases. In particular, the DA2SGMM with 4 bit feedback achieves a more than 200 % higher throughput in average compared to its best alternative. Compared to regular HARQ, the DA2SGMM reduces the maximum transmission latency by more than 72.4 %, while maintaining more than 75 % of the throughput in the no-blockage scenario. In the single-blockage scenario, DA2SGMM significantly increases the throughput for most of the evaluated Signal-to-Noise-Ratios (SNRs) compared to regular HARQ.
Baris Göktepe, Cornelius Hellge, Thomas Schierl, Slawomir Stanczak
IEEE Trans. Wirel. Commun.2
2023 L4S Congestion Control Algorithm for Interactive Low Latency Applications over 5G
abstract
In recent years, applications such as cloud gaming and virtual video conferencing have gained increasing popularity and new applications, such as immersive applications, have emerged that require very low latency in order to guarantee quality of service. Such applications can benefit from using the Low Latency, Low Loss, Scalable Throughput (L4S) specification that is currently being defined by the IETF. This paper presents a congestion control algorithm aiming at achieving a target latency over a 5G connection, which together with L4S can be used in delay-critical applications. The described algorithm has been implemented into WebRTC and uses ECN marking to adapt its sending rate. The developed algorithm has been compared to the Google Congestion Control (GCC) as a baseline on various data rate patterns and shown to be more responsive to throughput variations more successfully avoiding latency spikes that surpass the acceptable latency.
Jangwoo Son, Yago Sánchez de la Fuente, Christian Hampe, Dominik Schnieders, Thomas Schierl, Cornelius Hellge
ICME6
2023 On the Limits of HARQ Prediction for Short Deterministic Codes with Error Detection in Memoryless Channels
abstract
We provide a mathematical framework to analyze the limits of Hybrid Automatic Repeat reQuest (HARQ) and derive analytical expressions for the most powerful test for estimating the decodability under maximum-likelihood decoding and t-error decoding. Furthermore, we numerically approximate the most powerful test for sum-product decoding. We compare the performance of previously studied HARQ prediction schemes and show that none of the state-of-the-art HARQ prediction is most powerful to estimate the decodability of a partially received signal vector under maximum-likelihood decoding and sum-product decoding. Furthermore, we demonstrate that decoding in general is suboptimal for predicting the decodability.
Baris Göktepe, Cornelius Hellge, Tatiana Rykova, Thomas Schierl, Slawomir Stanczak
ISIT2
2023 Soft-Output PAC Decoder for Uplink Sparse Code Multiple Access System
abstract
In this paper, we investigate short packet transmissions in a Polar Coded Sparse Code Multiple Access (PC-SCMA) system. Since recently introduced by Arikan Polarization-Adjusted Convolutional (PAC) codes have shown remarkable performance at short block lengths with an ability of reaching theoretical limits, in this work we propose a PC-SCMA system based on the soft output PAC decoder. The proposed decoder is a combination of a sequential stack decoder and an enhanced Belief Propagation (BP) algorithm. The numerical analysis shows its ability to outperform the state-of-the-art PC-SCMA system with a Successive Cancellation List (SCL) decoder at short block lengths in terms of the bit error rate with smaller computational complexity. Furthermore, we enhance the SCL-based scheme by row-merging of the polarization kernel and evaluate its performance.
Tatiana Rykova, Baris Göktepe, Thomas Schierl, Cornelius Hellge
WiMob4
2022 A hybrid HARQ feedback prediction approach for Single- and Cloud-RANs in the sub-THz regime
abstract
In this work, we extend two autoencoder-based HARQ prediction schemes to exploit subcode-based features and SNR-based features jointly. We apply the proposed HARQ prediction schemes to Cloud-RAN (C-RAN) and Single-RAN (S-RAN) architectures. Furthermore, we conduct realistic link-level simulations to test the performance and compare to state-of-the-art prediction schemes that rely solely on either subcode-based features or SNR-based features. Compared to the state-of-the-art, we show that the proposed schemes reduce the transmitted redundancy at a target error rate of$5\cdot 10^{-5}$and$2\cdot 10^{-5}$by 12.3% - 27.3% in C-RAN architectures and 10.5% - 11.0% in S- RAN architectures, respectively.
Baris Göktepe, Cornelius Hellge, Thomas Schierl, Slawomir Stanczak
GLOBECOM2
2022 Latency Compensation Through Image Warping For Remote Rendering-Based Volumetric Video Streaming
abstract
Rendering multiple high-quality volumetric videos is still a challenge for today’s mobile devices. Remote rendering offloads complex rendering operations to a powerful server and provides the final result to the end device as a 2D video stream. A drawback of remote rendering is the significant increase of interaction latency that can degrade the user experience. We present a planar homography-based approach that compensates minor changes of the user’s head pose due to the interaction latency by warping the transmitted image on the client-side, just before it is sent to the display. In detail, we use the homography between the initial head pose when the image is rendered at the server and the latest available head pose of the user at the client. We perform controlled experiments using artificial camera traces to evaluate our approach. The results show that the proposed approach reduces the rendering errors significantly in terms of the mean-squared error between the rendered and reference images, especially combined with initial head motion prediction.
Serhan Gul, Cornelius Hellge, Peter Eisert
ICIP2
2022 Region-Of-Interest Coding Schemes For Http Adaptive Streaming With VVC
abstract
Independently coding areas in a video have potential for improved coding efficiency in Region-of-Interest (RoI) streaming applications where clients can freely switch between RoI and full view streams. This paper investigates three schemes for RoI coding in streaming scenarios based on VVC to enable harnessing open GOP coding efficiency gains. Two schemes are based on Motion-Constrained-Tile-Sets while a third scheme relies on the new VVC feature referred to as independent subpictures. The reported results show substantial bitrate gains for the RoI streams compared to naïve closed-GOP coding of up to -8.68% YUV BD-rate while for the full stream minor coding efficiency losses are reported.
Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl, Moncef Gabbouj
ICIP3
2021 Reproducibility Companion Paper: Kalman Filter-Based Head Motion Prediction for Cloud-Based Mixed Reality
abstract
In our MM'20 paper,, we presented a Kalman filter-based approach for prediction of head motion in 6DoF. The proposed approach was employed in our cloud-based volumetric video streaming system to reduce the interaction latency experienced by the user. In this companion paper, we present the dataset collected for our experiments and our simulation framework that reproduces the obtained experimental results. Our implementation is freely available on Github to facilitate further research.
Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge, Marc A. Kastner 0001, Jan Zahálka
ACM Multimedia5
2021 Open GOP Resolution Switching in HTTP Adaptive Streaming with VVC
abstract
The user experience in adaptive HTTP streaming relies on offering bitrate ladders with suitable operation points for all users and typically involves multiple resolutions. While open GOP coding structures are generally known to provide substantial coding efficiency benefit, their use in HTTP streaming has been precluded through lacking support of reference picture resampling (RPR) in AVC and HEVC. The newly emerging Versatile Video Coding (VVC) standard supports RPR, but only conversational scenarios were primarily investigated during the design of VVC. This paper aims at enabling usage of RPR in HTTP streaming scenarios through analysing the drift potential of VVC coding tools and presenting a constrained encoding method that avoids severe drift artefacts in resolution switching with open GOP coding in VVC. In typical live streaming configurations, the presented method achieves -8.7% BD-rate reduction compared to closed GOP coding while in a typical Video on Demand configuration, -1.89% BD-rate reduction is reported. The constraints penalty compared to regular open GOP coding is 0.65% BD-rate in the worst case. The presented method was integrated into the publicly available open source VVC encoder VVenC v0.3.
Robert Skupin, Christian Bartnik, Adam Wieckowski, Yago Sánchez de la Fuente, Benjamin Bross, Cornelius Hellge, Thomas Schierl
PCS6
2020 Feedback Prediction for Proactive HARQ in the Context of Industrial Internet of Things
abstract
In this work, we investigate proactive Hybrid Automatic Repeat reQuest (HARQ) using link-level simulations for multiple packet sizes, modulation orders, BLock Error Rate (BLER) targets and two delay budgets of 1 ms and 2 ms, in the context of Industrial Internet of Things (IIOT) applications. In particular, we propose an enhanced proactive HARQ protocol using a feedback prediction mechanism. We show that the enhanced protocol achieves a significant gain over the classical proactive HARQ in terms of energy efficiency for almost all evaluated BLER targets at least for sufficiently large feedback delays. Furthermore, we demonstrate that the proposed protocol clearly outperforms the classical proactive HARQ in all scenarios when taking a processing delay reduction due to the less complex prediction approach into account, achieving an energy efficiency gain in the range of 11% up to 15% for very stringent latency budgets of 1 ms at 10-2BLER and from 4% up to 7.5% for less stringent latency budgets of 2 ms at 10-3BLER. Furthermore, we show that power-constrained proactive HARQ with prediction even outperforms unconstrained reactive HARQ for sufficiently large feedback delays.
Baris Göktepe, Tatiana Rykova, Thomas Fehrenbach, Thomas Schierl, Cornelius Hellge
GLOBECOM5
2020 Rate Assignment in 360-Degree Video Tiled Streaming Using Random Forest Regression
abstract
Streaming of high-resolution 360-degree video is typically done in a viewport-dependent fashion such as in the tile-based viewport-dependent profile of MPEG OMAF wherein clients continuously adapt their tile selection according to the user viewport. From the perspective of a streaming service operator, tile rate assignment is crucial to ensure that a given target bitrate is obeyed while quality distribution among tiles leads to a favorable user experience. This paper addresses rate assignment in a distributed tile encoding system for such multi-resolution tiled streaming services based on the emerging Versatile Video Coding Standard. A model for rate assignment is derived based on random forest regression using spatio-temporal activity and encodings of a learning dataset.
Robert Skupin, Kai Bitterschulte, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl
ICASSP4
2020 Kalman Filter-based Head Motion Prediction for Cloud-based Mixed Reality
abstract
Volumetric video allows viewers to experience highly-realistic 3D content with six degrees of freedom in mixed reality (MR) environments. Rendering complex volumetric videos can require a prohibitively high amount of computational power for mobile devices. A promising technique to reduce the computational burden on mobile devices is to perform the rendering at a cloud server. However, cloud-based rendering systems suffer from an increased interaction (motion-to-photon) latency that may cause registration errors in MR environments. One way of reducing the effective latency is to predict the viewer's head pose and render the corresponding view from the volumetric video in advance.
Serhan Gul, Sebastian Bosse, Dimitri Podborski, Thomas Schierl, Cornelius Hellge
ACM Multimedia5
2020 Cloud rendering-based volumetric video streaming system for mixed reality services
abstract
Volumetric video is an emerging technology for immersive representation of 3D spaces that captures objects from all directions using multiple cameras and creates a dynamic 3D model of the scene. However, processing volumetric content requires high amounts of processing power and is still a very demanding task for today's mobile devices. To mitigate this, we propose a volumetric video streaming system that offloads the rendering to a powerful cloud/edge server and only sends the rendered 2D view to the client instead of the full volumetric content. We use 6DoF head movement prediction techniques, WebRTC protocol and hardware video encoding to ensure low-latency in different parts of the processing chain. We demonstrate our system using both a browser-based client and a Microsoft HoloLens client. Our application contains generic interfaces that allow for easy deployment of various augmented/mixed reality clients using the same server implementation.
Serhan Gul, Dimitri Podborski, Jangwoo Son, Gurdeep Singh Bhullar, Thomas Buchholz, Thomas Schierl, Cornelius Hellge
MMSys7
2020 Low-latency cloud-based volumetric video streaming using head motion prediction
abstract
Volumetric video is an emerging key technology for immersive representation of 3D spaces and objects. Rendering volumetric video requires lots of computational power which is challenging especially for mobile devices. To mitigate this, we developed a streaming system that renders a 2D view from the volumetric video at a cloud server and streams a 2D video stream to the client. However, such network-based processing increases the motion-to-photon (M2P) latency due to the additional network and processing delays. In order to compensate the added latency, prediction of the future user pose is necessary. We developed a head motion prediction model and investigated its potential to reduce the M2P latency for different look-ahead times. Our results show that the presented model reduces the rendering errors caused by the M2P latency compared to a baseline system in which no prediction is performed.
Serhan Gul, Dimitri Podborski, Thomas Buchholz, Thomas Schierl, Cornelius Hellge
NOSSDAV5
2019 Multi-Kernel Prediction Networks for Denoising of Burst Images
abstract
In low light or short-exposure photography the image is often corrupted by noise. While longer exposure helps reduce the noise, it can produce blurry results due to the object and camera motion. The reconstruction of a noise-less image is an ill posed problem. Recent approaches for image denoising aim to predict kernels which are convolved with a set of successively taken images (burst) to obtain a clear image. We propose a deep neural network based approach called Multi-Kernel Prediction Networks (MKPN) for burst image denoising. MKPN predicts kernels of not just one size but of varying sizes and performs fusion of these different kernels resulting in one kernel per pixel. The advantages of our method are two fold: (a) the different sized kernels help in extracting different information from the image which results in better reconstruction and (b) kernel fusion assures retaining of the extracted information while maintaining computational efficiency. Experimental results reveal that MKPN outperforms state-of-the-art on our synthetic datasets with different noise levels.
Talmaj Marinc, Vignesh Srinivasan, Serhan Gul, Cornelius Hellge, Wojciech Samek
ICIP4
2019 Encoding Configurations for Tile-Based 360° Video
abstract
360° video streaming has drawn a lot of attention in the last years. In order to provide a good quality of experience with the existing device capability and available bandwidths several viewport dependent techniques have been investigated and developed. One of these techniques, as specified in MPEG OMAF, is to use tile-based streaming, with the 360° video being offered in a tiled manner at various resolutions. Thus, each user can retrieve the tiles at different resolutions, thereby selecting the tiles that match its viewport at high-resolution. This paper provides an analysis on different configurations for tile-based 360° video streaming, comparing the quality that can be achieved by using them for different values of end-to-end delays.
Yago Sánchez de la Fuente, Gurdeep Singh Bhullar, Robert Skupin, Cornelius Hellge, Thomas Schierl
ISM4
2019 HTML5 MSE playback of MPEG 360 VR tiled streaming: JavaScript implementation of MPEG-OMAF viewport-dependent video profile with HEVC tiles
abstract
Virtual Reality (VR) and 360-degree video streaming have gained significant attention in recent years. First standards have been published in order to avoid market fragmentation. For instance, 3GPP released its first VR specification to enable 360-degree video streaming over 5G networks which relies on several technologies specified in ISO/IEC 23090-2, also known as MPEG-OMAF. While some implementations of OMAF-compatible players have already been demonstrated at several trade shows, so far, no web browser-based implementations have been presented. In this demo paper we describe a browser-based JavaScript player implementation of the most advanced media profile of OMAF: HEVC-based viewport-dependent OMAF video profile, also known as tile-based streaming, with multi-resolution HEVC tiles. We also describe the applied workarounds for the implementation challenges we encountered with state-of-the-art HTML5 browsers. The presented implementation was tested in the Safari browser with support of HEVC video through the HTML5 Media Source Extensions API. In addition, the WebGL API was used for rendering, using region-wise packing metadata as defined in OMAF.
Dimitri Podborski, Jangwoo Son, Gurdeep Singh Bhullar, Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl
MMSys6
2019 A Virtualized Video Surveillance System for Public Transportation
abstract
777
Talmaj Marinc, Serhan Gul, Cornelius Hellge, Peter Schüßler, Thomas Riegel, Peter Amon
ECML/PKDD (3)3
2019 Enhanced Machine Learning Techniques for Early HARQ Feedback Prediction in 5G
abstract
We investigate Early Hybrid Automatic Repeat reQuest (E-HARQ) feedback schemes enhanced by machine learning techniques as a path towards ultra-reliable and low-latency communication (URLLC). To this end, we propose machine learning methods to predict the outcome of the decoding process ahead of the end of the transmission. We discuss different input features and classification algorithms ranging from traditional methods to newly developed supervised autoencoders. These methods are evaluated based on their prospects of complying with the URLLC requirements of effective block error rates below 10-5at small latency overheads. We provide realistic performance estimates in a system model incorporating scheduling effects to demonstrate the feasibility of E-HARQ across different signal-to-noise ratios, subcode lengths, channel conditions and system loads, and show the benefit over regular HARQ and existing E-HARQ schemes without machine learning.
Nils Strodthoff, Baris Göktepe, Thomas Schierl, Cornelius Hellge, Wojciech Samek
IEEE J. Sel. Areas Commun.4
2018 Tile-Based Rate Assignment for 360-Degree Video Based on Spatio-Temporal Activity Metrics
abstract
Tile-based video systems have recently emerged as a viable solution to overcome the challenges of 360-degree video. For instance, the HEVC based viewport-dependent profile of MPEG OMAF allows serving clients independently coded tiles of the 360-degree video at varying resolution to enhance fidelity within the actual user viewport. During streaming, the client constantly adapts its tile selection and feeds a single merged bitstream to the video decoder. This paper addresses the open issue of rate assignment in a distributed encoding system in such a multi-resolution tiled streaming scenario. A model for tile rate assignment based on the spatio-temporal activity of the video is presented to reduce variance of the quality distribution and experimental results are reported.
Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl
ISM4
2018 Efficient Object Tracking in Compressed Video Streams with Graph Cuts
abstract
In this paper we present a compressed-domain object tracking algorithm for H.264/AVC compressed videos and integrate the proposed algorithm into an indoor vehicle tracking scenario at a car park. Our algorithm works by taking an initial segmentation map or bounding box of the target object in the first frame of the video sequence as input and applying Graph Cuts optimization based on a Markov Random Field model. Our algorithm does not rely on pixels (except for the first frame) and works by only using the codec motion vectors and block coding modes extracted from the H.264/AVC bitstream via inexpensive partial decoding. In this way, we manage to reduce the compute and storage requirements of our system significantly compared to “pixel-domain” tracking algorithms that first fully decode the video stream and work on reconstructed pixels. We demonstrate the quantitative performance of our algorithm over VOT2016 dataset and also integrate our algorithm into a camera-based parking management system and show qualitative results in a real application scenario. Results show that our compressed-domain algorithm provides a good compromise between high accuracy tracking and low-complexity processing showing that it is feasible for scenarios requiring large-scale object tracking in bandwidth-limited conditions.
Fernando Bombardelli da Silva, Serhan Gul, Daniel Becker 0004, Cornelius Hellge
MMSP5
2018 Visual object tracking in a parking garage using compressed domain analysis
abstract
Modern driver assistance systems enable a variety of use cases which rely on accurate localization information of all traffic participants. Due to the unavailability of satellite-based localization, the use of infrastructure cameras is a promising alternative in indoor spaces such as parking garages. This paper presents a parking management system which extends the previous work of the eValet system with a low-complexity tracking functionality on compressed video bitstreams (compressed-domain tracking). The advantages of this approach include the improved robustness to partial occlusions as well as a resource-efficient processing of compressed video bit-streams. We have separated the tasks into different modules which are integrated into a comprehensive architecture. The demonstrator setup includes a 2D visualizer illustrating the operation of the algorithms on a single camera stream and a 3D visualizer displaying the abstract object detections in a global reference frame.
Daniel Becker 0004, Fernando Bombardelli da Silva, Serhan Gul, Cornelius Hellge, Oliver Sawade, Ilja Radusch
MMSys5
2018 URLLC Services in 5G Low Latency Enhancements for LTE
abstract
5G is envisioned to support three broad categories of services: eMBB, URLLC, and mMTC. URLLC services refer to future applications which require reliable data communications from one end to another, while fulfilling ultra-low latency constraints. In this paper, we highlight the requirements and mechanisms that are necessary for URLLC in LTE. Design challenges faced when reducing the latency in LTE are shown. The performance of short processing time and frame structure enhancements are analyzed. Our proposed DCI Duplication method to increase LTE control channel reliability is presented and evaluated. The feasibility of achieving low latency and high reliability for the IMT-2020 submission of LTE is shown. We further anticipate the opportunities and technical design challenges when evolving 3GPP's LTE and designing the new 5G NR standard to meet the requirements of novel URLLC services.
Thomas Fehrenbach, Rohit Datta, Baris Göktepe, Thomas Wirth, Cornelius Hellge
VTC Fall5
2017 HEVC tile based streaming to head mounted displays
abstract
360° video streaming to clients using Virtual Reality head mounted displays is a challenge for traditional video delivery. As transmission of the complete content in a desirable quality sacrifices a large fraction of available client and network resources, adaptivity to the user viewport promises substantial benefits. An efficient way to achieve viewport adaptive streaming without per-user or per-orientation encoding, i.e. essentially transcoding, is to make use of motion-constrained HEVC tiles. DASH can be used for tiled streaming, where tiled content resides on the server at multiple resolutions. The DASH client selects the resolutions of each tile according to the current viewport. This demonstration paper presents an agile and responsive streaming prototype system for 360° video content. In order to achieve acceptable responsiveness, the DASH client relies on small buffer sizes and shifted random access points across tiles when suitable.
Robert Skupin, Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl
CCNC4
2017 Interpretable human action recognition in compressed domain
abstract
Compressed domain human action recognition algorithms are extremely efficient, because they only require a partial decoding of the video bit stream. However, the question what exactly makes these algorithms decide for a particular action is still a mystery. In this paper, we present a general method, Layer-wise Relevance Propagation (LRP), to understand and interpret action recognition algorithms and apply it to a state-of-the-art compressed domain method based on Fisher vector encoding and SVM classification. By using LRP, the classifiers decisions are propagated back every step in the action recognition pipeline until the input is reached. This methodology allows to identify where and when the important (from the classifier's perspective) action happens in the video. To our knowledge, this is the first work to interpret a compressed domain action recognition algorithm. We evaluate our method on the HMDB51 dataset and show that in many cases a few significant frames contribute most towards the prediction of the video to a particular class.
Vignesh Srinivasan, Sebastian Lapuschkin, Cornelius Hellge, Klaus-Robert Müller, Wojciech Samek
ICASSP3
2017 Random access point period optimization for viewport adaptive tile based streaming of 360° video
abstract
360° video streaming introduces stricter requirements to the established transmission chain than in traditional streaming services. Transmission of the complete 360° video in desirable quality can lead to multiple times UHD resolution which would waste a large fraction of network resources as the larger fraction of the video is not presented on the end device. Adaptivity to the user viewport promises substantial benefits and contrary to per-user or per-orientation encoding, tiled based streaming provides a scalable solution. We consider tiled streaming of 360° video using the cubic projection, where tiled content resides on the server at 2 different resolutions. User traces have been collected for a range of content and based on consumption patterns, tiles at the server have been optimized. More concretely, the paper optimizes the Random Access Point period that leads to the minimum transmitted bitrate, while ensuring that users watch most of the time high resolution content.
Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl
ICIP3
2017 Viewport-dependent 360 degree video streaming based on the emerging Omnidirectional Media Format (OMAF) standard
abstract
360 degree video streaming has gained much interest recently. The Omnidirectional MediA Format (OMAF) standard, currently in development by the Moving Picture Experts Group (MPEG), standardizes means for storage and delivery of 360 degree coded video based on a well-established standards ecosystem. This demonstration system uses the viewport-dependent OMAF media profile in which visual content within the current viewport is transmitted and displayed in higher fidelity than content outside the viewport to exceed the visual fidelity of viewport-independent solutions. This is achieved by dividing the video frame into tiles and using HEVC motion-constrained tile set (MCTS) encoding at multiple resolutions.
Robert Skupin, Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl
ICIP4
2017 Tile based panoramic streaming using shifted IDR representations
abstract
Panoramic streaming enables users to interactively navigate through high-spatial resolution videos and create an immersive and personalized user experience. Since transmission of high-resolution videos in desirable quality is not feasible given the limited throughput of access and home network links, our work is based on tile-based streaming, where only a spatial subset of the video is transmitted. In this paper, we propose a highly responsive DASH client algorithm that allows users to rapidly change the set of downloaded tiles. The high responsiveness is achieved by using small buffers, which are usually very sensitive to variations of throughput and instantaneous media bitrate and are therefore prone to playback interruptions. The proposed rate adaptation algorithm for DASH, in combination with a peak bitrate reducing RAP configuration referred to as shifted IDRs, outperforms the state of the art rate adaptation algorithms. Experiments report a decreased number of playback interruptions and quality changes while maintaining the average video quality.
Dimitri Podborski, Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl
ICME4
2017 Spatio-Temporal Activity based Tiling for Panorama Streaming
abstract
In panorama streaming an arbitrary Region-of-Interest (RoI) of a high-resolution video is transmitted, allowing users to navigate interactively within the videos. Transmitting the whole video becomes unfeasible due to the required high bitrates and sending a single video per user, which is encoded for that specific user (i.e. its RoI), has scalability issues. Tile based panoramic streaming overcomes the mentioned drawbacks by allowing users to receive a set of tiles that match their RoI instead of the whole set of tiles. However, optimal tiling -- so that the transmitted bitrate of the RoI content is minimized - is content dependent. In this paper, we propose a model based on a spatio-temporal activity metric so that optimization of the tiling process can be performed in a low complexity manner.
Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl
NOSSDAV3
2016 Multi-code Distributed Storage
abstract
Distributed storage systems need to guarantee reliable access to stored data. Resilience to node failures can be increased by using erasure encoding. A variety of erasure codes are discussed in literature and implemented in practice. This multiplicity of codes puts a heavy burden on existing systems. In scenarios such as multi-cloud file delivery or migration of data to a new erasure code, the ability to combine data from diverse erasure codes of multiple cloud systems is essential. The methods presented in this paper enable combining symbols of different erasure codes regardless of their underlying generator matrix, finite field size, and source block size. Mathematical approaches are discussed using Reed-Solomon and RLNC codes as example but without loss of generality. The presented approaches enable multi-cloud file delivery across diverse coding algorithms and permits graceful migration of a legacy erasure coding without the need of re-ingestion of existing data.
Cornelius Hellge, Muriel Médard
CLOUD1
2016 Peak bitrate reduction for multi-party video conferencing using SHVC
abstract
Real-time video applications, such as multi party video conferencing, involve the simultaneous transport of multiple and potentially multi-layered video sources to participating or interested parties. It is desirable to mix these multiple source videos into a single video stream at intermediary nodes in the network, e.g. at Multipoint Control Units (MCU). This has the advantage of reduced application and transport complexity on the client device while allowing single hardware decoder devices to consume the content. This paper proposes a solution, which uses the scalable extension of H.265/HEVC (SHVC), and that generates a single bitstream out of several sources by a low-complexity operation. In addition, the presented technique drastically reduces the peak bitrate at layout change events in comparison to state-of-the-art solutions.
Yago Sánchez de la Fuente, Robert Skupin, Cornelius Hellge, Thomas Schierl
ICIP3
2016 Shifted IDR Representations for Low Delay Live DASH Streaming Using HEVC Tiles
abstract
Dynamic Adaptive Streaming over HTTP (DASH) has gained a lot of popularity in the last years and has become a widespread solution for media delivery. It is mainly used for VoD services but has being also standardized and can be used for live scenarios. Although there has been a lot of work on low delay DASH services, it is still very challenging to cope with network variations. In fact, low delay configurations prevent DASH users from having a large media buffer to cope with network variations and they can easily lead to deadline misses, which results in a service of low Quality of Experience (QoE). In this paper we propose a solution based on HEVC Tiles and an advanced DASH adaptation algorithm that improves drastically the performance by heavily reducing the deadline misses.
Yago Sánchez de la Fuente, Dimitri Podborski, Cornelius Hellge, Thomas Schierl
ISM3
2016 Tile Based HEVC Video for Head Mounted Displays
abstract
360° Video services with resolutions of UHD and beyond for Virtual Reality head mounted displays are a challenging task due to limits of video decoders in constrained end devices. Adaptivity to the current user viewport is a promising approach but incurs significant encoding overhead when encoding per user or set of viewports. A more efficient way to achieve viewport adaptive streaming is to facilitate motion-constrained HEVC tiles. Original content resolution within the user viewport is preserved while content currently not presented to the user is delivered in lower resolution. A lightweight aggregation of varying resolution tiles into a single HEVC bitstream can be carried out on-the-fly and allows usage of a single decoder instance on the end device.
Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl
ISM3
2016 Hybrid video object tracking in H.265/HEVC video streams
abstract
In this paper we propose a hybrid tracking method which detects moving objects in videos compressed according to H.265/HEVC standard. Our framework largely depends on motion vectors (MV) and block types obtained by partially decoding the video bit stream and occasionally uses pixel domain information to distinguish between two objects. The compressed domain method is based on a Markov Random Field (MRF) model that captures spatial and temporal coherence of the moving object and is updated on a frame-to-frame basis. The hybrid nature of our approach stems from the usage of a pixel domain method that extracts the color information from the fully-decoded I frames and is updated only after completion of each Group-of-Pictures (GOP). We test the tracking accuracy of our method using standard video sequences and show that our hybrid framework provides better tracking accuracy than a state-of-the-art MRF model.
Serhan Gul, Jan Timo Meyer, Cornelius Hellge, Thomas Schierl, Wojciech Samek
MMSP3
2013 Enhancement of Pro-MPEG COP3 codes and application to layer-aware FEC protection of two-layered video transmission
abstract
In this paper we propose an enhancement of the Application-Layer FEC codes introduced by the Pro-MPEG Forum in its Code of Practice 3 r2 (Pro-MPEG COP3 codes) through allowing the introduction of a third dimension. The potential addition of an extra set of protection packets augments the number of possible combinations of data packets within a FEC block for parity packet computation. This enables a finer optimization process of the parameters of the FEC codes for a better adaptation to the specific conditions of the communication channel, increasing their capability. Additionally, we propose a Layer-Aware FEC scheme in which the enhanced Pro-MPEG COP3 codes are used to protect two-layered video streams. Experiment results reveal a gain in the introduction of this protection mechanism, when compared to the standard codes.
César Díaz, Cornelius Hellge, Julián Cabrera, Fernando Jaureguizar, Thomas Schierl
ICIP2
2012 Efficient HTTP-based streaming using Scalable Video Coding
Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec
Signal Process. Image Commun.3
2011 Improved caching for HTTP-based Video on Demand using Scalable Video Coding
abstract
HTTP-based delivery for Video on Demand (VoD) has been gaining popularity within recent years. Progressive Download over HTTP, typically used in VoD, takes advantage of the widely deployed network caches to release video servers from sending the same content to a high number of users in the same VoD service. However, due to the inherent heterogeneity of user demands, which may result in requesting the same video content in different resolutions or qualities, the caching efficiency is expected to decrease due to a higher variety in requested media files. The use of Scalable Video Coding allows different representations of the same content to be combined in a single file, whose parts, aka layers, are requested sequentially by a user up to the maximum desired quality. In this paper we show the benefits of using Scalable Video Coding to maintain the same set of possible video content representations, while at the same time maximizing the caching efficiency.
Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec
CCNC3
2011 Capacity improvement in EMBMS using SVC and layer-aware bearer allocation
abstract
Today's solution to realize wide coverage with the Release 9 MBMS specification is to transmit an H.264/AVC video service with a robustness determined by the worst reception conditions of the target area. This sacrifices service data rate for robustness due to lower modulation and FEC code-rate. The combination of bearer channel multiplex and SVC promises significant gains in terms of channel capacity by establishing unequal error protection on the physical layer. In comparison with single layer transmission, SVC achieves equal performance in terms of continuous playout while allocating fewer resources for the expensive robust channel. This benefit is gained by providing temporary lower quality video to users within bad reception conditions. Simulation results within an MBSFN Release 9 network show significant capacity gains over a given bandwidth with only minor quality degradation for users located in bad reception areas while preserving playout robustness of H.264/AVC single layer services.
Cornelius Hellge, Robert Skupin, Jaihyung Cho, Thomas Schierl, Thomas Wiegand 0001
ICIP1
2011 iDASH: improved dynamic adaptive streaming over HTTP using scalable video coding
abstract
HTTP-based delivery for Video on Demand (VoD) has been gaining popularity within recent years. Progressive Download over HTTP, typically used in VoD, takes advantage of the widely deployed network caches to relieve video servers from sending the same content to a high number of users in the same access network. However, due to a sharp increase in the requests at peak hours or due to cross-traffic within the network, congestion may arise in the cache feeder link or access link respectively. Since the connection characteristics may vary over the time, with Dynamic Adaptive Streaming over HTTP (DASH), a technique that has been recently proposed, video clients may dynamically adapt the requested video quality for ongoing video flows, to match their current download rate as good as possible. In this work we show the benefits of using the Scalable Video Coding (SVC) for such a DASH environment.
Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001, Dohy Hong, Danny De Vleeschauwer, Werner Van Leekwijck, Yannick Le Louédec
MMSys3
2011 Priority-based Media Delivery using SVC with RTP and HTTP streaming
Thomas Schierl, Yago Sánchez de la Fuente, Ralf Globisch, Cornelius Hellge, Thomas Wiegand 0001
Multim. Tools Appl.4
2011 Layer-Aware Forward Error Correction for Mobile Broadcast of Layered Media
abstract
The bitstream structure of layered media formats such as scalable video coding (SVC) or multiview video coding (MVC) opens up new opportunities for their distribution in Mobile TV services. Features like graceful degradation or the support of the 3-D experience in a backwards-compatible way are enabled. The reason is that parts of the media stream are more important than others with each part itself providing a useful media representation. Typically, the decoding of some parts of the bitstream is only possible, if the corresponding more important parts are correctly received. Hence, unequal error protection (UEP) can be applied protecting important parts of the bitstream more strongly than others. Mobile broadcast systems typically apply forward error correction (FEC) on upper layers to cope with transmission errors, which the physical layer FEC cannot correct. Today's FEC solutions are optimized to transmit single layer video. The exploitation of the dependencies in layered media codecs for UEP using FEC is the subject of this paper. The presented scheme, which is called layer-aware FEC (LA-FEC), incorporates the dependencies of the layered video codec into the FEC code construction. A combinatorial analysis is derived to show the potential theoretical gain in terms of FEC decoding probability and video quality. Furthermore, the implementation of LA-FEC as an extension of the Raptor FEC and the related signaling are described. The performance of layer-aware Raptor code with SVC is shown by experimental results in a DVB-H environment showing significant improvements achieved by LA-FEC.
Cornelius Hellge, David Gomez-Barquero, Thomas Schierl, Thomas Wiegand 0001
IEEE Trans. Multim.1
2010 P2P group communication using Scalable Video Coding
abstract
P2P-streaming has become of high interest in the last years, since it reduces the load on expensive servers, due to the participation of receivers in the media transmission. In this paper, P2P content delivery is shown to be a promising technique for video group communication, for which the main requirement is low-delay. Combining low-delay encoding and low-delay P2P Application Layer Multicast makes it possible to fulfill the delay constraints for interactive group communication applications. For such an application, congestion is a considerable problem, since it causes packet loss or late arrival of the packets, degrading the quality of the service. The results presented in this paper show how rate adaptation in combination with the Scalable Video Coding (SVC) helps to overcome problems in the network, providing a better solution than when non-adaptive single layer coding is transmitted.
Yago Sánchez de la Fuente, Thomas Schierl, Cornelius Hellge, Thomas Wiegand 0001
ICIP3
2010 Intra-burst layer aware FEC for scalable video coding delivery in DVB-H
abstract
This paper investigates how Scalable Video Coding (SVC) can benefit from different Forward Error Correction (FEC) and transmission schemes in mobile broadcast systems. Simulation are performed in DVB-H (Digital Video Broadcasting - Handheld) systems. In DVB-H, a differentiation in robustness for the different SVC layers can be achieved at the link layer using intra-burst MPE-FEC (multi-Protocol Encapsulation FEC). The paper evaluates the gain that can be achieved with the MPE-FEC using equal (EEP) and unequal error protection (UEP), and the performance improvements of an SVC layer-aware FEC (LA-FEC) approach that can be implemented in DVB-H either at the link layer with MPE-iFEC (inter-burst MPE-FEC) or at the application layer with AL-FEC. LA-FEC improves the SVC base layer robustness generating parity information across existing dependencies within the SVC video coding structure. Laboratory measurement results using a TU6 channel model show that the performance of an SVC service in DVB-H can be significantly increased by a proper link layer FEC scheme. It is shown that using SVC in combination with LA-FEC and a proper transmission scheduling does not only give a better performance in terms of PSNR and amount of video outages compared to a single layer service at the same service bit rate, but also gives an additional lower quality layer which can be used for applications like conditional access.
Cornelius Hellge, David Gomez-Barquero, Thomas Schierl, Thomas Wiegand 0001
ICME1
2010 Study of the performance of Scalable Video Coding over a DVB-SH satellite link
abstract
In this paper, we will present and discuss selected results obtained for the ESA study on “Scalable Video Coding Applications and Technologies for mobile satellite based hybrid networks”. The reference use case under investigation concerns the exploitation of heterogeneous reception conditions on a DVB-SH satellite link via scalability and unequal error protection.
Günther Liebl, Ktawut Tappayuthpijarn, Thiago Martins de Moraes, Karsten Grüneberg, Cornelius Hellge, Thomas Schierl, Cedric Keip, Holger Stadali, Nghia Pham
ICME5
2010 Improving P2P live-content delivery using SVC
abstract
P2P content delivery techniques for video transmission have become of high interest in the last years. With the involvement of client into the delivery process, P2P approaches can significantly reduce the load and cost on servers, especially for popular services. However, previous studies have already pointed out the unreliability of P2P-based live streaming approaches due to peer churn, where peers may ungracefully leave the P2P infrastructure, typically an overlay networks. Peers ungracefully leaving the system cause connection losses in the overlay, which require repair operations. During such repair operations, which typically take a few roundtrip times, no data is received from the lost connection. While taking low delay for fast-channel tune-in into account as a key feature for broadcast-like streaming applications, the P2P live streaming approach can only rely on a certain media pre-buffer during such repair operations. In this paper, multi-tree based Application Layer Multicast as a P2P overlay technique for live streaming is considered. The use of Flow Forwarding (FF), a.k.a. Retransmission, or Forward Error Correction (FEC) in combination with Scalable video Coding (SVC) for concealment during overlay repair operations is shown. Furthermore the benefits of using SVC over the use of AVC single layer transmission are presented.
Thomas Schierl, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Wiegand 0001
VCIP3
2008 Multidimensional Layered Forward Error Correction Using Rateless Codes
abstract
Modern layered or scalable video coding technologies generate a video bit stream with various inter layer dependencies due to references between the layers. This work proposes a method for extending forward error correction (FECs) codes following dependency structures within the media. The proposed layer-aware FEC (L-FEC) generates repair symbols so that protection of less important dependency layers can be used with protection of more important layers for combined error correction. The L-FEC approach is exemplary applied to rateless LT and Raptor codes. Gains for more important layers can be achieved without increasing the total FEC code rate. The performance gain of the L-FEC is shown by simulation results with receiver-driven layered multicast transmission using scalable video coding (SVC) with a Raptor-based L-FEC.
Cornelius Hellge, Thomas Schierl, Thomas Wiegand 0001
ICC1
2008 Temporal scalability and layered transmission
abstract
The deployment of mobile multimedia broadcast services like mobile TV over cellular networks has just started. The coverage in terms of delivered quality per receiver is an important measure to evaluate the experienced quality. In order to provide a reasonable quality for all receivers in a network cell, we analyze graceful degradation techniques in Rel. 6 of 3rd Generation Partnership Project's (3QPP) Multimedia Broadcast and Multicast Services (MBMS). We show the benefits of using H.264/AVC temporal scalability in an existing system for graceful degradation in order to add coverage for users in poor reception conditions, although at a reduced level of quality. For that we define assessment techniques for measuring the quality coverage describing the amount of users reached with a certain media play-out quality. We connect this quality coverage to a measure for used network cell capacity giving the used network resources for 3QPP Rel. 6 network cells. We evaluate graceful degradation under simulated 3QPP Rel. 6 network conditions showing the benefits of graceful degradation for 3QPP MBMS mobile television services.
Cornelius Hellge, Thomas Schierl, Jörg Huschke, Thomas Rusert, Markus Kampmann, Thomas Wiegand 0001
ICIP1
2008 Receiver driven layered multicast with layer-aware forward error correction
abstract
A wide range of different capabilities and connection qualities typically characterizes receivers of mobile television services. Receiver driven layered multicast (RDLM) offers an efficient way for providing different capabilities over such a broadcast channel. Scalable video coding (SVC) allows for the transmission of multiple video qualities within one media stream. Using SVC generates a video bit stream with various inter layer dependencies due to references between the layers. This work proposes a layer-aware forward error correction (L-FEC) approach in combination with SVC. L-FEC increases robustness of the more important layers by generating protection across layers following existing dependencies of the media stream. The L-FEC is integrated as an extension of a Raptor FEC implementation in a DVB-H broadcast system. It is shown by experimental results that L-FEC outperforms traditional UEP protection schemes.
Cornelius Hellge, Thomas Schierl, Thomas Wiegand 0001
ICIP1
2008 Mobile TV using scalable video coding and layer-aware forward error correction
abstract
A new approach to error protection for scalable media is presented for mobile TV applications. Mobile TV is typically characterized by a number of receiver capabilities and connection qualities. A broadcast service should preferably work for multiple receiver capabilities without the need for downscaling or transcoding at the battery-powered mobile devices. Moreover, a media quality that gracefully degrades with reception quality instead of a complete signal loss is also a desirable feature. The scalable video coding (SVC) extension of H.264/AVC offers an efficient way to support the aforementioned features. In mobile broadcast channels forward error correction (FEC) is used to overcome packet losses. This work proposes a layer-aware forward error correction (L-FEC) approach in combination with SVC. L-FEC increases robustness of the more important layers by generating protection across layers. L-FEC is integrated as an extension of an Raptor FEC implementation. It is shown by experimental results that L-FEC outperforms traditional FEC and UEP protection schemes.
Cornelius Hellge, Thomas Schierl, Thomas Wiegand 0001
ICME1
2007 Distributed Rate-Distortion Optimization for Rateless Coded Scalable Video in Mobile Ad Hoc Networks
abstract
Recent advances in forward error correction and scalable video coding enable new approaches for robust, distributed streaming in mobile ad hoc networks (MANETs). This work presents an approach for distribution of real time video by different uncoordinated peer-to-peer relay or source nodes in an overlay network on top of a MANET. The approach proposed here allows for distributed, rate-distortion optimized transmission-rate allocation for competing scalable video streams at relay nodes in the overlay network. Furthermore the approach has the desirable feature of path/source diversity for enhancing reliability in connectivity to serving nodes. Signaling overhead within the overlay network is kept at a minimum, since optimizations are done at relay nodes and clients rather than at servers.
Thomas Schierl, Stian Johansen, Cornelius Hellge, Thomas Stockhammer, Thomas Wiegand 0001
ICIP (6)3
2007 Using H.264/AVC-based Scalable Video Coding (SVC) for Real Time Streaming in Wireless IP Networks
abstract
A streaming system based on the scalable video coding (SVC) extension of H.264/AVC is shown. SVC allows for data rate adaptation without re-encoding just by dropping packets of the bit stream. By that, it enables for instance multicast services to clients of heterogeneous capabilities at the same time, while consuming less bit rate compared to simulcasting the services. Additionally, the robustness of a streaming connection against packet losses can be significantly increased if the different layers of the coded video stream are unequally protected by a forward error correction scheme. In this case, users will experience graceful degradation of image quality rather than visible errors or interruptions. The introduction of new services using SVC can also take advantage of a backward-compatible base layer. Different use cases for SVC streaming in wireless IP networks and selected simulation results of the improved error robustness for a DVB-H wireless transmission are presented.
Thomas Schierl, Cornelius Hellge, Shpend Mirta, Karsten Grüneberg, Thomas Wiegand 0001
ISCAS2
2006 Multi Source Streaming for Robust Video Transmission in Mobile Ad-Hoc Networks
abstract
Video transmission over networks is getting increasingly important, thus upcoming network types like mobile ad-hoc networks (MANETs) can also become a suitable platform for exchanging/sharing real-time video streams. We present an approach using different sources with linear independent representations of a layered video stream for increasing the robustness in transmission. This approach is based on the scalable video coding (SVC) extensions of H.264/MPEG-4 AVC with different layers for assigning importance for transmission. Additionally a novel unequal packet loss protection (UPLP) scheme based on Raptor forward error correction codes is employed. This scheme allows for reception from different sources and benefits from it. While the reception of a single stream guarantees base quality at least, the combined reception enables play-back of video of full quality and/or lower error rates.
Thomas Schierl, Cornelius Hellge, Karsten Gänger, Thomas Stockhammer, Thomas Wiegand 0001
ICIP2