VLDB 2026 Research / reviewers in the wild / expert
Mathias Wien
dblp:14/3667
· DBLP profile ↗
66ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-8724-2752ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 10 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entropy Coding for Non-Rectangular Transform Blocks Using Partitioned DCT Dictionaries for AV1abstractRecent video codecs, e.g. AV1, VVC, apply a Non-rectangular (NR) partitioning to combine prediction signals using a smooth blending around the boundary, followed by a rectangular transform (TX) on the whole block. TX on each NR residual separately is not yet supported. A recent NR TX technique [1] demonstrated promising gains in an experimental setup outside the reference software. This method employs the regular inverse-DCT at the decoder to reconstruct a rectangular signal while discarding the signal outside the region of interest. This design is appealing due to the minimal changes required at the decoder. The method uses a partitioned 2D DCT as a dictionary to find a sparse representation of the NR signal, with scaled representations serving as TX coefficients. These coefficients typically exhibit properties distinct from those of DCT TX coefficients. Therefore, the established entropy coding schemes in video codecs, which are primarily designed for DCT coefficients, are not well-suited for optimally encoding these TX coefficients. Priyanka Das 0005, Tim Classen, Mathias Wien |
DCC | 3 |
| 2025 | Adaptive Smoothing of Non-Rectangular Prediction Block Edges in the Wedge Mode of AVMabstractThis work introduces an additional boundary for extended wedge mode in the reference software of AVM. Wedge mode, introduced in AV1 and modified later, employs a non-rectangular partitioning mode to combine two prediction signals. However, the available wedge masks might not be sufficient to handle diverse video content. In this work, an adaptive boundary selection scheme to construct the wedge masks is proposed. A set of two boundaries are used for this purpose. Additionally, three signaling schemes are presented, with different levels of complexity, and one low-complexity design which limits the adaptivity to smaller block sizes while maintaining similar performance. Experimental results using AVM common testing conditions showed promising gains. In random access configuration, class A5 demonstrates 0.89% PSNR-Y BD rate gain, while classes A1, B1, and A3 show approximately 0.08% PSNR-Y BD rate gain. Priyanka Das 0005, Tim Classen, Mathias Wien |
ICIP | 3 |
| 2025 | Loop Filters and Edge Enhancement for Variable Resolution Video Coding
Tim Claßen, Xiang Li 0003, Priyanka Das 0005, Mathias Wien |
PCS | 4 |
| 2025 | Video Quality Assessment with Spatio-Temporal Interleaving applied for Film Grain Synthesis Evaluation
Luan Shkurti, Mathias Wien |
PCS | 2 |
| 2025 | A Method for Rate Point Determination for Visual Evaluation of Video Sequences
Mathias Wien, Adam Wieckowski, Elena Alshina, Edouard François, Pavel Nikitin, Kenneth Andersson |
PCS | 1 |
| 2025 | Parameter Dependent Wedge Boundary Switching in AVMabstractThe wedge mode in Alliance of Open Media Video Model (AVM) divides a block into two non-rectangular regions to combine two prediction signals. Rather than employing a sharp transition between these signals, the wedge mode utilizes a gradual transition around the boundary. Currently, there is only one such wedge boundary integrated in the AVM reference software, which is insufficient for handling diverse video content. Recently, an additional relatively smooth boundary has been explored using several signaling schemes. While this adaptive wedge boundary demonstrated promising improvements for short sequences, the signaling overhead associated with it restricted its overall potential. The current work uses this smooth boundary and proposes three simple switching schemes, where the boundary is selected based on specific parameters. While this approach reduces flexibility, it mitigates the requirement for signalling, which has proven to be more beneficial. This provides a PSNR-Y BD-rate gain of up to 0.05% in Random Access (RA) and 0.10% gain in Low Delay (LD) configuration. The improvements are particularly significant for high-resolution classes; in class A1, one scheme achieves 0.38% gain in RA, and another 0.24% gain in LD. A High Level Syntax scheme is presented, which combines the performance of 0.10% in RA and 0.12% in LD by signalling 2 bits per sequence. Priyanka Das 0005, Tim Classen, Mathias Wien |
VCIP | 4 |
| 2024 | Complexity Reduction of Template Matching-Based Reference Picture Padding in Video CodingabstractReference Picture Padding removes the restriction for motion vectors to point completely inside the reference picture. Removing this restriction increases the compression efficiency and is, therefore, applied in many video coding standards. However, artifacts can occur at the picture boundaries if the surroundings of the reference picture are not predicted accurately by the employed method. This especially becomes a problem in viewport-adaptive streaming scenarios if the independently decodable subpictures of Versatile Video Coding (VVC) are used. In this case, the subpicture boundaries behave equivalently to picture boundaries. The boundary-related artifacts then may become visible in the viewport. It has been shown that artifacts can be reduced by using a template matching-based padding algorithm. The main drawback of this algorithm is its high computational complexity. We propose a complexity reduction of the search step by reusing the results of previous searches and an advanced chroma handling. The proposed algorithm reduces the decoder runtime increase from 31.3% to 19.2% alongside slightly increased compression efficiency in a subpicture-coding scenario. An alternative variant of the algorithm reduces the decoder runtime increase to 12.9% while only sacrificing about 15% of the compression gains compared to the first variant. Nicolas Horst, Mathias Wien |
ICASSP | 2 |
| 2024 | Fast Template Matching-Based Reference Picture Padding for Video CodingabstractReference Picture Padding is utilized in a variety of video coding standards. It allows motion vectors to point partly outside the reference picture in inter prediction. This approach provides advantages in compression efficiency. The method employed in current standards is repetative padding, which copies the pixels at the picture border outward. The major advantage of this approach is its low complexity. However, it is not optimal in terms of compression efficiency. Recently, motion-compensated padding has been introduced, which utilizes already coded content in the padding process. Another padding method is template matching-based padding, with the main downside being its high computational complexity. In this work, we propose a complexity reduction of the method by applying different measures, including an improved virtual target block increase and an early termination of the search, among other improvements. As a result, we achieve a decoder runtime increase of 1% and a BD-rate of -0.33% in a low delay subpicture scenario, outperforming motion-compensated padding with a decoder runtime increase of 9% and a BD-rate of -0.19%. In the random-access configuration we can show a consistent improvement by combining the methods over the stand-alone methods. Nicolas Neumann, Priyanka Das 0005, Tim Classen, Mathias Wien |
ICIP | 4 |
| 2024 | Balancing Complexity of Template Matching-Based Reference Picture Padding for Video CodingabstractReference picture padding is the task of padding the outside of the reference picture for inter prediction. This task is done to accommodate for motion vectors that extend partly outside the picture, thereby increasing compression efficiency. The established approach involves simply duplicating the boundary pixels outward. While this solution boasts low complexity, it often results in suboptimal compression performance in numerous cases. Template matching-based reference picture padding presents itself as a promising method to further increase compression efficiency and reduce artifacts. However, its primary drawback lies in its high computational complexity. One potential solution entails restricting the number of considered candidates per search step to a very small set to maintain a feasibly low increase in decoder runtime. This, however, compromises the compression efficiency. In this study, we propose a novel approach that maintains the efficiency gains, while significantly reducing the increase in decoder runtime. This method primarily focuses the computational complexity on pixels more frequently utilized in inter prediction. Additionaly, we introduce an early stopping criterion that terminates the search if at least one similar candidate is found. Through these modifications we achieve a reduction in decoder runtime increase from 1252% to 159%, while maintaining −0.37% compared to −0.41% Bj⊘ntegaard delta rate in a Versatile Video Coding subpicture coding scenario. Nicolas Neumann, Priyanka Das 0005, Tim Claßen, Mathias Wien |
PCS | 4 |
| 2023 | Adaptive and Scalable Compression of Multispectral Images using VVCabstractThe VVC codec is applied to the task of multispectral image (MSI) compression using adaptive and scalable coding structures. In a “plain” VVC approach, concepts from picture-to-picture temporal prediction are employed for decorrelation along the MSI’s spectral dimension. The popular principle component analysis (PCA) for spectral decorrelation is further evaluated in combination with VVC intra-coding for spatial decorrelation. This approach is referred to as PCA-VVC. A novel adaptive MSI compression algorithm, named HPCLS, is introduced, that uses PCA and inter-prediction for spectral and VVC intra-coding for spatial decorrelation. Further, a novel adaptive scalable approach is proposed, that provides a separately decodable spectrally scaled preview of the MSI in the compressed file. Information contained in the preview is exploited in order to reduce the overall file size. All schemes are evaluated on images from the ARAD HS data set containing outdoor scenes with a high variety in brightness and color. We found that “Plain” VVC is outperformed by both PCA-VVC and HPCLS. HPCLS shows advantageous rate-distortion (RD) behavior compared to PCA-VVC for reconstruction quality above 51 dB PSNR. The performance of the scalable approach is compared to the combination of an independent RGB preview and one of HPCLS or PCA-VVC denoted as simulcast. The scalable approach shows significant benefit especially at higher preview qualities. A more detailed version of this article can be found on arXiv1. Philipp Seltsam, Priyanka Das 0005, Mathias Wien |
DCC | 3 |
| 2023 | A Template Matching Approach for Reference Picture Padding in Video CodingabstractReference picture padding is needed in areas close to the picture boundary. It lifts the restriction of motion vectors not to point over the boundary when using inter prediction. Especially, when independently decodable subpictures are used e.g. in viewport-adaptive streaming, many boundaries occur, where padding is necessary. In such cases, the prediction quality of reference (sub)picture padding has an increased impact on the coding performance. Template matching has shown to work well for texture prediction in video coding. The paper shows that it also improves coding performance when applied in the context of reference picture padding. Experimental results on a set of test sequences demonstrate consistent coding gains and an average Bjøntegaard delta rate reduction of −0.38% and −0.48% for two sets of sequences with frequent utilization of reference picture padding. A drawback of template matching is its computational complexity. A preliminary investigation shows that the complexity can be reduced by a factor of 330 while maintaining about 60% of the Bjøntegaard delta rate savings. Nicolas Horst, Priyanka Das 0005, Mathias Wien |
ICASSP | 3 |
| 2023 | Weighted Edge Sharpening Filtering for Upscaled Content in Adaptive Resolution CodingabstractReference Picture Resampling (RPR) is an essential tool for quickly adapting to varying network conditions. It enables a resolution change without the introduction of an intra random access point (IRAP). This feature is particularly crucial in real-time transmission scenarios and has proven effective for compressing high-resolution pictures. However, a significant challenge of Reference Picture Resampling is the loss of high-frequency information in the downscaling operation. This results in blurred pictures after upscaling, which significantly affects the viewer experience. To address this issue, we propose an additional enhancement step after upscaling. The proposed enhancement method involves a locally weighted adaptive filter specifically designed for edge sharpening. Through the application of local weighting, we can avoid the common problems associated with linear high-pass filters. Picture-wise content adaptivity helps in handling different scene characteristics. The proposed method is implemented into the enhanced compression model 8.0 (ECM) and achieves performance gains for the joint video experts team (JVET) common testing conditions for reference picture resampling (RPR), with a Bjøntegaard delta-rate (BD-rate) reduction of -7.27%, -0.57%, and -0.32% for the Y, Cb, and Cr channels, respectively. Tim Claßen, Mathias Wien |
VCIP | 2 |
| 2021 | D3dlo: Deep 3d Lidar OdometryabstractLiDAR odometry (LO) describes the task of finding an alignment of subsequent LiDAR point clouds. This alignment can be used to estimate the motion of the platform where the LiDAR sensor is mounted on. Currently, on the well-known KITTI Vision Benchmark Suite state-of-the-art algorithms are non-learning approaches. We propose a network architecture that learns LO by directly processing 3D point clouds. It is trained on the KITTI dataset in an end-to-end manner without the necessity of pre-defining corresponding pairs of points. An evaluation on the KITTI Vision Benchmark Suite shows similar performance to a previously published work, DeepCLR [1], even though our model uses only around 3.56% of the number of network parameters thereof. Furthermore, a plane point extraction is applied which leads to a marginal performance decrease while simultaneously reducing the input size by up to 50%. Philipp Adis, Nicolas Horst, Mathias Wien |
ICIP | 3 |
| 2021 | Adaptive Boundary Extension for Inter PredictionabstractBoundary extension refers to the extension of a picture boundary to enable inter prediction from regions outside the picture. In current video coding schemes, only non-adaptive approaches are used with a constant prediction which continues the boundary samples in the extension region. This causes artifacts in the region of the boundary which may be strongly visible. Especially when 360° video is coded using independently decodable subpictures, extension-related artifacts can occur at all subpicture boundaries and are not limited to the picture boundary area. Thereby, the handling of subpicture boundaries becomes more important. In this paper, an adaptive boundary extension method is investigated with explicit signaling that uses angular prediction for the extension task. It is shown that angular prediction modes are promising candidates for an extension by isolating the impact of the prediction improvement from the signaling cost. The scheme is implemented in the VVC test model, with a simple signaling method that leads to coding gain for over 40% of the subpictures. Preliminary results indicate Bj0ntegaard delta rate savings of about 0.1% when only selected subpictures are considered. This can be considered significant given that only a small area of the prediction signal is affected by the method. A major advantage of the explicit signaling approach is seen in the fact that the encoder can influence the predictions in the boundary region, such that subpicture transitions are more consistent. Nicolas Horst, Priyanka Das 0005, Mathias Wien |
PCS | 3 |
| 2021 | RDPlot - An Evaluation Tool for Video Coding SimulationsabstractRDPlot is an open source GUI application for plotting Rate-Distortion (RD)-curves and calculating Bjøntegaard Delta (BD) statistics [1]. It supports parsing the output of commonly used reference software packages, parsing *.csv-formatted files, and *.xml-formatted files. Once parsed, RDPlot offers the ability to evaluate video coding results interactively. Conceptually, several measures can be plotted over the bitrate and BD measurements can be conducted accordingly. Moreover, plots and corresponding BD statistics can be exported, and directly integrated into LaTeX documents. Jens Schneider 0001, Johannes Sauer, Mathias Wien |
VCIP | 3 |
| 2021 | Special issue on Open Media Compression: Overview, Design Criteria, and Outlook on Emerging StandardsabstractUniversal access to and provisioning of multimedia content is now a reality. It is easy to generate, distribute, share, and consume any multimedia content, anywhere, anytime, or any device. Open media standards took a crucial role toward enabling all these use cases leading to a plethora of applications and services that have now become a commodity in our daily life. Interestingly, most of these services adopt a streaming paradigm, are typically deployed over the open, unmanaged Internet, and account for most of today’s Internet traffic. Currently, the global video traffic is greater than 60% of all Internet traffic[1], and it is expected that this share will grow to more than 80% in the near future[2]. In addition, Nielsen’s law of Internet bandwidth states that the users’ bandwidth grows by 50% per year, which roughly fits data from 1983 to 2019[3]. Thus, the users’ bandwidth can be expected to reach approximately 1 Gb/s by 2022. At the same time, network applications will grow and utilize the bandwidth provided, just like programs and their data expand to fill the memory available in a computer system. Most of the available bandwidth today is consumed by video applications, and the amount of data is further increasing due to already established and emerging applications, e.g., ultrahigh definition, high dynamic range, or virtual, augmented, mixed realities, or immersive media applications in general. Christian Timmerer, Mathias Wien, Lu Yu 0003, Amy R. Reibman |
Proc. IEEE | 2 |
| 2020 | Coding Of Non-Rectangular Signals With Block-Based TransformsabstractThis paper presents a transform coding technique for non-rectangular 2-D signals by extending the signal into a rectangular block in order to enable conventional block-based transform coding. The technique could be suitable for coding residuals of prediction blocks using geometric partitioning which has been adopted into the draft Versatile Video Coding standard. The extension of the non-rectangular signal is found using a sparse solution set generated by applying Orthogonal Matching Pursuits using partitioned transform bases. The method developed in this paper is based on Discrete Cosine Transform. Results achieved in an experimental setup outside of the video coding loop are presented for signals of triangular and trapezoidal shape in comparison to the shape-adaptive DCT. Encouraging gains are observed specifically for larger block sizes and in dependency of the quantization parameter and the partitioning shape. Priyanka Das 0005, Nicolas Horst, Mathias Wien |
ICIP | 3 |
| 2020 | Versatile Video Coding - Algorithms and SpecificationabstractThe tutorial provides an overview on the latest emerging video coding standard VVC (Versatile Video Coding) to be jointly published by ITU-T and ISO/IEC. It has been developed by the Joint Video Experts Team (JVET), consisting of ITU-T Study Group 16 Question 6 (known as VCEG) and ISO/IEC JTC 1/SC 29/WG 11 (known as MPEG). VVC has been designed to achieve significantly improved compression capability compared to previous standards such as HEVC, and at the same time to be highly versatile for effective use in a broadened range of applications. Some key application areas for the use of VVC particularly include ultra-high-definition video (e.g. 4K or 8K resolution), video with a high dynamic range and wide colour gamut (e.g., with transfer characteristics specified in Rec. ITU-R BT.2100), and video for immersive media applications such as 360° omnidirectional video, in addition to the applications that have commonly been addressed by prior video coding standards. Important design criteria for VVC have been low computational complexity on the decoder side and friendliness for parallelization on various algorithmic levels. VVC is planned to be finalized by July 2020 and is expected to enter the market very soon.The tutorial details the video layer coding tools specified in VVC and develops the concepts behind the selected design choices. While many tools or variants thereof have been available before, the VVC design reveals many improvements compared to previous standards which result in compression gain and implementation friendliness. Furthermore, new tools such as the Adaptive Loop Filter, or Matrix-based Intra Prediction have been adopted which contribute significantly to the overall performance. The high-level syntax of VVC has been re-designed compared to previous standards such as HEVC, in order to enable dynamic sub-picture access as well as major scalability features already in version 1 of the specification. Mathias Wien, Benjamin Bross |
VCIP | 1 |
| 2018 | Adaptive Coding of Non-Negative Factorization Parameters with Application to Informed Source SeparationabstractInformed source separation (ISS) uses source separation for extracting audio objects out of their downmix given some pre-computed parameters. In recent years, non-negative tensor factorization (NTF) has proven to be a good choice for compressing audio objects at an encoding stage. At the decoding stage, these parameters are used to separate the downmix with Wiener-filtering. The quantized NTF parameters have to be encoded to a bit stream prior to transmission. In this paper, we propose to use context-based adaptive binary arithmetic coding (CABAC) for this task. CABAC is widely used in the video coding community and exploits local signal statistics. We adapt CABAC to the task of NTF-based ISS and show that our contribution outperforms reference coding methods. Max Bläser, Christian Rohlfing, Yingbo Gao, Mathias Wien |
ICASSP | 4 |
| 2018 | Pyramid Pooling of Convolutional Feature Maps for Image RetrievalabstractWe propose a novel method for content based image retrieval based on the features extracted from the convolutional layers of the deep neural network architecture. Some of the popular approaches form the feature vectors from the fully connected layers of the convolutional neural networks or directly concatenate the features from the convolutional layers. However, the main problem with the use of feature vectors from fully connected layers is that the spatial information about the objects are lost. This motivated us to use the features from the convolutional layer. We incorporate a pyramid pooling based approach to form more compact and location invariant feature vectors. We have measured the Mean Average Precision (MAP) on benchmark databases such as the Holidays and Oxford5K datasets using features extracted from the AlexNet model. The proposed method gives better retrieval results compared to other state-of-the-art approaches which use feature vectors from fully connected layers and convolutional layers without spatial pooling. Abin Jose, Ricard Durall, Iris Heisterklaus, Mathias Wien |
ICIP | 4 |
| 2018 | Geometry-based Partitioning for Predictive Video Coding with Transform AdaptationabstractRectangular block partitioning as it is used in state of the art video codecs such as HEVC can produce visually displeasing artifacts at low bitrates. This effect is particularly noticeable at moving object boundaries. This contribution presents a comprehensive geometry-based block partitioning framework in a post-HEVC codec for motion compensated prediction, intra-prediction and transform coding as a solution. The method is evaluated on the set of sequences defined by the Joint Call for Proposals on Video Compression with Capabilities beyond HEVC [1]. Our contribution aims at visually improving the quality of object boundaries and provides an objective BD-rate gain of 0.82% on average compared to the reference Joint Video Exploration Team (JVET) test model (JEM 7.0). Max Bläser, Jens Schneider 0001, Johannes Sauer, Mathias Wien |
PCS | 4 |
| 2018 | Motion-Distribution based Dynamic Texture Synthesis for Video CodingabstractIn this paper, a new approach for an improved video coding scheme is presented, which combines hybrid video coding and texture synthesis based on motion distribution statistics. Considering that the utilized texture synthesis approach provides high-quality visual results, while it is developed only for synthe- sizing the identified dynamic textures within a certain area, a new framework is presented, which allows to identify of areas for synthesis and combine conventional coding with synthesis. Also, a new representation and compression of synthesis parameters is presented, which is required due to the updated coding structure. When combining the proposed approach with conventional en- coder (HEVC reference software, HM 16.6), significantly reduced bit rates of the compressed video sequences with the texture replaced can be obtained. Moreover, because the synthesized textures have similar perceptual characteristics to those of the original textures, the video sequences with the texture replaced are also visually similar to the original sequences. Video results are provided online to allow assessing the visual quality of the tested content. Olena Chubach, Patrick Garus, Mathias Wien, Jens-Rainer Ohm |
PCS | 3 |
| 2018 | Image-Based Rendering using Point Cloud for 2D Video CompressionabstractThe main idea of this paper is to extract the 3D scene geometry for the observed scene and use it for synthesizing a more precise prediction using Image-Based Rendering (IBR) for motion compensation in a hybrid coding scheme. The proposed method first extracts camera parameters using Structure from Motion (SfM). Then, a Patch-based Multi-View Stereo (PMVS) technique is employed to generate the scene Point-Cloud (PC) only from already decoded key-frames. Since the PC could be really sparse in poorly reconstructed regions, a depth expansion mechanism is also used. This 3D information helps to properly warp textures from the key-frames to the target frame. This IBR-based prediction is then used as an additional reference for motion compensation. In this way, the encoder can choose between the rendered prediction and the regular reference pictures through a rate- distortion optimization. On average, the simulation results show about 2.16% bitrate reduction compared to the reference HEVC implementation, for tested dynamic and static scene video sequences. Hossein Bakhshi Golestani, Thibaut Meyer, Mathias Wien |
PCS | 3 |
| 2018 | Geometry-Corrected Deblocking Filter for 360° Video Coding using Cube RepresentationabstractIn 360° video, a complete scene is captured, as it can be seen from a single point in any direction. Since the captured 360 images are spherical, they cannot be converted to planar images without introducing geometric distortions. The nature of these distortion depends on the used projection format.This paper introduces an approach to reduce artifacts occurring when encoding 360° video which has been projected to the faces of a cube. In order to achieve this, the operation of the deblocking filter is modified such that the correct pixels with respect to the 3D geometry are used for filtering of edges.The method is evaluated on the set of sequences defined by the Joint Call for Proposals on Video Compression with Capability beyond HEVC. While the method has almost no impact on the objective coding performance, the visual quality is still clearly enhanced. Edges of the cube, previously visible as coding artifacts, are mostly removed with the proposed method. Johannes Sauer, Mathias Wien, Jens Schneider 0001, Max Bläser |
PCS | 2 |
| 2017 | Analysis/synthesis coding of dynamic textures based on motion distribution statisticsabstractThis paper presents improvements to a dynamic texture synthesis approach which is based on motion distribution statistics, able to produce high visual quality of synthesised dynamic textures. The aim is to recreate synthetically highly textured regions like water, leaves and smoke, instead of processing them with a conventional codec such as HEVC. The method involves two steps: analysis, where motion distribution statistics are computed, and synthesis, where the texture region is synthesized. Dense optical flow is utilized for estimating the random motion of dynamic textures. The performance of our dynamic texture analysis and synthesis approach is tested on cropped sequences, containing water, leaves and smoke. Simulation results show potential bitrate savings up to 50% on texture sequences at comparable visual quality. Olena Chubach, Patrick Garus, Mathias Wien, Jens-Rainer Ohm |
ICIP | 3 |
| 2017 | Synthesis of fine details in B picture for dynamic texturesabstractDynamic textures are characterized with irregular motions that are often challenging for motion compensation as applied in the state of the art video codecs. Due to rapid and randomly evolving nature of such a signal, it is accompanied with very high energy in the residual. As a result, B-pictures as used in HEVC layer are relative expensive to code. This leads to an overall increase in the bitrate. Further, increasing QPoffsetworsens the quality by forcing lower rate to these B-pictures, leading to strong blurring and blocking artefacts. In this paper, we exploit Steerable Pyramid (SP) for coding pictures with tid> 2 in a downsampled format. At the decoder side, details are synthesized for these low resolution pictures by adding back the high frequencies using motion compensation from the nearest key picture followed by an inverse SP transform. The paper synthesizes details for the dynamic textures that are expensive to code. Our investigation shows up to 31% saving in bitrate, while visual quality is kept acceptable. Uday Singh Thakur, Madhukar Bhat, Max Bläser, Mathias Wien, David Bull 0001, Jens-Rainer Ohm |
ICIP | 4 |
| 2017 | Point cloud estimation for 3D structure-based frame prediction in video codingabstract3D scene reconstruction from multi-view images has many practical applications, including games, virtual/augmented reality, and digital archives of cultural heritage. In this paper, we introduce a new application in video compression. The proposed idea is to have the decoder reconstruct a 3D scene model based on a subset of decoded frames and then reproject the 3D model to 2D for prediction or reconstruction of intermediate and/or future frames; this can also include a further motion compensation step in 2D. Structure from Motion (SfM) has been employed as a tool to estimate 3D point clouds and camera parameters. This approach has been integrated to generate additional reference pictures in an HEVC codec, and was tested so far on two 4K video sequences: A computer generated sequence with moving objects and a natural but stationary scene captured from a moving camera. Initial simulation results show around 0.8% bit-rate reduction compared to HEVC Test Model (HM16.7). It is asserted that the method offers headroom for further improvements by enhancing the reconstruction algorithms. Hossein Bakhshi Golestani, Jens Schneider 0001, Mathias Wien, Jens-Rainer Ohm |
ICME | 3 |
| 2017 | Improved motion compensation for 360° video projected to polytopesabstract360° video consists of images capturing the complete scene as seen from a single point in any direction. In contrast to conventional images, 360° images cannot be natively represented in a planar fashion. Consequently, video created from a sequence of 360° images suffers from geometric distortions. These are caused by mapping of the images from a sphere in 3D space to a planar image. This paper introduces an approach to improve motion compensation for 360° video by compensating the geometric distortions resulting from projection to the faces of a poly-tope. Each face is extended with content projected from faces connected to it. This provides better prediction candidates for motion compensation at face borders. Here, the applicability of the proposed method is demonstrated for 360° video projected to the faces of a cube. An evaluation of the method on two sets of sequences, static camera and non-static camera, verifies its effectiveness. It works especially well on non-static camera sequences, which have plenty of global motion. For the non-static camera sequences in the used test set rate savings of -2.14% were achieved. Johannes Sauer, Jens Schneider 0001, Mathias Wien |
ICME | 3 |
| 2017 | Geometry-adaptive motion partitioning using improved temporal predictionabstractCurrent state of the art video codecs such as HEVC are based on rectangular motion partitioning and compensation. To further enhance the compression performance, more flexible block partitioning strategies are needed. We present geometry based motion partitioning (GMP) in a post-HEVC framework with improved temporal prediction of GMP parameters. Our main contribution is a simple yet efficient temporal projection method for GMP parameters using available reference picture motion vectors, which allows the tracking of geometric partitioning lines along a motion trajectory. Average bit rate reductions by 1.35% are reported. Max Bläser, Cordula Heithausen, Mathias Wien |
VCIP | 3 |
| 2017 | Dictionary learning based high frequency inter-layer prediction for scalable HEVCabstractImage scale-up is a crucial task in resolution varying scalable video coding, as the coding costs for the enhancement layer depend heavily on the prediction signal generated by inter-layer prediction. In order to generate a suitable prediction signal the missing high frequencies in the base layer picture have to be reconstructed. For this purpose upscaling methods which go beyond the classical sampling theory are required. In this paper, an image scale-up method based on dictionary learning and sparse coding techniques for inter-layer prediction in scalable video coding is presented. Experimental results show that the proposed method outperforms state of the art scalable coding models in the case of 2x upscaling. In more detail, 2.35 % BD-rate savings against SHM 12.0 reference software are observed on average for an All Intra coding configuration. The maximum achieved rate savings were 6 15 % for the sequence PeopleOnStreet. Jens Schneider 0001, Johannes Sauer, Mathias Wien |
VCIP | 3 |
| 2017 | Standardization status of 360 degree video coding and deliveryabstractThe emergence of consumer level capturing and display devices for 360 degree video creates new and promising segments in entertainment, education, professional training, and other markets. In order to avoid market fragmentation and ensure interoperability of 360 degree video ecosystems, industry and academia cooperate in standardization efforts in this field. In the video coding domain, 360 degree video invalidates many established procedures, e.g., concerning evaluation of the visual quality, while the specific content characteristics offer potential for higher compression efficiency beyond the current standards. Likewise, 360 degree video puts stricter demands on the system level aspects of transmission but may also offer the potential to enhance existing transport schemes. The Joint Collaborative Team on Video Coding (JCT-VC) as well as the Joint Video Exploration Team (JVET) already started investigations into 360 degree video coding while numerous activities in the Systems subgroup of the Moving Picture Experts Group (MPEG) started to investigate application requirements and delivery aspects of 360 degree video. This paper reports on the current status of the outlined standardization efforts. Robert Skupin, Yago Sánchez de la Fuente, Ye-Kui Wang, Miska M. Hannuksela, Jill M. Boyce, Mathias Wien |
VCIP | 6 |
| 2016 | NMF-based informed source separationabstractInformed Source Separation (ISS) is a topic unifying the research fields of both source separation and source coding. Its main objective is to recover audio objects out of a mixture with a source separation step assisted by a set of compact parameters extracted with complete knowledge of the sources. ISS can be used for applications such as active listening and remixing of music (e.g. karaoke). In this paper, we propose a new ISS method which includes a semi-blind source separation (SBSS) step in the ISS decoder to decrease the amount of parameter bit rate. SBSS is conducted by factorizing the mixture in time-frequency domain by nonnegative matrix factorization (NMF). The transmitted parameters consist of a compact NMF initialization as well as residuals calculated in the NMF domain. We show in simulations that using SBSS in the decoder increases the separation quality and that our scheme improves the rate-distortion performance in comparison to a state-of-the art method. Christian Rohlfing, Julian Mathias Becker, Mathias Wien |
ICASSP | 3 |
| 2016 | Improved higher order motion compensation in HEVC with block-to-block translational shift compensationabstractConventionally, complex motion in video sequences is approximated by smaller block units in order to be representable by a translational motion model. This approximation results in a fine block partitioning and a high prediction error, both at cost of more data rate than potentially necessary. A worthwhile data reduction has been shown to be achievable by adding a higher order motion model to the most recent video coding standard, High Efficiency Video Coding (HEVC). The benefit of this additional option of inter-frame prediction is due to the more accurate motion compensation as well as the usage of larger block sizes. This paper deals with more efficient encoding of higher order motion parameters in this context. The geometrically accurate prediction of higher order motion parameters from a neighbored block needs to consider the dependency of the block-to-block parameter difference based on the spatial relation between two block centers. An algorithm is introduced for correcting the translational component and reducing the difference between the actual and the predicted motion when determining higher order parameters from neighbored blocks. Additionally, a further increase of the maximum block size up to 512×512 pixels is investigated. Cordula Heithausen, Max Bläser, Mathias Wien, Jens-Rainer Ohm |
ICIP | 3 |
| 2016 | Segmentation-based partitioning for motion compensated prediction in video codingabstractIn video coding standards such as Advanced Video Coding (AVC) and its successor High Efficiency Video Coding (HEVC), motion compensation is performed by partitioning each inter-predicted picture into square or rectangular regions. While HEVC introduces an efficient quad-tree based splitting of square-shaped coding blocks into prediction blocks by symmetric and asymmetric motion partitioning, the boundaries of natural moving objects can only be approximated by a fine block partitioning, resulting in a redundant representation of the motion and therefore a potential coding overhead. This paper studies the method of using a segmentation based partitioning of coding blocks into arbitrary shaped segments, where the segmentation is performed on coded reference pictures. Experimental results show that for cases where a reasonable segmentation can be obtained, bitrate reductions of around 2% can be achieved. Max Bläser, Cordula Heithausen, Mathias Wien |
PCS | 3 |
| 2016 | Motion-based analysis and synthesis of dynamic texturesabstractIn this paper, we propose an approach for synthesising motion of the identified dynamic textures, which is intended to replace conventional video coding of such content, i.e. encoding prediction residuals after block-wise translational motion estimation. Our algorithm first employs a perspective motion model to compensate for the global camera motion, such that only remaining dynamic motion is further analysed. Moreover, we use dense optical flow instead of block-based approach, which is important for representing the true motion of rigid objects and modeling the random motion of dynamic textures. The simulation results show that the proposed technique is able to synthesise visually plausible dynamic textures. Olena Chubach, Patrick Garus, Mathias Wien |
PCS | 3 |
| 2016 | Block adaptive selection of multiple core transforms for video codingabstractTransform coding tools in video coding have traditionally relied on the Discrete Cosine Transform Type II (DCT-II) to map residual signals to a new domain where quantization and entropy coding tools achieve a better coding efficiency than in the spatial domain. However, the DCT-II is not sufficient to model all different types of residual signals efficiently, especially in the intra-predicted blocks case. For this reason, the DST-VII was introduced in H.265/High Efficiency Video Coding (HEVC) in order to improve the compression performance of 4 × 4 intra-predicted blocks. In this paper we propose a multiple core transform approach, in which each transform is separable and generated by combining two one-dimensional transforms for the vertical and horizontal directions. The pair of 1-D transforms is selected from a set of three different types of Discrete Trigonometric Transforms and the Identity Transformation. Test results show that the proposed algorithm achieves bit rate reductions of 3% on average with respect to HEVC for intra-predicted residuals. Santiago De-Luxán-Hernández, Detlev Marpe, Heiko Schwarz, Klaus-Robert Müller, Mathias Wien, Jens-Rainer Ohm, Thomas Wiegand 0001 |
PCS | 5 |
| 2016 | Distance scaling of higher order motion parameters in an extension of HEVCabstractComplex motion in video sequences, such as rotation and scaling, conventionally approximated translationally, can more efficiently be represented by higher order motion models. An important contribution to the efficiency of the higher order motion compensation approach is the prediction and encoding of the additional motion parameters. The additional data rate caused by an increased amount of motion parameters has to be kept small for the higher order motion model to outperform the translational one within blocks of complex motion. Addressing an aspects of motion prediction yet to be adjusted to higher order motion, this paper proposes an advanced method of distance scaling of higher order motion parameters. Just as translational motion vector predictors require scaling whenever the block inheriting them has a different distance to its reference picture than the block predicted from, higher order motion parameter prediction can profit from such scaling as well. However, the distance scaling of higher order motion parameters cannot be applied to the elements of the motion model transformation matrix directly, but is performed on separate higher order motion components provided by a preceding transformation matrix decomposition. Accordingly, the objective of this work is to introduce and evaluate a procedure for higher order motion distance scaling. The resulting rate reduction of about 2% on average and up to over 6% attains a further improvement on the higher order motion compensation in an extension of HEVC. Cordula Heithausen, Max Bläser, Mathias Wien |
PCS | 3 |
| 2016 | Enhanced view synthesis prediction for coding of non-coplanar 3D video sequencesabstractIn many cases, view synthesis prediction (VSP) for 3D video coding suffers from inaccurate depth maps (DMs) caused by non-reliable point correspondences in low textured areas. In addition, occlusion handling is a major issue if prediction is performed from one reference view. In this paper, we address VSP for non-coplanar 3D video sequences modified by a linear DM enhancement model. In this model, points close to edges in the DM are treated as key points, on which depth estimation is assumed to work well. Other points are assumed to have unreliable depth information and the depth is modelled by the solution of the Laplace equation. As another preprocessing stage for DMs, the influence of guided image depthmap filtering is investigated. Moreover, we present an inpainting method for occluded areas, which is based on the solution of the Laplace equation and corresponding depth information. Combining these techniques and applying it to VSP for 3D video sequences, which are captured by an arc-arranged camera array, our approach outperforms state of the art models in coding efficiency. Experimental results show bit rate savings up to 4.3 % compared to HEVC-3D for coding two views and corresponding depth in an All Intra encoder configuration. Jens Schneider 0001, Johannes Sauer, Mathias Wien |
PCS | 3 |
| 2016 | Dynamic texture synthesis using linear phase shift interpolationabstractDynamic texture motions like flowing water, motion of leaves etc. have a complex random character. A sequence containing such a content is challenging to encode even when using state of the art High Efficiency Video Coding (HEVC) especially, if the available bandwidth is limited. It is observed that often when predicting dynamic textures, codec switches to intra prediction. At lower rates, dynamic texture content shows visually annoying blurring and blocking artifacts. For dynamic textures, both spatial and temporal details are perceptually of less importance. This property of the Human Visual System (HVS) can be exploited when coding dynamic texture content, as suggested in this paper. At the encoder, preprocessing is done by skipping even numbered B pictures. At the decoder side, skipped pictures are synthesized using linear phase shift interpolation of the complex wavelet coefficients, from the adjacent already decoded pictures. Subjective evaluation of proposed approach is done by using a pairwise comparison test between the proposed results and the conventional HEVC decoded bitstream at similar bitrates. The evaluation results show that viewers prefer the proposed result over conventional HEVC. Uday Singh Thakur, Karam Naser, Mathias Wien |
PCS | 3 |
| 2014 | Decoder complexity reduction for the scalable extension of HEVCabstractIn the current standardization process of the scalable extension to High Efficiency Video Coding (SHVC) a high level syntax multi-loop approach is close to completion. On the one hand this multi-loop approach offers a reasonable rate-distortion performance while only minimal modifications to the encoder and decoder in both layers are required. On the other hand this approach requires full reconstruction of all pictures of all layers at the decoder side which, in the case of quality scalability with two layers, doubles the decoder complexity. In this paper high layer modifications to the prediction structure similar to the scalable extension of H.264 - AVC are implemented in SHVC and studied. These modifications allow for an enhancement layer decoder implementation to skip a significant amount of motion compensation and deblocking operations in the base layer. It is shown that the decoder complexity can hereby be reduced up to 55% for the random access configuration and up to 64% for the low delay configuration compared to SHVC. An overall coding performance increase of 1.2% when decoding the enhancement layer is reported while when only decoding the base layer a drift can be observed between -0.16 dB for random access and -0.39 dB for low delay. Christian Feldmann, Fabian Jäger, Mathias Wien |
ICIP | 3 |
| 2014 | Simplified depth-based block partitioning and prediction merging in 3D video codingabstract3D video is an emerging technology that bundles depth information with texture videos to allow for view synthesis applications at the receiver. Depth discontinuities define object boundaries in both, depth maps and the collocated texture video. Therefore, depth segmentation can be utilized for a fine-grained motion field partitioning of the corresponding texture component. In this paper, depth information is used to increase coding efficiency for texture videos by deriving an arbitrarily shaped partitioning. By applying motion compensation to each partition independently and eventually merging the two prediction signals, highly accurate prediction signals can be produced that reduce the remaining texture residual signal significantly. Simulation results show bitrate savings of up to 2.8% for the dependent texture views and up to about 1.0% with respect to the total bitrate. Fabian Jäger, Mathias Wien |
VCIP | 2 |
| 2013 | Single-loop SNR scalability using binary residual refinement codingabstractIn scalable video coding, a stream contains multiple layers of different temporal resolution, spatial resolution, or quality. Each enhancement layer can utilize the information from the lower layers in order to improve the overall coding efficiency. In the draft scalable extension of HEVC the enhancement layer uses the lower layer reconstructed pixel values for prediction. While this dual-loop approach is efficient in terms of rate-distortion performance, it doubles the decoder complexity. In this paper a method is presented that allows for single-loop decoding using a new coefficient refinement scheme for the enhancement layer. The scheme conceptually allows for binary mapping between arbitrary quantizer step sizes in base and enhancement layer. Simulation results reveal a promising rate-distortion performance, especially for small quantizer deltas. Christian Feldmann, Fabian Jäger, Juliana Hsu, Mathias Wien |
ICIP | 4 |
| 2012 | Decoder-Side Motion Vector Derivation for Block-Based Video CodingabstractA decoder-side motion vector derivation algorithm for hybrid video coding is proposed. The algorithm is based on template matching and aims to reduce motion parameter bit-rate by re-estimating the applicable motion parameters at the decoder side. An average bit-rate savings of about 6%-8% is observed compared to the reference H.264/AVC. Decoder-side motion vector derivation was included in multiple proposals for the new High Efficiency Video Coding standard. This paper details and analyzes the algorithm and discusses its relation to other coding tools. Steffen Kamp, Mathias Wien |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Rate-Complexity-Distortion Optimization for Hybrid Video CodingabstractIn recent years, video applications on handheld devices became more and more popular. Due to limited computational capability and power supply in handheld devices, rate-complexity-distortion optimization (RCDO) algorithms at encoder side draw increasing attention. The target of RCDO is to obtain the best rate-distortion (R-D) performance under a constraint of complexity. Generally, there are three essential problems in RCDO. First, complexity needs to be properly mapped to a target in terms of coding parameters such that the control over complexity can be achieved. Second, the complexity budget should be efficiently distributed among frames or other coding units. Third, the allocated budget for each coding unit has to be effectively used to obtain good R-D performance. In this paper, these problems are well addressed. To obtain a large dynamic range in complexity control, medium-granularity control methods are presented. Then, a frame level complexity allocation algorithm is developed based on dependent rate-distortion function. Finally, an adaptive mode and reference searching method is proposed for motion compensation process. Comprehensive simulations verify the proposed algorithms. In the environment of the H.264/AVC reference software, an average gain of over 0.5 dB and 0.7 dB in BD-PSNR was achieved for nine sequences at low complexity when compared to two RCDO methods from literature. Moreover, experiments on x264 (a practical implementation of H.264/AVC) show that the proposed algorithms outperform predefined complexity levels by x264 in terms of both coding efficiency and computational scalability. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Optimized channel rate allocation for H.264/AVC scalable video multicast streaming over heterogeneous networksabstractWe present an algorithm to optimize the allocation of channel bitrate to different network abstraction layer (NAL) units of the H.264/AVC scalable video bitstreams for real-time multicast streaming over heterogeneous networks. We focus on the problem of achieving a high robustness of video streaming under varying channel conditions in terms of the reconstructed video qualities at different users. As an extension of our previous work for unicast streaming, the proposed algorithm can achieve an optimized allocation of channel bitrate for multicast streaming with any user distribution. Our simulations show that a good performance on the video qualities among the multicast users can be achieved for different user distributions. A gain in terms of the overall multicast PSNR can be achieved against the protection strategies targeting at users with medium channel qualities in our experiments. Bin Zhang 0018, Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
ICIP | 3 |
| 2010 | Decoder-side motion vector derivation for hybrid video inter codingabstractThe ongoing increase of computing performance facilitates a higher algorithmical complexity in video coding systems. The decoder may be able to estimate or derive prediction parameters based on the previously decoded signal. In this paper we present an extension to H.264/AVC where the explicit coding of motion parameters is adaptively replaced by a template matching algorithm that is performed identically at the encoder and decoder. The decision between explicit coding and derivation of motion parameters is done by the rate-distortion optimised mode decision and coded into the bitstream. Compared to previous work, the provided scheme has been extended to bidirectional prediction (B pictures). Simulation results show an improved coding efficiency over a wide range of test sequences, especially for higher spatial resolutions. Steffen Kamp, Mathias Wien |
ICME | 2 |
| 2010 | Rate-complexity-distortion evaluation for hybrid video codingabstractTo objectively evaluate the coding efficiency of video codecs, Bj⊘ntegaard Delta PSNR (BD-PSNR) was proposed. Based on the rate-distortion (R-D) curve fitting, BD-PSNR is able to provide a good evaluation of the R-D performance. However, BD-PSNR has a critical drawback: It doesn't take the coding complexity into account. Clearly for practical video applications, especially for those on handheld devices, coding complexity has to be considered when evaluating the overall coding performance. Therefore in this paper, a new coding efficiency measurement is developed by generalizing BD-PSNR from R-D curve fitting to rate-complexity-distortion (R-C-D) surface fitting. Simulations show that a comprehensive performance evaluation can easily be obtained with the proposed method. Moreover, the idea can be used for rate-distortion optimization for complexity-constrained video coding. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
ICME | 2 |
| 2010 | Medium-granularity computational complexity control for H.264/AVCabstractToday, video applications on handheld devices become more and more popular. Due to limited computational capability of handheld devices, complexity constrained video coding draws much attention. In this paper, a medium-granularity computational complexity control (MGCC) is proposed for H.264/AVC. First, a large dynamic range in complexity is achieved by taking 16×16 motion estimation in a single reference frame as the basic computational unit. Then a high coding efficiency is obtained by an adaptive computation allocation at MB level. Simulations show that coarse-granularity methods cannot work when the normalized complexity is below 15%. In contrast, the proposed MGCC performs well even when the complexity is reduced to 8.8%. Moreover, an average gain of 0.3 dB over coarse-granularity methods in BD-PSNR is obtained for 11 sequences when the complexity is around 20%. Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm |
PCS | 2 |
| 2009 | Synthesis-in-the-loop for video texture codingabstractIn this paper, we present an algorithm using dynamic texture synthesis for closed-loop video coding. Video textures, or so-called dynamic textures are video sequences with moving texture showing some stationarity properties over time, like water surfaces, whirlwind, clouds, crowds, or even parts of head-and-shoulder scenes. By learning the temporal statistics of such content, we can in principle synthesize the corresponding areas in future frames of the video. In this paper we show that this synthesized image content can also be used for prediction in a closed-loop hybrid video coding system, where the encoder decides about usage of such synthesized content and possible transmission of a residual error signal. This is done in an adaptive and rate-distortion optimized way, such that higher compression performance can be achieved for both high and low bitrates. We show that local adaptation of the algorithm can lead to better compression performance and reduce the computation complexity considerably. If PSNR is used as the quality criterion, savings of up to 15 % in bitrate have been observed experimentally. Aleksandar Stojanovic 0002, Mathias Wien, Thiow Keng Tan |
ICIP | 2 |
| 2009 | Fast decoder side motion vector derivation for inter frame video codingabstractDecoder-side motion vector derivation (DMVD) using template matching has been shown to improve coding efficiency of H.264/AVC based video coding. Instead of explicitly coding motion vectors into the bitstream, the decoder performs motion estimation in order to derive the motion vector used for motion compensated prediction. In previous works, DMVD was performed using a full template matching search in a limited search range. In this paper, a candidate based fast search algorithm replaces the full search. While the complexity reduction especially for the decoder is quite significant, the coding efficiency remains comparable. While for the full search algorithm BD-bitrate savings of 7.4% averaged over CIF and HD sequences according to the VCEG common conditions for IPPP high profile are observed, the proposed fast search achieves bitrate reductions of up to 7.5% on average. By further omitting sub-pel refinement, average savings observed for CIF and HD are still up to 7%. Steffen Kamp, Benjamin Bross, Mathias Wien |
PCS | 3 |
| 2008 | A context-aware architecture for QoS and transcoding management of multimedia streams in smart homesabstractCurrent trends in smart homes suggest that several multimedia services will soon converge towards common standards and platforms. However this rapid evolution gives rise to several issues related to the management of a large number of multimedia streams in the home communication infrastructure. An issue of particular relevance is how a context acquisition system can be used to support the management of such a large number of streams with respect to the Quality of Service (QoS), to their adaptation to the available bandwidth or to the capacity of the involved devices, and to their migration and adaptation driven by the users’ needs that are implicitly or explicitly notified to the system. Under this scenario this paper describes the experience of the INTERMEDIA project in the exploitation of context information to support QoS, migration, and adaptation of multimedia streams. Raffaele Bolla, Matteo Repetto, Saar De Zutter, Rik Van de Walle, Stefano Chessa, Francesco Furfari, Bernhard Reiterer, Hermann Hellwagner, Mark Asbach, Mathias Wien |
ETFA | 10 |
| 2008 | Decoder side motion vector derivation for inter frame video codingabstractIn this paper, a decoder side motion vector derivation scheme for inter frame video coding is proposed. Using a template matching algorithm, motion information is derived at the decoder instead of explicitly coding the information into the bitstream. Based on Lagrangian rate-distortion optimisation, the encoder locally signals whether motion derivation or forward motion coding is used. While our method exploits multiple reference pictures for improved prediction performance and bitrate reduction, only a small template matching search range is required. Derived motion information is reused to improve the performance of predictive motion vector coding in subsequent blocks. An efficient conditional signalling scheme for motion derivation in Skip blocks is employed. The motion vector derivation method has been implemented as an extension to H.264/AVC. Simulation results show that a bitrate reduction of up to 10.4% over H.264/AVC is achieved by the proposed scheme. Steffen Kamp, Michael Evertz, Mathias Wien |
ICIP | 3 |
| 2008 | Subjective performance evaluation of the SVC extension of H.264/AVCabstractThis contribution presents results of the MPEG verification test that was carried out for the new Scalable Video Coding (SVC) Amendment of H.264/AVC. The test consisted of a series of subjective comparisons of SVC and single layer H.264/AVC coding for different application scenarios including conversational applications, broadcasting over mobile channels, and HD broadcasting. The results show that a reasonable degree of spatial and quality scalability can be supported with a bit rate overhead of less than or about 10% and an indistinguishable visual quality compared to the state of the art single layer coding. This paper describes the coding conditions, the test procedure, and presents the results of the SVC verification test. Tobias Oelbaum, Heiko Schwarz, Mathias Wien, Thomas Wiegand 0001 |
ICIP | 3 |
| 2008 | Dynamic texture synthesis for H.264/AVC inter codingabstractDynamic textures are sequences of frames exhibiting certain stationarity properties over time; examples are sea-waves, whirlwind or moving crowds. We present an algorithm for dynamic texture extrapolation using only few training frames. A dynamic texture synthesizer using this algorithm has been integrated into a state-of-the-art H.264/AVC coding system, such that synthesized frames can be used by the encoder and decoder for inter prediction. For sequences not containing dynamic textures the same performance as with the conventional encoding system was achieved. In the case of sequences containing dynamic textures, intra coded macroblocks can be avoided by using the synthesized frame. Bitrate savings of up to 10% have been observed experimentally. Aleksandar Stojanovic 0002, Mathias Wien, Jens-Rainer Ohm |
ICIP | 2 |
| 2008 | System architecture for semantic annotation and adaptation in content sharing environments
Saar De Zutter, Mark Asbach, Sarah De Bruyne, Michael Unger 0001, Mathias Wien, Rik Van de Walle |
Vis. Comput. | 5 |
| 2007 | Extended Texture Prediction for H.264/AVC Intra CodingabstractEfficient intra prediction is an important aspect of video coding with high compression efficiency. H.264/AVC applies directional prediction from neighboring pixels on an adjustable block size for local decorrelation. In this paper, we present an extended prediction scheme in the context of H.264/AVC that comprises two additional prediction methods exploiting self-similar properties of the encoded texture. A new macroblock type is implemented, allowing for flexible selection of the available prediction methods for sub-partitions of the macroblock. Depending on the content of the encoded video sequence, substantial gains in rate-distortion performance are achieved. The results may indicate directions towards an enhanced intra coding scheme with improved rate-distortion performance. Jona Ballé, Mathias Wien |
ICIP (6) | 2 |
| 2007 | Real-Time System for Adaptive Video Streaming Based on SVCabstractThis paper presents the integration of scalable video coding (SVC) into a generic platform for multimedia adaptation. The platform provides a full MPEG-21 chain including server, adaptation nodes, and clients. An efficient adaptation framework using SVC and MPEG-21 digital item adaptation (DIA) is integrated and it is shown that SVC can seamlessly be adapted using DIA. For protection of packet losses in an error prone environment an unequal erasure protection scheme for SVC is provided. The platform includes a real-time SVC encoder capable of encoding CIF video with a QCIF base layer and fine grain scalable quality refinement at 12.5 fps on off-the-shelf high-end PCs. The reported quality degradation due to the optimization of the encoding algorithm is below 0.6 dB for the tested sequences. Mathias Wien, R. Cazoulat, Andreas Graffunder, Andreas Hutter, Peter Amon |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Performance Analysis of SVCabstractThis paper provides a performance analysis of the scalable video coding (SVC) extension of H.264/AVC. A short overview presenting the main functionalities of SVC is given and main issues in encoder control and bit stream extraction are outlined. Some aspects of rate-distortion optimization in the context of SVC are discussed and strategies for derivation of optimized configurations relative to the investigated scalability scenarios are presented. Based on these methods, rate-distortion results for several SVC configurations are presented and compared to rate-distortion optimized H.264/AVC single layer coding. For reference, a comparison to rate-distortion optimized MPEG-4 visual (advanced simple profile) coding results is provided. The results show that the performance gap between single layer coding and scalable video coding can be very small and that SVC clearly outperforms previous video coding technology such as MPEG-4 ASP. Mathias Wien, Heiko Schwarz, Tobias Oelbaum |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Erratum to "Performance Analysis of SVC"abstractIn the above titled article (ibid., vol. 17, no. 9 pp. 1194-1203, Sep 07), Fig. 6 was printed incorrectly. The correct figure is presented here. Mathias Wien, Heiko Schwarz, Tobias Oelbaum |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Variable block-size transforms for H.264/AVCabstractA concept for variable block-size transform coding is presented. It is called adaptive block-size transforms (ABT) and was proposed for coding of high resolution and interlaced video in the emerging video coding standard H.264/AVC. The basic idea of inter ABT is to align the block size used for transform coding of the prediction error to the block size used for motion compensation. Intra ABT employs variable block-size prediction and transforms for encoding. With ABT, the maximum feasible signal length is exploited for transform coding. Simulation results reveal a performance increase up to 12% overall rate savings and 0.9 dB in peak signal-to-noise ratio. Mathias Wien |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Look-ahead coding considering rate/distortion-optimizationabstractA new approach to combine R/D-optimization and lookahead coding is proposed. The dependent-R/D idea has been applied to blocks with no coded coefficients. This is an important case at low bit rates and whenever the motion model, which is used in virtually all modern video coders, fits accurately enough. The requirement of the motion model's applicability suggests the necessity to include spatial or temporal aliasing reducing filtering to support the proposed strategy. Markus Beermann, Mathias Wien, Jens-Rainer Ohm |
ICIP (1) | 2 |
| 2002 | Hybrid video coding using variable size block transforms
Mathias Wien, Achim Dahlhoff |
VCIP | 1 |
| 2001 | Adaptive block transforms for hybrid video coding
Mathias Wien, Claudia Mayer |
VCIP | 1 |
| 2000 | Hierarchical Wavelet Video Coding Using Warping PredictionabstractA hierarchical spatially scalable wavelet video coder is presented. The lowpass band of the wavelet decomposition is the base layer in the scalable scheme. It is encoded using forward motion compensation. Here, warping prediction is employed. Since warping prediction produces a smooth prediction error it is well suited for wavelet decomposition. The coder employs backward motion compensation in the enhancement layers, therefore no enhancement layer motion vectors have to be transmitted. The coarser levels of the wavelet decomposition of the current frame are used for motion estimation and motion compensation of the lowpass band of the next finer level. Experimental results show that the application of warping prediction is suited to improve the coding gain of the presented scheme significantly. Mathias Wien |
ICIP | 1 |
| 2000 | Segmentation and tracking of facial regions in color image sequences
Bernd Menser, Mathias Wien |
VCIP | 2 |
| 2000 | Adaptive scalable video coding using wavelet packets
Mathias Wien, Bernd Menser |
VCIP | 1 |