Jens-Rainer Ohm

dblp:33/2226 · DBLP profile ↗
← Back
66ranked-venue papers
17as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 16 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 GVC: efficient random access compression for gene sequence variations
abstract
BACKGROUND: In recent years, advances in high-throughput sequencing technologies have enabled the use of genomic information in many fields, such as precision medicine, oncology, and food quality control. The amount of genomic data being generated is growing rapidly and is expected to soon surpass the amount of video data. The majority of sequencing experiments, such as genome-wide association studies, have the goal of identifying variations in the gene sequence to better understand phenotypic variations. We present a novel approach for compressing gene sequence variations with random access capability: the Genomic Variant Codec (GVC). We use techniques such as binarization, joint row- and column-wise sorting of blocks of variations, as well as the image compression standard JBIG for efficient entropy coding. RESULTS: Our results show that GVC provides the best trade-off between compression and random access compared to the state of the art: it reduces the genotype information size from 758 GiB down to 890 MiB on the publicly available 1000 Genomes Project (phase 3) data, which is 21% less than the state of the art in random-access capable methods. CONCLUSIONS: By providing the best results in terms of combined random access and compression, GVC facilitates the efficient storage of large collections of gene sequence variations. In particular, the random access capability of GVC enables seamless remote data access and application integration. The software is open source and available at https://github.com/sXperfect/gvc/ .
Yeremia Gunawan Adhisantoso, Jan Voges, Christian Rohlfing, Viktor Tunev, Jens-Rainer Ohm, Jörn Ostermann
BMC Bioinform.5
2022 Deep Hashing with Hash Center Update for Efficient Image Retrieval
abstract
In this paper, we propose an approach for learning binary hash codes for image retrieval. Canonical Correlation Analysis (CCA) is used to design two loss functions for training a neural network such that the correlation between the two views to CCA is maximum. The main motivation for using CCA for feature space learning is that dimensionality reduction is possible and short binary codes could be generated. The first loss maximizes the correlation between the hash centers and the learned hash codes. The second loss maximizes the correlation between the class labels and the classification scores. In this paper, a novel weighted mean and thresholding-based hash center update scheme for adapting the hash centers is proposed. The training loss reaches the theoretical lower bound of the proposed loss functions, showing that the correlation coefficients are maximized during training and substantiating the formation of efficient feature space for retrieval. The measured mean average precision shows that the proposed approach outperforms other state-of-the-art methods.
Abin Jose, Daniel Filbert, Christian Rohlfing, Jens-Rainer Ohm
ICASSP4
2022 Optimized feature space learning for generating efficient binary codes for image retrieval
abstract
In this paper, a novel approach for learning a low-dimensional optimized feature space for image retrieval with minimum intra-class variance and maximum inter-class variance is proposed. The classical approach of Linear Discriminant Analysis (LDA) is generally used for generating an optimized low-dimensional feature space for single-labeled images. Since image retrieval involves images with multiple objects, LDA cannot be directly used for dimensionality reduction and feature space optimization. This problem is addressed by utilizing the relationship between LDA and Canonical Correlation Analysis (CCA) eigenvalues to generate an optimized feature space for both single-labeled and multi-labeled images. A CCA-based network architecture which correlates the low-dimensional feature vectors with the image label vectors is proposed. We design a novel loss function such that the correlation coefficients of CCA are maximized. Our experiments prove that we could train the neural network to reach the theoretical lower bound of loss corresponding to the negative sum of the correlation coefficients. Once the optimized feature space is generated, feature vectors are binarized with the Iterative Quantization (ITQ) approach. Finally, we propose an ensemble network to generate binary codes of desired bit length for retrieval. The measurement of mean average precision shows that the proposed approach outperforms the retrieval results of other single-labeled and multi-labeled image retrieval benchmarks at same bit numbers in a considerable number of cases.
Abin Jose, Erik Stefan Ottlik, Christian Rohlfing, Jens-Rainer Ohm
Signal Process. Image Commun.4
2021 3D Geometry-Based Global Motion Compensation For VVC
abstract
Since most 2D videos are initially captured in 3D environments, developing 3D motion models for video coding is beneficial. This paper introduces a method for extracting 3D geometry data from 2D videos and synthesizing 3D-based virtual Reference Pictures (RPs). These novel RPs are offered to the Versatile Video Coding (VVC) encoder for use in motion compensation. The proposed method generates 3D geometry data in the form of 3D meshes at the encoder and transmits it to the decoder as an overhead bitstream. However, this overhead could eat up the whole coding gain or even make it worse than the anchor VVC. This paper solves this problem by employing 3D mesh processing techniques, e.g., mesh de-noising, mesh decimation, and mesh compression. Simulation results show that the proposed method outperforms VVC up to ~ 3.8%.
Hossein Bakhshi Golestani, Johannes Sauer, Christian Rohlfing, Jens-Rainer Ohm
ICIP4
2021 Exploiting the 3D Structures Observed in 2D Video Sequences for Motion Compensation
abstract
Video coding standards widely use 2D translational and affine motion models to compensate for the motion between frames. Noting that most 2D videos are initially captured in a 3D space may lead us to a new direction: using 3D scene geometry for 2D motion compensation. We introduce a method for rendering virtual Reference Pictures (RPs) based on the 3D information extracted from 2D video sequences. The synthesized RPs are offered to the Versatile Video Coding (VVC) encoder to serve as motion compensation references. The 3D information (camera poses and 3D scene geometry) is estimated at the encoder and transmitted to the decoder as an overhead bitstream. Multiple techniques are discussed to minimize this overhead. Simulation results show up to 5% coding gain compared to our anchor (VTM 10.0), with an acceptable increase in decoding time.
Hossein Bakhshi Golestani, Jens-Rainer Ohm
PCS2
2021 Developments in International Video Coding Standardization After AVC, With an Overview of Versatile Video Coding (VVC)
abstract
In the last 17 years, since the finalization of the first version of the now-dominant H.264/Moving Picture Experts Group-4 (MPEG-4) Advanced Video Coding (AVC) standard in 2003, two major new generations of video coding standards have been developed. These include the standards known as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). HEVC was finalized in 2013, repeating the ten-year cycle time set by its predecessor and providing about 50% bit-rate reduction over AVC. The cycle was shortened by three years for the VVC project, which was finalized in July 2020, yet again achieving about a 50% bit-rate reduction over its predecessor (HEVC). This article summarizes these developments in video coding standardization after AVC. It especially focuses on providing an overview of the first version of VVC, including comparisons against HEVC. Besides further advances in hybrid video compression, as in previous development cycles, the broad versatility of the application domain that is highlighted in the title of VVC is explained. Included in VVC is the support for a wide range of applications beyond the typical standard- and high-definition camera-captured content codings, including features to support computer-generated/screen content, high dynamic range content, multilayer and multiview coding, and support for immersive media such as 360° video.
Benjamin Bross, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Ye-Kui Wang
Proc. IEEE3
2021 Guest Editorial Introduction to the Special Section on the VVC Standard
abstract
In this Special Section of the IEEE Transactions on Circuits and Systems for Video Technology, it is our honor to introduce the Versatile Video Coding (VVC) standard, the latest of the historic partnership collaborations between the International Telecommunication Union Telecommunication Standardization Sector (ITU-T), the International Organization for Standardization (ISO), and the International Electrotechnical Commission (IEC) in the field of video coding standardization.
Jill M. Boyce, Jianle Chen, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Overview of the Versatile Video Coding (VVC) Standard and its Applications
abstract
Versatile Video Coding (VVC) was finalized in July 2020 as the most recent international video coding standard. It was developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG) to serve an ever-growing need for improved video compression as well as to support a wider variety of today’s media content and emerging applications. This paper provides an overview of the novel technical features for new applications and the core compression technologies for achieving significant bit rate reductions in the neighborhood of 50% over its predecessor for equal video quality, the High Efficiency Video Coding (HEVC) standard, and 75% over the currently most-used format, the Advanced Video Coding (AVC) standard. It is explained how these new features in VVC provide greater versatility for applications. Highlighted applications include video with resolutions beyond standard- and high-definition, video with high dynamic range and wide color gamut, adaptive streaming with resolution changes, computer-generated and screen-captured video, ultralow-delay streaming, 360° immersive video, and multilayer coding e.g., for scalability. Furthermore, early implementations are presented to show that the new VVC standard is implementable and ready for real-world deployment.
Benjamin Bross, Ye-Kui Wang, Yan Ye 0003, Shan Liu 0001, Jianle Chen, Gary J. Sullivan, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.7
2020 Deep Subclass Linear Discriminant Analysis For Multimodal Feature Space Learning
abstract
In this work, we target a known problem in representation learning that is: beyond coarse classification, how can we better model fine-grained categorization? To address this problem, we introduce Deep Subclass Linear Discriminant Analysis (DeepSDA), which utilizes intra-class variation and inter-class similarity during training. We could achieve multimodal classification by maximizing the ratio of between-subclass scatter matrix and within-subclass scatter matrix. We maximize the eigenvalues along the discriminative eignevector directions. Hence the deep neural network is able to learn more discriminative representation space and thus has higher class separation in the linearly separable latent space. We show that DeepSDA leads to significant improvements on diverse fine-grained categorization and attribute learning benchmarks.
Abin Jose, Shen Yan 0008, Mi Zhang 0002, Jens-Rainer Ohm
ICIP4
2020 Guest Editorial Introduction to the Special Section on the Joint Call for Proposals on Video Compression With Capability Beyond HEVC
abstract
Standardization for digital video compression has shown significant evolution over the last three decades. Starting in 1988 with ITU-T H.261 as the first such standard that was practical for consumer use, ISO/IEC MPEG-1 and H.262/MPEG-2 video (the latter jointly standardized by ITU-T and ISO/IEC) were developed very soon thereafter, creating the first wave of broad usage of digital technology in consumer video, such as broadcast and disc player applications. Later, the H.264/MPEG-4 Advanced Video Coding (AVC) standard was again developed jointly by ITU-T and ISO/IEC experts, with its High Profile becoming dominant from 2004 in HD broadcast and storage, as well as network-based streaming services and private capture of video. With ever-increasing demands for higher quality and the advent of flat-panel displays, the H.265/MPEG-H High Efficiency Video Coding (HEVC) standard became the next generation of video compression standard; with its first version defined in 2013, HEVC has been especially instrumental for the recent deployment of Ultra High Definition (UHD, a.k.a. 4K video). As time has moved forward, video content has continued to become an increasing presence in our lives, with an ever-growing diversification of usage models and continuing demands for higher quality. For example, flat panels evolved towards support of high dynamic range (HDR) video with a wider color gamut, and new modalities for consuming video have appeared, such as head-mounted displays (HMDs).
Jill M. Boyce, Jianle Chen, Jens-Rainer Ohm, Gary J. Sullivan, Thomas Wiegand 0001, Yan Ye 0003
IEEE Trans. Circuits Syst. Video Technol.3
2020 General Video Coding Technology in Responses to the Joint Call for Proposals on Video Compression With Capability Beyond HEVC
abstract
After the development of the High-Efficiency Video Coding Standard (HEVC), ITU-T VCEG and ISO/IEC MPEG formed the Joint Video Exploration Team (JVET), which started exploring video coding technology with higher coding efficiency, including development of a Joint Exploration Model (JEM) algorithm and a corresponding software implementation. The technology explored in the last version of the JEM further increases the compression capabilities of the hybrid video coding approach by adding new tools, reaching up to 30% bit rate reduction compared to HEVC based on the Bjøntegaard delta bit rate (BD-rate) metric, and further improvement beyond that in terms of subjective visual quality. This provided enough evidence to issue a joint Call for Proposals (CfP) for a new standardization activity now known as Versatile Video Coding (VVC). All technology proposed in the responses to the CfP was based on the classic block-based hybrid video coding design, extending it by new elements of partitioning, intra- and inter-picture prediction, prediction signal filtering, transforms, quantization/scaling, entropy coding, and in-loop filtering. This article provides an overview of technology that was proposed in the responses to the CfP, with a focus on techniques that were not already explored in the JEM context.
Benjamin Bross, Kenneth Andersson, Max Bläser, Virginie Drugeon, Seung-Hwan Kim 0001, Jani Lainema, Shan Liu 0001, Jens-Rainer Ohm, Gary J. Sullivan, Ruoyang Yu
IEEE Trans. Circuits Syst. Video Technol.9
2020 The Joint Exploration Model (JEM) for Video Compression With Capability Beyond HEVC
abstract
This paper provides an overview of the coding algorithms of the Joint Exploration Model (JEM) for video compression with capability beyond HEVC, which was developed by the Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). The goal of the JEM development and experimentation was to provide evidence that sufficient coding efficiency improvement over the High Efficiency Video Coding (HEVC) standard can be achieved, which would justify the need for a new video coding standard with a compression capability significantly exceeding that of HEVC. The development of the JEM provided an ability to conduct studies toward that goal in a verifiable and collaborative manner and led to the launching of the project to develop the new Versatile Video Coding (VVC) standard. Objective metric gains exceeding 30% were measured for most of the tested high-resolution video content that represents current demanding new applications, and subjective testing using human observers showed even more benefit.
Jianle Chen, Marta Karczewicz, Yu-Wen Huang, Kiho Choi, Jens-Rainer Ohm, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.5
2019 Reference Picture Synthesis for Video Sequences Captured with a Monocular Moving Camera
abstract
Inter-frame prediction plays an important role in video coding by predicting the current frame from previously encoded pictures, called reference pictures. In the case of camera motion, the content of a current frame could be very different from its reference pictures and may consequently lead to a more difficult Motion Compensation (MC). The main idea of this paper is to process the input 2D video sequence in order to estimate the 3D geometry of the scene and then employ this data to virtually synthesize "geometrically compensated" reference pictures. Since these virtual reference pictures are more similar to the current frame, motion estimation and consequently coding efficiency could be enhanced. The proposed method is tested over six different video sequences and around 11% bitrate reduction is achieved compared to the High Efficiency Video Coding (HEVC) standard.
Hossein Bakhshi Golestani, Christian Rohlfing, Jens-Rainer Ohm
VCIP3
2018 Motion-Distribution based Dynamic Texture Synthesis for Video Coding
abstract
In this paper, a new approach for an improved video coding scheme is presented, which combines hybrid video coding and texture synthesis based on motion distribution statistics. Considering that the utilized texture synthesis approach provides high-quality visual results, while it is developed only for synthe- sizing the identified dynamic textures within a certain area, a new framework is presented, which allows to identify of areas for synthesis and combine conventional coding with synthesis. Also, a new representation and compression of synthesis parameters is presented, which is required due to the updated coding structure. When combining the proposed approach with conventional en- coder (HEVC reference software, HM 16.6), significantly reduced bit rates of the compressed video sequences with the texture replaced can be obtained. Moreover, because the synthesized textures have similar perceptual characteristics to those of the original textures, the video sequences with the texture replaced are also visually similar to the original sequences. Video results are provided online to allow assessing the visual quality of the tested content.
Olena Chubach, Patrick Garus, Mathias Wien, Jens-Rainer Ohm
PCS4
2017 Temporal Prediction of Motion Parameters with Interchangeable Motion Models
abstract
While the translational motion model remains predominant in motion compensation, better coding efficiency can be achieved by applying a higher order motion model in cases of non-translationally moving video content. But so far, temporal motion vector prediction has not been fully optimized in case of higher order motion compensation. In particular, if the motion model provided by the reference picture differs from the motion model of the current block, the temporal prediction is either not used at all or it is sub-optimal. Also, temporal predictors of a much smaller partition size might not provide the best suited motion parameter prediction to a larger current block. In order to overcome these issues, a new method of temporal motion prediction is introduced in this paper, allowing for flexible switching between motion models. The picture-wise dense translational motion vector field is calculated from both translational and higher order motion parameters with a configurable granularity of 4x4 pixel subpartitions down to pixelwise accuracy. Both a translational and a higher order motion parameter predictor are estimated from that vector field, thus giving the current block two alternatives to choose from. The proposed algorithm achieves rate reductions of about 2% on average compared to the previous higher order motion compensation system it is based on, now resulting in an average of around 20% efficiency gain for non-translational video content, compared to HEVC without such an option.
Cordula Heithausen, Maria Meyer, Max Bläser, Jens-Rainer Ohm
DCC4
2017 Analysis/synthesis coding of dynamic textures based on motion distribution statistics
abstract
This paper presents improvements to a dynamic texture synthesis approach which is based on motion distribution statistics, able to produce high visual quality of synthesised dynamic textures. The aim is to recreate synthetically highly textured regions like water, leaves and smoke, instead of processing them with a conventional codec such as HEVC. The method involves two steps: analysis, where motion distribution statistics are computed, and synthesis, where the texture region is synthesized. Dense optical flow is utilized for estimating the random motion of dynamic textures. The performance of our dynamic texture analysis and synthesis approach is tested on cropped sequences, containing water, leaves and smoke. Simulation results show potential bitrate savings up to 50% on texture sequences at comparable visual quality.
Olena Chubach, Patrick Garus, Mathias Wien, Jens-Rainer Ohm
ICIP4
2017 Synthesis of fine details in B picture for dynamic textures
abstract
Dynamic textures are characterized with irregular motions that are often challenging for motion compensation as applied in the state of the art video codecs. Due to rapid and randomly evolving nature of such a signal, it is accompanied with very high energy in the residual. As a result, B-pictures as used in HEVC layer are relative expensive to code. This leads to an overall increase in the bitrate. Further, increasing QPoffsetworsens the quality by forcing lower rate to these B-pictures, leading to strong blurring and blocking artefacts. In this paper, we exploit Steerable Pyramid (SP) for coding pictures with tid> 2 in a downsampled format. At the decoder side, details are synthesized for these low resolution pictures by adding back the high frequencies using motion compensation from the nearest key picture followed by an inverse SP transform. The paper synthesizes details for the dynamic textures that are expensive to code. Our investigation shows up to 31% saving in bitrate, while visual quality is kept acceptable.
Uday Singh Thakur, Madhukar Bhat, Max Bläser, Mathias Wien, David Bull 0001, Jens-Rainer Ohm
ICIP6
2017 Point cloud estimation for 3D structure-based frame prediction in video coding
abstract
3D scene reconstruction from multi-view images has many practical applications, including games, virtual/augmented reality, and digital archives of cultural heritage. In this paper, we introduce a new application in video compression. The proposed idea is to have the decoder reconstruct a 3D scene model based on a subset of decoded frames and then reproject the 3D model to 2D for prediction or reconstruction of intermediate and/or future frames; this can also include a further motion compensation step in 2D. Structure from Motion (SfM) has been employed as a tool to estimate 3D point clouds and camera parameters. This approach has been integrated to generate additional reference pictures in an HEVC codec, and was tested so far on two 4K video sequences: A computer generated sequence with moving objects and a natural but stationary scene captured from a moving camera. Initial simulation results show around 0.8% bit-rate reduction compared to HEVC Test Model (HM16.7). It is asserted that the method offers headroom for further improvements by enhancing the reconstruction algorithms.
Hossein Bakhshi Golestani, Jens Schneider 0001, Mathias Wien, Jens-Rainer Ohm
ICME4
2016 Coordinate selection for affine invariant feature description
abstract
In this paper, we present a method for affine invariant feature description. Based on the gradient distribution of an image region we calculate two basis vectors defining an affine invariant coordinate system, used to normalize the image region. The estimated basis vectors are non-orthogonal and allow for a precise representation of the gradient distribution. The proposed method can be combined with any feature detector and descriptor. Its performance is evaluated on globally affine transformed as well as on real world images and compared to state of the art methods for affine invariant feature description. The observed results outperform the results obtained by the SIFT feature detector and are comparable to the results obtained by ASIFT while having less computational complexity and being more flexibly applicable in case of local affine modifications.
Christopher Bulla, Jens-Rainer Ohm
ICIP2
2016 Improved higher order motion compensation in HEVC with block-to-block translational shift compensation
abstract
Conventionally, complex motion in video sequences is approximated by smaller block units in order to be representable by a translational motion model. This approximation results in a fine block partitioning and a high prediction error, both at cost of more data rate than potentially necessary. A worthwhile data reduction has been shown to be achievable by adding a higher order motion model to the most recent video coding standard, High Efficiency Video Coding (HEVC). The benefit of this additional option of inter-frame prediction is due to the more accurate motion compensation as well as the usage of larger block sizes. This paper deals with more efficient encoding of higher order motion parameters in this context. The geometrically accurate prediction of higher order motion parameters from a neighbored block needs to consider the dependency of the block-to-block parameter difference based on the spatial relation between two block centers. An algorithm is introduced for correcting the translational component and reducing the difference between the actual and the predicted motion when determining higher order parameters from neighbored blocks. Additionally, a further increase of the maximum block size up to 512×512 pixels is investigated.
Cordula Heithausen, Max Bläser, Mathias Wien, Jens-Rainer Ohm
ICIP4
2016 Invariance against local affine deformation for feature based object detection systems
abstract
In this paper, we present a method to increase invariance against affine deformations in feature based object detection systems. We use the gradient distribution of an image region to calculate two non-orthogonal basis vectors defining an affine invariant coordinate system, which is used to normalize the image region. The proposed method is an intermediate processing step subsequent to the feature detection and can be combined with any feature detector and descriptor combination. Its performance is evaluated on locally affine transformed as well as on real world images and compared to state of the art methods for affine invariant feature description. The observed results outperform the results obtained by SIFT, ASIFT or the Harris-Affine based feature normalization method, without introducing significant additional demands on the memory requirement or the computational complexity.
Christopher Bulla, Jens-Rainer Ohm
PCS2
2016 Block adaptive selection of multiple core transforms for video coding
abstract
Transform coding tools in video coding have traditionally relied on the Discrete Cosine Transform Type II (DCT-II) to map residual signals to a new domain where quantization and entropy coding tools achieve a better coding efficiency than in the spatial domain. However, the DCT-II is not sufficient to model all different types of residual signals efficiently, especially in the intra-predicted blocks case. For this reason, the DST-VII was introduced in H.265/High Efficiency Video Coding (HEVC) in order to improve the compression performance of 4 × 4 intra-predicted blocks. In this paper we propose a multiple core transform approach, in which each transform is separable and generated by combining two one-dimensional transforms for the vertical and horizontal directions. The pair of 1-D transforms is selected from a set of three different types of Discrete Trigonometric Transforms and the Identity Transformation. Test results show that the proposed algorithm achieves bit rate reductions of 3% on average with respect to HEVC for intra-predicted residuals.
Santiago De-Luxán-Hernández, Detlev Marpe, Heiko Schwarz, Klaus-Robert Müller, Mathias Wien, Jens-Rainer Ohm, Thomas Wiegand 0001
PCS6
2016 Introduction to the Special Issue on HEVC Extensions and Efficient HEVC Implementations
abstract
High Efficiency Video Coding (HEVC) is the most recent standard in the series of major video coding standards jointly produced by the ITU-T Video Coding Experts Group (VCEG) and the ISO/IEC Moving Picture Experts Group (MPEG). HEVC was first approved in 2013 in the ITU-T as Recommendation H.265 and in ISO/IEC as International Standard 23008-2, and it offers an unprecedented degree of compression capability for a very wide variety of applications. In the three years since its initial completion, it has been extended in several important ways to further broaden its scope. This special issue on HEVC features two sections: 1) HEVC extensions and 2) efficient HEVC implementations.
Jens-Rainer Ohm, Gary J. Sullivan, Vivienne Sze, Thomas Wiegand 0001, Madhukar Budagavi
IEEE Trans. Circuits Syst. Video Technol.1
2016 Video Quality Evaluation Methodology and Verification Testing of HEVC Compression Performance
abstract
The High Efficiency Video Coding (HEVC) standard (ITU-T H.265 and ISO/IEC 23008-2) has been developed with the main goal of providing significantly improved video compression compared with its predecessors. In order to evaluate this goal, verification tests were conducted by the Joint Collaborative Team on Video Coding of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29. This paper presents the subjective and objective results of a verification test in which the performance of the new standard is compared with its highly successful predecessor, the Advanced Video Coding (AVC) video compression standard (ITU-T H.264 and ISO/IEC 14496-10). The test used video sequences with resolutions ranging from 480p up to ultra-high definition, encoded at various quality levels using the HEVC Main profile and the AVC High profile. In order to provide a clear evaluation, this paper also discusses various aspects for the analysis of the test results. The tests showed that bit rate savings of 59% on average can be achieved by HEVC for the same perceived video quality, which is higher than a bit rate saving of 44% demonstrated with the PSNR objective quality metric. However, it has been shown that the bit rates required to achieve good quality of compressed content, as well as the bit rate savings relative to AVC, are highly dependent on the characteristics of the tested content.
Thiow Keng Tan, Rajitha Weerakkody, Marta Mrak, Naeem Ramzan, Vittorio Baroncini, Jens-Rainer Ohm, Gary J. Sullivan
IEEE Trans. Circuits Syst. Video Technol.6
2016 Overview of the Multiview and 3D Extensions of High Efficiency Video Coding
abstract
The High Efficiency Video Coding (HEVC) standard has recently been extended to support efficient representation of multiview video and depth-based 3D video formats. The multiview extension, MV-HEVC, allows efficient coding of multiple camera views and associated auxiliary pictures, and can be implemented by reusing single-layer decoders without changing the block-level processing modules since block-level syntax and decoding processes remain unchanged. Bit rate savings compared with HEVC simulcast are achieved by enabling the use of inter-view references in motion-compensated prediction. The more advanced 3D video extension, 3D-HEVC, targets a coded representation consisting of multiple views and associated depth maps, as required for generating additional intermediate views in advanced 3D displays. Additional bit rate reduction compared with MV-HEVC is achieved by specifying new block-level video coding tools, which explicitly exploit statistical dependencies between video texture and depth and specifically adapt to the properties of depth maps. The technical concepts and features of both extensions are presented in this paper.
Gerhard Tech, Ying Chen 0011, Karsten Müller 0001, Jens-Rainer Ohm, Anthony Vetro, Ye-Kui Wang
IEEE Trans. Circuits Syst. Video Technol.4
2012 Comparison of the Coding Efficiency of Video Coding Standards - Including High Efficiency Video Coding (HEVC)
abstract
The compression capability of several generations of video coding standards is compared by means of peak signal-to-noise ratio (PSNR) and subjective testing results. A unified approach is applied to the analysis of designs, including H.262/MPEG-2 Video, H.263, MPEG-4 Visual, H.264/MPEG-4 Advanced Video Coding (AVC), and High Efficiency Video Coding (HEVC). The results of subjective tests for WVGA and HD sequences indicate that HEVC encoders can achieve equivalent subjective reproduction quality as encoders that conform to H.264/MPEG-4 AVC when using approximately 50% less bit rate on average. The HEVC design is shown to be especially effective for low bit rates, high-resolution video content, and low-delay communication applications. The measured subjective improvement somewhat exceeds the improvement measured by the PSNR metric.
Jens-Rainer Ohm, Gary J. Sullivan, Heiko Schwarz, Thiow Keng Tan, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.1
2012 Overview of the High Efficiency Video Coding (HEVC) Standard
abstract
High Efficiency Video Coding (HEVC) is currently being prepared as the newest video coding standard of the ITU-T Video Coding Experts Group and the ISO/IEC Moving Picture Experts Group. The main goal of the HEVC standardization effort is to enable significantly improved compression performance relative to existing standards-in the range of 50% bit-rate reduction for equal perceptual video quality. This paper provides an overview of the technical features and characteristics of the HEVC standard.
Gary J. Sullivan, Jens-Rainer Ohm, Woojin Han 0001, Thomas Wiegand 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 Rate-Complexity-Distortion Optimization for Hybrid Video Coding
abstract
In recent years, video applications on handheld devices became more and more popular. Due to limited computational capability and power supply in handheld devices, rate-complexity-distortion optimization (RCDO) algorithms at encoder side draw increasing attention. The target of RCDO is to obtain the best rate-distortion (R-D) performance under a constraint of complexity. Generally, there are three essential problems in RCDO. First, complexity needs to be properly mapped to a target in terms of coding parameters such that the control over complexity can be achieved. Second, the complexity budget should be efficiently distributed among frames or other coding units. Third, the allocated budget for each coding unit has to be effectively used to obtain good R-D performance. In this paper, these problems are well addressed. To obtain a large dynamic range in complexity control, medium-granularity control methods are presented. Then, a frame level complexity allocation algorithm is developed based on dependent rate-distortion function. Finally, an adaptive mode and reference searching method is proposed for motion compensation process. Comprehensive simulations verify the proposed algorithms. In the environment of the H.264/AVC reference software, an average gain of over 0.5 dB and 0.7 dB in BD-PSNR was achieved for nine sequences at low complexity when compared to two RCDO methods from literature. Moreover, experiments on x264 (a practical implementation of H.264/AVC) show that the proposed algorithms outperform predefined complexity levels by x264 in terms of both coding efficiency and computational scalability.
Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.3
2010 Optimized channel rate allocation for H.264/AVC scalable video multicast streaming over heterogeneous networks
abstract
We present an algorithm to optimize the allocation of channel bitrate to different network abstraction layer (NAL) units of the H.264/AVC scalable video bitstreams for real-time multicast streaming over heterogeneous networks. We focus on the problem of achieving a high robustness of video streaming under varying channel conditions in terms of the reconstructed video qualities at different users. As an extension of our previous work for unicast streaming, the proposed algorithm can achieve an optimized allocation of channel bitrate for multicast streaming with any user distribution. Our simulations show that a good performance on the video qualities among the multicast users can be achieved for different user distributions. A gain in terms of the overall multicast PSNR can be achieved against the protection strategies targeting at users with medium channel qualities in our experiments.
Bin Zhang 0018, Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm
ICIP4
2010 Rate-complexity-distortion evaluation for hybrid video coding
abstract
To objectively evaluate the coding efficiency of video codecs, Bj⊘ntegaard Delta PSNR (BD-PSNR) was proposed. Based on the rate-distortion (R-D) curve fitting, BD-PSNR is able to provide a good evaluation of the R-D performance. However, BD-PSNR has a critical drawback: It doesn't take the coding complexity into account. Clearly for practical video applications, especially for those on handheld devices, coding complexity has to be considered when evaluating the overall coding performance. Therefore in this paper, a new coding efficiency measurement is developed by generalizing BD-PSNR from R-D curve fitting to rate-complexity-distortion (R-C-D) surface fitting. Simulations show that a comprehensive performance evaluation can easily be obtained with the proposed method. Moreover, the idea can be used for rate-distortion optimization for complexity-constrained video coding.
Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm
ICME3
2010 Medium-granularity computational complexity control for H.264/AVC
abstract
Today, video applications on handheld devices become more and more popular. Due to limited computational capability of handheld devices, complexity constrained video coding draws much attention. In this paper, a medium-granularity computational complexity control (MGCC) is proposed for H.264/AVC. First, a large dynamic range in complexity is achieved by taking 16×16 motion estimation in a single reference frame as the basic computational unit. Then a high coding efficiency is obtained by an adaptive computation allocation at MB level. Simulations show that coarse-granularity methods cannot work when the normalized complexity is below 15%. In contrast, the proposed MGCC performs well even when the complexity is reduced to 8.8%. Moreover, an average gain of 0.3 dB over coarse-granularity methods in BD-PSNR is obtained for 11 sequences when the complexity is around 20%.
Xiang Li 0003, Mathias Wien, Jens-Rainer Ohm
PCS3
2010 Special Section on the Joint Call for Proposals on High Efficiency Video Coding (HEVC) Standardization
abstract
The five papers in this special section were among those submitted in response to the joint call for proposals on high efficiency video coding (HEVC) standardization. Although at this point of development it is still unclear which specific elements the final HEVC standard will contain, the selection of the papers was made such that together they would cover most of the promising tools and technologies that seem likely to be included in the standard.
Thomas Wiegand 0001, Jens-Rainer Ohm, Gary J. Sullivan, Woojin Han 0001, Rajan L. Joshi, Thiow Keng Tan, Kemal Ugur
IEEE Trans. Circuits Syst. Video Technol.2
2008 Dynamic texture synthesis for H.264/AVC inter coding
abstract
Dynamic textures are sequences of frames exhibiting certain stationarity properties over time; examples are sea-waves, whirlwind or moving crowds. We present an algorithm for dynamic texture extrapolation using only few training frames. A dynamic texture synthesizer using this algorithm has been integrated into a state-of-the-art H.264/AVC coding system, such that synthesized frames can be used by the encoder and decoder for inter prediction. For sequences not containing dynamic textures the same performance as with the conventional encoding system was achieved. In the case of sequences containing dynamic textures, intra coded macroblocks can be avoided by using the synthesized frame. Bitrate savings of up to 10% have been observed experimentally.
Aleksandar Stojanovic 0002, Mathias Wien, Jens-Rainer Ohm
ICIP3
2007 Backward Drift Estimation with Application to Quality Layer Assignment in H.264/AVC Based Scalable Video Coding
abstract
We present an approach for accurate estimation of the reconstruction distortion in SNR scalable video coding with drift. Based on a linear model of predictive video coding, we derive an algorithm to quantify spatio-temporal drift properties subject to prediction structure and motion information. This allows for low-complex estimation of the reconstruction distortion on a per-block basis. The accuracy of the distortion estimation is experimentally verified. We then utilize the method for quality layer assignment within the framework of H.264/AVC scalable video coding "SVC", which is currently under standardization. The quality layers allow for bit stream truncation in a rate-distortion optimized sense. Compared to the quality layer assignment as implemented in the SVC test model, use of backward drift estimation allows for achieving equivalent coding efficiency with reduced complexity.
Thomas Rusert, Jens-Rainer Ohm
ICASSP (1)2
2007 Non-linear up-sampling for image coding in a spatial pyramid
abstract
A locally adaptive up-sampling method that improves the efficiency of a spatially scalable representation of images in a spatial pyramid is presented. While linear methods use a globally optimized up-sampling filter design, the method presented locally switches between enhancement of significant structures and smoothing of flat regions that are dominated by noise. It is based on a locally adaptive Wiener filter expression that can be implemented by the bilateral filter. The performance of the method is assessed in a scenario resembling its possible use in MPEG's and ITU-T's joint current activity on scalable video coding (SVC).
Markus Beermann, Jens-Rainer Ohm
VCIP2
2007 Introduction to the Special Issue on Scalable Video Coding-Standardization and Beyond
abstract
The thirteen papers in this special issue are devoted to the standardization and development of scalable video coding techniques and applications.
Thomas Wiegand 0001, Gary J. Sullivan, Jens-Rainer Ohm, Ajay Luthra
IEEE Trans. Circuits Syst. Video Technol.3
2006 Macroblock Based Bit Allocation for SNR Scalable Video Coding with Hierarchical B Pictures
abstract
We investigate SNR scalable video coding based on motion compensated temporal prediction with hierarchical B pictures. This structure is a fundamental part of the scalable extension of H.264/AVC (SVC), which is currently under standardization. Due to SNR scalability, reconstruction of inter blocks may cause error accumulation within the B picture hierarchy, which is called the drift effect. In this paper we investigate the error propagation considering coding control, motion information, and target quality. We first consider open loop coding control and generalize to the case of one or multiple loops. We present a practical algorithm to quantify locally varying drift properties, which is utilized for performing rate-distortion optimized macroblock based bit allocation. We compare the approach to picture based bit allocation within the SVC test model and observe coding gains up to 0.4 dB.
Thomas Rusert, Jens-Rainer Ohm
ICIP2
2005 Thresholded Weighted Median Filters for Ringing Reduction in Processed Images
abstract
In this paper, we present how thresholded weighted median filters (WMFs), can significantly improve visual as well as objective quality of images affected by ringing. Ringing is identified by its structural properties and WMFs are chosen according to these structures. Since WMFs are concerned with the relative values of their input but do not consider its absolute values, we explicitly enforce a limited filter output. In this paper we analyze the formation of ringing and match thresholded WMFs for example applications. For image resampling, we are able to obtain a significantly decreased amount of ringing. As a post-filter of a lossy wavelet-coder, typical PSNR gains around 0.4 dB are obtained while visually, ringing is reduced to a very pleasing extent
Markus Beermann, Adeel Jalil, Jens-Rainer Ohm
MMSP3
2005 Advances in Scalable Video Coding
abstract
Scalable video coding is attractive due to the capability of reconstructing lower resolution or lower quality signals from partial bit streams. This allows for simple solutions in adaptation to network and terminal capabilities. Different modalities of scalability are specified by video coding standards like MPEG-2 and MPEG-4. This paper gives a short overview over these techniques and analyzes in more detail the encoder/decoder drift problem, which is the major reason why scalable coding has been significantly less efficient than single-layer coding in most of these implementations. Only recently, new scalable video coding technology has evolved, which seems to close the gap of compression performance compared to state of the art single-layer video coding. New methods of efficient enhancement layer prediction were developed to improve traditional (motion-compensated hybrid) scalable coders, providing more flexible compromises on the drift problem. As a new technology trend, motion-compensated spatiotemporal wavelet coding has matured which entirely discards the drift and allows most flexible combinations of spatial, temporal, and signal-to-noise ratio (SNR) scalability with fine granularity over a broad range of data rates.
Jens-Rainer Ohm
Proc. IEEE1
2004 Interframe wavelet coding - motion picture representation for universal scalability
Jens-Rainer Ohm, Mihaela van der Schaar, John W. Woods
Signal Process. Image Commun.1
2004 Special issue on subband/wavelet interframe video coding
John W. Woods, Jens-Rainer Ohm
Signal Process. Image Commun.2
2004 Invertible temporal subband/wavelet filter banks with half-pixel-accurate motion compensation
abstract
Three-dimensional (3-D) subband/wavelet coding with motion compensation has been demonstrated to be an efficient technique for video coding applications in some recent research works. When motion compensation is performed with half-pixel accuracy, images need to be interpolated in both temporal subband analysis and synthesis stages. The resulting subband filter banks developed in these former algorithms were not invertible due to image interpolation. In this paper, an invertible temporal analysis/synthesis system with half-pixel-accurate motion compensation is presented. We look at temporal decomposition of image sequences as a kind of down-conversion of the sampling lattices. The earlier motion-compensated (MC) interlaced/progressive scan conversion scheme is extended for temporal subband analysis/synthesis. The proposed subband/wavelet filter banks allow perfect reconstruction of the decomposed video signal while retaining high energy compaction of subband transforms. The invertible filter banks are then utilized in our 3-D subband video coder. This video coding system does not contain the temporal DPCM loop employed in the conventional hybrid coder and the earlier MC 3-D subband coders. The experimental results show a significant PSNR improvement by the proposed method. The generalization of our algorithm for MC temporal filtering at arbitrary subpixel accuracy is also discussed.
Shih-Ta Hsiang, John W. Woods, Jens-Rainer Ohm
IEEE Trans. Image Process.3
2003 Shape retrieval with robustness against partial occlusion
abstract
Object shape features are powerful when used in similarity search-&-retrieval and object recognition because object shape is usually strongly linked to object functionality and identity. Many applications, including those concerned with visual object retrieval or indexing, are likely to use shape features. Those systems have to cope with scaling, rotation, deformation and partial occlusion of the objects to be described. The ISO standard MPEG-7 contains different shape descriptors; we focus especially on the region-shape descriptor. Since we have found the region-shape descriptor not to be very robust against partial occlusion, we propose a slightly changed feature extraction method, which is based on central-moments. Further, we compare our method with the original region-shape implementation and show that, by applying the proposed changes, the robustness of the region-shape descriptor against partial occlusions can be significantly increased.
Michael Hoeynck, Jens-Rainer Ohm
ICASSP (3)2
2003 Application of MPEG-7 descriptors for content-based indexing of sports videos
Michael Hoeynck, Thorsten Auweiler, Jens-Rainer Ohm
VCIP3
2003 Transition filtering and optimized quantization in interframe wavelet video coding
Thomas Rusert, Konstantin Hanke, Jens-Rainer Ohm
VCIP3
2002 Look-ahead coding considering rate/distortion-optimization
abstract
A new approach to combine R/D-optimization and lookahead coding is proposed. The dependent-R/D idea has been applied to blocks with no coded coefficients. This is an important case at low bit rates and whenever the motion model, which is used in virtually all modern video coders, fits accurately enough. The requirement of the motion model's applicability suggests the necessity to include spatial or temporal aliasing reducing filtering to support the proposed strategy.
Markus Beermann, Mathias Wien, Jens-Rainer Ohm
ICIP (1)3
2002 Bit plane quantization for scalable video coding
Claudia Mayer, Holger Crysandt, Jens-Rainer Ohm
VCIP3
2001 The MPEG-7 Visual Description Framework - Concepts, Accuracy, and Applications
Jens-Rainer Ohm
CAIP1
2001 Color and texture descriptors
abstract
This paper presents an overview of color and texture descriptors that have been approved for the Final Committee Draft of the MPEG-7 standard. The color and texture descriptors that are described in this paper have undergone extensive evaluation and development during the past two years. Evaluation criteria include effectiveness of the descriptors in similarity retrieval, as well as extraction, storage, and representation complexities. The color descriptors in the standard include a histogram descriptor that is coded using the Haar transform, a color structure histogram, a dominant color descriptor, and a color layout descriptor. The three texture descriptors include one that characterizes homogeneous texture regions and another that represents the local edge distribution. A compact descriptor that facilitates texture browsing is also defined. Each of the descriptors is explained in detail by their semantics, extraction and usage. The effectiveness is documented by experimental results.
B. S. Manjunath, Jens-Rainer Ohm, Vinod V. Vasudevan, Akio Yamada
IEEE Trans. Circuits Syst. Video Technol.2
2000 Low-Complexity Global Motion Estimation from P-Frame Motion Vectors for MPEG-7 Applications
abstract
We present an algorithm for low-complexity global motion estimation, that works with block-coded video (e.g. MPEG-2). A superimposed global motion model is fitted to decoded P-frame motion vector fields, providing real-time performance. Therefore a robust M-estimator is applied since a lot of outliers have to be expected in the motion vector fields. The algorithm is compared to a state-of-the-art global motion estimator. The main characteristics of the global motion are captured accurately in most cases, although estimation may fail. We also report the integration of the estimator into a complete MPEG-7 system for search-and-retrieval of video based on motion characteristics.
Aljoscha Smolic, Michael Hoeynck, Jens-Rainer Ohm
ICIP3
2000 Robust Global Motion Estimation Using a Simplified M-Estimator Approach
abstract
Global motion estimation is an important task in a variety of video processing applications, such as coding, segmentation, classification/indexing or mosaicing. Due to the possible presence of differently moving foreground objects and other sources of distortions, robust methods such as M-estimators have to be applied. We present a simplified implementation of a robust M-estimator for global motion estimation that does not increase the computational complexity significantly compared to a non-robust estimator, while providing excellent results in terms of estimation accuracy. Additionally, unstructured image regions are detected and rejected for the estimation. This avoids aperture problems, that can have an bad impact especially on robust estimators that rely on a certain error measure.
Aljoscha Smolic, Jens-Rainer Ohm
ICIP2
2000 Image-based rendering and 3D modeling: A complete framework
Ebroul Izquierdo, Jens-Rainer Ohm
Signal Process. Image Commun.2
2000 A set of visual feature descriptors and their combination in a low-level description scheme
Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier
Signal Process. Image Commun.1
1999 A multi-feature description scheme for image and video database retrieval
abstract
This paper reports about a description scheme for visual information content, which has been developed in the context of the forthcoming MPEG-7 standard. The system supports similarity-based retrieval of visual (image and video) data along feature axes like color, texture, shape/geometry and motion. The descriptors for these features have been developed in a way such that invariance against common transformations of visual material, e.g. filtering, contrast/color manipulation, resizing etc. is achieved, and that they are fitted to human perception properties. Furthermore, descriptors have been designed that allow a fast, hierarchical search procedure. A search engine has been developed on the basis of this description scheme, which allows similarity-based retrieval from an image or video database. The results show that efficient search and retrieval in visual database systems is possible based on a normative feature description such as MPEG-7.
Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier
MMSP1
1999 Incomplete 3-D multiview representation of video objects
abstract
This paper introduces a new form of representation for three-dimensional (3-D) video objects. We have developed a technique to extract disparity and texture data from video objects that are captured simultaneously with multiple-camera configurations. For this purpose, we derive an "area of interest" (AOI) for each of the camera views, which represents an area on the video object's surface that is best visible from this specific camera viewpoint. By combining all AOIs, we obtain the video object plane as an unwrapped surface of a 3-D object, containing all texture data visible from any of the cameras. This texture surface can be encoded like any 2-D video object plane, while the 3-D information is contained in the associated disparity map. It is then possible to reconstruct different viewpoints from the texture surface by simple disparity-based projection. The merits of the technique are efficient multiview encoding of single video objects and support for viewpoint adaptation functionality, which is desirable in mixing natural and synthetic images. We have performed experiments with the MPEG-4 video verification model, where the disparity map is encoded by use of the tools provided for grayscale alpha data encoding. Due to its simplicity, the technique is suitable for applications that require real-time viewpoint adaptation toward video objects.
Jens-Rainer Ohm, Karsten Müller 0001
IEEE Trans. Circuits Syst. Video Technol.1
1999 Long-term global motion estimation and its application for sprite coding, content description, and segmentation
abstract
We present a new technique for long-term global motion estimation of image objects. The estimated motion parameters describe the continuous and time-consistent motion over the whole sequence relatively to a fixed reference coordinate system. The proposed method is suitable for the estimation of affine motion parameters as well as for higher order motion models like the parabolic model-combining the advantages of feature matching and optical flow techniques. A hierarchical strategy is applied for the estimation, first translation, affine motion, and finally higher order motion parameters, which is robust and computationally efficient. A closed-loop prediction scheme is applied to avoid the problem of error accumulation in long-term motion estimation. The presented results indicate that the proposed technique is a very accurate and robust approach for long-term global motion estimation, which can be used for applications such as MPEG-4 sprite coding or MPEG-7 motion description. We also show that the efficiency of global motion estimation can be significantly increased if a higher order motion model is applied, and we present a new sprite coding scheme for on-line applications. We further demonstrate that the proposed estimator serves as a powerful tool for segmentation of video sequences.
Aljoscha Smolic, Thomas Sikora, Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.3
1998 A realtime hardware system for stereoscopic videoconferencing with viewpoint adaptation
Jens-Rainer Ohm, Karsten Grüneberg, Emile A. Hendriks, Ebroul Izquierdo, Dimitris Kalivas, Michael Karl 0001, Dionysis Papadimatos, André Redert
Signal Process. Image Commun.1
1997 Motion-Compensating Real-Time Format Converter for Video on Multimedia Displays
abstract
This paper introduces a high-quality, low-cost video converter for the conversion of interlaced TV signals into progressive display formats of the same or higher frame repetition rate. This conversion is performed by motion compensated filtering which is preceded by motion estimation. The applied motion estimation algorithm operates on blocks sized 4/spl times/4 pixels. For each of these blocks, a candidate vector is chosen out of temporal and spatial predecessors using a displayed field differences error criterion. The optimal candidate is updated by a series of pixel recursive steps. The filter algorithm employs a motion compensating median filter whose shape depends on the motion vector. A fallback mode is implemented to deal with areas for which no accurate motion vectors could be derived. Single-chip integration of the whole format conversion system is feasible.
Martin Hahn 0001, Jens-Rainer Ohm, Maati Talmi
ICIP (1)3
1997 Feature-based cluster segmentation of image sequences
abstract
One of the crucial points in object segmentation within image sequences is the interdependence of different features that classify some area as an object. This paper introduces a concept of cluster segmentation, which acquires different features on a pixel basis. Weighting of these features based on predefined rules is applied, in order to judge the evidence of each particular feature for the final classification. To determine the various clusters, we use a procedure which is similar to vector quantization. This allows the tracking of classification results over time, because cluster labels change only gradually from frame to frame. Furthermore, a technique for local feature analysis is applied for segment merging after global classification. The most common features used for object separation in image sequences are color and motion. The results indicate that reliable segmentation and tracking of objects can be accomplished, using this low-complexity technique.
Jens-Rainer Ohm, Phuong Ma
ICIP (3)1
1997 Variable-Raster Multiresolution VideoProcessing with Motion Compensation Techniques
abstract
The standard format for video acquisition, transmission and presentation uses an interlaced scanning technique. On the other hand for digital video services, the progressive format is more desirable in many situations, and different resolutions of the signal are required. The interlaced scan is based on a tight concatenation of vertical and temporal sampling effects. This paper introduces the technique of motion-compensated vertical filters within pairs of fields to solve the problem of sampling rate conversion between interlaced and progressive raster. It is shown that the motion-compensated vertical filter is a polyphase realization of a factor-2 up- or downsampling filterbank, with polyphase shift controlled by motion parameters. Possible applications are down- and up-conversion between progressive and interlaced formats, and multiresolution representation of video data within a spatio-temporal subband or wavelet framework.
Jens-Rainer Ohm, K. Rummler
ICIP (1)1
1997 An object-based system for stereoscopic viewpoint synthesis
abstract
This paper describes algorithms that were developed for a real-time stereoscopic videoconferencing systems with viewpoint adaptation. The goal is a real telepresence illusion, which is achieved by synthesis of intermediate views from a stereoscopic camera shot with a rather large baseline. The actual viewpoint will be adapted according to the head position of the viewer, such that the impression of motion parallax is produced. The object-based system first identifies foreground and background regions and applies disparity estimation to the foreground object. A hierarchical block matching algorithm is employed for this purpose which takes into account the position of high-activity feature points and the object/background border positions. Using the disparity estimator's output, it is possible to generate arbitrary intermediate views by projections from the left- and right-view images. For this purpose, we have also developed an object-based interpolation algorithm, taking into account a very simple convex-surface model of a person's face and body. Though the algorithms had to be held rather simple under the constraint of hardware feasibility, we obtain a good quality of the intermediate-view images. Finally, we describe the hardware concept for the disparity estimator, which is the most complicated part of the algorithm.
Jens-Rainer Ohm, Ebroul Izquierdo
IEEE Trans. Circuits Syst. Video Technol.1
1994 Motion-Compensated 3-D Subband Coding with Multiresolution Representation of Motion Parameters
abstract
This paper concentrates on the aspects of motion estimation and representation in the framework of motion-compensated 3-D subband coding. The motion vector field (MVF) is estimated hierarchically, based on a description obtained from a decimated field of support points. Displacements in between these points are interpolated. The motion parameters are encoded by use of a 3-D Laplacian pyramid structure, which can exploit both spatial and temporal redundancies in the motion vector field. The scheme can cope with non-translational and spatially/temporally variant motion, and supports a multiresolution description of the MVF. The representation of both information components-image and motion-has a hierarchical structure allowing efficient error protection as well as compatibility towards finer resolution and higher quality. More than that, with the improved motion estimation and representation procedures, noticeable enhancements were obtained in the encoding of full-motion video at low rates.>
Jens-Rainer Ohm
ICIP (3)1
1994 Three-dimensional subband coding with motion compensation
abstract
Three-dimensional (3-D) frequency coding is an alternative approach to hybrid coding concepts used in today's standards. The first part of this paper presents a study on concepts for temporal-axis frequency decomposition along the motion trajectory in video sequences. It is shown that, if a two-band split is used, it is possible to overcome the problem of spatial inhomogeneity in the motion vector field (MVF), which occurs at the positions of uncovered and covered areas. In these cases, original pixel values from one frame are placed into the lowpass-band signal, while displaced-frame-difference values are embedded into the highpass band. This technique is applicable with arbitrary MVF's; examples with block-matching and interpolative motion compensation are given. Derivations are first performed for the example of two-tap quadrature mirror filters (QMF's), and then generalized to any linear-phase QMF's. With two-band analysis and synthesis stages arranged as cascade structures, higher resolution frequency decompositions are realizable. In the second part of the paper, encoding of the temporal-axis subband signals is discussed. A parallel filterbank scheme was used for spatial subband decomposition, and adaptive lattice vector quantization was employed to approach the entropy rate of the 3-D subband samples. Coding results suggest that high-motion video sequences can be encoded at significantly lower rates than those achievable with conventional hybrid coders. Main advantages are the high energy compaction capability and the nonrecursive decoder structure. In the conclusion, the scheme is interpreted more generally, viewed as a motion-compensated short-time spectral analysis of video sequences, which can adapt to the quickness of changes. Although a 3-D multiresolution representation of the picture information is produced, a true multiresolution representation of motion information, based on spatio-temporal decimation and interpolation of the MVF, is regarded as the still-missing part.
Jens-Rainer Ohm
IEEE Trans. Image Process.1
1993 Advanced packet-video coding based on layered VQ and SBC techniques
abstract
The performances of two subband-coding-vector-quantization (SBC-VQ) schemes (a quadrature-mirror-filter-based SBC with trained-codebook VQ and a parallel-filterbank SBC with lattice VQ) are compared to that of a DCT-SQ (discrete-cosine-transform-scalar-quantization) scheme as it is used in present image-coding standards. Two-layer versions of these spatial coders are evaluated in a hybrid combination with motion-compensation prediction for interframe data compression. As the SBC scheme with lattice VQ is found to perform the best, this coding scheme is further investigated in combination with more sophisticated interframe-data-compression schemes. The motion-compensated 3-D SBC coder with lattice VQ was found to outperform the techniques used in current standard coders by several dBs in the compression of interlaced CCIR 601 sequences. The performance of this coder is extremely robust in the presence of asynchronous transfer mode (ATM) cell losses due to the nonrecursive decoder structure.>
Jens-Rainer Ohm
IEEE Trans. Circuits Syst. Video Technol.1
1992 Temporal domain sub-band video coding with motion compensation
abstract
Temporal domain subband coding with motion compensation (MC-SBC) is a new technique of interframe data compression in video coding applications. Perfect reconstruction is possible with one-pixel accuracy for motion parameters and a trivial first-order quadrature mirror filter (QMF). With comparable complexity, even this most simple type of MC-SBC outperforms MC prediction (MC-DPCM) techniques when blockwise motion estimation is used. The concept of MC-SBC is generalized to subpixel accuracy of MC and any even-length (odd-order) QMF filters. Motion estimation is performed in a hierarchical forward-backward procedure which gives better SNR and visual performance results than blockwise-independent estimation. The scheme is extremely error resistant due to its nonrecursive structure; the gain over MC-DPCM is remarkably high, especially in layered coding schemes as they are discussed for ATM video applications. Results of MC-SBC and other interframe coding schemes are compared using two different intrafracture schemes in layered and nonlayered coding applications.>
Jens-Rainer Ohm
ICASSP1
1990 Still image coding using predictive tree-VQ with sub-band decomposition
abstract
The combination of a linear-predictive still-image coding scheme with tree encoding and vector quantization (predictive tree-VQ, PTVQ) is applied to subband decomposed images. PTVQ allows the use of large codeword lengths and produces good image quality at bit rates around 0.3 b/pixel. Schemes with and without local bit-rate adaptation are studied. All combinations of PTVQ and subband coding (SBC) show good performance with regard to signal-to-noise-ratio (SNR) and perceptual quality criteria and outperform the original PTVQ as well as simple SBC-DPCM (differential pulse-code modulation) schemes, even at rates as low as 0.2 b/pixel.>
Jens-Rainer Ohm
ICASSP1