VLDB 2026 Research / reviewers in the wild / expert
Jörn Ostermann
dblp:o/JornOstermann
· DBLP profile ↗
166ranked-venue papers
12as first author
30since 2021 · last 2026
0000-0002-6743-3324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 146 · 12 first-author · 24 since 2021Artificial intelligence and machine learning · 31 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dataset for Automatic Vocal Mode Classification
Reemt Hinrichs, Sonja Stephan, Alexander Lange, Jörn Ostermann |
EvoMUSART | 4 |
| 2025 | Reliability of Lexical Richness Measures for ASR-Based Children's Speech AssessmentabstractEarly language acquisition is vital for child cognitive and social development. Timely identification of language delays enables effective interventions. Traditional Language Sample Analysis (LSA) methods, the gold standard for assessing language acquisition, are time-consuming, making automation a promising alternative. Yet, analyzing spontaneous speech for linguistic features is hard due to errors from spoken language processing (SLP) systems. This study uses the KidsTALC dataset to explore the link between SLP performance and LSA measures of early language acquisition. We use a child-adapted ASR model to assess LSA measures, focusing on lexical diversity, on both human and machine transcripts. Our findings show that some LSA measures, like HD-D and vocd-D, are robust to imperfect ASR, with correlations over 0.9. This suggests that ASR-driven LSA, with proper methods and metric selection, can offer valuable insights for early language evaluation and aid in therapy decisions. Imen Talbi, Christopher Gebauer, Lars Rumberg, Edith Beaulac, Hanna Ehlert, Jörn Ostermann |
ASRU | 6 |
| 2025 | Exploration of Sequence-wise Optimized Parameters for Low Complexity Enhancement Video Coding (LCEVC) on 4K ContentabstractThis paper explores the configuration space of Low Complexity Enhancement Video Coding (LCEVC) for 4K SDR content and the VVC reference software VTM as base layer codec. The configuration space is spanned using a sweep over multiple values for the base layer QP and the LCEVC parameters sublayer 2 step width, scaling mode, number of processed picture planes, transformation type, and sequence-wise optimized upscaling filter coefficients. The best configurations are selected using a convex hull approach. Compared to the default LCEVC configuration, we achieve average BD-rate gains of 10.92% and 2.06% using 1d and 2d upscaling, respectively. However, our best possible LCEVC configuration still yields an average YCbCr BD-rate loss of 15.26% compared to VTM at full resolution. Martin Benjak, Jörn Ostermann |
ICASSP | 2 |
| 2025 | HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video CodingabstractMost frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded frames, are intuitive but impose dual constraints on the output decoded frames, leading to suboptimal rate-distortion performance. In contrast, the hidden-to-hidden connection approaches, which propagate latent features within the RNN, offer greater flexibility but require large buffer sizes. To address these issues, we propose HyTIP, a learned video coding framework that combines both mechanisms. Our hybrid buffering strategy uses explicit decoded frames and a small number of implicit latent features to achieve competitive coding performance. Experimental results show that our HyTIP outperforms the sole use of either output-recurrence or hidden-to-hidden approaches. Furthermore, it achieves comparable performance to state-of-the-art methods but with a much smaller buffer size, and outperforms VTM 17.0 (Low-delay B) in terms of PSNR-RGB and MS-SSIM-RGB. The source code of HyTIP is available at https://github.com/NYCU-MAPL/HyTIP. Yi-Hsin Chen, Yi-Chen Yao, Kuan-Wei Ho, Chun-Hung Wu, Huu-Tai Phung, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
ICCV | 7 |
| 2025 | Learned Hybrid Video Coding for Human Perception and Multiple Machine Vision TasksabstractIn this work, we present a learned multi-task video codec that is optimized for human and machine vision. The codec consists of an encoder that maps images from the pixel domain to a latent representation and multiple decoders that map the latent to either an image for human consumption or multiple task-specific features for different machine vision tasks. This allows a single bitstream to be used for multiple tasks while also reducing the decoder complexity for machine vision tasks. Unlike most learned codecs, our method performs inter-coding at the latent level instead of the pixel domain. Experiments show that the proposed method achieves a compression performance for machine vision tasks comparable to other multi-task codecs designed for machine vision only, while also providing video reconstruction. The code is available at https://github.com/GreenAutoML4FAS/HybridMultiTaskCoding. Martin Benjak, Saifullah Khan, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
ICIP | 5 |
| 2025 | Conditional Residual Coding with Explicit-Implicit Temporal Buffering for Learned Video CompressionabstractThis work proposes a hybrid, explicit-implicit temporal buffering scheme for conditional residual video coding. Recent conditional coding methods propagate implicit temporal information for inter-frame coding, demonstrating superior coding performance to those relying exclusively on previously decoded frames (i.e. the explicit temporal information). However, these methods require substantial memory to store a large number of implicit features. This work presents a hybrid buffering strategy. For inter-frame coding, it buffers one previously decoded frame as the explicit temporal reference and a small number of learned features as implicit temporal reference. Our hybrid buffering scheme for conditional residual coding outperforms the single use of explicit or implicit information. Moreover, it allows the total buffer size to be reduced to the equivalent of two video frames with a negligible performance drop on 2K video sequences. The ablation experiment further sheds light on how these two types of temporal references impact the coding performance. Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
ICME | 4 |
| 2025 | Grammatical Error Detection on Spontaneous Children's Speech Using Iterative Pseudo Labeling
Christopher Gebauer, Lars Rumberg, Lars Köhn, Hanna Ehlert, Edith Beaulac, Jörn Ostermann |
INTERSPEECH | 6 |
| 2025 | Scalable COOL-CHIC: Dual-Resolution Images from a Single Bitstream
Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
PCS | 4 |
| 2025 | A Cross-Framework Study of Temporal Information Buffering Strategies for Learned Video Compression
Kuan-Wei Ho, Yi-Hsin Chen, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
PCS | 4 |
| 2025 | Progressive COOL-CHIC: Efficient Decoding for Dual-Resolution ImagesabstractIn this work, we propose Progressive Cool-Chic (PCC), a scalable overfitted neural image codec that can decode an image at two different resolutions from a single bitstream. Experiments show that our method reduces the necessary bitrate to encode two representations of the same image by up to 31.54% in terms of BD-rate compared to encoding both representations independently using Cool-Chic while also decreasing the necessary decoding time. The bitstream is structured in a way that the low-resolution image can already be decoded, when only a part of the bitstream has been received. The code is available at https://github.com/mbenjak/progressive-CC. Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
VCIP | 4 |
| 2024 | On the Rate-Distortion-Complexity Trade-Offs of Neural Video CodingabstractThis paper aims to delve into the rate-distortion-complexity trade-offs of modern neural video coding. Recent years have witnessed much research effort being focused on exploring the full potential of neural video coding. Conditional auto encoders have emerged as the mainstream approach to efficient neural video coding. The central theme of conditional auto encoders is to leverage both spatial and temporal information for better conditional coding. However, a recent study indicates that conditional coding may suffer from information bottlenecks, potentially performing worse than traditional residual coding. To address this issue, recent conditional coding methods incorporate a large number of high-resolution features as the condition signal, leading to a considerable increase in the number of multiply-accumulate operations, memory footprint, and model size. Taking DCVC as the common code base, we investigate how the newly proposed conditional residual coding, an emerging new school of thought, and its variants may strike a better balance among rate, distortion, and complexity. Yi-Hsin Chen, Kuan-Wei Ho, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng |
MMSP | 4 |
| 2024 | HiCMC: High-Efficiency Contact Matrix CompressorabstractBACKGROUND: Chromosome organization plays an important role in biological processes such as replication, regulation, and transcription. One way to study the relationship between chromosome structure and its biological functions is through Hi-C studies, a genome-wide method for capturing chromosome conformation. Such studies generate vast amounts of data. The problem is exacerbated by the fact that chromosome organization is dynamic, requiring snapshots at different points in time, further increasing the amount of data to be stored. We present a novel approach called the High-Efficiency Contact Matrix Compressor (HiCMC) for efficient compression of Hi-C data. RESULTS: By modeling the underlying structures found in the contact matrix, such as compartments and domains, HiCMC outperforms the state-of-the-art method CMC by approximately 8% and the other state-of-the-art methods cooler, LZMA, and bzip2 by over 50% across multiple cell lines and contact matrix resolutions. In addition, HiCMC integrates domain-specific information into the compressed bitstreams that it generates, and this information can be used to speed up downstream analyses. CONCLUSION: HiCMC is a novel compression approach that utilizes intrinsic properties of contact matrix, such as compartments and domains. It allows for a better compression in comparison to the state-of-the-art methods. HiCMC is available at https://github.com/sXperfect/hicmc . Yeremia Gunawan Adhisantoso, Tim Körner, Fabian Müntefering, Jörn Ostermann, Jan Voges |
BMC Bioinform. | 4 |
| 2024 | MaskCRT: Masked Conditional Residual Transformer for Learned Video CompressionabstractConditional coding has lately emerged as the mainstream approach to learned video compression. However, a recent study shows that it may perform worse than residual coding when the information bottleneck arises. Conditional residual coding was thus proposed, creating a new school of thought to improve on conditional coding. Notably, conditional residual coding relies heavily on the assumption that the residual frame has a lower entropy rate than that of the intra frame. Recognizing that this assumption is not always true due to dis-occlusion phenomena or unreliable motion estimates, we propose a masked conditional residual coding scheme. It learns a soft mask to form a hybrid of conditional coding and conditional residual coding in a pixel adaptive manner. We introduce a Transformer-based conditional autoencoder. Several strategies are investigated with regard to how to condition a Transformer-based autoencoder for inter-frame coding, a topic that is largely under-explored. Additionally, we propose a channel transform module (CTM) to decorrelate the image latents along the channel dimension, with the aim of using the simple hyperprior to approach similar compression performance to the channel-wise autoregressive model. Experimental results confirm the superiority of our masked conditional residual transformer (termed MaskCRT) to both conditional coding and conditional residual coding. On commonly used datasets, MaskCRT shows comparable BD-rate results to VTM-17.0 under the low delay P configuration in terms of PSNR-RGB and outperforms VTM-17.0 in terms of MS-SSIM-RGB. It also opens up a new research direction for advancing learned video compression. Yi-Hsin Chen, Hong-Sheng Xie, Cheng-Wei Chen, Zong-Lin Gao, Martin Benjak, Wen-Hsiao Peng, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Continuous Sign-Language Recognition using Transformers and Augmented Pose Estimation
Reemt Hinrichs, Angelo Jovin Yamachui Sitcheu, Jörn Ostermann |
ICPRAM | 3 |
| 2023 | Exploiting Diversity of Automatic Transcripts from Distinct Speech Recognition Techniques for Children's Speech
Christopher Gebauer, Lars Rumberg, Hanna Ehlert, Ulrike Lüdtke, Jörn Ostermann |
INTERSPEECH | 5 |
| 2023 | Uncertainty Estimation for Connectionist Temporal Classification Based Automatic Speech Recognition
Lars Rumberg, Christopher Gebauer, Hanna Ehlert, Maren Wallbaum, Ulrike Lüdtke, Jörn Ostermann |
INTERSPEECH | 6 |
| 2023 | Learning-Based Scalable Video Coding with Spatial and Temporal PredictionabstractIn this work, we propose a hybrid learning-based method for layered spatial scalability. Our framework consists of a base layer (BL), which encodes a spatially downsampled representation of the input video using Versatile Video Coding (VVC), and a learning-based enhancement layer (EL), which conditionally encodes the original video signal. The EL is conditioned by two fused prediction signals: a spatial inter-layer prediction signal, that is generated by spatially upsampling the output of the BL using super-resolution, and a temporal inter-frame prediction signal, that is generated by decoder-side motion compensation without signaling any motion vectors. We show that our method outperforms LCEVC and has comparable performance to full-resolution VVC for high-resolution content, while still offering scalability. Martin Benjak, Yi-Hsin Chen, Wen-Hsiao Peng, Jörn Ostermann |
VCIP | 4 |
| 2023 | Rate Adaptation for Learned Two-layer B-frame Coding without Signaling Motion InformationabstractThis paper explores the potential of a learned two-layer B-frame codec, known as TLZMC. TLZMC is one of the few early attempts that deviate from the hybrid-based coding architecture by skipping motion coding. With TLZMC, a low-resolution base layer is utilized to encode temporally unpredictable information. We address the question of whether adapting the base-layer bitrate can achieve better rate-distortion performance. We apply the feature map modulation technique to enable per-frame bitrate adaptation of the base layer. We then propose and compare three online search strategies for determining the base-layer rate parameter: per-level brute-force search, per-level greedy search, and per-frame greedy search. Experimental results show that our top-performing search strategy achieves 0.6%-15.8% Bjøntegaard-Delta rate savings over TLZMC. Hong-Sheng Xie, Yi-Hsin Chen, Wen-Hsiao Peng, Martin Benjak, Jörn Ostermann |
VCIP | 5 |
| 2023 | GVC: efficient random access compression for gene sequence variationsabstractBACKGROUND: In recent years, advances in high-throughput sequencing technologies have enabled the use of genomic information in many fields, such as precision medicine, oncology, and food quality control. The amount of genomic data being generated is growing rapidly and is expected to soon surpass the amount of video data. The majority of sequencing experiments, such as genome-wide association studies, have the goal of identifying variations in the gene sequence to better understand phenotypic variations. We present a novel approach for compressing gene sequence variations with random access capability: the Genomic Variant Codec (GVC). We use techniques such as binarization, joint row- and column-wise sorting of blocks of variations, as well as the image compression standard JBIG for efficient entropy coding. RESULTS: Our results show that GVC provides the best trade-off between compression and random access compared to the state of the art: it reduces the genotype information size from 758 GiB down to 890 MiB on the publicly available 1000 Genomes Project (phase 3) data, which is 21% less than the state of the art in random-access capable methods. CONCLUSIONS: By providing the best results in terms of combined random access and compression, GVC facilitates the efficient storage of large collections of gene sequence variations. In particular, the random access capability of GVC enables seamless remote data access and application integration. The software is open source and available at https://github.com/sXperfect/gvc/ . Yeremia Gunawan Adhisantoso, Jan Voges, Christian Rohlfing, Viktor Tunev, Jens-Rainer Ohm, Jörn Ostermann |
BMC Bioinform. | 6 |
| 2022 | Contact Matrix CompressorabstractThe study of three-dimensional folding of chromosomes is important to understand genomics processes. This is done through techniques, such as Hi-C, that analyze the spatial organization of chromosomes in a cell. The data coming from the study is a 2-dimensional quantitative maps with genomic coordinate systems. We present a novel approach called Contact Matrix Compressor(CMC) for the efficient compression of Hi-C data. By exploiting the properties of the data, such as diagonally dominant and symmetrical, CMC achieves a much higher compression. CMC outperforms the existing method Cooler, and also the generic compression methods LZMA as well as BZip2. Yeremia Gunawan Adhisantoso, Jörn Ostermann |
DCC | 2 |
| 2022 | Improved Compression of Artificial Neural Networks through Curvature-Aware TrainingabstractArtificial neural networks achieve state-of-the-art performance in many branches of engineering. As such they are used for all kinds of tasks and nowadays are desired to be used on mobile devices like smartphones. Due to limited hardware resources or limited channel capacity on mobile devices, compression of neural network models to reduce storage or transmission costs is desired. Furthermore, reduced complexity is of interest. This work investigates introducing the curvature of the loss surface in the training of artificial neural networks and analyzes its benefit for the compression of neural networks through quantization and pruning of its weights. As proof-of-concept, three small LeNet-based neural networks were trained using a novel loss function consisting of a weighted average of the cross-entropy loss and the Frobenius norm of the hessian matrix. That way both, the loss as well as the local curvature, are minimized concurrently. Using the proposed method, mean test accuracies on the MNIST and FashionMNIST datasets after quantization were considerably improved by up to about 47.6 % for 1 bit quantization on MNIST and about 27.8 % on FashionMNIST compared to quantization after training without curvature information. Additionally, pruning was found to benefit from introducing curvature in the training as well with an increase of up to about 14.6 % mean test accuracy compared to pruning after training without curvature except for isolated cases. Training the artificial neural networks first without curvature information and subsequent training by only one epoch using curvature information allowed to increase the mean test accuracy after quantization at 1 bit by about 16 %. The proposed method can potentially improve the accuracy after compression irrespective of the compression method applied. Reemt Hinrichs, Ze Lu, Jörn Ostermann |
IJCNN | 4 |
| 2022 | Improving Phonetic Transcriptions of Children's Speech by Pronunciation Modelling with Constrained CTC-Decoding
Lars Rumberg, Christopher Gebauer, Hanna Ehlert, Ulrike Lüdtke, Jörn Ostermann |
INTERSPEECH | 5 |
| 2022 | kidsTALC: A Corpus of 3- to 11-year-old German Children's Connected Natural Speech
Lars Rumberg, Christopher Gebauer, Hanna Ehlert, Maren Wallbaum, Lena Bornholt, Jörn Ostermann, Ulrike Lüdtke |
INTERSPEECH | 6 |
| 2022 | Neural Network-based Error Concealment for B-Frames in VVCabstractIn this paper we introduce an error concealment method for VVC that error-conceals B-frames based on the neural frame interpolation network RIFE. The network is trained using the BVI-DVC dataset to infer even full-HD frames. We integrate our proposed model in the VVC reference software VTM for its evaluation. The average error of a whole GOP with a single corrupted frame is decreased by 15% and 24% in terms of PSNR measurement compared to block matching and frame copy, respectively. To our knowledge, our approach is currently the best performing error concealment algorithm for single slice per B-frame settings. Martin Benjak, Niklas Aust, Yasser Samayoa, Jörn Ostermann |
ISCAS | 4 |
| 2021 | Relative Pose Consistency for Semi-Supervised Head Pose EstimationabstractHuman head pose estimation from images plays a vital role in applications like driver assistance systems and human behavior analysis. Head pose estimation networks are typically trained in a supervised manner. Unfortunately, manual/sensor-based annotations of head poses are prone to errors. A solution is supervised training on synthetic training data generated from 3D face models which can provide an infinite amount of perfect labels. However, computer generated face images only provide an approximation of real-world images which results in a domain gap between training and application domain. To date, domain adaptation is rarely addressed in current work on head pose estimation. In this work we propose relative pose consistency, a semi-supervised learning strategy for head pose estimation based on consistency regularization. It allows simultaneous learning on labeled synthetic data and unlabeled real-world data to overcome the domain gap, while keeping the advantages of synthetic data. Consistency regularization enforces consistent network predictions under random image augmentations. We address pose-preserving and pose-altering augmentations. Naturally, pose-altering augmentations cannot be used on unlabeled data. We therefore propose a strategy to exploit the relative pose introduced by pose-altering augmentations between augmented image pairs. This allows the network to benefit from relative pose labels during training on the unlabeled, real-world images. We evaluate our approach on a widely used benchmark (Biwi Kinect Head Pose) and outperform domain-adaptation SOTA. We are the first to present a consistency regularization framework for head pose estimation. Our experiments show that our approach improves head pose estimation accuracy for real-world images despite using only labels from synthetic images. Felix Kuhnke, Sontje Ihler, Jörn Ostermann |
FG | 3 |
| 2021 | Neural Network-Based Error Concealment For VVCabstractIn this paper we introduce an error concealment method for VVC based on deep recurrent neural networks, which employs the PredNet model to estimate missing video frames by using past decoded frames. The network is trained using the BVI-DVC data set to infer even full-HD frames. We integrated our proposed model in the VVC reference software VTM for its evaluation. It performs, in average, 6 dB or up to 5 dB better than the frame copy model in terms of PSNR measurements for a concealed I-frame or P-frame, respectively. Martin Benjak, Yasser Samayoa, Jörn Ostermann |
ICIP | 3 |
| 2021 | Age-Invariant Training for End-to-End Child Speech Recognition Using Adversarial Multi-Task Learning
Lars Rumberg, Hanna Ehlert, Ulrike Lüdtke, Jörn Ostermann |
Interspeech | 4 |
| 2021 | Contour-based Intra Coding Using Gaussian Processes and Neural NetworksabstractIntra prediction is an essential part of video coding. In this work, two methods are proposed for improving intra prediction. The first contribution is a stochastic contour model for modeling and extrapolation of contours detected in the reference area. A Gaussian process is used for the modeling and a multivariate Gaussian distribution is formulated for the extrapolation. The second contribution is a neural network-based method for sample value prediction. The neural networks are used to process the adjacent reference sample values and the results of contour modeling and extrapolation as input data to generate a prediction of the sample values of the block to be coded. The neural networks were designed with an auto-encoder architecture and trained to minimize the appproximated bit rate of the prediction error. The coding efficiency of the video codec HEVC is increased by up to 5%. Averaged over all 55 test sequences, the All Intra configuration resulted in BD-rates of -0.54% for high bit rates and -1.0% for low bit rates. Thorsten Laude, Jörn Ostermann |
PCS | 2 |
| 2021 | Multilinear Modelling of Faces and ExpressionsabstractIn this work, we present a new versatile 3D multilinear statistical face model, based on a tensor factorisation of 3D face scans, that decomposes the shapes into person and expression subspaces. Investigation of the expression subspace reveals an inherent low-dimensional substructure, and further, a star-shaped structure. This is due to two novel findings. (1) Increasing the strength of one emotion approximately forms a linear trajectory in the subspace. (2) All these trajectories intersect at a single point - not at the neutral expression as assumed by almost all prior works-but at an apathetic expression. We utilise these structural findings by reparameterising the expression subspace by the fourth-order moment tensor centred at the point of apathy. We propose a 3D face reconstruction method from single or multiple 2D projections by assuming an uncalibrated projective camera model. The non-linearity caused by the perspective projection can be neatly included into the model. The proposed algorithm separates person and expression subspaces convincingly, and enables flexible, natural modelling of expressions for a wide variety of human faces. Applying the method on independent faces showed that morphing between different persons and expressions can be performed without strong deformations. Stella Graßhof, Hanno Ackermann, Sami S. Brandt, Jörn Ostermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | An Introduction to MPEG-G: The First Open ISO/IEC Standard for the Compression and Exchange of Genomic Sequencing DataabstractThe development and progress of high-throughput sequencing technologies have transformed the sequencing of DNA from a scientific research challenge to practice. With the release of the latest generation of sequencing machines, the cost of sequencing a whole human genome has dropped to less than $\$ $ 600. Such achievements open the door to personalized medicine, where it is expected that genomic information of patients will be analyzed as a standard practice. However, the associated costs, related to storing, transmitting, and processing the large volumes of data, are already comparable to the costs of sequencing. To support the design of new and interoperable solutions for the representation, compression, and management of genomic sequencing data, the Moving Picture Experts Group (MPEG) jointly with working group 5 of ISO/TC276 “Biotechnology” has started to produce the ISO/IEC 23092 series, known as MPEG-G. MPEG-G does not only offer higher levels of compression compared with the state of the art but it also provides new functionalities, such as built-in support for random access in the compressed domain, support for data protection mechanisms, flexible storage, and streaming capabilities. MPEG-G only specifies the decoding syntax of compressed bitstreams, as well as a file format and a transport format. This allows for the development of new encoding solutions with higher degrees of optimization while maintaining compatibility with any existing MPEG-G decoder. Jan Voges, Mikel Hernaez, Marco Mattavelli, Jörn Ostermann |
Proc. IEEE | 4 |
| 2020 | Two-Stream Aural-Visual Affect Analysis in the WildabstractHuman affect recognition is an essential part of natural human-computer interaction. However, current methods are still in their infancy, especially for in-the-wild data. In this work, we introduce our submission to the Affective Behavior Analysis in-the-wild (ABAW) 2020 competition. We propose a two-stream aural-visual analysis model to recognize affective behavior from videos. Audio and image streams are first processed separately and fed into a convolutional neural network. Instead of applying recurrent architectures for temporal analysis we only use temporal convolutions. Furthermore, the model is given access to additional features extracted during face-alignment. At training time, we exploit correlations between different emotion representations to improve performance. Our model achieves promising results on the challenging Aff-Wild2 database. Felix Kuhnke, Lars Rumberg, Jörn Ostermann |
FG | 3 |
| 2020 | Non-Line-of-Sight Time-Difference-of-Arrival Localization with Explicit Inclusion of Geometry Information in a Simple Diffraction ScenarioabstractTime-difference-of-arrival (TDOA) localization is a technique for finding the position of a wave emitting object, e.g., a car horn. Many algorithms have been proposed for TDOA localization under line-of-sight (LOS) conditions. In the non-line-of-sight (NLOS) case the performance of these algorithms usually deteriorates. There are techniques to reduce the error introduced by the NLOS condition, which, however, do not directly take into account information on the geometry of the surroundings. In this paper a NLOS TDOA localization approach for a simple diffraction scenario is described, which includes information on the surroundings into the equation system. An experiment with three different loudspeaker positions was conducted to validate the proposed method. The localization error was less than 6.2 % of the distance from the source to the closest microphone position. Simulations show that the proposed method attains the Cramer-Rao-Lower-Bound for low enough TDOA noise levels. Sönke Südbeck, Thomas Krause 0003, Jörn Ostermann |
MMSP | 3 |
| 2020 | GABAC: an arithmetic coding solution for genomic dataabstractMOTIVATION: In an effort to provide a response to the ever-expanding generation of genomic data, the International Organization for Standardization (ISO) is designing a new solution for the representation, compression and management of genomic sequencing data: the Moving Picture Experts Group (MPEG)-G standard. This paper discusses the first implementation of an MPEG-G compliant entropy codec: GABAC. GABAC combines proven coding technologies, such as context-adaptive binary arithmetic coding, binarization schemes and transformations, into a straightforward solution for the compression of sequencing data. RESULTS: We demonstrate that GABAC outperforms well-established (entropy) codecs in a significant set of cases and thus can serve as an extension for existing genomic compression solutions, such as CRAM. AVAILABILITY AND IMPLEMENTATION: The GABAC library is written in C++. We also provide a command line application which exercises all features provided by the library. GABAC can be downloaded from https://github.com/mitogen/gabac. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jan Voges, Tom Paridaens, Fabian Müntefering, Liudmila S. Mainzer, Brian Bliss, Idoia Ochoa, Jan Fostier, Jörn Ostermann, Mikel Hernaez |
Bioinform. | 9 |
| 2020 | Analysis of Affine Motion-Compensated Prediction in Video CodingabstractMotion-compensated prediction is used in video coding standards like High Efficiency Video Coding (HEVC) as one key element of data compression. Commonly, a purely translational motion model is employed. In order to also cover non-translational motion types like rotation or scaling (zoom), e. g. contained in aerial video sequences such as captured from unmanned aerial vehicles (UAV), an affine motion model can be applied. In this work, a model for affine motion-compensated prediction in video coding is derived. Using the rate-distortion theory and the displacement estimation error caused by inaccurate affine motion parameter estimation, the minimum required bit rate for encoding the prediction error is determined. In this model, the affine transformation parameters are assumed to be affected by statistically independent estimation errors, which all follow a zero-mean Gaussian distributed probability density function (pdf). The joint pdf of the estimation errors is derived and transformed into the pdfof the location-dependent displacement estimation error in the image. The latter is related to the minimum required bit rate for encoding the prediction error. Similar to the derivations of the fully affine motion model, a four-parameter simplified affine model is investigated. Both models are of particular interest since they are considered for the upcoming video coding standard Versatile Video Coding (VVC) succeeding HEVC. Both models provide valuable information about the minimum bit rate for encoding the prediction error as a function of affine estimation accuracies. Holger Meuel, Jörn Ostermann |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Head Pose Estimation Using Synthetic Images and Partial Adversarial Domain Adaption for Continuous Label SpacesabstractHead pose estimation aims at predicting an accurate pose from an image. Current approaches rely on supervised deep learning, which typically requires large amounts of labeled data. Manual or sensor-based annotations of head poses are prone to errors. A solution is to generate synthetic training data by rendering 3D face models. However, the differences (domain gap) between rendered (source-domain) and real-world (target-domain) images can cause low performance. Advances in visual domain adaptation allow reducing the influence of domain differences using adversarial neural networks, which match the feature spaces between domains by enforcing domain-invariant features. While previous work on visual domain adaptation generally assumes discrete and shared label spaces, these assumptions are both invalid for pose estimation tasks. We are the first to present domain adaptation for head pose estimation with a focus on partially shared and continuous label spaces. More precisely, we adapt the predominant weighting approaches to continuous label spaces by applying a weighted resampling of the source domain during training. To evaluate our approach, we revise and extend existing datasets resulting in a new benchmark for visual domain adaption. Our experiments show that our method improves the accuracy of head pose estimation for real-world images despite using only labels from synthetic images. Felix Kuhnke, Jörn Ostermann |
ICCV | 2 |
| 2019 | HEVC Inter Coding using Deep Recurrent Neural Networks and Artificial Reference PicturesabstractThe efficiency of motion compensated prediction in modern video codecs highly depends on the available reference pictures. Occlusions and non-linear motion pose challenges for the motion compensation and often result in high bit rates for the prediction error. We propose the generation of artificial reference pictures using deep recurrent neural networks. Conceptually, a reference picture at the time instance of the currently coded picture is generated from previously reconstructed conventional reference pictures. Based on these artificial reference pictures, we propose a complete coding pipeline based on HEVC. By using the artificial reference pictures for motion compensated prediction, average BD-rate gains of 1.5% over HEVC are achieved. Thorsten Laude, Felix Haub, Jörn Ostermann |
PCS | 3 |
| 2019 | Occlusion-Aware Method for Temporally Consistent SuperpixelsabstractA wide variety of computer vision applications rely on superpixel or supervoxel algorithms as a preprocessing step. This underlines the overall importance that these approaches have gained in recent years. However, most methods show a lack of temporal consistency or fail in producing temporally stable superpixels. In this paper, we present an approach to generate temporally consistent superpixels for video content. Our method is formulated as a contour-evolving expectation-maximization framework, which utilizes an efficient label propagation scheme to encourage the preservation of superpixel shapes and their relative positioning over time. By explicitly detecting the occlusion of superpixels and the disocclusion of new image regions, our framework is able to terminate and create superpixels whose corresponding image region becomes hidden or newly appears. Additionally, the occluded parts of superpixels are incorporated in the further optimization. This increases the compliance of the superpixel flow with the optical flow present in the scene. Using established benchmark suites, we show the performance of our approach in comparison to state-of-the-art supervoxel and superpixel algorithms for video content. Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Backprojection Subimage Autofocus of Moving Ships for Synthetic Aperture RadarabstractWe propose a new autofocus approach for the backprojection reconstruction algorithm to compute high-quality synthetic aperture radar images of non-linearly moving and maneuvering ships. In contrast to the state-of-the-art autofocus techniques, our approach allows a long coherent processing interval even in the case of a rough sea, which improves the image quality. An improved image quality enables the classification of ships in airborne synthetic aperture radar (SAR) images. For this purpose, we decompose the image into subimages and estimate pulse-by-pulse a phase error for each subimage by maximizing subimage sharpness. A regularized Levenberg-Marquardt algorithm guarantees a smooth phase correction on subimage level. By correcting the subsequent range distances from the flight path to all pixels using the currently estimated phase errors, sharp images of maneuvering ships with arbitrary velocities can now be reconstructed. The evaluation of our proposed ship autofocus technique on the basis of real airborne X-band data shows that our approach leads to a visible improvement of image quality in comparison with the state-of-the-art techniques. Given these results, even an automatic ship classification based on radar images might be possible in the future. Aron Sommer, Jörn Ostermann |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Robust Fourier-Based Checkerboard Corner Detection for Camera Calibration
Benjamin Spitschan, Jörn Ostermann |
CIARP | 2 |
| 2018 | Lossy Compression of Quality Scores in Differential Gene Expression: A First Assessment and Impact AnalysisabstractHigh-throughput sequencing of RNA molecules has enabled the quantitative analysis of gene expression at the expense of storage space and processing power. To alleviate these problems, lossy compression methods of the quality scores associated to RNA sequencing data have recently been proposed, and the evaluation of their impact on downstream analyses is gaining attention. In this context, this work presents a first assessment of the impact of lossily compressed quality scores in RNA sequencing data on the performance of some of the most recent tools used for differential gene expression. Ana A. Hernandez-Lopez, Jan Voges, Claudio Alberti, Marco Mattavelli, Jörn Ostermann |
DCC | 5 |
| 2018 | Detail-Aware Image Decomposition for an HEVC-Based Texture Synthesis FrameworkabstractModern video coding standards like High Efficiency Video Coding (HEVC) provide superior coding efficiency. However, this does not state true for complex and hard to predict textures which require high bit rates to achieve a high quality. To overcome this limitation of HEVC, texture synthesis frameworks were proposed in previous works. However, these frameworks only result in good reconstruction quality if the decomposition into synthesizable and non-synthesizable regions is either known or trivial. The frameworks fail for more challenging content, e.g. for content with fine non-synthesizable details within synthesizable regions. To enable texture synthesis-based video coding with high quality for this content, we propose sophisticated detail-aware decomposition techniques in this paper. These techniques are based on an initial coarse segmentation step followed by a refinement step that detects even small differences in the previously segmented region. With this new approach, we are able to achieve average luma BD-rate gains of 13.77% over HEVC and 3.03% over the closest related work from the literature. Furthermore, the considerably improved visual quality in addition to the bit rate savings is confirmed by comprehensive subjective tests. Bastian Wandt, Thorsten Laude, Bodo Rosenhahn, Jörn Ostermann |
DCC | 4 |
| 2018 | Rate-Distortion Theory for Affine Global Motion Compensation in Video CodingabstractIn this work, we derive the rate-distortion function for video coding using affine global motion compensation. We model the displacement estimation error during motion estimation and obtain the bit rate after applying the rate-distortion theory. We assume that the displacement estimation error is caused by a perturbed affine transformation. The 6 affine transformation parameters are assumed statistically independent, with each of them having a zero-mean Gaussian distributed estimation error. Based on that, the joint p.d.f. of the displacement estimation errors is derived and related to the prediction error. Using the ratedistortion theory, we calculate the bit rate in dependence of the perturbation of the affine transformation parameters. Comparing with a translational motion model in video coding standards like HEVC, we determine accuracy boundaries for the affine transformation, with which a gain can be achieved. Holger Meuel, Stephan Ferenz, Yiqun Liu 0003, Jörn Ostermann |
ICIP | 4 |
| 2018 | A Comparison of JEM and AV1 with HEVC: Coding Tools, Coding Efficiency and ComplexityabstractThe current state-of-the-art for standardized video codecs is High Efficiency Video Coding (HEVC) which was developed jointly by ISO/IEC and ITU-T. Recently, the development of two contenders for the next generation of standardized video codecs began: ISO/IEC and ITU-T advance the development of the Joint Exploration Model (JEM), a possible successor of HEVC, while the Alliance for Open Media pushes forward the video codec AV1. It is asserted by both groups that their codecs achieve superior coding efficiency over the state-of-the-art. In this paper, we discuss the distinguishing features of JEM and AV1 and evaluate their coding efficiency and computational complexity under well-defined and balanced test conditions. Our main findings are that JEM considerably outperforms HM and AV1 in terms of coding efficiency while AV1 cannot transform increased complexity into competitiveness in terms of coding efficiency with neither of the competitors except for the all-intra configuration. Thorsten Laude, Yeremia Gunawan Adhisantoso, Jan Voges, Marco Munderloh, Jörn Ostermann |
PCS | 5 |
| 2018 | Scene-based KLT for Intra Coding in HEVCabstractTransform coding and quantization are part of the cornerstones in the current High Efficiency Video Coding (HEVC) standard. They are applied on the residuals from inter-frame or intra predictions. With specified transform matrices, HEVC enhances the coding efficiency vastly compared to Advanced Video Coding (AVC). However, there is still room for improvement. It is observed that the coding of transform coefficients occupies the majority of the bit rate in the stream, since transform matrices in HEVC can not offer the best energy compaction for prediction errors, especially for diagonal features. We introduce scene-based Karhunen-Loeve transform (KLT) in place of the conventional transform for the intra-predicted data for 8 × 8 and 16 × 16 Transform Units (TU). The transform matrices are adaptively designed and later applied according to the prediction modes, quantization steps as well as sizes. The simulation shows great prospect of reducing the bit rate further with KLT, as we gain 3.23%, 7.18% and 6.25% in terms of BD-Rate against HM-16.15 on average for class B, class C and BVI textures respectively, with All-Intra configuration. Yiqun Liu 0003, Jörn Ostermann |
PCS | 2 |
| 2018 | Physical High Dynamic Range Imaging with Conventional SensorsabstractThis paper aims at simplified high dynamic range (HDR) image generation with non-modified, conventional camera sensors. One typical HDR approach is exposure bracketing, e.g. with varying shutter speeds. It requires to capture the same scene multiple times at different exposure times. These pictures are then merged into a single HDR picture which typically is converted back to an 8-bit image by using tone-mapping. Existing works on HDR imaging focus on image merging and tone mapping whereas we aim at simplified image acquisition. The proposed algorithm can be used in consumer-level cameras without hardware modifications at sensor level. Based on intermediate samplings of each sensor element during the total (pre-defined) exposure time, we extrapolate the luminance of sensor elements which are saturated after the total exposure time. Compared to existing HDR approaches which typically require three different images with carefully determined exposure times, we only take one image at the longest exposure time. The shortened total time between start and end of image acquisition can reduce ghosting artifacts. The experimental evaluation demonstrates the effectiveness of the algorithm. Holger Meuel, Hanno Ackermann, Bodo Rosenhahn, Jörn Ostermann |
PCS | 4 |
| 2018 | Extending HEVC with a Texture Synthesis Framework using Detail-aware Image DecompositionabstractIn recent years, there has been a tremendous improvement in video coding algorithms. This improvement resulted in 2013 in the standardization of the first version of High Efficiency Video Coding (HEVC) which now forms the state-of-theart with superior coding efficiency. Nevertheless, the development of video coding algorithms did not stop as HEVC still has its limitations. Especially for complex textures HEVC reveals one of its limitations. As these textures are hard to predict, very high bit rates are required to achieve a high quality. Texture synthesis was proposed as solution for this limitation in previous works. However, previous texture synthesis frameworks only prevailed if the decomposition into synthesizable and non-synthesizable regions was either known or very easy. In this paper, we address this scenario with a texture synthesis framework based on detail-aware image decomposition techniques. Our techniques are based on a multiple-steps coarse-to-fine approach in which an initial decomposition is refined with awareness for small details. The efficiency of our approach is evaluated objectively and subjectively: BD-rate gains of up to 28.81% over HEVC and up to 12.75% over the closest related work were achieved. Our subjective tests indicate an improved visual quality in addition to the bit rate savings. Bastian Wandt, Thorsten Laude, Bodo Rosenhahn, Jörn Ostermann |
PCS | 4 |
| 2018 | Rate-Distortion Theory for Simplified Affine Motion Compensation Used in Video CodingabstractIn this work, we derive the rate-distortion function for video coding using the simplified affine, 4-parameter motion compensation model as it is used in the Joint Exploration Model (JEM) by the Joint Video Exploration Team (JVET) on Future Video coding. We model the displacement estimation error during motion estimation and obtain the bit rate by applying the rate-distortion theory. We assume that the displacement estimation error is caused by perturbed parameters of the simplified affine model. These transformation parameters are assumed statistically independent, with each of them having a zero-mean Gaussian distributed estimation error. The joint probability density function (p.d.f.) of the displacement estimation errors is derived and related to the prediction error. We calculate the bit rate as a function of the accuracy of the parameter estimation for the simplified affine motion model. Finally, we compare our results with a translational motion model as used in video coding standards like HEVC as well as with a full affine motion model with 6 degrees of freedom. For aerial sequences containing distinct affine motion, the minimum required bit rate to encode the prediction error can be significantly reduced from 2.5bit/sampleto 0.02 bit/sample for a reasonable operating point and a block size of 64×64 pel2. Holger Meuel, Stephan Ferenz, Yiqun Liu 0003, Jörn Ostermann |
VCIP | 4 |
| 2018 | Low-Cost Channel Sounder Design Based on Software-Defined Radio and OFDMabstractIn this paper we describe the design of a low-cost and portable channel sounding system. The baseband signal processing in both the transmitter and the receiver are done in software, while the up/down conversion are done by means of a Software Defined Radio (SDR) frontend. The modulation procedure is based on the Orthogonal Frequency Division Multiplexing (OFDM) technique due to its robustness against inter-symbol interference (ISI). The sounder is designed to overcome some well known important issues of any OFDM system. For example, it does not need extra resources neither to guarantee a ISI-free transmission by means of a guard interval, nor for the synchronization by means of a preamble. Moreover, the excitation signal is designed to have a peak-to-average power ratio (PAPR) close to 0 dB while maintaining a flat spectrum. The system is calibrated and validated under laboratory conditions prior to the measurement of an air-to-ground channel in the C frequency band, the results are presented and discussed. Yasser Samayoa, Markus Kock, Holger Blume, Jörn Ostermann |
VTC Fall | 4 |
| 2018 | CALQ: compression of quality values of aligned sequencing dataabstractMotivation: Recent advancements in high-throughput sequencing technology have led to a rapid growth of genomic data. Several lossless compression schemes have been proposed for the coding of such data present in the form of raw FASTQ files and aligned SAM/BAM files. However, due to their high entropy, losslessly compressed quality values account for about 80% of the size of compressed files. For the quality values, we present a novel lossy compression scheme named CALQ. By controlling the coarseness of quality value quantization with a statistical genotyping model, we minimize the impact of the introduced distortion on downstream analyses. Results: We analyze the performance of several lossy compressors for quality values in terms of trade-off between the achieved compressed size (in bits per quality value) and the Precision and Recall achieved after running a variant calling pipeline over sequencing data of the well-known NA12878 individual. By compressing and reconstructing quality values with CALQ, we observe a better average variant calling performance than with the original data while achieving a size reduction of about one order of magnitude with respect to the state-of-the-art lossless compressors. Furthermore, we show that CALQ performs as good as or better than the state-of-the-art lossy compressors in terms of variant calling Recall and Precision for most of the analyzed datasets. Availability and implementation: CALQ is written in C ++ and can be downloaded from https://github.com/voges/calq. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jan Voges, Jörn Ostermann, Mikel Hernaez |
Bioinform. | 2 |
| 2017 | Robust Long-Term Aerial Video Mosaicking by Weighted Feature-Based Global Motion Estimation
Holger Meuel, Stephan Ferenz, Florian Kluger, Jörn Ostermann |
CAIP (1) | 4 |
| 2017 | Differential Gene Expression with Lossy Compression of Quality Scores in RNA-Seq DataabstractHigh-throughput sequencing of RNA molecules has enabled the quantitative analysis of the expression of genes at the expense of storage space and processing power. To help alleviate these problems, lossy compression methods of the quality scores associated to RNA sequence data have recently been proposed, and the evaluation of their impact on downstream analysis is gaining attention. This work presents a first assessment of the impact of lossly compressed quality scores in RNA sequence data on the performance of some of the most recent tools used for differential gene expression. Ana A. Hernandez-Lopez, Jan Voges, Claudio Alberti, Marco Mattavelli, Jörn Ostermann |
DCC | 5 |
| 2017 | Apathy Is the Root of All ExpressionsabstractIn this paper, we present a new statistical model for human faces. Our approach is built upon a tensor factorisation model that allows controlled estimation, morphing and transfer of new facial shapes and expressions. We propose a direct parametrisation and regularisation for person and expression related terms so that the training database is well utilised. In contrast to existing works we are the first to reveal that the expression subspace is star shaped. This stems from the fact that increasing the strength of an expression approximately forms a linear trajectory in the expression subspace, and all these linear trajectories intersect in a single point which corresponds to the point of no expression or the point of apathy. After centring our analysis to this point, we then demonstrate how the dimensionality of the expression subspace can be further reduced by projection pursuit with the help of the fourth-order moment tensor. The results show that our method is able to achieve convincing separation of the person specific and expression subspaces as well as flexible, natural modelling of facial expressions for wide variety of human faces. By the proposed approach, one can morph between different persons and different expressions even if they do not exist in the database. In contrast to the state-of-the-art, the morphing works without causing strong deformations. In the application of expression classification, the results are also better. Stella Graßhof, Hanno Ackermann, Sami S. Brandt, Jörn Ostermann |
FG | 4 |
| 2017 | Visual speech synthesis from 3D mesh sequences driven by combined speech featuresabstractGiven a pre-registered 3D mesh sequence and accompanying phoneme-labeled audio, our system creates an animatable face model and a mapping procedure to produce realistic speech animations for arbitrary speech input. Mapping of speech features to model parameters is done using random forests for regression. We propose a new speech feature based on phonemic labels and acoustic features. The novel feature produces more expressive facial animation and it robustly handles temporal labeling errors. Furthermore, by employing a sliding window approach to feature extraction, the system is easy to train and allows for low-delay synthesis. We show that our novel combination of speech features improves visual speech synthesis. Our findings are confirmed by a subjective user study. Felix Kuhnke, Jörn Ostermann |
ICME | 2 |
| 2017 | Extending HEVC using texture synthesisabstractThe High Efficiency Video Coding (HEVC) standard provides superior coding efficiency compared to its predecessors. Nevertheless, the encoding of complex and thus hardly to predict textures either requires high bit rates or results in low quality of the reconstructed signal. To compensate for this limitation of HEVC, we propose a sophisticated texture synthesis framework which solves multiple lacks of previous texture synthesis approaches. By easing the bit rate cost for synthesizable regions and reallocating the freed bit rate resources to non-synthesizable regions, for high-value soccer content we are able to achieve average BD-rate gains of 21.9% for all-intra, 17.6% for low delay, and 16.3% for random access, respectively, while maintaining the same objective quality for the latter. Subjective tests for the synthesizable regions confirm the objectively measured convincing results. The general applicability of our method is confirmed for other types of content. Bastian Wandt, Thorsten Laude, Yiqun Liu 0003, Bodo Rosenhahn, Jörn Ostermann |
VCIP | 5 |
| 2016 | Predictive Coding of Aligned Next-Generation Sequencing DataabstractDue to novel high-throughput next-generation sequencing technologies, the sequencing of huge amounts of genetic information has become affordable. On account of this flood of data, IT costs have become a major obstacle compared to sequencing costs. High-performance compression of genomic data is required to reduce the storage size and transmission costs. The high coverage inherent in next-generation sequencing technologies produces highly redundant data. This paper describes a compression algorithm for aligned sequence reads. The proposed algorithm combines alignment information to implicitly assemble local parts of the donor genome in order to compress the sequence reads. In contrast to other algorithms, the proposed compressor does not need a reference to encode sequence reads. Compression is performed on-the-fly using solely a sliding window (i.e. a permanently updated short-time memory) as context for the prediction of sequence reads. The algorithm yields compression results on par or better than the state-of-the-art, compressing the data down to 1.9% of the original size at speeds of up to 60 MB/s and with a minute memory consumption of only several kilobytes,fitting in today's level 1 CPU caches. Jan Voges, Marco Munderloh, Jörn Ostermann |
DCC | 3 |
| 2016 | In-loop radial distortion compensation for long-term mosaicing of aerial videosabstractFor the generation of overview panoramic images from aerial surveillance videos, registered video frames are stitched together. Assuming a planar landscape, feature points can be detected and used to estimate a homography. However, if the features are affected by radial distortion, their mapping depends on their position within the frame and the resulting homography becomes inaccurate. As a result, the length of aerial panorama images is typically restricted to several hundred frames. To overcome this issue, we derive a model for the joint estimation of several homographies and one constant radial distortion. Due to the computational complexity of the solution, we propose a fast, iterative algorithm. Based on geometrical constraints, we regularize the projection of a jointly estimated picture group. We present panorama images from uncalibrated aerial videos with more than 1500 frames. Holger Meuel, Stephan Ferenz, Marco Munderloh, Hanno Ackermann, Jörn Ostermann |
ICIP | 5 |
| 2016 | Codec independent region of interest video coding using a joint pre- and postprocessing frameworkabstractFor low bit rate scenarios (video conferencing, aerial surveillance), conventional video coding is unable to meet the small bit rate and high quality requirements. In contrast to that Region of Interest (ROI) coding provides an efficient compression by improving the quality of ROIs at the expense of non-ROIs. We also transmit ROI only, but reconstruct non-ROI from already transmitted content by means of global motion compensation in order to provide a high quality for the full frame. Previous ROI coding systems modified the video codec to control the coding of individual blocks. We propose a codec and ROI detector independent pre- and postprocessing framework instead. This enables the usage of off-the-shelf hard-/software and an easy adaption to the latest video coding technology. Maintaining the performance of subsequent computer vision tasks, we reduce the bit rate by 90-95 % to less than 1 Mbit/s using HEVC for full HDTV videos. Holger Meuel, Marco Munderloh, Florian Kluger, Jörn Ostermann |
ICME | 4 |
| 2016 | Contour-based multidirectional intra coding for HEVCabstractIntra coding is an indispensable part of all video coding systems. However, the intra prediction in the HEVC standard has several limitations with respect to the available amount of information which is exploited for the prediction and the limited number of prediction directions per block. In this paper, we propose a novel contour-based multidirectional intra coding mode which combines computer vision technologies with conventional video coding technologies to overcome these limitations. Structural parts and smooth parts of the video signal are predicted separately. While the structural parts are predicted based on contour detection, parameterization and extrapolation algorithms, the smooth areas are filled by means of low complexity technologies like sample value continuation. Thereby, the coding efficiency is considerably increased, with weighted average BD-rate gains of up to 1.9% compared to HM-16.3. A stand-alone codec implementation achieves average bit rate savings of 29.7% over JPEG and outperforms related work. Thorsten Laude, Jörn Ostermann |
PCS | 2 |
| 2016 | Deep learning-based intra prediction mode decision for HEVCabstractThe High Efficiency Video Coding standard and its screen content coding extension provide superior coding efficiency compared to predecessor standards. However, this coding efficiency is achieved at the expense of very complex encoders. One major complexity driver is the comprehensive rate distortion (RD) optimization. In this paper, we present a deep learning-based encoder control which replaces the conventional RD optimization for the intra prediction mode with deep convolutional neural network (CNN) classifiers. Thereby, we save the RD optimization complexity. Our classifiers operate independently of any encoder decisions and reconstructed sample values. Thus, no additional systematic latency is introduced. Furthermore, the loss in coding efficiency is negligible with an average value of 0.52% over HM-16.6+SCM-5.2. Thorsten Laude, Jörn Ostermann |
PCS | 2 |
| 2016 | Moving object tracking for aerial video coding using linear motion prediction and block matchingabstractRegion of Interest (ROI) coding is a common method for data reduction in scenarios where bandwidth is crucial like in aerial video surveillance from Unmanned Aerial Vehicles (UAVs). In order to save bits, non-ROI areas are typically reduced in quality or not transmitted at all and thus, an accurate ROI classification is mandatory. Moving objects (MOs) are often considered as ROIs and consequently have to be accurately detected onboard. However, common detection approaches either rely on computationally demanding processing which is not available at small UAVs with only limited energy, are model based or cannot provide a sufficient detection precision. While not detected MOs lead to a degraded representation at the decoder, erroneously detected MOs lead to an unnecessary high bit rate. We tackle all these issues utilizing an efficient object proposal computation. Based on a dual-threshold strategy applied to image differences, we propose a linear prediction-supported block matcher. Compared to a simple thresholding approach, it shows superior performance and is robust to threshold tuning. By integrating superpixels into the framework, we further recover the complete shape of the MOs. Finally, an efficient tracking-by-detection system is employed to produce accurate detections from the proposals, thereby recovering missed MOs and denying wrong proposals, making the coding more efficient. We achieve an improved detection precision of up to 76 % compared to a simple difference image-based approach. By using a general ROI coding framework we reduce the bit rate of our test set by 70 % compared to common HEVC. Holger Meuel, Luis Angerstein, Roberto Henschel, Bodo Rosenhahn, Jörn Ostermann |
PCS | 5 |
| 2016 | Impact of hyperspectral image coding on subpixel detectionabstractFour lossy hyperspectral image data coding schemes are compared with regard to their aptitude for subpixel detection use. The coding standards H.265/HEVC and JPEG2000 are investigated with and without a PCA preprocessing. As evaluation criteria, both the ‘Area under Receiver Operation Curve’ as well as the ‘Peak Signal to Noise Ratio’ are calculated. The ‘Area under Reiceiver Operation Curve’ is based on the ‘Spectral Angle Mapper’. Under both criteria, the two coding schemes with PCA preprocessing are the best while the JPEG2000 coding scheme works significantly less efficient. Furthermore, it was shown why the classification is not monotonically improving over increasing data rate. The PCA&HEVC and PCA&JPEG2000 schemes are stable at data rates of 0.1 bit per pixel per band [bpppb] and above while achieving an ‘Area under ROC’ of at least 0.99. If a data link of 0.3 bpppb is available, even the HEVC coding scheme reaches an ‘Area under ROC’ of 0.99 or more. Thus, it depends on the available data link, whether the HEVC coding scheme can be applied or if one of the more complex coding schemes with PCA preprocessing is required. Ulrike Pestel-Schiller, Karsten Vogt, Jörn Ostermann, Wolfgang Gros |
PCS | 3 |
| 2016 | Adaptive Symbol Request Sharing Scheme for Mobile Cooperative Receivers in OFDM SystemsabstractIn this paper we introduce an improvement of the symbol request sharing (SRS) cooperative scheme, namely the adaptive SRS (A-SRS). Both schemes are designed for systems assuming a source and several receivers with one target receiver among them which is denoted as destination. In addition, the source or the receivers are not restricted to be static. These schemes follow a request-answer strategy, in which the destination requests specific information from the remaining receivers. With this strategy the schemes achieve spatial diversity by performing maximum ratio combining (MRC) on selected subcarriers of a coded OFDM-based system. The A-SRS complements its predecessor by adding to it more sophisticated steps in the algorithm toward its implementation on real systems. With these enhancements, the A-SRS scheme preserves the same performance as SRS for hostile source-receiver channels; nevertheless, it provides a notably enhancement for better channels conditions. In terms of BER, for instance, the A-SRS outperforms the SRS scheme by some dB's of gain and it tends to converge more slowly as the number of receivers increases. Furthermore, the A-SRS scheme adapts the length of the cooperation overhead. Therefore, it reaches the highest throughput, which is not the case for the SRS scheme. Yasser Samayoa, Jörn Ostermann |
VTC Fall | 2 |
| 2015 | Stereo mosaicking and 3D-video for singleview HDTV aerial sequences using a low bit rate ROI coding frameworkabstractLow bit rate coding systems for the transmission of high quality aerial surveillance videos captured from UAVs are of high interest. One way to achieve high quality low bit rate video is to assume a planar surface of the earth, which is valid for sequences captured at high flight altitudes. Those systems only transmit the area of the current frame not contained in the previous frames (New Area) and reconstruct the already known areas by means of Global Motion Compensation (GMC) at decoder side. Although the bit rate can be reduced significantly compared to standardized video coders, no reconstruction of stereo video is possible at the decoder since each image pixel is transmitted only once and thus no motion parallax of objects can be observed in the reconstructed video. In this paper we present a coding system for stereo video reconstruction at very low bit rates. On-board the UAV we employ the camera path estimated from the image data to create a second view of a virtual camera. We derive convenient baseline distances and demonstrate the resulting perceptively good stereo impression for different test sequences. Similar to the coding concept introduced above we transmit a second New Area 2 in addition to the New Area already introduced. By doubling the bit rate to about 2Mbit/sfor a reasonable video quality of more than 38 dB, still saving more than 85% BD-rate compared to common HEVC coding, we are able to reconstruct a full HDTV (30 fps) stereo video at the decoder. Holger Meuel, Marco Munderloh, Jörn Ostermann |
AVSS | 3 |
| 2015 | Feature Evaluation with High-Resolution Images
Kai Cordes, Lukas Grundmann, Jörn Ostermann |
CAIP (1) | 3 |
| 2015 | Globally optimized dynamic bit-allocation strategy for subband ADPCM-based low delay audio codingabstractThis paper presents an extension and global optimization of a bit-allocation strategy for use in a low delay subband ADPCM-based audio coding scheme. The framework employed for the assignment of bits to subband quantizers is based on a subband level estimation of the signals used for the prediction error normalization. It is modified to reduce switching artifacts and extended to contain a mapping to predefined allocation presets. Since our bit-allocation scheme as well as the subband coding involve several partially interacting parameters that are hard to adjust manually, a framework for their global optimization is presented. Experiments and their results show that in our coding scheme a significant improvement of audio quality without large signaling overhead can be achieved by the use of the modified bit-allocation. Furthermore, the global optimization allows for an additional gain in audio quality compared to manually adjusted parameters. Stephan Preihs, Jörn Ostermann |
ICASSP | 2 |
| 2015 | Continuous Pose Estimation with a Spatial Ensemble of Fisher RegressorsabstractIn this paper, we treat the problem of continuous pose estimation for object categories as a regression problem on the basis of only 2D training information. While regression is a natural framework for continuous problems, regression methods so far achieved inferior results with respect to 3D-based and 2D-based classification-and-refinement approaches. This may be attributed to their weakness to high intra-class variability as well as to noisy matching procedures and lack of geometrical constraints. We propose to apply regression to Fisher-encoded vectors computed from large cells by learning an array of Fisher regressors. Fisher encoding makes our algorithm flexible to variations in class appearance, while the array structure permits to indirectly introduce spatial context information in the approach. We formulate our problem as a MAP inference problem, where the likelihood function is composed of a generative term based on the prediction error generated by the ensemble of Fisher regressors as well as a discriminative term based on SVM classifiers. We test our algorithm on three publicly available datasets that envisage several difficulties, such as high intra-class variability, truncations, occlusions, and motion blur, obtaining state-of-the-art results. Michele Fenzi, Laura Leal-Taixé, Jörn Ostermann, Tinne Tuytelaars |
ICCV | 3 |
| 2015 | Copy mode for static screen content coding with HEVCabstractScreen content largely consists of static parts, e.g. static background. However, none of the available screen content coding tools fully employs this characteristic. In this paper we present the copy mode, a new coding mode specifically aiming at increased coding efficiency for static screen content. The basic principle of the copy mode is the direct copy of the collocated block from the reference frame. Mean weighted BD-Rate gains of 2.4% are achieved for JCT-VC test sequences compared to SCM-2.0. For sequences containing lots of static background, coding gains as high as 7.6% are observed. The new coding mode is further enhanced by several encoder optimizations, among them an early skip mechanism. Thereby, the encoder runtime complexity is reduced by up to 39%. Thorsten Laude, Jörn Ostermann |
ICIP | 2 |
| 2015 | Fast motion blur compensation in HEVC using fixed-length filterabstractMotion compensation is one of the most important elements in modern hybrid video coders. It utilizes temporal information to predict the current block and reduces thereby the redundancy of a video. The prediction accuracy depends on the similarity between the reference block and the current block. It is decreased by varying motion blur caused by the acceleration of the camera or certain objects in a scene. Thus, we employ fixed-length filters to compensate varying motion blur in hybrid video coding. While former approaches needed additional signaling for blurring filters or a second motion estimation, our algorithm derives the blurring filter only based on the motion vector and needs only one motion estimation. We implemented our approach in the High Efficiency Video Coding (HEVC) reference software HM-13.0. Compared to the reference HM-13.0, we gain 2.54% in terms of BD-Rate in average for JCT-VC test sequences and 4.51% for self-recorded sequence containing lots of varying motion blur, with limited increase in coding time. Yiqun Liu 0003, Jörn Ostermann |
ICIP | 2 |
| 2015 | Fast label propagation for real-time superpixels for video contentabstractMany recent superpixel algorithms for video content rely on dense optical flow vectors to propagate segmentation results from one frame to the next. In this paper, we assess the impact of the optical flow quality on the over-segmentation quality. Our evaluation shows that it is indispensable for videos with large object displacement and camera motion. But due to the high computational costs high-quality, dense optical flow is not suitable for real-time applications. Therefore, we propose a fast propagation scheme that is based on sparse feature tracking and mesh-based image warping. In a thorough evaluation, we compare our proposed scheme to the results of other state-of-the-art propagation methods using established benchmarks. The results show that our method speeds up the propagation process by a factor of 100 while producing a comparable segmentation quality. Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann |
ICIP | 4 |
| 2015 | Report on the evaluation of current and future image compression technologiesabstractThis document reports the conclusions, comments and recommendations resulting from the final panel discussion which happened at the Session on Evaluation of Current and Future Image Compression Technologies held at the Picture Coding Symposium, Cairns, Australia, in June 2015. Fernando Pereira 0001, Ralf Schaefer, Touradj Ebrahimi, Jörn Ostermann, Edward J. Delp |
PCS | 4 |
| 2015 | Motion blur compensation in HEVC using fixed-length adaptive filterabstractMotion compensation is one of the most important elements in modern hybrid video coders. It utilizes temporal information to predict the current block and reduces thereby the redundancy of a video. The accuracy of prediction depends on the similarity of the content between the reference block and the current block. With the change of velocity of the camera or certain objects in a scene, which is typically expected in action and sports movies, motion blur varies from frame to frame leading to a reduced prediction accuracy. We employ fixed-length filters to compensate varying motion blur in hybrid video coding. While former approaches needed additional signaling for blurring filters, our filter is derived only based on the motion vector. We implemented our approach in the High Efficiency Video Coding (HEVC) reference software HM 13.0. Compared to the reference we gain 2.15% in terms of BD-Rate in average for JCT-VC test sequences and 4.43% for self-recorded sequences containing lots of varying motion blur. Yiqun Liu 0003, Jörn Ostermann |
PCS | 3 |
| 2015 | Pose Estimation of Object Categories in Videos Using Linear ProgrammingabstractIn this paper we propose a method to consistently recover the pose of an object from a known class in a video sequence. As individual poses estimated from monocular images are rather noisy, we optimally aggregate pose evidence over all video frames. We construct a graph where nodes are values sampled from the pose posterior distributions computed by a continuous pose estimator in each frame of the sequence. We then find the globally optimum pose path through the graph that best explains the pose evidence for the whole sequence. As a result, we recover the correct object orientation at each frame even if single-frame pose evidence is sometimes inaccurate. We evaluate our approach on two publicly available car datasets, which encompass busy street scenarios and car races with significant changes in car orientation, blur and occlusions. We show that our method outperforms state-of-the-art approaches reducing the error by 40% on the challenging KITTI dataset. Michele Fenzi, Laura Leal-Taixé, Konrad Schindler, Jörn Ostermann |
WACV | 4 |
| 2014 | Superpixels for Video Content Using a Contour-Based EM Optimization
Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann |
ACCV (4) | 4 |
| 2014 | Localization accuracy of interest point detectors with different scale space representationsabstractThe detection of scale invariant image features is a fundamental task for computer vision applications like object recognition or re-identification. Features are localized by computing extrema of the gradients in the Laplacian of Gaussian (LoG) scale space. The most popular detector for scale invariant features is the SIFT detector which uses the Difference of Gaussians (DoG) pyramid as an approximation of the LoG. Recently, the alternative interest point (ALP) detector demonstrated its strength in fast computation on highly parallel architectures like the GPU. It uses the LoG scale space representation for the localization of interest points. This paper evaluates the localization accuracy of ALP in comparison to SIFT. By using synthetic images, it is demonstrated that both localization approaches show a systematic error which is dependent on the subpixel position of the feature. The error increases with the scale of the detected feature. However, using the LoG instead of the DoG representation reduces the maximum systematic error by 77 %. For the evaluation with natural images, benchmark data sets are used. The repeatability criterion evaluates the accuracy of the detectors. The LoG based detector results in up to 16 % higher repeatability. The comparisons are completed with a reference feature localization which uses a signal based approach for the gradient approximation. Based on this approach, a new feature selection criterion is proposed. Kai Cordes, Bodo Rosenhahn, Jörn Ostermann |
AVSS | 3 |
| 2014 | ASEV - Automatic situation assessment for event-driven video analysisabstractMany complex maneuvers involving aircraft, vehicles and persons are carried out at airport aprons. Manual video surveillance used for safety and security purposes is inefficient and privacy protection must be guaranteed. In this paper, we propose a system named ASEV that automatically assesses situations for airport surveillance. It combines four main components: a low-level image processing unit based on a new hardware implementation to extract features in real time, a high-level image processing unit for scene analysis, a real-time inference engine for scene understanding, and a data protection stage for log encryption. In addition, four often neglected aspects are successfully addressed: two-way communication between system and operator, power consumption, monitored people privacy and operator activity control. Extensive evaluation at a real airport shows that the proposed system improves the operator performance with sound and visual alerts based on the automatic assessment of various events. Michele Fenzi, Jörn Ostermann, Nico Mentzer, Guillermo Payá-Vayá, Holger Blume, Tu Ngoc Nguyen, Thomas Risse 0001 |
AVSS | 2 |
| 2014 | Embedding Geometry in Generative Models for Pose Estimation of Object Categories
Michele Fenzi, Jörn Ostermann |
BMVC | 2 |
| 2014 | Improved Inter-Layer Prediction for the Scalable Extensions of HEVCabstractSummary form only given. Upon the completion of the single-layer H.265/HEVC, scalable extensions of the H.265/HEVC standard, called Scalable High Efficiency Video Coding (SHVC), are currently under development. Compared to the simulcast solution that simply compresses each layer separately, SHVC offers higher coding efficiency by means of inter-layer prediction which is implemented by inserting inter-layer reference (ILR) pictures generated from reconstructed base layer (BL) pictures into the enhancement layer (EL) decoded picture buffer (DPB) for motion-compensated prediction of the collocated pictures in the EL. If the EL has a higher resolution than that of the BL, the reconstructed BL pictures need to be up-sampled to form the ILR pictures. Given that the ILR picture is generated based on the reconstructed BL picture, its suitability for an efficient inter-layer prediction may be limited due to the following reasons. Firstly, quantization is usually applied when coding the BL pictures. Quantization causes the BL reconstructed texture to contain undesired coding artifacts, such as blocking artifacts, ringing artifacts, and color artifacts. Secondly, in case of spatial scalability, a down-sampling process is used to create the BL pictures. To reduce aliasing, the high frequency information in the video signal is typically removed by the down-sampling process. As a result, the texture information in the ILR picture lacks certain high frequency information. In contrast to the ILR picture, the EL temporal reference pictures contain plentiful high frequency information, which could be extracted to enhance the quality of the ILR picture. To further improve the efficiency of inter-layer prediction, a low pass filter may be applied to the ILR picture to alleviate the quantization noise introduced by the BL coding process. In this paper, an ILR enhancement method is proposed to improve the quality of the ILR picture by combining the high frequency information extracted from the EL temporal reference pictures together with the low frequency information extracted from the ILR picture. Experimental results show that the proposed method can significantly increase the ILR efficiency for EL coding, under the Common Test Condition of SHVC, which defines a number of temporal prediction structures called Random Access (RA), Low-delay B (LD-B) and Low-delay P (LD-P), on average the proposed method provides {Y, U, V} BD-rate (BL+EL) gains of {2.0%, 7.1%, 8.2%}, {2.2%, 6.7%, 7.6%} and {4.0%, 7.4%, 8.4%} for RA, LD-B, and LD-P, respectively, in comparison to the performance of the SHVC reference software SHM-2.0. Thorsten Laude, Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Jörn Ostermann |
DCC | 6 |
| 2014 | Scalable extension of HEVC using enhanced inter-layer predictionabstractIn Scalable High Efficiency Video Coding (SHVC), inter-layer prediction efficiency may be degraded because much high frequency information can be removed during: 1) the down-sampling/up-sampling process and, 2) the base layer coding/quantization process. In this paper, we present a method to enhance the quality of the inter-layer reference (ILR) picture by combining the high frequency information from enhancement layer temporal reference pictures with the low frequency information from the up-sampled base layer picture. Experimental results show that on average 3.9% weighted BD-rate gain is achieved compared to SHM-2.0 under SHVC common test conditions. Thorsten Laude, Xiaoyu Xiu, Yuwen He, Yan Ye 0003, Jörn Ostermann |
ICIP | 6 |
| 2014 | Plant classification system for crop /weed discrimination without segmentationabstractThis paper proposes a machine vision approach for plant classification without segmentation and its application in agriculture. Our system can discriminate crop and weed plants growing in commercial fields where crop and weed grow close together and handles overlap between plants. Automated crop / weed discrimination enables weed control strategies with specific treatment of weeds to save cost and mitigate environmental impact. Instead of segmenting the image into individual leaves or plants, we use a Random Forest classifier to estimate crop/weed certainty at sparse pixel positions based on features extracted from a large overlapping neighborhood. These individual sparse results are spatially smoothed using a Markov Random Field and continuous crop/weed regions are inferred in full image resolution through interpolation. We evaluate our approach using a dataset of images captured in an organic carrot farm with an autonomous field robot under field conditions. Applying the plant classification system to images from our dataset and performing cross-validation in a leave one out scheme yields an average classification accuracy of 93.8 %. Sebastian Haug, Andreas Michaels, Peter Biber, Jörn Ostermann |
WACV | 4 |
| 2014 | Estimating layout of cluttered indoor scenes using trajectory-based priors
Muhammad Shoaib 0007, Michael Ying Yang, Bodo Rosenhahn, Jörn Ostermann |
Image Vis. Comput. | 4 |
| 2013 | Superpixel-based segmentation of moving objects for low bitrate ROI coding systemsabstractA major constraint for mobile surveillance systems (e.g. UAV-mounted) is the limited available bandwidth. Thus, ROI-based video coding, where only important image parts are transmitted with high data rate, seems to be a suitable solution for these applications. The challenging part in these systems is to reliably detect the regions of interest (ROI) and to select the corresponding macroblocks to be encoded in high quality. Recently, difference image-based moving object detectors have been used for this purpose. But as these pixel-based approaches fail for homogeneous parts of moving objects, which leads to artifacts in the decoded video stream, this paper proposes the use of a superpixel-based macroblock selection. It is shown that the proposed approach significantly reduces artifacts by raising the moving objects detection rate from 7 % to over 96 % while increasing the data rate by only 20% to 1.5-4Mbit/s (full HD resolution, 30 Hz) still meeting the given constraints. Moreover, other simple approaches (dilation) as well as state-of-the-art GraphCut-based approaches are outperformed. Holger Meuel, Matthias Reso, Jörn Jachalsky, Jörn Ostermann |
AVSS | 4 |
| 2013 | High-Resolution Feature Evaluation Benchmark
Kai Cordes, Bodo Rosenhahn, Jörn Ostermann |
CAIP (1) | 3 |
| 2013 | Class Generative Models Based on Feature Regression for Pose Estimation of Object CategoriesabstractIn this paper, we propose a method for learning a class representation that can return a continuous value for the pose of an unknown class instance using only 2D data and weak 3D labeling information. Our method is based on generative feature models, i.e., regression functions learned from local descriptors of the same patch collected under different viewpoints. The individual generative models are then clustered in order to create class generative models which form the class representation. At run-time, the pose of the query image is estimated in a maximum a posteriori fashion by combining the regression functions belonging to the matching clusters. We evaluate our approach on the EPFL car dataset and the Pointing'04 face dataset. Experimental results show that our method outperforms by 10% the state-of-the-art in the first dataset and by 9% in the second. Michele Fenzi, Laura Leal-Taixé, Bodo Rosenhahn, Jörn Ostermann |
CVPR | 4 |
| 2013 | Temporally Consistent SuperpixelsabstractSuper pixel algorithms represent a very useful and increasingly popular preprocessing step for a wide range of computer vision applications, as they offer the potential to boost efficiency and effectiveness. In this regards, this paper presents a highly competitive approach for temporally consistent super pixels for video content. The approach is based on energy-minimizing clustering utilizing a novel hybrid clustering strategy for a multi-dimensional feature space working in a global color subspace and local spatial subspaces. Moreover, a new contour evolution based strategy is introduced to ensure spatial coherency of the generated super pixels. For a thorough evaluation the proposed approach is compared to state of the art super voxel algorithms using established benchmarks and shows a superior performance. Matthias Reso, Jörn Jachalsky, Bodo Rosenhahn, Jörn Ostermann |
ICCV | 4 |
| 2013 | Inter-layer intra mode coding for the scalable extension of HEVCabstractIn this paper, two inter-layer intra mode coding tools are proposed for the all intra spatial scalable extension of High Efficiency Video Coding (HEVC). The proposed methods make use of the correlations between intra prediction mode information in base layer and enhancement layer to improve the intra mode coding in enhancement layer. The first approach directly utilizes the intra prediction mode from the corresponding prediction unit (PU) in base layer as the intra prediction mode of the enhancement layer PU, while the second approach uses the intra prediction mode of the corresponding base layer PU as an additional most probable mode candidate. Dyadic all intra spatial scalability with 2 layers is tested in our experiments. Compared to enhancement layer only, the experimental results show a maximal 1.2% luma BD-rate reduction for the test sequences relative to inter-layer intra prediction. Zhijie Zhao, Junyong Si, Jörn Ostermann |
ISCAS | 3 |
| 2013 | Motion blur compensation in scalable HEVC hybrid video codingabstractOne main element of modern hybrid video coders consists of motion compensated prediction. It employs spatial or temporal neighborhood to predict the current sample or block of samples, respectively. The quality of motion compensated prediction largely depends on the similarity of the reference picture block used for prediction and the current picture block. In case of varying blur in the scene, e.g. caused by accelerated motion between the camera and objects in the focal plane, the picture prediction is degraded. Since motion blur is a common characteristic in several application scenarios like action and sport movies we suggest the in-loop compensation of motion blur in hybrid video coding. Former approaches applied motion blur compensation in single layer coding with the drawback of needing additional signaling. In contrast to that we employ a scalable video coding framework. Thus, we can derive strength as well as the direction of motion of any block for the high quality enhancement layer by base-layer information. Hence, there is no additional signaling necessary neither for predefined filters nor for current filter coefficients. We implemented our approach in a scalable extension of the High Efficiency Video Coding (HEVC) reference software HM 8.1 and are able to provide up to 1% BD-Rate gain in the enhancement layer compared to the reference at the same PSNR-quality for JCT-VC test sequences and up to 2.5% for self-recorded sequences containing lots of varying motion blur. Thorsten Laude, Holger Meuel, Yiqun Liu 0003, Jörn Ostermann |
PCS | 4 |
| 2013 | Optical flow cluster filtering for ROI codingabstractCurrent Moving Object Detectors in airborne Region of Interest (ROI) coding systems for police surveillance applications used on-board of UAVs are often based on Global Motion Estimation (GME) techniques. Since in these scenarios the camera is moving, simple background removal approaches cannot be applied without a Global Motion Compensation (GMC). Common GMC algorithms assume the ground to be planar, allowing the pixels of the previous frame to be motion compensated into the current frame by applying a projective transformation. The difference image between the compensated frame and the current frame emphasis regions containing possible motion. Such moving object detectors are great in terms of run-time efficiency but are known to lack in terms of accuracy - especially for unstructured regions of moving objects - as well as the robustness against noise. Superpixel segmentation was recently proposed to overcome the issue of the imprecise region cuts given by the difference image. It provides a greatly improved true positive detection rate, but unintentionally also increases the area of false positives. This paper proposes the use of a mesh-based GME and GMC to detect the moving object regions wherein a cluster filter eliminates errors in the optical flow by assuming a smooth vector field as the global motion model. In doing so we improve the coding efficiency of the fully automatic ROI coding system by more than 24% for moving object areas conserving the detection benefits of the integration of superpixel segmentation. Holger Meuel, Marco Munderloh, Matthias Reso, Jörn Ostermann |
PCS | 4 |
| 2012 | Learning Object Appearance from Occlusions Using Structure and Motion Recovery
Kai Cordes, Björn Scheuermann 0002, Bodo Rosenhahn, Jörn Ostermann |
ACCV (3) | 4 |
| 2012 | Multi-scale Clustering of Frame-to-Frame Correspondences for Motion Segmentation
Ralf Dragon, Bodo Rosenhahn, Jörn Ostermann |
ECCV (2) | 3 |
| 2012 | Hierarchical Bayer-pattern based background subtraction for low resource devicesabstractAutomatic visual monitoring is one of the key research areas in recent years. Segmentation of moving objects is an important step toward monitoring and activity analysis. This paper proposes a real-time moving object segmentation technique using Bayer-pattern images. The proposed method takes care of low resource availability at embedded system. In order to reduce memory requirements, we model the background scene in single channel Bayer-pattern domain. In classification phase, we first define an approximate foreground using block-level information. The Approximate foreground is then refined using in-loop interpolated pixel-level RGB information. Refinement process also takes care of illumination changes. Our proposed method achieves almost the same accuracy as state-of-the-art RGB based pixel-level background subtraction methods, while using lower computational resources. Experimental results show that the proposed has a true positive rate of 87% and false positive rate 1.8% using low resources, quite suitable for implementation in real-time embedded systems that can be used for monitoring. Muhammad Shoaib 0007, Tobias Elbrandt, Evgeny Zaretskiy, Jörn Ostermann |
ISCAS | 4 |
| 2012 | Comparison between multiple description coding and forward error correction for scalable video coding with different burst lengthsabstractIn this paper, a comparison of a scalable multiple description coding (MDC) scheme and the forward error correction (FEC) coding using Raptor code for the scalable extension of H.264/AVC (SVC) is presented. Unequal and equal protection of the base and enhancement layer with different amount of redundancy at the average burst lengths of 2, 5 and 20 are considered. The experimental results show that scalable MDC is generally preferable at the low redundancy rate and long average burst length, while FEC using Raptor code is favorite in case of high redundancy rate and the channel with short average burst length for the resilient delivery of scalable video. Zhijie Zhao, Doug Young Suh, Jörn Ostermann |
MMSP | 4 |
| 2012 | Analysis of coding tools and improvement of text readability for screen contentabstractCurrent video coding standards perform well for video sequences captured by a real camera. The aperture of the camera's optical system smooths the content and attenuates higher frequencies. New application scenarios, enabled by the growing number of high bit rate internet gateways, however, make it necessary to take a closer look at the efficiency of such standards in handling artificial content. Remote desktop applications for example often include text parts. As a consequence, these content types contain sharp edges or high frequencies, which are considered less important in natural video and are therefore treated less carefully. The frequent result is an increased occurrence of artefacts or the loss of information that is actually important to the user. This paper gives an analysis of such artificially created video sequences, evaluates the performance of current coding tools for this type of content and proposes a simple, yet effective way to maintain readability of text within video material using only well considered encoder control and without the need of large additional modules. Holger Meuel, Julia Schmidt, Marco Munderloh, Jörn Ostermann |
PCS | 4 |
| 2011 | Realistic head motion synthesis for an image-based talking headabstractIn this paper, we present a novel approach to add flexible head motions to talking heads. First, head motion patterns are collected from original recordings. These head motion patterns are recorded video segments with different head motions, like nod and shake. The head motion is synthesized by selecting and concatenating appropriate head motion patterns according to the input text or head motion tags. In order to join these patterns, optical flow based morphing is used to smooth transitions without introducing noticeable discontinuities. Experimental results show that head motion synthesis is realistic, and animations with flexible head motions are rated with a higher average mean opinion score than the ones with repeated head motions. Kang Liu 0002, Jörn Ostermann |
FG | 2 |
| 2011 | Realistic head motion synthesis for an image-based talking headabstractIn this paper, we present a novel approach to add flexible head motions to talking heads. First, head motion patterns are collected from original recordings. These head motion patterns are recorded video segments with different head motions, like nod and shake. The head motion is synthesized by selecting and concatenating appropriate head motion patterns according to the input text or head motion tags. In order to join these patterns, optical flow based morphing is used to smooth transitions without introducing noticeable discontinuities. Experimental results show that head motion synthesis is realistic, and animations with flexible head motions are rated with a higher average mean opinion score than the ones with repeated head motions. Kang Liu 0002, Jörn Ostermann |
FG | 2 |
| 2011 | Realistic facial expression synthesis for an image-based talking headabstractThis paper presents an image-based talking head system that is able to synthesize realistic facial expressions accompanying speech, given arbitrary text input and control tags of facial expression. As an example of facial expression primitives, smile is used. First, three types of videos are recorded: a performer speaking without any expressions, smiling while speaking, and smiling after speaking. By analyzing the recorded audiovisual data, an expressive database is built and contains normalized neutral mouth images and smiling mouth images, as well as their associated features and expressive labels. The expressive talking head is synthesized by an unit selection algorithm, which selects and concatenates appropriate mouth image segments from the expressive database. Experimental results show that the smiles of talking heads are as realistic as the real ones objectively, and the viewers cannot distinguish the real smiles from the synthesized ones. Kang Liu 0002, Jörn Ostermann |
ICME | 2 |
| 2011 | Prediction of DCT coefficients considering motion compensation error distributionsabstractCurrent video coding techniques use a Discrete Cosine Transform (DCT) to reduce spatial correlations within the motion estimation residual. Often the correlation cannot be completely eliminated leaving the transform coefficients statistically dependent. The presented paper proposes a method to predict these coefficients on a block level by using the distribution of the prediction error variance to improve coding efficiency. First experiments lead to a reduction in bit rate by 1.83% when compared to the standard JM 17.2 implementation results. Julia Schmidt, Bernd Edler, Jörn Ostermann |
VCIP | 3 |
| 2010 | NF-Features - No-Feature-Features for Representing Non-textured Regions
Ralf Dragon, Muhammad Shoaib 0007, Bodo Rosenhahn, Jörn Ostermann |
ECCV (2) | 4 |
| 2010 | Block size dependent error model for motion compensationabstractCurrent video coding standards use block-based motion estimation and compensation algorithms to exploit dependencies between consecutive frames. It is a well-known fact that decreasing the block size reduces the motion-compensated frame difference, and thus reduces the data rate. However, no theoretical evaluations are available to model this relation. This paper derives a model for the prediction error variance of block-based motion compensation algorithms with respect to the block size. It is shown that the variance of the displaced frame difference of a block can be modelled with the pixel position and only three additional parameters. It can be observed that the variance increases almost linearly with the block size. Sven Klomp, Marco Munderloh, Jörn Ostermann |
ICIP | 3 |
| 2010 | Mesh-based decoder-side motion estimationabstractCurrent video coding standards like H.264|AVC perform a block-based motion estimation and compensation at the encoder to exploit temporal dependencies between consecutive frames. Current research proved that motion compensation can also be done profitably at the decoder. In this so-called decoder-side motion estimation (DSME), frames are interpolated at the decoder and inserted into the reference buffer as additional information for prediction. To gather the motion information for compensation, the current approach is based on a block matching algorithm estimating one motion vector for each block of the frame to be interpolated. Therefore, only translational movement is compensated. We evaluate the applicability of a mesh-based motion compensation to DSME which models affine motion in each patch of the mesh and is able to compensate camera zooming or panning, object rotation, or object deformation. Marco Munderloh, Sven Klomp, Jörn Ostermann |
ICIP | 3 |
| 2010 | Video streaming using standard-compatible scalable multiple description coding based on SVCabstractIn this paper, a scalable multiple description video coding method based on the scalable video extension of H.264/AVC (SVC) is presented. The proposed scheme employs a standard SVC encoder to generate several bit streams which have different bit rates. The spatial enhancement layers of these bit streams are quantized in an unbalanced fashion, while the spatial base layers are coded with the same quantization step size. Balanced multiple scalable descriptions are generated by mixing the pre-encoded scalable bit streams. Each description alone is totally decodable by a standard SVC decoder. A preprocessor before a SVC decoder is employed to extract the packets from the highest quality bit stream. The experimental results show that our proposed scheme achieves improvements of 3.5dB and 3.9dB when compared to the methods based on spatial downspamling and temporal downsampling at 10% packet loss rate. Zhijie Zhao, Jörn Ostermann |
ICIP | 2 |
| 2010 | Prediction Framework for Statistical Respiratory Motion Modeling
Tobias Klinder, Cristian Lorenz, Jörn Ostermann |
MICCAI (3) | 3 |
| 2010 | Subjective evaluation of scalable video coding for content distributionabstractThis paper investigates the influence of the combination of the scalability parameters in scalable video coding (SVC) schemes on the subjective visual quality. We aim at providing guidelines for an adaptation strategy of SVC that can select the optimal scalability options for resource-constrained networks. Extensive subjective tests are conducted by using two different scalable video codecs and high definition contents. The results are analyzed with respect to five dimensions, namely, codec, content, spatial resolution, temporal resolution, and frame quality. Jong-Seok Lee, Francesca De Simone, Naeem Ramzan, Zhijie Zhao, Engin Kurutepe, Thomas Sikora, Jörn Ostermann, Ebroul Izquierdo, Touradj Ebrahimi |
ACM Multimedia | 7 |
| 2010 | Decoder-side hierarchical motion estimation for dense vector fieldsabstractCurrent video coding standards perform motion estimation at the encoder to predict frames prior to coding them. Since the decoder does not possess the source frames, the estimated motion vectors have to be transmitted as additional side information. Recent research revealed that the data rate can be reduced by performing an additional motion estimation at the decoder. As only already decoded data is used, no additional data has to be transmitted. This paper addresses an improved hierarchical motion estimation algorithm to be used in a decoder-side motion estimation system. A special motion vector latching is used to be more robust for very small block sizes and to better adapt to object borders. With this technique, a dense motion vector field is estimated which reduces the rate by 6.9% in average compared to H.264 / AVC at the same quality. Sven Klomp, Marco Munderloh, Jörn Ostermann |
PCS | 3 |
| 2010 | 3D information codingabstractThere are several technologies available for the transmission and presentation of 3D information to the human user. The appropriate selection of technology depends to a large extent on the application as well as on the maturity of the required technology. 3D video was established in niche markets, including professional applications (e.g., scientific visualization) and entertainment (IMAX cinemas, 3D gaming) some time ago [1] and since 2008, it became main stream in digital cinemas with the consumer market getting ready for 3D video by introducing stereo TVs and related technology. Depending on the application, different presentations and data formats are required. For scientific visualization, 3D data formats are used. The 3D data is rendered at the server or it is transmitted to the receiver or client which renders the appropriate view as determined by the user. The correct display of the objects on monoscopic and stereoscopic displays is possible. 3D games typically require the 3D data at the client in order to enable low-latency interaction with the data. The most common 3D data sets consist of 3D geometry and 2D texture data. In case a 3D object is recorded from all allowed viewing angles, this dataset of images can be used for image-based rendering where always one of the recorded images is shown on the display. Image-based rendering may be combined with 3D geometry and animation[2]. The estimation of 3D geometry of natural scenes is a challenging task [3]. Therefore most applications rely on synthetic or manually created 3D data. For presentation of 3D movies, stereo displays requiring glasses to separate left and right views as well as autostereoscopic displays requiring no glasses exist. Independent of the viewer position and orientation, a stereo display presents one view to the left eye and one view to the right eye. These displays just need to receive two video streams typically recorded with a stereo camera. MPEG developed video coding standards to support these displays. The latest standard is MVC, an extension of AVC for coding of stereo sequences using frame reordering [4]. Due to the sudden popularity of stereoscopic video, service providers face the challenge of transmitting stereo using the existing video distribution and presentation infrastructure. One solution is referred to as Frame-packing (FP) because both video frames are copied typically side by side into one frame prior to encoding. Hence, the resulting stereo sequence has only half of the regular resolution. The set-top box decodes FP into one regular frame and signals FP to the display which upconverts each half of the decoded frame to full size for the left and right view. In MPEG-2 FP is signaled as frame packing using private data (USA, Japan) or a newly developed extension in the MPEG-2 Transport Stream. AVC, which is mainly used for HDTV distribution, uses SEI messages to signal FP. MPEG is currently working on an extension to AVC that enhances FP to full resolution stereo while still keeping the compatibility to AVC with FP. In the real world, the view of a scene depends on the head position and orientation of the viewer. Hence, the images presented to the two eyes should be changed depending on the eye position. To a limited extend, auto-stereoscopic displays support this feature by displaying simultaneously several views with a lens in front of the screen assuring that always two appropriate views are visible to the eyes of a viewer [4]. Starting with subjective tests in 2011, MPEG targets this 3DV application by coding several video streams and a depth map of the scene enabling the display to create the appropriate views by means of a view synthesis algorithm. The main challenges are the reliable estimation of a depth map, which may be manually supported for non-real-time applications, as well as the view-synthesis algorithm. Jörn Ostermann |
PCS | 1 |
| 2010 | Error concealment in the network abstraction layer for medium grain scalability of SVCabstractThis paper presents an error concealment method for medium grain scalability (MGS) of the scalable extension of H.264/AVC. Instead of reconstructing a lost frame using a typically compute-intensive error concealment method, in this method, a light-weight decoder preprocessor generates a valid bit stream from the available network abstraction layer (NAL) units in case that MGS layers are used. NAL unit header and slice header are parsed to detect the loss of a NAL unit from MGS layers. Modifications are applied to NAL unit header or slice header if some NAL units of MGS layers are lost. Hence, the number of additional packets that a decoder has to discard as a result of a packet loss is minimized. The proposed method requires low computational cost which is suitable for real-time video streaming. The experimental results demonstrate that the proposed method achieves an improvement of 2.16 dB compared to the NAL unit removal method at 5% packet loss rate. The proposed method can work complementary with other error concealment schemes for a whole-frame loss or serve as a preprocessor to a standard decoder. Zhijie Zhao, Jörn Ostermann |
VCIP | 2 |
| 2010 | SPC: Fast and Efficient Scalable Predictive Coding of Animated MeshesabstractAbstract Animated meshes are often represented by a sequence of static meshes with constant connectivity. Due to their frame‐based representation they usually occupy a vast amount of bandwidth or disk space. We present a fast and efficient scalable predictive coding (SPC) scheme for frame‐based representations of animated meshes. SPC decomposes animated meshes in spatial and temporal layers which are efficiently encoded in one pass through the animation. Coding is performed in a streamable and scalable fashion. Dependencies between neighbouring spatial and temporal layers are predictively exploited using the already encoded spatio‐temporal neighbourhood. Prediction is performed in the space of rotation‐invariant coordinates compensating local rigid motion. SPC supports spatial and temporal scalability, and it enables efficient compression as well as fast encoding and decoding. Parts of SPC were adopted in the MPEG‐4 FAMC standard. However, SPC significantly outperforms the streaming mode of FAMC with coding gains of over 33%, while in comparison to the scalable FAMC, SPC achieves coding gains of up to 15%. SPC has the additional advantage over FAMC of achieving real‐time encoding and decoding rates while having only low memory requirements. Compared to some other non‐scalable state‐of‐the‐art approaches, SPC shows superior compression performance with gains of over 16% in bit‐rate. Nikolce Stefanoski, Jörn Ostermann |
Comput. Graph. Forum | 2 |
| 2010 | Video-realistic image-based eye animation via statistically driven state machines
Axel Weissenfeld, Kang Liu 0002, Jörn Ostermann |
Vis. Comput. | 3 |
| 2009 | Minimized Database of Unit Selection in Visual Speech Synthesis without Loss of Naturalness
Kang Liu 0002, Jörn Ostermann |
CAIP | 2 |
| 2009 | Shadow detection for moving humans using gradient-based background subtractionabstractCast shadows cause serious problems in the functionality of vision-based applications, such as video surveillance, traffic monitoring and various other applications. Accurate detection and removal of cast shadows is a challenging task. Common shadow detection techniques normally use color information, which is not a reliable base in every scenario. This paper presents a novel scheme for real time detection of cast shadows using contour like structures of objects, which are obtained by gradient-based background subtraction. The scheme does not use any color information. Two basic rules are followed for shadow detection. The first rule is that shadows do not change the texture of the background. The second rule is a cast shadow lies outside the boundary of an object and has a relatively small common boundary with the object. Experimental results show the performance of the proposed scheme. Objective evaluation shows that the algorithm classifies 90 percent of the pixels of the objects and their shadow correctly. Muhammad Shoaib 0007, Ralf Dragon, Jörn Ostermann |
ICASSP | 3 |
| 2009 | Decoder-side Block Motion Estimation for H.264 / MPEG-4 AVC based Video CodingabstractIn video coding standards like H.264 / MPEG-4 AVC, the encoder performs motion estimation in order to utilise temporal dependencies within a sequence. In addition to the rate of the residue, the encoder has to allocate bits for motion vectors required to compensate the motion at the decoder. This bit rate increases for smaller block sizes, since more motion vectors need to be transmitted. Therefore, motion compensation using dense motion vector field is not feasible for such an architecture. This paper proposes to estimate motion for coding of B frames at the decoder. Using this decoder-side motion estimation, the transmission of the motion vectors is not necessary and the bit rate is reduced. Furthermore, prediction quality is higher in many cases resulting in a coding gain of up to 1.7 dB at low bit rates and 0.2 dB at higher bit rates. Sven Klomp, Marco Munderloh, Yuri Vatis, Jörn Ostermann |
ISCAS | 4 |
| 2009 | Low complexity multiple description coding for the scalable extension of H.264/AVCabstractIn this paper, we propose a scheme for the robust transmission of video in error prone environments using multiple description coding (MDC) based on the scalable video coding extension (SVC) of H.264/AVC. Due to the layer structure of SVC, a base layer is referenced by one or more enhancement layers. The proposed method produces multiple description base layers to achieve robust video communication over unreliable channels with reasonable redundancy. Two base layers are generated using residual data downsampling, which makes it possible that the two base layers have the same motion vectors. The proposed method combines the advantages of SVC and MDC. Experimental results show that the proposed algorithm outperforms the temporal splitting based multiple description scalable coding method in terms of PSNR by 0.3 dB. Zhijie Zhao, Jörn Ostermann, Hexin Chen |
PCS | 2 |
| 2009 | Automated model-based vertebra detection, identification, and segmentation in CT images
Tobias Klinder, Jörn Ostermann, Matthias Ehm, Astrid Franz, Reinhard Kneser, Cristian Lorenz |
Medical Image Anal. | 2 |
| 2009 | Fast Inter-Mode Decision in an H.264/AVC Encoder Using Mode and Lagrangian Cost CorrelationabstractIn this paper we present a novel algorithm to speed up the inter-mode decision process for the H.264/AVC encoding. The proposed inter-mode decision scheme determines the best coding mode of a given macroblock (MB) by predicting the best mode from neighboring MBs in time and in space and by estimating its rate-distortion (RD) cost from the MB in the previous frame. The performance of the proposed algorithm is evaluated in metrics such as the encoding time, the average peak signal-to-noise ratio and the coding bit-rate for test sequences. Simulation results demonstrate that the proposed algorithm can determine the best mode using only one or two rate-distortion cost computations for about half of the MBs resulting in up to 56% total encoding time reduction with on average 2.4% of bit rate increase at the same PSNR compared to H.264/AVC JM 12.1. Song-Hak Ri, Yuri Vatis, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Adaptive Interpolation Filter for H.264/AVCabstractIn order to reduce the bit-rate of video signals, current coding standards apply hybrid coding with motion-compensated prediction and transform coding of the prediction error. In former publications, it has been shown that aliasing components contained in an image signal, as well as motion blur are limiting the prediction efficiency obtained by motion compensation. In this paper, we show that the analytical calculation of an optimal interpolation filter at particular constraints is possible, resulting in total coding improvements of 20% at broadcast quality compared to the H.264/AVC High Profile. Furthermore, the spatial adaptation to local image characteristics enables further improvements of 0.15 dB for CIF sequences compared to globally adaptive filter or up to 0.6 dB, compared to the standard H.264/AVC. Additionally, we show that the presented approach is generally applicable, i.e., also motion blur can be exactly compensated, if particular constraints are fulfilled. Yuri Vatis, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Spatially and temporally scalable compression of animated 3D meshes with MPEG-4 / FAMCabstractWe introduce an efficient method for scalable compression of animated 3D meshes. The approach consists of a combination of inter-frame motion compensation and layer-wise predictive coding of remaining residuals. It is shown that motion compensated predictive coding can lead to a significant improvement in compression performance upon the current state of the art with gains of over 40%. In addition, the created bit stream is temporally and spatially scalable allowing an adaptation to network transfer rates and end-user devices. This compression method is currently standardized within MPEG as part of MPEG-4 AFX Amd. 2, where it is referred to as FAMC - Frame-based Animated Mesh Compression. Nikolce Stefanoski, Jörn Ostermann |
ICIP | 2 |
| 2008 | Robust AAM building for morphing in an image-based facial animation systemabstractFacial animation has been combined with text-to-speech synthesis to create innovative multimodal interfaces, such as Web stores and Web-based customer services. This paper presents an image-based facial animation system using active appearance models (AAM) for precisely detecting feature points in the human face, which are required for selecting mouth images from the database of the face model. In order to minimize the impact of human error when creating the training data, a new optimization method for building an AAM is produced. Optimized training set reduces the average feature point location error from 1.15 pixels to 0.17 pixel. The feature points are suitable for automatic morphing between mouth images with large visual differences. Subjective tests show that morphing improves the visual quality of the animation from "fair" to "good". Kang Liu 0002, Axel Weissenfeld, Jörn Ostermann, Xinghan Luo |
ICME | 3 |
| 2008 | Frame-based compression of animated meshes in MPEG-4abstractThis paper presents a new compression technique for 3D dynamic meshes, referred to as FAMC - frame-based animated mesh compression, promoted within the MPEG-4 standard as amendment 2 of part 16 AFX (animation framework extension). The FAMC approach combines a model-based motion compensation strategy, with transform/predictive coding of residual errors. First, a skinning motion compensation model is automatically computed from a frame-based representation and then encoded. Subsequently, either 1) DCT/lifting wavelets or 2) layer-based predictive coding is employed to exploit remaining spatio-temporal correlations in the residual signal. The proposed encoder offers high compression performances (gains in bit rate of 60% with respect to the previous MPEG-4 technique and of 20% to 40% with respect to state-of-the-art approaches) and is well suited for compressing both geometric and photometric (normal vectors, colors...) attributes. In addition, the FAMC method supports a rich set of functionalities including streaming, scalability (spatial, temporal and quality) and progressive transmission. Khaled Mamou, Titus Zaharia, Françoise J. Prêteux, Nikolce Stefanoski, Jörn Ostermann |
ICME | 5 |
| 2008 | Adaptive error protection for Scalable Video Coding extension of H.264/AVCabstractThis paper presents an adaptive error protection method which provides different packet correction capacities by using only one Reed-Solomon code. The proposed method can be applied separately for each data part in a bit stream. The adaption of the error correction capacity works on-the-fly and only based on the way of data interleaving. In this work, the error protection is applied unequally to data units in the Network Abstraction Layer (NAL) of the Scalable Video Coding (SVC) extension of H.264/AVC. Simulation results show that the video quality increases 6 dB in average with the total overhead of ca. 9%. The advantage of our method is the simpleness and flexibility to apply. Therefore, it is suitable for real-time streaming applications. Dieu Thanh Nguyen, Michio Hayashi, Jörn Ostermann |
ICME | 3 |
| 2008 | Variable bit rate DV streaming with TCP-friendly rate controlabstractDigital video (DV) is one of the most popular high quality video formats. Some applications benefit from real-time DV streaming. However, the huge data rate of this digital video format brings some challenges to deliver a digital video stream over the Internet or intranet. In this paper, a variable bit rate DV streaming system with TCP-friendly rate control is implemented. In order to reduce the sending rate when congestion occurs, the proposed scheme discards some high frequency coefficients of a DV frame without touching the audio and system data or reduces the frame rate. The data of one DV frame can be reduced up to 31.5%. Combining frame reduction with high frequency coefficients discarding, the proposed system can provide several data rate levels that the TCP-friendly rate control can select to adapt to the current available network bandwidth. Through the experimental results, we demonstrate the effectiveness of our DV streaming system. Zhijie Zhao, Dieu Thanh Nguyen, Jörn Ostermann |
ICME | 3 |
| 2008 | Realistic facial animation system for interactive servicesabstractThis paper presents the optimization of parameters of talking head for web-based applications with a talking head, such as Newsreader and E-commerce, in which the realistic talking head initiates a conversation with users. Our talking head system includes two parts: analysis and synthesis. The audiovisual analysis part creates a face model of a recorded human subject, which is composed of a personalized 3D mask as well as a large database of mouth images and their related information. The synthesis part generates facial animation by concatenating appropriate mouth images from the database. A critical issue of the synthesis is the unit selection which selects these appropriate mouth images from the database such that they match the spoken words of the talking head. In order to achieve a realistic facial animation, the unit selection has to be optimized. Objective criteria are proposed in this paper and the Pareto optimization is used to train the unit selection. Subjective tests are carried out in our web-based evaluation system. Experimental results show that most people cannot distinguish our facial animations from real videos. Kang Liu 0002, Jörn Ostermann |
INTERSPEECH | 2 |
| 2008 | Spine Segmentation Using Articulated Shape Models
Tobias Klinder, Robin Wolz, Cristian Lorenz, Astrid Franz, Jörn Ostermann |
MICCAI (1) | 5 |
| 2007 | Layered Predictive Coding of Time-Consistent Dynamic 3D Meshes using a Non-Linear PredictorabstractWe present a layered predictive compression approach for time-consistent dynamic 3D meshes. The algorithm decomposes each frame of a dynamic 3D mesh in layers employing patch-based mesh simplification techniques. This layered decomposition is consistent in time. Following the predictive coding paradigm, local temporal and spatial dependencies between layers and frames are exploited for compression. Prediction is performed vertex-wise from coarse to fine layers exploiting local linear and non-linear dependencies between vertex locations for compression. It is shown that a non-linear predictive exploitation of the proposed layered configuration of vertices can improve the compression performance upon other state-of-the-art approaches by more than 15% in domains relevant for applications. Nikolce Stefanoski, Patrick Klie, Xiaoliang Liu, Jörn Ostermann |
ICIP (5) | 4 |
| 2007 | Inverse Bit Plane Decoding Order for Turbo Code Based Distributed Video CodingabstractFor turbo code based Wyner-Ziv codecs, we propose to use an inverse bit plane decoding order. The investigations have shown that the knowledge about MSB's has only a marginal influence on LSB's, which can be regarded as noise, while the influence of LSB's on MSB's tends to be higher. As a results, we obtain a coding gain of up to 0.3 dB, especially for the sequences, which can only be decoded with a coding efficiency lower than H.264/AVC intra. Furthermore, the distribution of the requests for additional parity bits over the back channel is different compared to the conventional decoding order, resulting in a high number of requests for the last bit plane and a low number of requests for the other bit planes. Therefore, the decoder can request several puncturing levels for LSB's in one step, reducing the number of turbo decoding loops. Hence, transmitting the LSB's first, we can obtain a decoding speedup of 30%, thus making DVC more attractive for real-time applications. Yuri Vatis, Sven Klomp, Jörn Ostermann |
ICIP (2) | 3 |
| 2007 | Enhanced Reconstruction of the Quantised Transform Coefficients for WYNER-ZIV CodingabstractIn recent years, the Distributed Video Coding (DVC) has become a more and more popular research area. In the current work, the framework of transform domain Wyner-Ziv coding of video frames is considered, following the scheme developed in the EU Project DISCOVER. While some frames are conventionally coded as key frames, other frames are Wyner-Ziv encoded using turbo codes. After the generation of side information and turbo decoding process at the decoder, the quantisation indices of the quantised transform coefficients of Wyner-Ziv frames are available. In this paper, we devote our attention to the reconstruction of quantised transform coefficients resulting in a minimised quantisation error energy, where coding gains of up to 0.9 dB are obtained. Yuri Vatis, Sven Klomp, Jörn Ostermann |
ICME | 3 |
| 2007 | Automatic speech recognition with a cochlear implant front-endabstractToday, cochlear implants (CIs) are the treatment of choice in patients with profound hearing loss. However speech intelligibility with these devices is still limited. A factor that determines hearing performance is the processing method used in CIs. Therefore, research is focused on designing different speech processing methods. The evaluation of these strategies is subject to variability as it is usually performed with cochlear implant recipients. Hence, an objective method for the evaluation would give more robustness compared to the tests performed with CI patients. This paper proposes a method to evaluate signal processing strategies for CIs based on a hidden markov model speech recognizer. Two signal processing strategies for CIs, the Advanced Combinational Encoder (ACE) and the Psychoacoustic Advanced Combinational Encoder (PACE), have been compared in a phoneme recognition task. Results show that PACE obtained higher recognition scores than ACE as found with CI r ecipients. Waldo Nogueira, Tamás Harczos, Bernd Edler, Jörn Ostermann, Andreas Büchner |
INTERSPEECH | 4 |
| 2007 | Automated Model-Based Rib Cage Segmentation and Labeling in CT Images
Tobias Klinder, Cristian Lorenz, Jens von Berg, Sebastian P. M. Dries, Thomas Bülow, Jörn Ostermann |
MICCAI (2) | 6 |
| 2007 | Special issue on three-dimensional video and television
M. Reha Civanlar, Jörn Ostermann, Haldun M. Özaktas, Aljoscha Smolic, John Watson |
Signal Process. Image Commun. | 2 |
| 2007 | Introduction to the Special Section on Multiview Video CodingabstractThe 11 papers in this special section focus on multiview video coding. Jörn Ostermann, Masayuki Tanimoto, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Coding Algorithms for 3DTV - A SurveyabstractResearch efforts on 3DTV technology have been strengthened worldwide recently, covering the whole media processing chain from capture to display. Different 3DTV systems rely on different 3D scene representations that integrate various types of data. Efficient coding of these data is crucial for the success of 3DTV. Compression of pixel-type data including stereo video, multiview video, and associated depth or disparity maps extends available principles of classical video coding. Powerful algorithms and open international standards for multiview video coding and coding of video plus depth data are available and under development, which will provide the basis for introduction of various 3DTV systems and services in the near future. Compression of 3D mesh models has also reached a high level of maturity. For static geometry, a variety of powerful algorithms are available to efficiently compress vertices and connectivity. Compression of dynamic 3D geometry is currently a more active field of research. Temporal prediction is an important mechanism to remove redundancy from animated 3D mesh sequences. Error resilience is important for transmission of data over error prone channels, and multiple description coding (MDC) is a suitable way to protect data. MDC of still images and 2D video has already been widely studied, whereas multiview video and 3D meshes have been addressed only recently. Intellectual property protection of 3D data by watermarking is a pioneering research area as well. The 3D watermarking methods in the literature are classified into three groups, considering the dimensions of the main components of scene representations and the resulting components after applying the algorithm. In general, 3DTV coding technology is maturating. Systems and services may enter the market in the near future. However, the research area is relatively young compared to coding of other types of media. Therefore, there is still a lot of room for improvement and new development of algorithms. Aljoscha Smolic, Karsten Müller 0001, Nikolce Stefanoski, Jörn Ostermann, Atanas P. Gotchev, Gozde Bozdagi Akar, George A. Triantafyllidis, Alper Koz |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | 3-D Time-Varying Scene Capture Technologies - A SurveyabstractAdvances in image sensors and evolution of digital computation is a strong stimulus for development and implementation of sophisticated methods for capturing, processing and analysis of 3D data from dynamic scenes. Research on perspective time-varying 3D scene capture technologies is important for the upcoming 3DTV displays. Methods such as shape-from-texture, shape-from-shading, shape-from-focus, and shape-from-motion extraction can restore 3D shape information from a single camera data. The existing techniques for 3D extraction from single-camera video sequences are especially useful for conversion of the already available vast mono-view content to the 3DTV systems. Scene-oriented single-camera methods such as human face reconstruction and facial motion analysis, body modeling and body motion tracking, and motion recognition solve efficiently a variety of tasks. 3D multicamera dynamic acquisition and reconstruction, their hardware specifics including calibration and synchronization and software demands form another area of intensive research. Different classes of multiview stereo algorithms such as those based on cost function computing and optimization, fusing of multiple views, and feature-point reconstruction are possible candidates for dynamic 3D reconstruction. High-resolution digital holography and pattern projection techniques such as coded light or fringe projection for real-time extraction of 3D object positions and color information could manifest themselves as an alternative to traditional camera-based methods. Apart from all of these approaches, there also are some active imaging devices capable of 3D extraction such as the 3D time-of-flight camera, which provides 3D image data of its environment by means of a modulated infrared light source. Elena Stoykova, A. Aydin Alatan, Philip W. Benzie, Nikolaos Grammalidis, Sotiris Malassiotis, Jörn Ostermann, S. Piekh, Ventseslav Sainov, Christian Theobalt, T. Thevar, Xenophon Zabulis |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2006 | Fast Mode Decision for H.264/AVC Using Mode Prediction
Song-Hak Ri, Jörn Ostermann |
ACIVS | 2 |
| 2006 | Parameterization of Mouth Images by LLE and PCA for Image-Based Facial AnimationabstractThis paper describes parameterization of mouth images for an image-based facial animation system. The analysis part of the facial animation system produces a face model, which is composed of a personalized mask as well as a large database of mouth images and their related phonetic and visual information. Then, a photo-realistic talking head is synthesized by rendering a personalized mask textured with a mouth image, which is selected from the database. The selection is driven by a unit selection algorithm, which finds the appropriate mouth images from the database such that they match the words spoken by the talking head. The selection of mouth images is based on parameters describing the mouth images. Therefore, the parameterization of mouth images is the key part for creating a photo-realistic facial animation. Hereby the visual parameterization of mouth images by LLE (locally linear embedding) is investigated comprehensively and compared with PCA (principal component analysis). Experimental results show that the parameterization of mouth images by LLE performs better for an image-based facial animation system than by PCA Kang Liu 0002, Axel Weissenfeld, Jörn Ostermann |
ICASSP (5) | 3 |
| 2006 | Connectivity-Guided Predictive Compression of Dynamic 3D MeshesabstractWe introduce an efficient algorithm for real-time compression of temporally consistent dynamic 3D meshes. The algorithm uses mesh connectivity to determine the order of compression of vertex locations within a frame. Compression is performed in a frame to frame fashion using only the last decoded frame and the partly decoded current frame for prediction. Following the predictive coding paradigm, local temporal and local spatial dependencies between vertex locations are exploited. In this framework we present a novel angle preserving predictor and evaluate its performance against other state of the art predictors. It is shown that the proposed algorithm improves up to 25% upon the current state of the art for compression of temporally consistent dynamic 3D meshes. Nikolce Stefanoski, Jörn Ostermann |
ICIP | 2 |
| 2006 | Locally Adaptive Non-Separable Interpolation Filter for H.264/AVCabstractIn order to reduce the bit rate of video signals, current coding standards apply hybrid coding with motion-compensated prediction and transform coding of the prediction error. In former publications, it was shown that aliasing components contained in an image signal, as well as inaccuracies in description of motion, are limiting the prediction efficiency obtained by motion compensation. For H.264/AVC, we showed that the analytical calculation of an optimal interpolation filter at given constraints (6-tap, frame-based) is possible, resulting in total coding improvements of up to 0.9 dB for HDTV sequences and up to 0.5 dB for CIF sequences and main profile. However, applying an adaptive interpolation filter to the entire image enables only an average prediction improvement. Further improvements are possible if the filter is also adapted to local image characteristics. Thus, an optimal adaptive interpolation filter is calculated only for the macroblocks, which prediction is distorted by aliasing or by inaccurately estimated motion. This enables further improvements of up to 0.2 dB, compared to globally adaptive filter or up to 0.6 dB, compared to the standard H.264/AVC, for CIF sequences. Yuri Vatis, Jörn Ostermann |
ICIP | 2 |
| 2006 | Robust Rigid Head Motion Estimation Based on Differential EvolutionabstractIn this paper we present a system to robustly estimate the 3D position of a human head. Before the face model is positioned in the initial frame, it is adapted to the 3D scan of the tracked human head. Head tracking is achieved by minimizing a robust cost function with a stochastic optimization algorithm called differential evolution. This approach enables the estimation of large motions between consecutive frames. Furthermore, the algorithm can even handle a large number of outliers e.g. caused by occlusion and still estimate the precise position Axel Weissenfeld, Onay Urfalioglu, Kang Liu 0002, Jörn Ostermann |
ICME | 4 |
| 2005 | Motion-and aliasing-compensated prediction using a two-dimensional non-separable adaptive Wiener interpolation filterabstractIn the context of prediction with fractional-pel motion vector resolution it was shown, that aliasing components contained in an image signal are limiting the prediction accuracy obtained by motion compensation. In order to consider aliasing, quantisation and motion estimation errors, camera noise, etc., we analytically developed a two-dimensional (2D) non-separable interpolation filter, which is calculated for each frame independently by minimising the prediction error energy. For every fractional-pel position to be interpolated, an individual set of 2D filter coefficients is determined. As a result, a coding gain of up to 1,2 dB for HDTV-sequences and up to 0,5 dB for CIF-sequences compared to the standard H.264/AVC is obtained. Yuri Vatis, Bernd Edler, Dieu Thanh Nguyen, Jörn Ostermann |
ICIP (2) | 4 |
| 2005 | Personalized Unit Selection for an Image-based Facial Animation SystemabstractThis paper describes an image-based facial animation system, which consists of the audiovisual analysis of a human subject and the synthesis of a photo-realistic facial animation. The unit selection algorithm selects for a given audio output the best mouth samples from the database by assigning two costs, the phonetic context and the visual distance between two consecutive samples. Here a novel approach to adapt the unit selection algorithm to an individual human subject is presented, such that a photo-realistic facial animation can be generated Axel Weissenfeld, Kang Liu 0002, Sven Klomp, Jörn Ostermann |
MMSP | 4 |
| 2004 | Progressive coding for QoS-enabled streaming of dynamic 3-D meshesabstractWe introduce a new compressed representation for the geometry of dynamic polygonal 3-D meshes, suitable for streaming rate-controlled 3-D animations with fine-grained adaptation to the available network transmission rates. The new representation is based on both mesh simplification and dynamic geometry compression methods. This is in contrast to all other simplification algorithms reported to date, which are tailored to static meshes, whose vertex positions do not vary with time. Furthermore, recently reported methods for compressing dynamic meshes do not encode geometry in a progressive manner. In this work, we lay the foundations of progressive encoding of dynamic 3-D meshes, by formally expressing such models and their associated sequences of topological operations. We also propose a rate adaptation scheme for animated meshes, based on vertex pair candidate contractions and vertex splits, and show its efficiency compared to a typical quantisation-based approach. Socrates Varakliotis, Stephen Hailes, Jörn Ostermann |
ICC | 3 |
| 2003 | Optimally smooth error resilient streaming of 3D wireframe animations
Socrates Varakliotis, Stephen Hailes, Jörn Ostermann |
VCIP | 3 |
| 2003 | Lifelike talking faces for interactive servicesabstractLifelike talking faces for interactive services are an exciting new modality for man-machine interactions. Recent developments in speech synthesis and computer animation enable the real-time synthesis of faces that look and behave like real people, opening opportunities to make interactions with computers more like face-to-face conversations. This paper focuses on the technologies for creating lifelike talking heads, illustrating the two main approaches: model-based animations and sample-based animations. The traditional model-based approach uses three-dimensional wire-frame models, which can be animated from high-level parameters such as muscle actions, lip postures, and facial expressions. The sample-based approach, on the other hand, concatenates segments of recorded videos, instead of trying to model the dynamics of the animations in detail. Recent advances in image analysis enable the creation of large databases of mouth and eye images, suited for sample-based animations. The sample-based approach tends to generate more naturally looking animations at the expense of a larger size and less flexibility than the model-based animations. Beside lip articulation, a talking head must show appropriate head movements, in order to appear natural. We illustrate how such "visual prosody" is analyzed and added to the animations. Finally, we present four applications where the use of face animation in interactive services results in engaging user interfaces and an increased level of trust between user and machine. Using an RTP-based protocol, face animation can be driven with only 800 bits/s in addition to the rate for transmitting audio. Eric Cosatto, Jörn Ostermann, Hans Peter Graf, Juergen Schroeter |
Proc. IEEE | 2 |
| 2003 | Introduction to the special issue on image-based modeling, rendering, and animation
Harry Shum, Eric Petajan, Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Repair options for 3-D wireframe model animation sequencesabstractTo date, almost all work within the field of 3D animation has focused on raising perceptual appeal to an adequate level. However, it is a non-trivial task to match the requirements of real-time streams to the reality of the Internet. One of the key problems that must be addressed is that of how best to conceal errors in the event of packet loss. Thus we present the results of experiments designed to evaluate the effect of different possible schemes under different conditions of loss. We conclude that frame insertion methods show good performance at low loss rates, but that frame interpolation methods exhibit better performance at both low and high burst loss rates, though at the cost of added delay and complexity. Visualised traces subjectively validate these results, by exhibiting smoother animation and reduced artifacts on the animated wireframe models. Socrates Varakliotis, Stephen Hailes, Jörn Ostermann |
ICME (1) | 3 |
| 2001 | Coding Of Animated 3-D Wireframe Models For Internet Streaming ApplicationsabstractIn this paper we present a coding technique for 3-D animated wireframe models, suitable for Internet streaming. First we present the MPEG-4 like coding scheme with simple RTP packetisation, and provide details of the achieved compression efficiency. Then, we describe a distortion metric for such a signal and we conduct experiments to study the effect of packet loss, following a bursty packet loss model that approximates a simulated IP network. Our experiments examine the performance of the coding scheme for a simple streaming scenario with different sequence configurations. The results show that short-term and short average length burst losses of up to 30% have a logarithmic effect on the decrease of the animation smoothness in the case of simple differential coding. This logarithmic decrease is corrected to linear by inserting I-frames to the sequence at the expense of reduced compression. The result is smoother animation. Socrates Varakliotis, Jörn Ostermann, Vicky Hardman |
ICME | 2 |
| 2000 | Source Models for Content-Based Video CodingabstractDifferent source models for content-based video coding are presented. An object-based analysis-synthesis coder describes video objects by means of motion, shape and texture parameters. In contrast to block-based coding, the representation of arbitrarily shaped objects in a real scene is possible. Using knowledge about the scene content and about the object's behavior, object-based coding is extended to knowledge-based coding and semantic coding, respectively. Compared with other content-based video coding concepts like MBASIC and MPEG-4, the above mentioned coders provide an integrated and more efficient framework for coding video using 3D source models. Jörn Ostermann, Markus Kampmann |
ICIP | 1 |
| 2000 | Special Issue on Shape Coding for Emerging Multimedia Applications
Minoru Etoh, Jörn Ostermann, Thomas Sikora |
Signal Process. Image Commun. | 2 |
| 2000 | Face and 2-D mesh animation in MPEG-4abstractThis paper presents an overview of some of the synthetic visual objects supported by MPEG-4 version-1, namely animated faces and animated arbitrary 2D uniform and Delaunay meshes. We discuss both specification and compression of face animation and 2D-mesh animation in MPEG-4. Face animation allows to animate a proprietary face model or a face model downloaded to the decoder. We also address integration of the face animation tool with the text-to-speech interface (TTSI), so that face animation can be driven by text input. A. Murat Tekalp, Jörn Ostermann |
Signal Process. Image Commun. | 2 |
| 1999 | Subjective evaluation of animated talking facesabstractComputer simulation of human faces has been an active research area for a long time, resulting in a multitude of facial models and several animation systems. Current interest for this technology is clearly shown by its inclusion in the MPEG-4 standard. However, it is less clear what the actual applications of facial animation (FA) will be. We have therefore undertaken experiments on 190 subjects in order to explore the benefits of FA. Part of the experiment was aimed at exploring the objective benefits, i.e., to see if FA can help users to perform certain tasks more accurately or efficiently. The other part of the experiment was aimed at more subjective benefits, like raising the level of appeal to the user, gaining more users interest, filling in the waiting times for server access so the users do not get bored. We present the experiment design and the results, The results show that FA aids users in understanding spoken numbers in noisy conditions (error rates drop from 16% to 8%); that it can make waiting times more acceptable to the user; and that it makes services more attractive to the users. Jörn Ostermann, David R. Millen, Igor S. Pandzic |
MMSP | 1 |
| 1999 | Detection of Moving Cast Shadows for Object SegmentationabstractTo prevent moving shadows being misclassified as moving objects or parts of moving objects, this paper presents an explicit method for detection of moving cast shadows on a dominating scene background. Those shadows are generated by objects moving between a light source and the background. Moving cast shadows cause a frame difference between two succeeding images of a monocular video image sequence. For shadow detection, these frame differences are detected and classified into regions covered and regions uncovered by a moving shadow. The detection and classification assume plane background and a nonnegligible size and intensity of the light sources. A cast shadow is detected by temporal integration of the covered background regions while subtracting the uncovered background regions. The shadow detection method is integrated into an algorithm for two-dimensional (2-D) shape estimation of moving objects from the informative part of the description of the international standard ISO/MPEG-4. The extended segmentation algorithm compensates first apparent camera motion. Then, a spatially adaptive relaxation scheme estimates a change detection mask for two consecutive images. An object mask is derived from the change detection mask by elimination of changes due to background uncovered by moving objects and by elimination of changes due to background covered or uncovered by moving cast shadows. Results obtained with MPEG-4 test sequences and additional sequences show that the accuracy of object segmentation is substantially improved in presence of moving cast shadows. Objects and shadows are detected and tracked separately. Jürgen Stauder, Roland Mech, Jörn Ostermann |
IEEE Trans. Multim. | 3 |
| 1999 | User evaluation: Synthetic talking faces for interactive services
Igor S. Pandzic, Jörn Ostermann, David R. Millen |
Vis. Comput. | 2 |
| 1998 | Animation of Synthetic Faces in MPEG-4abstractMPEG-4 is the first international standard that standardizes true multimedia communication-including natural and synthetic audio, natural and synthetic video, as well as 3D graphics. Integrated into this standard is the capability to define and animate virtual humans consisting of synthetic heads and bodies. For the head, more than 70 model-independent animation parameters defining low-level actions like "move left mouth corner" up to high-level parameters like facial expressions and visemes are standardized In a communication application. The encoder can define the face model using MPEG-4 BIFS (BInary Format for Scenes) and transmit it to the decoder. Alternatively, the encoder can rely on a face model that is available at the decoder. The animation parameters are quantized, predictively encoded using an arithmetic encoder or a DCT. The decoder receives the model and the animation parameters in order to animate the model. Since MPEG-4 defines the minimum MPEG-4 terminal capabilities in profiles and levels, the encoder knows the quality of the animation at the decoder. Jörn Ostermann |
CA | 1 |
| 1998 | Natural and synthetic video in MPEG-4abstractThe ISO MPEG Committee, after successful completion of the MPEG-1 and the MPEG-2 standards, has recently completed the Committee Draft for MPEG-4, its third standard. MPEG-4 is designed to be an object-based standard for multimedia coding. The visual part of the standard specifies coding of both natural and synthetic video. The MPEG-4 visual standard supports coding of natural video not only in a conventional manner (using frames) but also as a collection of arbitrary shape objects (using video object planes). Further, it supports functionalities such as spatial and temporal scalability, both conventional and with arbitrary shape objects. It also supports error resilient coding for delivery of coded video on error prone channels. MPEG-4 visual standard also supports coding of synthetic video which includes still texture maps used in 3D graphics models, mesh geometry for object animation, and parameters for facial animation. Jörn Ostermann, Atul Puri |
ICASSP | 1 |
| 1998 | Efficient Encoding of Binary Shapes using MPEG-4abstractMPEG-4 Visual, that part of the upcoming MPEG-4 standard describing the coding of natural and synthetic video signals, allows the encoding of video objects using motion, texture and shape information. The MPEG-4 context-based arithmetic encoder for encoding binary shape information is presented in the context of a new MPEG-4 video encoder architecture. The encoder architecture enables us to efficiently encode lossless and lossy shape, motion and texture of a moving video object. Several non-normative choices for efficient computation and bit efficient encoding of arbitrarily shaped video objects are investigated in order to enable real-time encoding. Jörn Ostermann |
ICIP (1) | 1 |
| 1998 | Integration of talking heads and text-to-speech synthesizers for visual TTSabstractThe integration of text-to-speech (TTS) synthesis and animation of synthetic faces allows new applications like visual human computer interfaces using agents or avatars. The TTS informs the talking head when phonemes are spoken. The appropriate mouth shapes are animated and rendered while the TTS produces the sound. We call this integrated system of TTS and animation a Visual TTS (VTTS). This paper describes the architecture on an integrated VTTS synthesizer that allows defining facial expressions as bookmarks in the text that will be animated while the model is talking. The position of a bookmark in the text defines the start time for the facial expression. The bookmark itself names the expression, its amplitude and the duration during which the amplitude has to be reached by the face. A bookmark to face animation parameter (FAP) converter creates a curve defining the amplitude for the given FAP over time using Hermite functions of 3 rd order. 1. INTRODUCTION With the new generati... Jörn Ostermann, Marc C. Beutnagel, Ariel Fischer |
ICSLP | 1 |
| 1998 | MPEG-4 and rate-distortion-based shape-coding techniquesabstractWe address the problem of the efficient encoding of object boundaries. This problem is becoming increasingly important in applications such as content-based storage and retrieval, studio and television postproduction, and mobile multimedia applications. The MPEG-4 visual standard will allow the transmission of arbitrarily shaped video objects. The techniques developed for shape coding within the MPEG-4 standardization effort are described and compared first. A framework for the representation of shapes using their contours is presented next. Such representations are achieved using curves of various orders, and they are optimal in the rate-distortion sense. Finally, conclusions are drawn. Aggelos K. Katsaggelos, Lisimachos P. Kondi, Fabian W. Meier, Jörn Ostermann, Guido M. Schuster |
Proc. IEEE | 4 |
| 1998 | Image and video coding-emerging standards and beyondabstractDiscusses coding standards for still images and motion video. We first briefly discuss standards already in use, including: Group 3 and Group 4 for bilevel fax images; JPEG for still color images; and H.261, H.263, MPEG-1, and MPEG-2 for motion video. We then cover newly emerging standards such as JBIG1 and JBIG2 for bilevel fax images, JPEG-2000 for still color images, and H.263+ and MPEG-4 for motion video. Finally, we describe some directions beyond the standards such as hybrid coding of graphics/photo images, MPEG-7 for multimedia metadata, and possible new technologies. Barry G. Haskell, Paul G. Howard, Yann LeCun, Atul Puri, Jörn Ostermann, M. Reha Civanlar, Lawrence R. Rabiner, Léon Bottou, Patrick Haffner |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 1998 | Evaluation of mesh-based motion estimation in H.263-like codersabstractWe present two mesh-based motion estimation algorithms, and evaluate their performance when incorporated in an H.263-like block-based video coder. Both algorithms compute nodal motions in a hierarchical manner. Within each hierarchy level, the first algorithm (HMMA) minimizes the prediction error in the four elements surrounding each node, where the prediction is accomplished by a bilinear mapping. The optimal solution is obtained by a full search within a range defined by the topology of the mesh. The second algorithm (HBMA) minimizes the error in a block surrounding each node, assuming the motion in the block is constant. In both cases, bilinear mapping is used for motion-compensated prediction based on nodal displacements. The two algorithms are compared with an exhaustive block-matching algorithm (EBMA) by evaluating their performance in temporal prediction and in an H.263/TMN4 coder. For prediction only, the HMMA and HBMA algorithms yield visually more satisfactory results, even though the PSNRs of the predicted images are on average lower. The coded images also have lower PSNRs at similar bit rates. The coding artifacts are different: while the block-based method leads to more severe block distortions, the mesh-based method experiences some warping artifacts. The HMMA algorithm outperforms the HBMA slightly for certain sequences at the expense of higher computational complexity. Jörn Ostermann |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | Feedback loop for coder control in a block-based hybrid coder with mesh-based motion compensationabstractIn this paper, the motion estimation and motion compensation within a block-based hybrid coder is modified. Input to the motion estimator is the original frame and a representation of the previously decoded frame generated by means of a second prediction loop. The second prediction loop works in parallel to the prediction loop of the decoder. It distinguishes itself from the conventional coder prediction loop such that for blocks with transmitted DCT coefficients, the original image signal is fed into the second prediction loop, and for blocks without transmitted DCT coefficients, the motion compensated signal is fed into the prediction loop. The image generated by the second prediction loop is influenced by motion but not by the quantization noise of the DCT. The motion-compensated prediction image from the second predictor loop can be easily used for control of the encoder. This coder control is not influenced by the actual quantization selected by the encoder and hence very stable for a wide range of bit rates. Savings of 15% to 25% in bit rate can be observed without loss of subjective picture quality. Jörn Ostermann |
ICASSP | 1 |
| 1997 | Coding of Arbitrarily Shaped Video Object in MPEG-4abstractMPEG-4 Visual, that part of the upcoming MPEG-4 standard describing the coding of natural and synthetic video signals, allows the encoding of video objects using motion, texture and shape information. In this paper, the shape coding algorithms and related texture coding algorithms that were thoroughly evaluated within MPEG-4 are reviewed. The evaluation leading to the selection of a block-based shape coder based on a context-based arithmetic encoder/decoder with motion compensation as well as padding and shape adaptive DCT for texture coding are presented. Jörn Ostermann, Euee S. Jang, Jae-Seob Shin |
ICIP (1) | 1 |
| 1997 | Animated talking head with personalized 3D head modelabstractNatural Human-Computer Interface requires integration of realistic audio and visual information for perception and display. An example of such an interface is an animated talking head displayed on the computer screen in the form of a human-like computer agent. This system converts text to acoustic speech with synchronized animation of mouth movements. The talking head is based on a generic 3D human head model, but to improve realism, natural looking personalized models are necessary. In this paper we report results in adapting a generic head model to 3D range data of a human head obtained from a 3D laser range scanner. This personalized model is incorporated into the talking head system. With texture mapping, the personalized model offers a more natural and realistic look than the generic model. Lawrence S. Chen, Thomas S. Huang, Jörn Ostermann |
MMSP | 3 |
| 1997 | MPEG-4: Audio/video and synthetic graphics/audio for mixed media
Peter K. Doenges, Tolga K. Çapin, Fabio Lavagetto, Jörn Ostermann, Igor S. Pandzic, Eric Petajan |
Signal Process. Image Commun. | 4 |
| 1997 | Automatic adaptation of a face model in a layered coder with an object-based analysis-synthesis layer and a knowledge-based layer
Markus Kampmann, Jörn Ostermann |
Signal Process. Image Commun. | 2 |
| 1997 | Methodologies used for evaluation of video tools and algorithms in MPEG-4
Jörn Ostermann |
Signal Process. Image Commun. | 1 |
| 1996 | The European COST211ter activities-research towards advanced algorithms for coding of video signals at very low bit ratesabstractA significant effort of the COST211ter group activities is dedicated towards the contribution for the MPEG-4 video group activities. This paper provides an overview of the COST telecommunication framework and discusses the technical simulation model approach taken by the COST211ter group. Video coding with application to multimedia services is discussed. Thomas Sikora, Jörn Ostermann |
ICIP (3) | 2 |
| 1994 | Object-based analysis-synthesis coding based on the source model of moving rigid 3D objects
Jörn Ostermann |
Signal Process. Image Commun. | 1 |
| 1994 | Object-based analysis-synthesis coding (OBASC) based on the source model of moving flexible 3-D objectsabstractThis correspondence investigates object-based analysis-synthesis coding (OBASC) for the encoding of moving images at very low data rates. According to the source model, each moving object of an image is described and encoded by three parameter sets defining its motion, shape, and surface color. The parameter sets of each object are obtained by model-based image analysis. They are coded by an object-dependent parameter coding. Using the coded parameter sets, an image can be synthesized by model-based image synthesis. Here, OBASC based on the source model of "moving flexible 3-D objects with 3-D motion" (F3D) is introduced. The efficiency of this source model F3D is compared to the efficiency of OBASC based on the source model of "moving rigid 3-D objects with 3-D motion" (R3D). Compared to R3D, F3D requires the additional transmission of flexible-shape parameters. Therefore, the source model F3D is only applied in those areas of the image which cannot be described by the source model R3D. The new source model F3D reduces the bit rate from 64 to 56 kb/s, providing the same picture quality measured by the SNR of the encoded color parameters. Jörn Ostermann |
IEEE Trans. Image Process. | 1 |
| 1989 | Object-oriented analysis-synthesis coding of moving images
Hans Georg Musmann, Michael Hötter, Jörn Ostermann |
Signal Process. Image Commun. | 3 |