EDBT 2026 Demo / reviewers in the wild / expert
Karsten Müller 0001
dblp:43/508-1 · also Karsten Mueller 0001
· DBLP profile ↗
64ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0001-8611-7864ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 10 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Sparsity for Privacy in Collaborative InferenceabstractCollaborative inference (CI) is hampered by high communication costs and privacy risks, with existing defenses often forcing a trade-off between efficiency and formal privacy guarantees. In this work, we present a framework that leverages activation sparsity as a dual-purpose mechanism to address both challenges simultaneously. Our approach uses a lightweight Sparse Autoencoder (SAE) to learn a sparse representation, which is then protected by a novel two-channel noise mechanism grounded in information theory. This design provides a tunable privacy budget while remaining computationally inexpensive. Evaluations on CIFAR-10, Tiny-ImageNet, and FaceScrub show that our method achieves a state-of-the-art privacy-utility trade-off, sustaining high accuracy at sparsity levels of up to 97%, while offering superior resilience against strong model inversion attacks. Our results underline that sparsity can be transformed from an effective compression tool into a powerful and theoretically-grounded privacy defense, paving the way for more practical and trustworthy CI systems. We provide code at https://github.com/an7123/privacy_ci. Maximilian Andreas Hoefler, Karsten Müller 0001, Wojciech Samek |
WACV | 2 |
| 2025 | Efficient Federated Learning of Mixed-Token Transformers for Cellular Feature PredictionabstractIn telecommunications, Autonomous Networks (ANs) automatically adjust configurations based on specific requirements (e.g., bandwidth) and available resources. These networks rely on continuous monitoring and intelligent mechanisms for self-optimization, self-repair, and self-protection, nowa-days enhanced by Neural Networks (NNs) to enable predictive modeling and pattern recognition.In this work, we propose a neural architecture called Mixed-Token Transformer (MTT) that can process and predict both numerical and textual tokens representing various mobile network features (e.g., ping, SNR or band frequency). We train the models using a Federated Learning (FL) approach, enabling multiple AN cells — each equipped with an MTT — to collaboratively learn from cellular data while preserving data privacy.Since FL involves frequent transmission of neural network updates, it requires an efficient, standardized compression strategy for reliable communication. To address this, we investigate NNCodec, our implementation of the ISO/IEC Neural Network Coding (NNC) standard. Our experimental results on the Berlin V2X dataset demonstrate that NNCodec achieves transparent compression (i.e., negligible performance loss) while reducing communication overhead to below 1%, showing the effectiveness of combining NNC with FL in collaboratively learned autonomous mobile networks. Our proposed MTT architecture speeds up model inference by ≈ 5×. Code is available at: github.com/Jostarndt/NNCodec-FL-MixedTokenTransformer. Daniel Becking, Jost Arndt, Ingo Friese, Karsten Müller 0001, Jackie Ma, Thomas Buchholz, Mandy Galkow-Schneider, Wojciech Samek, Detlev Marpe |
GLOBECOM | 4 |
| 2025 | FedXDS: Leveraging Model Attribution Methods to Counteract Data Heterogeneity in Federated Learningabstract4572 Maximilian Andreas Hoefler, Karsten Müller 0001, Wojciech Samek |
ICCV | 2 |
| 2025 | A Privacy Preserving System for Movie Recommendations Using Federated LearningabstractRecommender systems have become ubiquitous in the past years. They solve the tyranny of choice problem faced by many users, and are utilized by many online businesses to drive engagement and sales. Besides other criticisms, like creating filter bubbles within social networks, recommender systems are often reproved for collecting considerable amounts of personal data. However, to personalize recommendations, personal information is fundamentally required. A recent distributed learning scheme called federated learning has made it possible to learn from personal user data without its central collection. Consequently, we present a recommender system for movie recommendations, which provides privacy and thus trustworthiness on multiple levels: First and foremost, it is trained using federated learning and thus, by its very nature, privacy-preserving, while still enabling users to benefit from global insights. Furthermore, a novel federated learning scheme, called FedQ, is employed, which not only addresses the problem of non-i.i.d.-ness and small local datasets, but also prevents input data reconstruction attacks by aggregating client updates early. Finally, to reduce the communication overhead, compression is applied, which significantly compresses the exchanged neural network parametrizations to a fraction of their original size. We conjecture that this may also improve data privacy through its lossy quantization stage. David Neumann, Andreas Lutz, Karsten Müller 0001, Wojciech Samek |
Trans. Recomm. Syst. | 3 |
| 2024 | Neural Network Coding of Difference Updates for Efficient Distributed Learning CommunicationabstractDistributed learning requires a frequent communication of neural network update data. For this, we present a set of new compression tools, jointly called differential neural network coding (dNNC). dNNC is specifically tailored to efficiently code incremental neural network updates and includes tools for federated BatchNorm folding (FedBNF), structured and unstructured sparsification, tensor row skipping, quantization optimization and temporal adaptation for improved context-adaptive binary arithmetic coding (CABAC). Furthermore, dNNC provides a new parameter update tree (PUT) mechanism, which allows to identify updates for different neural network parameter sub-sets and their relationship in synchronous and asynchronous neural network communication scenarios. Most of these tools have been included into the standardization process of the NNC standard (ISO/IEC 15938-17) edition 2. We benchmark dNNC in multiple federated and split learning scenarios using a variety of NN models and data including vision transformers and large-scale ImageNet experiments: It achieves compression efficiencies of 60% in comparison to the NNC standard edition 1 for transparent coding cases, i.e., without degrading the inference or training performance. This corresponds to a reduction in the size of the NN updates to less than 1% of their original size. Moreover, dNNC reduces the overall energy consumption required for communication in federated learning systems by up to 94%. Daniel Becking, Karsten Müller 0001, Paul Haase, Heiner Kirchhoffer, Gerhard Tech, Wojciech Samek, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Region-Based Template Matching Prediction for Intra CodingabstractCopy prediction is a renowned category of prediction techniques in video coding where the current block is predicted by copying the samples from a similar block that is present somewhere in the already decoded stream of samples. Motion-compensated prediction, intra block copy, template matching prediction etc. are examples. While the displacement information of the similar block is transmitted to the decoder in the bit-stream in the first two approaches, it is derived at the decoder in the last one by repeating the same search algorithm which was carried out at the encoder. Region-based template matching is a recently developed prediction algorithm that is an advanced form of standard template matching. In this method, the reference area is partitioned into multiple regions and the region to be searched for the similar block(s) is conveyed to the decoder in the bit-stream. Further, its final prediction signal is a linear combination of already decoded similar blocks from the given region. It was demonstrated in previous publications that region-based template matching is capable of achieving coding efficiency improvements for intra as well as inter-picture coding with considerably less decoder complexity than conventional template matching. In this paper, a theoretical justification for region-based template matching prediction subject to experimental data is presented. Additionally, the test results of the aforementioned method on the latest H.266/Versatile Video Coding (VVC) test model (version VTM-14.0) yield an average Bjøntegaard-Delta (BD) bit-rate savings of $-0.75\%$ using all intra (AI) configuration with 130% encoder run-time and 104% decoder run-time for a particular parameter selection. Gayathri Venugopal, Karsten Müller 0001, Jonathan Pfaff, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | History Dependent Significance Coding for Incremental Neural Network CompressionabstractThis paper presents an improved probability estimation scheme for the entropy coder of Incremental Neural Network Coding (INNC), which is currently under standardization in ISO/IEC MPEG. More specifically, the paper first analyzes the compression performance of INNC and how the bitstream size relates to the neural network (NN) layers. For the layers requiring the most bits, it analyzes the coded NN weight updates and their temporal dependencies. Major finding is that the probability of a significant (i.e., non-zero) update for a weight can depend considerably on whether the weight has been updated before. Based on this finding, the paper proposes a new probability estimation scheme: Depending on whether a significant update has been received before (i.e., based on the weight’s history), the entropy coder models the probability for a current significant update differently. This scheme achieves a bitstream size reduction of about 2% and 1% in a transfer and a federated learning scenario, respectively, without any accuracy loss or significant complexity increase. Therefore, MPEG adopted our history dependent significance probability (HDSP) scheme to its emerging standard for INNC. Gerhard Tech, Paul Haase, Daniel Becking, Heiner Kirchhoffer, Karsten Müller 0001, Jonathan Pfaff, Heiko Schwarz, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 5 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Encoder Optimizations For The NNR Standard On Neural Network CompressionabstractThe novel Neural Network Compression and Representation Standard (NNR), recently issued by ISO/IEC MPEG, achieves very high coding gains, compressing neural networks to 5% in size without accuracy loss. The underlying NNR encoder technology includes parameter quantization, followed by efficient arithmetic coding, namely DeepCABAC. In addition, NNR also allows very flexible adaptations, such as signaling specific local scaling values, setting quantization parameters per tensor rather than per network and supporting specific parameter fusion operations. This paper presents our new approach for optimally deriving these parameters, namely the derivation of parameters for local scaling adaptation (LSA), inference-optimized quantization (IOQ), and batch-norm folding (BNF). By allowing inference and fine tuning within the encoding process, quantization errors are reduced and the NNR coding efficiency is further improved to a compressed bitstream size of only 3% in comparison to the original model size. Paul Haase, Daniel Becking, Heiner Kirchhoffer, Karsten Müller 0001, Heiko Schwarz, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 4 |
| 2020 | Dependent Scalar Quantization For Neural Network CompressionabstractRecent approaches to compression of deep neural networks, like the emerging standard on compression of neural networks for multimedia content description and analysis (MPEG-7 part 17), apply scalar quantization and entropy coding of the quantization indexes. In this paper we present an advanced method for quantization of neural network parameters, which applies dependent scalar quantization (DQ) or trellis-coded quantization (TCQ), and an improved context modeling for the entropy coding of the quantization indexes. We show that the proposed method achieves 5.778% bitrate reduction and virtually no loss (0.37%) of network performance in average, compared to the baseline methods of the second test model (NCTM) of MPEG-7 part 17 for relevant working points. Paul Haase, Heiko Schwarz, Heiner Kirchhoffer, Simon Wiedemann, Talmaj Marinc, Arturo Marbán, Karsten Müller 0001, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 7 |
| 2020 | Deepcabac: Plug & Play Compression of Neural Network Weights and Weight UpdatesabstractAn increasing number of distributed machine learning applications require efficient communication of neural network parameterizations. DeepCABAC, an algorithm in the current working draft of the emerging MPEG-7 part 17 standard for compression of neural networks for multimedia content description and analysis, has demonstrated high compression gains for a variety of neural network models. In this paper we propose a method for employing DeepCABAC in a Federated Learning scenario for the exchange of intermediate differential parameterizations. Furthermore, we discuss the efficiency of DeepCABAC when compressing trained neural networks. Our experiments on large neural networks show that in both scenarios, DeepCABAC achieves competitive compression rates, without degrading the network accuracy. David Neumann, Felix Sattler, Heiner Kirchhoffer, Simon Wiedemann, Karsten Müller 0001, Heiko Schwarz, Thomas Wiegand 0001, Detlev Marpe, Wojciech Samek |
ICIP | 5 |
| 2020 | Region-Based Predictors For Intra Block CopyabstractIntra block copy is a prediction technique in intra coding, which has high compression performance for screen content or computer generated videos. Accordingly, this tool has become part of the upcoming video coding standard H.266/VVC (Versatile Video Coding). Intra block copy is analogous to motion-compensated prediction for the usual inter-picture case with the additional constraint that the reference picture is given by the current partially reconstructed picture. Moreover, the entropy coding of the displacement information of the reference block is carried out in the same way as in inter-picture motion compensation. This publication proposes a set of region-based predictors to improve the current intra block copy approach. The proposed method achieves up to -3.81% of Bjøntegaard-Delta bit-rate saving on top of the existing intra block copy for an all-intra picture configuration. Gayathri Venugopal, Santiago De-Luxán-Hernández, Karsten Müller 0001, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 3 |
| 2019 | Hardware-Friendly Intra Region-Based Template Matching for VVCabstractIn a template matching (TM) intra method, the neighboring samples of the current block are regarded as a template. The decoder searches for the best template match in the current reconstructed picture using an error minimizing metric like sum of squared differences (SSD). The prediction signal is generated by copying the samples of the adjacent block to the selected template match. The increased decoder complexity from the search algorithm makes it less attractive for modern applications like video calling. In a previous publication [1], we presented a region-based template matching (RTM) approach for intra coding. Compared to the conventional TM methods which searches for the template match in a complete search window, RTM searches in a region of the search window. Thus, RTM offers a better trade-off between coding efficiency and decoder complexity. Nevertheless, the memory requirements and number of computations to be carried at the decoder are still high, making RTM difficult for hardware realization. This paper aims to address these issues. Gayathri Venugopal, Philipp Helle, Karsten Müller 0001, Detlev Marpe, Thomas Wiegand 0001 |
DCC | 3 |
| 2019 | A Unified Region-Based Template Matching Approach for Intra and Inter Prediction in VVCabstractTemplate matching is a texture synthesis method which has found applications in video coding. However, the increased decoder complexity from its underlying search algorithm makes it less suitable for practical applications. In our previous paper, we presented region-based template matching where the decoder complexity is considerably reduced. This publication describes further improvements of the aforementioned tool for intra prediction and decoder side motion vector derivation. When applied together, it achieves an average coding gain of -1.04% with 107% decoder run-time for all intra and -1.60% with 113% decoder run-time for random access configurations respectively. The proposed method needs considerably fewer computations and memory requirements compared to the previous version of the algorithm. Gayathri Venugopal, Karsten Müller 0001, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 2 |
| 2018 | Partial Depth Image Based Re-Rendering for Synthesized View Distortion Computationabstract3D video systems transmit depth maps in order to render synthesized views (SVs) at a receiver. To anticipate this purpose when processing a depth map, a sender-side depth processing algorithm (DPA), e.g. a depth encoder, can also render the SVs, compute their SV distortion (SVD), and adapt to it. This requires a low-complexity algorithm as computational resources are usually limited. We propose such an algorithm in this paper. First, we discuss a measure that relates a depth change to an SVD change using rendering. Then, we present an optimized process combining basic rendering steps, as warping, occlusion handling, interpolation, hole filling, and blending. Furthermore, we analyze which parts of an SV are affected by a depth change and modify the process to re-render only them. The resulting algorithm is significantly less complex than an unoptimized rendering-based variant and quantifies the SVD more accurately than existing estimation methods. The algorithm is used by the 3D-High Efficiency Video Coding reference software encoder as the main method for distortion computation and can also be used by other DPAs. Gerhard Tech, Karsten Müller 0001, Heiko Schwarz, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Soccer player recognition using spatial constellation features and jersey number recognition
Sebastian Gerke, Antje Linnemann, Karsten Müller 0001 |
Comput. Vis. Image Underst. | 3 |
| 2016 | Efficient no-reference metric for sharpness mismatch artifact between stereoscopic views
Mohan Liu, Karsten Müller 0001, Alexander Raake |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Depth Intra Coding for 3D Video Based on Geometric PrimitivesabstractThis paper presents an advanced depth intra-coding approach for 3D video coding based on the High Efficiency Video Coding (HEVC) standard and the multiview video plus depth (MVD) representation. This paper is motivated by the fact that depth signals have specific characteristics that differ from those of natural signals, i.e., camera-view video. Our approach replaces conventional intra-picture coding for the depth component, targeting a consistent and efficient support of 3D video applications that utilize depth maps or polygon meshes or both, with a high depth coding efficiency in terms of minimal artifacts in rendered views and meshes with a minimal number of triangles for a given bit rate. For this purpose, we introduce intra-picture prediction modes based on geometric primitives along with a residual coding method in the spatial domain, substituting conventional intra-prediction modes and transform coding, respectively. The results show that our solution achieves the same quality of rendered or synthesized views with about the same bit rate as MVD coding with the 3D video extension of HEVC (3D-HEVC) for high-quality depth maps and with about 8% less overall bit rate as with 3D-HEVC without related depth tools. At the same time, the combination of 3D video with 3D computer graphics content is substantially simplified, as the geometry-based depth intra signals can be represented as a surface mesh with about 85% less triangles, generated directly in the decoding process as an alternative decoder output. Philipp Merkle, Karsten Müller 0001, Detlev Marpe, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Overview of the Multiview and 3D Extensions of High Efficiency Video CodingabstractThe High Efficiency Video Coding (HEVC) standard has recently been extended to support efficient representation of multiview video and depth-based 3D video formats. The multiview extension, MV-HEVC, allows efficient coding of multiple camera views and associated auxiliary pictures, and can be implemented by reusing single-layer decoders without changing the block-level processing modules since block-level syntax and decoding processes remain unchanged. Bit rate savings compared with HEVC simulcast are achieved by enabling the use of inter-view references in motion-compensated prediction. The more advanced 3D video extension, 3D-HEVC, targets a coded representation consisting of multiple views and associated depth maps, as required for generating additional intermediate views in advanced 3D displays. Additional bit rate reduction compared with MV-HEVC is achieved by specifying new block-level video coding tools, which explicitly exploit statistical dependencies between video texture and depth and specifically adapt to the properties of depth maps. The technical concepts and features of both extensions are presented in this paper. Gerhard Tech, Ying Chen 0011, Karsten Müller 0001, Jens-Rainer Ohm, Anthony Vetro, Ye-Kui Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | HEVC-Compatible Extensions for Advanced Coding of 3D and Multiview VideoabstractThis article provides an overview of standardized extensions of HEVC for the advanced coding of 3D and multiview video. In those extensions, new coding tools that better exploit the inter-view redundancy of the multiview texture videos have been developed. Additionally, dedicated tools for the improved coding of depth have been extensively studied and incorporated into the standard. In this paper, the performance of these extensions is assessed, and experimental results demonstrate notable gains in coding efficiency. Anthony Vetro, Ying Chen 0011, Karsten Müller 0001 |
DCC | 3 |
| 2015 | Fast image completion method using patch offset statisticsabstractImage completion is a technique that involves repairing damaged images or filling in missing regions in a visually pleasing manner. In this work, a novel fast image completion approach is proposed, which finds the best textures to be filled into the unknown areas through internal statistics of natural images. Firstly, we create statistics from spatial offsets (relative positions) of similar patches in the image. From these statistics, a sparse representation of the most frequent offsets is derived. These offsets are then used to fill the missing region by combining a stack of shifted images using a weighted means method. Hereby, the offsets are considered in descending order in the filling routine according to their importance. Experimental results show that the proposed method yields subjective improvements compared to the state-of-the-art. Furthermore, the method synthesizes missing regions faster than existing methods. Martin Köppel, Mehdi Ben Makhlouf, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP | 3 |
| 2013 | Coding of depth signals for 3D video using wedgelet block segmentation with residual adaptationabstractThis paper presents a new approach for the depth coding part of a 3D video coding extension based on the Multiview Video plus Depth (MVD) representation. Our approach targets a higher coding efficiency for the depth component and is motivated by the fact that depth signals have specific characteristics that differ from video. For this purpose we apply the method of wedgelet segmentation with residual adaptation for depth blocks by implementing a new set of coding and prediction modes and by optimizing the algorithms for efficient processing and signaling. The results show that a bit rate reduction of up to 6% is achieved for the depth component, using a 3D video codec based on the high-efficiency video coding (HEVC) technology. Apart from the depth coding gains, wedgelets lead to a considerably better rendered view quality. Philipp Merkle, Karsten Müller 0001, Thomas Wiegand 0001 |
ICME | 2 |
| 2013 | Edge aware disparity estimation for intermediate view synthesisabstractIn this paper we present a new local and multiscale disparity estimation algorithm. Results show that the proposed method can preserve arbitrarily shaped depth discontinuities. Some results obtained from this method are shown as depth maps and synthetic views produced using the estimated disparity maps. Miquel A. Farre, Haricharan Lakshman, Philipp Helle, Heiko Schwarz, Karsten Müller 0001, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 5 |
| 2013 | 3D High-Efficiency Video Coding for Multi-View Video and Depth DataabstractThis paper describes an extension of the high efficiency video coding (HEVC) standard for coding of multi-view video and depth data. In addition to the known concept of disparity-compensated prediction, inter-view motion parameter, and inter-view residual prediction for coding of the dependent video views are developed and integrated. Furthermore, for depth coding, new intra coding modes, a modified motion compensation and motion vector coding as well as the concept of motion parameter inheritance are part of the HEVC extension. A novel encoder control uses view synthesis optimization, which guarantees that high quality intermediate views can be generated based on the decoded data. The bitstream format supports the extraction of partial bitstreams, so that conventional 2D video, stereo video, and the full multi-view video plus depth format can be decoded from a single bitstream. Objective and subjective results are presented, demonstrating that the proposed approach provides 50% bit rate savings in comparison with HEVC simulcast and 20% in comparison with a straightforward multi-view extension of HEVC without the newly developed coding tools. Karsten Müller 0001, Heiko Schwarz, Detlev Marpe, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Extension of High Efficiency Video Coding (HEVC) for multiview video and depth dataabstractThis paper presents an approach for 3D video coding that uses a format in which a small number of views as well as associated depth maps are coded and transmitted. At the receiver side, additional views required for displaying the 3D video on an autostereoscopic display can be generated based on the corresponding decoded signals by using depth image based rendering (DIBR) techniques. In terms of coding technology, the proposed coding scheme represents an extension of High Efficiency Video Coding (HEVC), similar to the Multiview Coding (MVC) extension of H.264/AVC. Besides the well-known disparity-compensated prediction, advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been developed and integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% average bit rate savings for a whole test set when configured to comply with a 2- and 3-view scenario, respectively. The proposed codec was submitted as response to a Call for Proposals on 3D Video Technology issued by the ISO/IEC Moving Picture Experts Group (MPEG) and it was ranked as the overall best performing HEVC-based proposal in the related subjective tests. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 8 |
| 2012 | Synthesized View Distortion Based 3D Video Coding for Extrapolation and Interpolation of ViewsabstractIn 3D video coding, the Multi-View Video plus Depth (MVD) representation format enables the interpolation of virtual intermediate views in between the original recorded camera views, as well as the extrapolation of views outside this range using view synthesis algorithms. For this, MVD includes depth maps as additional information to be transmitted. The performance of the associated depth data coding can be improved by considering the distortion of synthesized views in the rate-distortion optimization of the encoder. Therefore, this paper evaluates coding gains achieved with encoder optimization with respect to virtual view positions. The impact of the number and the positions of virtual views used as reference in the encoder side optimization process on the quality of the synthesized views is analyzed. Coding results for view interpolation and view extrapolation are provided. Moreover, the encoder-side complexity with respect to the number of reference views is discussed. Gerhard Tech, Heiko Schwarz, Karsten Müller 0001, Thomas Wiegand 0001 |
ICME | 3 |
| 2012 | 3D video: Depth coding based on inter-component prediction of block partitionsabstractThis paper presents a new approach for 3D video coding, where the video and the depth component of an MVD representation are jointly coded in an integrated framework. This enables a new type of prediction for exploiting the correlations of video and depth signals in addition to existing methods for temporal and inter-view prediction. Our new method is referred to as inter-component prediction and we adopt it for predicting non-rectangular partitions in depth blocks. By dividing the block into two regions, each represented with a constant value, such block partitions are well-adapted to the characteristics of depth maps. The results show that this approach reduces the bit rate of the depth component by up to 11% and leads to an increased quality of rendered views. Philipp Merkle, Christian Bartnik, Karsten Müller 0001, Detlev Marpe, Thomas Wiegand 0001 |
PCS | 3 |
| 2012 | 3D video coding using advanced prediction, depth modeling, and encoder control methodsabstractThe presented approach for 3D video coding uses the multiview video plus depth format, in which a small number of video views as well as associated depth maps are coded. Based on the coded signals, additional views required for displaying the 3D video on an autostereoscopic display can be generated by depth image based rendering techniques. The developed coding scheme represents an extension of HEVC, similar to the MVC extension of H.264/AVC. However, in addition to the well-known disparity-compensated prediction advanced techniques for inter-view and inter-component prediction, the representation of depth blocks, and the encoder control for depth signals have been integrated. In comparison to simulcasting the different signals using HEVC, the proposed approach provides about 40% and 50% bit rate savings for the tested configurations with 2 and 3 views, respectively. Bit rate reductions of about 20% have been obtained in comparison to a straightforward multiview extension of HEVC without the newly developed coding tools. Heiko Schwarz, Christian Bartnik, Sebastian Bosse, Heribert Brust, Tobias Hinz, Haricharan Lakshman, Detlev Marpe, Philipp Merkle, Karsten Müller 0001, Hunn Rhee, Gerhard Tech, Martin Winken, Thomas Wiegand 0001 |
PCS | 9 |
| 2012 | 3D video coding using the synthesized view distortion changeabstractIn 3D video, texture and supplementary depth data are coded to enable the interpolation of a required number of synthesized views for multi-view displays in the range of the original camera views. The coding of the depth data can be improved by analyzing the distortion of synthesized video views instead of the depth map distortion. Therefore, this paper introduces a new distortion metric for 3D video coding, which relates changes in the depth map directly to changes of the overall synthesized view distortion. It is shown how the new metric can be integrated into the rate-distortion optimization (RDO) process of an encoder, that is based on high-efficiency video coding technology. An evaluation of the modified encoder is conducted using different view synthesis algorithms and shows about 50% rate savings for the depth data or 0.6 dB PSNR gains for the synthesized view. Gerhard Tech, Heiko Schwarz, Karsten Müller 0001, Thomas Wiegand 0001 |
PCS | 3 |
| 2011 | Challenges in 3D video standardizationabstractStereoscopic video transmission systems have now evolved from 2D video systems and have been commercialized for a number of application areas, driven by developments in stereo capturing and display technology. With the new developments in autostereoscopic display technology, these stereo systems need to further advance towards 3D video systems. In contrast to all previous video coding technologies, 3D video data requires a number of new assumptions and novel technology developments. This paper discusses the evolution from 2D to stereo video and finally to 3D video systems and highlights some of the major new challenges for 3D Video. Finally, an evaluation framework for 3D video technology is shown, that addresses these challenges and is used by ISO- MPEG for standardizing the best 3D video coding solution. Karsten Müller 0001, Philipp Merkle |
VCIP | 1 |
| 2011 | 3-D Video Representation Using Depth MapsabstractCurrent 3-D video (3DV) technology is based on stereo systems. These systems use stereo video coding for pictures delivered by two input cameras. Typically, such stereo systems only reproduce these two camera views at the receiver and stereoscopic displays for multiple viewers require wearing special 3-D glasses. On the other hand, emerging autostereoscopic multiview displays emit a large numbers of views to enable 3-D viewing for multiple users without requiring 3-D glasses. For representing a large number of views, a multiview extension of stereo video coding is used, typically requiring a bit rate that is proportional to the number of views. However, since the quality improvement of multiview displays will be governed by an increase of emitted views, a format is needed that allows the generation of arbitrary numbers of views with the transmission bit rate being constant. Such a format is the combination of video signals and associated depth maps. The depth maps provide disparities associated with every sample of the video signal that can be used to render arbitrary numbers of additional views via view synthesis. This paper describes efficient coding methods for video and depth data. For the generation of views, synthesis methods are presented, which mitigate errors from depth estimation and coding. Karsten Müller 0001, Philipp Merkle, Thomas Wiegand 0001 |
Proc. IEEE | 1 |
| 2011 | Depth Image-Based Rendering With Advanced Texture Synthesis for 3-D VideoabstractA depth image-based rendering (DIBR) approach with advanced inpainting methods is presented. The DIBR algorithm can be used in 3-D video applications to synthesize a number of different perspectives of the same scene, e.g., from a multiview-video-plus-depth (MVD) representation. This MVD format consists of video and depth sequences for a limited number of original camera views of the same natural scene. Here, DIBR methods allow the computation of additional new views. An inherent problem of the view synthesis concept is the fact that image information which is occluded in the original views may become visible, especially in extrapolated views beyond the viewing range of the original cameras. The presented algorithm synthesizes these occluded textures. The synthesizer achieves visually satisfying results by taking spatial and temporal consistency measures into account. Detailed experiments show significant objective and subjective gains of the proposed method in comparison to the state-of-the-art methods. Patrick Ndjiki-Nya, Martin Köppel, Dimitar Doshkov, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Thomas Wiegand 0001 |
IEEE Trans. Multim. | 6 |
| 2010 | Temporally consistent handling of disocclusions with texture synthesis for depth-image-based renderingabstractDepth-image-based rendering (DIBR) is used to generate additional views of a real-world scene from images or videos and associated per-pixel depth information. An inherent problem of the view synthesis concept is the fact that image information which is occluded in the original view may become visible in the “virtual” image. The resulting question is: how can these disocclusions be covered in a visually plausible manner? In this paper, a new temporally and spatially consistent hole filling method for DIBR is presented. In a first step, disocclusions in the depth map are filled. Then, a background sprite is generated and updated with every frame using the original and synthesized information from previous frames to achieve temporally consistent results. Next, small holes resulting from depth estimation inaccuracies are closed in the textured image, using methods that are based on solving Laplace equations. The residual disoccluded areas are coarsely initialized and subsequently refined by patch-based texture synthesis. Experimental results are presented, highlighting that gains in objective and visual quality can be achieved in comparison to the latest MPEG view synthesis reference software (VSRS). Martin Köppel, Patrick Ndjiki-Nya, Dimitar Doshkov, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP | 6 |
| 2010 | Correlation histogram analysis of depth-enhanced 3D video codingabstractThis paper introduces a correlation histogram method for analyzing the different components of depth-enhanced 3D video representations. Depth-enhanced 3D representations such as multi-view video plus depth consist of two components: video and depth map sequences. As depth maps represent the scene geometry, their characteristics differ from the video data. We present a comparative analysis that identifies the significant characteristics of the two components via correlation histograms. These characteristics are of special importance for compression. Modern video codecs like H.264/AVC are highly optimized to the statistical properties of natural video. Therefore the effect of compressing the two components using the MVC extension of H.264/AVC is evaluated in the second part of the analysis. The presented results show that correlation histograms are a powerful and well-suited method for analyzing the impact of processing on the characteristics of depth-enhanced 3D video. Philipp Merkle, Jordi Bayo Singla, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP | 3 |
| 2010 | 3D video formats and coding methodsabstractThe introduction of first 3D systems for digital cinema and home entertainment is based on stereo technology. For efficiently supporting new display types, depth-enhanced formats and coding technology is required, as introduced in this overview paper. First, we discuss the necessity for a generic 3D video format, as the current state-of-the-art in multi-view video coding cannot support different types of multi-view displays at the same time. Therefore, a generic depth-enhanced 3D format is developed, where any number of views can be generated from one bit stream. This, however, requires a complex framework for 3D video, where not only the 3D format and new coding methods are investigated, but also view synthesis and the provision of high-quality depth maps, e.g. via depth estimation. We present this framework and discuss the interdependencies between the different modules. Karsten Müller 0001, Philipp Merkle, Gerhard Tech, Thomas Wiegand 0001 |
ICIP | 1 |
| 2010 | Depth image based rendering with advanced texture synthesisabstractIn free viewpoint television or 3D video, depth image based rendering (DIBR) is used to generate virtual views based on a textured image and its associated depth information. In doing so, image regions which are occluded in the original view may become visible in the virtual image. One of the main challenges in DIBR is to extrapolate known textures into the disoccluded area without inserting subjective annoyance. In this paper, a new hole filling approach for DIBR using texture synthesis is presented. Initially, the depth map in the virtual view is filled at disoccluded locations. Then, in the textured image, holes of limited spatial extent are closed by solving Laplace equations. Larger disoccluded regions are initialized via median filtering and subsequently refined by patch-based texture synthesis. Experimental results show that the proposed approach provides improved rendering results in comparison to the latest MPEG view synthesis reference software (VSRS) version 3.6. Patrick Ndjiki-Nya, Martin Köppel, Dimitar Doshkov, Haricharan Lakshman, Philipp Merkle, Karsten Müller 0001, Thomas Wiegand 0001 |
ICME | 6 |
| 2010 | Immersive future media technologies: from 3D video to sensory experiencesabstractIn this tutorial we present immersive future media technologies ranging from 3D video to sensory experiences. The former targets stereo and multi-view video technologies whereas the latter aims at stimulating other senses than vision or audition enabling an advanced user experiences through sensory effects. Christian Timmerer, Karsten Müller 0001 |
ACM Multimedia | 2 |
| 2010 | Diffusion filtering of depth maps in stereo video codingabstractA method for removing irrelevant information from depth maps in Video plus Depth coding is presented. The depth map is filtered in several iterations using a diffusional approach. In each iteration smoothing is carried out in local sample neighborhoods considering the distortion introduced to a rendered view. Smoothing is only applied when the rendered view is not affected. Therefore irrelevant edges and features in the depth map can be damped while the quality of the rendered view is retained. The processed depth maps can be coded at a reduced rate compared to unaltered data. Coding experiments show gains up to 0.5 dB for the rendered view at the same bit rate. Gerhard Tech, Karsten Müller 0001, Thomas Wiegand 0001 |
PCS | 2 |
| 2010 | 3D video coding: an overview of present and upcoming standardsabstractAn overview of existing and upcoming 3D video coding standards is given. Various different 3D video formats are available, each with individual pros and cons. The 3D video formats can be separated into two classes: video-only formats (such as stereo and multiview video) and depth-enhanced formats (such as video plus depth and multiview video plus depth). Since all these formats exist of at least two video sequences and possibly additional depth data, efficient compression is essential for the success of 3D video applications and technologies. For the video-only formats the H.264 family of coding standards already provides efficient and widely established compression algorithms: H.264/AVC simulcast, H.264/AVC stereo SEI message, and H.264/MVC. For the depth-enhanced formats standardized coding algorithms are currently being developed. New and specially adapted coding approaches are necessary, as the depth or disparity information included in these formats has significantly different characteristics than video and is not displayed directly, but used for rendering. Motivated by evolving market needs, MPEG has started an activity to develop a generic 3D video standard within the 3DVC ad-hoc group. Key features of the standard are efficient and flexible compression of depth-enhanced 3D video representations and decoupling of content creation and display requirements. Philipp Merkle, Karsten Müller 0001, Thomas Wiegand 0001 |
VCIP | 2 |
| 2010 | Multi-camera imaging, coding and innovative display: techniques and systems
Minh N. Do, Chang-Su Kim 0001, Karsten Müller 0001, Masayuki Tanimoto, Anthony Vetro |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Coding and intermediate view synthesis of multiview video plus depthabstractFor advanced 3D Video (3DV) applications, efficient data representations are investigated, which only transmit a subset of the views that are required for 3D visualization. From this subset, all intermediate views are synthesized from sample-dense color and depth data. In this paper, the method for reliability-based view synthesis from compressed multi-view + depth data (MVD) is investigated and corresponding results are shown. The initial problem in such 3DV systems is the interdependency between view capturing, coding and view synthesis. For evaluating each component separately, we first generate results from the coding stage only, where color and depth coding is carried out separately. In the next step, we add the view synthesis stage with reliability-based view synthesis and show, how the separate coding results influence the view synthesis quality and what type of artifacts are produced. Efficient bit rate distribution between color and depth is investigated by objective as well as subjective evaluations. Furthermore, quality characteristics across the viewing range for different bit rate distributions are analyzed. Finally, the robustness of the reliability-based view synthesis to coding artifacts is presented. Karsten Müller 0001, Aljoscha Smolic, Kristina Dix, Philipp Merkle, Thomas Wiegand 0001 |
ICIP | 1 |
| 2009 | An overview of available and emerging 3D video formats and depth enhanced stereo as efficient generic solutionabstractRecently, popularity of 3D video has been growing significantly and it may turn into a home user mass market in the near future. However, diversity of 3D video content formats is still hampering wide success. An overview of available and emerging 3D video formats and standards is given, which are mostly related to specific types of applications and 3D displays. This includes conventional stereo video, multiview video, video plus depth, multiview video plus depth and layered depth video. Features and limitations are explained. Finally, depth enhanced stereo (DES) is introduced as a flexible, generic, and efficient 3D video format that can unify all others and serve as universal 3D video format in the future. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Peter Kauff, Thomas Wiegand 0001 |
PCS | 2 |
| 2009 | The effects of multiview depth video compression on multiview rendering
Philipp Merkle, Yannick Morvan, Aljoscha Smolic, Dirk Farin, Karsten Müller 0001, Peter H. N. de With, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 5 |
| 2008 | Efficient representation and coding of prediction residuals and parameters in frame-based animated mesh compressionabstractFor compression of 3-D dynamic meshes, the novel framework of so-called frame-based animated mesh compression (FAMC) has been introduced recently. In this context, we propose an efficient scheme for representation and statistical coding which is conceptually based on our previous work on context-based adaptive binary arithmetic coding (CABAC). After reviewing the basic principles of both CABAC and FAMC, we present suitable modifications and adaptations of both concepts in order to build an integrated solution with a high degree of coding efficiency. To this end, particular focus of our study has been put on the design of appropriate binarization and context modeling schemes. In our experiments, we obtained average bit-rate savings of more than 30% for a typical test set of dynamic meshes, when comparing the final design of our CABAC enriched FAMC scheme to the original version of FAMC using a conventional N-ary arithmetic coder. Our integrated approach has been adopted recently as part of the MPEG-4 Animated Framework extension (AFX). Detlev Marpe, Heiner Kirchhoffer, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP | 3 |
| 2008 | Intermediate view interpolation based on multiview video plus depth for advanced 3D video systemsabstractA system for video on multiscopic 3D displays is considered where the data representation consists of multiview video plus scene depth. At most, 3 multiview video signals are being transmitted and used together with the depth data to generate intermediate views at the receiver. The paper presents an approach to such an intermediate view interpolation that separates unreliable image regions along depth discontinuities from reliable image regions. These image regions are processed with different algorithms and then fused to obtain the final interpolated view. In contrast to previous layered approaches, two boundary layers and one reliable layer is used. Moreover, the presented technique does not rely on 3D graphics support but uses image-based 3D warping instead. For enhanced quality intermediate view generation, hole-filling and filtering methods are described. As a result, high quality intermediate views for an existing 9-view auto-stereoscopic display are presented, which prove the suitability of the approach for advanced 3D video (3DV) systems. Aljoscha Smolic, Karsten Müller 0001, Kristina Dix, Philipp Merkle, Peter Kauff, Thomas Wiegand 0001 |
ICIP | 2 |
| 2008 | Context-adaptive binary arithmetic coding for frame-based animated mesh compressionabstractContext-based adaptive binary arithmetic coding (CABAC) has proven to be an efficient technique in the area of video coding. This paper presents an approach for integrating CABAC into the framework of frame-based animated mesh compression (FAMC). It presents the specific modifications and adaptations that have been worked out to adapt CABAC to the specific requirements of FAMC in order to build a solution with a higher degree of coding efficiency. For a typical test set of animated meshes, average bit rate savings of 25% have been observed for the combination of CABAC and FAMC when compared to a previous version of FAMC using a conventional N-ary arithmetic coder. The presented approach has been recently adopted as part of the MPEG-4 AFX standard. Heiner Kirchhoffer, Detlev Marpe, Karsten Müller 0001, Thomas Wiegand 0001 |
ICME | 3 |
| 2008 | Reliability-based generation and view synthesis in layered depth videoabstractIn this paper, a system for video rendering on multiscopic 3D displays is considered where the data is represented as layered depth video (LDV). This representation consists of one full or central video with associated per-pixel depth and additional residual layers. Thus, only one full view with additional residual data needs to be transmitted. The LDV data is used at the receiver to generate all intermediate views for the display. The paper presents the LDV layer extraction as well as the view synthesis, using a scene reliability-driven approach. Here, unreliable image regions are detected and in contrast to previous approaches the residual data is enlarged to reduce artifacts in unreliable areas during rendering. To provide maximum data coverage, the residual data remains at its original positions and will not be projected towards the central view. The view synthesis process also uses this reliability analysis to provide higher quality intermediate views than previous approaches. As a final result, high quality intermediate views for an existing 9-view auto-stereoscopic display are presented, which prove the suitability of the LDV approach for advanced 3D video (3DV) systems. Karsten Müller 0001, Aljoscha Smolic, Kristina Dix, Peter Kauff, Thomas Wiegand 0001 |
MMSP | 1 |
| 2007 | Multi-View Video Plus Depth Representation and CodingabstractA study on the video plus depth representation for multi-view video sequences is presented. Such a 3D representation enables functionalities like 3D television and free viewpoint video. Compression is based on algorithms for multi-view video coding, which exploit statistical dependencies from both temporal and inter-view reference pictures for prediction of both color and depth data. Coding efficiency of prediction structures with and without inter-view reference pictures is analyzed for multi-view video plus depth data, reporting gains in luma PSNR of up to 0.5 dB for depth and 0.3 dB for color. The main benefit from using a multi-view video plus depth representation is that intermediate views can be easily rendered. Therefore the impact on image quality of rendered arbitrary intermediate views is investigated and analyzed in a second part, comparing compressed multi-view video plus depth data at different bit rates with the uncompressed original. Philipp Merkle, Aljoscha Smolic, Karsten Müller 0001, Thomas Wiegand 0001 |
ICIP (1) | 3 |
| 2007 | Scene Representation Technologies for 3DTV - A Surveyabstract3-D scene representation is utilized during scene extraction, modeling, transmission and display stages of a 3DTV framework. To this end, different representation technologies are proposed to fulfill the requirements of 3DTV paradigm. Dense point-based methods are appropriate for free-view 3DTV applications, since they can generate novel views easily. As surface representations, polygonal meshes are quite popular due to their generality and current hardware support. Unfortunately, there is no inherent smoothness in their description and the resulting renderings may contain unrealistic artifacts. NURBS surfaces have embedded smoothness and efficient tools for editing and animation, but they are more suitable for synthetic content. Smooth subdivision surfaces, which offer a good compromise between polygonal meshes and NURBS surfaces, require sophisticated geometry modeling tools and are usually difficult to obtain. One recent trend in surface representation is point-based modeling which can meet most of the requirements of 3DTV, however the relevant state-of-the-art is not yet mature enough. On the other hand, volumetric representations encapsulate neighborhood information that is useful for the reconstruction of surfaces with their parallel implementations for multiview stereo algorithms. Apart from the representation of 3-D structure by different primitives, texturing of scenes is also essential for a realistic scene rendering. Image-based rendering techniques directly render novel views of a scene from the acquired images, since they do not require any explicit geometry or texture representation. 3-D human face and body modeling facilitate the realistic animation and rendering of human figures that is quite crucial for 3DTV that might demand real-time animation of human bodies. Physically based modeling and animation techniques produce impressive results, thus have potential for use in a 3DTV framework for modeling and animating dynamic scenes. As a concluding remark, it can be argued that 3-D scene and texture representation techniques are mature enough to serve and fulfill the requirements of 3-D extraction, transmission and display sides in a 3DTV scenario. A. Aydin Alatan, Yücel Yemez, Ugur Güdükbay, Xenophon Zabulis, Karsten Müller 0001, Çigdem Eroglu Erdem, C. Weigel, Aljoscha Smolic |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2007 | Efficient Prediction Structures for Multiview Video CodingabstractAn experimental analysis of multiview video coding (MVC) for various temporal and inter-view prediction structures is presented. The compression method is based on the multiple reference picture technique in the H.264/AVC video coding standard. The idea is to exploit the statistical dependencies from both temporal and inter-view reference pictures for motion-compensated prediction. The effectiveness of this approach is demonstrated by an experimental analysis of temporal versus inter-view prediction in terms of the Lagrange cost function. The results show that prediction with temporal reference pictures is highly efficient, but for 20% of a picture's blocks on average prediction with reference pictures from adjacent views is more efficient. Hierarchical B pictures are used as basic structure for temporal prediction. Their advantages are combined with inter-view prediction for different temporal hierarchy levels, starting from simulcast coding with no inter-view prediction up to full level inter-view prediction. When using inter-view prediction at key picture temporal levels, average gains of 1.4-dB peak signal-to-noise ratio (PSNR) are reported, while additionally using inter-view prediction at nonkey picture temporal levels, average gains of 1.6-dB PSNR are reported. For some cases, gains of more than 3 dB, corresponding to bit-rate savings of up to 50%, are obtained. Philipp Merkle, Aljoscha Smolic, Karsten Müller 0001, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Coding Algorithms for 3DTV - A SurveyabstractResearch efforts on 3DTV technology have been strengthened worldwide recently, covering the whole media processing chain from capture to display. Different 3DTV systems rely on different 3D scene representations that integrate various types of data. Efficient coding of these data is crucial for the success of 3DTV. Compression of pixel-type data including stereo video, multiview video, and associated depth or disparity maps extends available principles of classical video coding. Powerful algorithms and open international standards for multiview video coding and coding of video plus depth data are available and under development, which will provide the basis for introduction of various 3DTV systems and services in the near future. Compression of 3D mesh models has also reached a high level of maturity. For static geometry, a variety of powerful algorithms are available to efficiently compress vertices and connectivity. Compression of dynamic 3D geometry is currently a more active field of research. Temporal prediction is an important mechanism to remove redundancy from animated 3D mesh sequences. Error resilience is important for transmission of data over error prone channels, and multiple description coding (MDC) is a suitable way to protect data. MDC of still images and 2D video has already been widely studied, whereas multiview video and 3D meshes have been addressed only recently. Intellectual property protection of 3D data by watermarking is a pioneering research area as well. The 3D watermarking methods in the literature are classified into three groups, considering the dimensions of the main components of scene representations and the resulting components after applying the algorithm. In general, 3DTV coding technology is maturating. Systems and services may enter the market in the near future. However, the research area is relatively young compared to coding of other types of media. Therefore, there is still a lot of room for improvement and new development of algorithms. Aljoscha Smolic, Karsten Müller 0001, Nikolce Stefanoski, Jörn Ostermann, Atanas P. Gotchev, Gozde Bozdagi Akar, George A. Triantafyllidis, Alper Koz |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Rate-Distortion Optimization in Dynamic Mesh CompressionabstractRecent developments in the compression of dynamic meshes or mesh sequences have shown that the statistical dependencies within a mesh sequence can be exploited well by predictive coding approaches. Coders introduced so far use experimentally determined or heuristic thresholds for tuning the algorithms. In video coding rate-distortion (RD) optimization is often used to avoid fixing of thresholds and to select a coding mode. We applied these ideas and present here an RD-optimized mesh coder. It includes different prediction modes as well as an RD cost computation that controls the mode selection across all possible spatial partitions of a mesh to find the clustering structure together with the associated prediction modes. The structure of the RD-optimized D3DMC coder is presented, followed by comparative results with mesh sequences at different resolutions. Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Thomas Wiegand 0001 |
ICIP | 1 |
| 2006 | Increasing the Accuracy of the Space-Sweeping Approach to Stereo Reconstruction, using Spherical Backprojection SurfacesabstractIn this paper, interest is focused on the accurate and time-efficient stereo reconstruction, for the purpose of generating 3D animated scenes from multiple synchronized videos. The plane-sweeping approach is reviewed as relevant to the goal of time-efficiency, since its execution can be optimized on a GPU. A method compatible for optimization on the GPU is proposed as a more accurate alternative to plane sweeping and to the derived visibility computation. The method is compared to plane sweeping as to its accuracy, by evaluating the backprojected 3D model against independent views and using n-fold cross validation to estimate the peak signal to noise ratio (PSNR). Finally, the method's output is casted integratable with multicamera stereo reconstruction frameworks. Xenophon Zabulis, Georgios Kordelas, Karsten Müller 0001, Aljoscha Smolic |
ICIP | 3 |
| 2006 | Efficient Compression of Multi-View Video Exploiting Inter-View Dependencies Based on H.264/MPEG4-AVCabstractEfficient multi-view coding requires coding algorithms that exploit temporal, as well as inter-view dependencies between adjacent cameras. Based on a spatiotemporal analysis on the multi-view data set, we present a coding scheme utilizing an H.264/MPEG4-AVC codec. To handle the specific requirements of multi-view datasets, namely temporal and inter-view correlation, two main features of the coder are used: hierarchical B pictures for temporal dependencies and an adapted prediction scheme to exploit inter-view dependencies. Both features are set up in the H.264/MPEG4-AVC configuration file, such that coding and decoding is purely based on standardized software. Additionally, picture reordering before coding to optimize coding efficiency and inverse reordering after decoding to obtain individual views are applied. Finally, coding results are shown for the proposed multi-view coder and compared to simulcast anchor and simulcast hierarchical B picture coding Philipp Merkle, Karsten Müller 0001, Aljoscha Smolic, Thomas Wiegand 0001 |
ICME | 2 |
| 2006 | 3D Video and Free Viewpoint Video - Technologies, Applications and MPEG StandardsabstractAn overview of 3D and free viewpoint video is given in this paper with special focus on related standardization activities in MPEG. Free viewpoint video allows the user to freely navigate within real world visual scenes, as known from virtual worlds in computer graphics. Examples are shown, highlighting standards conform realization using MPEG-4. Then the principles of 3D video are introduced providing the user with a 3D depth impression of the observed scene. Example systems are described again focusing on their realization based on MPEG-4. Finally multi-view video coding is described as a key component for 3D and free viewpoint video systems. The conclusion is that the necessary technology including standard media formats for 3D and free viewpoint is available or will be available in the near future, and that there is a clear demand from industry and user side for such applications. 3D TV at home and free viewpoint video on DVD will be available soon, and will create huge new markets Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Christoph Fehn, Peter Kauff, Peter Eisert, Thomas Wiegand 0001 |
ICME | 2 |
| 2006 | Rate-distortion-optimized predictive compression of dynamic 3D mesh sequences
Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
Signal Process. Image Commun. | 1 |
| 2005 | Predictive compression of dynamic 3D meshesabstractAn efficient algorithm for compression of dynamic time-consistent 3D meshes is presented. Such a sequence of meshes contains a large degree of temporal statistical dependencies that can be exploited for compression using DPCM. The vertex positions are predicted at the encoder from a previously decoded mesh. The difference vectors are further clustered in an octree approach. Only a representative for a cluster of difference vectors is further processed providing a significant reduction of data rate. The representatives are scaled and quantized and finally entropy coded using CABAC, the arithmetic coding technique used in H.264/MPEG4-AVC. The mesh is then reconstructed at the encoder for prediction of the next mesh. In our experiments we compare the efficiency of the proposed algorithm in terms of bit-rate and quality compared to static mesh coding and interpolator compression indicating a significant improvement in compression efficiency. Karsten Müller 0001, Aljoscha Smolic, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP (1) | 1 |
| 2005 | 3-D Reconstruction of a Dynamic Environment With a Fully Calibrated Background for Traffic ScenesabstractVision-based traffic surveillance systems are more and more employed for traffic monitoring, collection of statistical data and traffic control. We present an extension of such a system that additionally uses the captured image content for 3-D scene modeling and reconstruction. A basic goal of surveillance systems is to get a good coverage of the observed area with as few cameras as possible to keep the costs low. Therefore, the 3-D reconstruction has to be done from only a few original views with limited overlap and different lighting conditions. To cope with these specific restrictions we developed a model-based 3-D reconstruction scheme that exploits a priori knowledge about the scene. The system is fully calibrated offline by estimating camera parameters from measured 3-D-2-D correspondences. Then the scene is divided into static parts, which are modeled offline and dynamic parts, which are processed online. Therefore, we segment all views into moving objects and static background. The background is modeled as multitexture planes using the original camera textures. Moving objects are segmented and tracked in each view. All segmented views of a moving object are combined to a 3-D object, which is positioned and tracked in 3-D. Here we use predefined geometric primitives and map the original textures onto them. Finally the static and dynamic elements are combined to create the reconstructed 3-D scene, where the user can freely navigate, i.e., choose an arbitrary viewpoint and direction. Additionally, the system allows analyzing the 3-D properties of the scene and the moving objects. Karsten Müller 0001, Aljoscha Smolic, Michael Drose, Patrick Voigt, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Free viewpoint video extraction, representation, coding, and renderingabstractFree viewpoint video provides the possibility to freely navigate within dynamic real world video scenes by choosing arbitrary viewpoints and view directions. So far, related work only considered free viewpoint video extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient free viewpoint video extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are coded using an H.264/AVC codec. Our algorithms for view-dependent texture mapping have been adopted as an extension of MPEG-4 AFX. The presented results illustrate that based on the proposed methods a complete transmission system for efficient free viewpoint video can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
ICIP | 2 |
| 2004 | Representation, coding, and rendering of 3D video objects with MPEG-4 and H.264/AVCabstract3D video objects provide the same functionalities as virtual computer graphics objects but depict the motion and appearance of real world moving objects. They can be viewed interactively from any direction and integrated in complete 3D scenes with other virtual and real world elements. So far, related work only considered extraction, representation, and rendering methods. Compression and transmission has not yet been studied in detail and combined with the other components into one complete system. In this paper, we present such a complete system for efficient 3D video object extraction, representation, coding, and interactive rendering. Data representation is based on 3D mesh models and view-dependent texture mapping using video textures. The geometry extraction is based on a shape-from-silhouette algorithm. The resulting voxel models are converted into 3D meshes that are coded using MPEG-4 SNHC tools. The corresponding video textures are preprocessed taking the object's shape into account and coded using an H.264/AVC codec. The presented results illustrate that based on the proposed methods a complete transmission system for 3D video objects can be built. Aljoscha Smolic, Karsten Müller 0001, Philipp Merkle, Tobias Rein, Matthias Kautzner, Peter Eisert, Thomas Wiegand 0001 |
MMSP | 2 |
| 2003 | Multi-texture modeling of 3D traffic scenesabstractWe present a system for 3D reconstruction of traffic scenes. Traffic surveillance is a challenging scenario for 3D reconstruction in cases, where only a small number of views is available that do not contain much overlap. We address the possibilities and restrictions for modeling such scenarios with only a few cameras and introduce a compositor that allows rendering of the semi automatically generated 3D scenes. Some of the occurring problems concern camera images, which might show a common background area, but can still differ drastically in lighting effects. For foreground objects nearly no common visual information might be available, as angles between cameras may exceed even 90/spl deg/. Karsten Müller 0001, Aljoscha Smolic, Michael Drose, Patrick Voigt, Thomas Wiegand 0001 |
ICME | 1 |
| 2000 | A set of visual feature descriptors and their combination in a low-level description scheme
Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier |
Signal Process. Image Commun. | 5 |
| 1999 | A multi-feature description scheme for image and video database retrievalabstractThis paper reports about a description scheme for visual information content, which has been developed in the context of the forthcoming MPEG-7 standard. The system supports similarity-based retrieval of visual (image and video) data along feature axes like color, texture, shape/geometry and motion. The descriptors for these features have been developed in a way such that invariance against common transformations of visual material, e.g. filtering, contrast/color manipulation, resizing etc. is achieved, and that they are fitted to human perception properties. Furthermore, descriptors have been designed that allow a fast, hierarchical search procedure. A search engine has been developed on the basis of this description scheme, which allows similarity-based retrieval from an image or video database. The results show that efficient search and retrieval in visual database systems is possible based on a normative feature description such as MPEG-7. Jens-Rainer Ohm, F. Bunjamin, Wolfram Liebsch, Bela Makai, Karsten Müller 0001, Aljoscha Smolic, D. Zier |
MMSP | 5 |
| 1999 | Incomplete 3-D multiview representation of video objectsabstractThis paper introduces a new form of representation for three-dimensional (3-D) video objects. We have developed a technique to extract disparity and texture data from video objects that are captured simultaneously with multiple-camera configurations. For this purpose, we derive an "area of interest" (AOI) for each of the camera views, which represents an area on the video object's surface that is best visible from this specific camera viewpoint. By combining all AOIs, we obtain the video object plane as an unwrapped surface of a 3-D object, containing all texture data visible from any of the cameras. This texture surface can be encoded like any 2-D video object plane, while the 3-D information is contained in the associated disparity map. It is then possible to reconstruct different viewpoints from the texture surface by simple disparity-based projection. The merits of the technique are efficient multiview encoding of single video objects and support for viewpoint adaptation functionality, which is desirable in mixing natural and synthetic images. We have performed experiments with the MPEG-4 video verification model, where the disparity map is encoded by use of the tools provided for grayscale alpha data encoding. Due to its simplicity, the technique is suitable for applications that require real-time viewpoint adaptation toward video objects. Jens-Rainer Ohm, Karsten Müller 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |