VLDB 2026 Research / reviewers in the wild / expert
André Kaup
dblp:00/6329
· DBLP profile ↗
328ranked-venue papers
7as first author
104since 2021 · last 2026
0000-0002-0929-5074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 307 · 7 first-author · 93 since 2021Systems, architecture and hardware · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEE: An Efficient Feature-Based Framework for Performance Prediction in Semantic Segmentation
Felix Deichsel, Yongxu Ren, Philipp Beckerle, Jürgen Seiler, André Kaup |
ISCAS | 5 |
| 2026 | Perceptually-Weighted Video Quality Metric for Asymmetric Encoded Sports Videos
Anna Meyer, Jonas Janzen, Diwakara Reddy, Alexander Kopte, Simon Deniffel, Paul Wawerek-López, Marc Windsheimer, André Kaup |
QoMEX | 8 |
| 2026 | Frequency-spatial decoupled co-modeling transformer for fine-grained remote sensing image segmentation
Xin Li 0090, Shangtuo Qian, Xin Lyu 0001, Yongze Song, Fan Liu 0003, Yiwei Fang, Zhennan Xu, André Kaup |
Inf. Sci. | 8 |
| 2026 | A Dual Domain Collaborative Network for Polyp SegmentationabstractAccurate polyp segmentation in colonoscopy images is essential for early colorectal cancer detection but remains a challenging problem due to the limitations in existing methods for optimizing boundary features and aligning cross-level representations. Specifically, the indistinct polyp boundaries and scale variations across different feature levels pose significant challenges for segmentation accuracy. To address these issues, we propose a dual domain collaborative network (DDCNet) that introduces two novel modules: a frequency context enhancement module (FCEM), which operates in the frequency domain to refine high- and low-frequency features, and a cross-level shift-recalibrated fusion module (CSFM), which improves multi-scale feature alignment in the spatial domain. The FCEM improves boundary precision by adaptively refining high-frequency boundary features and enhancing low-frequency contextual information, while the CSFM mitigates cross-level feature misalignment by dynamically recalibrating multi-scale features throughout the encoder-decoder architecture. Additionally, we design a hybrid loss function that integrates boundary, cross-entropy, and frequency consistency losses to further boost segmentation performance. Experimental results on three benchmark datasets (Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB) demonstrate that DDCNet achieves state-of-the-art performance, with Dice coefficients of 0.9343, 0.9447, and 0.8155, respectively. These results represent improvements of 1.0%-1.5% over the best existing methods. Ablation studies further validate the individual contributions of FCEM, CSFM, and the hybrid loss function. Additionally, we compared the proposed loss function with three commonly used functions. Zuojian Zhou, Kongfa Hu, Tao Yang 0048, André Kaup, Xin Li 0090 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | A spectrum-enhanced attention model for semantic segmentation of remote sensing imagesabstractSemantic segmentation of remote sensing images (RSIs) is essential for applications such as environmental monitoring, urban planning, and disaster management. Convolutional Neural Networks (CNNs) and their variants struggle to capture comprehensive spectral context for learning discriminative representations. In this paper, we propose a Spectrum-Enhanced Network (SPENet) that leverages the Frequency Transformer Block (FTB) to capture rich spectral context. FTB integrates Spectrum-Enhanced Attention (SEA) with Multi-Head Frequency Self-Attention (MH-FSA), incorporating more informative contextual cues. Specifically, SEA aggregates spectral statistics through covariance matrix normalization before applying channel-wise attention. By projecting feature maps onto the frequency domain, MH-FSA provides the network with a broader context, extending beyond the low-frequency focus of standard self-attention mechanisms. Extensive experiments on the ISPRS Potsdam and LoveDA datasets show that SPENet significantly outperforms state-of-the-art methods. Besides, the proposed SEA module notably rises average F1-score/overall accuracy/mean insert over union wiht more than 2.5/2.6%/2.3%, as demonstrated by ablation study. Xin Li 0090, Feng Xu 0008, Feifei Tao, Xin Lyu 0001, Jianyi Zhong, André Kaup |
ICASSP | 7 |
| 2025 | Improved Motion Plane Adaptive 360-Degree Video Compression Using Affine Motion ModelsabstractEfficient compression of 360-degree video content requires the application of advanced motion models for inter-frame prediction. The Motion Plane Adaptive (MPA) motion model projects the frames on multiple perspective planes in the 3D space. It improves the motion compensation by estimating the motion on those planes with a translational diamond search. In this work, we enhance this motion model with an affine parameterization and motion estimation method. Thereby, we find a feasible trade-off between the quality of the reconstructed frames and the computational cost. The affine motion estimation is hereby done with the inverse compositional Lucas-Kanade algorithm. With the proposed method, it is possible to improve the motion compensation significantly, so that the motion compensated frame has a Weighted-to-Spherically-uniform Peak Signal-to-Noise Ratio (WS-PSNR) which is about 1.6 dB higher than with the conventional MPA. In a basic video codec, the improved inter prediction can lead to Bjøntegaard Delta (BD) rate savings between 9 % and 35 % depending on the block size (BS) and number of motion parameters. Marina Ritthaler, Andy Regensky, André Kaup |
ICASSP | 3 |
| 2025 | OSLO-IC: On-the-Sphere Learned Omnidirectional Image Compression with Attention Modules and Spatial ContextabstractDeveloping effective 360-degree (spherical) image compression techniques is crucial for technologies like virtual reality and automated driving. This paper advances the state-of-the-art in on-the-sphere learning (OSLO) for omnidirectional image compression framework by proposing spherical attention modules, residual blocks, and a spatial autoregressive context model. These improvements achieve a 23.1% bit rate reduction in terms of WS-PSNR BD rate. Additionally, we introduce a spherical transposed convolution operator for upsampling, which reduces trainable parameters by a factor of four compared to the pixel shuffling used in the OSLO framework, while maintaining similar compression performance. Therefore, in total, our proposed method offers significant rate savings with a smaller architecture and can be applied to any spherical convolutional application. Paul Wawerek-López, Navid Mahmoudian Bidgoli, Pascal Frossard, André Kaup, Thomas Maugey |
ICASSP | 4 |
| 2025 | Beyond Perspective: Neural 360-Degree Video Compression
Andy Regensky, Marc Windsheimer, Fabian Brand, André Kaup |
ICCV | 4 |
| 2025 | Compact Latent Representation for Image Compression (CLRIC)abstractCurrent image compression models often require separate models for each quality level, making them resource-intensive in terms of both training and storage. To address these limitations, we propose an innovative approach that utilizes latent variables from pre-existing trained models (such as the Stable Diffusion Variational Autoencoder) for perceptual image compression. Our method eliminates the need for distinct models dedicated to different quality levels. We employ overfitted learnable functions to compress the latent representation from the target model at any desired quality level. These overfitted functions operate in the latent space, ensuring low computational complexity, around 25.5 MAC/pixel for a forward pass on images with dimensions (1363×2048) pixels. This approach efficiently utilizes resources during both training and decoding. Our method achieves comparable perceptual quality to state-of-the-art learned image compression models while being both model-agnostic and resolution-agnostic. This opens up new possibilities for the development of innovative image compression methods. Ayman A. Ameen, Thomas Richter 0005, André Kaup |
ICIP | 3 |
| 2025 | Optimized Learned Image Compression for Facial Expression RecognitionabstractEfficient data compression is crucial for the storage and transmission of visual data. However, in facial expression recognition (FER) tasks, lossy compression often leads to feature degradation and reduced accuracy. To address these challenges, this study proposes an end-to-end model designed to preserve critical features and enhance both compression and recognition performance. A custom loss function is introduced to optimize the model, tailored to balance compression and recognition performance effectively. This study also examines the influence of varying loss term weights on this balance. Experimental results indicate that fine-tuning the compression model alone improves classification accuracy by 0.71 % and compression efficiency by 49.32 %, while joint optimization achieves significant gains of 4.04 % in accuracy and 89.12 % in efficiency. Moreover, the findings demonstrate that the jointly optimized classification model maintains high accuracy on both compressed and uncompressed data, while the compression model reliably preserves image details, even at high compression rates. Xiumei Li, Marc Windsheimer, Misha Sadeghi, Björn M. Eskofier, André Kaup |
ICIP | 5 |
| 2025 | Variable Rate Learned Wavelet Video Coding Using Temporal Layer AdaptivityabstractLearned wavelet video coders provide an explainable framework by performing discrete wavelet transforms in temporal, horizontal, and vertical dimensions. With a temporal transform based on motion-compensated temporal filtering (MCTF), spatial and temporal scalability is obtained. In this paper, we introduce variable rate support and a mechanism for quality adaption to different temporal layers for a higher coding efficiency. Moreover, we propose a multi-stage training strategy that allows training with multiple temporal layers. Our experiments demonstrate Bjøntegaard Delta bitrate savings of at least -32% compared to a learned MCTF model without these extensions. Training and inference code is available at: https://github.com/FAU-LMS/Learned-pMCTF. Anna Meyer, André Kaup |
ICIP | 2 |
| 2025 | Sliding Window Attention for Learned Video Compression
Alexander Kopte, André Kaup |
PCS | 2 |
| 2025 | TreeNet: A Light Weight Model for Low Bitrate Image Compression
Mahadev Prasad Panda, Purnachandra Rao Makkena, Srivatsa Prativadibhayankaram, Siegfried Fößel, André Kaup |
PCS | 5 |
| 2025 | A High-Level Feature Model to Predict the Encoding Energy of a Hardware Video Encoder
Diwakara Reddy, Christian Herglotz, André Kaup |
PCS | 3 |
| 2025 | Domain Adaptation for Camera-Specific Image Characteristics using Shallow DiscriminatorsabstractEach image acquisition setup leads to its own camera-specific image characteristics degrading the image quality. In learning-based perception algorithms, characteristics occurring during the application phase, but absent in the training data, lead to a domain gap impeding the performance. Previously, pixel-level domain adaptation through unpaired learning of the pristine-to-distorted mapping function has been proposed. In this work, we propose shallow discriminator architectures to address limitations of these approaches. We show that a smaller receptive field size improves learning of unknown image distortions by more accurately reproducing local distortion characteristics at a low network complexity. In a domain adaptation setup for instance segmentation, we achieve mean average precision increases over previous methods of up to 0.15 for individual distortions and up to 0.16 for camera-specific image characteristics in a simplified camera model. In terms of number of parameters, our approach matches the complexity of one state of the art method while reducing complexity by a factor of 20 compared to another, demonstrating superior efficiency without compromising performance. Maximiliane Gruber, Jürgen Seiler, André Kaup |
VCIP | 3 |
| 2025 | Towards Object Segmentation Mask Selection Using Specular ReflectionsabstractSpecular reflections pose a significant challenge for object segmentation, as their sharp intensity transitions often mislead both conventional algorithms and deep learning based methods. However, as the specular reflection must lie on the surface of the object, this fact can be exploited to improve the segmentation masks. By identifying the largest region containing the reflection as the object, we derive a more accurate object mask without requiring specialized training data or model adaption. We evaluate our method on both synthetic and real world images and compare it against established and state-of-the-art techniques including Otsu thresholding, YOLO, and SAM2. Compared to the best performing baseline SAM2, our approach achieves up to 26.7% improvement in IoU, 22.3% in DSC, and 9.7% in pixel accuracy. Qualitative evaluations on real world images further confirm the robustness and generalizability of the proposed approach. Katja Kossira, Yunxuan Zhu, Jürgen Seiler, André Kaup |
VCIP | 4 |
| 2025 | Anti-Aliasing Snapshot HDR Imaging Using Non-Regular SensingabstractSnapshot HDR imaging is essential to capture the full dynamic range of a scene in a single exposure, making it essential for video and dynamic environments where motion prevents the use of multi-exposure techniques or complex hardware set-ups. This work presents a snapshot HDR imaging sensor that is based on spatially varying apertures, implemented by combining two differently sized prototype pixels. The different light integration areas physically extend the dynamic range towards the lower end, compared to a standard high resolution sensor. A non-regular pixel arrangement is suggested, to mitigate aliasing and overcome a loss in spatial resolution that is associated with increased light integration area of the larger prototype pixel. Subsequent reconstruction in the Fourier domain, where natural images can be sparsely represented allows to recover the image with high detail. The image acquisition approach with the proposed non-regular HDR sensor is simulated and analysed with special emphasis on the spatial resolution. The results suggest the snapshot HDR sensor layout to be an effective way to acquire images with high dynamic range and free from aliasing artefacts. Teresa Stürzenhofäcker, Moritz Klimm, Jürgen Seiler, André Kaup |
VCIP | 4 |
| 2025 | Boosting Neural Image Compression for Machines Using Latent Space MaskingabstractToday, many image coding scenarios do not have a human as final intended user, but rather a machine fulfilling computer vision tasks on the decoded image. Thereby, the primary goal is not to keep visual quality but maintain the task accuracy of the machine for a given bitrate. Due to the tremendous progress of deep neural networks setting benchmarking results, mostly neural networks are employed to solve the analysis tasks at the decoder side. Moreover, neural networks have also found their way into the field of image compression recently. These two developments allow for an end-to-end training of the neural compression network for an analysis network as information sink. Therefore, we first roll out such a training with a task-specific loss to enhance the coding performance of neural compression networks. Compared to the standard VVC, 41.4% of bitrate are saved by this method for Mask R-CNN as analysis network on the uncompressed Cityscapes dataset. As a main contribution, we propose LSMnet, a network that runs in parallel to the encoder network and masks out elements of the latent space that are presumably not required for the analysis network. By this approach, additional 27.3% of bitrate are saved compared to the basic neural compression network optimized with the task loss. In addition, we are the first to utilize a feature-based distortion in the training loss within the context of machine-to-machine communication, which allows for a training without annotated data. We provide extensive analyses on the Cityscapes dataset including cross-evaluation with different analysis networks and present exemplary visual results. Kristian Fischer 0001, Fabian Brand, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Overview of Variable Rate Coding in JPEG AIabstractEmpirical evidence has demonstrated that learning-based image compression can outperform classical compression frameworks. This has led to the ongoing standardization of learned-based image codecs, namely Joint Photographic Experts Group (JPEG) AI. The objective of JPEG AI is to enhance compression efficiency and provide a software and hardware-friendly solution. Based on our research, JPEG AI represents the first standardization that can facilitate the implementation of a learned image codec on a mobile device. This article presents an overview of the variable rate coding functionality in JPEG AI, which includes three variable rate adaptations: a three-dimensional quality map, a fast bit rate matching algorithm, and a training strategy. The variable rate adaptations offer a continuous rate function up to 2.0 bpp, exhibiting a high level of performance, a flexible bit allocation between different color components, and a region of interest function for the specified use case. The evaluation of performance encompasses both objective and subjective results. With regard to the objective bit rate matching, the main profile with low complexity yielded a 13.1% BD-rate gain over VVC intra, while the high profile with high complexity achieved a 19.2% BD-rate gain over VVC intra. The BD-rate result is calculated as the mean of the seven perceptual metrics defined in the JPEG AI common test conditions. With respect to subjective results, the example of improving the quality of the region of interest is illustrated. Panqi Jia, Fabian Brand, Dequan Yu, Alexander Karabutov, Elena Alshina, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation of remote sensing images (RSIs) plays a pivotal role in advancing geospatial analyses and applications across diverse fields, such as urban planning and environmental monitoring. Traditional learning paradigms predominantly utilize Euclidean spaces for feature extraction. This approach can introduce spatial distortions when representing objects, as Euclidean architectures typically focus on locality and are optimized for grid data, not always yielding optimal geometrical representations for data structured in non-Euclidean spaces. To address these problems, we propose EAAHNet, the first fully hyperbolic neural network designed for semantic segmentation of RSIs. EAAHNet employs the Lorentz model to reformalize conventional Euclidean-based neural network operations, ensuring the preservation of hyperbolic properties. Furthermore, to account for the inherently Euclidean nature of ground objects, we propose a Euclidean affinity-augmented hyperbolic attention module (EAAHAM) that enriches contextual dependencies through an attention fusion manner. This enhancement significantly improves the network’s capacity to discern pixel-wise semantics. Extensive experiments conducted on the ISPRS Vaihingen, ISPRS Potsdam, and LoveDA datasets demonstrate EAAHNet’s superior performance over several state-of-the-art methods. Additionally, the ablation study verifies the impacts of EAAHAM. Xin Li 0090, Feng Xu 0008, Fan Liu 0003, Xin Lyu 0001, Hongmin Gao 0001, Jun Zhou 0011, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Multispectral Snapshot Image Registration Using Learned Cross Spectral Disparity Estimation and a Deep Guided Occlusion Reconstruction NetworkabstractMultispectral imaging aims at recording images in different spectral bands. This is extremely beneficial in diverse discrimination applications, for example in agriculture, recycling or healthcare. One approach for snapshot multispectral imaging, which is capable of recording multispectral videos, is by using camera arrays, where each camera records a different spectral band. Since the cameras are at different spatial positions, a registration procedure is necessary to map every camera to the same view. In this paper, we present a multispectral snapshot image registration with three novel components. First, a cross spectral disparity estimation network is introduced, which is trained on a popular stereo database using pseudo spectral data augmentation. Subsequently, this disparity estimation is used to accurately detect occlusions by warping the disparity map in a layer-wise manner. Finally, these detected occlusions are reconstructed by a learned deep guided neural network, which leverages the structure from other spectral components. It is shown that each element of this registration process as well as the final result is superior to the current state of the art. In terms of PSNR, our registration achieves an improvement of over 3 dB. At the same time, the runtime is decreased by a factor of over 3 on a CPU. Additionally, the registration is executable on a GPU, where the runtime can be decreased by a factor of 113. The source code and the data is available at https://github.com/FAU-LMS/MSIR. Frank Sippel, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 3 |
| 2024 | SLIC: A Learned Image Codec Using Structure and ColorabstractWe propose the structure and color based learned image codec (SLIC) in which the task of compression is split into that of luminance and chrominance. The deep learning model is built with a novel multi-scale architecture for Y and UV channels in the encoder, where the features from various stages are combined to obtain the latent representation. An autoregressive context model is employed for backward adaptation and a hyperprior block for forward adaptation. Various experiments are carried out to study and analyze the performance of the proposed model, and to compare it with other image codecs. We also illustrate the advantages of our method through the visualization of channel impulse responses, latent channels and various ablation studies. The model achieves Bjøntegaard delta bitrate gains of 7.5% and 4.66% in terms of MS-SSIM and CIEDE2000 metrics with respect to other state-of-the-art reference codecs. Srivatsa Prativadibhayankaram, Mahadev Prasad Panda, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel, André Kaup |
DCC | 6 |
| 2024 | Encoding Time and Energy Model for SVT-AV1 Based on Video ComplexityabstractThe share of online video traffic in global carbon dioxide emissions is growing steadily. To comply with the demand for video media, dedicated compression techniques are continuously optimized, but at the expense of increasingly higher computational demands and thus rising energy consumption at the video encoder side. In order to find the best trade-off between compression and energy consumption, modeling encoding energy for a wide range of encoding parameters is crucial. We propose an encoding time and energy model for SVT-AV1 based on empirical relations between the encoding time and video parameters as well as encoder configurations. Furthermore, we model the influence of video content by established content descriptors such as spatial and temporal information. We then use the predicted encoding time to estimate the required energy demand and achieve a prediction error of 19.6% for encoding time and 20.9% for encoding energy. Lena Eichermüller, Gaurang Chaudhari, Ioannis Katsavounidis, Zhijun Lei, Hassene Tmar, Christian Herglotz, André Kaup |
ICASSP | 7 |
| 2024 | Quantized Decoder in Learned Image Compression for Deterministic ReconstructionabstractLearned image compression has a problem of non-bit-exact reconstruction due to different calculations of floating point arithmetic on different devices. This paper shows a method to achieve a deterministic reconstructed image by quantizing only the decoder of the learned image compression model. From the implementation perspective of an image codec, it is beneficial to have the results reproducible when decoded on different devices. In this paper, we study quantization of weights and activations without overflow of accumulator in all decoder subnetworks. We show that the results are bit-exact at the output, and the resulting BD-rate loss of quantization of decoder is 0.5% in the case of 16-bit weights and 16-bit activations, and 7.9% in the case of 8-bit weights and 16-bit activations. Esin Koyuncu, Timofey Solovyev, Johannes Sauer, Elena Alshina, André Kaup |
ICASSP | 5 |
| 2024 | Enhanced Color Palette Modeling For Lossless Screen Content CompressionabstractSoft context formation is a lossless image coding method for screen content. It encodes images pixel by pixel via arithmetic coding by collecting statistics for probability distribution estimation. Its main pipeline includes three stages, namely a context model based stage, a color palette stage and a residual coding stage. Each subsequent stage is only employed if the previous stage can not be applied since necessary statistics, e.g. colors or contexts, have not been learned yet. We propose the following enhancements: First, information from previous stages is used to remove redundant color palette entries and prediction errors in subsequent stages. Additionally, implicitly known stage decision signals are no longer explicitly transmitted. These enhancements lead to an average bit rate decrease of 1.07% on the evaluated data. Compared to VVC and HEVC, the proposed method needs roughly 0.44 and 0.17 bits per pixel less on average for 24-bit screen content images, respectively. Hannah Och, Shabhrish Reddy Uddehal, Tilo Strutz, André Kaup |
ICASSP | 4 |
| 2024 | Improved Screen Content Coding in VVC Using Soft Context FormationabstractScreen content images typically contain a mix of natural and synthetic image parts. Synthetic sections usually are comprised of uniformly colored areas and repeating colors and patterns. In the VVC standard, these properties are exploited using Intra Block Copy and Palette Mode. In this paper, we show that pixel-wise lossless coding can outperform lossy VVC coding in such areas. We propose an enhanced VVC coding approach for screen content images using the principle of soft context formation. First, the image is separated into two layers in a block-wise manner using a learning-based method with four block features. Synthetic image parts are coded losslessly using soft context formation, the rest with VVC. We modify the available soft context formation coder to incorporate information gained by the decoded VVC layer for improved coding efficiency. Using this approach, we achieve Bjontegaard-Delta-rate gains of 4.98% on the evaluated data sets compared to VVC. Hannah Och, Shabhrish Reddy Uddehal, Tilo Strutz, André Kaup |
ICASSP | 4 |
| 2024 | Geometry-Corrected Geodesic Motion Modeling with Per-Frame Camera Motion for 360-Degree Video CompressionabstractThe large amounts of data associated with 360-degree video require highly effective compression techniques for efficient storage and distribution. The development of improved motion models for 360-degree motion compensation has shown significant improvements in compression efficiency. A geodesic motion model representing translational camera motion proved to be one of the most effective models. In this paper, we propose an improved geometry-corrected geodesic motion model that outperforms the state of the art at reduced complexity. We additionally propose the transmission of per-frame camera motion information, where prior work assumed the same camera motion for all frames of a sequence. Our approach yields average Bjøntegaard Delta rate savings of 2.27% over H.266/VVC, outperforming the original geodesic motion model by 0.32 percentage points at reduced computational complexity. Andy Regensky, André Kaup |
ICASSP | 2 |
| 2024 | Color Agnostic Cross-Spectral Disparity EstimationabstractSince camera modules become more and more affordable, multi-spectral camera arrays have found their way from special applications to the mass market, e.g., in automotive systems, smartphones, or drones. Due to multiple modalities, the registration of different viewpoints and the required cross-spectral disparity estimation is up to the present extremely challenging. To overcome this problem, we introduce a novel spectral image synthesis in combination with a color agnostic transform. Thus, any recently published stereo matching network can be turned to a cross-spectral disparity estimator. Our novel algorithm requires only RGB stereo data to train a cross-spectral disparity estimator and a generalization from artificial training data to camera-captured images is obtained. The theoretical examination of the novel color agnostic method is completed by an extensive evaluation compared to state of the art including self-recorded multispectral data and a reference implementation. The novel color agnostic disparity estimation improves cross-spectral as well as conventional color stereo matching by reducing the average end-point error by 41 % for cross-spectral and by 22 % for mono-modal content, respectively. Frank Sippel, Nils Genser, Hannah Och, Jürgen Seiler, André Kaup |
ICASSP | 5 |
| 2024 | A Guided Upsampling Network for Short wave Infrared Images Using Graph RegularizationabstractExploiting the infrared area of the spectrum for classification problems is getting increasingly popular, because many materials have characteristic absorption bands in this area. However, sensors in the short wave infrared (SWIR) area and even higher wavelengths have a very low spatial resolution in comparison to classical cameras that operate in the visible wavelength area. Thus, in this paper an upsampling method for SWIR images guided by a visible image is presented. For that, the proposed guided upsampling network (GUNet) uses a graph-regularized optimization problem based on learned affinities is presented. The evaluation is based on a novel synthetic near-field visible-SWIR stereo database. Different guided upsampling methods are evaluated, which shows an improvement of nearly 1 dB on this database for the proposed upsampling method in comparison to the second best guided upsampling network. Furthermore, a visual example of an upsampled SWIR image of a real-world scene is depicted for showing real-world applicability. Frank Sippel, Jürgen Seiler, André Kaup |
ICASSP | 3 |
| 2024 | Multiscale Augmented Normalizing Flows for Image CompressionabstractMost learning-based image compression methods lack efficiency for high image quality due to their non-invertible design. The decoding function of the frequently applied compressive autoencoder architecture is only an approximated inverse of the encoding transform. This issue can be resolved by using invertible latent variable models, which allow a perfect reconstruction if no quantization is performed. Furthermore, many traditional image and video coders apply dynamic block partitioning to vary the compression of certain image regions depending on their content. Inspired by this approach, hierarchical latent spaces have been applied to learning-based compression networks. In this paper, we present a novel concept, which adapts the hierarchical latent space for augmented normalizing flows, an invertible latent variable model. Our best performing model achieves significant rate savings of more than 7% over comparable single-scale models. Marc Windsheimer, Fabian Brand, André Kaup |
ICASSP | 3 |
| 2024 | Conditional Optimal Filter Selection For Multispectral Object ClassificationabstractCapturing images using multispectral camera arrays has gained importance in medical, agricultural and environmental processes. However, using all available spectral bands is infeasible and produces much data, while only a fraction is needed for a given task. Nearby bands may contain similar information, therefore redundant spectral bands should not be considered in the evaluation process to keep complexity and the data load low. In current methods, a restricted and pre-determined number of spectral bands is selected. Our approach improves this procedure by including preset conditions such as noise or the bandwidth of available filters, minimizing spectral redundancy. Furthermore, a minimal filter selection can be conducted, keeping the hardware setup at low costs, while still obtaining all important spectral information. In comparison to the fast binary search filter band selection method, we managed to reduce the amount of misclassified objects of the SMM dataset from 318 to 124 using a random forest classifier. Katja Kossira, David Schön, Jürgen Seiler, André Kaup |
ICIP | 4 |
| 2024 | Efficient Learned Wavelet Image and Video CodingabstractLearned wavelet image and video coding approaches provide an explainable framework with a latent space corresponding to a wavelet decomposition. The wavelet image coder iWave++ achieves state-of-the-art performance and has been employed for various compression tasks, including lossy as well as lossless image, video, and medical data compression. However, the approaches suffer from slow decoding speed due to the autoregressive context model used in iWave++. In this paper, we show how a parallelized context model can be integrated into the iWave++ framework. Our experimental results demonstrate a speedup factor of over 350 and 240 for image and video compression, respectively. At the same time, the rate-distortion performance in terms of Bjøntegaard delta bitrate is slightly worse by 1.5% for image coding and 1% for video coding. In addition, we analyze the learned wavelet decomposition by visualizing its subband impulse responses. Anna Meyer, Srivatsa Prativadibhayankaram, André Kaup |
ICIP | 3 |
| 2024 | End-to-End Learned Lossy Dynamic Point Cloud Attribute CompressionabstractRecent advancements in point cloud compression have primarily emphasized geometry compression while comparatively fewer efforts have been dedicated to attribute compression. This study introduces an end-to-end learned dynamic lossy attribute coding approach, utilizing an efficient high-dimensional convolution to capture extensive inter-point dependencies. This enables the efficient projection of attribute features into latent variables. Subsequently, we employ a context model that leverage previous latent space in conjunction with an auto-regressive context model for encoding the latent tensor into a bitstream. Evaluation of our method on widely utilized point cloud datasets from the MPEG and Microsoft demonstrates its superior performance compared to the core attribute compression module Region-Adaptive Hierarchical Transform method from MPEG Geometry Point Cloud Compression with $38.1 \%$ Bjontegaard Delta-rate saving in average while ensuring a low-complexity encoding/decoding. Dat Thanh Nguyen, Daniel Zieger, Marc Stamminger, André Kaup |
ICIP | 4 |
| 2024 | A Study on the Effect of Color Spaces in Learned Image CompressionabstractIn this work, we present a comparison between color spaces namely YUV, LAB, RGB and their effect on learned image compression. For this we use the structure and color based learned image codec (SLIC) from our prior work, which consists of two branches - one for the luminance component (Y or L) and another for chrominance components (UV or AB). However, for the RGB variant we input all 3 channels in a single branch, similar to most learned image codecs operating in RGB. The models are trained for multiple bitrate configurations in each color space. We report the findings from our experiments by evaluating them on various datasets and compare the results to state-of-the-art image codecs. The YUV model performs better than the LAB variant in terms of MSSSIM with a Bjøntegaard delta bitrate (BD-BR) gain of 7.5% using VTM intra-coding mode as the baseline. Whereas the LAB variant has a better performance than YUV model in terms of CIEDE2000 having a BD-BR gain of 8%. Overall, the RGB variant of SLIC achieves the best performance with a BD-BR gain of 13.14% in terms of MS-SSIM and a gain of 17.96% in CIEDE2000 at the cost of a higher model complexity. Srivatsa Prativadibhayankaram, Mahadev Prasad Panda, Jürgen Seiler, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel, André Kaup |
ICIP | 7 |
| 2024 | Fast Edge-Aware Occlusion Detection In The Context of Multispectral Camera ArraysabstractMultispectral imaging is very beneficial in diverse applications, like healthcare and agriculture, since it can capture absorption bands of molecules in different spectral areas. A promising approach for multispectral snapshot imaging are camera arrays. Image processing is necessary to warp all different views to the same view to retrieve a consistent multispectral datacube. This process is also called multispectral image registration. After a cross spectral disparity estimation, an occlusion detection is required to find the pixels that were not recorded by the peripheral cameras. In this paper, a novel fast edge-aware occlusion detection is presented, which is shown to reduce the runtime by at least a factor of 12. Moreover, an evaluation on ground truth data reveals better performance in terms of precision and recall. Finally, the quality of a final multispectral datacube can be improved by more than 1.5 dB in terms of PSNR as well as in terms of SSIM in an existing multispectral registration pipeline. The source code is available at https://github.com/FAU-LMS/fast-occlusion-detection. Frank Sippel, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2024 | ON Annotation-Free Optimization of Video Coding for MachinesabstractToday, image and video data is not only viewed by humans, but also automatically analyzed by computer vision algorithms. However, current coding standards are optimized for human perception. Emerging from this, research on video coding for machines tries to develop coding methods designed for machines as information sink. Since many of these algorithms are based on neural networks, most proposals for video coding for machines build upon neural compression. So far, optimizing the compression by applying the task loss of the analysis network, for which ground truth data is needed, is achieving the best coding performance. But ground truth data is difficult to obtain and thus an optimization without ground truth is preferred. In this paper, we present an annotation-free optimization strategy for video coding for machines. We measure the distortion by calculating the task loss of the analysis network. Therefore, the predictions on the compressed image are compared with the predictions on the original image, instead of the ground truth data. Our results show that this strategy can even outperform training with ground truth data with rate savings of up to 7.5 %. By using the non-annotated training data, the rate gains can be further increased up to 8.2 %. Marc Windsheimer, Fabian Brand, André Kaup |
ICIP | 3 |
| 2024 | Decoding Energy Optimization for Video Coding Using Model-Driven Gradient DescentabstractNowadays, a large part of the global energy consumption caused by video communications can be attributed to end-user devices such as smartphones, tablet PCs, and TV sets. In this paper, we present a method to increase the performance of an existing algorithm dedicated to reduce the end-user side energy consumption during video streaming. The algorithm, which is called decoding-energy-rate-distortion optimization (DERDO), exploits a decoding energy model during encoding and chooses coding modes in such a way that the software decoding energy is minimized. In this paper, we develop a dedicated gradient descent approach for DERDO that refines specific energy coefficients used for decoding energy modeling. We find that this approach boosts the performance of DERDO by increasing the energy savings by at least 5% with respect to standard DERDO. As a consequence, we observe decoding energy savings of more than 40% and more than 7% for practical encoder and decoder implementations of HEVC and H.264/AVC, respectively, when compared to standard encoding using classic rate-distortion optimization. Christian Herglotz, Matthias Kränzler, Bide Xu, André Kaup |
MMSP | 4 |
| 2024 | Inter-Camera Color Correction for Multispectral Imaging with Camera Arrays Using a Consensus ImageabstractThis paper introduces a novel method for inter-camera color calibration for multispectral imaging with camera arrays using a consensus image. Capturing images using multispectral camera arrays has gained importance in medical, agricultural, and environmental processes. Due to fabrication differences, noise, or device altering, varying pixel sensitivities occur, influencing classification processes. Therefore, color calibration between the cameras is necessary. In existing methods, one of the camera images is chosen and considered as a reference, ignoring the color information of all other recordings. Our new approach does not just take one image as reference, but uses statistical information such as the location parameter to generate a consensus image as basis for calibration. This way, we managed to improve the PSNR values for the linear regression color correction algorithm by 1.15 dB and the improved color difference (iCID) values by 2.81. Katja Kossira, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2024 | Modeling the Energy Consumption of the HEVC Software Encoding Process Using Processor EventsabstractDeveloping energy-efficient video encoding algorithms is highly important due to the high processing complexities and, consequently, the high energy demand of the encoding process. To accomplish this, the energy consumption of the video encoders must be studied, which is only possible with a complex and dedicated energy measurement setup. This emphasizes the need for simple energy estimation models, which estimate the energy required for the encoding. Our paper investigates the possibility of estimating the energy demand of a HEVC software CPU-encoding process using processor events. First, we perform energy measurements and obtain processor events using dedicated profiling software. Then, by using the measured energy demand of the encoding process and profiling data, we build an encoding energy estimation model that uses the processor events of the ultrafast encoding preset to obtain the energy estimate for complex encoding presets with a mean absolute percentage error of 5.36% when averaged over all the presets. Additionally, we present an energy model that offers the possibility of obtaining energy distribution among various encoding sub-processes. Geetha Ramasubbu, André Kaup, Christian Herglotz |
MMSP | 2 |
| 2024 | Adaptive Variance-Threshold-Based Skip Modes for Learned Video Compression Using a Motion Complexity CriterionabstractSkip modes are a powerful tool to reduce the rate in video compression. The main idea is that the residual areas where the prediction performs well are not transmitted since the prediction quality is good enough that the prediction signal itself can be used as the reconstruction signal. This is commonly used, e.g., in the compression standard VVC, where a skip flag can be transmitted for inter blocks under certain conditions. The skipped residual block is then not transmitted and the content is instead inferred to be zero at the decoder. Current learning-based methods use different kinds of skip modes. One possibility here arises from the fact that the coders estimate and transmit the variance for each transmitted symbol. It has been proposed to use this estimated variance to derive a skip mode. When the variance falls below a threshold, the symbol is not transmitted. In this paper we propose an extension to this method. By classifying each position in the latent space according to the local motion complexity, we can transmit adaptive thresholds for each class. That way, we can employ motion information to refine the granularity of the skip mode. When we implement this method in FVC, we are able to save up 2.11% rate on a GOP 20 sequence. We also discuss the behavior of increasingly adaptive skip modes in scenarios with larger GOP size, where error-propagation becomes a larger issue. Fabian Brand, Jürgen Seiler, Johannes Sauer, Elena Alshina, André Kaup |
PCS | 5 |
| 2024 | Complexity Metrics for VVC Decoder Power Reduction in Green MetadataabstractThis paper discusses the new complexity metrics for VVC decoder power reduction introduced in the 3rdedition of the Green Metadata standard. The standard defines dedicated syntax elements that represent the expected software decoding complexity with a high accuracy. Using a simple complexity model, which can be trained for any VVC decoder software implementation, the receiver of the video can estimate the processing complexity to decode the subsequent video segment. Afterwards, it can adjust the clock frequency of the processor to keep the real-time playback constraint. By reducing the frequency of the processor, a significant amount of energy is saved. In this paper, we present the syntax elements, their meaning, and show that they can accurately estimate the processing complexity of various software decoder implementations with errors below 11%. Furthermore, we present a processor-frequency-control algorithm and apply it to a development board performing VVC video decoding. Measurements reveal that the complexity metric signaling can lead to up to 30% of energy savings. Christian Herglotz, Matthias Kränzler, André Kaup |
PCS | 4 |
| 2024 | Bit Rate Matching Algorithm Optimization in JPEG-AI Verification ModelabstractThe research on neural network (NN) based image compression has shown superior performance compared to classical compression frameworks. Unlike the hand-engineered transforms in the classical frameworks, NN-based models learn the non-linear transforms providing more compact bit represen-tations, and achieve faster coding speed on parallel devices over their classical counterparts. Those properties evoked the attention of both scientific and industrial communities, resulting in the standardization activity JPEG-AI. The verification model for the standardization process of JPEG-AI is already in development and has surpassed the advanced VVC intra codec. To generate reconstructed images with the desired bits per pixel and assess the BD-rate performance of both the JPEG-AI verification model and VVC intra, bit rate matching is employed. However, the current state of the JPEG-AI verification model experiences significant slowdowns during bit rate matching, resulting in suboptimal performance due to an unsuitable model. The proposed methodology offers a gradual algorithmic optimization for matching bit rates, resulting in a fourfold acceleration and over 1% improvement in BD-rate at the base operation point. At the high operation point, the acceleration increases up to sixfold. Panqi Jia, Ahmet Burakhan Koyuncu, Jue Mao, Ze Cui, Tiansheng Guo, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Jing Wang 0194, Elena Alshina, André Kaup |
PCS | 12 |
| 2024 | A Comprehensive Review of Software and Hardware Energy Efficiency of Video DecodersabstractEnergy and compression efficiency are two essential parts of modern video decoder implementations that have to be considered. This work comprehensively studies the following six video coding formats regarding compression and decoding energy efficiency: AVC, VP9, HEVC, AV1, VVC, and AVM. We first evaluate the energy demand of reference and optimized software decoder implementations. Furthermore, we consider the influence of the usage of SIMD instructions on those decoder implementations. We find that AV1 is a sweet spot for optimized software decoder implementations with an additional energy demand of 16.55% and bitrate savings of -43.95% compared to VP9. We furthermore evaluate the hardware decoding energy demand of four video coding formats. Thereby, we show that AV1 has energy demand increases by 117.50% compared to VP9. For HEVC, we found a sweet spot in terms of energy demand with an increase of 6.06% with respect to VP9. Relative to their optimized software counterparts, hardware video decoders reduce the energy consumption to less than 9% compared to software decoders. Matthias Kränzler, Christian Herglotz, André Kaup |
PCS | 3 |
| 2024 | Analysis of Neural Video Compression Networks for 360-Degree Video CodingabstractWith the increasing efforts of bringing high-quality virtual reality technologies into the market, efficient 360-degree video compression gains in importance. As such, the state-of-the-art H.266 NVC video coding standard integrates dedicated tools for 360-degree video, and considerable efforts have been put into designing 360-degree projection formats with improved compression efficiency. For the fast-evolving field of neural video compression networks (NVCs), the effects of different 360-degree projection formats on the overall compression performance have not yet been investigated. It is thus unclear, whether a resampling from the conventional equirectangular projection (ERP) to other projection formats yields similar gains for NVCs as for hybrid video codecs, and which formats perform best. In this paper, we analyze several generations of NVCs and an extensive set of 360-degree projection formats with respect to their compression performance for 360-degree video. Based on our analysis, we find that projection format resampling yields significant improvements in compression performance also for NVCs. The adjusted cubemap projection (ACP) and equatorial cylindrical projection (ECP) show to perform best and achieve rate savings of more than 55% compared to ERP based on WS-PSNR for the most recent NVC. Remarkably, the observed rate savings are higher than for H.266/VVC, emphasizing the importance of projection format resampling for NVCs. Andy Regensky, Fabian Brand, André Kaup |
PCS | 3 |
| 2024 | Towards Video Codec Performance Evaluation: A Rate-Energy-Distortion PerspectiveabstractThe Bjøntegaard Delta rate (BD-rate) objectively assesses the coding efficiency of video codecs using the rate-distortion (R-D) performance but overlooks encoding energy, which is crucial in practical applications, especially for those on handheld devices. Although R-D analysis can be extended to incorporate encoding energy as energy-distortion (E-D), it fails to integrate all three parameters seamlessly. This work proposes a novel approach to address this limitation by introducing a 3D representation of rate, encoding energy, and distortion through surface fitting. In addition, we evaluate various surface fitting techniques based on their accuracy and investigate the proposed 3D representation and its projections. The overlapping areas in projections help in encoder selection and recommend avoiding the slow presets of the older encoders (x264, x265), as the recent encoders (x265, VVenC) offer higher quality for the same bitrate-energy performance and provide a lower rate for the same energy-distortion performance. Geetha Ramasubbu, André Kaup, Christian Herglotz |
QoMEX | 2 |
| 2024 | Networked Systems Diagnostics: A Fusion of Failure Mode and Effects Analysis and a Delphi Expert StudyabstractNetworked devices, especially those comprising multiple identical devices, are extensively utilized in industrial scenarios. However, their complexity poses unique challenges in diagnostic processes, demanding efficient methodologies to identify and assess risks. The application of Failure Mode and Effects Analysis (FMEA) for analyzing complex systems, especially those consisting of networked devices, appears to be limited, particularly in identifying critical risk factors. In this paper, we propose a novel pipeline to diagnose networked systems by fusing FMEA with a Delphi (expert) Study. Our approach leverages the collective knowledge of a group of experts through a structured Delphi Study enabling them to contribute individually but also to interact. We demonstrate the applicability of our approach through a case study involving a system of networked mobile robots. Based on a Fault Tree Analysis (FTA), Risk Priority Numbers (RPN) of minimal fault tree cut sets are calculated to identify the most critical mechatronic failures. Our findings show that our methodology provides an RPN ranking that closely aligns with expert insights, highlighting its efficacy in accurately assessing risk in complex networked systems. Yongxu Ren, Felix Deichsel, Valentin Hopf, Jürgen Seiler, André Kaup, Philipp Beckerle |
SMC | 5 |
| 2024 | Bit Distribution Study and Implementation of Spatial Quality Map in the JPEG-AI StandardizationabstractCurrently, there is a high demand for neural network-based image compression codecs. These codecs employ non-linear transforms to create compact bit representations and facilitate faster coding speeds on devices compared to the handcrafted transforms used in classical frameworks. The scientific and industrial communities are highly interested in these properties, leading to the standardization effort of JPEG-AI. The JPEG-AI verification model has been released and is currently under development for standardization. Utilizing neural networks, it can outperform the classic codec VVC intra by over 10% BD-rate operating at base operation point. Researchers attribute this success to the flexible bit distribution in the spatial domain, in contrast to VVC intra’s anchor that is generated with a constant quality point. However, our study reveals that VVC intra displays a more adaptable bit distribution structure through the implementation of various block sizes. As a result of our observations, we have proposed a spatial bit allocation method to optimize the JPEG-AI verification model’s bit distribution and enhance the visual quality. Furthermore, by applying the VVC bit distribution strategy, the objective performance of JPEG-AI verification mode can be further improved, resulting in a maximum gain of 0.45 dB in PSNR-Y. Panqi Jia, Jue Mao, Esin Koyuncu, Ahmet Burakhan Koyuncu, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Elena Alshina, André Kaup |
VCIP | 9 |
| 2024 | Forensic analysis of AI-compression traces in spatial and frequency domainabstractThe classical JPEG compression is a rich source of cues for forensic image analysis. However, this compression standard will in the near future be complemented by a new, highly efficient learning-based compression standard called JPEG-AI. JPEG-AI is fundamentally different from classical JPEG. Hence, its forensic traces can also be expected to be fundamentally different. We argue that there is a pressing need for image forensics research to investigate these traces. In this work, we characterize forensic compression traces of different AI compression algorithms. Our analysis investigates AI compression artifacts in frequency domain and in spatial domain. Both domains exhibit similar artifacts that likely stem from upsampling operations of the decoders. Additionally, we report for one AI codec another artifact in homogeneous regions. We also investigate the artifact detectability in several scenarios including unseen AI compression traces and postprocessing. Here, frequency and autocorrelation features are better on additive noise and classical JPEG post-compression, while RGB features perform better on blurred and downsampled images. Sandra Bergmann, Denise Moussa, Fabian Brand, André Kaup, Christian Riess |
Pattern Recognit. Lett. | 4 |
| 2024 | Conditional Residual Coding: A Remedy for Bottleneck Problems in Conditional Inter Frame CodingabstractConditional coding is a new video coding paradigm enabled by neural-network-based compression. It can be shown that conditional coding is in theory better than the traditional residual coding, which is widely used in video compression standards like HEVC or VVC. However, on closer inspection, it becomes clear that conditional coders can suffer from information bottlenecks in the prediction path, i.e., that due to the data processing inequality not all information from the prediction signal can be passed to the reconstructed signal, thereby impairing the coder performance. In this paper we propose the conditional residual coding concept, which we derive from information theoretical properties of the conditional coder. This coder significantly reduces the influence of bottlenecks, while maintaining the theoretical performance of the conditional coder. We provide a theoretical analysis of the coding paradigm and demonstrate the performance of the conditional residual coder in a practical example. We show that conditional residual coders alleviate the disadvantages of conditional coders while being able to maintain their advantages over residual coders. In the spectrum of residual and conditional coding, we can therefore consider them as “the best from both worlds”. Fabian Brand, Jürgen Seiler, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Cross-Modal Prealigned Method With Global and Local Information for Remote Sensing Image and Text RetrievalabstractIn recent years, remote sensing cross-modal text-image retrieval (RSCTIR) has attracted considerable attention owing to its convenience and information mining capabilities. However, two significant challenges persist: effectively integrating global and local information during feature extraction due to substantial variations in remote sensing imagery, and the failure of existing methods to adequately consider feature prealignment before modal fusion, resulting in complex modal interactions that adversely impact retrieval accuracy and efficiency. To address these challenges, we propose a cross-modal prealigned method with global and local information (CMPAGL) for remote sensing imagery. Specifically, we design a global-Swin (Gswin) Transformer block, which introduces a global information window on top of the local window attention mechanism, synergistically combining local window self-attention and global-local window cross-attention to effectively capture multiscale features of remote sensing images. In addition, our approach incorporates a prealignment mechanism to mitigate the training difficulty of modal fusion, thereby enhancing retrieval accuracy. Moreover, we propose a similarity matrix reweighting (SMR) reranking algorithm to deeply exploit information from the similarity matrix during the retrieval process. This algorithm combines forward and backward ranking, extreme difference ratio, and other factors to reweight the similarity matrix, thereby further enhancing retrieval accuracy. Finally, we optimize the triplet loss function by introducing an intraclass distance term for matched image-text pairs, not only focusing on the relative distance between matched and unmatched pairs but also minimizing the distance within matched pairs. Experiments on four public remote sensing text-image datasets, including RSICD, RSITMD, UCM-Captions, and Sydney-Captions, demonstrate the effectiveness of our proposed method, achieving improvements over state-of-the-art methods, such as a 2.28% increase in mean Recall (mR) on the RSITMD dataset and a significant 4.65% improvement in R@1. The code is available athttps://github.com/ZbaoSun/CMPAGL. Zengbao Sun, Ming Zhao 0009, Gaorui Liu, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | The Bjøntegaard Bible Why Your Way of Comparing Video Codecs May Be WrongabstractIn this paper, we provide an in-depth assessment on the Bjøntegaard Delta. We construct a large data set of video compression performance comparisons using a diverse set of metrics including PSNR, VMAF, bitrate, and processing energies. These metrics are evaluated for visual data types such as classic perspective video, 360° video, point clouds, and screen content. As compression technology, we consider multiple hybrid video codecs as well as state-of-the-art neural network based compression methods. Using additional supporting points in-between standard points defined by parameters such as the quantization parameter, we assess the interpolation error of the Bjøntegaard-Delta (BD) calculus and its impact on the final BD value. From the analysis, we find that the BD calculus is most accurate in the standard application of rate-distortion comparisons with mean errors below 0.5 percentage points. For other applications and special cases, e.g., VMAF quality, energy considerations, or inter-codec comparisons, the errors are higher (up to 5 percentage points), but can be halved by using a higher number of supporting points. We finally come up with recommendations on how to use the BD calculus such that the validity of the resulting BD-values is maximized. Main recommendations are as follows: First, relative curve differences should be plotted and analyzed. Second, the logarithmic domain should be used for saturating metrics such as SSIM and VMAF. Third, BD values below a certain threshold indicated by the subset error should not be used to draw recommendations. Fourth, using two supporting points is sufficient to obtain rough performance estimates. Christian Herglotz, Hannah Och, Anna Meyer, Geetha Ramasubbu, Lena Eichermüller, Matthias Kränzler, Fabian Brand, Kristian Fischer 0001, Dat Thanh Nguyen, Andy Regensky, André Kaup |
IEEE Trans. Image Process. | 11 |
| 2023 | Image Super-Resolution Using T-Tetromino PixelsabstractFor modern high-resolution imaging sensors, pixel binning is performed in low-lighting conditions and in case high frame rates are required. To recover the original spatial resolution, single-image super-resolution techniques can be applied for upscaling. To achieve a higher image quality after upscaling, we propose a novel binning concept using tetromino-shaped pixels. It is embedded into the field of compressed sensing and the coherence is calculated to motivate the sensor layouts used. Next, we investigate the reconstruction quality using tetromino pixels for the first time in literature. Instead of using different types of tetrominoes as proposed elsewhere, we show that using a small repeating cell consisting of only four T-tetrominoes is sufficient. For reconstruction, we use a locally fully connected reconstruction (LFCR) network as well as two classical reconstruction methods from the field of compressed sensing. Using the LFCR network in combination with the proposed tetromino layout, we achieve superior image quality in terms of PSNR, SSIM, and visually compared to conventional single-image super-resolution using the very deep super-resolution (VDSR) network. For PSNR, a gain of up to +1.92 dB is achieved. Simon Grosche, Andy Regensky, Jürgen Seiler, André Kaup |
CVPR | 4 |
| 2023 | Saliency-Driven Hierarchical Learned Image Coding for MachinesabstractWe propose to employ a saliency-driven hierarchical neural image compression network for a machine-to-machine communication scenario following the compress-then-analyze paradigm. By that, different areas of the image are coded at different qualities depending on whether salient objects are located in the corresponding area. Areas without saliency are transmitted in latent spaces of lower spatial resolution in order to reduce the bitrate. The saliency information is explicitly derived from the detections of an object detection network. Furthermore, we propose to add saliency information to the training process in order to further specialize the different latent spaces. All in all, our hierarchical model with all proposed optimizations achieves 77.1 % bitrate savings over the latest video coding standard VVC on the Cityscapes dataset and with Mask R-CNN as analysis network at the decoder side. Thereby, it also outperforms traditional, non-hierarchical compression networks. Kristian Fischer 0001, Fabian Brand, Christian Blum 0004, André Kaup |
ICASSP | 4 |
| 2023 | A Novel Cross-Component Context Model for End-to-End Wavelet Image CodingabstractIn contrast to traditional compression techniques performing linear transforms, the latent space of popular compressive autoencoders is obtained from a learned nonlinear mapping and hard to interpret. In this paper, we explore a promising alternative approach for neural compression, with an autoencoder whose latent space represents a nonlinear wavelet decomposition. Previous work has shown that neural wavelet image coding can outperform HEVC. However, the approach codes color components independently, thereby ignoring inter-component dependencies. Hence, we propose a novel crosscomponent context model (CCM). With CCM, the entropy model for the chroma latent space can be conditioned on previously coded components exploiting correlations in the learned wavelet space. The proposed CCM outperforms the baseline model with average Bjøntegaard delta rate savings of 2.6 % and 1.6 % for the Kodak and Tecnick image sets. Also, our method is competitive with VVC and learning-based methods. Anna Meyer, André Kaup |
ICASSP | 2 |
| 2023 | Deep Probabilistic Model for Lossless Scalable Point Cloud Attribute CompressionabstractIn recent years, several point cloud geometry compression methods that utilize advanced deep learning techniques have been proposed, but there are limited works on attribute compression, especially lossless compression. In this work, we build an end-to-end multiscale point cloud attribute coding method (MNeT) that progressively projects the attributes onto multiscale latent spaces. The multiscale architecture provides an accurate context for the attribute probability modeling and thus minimizes the coding bitrate with a single network prediction. Besides, our method allows scalable coding that lower quality versions can be easily extracted from the losslessly compressed bitstream. We validate our method on a set of point clouds from MVUB and MPEG and show that our method outperforms recently proposed methods and on par with the latest G-PCC version 14. Besides, our coding time is substantially faster than G-PCC. Dat Thanh Nguyen, Kamal Gopikrishnan Nambiar, André Kaup |
ICASSP | 3 |
| 2023 | Image Segmentation for Improved Lossless Screen Content CompressionabstractIn recent years, it has been found that screen content images (SCI) can be effectively compressed based on appropriate probability modelling and suitable entropy coding methods such as arithmetic coding. The key objective is determining the best probability distribution for each pixel position. This strategy works particularly well for images with synthetic (textual) content. However, usually screen content images not only consist of synthetic but also pictorial (natural) regions. These images require diverse models of probability distributions to be optimally compressed. One way to achieve this goal is to separate synthetic and natural regions. This paper proposes a segmentation method that identifies natural regions enabling better adaptive treatment. It supplements a compression method known as Soft Context Formation (SCF) and operates as a pre-processing step. If at least one natural segment is found within the SCI, it is split into two subimages (natural and synthetic parts) and the process of modelling and coding is performed separately for both. For SCIs with natural regions, the proposed method achieves a bit-rate reduction of up to 11.6% and 1.52% with respect to HEVC and the previous version of the SCF. Shabhrish Reddy Uddehal, Tilo Strutz, Hannah Och, André Kaup |
ICASSP | 4 |
| 2023 | Spatially-Adaptive Learning-Based Image Compression with Hierarchical Multi-Scale Latent SpacesabstractAdaptive block partitioning is responsible for large gains in current image and video compression systems. This method is able to compress large stationary image areas with only a few symbols, while maintaining a high level of quality in more detailed areas. Current state-of-the-art neural-network-based image compression systems however use only one scale to transmit the latent space. In previous publications, we proposed RDONet, a scheme to transmit the latent space in multiple spatial resolutions. Following this principle, we extend a state-of-the-art compression network by a second hierarchical latent-space level to enable multi-scale processing. We extend the existing rate variability capabilities of RDONet by a gain unit. With that we are able to outperform an equivalent traditional autoencoder by 7% rate savings. Furthermore, we show that even though we add an additional latent space, the complexity only increases marginally and the decoding time can potentially even be decreased. Fabian Brand, Alexander Kopte, Kristian Fischer 0001, André Kaup |
ICIP | 4 |
| 2023 | Encoder Complexity Control in SVT-AV1 by Speed-Adaptive Preset SwitchingabstractCurrent developments in video encoding technology lead to continuously improving compression performance but at the expense of increasingly higher computational demands. Regarding the online video traffic increases during the last years and the concomitant need for video encoding, encoder complexity control mechanisms are required to restrict the processing time to a sufficient extent in order to find a reasonable trade-off between performance and complexity. We present a complexity control mechanism in SVT-AV1 by using speed-adaptive preset switching to comply with the remaining time budget. This method enables encoding with a user-defined time constraint within the complete preset range with an average precision of 8.9 % without introducing any additional latencies. Lena Eichermüller, Gaurang Chaudhari, Ioannis Katsavounidis, Zhijun Lei, Hassene Tmar, André Kaup, Christian Herglotz |
ICIP | 6 |
| 2023 | Processing Energy Modeling For Neural Network Based Image CompressionabstractNowadays, the compression performance of neural-network-based image compression algorithms outperforms state-of-the-art compression approaches such as JPEG or HEIC-based image compression. Unfortunately, most neural-network based compression methods are executed on GPUs and consume a high amount of energy during execution. Therefore, this paper performs an in-depth analysis on the energy consumption of state-of-the-art neural-network based compression methods on a GPU and show that the energy consumption of compression networks can be estimated using the image size with mean estimation errors of less than 7%. Finally, using a correlation analysis, we find that the number of operations per pixel is the main driving force for energy consumption and deduce that the network layers up to the second downsampling step are consuming most energy. Christian Herglotz, Fabian Brand, Andy Regensky, Felix Rievel, André Kaup |
ICIP | 5 |
| 2023 | Motion Plane Adaptive Motion Modeling for Spherical Video Coding in H.266/VVCabstractMotion compensation is one of the key technologies enabling the high compression efficiency of modern video coding standards. To allow compression of spherical video content, special mapping functions are required to project the video to the 2D image plane. Distortions inevitably occurring in these mappings impair the performance of classical motion models. In this paper, we propose a novel motion plane adaptive motion modeling technique (MPA) for spherical video that allows to perform motion compensation on different motion planes in 3D space instead of having to work on the - in theory arbitrarily mapped - 2D image representation directly. The integration of MPA into the state-of-the-art H.266/VVC video coding standard shows average Bjøntegaard Delta rate savings of 1.72% with a peak of 3.37% based on PSNR and 1.55% with a peak of 2.92% based on WS-PSNR compared to VTM-14.2. Andy Regensky, Christian Herglotz, André Kaup |
ICIP | 3 |
| 2023 | Improving Spherical Image Resampling Through Viewport-AdaptivityabstractThe conversion between different spherical image and video projection formats requires highly accurate resampling techniques in order to minimize the inevitable loss of information. Suitable resampling algorithms such as nearest neighbor, linear or cubic resampling are readily available. However, no generally applicable resampling technique exploits the special properties of spherical images so far. Thus, we propose a novel viewport-adaptive resampling (VAR) technique that takes the spherical characteristics of the underlying resampling problem into account. VAR can be applied to any mesh-to-mesh capable resampling algorithm and shows significant gains across all tested techniques. In combination with frequency-selective resampling, VAR outperforms conventional cubic resampling by more than 2 dB in terms of WS-PSNR. A visual inspection and the evaluation of further metrics such as PSNR and SSIM support the positive results. Andy Regensky, Viktoria Heimann, André Kaup |
ICIP | 4 |
| 2023 | Cross Spectral Image Reconstruction Using a Deep Guided Neural NetworkabstractCross spectral camera arrays, where each camera records different spectral content, are becoming increasingly popular for RGB, multispectral and hyperspectral imaging, since they are capable of a high resolution in every dimension using off-the-shelf hardware. For these, it is necessary to build an image processing pipeline to calculate a consistent image data cube, i.e., it should look like as if every camera records the scene from the center camera. Since the cameras record the scene from a different angle, this pipeline needs a reconstruction component for pixels that are not visible to peripheral cameras. For that, a novel deep guided neural network (DGNet) is presented. Since only little cross spectral data is available for training, this neural network is highly regularized. Furthermore, a new data augmentation process is introduced to generate the cross spectral content. On synthetic and real multispectral camera array data, the proposed network out-performs the state of the art by up to 2 dB in terms of PSNR on average. Besides, DGNet also tops its best competitor in terms of SSIM as well as in runtime by a factor of nearly 12. Moreover, a qualitative evaluation reveals visually more appealing results for real camera array data. Frank Sippel, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2023 | End-to-End Lidar-Camera Self-Calibration for Autonomous VehiclesabstractAutonomous vehicles are equipped with a multi-modal sensor setup to enable the car to drive safely. The initial calibration of such perception sensors is a highly matured topic and is routinely done in an automated factory environment. However, an intriguing question arises on how to maintain the calibration quality throughout the vehicle’s operating duration. Another challenge is to calibrate multiple sensors jointly to ensure no propagation of systemic errors. In this paper, we propose Camera Lidar Calibration Network (CaLiCaNet), an end-to-end deep self-calibration network which addresses the automatic calibration problem for pinhole camera and Lidar. We jointly predict the camera intrinsic parameters (focal length and distortion) as well as Lidar-Camera extrinsic parameters (rotation and translation), by regressing feature correlation between the camera image and the Lidar point cloud. The network is arranged in a Siamese-twin structure to constrain the network features learning to a mutually shared feature in both point cloud and camera (Lidar-camera constraint). Evaluation using KITTI datasets shows that we achieve 0.154° and 0.059 m accuracy with a reprojection error of 0.028 pixel with a single-pass inference. We also provide an ablative study of how our end-to-end learning architecture offers lower terminal loss (21% decrease in rotation loss) compared to isolated calibration. Arya Rachman, Jürgen Seiler, André Kaup |
IV | 3 |
| 2023 | On Interpolation of Subjective Rate-Distortion Curves for Video Coder ComparisonabstractWhen comparing two video coders, in order to obtain meaningful and unambiguous results, it is necessary to interpolate the rate-distortion curves. Typical methods like the BjØntegaard delta rate use piece-wise cubic interpolation for that purpose. This works well for the unbounded quality metric PSNR. For fitting the curve of saturating quality metrics, which are common when measuring subjective or perceptual quality, the use of a logistic fitting function was proposed with SCENIC earlier. However, the interpolation accuracy has not been validated. In this paper, we compare different methods for interpolating rate-distortion curves of subjective metrics. We come to the conclusion that SCENIC is in fact better than piece-wise cubic interpolation for four supporting points. With more supporting points, however it is restricted by the number of free parameters, and piece-wise cubic interpolation should be used. We furthermore show that logarithmic rescaling has a small benefit for the interpolation accuracy of MOS values. Fabian Brand, Christian Herglotz, André Kaup |
QoMEX | 3 |
| 2023 | Power Reduction Opportunities on End-User Devices in Quality-Steady Video StreamingabstractThis paper uses a crowdsourced dataset of online video streaming sessions to investigate opportunities to reduce the power consumption while considering QoE. For this, we base our work on prior studies which model both the end-user's QoE and the end-user device's power consumption with the help of high-level video features such as the bitrate, the frame rate, and the resolution. On top of existing research, which focused on reducing the power consumption at the same QoE optimizing video parameters, we investigate potential power savings by other means such as using a different playback device, a different codec, or a predefined maximum quality level. We find that based on the power consumption of the streaming sessions from the crowdsourcing dataset, devices could save more than 55% of power if all participants adhere to low-power settings. Christian Herglotz, Werner Robitza, Alexander Raake, Tobias Hoßfeld, André Kaup |
QoMEX | 5 |
| 2023 | Lossless Point Cloud Geometry and Attribute Compression Using a Learned Conditional Probability ModelabstractIn recent years, we have witnessed the presence of point cloud data in many aspects of our life, from immersive media, autonomous driving to healthcare, although at the cost of a tremendous amount of data. In this paper, we present an efficient lossless point cloud compression method that uses sparse tensor-based deep neural networks to learn point cloud geometry and color probability distributions. Our method represents a point cloud with both occupancy feature and three attribute features at different bit depths in a unified sparse representation. This allows us to efficiently exploit feature-wise and point-wise dependencies within point clouds using a sparse tensor-based neural network and thus build an accurate auto-regressive context model for an arithmetic coder. To the best of our knowledge, this is the first learning-based lossless point cloud geometry and attribute compression approach. Compared with the-state-of-the-art lossless point cloud compression method from Moving Picture Experts Group (MPEG), our method achieves 22.6% reduction in total bitrate on a diverse set of test point clouds while having 49.0% and 18.3% rate reduction on geometry and color attribute component, respectively. Dat Thanh Nguyen, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multitask Learning for SAR Ship Detection With Gaussian-Mask Joint SegmentationabstractDetecting ships from synthetic aperture radar (SAR) images is inherently subject to the limitations of SAR’s imaging mechanism. SAR object detection technology has rapidly advanced in recent years due to deep learning based techniques for detecting objects from optical images. SAR ship detection still faces some challenges due to the strong speckle noise, complex surroundings, and variety of scales. This paper proposes a multitask learning framework for object detection (MLDet) detect ships in SAR images. The proposed end-to-end framework consists of object detection task, speckle supression task and target segmentation task. Firstly, an angle classification loss with aspect ratio weighting is explored during object detection to improve the accuracy by making the detector sensitive to the periodicity of angular and the aspect ratio of objects. Secondly, the speckle supression task employs a dual-feature fusion attention mechanism to suppress noisy background information and fuse shallow features and denoising features, which helps MLDet be more robust to speckle noise. Thirdly, the target segmentation task with rotated Gaussian-mask is explored to further help the detection network to extract the regions of intersect from the cluttered background, as well as improving the detection efficiency through pixel-by-pixel prediction. The rotated Gaussian-mask for ship modeling ensures that the center of a ship has the highest probablities to be labeled as an object, and the probablities of the remaining regions are gradually reduced under a Gaussian distribution. Assisted by these two subtasks, the shallow level features are robust to speckle noise and reliably support deep level feature learning. In addition, the weighted rotated boxes fusion (WRBF) strategy is adopted to combine the predictions of multi-direction anchors for rotated objects, and eliminate the anchors beyond the boundary as well as the anchors with high overlap rates but low scores. A large number of experiments and comprehensive evaluations on SAR ship detection datasets SSDD+ and HRSID have shown the effectiveness and superiority of the proposed method. The code is available from https://github.com/zx152/MLDet. Ming Zhao 0009, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Learning True Rate-Distortion-Optimization for End-To-End Image CompressionabstractEven though rate-distortion optimization is a crucial part of traditional image and video compression, not many approaches exist which transfer this concept to end-to-end-trained image compression. Most frameworks contain static compression and decompression models which are fixed after training, so efficient rate-distortion optimization is not possible. In a previous work, we proposed RDONet [1], which enables an RDO approach comparable to adaptive block partitioning in HEVC. In this paper, we enhance the training and boost the model performance by introducing low-complexity estimations of the RDO result into the training. It is well known that the setup during the training should be as close as possible to the setup during inference. Since including an RDO search into the training is computationally not feasible, we propose a fast variance-based criterion which we can use to approximate the RDO behavior during training. Additionally, we use the same criterion to propose a variance-adaptive RDO initialization which converges faster, needs fewer RDO passes. We can therefore decrease the inference runtime significantly. With our novel training method, we achieve average Bjøntegaard rate savings of 19.6% in MS-SSIM over the previous RDONet model [1], which equals rate savings of 27.3% over a comparable conventional deep image coder, similar to [2]. With our novel initialization method, we can reduce the number of RDO passes to one. Therefore, we need only half the time for RDO, while still saving 26.8% rate. When we do not perform an RDO search but instead only rely on the initial estimation, we still obtain remarkable rate-savings of 23.6%, needing no additional time for an RDO search. The full paper is available on arXiv [3]. Fabian Brand, Kristian Fischer 0001, Alexander Kopte, André Kaup |
DCC | 4 |
| 2022 | A Low-Parametric Model for Bit-Rate Estimation of VVC Residual CodingabstractThere are many tasks within video compression which re-quire fast bit rate estimation. As an example, rate-control algorithms are only feasible because it is possible to estimate the required bit rate without needing to encode the en-tire block. With residual coding technology becoming more and more sophisticated, the corresponding bit rate models re-quire more advanced features. In this work, we propose a set of four features together with a linear model, which is able to estimate the rate of arbitrary residual blocks which were compressed using the VVC standard. Our method out-performs other methods which were used for the same task both in terms of mean absolute error and mean relative error. Our model deviates by less than 4 bit on average over a large dataset of natural images. Fabian Brand, Christian Herglotz, André Kaup |
ICASSP | 3 |
| 2022 | Evaluation of Video Coding for Machines without Ground TruthabstractIn the emerging field of video coding for machines, video datasets with pristine video quality and high-quality annotations are required for a comprehensive evaluation. However, existing video datasets with detailed annotations are severely limited in size and video quality. Thus, current methods have to either evaluate their codecs on still images or on already compressed data. To mitigate this problem, we propose an evaluation method based on pseudo ground-truth data from the field of semantic segmentation to the evaluation of video coding for machines. Through extensive evaluation, this paper shows that the proposed ground-truth-agnostic evaluation method results in an acceptable absolute measurement error below 0.7 percentage points on the Bjøntegaard Delta Rate compared to using the true ground truth for mid-range bitrates. We evaluate on the three tasks of semantic segmentation, instance segmentation, and object detection. Lastly, we utilize the ground-truth-agnostic method to measure the coding performances of the VVC compared against HEVC on the Cityscapes sequences. This reveals that the coding position has a significant influence on the task performance. Kristian Fischer 0001, Markus Hofbauer, Christopher B. Kuhn, Eckehard G. Steinbach, André Kaup |
ICASSP | 5 |
| 2022 | P-Frame Coding with Generalized Difference: A Novel Conditional Coding ApproachabstractMotion compensated inter frame prediction is a common component of all video coders and greatly reduces temporal redundancy. With the rise of deep learning-based image and video compression, this concept has been successfully taken over from traditional coding approaches. These approaches offer a larger flexibility than traditional transform coding and therefore enable efficient conditional coding. In this work, we develop a novel conditional coding approach based on the generalized difference and generalized sum operators. This approach is a special case of a general conditional coder and has a very small complexity overhead. We also propose an extension which enables dynamic content-adaptive switching between conditional and residual coding. We show that the extended generalized difference coding outperforms both residual and conditional coding, saving 27.8% Bjøntegaard delta rate compared to the former. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2022 | Learning Frequency-Specific Quantization Scaling in VVC for Standard-Compliant Task-Driven Image CodingabstractToday, visual data is often analyzed by a neural network without any human being involved, which demands for specialized codecs. For standard-compliant codec adaptations towards certain information sinks, HEVC or VVC provide the possibility of frequency-specific quantization with scaling lists. This is a well-known method for the human visual system, where scaling lists are derived from psycho-visual models. In this work, we employ scaling lists when performing VVC intra coding for neural networks as information sink. To this end, we propose a novel data-driven method to obtain optimal scaling lists for arbitrary neural networks. Experiments with Mask R-CNN as information sink reveal that coding the Cityscapes dataset with the proposed scaling lists result in peak bitrate savings of 8.9 % over VVC with constant quantization. By that, our approach also outperforms scaling lists optimized for the human visual system. The generated scaling lists can be found under https://github.com/FAU-LMS/VCM_scaling_lists. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
ICIP | 4 |
| 2022 | Domain Adaptation for Unknown Image Distortions in Instance SegmentationabstractData-driven techniques for machine vision heavily depend on the training data to sufficiently resemble the data occurring during test and application. However, in practice unknown distortion can lead to a domain gap between training and test data, impeding the performance of a machine vision system. With our proposed approach this domain gap can be closed by unpaired learning of the pristine-to-distortion mapping function of the unknown distortion. This learned mapping function may then be used to emulate the unknown distortion in the training data. Employing a fixed setup, our approach is independent from prior knowledge of the distortion. Within this work, we show that we can effectively learn unknown distortions at arbitrary strengths. When applying our approach to instance segmentation in an autonomous driving scenario, we achieve results comparable to an oracle with knowledge of the distortion. An average gain in mean Average Precision (mAP) of up to 0.19 can be achieved. Maximiliane Gruber, Fabian Brand, Alina Mosebach, Jürgen Seiler, André Kaup |
ICIP | 5 |
| 2022 | Frequency-Selective Geometry Upsampling of Point CloudsabstractThe demand for high-resolution point clouds has increased throughout the last years. However, capturing high-resolution point clouds is expensive and thus, frequently replaced by upsampling of low-resolution data. Most state-of-the-art methods are either restricted to a rastered grid, incorporate normal vectors, or are trained for a single use case. We propose to use the frequency selectivity principle, where a frequency model is estimated locally that approximates the surface of the point cloud. Then, additional points are inserted into the approximated surface. Our novel frequency-selective geometry upsampling shows superior results in terms of subjective as well as objective quality compared to state-of-the-art methods for scaling factors of 2 and 4. On average, our proposed method shows a 4.4 times smaller point-to-point error than the second best state-of-the-art PU-Net for a scale factor of 4. Viktoria Heimann, Andreas Spruck, André Kaup |
ICIP | 3 |
| 2022 | Beyond Bjøntegaard: Limits of Video Compression Performance ComparisonsabstractFor 20 years, the gold standard to evaluate the performance of video codecs is to calculate average differences between rate-distortion curves, also called the "Bjøntegaard Delta". With the help of this tool, the compression performance of codecs can be compared. In the past years, we could observe that the calculus was also deployed for other metrics than bitrate and distortion in terms of peak signal-to-noise ratio, for example other quality metrics such as video multi-method assessment fusion or hardware-dependent metrics such as the decoding energy. However, it is unclear whether the Bjøntegaard Delta is a valid way to evaluate these metrics. To this end, this paper reviews several interpolation methods and evaluates their accuracy using different performance metrics. As a result, we propose to use a novel approach based on Akima interpolation, which returns the most accurate results for a large variety of performance metrics. The approximation accuracy of this new method is determined to be below a bound of 1.5%. Christian Herglotz, Matthias Kränzler, Ruben Mons, André Kaup |
ICIP | 4 |
| 2022 | Optimized Decoding-Energy-Aware Encoding In Practical VVC ImplementationsabstractThe optimization of the energy demand is crucial for modern video codecs. Previous studies show that the energy demand of VVC decoders can be improved by more than 50% if specific coding tools are disabled in the encoder. However, those approaches increase the bit rate by over 20% if the concept is applied to practical encoder implementations such as VVenC. Therefore, in this work, we investigate VVenC and study possibilities to reduce the additional bit rate, while still achieving low-energy decoding at reasonable encoding times. We show that encoding using our proposed coding tool profiles, the decoding energy efficiency is improved by over 25% with a bit rate increase of less than 5% with respect to standard encoding. Furthermore, we propose a second coding tool profile targeting maximum energy savings, which achieves 34% of energy savings at bitrate increases below 15%. Matthias Kränzler, Adam Wieckowski, Geetha Ramasubbu, Benjamin Bross, André Kaup, Detlev Marpe, Christian Herglotz |
ICIP | 5 |
| 2022 | Learning-Based Lossless Point Cloud Geometry Coding Using Sparse TensorsabstractMost point cloud compression methods operate in the voxel or octree domain which is not the original representation of point clouds. Those representations either remove the geometric information or require high computational power for processing. In this paper, we propose a context-based lossless point cloud geometry compression that directly processes the point representation. Operating on a point representation allows us to preserve geometry correlation between points and thus to obtain an accurate context model while significantly reduce the computational cost. Specifically, our method uses a sparse convolution neural network to estimate the voxel occupancy sequentially from the x, y, z input data. Experimental results show that our method outperforms the state-of-the-art geometry compression standard from MPEG with average rate savings of 52% on a diverse set of point clouds from four different datasets. Dat Thanh Nguyen, André Kaup |
ICIP | 2 |
| 2022 | Camera Self-Calibration: Deep Learning from Driving ScenesabstractPrior to driving, cameras embedded in an autonomous driving system need to be calibrated intrinsically. Calibration is crucial to ensure that safety-related perception functions can reliably perceive the environment. Vehicle cameras are also exposed to mechanical perturbations requiring periodic re-calibration with regular uses. The current widely-accepted calibration approaches are based on robust but potentially demanding target-based methods. Such methods require a car to be taken offline and rely on static infrastructure and operators. Targetless online calibration approaches exist but remain largely unadopted due to the accuracy gaps compared to the classical methods. We propose a deep-learning-based self-calibration strategy for the vehicular camera that learns from driving scenes—they make an inherently large-scale dataset—and is validated back-to-back against checkerboard reprojection error. Our approach results in a 2.5% decrease in subpixel reprojection error compared to the existing deep-learning-based approaches. We also demonstrate its practical application in the automotive domain. Arya Rachman, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2022 | Modeling the HEVC Encoding Energy Using the Encoder Processing TimeabstractThe global significance of energy consumption of video communication renders research on the energy need of video coding an important task. To do so, usually, a dedicated setup is needed that measures the energy of the encoding and decoding system. However, such measurements are costly and complex. To this end, this paper presents the results of an exhaustive measurement series using the x265 encoder implementation of HEVC and analyzes the relation between encoding time and encoding energy. Finally, we introduce a simple encoding energy estimation model which employs the encoding time of a lightweight encoding process to estimate the encoding energy of complex encoding configurations. The proposed model reaches a mean estimation error of 11.35% when averaged over all presets. The results from this work are useful when the encoding energy estimate is required to develop new energy-efficient video compression algorithms. Geetha Ramasubbu, André Kaup, Christian Herglotz |
ICIP | 2 |
| 2022 | Increasing the Accuracy of a Neural Network Using Frequency Selective Mesh-to-Grid ResamplingabstractNeural networks are widely used for almost any task of recognizing image content. Even though much effort has been put into investigating efficient network architectures, optimizers, and training strategies, the influence of image interpolation on the performance of neural networks is not well studied. Furthermore, research has shown that neural networks are often sensitive to minor changes in the input image leading to drastic drops of their performance. Therefore, we propose the use of keypoint agnostic frequency selective mesh-to-grid resampling (FSMR) for the processing of input data for neural networks in this paper. This model-based interpolation method already showed that it is capable of outperforming common interpolation methods in terms of PSNR. Using an extensive experimental evaluation we show that depending on the network architecture and classification task the application of FSMR during training aids the learning process. Furthermore, we show that the usage of FSMR in the application phase is beneficial. The classification accuracy can be increased by up to 4.31 percentage points for ResNet50 and the Oxtlower17 dataset. Andreas Spruck, Viktoria Heimann, André Kaup |
ISCAS | 3 |
| 2022 | Jointly Resampling and Reconstructing Corrupted Images for Image Classification using Frequency-Selective Mesh-to-Grid ResamplingabstractNeural networks became the standard technique for image classification throughout the last years. They are extracting image features from a large number of images in a training phase. In a following test phase, the network is applied to the problem it was trained for and its performance is measured. In this paper, we focus on image classification. The amount of visual data that is interpreted by neural networks grows with the increasing usage of neural networks. Mostly, the visual data is transmitted from the application side to a central server where the interpretation is conducted. If the transmission is disturbed, losses occur in the transmitted images. These losses have to be reconstructed using postprocessing. In this paper, we incorporate the widely applied bilinear and bicubic interpolation and the high-quality reconstruction Frequency-Selective Reconstruction (FSR) for the reconstruction of corrupted images. However, we propose to use Frequency-Selective Mesh-to-Grid Resampling (FSMR) for the joint reconstruction and resizing of corrupted images. The performance in terms of classification accuracy of EfficientNetB0, DenseNet121, DenseNet201, ResNet50 and ResNet152 is examined. Results show that the reconstruction with FSMR leads to the highest classification accuracy for most networks. Average improvements of up to 6.7 percentage points are possible for DenseNet121. Viktoria Heimann, Andreas Spruck, André Kaup |
MMSP | 3 |
| 2022 | Optimal Filter Selection for Multispectral Object Classification Using Fast Binary SearchabstractWhen designing multispectral imaging systems for classifying different spectra it is necessary to choose a small number of filters from a set with several hundred different ones. Tackling this problem by full search leads to a tremendous number of possibilities to check and is NP-hard. In this paper we introduce a novel fast binary search for optimal filter selection that guarantees a minimum distance metric between the different spectra to classify. In our experiments, this procedure reaches the same optimal solution as with full search at much lower complexity. The desired number of filters influences the full search in factorial order while the fast binary search stays constant. Thus, fast binary search allows to find the optimal solution of all combinations in an adequate amount of time and avoids prevailing heuristics. Moreover, our fast binary search algorithm outperforms other filter selection techniques in terms of misclassified spectra in a real-world classification problem. Frank Sippel, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2022 | On Benefits and Challenges of Conditional Interframe Video Coding in Light of Information TheoryabstractThe rise of variational autoencoders for image and video compression has opened the door to many elaborate coding techniques. One example here is the possibility of conditional interframe coding. Here, instead of transmitting the residual between the original frame and the predicted frame (often obtained by motion compensation), the current frame is transmitted under the condition of knowing the prediction signal. In practice, conditional coding can be straightforwardly implemented using a conditional autoencoder, which has also shown good results in recent works. In this paper, we provide an information theoretical analysis of conditional coding for inter frames and show in which cases gains compared to traditional residual coding can be expected. We also show the effect of information bottlenecks which can occur in practical video coders in the prediction signal path due to the network structure, as a consequence of the data-processing theorem or due to quantization. We demonstrate that conditional coding has theoretical benefits over residual coding but that there are cases in which the benefits are quickly canceled by small information bottlenecks of the prediction signal. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 3 |
| 2022 | Learning-Based Conditional Image Coder Using Color SeparationabstractRecently, image compression codecs based on Neural Networks (NN) outperformed the state-of-art classic ones such as BPG, an image format based on HEVC intra. However, the typical NN codec has high complexity, and it has limited options for parallel data processing. In this work, we propose a conditional separation principle that aims to improve parallelization and lower the computational requirements of an NN codec. We present a Conditional Color Separation (CCS) codec which follows this principle. The color components of an image are split into primary and non-primary ones. The processing of each component is done separately, by jointly trained networks. Our approach allows parallel processing of each component, flexibility to select different channel numbers, and an overall complexity reduction. The CCS codec uses over 40% less memory, has 2x faster encoding and 22% faster decoding speed, with only 4% BD-rate loss in RGB PSNR compared to our baseline model over BPG. Panqi Jia, Ahmet Burakhan Koyuncu, Georgii Gaikov, Alexander Karabutov, Elena Alshina, André Kaup |
PCS | 6 |
| 2022 | Device Interoperability for Learned Image Compression with Weights and Activations QuantizationabstractLearning-based image compression has improved to a level where it can outperform traditional image codecs such as HEVC and VVC in terms of coding performance. In addition to good compression performance, device interoperability is essential for a compression codec to be deployed, i.e., encoding and decoding on different CPUs or GPUs should be error-free and with negligible performance reduction. In this paper, we present a method to solve the device interoperability problem of a state-of-the-art image compression network. We implement quantization to entropy networks which output entropy parameters. We suggest a simple method which can ensure cross-platform encoding and decoding, and can be implemented quickly with minor performance deviation, of 0.3% BD-rate, from floating point model results. Esin Koyuncu, Timofey Solovyev, Elena Alshina, André Kaup |
PCS | 4 |
| 2022 | Advanced Design Space Exploration for Joint Energy and Quality optimization for VVCabstractIn recent studies, it could be shown that the energy demand of Versatile Video Coding (VVC) decoders can be twice as high as comparable High Efficiency Video Coding (HEVC) decoders. A significant part of this increase in complexity is attributed to the usage of new coding tools. By using a design space exploration algorithm, it was shown that the energy demand of VVC-coded sequences could be reduced if different coding tool profiles were used for the encoding process. This work extends the algorithm with several optimization strategies, methodological adjustments to optimize perceptual quality, and a new minimization criterion. As a result, we significantly improve the Pareto front, and the rate-distortion and energy efficiency of the state-of-the-art design space exploration. Therefore, we show an energy demand reduction of up to 47% with less than 30% additional bit rate, or a reduction of over 35% with approximately 6% additional bit rate. Matthias Kränzler, André Kaup, Christian Herglotz |
PCS | 2 |
| 2022 | A Bit Stream Feature-Based Energy Estimator for HEVC Software EncodingabstractThe total energy consumption of today’s video coding systems is globally significant and emphasizes the need for sustainable video coder applications. To develop such sustainable video coders, the knowledge of the energy consumption of state-of-the-art video coders is necessary. For that purpose, we need a dedicated setup that measures the energy of the encoding and decoding system. However, such measurements are costly and laborious. To this end, this paper presents an energy estimator that uses a subset of bit stream features to accurately estimate the energy consumption of the HEVC software encoding process. The proposed model reaches a mean estimation error of 4.88% when averaged over presets of the x265 encoder implementation. The results from this work help to identify properties of encoding energy-saving bit streams and, in turn, are useful for developing new energy-efficient video coding algorithms. Geetha Ramasubbu, André Kaup, Christian Herglotz |
PCS | 2 |
| 2022 | Modeling of Energy Consumption and Streaming Video QoE using a Crowdsourcing DatasetabstractIn the past decade, we have witnessed an enormous growth in the demand for online video services. Recent studies estimate that nowadays, more than 1% of the global greenhouse gas emissions can be attributed to the production and use of devices performing online video tasks. As such, research on the true power consumption of devices and their energy efficiency during video streaming is highly important for a sustainable use of this technology. At the same time, over-the-top providers strive to offer high-quality streaming experiences to satisfy user expectations. Here, energy consumption and QoE partly depend on the same system parameters. Hence, a joint view is needed for their evaluation. In this paper, we perform a first analysis of both end-user power efficiency and Quality of Experience of a video streaming service. We take a crowdsourced dataset comprising 447,000 streaming events from YouTube and estimate both the power consumption and perceived quality. The power consumption is modeled based on previous work which we extended towards predicting the power usage of different devices and codecs. The user-perceived QoE is estimated using a standardized model. Our results indicate that an intelligent choice of streaming parameters can optimize both the QoE and the power efficiency of the end user device. Further, the paper discusses limitations of the approach and identifies directions for future research. Christian Herglotz, Werner Robitza, Matthias Kränzler, André Kaup, Alexander Raake |
QoMEX | 4 |
| 2022 | Enhancement Layer Coding for Chroma Sub-Sampled Screen Content VideoabstractPrevalent video codec implementations deployed in the field often solely support chrominance sub-sampled video data represented by the YCbCr 4:2:0 format. For certain applications like screen sharing, however, chroma sub-sampling leads to disturbing artifacts, especially for text or graphics with thin lines. It is desirable to reduce these artifacts while maintaining compatibility with all user devices. For this reason, an enhancement layer coding framework for chroma-format scalable coding with specific focus on screen content is proposed in this paper. Based on an analysis of screen content data characteristics, the enhancement layer codec is optimized specifically for this content class, is of low algorithmic complexity, and applicable with any image or video codec for base layer compression in a joint coding system. The system is intentionally not designed as a general purpose lossless YCbCr 4:4:4 coding scheme, but instead chooses to close a quality gap that prevalent video codecs do not address. Experimental analysis reveals an average BD-PSNR gain in the RGB domain of 1.0 dB comparing the proposed two-layer scalable coding approach to single layer compression using base layer coding only. In relation to simulcast coding, an average BDR RGB difference of 8.1% is observed. Andreas Heindel, Benjamin Prestele, Alexander Gehlert, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Rate-Distortion Optimal Transform Coefficient Selection for Unoccupied Regions in Video-Based Point Cloud CompressionabstractThis paper presents a novel method to determine rate-distortion optimized transform coefficients for efficient compression of videos generated from point clouds. The method exploits a generalized frequency selective extrapolation approach that iteratively determines rate-distortion-optimized coefficients for all basis functions of two-dimensional discrete cosine and sine transforms. The method is applied to blocks containing both occupied and unoccupied pixels in video based point cloud compression for HEVC encoding. In the proposed algorithm, only the values of the transform coefficients are changed such that resulting bit streams are compliant to the V-PCC standard. For all-intra coded point clouds, bitrate savings of more than 4% for geometry and more than 6% for texture error metrics with respect to standard encoding can be observed. These savings are more than twice as high as savings obtained using competing methods from literature. In the randomaccess case, our proposed method outperforms competing V-PCC methods by more than 0.5%. Christian Herglotz, Nils Genser, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Energy Efficient Video Decoding for VVC Using a Greedy Strategy-Based Design Space ExplorationabstractIP traffic has increased significantly in recent years, and it is expected that this progress will continue. Recent studies report that the viewing of online video content accounts for a share of 1% of the global greenhouse gas emissions. To reduce the data traffic of video streaming, the new standard Versatile Video Coding (VVC) has been finalized in 2020. In this paper, the energy efficiency of two different VVC decoders is analyzed in detail. Furthermore, we propose a design space exploration that uses an algorithm based on a greedy strategy to derive coding tool profiles that optimize the energy demand of the decoder. We show that the algorithm derives optimal coding tool profiles for a subset of coding tools. Additionally, we propose profiles that reduce the energy demand of VVC decoders and provide energy savings of more than 50% for sequences with 4K resolution. Thereby, we will also show that the proposed profiles can have a lower decoding energy demand than comparable HEVC-encoded bit streams while also having a significantly lower bit rate. Matthias Kränzler, Christian Herglotz, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Saliency-Driven Versatile Video Coding for Neural Object DetectionabstractSaliency-driven image and video coding for humans has gained importance in the recent past. In this paper, we pro-pose such a saliency-driven coding framework for the video coding for machines task using the latest video coding standard Versatile Video Coding (VVC). To determine the salient regions before encoding, we employ the real-time-capable object detection network You Only Look Once (YOLO) in combination with a novel decision criterion. To measure the coding quality for a machine, the state-of-the-art object segmentation network Mask R-CNN was applied to the decoded frame. From extensive simulations we find that, compared to the reference VVC with a constant quality, up to 29 % of bitrate can be saved with the same detection accuracy at the decoder side by applying the proposed saliency-driven framework. Besides, we compare YOLO against other, more traditional saliency detection methods. Kristian Fischer 0001, Felix Fleckenstein, Christian Herglotz, André Kaup |
ICASSP | 4 |
| 2021 | Frame Rate Up-Conversion Using Key Point Agnostic Frequency-Selective Mesh-to-Grid ResamplingabstractHigh frame rates are desired in many fields of application. As in many cases the frame repetition rate of an already captured video has to be increased, frame rate up-conversion (FRUC) is of high interest. We conduct a motion compensated approach. From two neighboring frames, the motion is estimated and the neighboring pixels are shifted along the motion vector into the frame to be reconstructed. For displaying, these irregularly distributed mesh pixels have to be resampled onto regularly spaced grid positions. We use the model-based key point agnostic frequency-selective mesh-to-grid resampling (AFSMR) for this task and show that AFSMR works best for applications that contain irregular meshes with varying densities. AFSMR gains up to 3.2 dB in contrast to the already high performing frequency-selective mesh-to-grid resampling (FSMR). Additionally, AFSMR increases the run time by a factor of 11 relative to FSMR. Viktoria Heimann, Andreas Spruck, André Kaup |
ICASSP | 3 |
| 2021 | A Novel Viewport-Adaptive Motion Compensation Technique for Fisheye Video
Andy Regensky, Christian Herglotz, André Kaup |
ICASSP | 3 |
| 2021 | Intra To Inter: Towards Intra Prediction for Learning-Based Video Coders Using Optical FlowabstractTraditional video coders often rely on a block structure for transmission. Here each block is coded separately and sequentially and for each block the encoder can decide whether to use intra or inter prediction. This way, inter and intra prediction can be mixed within a single frame. This has advantages when new areas are uncovered, which were not present in the reference frame, and can hence not be predicted well. These areas are typically predicted using intra prediction. Currently much research goes into end-to-end-trained video coders which do not operate on a block level and typically use dense motion fields for inter prediction. There it is more difficult to incorporate intra prediction for uncovered regions. In this paper we propose a novel concept which enables us to reinterpret classical angular intra prediction in a way that we can transmit it as part of the dense motion field. We can save an average of 18% rate for the transmission of the motion vectors for the same quality of the prediction image. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2021 | Analysis Of Neural Image Compression Networks For Machine-To-Machine CommunicationabstractVideo and image coding for machines (VCM) is an emerging field that aims to develop compression methods resulting in optimal bitstreams when the decoded frames are analyzed by a neural network. Several approaches already exist improving classic hybrid codecs for this task. However, neural compression networks (NCNs) have made an enormous progress in coding images over the last years. Thus, it is reasonable to consider such NCNs, when the information sink at the decoder side is a neural network as well. Therefore, we build-up an evaluation framework analyzing the performance of four state-of-the-art NCNs, when a Mask R-CNN is segmenting objects from the decoded image. The compression performance is measured by the weighted average precision for the Cityscapes dataset. Based on that analysis, we find that networks with leaky ReLU as non-linearity and training with SSIM as distortion criteria results in the highest coding gains for the VCM task. Furthermore, it is shown that the GAN-based NCN architecture achieves the best coding performance and even out-performs the recently standardized Versatile Video Coding (VVC) for the given scenario. Kristian Fischer 0001, Christian Forsch, Christian Herglotz, André Kaup |
ICIP | 4 |
| 2021 | Novel Consistency Check For Fast Recursive Reconstruction Of Non-Regularly Sampled Video DataabstractQuarter sampling is a novel sensor design that allows for an acquisition of higher resolution images without increasing the number of pixels. When being used for video data, one out of four pixels is measured in each frame. Effectively, this leads to a non-regular spatio-temporal sub-sampling. Compared to purely spatial or temporal sub-sampling, this allows for an increased reconstruction quality, as aliasing artifacts can be reduced. For the fast reconstruction of such sensor data with a fixed mask, recursive variant of frequency selective reconstruction (FSR) was proposed. Here, pixels measured in previous frames are projected into the current frame to support its reconstruction. In doing so, the motion between the frames is computed using template matching. Since some of the motion vectors may be erroneous, it is important to perform a proper consistency checking. In this paper, we propose faster consistency checking methods as well as a novel recursive FSR that uses the projected pixels different than in literature and can handle dynamic masks. Altogether, we are able to significantly increase the reconstruction quality by + 1.01 dB compared to the state-of-the-art recursive reconstruction method using a fixed mask. Compared to a single frame reconstruction, an average gain of about + 1.52 dB is achieved for dynamic masks. At the same time, the computational complexity of the consistency checks is reduced by a factor of 13 compared to the literature algorithm. Simon Grosche, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2021 | Robust Deep Neural Object Detection and Segmentation for Automotive Driving Scenario with Compressed Image DataabstractDeep neural object detection or segmentation networks are commonly trained with pristine, uncompressed data. However, in practical applications the input images are usually deteriorated by compression that is applied to efficiently transmit the data. Thus, we propose to add deteriorated images to the training process in order to increase the robustness of the two state-of-the-art networks Faster and Mask R-CNN. Throughout our paper, we investigate an autonomous driving scenario by evaluating the newly trained models on the Cityscapes dataset that has been compressed with the upcoming video coding standard Versatile Video Coding (VVC). When employing the models that have been trained with the proposed method, the weighted average precision of the R-CNNs can be increased by up to 3.68 percentage points for compressed input images, which corresponds to bitrate savings of nearly 48 %. Kristian Fischer 0001, Christian Blum 0004, Christian Herglotz, André Kaup |
ISCAS | 4 |
| 2021 | A Novel End-To-End Network for Reconstruction of Non-Regularly Sampled Image Data Using Locally Fully Connected LayersabstractQuarter sampling and three-quarter sampling are novel sensor concepts that enable the acquisition of higher resolution images without increasing the number of pixels. This is achieved by non-regularly covering parts of each pixel of a low-resolution sensor such that only one quadrant or three quadrants of the sensor area of each pixel is sensitive to light. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be further enhanced using super resolution algorithms such as the very deep super resolution network (VDSR). In this paper, we propose a novel end-to-end neural network to reconstruct high resolution images from non-regularly sampled sensor data. The network is a concatenation of a locally fully connected reconstruction network (LFCR) and a standard VDSR network. Altogether, using a three-quarter sampling sensor with our novel neural network layout, the image quality in terms of PSNR for the Urban100 dataset can be increased by 2.96 dB compared to the state-of-the-art approach. Compared to a low-resolution sensor with VDSR, a gain of 1.11 dB is achieved. Simon Grosche, Fabian Brand, André Kaup |
MMSP | 3 |
| 2021 | Frequency-Selective Mesh-to-Mesh Resampling for Color Upsampling of Point CloudsabstractWith the increased use of virtual and augmented reality applications, the importance of point cloud data rises. High-quality capturing of point clouds is still expensive and thus, the need for point cloud super-resolution or point cloud upsampling techniques emerges. In this paper, we propose an interpolation scheme for color upsampling of three-dimensional color point clouds. As a point cloud represents an object’s surface in three-dimensional space, we first conduct a local transform of the surface into a two-dimensional plane. Secondly, we propose to apply a novel Frequency-Selective Mesh-to-Mesh Resampling (FSMMR) technique for the interpolation of the points in 2D. FSMMR generates a model of weighted superpositions of basis functions on scattered points. This model is then evaluated for the final points in order to increase the resolution of the original point cloud. Evaluation shows that our approach outperforms common interpolation schemes. Visual comparisons of the jaguar point cloud underlines the quality of our upsampling results. The high performance of FSMMR holds for various sampling densities of the input point cloud. Viktoria Heimann, Andreas Spruck, André Kaup |
MMSP | 3 |
| 2021 | Hyperspectral Image Reconstruction from Multispectral Images Using Non-Local FilteringabstractUsing light spectra is an essential element in many applications, for example, in material classification. Often this information is acquired by using a hyperspectral camera. Unfortunately, these cameras have some major disadvantages like not being able to record videos. Therefore, multispectral cameras with wide-band filters are used, which are much cheaper and are often able to capture videos. However, using multispectral cameras requires an additional reconstruction step to yield spectral information. Usually, this reconstruction step has to be done in the presence of imaging noise, which degrades the reconstructed spectra severely. Typically, same or similar pixels are found across the image with the advantage of having independent noise. In contrast to state-of-the-art spectral reconstruction methods which only exploit neighboring pixels by block-based processing, this paper introduces non-local filtering in spectral reconstruction. First, a block-matching procedure finds similar non-local multispectral blocks. Thereafter, the hyperspectral pixels are reconstructed by filtering the matched multispectral pixels collaboratively using a reconstruction Wiener filter. The proposed novel procedure even works under very strong noise. The method is able to lower the spectral angle up to 18% and increase the peak signal-to-noise-ratio up to 1.1dB in noisy scenarios compared to state-of-the-art methods. Moreover, the visual results are much more appealing. Frank Sippel, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2021 | Switchable Motion Models for Non-Block-Based Inter Prediction in Learning-Based Video CodingabstractMost state-of-the-art video coders rely on a block structure. For inter-frame prediction, motion vectors are transmitted per block. For example in VVC, the coder can choose between a translational or an affine motion model on a block level, depending on the content. In non-block-based coding, which is on the rise since the development of end-to-end learning based image compression, the motion vectors have to be transmitted differently. Due to the missing inherent block structure, switching between different motion models presents a challenge, but also an opportunity. In this paper, we propose an alternative approach to efficiently signal additional information regarding the motion model to improve the quality of the motion compensated image. Using our methods, we are able to increase the quality of the prediction image in our scenario by 0.40 dB on average and by up to 0.88 dB for sequences with strong and complex motion at the same rate. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 3 |
| 2021 | Optimization of Probability Distributions for Residual Coding of Screen ContentabstractProbability distribution modeling is the basis for most competitive methods for lossless coding of screen content. One such state-of-the-art method is known as soft context formation (SCF). For each pixel to be encoded, a probability distribution is estimated based on the neighboring pattern and the occurrence of that pattern in the already encoded image. Using an arithmetic coder, the pixel color can thus be encoded very efficiently, provided that the current color has been observed before in association with a similar pattern. If this is not the case, the color is instead encoded using a color palette or, if it is still unknown, via residual coding. Both palette-based coding and residual coding have significantly worse compression efficiency than coding based on soft context formation. In this paper, the residual coding stage is improved by adaptively trimming the probability distributions for the residual error. Furthermore, an enhanced probability modeling for indicating a new color depending on the occurrence of new colors in the neighborhood is proposed. These modifications result in a bitrate reduction of up to 2.9 % on average. Compared to HEVC (HM-16.21 + SCM-8.8) and FLIF, the improved SCF method saves on average about 11 % and 18 % rate, respectively. Hannah Och, Tilo Strutz, André Kaup |
VCIP | 3 |
| 2021 | Spatio-spectral Image Reconstruction Using Non-local FilteringabstractIn many image processing tasks it occurs that pixels or blocks of pixels are missing or lost in only some channels. For example during defective transmissions of RGB images, it may happen that one or more blocks in one color channel are lost. Nearly all modern applications in image processing and transmission use at least three color channels, some of the applications employ even more bands, for example in the infrared and ultraviolet area of the light spectrum. Typically, only some pixels and blocks in a subset of color channels are distorted. Thus, other channels can be used to reconstruct the missing pixels, which is called spatio-spectral reconstruction. Current state-of-the-art methods purely rely on the local neighborhood, which works well for homogeneous regions. However, in high-frequency regions like edges or textures, these methods fail to properly model the relationship between color bands. Hence, this paper introduces non-local filtering for building a linear regression model that describes the inter-band relationship and is used to reconstruct the missing pixels. Our novel method is able to increase the PSNR on average by 2 dB and yields visually much more appealing images in high-frequency regions. Frank Sippel, Jürgen Seiler, André Kaup |
VCIP | 3 |
| 2020 | Anefficient Alternative to Network Pruning Through Ensemble LearningabstractConvolutional Neural Networks (CNNs) currently represent the best tool for classification of image content. CNNs are trained in order to develop generalized expressions in form of unique features to distinguish different classes. During this process, one or more filter weights might develop the same or similar values. In this case, the redundant filters can be pruned without damaging accuracy.Unlike normal pruning methods, we investigate the possibility of replacing a full-sized convolutional neural network with an ensemble of its narrow versions. Empirically, we show that the combination two narrow networks, which has only half of the original parameter count in total, is able to contest and even outperform the full-sized original network. In other words, a pruning rate of 50% can be achieved without any compromise on accuracy. Furthermore, we introduce a novel approach on ensemble learning called ComboNet, which increases the accuracy of the ensemble even further. Martin Pöllot, André Kaup |
ICASSP | 3 |
| 2020 | On Intra Video Coding And In-Loop Filtering For Neural Object Detection NetworksabstractClassical video coding for satisfying humans as the final user is a widely investigated field of studies for visual content, and common video codecs are all optimized for the human visual system (HVS). But are the assumptions and optimizations also valid when the compressed video stream is analyzed by a machine? To answer this question, we compared the performance of two state-of-the-art neural detection networks when being fed with deteriorated input images coded with HEVC and VVC in an autonomous driving scenario using intra coding. Additionally, the impact of the three VVC in-loop filters when coding images for a neural network is examined. The results are compared using the mean average precision metric to evaluate the object detection performance for the compressed inputs. Throughout these tests, we found that the Bjøntegaard Delta Rate savings with respect to PSNR of 22.2 % using VVC instead of HEVC cannot be reached when coding for object detection networks with only 13.6 % in the best case. Besides, it is shown that disabling the VVC in-loop filters SAO and ALF results in bitrate savings of 6.4 % compared to the standard VTM at the same mean average precision. Kristian Fischer 0001, Christian Herglotz, André Kaup |
ICIP | 3 |
| 2020 | Joint Content-Adaptive Dictionary Learning And Sparse Selective Extrapolation For Cross-Spectral Image ReconstructionabstractNumerous applications deal with distorted images, e.g., during transmission over lossy channels in image coding, or to reconstruct occlusions in multi-color multi-view imaging scenarios. In many cases, not all spectral channels are distorted, or the losses distribute differently between the channels. Recently, efforts were made to develop reference guided approaches to reconstruct distorted spectral content. However, these methods are only able to exploit information from a single reference, even if there are multiple spectral components available that could be used for reconstruction. To overcome this limitation, a novel method is proposed in this paper, which introduces a content-adaptive dictionary learning approach that is applicable to an arbitrary number of references. With the novel approach, an average PSNR gain of approx. 1.5 dB is achieved in comparison to the best recently published state-of-the-art methods for a single reference component. Moreover, the presence of multiple guidance channels can enhance the reconstruction by another 6 dB on average. At the same time, the complexity of the novel algorithm is significantly lower than for previously published methods. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2020 | Deep Learning Based Cross-Spectral Disparity Estimation For Stereo ImagingabstractRecently, cross-spectral stereo-camera setups found their way from special applications to mass market, especially in smartphones, automotive systems, or drones. In the following, a novel concept is introduced to bring stereo cameras and cross-spectral disparity estimation together. So far, either monomodal stereo algorithms exist that are not suitable for cross-spectral image registration, or structural template matching is applied that achieves a low quality. To overcome these limitations, a technique is proposed to synthesize arbitrary spectral components from widely available color stereo databases, and to retrain mono-modal deep learning methods. In this contribution, the estimation of spectral bands based on random processes is shown together with noise models, which also allow for a robust registration of narrowband components. The theoretical examination is completed by an extensive evaluation, including a self-manufactured cross-spectral camera setup. In comparison to state-of-the-art techniques, the end-point error is on average reduced by a factor of seven. Nils Genser, Andreas Spruck, Jürgen Seiler, André Kaup |
ICIP | 4 |
| 2020 | Enhanced Image Reconstruction From Quarter Sampling Measurements Using An Adapted Very Deep Super Resolution NetworkabstractQuarter sampling is a novel sensor concept that enables the acquisition of higher resolution images without increasing the number of pixels. This is achieved by covering three quarters of each pixel of a low-resolution sensor such that only one quadrant of the sensor area of each pixel is sensitive to light. By randomly masking different parts, effectively a non-regular sampling of a higher resolution image is performed. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be enhanced using super resolution algorithms. Recently, algorithms based on machine learning such as the Very Deep Super Resolution network (VDSR) proofed to be successful for this task. In this work, we transfer the concepts of VDSR to the special case of quarter sampling. Besides adapting the network layout to take advantage of the case of quarter sampling, we introduce a novel data augmentation technique enabled by quarter sampling. Altogether, using the quarter sampling sensor, the image quality in terms of PSNR can be increased by + 0.67 dB for the Urban 100 dataset compared to using a low-resolution sensor with VDSR. Simon Grosche, Kristian Fischer 0001, Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 5 |
| 2020 | Decoding Energy Modeling For Versatile Video CodingabstractIn previous research, it was shown that the software decoding energy demand of High Efficiency Video Coding (HEVC) can be reduced by 15% by using a decoding-energy-ratedistortion optimization algorithm. To achieve this, the energy demand of the decoder has to be modeled by a bit stream feature-based model with sufficiently high accuracy. Therefore, we propose two bit stream feature-based models for the upcoming Versatile Video Coding (VVC) standard. The newly introduced models are compared with models from literature, which are used for HEVC. An evaluation of the proposed models reveals that the mean estimation error is similar to the results of the literature and yields an estimation error of 1.85% with 10-fold cross-validation. Matthias Kränzler, Christian Herglotz, André Kaup |
ICIP | 3 |
| 2020 | A Triangulation-Based Backward Adaptive Motion Field Subsampling SchemeabstractOptical flow procedures are used to generate dense motion fields which approximate true motion. Such fields contain a large amount of data and if we need to transmit such a field, the raw data usually exceeds the raw data of the two images it was computed from. In many scenarios, however, it is of interest to transmit a dense motion field efficiently. Most prominently this is the case in inter prediction for video coding. In this paper we propose a transmission scheme based on subsampling the motion field. Since a field which was subsampled with a regularly spaced pattern usually yields suboptimal results, we propose an adaptive subsampling algorithm that preferably samples vectors at positions where changes in motion occur. The subsampling pattern is fully reconstructable without the need for signaling of position information. We show an average gain of 2.95 dB in average end point error compared to regular subsampling. Furthermore we show that an additional prediction stage can improve the results by an additional 0.43 dB, gaining 3.38 dB in total. Fabian Brand, Jürgen Seiler, Elena Alshina, André Kaup |
MMSP | 4 |
| 2020 | Video Coding for Machines with Feature-Based Rate-Distortion OptimizationabstractCommon state-of-the-art video codecs are optimized to deliver a low bitrate by providing a certain quality for the final human observer, which is achieved by rate-distortion optimization (RDO). But, with the steady improvement of neural networks solving computer vision tasks, more and more multimedia data is not observed by humans anymore, but directly analyzed by neural networks. In this paper, we propose a standard-compliant feature-based RDO (FRDO) that is designed to increase the coding performance, when the decoded frame is analyzed by a neural network in a video coding for machine scenario. To that extent, we replace the pixel-based distortion metrics in conventional RDO of VTM-8.0 with distortion metrics calculated in the feature space created by the first layers of a neural network. Throughout several tests with the segmentation network Mask R-CNN and single images from the Cityscapes dataset, we compare the proposed FRDO and its hybrid version HFRDO with different distortion measures in the feature space against the conventional RDO. With HFRDO, up to 5.49% bitrate can be saved compared to the VTM-8.0 implementation in terms of Bjøntegaard Delta Rate and using the weighted average precision as quality metric. Additionally, allowing the encoder to vary the quantization parameter results in coding gains for the proposed HFRDO of up 9.95% compared to conventional VTM. Kristian Fischer 0001, Fabian Brand, Christian Herglotz, André Kaup |
MMSP | 4 |
| 2020 | Key Point Agnostic Frequency-Selective Mesh-to-Grid Image Resampling using Spectral WeightingabstractMany applications in image processing require re-sampling of arbitrarily located samples onto regular grid positions. This is important in frame-rate up-conversion, super-resolution, and image warping among others. A state-of-the-art high quality model-based resampling technique is frequency-selective mesh-to-grid resampling which requires pre-estimation of key points. In this paper, we propose a new key point agnostic frequency-selective mesh-to-grid resampling that does not depend on pre-estimated key points. Hence, the number of data points that are included is reduced drastically and the run time decreases significantly. To compensate for the key points, a spectral weighting function is introduced that models the optical transfer function in order to favor low frequencies more than high ones. Thereby, resampling artefacts like ringing are supressed reliably and the resampling quality increases. On average, the new AFSMR is conceptually simpler and gains up to 1.2 dB in terms of PSNR compared to the original mesh-to-grid resampling while being approximately 14.5 times faster. Viktoria Heimann, Nils Genser, André Kaup |
MMSP | 3 |
| 2020 | Decoding-Energy Optimal Video Encoding For x265abstractThis paper presents optimal x265-encoder configurations and an enhanced optimization algorithm for minimizing the software decoding energy of HEVC-coded videos. We reach this goal with two contributions. First, we perform a detailed analysis on the influence of various encoder settings on the decoding energy. Second, we include an enhanced version of an algorithm called decoding-energy-rate-distortion optimization into x265, which we optimize for fast and efficient encoding. This algorithm introduces the estimated decoding energy as an additional optimization criterion into the rate-distortion cost function. We evaluate the extended encoder in terms of bitrate, distortion, and decoding energy, where we perform energy measurements to prove the superior energy efficiency. We find that the combination of the `fastdecoding' tuning option of x265 with the enhanced decoding-energy-rate-distortion optimization leads to 27.2% and 26.0% of decoding energy savings for OpenHEVC and HM decoding, respectively. At the same time, compression efficiency losses of 38.2% and negligible decreases in encoder runtime of 0.39% can be observed. Christian Herglotz, Marco Bader, Kristian Fischer 0001, André Kaup |
MMSP | 4 |
| 2020 | A Comparative Analysis of the Time and Energy Demand of Versatile Video Coding and High Efficiency Video Coding Reference DecodersabstractThis paper investigates the decoding energy and decoding time demand of VTM-7.0 in relation to HM-16.20. We present the first detailed comparison of two video codecs in terms of software decoder energy consumption. The evaluation shows that the energy demand of the VTM decoder is increased significantly compared to HM and that the increase depends on the coding configuration. For the coding configuration randomaccess, we find that the decoding energy is increased by over 80% at a decoding time increase of over 70%. Furthermore, results indicate that the energy demand increases by up to 207% when Single Instruction Multiple Data (SIMD) instructions are disabled, which corresponds to the HM implementation style. By measurements, it is revealed that the coding tools MIP, AMVR, TPM, LFNST, and MTS increase the energy efficiency of the decoder. Furthermore, we propose a new coding configuration based on our analysis, which reduces the energy demand of the VTM decoder by over 17% on average. Matthias Kränzler, Christian Herglotz, André Kaup |
MMSP | 3 |
| 2020 | Optimizing Rate-Distortion Performance of Motion Compensated Wavelet Lifting with Denoised Prediction and UpdateabstractEfficient lossless coding of medical volume data with temporal axis can be achieved by motion compensated wavelet lifting. As side benefit, a scalable bit stream is generated, which allows for displaying the data at different resolution layers, highly demanded for telemedicine applications. Additionally, the similarity of the temporal base layer to the input sequence is preserved by the use of motion compensated temporal filtering. However, for medical sequences the overall rate is increased due to the specific noise characteristics of the data. The use of denoising filters inside the lifting structure can improve the compression efficiency significantly without endangering the property of perfect reconstruction. However, the design of an optimum filter is a crucial task. In this paper, we present a new method for selecting the optimal filter strength for a certain denoising filter in a rate-distortion sense. This allows to minimize the required rate based on a single input parameter for the encoder to control the requested distortion of the temporal base layer. Daniela Lanz, André Kaup |
MMSP | 2 |
| 2020 | Multispectral Image Compression Based on HEVC Using Pel-Recursive Inter-Band PredictionabstractRecent developments in optical sensors enable a wide range of applications for multispectral imaging, e.g., in surveillance, optical sorting, and life-science instrumentation. Increasing spatial and spectral resolution allows creating higher quality products, however, it poses challenges in handling such large amounts of data. Consequently, specialized compression techniques for multispectral images are required. High Efficiency Video Coding (HEVC) is known to be the state of the art in efficiency for both video coding and still image coding. In this paper, we propose a cross-spectral compression scheme for efficiently coding multispectral data based on HEVC. Extending intra picture prediction by a novel inter-band predictor, spectral as well as spatial redundancies can be effectively exploited. Dependencies among the current band and further spectral references are considered jointly by adaptive linear regression modeling. The proposed backward prediction scheme does not require additional side information for decoding. We show that our novel approach is able to outperform state-of-the-art lossy compression techniques in terms of rate-distortion performance. On different data sets, average Bjøntegaard delta rate savings of 82 % and 55 % compared to HEVC and a reference method from literature are achieved, respectively. Anna Meyer, Nils Genser, André Kaup |
MMSP | 3 |
| 2020 | Real-Time Frequency Selective Reconstruction through Register-Based Argmax CalculationabstractFrequency Selective Reconstruction (FSR) is a state-of-the-art algorithm for solving diverse image reconstruction tasks, where a subset of pixel values in the image is missing. However, it entails a high computational complexity due to its iterative, blockwise procedure to reconstruct the missing pixel values. Although the complexity of FSR can be considerably decreased by performing its computations in the frequency domain, the reconstruction procedure still takes multiple seconds up to multiple minutes depending on the parameterization. However, FSR has the potential for a massive parallelization greatly improving its reconstruction time. In this paper, we introduce a novel highly parallelized formulation of FSR adapted to the capabilities of modern GPUs and propose a considerably accelerated calculation of the inherent argmax calculation. Altogether, we achieve a 100-fold speed-up, which enables the usage of FSR for real-time applications. Andy Regensky, Simon Grosche, Jürgen Seiler, André Kaup |
MMSP | 4 |
| 2020 | On Versatile Video Coding at UHD with Machine-Learning-Based Super-ResolutionabstractCoding 4K data has become of vital interest in recent years, since the amount of 4K data is significantly increasing. We propose a coding chain with spatial down- and upscaling that combines the next-generation VVC codec with machine learning based single image super-resolution algorithms for 4K. The investigated coding chain, which spatially downscales the 4K data before coding, shows superior quality than the conventional VVC reference software for low bitrate scenarios. Throughout several tests, we find that up to 12 % and 18% Bj⊘ntegaard delta rate gains can be achieved on average when coding 4K sequences with VVC and QP values above 34 and 42, respectively. Additionally, the investigated scenario with up- and downscaling helps to reduce the loss of details and compression artifacts, as it is shown in a visual example. Kristian Fischer 0001, Christian Herglotz, André Kaup |
QoMEX | 3 |
| 2020 | Matched Quality Evaluation of Temporally Downsampled Videos with Non-Integer FactorsabstractRecent research has shown that temporal downsampling of high-frame-rate sequences can be exploited to improve the rate-distortion performance in video coding. However, until now, research only targeted downsampling factors of powers of two, which greatly restricts the potential applicability of temporal downsampling. A major reason is that traditional, objective quality metrics such as peak signal-to-noise ratio or more recent approaches, which try to mimic subjective quality, can only be evaluated between two sequences whose frame rate ratio is an integer value. To relieve this problem, we propose a quality evaluation method that allows calculating the distortion between two sequences whose frame rate ratio is fractional. The proposed method can be applied to any full-reference quality metric. Christian Herglotz, Geetha Ramasubbu, André Kaup |
QoMEX | 3 |
| 2020 | Introducing Latent Space Correlation to Conditional Autoencoders for Intra PredictionabstractIntra prediction has been an integral part of image and video coders for a long time. A predominant method is angular prediction that extends the reference area in a certain angle into the block. Recently many deep-learning-based methods have been proposed. Since intra prediction uses multiple modes this usually requires training a large number of networks. With a conditional autoencoder we are able to generate an arbitrary number of modes with only one network. In this paper we introduce a novel loss function enforcing a spatially correlated latent space and extend the network structure to the same end. Thereby we are able to propose a simple spatial mode prediction scheme using most-probable-mode lists. By replacing matrix-based intra prediction in VVC with our method, we obtain average rate savings of 0.84% with peak gains of 2.37%. Fabian Brand, Jürgen Seiler, André Kaup |
VCIP | 3 |
| 2020 | DENESTO: A Tool for Video Decoding Energy Estimation and VisualizationabstractIn previous research, it is shown that the decoding energy demand of several video codecs can be estimated accurately by using bit stream feature-based models. Therefore, we show in this paper that the visualization with the Decoding Energy Estimation Tool (DENESTO) can help to improve the understanding of the energy demand of the decoder. Matthias Kränzler, Christian Herglotz, André Kaup |
VCIP | 3 |
| 2020 | FishUI: Interactive Fisheye Distortion VisualizationabstractFisheye lenses provide major benefits for many applications due to their large field of view. However, they come at the cost of strong radial distortions leading to problems in a variety of signal processing tasks which have been developed with perspective lenses in mind. As such, while state-of-the-art image and video codecs excel in reducing redundancy and irrelevance in content captured with perspective lenses, the coding gain reduces significantly when fisheye lenses are applied. To improve the understanding of distortions introduced by fisheye lenses with respect to perspective lenses, we provide an interactive user interface for the visualization of fisheye block distortions. Andy Regensky, Christian Herglotz, André Kaup |
VCIP | 3 |
| 2020 | Toward Bridging the Simulated-to-Real Gap: Benchmarking Super-Resolution on Real DataabstractCapturing ground truth data to benchmark super-resolution (SR) is challenging. Therefore, current quantitative studies are mainly evaluated on simulated data artificially sampled from ground truth images. We argue that such evaluations overestimate the actual performance of SR methods compared to their behavior on real images. Toward bridging this simulated-to-real gap, we introduce the Super-Resolution Erlangen (SupER) database, the first comprehensive laboratory SR database of all-real acquisitions with pixel-wise ground truth. It consists of more than 80k images of 14 scenes combining different facets: CMOS sensor noise, real sampling at four resolution levels, nine scene motion types, two photometric conditions, and lossy video coding at five levels. As such, the database exceeds existing benchmarks by an order of magnitude in quality and quantity. This paper also benchmarks 19 popular single-image and multi-frame algorithms on our data. The benchmark comprises a quantitative study by exploiting ground truth data and qualitative evaluations in a large-scale observer study. We also rigorously investigate agreements between both evaluations from a statistical perspective. One interesting result is that top-performing methods on simulated data may be surpassed by others on real data. Our insights can spur further algorithm development, and the publicy available dataset can foster future evaluations. Thomas Köhler 0004, Michel Bätz, Farzad Naderi, André Kaup, Andreas K. Maier, Christian Riess |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Camera Array for Multi-Spectral ImagingabstractRecently, many new applications arose for multispectral and hyper-spectral imaging. Besides modern biometric systems for identity verification, also agricultural and medical applications came up, which measure the health condition of plants and humans. Despite the growing demand, the acquisition of multi-spectral data is up to the present complicated. Often, expensive, inflexible, or low resolution acquisition setups are only obtainable for specific professional applications. To overcome these limitations, a novel camera array for multi-spectral imaging is presented in this article for generating consistent multispectral videos. As differing spectral images are acquired at various viewpoints, a geometrically constrained multi-camera sensor layout is introduced, which enables the formulation of novel registration and reconstruction algorithms to globally set up robust models. On average, the novel acquisition approach achieves a gain of 2.5 dB PSNR compared to recently published multi-spectral filter array imaging systems. At the same time, the proposed acquisition system ensures not only a superior spatial, but also a high spectral, and temporal resolution, while filters are flexibly exchangeable by the user depending on the application. Moreover, depth information is generated, so that 3D imaging applications, e.g., for augmented or virtual reality, become possible. The proposed camera array for multi-spectral imaging can be set up using off-the-shelf hardware, which allows for a compact design and employment in, e.g., mobile devices or drones, while being cost-effective. Nils Genser, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 3 |
| 2020 | Boosting Compressed Sensing Using Local Measurements and Sliding Window ReconstructionabstractIn the framework of compressed sensing, image data is measured using less measurements than the total number of pixels. Each measurement consists of a (random) linear combination of all pixels. Since image data is approximately sparse in an appropriate transform domain, a reasonable reconstruction is possible for many measurement matrices, especially for i.i.d. Gaussian measurement matrices. In a seemingly different field, non-regular sampling techniques such as three-quarter sampling have shown promising results to enhance the resolution of an imaging sensor by effectively sub-sampling a higher resolution image. Here, the measurements can be described as linear combinations of only three pixels, which can also be seen as a (spectral) compressed sensing measurement. Since each measurement is spatially localized, the reconstruction can be performed in overlapping sliding windows. In this work, we show that compressed sensing reconstruction algorithms can greatly benefit from such an overlapping sliding window reconstruction. Compared to conventional block-wise compressed sensing with i.i.d. Gaussian measurement matrices, the reconstruction quality in terms of the PSNR increases up to +5dB using small, local i.d.d. Gaussian measurement blocks. Additionally, we propose a local joint sparse deconvolution and extrapolation (L-JSDE) to reconstruct images from arbitrary local measurements. For several applications with local measurements we show that L-JSDE increases the PSNR by +2.2dB relative to conventional block-wise i.i.d. Gaussian measurements reconstructed with the state-of-the-art reconstruction algorithm D-AMP using the same overall sampling density. Simon Grosche, Andy Regensky, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 4 |
| 2020 | Graph-Based Compensated Wavelet Lifting for Scalable Lossless Coding of Dynamic Medical DataabstractLossless compression of dynamic 2-D+t and 3-D+t medical data is challenging regarding the huge amount of data, the characteristics of the inherent noise, and the high bit depth. Beyond that, a scalable representation is often required in telemedicine applications. Motion Compensated Temporal Filtering works well for lossless compression of medical volume data and additionally provides temporal, spatial, and quality scalability features. To achieve a high quality lowpass subband, which shall be used as a downscaled representative of the original data, graph-based motion compensation was recently introduced to this framework. However, encoding the motion information, which is stored in adjacency matrices, is not well investigated so far. This work focuses on coding these adjacency matrices to make the graph-based motion compensation feasible for data compression. We propose a novel coding scheme based on constructing so-called motion maps. This allows for the first time to compare the performance of graph-based motion compensation to traditional block-and mesh-based approaches. For high quality lowpass subbands our method is able to outperform the block-and mesh-based approaches by increasing the visual quality in terms of PSNR by 0.53 dB and 0.28 dB for CT data, as well as 1.04 dB and 1.90 dB for MR data, respectively, while the bit rate is reduced at the same time. Daniela Lanz, André Kaup |
IEEE Trans. Image Process. | 2 |
| 2019 | Deep Counting Model Extensions with Segmentation for Person DetectionabstractApplications like autonomous driving, surveillance, or any application that demands scene analysis requires object detection, semantic segmentation and instance segmentation. In this paper, we focus on the problem of detecting each instance of a specific category of objects, specifically persons. A novel method for object detection is proposed based on a deep counting model. The feature extractor of the deep counting model is extended with additional layers for segmenting specific instances. While the feature extractor of the deep counting model already focuses on the persons in the scene, the segmentation layers help to get a more accurate estimation of the foreground with persons and the instance segmentation is able to estimate separate instances of persons. Our proposed method outperforms other methods on the CUHK08 dataset with an Average Miss Rate (AMR) of 14% and on the PETS09 dataset with an AMR of 41%. Sanjukta Ghosh, Peter Amon, Andreas Hutter, André Kaup |
ICASSP | 4 |
| 2019 | Improving the Rate-distortion Model of HEVC Intra by Integrating the Maximum Absolute ErrorabstractNormally, the mean squared error in conjunction with the rate is used to optimize the compression in hybrid video coding. However, in some areas, such as medical image coding, not only the average error but also the maximum error should be considered. Recently, it has been shown that incorporating the maximum absolute error into the calculation of the rate-distortion optimization (RDO) of the HEVC encoder can improve image quality, while preserving the average error and the rate. In this paper, we optimize the inclusion of the maximum absolute error in the RDO by considering the whole quantization parameter range as well as the different prediction unit sizes. Compared to the existing extended RDO, up to 1.5 % bitrate can be saved for the same maximum absolute error reduction. Karina Jaskolka, André Kaup |
ICASSP | 2 |
| 2019 | Content Adaptive Wavelet Lifting for Scalable Lossless Video CodingabstractScalable lossless video coding is an important aspect for many professional applications. Wavelet-based video coding decomposes an input sequence into a lowpass and a highpass subband by filtering along the temporal axis. The lowpass subband can be used for previewing purposes, while the highpass subband provides the residual content for lossless reconstruction of the original sequence. The recursive application of the wavelet transform to the lowpass subband of the previous stage yields coarser temporal resolutions of the input sequence. This allows for lower bit rates, but also affects the visual quality of the lowpass subband. So far, the number of total decomposition levels is determined for the entire input sequence in advance. However, if the motion in the video sequence is strong or if abrupt scene changes occur, a further decomposition leads to a low-quality lowpass subband. Therefore, we propose a content adaptive wavelet transform, which locally adapts the depth of the decomposition to the content of the input sequence. Thereby, the visual quality of the low-pass subband is increased by up to 10.28 dB compared to a uniform wavelet transform with the same number of total decomposition levels, while the required rate is reduced by 1.06% additionally. Daniela Lanz, Christian Herbert, André Kaup |
ICASSP | 3 |
| 2019 | Motion-adapted Three-dimensional Frequency Selective ExtrapolationabstractIt has been shown, that high resolution images can be acquired using a low resolution sensor with non-regular sampling. Therefore, post-processing is necessary. In terms of video data, not only the spatial neighborhood can be used to assist the reconstruction, but also the temporal neighbor-hood. A popular and well performing algorithm for this kind of problem is the three-dimensional frequency selective extrapolation (3D-FSE) for which a motion adapted version is introduced in this paper. This proposed extension solves the problem of changing content within the area considered by the 3D-FSE, which is caused by motion within the sequence. Because of this motion, it may happen that regions are emphasized during the reconstruction that are not present in the original signal within the considered area. By that, false content is introduced into the extrapolated sequence, which affects the resulting image quality negatively. The novel extension, presented in the following, incorporates motion data of the sequence in order to adapt the algorithm accordingly, and compensates changing content, resulting in gains of up to 1.75 dB compared to the existing 3D-FSE. Andreas Spruck, Markus Jonscher, Jürgen Seiler, André Kaup |
ICASSP | 4 |
| 2019 | Joint Regression Modeling and Sparse Spatial Refinement for High-Quality Reconstruction of Distorted Color ImagesabstractHigh quality algorithms are demanded to reconstruct distorted color images in a variety of applications. For example, distortions can result during transmission over lossy channels in image coding or in multi-view imaging scenarios. In general, not all color channels are equally affected and the losses distribute differently in-between channels. However, state-ofthe-art methods process color channels independently and do not take the cross color information into account. Thus, a novel and powerful reconstruction algorithm is formulated in this contribution that exploits color as well as spatial information. Therefore, an initial model is estimated for the distorted area using a reference channel. Then, its quality is estimated and a spatial weighting model is set-up. Afterwards, the initial inter channel prediction is refined by generating a sparse model that takes the spatial correlations into account, as well. Consequently, the proposed method achieves an outstanding quality compared to state-of-the-art methods. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2019 | Deep Network Pruning for Object DetectionabstractWith the increasing success of deep learning in various applications, there is an increasing need to have deep models that can be used for deployment in real-time and/or resource constrained scenarios. In this context, this paper analyzes the pruning of deep models for object detection in order to reduce the number of weights and hence the number of computations. Very deep networks based on ResNet like architectures, like YOLOv3 have unique challenges when attempting to prune them. This paper proposes a network pruning technique based on agglomerative clustering for the feature extractor and using mutual information for the detector. The performance of the proposed techniques is also compared with that of a relatively shallow network, i.e., YOLOv2. A compression percentage of around 30% results in a 10% drop of mean average precision (mAP) in YOLOv3, whereas in YOLOv2 the drop was around 6% on the COCO dataset. Sanjukta Ghosh, Shashi K. K. Srinivasa, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 5 |
| 2019 | Intra Frame Prediction for Video Coding Using a Conditional Autoencoder ApproachabstractIntra prediction is a vital component of most modern image and video codecs. State of the art video codecs like High Efficiency Video Coding (HEVC) or the upcoming Versatile Video Coding (VVC) use a high number of directional modes. With the recent advances in deep learning, it is now possible to use artificial neural networks for intra frame prediction. Previously published approaches usually add additional ANN based modes or replace all modes by training several networks. In our approach, we use a single autoencoder network to first compress the original with help of already transmitted pixels to four parameters. We then use the parameters together with this support area to generate a prediction for the block. This way, we are able to replace all angular intra modes by a single ANN. In the experiments we compare our method with the intra prediction method currently used in the VVC Test Model (VTM). Using our method, we are able to gain up to 0.85 dB prediction PSNR with a comparable amount of side information or reduce the amount of side information by 2 bit per prediction unit with similar PSNR. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 3 |
| 2019 | Extending Video Decoding Energy Models for 360° and HDR Video Formats in HEVCabstractResearch has shown that decoder energy models are helpful tools for improving the energy efficiency in video playback applications. For example, an accurate feature-based bit stream model can reduce the energy consumption of the decoding process. However, until now only sequences of the SDR video format were investigated. Therefore, this paper shows that the decoding energy of HEVC-coded bit streams can be estimated precisely for different video formats and coding bit depths. Therefore, we compare a state-of-the-art model from the literature with a proposed model. We show that bit streams of the 360°, HDR, and fisheye video format can be estimated with a mean estimation error lower than 3.88% if the setups have the same coding bit depth. Furthermore, it is shown that on average, the energy demand for the decoding of bit streams with a bit depth of 10-bit is 55% higher than with 8-bit. Matthias Kränzler, Christian Herglotz, André Kaup |
PCS | 3 |
| 2019 | Scalable Lossless Coding of Dynamic Medical CT Data Using Motion Compensated Wavelet Lifting with Denoised Prediction and UpdateabstractProfessional applications like telemedicine often require scalable lossless coding of sensitive data. 3-D subband coding has turned out to offer good compression results for dynamic CT data and additionally provides a scalable representation in terms of low- and highpass subbands. To improve the visual quality of the lowpass subband, motion compensation can be incorporated into the lifting structure, but leads to inferior compression results at the same time. Prior work has shown that a denoising filter in the update step can improve the compression ratio. In this paper, we present a new processing order of motion compensation and denoising in the update step and additionally introduce a second denoising filter in the prediction step. This allows for reducing the overall file size by up to 4.4%, while the visual quality of the lowpass subband is kept nearly constant. Daniela Lanz, Franz Schilling, André Kaup |
PCS | 3 |
| 2019 | Motion Estimation for Fisheye Video With an Application to Temporal Resolution EnhancementabstractSurveying wide areas with only one camera is a typical scenario in surveillance and automotive applications. Ultra wide-angle fisheye cameras employed to that end produce video data with characteristics that differ significantly from conventional rectilinear imagery as obtained by perspective pinhole cameras. Those characteristics are not considered in typical image and video processing algorithms such as motion estimation, where translation is assumed to be the predominant kind of motion. This contribution introduces an adapted technique for use in block-based motion estimation that takes into the account the projection function of fisheye cameras and thus compensates for the non-perspective properties of fisheye videos. By including suitable projections, the translational motion model that would otherwise only hold for perspective material is exploited, leading to improved motion estimation results without altering the source material. In addition, we discuss extensions that allow for a better prediction of the peripheral image areas, where motion estimation falters due to spatial constraints, and further include calibration information to account for lens properties deviating from the theoretical function. Simulations and experiments are conducted on synthetic as well as real-world fisheye video sequences that are part of a data set created in the context of this paper. Average synthetic and real-world gains of 1.45 and 1.51 dB in luminance PSNR are achieved compared against conventional block matching. Furthermore, the proposed fisheye motion estimation method is successfully applied to motion compensated temporal resolution enhancement, where average gains amount to 0.79 and 0.76 dB. Andrea Eichenseer, Michel Bätz, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Decoding-Energy-Rate-Distortion Optimization for Video CodingabstractThis paper presents a method for generating coded video bit streams requiring less decoding energy than conventionally coded bit streams. To this end, we propose extending the standard rate-distortion optimization approach to also consider the decoding energy. In the encoder, the decoding energy is estimated during runtime using a feature-based energy model. These energy estimates are then used to calculate decoding-energy-rate-distortion costs that are minimized by the encoder. This ultimately leads to optimal tradeoffs between these three parameters. Therefore, we introduce the mathematical theory for describing decoding-energy-rate-distortion optimization and the proposed encoder algorithm is explained in detail. For rate-energy control, a new encoder parameter is introduced. Finally, measurements of the software decoding process for HEVC-coded bit streams are performed. Results show that this approach can lead to up to 30% of decoding energy reduction at a constant visual objective quality when accepting a bit rate increase at the same order of magnitude. Christian Herglotz, Andreas Heindel, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Dynamic Non-Regular Sampling Sensor Using Frequency Selective ReconstructionabstractBoth a high spatial and a high temporal resolution of images and videos are desirable in many applications, such as entertainment systems, monitoring manufacturing processes, or video surveillance. Due to the limited throughput of pixels per second, however, there is always a tradeoff between acquiring sequences with a high spatial resolution at a low temporal resolution or vice versa. In this paper, a modified sensor concept is proposed which is able to acquire both a high spatial and a high temporal resolution. This is achieved by dynamically reading out only a subset of pixels in a non-regular order to obtain a high temporal resolution. A full high spatial resolution is then obtained by performing a subsequent 3D reconstruction of the partially acquired frames. The main benefit of the proposed dynamic readout is that for each frame, different sampling points are available, which is advantageous since this information can significantly enhance the reconstruction quality of the proposed reconstruction algorithm. Using the proposed dynamic readout strategy, gains in the peak-signal-to-noise ratio (PSNR) of up to 8.55 dB are achieved compared with a static readout strategy. Compared with the other state-of-the-art techniques, such as frame rate up-conversion or super-resolution, which are also able to reconstruct sequences with both a high spatial and a high temporal resolution, average gains in PSNR of up to 6.58 dB are possible. Markus Jonscher, Jürgen Seiler, Daniela Lanz, Michael Schöberl, Michel Bätz, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Increasing Imaging Resolution by Non-Regular Sampling and Joint Sparse Deconvolution and ExtrapolationabstractIncreasing the resolution of image sensors has been a never ending struggle since many years. In this paper, we propose a novel image sensor layout, which allows for the acquisition of images at a higher resolution and improved quality. For this, the image sensor makes use of non-regular sampling, which reduces the impact of aliasing. Therewith, it allows for capturing details, which would not be possible with state-of-the-art sensors of the same number of pixels. The non-regular sampling is achieved by rotating prototype pixel cells in a non-regular fashion. As not the whole area of the pixel cell is sensitive to light, a non-regular spatial integration of the incident light is obtained. Based on the sensor output data, a high-resolution image can be reconstructed by performing a deconvolution with respect to the integration area and an extrapolation of the information to the insensitive regions of the pixels. To solve this challenging task, we introduce a novel joint sparse deconvolution and extrapolation algorithm. The union of non-regular sampling and the proposed reconstruction allows for achieving a higher resolution and therewith an improved imaging quality. Jürgen Seiler, Markus Jonscher, Thomas Ußmüller, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Online Decomposition of Compressive Streaming Data Using n-l1 Cluster-Weighted MinimizationabstractWe consider a decomposition method for compressive streaming data in the context of online compressive Robust Principle Component Analysis (RPCA). The proposed decomposition solves an n-ℓ1 cluster-weighted minimization to decompose a sequence of frames (or vectors), into sparse and low-rank components from compressive measurements. Our method processes a data vector of the stream per time instance from a small number of measurements in contrast to conventional batch RPCA, which needs to access full data. The n-ℓ1 cluster-weighted minimization leverages the sparse components along with their correlations with multiple previously-recovered sparse vectors. Moreover, the proposed minimization can exploit the structures of sparse components via clustering and re-weighting iteratively. The method outperforms the existing methods for both numerical data and actual video data. Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup |
DCC | 4 |
| 2018 | Robustness of Deep Convolutional Neural Networks for Image DegradationsabstractDeep convolutional neural networks (CNNs) have achieved tremendous success in image recognition tasks. However, the performance of CNNs degrade in situations where the input image is degraded by compression artifacts, blur or noise. In this paper, we analyze some of the common CNNs for degradations in images caused by Gaussian noise, blur as well as compression using JPEG and JPEG 2000 for the full range of quality factors. Moreover, we propose a method to improve the performance of CNNs for image classification in the presence of input images with degradations based on a master-slave architecture. Our method was found to perform well for individual and combined degradations. Sanjukta Ghosh, Rohan Shet, Peter Amon, Andreas Hutter, André Kaup |
ICASSP | 5 |
| 2018 | Iterative Optimization of Quarter Sampling Masks for Non-Regular Sampling SensorsabstractNon-regular sampling can reduce aliasing at the expense of noise. Recently, it has been shown that non-regular sampling can be carried out using a conventional regular imaging sensor when the surface of its individual pixels is partially covered. This technique is called quarter sampling (also 1/4 sampling), since only one quarter of each pixel is sensitive to light. For this purpose, the choice of a proper sampling mask is crucial to achieve a high reconstruction quality. In the scope of this work, we present an iterative algorithm to improve an arbitrary quarter sampling mask which results in a continuous increase of the reconstruction quality. In terms of the reconstruction algorithms, we test two simple algorithms, namely, linear interpolation and nearest neighbor interpolation, as well as two more sophisticated algorithms, namely, steering kernel regression and frequency selective extrapolation. Besides PSNR gains of +0.31 dB to +0.68 dB relative to a random quarter sampling mask resulting from our optimized mask, visually noticeable enhancements are perceptible. Simon Grosche, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2018 | Decoding Energy Estimation of an HEVC Hardware DecoderabstractThis paper investigates the energy a hardware decoder needs for decoding and displaying an HEVC-coded bit stream. Therefore, energy measurements are performed on a high amount of different bit streams. The results are used to validate the estimation accuracy of various decoder energy models. Furthermore, from power measurements, two new models are derived and proposed. The results show that all models are capable of estimating the decoding energy with a mean estimation error below 5%, where the lowest error of 1.04% was found for one of the proposed models. Furthermore, trained parameter values are given quantifying the influence of various sequence properties on the decoding energy: the resolution, the frame rate, the playback time, and the file size of the bit stream. Christian Herglotz, André Kaup |
ISCAS | 2 |
| 2018 | Improving HEVC Encoding of Rendered Video Data Using True Motion InformationabstractThis paper shows that motion vectors representing the true motion of an object in a scene can be exploited to improve the encoding process of computer generated video sequences. Therefore, a set of sequences is presented for which the true motion vectors of the corresponding objects were generated on a per-pixel basis during the rendering process. In addition to conventional motion estimation methods, it is proposed to exploit the computer generated motion vectors to enhance the rate-distortion performance. To this end, a motion vector mapping method including disocclusion handling is presented. It is shown that mean rate savings of 3.78% can be achieved. Christian Herglotz, David Muller, Andreas Weinlich, Frank Bauer 0001, Michael Ortner, Marc Stamminger, André Kaup |
ISM | 7 |
| 2018 | Sparse Hartley Modeling for Fast Image ExtrapolationabstractIn many cases, image and video signal processing demands for high quality extrapolation algorithms, e.g., to solve inpainting problems or to increase image resolution. Indeed, a high computational load goes hand in hand with a good reconstruction quality as expensive models are calculated to estimate the missing data. To overcome this, the high-speed sparse Hartley modeling is introduced in this paper. This algorithm is based on Frequency Selective Extrapolation. In contrast to that, the model generation is carried out in the Hartley domain to exploit its real-valued transform properties. Due to this, it is possible to reduce the computational complexity significantly as no complex-valued arithmetic operations have to be conducted. In other words, a slightly higher reconstruction quality is obtained, while the proposed method is more than three times faster than the competing Frequency Selective Extrapolation. Nils Genser, Simon Grosche, Jürgen Seiler, André Kaup |
MMSP | 4 |
| 2018 | Signal and Loss Geometry Aware Frequency Selective Extrapolation for Error ConcealmentabstractThe concealment of errors is an important task in image and video signal processing. Often, complex models are calculated to reconstruct the missing samples, which results in a long computation time. One method that achieves a very high reconstruction quality, but demands a moderate computational complexity only, is the block based Frequency Selective Extrapolation. Nevertheless, the reconstruction of a Full HD image can still take several minutes depending on the error pattern. To accelerate the computation, a novel algorithm is introduced in this paper that analyzes the adjacent, undistorted samples and optimizes the reconstruction parameters accordingly. Moreover, the analyzation is further used to adapt the partitioning of the blocks and the processing order. Similar to modern video codecs, e.g., High Efficiency Video Coding, a content based partitioning and processing is proposed as it takes the signal characteristics into account. Thus, the novel algorithm is on average four times faster than the state-of-the-art method and up to 25× quicker at best, while achieving a slightly higher reconstruction quality as well. Nils Genser, Jürgen Seiler, Franz Schilling, André Kaup |
PCS | 4 |
| 2018 | Decoding Energy Modeling For The Next Generation Video Codec Based On JemabstractThis paper shows that the processing energy of the decoder software for the next generation video codec can be accurately estimated using a feature based model. Therefore, a model from the literature is taken and extended to account for a high amount of the newly introduced coding modes. It is shown that using a selected set of 60 features, for a large set of more than 800 coded bit streams, a mean estimation error below 5% can be reached. Using the trained parameters of the model, the energy consumption of the decoder can be analyzed in detail such that, e.g., the coding modes consuming most processing energy can be identified. The model can be used inside the encoder for decoding- energy-rate-distortion optimization to generate decoding energy saving bit streams. Christian Herglotz, Matthias Kränzler, André Kaup |
PCS | 3 |
| 2018 | Joint Optimization of Rate, Distortion, and Maximum Absolute Error for Compression of Medical Volumes Using HEVC IntraabstractMany visual quality metrics are used to measure the quality of lossy compressed images and videos, and are integrated in the rate-distortion optimization of hybrid video codecs. However, most of the metrics focus on the average objective quality in a picture. In certain applications, like medical image processing, the maximum absolute error should be more weighted. In this paper, the rate-distortion optimization of HEVC is extended by integrating this error metric. Thus, rate, average error, and maximum absolute error are jointly optimized. Furthermore, a weighting factor α is included into the calculation of the optimization for balancing the ratio between average and maximum absolute error. For HEVC intra with α=0.25 an average maximum absolute error reduction of -25.63 can be achieved, while the bitrate increases slightly by 0.59%. Furthermore, the visual quality of the medical volumes improves and the data fidelity increases, i.e. less block artifacts appear and less structure disappear. Karina Jaskolka, André Kaup |
PCS | 2 |
| 2018 | Compression of Dynamic Medical CT Data Using Motion Compensated Wavelet Lifting with Denoised UpdateabstractFor the lossless compression of dynamic $3-\mathrm {D}+\mathrm {t}$ volumes as produced by medical devices like Computed Tomography, various coding schemes can be applied. This paper shows that 3-D subband coding outperforms lossless HEVC coding and additionally provides a scalable representation, which is often required in telemedicine applications. However, the resulting lowpass subband, which shall be used as a downscaled representative of the whole original sequence, contains a lot of ghosting artifacts. This can be alleviated by incorporating motion compensation methods into the subband coder. This results in a high quality lowpass subband but also leads to a lower compression ratio. In order to cope with this, we introduce a new approach for improving the compression efficiency of compensated 3-D wavelet lifting by performing denoising in the update step. We are able to reduce the file size of the lowpass subband by up to 1.64%, while the lowpass subband is still applicable for being used as a downscaled representative of the whole original sequence. Daniela Lanz, Jürgen Seiler, Karina Jaskolka, André Kaup |
PCS | 4 |
| 2018 | Sparse signal recovery with multiple prior information: Algorithm and measurement bounds
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup |
Signal Process. | 5 |
| 2018 | Modeling the Energy Consumption of the HEVC Decoding ProcessabstractIn this paper, we present a bit stream feature-based energy model that accurately estimates the energy required to decode a given High Efficiency Video Coding-coded bit stream. Therefore, we take a model from literature and extend it by explicitly modeling the in-loop filters, which was not done before. Furthermore, to prove its superior estimation performance, it is compared with seven different energy models from the literature. By using a unified evaluation framework, we show how accurately the required decoding energy for different decoding systems can be approximated. We give thorough explanations on the model parameters and explain how the model variables are derived. To show the modeling capabilities in general, we test the estimation performance for different decoding software and hardware solutions, where we find that the proposed model outperforms the models from the literature by reaching framewise mean estimation errors of less than 7% for software and less than 15% for hardware-based systems. Christian Herglotz, Dominic Springer, Marc Reichenbach, Benno Stabernack, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Compressive Online Robust Principal Component Analysis via n-ℓ1 MinimizationabstractThis paper considers online robust principal component analysis (RPCA) in time-varying decomposition problems such as video foreground-background separation. We propose a compressive online RPCA algorithm that decomposes recursively a sequence of data vectors (e.g., frames) into sparse and low-rank components. Different from conventional batch RPCA, which processes all the data directly, our approach considers a small set of measurements taken per data vector (frame). Moreover, our algorithm can incorporate multiple prior information from previous decomposed vectors via proposing an - minimization method. At each time instance, the algorithm recovers the sparse vector by solving the - minimization problem-which promotes not only the sparsity of the vector but also its correlation with multiple previously recovered sparse vectors-and, subsequently, updates the low-rank component using incremental singular value decomposition. We also establish theoretical bounds on the number of measurements required to guarantee successful compressive separation under the assumptions of static or slowly changing low-rank components. We evaluate the proposed algorithm using numerical experiments and online video foreground-background separation experiments. The experimental results show that the proposed method outperforms the existing methods. Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup |
IEEE Trans. Image Process. | 5 |
| 2018 | Temporal Scalability of Dynamic Volume Data Using Mesh Compensated Wavelet LiftingabstractDue to their high resolution, dynamic medical 2D+t and 3D+t volumes from computed tomography (CT) and magnetic resonance tomography (MR) reach a size which makes them very unhandy for teleradiologic applications. A lossless scalable representation offers the advantage of a down-scaled version which can be used for orientation or previewing, while the remaining information for reconstructing the full resolution is transmitted on demand. The wavelet transform offers the desired scalability. A very high quality of the lowpass sub-band is crucial in order to use it as a down-scaled representation. We propose an approach based on compensated wavelet lifting for obtaining a scalable representation of dynamic CT and MR volumes with very high quality. The mesh compensation is feasible to model the displacement in dynamic volumes which is mainly given by expansion and contraction of tissue over time. To achieve this, we propose an optimized estimation of the mesh compensation parameters to optimally fit for dynamic volumes. Within the lifting structure, the inversion of the motion compensation is crucial in the update step. We propose to take this inversion directly into account during the estimation step and can improve the quality of the lowpass sub-band by 0.63 and 0.43 dB on average for our tested dynamic CT and MR volumes at the cost of an increase of the rate by 2.4% and 1.2% on average. Wolfgang Schnurrer, Niklas Pallast, Thomas Richter 0001, André Kaup |
IEEE Trans. Image Process. | 4 |
| 2018 | Automatic Registration of Images With Inconsistent Content Through Line-Support Region Segmentation and Geometrical Outlier RemovalabstractThe implementation of automatic image registration is still difficult in various applications. In this paper, an automatic image registration approach through line-support region segmentation and geometrical outlier removal is proposed. This new approach is designed to address the problems associated with the registration of images with affine deformations and inconsistent content, such as remote sensing images with different spectral content or noise interference, or map images with inconsistent annotations. To begin with, line-support regions, namely a straight region whose points share roughly the same image gradient angle, are extracted to address the issues of inconsistent content existing in images. To alleviate the incompleteness of line segments, an iterative strategy with multi-resolution is employed to preserve global structures that are masked at full resolution by image details or noise. Then, geometrical outlier removal is developed to provide reliable feature point matching, which is based on affine-invariant geometrical classifications for corresponding matches initialized by scale invariant feature transform. The candidate outliers are selected by comparing the disparity of accumulated classifications among all matches, instead of conventional methods which only rely on local geometrical relations. Various image sets have been considered in this paper for the evaluation of the proposed approach, including aerial images with simulated affine deformations, remote sensing optical and synthetic aperture radar images taken at different situations (multispectral, multisensor, and multitemporal), and map images with inconsistent annotations. Experimental results demonstrate the superior performance of the proposed method over the existing approaches for the whole data set. Ming Zhao 0009, Yongpeng Wu 0001, Shengda Pan, Bowen An, André Kaup |
IEEE Trans. Image Process. | 6 |
| 2017 | Motion compensated frame rate up-conversion using 3D frequency selective extrapolation and a multi-layer consistency checkabstractA high temporal resolution is desirable in many applications such as entertainment systems, automotive systems, or video surveillance. Apart from using cameras with a higher temporal resolution, it is also possible to employ frame rate up-conversion methods to obtain an enhanced temporal resolution. In principle, those algorithms can be grouped into approaches that rely on a motion estimation and approaches that do not. Both strategies typically process a video sequence frame by frame and take into account only the directly adjacent frames to compute the intermediate frame. In this paper, we propose a frame rate up-conversion technique that employs a motion compensated three-dimensional reconstruction algorithm. As a result, the proposed method takes into account more than two frames and is capable of jointly reconstructing up to a certain amount of missing frames in a video sequence. Furthermore, we present a multi-layer consistency check to further improve the reconstruction. On average, simulation results show a luminance PSNR gain compared to a conventional frame rate up-conversion method of 0.5 dB. Visual examples substantiate our objective results. Michel Bätz, Fabian Brand, Andrea Eichenseer, André Kaup |
ICASSP | 4 |
| 2017 | Improving mesh-based motion compensation by using edge adaptive graph-based compensated wavelet lifting for medical data setsabstractMedical applications like Computed Tomography (CT) or Magnetic Resonance Tomography (MRT) often require an efficient scalable representation of their huge output volumes in the further processing chain of medical routine. A downscaled version of such a signal can be obtained by using image and video coders based on wavelet transforms. The visual quality of the resulting lowpass band, which shall be used as a representative, can be improved by applying motion compensation methods during the transform. This paper presents a new approach of using the distorted edge lengths of a mesh-based compensated grid instead of the approximated intensity values of the underlying frame to perform a motion compensation. We will show that an edge adaptive graph-based compensation and its usage for compensated wavelet lifting improves the visual quality of the lowpass band by approximately 2.5 dB compared to the traditional mesh-based compensation, while the additional filesize required for coding the motion information doesn't change. Daniela Lanz, André Kaup |
ICASSP | 2 |
| 2017 | Super-resolution for differently exposed mixed-resolution multi-view images adapted by a histogram matching methodabstractSuper-resolution is an important task in the image and video processing domain. In mixed-resolution multi-view scenarios, neighboring high-resolution reference perspectives can be used to increase the image quality of a given low-resolution target view. By using corresponding depth information, the required high-frequency part can be projected from a reference view onto the image plane of the target perspective. However, the contrast and thus the amount of high-frequency information in a reference view varies with the cameras exposure settings. As a consequence, the resulting super-resolution quality drops in case of exposure time variations between the different views. By incorporating a histogram matching method, the required high-frequency part can be efficiently adapted to the exposure settings of the target view. The simulation results show that the proposed adaption leads to an average PSNR gain of 0.63 dB for differently exposed mixed-resolution multi-view setups. Thomas Richter 0001, André Kaup |
ICASSP | 2 |
| 2017 | Demonstration of rapid frequency selective reconstruction for image resolution enhancementabstractWe propose to adapt the local statistics estimation, which has been designed for rapid error concealment originally. By merging the approaches, the reconstruction of images is fastened significantly. We will show in this demonstration that several frames per second (fps) can be processed using conventional computers and HD sized image data. Nils Genser, Jürgen Seiler, Markus Jonscher, André Kaup |
ICIP | 4 |
| 2017 | Scaled fixed-point frequency selective extrapolation for fast image error concealmentabstractImage and video signal processing demands in many areas for error concealment algorithms, whereby the execution time plays an important role. Within this paper, we introduce the scaled fixed-point Frequency Selective Extrapolation for fast image error concealment. This algorithm is based on the complex-valued Frequency Selective Extrapolation, but reduces computational complexity. It is shown that a fixed-point realization of the state-of-the-art algorithm requires an impracticable word-length, which makes it necessary to design a novel, scaled Frequency Selective Extrapolation that reduces the required word-length. Regarding the new approach, execution time can be reduced on average by 22.84 % and at best by up to 40.45 % compared to the state-of-the-art floating point Frequency Selective Extrapolation at the same reconstruction quality. Moreover, platforms can be exploited, which support fixed-point arithmetic only. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2017 | Reliable pedestrian detection using a deep neural network trained on pedestrian countsabstractPedestrian detection is an important task for applications like surveillance, driver assistance systems and autonomous driving. We present a novel approach for detecting pedestrians using a deep convolutional neural network (CNN) trained for counting pedestrians. Our method avoids the need for annotation of the position of the pedestrians in the training data via bounding boxes. The deconvolved outputs of the filters of the trained counting model are used to detect the pedestrians. The average miss rate values on the tested datasets were found to be in the same range as other methods in spite of a simpler training using only pedestrian counts. This method is found to be suitable for detecting pedestrians in crowded scenes with occlusion as well as less crowded scenes. Sanjukta Ghosh, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 4 |
| 2017 | A low-complexity metric for the estimation of perceived chrominance sub-sampling errors in screen content imagesabstractIn web conferencing and remote desktop applications screen sharing is a popular feature. Due to the limited data rate, image and video compression technologies are used to enable this functionality over the internet. However, the commonly used basic profiles of modern video codecs typically utilize the YCbCr color space involving 4:2:0 chrominance sub-sampling. This sub-sampling introduces visually disturbing artifacts for screen content, which may be avoided or reduced if they could be automatically detected. However, these artifacts are shown to be hardly captured by conventional image quality metrics. To address this problem, we developed the Perceived Chrominance Sub-sampling Error (PCSE) metric. Moreover, we created a subjectively evaluated screen content image data set, which is employed for the performance evaluation. The PCSE achieves absolute correlation coefficient values of almost 0.8. Thus, conventional quality metrics are significantly outperformed. Andreas Heindel, Eugen Wige, Felix Fleckenstein, Benjamin Prestele, Alexander Gehlert, André Kaup |
ICIP | 6 |
| 2017 | Video decoding energy estimation using processor eventsabstractIn this paper, we show that processor events like instruction counts or cache misses can be used to accurately estimate the processing energy of software video decoders. Therefore, we perform energy measurements on an ARM-based evaluation platform and count processor level events using a dedicated profiling software. Measurements are performed for various codecs and decoder implementations to prove the general viability of our observations. Using the estimation method proposed in this paper, the true decoding energy for various recent video coding standards including HEVC and VP9 can be estimated with a mean estimation error that is smaller than 6%. Christian Herglotz, André Kaup |
ICIP | 2 |
| 2017 | Disparity estimation for fisheye images with an application to intermediate view synthesisabstractTo obtain depth information from a stereo camera setup, a common way is to conduct disparity estimation between the two views; the disparity map thus generated may then also be used to synthesize arbitrary intermediate views. A straightforward approach to disparity estimation is block matching, which performs well with perspective data. When dealing with non-perspective imagery such as obtained from ultra wide-angle fisheye cameras, however, block matching meets its limits. In this paper, an adapted disparity estimation approach for fisheye images is introduced. The proposed method exploits knowledge about the fisheye projection function to transform the fisheye coordinate grid to a corresponding perspective mesh. Offsets between views can thus be determined more accurately, resulting in more reliable disparity maps. By re-projecting the perspective mesh to the fisheye domain, the original fisheye field of view is retained. The benefit of the proposed method is demonstrated in the context of intermediate view synthesis, for which both objectively evaluated as well as visually convincing results are provided. Andrea Eichenseer, Michel Bätz, André Kaup |
MMSP | 3 |
| 2017 | Iterative denoising-based mesh-to-grid reconstruction with hyperparametric adaptationabstractThis paper presents a new method for the reconstruction of images from samples located at non-integer mesh positions. This is a common scenario for many image processing applications such as multi-image super-resolution, frame-rate up-conversion, or virtual view synthesis in multi-camera systems. The proposed method consists of an iterative procedure that employs adaptive denoising in order to reduce the reconstruction error. The denoising strength is controlled by a novel hyper-parametric adaptation mechanism that aims at maximizing the reconstruction quality in each iteration. Furthermore, the usage of the proposed hyperparametric model drastically reduces the number of necessary parameters that need to be trained or stored. In terms of luminance PSNR, the proposed approach improves the reconstruction quality by more than 1 dB with respect to the initial estimate and outperforms current state-of-the-art denoising-based reconstruction schemes by up to 0.5 dB. Ján Koloda, Michel Bätz, André Kaup |
MMSP | 3 |
| 2017 | Detecting closely spaced and occluded pedestrians using specialized deep models for countingabstractPedestrian detection is an important task in surveillance applications and becomes particularly challenging when pedestrians are close together or occluding one another. This paper presents a novel approach to detect pedestrians in such challenging scenarios. A deep convolutional neural network trained for counting is specialized to count one pedestrian. The feature extractor learned thereby is exploited to detect one pedestrian at a time iteratively. For the base counting model and the specialization, extensive annotation efforts are not required since only a single number at the image level is used. Use of our method on pedestrian datasets with occlusion showed an improvement in the average miss rate values as compared to other methods for handling occlusion. Sanjukta Ghosh, Peter Amon, Andreas Hutter, André Kaup |
VCIP | 4 |
| 2017 | Scalable near-lossless video compression based on HEVCabstractLossless or near-lossless video compression with a certain maximum absolute reconstruction error is required in many professional applications like archiving or medical imaging. However, lossless coding resulting in rather high bit rates and near-lossless coding being traditionally used and especially optimized for still image coding prevents these techniques from practical application. Consequently, this paper proposes a system for scalable lossy to near-lossless video compression, where the inter picture correlations are exploited by the base layer codec and near-lossless enhancement layer compression delivers the required final reconstruction fidelity in terms of maximum absolute reconstruction error. For the proposed near-lossless enhancement layer codec we extended the Sample-based weighted prediction for Enhancement Layer Coding (SELC) scheme, and the HEVC reference software is used for base layer coding. The experimental results show significant average savings compared to both JPEG-LS single layer coding (27%) and also JPEG-LS for enhancement layer coding (about 14 percentage points). Andreas Heindel, André Kaup |
VCIP | 2 |
| 2017 | Low-Complexity Enhancement Layer Compression for Scalable Lossless Video Coding Based on HEVCabstractLossless compression is desired, especially for professional applications like medical imaging. However, scalability may be necessary to transmit a fast preview of pictures when the channel capacity is limited. Furthermore, there may be the need for random access to single frames of a losslessly reconstructed video. We present and evaluate a lossy-to-lossless scalable video coding system in which the scalability is achieved by using a lossy base layer (BL) in conjunction with lossless compression of the reconstruction error in the enhancement layer (EL). High Efficiency Video Coding (HEVC) is employed for encoding the BL. For compression of the EL, we propose a low-complexity scheme called sample-based weighted prediction for EL coding (SELC). Furthermore, EL compression using JPEG-LS and the scalable extension of HEVC (SHVC) are evaluated. The performance of scalable coding using each of these three methods is compared with lossless single-layer (SL) coding using the intra-main RExt configuration. Experimental results show that the scalable coding with SELC achieves average bitrate savings of 7.3% compared with the lossless SL coding, while average bitrate reductions of only 6.3% and 3.9% are obtained using JPEG-LS and SHVC, respectively. Concerning runtime, the additional encoder runtimes of SELC and JPEG-LS compared with the BL coding are negligible. In contrast to that, the runtime of the SHVC EL encoder is much higher and similar to the runtime of the BL encoder. Andreas Heindel, Eugen Wige, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | RFVTM: A Recovery and Filtering Vertex Trichotomy Matching for Remote Sensing Image RegistrationabstractReliable feature point matching is a vital yet challenging process in feature-based image registration. In this paper, a robust feature point matching algorithm, which is called recovery and filtering vertex trichotomy matching, is proposed to remove outliers and retain sufficient inliers for remote sensing images. A novel affine-invariant descriptor, which is called the vertex trichotomy descriptor, is proposed on the basis of that geometrical relations between any of vertices and lines are preserved after affine transformations, which is constructed by mapping each vertex into trichotomy sets. The outlier removals in vertex trichotomy matching (VTM) are implemented by iteratively comparing the disparity of the corresponding vertex trichotomy descriptors. Some inliers mistakenly validated by a large number of outliers are removed in VTM iterations, and several residual outliers that are close to the correct locations cannot be excluded with the same graph structures. Therefore, a recovery and filtering strategy is designed to recover some inliers based on identical vertex trichotomy descriptors and restricted transformation errors. Assisted with the additional recovered inliers, residual outliers can be also filtered out during the process of reaching identical graphs for the expanded vertex sets. Experimental results demonstrate the superior performance on precision and stability of this algorithm under various conditions, such as remote sensing images with large transformations, duplicated patterns, or inconsistent spectral content. Ming Zhao 0009, Bowen An, Yongpeng Wu 0001, Huynh Van Luong, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Frequency-Selective Mesh-to-Grid Resampling for Image CommunicationabstractThis paper presents a novel approach for image reconstruction from pixels located at arbitrary noninteger positions, called mesh. This task forms an intrinsic part of various multimedia applications, including superresolution, fisheye imaging, or generations of new views in multicamera systems, among others. We propose a new frequency-selective mesh-to-grid resampling algorithm that aims at producing high quality reconstructions. It is inspired by the existing frequency-selective reconstruction (FSR) algorithm that is known to exhibit high performance when pixels are located on the regular 2D grid. However, if samples that are located at noninteger positions are involved, a severe overfitting problem arises from the fact that nonorthogonal weighted bases sampled at noninteger positions are used for signal modeling. In order to overcome this issue, we propose a novel stabilizing mechanism that is based on a set of adaptively weighted initial estimates, called key points. We also show that Fourier basis, used in the classic grid-based FSR, yields complex signals when noninteger positions are involved. Since digital images are real valued, we propose to employ a 2D cosine transform basis. Experimental results show the superiority of the proposed approach over a wide range of existing reconstruction techniques. Ján Koloda, Jürgen Seiler, André Kaup |
IEEE Trans. Multim. | 3 |
| 2016 | Multi-mode Kernel-Based Minimum Mean Square Error Estimator for Accelerated Image Error ConcealmentabstractSummary form only given. In this paper, we propose a novel multi-mode error concealment algorithm that aims at obtaining high quality reconstructions with reduced computational burden. Block-based coding schemes in packet loss-environment are considered. The proposed technique exploits the excellent reconstructing abilities of the kernel-based minimum mean square error (K-MMSE) estimator [1]. The complexity of our technique is dynamically adapted to the visual complexity of the area being reconstructed. The technique outperforms other state of the art algorithms and produces high quality reconstructions, equivalent to K-MMSE, while requiring less than one fourth of its computational time. Ján Koloda, Jürgen Seiler, Antonio M. Peinado, André Kaup |
DCC | 4 |
| 2016 | A Reconstruction Algorithm with Multiple Side Information for Distributed Compression of Sparse SourcesabstractWe consider the task of reconstructing target signals which are processed as sparse sources for a distributed compression scenario, where communication between the sources is prohibited, however, correlation of information among sources can be utilized at the decoder. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI) for such distributed sparse sources. The proposed algorithm takes advantage of both a compressive sensing (CS) reconstruction with SI and an iteratively weighted ℓ1-norm minimization by solving a general weighted multi-ℓ1(or n-ℓ1) minimization. To utilize the known multiple SIs, the algorithm computes optimal weights on not only each individual SI but among SIs where the weights are adaptively updated according to changes at every iteration of the reconstruction. By this optimization, the proposed reconstruction algorithm with multiple SI (RAMSI) can robustly exploit the multiple SIs with different qualities. We experimentally demonstrate our algorithm on compressing feature histograms as sparse sources which are extracted from a multi-view image database for multi-view recognition. The results show that the RAMSI with multiple SIs efficiently outperforms the ℓ1minimization and also the CS reconstruction with only one SI. Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer |
DCC | 3 |
| 2016 | A data set providing synthetic and real-world fisheye video sequencesabstractIn video surveillance as well as automotive applications, so-called fisheye cameras are often employed to capture a very wide angle of view. As such cameras depend on projections quite different from the classical perspective projection, the resulting fisheye image and video data correspondingly exhibits non-rectilinear image characteristics. Typical image and video processing algorithms, however, are not designed for these fisheye characteristics. To be able to develop and evaluate algorithms specifically adapted to fisheye images and videos, a corresponding test data set is therefore introduced in this paper. The first of those sequences were generated during the authors' own work on motion estimation for fish-eye videos and further sequences have gradually been added to create a more extensive collection. The data set now comprises synthetically generated fisheye sequences, ranging from simple patterns to more complex scenes, as well as fisheye video sequences captured with an actual fisheye camera. For the synthetic sequences, exact information on the lens employed is available, thus facilitating both verification and evaluation of any adapted algorithms. For the real-world sequences, we provide calibration data as well as the settings used during acquisition. The sequences are freely available via www.lms.lnt.de/fisheyedataset/. Andrea Eichenseer, André Kaup |
ICASSP | 2 |
| 2016 | Multi-image super-resolution using a dual weighting scheme based on Voronoi tessellationabstractIncreasing spatial resolution is often required in many applications such as entertainment systems or video surveillance. Apart from using higher resolution sensors, it is also possible to apply superresolution algorithms to realize an increased resolution. Those methods can be divided into approaches that rely on only a single low resolution image or on multiple low resolution video frames. While incorporating more frames into the super-resolution is beneficial for the resolution enhancement in principle, it is also likely to introduce more artifacts from inaccurate motion estimation. To alleviate this problem, various weightings have been proposed in the literature. In this paper, we propose an extended dual weighting scheme for an interpolation-based super-resolution method based on Voronoi tessellation that relies on both a motion confidence weight and a distance weight. Compared to non-weighted super-resolution, the proposed method yields an average gain in luminance PSNR of up to 1.29 dB and 0.61 dB for upscaling factors of 2 and 4, respectively. Visual comparisons substantiate the objective results. Michel Bätz, Andrea Eichenseer, André Kaup |
ICIP | 3 |
| 2016 | Motion estimation for fisheye video sequences combining perspective projection with camera calibration informationabstractFisheye cameras prove a convenient means in surveillance and automotive applications as they provide a very wide field of view for capturing their surroundings. Contrary to typical rectilinear imagery, however, fisheye video sequences follow a different mapping from the world coordinates to the image plane which is not considered in standard video processing techniques. In this paper, we present a motion estimation method for real-world fisheye videos by combining perspective projection with knowledge about the underlying fisheye projection. The latter is obtained by camera calibration since actual lenses rarely follow exact models. Furthermore, we introduce a re-mapping for ultra-wide angles which would otherwise lead to wrong motion compensation results for the fisheye boundary. Both concepts extend an existing hybrid motion estimation method for equisolid fisheye video sequences that decides between traditional and fisheye block matching in a block-based manner. Compared to that method, the proposed calibration and re-mapping extensions yield gains of up to 0.58 dB in luminance PSNR for real-world fisheye video sequences. Overall gains amount to up to 3.32 dB compared to traditional block matching. Andrea Eichenseer, Michel Bätz, André Kaup |
ICIP | 3 |
| 2016 | Two-stage exclusion of angular intra prediction modes for fast mode decision in HEVCabstractThe latest video compression standard HEVC sets new benchmarks concerning the efficiency for both video coding and also still image coding, i.e., pure intra picture coding. Nevertheless, its high complexity created by the rate-distortion optimization procedure is a serious drawback. To reduce this computational burden, several algorithms for fast mode decision have been proposed. However, most of these methods exclude modes based on analyzing the respective prediction errors. In this paper we present an algorithm for individual angular prediction mode exclusion which intervenes earlier, namely before prediction, almost solely based on the analysis of the reference samples. It is proposed to use this approach as an extension of the global angular intra mode exclusion scheme based on the reference samples from our previous work. The combined algorithm achieves average encoding time savings of approximately 25%, accompanied by only about 1.0% bit rate increase. Andreas Heindel, Christoph Pylinski, André Kaup |
ICIP | 3 |
| 2016 | Joint optimization of rate, distortion, and decoding energy for HEVC intraframe codingabstractThis paper presents a novel algorithm that aims at minimizing the required decoding energy by exploiting a general energy model for HEVC-decoder solutions. We incorporate the energy model into the HEVC encoder such that it is capable of constructing a bit stream whose decoding process consumes less energy than the decoding process of a conventional bit stream. To achieve this, we propose to extend the traditional Rate-Distortion-Optimization scheme to a Decoding-Energy-Rate-Distortion approach. To obtain fast encoding decisions in the optimization process, we derive a fixed relation between the quantization parameter and the Lagrange multiplier for energy optimization. Our experiments show that this concept is applicable for intraframe-coded videos and that for local playback as well as online streaming scenarios, up to 15% of the decoding energy can be saved at the expense of a bitrate increase of approximately the same magnitude. Christian Herglotz, André Kaup |
ICIP | 2 |
| 2016 | Sparse signal reconstruction with multiple side information using adaptive weights for multiview sourcesabstractThis work considers reconstructing a target signal in a context of distributed sparse sources. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI). The proposed algorithm takes advantage of compressive sensing (CS) with SI and adaptive weights by solving a proposed weighted n-ℓ1minimization. The proposed algorithm computes the adaptive weights in two levels, first each individual intra-SI and then inter-SI weights are iteratively updated at every reconstructed iteration. This two-level optimization leads the proposed reconstruction algorithm with multiple SI using adaptive weights (RAMSIA) to robustly exploit the multiple SIs with different qualities. We experimentally perform our algorithm on generated sparse signals and also correlated feature histograms as multiview sparse sources from a multiview image database. The results show that RAMSIA significantly outperforms both classical CS and CS with single SI, and RAMSIA with higher number of SIs gained more than the one with smaller number of SIs. Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer |
ICIP | 3 |
| 2016 | Fast exclusion of angular intra prediction modes in HEVC using reference sample varianceabstractIntra picture coding using HEVC is very efficient and is applied to video as well as single image compression. However, the cost for this high compression efficiency is the complexity caused by the high number of 35 available coding modes. Existing methods for fast mode decision estimate the mode costs based on the prediction error samples. This paper proposes a smart method to exclude the 33 angular prediction modes of HEVC even before prediction, if they are unlikely to deliver a small prediction error. The sensitivity of excluding these modes can be adjusted by a threshold. Completely disregarding the angular modes leads to a very high bitrate increase compared to the unmodified encoder of almost 18% with a time saving of 51%. However, excluding the modes using an exemplary threshold value, the encoding time can be reduced by more than 20%, accompanied by a bitrate increase of only 0.7%. Andreas Heindel, André Kaup |
ISCAS | 2 |
| 2016 | Multi-image super-resolution using a locally adaptive denoising-based refinementabstractSpatial resolution enhancement is of particular interest in many applications such as entertainment, surveillance, or automotive systems. Besides using a more expensive, higher resolution sensor, it is also possible to apply super-resolution techniques on the low resolution content. Super-resolution methods can be basically classified into single-image and multi-image super-resolution. In this paper, we propose the integration of a novel locally adaptive de noising-based refinement step as an intermediate processing step in a multi-image super-resolution framework. The idea is to be capable of removing reconstruction artifacts while preserving the details in areas of interest such as text. Simulation results show an average gain in luminance PSNR of up to 0.2 dB and 0.3 dB for an up scaling of 2 and 4, respectively. The objective results are substantiated by the visual impression. Michel Bätz, Ján Koloda, Andrea Eichenseer, André Kaup |
MMSP | 4 |
| 2016 | Reliability-based mesh-to-grid image reconstructionabstractThis paper presents a novel method for the reconstruction of images from samples located at non-integer positions, called mesh. This is a common scenario for many image processing applications, such as super-resolution, warping or virtual view generation in multi-camera systems. The proposed method relies on a set of initial estimates that are later refined by a new reliability-based content-adaptive framework that employs denoising in order to reduce the reconstruction error. The reliability of the initial estimate is computed so stronger denoising is applied to less reliable estimates. The proposed technique can improve the reconstruction quality by more than 2 dB (in terms of PSNR) with respect to the initial estimate and it outperforms the state-of-the-art denoising-based refinement by up to 0.7 dB. Ján Koloda, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2016 | Adaptive frequency prior for frequency selective reconstruction of images from non-regular subsamplingabstractImage signals typically are defined on a rectangular two-dimensional grid. However, there exist scenarios where this is not fulfilled and where the image information only is available for a non-regular subset of pixel position. For processing, transmitting or displaying such an image signal, a re-sampling to a regular grid is required. Recently, Frequency Selective Reconstruction (FSR) has been proposed as a very effective sparsity-based algorithm for solving this under-determined problem. For this, FSR iteratively generates a model of the signal in the Fourier-domain. In this context, a fixed frequency prior inspired by the optical transfer function is used for favoring low-frequency content. However, this fixed prior is often too strict and may lead to a reduced reconstruction quality. To resolve this weakness, this paper proposes an adaptive frequency prior which takes the local density of the available samples into account. The proposed adaptive prior allows for a very high reconstruction quality, yielding gains of up to 0.6 dB PSNR over the fixed prior, independently of the density of the available samples. Compared to other state-of-the-art algorithms, visually noticeable gains of several dB are possible. Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2016 | Joint shape and centroid adaptive frequency selective extrapolation for the reconstruction of arbitrarily shaped loss areasabstractReconstructing missing areas of arbitrary shape and size is particularly important in error-prone communication as well as in applications where motion compensation is conducted such as multi-image super-resolution or framerate up-conversion. To that end, frequency selective extrapolation is an effective image reconstruction technique. This approach was originally designed for block losses and has recently been enhanced by a centroid adaptation to improve the reconstruction quality in the case of arbitrarily shaped loss areas. In this paper, we reuse the idea of centroid adaptation and introduce a novel shape adaptation so as to better assign higher weights to more relevant pixels. Moreover, we propose to combine both the shape adaptation and the centroid adaptation into a joint solution to further improve the reconstruction quality. To evaluate the proposed method, three different loss patterns are used. Simulation results yield an average gain in luminance PSNR of up to 0.2 dB for the high quality profile and 3.4 dB for the high efficiency profile, respectively. A visual comparison confirms these results. Michel Bätz, Wolfgang Schnurrer, Ján Koloda, Andrea Eichenseer, André Kaup |
PCS | 5 |
| 2016 | Fast CU split decisions for HEVC inter coding using support vector machinesabstractHigh Efficiency Video Coding (HEVC) is known to be the state of the art in video compression. However, aiming at real-time applications, its very high number of possible coding options makes it necessary to approximate the selection of coding modes by fast algorithms. For this purpose Coding Unit (CU) split decisions are made by Support Vector Machine (SVM) classifiers in this paper. We use one SVM model per depth and the traditional Rate-Distortion Optimization (RDO) is used as fallback method in the case of uncertain SVM decisions. Regarding the sequence classes ClassB to ClassE, the proposed algorithm achieves encoding time reductions of more than 60%, accompanied by only 1.8% additional bitrate (low delay main configuration). Considering all sequence classes from ClassA to ClassF, the additional bitrate increases to 4% for the lowdelay main and about 3% for the random access main configuration, with a similar encoding time reduction of more than 60%. Andreas Heindel, Thomas Haubner, André Kaup |
PCS | 3 |
| 2016 | Multi-objective design space exploration for the optimization of the HEVC mode decision processabstractFinding the best possible encoding decisions for compressing a video sequence is a highly complex problem. In this work, we propose a multi-objective Design Space Exploration (DSE) method to automatically find HEVC encoder implementations that are optimized for several different criteria. The DSE shall optimize the coding mode evaluation order of the mode decision process and jointly explore early skip conditions to minimize the four objectives a) bitrate, b) distortion, c) encoding time, and d) decoding energy. In this context, we use a SystemC-based actor model of the HM test model encoder for the evaluation of each explored solution. The evaluation that is based on real measurements shows that our framework can automatically generate encoder solutions that save more than 60% of encoding time or 3% of decoding energy when accepting bitrate increases of around 3%. Christian Herglotz, Rafael Rosales, Michael Glaß, Jürgen Teich, André Kaup |
PCS | 5 |
| 2016 | A bitstream feature based model for video decoding energy estimationabstractIn this paper we show that a small amount of bit stream features can be used to accurately estimate the energy consumption of state-of-the-art software and hardware accelerated decoder implementations for four different video codecs. By testing the estimation performance on HEVC, H.264, H.263, and VP9 we show that the proposed model can be used for any hybrid video codec. We test our approach on a high amount of different test sequences to prove the general validity. We show that less than 20 features are sufficient to obtain mean estimation errors that are smaller than 8%. Finally, an example will show the performance trade-offs in terms of rate, distortion, and decoding energy for all tested codecs. Christian Herglotz, Yongjun Wen, Bowen Dai, Matthias Kränzler, André Kaup |
PCS | 5 |
| 2016 | Recursive frequency selective reconstruction of non-regularly sampled video dataabstractHigh resolution images can be acquired using a non-regular sampling sensor which consists of an underlying low resolution sensor that is covered with a non-regular sampling mask. The reconstructed high resolution image is then obtained during post-processing. Recently, it has been shown that the temporal correlation between neighboring frames can be exploited in order to enhance the reconstruction quality of non-regularly sampled video data. In this paper, a new recursive multi-frame reconstruction approach is proposed in order to further increase the reconstruction quality. By using a new reference order, previously reconstructed frames can be used for the subsequent motion estimation and a new weighting function allows for the incorporation of multiple pixels projected onto the same position. With the new recursive multi-frame approach, a visually noticeable average gain in PSNR of up to 1.13 dB with respect to a state-of-the-art single-frame reconstruction approach can be achieved. Compared to the existing multi-frame approach, a gain of 0.31 dB is possible. SSIM results show the same behavior as PSNR results. Additionally, the pre-reconstruction step of the existing multi-frame approach can be avoided and the new algorithm is, in general, capable of real-time processing. Markus Jonscher, Karina Jaskolka, Jürgen Seiler, André Kaup |
PCS | 4 |
| 2016 | Texture-dependent frequency selective reconstruction of non-regularly sampled imagesabstractThere exist many scenarios where pixel information is available only on a non-regular subset of pixel positions. For further processing, however, it is required to reconstruct such images on a regular grid. Besides many other algorithms, frequency selective reconstruction can be applied for this task. It performs a block-wise generation of a sparse signal model as an iterative superposition of Fourier basis functions and uses this model to replace missing or corrupted pixels in an image. In this paper, it is shown that it is not required to spend the same amount of iterations on both homogeneous and heterogeneous regions. Hence, a new texture-dependent approach for frequency selective reconstruction is introduced that distributes the number of iterations depending on the texture of the regions to be reconstructed. Compared to the original frequency selective reconstruction and depending on the number of iterations, visually noticeable gains in PSNR of up to 1.47 dB can be achieved. Markus Jonscher, Jürgen Seiler, André Kaup |
PCS | 3 |
| 2016 | Graph-based compensated wavelet lifting for 3-D+t medical CT dataabstractAn efficient scalable data representation is an important task especially in the medical area, e.g. for volumes from Computed Tomography (CT) or Magnetic Resonance Tomography (MRT), when a downscaled version of the original signal is needed. Image and video coders based on wavelet transforms provide an adequate way to naturally achieve scalability. This paper presents a new approach for improving the visual quality of the lowpass band by using a novel graph-based method for motion compensation, which is an important step considering data compression. We compare different kinds of neighborhoods for graph construction and demonstrate that a higher amount of referenced nodes increases the quality of the lowpass band while the mean energy of the highpass band decreases. We show that for cardiac CT data the proposed method outperforms a traditional mesh-based approach of motion compensation by approximately 11 dB in terms of PSNR of the lowpass band. Also the mean energy of the highpass band decreases by around 30%. Daniela Lanz, André Kaup |
PCS | 2 |
| 2016 | Distributed coding of multiview sparse sources with joint recoveryabstractIn support of applications involving multiview sources in distributed object recognition using lightweight cameras, we propose a new method for the distributed coding of sparse sources as visual descriptor histograms extracted from multiview images. The problem is challenging due to the computational and energy constraints at each camera as well as the limitations regarding inter-camera communication. Our approach addresses these challenges by exploiting the sparsity of the visual descriptor histograms as well as their intra- and inter-camera correlations. Our method couples distributed source coding of the sparse sources with a new joint recovery algorithm that incorporates multiple side information signals, where prior knowledge (low quality) of all the sparse sources is initially sent to exploit their correlations. Experimental evaluation using the histograms of shift-invariant feature transform (SIFT) descriptors extracted from multiview images shows that our method leads to an average bit-rate saving of 30.7% compared to the state-of-the-art distributed compressed sensing method with independent encoding of the sources. Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup |
PCS | 4 |
| 2016 | Optimized processing order for 3D hole filling in video sequences using frequency selective extrapolationabstractA problem often arising in video communication is the reconstruction of missing or distorted areas in a video sequence. Such holes of unavailable pixels may be caused for example by transmission errors of coded video data or undesired objects like logos. In order to close the holes given neighboring available content, a signal extrapolation has to be performed. The best quality can be achieved, if spatial as well as temporal information is used for the reconstruction. However, the question always is in which order to process the extrapolation to obtain the best result. In this paper, an optimized processing order is introduced for improving the extrapolation quality of Three-dimensional Frequency Selective Extrapolation. Using the proposed optimized order, holes in video sequences can be closed from the outer margin to the center, leading to a higher reconstruction quality, and visually noticeable gains of more than 0.5 dB PSNR are possible. Jürgen Seiler, Susanne Scholl, Wolfgang Schnurrer, André Kaup |
PCS | 4 |
| 2016 | Robust Super-Resolution for Mixed-Resolution Multiview Image Plus Depth DataabstractIncreasing spatial resolution and thus improving the image quality is a key issue in the mixed-resolution multiview image and video processing domain. Given adjacent camera perspectives with various spatial resolutions and their corresponding depth information, the high-frequency part of a high-resolution view can be used for increasing the image quality of a neighboring low-resolution camera perspective. However, a reasonable projection of high-frequency information onto the image plane of a neighboring low-resolution view typically requires pixel-wise error-free depth data for the high-resolution reference image. Starting from this, a novel image super-resolution approach is proposed that is robust against both inaccurate depth acquisition and nonperfect calibration of spatially low-resolution depth sensors. The algorithm is based on displacement-compensated high-frequency synthesis and aims at correcting the projection errors introduced by inaccurate depth information. The proposed approach is further effectively extended by a signal extrapolation technique. For a wide range of proper scenarios, the proposed framework achieves substantial objective and visual gains compared with the considered reference approaches. The improvement of quality is shown for both simulated and self-recorded experimental data. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Probability Distribution Estimation for Autoregressive Pixel-Predictive Image CodingabstractPixelwise linear prediction using backward-adaptive least-squares or weighted least-squares estimation of prediction coefficients is currently among the state-of-the-art methods for lossless image compression. While current research is focused on mean intensity prediction of the pixel to be transmitted, best compression requires occurrence probability estimates for all possible intensity values. Apart from common heuristic approaches, we show how prediction error variance estimates can be derived from the (weighted) least-squares training region and how a complete probability distribution can be built based on an autoregressive image model. The analysis of image stationarity properties further allows deriving a novel formula for weight computation in weighted least-squares proofing and generalizing ad hoc equations from the literature. For sparse intensity distributions in non-natural images, a modified image model is presented. Evaluations were done in the newly developed C++ framework volumetric, artificial, and natural image lossless coder (Vanilc), which can compress a wide range of images, including 16-bit medical 3D volumes or multichannel data. A comparison with several of the best available lossless image codecs proofs that the method can achieve very competitive compression ratios. In terms of reproducible research, the source code of Vanilc has been made public. Andreas Weinlich, Peter Amon, Andreas Hutter, André Kaup |
IEEE Trans. Image Process. | 4 |
| 2015 | Hybrid super-resolution combining example-based single-image and interpolation-based multi-image reconstruction approachesabstractAchieving a higher spatial resolution is of particular interest in many applications such as video surveillance and can be realized by employing higher resolution sensors or applying super-resolution methods. Traditional super-resolution algorithms are based on either a single low resolution image or on multiple low resolution frames. In this paper, a hybrid super-resolution method is proposed which combines both a single-image and a multi-image approach using a soft decision mask. The mask is computed from the motion information utilized in the multi-image super-resolution part. This concept is shown to work for one particular setup but is also extensible toward other combinations of single-image and multi-image super-resolution algorithms as well as other merging metrics. Simulation results show an average luminance PSNR gain of up to 0.85 dB and 0.59 dB for upscaling factors of 2 and 4, respectively. Visual results substantiate the objective results. Michel Bätz, Andrea Eichenseer, Jürgen Seiler, Markus Jonscher, André Kaup |
ICIP | 5 |
| 2015 | A hybrid motion estimation technique for fisheye video sequences based on equisolid re-projectionabstractCapturing large fields of view with only one camera is an important aspect in surveillance and automotive applications, but the wide-angle fisheye imagery thus obtained exhibits very special characteristics that may not be very well suited for typical image and video processing methods such as motion estimation. This paper introduces a motion estimation method that adapts to the typical radial characteristics of fisheye video sequences by making use of an equisolid re-projection after moving part of the motion vector search into the perspective domain via a corresponding back-projection. By combining this approach with conventional translational motion estimation and compensation, average gains in luminance PSNR of up to 1.14 dB are achieved for synthetic fish-eye sequences and up to 0.96 dB for real-world data. Maximum gains for selected frame pairs amount to 2.40 dB and 1.39 dB for synthetic and real-world data, respectively. Andrea Eichenseer, Michel Bätz, Jürgen Seiler, André Kaup |
ICIP | 4 |
| 2015 | Denoising-based image reconstruction from pixels located at non-integer positionsabstractDigital images are commonly represented as regular 2D arrays, so pixels are organized in form of a matrix addressed by integers. However, there are many image processing operations, such as rotation or motion compensation, that produce pixels at non-integer positions. Typically, image reconstruction techniques cannot handle samples at non-integer positions. In this paper, we propose to use triangulation-based reconstruction as initial estimate that is later refined by a novel adaptive denoising framework. Simulations reveal that improvements of up to more than 1.8 dB (in terms of PSNR) are achieved with respect to the initial estimate. Ján Koloda, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2015 | Super-resolution for mixed-resolution multiview image plus depth data using a novel two-stage high-frequency extrapolation method for occluded areasabstractMixed-resolution multiview setups offer great savings in both, costs regarding the equipment acquisition and complexity refering to data transmission and storage. However, for applications like free viewpoint television, high-resolution images are desired from all available camera perspectives. Therefore, high-frequency information of adjacent high-resolution cameras can be used to increase the visual quality of a low-resolution camera perspective. However, due to occlusions, some parts of the scene are invisible in the high-resolution views and cannot be synthesized from the reference images. In this paper, a novel high-frequency extrapolation method is proposed, utilizing additional sampling points pre-estimated by a low-resolution block-matching approach. Averaged over all considered test sets, the proposed method achieves a PSNR gain of 0.94 dB with respect to the unrefined signal extrapolation algorithm. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
ICIP | 4 |
| 2015 | Estimating the HEVC decoding energy using the decoder processing timeabstractThis paper presents a method to accurately estimate the required decoding energy for a given HEVC software decoding solution. We show that the decoder's processing time as returned by common C++ and UNIX functions is a highly suitable parameter to obtain valid estimations for the actual decoding energy. We verify this hypothesis by performing an exhaustive measurement series using different decoder setups and video bit streams. Our findings can be used by developers and researchers in the search for new energy saving video compression algorithms. Christian Herglotz, Elisabeth Walencik, André Kaup |
ISCAS | 3 |
| 2015 | Compressed domain moving object detection by spatio-temporal analysis of H.264/AVC syntax elementsabstractIn this paper, we present a novel moving object detection algorithm for H.264/AVC-compressed video streams. The algorithm does not require full decoding up to the pixel domain but only parsing the compressed bit streams. Thereby, only syntax elements for reconstructing (sub-)macroblock types and quantization parameters are extracted. These features are used to segment the video frames into foreground and background and, according to this segmentation, to identify regions containing moving objects. In a first step, (sub-)macroblock types are analyzed to create initial maps indicating for each block the “weight” for the presence of a moving object. These maps serve as input for our novel spatio-temporal detection algorithm to refine the weight indicating the level of motion for each block. Finally, quantization parameters of macroblocks are used to apply individual thresholds to the block weights to segment the video frames. Experimental results show that our approach efficiently identifies regions containing moving objects and that the presented algorithm is suitable for processing a large number of video streams in parallel. Marcus Laumer, Peter Amon, Andreas Hutter, André Kaup |
PCS | 4 |
| 2015 | Reconstruction of videos taken by a non-regular sampling sensorabstractRecently, it has been shown that a high resolution image can be obtained without the usage of a high resolution sensor. The main idea has been that a low resolution sensor is covered with a non-regular sampling mask followed by a reconstruction of the incomplete high resolution image captured this way. In this paper, a multi-frame reconstruction approach is proposed where a video is taken by a non-regular sampling sensor and fully reconstructed afterwards. By utilizing the temporal correlation between neighboring frames, the reconstruction quality can be further enhanced. Compared to a state-of-the-art single-frame reconstruction approach, this leads to a visually noticeable gain in PSNR of up to 1.19 dB on average. Markus Jonscher, Jürgen Seiler, Michel Bätz, Thomas Richter 0001, Wolfgang Schnurrer, André Kaup |
VCIP | 6 |
| 2015 | Super-resolution for mixed-resolution multiview images using a relative frequency response estimation methodabstractIn mixed-resolution multiview setups, spatially neighboring high-resolution views can be used to increase the image quality of an intermediate low-resolution camera perspective. By utilizing corresponding depth data, high-frequency information can be projected from the reference perspectives onto the image plane of the low-resolution view, creating the desired super-resolved image. In this context, the high-frequency part is extracted by subtracting an estimated low-resolution image from a high-resolution reference view. Starting from this, an approach is proposed that aims at first estimating the frequency responses of the involved cameras. As a consequence, the low-frequency part of a neighboring view can be adapted to the image characteristics of the intermediate low-resolution camera perspective. Exploiting this knowledge, the approach finally leads to a high-frequency part that better fits the target low-resolution view and thus to a better super-resolution quality. For a downsampling factor of 4, the proposed estimation leads to an average gain of 0.53 dB. Thomas Richter 0001, Annelie Habermann, André Kaup |
VCIP | 3 |
| 2015 | Centroid adapted frequency selective extrapolation for reconstruction of lost image areasabstractLost image areas with different size and arbitrary shape can occur in many scenarios such as error-prone communication, depth-based image rendering or motion compensated wavelet lifting. The goal of image reconstruction is to restore these lost image areas as close to the original as possible. Frequency selective extrapolation is a block-based method for efficiently reconstructing lost areas in images. So far, the actual shape of the lost area is not considered directly. We propose a centroid adaption to enhance the existing frequency selective extrapolation algorithm that takes the shape of lost areas into account. To enlarge the test set for evaluation we further propose a method to generate arbitrarily shaped lost areas. On our large test set, we obtain an average reconstruction gain of 1.29 dB. Wolfgang Schnurrer, Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, Michel Bätz, André Kaup |
VCIP | 6 |
| 2015 | Resampling Images to a Regular Grid From a Non-Regular Subset of Pixel Positions Using Frequency Selective ReconstructionabstractEven though image signals are typically defined on a regular 2D grid, there also exist many scenarios where this is not the case and the amplitude of the image signal only is available for a non-regular subset of pixel positions. In such a case, a resampling of the image to a regular grid has to be carried out. This is necessary since almost all algorithms and technologies for processing, transmitting or displaying image signals rely on the samples being available on a regular grid. Thus, it is of great importance to reconstruct the image on this regular grid, so that the reconstruction comes closest to the case that the signal has been originally acquired on the regular grid. In this paper, Frequency Selective Reconstruction is introduced for solving this challenging task. This algorithm reconstructs image signals by exploiting the property that small areas of images can be represented sparsely in the Fourier domain. By further considering the basic properties of the optical transfer function of imaging systems, a sparse model of the signal is iteratively generated. In doing so, the proposed algorithm is able to achieve a very high reconstruction quality, in terms of peak signal-to-noise ratio (PSNR) and structural similarity measure as well as in terms of visual quality. The simulation results show that the proposed algorithm is able to outperform state-of-the-art reconstruction algorithms and gains of more than 1 dB PSNR are possible. Jürgen Seiler, Markus Jonscher, Michael Schöberl, André Kaup |
IEEE Trans. Image Process. | 4 |
| 2014 | Frequency selective extrapolation with residual filtering for image error concealmentabstractThe purpose of signal extrapolation is to estimate unknown signal parts from known samples. This task is especially important for error concealment in image and video communication. For obtaining a high quality reconstruction, assumptions have to be made about the underlying signal in order to solve this underdetermined problem. Among existent reconstruction algorithms, frequency selective extrapolation (FSE) achieves high performance by assuming that image signals can be sparsely represented in the frequency domain. However, FSE does not take into account the low-pass behaviour of natural images. In this paper, we propose a modified FSE that takes this prior knowledge into account for the modelling, yielding significant PSNR gains. Ján Koloda, Jürgen Seiler, André Kaup, Victoria E. Sánchez, Antonio M. Peinado |
ICASSP | 3 |
| 2014 | Reconstruction of multiview images taken with non-regular sampling sensorsabstractIncreasing spatial image resolution is a widely discussed area in the field of image processing. In this paper, we present an efficient reconstruction approach for high-resolution images, taken with irregularly shielded low-resolution sensors in a multiview setup. The approach is based on the sparsity assumption, meaning that natural images can be efficiently represented in a transform-domain using only few coefficients. Utilizing information from adjacent cameras results in a better reconstruction quality for the central high-resolution view. Since neighboring camera perspectives might differ in illumination, the information from adjacent views has to be adapted to the view to be reconstructed. The simulation results show that a proper incorporation of information from neighboring views leads to a PSNR gain of up to 2.20 dB compared to a state-of-the-art singleview reconstruction approach. Thomas Richter 0001, Markus Jonscher, Wolfgang Schnurrer, Jürgen Seiler, André Kaup |
ICASSP | 5 |
| 2014 | Coding of distortion-corrected fisheye video sequences using H.265/HEVCabstractImages and videos captured by fisheye cameras exhibit strong radial distortions due to their large field of view. Conventional intra-frame as well as inter-frame prediction techniques as employed in hybrid video coding schemes are not designed to cope with such distortions, however. So far, captured fish-eye data has been coded and stored without consideration to any loss in efficiency resulting from radial distortion. This paper investigates the effects on the coding efficiency when applying distortion correction as a pre-processing step as opposed to the state-of-the-art method of post-processing. Both methods make use of the latest video coding standard H.265/HEVC and are compared with regard to objective as well as subjective video quality. It is shown that a maximum PSNR gain of 1.91 dB for intra-frame and 1.37 dB for inter-frame coding is achieved when using the pre-processing method. Average gains amount to 1.16 dB and 0.95 dB for intra-frame and inter-frame coding, respectively. Andrea Eichenseer, André Kaup |
ICIP | 2 |
| 2014 | Sample-based Weighted Prediction for lossless enhancement layer coding in SHVCabstractMany professional applications require video content to be compressed in a lossless manner. However, often it is also advantageous to have access to a lossy coded version of the same video content, for example for preview purposes. In this paper the Sample-based Weighted Prediction (SWP) algorithm is applied for two-layer lossy to lossless scalable video compression with SHVC. This means that lossy compression is used for the first layer, the so-called base layer, and lossless compression is used for the second layer, also referred to as enhancement layer. In order to increase the coding efficiency of the enhancement layer the SWP algorithm is used for sample-wise prediction. Depending on the quantization parameter for the base layer, significant bitrate savings considering both layers of up to 8.38% (up to 6.20% on average) compared to the unmodified SHVC reference software can be achieved. Andreas Heindel, Eugen Wige, André Kaup |
ICIP | 3 |
| 2014 | Modeling the energy consumption of HEVC P- and B-frame decodingabstractVideo decoding on portable devices such as smartphones or tablet PCs requires a considerable amount of battery power, shortening their operating time significantly. Hence, tools aiming at minimizing the decoding energy are of special interest. To this end, this paper presents a novel model capable of estimating the energy consumed when decoding an HEVC-coded video. This information can be used to optimize implementations and improve the encoding procedure. An accurate and a more convenient simple model is proposed that, for the evaluated set of test videos, achieved an average relative estimation error of 2.34% and 3.63%, respectively. Christian Herglotz, Dominic Springer, André Kaup |
ICIP | 3 |
| 2014 | Reconstruction of images taken by a pair of non-regular sampling sensors using correlation based matchingabstractMulti-view image acquisition systems with two or more cameras can be rather costly due to the number of high resolution image sensors that are required. Recently, it has been shown that by covering a low resolution sensor with a non-regular sampling mask and by using an efficient algorithm for image reconstruction, a high resolution image can be obtained. In this paper, a stereo image reconstruction setup for multi-view scenarios is proposed. A scene is captured by a pair of non-regular sampling sensors and by incorporating information from the adjacent view, the reconstruction quality can be increased. Compared to a state-of-the-art single-view reconstruction algorithm, this leads to a visually noticeable average gain in PSNR of 0.74 dB. Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, Michel Bätz, André Kaup |
ICIP | 5 |
| 2014 | An error-based recursive filling ordering for image error concealmentabstractMany image error concealment (EC) algorithms can be employed to reconstruct a lost region by dividing it into a set of (sub)blocks which are estimated from a set of known context pixels. In turn, these former pixels may have been previously obtained by estimation. In this situation, the order in which the lost region is recursively filled will clearly condition the resulting reconstruction. This paper proposes a novel filling order aimed to improve the performance of EC algorithms that are applied recursively over the lost area. It takes into account the reconstruction quality of the already concealed blocks in order to determine the filling order. Blocks surrounded by known pixels or by high quality reconstructions are prioritized. The proposed technique is applicable to a wide range of EC algorithms and achieves an improvement of up to 1dB (in PSNR) with negligible additional computational load. Ján Koloda, Jürgen Seiler, André Kaup, Victoria E. Sánchez, Antonio M. Peinado |
ICIP | 3 |
| 2014 | 3-D mesh compensated wavelet lifting for 3-D+t medical CT dataabstractFor scalable coding, a high quality of the lowpass band of a wavelet transform is crucial when it is used as a downscaled version of the original signal. However, blur and motion can lead to disturbing artifacts. By incorporating feasible compensation methods directly into the wavelet transform, the quality of the lowpass band can be improved. The displacement in dynamic medical 3-D+t volumes from Computed Tomography is mainly given by expansion and compression of tissue over time and can be modeled well by mesh-based methods. We extend a 2-D mesh-based compensation method to three dimensions to obtain a volume compensation method that can additionally compensate deforming displacements in the third dimension. We show that a 3-D mesh can obtain a higher quality of the lowpass band by 0.28 dB with less than 40% of the model parameters of a comparable 2-D mesh. Results from lossless coding with JPEG 2000 3D and SPECK3D show that the compensated subbands using a 3-D mesh need about 6% less data compared to using a 2-D mesh. Wolfgang Schnurrer, Thomas Richter 0001, Jürgen Seiler, Christian Herglotz, André Kaup |
ICIP | 5 |
| 2014 | Open source HEVC analyzer for rapid prototyping (HARP)abstractThe design of new HEVC extensions comes with the need for careful analysis of internal HEVC codec decisions. Several bitstream analyzers have evolved for this purpose and provide a visualization of encoder decisions as seen from a decoder viewpoint. None of the existing solutions is able to provide actual insight into the encoder and its RDO decision process. With one exception, all solutions are closed source and make adaption of their code to specific implementation needs impossible. Overall, development with the HM code base remains a time-consuming task. Here, we present the HEVC Analyzer for Rapid Prototyping (HARP), which directly addresses the above issues and is freely available under www.lms.lnt.de/HARP. Dominic Springer, Wolfgang Schnurrer, Andreas Weinlich, Andreas Heindel, Jürgen Seiler, André Kaup |
ICIP | 6 |
| 2014 | Accelerated hybrid image reconstruction for non-regular sampling color sensorsabstractIncreasing the spatial resolution is an ongoing research topic in image processing. A recently presented approach applies a non-regular sampling mask on a low resolution sensor and subsequently reconstructs the masked area via an extrapolation algorithm to obtain a high resolution image. This paper introduces an acceleration of this approach for use with full color sensors. Instead of employing the effective, yet computationally expensive extrapolation algorithm on each of the three RGB channels, a color space conversion is performed and only the luminance channel is then reconstructed using this algorithm. As natural images contain much less information in the chrominance channels, a fast linear interpolation technique can here be used to accelerate the whole reconstruction procedure. Simulation results show that an average speed up factor of 2.9 is thus achieved, while the loss in visual quality stays imperceptible. Comparisons of PSNR results confirm this. Michel Bätz, Andrea Eichenseer, Markus Jonscher, Jürgen Seiler, André Kaup |
VCIP | 5 |
| 2014 | Reducing randomness of non-regular sampling masks for image reconstructionabstractIncreasing spatial image resolution is an often required, yet challenging task in image acquisition. Recently, it has been shown that it is possible to obtain a high resolution image by covering a low resolution sensor with a non-regular sampling mask. Due to the masking, however, some pixel information in the resulting high resolution image is not available and has to be reconstructed by an efficient image reconstruction algorithm in order to get a fully reconstructed high resolution image. In this paper, the influence of different sampling masks with a reduced randomness of the non-regularity on the image reconstruction process is evaluated. Simulation results show that it is sufficient to use sampling masks that are non-regular only on a smaller scale. These sampling masks lead to a visually noticeable gain in PSNR compared to arbitrary chosen sampling masks which are non-regular over the whole image sensor size. At the same time, they simplify the manufacturing process and allow for efficient storage. Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, André Kaup |
VCIP | 4 |
| 2014 | Efficient lossless coding of highpass bands from block-based motion compensated wavelet lifting using JPEG 2000abstractLossless image coding is a crucial task especially in the medical area, e.g., for volumes from Computed Tomography or Magnetic Resonance Tomography. Besides lossless coding, compensated wavelet lifting offers a scalable representation of such huge volumes. While compensation methods increase the details in the lowpass band, they also vary the characteristics of the wavelet coefficients, so an adaption of the coefficient coder should be considered. We propose a simple invertible extension for JPEG 2000 that can reduce the filesize for lossless coding of the highpass band by 0.8% on average with peak rate saving of 1.1%. Wolfgang Schnurrer, Tobias Tröger, Thomas Richter 0001, Jürgen Seiler, André Kaup |
VCIP | 5 |
| 2014 | High dynamic range video reconstruction from a stereo camera setup
Michel Bätz, Thomas Richter 0001, Jens-Uwe Garbas, Anton Papst, Jürgen Seiler, André Kaup |
Signal Process. Image Commun. | 6 |
| 2014 | In-Loop Noise-Filtered Prediction for High Efficiency Video CodingabstractIn this paper, we focus on optimization and enhancement of High Efficiency Video Coding (HEVC) with respect to professional applications. In most of the professional video applications, noise dominates the compression performance. We therefore theoretically and practically analyze the denoising performance of an HEVC codec and show that, especially for low to medium quantization parameters (QPs), the source noise is still present in the reconstructed frame. This motivates us to use an explicit in-loop reference frame denoising filter for this quality range. We show how the noise, which badly influences the prediction, can be modeled and estimated. Two approaches are introduced, in which either the noise is calculated from prior knowledge or estimated automatically without prior knowledge. A denoising filter that exploits the HEVC core transform is introduced. Finally, a coding mode adaptive scheme for motion compensation using the noise-filtered reference frame is described. The compression results show that especially for low to medium QP settings as well as lossless compression considerable bitrate savings can be achieved using the proposed schemes. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Novel Similarity-Invariant Line Descriptor and Matching Algorithm for Global Motion EstimationabstractWe present a new and fast line descriptor and matching algorithm to geometrically register video frames that are deformed by a similarity transformation and contain line segments but with very low texture details, such as navigation maps. Line segments are extracted from the frames, and each line is described with a novel line descriptor that does not depend on pixel intensities. The described lines are then matched together and the matches are input to an outlier removal algorithm in order to estimate the parameters of the transformation describing the global motion between the frames. We propose a method for fast parameter estimation of the transformation using line segments instead of points. Additionally, we apply this algorithm in the testing and error detection context of navigation systems, in which we show how to detect map jump artifacts that could occur during the development of these systems, leading to a jerky and unsmooth motion between the frames. The proposed descriptor and its matching algorithm are shown to be fast enough for online use and very robust against a wide range of translation, rotation, and scale changes. Furthermore, the error detection algorithm allows to detect almost all map jump artifacts while maintaining a very low number of false alarms. Gilbert Yammine, Eugen Wige, Franz Simmet, Dieter Niederkorn, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Multiple Description Coding With Randomly and Uniformly Offset QuantizersabstractIn this paper, two multiple description coding schemes are developed, based on prediction-induced randomly offset quantizers and unequal-deadzone-induced near-uniformly offset quantizers, respectively. In both schemes, each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. In the first method, due to predictive coding, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped. The optimal reconstruction is obtained by finding the intersection of all received bins. In the second method, joint dequantization is also used, but near-uniform offsets are created among different low-rate quantizers by quantizing the predictions and by employing unequal deadzones. By generalizing the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained for the first method, and a lower bound is obtained for the second method. The schemes are then applied to lapped transform-based multiple description image coding. The closed-form expressions enable the optimization of the lapped transform. An iterative algorithm is also developed to facilitate the optimization. Theoretical analyzes and image coding results show that both schemes achieve better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
IEEE Trans. Image Process. | 6 |
| 2013 | M-channel multiple description coding based on uniformly offset quantizers with optimal deadzoneabstractThis paper proposes an improved source-splitting-based two-rate M-channel multiple description coding scheme, where the source is split into M subsets. In each description, one subset is coded at a high rate, and others are predictively coded at a low rate. Uniform offsets among low-rate quantizers of different descriptions are achieved by employing unequal deadzones and by quantizing the predictions. When several descriptions are received, the optimal reconstruction of each subset is achieved by finding the intersection of all received quantization bins. The closed-form expression of the expected distortion is obtained. The proposed scheme is applied to lapped transform-based multiple description image coding and achieves improved performance. The optimal deadzone selection and its impact are also given in this paper. Lili Meng, Jie Liang 0001, Yao Zhao 0001, Huihui Bai 0001, Chunyu Lin, André Kaup |
ICASSP | 6 |
| 2013 | Robust Rotational Motion Estimation for efficient HEVC compression of 2D and 3D navigation video sequencesabstractIn the context of test automation for automobiles, the compressed video recording of infotainment system components like navigation devices is a required practice. These recordings are then analyzed, archived, and forwarded to the responsible engineering teams. In order to compress navigation video sequences efficiently, the dominant rotational motion must be compensated properly. However, the process of Rotational Motion Estimation (RME) is hindered by the presence of static areas like info boxes and overlay graphics. We analyze this problem and show how to build masks for static areas in order to allow high speed feature transforms to be applied. With the acquired fast and accurate RME, we then demonstrate how to significantly reduce the required bitrate during HEVC encoding of navigation sequences. Dominic Springer, Franz Simmet, Dieter Niederkorn, André Kaup |
ICASSP | 4 |
| 2013 | Comparative image quality assessment using free energy minimizationabstractIt is a straightforward task for human observers to judge the relative quality of two visual signals of the same content, but subject to different type/level of distortions. However, this comparative image quality assessment (C-IQA) problem remains a difficult challenge for the current research of image quality assessment (IQA). In this paper, we propose a C-IQA approach to predict the relative perceptual quality of a pair of images that are possibly subject to different artifact types/levels. The C-IQA algorithm is designed to emulate the process of comparing the relative quality of two visual stimuli as performed by the human visual system (HVS) within the framework of free energy minimization. The brain's internal generative models initialized on the inputs are used to explain the two images. And their relative quality can then be determined through comparing the free energy level of this model-data fitting process. In the existing work of IQA, the full-reference (FR) and reduced-reference (RR) methods need the prior knowledge of the original images while the no-reference (NR) algorithms usually work with a single input image. The C-IQA approach is inherently different from those existing methods in that it takes as input an image pair and predicts their relative quality without using any knowledge about the original image. A computationally efficient solution to the proposed C-IQA scheme based on a linear autoregressive image model is also introduced. Experimental results show that the proposed method achieves about 98% accuracy in line with the subjective ratings when applied on over 300,000 image pairs sampled from the LIVE database, outperforming the FR metrics such as PSNR, SSIM, and some of the most advanced NR IQA algorithms. Guangtao Zhai, André Kaup |
ICASSP | 2 |
| 2013 | Spatio-temporal error concealment in video by denoised temporal extrapolation refinementabstractIn video communication, the concealment of distortions caused by transmission errors is important for allowing for a pleasant visual quality and for reducing error propagation. In this article, Denoised Temporal Extrapolation Refinement is introduced as a novel spatiotemporal error concealment algorithm. The algorithm operates in two steps. First, temporal error concealment is used for obtaining an initial estimate. Afterwards, a spatial denoising algorithm is used for reducing the imperfectness of the temporal extrapolation. For this, Non-Local Means denoising is used which is extended by a spiral scan processing order and is improved by an adaptation step for taking the preliminary temporal extrapolation into account. In doing so, a spatio-temporal error concealment results. By making use of the refinement, a visually noticeable average gain of 1 dB over pure temporal error concealment is possible. With this, the algorithm also is able to clearly outperform other spatio-temporal error concealment algorithms. Jürgen Seiler, Michael Schöberl, André Kaup |
ICIP | 3 |
| 2013 | Motion vector analysis based homography estimation for efficient HEVC compression of 2D and 3D navigation video sequencesabstractNavigation systems have become complex devices in automobiles nowadays. As part of large in-car infotainment systems, these devices undergo extensive hardware and software tests in order to assert correct system behavior under all circumstances. During field tests, the display output is typically recorded in compressed form for days or weeks, followed by a thorough analysis of the video data. In this paper, we demonstrate how to setup an HEVC-based compression solution specifically designed for navigation sequence content. We show how rotational motion, which is a dominant characteristic, can be estimated and compensated in an efficient way. We avoid any complex feature-based approaches for global motion estimation but find precise motion parameters by analyzing and filtering motion vector sets produced by HEVC during encoding. While achieving compression efficiency similar to a feature-based approach, processing time for global motion estimation can be significantly reduced. Dominic Springer, Franz Simmet, Dieter Niederkorn, André Kaup |
ICIP | 4 |
| 2013 | Massively parallel lossless compression of medical images using least-squares prediction and arithmetic codingabstractMedical imaging in hospitals requires fast and efficient image compression to support the clinical work flow and to save costs. Least-squares autoregressive pixel prediction methods combined with arithmetic coding constitutes the state of the art in lossless image compression. However, a high computational complexity of both prevents the application of respective CPU implementations in practice. We present a massively parallel compression system for medical volume images which runs on graphics cards. Image blocks are processed independently by separate processing threads. After pixel prediction with specialized border treatment, prediction errors are entropy coded with an adaptive binary arithmetic coder. Both steps are designed to match particular demands of the parallel hardware architecture. Comparisons with current image and video coders show efficiency gains of 3.3-13.6% while compression times can be reduced to a few seconds. Andreas Weinlich, Johannes Rehm, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 5 |
| 2013 | Pixel-based averaging predictor for HEVC lossless codingabstractThis paper presents an intra-frame prediction scheme designed for lossless coding using HEVC. The proposed coding method comprises a pixel-wise prediction based on original samples. It is realized as a separate intra prediction mode, which replaces the PLANAR mode. In order to perform the prediction, a four-sample template around the pixel that is to be predicted is compared to the respective template of a four-pixel neighborhood. For each reference template, the sum of absolute differences (SAD) is determined. A table look-up of the SAD value gives the respective weighting factor for each neighborhood pixel. The predictor for the current pixel is calculated as the weighted average of the neighborhood pixels. In comparison to the unmodified HEVC Test Model HM-9.1 configured for lossless coding by disabling/bypassing transformation, quantization, and in-loop filters, the proposed method provides average bitrate savings up to 10.88% for intra-only coding at similar computational complexity. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 5 |
| 2013 | A novel similarity-invariant line descriptor for geometric map registrationabstractIn this paper we present a new line descriptor that can be used in the process of tracking 2D navigation maps that are low in texture details. The problem of tracking this kind of maps is very important in the development and testing phases of a navigation system. In order to track the motion of the map, which could be a translation, rotation, or scale (also known as similarity transform), line segments are used as features and are matched between two consecutive frames using the proposed descriptor in order to determine the motion between the frames. The descriptor is shown to be very robust against typical translation, rotation, and scale transformations, while requiring an acceptable processing time. Gilbert Yammine, Eugen Wige, Franz Simmet, Dieter Niederkorn, André Kaup |
ICIP | 5 |
| 2013 | Multiple description coding with randomly offset quantizersabstractA multiple description coding scheme based on prediction-induced randomly offset quantizers is proposed, where each description encodes one source subset with a small quantization stepsize, and other subsets are predictively coded with a large quantization stepsize. Due to the prediction, the quantization bins that a coefficient belongs to in different descriptions are randomly overlapped with each others. The optimal reconstruction is obtained by finding the intersection of all received quantization bins. Using the recently developed random quantization theory, the closed-form expression of the expected distortion is obtained. The proposed scheme is then applied to lapped transform-based multiple-description image coding, and an iterative optimization scheme is developed to find the optimal lapped transform. Experimental results show that the proposed scheme achieves better performance than other methods in this category. Lili Meng, Jie Liang 0001, Upul Samarawickrama, Yao Zhao 0001, Huihui Bai 0001, André Kaup |
ISCAS | 6 |
| 2013 | Spiral search based fast rotation estimation for efficient HEVC compression of navigation video sequencesabstractDuring the development and testing of navigation systems for modern cars, the video streams of the systems-under-test are carefully observed for errors of any kind. While this observation is typically carried out manually by software engineers monitoring the video output, the complexity and variety of the latest navigation systems demand an automated analysis of system functionality. Recently published work on automated error detection for navigation systems [1] can only enfold its full functionality if video sequences can be recorded and compressed in real-time, so that the full test program can be run under lab conditions after recording. However, the rotational motion of these sequences makes efficient compression a difficult task. In this paper, we present a fast and efficient method for rotation estimation, so that rotational motion compensation and thus efficient encoding can take place. Compared to existing state of the art approaches for SIFT- or SURF-based global motion estimation, our scheme requires 1/16th of the original processing time while providing almost identical quality gains of up to 2.5dB. Dominic Springer, Martin Frank 0001, Franz Simmet, Dieter Niederkorn, André Kaup |
PCS | 5 |
| 2013 | Volumetric deformation compensation in CUDA for coding of dynamic cardiac imagesabstractA new approach for volumetric deformation compensation in temporally predictive coding of dynamic medical heart images is presented. Instead of using conventional vectors, motion is represented by deformation values to model 3-D muscle contractions. In this way, estimated motion is more homogeneous among the image domain and predictions do not contain disturbing blocking artifacts making the approach suitable also for non-block-based transform coding. It is shown that with equal numbers of motion values the method can achieve prediction accuracies similar to cube-based motion estimation. The run time of the described parallel implementation in Nvidia CUDA is shown to be shorter than for an equivalent implementation of cube-based motion estimation. Andreas Weinlich, Michel Bätz, Peter Amon, Andreas Hutter, André Kaup |
PCS | 5 |
| 2013 | Sample-based Weighted Prediction with Directional Template Matching for HEVC lossless codingabstractThe recently introduced High Efficiency Video Coding (HEVC) standard is currently further investigated for potential use in professional applications. The considered Range Extensions should on the one hand introduce higher bit depths and additional color formats, and on the other hand the coding efficiency of HEVC for high fidelity compression as well as lossless compression is to be improved. In this paper we investigate and improve the recently introduced Sample-based Weighted Prediction (SWP) for HEVC lossless coding. Although being very efficient for natural video content, the SWP algorithm can be further improved for screen content by using a directional template predictor in cases where the SWP algorithm yields worse prediction. The mainly introduced predictor improves the lossless coding results by up to 9.9% compared to the unmodified HEVC reference software for lossless compression. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
PCS | 5 |
| 2013 | A dual-model approach to blind quality assessment of noisy imagesabstractPhysiological and psychological evidences exist that the human visual system (HVS) has different behavioral patterns under low and high noise/artifact levels. We propose in this paper a dual-model approach to blind or no-reference (NR) image quality assessment (IQA) of noisy images through differentiating near-threshold and suprathreshold noise conditions. The underlying assumption for the proposed dual-model method is that for images with low level near-threshold noise, HVS tries to gauge the strength of the noise, so the image quality can be well approximated via measuring strength of the noise. And for images with their contents overwhelmed by high level suprathreshold noise, the HVS tries to recover meaningful structure from the noisy pixels using past experiences and prior knowledge encoded into an internal generative model of the brain. So image quality is closely related to the agreement between the noisy observation and the internal generative model explainable part of the image. More specifically, under the near-threshold noise condition, a noise level estimation algorithm based on natural image statistics is used, while under suprathreshold condition, an active inference model based on the free energy principle is adopted. The near-and suprathreshold models can be seamlessly integrated through a transformation between both estimates. The proposed dual-model algorithm has been tested on additive Gaussian noise contaminated images. Experimental results and comparative studies suggest that although being a no-reference approach, the proposed algorithm has prediction accuracy comparable to some of the best full-reference (FR) IQA methods. Guangtao Zhai, André Kaup, Jia Wang 0004, Xiaokang Yang 0001 |
PCS | 2 |
| 2013 | Retina model inspired image quality assessmentabstractWe proposed in this paper a retina model based approach for image quality assessment. The retinal model is consisted of an optical modulation transfer module and an adaptive low-pass filtering module. We treat the model as a black box and design the adaptive filter using an information theoretical approach. Since the information rate of visual signals is far beyond the processing power of the human visual system, there must be an effective data reduction stage in human visual brain. Therefore, the underlying assumption for the retina model is that the retina reduces the data amount of the visual scene while retaining as much useful information as possible. For full reference image quality assessment, the original and distorted images pass through the retinal filter before some kind of distance is calculated between the images. Retina filtering can serve as a general preprocessing stage for most existing image quality metrics. We show in this paper that retina model based MSE/PSNR, though being straightforward, has already state of the art performance on several image quality databases. Guangtao Zhai, André Kaup, Jia Wang 0004, Xiaokang Yang 0001 |
VCIP | 2 |
| 2012 | High dynamic range video by spatially non-regular optical filteringabstractWe present a new method for capturing high dynamic range video (HDRV). Our method is based on spatially varying exposures, where individual pixels are covered with filters for different optical attenuation. For preventing the loss in resolution we use a new non-regular arrangement of the attenuation pattern. Subsequent image reconstruction based on the sparsity assumption allows the reconstruction of natural images with high detail. Michael Schöberl, Alexander Belz, Jürgen Seiler, Siegfried Fößel, André Kaup |
ICIP | 5 |
| 2012 | Robust super-resolution in a multiview setup based on refined high-frequency synthesisabstractIncreasing image sharpness and thus improving the visual quality is an important task in multiview image and video processing. We propose a novel super-resolution approach for multiview images in a mixed-resolution setup that is robust to various depth map distortions. The considered distortion scenarios may be caused by an inaccurate calibration of the depth camera or a limitation of depth range. Our method is based on a refined high-frequency synthesis that relies on a blockwise and depth-dependant low-frequency registration. The refinement step efficiently adapts the high-frequency content from a neighboring high-resolution camera to a low-resolution view and thereby compensates the displacement caused by depth inaccuracies. In case of undistorted depth maps, the results show that our algorithm leads to a PSNR gain of up to 1.33 dB with respect to a comparable unrefined super-resolution approach for a mixed-resolution multiview video plus depth format. Compared to the initial low-resolution view, a PSNR gain of up to 2.61 dB is obtained. In case of distorted depth maps, a PSNR gain of even 4.78 dB is achieved with respect to the reference superresolution algorithm. The PSNR gains get confirmed by the corresponding SSIM values which manifest a similar behaviour. The improvement of visual quality is also convincingly for all considered scenarios. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
MMSP | 4 |
| 2012 | Analysis of mesh-based motion compensation in wavelet lifting of dynamical 3-D+t CT dataabstractFactorized in the lifting structure, the wavelet transform can easily be extended by arbitrary compensation methods. Thereby, the transform can be adapted to displacements in the signal without losing the ability of perfect reconstruction. This leads to an improvement of scalability. In temporal direction of dynamic medical 3-D+t volumes from Computed Tomography, displacement is mainly given by expansion and compression of tissue. We show that these smooth movements can be well compensated with a mesh-based method. We compare the properties of triangle and quadrilateral meshes. We also show that with a mesh-based compensation approach coding results are comparable to the common slice wise coding with JPEG 2000 while a scalable representation in temporal direction can be achieved. Wolfgang Schnurrer, Thomas Richter 0001, Jürgen Seiler, André Kaup |
MMSP | 4 |
| 2012 | Freeze detection in 2D navigation video sequences overlaid with real satellite imagesabstractThis paper focuses on the detection of freezing artifacts that could occur in the navigation system of the car during field tests. Knowing that the motion in the navigation system could also stop when the car stops, it is crucial to differentiate between this situation and a real freezing artifact. We propose an algorithm that reliably detects map jumps which occur after the freezing artifacts. The proposed algorithm extracts invariant features from the possibly frozen frame and the frame following the freeze event and matches the features in order to output a geometric transformation that describes the motion between both frames. If the motion is larger than normal, a map jump is detected and a freeze event is signaled. Gilbert Yammine, Ali Khairat, Franz Simmet, Dieter Niederkorn, André Kaup |
MMSP | 5 |
| 2012 | On the influence of clipping in lossless predictive and wavelet coding of noisy imagesabstractEspecially in lossless image coding the obtainable compression ratio strongly depends on the amount of noise included in the data as all noise has to be coded, too. Different approaches exist for lossless image coding. We analyze the compression performance of three kinds of approaches, namely direct entropy, predictive and wavelet-based coding. The results from our theoretical model are compared to simulated results from standard algorithms that base on the three approaches. As long as no clipping occurs with increasing noise more bits are needed for lossless compression. We will show that for very noisy signals it is more advantageous to directly use an entropy coder without advanced preprocessing steps. Wolfgang Schnurrer, Jürgen Seiler, Michael Schöberl, André Kaup |
PCS | 4 |
| 2012 | Compression of 2D and 3D navigation video sequences using skip mode masking of static areasabstractIn order to assert correct behavior of electronics in modern automobiles, extensive tests are conducted. Part of these tests focus on the correct rendering of the navigation systems, including map rotation, content of info boxes and smoothness of frame updates. Test engineers have the need to record the rendered navigation system in live setups (e.g. in moving cars) and evaluate them afterwards. Traditional video encoding produces significant bitrate overhead due to the rotating characteristics of the navigation since rotation is approximated with small macroblock partitioning. In this paper, we show how to construct an encoding scheme specifically designed for encoding of 2D and 3D navigation video sequences. For this purpose we develop a Global Motion Estimation (GME) based on feature matching and model parameter estimation and combine it with an H.264/AVC video encoder as backend. By using skip mode information from the H.264/AVC rate distortion optimization, we are able to stabilize the parameter estimation process even in the presence of large static areas. Dominic Springer, Franz Simmet, Dieter Niederkorn, André Kaup |
PCS | 4 |
| 2012 | Blind frame freeze detection in coded videosabstractIn multimedia systems, system errors and artifacts should be avoided in order to keep the Quality of Experience (QoE) of the user as high as possible. For that, the system should be tested and monitored for a long time to assure normal operation. In this paper, we present a new algorithm that automatically detects freezing artifacts in coded videos. The system output is captured and encoded, and the error/artifact detection is run later on the decoded video. One major constraint on our detection algorithm is that every freeze should be detected in order to analyze the reason of its occurrence. Such a constraint entails a high number of false alarms when using the standard freeze detection algorithms. For that, we present a new algorithm that keeps the number of true positives to its maximum, while minimizing the number of false positives. Gilbert Yammine, Eugen Wige, Franz Simmet, Dieter Niederkorn, André Kaup |
PCS | 5 |
| 2012 | Multiview super-resolution using high-frequency synthesis in case of low-framerate depth informationabstractIncreasing the image sharpness of low-resolution views is a key issue in the multiview image and video processing domain. Thereby, a low-resolution view gets refined by high-frequency content that can either be obtained from temporally or spatially adjacent highresolution reference images. We propose a refined super-resolution algorithm for multiview images that is robust to the usage of temporally highly misaligned depth maps. The temporal misalignment may be caused either by a temporal subsampling or a lower framerate of the depth camera with respect to the image cameras. Our refinement step is based on a blockwise low-frequency registration in order to efficiently adapt the high-frequency content of the highresolution reference to the low-resolution destination view. The simulation results show that our proposed algorithm leads to a peak PSNR gain of up to 1.21 dB with respect to a comparable unrefined super-resolution approach. On average, our approach outperforms the unrefined super-resolution algorithm by 0.61 dB. For all considered scenarios the improvement of visual quality is also convincingly. Thomas Richter 0001, André Kaup |
VCIP | 2 |
| 2012 | Analysis of displacement compensation methods for wavelet lifting of medical 3-D thorax CT volume dataabstractA huge advantage of the wavelet transform in image and video compression is its scalability. Wavelet-based coding of medical computed tomography (CT) data becomes more and more popular. While much effort has been spent on encoding of the wavelet coefficients, the extension of the transform by a compensation method as in video coding has not gained much attention so far. We will analyze two compensation methods for medical CT data and compare the characteristics of the displacement compensated wavelet transform with video data. We will show that for thorax CT data the transform coding gain can be improved by a factor of 2 and the quality of the lowpass band can be improved by 8 dB in terms of PSNR compared to the original transform without compensation. Wolfgang Schnurrer, Jürgen Seiler, Eugen Wige, André Kaup |
VCIP | 4 |
| 2012 | Edge modeling prediction for computed tomography imagesabstractPredictive coding is applied in many state-of-the-art lossless image compression algorithms like JPEG-LS, CALIC, or least-squares-based methods. We propose a new approach for accurate intensity prediction in pixel-predictive coding of computed tomography (CT) images. Exploiting their particular edge characteristic, the method only relies on a small twelve-pixel context. It does neither require adaptation to larger-region image characteristics nor the transmission of side-information and therefore may be particularly suitable for compression of small images like in region-of-interest coding. While applying simple linear prediction with fixed weights in homogeneous regions, a Gauss error model-function is fit to given contexts in edge regions and then sampled at the position corresponding to the pixel to be predicted in order to obtain prediction values. By the example of CALIC, it is shown that for CT data the edge modeling prediction (EMP) approach can yield an even smaller prediction error than other methods relying on context modeling. Andreas Weinlich, Peter Amon, Andreas Hutter, André Kaup |
VCIP | 4 |
| 2012 | Mode adaptive reference frame denoising for high fidelity compression in HEVCabstractA new video coding standard, High Efficiency Video Coding (HEVC), is currently under development. In this paper we propose two relatively low-complex adaptive Wiener schemes for efficient P-frame coding of noisy videos with HEVC. In the proposed in-loop denoising framework the reference frame is noise filtered for P-frame prediction. The introduced algorithms adapt to the HEVC coding structure and thus can efficiently model the noise within the reference frame. The simulation results show that on the one hand considerable compression gains can be achieved using the proposed in-loop denoising framework. On the other hand the developed algorithms decrease the encoder runtime, but at the same time increase the decoder runtime. Eugen Wige, Gilbert Yammine, Wolfgang Schnurrer, André Kaup |
VCIP | 4 |
| 2011 | Sparsity-based defect pixel compensation for arbitrary camera raw imagesabstractIn high quality imaging even tiny distortions as small as a single pixel are visible and can not be accepted. Although the production quality of CMOS image sensors is very high, for reasonable yields we still need to accept some defect pixels and clusters of defects in large image sensors. In this paper we will compare compensation algorithms for raw image sensor data. We propose a new approach based on the sparsity assumption that outperforms existing defect compensation algorithms. Furthermore, our proposed interpolation algorithm is universal and not at all adapted to Bayer pattern images. It can directly be applied to any regular color filter pattern or gray scale image. Our examples show, that image sensors with large clusters of defects can still be used for the generation of high quality images. Michael Schöberl, Jürgen Seiler, Bernhard Kasper, Siegfried Fößel, André Kaup |
ICASSP | 5 |
| 2011 | Reusing the H.264/AVC deblocking filter for efficient spatio-temporal prediction in video codingabstractThe prediction step is a very important part of hybrid video codecs for effectively compressing video sequences. While existing video codecs predict either in temporal or in spatial direction only, the compression efficiency can be increased by a combined spatio-temporal prediction. In this paper we propose an algorithm for reusing the H.264/AVC deblocking filter for spatio-temporal prediction. Reusing this highly op timized filter allows for a very low computational complexity of this prediction mode and an average rate reduction of up to 7.2% can be achieved. Jürgen Seiler, André Kaup |
ICASSP | 2 |
| 2011 | Difference image extrapolation for spectral completion in inter-sequence error concealmentabstractMobile reception of digital TV often suffers from lost blocks and slices due to suboptimal channels. Inter-sequence error concealment reconstructs lost image blocks of distorted high-resolution TV signals by inserting corresponding error-free blocks of low-resolution reference signals. It is well-suited for automotive multi-broadcast receivers and can outperform state-of-the-art methods by up to 15 dB PSNR. In this contribution, a novel spectral completion scheme is proposed which robustly estimates missing spatial frequencies of reconstructed blocks by extrapolation of difference images. Simulation results show that the proposed scheme can increase the re construction quality of inter-sequence error concealment by up to 3.2 dB. Also, the subjective quality is enhanced by significantly reducing blurring artifacts. Tobias Tröger, André Kaup |
ICASSP | 2 |
| 2011 | Boosting based object detection using a geometric modelabstractIn this paper we present a new method for automatic object detection in images and video sequences. As a classifier the popular Ad aBoost algorithm is used, that combines several weak classifiers into one strong classifier. To create a detector based on this classifier, the weak classifiers are set into relation during boosting by using a geometric model. All votes of the weak detectors are evaluated in a voting space. The voting space allows a detection with combinations of different object features. We trained and tested the proposed method with SIFT and kAS features and combinations of these. The learned detector is then used to localize objects in images and video sequences. The performance of the algorithm is examined based on selected image data. Katharina Quast, Christoph Seeger, Mohan M. Trivedi, André Kaup |
ICIP | 4 |
| 2011 | Increasing imaging resolution by covering your sensorabstractUp to now, an increase in camera resolution required image sensors with more and more pixels. However, acquisition systems are limited in their pixels per second throughput given as power and complexity constraints. Simply capturing more pixels in a given system is often not possible. We propose a new non-regular imaging architecture that samples only few pixels and reconstructs a high resolution image afterwards. Our sampling is optimized to provide non-regular spatial sampling from a sensor with regular readout circuits. An existing slow image acquisition system can then be used to capture the data. The image reconstruction is performed with a local sparsity-based approach. The result is a high resolution image that requires a much smaller effort during acquisition. Michael Schöberl, Jürgen Seiler, Siegfried Fößel, André Kaup |
ICIP | 4 |
| 2011 | Increasing camera dynamic range through in-sensor multi-exposure white balancingabstractIn typical image sensors the spectral sensitivity of color channels is fixed. The illumination spectrum in natural scenes can vary to a great extent. This leads to an unbalanced response in color channels and hence a reduction in dynamic range. We propose a new method to adjust the relative sensitivity of color channels based on multi-exposure frame combination. Instead of a single long exposure we capture a different number of gapless exposures in each color channel and combine them. This offers a digital option for reducing sensitivity for some color channels and aligns the color channels. The method preserves motion blur and can be used in any long exposure or moving picture photography. We can now get a higher dynamic range from a camera system under any unfavorable illumination conditions with very little effort. Michael Schöberl, Wolfgang Schnurrer, Siegfried Fößel, André Kaup |
ICIP | 4 |
| 2011 | Temporal adaptation strategies for spatio-temporal image alignment in inter-sequence error concealment of digital TVabstractIn multi-broadcast scenarios, inter-sequence error concealment replaces lost image blocks of digital TV signals with error-free ones of low-resolution reference signals. Although single reference frames are misaligned spatio-temporally, the reconstruction quality is significantly higher than for conventional techniques. However, it can be further increased by reusing the alignment parameters of previous frames in case of misalignment of the current frame. Two strategies are proposed for this temporal adaptation of the alignment parameters either minimizing the additional complexity or maximizing the reconstruction quality. Overall, a gain of up to 5.0 dB PSNR Y is obtained while the computation time does not rise by more than 1.8%. Tobias Tröger, André Kaup |
ICIP | 2 |
| 2011 | Efficient coding of video sequences by non-local in-loop denoising of reference framesabstractThe compression efficiency of coding noisy image sequences is highly dependent on the noise itself and on the used quantization parameter. For very high quality near lossless compression a promising approach to save bitrate is to introduce an in-loop denoising filter in the codec. In this paper, we enhance our in-loop denoising scheme by a low complexity noise variance estimation algorithm. Additionally, we compare three different algorithms for denoising of the reference frame for medium to high quality compression of noisy videos. We use a low complexity adaptive Wiener filter, the NonLocal Means algorithm, and the BM3D algorithm for denoising of the reference frame only. We show that maximum bitrate savings up to 18% can be achieved using a sophisticated algorithm like the BM3D and up to 12% using a low complexity algorithm like the adaptive Wiener filter. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 5 |
| 2011 | A compressed domain change detection algorithm for RTP streams in video surveillance applicationsabstractThis paper presents a novel change detection algorithm for the compressed domain. Many video surveillance systems in practical use transmit their video data over a network by using the Real-time Transport Protocol (RTP). Therefore, the presented algorithm concentrates on analyzing RTP streams to detect major changes within contained video content. The paper focuses on a reliable preselection for further analysis modules by decreasing the number of events to be investigated. The algorithm is designed to work on scenes with mainly static background, like in indoor video surveillance streams. The extracted stream elements are RTP timestamps and RTP packet sizes. Both values are directly accessible by efficient byte-reading operations without any further decoding of the video content. Hence, the proposed approach is codec-independent, while at the same time its very low complexity enables the use in extensive video surveillance systems. About 40,000 frames per second of a single RTP stream can be processed on an Intel®Core™ 2 Duo CPU at 2 GHz and 2 GB RAM, without decreasing the efficiency of the algorithm. Marcus Laumer, Peter Amon, Andreas Hutter, André Kaup |
MMSP | 4 |
| 2011 | Adaptive in-loop noise-filtered prediction for High Efficiency Video CodingabstractCompression of noisy image sequences is a hard challenge in video coding. Especially for high quality compression the preprocessing of videos is not possible, as it decreases the objective quality of the videos. In order to overcome this problem, this paper presents an in-loop denoising framework for efficient medium to high fidelity compression of noisy video data. It is shown that using low complexity in-loop noise estimation and noise filtering as well as adaptive selection of the denoised inter frame predictors can improve the compression performance. The proposed algorithm for adaptive selection of the denoised predictor is based on the actual HEVC reference model. The different inter frame prediction modes within the current HEVC reference model are exploited for adaptive selection of denoised prediction by transmission of some side information in combination with decoder side estimation for denoised prediction. The simulation results show considerable gains using the proposed in-loop denoising framework with adaptive selection. In addition the theoretical bounds for the compression efficiency, if we could perfectly estimate the adaptive selection of the denoised prediction in the decoder, are shown in the simulation results. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
MMSP | 5 |
| 2011 | Motion Compensated Three-Dimensional Frequency Selective Extrapolation for improved error concealment in video communication
Jürgen Seiler, André Kaup |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Methods and Tools for Wavelet-Based Scalable Multiview Video CodingabstractA wavelet-based multiview video coding scheme is presented in this paper. It uses a 4-D wavelet transform, which is composed of a 1-D temporal wavelet transform, namely motion compensated temporal filtering, a 1-D view-directional wavelet transform, namely disparity compensated view filtering and a 2-D spatial wavelet transform. Since the presented framework can make use of the inherent scalability properties of the wavelet transforms involved, it allows full scalability of the coded bitstream in the temporal, view, spatial, and quality dimensions. Coding performance close to the H.264/advanced video coding based standard multiview video codec is shown. Enhancements of the view transform, in order to better account for brightness and color variations across views are introduced. Additionally, the use of a signal adaptive anisotropic wavelet packet (WP) transform as a generalization of WP transforms for the spatial decomposition is proposed. Both enhancements lead to a decrease of bit rate of up to 11% compared with the baseline version of the codec. Jens-Uwe Garbas, Béatrice Pesquet-Popescu, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Multiple Selection Approximation for improved spatio-temporal prediction in video codingabstractIn this contribution, a novel spatio-temporal prediction algorithm for video coding is introduced. This algorithm exploits temporal as well as spatial redundancies for effectively predicting the signal to be encoded. To achieve this, the algorithm operates in two stages. Initially, motion compensated prediction is applied on the block being encoded. Afterwards this preliminary temporal prediction is refined by forming a joint model of the initial predictor and the spatially adjacent already transmitted blocks. The novel algorithm is able to outperform earlier refinement algorithms in speed and prediction quality. Compared to pure motion compensated prediction, the mean data rate can be reduced by up to 15% and up to 1.16 dB gain in PSNR can be achieved for the considered sequences. Jürgen Seiler, André Kaup |
ICASSP | 2 |
| 2010 | Fixed pattern noise column drift compensation (CDC) for digital moving picture camerasabstractIn CMOS image sensors tiny semiconductor variations cause a distortion of the image known as fixed pattern noise (FPN). With a good FPN compensation CMOS sensors can deliver very good image quality. Compensation algorithms differ in complexity, maximum frame rate, and compensation quality with thermal drift. This paper analyzes existing approaches to offset FPN compensation and their drawbacks. We also show a new compensation algorithm that offers fine grained per-pixel compensation and is able to compensate temperature variations while still operating at the maximum sensor frame rate. This enables the construction of motion picture cameras without active cooling or even temperature measurements while still delivering a high image quality at high frame rates. Michael Schöberl, Siegfried Fößel, André Kaup |
ICIP | 3 |
| 2010 | Dimensioning of optical birefringent anti-alias filters for digital camerasabstractA digital camera samples the continuous real world. As with any sampling process, questions of aliasing for certain sampling frequencies and the prevention thereof arise. In this paper we will discuss the spatial domain sampling and prevention of aliasing in digital cameras. We focus on the widely used birefringent anti alias filters that are often called optical low pass filters (OLPF). We show 2D models for all contributions to spatial domain sampling and derive optimum filter parameters for minimum aliasing and best possible image sharpness. Compared to previously used selection rules, we can show that the optimum selection of filter parameters can easily deliver more sharpness and reduce aliasing by a factor of 2. The simulated results are finally confirmed in real world experiments. Michael Schöberl, Wolfgang Schnurrer, Alexander Oberdörster, Siegfried Fößel, André Kaup |
ICIP | 5 |
| 2010 | Content-Adaptive Motion Compensated Frequency Selective Extrapolation for error concealment in video communicationabstractIf digital video data is transmitted over unreliable channels such as the internet or wireless terminals, the risk of severe image distortion due to transmission errors is ubiquitous. To cope with this, error concealment can be applied on the distorted data at the receiver. In this contribution we propose a novel spatio-temporal error concealment algorithm, the Content-Adaptive Motion Compensated Frequency Selective Extrapolation. The algorithm operates in two stages, whereas at first the motion in a distorted sequence is estimated. After that, a model of the signal is generated for concealing the distortion. The novel algorithm is based on an already existent error concealment algorithm. But by adapting the model generation to the content of a sequence, the novel algorithm is able to exploit the remaining information, which is still available in the distorted sequence, more effectively compared to the original algorithm. In doing so, a visually noticeable gain of up to 0.51 dB PSNR compared to the underlying algorithm and more than 3 dB compared to other error concealment algorithms can be achieved. Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2010 | Fast and robust spatio-temporal image alignment for inter-sequence error concealmentabstractTerrestrial broadcasting of digital TV is often error-prone due to suboptimal channels. Inter-sequence error concealment is a novel technique to reconstruct lost image blocks of video signals by inserting error-free blocks of reference signals. Considering multi-broadcast receivers, reference signals are typically scaled, cropped and delayed. Spatio-temporal image alignment therefore is the crucial point of inter-sequence error concealment. A well-known intensity-based approach utilizing a numerical optimization technique shows supreme reconstruction quality outperforming state-of-the-art concealment methods by up to 15 dB PSNRY. In this paper, we present a novel feature-based approach which significantly reduces complexity. By using scale-invariant features for image alignment, the computation time can be reduced by a factor of up to 32.4. As reconstruction quality is almost kept constant, robustness of our algorithm is high. The maximum loss is 0.8 dB in terms of PSNRY. Tobias Tröger, André Kaup |
ICIP | 2 |
| 2010 | In-loop denoising of reference frames for lossless coding of noisy image sequencesabstractThe major gain in video coding applications compared to single image coding is the use of temporal prediction, which exploits the correlation between adjacent frames. However, in high quality video coding, especially lossless video coding, the compression gain of P-frames over I-Frames becomes very small. The reason for that is that the reference frame for inter prediction is not good enough, and therefore the amount of inter-predicted blocks in a P-Frame becomes relatively small compared to the number of intra-predicted blocks. In order to generate a better predictor for inter-prediction, we propose to remove additive noise from the reference frame using an adaptive Wiener filter. This way, we could achieve a maximum compression gain of 4.6% and an average compression gain of 3.3% in contrast to the H.264/AVC standard for lossless coding of high quality image sequences without affecting the encoding time noticeably. Eugen Wige, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 4 |
| 2010 | A no-reference Blocking Artifacts Visibility Estimator in imagesabstractIn this paper, we present a new low-complexity no-reference metric that assesses the visibility of blocking artifacts in DCT-coded images. The metric works in the spatial domain as well as in the gradient image domain without remarkable changes. Our conducted simulations showed that the metric is highly robust among many image distortion databases and extremely consistent with subjective mean opinion scores (MOS). The metric also provides a blocking visibility map that can be used for adaptive filtering of blocking artifacts. Gilbert Yammine, Eugen Wige, André Kaup |
ICIP | 3 |
| 2010 | Adaptive quantization parameter cascading for hierarchical video codingabstractQuantization parameter (QP) cascaded hierarchical prediction structures have been proved as efficient techniques in hybrid video coding. However, the current QP cascading method is empirical and not adaptive. The reason for the higher coding efficiency of this method has not been fully explored so far. In this paper, the rate-distortion performance of QP cascaded hierarchical video coding is first analyzed with dependent rate-distortion function. Then the optimal offset in linear QP cascading scheme is derived theoretically. It is shown that the widely accepted empirical QP cascading method is actually an approximation to the theoretical solution in fast movement environment. For slow sequences, an average gain of 0.43 dB can be achieved by the combination of two proposed adaptive algorithms. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ISCAS | 4 |
| 2010 | Improved mode selection in hybrid error concealment for multi-broadcast-receptionabstractDigital TV signals are often degraded by lost macro blocks or slices if received with mobile devices. Recently, we introduced a technique for inter-sequence error concealment in multi-broadcast scenarios which reconstruct lost image blocks of video signals by inserting error-free blocks of corresponding reference signals. The algorithm outperforms state-of-the-art techniques by up to 15 dB PSNRY. The quality of reconstructed image blocks can he further enhanced by combining inter-sequence error concealment with a temporal concealment approach as we showed in 2009. To enhance the robustness of this hybrid error concealment, we propose a novel metric for mode selection in this paper which takes into account local image statistics. While keeping the complexity of hybrid error concealment nearly constant, a mean gain of up to 2.1 dB and a peak gain of up to 10.1 dB in terms of PSNRYcan be achieved additionally. Tobias Tröger, Henning Heiber, Andreas Schmitt, André Kaup |
ISCAS | 4 |
| 2010 | Spatially refined inter-sequence error concealment for a multi-broadcast receiver using frequency selective approximationabstractMobile reception of digital TV often suffers from severe signal degradations. Inter-sequence error concealment reconstructs lost image blocks of a distorted high-resolution TV signal by inserting corresponding error-free blocks from a low-resolution reference TV signal. It is well-suited for application in future automotive multi-broadcast receivers and can outperform state-of-the-art methods by up to 15 dB PSNRY depending on the quality of the reference signal. In this contribution, we show that inter-sequence error concealment can be improved by approximating inserted blocks jointly with neighboring pixels according to a well-known frequency selective method. The reconstruction quality can be significantly increased especially in case of low-bitrate reference signals. Inserted blocks can be further refined also for high bitrates. On average, a gain of 1.7 dB PSNRY can be achieved. The peak gain which is evaluated on frame basis even reaches 5.6 dB PSNRY. Tobias Tröger, Jürgen Seiler, André Kaup |
ACM Multimedia | 3 |
| 2010 | Spatio-temporal prediction in video coding by non-local means refined motion compensationabstractThe prediction step is a very important part of hybrid video codecs. In this contribution, a novel spatio-temporal prediction algorithm is introduced. For this, the prediction is carried out in two steps. Firstly, a preliminary temporal prediction is conducted by motion compensation. Afterwards, spatial refinement is carried out for incorporating spatial redundancies from already decoded neighboring blocks. Thereby, the spatial refinement is achieved by applying Non-Local Means de-noising to the union of the motion compensated block and the already decoded blocks. Including the spatial refinement into H.264/AVC, a rate reduction of up to 14 % or respectively a gain of up to 0.7 dB PSNR compared to unrefined motion compensated prediction can be achieved. Jürgen Seiler, Thomas Richter 0001, André Kaup |
PCS | 3 |
| 2010 | Analysis of in-loop denoising in lossy transform codingabstractWhen compressing noisy image sequences, the compression efficiency is limited by the noise amount within these image sequences as the noise part cannot be predicted. In this paper, we investigate the influence of noise within the reference frame on lossy video coding of noisy image sequences. We estimate how much noise is left within a lossy coded reference frame. Therefore we analyze the transform and quantization step inside a hybrid video coder, specifically H.264/AVC. The noise power after transform, quantization, and inverse transform is calculated analytically. We use knowledge of the noise power within the reference frame in order to improve the inter frame prediction. For noise filtering of the reference frame, we implemented a simple denoising algorithm inside the H.264/AVC reference software JM15.1. We show that the bitrate can be decreased by up to 8.1% compared to the H.264/AVC standard for high resolution noisy image sequences. Eugen Wige, Gilbert Yammine, Peter Amon, Andreas Hutter, André Kaup |
PCS | 5 |
| 2010 | Blind GOP structure analysis of MPEG-2 and H.264/AVC decoded videoabstractIn this paper, we provide a simple method for analyzing the GOP structure of an MPEG-2 or H.264/AVC decoded video without having access to the bitstream. Noise estimation is applied on the decoded frames and the variance of the noise in the different I-, P-, and B-frames is measured. After the encoding process, the noise variance in the video sequence shows a periodic pattern, which helps in the extraction of the GOP period, as well as the type of frames. This algorithm can be used along with other algorithms to blindly analyze the encoding history of a video sequence. The method has been tested on several MPEG-2 DVB and DVD streams, as well as on H.264/AVC encoded sequences, and shows successful results in both cases. Gilbert Yammine, Eugen Wige, André Kaup |
PCS | 3 |
| 2010 | Complex-Valued Frequency Selective Extrapolation for Fast Image and Video Signal ExtrapolationabstractSignal extrapolation tasks arise in miscellaneous manners in the field of image and video signal processing. But, due to the widespread use of low-power and mobile devices, the computational complexity of an algorithm plays a crucial role in selecting an algorithm for a given problem. Within the scope of this contribution, we introduce the complex-valued Frequency Selective Extrapolation for fast image and video signal extrapolation. This algorithm iteratively generates a generic complex-valued model of the signal to be extrapolated as weighted superposition of Fourier basis functions. We further show that this algorithm is up to 10 times faster than the existent real-valued Frequency Selective Extrapolation that takes the real-valued nature of the input signals into account during the model generation. At the same time, the quality which is achievable by the complex-valued model generation is similar to the quality of the real-valued model generation. Jürgen Seiler, André Kaup |
IEEE Signal Process. Lett. | 2 |
| 2009 | One-pass multi-layer rate-distortion optimization for quality scalable video codingabstractIn this paper, a one-pass multi-layer rate-distortion optimization algorithm is proposed for quality scalable video coding. To improve the overall coding efficiency, the MB mode in the base layer is selected not only based on its rate-distortion performance relative to this layer but also according to its impact on the enhancement layer. Moreover, the optimization module for residues is also improved to benefit inter-layer prediction. Simulations show that the proposed algorithm outperforms the most recent SVC reference software. For eight test sequences, a gain of 0.35 dB on average and 0.75 dB at maximum is achieved at a cost of less than 8% increase of the total coding time. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ICASSP | 4 |
| 2009 | Model based analysis for quantization parameter cascading in hierarchical video codingabstractOriginally, hierarchical prediction structures were proposed to achieve temporal scalability. Soon after, it was realized that with a proper quantization parameter cascading (QPC) scheme the general performance can be significantly improved by hierarchical coding. However, the theory behind the gain has not been explored so far. In this paper, the QPC in hierarchical coding is investigated by model based emulations. From the analysis, it is noticed that a parameter β which represents the error propagation in a group of pictures greatly affects the performance of QPC in hierarchical coding. Based on β, a simple adaptive QPC algorithm is designed. Simulations verify the efficiency of this algorithm: a gain up to 0.89 dB is obtained over the most recent SVC reference software. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
ICIP | 4 |
| 2009 | Modeling of image shutters and motion blur in analog and digital camera systemsabstractFor motion imaging the perceived smoothness of a sequence highly depends on motion blur. The exposure for each frame is started and ended with a shutter mechanism. There are different implementations and some of them have non-ideal behavior which can introduce artifacts. In this paper an overview of real-world shutter implementations both for analog and digital camera systems is shown. We develop a general description that models all types of shutters and their imperfections. Specific models for common shutter types are presented. Measurements are used to estimate unknown parameters. The modeled shutters are finally used for a virtual camera simulation. Typical artifacts can be simulated and directly compared for different shutter types and parameters. This powerful tool is useful for the construction of camera systems and allows design decisions to be directly compared before building the camera system. Michael Schöberl, Siegfried Fößel, Hans Bloß, André Kaup |
ICIP | 4 |
| 2009 | Enhancing coding efficiency in spatial scalable multiview video coding with waveletsabstractScalable multiview video coding might become important in future immersive telecommunication scenarios. Therefore, we present a fully scalable multiview video coding framework that is based on wavelets. It offers scalable decoding of the bitstream in the time, view, spatial, and quality dimensions. Whereas its performance is already comparable to state-of-the-art multiview video coding at the full spatial resolution, the quality of the downscaled spatial resolution is not satisfying. After reviewing the reasons for this, we propose an extended coding scheme where a spatial wavelet transform is performed first and the high spatial resolution layers of the multiview video sequence are predicted from the lower ones in the wavelet domain. The structure and performance of this in-band prediction are analyzed and several enhancements are proposed. Finally, it can be shown that efficient spatial scalability can be achieved in this framework with only minor performance loss at the full resolution. Jens-Uwe Garbas, André Kaup |
MMSP | 2 |
| 2009 | One-pass frame level budget allocation in video coding using inter-frame dependencyabstractIn this paper, a one-pass budget allocation algorithm is proposed for hybrid video coding. Taking the percentage of skipped MBs as the measure of inter-frame dependency, the optimal budget allocation is first modeled for a two-frame case. Then this model is extended to a practical method in slow movement scenario, where the information of inter-frame dependency is predicted based on the previously coded frames. Simulations show that the proposed algorithm outperforms the recommended MB and frame level rate control algorithms in H.264/AVC reference software JM 15.1. For eight CIF sequences with slow movement, significant gains of 0.91 dB and 0.43 dB on average (1.66 dB and 1.49 dB at maximum) were obtained over the reference MB and frame level rate control, respectively. Considering that the computational complexity by the proposed algorithm is quite low (less than 1% of the total coding time when fast motion estimation algorithm is enabled), it is quite appealing for real-time video applications. Xiang Li 0003, Andreas Hutter, André Kaup |
MMSP | 3 |
| 2009 | Multiple Selection Extrapolation for improved spatial error concealmentabstractThis contribution introduces a novel signal extrapolation algorithm and its application to image error concealment. The signal extrapolation is carried out by iteratively generating a model of the signal suffering from distortion. Thereby, the model results from a weighted superposition of two-dimensional basis functions whereas in every iteration step a set of these is selected and the approximation residual is projected onto the subspace they span. The algorithm is an improvement to the Frequency Selective Extrapolation that has proven to be an effective method for concealing lost or distorted image regions. Compared to this algorithm, the novel algorithm is able to reduce the processing time by a factor larger than three, by still preserving the very high extrapolation quality. Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2009 | Low-complexity inter-sequence error concealment based on scale-invariant feature transformabstractInter-sequence error concealment defines a novel class of concealment techniques which utilize two or more representations of a particular video sequence for image reconstruction. An image-based approach introduced recently provides superior reconstruction quality of concealed image parts and outperforms well-known intra methods by up to 15 dB PSNRY. However, computational complexity is high due to a numerical solution of the image registration problem. In this paper, we propose a feature-based technique for image registration which is included in the concept of inter-sequence error concealment. Based on the determination of scale-invariant features, the processing time is reduced on average by a factor of up to 28. At the same time, the loss in terms of objective image quality compared to the image-based approach is less than 1 dB PSNRY. Tobias Tröger, Henning Heiber, Andreas Schmitt, André Kaup |
MMSP | 4 |
| 2009 | Lagrange multiplier selection for rate-distortion optimization in SVCabstractThe Lagrangian multiplier based rate-distortion optimization (RDO) has been widely employed in single layer video coding. During the development of scalable video coding (SVC) extension of H.264/AVC, it was directly applied in a multilayer scenario. However, such an application is not very efficient since the correlation between layers is not considered in the Lagrange multiplier selection. To improve the overall performance, in this paper a new selection algorithm is presented for RDO in SVC. Simulations show that the proposed method outperforms the recent SVC reference software. With a tiny computational cost, average gains of 0.22 dB and 0.35 dB were achieved in the tests of four-layer quality scalability and three-layer spatial scalability, respectively. Xiang Li 0003, Peter Amon, Andreas Hutter, André Kaup |
PCS | 4 |
| 2009 | Spatio-temporal prediction in video coding by best approximationabstractWithin the scope of this contribution we propose a novel efficient spatio-temporal prediction algorithm for video coding. The algorithm operates in two stages. First, motion compensation is performed on the block to be predicted in order to exploit temporal correlations. Afterwards, in order to exploit spatial correlations, this preliminary estimate is spatially refined by forming a joint model of the motion compensated block and spatially adjacent already decoded blocks. Compared to an earlier refinement algorithm, the novel one only needs very little iteration, leading to a speedup of factor 17. The implementation of this new algorithm into the H.264/AVC leads to a maximum reduction in data rate of up to nearly 13% for the considered sequences. Jürgen Seiler, Haricharan Lakshman, André Kaup |
PCS | 3 |
| 2009 | Efficient one-pass frame level rate control for H.264/AVC
Xiang Li 0003, Andreas Hutter, André Kaup |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Laplace Distribution Based Lagrangian Rate Distortion Optimization for Hybrid Video CodingabstractIn today's hybrid video coding, Rate-Distortion Optimization (RDO) plays a critical role. It aims at minimizing the distortion under a constraint on the rate. Currently, the most popular RDO algorithm for one-pass coding is the one recommended in the H.264/AVC reference software. It, or HR-$\lambda $for convenience, is actually a kind of universal method which performs the optimization only according to the quantization process while ignoring the properties of input sequences. Intuitively, it is not efficient all the time and an adaptive scheme should be better. Therefore, a new algorithm Lap-$\lambda $is presented in this paper. Based on the Laplace distribution of transformed residuals, the proposed Lap-$\lambda $is able to adaptively optimize the input sequences so that the overall coding efficiency is improved. Cases which cannot be well captured by the proposed models are considered via escape methods. Comprehensive simulations verify that compared with HR-$\lambda $, Lap-$\lambda $shows a much better or similar performance in all scenarios. Particularly, significant gains of 1.79 dB and 1.60 dB in PSNR are obtained for slow sequences and B-frames, respectively. Xiang Li 0003, Norbert Oertel, Andreas Hutter, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Spatial Scalable Region of Interest Transcoding of JPEG2000 for Video SurveillanceabstractA new method for spatial scalable region of interest transcoding of JPEG2000 images for video surveillance applications is proposed in this paper. The method transcodes HD images into quarter HD images with a ROI in HD resolution. The transcoding method is based on the precinct feature as well as on the empty packet feature of the JPEG2000 standard. Based on a user defined ROI the transcoder extracts all packets which belong to the ROI or to the lower resolution of the background. All non-ROI packets of higher resolution levels are deleted and replaced by empty-packets. Thus, the extracted codestream contains no high frequency information for image regions outside the ROI. Because the transcoder operates on packets and not on codeblocks an expensive re-encoding of the codestream is not required. This leads to a short processing time and a low complexity of the transcoding algorithm. Therefore, our method is appropriate for real-time applications like video surveillance. Katharina Quast, André Kaup |
AVSS | 2 |
| 2008 | Fast orthogonality deficiency compensation for improved frequency selective image extrapolationabstractThe purpose of this paper is to introduce a very efficient algorithm for signal extrapolation. It can widely be used in many applications in image and video communication, e. g. for concealment of block errors caused by transmission errors or for prediction in video coding. The signal extrapolation is performed by extending a signal from a limited number of known samples into areas beyond these samples. Therefore a finite set of orthogonal basis functions is used and the known part of the signal is projected onto them. Since the basis functions are not orthogonal regarding the area of the known samples, the projection does not lead to the real portion a basis function has of the signal. The proposed algorithm efficiently copes with this non-orthogonality resulting in very good objective and visual extrapolation results for edges, smooth areas, as well as structured areas. Compared to an existent implementation, this algorithm has a significantly lower computational complexity without any degradation in quality. The processing time can be reduced by a factor larger than 100. Jürgen Seiler, André Kaup |
ICASSP | 2 |
| 2008 | Wavelet-based multi-view video coding with joint best basis wavelet packetsabstractAn approach to scalable multi-view video coding with joint best basis wavelet packets is examined in this paper. A 4-D wavelet transform is used to decorrelate the multi-view video data temporally, view-directionally, and spatially for efficient scalable compression. Motion compensated temporal filtering (MCTF) is used for temporal, and disparity compensated view filtering (DCVF) for view- directional decomposition. Adaptive wavelet packets as a generalized wavelet decomposition are presented for spatial decomposition. Two algorithms to find the best basis wavelet packets are evaluated and compared with classical dyadic wavelet transform: a low complexity entropy based best basis search and a search algorithm in a rate-distortion framework. In both cases, the joint best basis is determined for a group of frames rather than for each frame individually. Therefore, the rate to spend for the tree description is minimal while advantage is taken of the similarity of frames within a temporal-view-directional subband. Jens-Uwe Garbas, Béatrice Pesquet-Popescu, Maria Trocan, André Kaup |
ICIP | 4 |
| 2008 | Spatio-temporal prediction in video coding by spatially refined motion compensationabstractThe purpose of this contribution is to introduce a new method of signal prediction in video coding. Unlike most existent prediction methods that either use temporal or use spatial correlations to generate the prediction signal, the proposed method uses spatial and temporal correlations at the same time. The spatio-temporal prediction is obtained by first performing motion compensation for a macroblock, followed by a refinement step that pays attention to the correlations between the macroblock and its surroundings. At the decoder, the refinement step can be performed in the same manner, thus no additional side information has to be transmitted. Implementation of the spatial refinement step into the H.264/AVC video codec leads to reduction in data rate of up to nearly 15% and increase in PSNR of up to 0.75 dB, compared to pure motion compensated prediction. Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2008 | Analysis of spatio-temporal prediction methods in 4D volumetric medical image datasetsabstractDue to the huge amount of data and the increasing utilization of 4D medical image processing, compression of such data sets is essential. Unlike moving 3D objects in computer graphic applications, medical 4D datasets consist of a number of sampled volume elements, varying in time. Based on an analysis of H.264 compression for such data, this paper presents a spatio-temporal prediction scheme leveraging effective block-based prediction for 4D volumetric image data. Experimental results show that this new spatiotemporal approach achieves better prediction results than pure spatial or temporal schemes. Uwe-Erik Martin, André Kaup |
ICME | 2 |
| 2008 | 4-D frequency selective extrapolation for error concealment in multi-view videoabstractPractical applications for multi-view video such as three-dimensional television may require to transmit the data over error-prone channels. Even when channel coding is used, it is likely that parts of the image information are lost after decoding. To improve the image quality in this case, an efficient algorithm for multi-view error concealment is presented. Extending and enhancing a previously known method for single-view concealment, the algorithm simultaneously uses information from surrounding image parts, from temporally preceding and succeeding frames and from neighbouring camera views for extrapolating known image samples into the lost area. It is shown that the proposed algorithm leads to convincing results in terms of PSNR as well as to a good subjective quality. Compared to single-view concealment, the result is improved by up to 0.75 dB by adding cross-view information. Ulrich Fecker, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2008 | Adaptive joint spatio-temporal error concealment for video communicationabstractIn the past years, video communication has found its application in an increasing number of environments. Unfortunately, some of them are error-prone and the risk of block losses caused by transmission errors is ubiquitous. To reduce the effects of these block losses, a new spatio-temporal error concealment algorithm is presented. The algorithm uses spatial as well as temporal information for extrapolating the signal into the lost areas. The extrapolation is carried out in two steps, first a preliminary temporal extrapolation is performed which then is used to generate a model of the original signal, using the spatial neighborhood of the lost block. By applying the spatial refinement a significantly higher concealment quality can be achieved resulting in a gain of up to 5.2 dB in PSNR compared to the unrefined underlying pure temporal extrapolation. Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2008 | Histogram-Based Prefiltering for Luminance and Chrominance Compensation of Multiview VideoabstractSignificant advances have recently been made in the coding of video data recorded with multiple cameras. However, luminance and chrominance variations between the camera views may deteriorate the performance of multiview codecs and image-based rendering algorithms. A histogram matching algorithm can be applied to efficiently compensate for these differences in a prefiltering step. A mapping function is derived which adapts the cumulative histogram of a distorted sequence to the cumulative histogram of a reference sequence. If all camera views of a multiview sequence are adapted to a common reference using histogram matching, the spatial prediction across camera views is improved. The basic algorithm is extended in three ways: a time-constant calculation of the mapping function, RGB color conversion, and the use of global disparity compensation. The best coding results are achieved when time-constant histogram calculation and RGB color conversion are combined. In this case, the usage of histogram matching prior to multiview encoding leads to substantial gains in the coding efficiency of up to 0.7 dB for the luminance component and up to 1.9 dB for the chrominance components. This prefiltering step can be combined with block-based illumination compensation techniques that modify the coder and decoder themselves, especially with the approach implemented in the multiview reference software of the joint video team (JVT). Additional coding gains up to 0.4 dB can be observed when both methods are combined. Ulrich Fecker, Marcus Barkowsky, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Low-Complexity Heterogeneous Video Transcoding Using Data MiningabstractRecent developments have given birth to H.264/AVC: a video coding standard offering better bandwidth to video quality ratios than previous standards (such as H.263, MPEG-2, MPEG-4, etc.), due to its improved inter- and intraprediction modes at the expense of higher computation complexity. It is expected that H.264/AVC will take over the digital video market, replacing the use of previous standards in most digital video applications. This creates an important need for heterogeneous video transcoding technologies from older standards to H.264. In this paper, we focus our attention on the interframe prediction, the most computationally intensive task involved in the heterogeneous video transcoding process. This paper presents a novel macroblock (MB) mode decision algorithm for interframe prediction based on data mining techniques to be used as part of a very low complexity heterogeneous video transcoder. The proposed approach is based on the hypothesis that MB coding mode decisions in H.264 video have a correlation with the distribution of the motion compensated residual in the decoded video. We use data mining tools to exploit the correlation and derive decision trees to classify the incoming decoded MBs into one of the several coding modes in H.264. The proposed approach reduces the H.264 MB mode computation process into a decision tree lookup with very low complexity. For general validation purposes, we apply our algorithm to two of the most important heterogeneous video transcoders: MPEG-2 to H.264 and H.263 to H.264. Our results show that the our data-mining based transcoding algorithm is able to maintain a good video quality while considerably reducing the computational complexity by 72% on average when applied in MPEG-2 to H.264 transcoders, and by 62% on average when applied in H.263 to H.264 transcoders. Finally, we conduct a comparative study with some of the most prominent fast interprediction methods for H.264 presented in the literature. Our results show that the proposed data mining-based approach achieves the best results for video transcoding applications. Gerardo Fernández-Escribano, Jens Bialkowski, José A. Gámez 0001, Hari Kalva, Pedro Cuenca 0001, Luis Orozco-Barbosa, André Kaup |
IEEE Trans. Multim. | 7 |
| 2007 | H.263 to H.264 Transconding using Data MiningabstractIn this paper, we propose the use of data mining algorithms to create a macroblock partition mode decision algorithm for inter-frame prediction, to be used as part of a high-efficient H.263 to H.264 transcoder. We use machine learning tools to exploit the correlation and derive decision trees to classify the incoming H.263 MC residual into one of the several coding modes in H.264. The proposed approach reduces the H.264 MB mode computation process into a decision tree lookup with very low complexity. Experimental results show that the proposed approach reduces the inter-prediction complexity by as much as 60% while maintaining the coding efficiency. Gerardo Fernández-Escribano, Jens Bialkowski, Hari Kalva, Pedro Cuenca 0001, Luis Orozco-Barbosa, André Kaup |
ICIP (4) | 6 |
| 2007 | Spatio-Temporal Defect Pixel Interpolation using 3-D Frequency Selective ExtrapolationabstractFlat panel X-ray detectors allow the immediate availability of the acquired images for display. However, they provide images with defective areas amongst others due to manufacturing problems. In this contribution, we present a frequency selective extrapolation method in order to restore these defects by extending the surrounding signal into the defective area. In case of static radiographs, the spatial surrounding is evaluated by a 2-D approach. Defects in sequences acquired by cine-angiography or fluoroscopy are processed by 3-D extrapolation. The defects are replaced by extrapolating the signal from the spatial and at the same time temporal surrounding, taking previous and subsequent frames into account. Hence, inherent motion compensation is accomplished. The application of the spatial and where applicable spatio-temporal extrapolation approach allows to restore smooth areas, edges, patterns as well as noise. The ability to restore noise is especially important for medical images because it leads to a natural appearance of the concealed defective areas. Katrin Meisinger, Til Aach, André Kaup |
ICIP (4) | 3 |
| 2007 | Advanced Lagrange Multiplier Selection for Hybrid Video CodingabstractThe Lagrangian multiplier based rate-distortion optimization has been proved to be an effective way in hybrid video coding. In this paper, an advanced Lagrange multiplier selection method is presented. Based on Laplace distribution, the variance of transformed residuals is introduced into the rate and distortion models. Moreover, inspired by the ρ-domain method, the percentage of non-zeros among quantized residuals is also considered to refine the rate model. Thanks to the more accurate models, the proposed method is able to adaptively optimize the encoding process for videos with different properties. It outperforms the algorithm used in the reference software of H.264/AVC. According to the simulation, a gain up to 1.1dB was achieved. Xiang Li 0003, Norbert Oertel, Andreas Hutter, André Kaup |
ICME | 4 |
| 2007 | Adaptive Lagrange Multiplier Selection for Intra-Frame Video CodingabstractThe Lagrangian technique proves to be an effective way in Rate-Distortion optimization for hybrid video coding. In this paper, an new Lagrange multiplier selection method for Intra-Frame coding is presented. Based on Laplacian distribution, the variance of transformed residual coefficients is introduced into the rate and distortion model, so that the proposed method is able to adaptively optimize the encoding process for different types of videos. It outperforms the algorithm used in the reference software of H.264/AVC, especially for computer animations and videos with small movements. According to the simulations, a gain up to 0.3 dB was achieved. Xiang Li 0003, Norbert Oertel, André Kaup |
ISCAS | 3 |
| 2007 | Wavelet-based multi-view video coding with full scalability and illumination compensationabstractA wavelet-based approach to scalable multi-view video coding (MVC) is examined in this paper. A 4-D wavelet transform is used to decorrelate the multi-view video data temporally, view-directionally, and spatially for efficient compression. Motion compensated temporal filtering (MCTF) is applied to each video sequence of each camera to exploit temporal correlation and inter-view dependencies are exploited with disparity compensated view filtering (DCVF). Together with a 2-D spatial wavelet transform the 4-D wavelet transform is constituted. The open-loop structure of the decomposition, together with embedded quantization and coding, allows for the construction of a bitstream which is fully scalable in temporal, view-directional, spatial and quality dimension. Results for scalable decoding in all dimensions are given, as well as comparisons to simulcast coding and state-of-the-art non-scalable multi-view video coding. Further, a significant coding gain is shown when the sequences are pre-processed with an algorithm for cross-view illumination compensation based on histogram matching. Jens-Uwe Garbas, Ulrich Fecker, André Kaup |
ACM Multimedia | 3 |
| 2007 | Temporal registration using 3D phase correlation and a maximum likelihood approach in the perceptual evaluation of video qualityabstractThe estimation of the video quality is often performed using a full reference approach. One of the most important steps in a video quality measurement algorithm is to find the corresponding frames between the reference and the distorted video sequence. In this paper an algorithm with three steps is proposed. First, an extended version of the phase correlation is used to find candidate images with an arbitrary temporal offset, spatial scaling or spatial shift. Based on the assumption that the spatial scaling and spatial shift does not change during the sequence a set of probable parameters is selected. Finally, a maximum likelihood estimation is applied to select those temporal offsets which support the smoothest playback. A set of video sequences degraded with several distortions which are typical for multimedia scenarios are used to compare the performance to other algorithms. Marcus Barkowsky, Jens Bialkowski, Roland Bitto, André Kaup |
MMSP | 4 |
| 2007 | Wavelet-Based Multi-View Video Coding with Spatial ScalabilityabstractIn this paper, we propose two wavelet-based frameworks which allow fully scalable multi-view video coding. Using a 4-D wavelet transform, both schemes generate a bitstream that can be truncated to achieve a temporally, view-directionally, and/or spatially downscaled representation of the coded multi-view video sequence. Well-known wavelet-based scalable coding schemes for single-view video sequences have been adopted and extended to match the specific needs of scalable multi-view video coding. Motion compensated temporal filtering (MCTF) is applied to each Video sequence of each camera to exploit temporal correlation and inter-view dependencies are exploited with disparity compensated view filtering (DCVF). A spatial wavelet transform is utilized either before and after temporal-view-directional decomposition (2D+T+V+2D scheme) or only after the temporal-view-directional decomposition (T+V+2D scheme) for spatial decorrelation. The influence of the two different approaches on spatial scalability is shown in this paper as well as the superior coding efficiency of both codecs compared with simulcast coding. Jens-Uwe Garbas, André Kaup |
MMSP | 2 |
| 2007 | Fast video transcoding from H.263 to H.264/MPEG-4 AVC
Jens Bialkowski, Marcus Barkowsky, André Kaup |
Multim. Tools Appl. | 3 |
| 2007 | Spatiotemporal Selective Extrapolation for 3-D Signals and Its Applications in Video CommunicationsabstractIn this paper, we derive a spatiotemporal extrapolation method for 3-D discrete signals. Extending a discrete signal beyond a limited number of known samples is commonly referred to as discrete signal extrapolation. Extrapolation problems arise in many applications in video communications. Transmission errors in video communications may cause data losses which are concealed by extrapolating the surrounding video signal into the missing area. The same principle is applied for TV logo removal. Prediction in hybrid video coding is also interpreted as an extrapolation problem. Conventionally, the unknown areas in the video sequence are estimated from either the spatial or temporal surrounding. Our approach considers the spatiotemporal signal including the missing area in a volume and replaces the unknown samples by extrapolating the surrounding signal from spatial, as well as temporal direction. By exploiting spatial and temporal correlations at the same time, it is possible to inherently compensate motion. Deviations in luminance occurring from frame to frame can be compensated, too. Katrin Meisinger, André Kaup |
IEEE Trans. Image Process. | 2 |
| 2006 | Influence of the Presentation Time on Subjective Votings of Coded Still ImagesabstractThe quality of coded images is often assessed by a subjective test. Usually the viewers get as much time as they need to find a stable result. In video sequences however, the viewer has to judge the quality in a shorter time that is defined by the changing content or a following scene cut. Therefore it is desirable to know the influence of a shorter presentation time on the perceptibility of distortions. In this paper we present the results of a suitable subjective test on coded still images. The images were presented for six different durations, ranging from 200 ms to 3 s. Special care was taken to avoid the memorization effect usually present after short presentations. The results show that the viewers tend to avoid extreme votings at short durations. The variance of the votings is also discussed in detail. Based on the result of the voting for the longest presentation time, we propose a prediction model for the voting of the shorter durations using a logistic curve fit. This presentation time model (PTM) is presented and analysed in detail. Marcus Barkowsky, Björn M. Eskofier, Jens Bialkowski, André Kaup |
ICIP | 4 |
| 2006 | Low-Complexity Transcoding of Inter Coded Video Frames from H.264 to H.263abstractThe presented work addresses the reduction of computational complexity for transcoding of interframes from H.264 to H.263 baseline profiles maintaining the quality of a full search approach. This scenario aims to achieve fast backward compatible interoperability inbetween new and existing video coding platforms, e.g. between DVB-H and UMTS. By exploiting side information of the H.264 input bitstream the encoding complexity of the motion estimation is strongly reduced. Due to the possibility to divide a macroblock (MB) into partitions with different motion vectors (MV), one single MV has to be selected for H.263. It will be shown, that this vector is suboptimal for all sequences, even if all existing MVs of a MB of H.264 are compared as candidate. Also motion vector refinement with a fixed ½-pel refinement window as used by transcoders throughout the literature is not sufficient for scenes with fast movement. We propose an algorithm for selecting a suitable vector candidate from the input bitstream and this MV is then refined using an adaptive window. Using this technique, the complexity is still low at nearly optimum rate-distortion results compared to an exhaustive full-search approach. Jens Bialkowski, Marcus Barkowsky, Florian Leschka, André Kaup |
ICIP | 4 |
| 2006 | Spatio-Temporal Fading Scheme for Error Concealment in Block-Based Video Decoding SystemsabstractIn this contribution we propose a new spatio-temporal fading scheme for block loss recovery in block-based video decoding systems. Based on a boundary error criterion obtained from temporal error concealment, either spatial, temporal, or fading of both methods is used for recovering lost image samples in one macroblock. The weights for fading are interpolated from the boundary error. A weighted absolute difference between well received macroblock boundary samples from the current frame and motion compensated macroblock boundary samples from the previous frame represents the boundary error. It is shown, that in case of transmission errors this method can successfully be used for block loss recovery. Markus Friebe, André Kaup |
ICIP | 2 |
| 2006 | 2D Frequency Selective Extrapolation for Spatial Error Concealment in H.264/AVC Video CodingabstractThe frequency selective extrapolation extends an image signal beyond a limited number of known samples. This problem arises in image and video communication in error prone environments where transmission errors may lead to data losses. In order to estimate the lost image areas, the missing pixels are extrapolated from the available correctly received surrounding area which is approximated by a weighted linear combination of basis functions. In this contribution, we integrate the frequency selective extrapolation into the H.264/AVC coder as spatial concealment method. The decoder reference software uses spatial concealment only for I frames. Therefore, we investigate the performance of our concealment scheme for I frames and its impact on following P frames caused by error propagation due to predictive coding. Further, we compare the performance for coded video sequences in TV quality against the non-normative concealment feature of the decoder reference software. The investigations are done for slice patterns causing chequerboard and raster scan losses enabled by flexible macroblock ordering (FMO). Katrin Meisinger, André Kaup |
ICIP | 2 |
| 2006 | Overview of Low-Complexity Video Transcoding from H.263 to H.264abstractWith the standardization of H.264/AVC by ITU-T and ISO/IEC and the adaptatation into new hardware, the necessity of transcoding between existing standards and H.264 will arise to achieve interoperability between hardware devices. Because of the many new prediction parameters as well as the pixel-based deblocking filter and the new transform of H.264 this is a difficult task to perform. In our work we propose a fast cascaded pixel-domain transcoder from H.263 to H.264 for both intra- and inter-frame coding. The rate-distortion (RD) performance of the encoded bitstreams is compared to an exhaustive full-search approach. Our approach leads to 9% higher data rate in average, but the computational complexity for the prediction can be reduced by 90% and more. It will be shown that the algorithms proposed for H.263 are applicable for transcoding MPEG-2 to H.264, too Jens Bialkowski, Marcus Barkowsky, André Kaup |
ICME | 3 |
| 2006 | Spatio-Bi-Temporal Error Concealment in Block-Based Video Decoding SystemsabstractIn this paper we present a spatio-bi-temporal fading scheme for block loss recovery in block-based video decoding systems. In the first part of the algorithm, based on two different boundary error criterions obtained from bi-temporal error concealment, either the previous, the future, or fading between both temporal methods is used for bi-temporal macroblock estimation. A weighted absolute difference between motion compensated image samples and macroblock boundary samples of the current frame represents one boundary error. In the second part of the algorithm, based on a boundary error criterion obtained from bi-temporal concealment, spatial, bi-temporal, or fading between both methods is used for recovering a lost macroblock. The advantage of this method is that one lost macroblock can be recovered pelwise spatially from the current or bi-temporally from the previous and the future frame by weighted averaging both error concealment results. The simulation results have shown that for recovering a lost macroblock this method outperforms the reference methods both in subjective and objective video quality Markus Friebe, André Kaup |
MMSP | 2 |
| 2006 | 4D Scalable Multi-View Video Coding Using Disparity Compensated View Filtering and Motion Compensated Temporal FilteringabstractIn this paper, a novel framework for scalable multi-view video coding is described. A well known wavelet based scalable coding scheme for single-view video sequences has been adopted and extended to match the specific needs of scalable multi-view video coding. Motion compensated temporal filtering (MCTF) is applied to each video sequence of each camera. The use of a wavelet lifting structure guarantees perfect invertibility of this step, and as a consequence of its open-loop architecture, SNR and temporal scalability are attained. Correlations between the temporal subbands of adjacent cameras are reduced by a novel disparity compensated view filtering (DCVF), method which is also lifting based and open-loop to enable view scalability. Spatial scalability and entropy coding are achieved by the JPEG2000 spatial wavelet transform and EBCOT coding, respectively. Rate allocation along the temporal-view-filtered subbands is done by means of an RD-optimal algorithm. Experimental results show the high scaling capability in terms of SNR, temporal and view scalability Jens-Uwe Garbas, Ulrich Fecker, Tobias Tröger, André Kaup |
MMSP | 4 |
| 2006 | Gradient Intra Prediction for Coding of Computer Animated VideosabstractThe increasing interest in computer animations initiated the need of efficient coding for such applications. This paper proposes two gradient based approaches to improve the efficiency of Intra Prediction. Taking advantage of a special property of computer animations, namely the gradient distribution of intensity on object surfaces, the devised methods provided a better performance over H.264/AVC. According to the simulations, up to 0.3 dB gain in PSNR has been achieved. Xiang Li 0003, Norbert Oertel, André Kaup |
MMSP | 3 |
| 2006 | Spatio-Temporal Concealment in H.264/AVC Video Coding by 3-D Selective ExtrapolationabstractIn this contribution, we derive a spatio-temporal extrapolation method for 3-D signals. Lost areas in video signals caused by transmission errors in video communications are concealed by extrapolating the surrounding video signal into the missing area. The method is integrated into the H.264/AVC coder as concealment feature. By exploiting spatial and temporal correlations at the same time, it is possible to inherently compensate motion. Deviations in luminance occurring from frame to frame can be compensated, too. Since the spatio-temporal selective extrapolation method does not rely on motion vectors, the algorithm can be also easily applied to Intra concealment by additionally exploiting temporal correlations. Further, the frequency selective extrapolation shows especially convincing results in connection with flexible macroblock ordering (FMO) causing a chequerboard like loss pattern in case of transmission errors Katrin Meisinger, Sandra Martin, André Kaup |
MMSP | 3 |
| 2005 | Low complexity streak noise reduction for mobile TV using line selective interpolation of field informationabstractThis contribution presents a method for quality enhancement of mobile received analog TV signals using line selective interpolation of field information (LSI-FI). Using combining techniques of neighboring lines one can restaurate mobile received images successfully. Spatial statistical methods are used to detect distorted lines of an image. Distorted lines are interpolated by correctly received neighboring lines. The low complexity of the algorithm makes it possible to enable real-time implementation in a DSP or FPGA. Markus Friebe, André Kaup |
ICIP (3) | 2 |
| 2005 | On Requantization in Intra-Frame Video Transcoding with Different Transform Block SizesabstractTranscoding is a technique to convert one video bit-stream into another. While homogeneous transcoding is done at the same coding standard, inhomogeneous transcoding converts from one standard format to another standard. Inhomogeneous transcoding between MPEG-2, MPEG-4 or H.263 was performed using the same transform. With the standardisation of H.264 also a new transform basis and different block size was defined. For requantization from block size 8times8 to 4times4 this leads to the effect that the quantization error of one coefficient in a block of size 8times8 is distributed over multiple coefficients in blocks of size 4times4. In our work, we analyze the requantization process for inhomogeneous transcoding with different transforms. The deduced equations result in an expression for the correlation of the error contributions from the coefficients of block size 8times8 at each coefficient of block size 4times4. We then compare the mathematical analysis to simulations on real sequences. The reference to the requantization process is the direct quantization of the undistorted signal. It will be shown that the loss is as high as 3 dB PSNR at equivalent step size for input and output bitstream. Also an equation for the choice of the second quantization step size in dependency of the requantization loss is deduced. The model is then extended from the DCT to the integer-based transform as defined in H.264 Jens Bialkowski, Marcus Barkowsky, André Kaup |
MMSP | 3 |
| 2004 | Spatial error concealment of corrupted image data using frequency selective extrapolationabstractThe paper introduces a method for spatial error concealment of lost image data in erroneous image transmission. The image content of the correctly received surrounding blocks is successively approximated by a weighted linear combination of basis functions and the missing block is obtained by extrapolation. An implementation in the frequency domain allows an efficient realization. Investigations show that particularly 2D DFT basis functions are suited for signal extrapolation in order to be able to reconstruct both monotone areas and edges. Katrin Meisinger, André Kaup |
ICASSP (3) | 2 |
| 2004 | Fast transcoding of intra frames between H.263 and H.264
Jens Bialkowski, André Kaup, Klaus Illgner |
ICIP | 2 |
| 2004 | Minimizing a weighted error criterion for spatial error concealment of missing image dataabstractIn this contribution we present an algorithm for spatial error concealment of lost image data caused by transmission of images in error prone environments. The surrounding correctly received image signal is approximated by a weighted linear combination of basis functions and the missing image data is obtained by a frequency selective extrapolation. During the approximation a novel weighted error criterion is minimized. We use an isotropic correlation model for the weighting function taking the correlation among pixels into account and emphasizing pixels which are closer to the missing area. 2D DFT basis functions are especially suited for the signal extrapolation in order to be able to reconstruct monotone areas, edges and noisy regions and allow an efficient realization of the algorithm in the frequency domain. Due to the weighting function we could improve the concealing performance of the algorithm considerably while halving the required FFT size. Katrin Meisinger, André Kaup |
ICIP | 2 |
| 2004 | Temporal video segmentation using global motion estimation and discrete curve evolutionabstractThe identification of syntactic or semantic temporal segments is an important process of video-content analysis. The paper proposes a temporal video segmentation method based on global motion in order to analyze meaningful temporal substructures of camera shots. To ensure that the detected segment optimally contributes to the shot global characteristic, the proposed method exploits a state-of-the-art discrete curve evolution. This technique leads to a subdivision of the global motion trajectory, where each segment of the subdivision has a constant relevant global motion. Experimental results based on standard test sequences acknowledge the method functionality, especially for shots characterized by pronounced camera motion. Siripong Treetasanatavorn, Jörg Heuer, Uwe Rauschenbach, Klaus Illgner, André Kaup |
ICIP | 5 |
| 2002 | An MPEG-7 tool for compression and streaming of XML dataabstractIn the course of work on the MPEG-7 standard, a binary format with special features for the encoding of XML data was required. These required key features are a high data compression ratio, provision for streaming, dynamic update of the document structure and fast random access of data entities in the compressed stream. To support these features, we propose a novel, schema-aware approach which exploits the knowledge of the standardized MPEG-7 syntax definition of the encoded XML document on the encoder and decoder side. The technique is part of the MPEG-7 standard. This paper gives an overview of the coding algorithm, including a comparison to standard (XML) compression tools. Ulrich Niedermeier, Jörg Heuer, Andreas Hutter, Walter Stechele, André Kaup |
ICME (1) | 5 |
| 2000 | Error concealment for SNR scalable video coding in wireless communication
André Kaup |
VCIP | 1 |
| 2000 | Image restoration for frame- and object-based video coding using an adaptive constrained least-squares approach
André Kaup |
Signal Process. | 1 |
| 1999 | Global motion estimation in image sequences using robust motion vector field segmentationabstractIn this paper we propose an algorithm for the purpose of video indexing which estimates the global motion caused by camera movement. Since motion estimation is a computationally expensive process, we are using information already computed in block based video encoders such as the displacement vectors of the blocks and the computed SAD values of the blockmatching process. With this approach we could significantly speed up the determination of the parameters of a chosen global motion model. The accuracy of the algorithm is improved further by introducing a reliability measure. To achieve a robust computation with respect to object motion, an effective and straightforward segmentation of this vector field is implemented. Finally, the modeling of the global motion is evaluated using MPEG-7 test data. Jörg Heuer, André Kaup |
ACM Multimedia (1) | 2 |
| 1999 | Object-based texture coding of moving video in MPEG-4abstractThis paper describes some of the most promising segment-based coding techniques which have been investigated in the course of the MPEG-4 standardization process. Padding methods aim at extending arbitrarily shaped image segments to a regular block grid such that common hybrid block-based coding techniques can be applied. A simple and efficient padding technique employing low-pass extrapolation is outlined which yields a signal extension with high energy concentration in the low-frequency area. Simulations indicate that this method is well suited for block-based video coding, and clearly outperforms other low-complexity extrapolation methods with respect to coding efficiency. In contrast to padding techniques, shape-adaptive methods take advantage of the shape information available at the decoder side. A well-known representative of this class is the SA-DCT. However, having been primarily designed for intraframe coding, it is shown that the transform is suboptimal when applied to interframe coding. Using a suitable covariance model, it is demonstrated that a rescaled, orthonormalized transform much closer approximates the optimal shape-adaptive eigentransform of motion-compensated frame difference images. Rate distortion curves verify that orthonormalization improves coding efficiency in interframe coding by up to 2 dB while not adding to complexity. In a comparison, it is finally shown that extrapolation and SA-DCT perform very closely in the case of low data rates, while there is a clear advantage for the shape-adaptive transform in the case of high-quality video coding. André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Coding of segmented images using shape-independent basis functionsabstractAn important issue in object-based image coding is the efficient description of image segments having an arbitrary shape. This paper outlines a solution to this task using basic principles of transform image coding that are generalized for the case of arbitrarily shaped image segments. The texture inside each region is successively approximated using two-dimensional (2-D) shape-independent basis functions defined on a rectangle circumscribing the given image segment. The resulting texture description exhibits a high energy compactness and is well suited for low bit rate image coding. Unlike other approaches aimed at segment-oriented image coding, the proposed concept does not couple texture description with contour coding. The computational load is kept low, especially for the decoding process. André Kaup, Til Aach |
IEEE Trans. Image Process. | 1 |
| 1997 | Adaptive constrained least squares restoration for removal of blocking artifacts in low bit rate video codingabstractFor high compression ratios current video coding standards produce noticeable blocking and ringing noise due to a rigid block structure and coarse quantization. We propose a new method for reduction of these coding artifacts based on spatially adaptive constrained least squares restoration. The proposal is numerically simple and yields visually convincing results for intra as well as inter coded images. As a post-processing technique it is compatible to all existing image and video coding standards. André Kaup |
ICASSP | 1 |
| 1995 | On texture analysis: Local energy transforms versus quadrature filters
Til Aach, André Kaup, Rudolf Mester |
Signal Process. | 2 |
| 1995 | Models for region-based image representation
André Kaup |
Signal Process. | 1 |
| 1995 | Bayesian algorithms for adaptive change detection in image sequences using Markov random fields
Til Aach, André Kaup |
Signal Process. Image Commun. | 2 |
| 1994 | Efficient prediction of uncovered background in interframe coding using spatial extrapolationabstractThe focus of this paper is on prediction of uncovered background in moving video coding. A new algorithm for estimating the background texture of uncovered areas is proposed which exploits spatial correlation in addition to temporal correlation and motion information. The surrounding known background area is analysed and spatially extrapolated using a highly efficient spectral domain extrapolation algorithm. The method avoids inherent insufficiencies of conventional static background memories such as strong sensitivity to camera zoom and pan and it can be incorporated into any existing motion compensating coding concept. Several prediction results for typical moving image sequences demonstrate the applicability of the proposed concept.> André Kaup, Til Aach |
ICASSP (5) | 1 |
| 1994 | Disparity-based segmentation of stereoscopic foreground/background image sequencesabstractDescribes a method for displacement estimation in stereoscopic images, which is closely coupled with a segmentation of the pictures into homogeneously displaced regions. The technique is driven by a statistical optimization criterion which assesses the quality of the disparity estimate and of the segmentation, thus improving both of these simultaneously. In addition, the optimization criterion explicitly takes occluded areas into consideration. With the additional help of two constraints, this enables the algorithm to locate regions corresponding to occlusions accurately.> Til Aach, André Kaup |
IEEE Trans. Commun. | 2 |
| 1993 | Statistical model-based change detection in moving video
Til Aach, André Kaup, Rudolf Mester |
Signal Process. | 2 |
| 1990 | Combined displacement estimation and segmentation of stereo image pairs based on Gibbs random fieldsabstractA technique for displacement estimation in stereo image pairs, which is closely coupled with a segmentation of the images into homogeneously displaced regions, is described. An arbitrary initial displacement field is modified by a displacement vector relaxation, evaluating a statistical optimization criterion. Since the displacements are considered vector by vector during the relaxation, the method provides accurate locations of the boundaries between differently displaced regions. The proposed technique derives its power from the fact that the spatial array of displacements-and hence the resulting segmentation-is modeled by a Gibbs-Markov random field.> Til Aach, André Kaup, Rudolf Mester |
ICASSP | 2 |