EDBT 2026 Demo / reviewers in the wild / expert
Jürgen Seiler
dblp:45/46
· DBLP profile ↗
98ranked-venue papers
16as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 16 first-author · 26 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEE: An Efficient Feature-Based Framework for Performance Prediction in Semantic Segmentation
Felix Deichsel, Yongxu Ren, Philipp Beckerle, Jürgen Seiler, André Kaup |
ISCAS | 4 |
| 2025 | Domain Adaptation for Camera-Specific Image Characteristics using Shallow DiscriminatorsabstractEach image acquisition setup leads to its own camera-specific image characteristics degrading the image quality. In learning-based perception algorithms, characteristics occurring during the application phase, but absent in the training data, lead to a domain gap impeding the performance. Previously, pixel-level domain adaptation through unpaired learning of the pristine-to-distorted mapping function has been proposed. In this work, we propose shallow discriminator architectures to address limitations of these approaches. We show that a smaller receptive field size improves learning of unknown image distortions by more accurately reproducing local distortion characteristics at a low network complexity. In a domain adaptation setup for instance segmentation, we achieve mean average precision increases over previous methods of up to 0.15 for individual distortions and up to 0.16 for camera-specific image characteristics in a simplified camera model. In terms of number of parameters, our approach matches the complexity of one state of the art method while reducing complexity by a factor of 20 compared to another, demonstrating superior efficiency without compromising performance. Maximiliane Gruber, Jürgen Seiler, André Kaup |
VCIP | 2 |
| 2025 | Towards Object Segmentation Mask Selection Using Specular ReflectionsabstractSpecular reflections pose a significant challenge for object segmentation, as their sharp intensity transitions often mislead both conventional algorithms and deep learning based methods. However, as the specular reflection must lie on the surface of the object, this fact can be exploited to improve the segmentation masks. By identifying the largest region containing the reflection as the object, we derive a more accurate object mask without requiring specialized training data or model adaption. We evaluate our method on both synthetic and real world images and compare it against established and state-of-the-art techniques including Otsu thresholding, YOLO, and SAM2. Compared to the best performing baseline SAM2, our approach achieves up to 26.7% improvement in IoU, 22.3% in DSC, and 9.7% in pixel accuracy. Qualitative evaluations on real world images further confirm the robustness and generalizability of the proposed approach. Katja Kossira, Yunxuan Zhu, Jürgen Seiler, André Kaup |
VCIP | 3 |
| 2025 | Anti-Aliasing Snapshot HDR Imaging Using Non-Regular SensingabstractSnapshot HDR imaging is essential to capture the full dynamic range of a scene in a single exposure, making it essential for video and dynamic environments where motion prevents the use of multi-exposure techniques or complex hardware set-ups. This work presents a snapshot HDR imaging sensor that is based on spatially varying apertures, implemented by combining two differently sized prototype pixels. The different light integration areas physically extend the dynamic range towards the lower end, compared to a standard high resolution sensor. A non-regular pixel arrangement is suggested, to mitigate aliasing and overcome a loss in spatial resolution that is associated with increased light integration area of the larger prototype pixel. Subsequent reconstruction in the Fourier domain, where natural images can be sparsely represented allows to recover the image with high detail. The image acquisition approach with the proposed non-regular HDR sensor is simulated and analysed with special emphasis on the spatial resolution. The results suggest the snapshot HDR sensor layout to be an effective way to acquire images with high dynamic range and free from aliasing artefacts. Teresa Stürzenhofäcker, Moritz Klimm, Jürgen Seiler, André Kaup |
VCIP | 3 |
| 2025 | Multispectral Snapshot Image Registration Using Learned Cross Spectral Disparity Estimation and a Deep Guided Occlusion Reconstruction NetworkabstractMultispectral imaging aims at recording images in different spectral bands. This is extremely beneficial in diverse discrimination applications, for example in agriculture, recycling or healthcare. One approach for snapshot multispectral imaging, which is capable of recording multispectral videos, is by using camera arrays, where each camera records a different spectral band. Since the cameras are at different spatial positions, a registration procedure is necessary to map every camera to the same view. In this paper, we present a multispectral snapshot image registration with three novel components. First, a cross spectral disparity estimation network is introduced, which is trained on a popular stereo database using pseudo spectral data augmentation. Subsequently, this disparity estimation is used to accurately detect occlusions by warping the disparity map in a layer-wise manner. Finally, these detected occlusions are reconstructed by a learned deep guided neural network, which leverages the structure from other spectral components. It is shown that each element of this registration process as well as the final result is superior to the current state of the art. In terms of PSNR, our registration achieves an improvement of over 3 dB. At the same time, the runtime is decreased by a factor of over 3 on a CPU. Additionally, the registration is executable on a GPU, where the runtime can be decreased by a factor of 113. The source code and the data is available at https://github.com/FAU-LMS/MSIR. Frank Sippel, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 2 |
| 2024 | Color Agnostic Cross-Spectral Disparity EstimationabstractSince camera modules become more and more affordable, multi-spectral camera arrays have found their way from special applications to the mass market, e.g., in automotive systems, smartphones, or drones. Due to multiple modalities, the registration of different viewpoints and the required cross-spectral disparity estimation is up to the present extremely challenging. To overcome this problem, we introduce a novel spectral image synthesis in combination with a color agnostic transform. Thus, any recently published stereo matching network can be turned to a cross-spectral disparity estimator. Our novel algorithm requires only RGB stereo data to train a cross-spectral disparity estimator and a generalization from artificial training data to camera-captured images is obtained. The theoretical examination of the novel color agnostic method is completed by an extensive evaluation compared to state of the art including self-recorded multispectral data and a reference implementation. The novel color agnostic disparity estimation improves cross-spectral as well as conventional color stereo matching by reducing the average end-point error by 41 % for cross-spectral and by 22 % for mono-modal content, respectively. Frank Sippel, Nils Genser, Hannah Och, Jürgen Seiler, André Kaup |
ICASSP | 4 |
| 2024 | A Guided Upsampling Network for Short wave Infrared Images Using Graph RegularizationabstractExploiting the infrared area of the spectrum for classification problems is getting increasingly popular, because many materials have characteristic absorption bands in this area. However, sensors in the short wave infrared (SWIR) area and even higher wavelengths have a very low spatial resolution in comparison to classical cameras that operate in the visible wavelength area. Thus, in this paper an upsampling method for SWIR images guided by a visible image is presented. For that, the proposed guided upsampling network (GUNet) uses a graph-regularized optimization problem based on learned affinities is presented. The evaluation is based on a novel synthetic near-field visible-SWIR stereo database. Different guided upsampling methods are evaluated, which shows an improvement of nearly 1 dB on this database for the proposed upsampling method in comparison to the second best guided upsampling network. Furthermore, a visual example of an upsampled SWIR image of a real-world scene is depicted for showing real-world applicability. Frank Sippel, Jürgen Seiler, André Kaup |
ICASSP | 2 |
| 2024 | Conditional Optimal Filter Selection For Multispectral Object ClassificationabstractCapturing images using multispectral camera arrays has gained importance in medical, agricultural and environmental processes. However, using all available spectral bands is infeasible and produces much data, while only a fraction is needed for a given task. Nearby bands may contain similar information, therefore redundant spectral bands should not be considered in the evaluation process to keep complexity and the data load low. In current methods, a restricted and pre-determined number of spectral bands is selected. Our approach improves this procedure by including preset conditions such as noise or the bandwidth of available filters, minimizing spectral redundancy. Furthermore, a minimal filter selection can be conducted, keeping the hardware setup at low costs, while still obtaining all important spectral information. In comparison to the fast binary search filter band selection method, we managed to reduce the amount of misclassified objects of the SMM dataset from 318 to 124 using a random forest classifier. Katja Kossira, David Schön, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2024 | A Study on the Effect of Color Spaces in Learned Image CompressionabstractIn this work, we present a comparison between color spaces namely YUV, LAB, RGB and their effect on learned image compression. For this we use the structure and color based learned image codec (SLIC) from our prior work, which consists of two branches - one for the luminance component (Y or L) and another for chrominance components (UV or AB). However, for the RGB variant we input all 3 channels in a single branch, similar to most learned image codecs operating in RGB. The models are trained for multiple bitrate configurations in each color space. We report the findings from our experiments by evaluating them on various datasets and compare the results to state-of-the-art image codecs. The YUV model performs better than the LAB variant in terms of MSSSIM with a Bjøntegaard delta bitrate (BD-BR) gain of 7.5% using VTM intra-coding mode as the baseline. Whereas the LAB variant has a better performance than YUV model in terms of CIEDE2000 having a BD-BR gain of 8%. Overall, the RGB variant of SLIC achieves the best performance with a BD-BR gain of 13.14% in terms of MS-SSIM and a gain of 17.96% in CIEDE2000 at the cost of a higher model complexity. Srivatsa Prativadibhayankaram, Mahadev Prasad Panda, Jürgen Seiler, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel, André Kaup |
ICIP | 3 |
| 2024 | Fast Edge-Aware Occlusion Detection In The Context of Multispectral Camera ArraysabstractMultispectral imaging is very beneficial in diverse applications, like healthcare and agriculture, since it can capture absorption bands of molecules in different spectral areas. A promising approach for multispectral snapshot imaging are camera arrays. Image processing is necessary to warp all different views to the same view to retrieve a consistent multispectral datacube. This process is also called multispectral image registration. After a cross spectral disparity estimation, an occlusion detection is required to find the pixels that were not recorded by the peripheral cameras. In this paper, a novel fast edge-aware occlusion detection is presented, which is shown to reduce the runtime by at least a factor of 12. Moreover, an evaluation on ground truth data reveals better performance in terms of precision and recall. Finally, the quality of a final multispectral datacube can be improved by more than 1.5 dB in terms of PSNR as well as in terms of SSIM in an existing multispectral registration pipeline. The source code is available at https://github.com/FAU-LMS/fast-occlusion-detection. Frank Sippel, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2024 | Inter-Camera Color Correction for Multispectral Imaging with Camera Arrays Using a Consensus ImageabstractThis paper introduces a novel method for inter-camera color calibration for multispectral imaging with camera arrays using a consensus image. Capturing images using multispectral camera arrays has gained importance in medical, agricultural, and environmental processes. Due to fabrication differences, noise, or device altering, varying pixel sensitivities occur, influencing classification processes. Therefore, color calibration between the cameras is necessary. In existing methods, one of the camera images is chosen and considered as a reference, ignoring the color information of all other recordings. Our new approach does not just take one image as reference, but uses statistical information such as the location parameter to generate a consensus image as basis for calibration. This way, we managed to improve the PSNR values for the linear regression color correction algorithm by 1.15 dB and the improved color difference (iCID) values by 2.81. Katja Kossira, Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2024 | Adaptive Variance-Threshold-Based Skip Modes for Learned Video Compression Using a Motion Complexity CriterionabstractSkip modes are a powerful tool to reduce the rate in video compression. The main idea is that the residual areas where the prediction performs well are not transmitted since the prediction quality is good enough that the prediction signal itself can be used as the reconstruction signal. This is commonly used, e.g., in the compression standard VVC, where a skip flag can be transmitted for inter blocks under certain conditions. The skipped residual block is then not transmitted and the content is instead inferred to be zero at the decoder. Current learning-based methods use different kinds of skip modes. One possibility here arises from the fact that the coders estimate and transmit the variance for each transmitted symbol. It has been proposed to use this estimated variance to derive a skip mode. When the variance falls below a threshold, the symbol is not transmitted. In this paper we propose an extension to this method. By classifying each position in the latent space according to the local motion complexity, we can transmit adaptive thresholds for each class. That way, we can employ motion information to refine the granularity of the skip mode. When we implement this method in FVC, we are able to save up 2.11% rate on a GOP 20 sequence. We also discuss the behavior of increasingly adaptive skip modes in scenarios with larger GOP size, where error-propagation becomes a larger issue. Fabian Brand, Jürgen Seiler, Johannes Sauer, Elena Alshina, André Kaup |
PCS | 2 |
| 2024 | Networked Systems Diagnostics: A Fusion of Failure Mode and Effects Analysis and a Delphi Expert StudyabstractNetworked devices, especially those comprising multiple identical devices, are extensively utilized in industrial scenarios. However, their complexity poses unique challenges in diagnostic processes, demanding efficient methodologies to identify and assess risks. The application of Failure Mode and Effects Analysis (FMEA) for analyzing complex systems, especially those consisting of networked devices, appears to be limited, particularly in identifying critical risk factors. In this paper, we propose a novel pipeline to diagnose networked systems by fusing FMEA with a Delphi (expert) Study. Our approach leverages the collective knowledge of a group of experts through a structured Delphi Study enabling them to contribute individually but also to interact. We demonstrate the applicability of our approach through a case study involving a system of networked mobile robots. Based on a Fault Tree Analysis (FTA), Risk Priority Numbers (RPN) of minimal fault tree cut sets are calculated to identify the most critical mechatronic failures. Our findings show that our methodology provides an RPN ranking that closely aligns with expert insights, highlighting its efficacy in accurately assessing risk in complex networked systems. Yongxu Ren, Felix Deichsel, Valentin Hopf, Jürgen Seiler, André Kaup, Philipp Beckerle |
SMC | 4 |
| 2024 | Conditional Residual Coding: A Remedy for Bottleneck Problems in Conditional Inter Frame CodingabstractConditional coding is a new video coding paradigm enabled by neural-network-based compression. It can be shown that conditional coding is in theory better than the traditional residual coding, which is widely used in video compression standards like HEVC or VVC. However, on closer inspection, it becomes clear that conditional coders can suffer from information bottlenecks in the prediction path, i.e., that due to the data processing inequality not all information from the prediction signal can be passed to the reconstructed signal, thereby impairing the coder performance. In this paper we propose the conditional residual coding concept, which we derive from information theoretical properties of the conditional coder. This coder significantly reduces the influence of bottlenecks, while maintaining the theoretical performance of the conditional coder. We provide a theoretical analysis of the coding paradigm and demonstrate the performance of the conditional residual coder in a practical example. We show that conditional residual coders alleviate the disadvantages of conditional coders while being able to maintain their advantages over residual coders. In the spectrum of residual and conditional coding, we can therefore consider them as “the best from both worlds”. Fabian Brand, Jürgen Seiler, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Image Super-Resolution Using T-Tetromino PixelsabstractFor modern high-resolution imaging sensors, pixel binning is performed in low-lighting conditions and in case high frame rates are required. To recover the original spatial resolution, single-image super-resolution techniques can be applied for upscaling. To achieve a higher image quality after upscaling, we propose a novel binning concept using tetromino-shaped pixels. It is embedded into the field of compressed sensing and the coherence is calculated to motivate the sensor layouts used. Next, we investigate the reconstruction quality using tetromino pixels for the first time in literature. Instead of using different types of tetrominoes as proposed elsewhere, we show that using a small repeating cell consisting of only four T-tetrominoes is sufficient. For reconstruction, we use a locally fully connected reconstruction (LFCR) network as well as two classical reconstruction methods from the field of compressed sensing. Using the LFCR network in combination with the proposed tetromino layout, we achieve superior image quality in terms of PSNR, SSIM, and visually compared to conventional single-image super-resolution using the very deep super-resolution (VDSR) network. For PSNR, a gain of up to +1.92 dB is achieved. Simon Grosche, Andy Regensky, Jürgen Seiler, André Kaup |
CVPR | 3 |
| 2023 | Cross Spectral Image Reconstruction Using a Deep Guided Neural NetworkabstractCross spectral camera arrays, where each camera records different spectral content, are becoming increasingly popular for RGB, multispectral and hyperspectral imaging, since they are capable of a high resolution in every dimension using off-the-shelf hardware. For these, it is necessary to build an image processing pipeline to calculate a consistent image data cube, i.e., it should look like as if every camera records the scene from the center camera. Since the cameras record the scene from a different angle, this pipeline needs a reconstruction component for pixels that are not visible to peripheral cameras. For that, a novel deep guided neural network (DGNet) is presented. Since only little cross spectral data is available for training, this neural network is highly regularized. Furthermore, a new data augmentation process is introduced to generate the cross spectral content. On synthetic and real multispectral camera array data, the proposed network out-performs the state of the art by up to 2 dB in terms of PSNR on average. Besides, DGNet also tops its best competitor in terms of SSIM as well as in runtime by a factor of nearly 12. Moreover, a qualitative evaluation reveals visually more appealing results for real camera array data. Frank Sippel, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2023 | End-to-End Lidar-Camera Self-Calibration for Autonomous VehiclesabstractAutonomous vehicles are equipped with a multi-modal sensor setup to enable the car to drive safely. The initial calibration of such perception sensors is a highly matured topic and is routinely done in an automated factory environment. However, an intriguing question arises on how to maintain the calibration quality throughout the vehicle’s operating duration. Another challenge is to calibrate multiple sensors jointly to ensure no propagation of systemic errors. In this paper, we propose Camera Lidar Calibration Network (CaLiCaNet), an end-to-end deep self-calibration network which addresses the automatic calibration problem for pinhole camera and Lidar. We jointly predict the camera intrinsic parameters (focal length and distortion) as well as Lidar-Camera extrinsic parameters (rotation and translation), by regressing feature correlation between the camera image and the Lidar point cloud. The network is arranged in a Siamese-twin structure to constrain the network features learning to a mutually shared feature in both point cloud and camera (Lidar-camera constraint). Evaluation using KITTI datasets shows that we achieve 0.154° and 0.059 m accuracy with a reprojection error of 0.028 pixel with a single-pass inference. We also provide an ablative study of how our end-to-end learning architecture offers lower terminal loss (21% decrease in rotation loss) compared to isolated calibration. Arya Rachman, Jürgen Seiler, André Kaup |
IV | 2 |
| 2022 | Reliability Scoring for the Recognition of Degraded License Plates*abstractCriminal investigations oftentimes need the identification of license plates of escape vehicles. The vehicles may be recorded by low-quality cameras in the wild. Their license plates may be unreadable for police officers. Recent efforts aim to use machine learning to forensically decipher license plates from such low-quality images. These methods operate near the information-theoretic limit of recognition and hence show quite high error rates. Unfortunately, it is unclear when such prediction errors occur, which makes it difficult to use these methods in practice. In this work, we propose a Bayesian Neural Network to inherently incorporate a reliability measure into the classifier. We additionally propose to integrate multiple estimations with an entropy weight to further improve the reliability. Our experiments show that this uncertainty metric dramatically reduces the number of false predictions while preserving most of the true predictions. Anatol Maier, Denise Moussa, Andreas Spruck, Jürgen Seiler, Christian Riess |
AVSS | 4 |
| 2022 | P-Frame Coding with Generalized Difference: A Novel Conditional Coding ApproachabstractMotion compensated inter frame prediction is a common component of all video coders and greatly reduces temporal redundancy. With the rise of deep learning-based image and video compression, this concept has been successfully taken over from traditional coding approaches. These approaches offer a larger flexibility than traditional transform coding and therefore enable efficient conditional coding. In this work, we develop a novel conditional coding approach based on the generalized difference and generalized sum operators. This approach is a special case of a general conditional coder and has a very small complexity overhead. We also propose an extension which enables dynamic content-adaptive switching between conditional and residual coding. We show that the extended generalized difference coding outperforms both residual and conditional coding, saving 27.8% Bjøntegaard delta rate compared to the former. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2022 | Domain Adaptation for Unknown Image Distortions in Instance SegmentationabstractData-driven techniques for machine vision heavily depend on the training data to sufficiently resemble the data occurring during test and application. However, in practice unknown distortion can lead to a domain gap between training and test data, impeding the performance of a machine vision system. With our proposed approach this domain gap can be closed by unpaired learning of the pristine-to-distortion mapping function of the unknown distortion. This learned mapping function may then be used to emulate the unknown distortion in the training data. Employing a fixed setup, our approach is independent from prior knowledge of the distortion. Within this work, we show that we can effectively learn unknown distortions at arbitrary strengths. When applying our approach to instance segmentation in an autonomous driving scenario, we achieve results comparable to an oracle with knowledge of the distortion. An average gain in mean Average Precision (mAP) of up to 0.19 can be achieved. Maximiliane Gruber, Fabian Brand, Alina Mosebach, Jürgen Seiler, André Kaup |
ICIP | 4 |
| 2022 | Forensic License Plate Recognition with Compression-Informed TransformersabstractForensic license plate recognition (FLPR) remains an open challenge in legal contexts such as criminal investigations, where unreadable license plates (LPs) need to be deciphered from highly compressed and/or low resolution footage, e.g., from surveillance cameras. In this work, we propose a side-informed Transformer architecture that embeds knowledge on the input compression level to improve recognition under strong compression. We show the effectiveness of Transformers for license plate recognition (LPR) on a low-quality real-world dataset. We also provide a synthetic dataset that includes strongly degraded, illegible LP images and analyze the impact of knowledge embedding on it. The network outperforms existing FLPR methods and standard state-of-the art image recognition models while requiring less parameters. For the severest degraded images, we can improve recognition by up to 8.9 percent points.1 Denise Moussa, Anatol Maier, Andreas Spruck, Jürgen Seiler, Christian Riess |
ICIP | 4 |
| 2022 | Camera Self-Calibration: Deep Learning from Driving ScenesabstractPrior to driving, cameras embedded in an autonomous driving system need to be calibrated intrinsically. Calibration is crucial to ensure that safety-related perception functions can reliably perceive the environment. Vehicle cameras are also exposed to mechanical perturbations requiring periodic re-calibration with regular uses. The current widely-accepted calibration approaches are based on robust but potentially demanding target-based methods. Such methods require a car to be taken offline and rely on static infrastructure and operators. Targetless online calibration approaches exist but remain largely unadopted due to the accuracy gaps compared to the classical methods. We propose a deep-learning-based self-calibration strategy for the vehicular camera that learns from driving scenes—they make an inherently large-scale dataset—and is validated back-to-back against checkerboard reprojection error. Our approach results in a 2.5% decrease in subpixel reprojection error compared to the existing deep-learning-based approaches. We also demonstrate its practical application in the automotive domain. Arya Rachman, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2022 | Optimal Filter Selection for Multispectral Object Classification Using Fast Binary SearchabstractWhen designing multispectral imaging systems for classifying different spectra it is necessary to choose a small number of filters from a set with several hundred different ones. Tackling this problem by full search leads to a tremendous number of possibilities to check and is NP-hard. In this paper we introduce a novel fast binary search for optimal filter selection that guarantees a minimum distance metric between the different spectra to classify. In our experiments, this procedure reaches the same optimal solution as with full search at much lower complexity. The desired number of filters influences the full search in factorial order while the fast binary search stays constant. Thus, fast binary search allows to find the optimal solution of all combinations in an adequate amount of time and avoids prevailing heuristics. Moreover, our fast binary search algorithm outperforms other filter selection techniques in terms of misclassified spectra in a real-world classification problem. Frank Sippel, Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2022 | On Benefits and Challenges of Conditional Interframe Video Coding in Light of Information TheoryabstractThe rise of variational autoencoders for image and video compression has opened the door to many elaborate coding techniques. One example here is the possibility of conditional interframe coding. Here, instead of transmitting the residual between the original frame and the predicted frame (often obtained by motion compensation), the current frame is transmitted under the condition of knowing the prediction signal. In practice, conditional coding can be straightforwardly implemented using a conditional autoencoder, which has also shown good results in recent works. In this paper, we provide an information theoretical analysis of conditional coding for inter frames and show in which cases gains compared to traditional residual coding can be expected. We also show the effect of information bottlenecks which can occur in practical video coders in the prediction signal path due to the network structure, as a consequence of the data-processing theorem or due to quantization. We demonstrate that conditional coding has theoretical benefits over residual coding but that there are cases in which the benefits are quickly canceled by small information bottlenecks of the prediction signal. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 2 |
| 2021 | Intra To Inter: Towards Intra Prediction for Learning-Based Video Coders Using Optical FlowabstractTraditional video coders often rely on a block structure for transmission. Here each block is coded separately and sequentially and for each block the encoder can decide whether to use intra or inter prediction. This way, inter and intra prediction can be mixed within a single frame. This has advantages when new areas are uncovered, which were not present in the reference frame, and can hence not be predicted well. These areas are typically predicted using intra prediction. Currently much research goes into end-to-end-trained video coders which do not operate on a block level and typically use dense motion fields for inter prediction. There it is more difficult to incorporate intra prediction for uncovered regions. In this paper we propose a novel concept which enables us to reinterpret classical angular intra prediction in a way that we can transmit it as part of the dense motion field. We can save an average of 18% rate for the transmission of the motion vectors for the same quality of the prediction image. Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2021 | Novel Consistency Check For Fast Recursive Reconstruction Of Non-Regularly Sampled Video DataabstractQuarter sampling is a novel sensor design that allows for an acquisition of higher resolution images without increasing the number of pixels. When being used for video data, one out of four pixels is measured in each frame. Effectively, this leads to a non-regular spatio-temporal sub-sampling. Compared to purely spatial or temporal sub-sampling, this allows for an increased reconstruction quality, as aliasing artifacts can be reduced. For the fast reconstruction of such sensor data with a fixed mask, recursive variant of frequency selective reconstruction (FSR) was proposed. Here, pixels measured in previous frames are projected into the current frame to support its reconstruction. In doing so, the motion between the frames is computed using template matching. Since some of the motion vectors may be erroneous, it is important to perform a proper consistency checking. In this paper, we propose faster consistency checking methods as well as a novel recursive FSR that uses the projected pixels different than in literature and can handle dynamic masks. Altogether, we are able to significantly increase the reconstruction quality by + 1.01 dB compared to the state-of-the-art recursive reconstruction method using a fixed mask. Compared to a single frame reconstruction, an average gain of about + 1.52 dB is achieved for dynamic masks. At the same time, the computational complexity of the consistency checks is reduced by a factor of 13 compared to the literature algorithm. Simon Grosche, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2021 | Hyperspectral Image Reconstruction from Multispectral Images Using Non-Local FilteringabstractUsing light spectra is an essential element in many applications, for example, in material classification. Often this information is acquired by using a hyperspectral camera. Unfortunately, these cameras have some major disadvantages like not being able to record videos. Therefore, multispectral cameras with wide-band filters are used, which are much cheaper and are often able to capture videos. However, using multispectral cameras requires an additional reconstruction step to yield spectral information. Usually, this reconstruction step has to be done in the presence of imaging noise, which degrades the reconstructed spectra severely. Typically, same or similar pixels are found across the image with the advantage of having independent noise. In contrast to state-of-the-art spectral reconstruction methods which only exploit neighboring pixels by block-based processing, this paper introduces non-local filtering in spectral reconstruction. First, a block-matching procedure finds similar non-local multispectral blocks. Thereafter, the hyperspectral pixels are reconstructed by filtering the matched multispectral pixels collaboratively using a reconstruction Wiener filter. The proposed novel procedure even works under very strong noise. The method is able to lower the spectral angle up to 18% and increase the peak signal-to-noise-ratio up to 1.1dB in noisy scenarios compared to state-of-the-art methods. Moreover, the visual results are much more appealing. Frank Sippel, Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2021 | Switchable Motion Models for Non-Block-Based Inter Prediction in Learning-Based Video CodingabstractMost state-of-the-art video coders rely on a block structure. For inter-frame prediction, motion vectors are transmitted per block. For example in VVC, the coder can choose between a translational or an affine motion model on a block level, depending on the content. In non-block-based coding, which is on the rise since the development of end-to-end learning based image compression, the motion vectors have to be transmitted differently. Due to the missing inherent block structure, switching between different motion models presents a challenge, but also an opportunity. In this paper, we propose an alternative approach to efficiently signal additional information regarding the motion model to improve the quality of the motion compensated image. Using our methods, we are able to increase the quality of the prediction image in our scenario by 0.40 dB on average and by up to 0.88 dB for sequences with strong and complex motion at the same rate. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 2 |
| 2021 | Spatio-spectral Image Reconstruction Using Non-local FilteringabstractIn many image processing tasks it occurs that pixels or blocks of pixels are missing or lost in only some channels. For example during defective transmissions of RGB images, it may happen that one or more blocks in one color channel are lost. Nearly all modern applications in image processing and transmission use at least three color channels, some of the applications employ even more bands, for example in the infrared and ultraviolet area of the light spectrum. Typically, only some pixels and blocks in a subset of color channels are distorted. Thus, other channels can be used to reconstruct the missing pixels, which is called spatio-spectral reconstruction. Current state-of-the-art methods purely rely on the local neighborhood, which works well for homogeneous regions. However, in high-frequency regions like edges or textures, these methods fail to properly model the relationship between color bands. Hence, this paper introduces non-local filtering for building a linear regression model that describes the inter-band relationship and is used to reconstruct the missing pixels. Our novel method is able to increase the PSNR on average by 2 dB and yields visually much more appealing images in high-frequency regions. Frank Sippel, Jürgen Seiler, André Kaup |
VCIP | 2 |
| 2020 | Joint Content-Adaptive Dictionary Learning And Sparse Selective Extrapolation For Cross-Spectral Image ReconstructionabstractNumerous applications deal with distorted images, e.g., during transmission over lossy channels in image coding, or to reconstruct occlusions in multi-color multi-view imaging scenarios. In many cases, not all spectral channels are distorted, or the losses distribute differently between the channels. Recently, efforts were made to develop reference guided approaches to reconstruct distorted spectral content. However, these methods are only able to exploit information from a single reference, even if there are multiple spectral components available that could be used for reconstruction. To overcome this limitation, a novel method is proposed in this paper, which introduces a content-adaptive dictionary learning approach that is applicable to an arbitrary number of references. With the novel approach, an average PSNR gain of approx. 1.5 dB is achieved in comparison to the best recently published state-of-the-art methods for a single reference component. Moreover, the presence of multiple guidance channels can enhance the reconstruction by another 6 dB on average. At the same time, the complexity of the novel algorithm is significantly lower than for previously published methods. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2020 | Deep Learning Based Cross-Spectral Disparity Estimation For Stereo ImagingabstractRecently, cross-spectral stereo-camera setups found their way from special applications to mass market, especially in smartphones, automotive systems, or drones. In the following, a novel concept is introduced to bring stereo cameras and cross-spectral disparity estimation together. So far, either monomodal stereo algorithms exist that are not suitable for cross-spectral image registration, or structural template matching is applied that achieves a low quality. To overcome these limitations, a technique is proposed to synthesize arbitrary spectral components from widely available color stereo databases, and to retrain mono-modal deep learning methods. In this contribution, the estimation of spectral bands based on random processes is shown together with noise models, which also allow for a robust registration of narrowband components. The theoretical examination is completed by an extensive evaluation, including a self-manufactured cross-spectral camera setup. In comparison to state-of-the-art techniques, the end-point error is on average reduced by a factor of seven. Nils Genser, Andreas Spruck, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2020 | Enhanced Image Reconstruction From Quarter Sampling Measurements Using An Adapted Very Deep Super Resolution NetworkabstractQuarter sampling is a novel sensor concept that enables the acquisition of higher resolution images without increasing the number of pixels. This is achieved by covering three quarters of each pixel of a low-resolution sensor such that only one quadrant of the sensor area of each pixel is sensitive to light. By randomly masking different parts, effectively a non-regular sampling of a higher resolution image is performed. Combining a properly designed mask and a high-quality reconstruction algorithm, a higher image quality can be achieved than using a low-resolution sensor and subsequent upsampling. For the latter case, the image quality can be enhanced using super resolution algorithms. Recently, algorithms based on machine learning such as the Very Deep Super Resolution network (VDSR) proofed to be successful for this task. In this work, we transfer the concepts of VDSR to the special case of quarter sampling. Besides adapting the network layout to take advantage of the case of quarter sampling, we introduce a novel data augmentation technique enabled by quarter sampling. Altogether, using the quarter sampling sensor, the image quality in terms of PSNR can be increased by + 0.67 dB for the Urban 100 dataset compared to using a low-resolution sensor with VDSR. Simon Grosche, Kristian Fischer 0001, Fabian Brand, Jürgen Seiler, André Kaup |
ICIP | 4 |
| 2020 | A Triangulation-Based Backward Adaptive Motion Field Subsampling SchemeabstractOptical flow procedures are used to generate dense motion fields which approximate true motion. Such fields contain a large amount of data and if we need to transmit such a field, the raw data usually exceeds the raw data of the two images it was computed from. In many scenarios, however, it is of interest to transmit a dense motion field efficiently. Most prominently this is the case in inter prediction for video coding. In this paper we propose a transmission scheme based on subsampling the motion field. Since a field which was subsampled with a regularly spaced pattern usually yields suboptimal results, we propose an adaptive subsampling algorithm that preferably samples vectors at positions where changes in motion occur. The subsampling pattern is fully reconstructable without the need for signaling of position information. We show an average gain of 2.95 dB in average end point error compared to regular subsampling. Furthermore we show that an additional prediction stage can improve the results by an additional 0.43 dB, gaining 3.38 dB in total. Fabian Brand, Jürgen Seiler, Elena Alshina, André Kaup |
MMSP | 2 |
| 2020 | Real-Time Frequency Selective Reconstruction through Register-Based Argmax CalculationabstractFrequency Selective Reconstruction (FSR) is a state-of-the-art algorithm for solving diverse image reconstruction tasks, where a subset of pixel values in the image is missing. However, it entails a high computational complexity due to its iterative, blockwise procedure to reconstruct the missing pixel values. Although the complexity of FSR can be considerably decreased by performing its computations in the frequency domain, the reconstruction procedure still takes multiple seconds up to multiple minutes depending on the parameterization. However, FSR has the potential for a massive parallelization greatly improving its reconstruction time. In this paper, we introduce a novel highly parallelized formulation of FSR adapted to the capabilities of modern GPUs and propose a considerably accelerated calculation of the inherent argmax calculation. Altogether, we achieve a 100-fold speed-up, which enables the usage of FSR for real-time applications. Andy Regensky, Simon Grosche, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2020 | Introducing Latent Space Correlation to Conditional Autoencoders for Intra PredictionabstractIntra prediction has been an integral part of image and video coders for a long time. A predominant method is angular prediction that extends the reference area in a certain angle into the block. Recently many deep-learning-based methods have been proposed. Since intra prediction uses multiple modes this usually requires training a large number of networks. With a conditional autoencoder we are able to generate an arbitrary number of modes with only one network. In this paper we introduce a novel loss function enforcing a spatially correlated latent space and extend the network structure to the same end. Thereby we are able to propose a simple spatial mode prediction scheme using most-probable-mode lists. By replacing matrix-based intra prediction in VVC with our method, we obtain average rate savings of 0.84% with peak gains of 2.37%. Fabian Brand, Jürgen Seiler, André Kaup |
VCIP | 2 |
| 2020 | Camera Array for Multi-Spectral ImagingabstractRecently, many new applications arose for multispectral and hyper-spectral imaging. Besides modern biometric systems for identity verification, also agricultural and medical applications came up, which measure the health condition of plants and humans. Despite the growing demand, the acquisition of multi-spectral data is up to the present complicated. Often, expensive, inflexible, or low resolution acquisition setups are only obtainable for specific professional applications. To overcome these limitations, a novel camera array for multi-spectral imaging is presented in this article for generating consistent multispectral videos. As differing spectral images are acquired at various viewpoints, a geometrically constrained multi-camera sensor layout is introduced, which enables the formulation of novel registration and reconstruction algorithms to globally set up robust models. On average, the novel acquisition approach achieves a gain of 2.5 dB PSNR compared to recently published multi-spectral filter array imaging systems. At the same time, the proposed acquisition system ensures not only a superior spatial, but also a high spectral, and temporal resolution, while filters are flexibly exchangeable by the user depending on the application. Moreover, depth information is generated, so that 3D imaging applications, e.g., for augmented or virtual reality, become possible. The proposed camera array for multi-spectral imaging can be set up using off-the-shelf hardware, which allows for a compact design and employment in, e.g., mobile devices or drones, while being cost-effective. Nils Genser, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 2 |
| 2020 | Boosting Compressed Sensing Using Local Measurements and Sliding Window ReconstructionabstractIn the framework of compressed sensing, image data is measured using less measurements than the total number of pixels. Each measurement consists of a (random) linear combination of all pixels. Since image data is approximately sparse in an appropriate transform domain, a reasonable reconstruction is possible for many measurement matrices, especially for i.i.d. Gaussian measurement matrices. In a seemingly different field, non-regular sampling techniques such as three-quarter sampling have shown promising results to enhance the resolution of an imaging sensor by effectively sub-sampling a higher resolution image. Here, the measurements can be described as linear combinations of only three pixels, which can also be seen as a (spectral) compressed sensing measurement. Since each measurement is spatially localized, the reconstruction can be performed in overlapping sliding windows. In this work, we show that compressed sensing reconstruction algorithms can greatly benefit from such an overlapping sliding window reconstruction. Compared to conventional block-wise compressed sensing with i.i.d. Gaussian measurement matrices, the reconstruction quality in terms of the PSNR increases up to +5dB using small, local i.d.d. Gaussian measurement blocks. Additionally, we propose a local joint sparse deconvolution and extrapolation (L-JSDE) to reconstruct images from arbitrary local measurements. For several applications with local measurements we show that L-JSDE increases the PSNR by +2.2dB relative to conventional block-wise i.i.d. Gaussian measurements reconstructed with the state-of-the-art reconstruction algorithm D-AMP using the same overall sampling density. Simon Grosche, Andy Regensky, Jürgen Seiler, André Kaup |
IEEE Trans. Image Process. | 3 |
| 2019 | Motion-adapted Three-dimensional Frequency Selective ExtrapolationabstractIt has been shown, that high resolution images can be acquired using a low resolution sensor with non-regular sampling. Therefore, post-processing is necessary. In terms of video data, not only the spatial neighborhood can be used to assist the reconstruction, but also the temporal neighbor-hood. A popular and well performing algorithm for this kind of problem is the three-dimensional frequency selective extrapolation (3D-FSE) for which a motion adapted version is introduced in this paper. This proposed extension solves the problem of changing content within the area considered by the 3D-FSE, which is caused by motion within the sequence. Because of this motion, it may happen that regions are emphasized during the reconstruction that are not present in the original signal within the considered area. By that, false content is introduced into the extrapolated sequence, which affects the resulting image quality negatively. The novel extension, presented in the following, incorporates motion data of the sequence in order to adapt the algorithm accordingly, and compensates changing content, resulting in gains of up to 1.75 dB compared to the existing 3D-FSE. Andreas Spruck, Markus Jonscher, Jürgen Seiler, André Kaup |
ICASSP | 3 |
| 2019 | Joint Regression Modeling and Sparse Spatial Refinement for High-Quality Reconstruction of Distorted Color ImagesabstractHigh quality algorithms are demanded to reconstruct distorted color images in a variety of applications. For example, distortions can result during transmission over lossy channels in image coding or in multi-view imaging scenarios. In general, not all color channels are equally affected and the losses distribute differently in-between channels. However, state-ofthe-art methods process color channels independently and do not take the cross color information into account. Thus, a novel and powerful reconstruction algorithm is formulated in this contribution that exploits color as well as spatial information. Therefore, an initial model is estimated for the distorted area using a reference channel. Then, its quality is estimated and a spatial weighting model is set-up. Afterwards, the initial inter channel prediction is refined by generating a sparse model that takes the spatial correlations into account, as well. Consequently, the proposed method achieves an outstanding quality compared to state-of-the-art methods. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2019 | Intra Frame Prediction for Video Coding Using a Conditional Autoencoder ApproachabstractIntra prediction is a vital component of most modern image and video codecs. State of the art video codecs like High Efficiency Video Coding (HEVC) or the upcoming Versatile Video Coding (VVC) use a high number of directional modes. With the recent advances in deep learning, it is now possible to use artificial neural networks for intra frame prediction. Previously published approaches usually add additional ANN based modes or replace all modes by training several networks. In our approach, we use a single autoencoder network to first compress the original with help of already transmitted pixels to four parameters. We then use the parameters together with this support area to generate a prediction for the block. This way, we are able to replace all angular intra modes by a single ANN. In the experiments we compare our method with the intra prediction method currently used in the VVC Test Model (VTM). Using our method, we are able to gain up to 0.85 dB prediction PSNR with a comparable amount of side information or reduce the amount of side information by 2 bit per prediction unit with similar PSNR. Fabian Brand, Jürgen Seiler, André Kaup |
PCS | 2 |
| 2019 | Dynamic Non-Regular Sampling Sensor Using Frequency Selective ReconstructionabstractBoth a high spatial and a high temporal resolution of images and videos are desirable in many applications, such as entertainment systems, monitoring manufacturing processes, or video surveillance. Due to the limited throughput of pixels per second, however, there is always a tradeoff between acquiring sequences with a high spatial resolution at a low temporal resolution or vice versa. In this paper, a modified sensor concept is proposed which is able to acquire both a high spatial and a high temporal resolution. This is achieved by dynamically reading out only a subset of pixels in a non-regular order to obtain a high temporal resolution. A full high spatial resolution is then obtained by performing a subsequent 3D reconstruction of the partially acquired frames. The main benefit of the proposed dynamic readout is that for each frame, different sampling points are available, which is advantageous since this information can significantly enhance the reconstruction quality of the proposed reconstruction algorithm. Using the proposed dynamic readout strategy, gains in the peak-signal-to-noise ratio (PSNR) of up to 8.55 dB are achieved compared with a static readout strategy. Compared with the other state-of-the-art techniques, such as frame rate up-conversion or super-resolution, which are also able to reconstruct sequences with both a high spatial and a high temporal resolution, average gains in PSNR of up to 6.58 dB are possible. Markus Jonscher, Jürgen Seiler, Daniela Lanz, Michael Schöberl, Michel Bätz, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Increasing Imaging Resolution by Non-Regular Sampling and Joint Sparse Deconvolution and ExtrapolationabstractIncreasing the resolution of image sensors has been a never ending struggle since many years. In this paper, we propose a novel image sensor layout, which allows for the acquisition of images at a higher resolution and improved quality. For this, the image sensor makes use of non-regular sampling, which reduces the impact of aliasing. Therewith, it allows for capturing details, which would not be possible with state-of-the-art sensors of the same number of pixels. The non-regular sampling is achieved by rotating prototype pixel cells in a non-regular fashion. As not the whole area of the pixel cell is sensitive to light, a non-regular spatial integration of the incident light is obtained. Based on the sensor output data, a high-resolution image can be reconstructed by performing a deconvolution with respect to the integration area and an extrapolation of the information to the insensitive regions of the pixels. To solve this challenging task, we introduce a novel joint sparse deconvolution and extrapolation algorithm. The union of non-regular sampling and the proposed reconstruction allows for achieving a higher resolution and therewith an improved imaging quality. Jürgen Seiler, Markus Jonscher, Thomas Ußmüller, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Iterative Optimization of Quarter Sampling Masks for Non-Regular Sampling SensorsabstractNon-regular sampling can reduce aliasing at the expense of noise. Recently, it has been shown that non-regular sampling can be carried out using a conventional regular imaging sensor when the surface of its individual pixels is partially covered. This technique is called quarter sampling (also 1/4 sampling), since only one quarter of each pixel is sensitive to light. For this purpose, the choice of a proper sampling mask is crucial to achieve a high reconstruction quality. In the scope of this work, we present an iterative algorithm to improve an arbitrary quarter sampling mask which results in a continuous increase of the reconstruction quality. In terms of the reconstruction algorithms, we test two simple algorithms, namely, linear interpolation and nearest neighbor interpolation, as well as two more sophisticated algorithms, namely, steering kernel regression and frequency selective extrapolation. Besides PSNR gains of +0.31 dB to +0.68 dB relative to a random quarter sampling mask resulting from our optimized mask, visually noticeable enhancements are perceptible. Simon Grosche, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2018 | Sparse Hartley Modeling for Fast Image ExtrapolationabstractIn many cases, image and video signal processing demands for high quality extrapolation algorithms, e.g., to solve inpainting problems or to increase image resolution. Indeed, a high computational load goes hand in hand with a good reconstruction quality as expensive models are calculated to estimate the missing data. To overcome this, the high-speed sparse Hartley modeling is introduced in this paper. This algorithm is based on Frequency Selective Extrapolation. In contrast to that, the model generation is carried out in the Hartley domain to exploit its real-valued transform properties. Due to this, it is possible to reduce the computational complexity significantly as no complex-valued arithmetic operations have to be conducted. In other words, a slightly higher reconstruction quality is obtained, while the proposed method is more than three times faster than the competing Frequency Selective Extrapolation. Nils Genser, Simon Grosche, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2018 | Signal and Loss Geometry Aware Frequency Selective Extrapolation for Error ConcealmentabstractThe concealment of errors is an important task in image and video signal processing. Often, complex models are calculated to reconstruct the missing samples, which results in a long computation time. One method that achieves a very high reconstruction quality, but demands a moderate computational complexity only, is the block based Frequency Selective Extrapolation. Nevertheless, the reconstruction of a Full HD image can still take several minutes depending on the error pattern. To accelerate the computation, a novel algorithm is introduced in this paper that analyzes the adjacent, undistorted samples and optimizes the reconstruction parameters accordingly. Moreover, the analyzation is further used to adapt the partitioning of the blocks and the processing order. Similar to modern video codecs, e.g., High Efficiency Video Coding, a content based partitioning and processing is proposed as it takes the signal characteristics into account. Thus, the novel algorithm is on average four times faster than the state-of-the-art method and up to 25× quicker at best, while achieving a slightly higher reconstruction quality as well. Nils Genser, Jürgen Seiler, Franz Schilling, André Kaup |
PCS | 2 |
| 2018 | Compression of Dynamic Medical CT Data Using Motion Compensated Wavelet Lifting with Denoised UpdateabstractFor the lossless compression of dynamic $3-\mathrm {D}+\mathrm {t}$ volumes as produced by medical devices like Computed Tomography, various coding schemes can be applied. This paper shows that 3-D subband coding outperforms lossless HEVC coding and additionally provides a scalable representation, which is often required in telemedicine applications. However, the resulting lowpass subband, which shall be used as a downscaled representative of the whole original sequence, contains a lot of ghosting artifacts. This can be alleviated by incorporating motion compensation methods into the subband coder. This results in a high quality lowpass subband but also leads to a lower compression ratio. In order to cope with this, we introduce a new approach for improving the compression efficiency of compensated 3-D wavelet lifting by performing denoising in the update step. We are able to reduce the file size of the lowpass subband by up to 1.64%, while the lowpass subband is still applicable for being used as a downscaled representative of the whole original sequence. Daniela Lanz, Jürgen Seiler, Karina Jaskolka, André Kaup |
PCS | 2 |
| 2018 | Sparse signal recovery with multiple prior information: Algorithm and measurement bounds
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup |
Signal Process. | 3 |
| 2018 | Compressive Online Robust Principal Component Analysis via n-ℓ1 MinimizationabstractThis paper considers online robust principal component analysis (RPCA) in time-varying decomposition problems such as video foreground-background separation. We propose a compressive online RPCA algorithm that decomposes recursively a sequence of data vectors (e.g., frames) into sparse and low-rank components. Different from conventional batch RPCA, which processes all the data directly, our approach considers a small set of measurements taken per data vector (frame). Moreover, our algorithm can incorporate multiple prior information from previous decomposed vectors via proposing an - minimization method. At each time instance, the algorithm recovers the sparse vector by solving the - minimization problem-which promotes not only the sparsity of the vector but also its correlation with multiple previously recovered sparse vectors-and, subsequently, updates the low-rank component using incremental singular value decomposition. We also establish theoretical bounds on the number of measurements required to guarantee successful compressive separation under the assumptions of static or slowly changing low-rank components. We evaluate the proposed algorithm using numerical experiments and online video foreground-background separation experiments. The experimental results show that the proposed method outperforms the existing methods. Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup |
IEEE Trans. Image Process. | 3 |
| 2017 | Demonstration of rapid frequency selective reconstruction for image resolution enhancementabstractWe propose to adapt the local statistics estimation, which has been designed for rapid error concealment originally. By merging the approaches, the reconstruction of images is fastened significantly. We will show in this demonstration that several frames per second (fps) can be processed using conventional computers and HD sized image data. Nils Genser, Jürgen Seiler, Markus Jonscher, André Kaup |
ICIP | 2 |
| 2017 | Scaled fixed-point frequency selective extrapolation for fast image error concealmentabstractImage and video signal processing demands in many areas for error concealment algorithms, whereby the execution time plays an important role. Within this paper, we introduce the scaled fixed-point Frequency Selective Extrapolation for fast image error concealment. This algorithm is based on the complex-valued Frequency Selective Extrapolation, but reduces computational complexity. It is shown that a fixed-point realization of the state-of-the-art algorithm requires an impracticable word-length, which makes it necessary to design a novel, scaled Frequency Selective Extrapolation that reduces the required word-length. Regarding the new approach, execution time can be reduced on average by 22.84 % and at best by up to 40.45 % compared to the state-of-the-art floating point Frequency Selective Extrapolation at the same reconstruction quality. Moreover, platforms can be exploited, which support fixed-point arithmetic only. Nils Genser, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2017 | Frequency-Selective Mesh-to-Grid Resampling for Image CommunicationabstractThis paper presents a novel approach for image reconstruction from pixels located at arbitrary noninteger positions, called mesh. This task forms an intrinsic part of various multimedia applications, including superresolution, fisheye imaging, or generations of new views in multicamera systems, among others. We propose a new frequency-selective mesh-to-grid resampling algorithm that aims at producing high quality reconstructions. It is inspired by the existing frequency-selective reconstruction (FSR) algorithm that is known to exhibit high performance when pixels are located on the regular 2D grid. However, if samples that are located at noninteger positions are involved, a severe overfitting problem arises from the fact that nonorthogonal weighted bases sampled at noninteger positions are used for signal modeling. In order to overcome this issue, we propose a novel stabilizing mechanism that is based on a set of adaptively weighted initial estimates, called key points. We also show that Fourier basis, used in the classic grid-based FSR, yields complex signals when noninteger positions are involved. Since digital images are real valued, we propose to employ a 2D cosine transform basis. Experimental results show the superiority of the proposed approach over a wide range of existing reconstruction techniques. Ján Koloda, Jürgen Seiler, André Kaup |
IEEE Trans. Multim. | 2 |
| 2016 | Multi-mode Kernel-Based Minimum Mean Square Error Estimator for Accelerated Image Error ConcealmentabstractSummary form only given. In this paper, we propose a novel multi-mode error concealment algorithm that aims at obtaining high quality reconstructions with reduced computational burden. Block-based coding schemes in packet loss-environment are considered. The proposed technique exploits the excellent reconstructing abilities of the kernel-based minimum mean square error (K-MMSE) estimator [1]. The complexity of our technique is dynamically adapted to the visual complexity of the area being reconstructed. The technique outperforms other state of the art algorithms and produces high quality reconstructions, equivalent to K-MMSE, while requiring less than one fourth of its computational time. Ján Koloda, Jürgen Seiler, Antonio M. Peinado, André Kaup |
DCC | 2 |
| 2016 | A Reconstruction Algorithm with Multiple Side Information for Distributed Compression of Sparse SourcesabstractWe consider the task of reconstructing target signals which are processed as sparse sources for a distributed compression scenario, where communication between the sources is prohibited, however, correlation of information among sources can be utilized at the decoder. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI) for such distributed sparse sources. The proposed algorithm takes advantage of both a compressive sensing (CS) reconstruction with SI and an iteratively weighted ℓ1-norm minimization by solving a general weighted multi-ℓ1(or n-ℓ1) minimization. To utilize the known multiple SIs, the algorithm computes optimal weights on not only each individual SI but among SIs where the weights are adaptively updated according to changes at every iteration of the reconstruction. By this optimization, the proposed reconstruction algorithm with multiple SI (RAMSI) can robustly exploit the multiple SIs with different qualities. We experimentally demonstrate our algorithm on compressing feature histograms as sparse sources which are extracted from a multi-view image database for multi-view recognition. The results show that the RAMSI with multiple SIs efficiently outperforms the ℓ1minimization and also the CS reconstruction with only one SI. Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer |
DCC | 2 |
| 2016 | Sparse signal reconstruction with multiple side information using adaptive weights for multiview sourcesabstractThis work considers reconstructing a target signal in a context of distributed sparse sources. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI). The proposed algorithm takes advantage of compressive sensing (CS) with SI and adaptive weights by solving a proposed weighted n-ℓ1minimization. The proposed algorithm computes the adaptive weights in two levels, first each individual intra-SI and then inter-SI weights are iteratively updated at every reconstructed iteration. This two-level optimization leads the proposed reconstruction algorithm with multiple SI using adaptive weights (RAMSIA) to robustly exploit the multiple SIs with different qualities. We experimentally perform our algorithm on generated sparse signals and also correlated feature histograms as multiview sparse sources from a multiview image database. The results show that RAMSIA significantly outperforms both classical CS and CS with single SI, and RAMSIA with higher number of SIs gained more than the one with smaller number of SIs. Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer |
ICIP | 2 |
| 2016 | Reliability-based mesh-to-grid image reconstructionabstractThis paper presents a novel method for the reconstruction of images from samples located at non-integer positions, called mesh. This is a common scenario for many image processing applications, such as super-resolution, warping or virtual view generation in multi-camera systems. The proposed method relies on a set of initial estimates that are later refined by a new reliability-based content-adaptive framework that employs denoising in order to reduce the reconstruction error. The reliability of the initial estimate is computed so stronger denoising is applied to less reliable estimates. The proposed technique can improve the reconstruction quality by more than 2 dB (in terms of PSNR) with respect to the initial estimate and it outperforms the state-of-the-art denoising-based refinement by up to 0.7 dB. Ján Koloda, Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2016 | Adaptive frequency prior for frequency selective reconstruction of images from non-regular subsamplingabstractImage signals typically are defined on a rectangular two-dimensional grid. However, there exist scenarios where this is not fulfilled and where the image information only is available for a non-regular subset of pixel position. For processing, transmitting or displaying such an image signal, a re-sampling to a regular grid is required. Recently, Frequency Selective Reconstruction (FSR) has been proposed as a very effective sparsity-based algorithm for solving this under-determined problem. For this, FSR iteratively generates a model of the signal in the Fourier-domain. In this context, a fixed frequency prior inspired by the optical transfer function is used for favoring low-frequency content. However, this fixed prior is often too strict and may lead to a reduced reconstruction quality. To resolve this weakness, this paper proposes an adaptive frequency prior which takes the local density of the available samples into account. The proposed adaptive prior allows for a very high reconstruction quality, yielding gains of up to 0.6 dB PSNR over the fixed prior, independently of the density of the available samples. Compared to other state-of-the-art algorithms, visually noticeable gains of several dB are possible. Jürgen Seiler, André Kaup |
MMSP | 1 |
| 2016 | Recursive frequency selective reconstruction of non-regularly sampled video dataabstractHigh resolution images can be acquired using a non-regular sampling sensor which consists of an underlying low resolution sensor that is covered with a non-regular sampling mask. The reconstructed high resolution image is then obtained during post-processing. Recently, it has been shown that the temporal correlation between neighboring frames can be exploited in order to enhance the reconstruction quality of non-regularly sampled video data. In this paper, a new recursive multi-frame reconstruction approach is proposed in order to further increase the reconstruction quality. By using a new reference order, previously reconstructed frames can be used for the subsequent motion estimation and a new weighting function allows for the incorporation of multiple pixels projected onto the same position. With the new recursive multi-frame approach, a visually noticeable average gain in PSNR of up to 1.13 dB with respect to a state-of-the-art single-frame reconstruction approach can be achieved. Compared to the existing multi-frame approach, a gain of 0.31 dB is possible. SSIM results show the same behavior as PSNR results. Additionally, the pre-reconstruction step of the existing multi-frame approach can be avoided and the new algorithm is, in general, capable of real-time processing. Markus Jonscher, Karina Jaskolka, Jürgen Seiler, André Kaup |
PCS | 3 |
| 2016 | Texture-dependent frequency selective reconstruction of non-regularly sampled imagesabstractThere exist many scenarios where pixel information is available only on a non-regular subset of pixel positions. For further processing, however, it is required to reconstruct such images on a regular grid. Besides many other algorithms, frequency selective reconstruction can be applied for this task. It performs a block-wise generation of a sparse signal model as an iterative superposition of Fourier basis functions and uses this model to replace missing or corrupted pixels in an image. In this paper, it is shown that it is not required to spend the same amount of iterations on both homogeneous and heterogeneous regions. Hence, a new texture-dependent approach for frequency selective reconstruction is introduced that distributes the number of iterations depending on the texture of the regions to be reconstructed. Compared to the original frequency selective reconstruction and depending on the number of iterations, visually noticeable gains in PSNR of up to 1.47 dB can be achieved. Markus Jonscher, Jürgen Seiler, André Kaup |
PCS | 2 |
| 2016 | Optimized processing order for 3D hole filling in video sequences using frequency selective extrapolationabstractA problem often arising in video communication is the reconstruction of missing or distorted areas in a video sequence. Such holes of unavailable pixels may be caused for example by transmission errors of coded video data or undesired objects like logos. In order to close the holes given neighboring available content, a signal extrapolation has to be performed. The best quality can be achieved, if spatial as well as temporal information is used for the reconstruction. However, the question always is in which order to process the extrapolation to obtain the best result. In this paper, an optimized processing order is introduced for improving the extrapolation quality of Three-dimensional Frequency Selective Extrapolation. Using the proposed optimized order, holes in video sequences can be closed from the outer margin to the center, leading to a higher reconstruction quality, and visually noticeable gains of more than 0.5 dB PSNR are possible. Jürgen Seiler, Susanne Scholl, Wolfgang Schnurrer, André Kaup |
PCS | 1 |
| 2016 | Robust Super-Resolution for Mixed-Resolution Multiview Image Plus Depth DataabstractIncreasing spatial resolution and thus improving the image quality is a key issue in the mixed-resolution multiview image and video processing domain. Given adjacent camera perspectives with various spatial resolutions and their corresponding depth information, the high-frequency part of a high-resolution view can be used for increasing the image quality of a neighboring low-resolution camera perspective. However, a reasonable projection of high-frequency information onto the image plane of a neighboring low-resolution view typically requires pixel-wise error-free depth data for the high-resolution reference image. Starting from this, a novel image super-resolution approach is proposed that is robust against both inaccurate depth acquisition and nonperfect calibration of spatially low-resolution depth sensors. The algorithm is based on displacement-compensated high-frequency synthesis and aims at correcting the projection errors introduced by inaccurate depth information. The proposed approach is further effectively extended by a signal extrapolation technique. For a wide range of proper scenarios, the proposed framework achieves substantial objective and visual gains compared with the considered reference approaches. The improvement of quality is shown for both simulated and self-recorded experimental data. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Hybrid super-resolution combining example-based single-image and interpolation-based multi-image reconstruction approachesabstractAchieving a higher spatial resolution is of particular interest in many applications such as video surveillance and can be realized by employing higher resolution sensors or applying super-resolution methods. Traditional super-resolution algorithms are based on either a single low resolution image or on multiple low resolution frames. In this paper, a hybrid super-resolution method is proposed which combines both a single-image and a multi-image approach using a soft decision mask. The mask is computed from the motion information utilized in the multi-image super-resolution part. This concept is shown to work for one particular setup but is also extensible toward other combinations of single-image and multi-image super-resolution algorithms as well as other merging metrics. Simulation results show an average luminance PSNR gain of up to 0.85 dB and 0.59 dB for upscaling factors of 2 and 4, respectively. Visual results substantiate the objective results. Michel Bätz, Andrea Eichenseer, Jürgen Seiler, Markus Jonscher, André Kaup |
ICIP | 3 |
| 2015 | A hybrid motion estimation technique for fisheye video sequences based on equisolid re-projectionabstractCapturing large fields of view with only one camera is an important aspect in surveillance and automotive applications, but the wide-angle fisheye imagery thus obtained exhibits very special characteristics that may not be very well suited for typical image and video processing methods such as motion estimation. This paper introduces a motion estimation method that adapts to the typical radial characteristics of fisheye video sequences by making use of an equisolid re-projection after moving part of the motion vector search into the perspective domain via a corresponding back-projection. By combining this approach with conventional translational motion estimation and compensation, average gains in luminance PSNR of up to 1.14 dB are achieved for synthetic fish-eye sequences and up to 0.96 dB for real-world data. Maximum gains for selected frame pairs amount to 2.40 dB and 1.39 dB for synthetic and real-world data, respectively. Andrea Eichenseer, Michel Bätz, Jürgen Seiler, André Kaup |
ICIP | 3 |
| 2015 | Denoising-based image reconstruction from pixels located at non-integer positionsabstractDigital images are commonly represented as regular 2D arrays, so pixels are organized in form of a matrix addressed by integers. However, there are many image processing operations, such as rotation or motion compensation, that produce pixels at non-integer positions. Typically, image reconstruction techniques cannot handle samples at non-integer positions. In this paper, we propose to use triangulation-based reconstruction as initial estimate that is later refined by a novel adaptive denoising framework. Simulations reveal that improvements of up to more than 1.8 dB (in terms of PSNR) are achieved with respect to the initial estimate. Ján Koloda, Jürgen Seiler, André Kaup |
ICIP | 2 |
| 2015 | Super-resolution for mixed-resolution multiview image plus depth data using a novel two-stage high-frequency extrapolation method for occluded areasabstractMixed-resolution multiview setups offer great savings in both, costs regarding the equipment acquisition and complexity refering to data transmission and storage. However, for applications like free viewpoint television, high-resolution images are desired from all available camera perspectives. Therefore, high-frequency information of adjacent high-resolution cameras can be used to increase the visual quality of a low-resolution camera perspective. However, due to occlusions, some parts of the scene are invisible in the high-resolution views and cannot be synthesized from the reference images. In this paper, a novel high-frequency extrapolation method is proposed, utilizing additional sampling points pre-estimated by a low-resolution block-matching approach. Averaged over all considered test sets, the proposed method achieves a PSNR gain of 0.94 dB with respect to the unrefined signal extrapolation algorithm. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
ICIP | 2 |
| 2015 | Reconstruction of videos taken by a non-regular sampling sensorabstractRecently, it has been shown that a high resolution image can be obtained without the usage of a high resolution sensor. The main idea has been that a low resolution sensor is covered with a non-regular sampling mask followed by a reconstruction of the incomplete high resolution image captured this way. In this paper, a multi-frame reconstruction approach is proposed where a video is taken by a non-regular sampling sensor and fully reconstructed afterwards. By utilizing the temporal correlation between neighboring frames, the reconstruction quality can be further enhanced. Compared to a state-of-the-art single-frame reconstruction approach, this leads to a visually noticeable gain in PSNR of up to 1.19 dB on average. Markus Jonscher, Jürgen Seiler, Michel Bätz, Thomas Richter 0001, Wolfgang Schnurrer, André Kaup |
VCIP | 2 |
| 2015 | Centroid adapted frequency selective extrapolation for reconstruction of lost image areasabstractLost image areas with different size and arbitrary shape can occur in many scenarios such as error-prone communication, depth-based image rendering or motion compensated wavelet lifting. The goal of image reconstruction is to restore these lost image areas as close to the original as possible. Frequency selective extrapolation is a block-based method for efficiently reconstructing lost areas in images. So far, the actual shape of the lost area is not considered directly. We propose a centroid adaption to enhance the existing frequency selective extrapolation algorithm that takes the shape of lost areas into account. To enlarge the test set for evaluation we further propose a method to generate arbitrarily shaped lost areas. On our large test set, we obtain an average reconstruction gain of 1.29 dB. Wolfgang Schnurrer, Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, Michel Bätz, André Kaup |
VCIP | 3 |
| 2015 | Resampling Images to a Regular Grid From a Non-Regular Subset of Pixel Positions Using Frequency Selective ReconstructionabstractEven though image signals are typically defined on a regular 2D grid, there also exist many scenarios where this is not the case and the amplitude of the image signal only is available for a non-regular subset of pixel positions. In such a case, a resampling of the image to a regular grid has to be carried out. This is necessary since almost all algorithms and technologies for processing, transmitting or displaying image signals rely on the samples being available on a regular grid. Thus, it is of great importance to reconstruct the image on this regular grid, so that the reconstruction comes closest to the case that the signal has been originally acquired on the regular grid. In this paper, Frequency Selective Reconstruction is introduced for solving this challenging task. This algorithm reconstructs image signals by exploiting the property that small areas of images can be represented sparsely in the Fourier domain. By further considering the basic properties of the optical transfer function of imaging systems, a sparse model of the signal is iteratively generated. In doing so, the proposed algorithm is able to achieve a very high reconstruction quality, in terms of peak signal-to-noise ratio (PSNR) and structural similarity measure as well as in terms of visual quality. The simulation results show that the proposed algorithm is able to outperform state-of-the-art reconstruction algorithms and gains of more than 1 dB PSNR are possible. Jürgen Seiler, Markus Jonscher, Michael Schöberl, André Kaup |
IEEE Trans. Image Process. | 1 |
| 2014 | Frequency selective extrapolation with residual filtering for image error concealmentabstractThe purpose of signal extrapolation is to estimate unknown signal parts from known samples. This task is especially important for error concealment in image and video communication. For obtaining a high quality reconstruction, assumptions have to be made about the underlying signal in order to solve this underdetermined problem. Among existent reconstruction algorithms, frequency selective extrapolation (FSE) achieves high performance by assuming that image signals can be sparsely represented in the frequency domain. However, FSE does not take into account the low-pass behaviour of natural images. In this paper, we propose a modified FSE that takes this prior knowledge into account for the modelling, yielding significant PSNR gains. Ján Koloda, Jürgen Seiler, André Kaup, Victoria E. Sánchez, Antonio M. Peinado |
ICASSP | 2 |
| 2014 | Reconstruction of multiview images taken with non-regular sampling sensorsabstractIncreasing spatial image resolution is a widely discussed area in the field of image processing. In this paper, we present an efficient reconstruction approach for high-resolution images, taken with irregularly shielded low-resolution sensors in a multiview setup. The approach is based on the sparsity assumption, meaning that natural images can be efficiently represented in a transform-domain using only few coefficients. Utilizing information from adjacent cameras results in a better reconstruction quality for the central high-resolution view. Since neighboring camera perspectives might differ in illumination, the information from adjacent views has to be adapted to the view to be reconstructed. The simulation results show that a proper incorporation of information from neighboring views leads to a PSNR gain of up to 2.20 dB compared to a state-of-the-art singleview reconstruction approach. Thomas Richter 0001, Markus Jonscher, Wolfgang Schnurrer, Jürgen Seiler, André Kaup |
ICASSP | 4 |
| 2014 | Reconstruction of images taken by a pair of non-regular sampling sensors using correlation based matchingabstractMulti-view image acquisition systems with two or more cameras can be rather costly due to the number of high resolution image sensors that are required. Recently, it has been shown that by covering a low resolution sensor with a non-regular sampling mask and by using an efficient algorithm for image reconstruction, a high resolution image can be obtained. In this paper, a stereo image reconstruction setup for multi-view scenarios is proposed. A scene is captured by a pair of non-regular sampling sensors and by incorporating information from the adjacent view, the reconstruction quality can be increased. Compared to a state-of-the-art single-view reconstruction algorithm, this leads to a visually noticeable average gain in PSNR of 0.74 dB. Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, Michel Bätz, André Kaup |
ICIP | 2 |
| 2014 | An error-based recursive filling ordering for image error concealmentabstractMany image error concealment (EC) algorithms can be employed to reconstruct a lost region by dividing it into a set of (sub)blocks which are estimated from a set of known context pixels. In turn, these former pixels may have been previously obtained by estimation. In this situation, the order in which the lost region is recursively filled will clearly condition the resulting reconstruction. This paper proposes a novel filling order aimed to improve the performance of EC algorithms that are applied recursively over the lost area. It takes into account the reconstruction quality of the already concealed blocks in order to determine the filling order. Blocks surrounded by known pixels or by high quality reconstructions are prioritized. The proposed technique is applicable to a wide range of EC algorithms and achieves an improvement of up to 1dB (in PSNR) with negligible additional computational load. Ján Koloda, Jürgen Seiler, André Kaup, Victoria E. Sánchez, Antonio M. Peinado |
ICIP | 2 |
| 2014 | 3-D mesh compensated wavelet lifting for 3-D+t medical CT dataabstractFor scalable coding, a high quality of the lowpass band of a wavelet transform is crucial when it is used as a downscaled version of the original signal. However, blur and motion can lead to disturbing artifacts. By incorporating feasible compensation methods directly into the wavelet transform, the quality of the lowpass band can be improved. The displacement in dynamic medical 3-D+t volumes from Computed Tomography is mainly given by expansion and compression of tissue over time and can be modeled well by mesh-based methods. We extend a 2-D mesh-based compensation method to three dimensions to obtain a volume compensation method that can additionally compensate deforming displacements in the third dimension. We show that a 3-D mesh can obtain a higher quality of the lowpass band by 0.28 dB with less than 40% of the model parameters of a comparable 2-D mesh. Results from lossless coding with JPEG 2000 3D and SPECK3D show that the compensated subbands using a 3-D mesh need about 6% less data compared to using a 2-D mesh. Wolfgang Schnurrer, Thomas Richter 0001, Jürgen Seiler, Christian Herglotz, André Kaup |
ICIP | 3 |
| 2014 | Open source HEVC analyzer for rapid prototyping (HARP)abstractThe design of new HEVC extensions comes with the need for careful analysis of internal HEVC codec decisions. Several bitstream analyzers have evolved for this purpose and provide a visualization of encoder decisions as seen from a decoder viewpoint. None of the existing solutions is able to provide actual insight into the encoder and its RDO decision process. With one exception, all solutions are closed source and make adaption of their code to specific implementation needs impossible. Overall, development with the HM code base remains a time-consuming task. Here, we present the HEVC Analyzer for Rapid Prototyping (HARP), which directly addresses the above issues and is freely available under www.lms.lnt.de/HARP. Dominic Springer, Wolfgang Schnurrer, Andreas Weinlich, Andreas Heindel, Jürgen Seiler, André Kaup |
ICIP | 5 |
| 2014 | Accelerated hybrid image reconstruction for non-regular sampling color sensorsabstractIncreasing the spatial resolution is an ongoing research topic in image processing. A recently presented approach applies a non-regular sampling mask on a low resolution sensor and subsequently reconstructs the masked area via an extrapolation algorithm to obtain a high resolution image. This paper introduces an acceleration of this approach for use with full color sensors. Instead of employing the effective, yet computationally expensive extrapolation algorithm on each of the three RGB channels, a color space conversion is performed and only the luminance channel is then reconstructed using this algorithm. As natural images contain much less information in the chrominance channels, a fast linear interpolation technique can here be used to accelerate the whole reconstruction procedure. Simulation results show that an average speed up factor of 2.9 is thus achieved, while the loss in visual quality stays imperceptible. Comparisons of PSNR results confirm this. Michel Bätz, Andrea Eichenseer, Markus Jonscher, Jürgen Seiler, André Kaup |
VCIP | 4 |
| 2014 | Reducing randomness of non-regular sampling masks for image reconstructionabstractIncreasing spatial image resolution is an often required, yet challenging task in image acquisition. Recently, it has been shown that it is possible to obtain a high resolution image by covering a low resolution sensor with a non-regular sampling mask. Due to the masking, however, some pixel information in the resulting high resolution image is not available and has to be reconstructed by an efficient image reconstruction algorithm in order to get a fully reconstructed high resolution image. In this paper, the influence of different sampling masks with a reduced randomness of the non-regularity on the image reconstruction process is evaluated. Simulation results show that it is sufficient to use sampling masks that are non-regular only on a smaller scale. These sampling masks lead to a visually noticeable gain in PSNR compared to arbitrary chosen sampling masks which are non-regular over the whole image sensor size. At the same time, they simplify the manufacturing process and allow for efficient storage. Markus Jonscher, Jürgen Seiler, Thomas Richter 0001, André Kaup |
VCIP | 2 |
| 2014 | Efficient lossless coding of highpass bands from block-based motion compensated wavelet lifting using JPEG 2000abstractLossless image coding is a crucial task especially in the medical area, e.g., for volumes from Computed Tomography or Magnetic Resonance Tomography. Besides lossless coding, compensated wavelet lifting offers a scalable representation of such huge volumes. While compensation methods increase the details in the lowpass band, they also vary the characteristics of the wavelet coefficients, so an adaption of the coefficient coder should be considered. We propose a simple invertible extension for JPEG 2000 that can reduce the filesize for lossless coding of the highpass band by 0.8% on average with peak rate saving of 1.1%. Wolfgang Schnurrer, Tobias Tröger, Thomas Richter 0001, Jürgen Seiler, André Kaup |
VCIP | 4 |
| 2014 | High dynamic range video reconstruction from a stereo camera setup
Michel Bätz, Thomas Richter 0001, Jens-Uwe Garbas, Anton Papst, Jürgen Seiler, André Kaup |
Signal Process. Image Commun. | 5 |
| 2013 | Spatio-temporal error concealment in video by denoised temporal extrapolation refinementabstractIn video communication, the concealment of distortions caused by transmission errors is important for allowing for a pleasant visual quality and for reducing error propagation. In this article, Denoised Temporal Extrapolation Refinement is introduced as a novel spatiotemporal error concealment algorithm. The algorithm operates in two steps. First, temporal error concealment is used for obtaining an initial estimate. Afterwards, a spatial denoising algorithm is used for reducing the imperfectness of the temporal extrapolation. For this, Non-Local Means denoising is used which is extended by a spiral scan processing order and is improved by an adaptation step for taking the preliminary temporal extrapolation into account. In doing so, a spatio-temporal error concealment results. By making use of the refinement, a visually noticeable average gain of 1 dB over pure temporal error concealment is possible. With this, the algorithm also is able to clearly outperform other spatio-temporal error concealment algorithms. Jürgen Seiler, Michael Schöberl, André Kaup |
ICIP | 1 |
| 2012 | High dynamic range video by spatially non-regular optical filteringabstractWe present a new method for capturing high dynamic range video (HDRV). Our method is based on spatially varying exposures, where individual pixels are covered with filters for different optical attenuation. For preventing the loss in resolution we use a new non-regular arrangement of the attenuation pattern. Subsequent image reconstruction based on the sparsity assumption allows the reconstruction of natural images with high detail. Michael Schöberl, Alexander Belz, Jürgen Seiler, Siegfried Fößel, André Kaup |
ICIP | 3 |
| 2012 | Robust super-resolution in a multiview setup based on refined high-frequency synthesisabstractIncreasing image sharpness and thus improving the visual quality is an important task in multiview image and video processing. We propose a novel super-resolution approach for multiview images in a mixed-resolution setup that is robust to various depth map distortions. The considered distortion scenarios may be caused by an inaccurate calibration of the depth camera or a limitation of depth range. Our method is based on a refined high-frequency synthesis that relies on a blockwise and depth-dependant low-frequency registration. The refinement step efficiently adapts the high-frequency content from a neighboring high-resolution camera to a low-resolution view and thereby compensates the displacement caused by depth inaccuracies. In case of undistorted depth maps, the results show that our algorithm leads to a PSNR gain of up to 1.33 dB with respect to a comparable unrefined super-resolution approach for a mixed-resolution multiview video plus depth format. Compared to the initial low-resolution view, a PSNR gain of up to 2.61 dB is obtained. In case of distorted depth maps, a PSNR gain of even 4.78 dB is achieved with respect to the reference superresolution algorithm. The PSNR gains get confirmed by the corresponding SSIM values which manifest a similar behaviour. The improvement of visual quality is also convincingly for all considered scenarios. Thomas Richter 0001, Jürgen Seiler, Wolfgang Schnurrer, André Kaup |
MMSP | 2 |
| 2012 | Analysis of mesh-based motion compensation in wavelet lifting of dynamical 3-D+t CT dataabstractFactorized in the lifting structure, the wavelet transform can easily be extended by arbitrary compensation methods. Thereby, the transform can be adapted to displacements in the signal without losing the ability of perfect reconstruction. This leads to an improvement of scalability. In temporal direction of dynamic medical 3-D+t volumes from Computed Tomography, displacement is mainly given by expansion and compression of tissue. We show that these smooth movements can be well compensated with a mesh-based method. We compare the properties of triangle and quadrilateral meshes. We also show that with a mesh-based compensation approach coding results are comparable to the common slice wise coding with JPEG 2000 while a scalable representation in temporal direction can be achieved. Wolfgang Schnurrer, Thomas Richter 0001, Jürgen Seiler, André Kaup |
MMSP | 3 |
| 2012 | On the influence of clipping in lossless predictive and wavelet coding of noisy imagesabstractEspecially in lossless image coding the obtainable compression ratio strongly depends on the amount of noise included in the data as all noise has to be coded, too. Different approaches exist for lossless image coding. We analyze the compression performance of three kinds of approaches, namely direct entropy, predictive and wavelet-based coding. The results from our theoretical model are compared to simulated results from standard algorithms that base on the three approaches. As long as no clipping occurs with increasing noise more bits are needed for lossless compression. We will show that for very noisy signals it is more advantageous to directly use an entropy coder without advanced preprocessing steps. Wolfgang Schnurrer, Jürgen Seiler, Michael Schöberl, André Kaup |
PCS | 2 |
| 2012 | Analysis of displacement compensation methods for wavelet lifting of medical 3-D thorax CT volume dataabstractA huge advantage of the wavelet transform in image and video compression is its scalability. Wavelet-based coding of medical computed tomography (CT) data becomes more and more popular. While much effort has been spent on encoding of the wavelet coefficients, the extension of the transform by a compensation method as in video coding has not gained much attention so far. We will analyze two compensation methods for medical CT data and compare the characteristics of the displacement compensated wavelet transform with video data. We will show that for thorax CT data the transform coding gain can be improved by a factor of 2 and the quality of the lowpass band can be improved by 8 dB in terms of PSNR compared to the original transform without compensation. Wolfgang Schnurrer, Jürgen Seiler, Eugen Wige, André Kaup |
VCIP | 2 |
| 2011 | Sparsity-based defect pixel compensation for arbitrary camera raw imagesabstractIn high quality imaging even tiny distortions as small as a single pixel are visible and can not be accepted. Although the production quality of CMOS image sensors is very high, for reasonable yields we still need to accept some defect pixels and clusters of defects in large image sensors. In this paper we will compare compensation algorithms for raw image sensor data. We propose a new approach based on the sparsity assumption that outperforms existing defect compensation algorithms. Furthermore, our proposed interpolation algorithm is universal and not at all adapted to Bayer pattern images. It can directly be applied to any regular color filter pattern or gray scale image. Our examples show, that image sensors with large clusters of defects can still be used for the generation of high quality images. Michael Schöberl, Jürgen Seiler, Bernhard Kasper, Siegfried Fößel, André Kaup |
ICASSP | 2 |
| 2011 | Reusing the H.264/AVC deblocking filter for efficient spatio-temporal prediction in video codingabstractThe prediction step is a very important part of hybrid video codecs for effectively compressing video sequences. While existing video codecs predict either in temporal or in spatial direction only, the compression efficiency can be increased by a combined spatio-temporal prediction. In this paper we propose an algorithm for reusing the H.264/AVC deblocking filter for spatio-temporal prediction. Reusing this highly op timized filter allows for a very low computational complexity of this prediction mode and an average rate reduction of up to 7.2% can be achieved. Jürgen Seiler, André Kaup |
ICASSP | 1 |
| 2011 | Increasing imaging resolution by covering your sensorabstractUp to now, an increase in camera resolution required image sensors with more and more pixels. However, acquisition systems are limited in their pixels per second throughput given as power and complexity constraints. Simply capturing more pixels in a given system is often not possible. We propose a new non-regular imaging architecture that samples only few pixels and reconstructs a high resolution image afterwards. Our sampling is optimized to provide non-regular spatial sampling from a sensor with regular readout circuits. An existing slow image acquisition system can then be used to capture the data. The image reconstruction is performed with a local sparsity-based approach. The result is a high resolution image that requires a much smaller effort during acquisition. Michael Schöberl, Jürgen Seiler, Siegfried Fößel, André Kaup |
ICIP | 2 |
| 2011 | Motion Compensated Three-Dimensional Frequency Selective Extrapolation for improved error concealment in video communication
Jürgen Seiler, André Kaup |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Multiple Selection Approximation for improved spatio-temporal prediction in video codingabstractIn this contribution, a novel spatio-temporal prediction algorithm for video coding is introduced. This algorithm exploits temporal as well as spatial redundancies for effectively predicting the signal to be encoded. To achieve this, the algorithm operates in two stages. Initially, motion compensated prediction is applied on the block being encoded. Afterwards this preliminary temporal prediction is refined by forming a joint model of the initial predictor and the spatially adjacent already transmitted blocks. The novel algorithm is able to outperform earlier refinement algorithms in speed and prediction quality. Compared to pure motion compensated prediction, the mean data rate can be reduced by up to 15% and up to 1.16 dB gain in PSNR can be achieved for the considered sequences. Jürgen Seiler, André Kaup |
ICASSP | 1 |
| 2010 | Content-Adaptive Motion Compensated Frequency Selective Extrapolation for error concealment in video communicationabstractIf digital video data is transmitted over unreliable channels such as the internet or wireless terminals, the risk of severe image distortion due to transmission errors is ubiquitous. To cope with this, error concealment can be applied on the distorted data at the receiver. In this contribution we propose a novel spatio-temporal error concealment algorithm, the Content-Adaptive Motion Compensated Frequency Selective Extrapolation. The algorithm operates in two stages, whereas at first the motion in a distorted sequence is estimated. After that, a model of the signal is generated for concealing the distortion. The novel algorithm is based on an already existent error concealment algorithm. But by adapting the model generation to the content of a sequence, the novel algorithm is able to exploit the remaining information, which is still available in the distorted sequence, more effectively compared to the original algorithm. In doing so, a visually noticeable gain of up to 0.51 dB PSNR compared to the underlying algorithm and more than 3 dB compared to other error concealment algorithms can be achieved. Jürgen Seiler, André Kaup |
ICIP | 1 |
| 2010 | Spatially refined inter-sequence error concealment for a multi-broadcast receiver using frequency selective approximationabstractMobile reception of digital TV often suffers from severe signal degradations. Inter-sequence error concealment reconstructs lost image blocks of a distorted high-resolution TV signal by inserting corresponding error-free blocks from a low-resolution reference TV signal. It is well-suited for application in future automotive multi-broadcast receivers and can outperform state-of-the-art methods by up to 15 dB PSNRY depending on the quality of the reference signal. In this contribution, we show that inter-sequence error concealment can be improved by approximating inserted blocks jointly with neighboring pixels according to a well-known frequency selective method. The reconstruction quality can be significantly increased especially in case of low-bitrate reference signals. Inserted blocks can be further refined also for high bitrates. On average, a gain of 1.7 dB PSNRY can be achieved. The peak gain which is evaluated on frame basis even reaches 5.6 dB PSNRY. Tobias Tröger, Jürgen Seiler, André Kaup |
ACM Multimedia | 2 |
| 2010 | Spatio-temporal prediction in video coding by non-local means refined motion compensationabstractThe prediction step is a very important part of hybrid video codecs. In this contribution, a novel spatio-temporal prediction algorithm is introduced. For this, the prediction is carried out in two steps. Firstly, a preliminary temporal prediction is conducted by motion compensation. Afterwards, spatial refinement is carried out for incorporating spatial redundancies from already decoded neighboring blocks. Thereby, the spatial refinement is achieved by applying Non-Local Means de-noising to the union of the motion compensated block and the already decoded blocks. Including the spatial refinement into H.264/AVC, a rate reduction of up to 14 % or respectively a gain of up to 0.7 dB PSNR compared to unrefined motion compensated prediction can be achieved. Jürgen Seiler, Thomas Richter 0001, André Kaup |
PCS | 1 |
| 2010 | Complex-Valued Frequency Selective Extrapolation for Fast Image and Video Signal ExtrapolationabstractSignal extrapolation tasks arise in miscellaneous manners in the field of image and video signal processing. But, due to the widespread use of low-power and mobile devices, the computational complexity of an algorithm plays a crucial role in selecting an algorithm for a given problem. Within the scope of this contribution, we introduce the complex-valued Frequency Selective Extrapolation for fast image and video signal extrapolation. This algorithm iteratively generates a generic complex-valued model of the signal to be extrapolated as weighted superposition of Fourier basis functions. We further show that this algorithm is up to 10 times faster than the existent real-valued Frequency Selective Extrapolation that takes the real-valued nature of the input signals into account during the model generation. At the same time, the quality which is achievable by the complex-valued model generation is similar to the quality of the real-valued model generation. Jürgen Seiler, André Kaup |
IEEE Signal Process. Lett. | 1 |
| 2009 | Multiple Selection Extrapolation for improved spatial error concealmentabstractThis contribution introduces a novel signal extrapolation algorithm and its application to image error concealment. The signal extrapolation is carried out by iteratively generating a model of the signal suffering from distortion. Thereby, the model results from a weighted superposition of two-dimensional basis functions whereas in every iteration step a set of these is selected and the approximation residual is projected onto the subspace they span. The algorithm is an improvement to the Frequency Selective Extrapolation that has proven to be an effective method for concealing lost or distorted image regions. Compared to this algorithm, the novel algorithm is able to reduce the processing time by a factor larger than three, by still preserving the very high extrapolation quality. Jürgen Seiler, André Kaup |
MMSP | 1 |
| 2009 | Spatio-temporal prediction in video coding by best approximationabstractWithin the scope of this contribution we propose a novel efficient spatio-temporal prediction algorithm for video coding. The algorithm operates in two stages. First, motion compensation is performed on the block to be predicted in order to exploit temporal correlations. Afterwards, in order to exploit spatial correlations, this preliminary estimate is spatially refined by forming a joint model of the motion compensated block and spatially adjacent already decoded blocks. Compared to an earlier refinement algorithm, the novel one only needs very little iteration, leading to a speedup of factor 17. The implementation of this new algorithm into the H.264/AVC leads to a maximum reduction in data rate of up to nearly 13% for the considered sequences. Jürgen Seiler, Haricharan Lakshman, André Kaup |
PCS | 1 |
| 2008 | Fast orthogonality deficiency compensation for improved frequency selective image extrapolationabstractThe purpose of this paper is to introduce a very efficient algorithm for signal extrapolation. It can widely be used in many applications in image and video communication, e. g. for concealment of block errors caused by transmission errors or for prediction in video coding. The signal extrapolation is performed by extending a signal from a limited number of known samples into areas beyond these samples. Therefore a finite set of orthogonal basis functions is used and the known part of the signal is projected onto them. Since the basis functions are not orthogonal regarding the area of the known samples, the projection does not lead to the real portion a basis function has of the signal. The proposed algorithm efficiently copes with this non-orthogonality resulting in very good objective and visual extrapolation results for edges, smooth areas, as well as structured areas. Compared to an existent implementation, this algorithm has a significantly lower computational complexity without any degradation in quality. The processing time can be reduced by a factor larger than 100. Jürgen Seiler, André Kaup |
ICASSP | 1 |
| 2008 | Spatio-temporal prediction in video coding by spatially refined motion compensationabstractThe purpose of this contribution is to introduce a new method of signal prediction in video coding. Unlike most existent prediction methods that either use temporal or use spatial correlations to generate the prediction signal, the proposed method uses spatial and temporal correlations at the same time. The spatio-temporal prediction is obtained by first performing motion compensation for a macroblock, followed by a refinement step that pays attention to the correlations between the macroblock and its surroundings. At the decoder, the refinement step can be performed in the same manner, thus no additional side information has to be transmitted. Implementation of the spatial refinement step into the H.264/AVC video codec leads to reduction in data rate of up to nearly 15% and increase in PSNR of up to 0.75 dB, compared to pure motion compensated prediction. Jürgen Seiler, André Kaup |
ICIP | 1 |
| 2008 | 4-D frequency selective extrapolation for error concealment in multi-view videoabstractPractical applications for multi-view video such as three-dimensional television may require to transmit the data over error-prone channels. Even when channel coding is used, it is likely that parts of the image information are lost after decoding. To improve the image quality in this case, an efficient algorithm for multi-view error concealment is presented. Extending and enhancing a previously known method for single-view concealment, the algorithm simultaneously uses information from surrounding image parts, from temporally preceding and succeeding frames and from neighbouring camera views for extrapolating known image samples into the lost area. It is shown that the proposed algorithm leads to convincing results in terms of PSNR as well as to a good subjective quality. Compared to single-view concealment, the result is improved by up to 0.75 dB by adding cross-view information. Ulrich Fecker, Jürgen Seiler, André Kaup |
MMSP | 2 |
| 2008 | Adaptive joint spatio-temporal error concealment for video communicationabstractIn the past years, video communication has found its application in an increasing number of environments. Unfortunately, some of them are error-prone and the risk of block losses caused by transmission errors is ubiquitous. To reduce the effects of these block losses, a new spatio-temporal error concealment algorithm is presented. The algorithm uses spatial as well as temporal information for extrapolating the signal into the lost areas. The extrapolation is carried out in two steps, first a preliminary temporal extrapolation is performed which then is used to generate a model of the original signal, using the spatial neighborhood of the lost block. By applying the spatial refinement a significantly higher concealment quality can be achieved resulting in a gain of up to 5.2 dB in PSNR compared to the unrefined underlying pure temporal extrapolation. Jürgen Seiler, André Kaup |
MMSP | 1 |