Dietmar Saupe

dblp:40/374 · DBLP profile ↗
← Back
82ranked-venue papers
10as first author
18since 2021 · last 2025
0000-0001-6735-5103ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 10 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 20 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Fine-Grained Subjective Visual Quality Assessment for High-Fidelity Compressed Images
abstract
Advances in image compression, storage, and display technologies have made high-quality images and videos widely accessible. At this level of quality, distinguishing between compressed and original content becomes difficult, highlighting the need for assessment methodologies that are sensitive to even the smallest visual quality differences. Conventional subjective visual quality assessments often use absolute category rating scales, ranging from “excellent” to “bad”. While suitable for evaluating more pronounced distortions, these scales are inadequate for detecting subtle visual differences. The JPEG standardization project AIC is currently developing a subjective image quality assessment methodology for high-fidelity images. This paper presents the proposed assessment methods, a dataset of high-quality compressed images, and their corresponding crowdsourced visual quality ratings. It also outlines a data analysis approach that reconstructs quality scale values in just noticeable difference (JND) units. The assessment method uses boosting techniques on visual stimuli to help observers detect compression artifacts more clearly. This is followed by a rescaling process that adjusts the boosted quality values back to the original perceptual scale. This reconstruction yields a fine-grained, high-precision quality scale in JND units, providing more informative results for practical applications. The dataset and code to reproduce the results will be available at https://github.com/jpeg-aic/dataset-BTC-PTC-24.
Michela Testolina, Mohsen Jenadeleh, Shima Mohammadi, Shaolin Su, João Ascenso, Touradj Ebrahimi, Jon Sneyers, Dietmar Saupe
DCC8
2025 Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and Resolution
abstract
Image Quality Assessment (IQA) measures and predicts perceived image quality by human observers. Although recent studies have highlighted the critical influence that variations in the scale of an image have on its perceived quality, this relationship has not been systematically quantified. To bridge this gap, we introduce the Image Intrinsic Scale (IIS), defined as the largest scale where an image exhibits its highest perceived quality. We also present the Image Intrinsic Scale Assessment (IISA) task, which involves subjectively measuring and predicting the IIS based on human judgments. We develop a subjective annotation methodology and create the IISA-DB dataset, comprising 785 image-IIS pairs annotated by experts in a rigorously controlled crowdsourcing study. Furthermore, we propose WIISA (Weak-labeling for Image Intrinsic Scale Assessment), a strategy that leverages how the IIS of an image varies with downscaling to generate weak labels. Experiments show that applying WIISA during the training of several IQA methods adapted for IISA consistently improves the performance compared to using only ground-truth labels. The code, dataset, and pre-trained models are available at https://github.com/SonyResearch/IISA.
Vlad Hosu, Lorenzo Agnolucci, Daisuke Iso, Dietmar Saupe
ICCV4
2025 Subjective Visual Quality Assessment for High-Fidelity Learning-Based Image Compression
abstract
Learning-based image compression methods offer a promising alternative to traditional codecs by improving rate-distortion performance. JPEG AI is the first standard in this domain and leverages deep neural networks to achieve high-fidelity image reconstruction. In this work, we present a comprehensive subjective visual quality assessment of JPEG AI-compressed images using the JPEG AIC-3 methodology, which quantifies perceptual differences using Just Noticeable Difference (JND) units. We created a dataset of 50 compressed images with fine-grained distortion levels from five diverse source images and conducted a large-scale crowdsourced experiment that collected 96,200 triplet responses from 459 participants. We reconstructed JND-based quality scales using a unified model based on both boosted and plain triplet comparisons. We also evaluated how well objective image quality metrics align with human perception in the high-fidelity range. The CVVDP metric achieved the highest overall performance, however, most metrics, including CVVDP, were overly optimistic in estimating image quality, emphasizing the need for rigorous subjective evaluation. We introduced the Meng–Rosenthal–Rubin test into Quality of Experience research to compare metric correlations with a shared ground truth. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Panqi Jia, Shima Mohammadi, João Ascenso, Dietmar Saupe
QoMEX6
2025 Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
abstract
High dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe
QoMEX13
2025 Robustness and accuracy of mean opinion scores with hard and soft outlier detection
abstract
In subjective assessment of image and video quality, observers rate or compare selected stimuli. Before calculating the mean opinion scores (MOS) for these stimuli from the ratings, it is recommended to identify and deal with outliers that may have given unreliable ratings. Several methods are available for this purpose, some of which have been standardized. These methods are typically based on statistics and sometimes tested by introducing synthetic ratings from artificial outliers, such as random clickers. However, a reliable and comprehensive approach is lacking for comparative performance analysis of outlier detection methods. To fill this gap, this work proposes and applies an empirical worst-case analysis as a general solution. Our method involves evolutionary optimization of an adversarial black-box attack on outlier detection algorithms, where the adversary maximizes the distortion of scale values with respect to ground truth. We apply our analysis to several hard and soft outlier detection methods for absolute category ratings and show their differing performance in this stress test. In addition, we propose two new outlier detection methods with low complexity and excellent worst-case performance. Software for adversarial attacks and data analysis is available.1
Dietmar Saupe, Tim Bleile
QoMEX1
2024 An Image Quality Dataset with Triplet Comparisons for Multi-dimensional Scaling
abstract
In the early days of perceptual image quality research more than 30 years ago, the multidimensionality of distortions in perceptual space was considered important. However, research focused on scalar quality as measured by mean opinion scores. With our work, we intend to revive interest in this relevant area by presenting a first pilot dataset of annotated triplet comparisons for image quality assessment. It contains one source stimulus together with distorted versions derived from 7 distortion types at 12 levels each. Our crowdsourced and curated dataset contains roughly 50,000 responses to 7,000 triplet comparisons. We show that the multidimensional embedding of the dataset poses a challenge for many established triplet embedding algorithms. Finally, we propose a new reconstruction algorithm, dubbed logistic triplet embedding (LTE) with Tikhonov regularization. It shows promising performance. This study helps researchers to create larger datasets and better embedding techniques for multidimensional image quality. The dataset includes images and ratings and can be accessed at https://github.com/jenadeleh/multidimensionalIQA-dataset/tree/main.
Mohsen Jenadeleh, Frederik L. Dennig, René Cutura, Quynh Quang Ngo, Daniel A. Keim, Michael Sedlmair, Dietmar Saupe
QoMEX7
2024 Impact of feedback on crowdsourced visual quality assessment with paired comparisons
abstract
This paper presents a comprehensive investigation into the effects of immediate feedback on crowdworkers’ performance in subjective image quality assessment tasks using paired comparisons. The study is motivated by the need for reliable and efficient crowdsourcing tasks for image quality assessment. A large-scale experiment involving 200 participants was conducted, where participants completed 120 paired comparisons with and without feedback. The feedback informed the workers of the correctness of their responses to comparisons. Almost all of the participants (97%) preferred receiving feedback. The results indicate that feedback reduced response time, improved user experience, and did not cause a bias in the estimation of the just noticeable difference (JND). On the other hand, feedback did not significantly affect accuracy, correlation with the ground truth, or create a learning effect. This study contributes to the field by being one of the first to examine the impact of feedback on crowdworker performance in subjective image quality assessment tasks. The dataset which includes the images and ratings can be accessed at https://database.mmsp-kn.de/feedback-study-dataset.html.
Mohsen Jenadeleh, Alexander Heß, Simon Hviid Del Pin, Edwin Gamboa, Matthias Hirth, Dietmar Saupe
QoMEX6
2024 National differences in image quality assessment: An investigation on three large-scale IQA datasets
abstract
This paper investigates the potential effects of national differences on image and video quality assessment using discrete rating scales. Drawing on cultural psychology, we hypothesize that observers from different countries may exhibit distinct response styles in interpreting and applying the five-level absolute and degradation category rating scales (ACR, DCR). For our study, we adapt state-of-the-art statistical models for three large-scale image quality datasets (KonIQ-10k, KADID-10k, and NIVD). Our models include country-specific components such as variable rating category thresholds and the probability for extreme ratings on these scales. We found statistically significant differences between ratings collected in different countries. Our results have implications for the analysis and design of current, respectively future datasets and contribute to a more comprehensive understanding of image quality in a global context. We also propose to include lapse rates into statistical models for categorical judgements. Lapse rates model unintentional erroneous responses of subjects in a quality assessment study and provide a regularization mechanism for the scale estimation by maximum likelihood estimation.
Dietmar Saupe, Simon Hviid Del Pin
QoMEX1
2024 Maximum Entropy and Quantized Metric Models for Absolute Category Ratings
abstract
The datasets of most image quality assessment studies contain ratings on a categorical scale with five levels, from bad (1) to excellent (5). For each stimulus, the number of ratings from 1 to 5 is summarized and given in the form of the mean opinion score. In this study, we investigate families of multinomial probability distributions parameterized by mean and variance that are used to fit the empirical rating distributions. To this end, we consider quantized metric models based on continuous distributions that model perceived stimulus quality on a latent scale. The probabilities for the rating categories are determined by quantizing the corresponding random variables using threshold values. Furthermore, we introduce a novel discrete maximum entropy distribution for a given mean and variance. We compare the performance of these models and the state of the art given by the generalized score distribution for two large data sets, KonIQ-10k and VQEG HDTV. Given an input distribution of ratings, our fitted two-parameter models predict unseen ratings better than the empirical distribution. In contrast to empirical distributions of absolute category ratings and their discrete models, our continuous models can provide fine-grained estimates of quantiles of quality of experience that are relevant to service providers to satisfy a certain fraction of the user population.
Dietmar Saupe, Krzysztof Rusek, David Hägele, Daniel Weiskopf, Lucjan Janowski
IEEE Signal Process. Lett.1
2024 Crowdsourced Estimation of Collective Just Noticeable Difference for Compressed Video With the Flicker Test and QUEST+
abstract
The concept of videowise just noticeable difference (JND) was recently proposed for determining the lowest bitrate at which a source video can be compressed without perceptible quality loss with a given probability. This bitrate is usually obtained from estimates of the satisfied used ratio (SUR) at different encoding quality parameters. The SUR is the probability that the distortion corresponding to the quality parameter is not noticeable. Commonly, the SUR is computed experimentally by estimating the subjective JND threshold of each subject using a binary search, fitting a distribution model to the collected data, and creating the complementary cumulative distribution function of the distribution. The subjective tests consist of paired comparisons between the source video and compressed versions. However, as shown in this paper, this approach typically overestimates or underestimates the SUR. To address this shortcoming, we directly estimate the SUR function by considering the entire population as a collective observer. In our method, the subject for each paired comparison is randomly chosen, and a state-of-the-art Bayesian adaptive psychometric method (QUEST+) is used to select the compressed video in the paired comparison. Our simulations show that this collective method yields more accurate SUR results using fewer comparisons than traditional methods. We also perform a subjective experiment to assess the JND and SUR for compressed video. In the paired comparisons, we apply a flicker test that compares a video interleaving the source video and its compressed version with the source video. Analysis of the subjective data reveals that the flicker test provides, on average, greater sensitivity and precision in the assessment of the JND threshold than does the usual test, which compares compressed versions with the source video. Using crowdsourcing and the proposed approach, we build a JND dataset for 45 source video sequences that are encoded with both advanced video coding (AVC) and versatile video coding (VVC) at all available quantization parameters. Our dataset and the source code have been made publicly available at http://database.mmsp-kn.de/flickervidset-database.html.
Mohsen Jenadeleh, Raouf Hamzaoui, Ulf-Dietrich Reips, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.4
2024 Going the Extra Mile in Face Image Quality Assessment: A Novel Database and Model
abstract
An accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer vision research, the lack of face IQA data and models limits the precision of current IQA metrics on face image processing tasks such as face superresolution, face enhancement, and face editing. To narrow this gap, in this article, we first introduce the largest annotated IQA database developed to date, which contains 20,000 human faces – an order of magnitude larger than all existing rated datasets of faces – of diverse individuals in highly varied circumstances. Based on the database, we further propose a novel deep learning model to accurately predict face image quality, which, for the first time, explores the use of generative priors for IQA. By taking advantage of rich statistics encoded in well pretrained off-the-shelf generative models, we obtain generative prior information and use it as latent references to facilitate blind IQA. The experimental results demonstrate both the value of the proposed dataset for face IQA and the superior performance of the proposed model.
Shaolin Su, Hanhe Lin, Vlad Hosu, Oliver Wiedemann, Jinqiu Sun, Yu Zhu 0004, Hantao Liu, Yanning Zhang 0001, Dietmar Saupe
IEEE Trans. Multim.9
2023 Localization of Just Noticeable Difference for Image Compression
abstract
The just noticeable difference (JND) is the minimal difference between stimuli that can be detected by a person. The picture-wise just noticeable difference (PJND) for a given reference image and a compression algorithm represents the minimal level of compression that causes noticeable differences in the reconstruction. These differences can only be observed in some specific regions within the image, dubbed as JND-critical regions. Identifying these regions can improve the development of image compression algorithms. Due to the fact that visual perception varies among individuals, determining the PJND values and JND-critical regions for a target population of consumers requires subjective assessment experiments involving a sufficiently large number of observers. In this paper, we propose a novel framework for conducting such experiments using crowdsourcing. By applying this framework, we created a novel PJND dataset, KonJND++, consisting of 300 source images, compressed versions thereof under JPEG or BPG compression, and an average of 43 ratings of PJND and 129 self-reported locations of JND-critical regions for each source image. Our experiments demonstrate the effectiveness and reliability of our proposed framework, which is easy to be adapted for collecting a large-scale dataset. The source code and dataset are available at https://github.com/angchen-dev/LocJND.
Guangan Chen, Hanhe Lin, Oliver Wiedemann, Dietmar Saupe
QoMEX4
2023 Relaxed forced choice improves performance of visual quality assessment methods
abstract
In image quality assessment, a collective visual quality score for an image or video is obtained from the individual ratings of many subjects. One commonly used format for these experiments is the two-alternative forced choice method. Two stimuli with the same content but differing visual quality are presented sequentially or side-by-side. Subjects are asked to select the one of better quality, and when uncertain, they are required to guess. The relaxed alternative forced choice format aims to reduce the cognitive load and the noise in the responses due to the guessing by providing a third response option, namely, “not sure”. This work presents a large and comprehensive crowdsourcing experiment to compare these two response formats: the one with the “not sure” option and the one without it. To provide unambiguous ground truth for quality evaluation, subjects were shown pairs of images with differing numbers of dots and asked each time to choose the one with more dots. Our crowdsourcing study involved 254 participants and was conducted using a within-subject design. Each participant was asked to respond to 40 pair comparisons with and without the “not sure” response option and completed a questionnaire to evaluate their cognitive load for each testing condition. The experimental results show that the inclusion of the “not sure” response option in the forced choice method reduced mental load and led to models with better data fit and correspondence to ground truth. We also tested for the equivalence of the models and found that they were different. The dataset is available at http://database.mmsp-kn.de/cogvqa-database.html.
Mohsen Jenadeleh, Johannes Zagermann, Harald Reiterer, Ulf-Dietrich Reips, Raouf Hamzaoui, Dietmar Saupe
QoMEX6
2023 JPEG AIC-3 Dataset: Towards Defining the High Quality to Nearly Visually Lossless Quality Range
abstract
Visual data play a crucial role in modern society, and the rate at which images and videos are acquired, stored, and exchanged every day is rapidly increasing. Image compression is the key technology that enables storing and sharing of visual content in an efficient and cost-effective manner, by removing redundant and irrelevant information. On the other hand, image compression often introduces undesirable artifacts that reduce the perceived quality of the media. Subjective image quality assessment experiments allow for the collection of information on the visual quality of the media as perceived by human observers, and therefore quantifying the impact of such distortions. Nevertheless, the most commonly used subjective image quality assessment methodologies were designed to evaluate compressed images with visible distortions, and therefore are not accurate and reliable when evaluating images having higher visual qualities. In this paper, we present a dataset of compressed images with quality levels that range from high to nearly visually lossless, with associated quality scores in JND units. The images were subjectively evaluated by expert human observers, and the results were used to define the range from high to nearly visually lossless quality. The dataset is made publicly available to researchers, providing a valuable resource for the development of novel subjective quality assessment methodologies or compression methods that are more effective in this quality range.
Michela Testolina, Vlad Hosu, Mohsen Jenadeleh, Davi Lazzarotto, Dietmar Saupe, Touradj Ebrahimi
QoMEX5
2022 Crowdsourced Quality Assessment of Enhanced Underwater Images - a Pilot Study
abstract
Underwater image enhancement (UIE) is essential for a high-quality underwater optical imaging system. While a number of UIE algorithms have been proposed in recent years, there is little study on image quality assessment (IQA) of enhanced underwater images. In this paper, we conduct the first crowdsourced subjective IQA study on enhanced underwater images. We chose ten state-of-the-art UIE algorithms and applied them to yield enhanced images from an underwater image benchmark. Their latent quality scales were reconstructed from pair comparison. We demonstrate that the existing IQA metrics are not suitable for assessing the perceived quality of enhanced underwater images. In addition, the overall performance of 10 UIE algorithms on the benchmark is ranked by the newly proposed simulated pair comparison of the methods.
Hanhe Lin, Hui Men, Yijun Yan, Jinchang Ren, Dietmar Saupe
QoMEX5
2022 TranSalNet: Towards perceptually relevant visual saliency prediction
abstract
Convolutional neural networks (CNNs) have significantly advanced computational modelling for saliency prediction. However, accurately simulating the mechanisms of visual attention in the human cortex remains an academic challenge. It is critical to integrate properties of human vision into the design of CNN architectures, leading to perceptually more relevant saliency prediction. Due to the inherent inductive biases of CNN architectures, there is a lack of sufficient long-range contextual encoding capacity. This hinders CNN-based saliency models from capturing properties that emulate viewing behaviour of humans. Transformers have shown great potential in encoding long-range information by leveraging the self-attention mechanism. In this paper, we propose a novel saliency model that integrates transformer components to CNNs to capture the long-range contextual visual information. Experimental results show that the transformers provide added value to saliency prediction, enhancing its perceptual relevance in the performance. Our proposed saliency model using transformers has achieved superior results on public benchmarks and competitions for saliency prediction models. The source code of our proposed saliency model TranSalNet is available at: https://github.com/LJOVO/TranSalNet.
Jianxun Lou, Hanhe Lin, David Marshall 0001, Dietmar Saupe, Hantao Liu
Neurocomputing4
2022 Large-Scale Crowdsourced Subjective Assessment of Picturewise Just Noticeable Difference
abstract
The picturewise just noticeable difference (PJND) for a given image, compression scheme, and subject is the smallest distortion level that the subject can perceive when the image is compressed with this compression scheme. The PJND can be used to determine the compression level at which a given proportion of the population does not notice any distortion in the compressed image. To obtain accurate and diverse results, the PJND must be determined for a large number of subjects and images. This is particularly important when experimental PJND data are used to train deep learning models that can predict a probability distribution model of the PJND for a new image. To date, such subjective studies have been carried out in laboratory environments. However, the number of participants and images in all existing PJND studies is very small because of the challenges involved in setting up laboratory experiments. To address this limitation, we develop a framework to conduct PJND assessments via crowdsourcing. We use a new technique based on slider adjustment and a flicker test to determine the PJND. A pilot study demonstrated that our technique could decrease the study duration by 50% and double the perceptual sensitivity compared to the standard binary search approach that successively compares a test image side by side with its reference image. Our framework includes a robust and systematic scheme to ensure the reliability of the crowdsourced results. Using 1,008 source images and distorted versions obtained with JPEG and BPG compression, we apply our crowdsourcing framework to build the largest PJND dataset, KonJND-1k (Konstanz just noticeable difference 1k dataset). A total of 503 workers participated in the study, yielding 61,030 PJND samples that resulted in an average of 42 samples per source image. The KonJND-1k dataset is available athttp://database.mmsp-kn.de/konjnd-1k-database.html
Hanhe Lin, Guangan Chen, Mohsen Jenadeleh, Vlad Hosu, Ulf-Dietrich Reips, Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.7
2021 KonIQ++: Boosting No-Reference Image Quality Assessment in the Wild by Jointly Predicting Image Quality and Defects
Shaolin Su, Vlad Hosu, Hanhe Lin, Yanning Zhang 0001, Dietmar Saupe
BMVC5
2020 Deep Learning VS. Traditional Algorithms for Saliency Prediction of Distorted Images
abstract
Saliency has been widely studied in relation to image quality assessment (IQA). The optimal use of saliency in IQA metrics, however, is nontrivial and largely depends on whether saliency can be accurately predicted for images containing various distortions. Although tremendous progress has been made in saliency modelling, very little is known about whether and to what extent state-of-the-art methods are beneficial for saliency prediction of distorted images. In this paper, we analyse the ability of deep learning versus traditional algorithms in predicting saliency, based on an IQA-aware saliency benchmark, the SIQ288 database. Building off the variations in model performance, we make recommendations for model selections for IQA applications.
Hanhe Lin, Dietmar Saupe, Hantao Liu
ICIP4
2020 ATQAM/MAST'20: Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends
abstract
The Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/ MAST) aims to bring together researchers and professionals working in fields ranging from computer vision, multimedia computing, multimodal signal processing to psychology and social sciences. It is divided into two tracks: ATQAM and MAST. ATQAM track: Visual quality assessment techniques can be divided into image and video technical quality assessment (IQA and VQA, or broadly TQA) and aesthetics quality assessment (AQA). While TQA is a long-standing field, having its roots in media compression, AQA is relatively young. Both have received increased attention with developments in deep learning. The topics have mostly been studied separately, even though they deal with similar aspects of the underlying subjective experience of media. The aim is to bring together individuals in the two fields of TQA and AQA for the sharing of ideas and discussions on current trends, developments, issues, and future directions. MAST track: The research area of media content analytics has been traditionally used to refer to applications involving inference of higher-level semantics from multimedia content. However, multimedia is typically created for human consumption, and we believe it is necessary to adopt a human-centered approach to this analysis, which would not only enable a better understanding of how viewers engage with content but also how they impact each other in the process.
Tanaya Guha, Vlad Hosu, Dietmar Saupe, Bastian Goldlücke, Naveen Kumar 0004, Weisi Lin, Victor R. Martinez, Krishna Somandepalli, Shri Narayanan, Wen-Huang Cheng, Kree Cole-McLaughlin, Hartwig Adam, John See, Lai-Kuan Wong
ACM Multimedia3
2020 Visual Quality Assessment for Interpolated Slow-Motion Videos Based on a Novel Database
abstract
Professional video editing tools can generate slow-motion video by interpolating frames from video recorded at a standard frame rate. Thereby the perceptual quality of such interpolated slow-motion videos strongly depends on the underlying interpolation techniques. We built a novel benchmark database that is specifically tailored for interpolated slow-motion videos (KoSMo-1k). It consists of 1,350 interpolated video sequences, from 30 different content sources, along with their subjective quality ratings from up to ten subjective comparisons per video pair. Moreover, we evaluated the performance of twelve existing full-reference (FR) image/video quality assessment (I/VQA) methods on the benchmark. In this way, we are able to show that specifically tailored quality assessment methods for interpolated slow-motion videos are needed, since the evaluated methods — despite their good performance on real-time video databases — do not give satisfying results when it comes to frame interpolation.
Hui Men, Vlad Hosu, Hanhe Lin, Andrés Bruhn, Dietmar Saupe
QoMEX5
2020 Foveated Video Coding for Real-Time Streaming Applications
abstract
Video streaming under real-time constraints is an increasingly widespread application. Many recent video encoders are unsuitable for this scenario due to theoretical limitations or run time requirements. In this paper, we present a framework for the perceptual evaluation of foveated video coding schemes. Foveation describes the process of adapting a visual stimulus according to the acuity of the human eye. In contrast to traditional region-of-interest coding, where certain areas are statically encoded at a higher quality, we utilize feedback from an eye-tracker to spatially steer the bit allocation scheme in real-time. We evaluate the performance of an H.264 based foveated coding scheme in a lab environment by comparing the bitrates at the point of just noticeable distortion (JND). Furthermore, we identify perceptually optimal codec parameterizations. In our trials, we achieve an average bitrate savings of 63.24% at the JND in comparison to the unfoveated baseline.
Oliver Wiedemann, Vlad Hosu, Hanhe Lin, Dietmar Saupe
QoMEX4
2020 KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment
abstract
Deep learning methods for image quality assessment (IQA) are limited due to the small size of existing datasets. Extensive datasets require substantial resources both for generating publishable content and annotating it accurately. We present a systematic and scalable approach to creating KonIQ-10k, the largest IQA dataset to date, consisting of 10,073 quality scored images. It is the first in-the-wild database aiming for ecological validity, concerning the authenticity of distortions, the diversity of content, and quality-related indicators. Through the use of crowdsourcing, we obtained 1.2 million reliable quality ratings from 1,459 crowd workers, paving the way for more general IQA models. We propose a novel, deep learning model (KonCept512), to show an excellent generalization beyond the test set (0.921 SROCC), to the current state-of-the-art database LIVE-in-the-Wild (0.825 SROCC). The model derives its core performance from the InceptionResNet architecture, being trained at a higher resolution than previous models (512 × 384). Correlation analysis shows that KonCept512 performs similar to having 9 subjective scores for each test image.
Vlad Hosu, Hanhe Lin, Tamás Szirányi, Dietmar Saupe
IEEE Trans. Image Process.4
2019 Effective Aesthetics Prediction With Multi-Level Spatially Pooled Features
abstract
We propose an effective deep learning approach to aesthetics quality assessment that relies on a new type of pre-trained features, and apply it to the AVA data set, the currently largest aesthetics database. While previous approaches miss some of the information in the original images, due to taking small crops, down-scaling or warping the originals during training, we propose the first method that efficiently supports full resolution images as an input, and can be trained on variable input sizes. This allows us to significantly improve upon the state of the art, increasing the Spearman rank-order correlation coefficient (SRCC) of ground-truth mean opinion scores (MOS) from the existing best reported of 0.612 to 0.756. To achieve this performance, we extract multi-level spatially pooled (MLSP) features from all convolutional blocks of a pre-trained InceptionResNet-v2 network, and train a custom shallow Convolutional Neural Network (CNN) architecture on these new features.
Vlad Hosu, Bastian Goldlücke, Dietmar Saupe
CVPR3
2019 SUR-Net: Predicting the Satisfied User Ratio Curve for Image Compression with Deep Learning
abstract
The Satisfied User Ratio (SUR) curve for a lossy image compression scheme, e.g., JPEG, characterizes the probability distribution of the Just Noticeable Difference (JND) level, the smallest distortion level that can be perceived by a subject. We propose the first deep learning approach to predict such SUR curves. Instead of the direct approach of regressing the SUR curve itself for a given reference image, our model is trained on pairs of images, original and compressed. Relying on a Siamese Convolutional Neural Network (CNN), feature pooling, a fully connected regression-head, and transfer learning, we achieved a good prediction performance. Experiments on the MCL-JCI dataset showed a mean Bhattacharyya distance between the predicted and the original JND distributions of only 0.072.
Chunling Fan, Hanhe Lin, Vlad Hosu, Yun Zhang 0002, Qingshan Jiang, Raouf Hamzaoui, Dietmar Saupe
QoMEX7
2019 KADID-10k: A Large-scale Artificially Distorted IQA Database
abstract
Current artificially distorted image quality assessment (IQA) databases are small in size and limited in content. Larger IQA databases that are diverse in content could benefit the development of deep learning for IQA. We create two datasets, the Konstanz Artificially Distorted Image quality Database (KADID-10k) and the Konstanz Artificially Distorted Image quality Set (KADIS-700k). The former contains 81 pristine images, each degraded by 25 distortions in 5 levels. The latter has 140,000 pristine images, with 5 degraded versions each, where the distortions are chosen randomly. We conduct a subjective IQA crowdsourcing study on KADID-10k to yield 30 degradation category ratings (DCRs) per image. We believe that the annotated set KADID-10k, together with the unlabelled set KADIS-700k, can enable the full potential of deep learning based IQA methods by means of weakly-supervised learning.
Hanhe Lin, Vlad Hosu, Dietmar Saupe
QoMEX3
2019 Visual Quality Assessment for Motion Compensated Frame Interpolation
abstract
Current benchmarks for optical flow algorithms evaluate the estimation quality by comparing their predicted flow field with the ground truth, and additionally may compare interpolated frames, based on these predictions, with the correct frames from the actual image sequences. For the latter comparisons, objective measures such as mean square errors are applied. However, for applications like image interpolation, the expected user's quality of experience cannot be fully deduced from such simple quality measures. Therefore, we conducted a subjective quality assessment study by crowdsourcing for the interpolated images provided in one of the optical flow benchmarks, the Middlebury benchmark. We used paired comparisons with forced choice and reconstructed absolute quality scale values according to Thurstone's model using the classical least squares method. The results give rise to a re-ranking of 141 participating algorithms w.r.t. visual quality of interpolated frames mostly based on optical flow estimation. Our re-ranking result shows the necessity of visual quality assessment as another evaluation metric for optical flow and frame interpolation benchmarks.
Hui Men, Hanhe Lin, Vlad Hosu, Daniel Maurer 0002, Andrés Bruhn, Dietmar Saupe
QoMEX6
2019 Quantifying Visual Abstraction Quality for Computer-Generated Illustrations
abstract
We investigate how the perceived abstraction quality of computer-generated illustrations is related to the number of primitives (points and small lines) used to create them. Since it is difficult to find objective functions that quantify the visual quality of such illustrations, we propose an approach to derive perceptual models from a user study. By gathering comparative data in a crowdsourcing user study and employing a paired comparison model, we can reconstruct absolute quality values. Based on an exemplary study for stippling, we show that it is possible to model the perceived quality of stippled representations based on the properties of an input image. The generalizability of our approach is demonstrated by comparing models for different stippling methods. By showing that our proposed approach also works for small lines, we demonstrate its applicability toward quantifying different representational drawing elements. Our results can be related to Weber--Fechner’s law from psychophysics and indicate a logarithmic relationship between number of rendering primitives in an illustration and the perceived abstraction quality thereof.
Marc Spicker, Franz Götz-Hahn, Thomas Lindemeier, Dietmar Saupe, Oliver Deussen
ACM Trans. Appl. Percept.4
2018 Deeprn: A Content Preserving Deep Architecture for Blind Image Quality Assessment
abstract
This paper presents a blind image quality assessment (BIQA) method based on deep learning with convolutional neural networks (CNN). Our method is trained on full and arbitrarily sized images rather than small image patches or resized input images as usually done in CNNs for image classification and quality assessment. The resolution independence is achieved by pyramid pooling. This work is the first that applies a fine-tuned residual deep learning network (ResNet-101) to BIQA. The training is carried out on a new and very large, labeled dataset of 10, 073 images (KonIQ-10k) that contains quality rating histograms besides the mean opinion scores (MOS). In contrast to previous methods we do not train to approximate the MOS directly, but rather use the distributions of scores. Experiments were carried out on three benchmark image quality databases. The results showed clear improvements of the accuracy of the estimated MOS values, compared to current state-of-the-art algorithms. We also report on the quality of the estimation of the score distributions.
Domonkos Varga, Dietmar Saupe, Tamás Szirányi
ICME2
2018 Expertise screening in crowdsourcing image quality
abstract
We propose a screening approach to find reliable and effectively expert crowd workers in image quality assessment (IQA). Our method measures the users' ability to identify image degradations by using test questions, together with several relaxed reliability checks. We conduct multiple experiments, obtaining reproducible results with a high agreement between the expertise-screened crowd and the freelance experts of 0.95 Spearman rank order correlation (SROCC), with one restriction on the image type. Our contributions include a reliability screening method for uninformative users, a new type of test questions that rely on our proposed database1of pristine and artificially distorted images, a group agreement extrapolation method and an analysis of the crowdsourcing experiments.
Vlad Hosu, Hanhe Lin, Dietmar Saupe
QoMEX3
2018 Spatiotemporal Feature Combination Model for No-Reference Video Quality Assessment
abstract
One of the main challenges in no-reference video quality assessment is temporal variation in a video. Methods typically were designed and tested on videos with artificial distortions, without considering spatial and temporal variations simultaneously. We propose a no-reference spatiotemporal feature combination model which extracts spatiotemporal information from a video, and tested it on a database with authentic distortions. Comparing with other methods, our model gave satisfying performance for assessing the quality of natural videos.
Hui Men, Hanhe Lin, Dietmar Saupe
QoMEX3
2018 Disregarding the Big Picture: Towards Local Image Quality Assessment
abstract
Image quality has been studied almost exclusively as a global image property. It is common practice for IQA databases and metrics to quantify this abstract concept with a single number per image. We propose an approach to blind IQA based on a convolutional neural network (patchnet) that was trained on a novel set of 32,000 individually annotated patches of 64×64 pixel. We use this model to generate spatially small local quality maps of images taken from KonIQ-10k, a large and diverse in-the-wild database of authentically distorted images. We show that our local quality indicator correlates well with global MOS, going beyond the predictive ability of quality related attributes such as sharpness. Averaging of patchnet predictions already outperforms classical approaches to global MOS prediction that were trained to include global image features. We additionally experiment with a generic second-stage aggregation CNN to estimate mean opinion scores. Our latter model performs comparable to the state of the art with a PLCC of 0.81 on KonIQ-10k.
Oliver Wiedemann, Vlad Hosu, Hanhe Lin, Dietmar Saupe
QoMEX4
2017 The Konstanz natural video database (KoNViD-1k)
abstract
Subjective video quality assessment (VQA) strongly depends on semantics, context, and the types of visual distortions. Currently, all existing VQA databases include only a small number of video sequences with artificial distortions. The development and evaluation of objective quality assessment methods would benefit from having larger datasets of real-world video sequences with corresponding subjective mean opinion scores (MOS), in particular for deep learning purposes. In addition, the training and validation of any VQA method intended to be ‘general purpose’ requires a large dataset of video sequences that are representative of the whole spectrum of available video content and all types of distortions. We report our work on KoNViD-1k, a subjectively annotated VQA database consisting of 1,200 public-domain video sequences, fairly sampled from a large public video dataset, YFCC100m. We present the challenges and choices we have made in creating such a database aimed at ‘in the wild’ authentic distortions, depicting a wide variety of content.
Vlad Hosu, Franz Götz-Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li 0001, Dietmar Saupe
QoMEX8
2017 Empirical evaluation of no-reference VQA methods on a natural video quality database
abstract
No-Reference (NR) Video Quality Assessment (VQA) is a challenging task since it predicts the visual quality of a video sequence without comparison to some original reference video. Several NR-VQA methods have been proposed. However, all of them were designed and tested on databases with artificially distorted videos. Therefore, it remained an open question how well these NR-VQA methods perform for natural videos. We evaluated two popular VQA methods on our newly built natural VQA database KoNViD-1k. In addition, we found that merely combining five simple VQA-related features, i.e., contrast, colorfulness, blurriness, spatial information, and temporal information, already gave a performance about as well as those of the established NR-VQA methods. However, for all methods we found that they are unsatisfying when assessing natural videos (correlation coefficients below 0.6). These findings show that NR-VQA is not yet matured and in need of further substantial improvement.
Hui Men, Hanhe Lin, Dietmar Saupe
QoMEX3
2016 Saliency-driven image coding improves overall perceived JPEG quality
abstract
Saliency-driven image coding is well worth pursuing. Previous studies on JPEG and JPEG2000 have suggested that region-of-interest coding brings little overall benefit compared to the standard implementation. We show that our saliency-driven variable quantization JPEG coding method significantly improves perceived image quality. To validate our findings, we performed large crowdsourcing experiments involving several hundred contributors, on 44 representative images. To quantify the level of improvement, we devised an approach to equate Likert-type opinions to bitrate differences. Our saliency-driven coding showed 11% bpp average benefit over the standard JPEG.
Vlad Hosu, Franz Götz-Hahn, Oliver Wiedemann, Sung-Hwan Jung, Dietmar Saupe
PCS5
2016 Detection of Fragmented Rectangular Enclosures in Very High Resolution Remote Sensing Images
abstract
We develop an approach for the detection of ruins of livestock enclosures (LEs) in alpine areas captured by high-resolution remotely sensed images. These structures are usually of approximately rectangular shape and appear in images as faint fragmented contours in complex background. We address this problem by introducing a rectangularity feature that quantifies the degree of alignment of an optimal subset of extracted linear segments with a contour of rectangular shape. The rectangularity feature has high values not only for perfectly regular enclosures but also for ruined ones with distorted angles, fragmented walls, or even a completely missing wall. Furthermore, it has a zero value for spurious structures with less than three sides of a perceivable rectangle. We show how the detection performance can be improved by learning a linear combination of the rectangularity and size features from just a few available representative examples and a large number of negatives. Our approach allowed detection of enclosures in the Silvretta Alps that were previously unknown. A comparative performance analysis is provided. Among other features, our comparison includes the state-of-the-art features that were generated by pretrained deep convolutional neural networks (CNNs). The deep CNN features, although learned from a very different type of images, provided the basic ability to capture the visual concept of the LEs. However, our handcrafted rectangularity-size features showed considerably higher performance.
Igor Zingman, Dietmar Saupe, Otávio A. B. Penatti, Karsten Lambers
IEEE Trans. Geosci. Remote. Sens.2
2015 No-Reference Video Quality Assessment Based on Artifact Measurement and Statistical Analysis
abstract
A discrete cosine transform (DCT)-based no-reference video quality prediction model is proposed that measures artifacts and analyzes the statistics of compressed natural videos. The model has two stages: 1) distortion measurement and 2) nonlinear mapping. In the first stage, an unsigned ac band, three frequency bands, and two orientation bands are generated from the DCT coefficients of each decoded frame in a video sequence. Six efficient frame-level features are then extracted to quantify the distortion of natural scenes. In the second stage, each frame-level feature of all frames is transformed to a corresponding video-level feature via a temporal pooling, then a trained multilayer neural network takes all video-level features as inputs and outputs, a score as the predicted quality of the video sequence. The proposed method was tested on videos with various compression types, content, and resolution in four databases. We compared our model with a linear model, a support-vector-regression-based model, a state-of-the-art training-based model, and a four popular full-reference metrics. Detailed experimental results demonstrate that the results of the proposed method are highly correlated with the subjective assessments.
Kongfeng Zhu, Chengqing Li, Vijayan K. Asari, Dietmar Saupe
IEEE Trans. Circuits Syst. Video Technol.4
2014 Optimizing feature pooling and prediction models of VQA algorithms
abstract
In this paper, we propose a strategy to optimize feature pooling and prediction models of video quality assessment (VQA) algorithms with a much smaller number of parameters than methods based on machine learning, such as neural networks. Based on optimization, the proposed mapping strategy is composed of a global linear model for pooling extracted features, a simple linear model for local alignment in which local factors depend on source videos, and a non-linear model for quality calibration. Also, a reduced-reference VQA algorithm is proposed to predict the local factors from the source video. In the IRCCyN/IVC video database of content influence and the LIVE mobile video database, the performance of VQA algorithms is improved significantly by local alignment. The proposed mapping strategy with prediction of local factors outperforms one no-reference VQA metric and is comparable to one full-reference VQA metric. Thus predicting the local factors in local alignment based on video content will be a promising new approach for VQA.
Kongfeng Zhu, Marcus Barkowsky, Minmin Shen, Patrick Le Callet, Dietmar Saupe
ICIP5
2014 A morphological approach for distinguishing texture and individual features in images
Igor Zingman, Dietmar Saupe, Karsten Lambers
Pattern Recognit. Lett.2
2013 A no-reference video quality assessment based on Laplacian pyramids
abstract
This paper presents an approach to predict the quality of compressed videos with content of natural scenes. The method is focused on measuring the distortion of compressed video without reference. There are two main steps of the proposed method: measuring distortion and predicting video quality. Each frame of the distorted video sequence is first decomposed to an N-subband Laplacian pyramid, then their intra-subband and inter-subband statistical features are fully exploited. Three intra-subband features and three inter-subband features are taken as inputs of the prediction model. Its output is a single score as the predicted video quality. The performance of the proposed method is evaluated on the LIVE video database and the LIVE mobile video database. Results show that the predicted quality scores are well correlated with the mean opinion scores associated to the subjective assessment.
Kongfeng Zhu, Keigo Hirakawa, Vijayan K. Asari, Dietmar Saupe
ICIP4
2012 Morphological operators for segmentation of high contrast textured regions in remotely sensed imagery
abstract
We develop a transformation based on morphological filters that measures the contrast of image texture. This transformation is proportional to texture contrast, but insensitive to its specific type. Though the transformation provides a high response in textured areas, it suppresses individual high contrast features that stand apart from textured areas. It can serve as an effective texture descriptor for unsupervised or supervised segmentation of textured regions, provides high accuracy of localization and does not involve heavy computations. The method is robust to variations of illumination and works on different types of images without needing to be tuned. The only parameter is a scale related parameter. We illustrate the use of the proposed method on satellite and aerial images.
Igor Zingman, Dietmar Saupe, Karsten Lambers
IGARSS2
2011 Recovering missing coefficients in DCT-transformed images
abstract
A general method for recovering missing DCT coefficients in DCT-transformed images is presented in this work. We model the DCT coefficients recovery problem as an optimization problem and recover all missing DCT coefficients via linear programming. The visual quality of the recovered image gradually decreases as the number of missing DCT coefficients increases. For some images, the quality is surprisingly good even when more than 10 most significant DCT coefficients are missing. When only the DC coefficient is missing, the proposed algorithm outperforms existing methods according to experimental results conducted on 200 test images. The proposed recovery method can be used for cryptanalysis of DCT based selective encryption schemes and other applications.
Shujun Li 0001, Andreas Karrenbauer, Dietmar Saupe, C.-C. Jay Kuo
ICIP3
2010 An improved DC recovery method from AC coefficients of DCT-transformed images
abstract
Motivated by the work of Uehara et al. [1], an improved method to recover DC coefficients from AC coefficients of DCT-transformed images is investigated in this work, which finds applications in cryptanalysis of selective multimedia encryption. The proposed under/over-flow rate minimization (FRM) method employs an optimization process to get a statistically more accurate estimation of unknown DC coefficients, thus achieving a better recovery performance. It was shown by experimental results based on 200 test images that the proposed DC recovery method significantly improves the quality of most recovered images in terms of the PSNR values and several state-of-the-art objective image quality assessment (IQA) metrics such as SSIM and MS-SSIM.
Shujun Li 0001, Junaid Jameel Ahmad, Dietmar Saupe, C.-C. Jay Kuo
ICIP3
2010 Spectral-Driven Isometry-Invariant Matching of 3D Shapes
Mauro R. Ruggeri, Giuseppe Patanè 0001, Michela Spagnuolo, Dietmar Saupe
Int. J. Comput. Vis.4
2010 Evaluation of texture registration by epipolar geometry
Ioan Cleju, Dietmar Saupe
Vis. Comput.2
2009 Highly-Automatic MI Based Multiple 2D/3D Image Registration Using Self-initialized Geodesic Feature Correspondences
Hongwei Zheng 0001, Ioan Cleju, Dietmar Saupe
ACCV (3)3
2008 Image-Based Surface Compression
abstract
Abstract We present a generic framework for compression of densely sampled three‐dimensional (3D) surfaces in order to satisfy the increasing demand for storing large amounts of 3D content. We decompose a given surface into patches that are parameterized as elevation maps over planar domains and resampled on regular grids. The resulting shaped images are encoded using a state‐of‐the‐art wavelet image coder. We show that our method is not only applicable to mesh‐ and point‐based geometry, but also outperforms current surface encoders for both primitives.
Tilo Ochotta, Dietmar Saupe
Comput. Graph. Forum2
2006 Compression of textured surfaces represented as surfel sets
Tal Darom, Mauro R. Ruggeri, Dietmar Saupe, Nahum Kiryati
Signal Process. Image Commun.3
2005 Processing of textured surfaces represented as surfel sets: representation, compression and geodesic paths
abstract
A method for representation and lossy compression of textured surfaces is presented. The input surfaces are represented by surfels (surface elements), i.e., by a set of colored, oriented, and sized disks. The position and texture of each surfel are mapped onto a sphere. The mapping is optimized for preservation of geodesic distances. The components of the resulting spherical vector-valued function are decorrelated by the Karhunen-Loeve transform and represented by spherical wavelets. Successful representation and reconstruction is demonstrated. Methods for geodesic distance computation on surfaces represented by surfels are presented.
Tal Darom, Mauro R. Ruggeri, Dietmar Saupe, Nahum Kiryati
ICIP (1)3
2004 Using entropy impurity for improved 3D object similarity search
abstract
Similarity search in 3D object databases is becoming an important problem in multimedia retrieval, with many practical applications. We investigate methods for improving the effectiveness in a retrieval system that implements multiple feature extraction algorithms to choose from. Our techniques are based on the entropy impurity measure, widely used in the context of decision trees. We propose a method for the a priori estimation of individual feature vector performance, given a query. We then define two approaches that use this estimator to improve the retrieval effectiveness. Our experimental results show that significant improvements are achievable using these methods.
Benjamin Bustos, Daniel A. Keim, Dietmar Saupe, Tobias Schreck, Dejan V. Vranic
ICME3
2004 Image based rendering of iterated function systems
Jarke J. van Wijk, Dietmar Saupe
Comput. Graph.2
2003 Influence of channel fluctuations on optimal real-time scalable image transmission
abstract
Joint source-channel coding systems using scalable source codes and forward error correction allow reliable transmission of multimedia data over noisy channels. The performance of such systems highly depends on the source-channel bit allocation strategy. Rate-based error protection schemes, which maximize the expected source rate are very attractive for real-time applications because the optimization can be done very quickly and is independent of the source. In real-world communication, channel conditions are varying in time. Thus, it is important to frequently update the error protection. For two state-of-the-art joint source-channel coding systems, we show that a channel mismatch can lead to a poor performance. We study theoretically and experimentally the dependency of a rate-based optimal protection on the channel statistics and provide an efficient strategy for adjusting the error protection when a channel mismatch occurs.
Vladimir Stankovic 0001, Raouf Hamzaoui, Dietmar Saupe
ICIP (1)3
2003 Model-based real-time progressive transmission of images over noisy channels
abstract
Many unequal error protection algorithms used in image communication systems need the operational distortion-rate (D/R) curve of the source coder whose computation is time-consuming. We study the use of parametric models instead of the true D/R curves for wavelet-based embedded image and video coders. We propose a Weibull model and show its superiority to the previous models for real-time applications. For unequal error protection over binary symmetric and packet erasure channels, the Weibull model yielded performance similar to the one obtained with the true D/R curve while satisfying the real-time constraint.
Youssef Charfi, Raouf Hamzaoui, Dietmar Saupe
WCNC3
2003 Fast algorithm for rate-based optimal error protection of embedded codes
abstract
Embedded image codes are very sensitive to channel noise because a single bit error can lead to an irreversible loss of synchronization between the encoder and the decoder. P.G. Sherwood and K. Zeger (see IEEE Signal Processing Lett., vol.4, p.191-8, 1997) introduced a powerful system that protects an embedded wavelet image code with a concatenation of a cyclic redundancy check coder for error detection and a rate-compatible punctured convolutional coder for error correction. For such systems, V. Chande and N. Farvardin (see IEEE J. Select. Areas Commun., vol.18, p.850-60, 2000) proposed an unequal error protection strategy that maximizes the expected number of correctly received source bits subject to a target transmission rate. Noting that an optimal strategy protects successive source blocks with the same channel code, we give an algorithm that accelerates the computation of the optimal strategy of Chande and Farvardin by finding an explicit formula for the number of occurrences of the same channel code. Experimental results with two competitive channel coders and a binary symmetric channel showed that the speed-up factor over the approach of Chande and Farvardin ranged from 2.82 to 44.76 for transmission rates between 0.25 and 2 bits per pixel.
Vladimir Stankovic 0001, Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Commun.3
2002 Description of 3D-shape using a complex function on the sphere
abstract
We propose a novel feature vector suitable for searching collections of 3D-object images by shape similarity. In this search, a polygonal mesh model serves as a query. For each model, feature vectors are automatically extracted and stored. Shape similarity between 3D-objects in the search space is determined by finding and ranking nearest neighbors in the feature vector space. Ranked objects are retrieved for inspection, selection, and processing. The feature vector is obtained by forming a complex function on the sphere. Afterwards, we apply the fast Fourier transform (FFT) on the sphere and obtain Fourier coefficients for spherical harmonics. The absolute values of the coefficients form the feature vector. Retrieval efficiency of the new approach is evaluated by constructing precision/recall diagrams and using two different 3D-model databases. We compared the approach with two methods based on real functions on the sphere. Our empirical comparison showed that the complex feature vector performed best. We also prepared a Web-based retrieval system for testing the methods discussed.
Dejan V. Vranic, Dietmar Saupe
ICME (1)2
2001 Progressive fractal coding
abstract
Progressive coding is an important feature of compression schemes. Wavelet coders are well suited for this purpose because the wavelet coefficients can be naturally ordered according to decreasing importance. Progressive fractal coding is feasible, but it was proposed only for hybrid fractal-wavelet schemes. We introduce a progressive fractal image coder in the spatial domain. A Lagrange optimization based on rate-distortion performance estimates determines an optimal ordering of the code bits. The optimality is in the sense that the reconstruction error is monotonically decreasing and minimum at intermediate rates. The decoder recovers this ordering without side information. As a side effect, our work motivates improved bit allocation strategies for fractal coding.
Iván Kopilovic, Dietmar Saupe, Raouf Hamzaoui
ICIP (1)2
2001 Rate-distortion unequal error protection for fractal image codes
abstract
Fractal image codes are very sensitive to bit errors because the decoding of a block is dependent not only on the code information associated to this block but also on the code information associated to other blocks. We analyze the sensitivity of a fractal code to transmission errors in a binary symmetric channel. We provide two rate-distortion unequal error protection techniques that allocate the code bits to protection classes in a nearly optimal way. We give an implementation for BCH and RCPC channel codes and show that rate-compatible punctured convolutional (RCPC) codes are preferable. For a binary symmetric channel bit error probability of 0.1 and a total code rate of 0.5 bpp, the loss in reconstruction quality with our best implementation was about 3.38 dB for the 512 /spl times/ 512 Lenna image, yielding a PSNR of 27.12 dB.
Vladimir Stankovic 0001, Dietmar Saupe, Raouf Hamzaoui
ICIP (1)2
2001 Tools for 3D-object retrieval: Karhunen-Loeve transform and spherical harmonics
abstract
We present tools for 3D object retrieval in which a model, a polygonal mesh, serves as a query and similar objects are retrieved from a collection of 3D objects. Algorithms proceed first by a normalization step (pose estimation) in which models are transformed into a canonical coordinate frame. Second, feature vectors are extracted and compared with those derived from normalized models in the search space. Using a metric in the feature vector space nearest neighbors are computed and ranked. Objects thus retrieved are displayed for inspection, selection, and processing. For the pose estimation we introduce a modified Karhunen-Loeve transform that takes into account not only vertices or polygon centroids from the 3D models but all points in the polygons of the objects. Some feature vectors can be regarded as samples of functions on the 2-sphere. We use Fourier expansions of these functions as uniform representations allowing embedded multi-resolution feature vectors. Our implementation demonstrates and visualizes these tools.
Dejan V. Vranic, Dietmar Saupe
MMSP2
2001 Rapid high quality compression of volume data for visualization
abstract
Volume data sets resulting from, e.g., computerized tomography (CT) or magnetic resonance (MR) imaging modalities require enormous storage capacity even at moderate resolution levels. Such large files may require compression for processing in CPU memory which, however, comes at the cost of decoding times and some loss in reconstruction quality with respect to the original data. For many typical volume visualization applications (rendering of volume slices, subvolumes of interest, or isosurfaces) only a part of the volume data needs to be decoded. Thus, efficient compression techniques are needed that provide random access and rapid decompression of arbitrary parts the volume data. We propose a technique which is block based and operates in the wavelet transformed domain. We report performance results which compare favorably with previously published methods yielding large reconstruction quality gains from about 6 to 12 dB in PSNR for a5123 -volume extracted from the Visible Human data set. In terms of compression our algorithm compressed the data 6 times as much as the previous state-of-the-art block based coder for a given PSNR quality.
Ky Giang Nguyen, Dietmar Saupe
Comput. Graph. Forum2
2001 Distortion Minimization with Fast Local Search for Fractal Image Compression
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
J. Vis. Commun. Image Represent.2
2000 RD-Optimization of Hierarchical Structured Adaptive Vector Quantization for Video Coding
abstract
Summary form only given. This poster contains two contributions to very-low-bitrate video coding. First, we show that in contrast to common practice incremental techniques for rate-distortion optimization such as the generalized BFOS algorithm may clearly outperform the standard technique based on Lagrangian multipliers. This is relevant in cases where the computation of RD points has a low complexity. An implementation independent performance measure is used for comparison and run-time experiments are provided. Second, we report on recent progress of our ongoing research evaluating the prospects of adaptive vector quantization (AVQ) for very-low-bitrate video coding. In contrast to conventional state-of-the-art video coding based on entropy coding of motion compensated residual frames in the frequency domain, adaptive vector quantization offers the potential to adapt its codebooks to the changing statistics of image sequences. The basic building blocks of our current AVQ video codec are: (1) block-based coding in the wavelet domain where wavelet coefficients correspond to (overlapping) spatial regions; (2) hierarchical organization of the wavelet coefficients using quadtree structures; (3) three way coding mode decision for each block (block replenishment, product code vector quantization, new VQ block with codebook update); and (4) rigorous rate/distortion optimization for all coding choices (image partition and block coding mode). This video codec does not apply motion compensation, however. A comparison with standard transform coding (H.263) shows that in spite of the improvements of our coder over previously published AVQ video coders it still shows a performance gap of about 1 dB for some test sequences.
Marcel Wagner, Dietmar Saupe
Data Compression Conference2
2000 Fast Code Enhancement with Local Search for Fractal Image Compression
abstract
Optimal fractal coding consists of finding, in a finite set of contractive affine mappings, one whose unique fixed point is closest to the original image. Optimal fractal coding is an NP-hard combinatorial optimization problem. Conventional coding is based on a greedy suboptimal algorithm known as collage coding. In a previous study, we proposed a local search algorithm that significantly improves on collage coding. However the algorithm, which requires the computation of many fixed points, is computationally expensive. In this paper we provide techniques that drastically reduce the time complexity of the algorithm.
Raouf Hamzaoui, Dietmar Saupe, Michael Hiller
ICIP2
2000 Adaptive Post-Processing for Fractal Image Compression
abstract
Fractal image compression is a block based technique. The blocks of the underlying image partition are called range blocks. Like all block based techniques, fractal coding suffers from blocking artifacts, in particular at high compression ratios. Moreover, due to the recursive nature of the fractal decoder blocking artifacts also show up in the interior of the range blocks. This paper proposes an effective post-processing to reduce both kinds of these artifacts. We apply a smoothing filter adapted to the partition before each iteration of the decoding process which reduces the blocking artifacts in the interior of the range blocks. After decoding the reconstructed image is enhanced by applying another adaptive filter along the range block boundaries. This filter is designed to reduce blocking artifacts while maintaining edges and texture of the original image. The experimental results show that the artifacts are reduced and significant improvements can be achieved. Our method compares favourably with previously published post-processing techniques for fractal image compression.
Ky Giang Nguyen, Dietmar Saupe
ICIP2
2000 Local iterative improvement of fractal image codes
Raouf Hamzaoui, Hannes Hartenstein, Dietmar Saupe
Image Vis. Comput.3
2000 Lossless acceleration of fractal image encoding via the fast Fourier transform
Hannes Hartenstein, Dietmar Saupe
Signal Process. Image Commun.2
2000 Combining fractal image compression and vector quantization
abstract
In fractal image compression, the code is an efficient binary representation of a contractive mapping whose unique fixed point approximates the original image. The mapping is typically composed of affine transformations, each approximating a block of the image by another block (called domain block) selected from the same image. The search for a suitable domain block is time-consuming. Moreover, the rate distortion performance of most fractal image coders is not satisfactory. We show how a few fixed vectors designed from a set of training images by a clustering algorithm accelerates the search for the domain blocks and improves both the rate-distortion performance and the decoding speed of a pure fractal coder, when they are used as a supplementary vector quantization codebook. We implemented two quadtree-based schemes: a fast top-down heuristic technique and one optimized with a Lagrange multiplier method. For the 8 bits per pixel (bpp) luminance part of the 512 x 512 Lena image, our best scheme achieved a peak-signal-to-noise ratio of 32.50 dB at 0.25 bpp.
Raouf Hamzaoui, Dietmar Saupe
IEEE Trans. Image Process.2
2000 Region-based fractal image compression
abstract
A fractal coder partitions an image into blocks that are coded via self-references to other parts of the image itself. We present a fractal coder that derives highly image-adaptive partitions and corresponding fractal codes in a time-efficient manner using a region-merging approach. The proposed merging strategy leads to improved rate-distortion performance compared to previously reported pure fractal coders, and it is faster than other state-of-the-art fractal coding methods.
Hannes Hartenstein, Matthias Ruhl, Dietmar Saupe
IEEE Trans. Image Process.3
1999 A Video Codec Based on R/D-Optimized Adaptive Vector Quantization
abstract
Summary form only given. We present a new AVQ-based video coder for very low bitrates. To encode a block from a frame, the encoder offers three modes: (1) a block from the same position in the last frame can be taken; (2) the block can be represented with a vector from the codebook; or (3) a new vector, that sufficiently represents a block, can be inserted into the codebook. For mode 2 a mean-removed VQ scheme is used. The decision on how blocks are encoded and how the codebook is updated is done in an rate-distortion (R-D) optimized fashion. The codebook of shape blocks is updated once per frame. First results for an implementation of such a scheme have been reported previously. Here we extend the method to incorporate a wavelet image transform before coding in order to enhance the compression performance. In addition the rate-distortion optimization is comprehensively discussed. Our R-D optimization is based on an efficient convex-hull computation. This method is compared to common R-D optimizations that use a Lagrangian multiplier approach. In the discussion of our R-D method we show the similarities and differences between our scheme and the generalized threshold replenishment (GTR) method of Fowler et al. (1997). Furthermore, we demonstrate that the translation of our R-D optimized AVQ into the wavelet domain leads to an improved coding performance. We present coding results that show that one can achieve the same encoding quality as with comparable standard transform coding (H.263). In addition we offer an empirical analysis of the short- and long-term behavior of the adaptive codebook. This analysis indicates that the AVQ method uses the vectors in its codebook for some kind of long-term prediction.
Marcel Wagner, Ralf Herz, Hannes Hartenstein, Raouf Hamzaoui, Dietmar Saupe
Data Compression Conference5
1998 Rate-Distortion based Video Coding with Adaptive Mean-Removed Vector Quantization
Raouf Hamzaoui, Dietmar Saupe, Marcel Wagner
ICIP (3)2
1998 Optimal Hierarchical Partitions for Fractal Image Compression
abstract
In fractal image compression a partitioning of the image is required. In this paper we discuss the construction of rate-distortion optimal partitions. We begin with a fine scale partition which gives a fractal encoding with a high bit rate and a low distortion. The partition is hierarchical, thus, corresponds to a tree. We employ a pruning strategy based on the generalized BFOS algorithm. It extracts subtrees corresponding to partitions and fractal encodings which are optimal in the rate-distortion sense. First results are included for the case of fractal encodings based on rectangular (HV) partitions. We also provide a comparison with greedy partitions based on the traditional collage error criterion or just using block variance.
Dietmar Saupe, Matthias Ruhl, Raouf Hamzaoui, Luigi Grandi, Daniele Marini
ICIP (1)1
1998 ATLAS2000 - Atlases of the future on the Internet
Matthias Friedrich, Mario Melle, Dietmar Saupe
Comput. Graph.3
1997 Quadtree Based Variable Rate Oriented Mean Shape-Gain Vector Quantization
abstract
Mean shape-gain vector quantization (MSGVQ) is extended to include negative gains and square isometries. Square isometries together with a classification technique based on average block intensities enable us to enlarge the MSGVQ codebook size without any additional storage requirements while keeping the complexity of both the codebook generation and the encoding manageable. Variable rate codes are obtained with a quadtree segmentation based on a rate-distortion criterion. Experimental results show that our scheme performs favorably when compared to previous product code techniques or quadtree based VQ methods.
Raouf Hamzaoui, Bertram Ganz, Dietmar Saupe
Data Compression Conference3
1997 VQ-encoding of luminance parameters in fractal coding schemes
abstract
This paper is concerned with the efficient storage of the luminance parameters in a fractal code by means of vector quantization (VQ). For a given image block (range) the collage error as a function of the luminance parameters is a quadratic function with ellipsoid contour lines. We demonstrate how these functions should be used in an optimal codebook design algorithm leading to a non-standard VQ-scheme. In addition we present results and an evaluation of this approach. The analysis of the quadratic error functions also provides guidance for optimal scalar quantization.
Hannes Hartenstein, Dietmar Saupe, Kai Uwe Barthel
ICASSP2
1997 Adaptive Partitionings for Fractal Image Compression
abstract
In fractal image compression a partitioning of the image into ranges is required. Saupe and Ruhl (1996) proposed to find good partitionings by means of a split-and-merge process guided by evolutionary computing. In this approach ranges are connected sets of small square image blocks. Far better rate-distortion curves can be obtained as compared to traditional quadtree partitionings, however, at the expense of an increase of computing time. In this paper we show how conventional acceleration techniques and a deterministic version of the evolution reduce the time-complexity of the method without degrading the encoding quality. Furthermore, we report on techniques to improve the rate-distortion performance and evaluate the results visually.
Matthias Ruhl, Hannes Hartenstein, Dietmar Saupe
ICIP (2)3
1997 Real-Time Very Low Bit Rate Video Coding with Adaptive Mean-Removed Vector Quantization
abstract
This article examines very low bit rate video coding in real-time and software-only implementation on personal computers. The method discussed is based on frame replenishment with block coding using mean-removed vector quantization. No motion compensation is employed and the VQ codebook is small. Algorithms are developed to minimize the bit rate and to reduce the search complexity for the vector quantizer. The codebook is adaptive ensuring good overall quality encoding for head-and-shoulder image sequences. Results are provided for some test sequences and compared to some other state-of-the-art techniques.
Dietmar Saupe, B. Butz
ICIP (1)1
1997 Interactive Visualization of Implicit Surfaces with Singularities
abstract
This paper presents work on two methods for interactive visualization of implicit surfaces: physically‐based sampling using particle systems and polygonization followed by physically‐based mesh improvement which explicitly makes use of the surface‐defining equation. While most previous work applied to bounded manifolds without singularities and without boundary (topological spheres) we broaden the scope of the methods to include surfaces with such features, in particular cusp points and surface self‐intersections. These aspects are not (yet) essential for computer graphics modelling with implicit surfaces but they naturally occur in simulations of interest in mathematical visualization. In this paper we use the Kummer family of algebraic surfaces as an example.
Haripriya Rangan, Matthias Ruhl, Dietmar Saupe
Comput. Graph. Forum3
1996 VQ-enhanced fractal image compression
abstract
A novel hybrid scheme combining fractal image compression with mean-removed shape-gain vector quantization is presented. The scheme uses a small set of VQ codebook blocks as a block classifier for the domain blocks and as an alternative means of coding when able to provide a satisfying distortion. Our scheme is shown to improve the performance of conventional fractal coding in all its aspects. The rate-distortion curve is ameliorated, and both the encoding and the decoding are faster.
Raouf Hamzaoui, Dietmar Saupe
ICIP (1)3
1996 The futility of square isometries in fractal image compression
abstract
In fractal image compression an image is partitioned into a set of image blocks, called ranges. The ranges are matched with blocks taken from a codebook of filtered and subsampled domain image blocks up to an affine transformation of intensity values. It is common practise in fractal image compression to include all 8 isometric versions of a codebook block in the codebook. It is reasoned that such enlarged domain pools yield better rate-distortion curves. However, this is not a valid argument supporting the use of isometries. A fair test must compare the performance of the method using a codebook including isometries with that obtained when using a plain codebook of the same size. We have performed such analysis and our results show that codebooks with isometries offer no advantages in terms of fidelity-in contrast to the prevalent belief. A similar study is carried out for the effects of including respectively excluding negative scaling factors in the fractal code.
Dietmar Saupe
ICIP (1)1
1996 Lossless acceleration of fractal image compression by fast convolution
abstract
In fractal image compression the encoding step is computationally expensive. We present a new technique for reducing the computational complexity. It is lossless, i.e., it does not sacrifice any image quality for the sake of the speedup. It is based on a codebook coherence characteristic to fractal image compression and leads to a novel application of the fast Fourier transform-based convolution. The method provides a new conceptual view of fractal image compression. This paper focuses on the implementation issues and presents the first empirical experiments analyzing the performance benefits of the convolution approach to fractal image compression depending on image size, range size, and codebook size. The results show acceleration factors for large ranges up to 23 (larger factors possible), outperforming all other currently known lossless acceleration methods for such range sizes.
Dietmar Saupe, Hannes Hartenstein
ICIP (1)1
1996 Evolutionary fractal image compression
abstract
This paper introduces evolutionary computing to fractal image compression. In fractal image compression a partitioning of the image into ranges is required. We propose to use evolutionary computing to find good partitionings. Here ranges are connected sets of small square image blocks. Populations consist of N/sub p/ configurations, each of which is a partitioning with a fractal code. In the evolution each configuration produces /spl sigma/ children who inherit their parent partitionings except for two random neighboring ranges which are merged. From the offspring the best ones are selected for the next generation population based on a fitness criterion (collage error). We show that a far better rate-distortion curve can be obtained with this approach as compared to traditional quad-tree partitionings.
Dietmar Saupe, Matthias Ruhl
ICIP (1)1
1996 A new view of fractal image compression as convolution transform coding
abstract
The article presents a new conceptual view of fractal image compression. It is based on image cross-correlation and the codebook coherence characteristic of fractal image compression. Our interpretation provides a close relationship to transform coding. As a first application, we demonstrate a technique for accelerating fractal encoding. It is lossless; i.e., it does not sacrifice any image quality and leads to a novel application of the fast Fourier transform-based convolution.
Dietmar Saupe
IEEE Signal Process. Lett.1
1995 Accelerating Fractal Image Compression by Multi-Dimensional Nearest Neighbor Search
abstract
In fractal image compression the encoding step is computationally expensive. A large number of sequential searches through a list of domains (portions of the image) are carried out while trying to find the best match for another image portion. Our theory developed here shows that this basic procedure of fractal image compression is equivalent to multi-dimensional nearest neighbor search. This result is useful for accelerating the encoding procedure in fractal image compression. The traditional sequential search takes linear time whereas the nearest neighbor search can be organized to require only logarithmic time. The fast search has been integrated into an existing state-of-the-art classification method thereby accelerating the searches carried out in the individual domain classes. In this case we record acceleration factors from 1.3 up to 11.5 depending on image and domain pool size with negligible or minor degradation in both image quality and compression ratio. Furthermore, as compared to plain classification our method is demonstrated to be able to search through larger portions of the domain pool without increased the computation time.
Dietmar Saupe
Data Compression Conference1