Thomas Richter 0005

dblp:11/4551-5 · DBLP profile ↗
← Back
39ranked-venue papers
26as first author
8since 2021 · last 2025
0000-0001-7721-0426ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 26 first-author · 7 since 2021Databases, data management, data science and information retrieval · 14 · 12 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Compact Latent Representation for Image Compression (CLRIC)
abstract
Current image compression models often require separate models for each quality level, making them resource-intensive in terms of both training and storage. To address these limitations, we propose an innovative approach that utilizes latent variables from pre-existing trained models (such as the Stable Diffusion Variational Autoencoder) for perceptual image compression. Our method eliminates the need for distinct models dedicated to different quality levels. We employ overfitted learnable functions to compress the latent representation from the target model at any desired quality level. These overfitted functions operate in the latent space, ensuring low computational complexity, around 25.5 MAC/pixel for a forward pass on images with dimensions (1363×2048) pixels. This approach efficiently utilizes resources during both training and decoding. Our method achieves comparable perceptual quality to state-of-the-art learned image compression models while being both model-agnostic and resolution-agnostic. This opens up new possibilities for the development of innovative image compression methods.
Ayman A. Ameen, Thomas Richter 0005, André Kaup
ICIP2
2025 Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
abstract
High dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe
QoMEX9
2024 SLIC: A Learned Image Codec Using Structure and Color
abstract
We propose the structure and color based learned image codec (SLIC) in which the task of compression is split into that of luminance and chrominance. The deep learning model is built with a novel multi-scale architecture for Y and UV channels in the encoder, where the features from various stages are combined to obtain the latent representation. An autoregressive context model is employed for backward adaptation and a hyperprior block for forward adaptation. Various experiments are carried out to study and analyze the performance of the proposed model, and to compare it with other image codecs. We also illustrate the advantages of our method through the visualization of channel impulse responses, latent channels and various ablation studies. The model achieves Bjøntegaard delta bitrate gains of 7.5% and 4.66% in terms of MS-SSIM and CIEDE2000 metrics with respect to other state-of-the-art reference codecs.
Srivatsa Prativadibhayankaram, Mahadev Prasad Panda, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel, André Kaup
DCC3
2024 A Study on the Effect of Color Spaces in Learned Image Compression
abstract
In this work, we present a comparison between color spaces namely YUV, LAB, RGB and their effect on learned image compression. For this we use the structure and color based learned image codec (SLIC) from our prior work, which consists of two branches - one for the luminance component (Y or L) and another for chrominance components (UV or AB). However, for the RGB variant we input all 3 channels in a single branch, similar to most learned image codecs operating in RGB. The models are trained for multiple bitrate configurations in each color space. We report the findings from our experiments by evaluating them on various datasets and compare the results to state-of-the-art image codecs. The YUV model performs better than the LAB variant in terms of MSSSIM with a Bjøntegaard delta bitrate (BD-BR) gain of 7.5% using VTM intra-coding mode as the baseline. Whereas the LAB variant has a better performance than YUV model in terms of CIEDE2000 having a BD-BR gain of 8%. Overall, the RGB variant of SLIC achieves the best performance with a BD-BR gain of 13.14% in terms of MS-SSIM and a gain of 17.96% in CIEDE2000 at the cost of a higher model complexity.
Srivatsa Prativadibhayankaram, Mahadev Prasad Panda, Jürgen Seiler, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel, André Kaup
ICIP4
2024 Evaluating Visually Lossless Compression of JPEG XS, JPEG 2000, HEVC and AV1 in Selected Medical Imaging Modalities
abstract
The objective of this study is to evaluate the effectiveness of state-of-the-art codecs in compressing selected medical imaging modalities while maintaining visual quality and reducing file sizes. To achieve this, a detailed comparative analysis is conducted comparing the performance of JPEG XS, JPEG 2000, HEVC, and AV1. The analysis takes into consideration compression efficiency, codec complexity, and visual fidelity in the context of medical imaging. Advanced evaluation methods, including the AIC-2 Flicker test, are utilized to determine the visually lossless threshold, which is crucial for preserving diagnostically important details. Additionally, the study explores the potential of crowd-sourcing as a means of assessing the visual quality of compressed medical images. Subjective lab and crowd-sourcing tests reveal varying proportions of correctly identifying the reference images among participants. Furthermore, the study proposes outlier detection methods to improve the reliability of the subjective evaluation and employs kappa analysis to measure the inter-rater agreements. The study analyzes the encoding time taken on a consumer-level CPU, and the results reveal that JPEG XS maintains a fast and consistent speed across different compression levels. The results also indicate that JPEG XS achieves visually lossless performance for diagnostic purposes at 2 BPP, JPEG 2000 at 1.5 BPP, HEVC, and AV1 at 1 BPP.
Bassem Elmeligy, Thomas Richter 0005, Rakesh Rao Ramachandra Rao, Siegfried Fößel, Alexander Raake
QoMEX2
2023 Color Learning for Image Compression
abstract
Deep learning based image compression has gained a lot of momentum in recent times. To enable a method that is suitable for image compression and subsequently extended to video compression, we propose a novel deep learning model architecture, where the task of image compression is divided into two sub-tasks, learning structural information from luminance channel and color from chrominance channels. The model has two separate branches to process the luminance and chrominance components. The color difference metric CIEDE2000 is employed in the loss function to optimize the model for color fidelity. We demonstrate the benefits of our approach and compare the performance to other codecs. Additionally, the visualization and analysis of latent channel impulse response is performed.
Srivatsa Prativadibhayankaram, Thomas Richter 0005, Heiko Sparenberg, Siegfried Fößel
ICIP2
2021 JPEG XS - A New Standard for Visually Lossless Low-Latency Lightweight Image Coding
abstract
Joint Photographic Experts Group (JPEG) XS is a new International Standard from the JPEG Committee (formally known as ISO/International Electrotechnical Commission (IEC) JTC1/SC29/WG1). It defines an interoperable, visually lossless low-latency lightweight image coding that can be used for mezzanine compression within any AV market. Among the targeted use cases, one can cite video transport over professional video links (serial digital interface (SDI), internet protocol (IP), and Ethernet), real-time video storage, memory buffers, omnidirectional video capture and rendering, and sensor compression (for example, in cameras and the automotive industry). The core coding system is composed of an optional color transform, a wavelet transform, and a novel entropy encoder, processing groups of coefficients by coding their magnitude level and packing the magnitude refinement. Such a design allows for visually transparent quality at moderate compression ratios, scalable end-to-end latency that ranges from less than one line to a maximum of 32 lines of the image, and a low-complexity real-time implementation in application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), central processing unit (CPU), and graphics processing unit (GPU). This article details the key features of this new standard and the profiles and formats that have been defined so far for the various applications. It also gives a technical description of the core coding system. Finally, the latest performance evaluation results of recent implementations of the standard are presented, followed by the current status of the ongoing standardization process and future milestones.
Antonin Descampe, Thomas Richter 0005, Touradj Ebrahimi, Siegfried Fößel, Joachim Keinert, Tim Bruylants, Pascal Pellegrin, Charles Buysschaert, Gaël Rouvroy
Proc. IEEE2
2021 Bayer CFA Pattern Compression With JPEG XS
abstract
While traditional image compression algorithms take a full three-component color representation of an image as input, capturing of such images is done in many applications with Bayer CFA pattern sensors that provide only a single color information per sensor element and position. In order to avoid additional complexity at the encoder side, such CFA pattern images can be compressed directly without prior conversion to a full color image. In this paper, we describe a recent activity of the JPEG committee (ISO SC 29 WG 1) to develop such a compression algorithm in the framework of JPEG XS. It turns out that it is important to understand the "development process" from CFA patterns to full color images in order to optimize the image quality of such a compression algorithm, which we will also describe shortly. We introduce (1) a novel decorrelation step upfront processing (the so-called Star-Tetrix transform), along with (2) a pre-emphasis function to improve the compression efficiency of the subsequent compression algorithm (here, JPEG XS). Our experiments clearly indicate a gain over a RGB compression workflow in terms of complexity and quality (between 1.5dB and more than 4dB depending on the target bitrate). A comparison is also made with other state-of-the-art CFA compression techniques.
Thomas Richter 0005, Siegfried Fößel, Antonin Descampe, Gaël Rouvroy
IEEE Trans. Image Process.1
2019 Bayer Pattern Compression with JPEG XS
abstract
Image sensors in digital cameras use a technology called "Bayer Patterns" allowing color photography with a planar arrangements of photosensitive elements. Alternating arrangements of red, green and blue filter masks on top of a rectangular grid of such elements allow capturing of color information, but also require a de-mosaicing algorithm to reconstruct a full-resolution color image from the sensor data. In high-speed applications, or applications where the system design requires low-latency, low-complexity compression, the JPEG XS standard of the JPEG committee offers an elegant solution to compress Bayer pattern images close to the sensor, and to transmit the compressed data over a lower bandwidth connection while maintaining visually lossless quality. This paper presents contributions to JPEG XS that are currently under discussion in the JPEG committee (SC29WG1) as parts of an amendment for Bayer pattern compression.
Thomas Richter 0005, Siegfried Fößel
ICIP1
2018 Entropy Coding and Entropy Coding Improvements of JPEG XS
abstract
JPEG XS is a new standard for low-latency and low-complexity coding designed by the JPEG committee. Unlike former developments, optimal rate distortion performance is only a secondary goal; the focus of JPEG~XS is to enable cost-efficient, easy to parallelize implementations suitable for FPGAs or GPUs. In this article, we shed some light on the entropy coding back-end of JPEG~XS and introduce modifications of the entropy coding stage currently under discussion that improve objective and subjective quality of the compressed images without compromising the parallelism of the original algorithm.
Thomas Richter 0005, Joachim Keinert, Antonin Descampe, Gaël Rouvroy
DCC1
2018 Decoding JPEG XS on a GPU
abstract
JPEG XS is an upcoming lightweight image compression standard that is especially developed to meet the requirements of compressed video-over-IP use cases. It is designed with not only CPU, FPGA or ASIC platforms in mind, but explicitly also targets GPUs. Though not yet finished, the codec is now sufficiently mature to present a first NVIDIA CUDA-based GPU decoder architecture and preliminary performance results. On a 2014 mid-range GPU with 640 cores a 12 bit UHD 4:2:2 (4:4:4) can be decoded with 54 (42) fps. The algorithm scales very well: on a 2017 high-end GPU with 2560 cores the throughput increases to 190 (150) fps. In contrast, an optimized GPU-accelerated JPEG 2000 decoder takes 2x as long for high compression ratios that yield a PSNR of 40 dB and 3x as long for lower compression ratios with a PSNR of over 50 dB.
Volker Bruns, Thomas Richter 0005, Joachim Keinert, Siegfried Fößel
PCS2
2017 Error Bounds for HDR Image Coding with JPEG XT
abstract
ISO recently published a new image compression standard, JPEG XT, which extends the popular JPEG standard towards higher dynamic range, compression of alpha channels and lossless coding. In part 7 ofJPEG XT, a two-layer lossy image compression for HDR images isintroduced that reconstructs HDR signals by the combination of a baselayer following the legacy JPEG standard, and an extension layer thatenlarges the dynamic range to that of the original image. While bothbase and extension layer are entropy coded by the mechanisms specifiedin the legacy JPEG standard, only the base layer is visible to legacyJPEG decoders. The original JPEG standard now enforces in its part 2error bounds conforming implementations shall satisfy. The questionanswered in this article is in how far these error bounds carry overto error bounds of JPEG~XT for coding of HDR image signals.
Thomas Richter 0005
DCC1
2017 Multi-generation-robust Coding with JPEG XS
abstract
The JPEG committee (formally, ISO SC29 WG1) is currently standardizing a lightweight mezzanine codec for video over IP transport under the name JPEG XS. A particular challenging design constraint of this codec is multi-generation robustness, that is the necessity to minimize the error built-up under multiple re-compression cycles. In this paper, we discuss the sources of such errors, how they are avoided in the JPEG XS design and compare the multi-generation robustness of JPEG XS with that of other codecs.
Thomas Richter 0005, Joachim Keinert, Antonin Descampe, Gaël Rouvroy, Alexandre Willème
ISM1
2016 Fine-tuning JPEG-XT compression performance using large-scale objective quality testing
abstract
The upcoming JPEG XT standard for High Dynamic Range (HDR) images defines a common framework for the lossy and lossless representation of high-dynamic range images. It describes the decoding process as the combination of various processing tools that can be combined freely. In this paper we analyze the coding efficiency of different decoding tools through a large scale objective quality testing using the HDR-VDP 2.2 objective metric. This evaluation is performed on a large database of 337 images, testing the effect of global and local tone mapping operators for various configurations, and for multiple combinations of quality parameters. The main findings are that using an inverse tone mapping operator for creating an HDR precursor image works well for global, but not for local operators, and that including refinement scans to increase the bit-depth of the extension layer provides substantial improvements for one of the encoding profiles and higher bit-rates.
Rafal Mantiuk, Thomas Richter 0005, Alessandro Artusi
ICIP2
2016 JPEG on STEROIDS: Common optimization techniques for JPEG image compression
abstract
Despite its age, JPEG (formally, Rec. ITU-T T.81 - ISO/IEC 10918-1) is still the omnipresent image file format for lossy compression of photographic images. While its rate-distortion performance is not competitive with state-of-the-art schemes like JPEG 2000 or HEVC, manifold techniques have been developed over the years to improve its compression performance. This article provides a short review of the known technologies and evaluates them on the basis of the JPEG XT demo implementation available on the home page of the JPEG committee. It also puts the compression gains into perspective of more modern compression formats such as JPEG 2000.
Thomas Richter 0005
ICIP1
2015 Lossless Coding Extensions for JPEG
abstract
The issue of backwards compatible image and video coding gained some attention in both MPEG and JPEG, let it be as extension for HEVC, let it be as the JPEG XT standardization initiative of the SC29WG1 committee. The coding systems work all on the principle of a base layer operating in the low-dynamic range regime, using a tone-mapped version of the HDR material as input, and an extension layer invisible to legacy applications. The extension layer allows implementations conforming to the full standard to reconstruct the original image in the high-dynamic range regime. What is also common to all approaches is the rate-allocation problem: How can one split the rate between base and extension layer to ensure optimal coding? In this work, an explicit answer is derived for a simplified model of a two-layer compression system in the high bit-rate approximation. For a HDR to LDR tone mapping that approximates the well-known sRGB non-linearity of ? = 2.4 and a Laplacian probability density function, explicit results in the form of the Lambert-W-function are derived. The theoretical results are then verified in experiments using a JPEG XT demo implementation.
Thomas Richter 0005
DCC1
2014 Rate Allocation in a Two Quantizer Coding System
abstract
The issue of backwards compatible image and video coding gained some attention in both MPEG and JPEG, let it be as extension for HEVC, let it be as the JPEG XT standardization initiative of the SC29WG1 committee. The coding systems work all on the principle of a base layer, perating in the low-dynamic range regime, using a one-mapped version of the HDR material as input, and an extension layer invisible to legacy applications. The extension layer allows implementations conforming to the full standard to reconstruct the original image in the high-dynamic range regime. What is also common to all approaches is the rate-allocation problem: How can one split the rate between base and extension layer to ensure optimal coding? In this work, an explicit answer is derived for a simplified model of a two-layer compression system in the high bit-rate approximation. For a HDR to LDR tone mapping that approximates the well-known sRGB non-linearity of gamma = 2.4 and a Laplacian probability density function, explicit results in the form of the Lambert-W-function are derived. The theoretical results are then verified in experiments using a JPEG XT demo implementation.
Thomas Richter 0005
DCC1
2013 Backwards Compatible Coding of High Dynamic Range Images with JPEG
abstract
In its Paris meeting, the JPEG committee decided to work on a backwards compatible extension of the popular JPEG (10918-1) standard enabling lossy and lossless coding of high-dynamic range (HDR) images, the new standard shall allow legacy applications to decompress new code streams into a tone mapped version of the HDR image while codecs aware of the extensions will decompress the stream with full dynamic range. This paper proposes a set of extensions that have rather low implementation complexity, and use - whenever possible - functional design blocks already present in 10918-1. It is seen that, despite its simplicity, the proposed extension performs close to JPEG 2000 (15444-2) and JPEG XR (29199-2) on the HDR test image set of the JPEG for high bit-rates.
Thomas Richter 0005
DCC1
2013 High Throughput Coding of Video Signals
abstract
As the resolution of monitors and TVs continue to increase, the available bandwidth between host system and monitor becomes more and more a bottleneck. The Video Electronics Standards Association (VESA) is currently developing standards for screen resolutions beyond 4K and, facing the problem of not having enough bandwidth available on traditional copper wires, contacted the JPEG to develop a low complexity, high-throughput still image coder for lossy transmission of video signals. This article describes two approaches to address this predicament, a simple SPIHT based coded and a Hadamard based embedded codec requiring only minimal buffering at encoder and decoder side, and avoiding any pixel-based feedback loops limiting the operating frequency of hardware implementations. Analyzing the details of both implementations reveals an interesting connection between run-length coding, as found in the progressive mode of traditional JPEG coding, and SPIHT/EZW coding - a technique popular in wavelet based compression techniques.
Thomas Richter 0005, Sven Simon 0001
DCC1
2013 A global image fidelity metric: Visual distance and its properties
abstract
The purpose of full reference image quality indices like Mean Square Error (MSE) or SSIM is to predict the judgement of human observers in subjective quality assessment tasks. More advanced indices like SSIM or VIF are, however, rarely metrics in the strict sense, i.e. they don't define a distance between pairs of images that would describe how far these images are related. It was shown in a recent work, however, that SSIM can be "integrated" to such a global metric, in the following called Visual Distance, which behaves locally like SSIM, but globally like a distance in a curved space. It was also seen that the Visual Distance between two images can be interpreted as the number of almost invisible image deformations to transform one image into another. In this work, properties of Visual Distances will be discussed; these results will allow to extend the result from SSIM to its multi-scale variant MS-SSIM. It will also seen that human judgement will typically not define a metric, but it is conjectured that scores and Visual Distances are related by a monotonie Judgement Function. If so, it will be seen that the underlying Visual Distance can always be reconstructed from the scores up to a proportionality factor defined by the scale of the score.
Thomas Richter 0005
ICIP1
2013 On the standardization of the JPEG XT image compression
abstract
The ISO recently started a standardization initiative on a forwards-compatible extension of its popular JPEG (ISO/IEC 10918-1) standard. This new standard aims at carefully extending the feature set of the existing technology while preserving the established tool chain built around its 20 years old predecessor. In this work, an overview on JPEG XT, its goals and the underlying technology will be given, and two of the proposed coding technologies for JPEG XT are compared side by side.
Thomas Richter 0005
PCS1
2013 Coding strategies and performance analysis of GPU accelerated image compression
abstract
Graphics Processing Units (GPUs) are freely programmable massively parallel general purpose processing units and thus offer the opportunity to off-load heavy computations from the CPU to the GPU. One application for GPU programming is image compression, where the massively parallel nature of GPUs promises high speed benefits. However, measurements with competative highly optimized CPU implementations show that GPU based codes are usually not considerably faster, or perform only with less than ideal rate-distortion performance. This article presents the predicaments of data-parallel image coding by first presenting a series of theoretical arguments that limit the performance of such implementations before advancing to existing GPU implementations demonstrating the challenges of parallel image coding. It will be argued and seen on experiments that either parts of the entropy coding and bitstream build-up must remain serial, or rate-distortion penalties must be paid when offloading all computations on the GPU.
Thomas Richter 0005, Sven Simon 0001
PCS1
2012 Fast and Context-Free Lossless Image Compression Algorithm Based on JPEG-LS
abstract
While the context-based entropy coding and bias cancellation steps in the JPEG-LS standard are key features to its compression performance, these steps also enlarge the memory footprint, create dependencies in the data path of implementations and hence limit parallelism in modern multi-core or GPU architectures, and the throughput in hardware implementations. In the proposed modification of JPEG-LS, such most expensive parts with respect to memory space requirements and computational complexity are omitted.
Yurij Gera, Zhe Wang 0008, Sven Simon 0001, Thomas Richter 0005
DCC4
2012 Compressing JPEG 2000 JPIP Cache State Information
abstract
JPEG 2000 part 9, or short JPIP, is an interactive image browsing protocol that allows the selective delivery of image regions, components or scales from JPEG 2000 image. Typical applications are browsing tools for medical databases where transmitting huge images from server to client in total would be uneconomical. Instead, JPIP allows extracting only the desired image parts for analysis by an http type request syntax. Such a JPIP connection may either operate in a session within which the server remains aware of the image data already cached at the client and it hence doesn't have to transmit again, or it may operate in a stateless mode in which the server has no model of the data already available on the client. In such cases, the client may include a description of its cache model within a proceeding request to avoid retransmission of data already buffered. Unfortunately, the standard defined methods how such cache models are described are very inefficient, and a single request including a cache model may grow several KBytes large for typical images and requests, making the deployment of a JPIP server on top of existing http server infrastructure rather inconvenient. In this work, a lossy and loss less embedded compression scheme for such JPIP cache model adjustment requests based on a modified zero-tree algorithm is proposed, this algorithm works even in constraint environments where request size must remain limited. The proposed algorithm losslessly compresses such cache model adjustment requests often better than by a factor of 1:8, but may even perform a 1:8000 compression in cases where the cache model has to describe a large number of precincts.
Thomas Richter 0005
DCC1
2012 On the JPEG 2000 ultrafast mode
abstract
Recently, the JPEG committee discussed the introduction of an “ultrafast” mode for JPEG 2000 encoding. This considered extension of the JPEG 2000 framework replaces the EBCOT coding by a combined Huffman-Runlength code, and adds an optional additional prediction step after quantization. While the resulting codec is not compatible with existing JPEG 2000, it still allows lossless transcoding from JPEG 2000 and back, and performance measurements show that it offers nearly the quality of JPEG 2000 and similar quality than JPEG XR at a much lower complexity comparable to the complexity of the IJG JPEG software. This work introduces the extension, and compares its performance with other JPEG standards and other extensions of JPEG 2000 currently under standardization.
Thomas Richter 0005, Sven Simon 0001
ICIP1
2012 SSPQ - spatial domain perceptual image codec based on subsampling and perceptual quantization
abstract
A spatial domain perceptual image codec based on subsampling and perceptual quantization (SSPQ) guided by the just-noticeable distortion (JND) profile is proposed. SSPQ integrates perceptual image coding and progressive transmission in one framework. The input image is first subsampled by a factor of two in both dimensions and the subsampled image is compressed without loss. The subsampled image provides a basis for both predicting the input pixels by interpolation and estimating the JND values for each pixel. Residual quantization thresholds are set to the estimated JND values for a perceptually tuned compression. Quantized residuals are progressively encoded by a context-based Golomb coder with run-length coding capacity. Experimental results show over 50% improvement in compression performance on average for the proposed SSPQ codec compared to the lossless JPEG-LS.
Zhe Wang 0008, Sven Simon 0001, Michael J. Klaiber, Silvia Ahmed, Thomas Richter 0005
ICIP5
2012 A Standardized Metadata Set for Annotation of Virtual and Remote Laboratories
abstract
Online Laboratories and Virtual Experiments start to play an increasingly important role in the education of Engineering and Science Education. While several repositories for online and virtual experiments are available, a common method for annotating experiments to simplify their discovery is not yet available and accepted. In 2010, an international group of online lab providers formed the Global Online Lab Consortium (GOLC) to address the issues of interoperability between online laboratories and laboratory compilations, one of its activities is the establishment of an ontology and a common metadata set that addresses not only the needs of typical lab providers and lab users, but also of storage and archival institutions such as libraries. This article describes the current status of the GOLC activities in the metadata subcommittee, lists the requirements of various user groups of the metadata set and provides insight into both the underlying ontology and the metadata specifications themselves.
Thomas Richter 0005, Per Pascal Grube, Danilo Garbi Zutin
ISM1
2012 Memory efficient lossless compression of image sequences with JPEG-LS and temporal prediction
abstract
In this paper, a lossless encoder for image sequences based on JPEG-LS defined for still images with temporal-extended prediction and context modeling is proposed. As embedded systems are one important field of application of the codec, on-line lossy reference frame compression is used to reduce the encoder's memory requirement. Variations of the pixel values in the reference frame due to lossy compression are acceptable since the predictor provides only estimations of the pixel values being encoded in the current frame. Larger variations decrease the final lossless compression performance of the encoder such that a trade-off between the memory requirement and the overall compression ratio is required. Different compression algorithms for the reference frame, including JPEG, JPEG 2000 and near-lossless JPEG-LS, and their impacts on the memory requirement and the overall lossless compression ratio have been studied. Experimental results show 9.6% or more gain in lossless compression ratio compared to applying the standard JPEG-LS frame-by-frame and 80% reduction in the encoder buffer size compared to storing the uncompressed reference frame.
Zhe Wang 0008, Debasish Chanda, Sven Simon 0001, Thomas Richter 0005
PCS4
2011 Deadzone Based Rate Allocation for JPEG XR
abstract
The JPEG XR image compression solely controls the image quality loss and hence the output rate by means of the quantizer bucket sizes; a precise rate control mechanism like the EBCOT rate allocation algorithm in JPEG 2000 is not specified, and hence rate-distortion optimality of the quantizer is, in general, not given. A simple rate-control mechanism for JPEG XR is introduced that allows an efficient control of the quantizer towards rate-distortion optimality. It was seen in an earlier work that the additional side information required for the spatial varying quantization mechanism of the standard almost compensates the PSNR gain and complicates the rate allocation process by requiring an additional quantizer allocation step.
Thomas Richter 0005
DCC1
2010 Spatial Constant Quantization in JPEG XR is Nearly Optimal
abstract
The JPEG XR image compression standard, originally developed under the name HD-Photo by Microsoft, offers the feature of spatial variably quantization; its codestream syntax allows to select one out of a limited set of possible quantizers per macro block and per frequency band. In this paper, an algorithm is presented that finds the rate-distortion optimal set of quantizers, and the optimal quantizer choice for each macro block. Even though it seems plausible that this feature may provide a huge improvement for images whose statistics is non-stationary, e.g. compound images, it is demonstrated that the PSNR improvement is not larger than 0.3 dB for a two-step heuristics of feasible complexity, but improvements of up to 0.8 dB for compound images are possible by a much more complex optimization strategy.
Thomas Richter 0005
DCC1
2010 On the duality of rate allocation and quality indices
abstract
In a recent work, the author proposed to study the performance of still image quality indices such as the SSIM by using them as objective function of rate allocation algorithms. The outcome of that work was not only a multi-scale SSIM optimal JPEG 2000 implementation, but also a first-order approximation of the MS-SSIM that is surprisingly similar to more traditional contrast-sensitivity and visual masking based approaches. It will be seen in this work that the only difference between the latter works and the MS-SSIM index is the choice of the exponent of the masking term, and furthermore, that a slight modification of the SSIM definition reproducing the traditional exponent is able to improve the performance of the index at or below the visual threshold. It is hence demonstrated that the duality of quality indices and rate allocation helps to improve both the visual performance of the compression codec and the performance of the index.
Thomas Richter 0005
PCS1
2010 A Comparison of Three Image Fidelity Metrics of Different Computational Principles for JPEG2000 Compressed Abdomen CT Images
abstract
This study aimed to evaluate three image fidelity metrics of different computational principles--peak signal-to-noise ratio (PSNR), high-dynamic range visual difference predictor (HDR-VDP), and multiscale structural similarity (MS-SSIM)--in measuring the fidelity of JPEG2000 compressed abdomen computed tomography images from a viewpoint of visually lossless compression. Three hundred images with 0.67- or 5-mm section thickness were compressed to one of five compression ratios ranging from reversible compression to 15:1. The fidelity of each compressed image was measured by five radiologists' visual analyses (distinguishable or indistinguishable from the original) and the three metrics. The Spearman rank correlation coefficients of the PSNR, HDR-VDP, and MS-SSIM values with the number of readers responding as indistinguishable were 0.86, 0.94, and 0.86, respectively. Using the pooled readers' responses as the reference standard, the area under the receiver-operating-characteristic curve for the HDR-VDP (0.99) was significantly greater than that for the PSNR (0.95) (p < 0.001) and for the MS-SSIM (0.96) (p = 0.003), and there was no significant difference between the PSNR and MS-SSIM (p = 0.70). In measuring the image fidelity, the HDR-VDP outperforms the PSNR and MS-SSIM, and the MS-SSIM and PSNR are comparable.
Kil Joong Kim, Bo Hyoung Kim, Rafal Mantiuk, Thomas Richter 0005, Hyunna Lee, Heung Sik Kang, Jinwook Seo, Kyoung Ho Lee
IEEE Trans. Medical Imaging4
2009 A MS-SSIM Optimal JPEG 2000 Encoder
abstract
In this work, we present a SSIM optimal JPEG 2000 rate allocation algorithm. However, our aim is less improving the visual performance of JPEG 2000, but more the study of the performance of the SSIM full reference metric by means beyond correlation measurements.Full reference image quality metrics assign a quality index to a pair of a reference and distorted image. The performance of a metric is then measured by the degree of correlation between the scores obtained from the metric and those from subjective tests. It is the aim of a rate allocation algorithm to minimize the distortion created by a lossy image compression scheme under a rate constraint.Noting this relation between objective function and performance evaluation allows us now to define an alternative approach to evaluate the usefulness of a candidate metric: we want to judge the quality of a metric by its ability to define an objective function for rate control purposes, and evaluate images compressed in this scheme subjectively. It turns out that deficiencies of image quality metrics become much easier visible - even in the literal sense - than under traditional correlation experiments.Our candidate metric in this work is the SSIM index proposed by Sheik and Bovik which is both simple enough to be implemented efficiently in rate control algorithms, but yet correlates better to visual quality than MSE; our candidate compression scheme is the highly flexible JPEG 2000 standard.
Thomas Richter 0005, Kil Joong Kim
DCC1
2009 Evaluation of floating point image compression
abstract
Recently, compression of High Dynamic Range (HDR) photography gained attention in the standardization of the Microsoft HDPhoto compression scheme as JPEG-XR. While integer data of 16 bits/pixel (bpp) in scRGB color-space can represent images up to a dynamic range of about 3.5 magnitudes in luminance, even higher ranges are more efficiently represented by floating-point number formats. In this work, the author presents the approach taken by JPEG-XR for compressing such data, and shows that this method when applied to JPEG 2000 generates equally good results. Furthermore, it is shown that the method performs nearly optimal under a mathematical quality index that is closely related to SSIM, and to the mean square error of the HDR images when rendered to the LDR regime.
Thomas Richter 0005
ICIP1
2009 Perceptual image coding by standard-constraint codecs
abstract
A perceptual image compression codec exploits the characteristics of the human senses to minimize the perceivable quality loss of digital images under compression. Such a codec has an even higher value if the resulting code-streams are compatible to an existing standard, and are thus decode-able by all-day, existing applications. This work describes strategies how to implement perceptual coding in standardized environments, namely JPEG, JPEG 2000 and JPEG-XR.
Thomas Richter 0005
PCS1
2008 Effective Visual Masking Techniques in JPEG2000
abstract
Rate allocation in the JPEG2000 image compression algorithm is performed by the EBCOT algorithm, measures file size and distortion, defined as mean square error (MSE). Since MSE correlates only mediocre to visual quality, more advanced metrics like the M-SSIM have been proposed. One exploitable effect of the human visual system is that of visual masking: If a structure of a fixed amplitude is overlayed by a texture, it becomes masked and less visible. This can be addressed in JPEG2000 by multiplying the MSE contribution of a codeblock by a factor mu computed from the neighbourhood of the data. Most of these techniques require, however, complex operations on the coefficients.
Thomas Richter 0005
DCC1
2008 Subjective and Objective Assesment of Visual Image Quality Metrics and Still Image Codecs
abstract
Summary form only given. Objective quality assessment of lossy image compression codecs have become an important part of the recent call of the JPEG committee for advanced image coding. We evaluated JPEG with Huffman and arithmetic coding option, a visual and PSNR optimal JPEG2000 version, H.264/AVC and the recently proposed HDPhoto format by Microsoft. For objective evaluation, we use a color version of the M-SSIM metric and the high-dynamic range version of VDP. The results obtained from these tests are compared to subjective testing obtained from an ordering test run by 15 observers in two sessions. Subjective results are compiled to Mean Opinion Score (MOS) and passed through a Kurtosis test to verify their validity and to reject outliers.
Thomas Richter 0005, Mohamed-Chaker Larabi
DCC1
2008 Effective visual masking techniques in JPEG200
abstract
This paper introduces a very low complexity visual masking algorithm for the JPEG2000 image compression standard and evaluates its impact on the visual image quality by means of the multi-scale SSIM index. The algorithm derives suitable weighting masks indirectly from a statistical model of the wavelet data which is defined from the second moment and the average absolute amplitude of the data. If combined with an a priori rate allocation algorithm, the computation of the visual masking weights has almost no overhead at all.
Thomas Richter 0005
ICIP1
2008 Visual quality improvement techniques of HDPhoto/JPEG-XR
abstract
Microsoft's recently proposed new image compression codec HDPhoto is currently undergoing ISO standardization as JPEG-XR. Even though performance measurements carried out by the JPEG committee indicated that the PSNR performance of HDPhoto is competitive, the visual performance of HDPhoto showed notable deficits, both in subjective and objective tests. This paper introduces various techniques that improve the visual performance of HDPhoto without leaving the current codestream definition. Objective measurements performed by the author indicate that the modified encoder, while staying backwards compatible to the current standard proposition, improves visual performance significantly, and the performance of the modified encoder is similar to JPEG.
Thomas Richter 0005
ICIP1