Søren Forchhammer

dblp:89/3548 · DBLP profile ↗
← Back
120ranked-venue papers
26as first author
13since 2021 · last 2025
0000-0002-6698-8870ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 96 · 16 first-author · 9 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-author · 1 since 2021Computer networks · 10 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Theory of computation · 5 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing
abstract
Text-guided image editing using Text-to-Image (T2I) models often fails to yield satisfactory results, frequently introducing unintended modifications, such as the loss of local detail and color changes. In this paper, we analyze these failure cases and attribute them to the indiscriminate optimization across all frequency bands, even though only specific frequencies may require adjustment. To address this, we introduce a simple yet effective approach that enables the selective optimization of specific frequency bands within localized spatial regions for precise edits. Our method leverages wavelets to decompose images into different spatial resolutions across multiple frequency bands, enabling precise modifications at various levels of detail. To extend the applicability of our approach, we provide a comparative analysis of different frequency-domain techniques. Additionally, we extend our method to 3D texture editing by performing frequency decomposition on the triplane representation, enabling frequency-aware adjustments for 3D textures. Quantitative evaluations and user studies demonstrate the effectiveness of our method in producing high-quality and precise edits. Further details are available on our project website: https://ivrl.github.io/fds-webpage/
Yufan Ren, Zicong Jiang, Tong Zhang 0023, Søren Forchhammer, Sabine Süsstrunk
CVPR4
2025 Variable-Rate Learned HDR Image Compression
abstract
Variable-rate learning excels in standard dynamic range (SDR) image compression, but extending it to high dynamic range (HDR) images is challenging. We propose an end-to-end Variable-Rate Learned HDR (VRLHDR) compression framework.
Claire Mantel, Søren Forchhammer
DCC3
2024 Low-complexity ℓ∞-compression of light field images with a deep-decompression stage
M. Umair Mukati, Xi Zhang 0019, Xiaolin Wu 0001, Søren Forchhammer
J. Vis. Commun. Image Represent.4
2023 Hardware Architecture of Channel Encoding for 5G New Radio Physical Downlink Control Channel
abstract
In this article, we propose a flexible and parallelizable hardware architecture of the channel encoding chain for the fifth generation new radio (5G NR) physical downlink control channel (PDCCH). We propose a new polar encoder architecture based on the radix-k processing and fast Fourier transform (FFT) concepts. We also introduce the hardware architectures for cyclic redundancy check (CRC) interleaver and rate matcher for 5G NR PDCCH. We synthesized this complete channel encoding chain on a Virtex Ultrascale+ field-programmable gate-array (FPGA) and show that with the proposed architecture, a codeword throughput of 4.26 Gbps can be realized while consuming as little as 3% of FPGAs resources. The proposed polar encoding architecture can encode from 84 up to 164 resource blocks in the 5G NR frame structure. Encoding of multiple resource blocks can be systematically applied to highly dense (time and frequency) 5G NR fronthaul links supporting multiple antennas.
Shajeel Iqbal, Anders Lund, Metodi Yankov, Thomas G. Nørgaard, Søren Forchhammer
ICC5
2022 Capacity and Achievable Rates of Fading Few-mode MIMO IM/DD Optical Fiber Channels
abstract
The optical fiber multiple-input multiple-output (MIMO) channel with intensity modulation and direct detection (IM/DD) per spatial path is treated. The spatial dimensions represent the multiple modes employed for transmission and the cross-talk between them originates in the multiplexers and demultiplexers, which are polarization dependent and thus time-varying. The upper bounds from free-space IM/DD MIMO channels are adapted to the fiber case, and the constellation constrained capacity is constructively estimated using the Blahut-Arimoto algorithm. An autoencoder is then proposed to optimize a practical MIMO transmission in terms of pre-coder and detector assuming channel distribution knowledge at the transmitter. The pre-coders are shown to be robust to changes in the channel.
Metodi Yankov, Francesco Da Ros, Søren Forchhammer, Lars Grüner-Nielsen
ICC3
2022 How bright should a virtual object be to appear opaque in optical see-through AR?
abstract
Reproduction of occlusions and opaque surfaces are the major challenges of additive optical see-through (OST) displays. This is because the user of an OST display sees a linear mixture of display and environment light, which creates an impression of transparency unless the displayed color is sufficiently bright. The primary goal of this work is to determine how bright a displayed surface needs to be in relation to environment light to be perceived as opaque. We test multiple factors that could affect the perception of opacity: background luminance, contrast, spatial frequency, and accommodation depth in foveal vision. The subjective results, collected on a high-dynamic-range multi-focal stereo display, indicate that a virtual object needs to be, on average, 60 times brighter than the background environment light to be perceived as opaque. A higher contrast of the texture of the virtual object and a background that is out of focus can reduce the required luminance ratio. We demonstrate that a model of visual perception based on Weber’s law and accounting for contrast masking and defocus blur can predict the experimental data with an averaged prediction error of 8.29%. Existing perceptual image difference metrics (PSNR, FovVideoVDP and HDR-VDP-3) can also predict the effect of major factors, but with lower accuracy (e.g. prediction error of 34% for PSNR with PU21 encoding).
Akshay Jindal, Claire Mantel, Søren Forchhammer, Rafal Mantiuk
ISMAR4
2022 Rate-Adaptive Concatenated Multi-Level Coding With Novel Probabilistic Amplitude Shaping
abstract
This paper proposes a new probabilistic amplitude shaping (PAS) approach for concatenated two-level multi-level coding (MLC). The proposed system is based on a concatenated forward error correction (FEC) scheme where outer codes are serially concatenated with inner two-level MLC. This concatenated two-level MLC scheme has recently been shown to have a potential for achieving better performance-complexity trade-offs than the conventional bit-interleaved coded modulation (BICM). Meanwhile, PAS has recently been demonstrated to offer remarkable performance gains as well as rate adaptivity. However, the majority of existing works on PAS assume the use of the binary reflected Gray code as a bit-labeling, and its application to coded modulation schemes with other bit-labelings, such as two-level MLC, may not be straightforward. In this paper, we devise a bit-labeling scheme and propose a new PAS structure for an efficient integration of PAS with two-level MLC systems. More specifically, we propose to generatesignedamplitude symbols with the distribution matcher (DM) for maximizing both coding and shaping gains achieved by two-level MLC and PAS, respectively, while the conventional PAS generatesunsignedamplitude symbols. It is demonstrated by simulation results that, with 256QAM and inner polar codes, the proposed two-level MLC with PAS simultaneously offers 75% reduction in the number of required inner encoding and soft-decision (SD) decoding operations for given outer and inner FEC code lengths, and up to 0.3 dB performance gain over the conventional PAS scheme.
Toshiki Matsumine, Metodi Yankov, Tayyab Mehmood, Søren Forchhammer
IEEE Trans. Commun.4
2021 A Simulation System for Scene Synthesis in Virtual Reality
Claire Mantel, Florian Schweiger, Søren Forchhammer
EuroXR4
2021 Light-Field View Synthesis Using A Convolutional Block Attention Module
abstract
Consumer light-field (LF) cameras suffer from a low or limited resolution because of the angular-spatial trade-off. To alleviate this drawback, we propose a novel learning-based approach utilizing attention mechanism to synthesize novel views of a light-field image using a sparse set of input views (i.e., 4 corner views) from a camera array. In the proposed method, we divide the process into three stages, stereo-feature extraction, disparity estimation, and final image refinement. We use three sequential convolutional neural networks for each stage. A residual convolutional block attention module (CBAM) is employed for final adaptive image refinement. Attention modules are helpful in learning and focusing more on the important features of the image and are thus sequentially applied in the channel and spatial dimensions. Experimental results show the robustness of the proposed method. Our proposed network outperforms the state-of-the-art learning-based light-field view synthesis methods on two challenging real-world datasets by 0.5 dB on average. Furthermore, we provide an ablation study to substantiate our findings.
Muhammad Shahzeb Khan Gul, M. Umair Mukati, Michel Bätz, Søren Forchhammer, Joachim Keinert
ICIP4
2021 Perception-Driven Hybrid Foveated Depth of Field Rendering for Head-Mounted Displays
abstract
In this paper, we present a novel perception-driven hybrid rendering method leveraging the limitation of the human visual system (HVS). Features accounted in our model include: foveation from the visual acuity eccentricity (VAE), depth of field (DOF) from vergence & accommodation, and longitudinal chromatic aberration (LCA) from color vision. To allocate computational workload efficiently, first we apply a gaze-contingent geometry simplification. Then we convert the coordinates from screen space to polar space with a scaling strategy coherent with VAE. Upon that, we apply a stochastic sampling based on DOF. Finally, we post-process the Bokeh for DOF, which can at the same time achieve LCA and anti-aliasing. A virtual reality (VR) experiment on 6 Unity scenes with a head-mounted display (HMD) HTC VIVE Pro Eye yields frame rates range from 25.2 to 48.7 fps. Objective evaluation with FovVideoVDP - a perceptual based visible difference metric - suggests that the proposed method gives satisfactory just-objectionable-difference (JOD) scores across 6 scenes from 7.61 to 8.69 (in a 10 unit scheme). Our method achieves better performance compared with the existing methods while having the same or better level of quality scores.
Claire Mantel, Søren Forchhammer
ISMAR3
2021 Perceptual Evaluation of 360 Audiovisual Quality and Machine Learning Predictions
abstract
In an earlier study, we gathered perceptual evaluations of the audio, video, and audiovisual quality for 360 audiovisual content. This paper investigates perceived audiovisual quality prediction based on objective quality metrics and subjective scores of 360 video and spatial audio content. Thirteen objective video quality metrics and three objective audio quality metrics were evaluated for five stimuli for each coding parameter. Four regression-based machine learning models were trained and tested here, i.e., multiple linear regression, decision tree, random forest, and support vector machine. Each model was constructed using a combination of audio and video quality metrics and two cross-validation methods (k-Fold and Leave-One-Out) were investigated and produced 312 predictive models. The results indicate that the model based on the evaluation of VMAF and AMBIQUAL is better than other combinations of audio-video quality metric. In this study, support vector machine provides higher performance using k-Fold (PCC = 0.909, SROCC = 0.914, and RMSE = 0.416). These results can provide insights for the design of multimedia quality metrics and the development of predictive models for audiovisual omnidirectional media.
Randy Frans Fela, Nick Zacharov, Søren Forchhammer
MMSP3
2021 Learning-based lossless light field compression
abstract
We propose a learning-based method for lossless light field compression. The approach consists of two steps: first, the view to be compressed is synthesized based on previously decoded views; then, the synthesized view is used as a context to predict probabilities of the residual signal for adaptive arithmetic coding. We leverage recent advances in deep-learning-based view synthesis and generative modeling. Specifically, we evaluate two strategies for entropy modeling: a fully parallel probability estimation, where all pixel probabilities are estimated simultaneously; and a partially auto-regressive estimation, in which groups of pixels are predicted sequentially. Our results show that the latter approach provides the best coding gains compared to the state of the art, while keeping the computational complexity competitive.
Milan Stepanov, M. Umair Mukati, Giuseppe Valenzise, Søren Forchhammer, Frédéric Dufaux
MMSP4
2021 Flexible Multilevel Coding With Concatenated Polar-Staircase Codes for M-QAM
abstract
In this work, a multilevel coding (MLC) based coded modulation scheme with two degrees of freedom in rate flexibility is proposed and compared with a bit-interleaved coded modulation (BICM) scheme from a performance versus complexity perspective. The proposed MLC scheme is based on a rate flexible inner soft-decision polar code and utilizes an outer hard-decision staircase code structure as in the 400ZR concatenated forward error-correcting code. The performance of the MLC scheme is investigated for a range of inner code lengths, inner decoder list sizes, and signaling with 16 and 64 quadrature amplitude modulation, respectively. The MLC is designed such that a portion of the staircase encoded bits can bypass the inner code. The number of required inner soft-decision decoders can thus be reduced, thereby saving computational complexity. The proposed MLC scheme simultaneously offers up to a 53.7% reduction in the number of inner decoders and up to 0.55 dB of performance improvement when compared with the similar BICM approach.
Tayyab Mehmood, Metodi Yankov, Shajeel Iqbal, Søren Forchhammer
IEEE Trans. Commun.4
2020 EPIC: Context Adaptive Lossless Light Field Compression using Epipolar Plane Images
abstract
This paper proposes extensions of CALIC for lossless compression of light field (LF) images. The overall prediction process is improved by exploiting the linear structure of Epipolar Plane Images (EPI) in a slope based prediction scheme. The prediction is improved further by averaging predictions made using horizontal and verticals EPIs. Besides this, the difference in these predictions is included in the error energy function, and the texture context is redefined to improve the overall compression ratio. The results using the proposed method shows significant bitrate-savings in comparison to standard lossless coding schemes and offers significant reduction in computational complexity in comparison to the state-of-the-art compression schemes.
M. Umair Mukati, Søren Forchhammer
DCC2
2020 Fast SD-Hamming Decoding in FPGA for High-Speed Concatenated FEC for Optical Communication
abstract
In this paper, we consider fast decoding of soft-decision (SD) Hamming codes as inner codes in concatenated forward error-correction (FEC) schemes for high-speed optical communication. The goal is single FPGA implementations at speeds of 400 Gb/s and beyond. A low complexity maximum a posteriori (MAP) probability decoding is applied to a (128,120) Hamming code. Chase decoding of a (128,119) Hamming code is also implemented. The VHDL designs for both decoding schemes are presented. The FEC performance and FPGA resource utilization are investigated and compared. Synthesis results indicate that, both the Chase and the MAP decoder leave sufficient resources available to also accommodate a powerful outer hard decision code, on a single FPGA. Furthermore, MAP decoding of (128,120) Hamming code features lower hardware complexity and provides a higher data throughput.
Søren Forchhammer, Jakob Dahl Andersen, Tayyab Mehmood, Metodi Yankov, Knud J. Larsen
GLOBECOM2
2020 Towards a Perceived Audiovisual Quality Model for Immersive Content
abstract
This paper studies the quality of multimedia content focusing on 360 video and ambisonic spatial audio reproduced using a head-mounted display and a multichannel loudspeaker setup. Encoding parameters following basic video quality test conditions for 360 videos were selected and a low-bitrate codec was used for the audio encoder. Three subjective experiments were performed for the audio, video, and audiovisual respectively. Peak signal-to-noise ratio (PSNR) and its variants for 360 videos were computed to obtain objective quality metrics and subsequently correlated with the subjective video scores. This study shows that a Cross-Format SPSNR-NN has a slightly higher linear and monotonic correlation over all video sequences. Based on the audiovisual model, a power model shows a highest correlation between test data and predicted scores. We concluded that to enable the development of superior predictive model, a high quality, critical, synchronized audiovisual database is required. Furthermore, comprehensive assessor training may be beneficial prior to the testing to improve the assessors' discrimination ability particularly with respect to multichannel audio reproduction. In order to further improve the performance of audiovisual quality models for immersive content, in addition to developing broader and critical audiovisual databases, the subjective testing methodology needs to be evolved to provide greater resolution and robustness.
Randy Frans Fela, Nick Zacharov, Søren Forchhammer
QoMEX3
2020 UAV image analysis for leakage detection in district heating systems using machine learning
Kabir Hossain, Frederik Villebro, Søren Forchhammer
Pattern Recognit. Lett.3
2020 Fingerprint Entropy and Identification Capacity Estimation Based on Pixel-Level Generative Modelling
abstract
A family of texture-based generative models for fingerprint images is proposed. The generative models are used to estimate upper bounds on the image entropy for systems with small sensor acquisition. The identification capacity of such systems is then estimated using the mutual information between different samples from the same finger. Similar to the generative model for entropy estimation, pixel-level model families are proposed for estimating the similarity between fingerprint images with a given global affine transformation. These models are used for mutual information estimation, and are also adopted to compensate for local deformations between samples. Finally, it is shown that sensor sizes as small as 52 × 52 pixels are potentially sufficient to discriminate populations as large as the entire world population that ever lived, given that the complexity-unconstrained recognition algorithm is available which operates on the lowest possible pixel level.
Metodi Yankov, Martin Aastrup Olsen, Mikkel B. Stegmann, Søren Skovgaard Christensen, Søren Forchhammer
IEEE Trans. Inf. Forensics Secur.5
2019 Evaluation of Prediction of Quality Metrics for IR Images for UAV Applications
abstract
This study presents a framework to predict, in a No Reference (NR) manner, Full Reference (FR) objective quality metrics. The methods are applied to infrared (IR) images acquired by Unmanned Aerial Vehicle (UAV) and compressed on-board and then streamed to a ground computer. The proposed method computes two kinds of features, namely Bitstream Based (BB) features which are estimated from the H.264 bitstream and Pixel Based (PB) features which are estimated from the decoded images. Two BB features are computed using the H.264 Quantization Parameter (QP) and estimated PSNR [1]. A total of 53 PB features are calculated based on spatial information and the rest of the features are based on NR quality assessment methods [1, 2, 3]. The most relevant ones are selected and nally mapped to predict FR objective scores using Support Vector Regression. For the performance evaluation, the proposed method is trained to predict scores of 6 FR image quality metrics (SSIM, NQM, MSSIM, FSIM, MAD and PSNR-HMA) using a set of 250 IR aerial images compressed at 4 levels with H.264/AVC as I-frames. For the SVR mapping, 80% of the contents are used for training (200 contents or 800 images) and the remaining 200 images (20%) for testing. We have evaluated our model for three cases; all features, only BB features and finally excluding BB features. The average SROCC values obtained are 0.970, 0.962 and 0.943, respectively. The BB only version achieves very close results to that of using all features. Thus the presented NR BB Image Quality Assessment (IQA) method for the considered IR image material is very ecient. We have compared our method with three NR methods [1, 2, 3]. The proposed method is competitive compared to the state-of-the-art NR algorithms.
Kabir Hossain, Claire Mantel, Søren Forchhammer
DCC3
2019 An Efficient Storage of Infrared Video of Drone Inspections via Iterative Aerial Map Construction
abstract
In this letter, we present a novel compression algorithm of infrared video sequences captured during drone inspections based on iterative aerial map construction. In our approach, we first apply a stitching algorithm to construct a map of an inspected area assuming that a drone is flying at the same altitude by trajectory close to meander, so that each frame can have a partial overlap with other frame captured much earlier or later. Then, we extract position and rotation angle within the map for each frame and use them as a side information for the video coding. In order to compress an input video sequence, we utilize a multi-view H.265/HEVC with two views. First view is a virtual view generated utilizing the decoded frames of the second view and the side information, whereas the input video is considered as the second view, which is encoded utilizing the virtual view as a reference for the inter-view prediction. The proposed approach has two main benefits. First, the aerial map is generated during decoding utilizing the side information, i.e., the map is not embedded into a bit stream. Second, the inter-view prediction allows to exploit an additional redundancy, which is typical for a drone video. Experimental results show that the proposed algorithm provides 1.4%-2.4% bit rate savings comparing to H.265/HEVC. The maximum possible bit rate savings are estimated from 15.5% to 18.9% assuming that the drone is repeatedly flying many times at exactly the same trajectory.
Eugeniy Belyaev, Søren Forchhammer
IEEE Signal Process. Lett.2
2018 Online Decomposition of Compressive Streaming Data Using n-l1 Cluster-Weighted Minimization
abstract
We consider a decomposition method for compressive streaming data in the context of online compressive Robust Principle Component Analysis (RPCA). The proposed decomposition solves an n-ℓ1 cluster-weighted minimization to decompose a sequence of frames (or vectors), into sparse and low-rank components from compressive measurements. Our method processes a data vector of the stream per time instance from a small number of measurements in contrast to conventional batch RPCA, which needs to access full data. The n-ℓ1 cluster-weighted minimization leverages the sparse components along with their correlations with multiple previously-recovered sparse vectors. Moreover, the proposed minimization can exploit the structures of sparse components via clustering and re-weighting iteratively. The method outperforms the existing methods for both numerical data and actual video data.
Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup
DCC3
2018 Drone HDR Infrared Video Coding via Aerial Map Prediction
abstract
In this paper, we present a novel drone high dynamic range infrared video coding algorithm based on aerial map prediction. First, at the encoder side we accumulate input frames in a buffer and use them to build an aerial map. Then we apply global motion estimation to extract the most similar frame from the aerial map and use it as an extra reference frame in the list of reference frames of H.265/HEVC video coding standard. The map is compressed by H.265/HEVC Intra and included into the overall bit stream. The global motion estimation may reflect the camera rotation which is typical for UAV video sequences. As a result, the overall coding performance is improved. Experimental results show that for a test video with camera rotation the proposed algorithm provides 3-35% bit rate savings comparing to the H.265/HEVC.
Eugeniy Belyaev, Søren Forchhammer
ICIP2
2018 Sparse signal recovery with multiple prior information: Algorithm and measurement bounds
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup
Signal Process.4
2018 Compressive Online Robust Principal Component Analysis via n-ℓ1 Minimization
abstract
This paper considers online robust principal component analysis (RPCA) in time-varying decomposition problems such as video foreground-background separation. We propose a compressive online RPCA algorithm that decomposes recursively a sequence of data vectors (e.g., frames) into sparse and low-rank components. Different from conventional batch RPCA, which processes all the data directly, our approach considers a small set of measurements taken per data vector (frame). Moreover, our algorithm can incorporate multiple prior information from previous decomposed vectors via proposing an - minimization method. At each time instance, the algorithm recovers the sparse vector by solving the - minimization problem-which promotes not only the sparsity of the vector but also its correlation with multiple previously recovered sparse vectors-and, subsequently, updates the low-rank component using incremental singular value decomposition. We also establish theoretical bounds on the number of measurements required to guarantee successful compressive separation under the assumptions of static or slowly changing low-rank components. We evaluate the proposed algorithm using numerical experiments and online video foreground-background separation experiments. The experimental results show that the proposed method outperforms the existing methods.
Huynh Van Luong, Nikos Deligiannis, Jürgen Seiler, Søren Forchhammer, André Kaup
IEEE Trans. Image Process.4
2017 A configurable FPGA FEC unit for Tb/s optical communication
abstract
Decoding of FEC (forward error correction) for optical communication beyond 1 Tb/s is investigated. A configurable single FPGA solution is presented having configurations supporting bit-rates in the range from 40 Gb/s to 1.6 Tb/s. The design allows for trade-offs of bit-rate, footprint, and latency within the resources of the FPGA. A proof-of-concept lab experiment at 40 Gb/s was conducted and pre-FEC — post-FEC performance validated with simulated results.
Jakob Dahl Andersen, Knud J. Larsen, Christian Bering Bogh, Søren Forchhammer, Francesco Da Ros, Kjeld Dalgaard, Shajeel Iqbal
ICC4
2017 Viewpoint adaptive display of HDR images
abstract
In this paper viewpoint adaptive display of HDR images incorporating the effects of ambient light is presented and evaluated. LED backlight displays may render HDR images, but while at a global scale a high dynamic range may be achieved, locally the contrast is limited by the leakage of light through the LC elements of the display. To render high quality images, the display with backlight dimming can compute the values of the LED backlight and LC elements based on the input image, information about the viewpoint of the observer(s) and information of the ambient light. The goal is to achieve the best perceptual reproduction of the specified target image derived from the HDR input image in the specific viewing situation including multiple viewers, possibly having different preferences. An optimization based approach is presented. Some tests with reproduced images are also evaluated subjectively and by image quality metrics as HDR-VDP-2.
Søren Forchhammer, Claire Mantel
ICIP1
2017 Low-complexity compression of high dynamic range infrared images with JPEG compatibility
abstract
We propose a low-complexity High Dynamic Range (HDR) infrared image (IR) coding algorithm assuming the typical case of IR images with an active range of more than 8 bit depth, but less than 16 bit depth. First, we separate an input image into base and residual images with maximum 8 bit depth each. Then we compress each image by a JPEG baseline encoder and include the residual image bit stream into the application part of JPEG header of the base image. As a result, the base image can be reconstructed by JPEG baseline decoder. If the JPEG bit stream size of the residual image is higher than the raw data size, then we include the raw residual image instead. If the residual image contains only zero values or the quality factor for it is 0 then we do not include the residual image into the header. Experimental results show that compared with JPEG-XT Part 6 with `global Reinhard' tone-mapping, the proposed approach has lower complexity and similar rate-distortion performance on IR test images.
Eugeniy Belyaev, Claire Mantel, Søren Forchhammer
VCIP3
2017 No-reference pixel based video quality assessment for HEVC decoded video
Xin Huang 0004, Jacob Søgaard, Søren Forchhammer
J. Vis. Commun. Image Represent.3
2017 An Adaptive Multialphabet Arithmetic Coding Based on Generalized Virtual Sliding Window
abstract
We propose a novel efficient multialphabet multiplication-free adaptive arithmetic coder. First, we generalize probability estimation via virtual sliding window for the multialphabet case and show that it does not require multiplications and provides a tradeoff between the probability adaptation speed and the precision of the probability estimation. Second, we show how the generalized virtual sliding window can be used to eliminate multiplications and divisions. Finally, we demonstrate that the proposed arithmetic coder provides better compression performance than existing implementations based on state-of-the-art multiplication-free binary arithmetic coders.
Eugeniy Belyaev, Søren Forchhammer, Kai Liu 0021
IEEE Signal Process. Lett.2
2016 A Reconstruction Algorithm with Multiple Side Information for Distributed Compression of Sparse Sources
abstract
We consider the task of reconstructing target signals which are processed as sparse sources for a distributed compression scenario, where communication between the sources is prohibited, however, correlation of information among sources can be utilized at the decoder. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI) for such distributed sparse sources. The proposed algorithm takes advantage of both a compressive sensing (CS) reconstruction with SI and an iteratively weighted ℓ1-norm minimization by solving a general weighted multi-ℓ1(or n-ℓ1) minimization. To utilize the known multiple SIs, the algorithm computes optimal weights on not only each individual SI but among SIs where the weights are adaptively updated according to changes at every iteration of the reconstruction. By this optimization, the proposed reconstruction algorithm with multiple SI (RAMSI) can robustly exploit the multiple SIs with different qualities. We experimentally demonstrate our algorithm on compressing feature histograms as sparse sources which are extracted from a multi-view image database for multi-view recognition. The results show that the RAMSI with multiple SIs efficiently outperforms the ℓ1minimization and also the CS reconstruction with only one SI.
Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer
DCC4
2016 Sparse signal reconstruction with multiple side information using adaptive weights for multiview sources
abstract
This work considers reconstructing a target signal in a context of distributed sparse sources. We propose an efficient reconstruction algorithm with the aid of other given sources as multiple side information (SI). The proposed algorithm takes advantage of compressive sensing (CS) with SI and adaptive weights by solving a proposed weighted n-ℓ1minimization. The proposed algorithm computes the adaptive weights in two levels, first each individual intra-SI and then inter-SI weights are iteratively updated at every reconstructed iteration. This two-level optimization leads the proposed reconstruction algorithm with multiple SI using adaptive weights (RAMSIA) to robustly exploit the multiple SIs with different qualities. We experimentally perform our algorithm on generated sparse signals and also correlated feature histograms as multiview sparse sources from a multiview image database. The results show that RAMSIA significantly outperforms both classical CS and CS with single SI, and RAMSIA with higher number of SIs gained more than the one with smaller number of SIs.
Huynh Van Luong, Jürgen Seiler, André Kaup, Søren Forchhammer
ICIP4
2016 Distributed coding of multiview sparse sources with joint recovery
abstract
In support of applications involving multiview sources in distributed object recognition using lightweight cameras, we propose a new method for the distributed coding of sparse sources as visual descriptor histograms extracted from multiview images. The problem is challenging due to the computational and energy constraints at each camera as well as the limitations regarding inter-camera communication. Our approach addresses these challenges by exploiting the sparsity of the visual descriptor histograms as well as their intra- and inter-camera correlations. Our method couples distributed source coding of the sparse sources with a new joint recovery algorithm that incorporates multiple side information signals, where prior knowledge (low quality) of all the sparse sources is initially sent to exploit their correlations. Experimental evaluation using the histograms of shift-invariant feature transform (SIFT) descriptors extracted from multiview images shows that our method leads to an average bit-rate saving of 30.7% compared to the state-of-the-art distributed compressed sensing method with independent encoding of the sources.
Huynh Van Luong, Nikos Deligiannis, Søren Forchhammer, André Kaup
PCS3
2016 Low complexity video encoding for UAV inspection
abstract
In this work we present several methods for fast integer motion estimation of videos recorded aboard an Unmanned Aerial Vehicle (UAV). Different from related work, the field depth is not considered to be consistent. The novel methods designed for low complexity MV prediction in H.264/AVC and analysis hereof include histogram-based prediction, constant global motion, and modification of the candidate sets in the Enhanced Predictive Zonal Search (EPZS). The results verify the applicability of the methods with a speed-up factor of integer motion estimation up to 2-4. Initial results for UAV infrared (IR) video are also provided.
Jacob Søgaard, Ruo Zhang, Søren Forchhammer, Kabir Hossain
PCS3
2016 Modeling the Quality of Videos Displayed With Local Dimming Backlight at Different Peak White and Ambient Light Levels
abstract
This paper investigates the impact of ambient light and peak white (maximum brightness of a display) on the perceived quality of videos displayed using local backlight dimming. Two subjective tests providing quality evaluations are presented and analyzed. The analyses of variance show significant interactions of the factors peak white and ambient light with the perceived quality. Therefore, we proceed to predict the subjective quality grades with objective measures. The rendering of the frames on liquid crystal displays with light emitting diodes backlight at various ambient light and peak white levels is computed using a model of the display. Widely used objective quality metrics are applied based on the rendering models of the videos to predict the subjective evaluations. As these predictions are not satisfying, three machine learning methods are applied: partial least square regression, elastic net, and support vector regression. The elastic net method obtains the best prediction accuracy with a spearman rank order correlation coefficient of 0.71, and two features are identified as having a major influence on the visual quality.
Claire Mantel, Jacob Søgaard, Soren Bech, Jari Korhonen, Jesper Melgaard Pedersen, Søren Forchhammer
IEEE Trans. Image Process.6
2015 Approximating the constellation constrained capacity of the MIMO channel with discrete input
abstract
In this paper the capacity of a Multiple Input Multiple Output (MIMO) channel is considered, subject to average power constraint, for multi-dimensional discrete input, in the case when no channel state information is available at the transmitter. We prove that when the constellation size grows, the QAM constrained capacity converges to Gaussian capacity, directly extending the AWGN result from [1]. Simulations show that for a given constellation size, a rate close to the Gaussian capacity can be achieved up to a certain SNR point, which can be found efficiently by optimizing the constellation for the equivalent orthogonal channel, obtained by the singular value decomposition. Furthermore, lower bounds on the constrained capacity are derived for the cases of square and tall MIMO matrix, by optimizing the constellation for the equivalent channel, obtained by QR decomposition.
Metodi Yankov, Søren Forchhammer, Knud J. Larsen, Lars P. B. Christensen
ICC2
2015 No-reference video quality assessment by HEVC codec analysis
abstract
This paper proposes a No-Reference (NR) Video Quality Assessment (VQA) method for videos subject to the distortion given by High Efficiency Video Coding (HEVC). The proposed assessment can be performed either as a Bitstream-Based (BB) method or as a Pixel-Based (PB). It extracts or estimates the transform coefficients, estimates the distortion, and assesses the video quality. The proposed scheme generates VQA features based on Intra coded frames, and then maps features using an Elastic Net to predict subjective video quality. A set of HEVC coded 4K UHD sequences are tested. Results show that the quality scores computed by the proposed method are highly correlated with the subjective assessment.
Xin Huang 0004, Jacob Søgaard, Søren Forchhammer
VCIP3
2015 No-Reference Video Quality Assessment Using Codec Analysis
abstract
A no-reference (NR) video quality assessment (VQA) method is presented for videos distorted by H.264/Advanced Video Coding (AVC) and MPEG-2. The assessment is performed without access to the bitstream. Instead, we analyze and estimate coefficients based on decoded pixels. The approach involves distinguishing between the two types of videos, estimating the level of quantization used in the I-frames, and exploiting this information to assess the video quality. To do this for H.264/AVC, the distribution of the discrete cosine transform-coefficients after intra-prediction and deblocking are modeled. To obtain VQA features for H.264/AVC, we propose a novel estimation method of the quantization in H.264/AVC videos without bitstream access, which can also be used for peak signal-to-noise ratio estimation. The results from the MPEG-2 and H.264/AVC analysis are mapped to a perceptual measure of video quality by support vector regression. For validation purposes, the proposed method was tested on two databases. In both cases, a good performance compared with state of the art full, reduced, and NR VQA algorithms was achieved.
Jacob Søgaard, Søren Forchhammer, Jari Korhonen
IEEE Trans. Circuits Syst. Video Technol.2
2015 Modeling the Subjective Quality of Highly Contrasted Videos Displayed on LCD With Local Backlight Dimming
abstract
Local backlight dimming is a technology aiming at both saving energy and improving visual quality on television sets. As the rendition of the image is specified locally, the numerical signal corresponding to the displayed image needs to be computed through a model of the display. This simulated signal can then be used as input to objective quality metrics. The focus of this paper is on determining which characteristics of locally backlit displays influence quality assessment. A subjective experiment assessing the quality of highly contrasted videos displayed with various local backlight-dimming algorithms is set up. Subjective results are then compared with both objective measures and objective quality metrics using different display models. The first analysis indicates that the most significant objective features are temporal variations, power consumption (probably representing leakage), and a contrast measure. The second analysis shows that modeling of leakage is necessary for objective quality assessment of sequences displayed with local backlight dimming.
Claire Mantel, Soren Bech, Jari Korhonen, Søren Forchhammer, Jesper Melgaard Pedersen
IEEE Trans. Image Process.4
2014 Rate-adaptive constellation shaping for near-capacity achieving turbo coded BICM
abstract
In this paper the problem of constellation shaping is considered. Mapping functions are designed for a many-to-one signal shaping strategy, combined with a turbo coded Bit-interleaved Coded Modulation (BICM), based on symmetric Huffman codes with binary reflected Gray-like properties. An algorithm is derived for finding the Huffman code with such properties for a variety of alphabet sizes, and near-capacity performance is achieved for a wide SNR region by dynamically choosing the optimal code rate, constellation size and mapping function based on the operating SNR point and assuming perfect channel quality estimation. Gains of more than 1dB are observed for high SNR compared to conventional turbo coded BICM, and it is shown that the mapping functions designed here significantly outperform current state of the art Turbo-Trellis Coded Modulation and other existing constellation shaping methods.
Metodi Yankov, Søren Forchhammer, Knud J. Larsen, Lars P. B. Christensen
ICC2
2014 Low delay Wyner-Ziv coding using optical flow
abstract
Distributed Video Coding (DVC) is a video coding paradigm that exploits the source statistics at the decoder based on the availability of the Side Information (SI). The SI can be seen as a noisy version of the source, and the lower the noise the higher the RD performance of the decoder. The SI is usually generated by means of interpolation-based methods, which rely on the availability of a preceding and a following frame with respect to the to-be-decoded one. These methods lead to relatively high RD performance but also high delays. This work is focused on a low-delay DVC codec, relying only on preceding frames for the generation of the SI by means of Optical Flow (OF), which is also used in the refinement step of the SI for enhanced RD performance. Compared with a state-of-the-art extrapolation-based decoder the proposed solution achieves RD BjØntegaard gains up to 1.3 dB.
Matteo Salmistraro, Søren Forchhammer
ICIP2
2014 Maximizing entropy of pickard random fields for 2×2 binary constraints
abstract
This paper considers the problem of maximizing the entropy of two-dimensional (2D) Pickard Random Fields (PRF) subject to constraints. We consider binary Pickard Random Fields, which provides a 2D causal finite context model and use it to define stationary probabilities for 2×2 squares, thus allowing us to calculate the entropy of the field. All possible binary 2×2 constraints are considered and all constraints are categorized into groups according to their properties. For constraints which can be modeled by a PRF approach and with positive entropy, we characterize and provide statistics of the maximum PRF entropy. As examples, we consider the well known hard square constraint along with a few other constraints.
Jacob Søgaard, Søren Forchhammer
ISIT2
2014 Comparing subjective and objective quality assessment of HDR images compressed with JPEG-XT
abstract
In this paper a subjective test in which participants evaluate the quality of JPEG-XT compressed HDR images is presented. Results show that for the selected test images and display, the subjective quality reached its saturation point starting around 3bpp. Objective evaluations are obtained by applying a model of the display and providing the modeled images to three objective metrics dedicated to HDR content. Objective grades are compared with subjective data both in physical domain and using a gamma correction to approximate perceptually uniform luminance coding. The MRSE metric obtains the best performance with the limit that it does not capture the quality saturation. The usage of the gamma correction prior to applying metrics depends on the characteristics of each objective metric.
Claire Mantel, Stefan Catalin Ferchiu, Søren Forchhammer
MMSP3
2014 Edge-preserving Intra mode for efficient depth map coding based on H.264/AVC
Marco Zamarin, Søren Forchhammer
Signal Process. Image Commun.2
2014 Re-estimation of Motion and Reconstruction for Distributed Video Coding
abstract
Transform domain Wyner-Ziv (TDWZ) video coding is an efficient approach to distributed video coding (DVC), which provides low complexity encoding by exploiting the source statistics at the decoder side. The DVC coding efficiency depends mainly on side information and noise modeling. This paper proposes a motion re-estimation technique based on optical flow to improve side information and noise residual frames by taking partially decoded information into account. To improve noise modeling, a noise residual motion re-estimation technique is proposed. Residual motion compensation with motion updating is used to estimate a current residue based on previously decoded frames and correlation between estimated side information frames. In addition, a generalized reconstruction algorithm to optimize a multihypothesis reconstruction is proposed. The proposed techniques using motion and reconstruction re-estimation (MORE) are integrated in the SING TDWZ codec, which uses side information and noise learning. For Wyner-Ziv frames using GOP size 2, the MORE codec significantly improves the TDWZ coding efficiency with an average (Bjøntegaard) PSNR improvement of 2.5 dB and up to 6 dB improvement compared with DISCOVER.
Huynh Van Luong, Lars Lau Rakêt, Søren Forchhammer
IEEE Trans. Image Process.3
2013 Adaptive deblocking and deringing of H.264/AVC video sequences
abstract
We present a method to reduce blocking and ringing artifacts in H.264/AVC video sequences. For deblocking, the proposed method uses a quality measure of a block based coded image to find filtering modes. Based on filtering modes, the images are segmented to three classes and a specific deblocking filter is applied to each class. Deringing is obtained by an adaptive bilateral filter; spatial and intensity spread parameters are selected adaptively using texture and edge mapping. The analysis of objective and subjective experimental results shows that the proposed algorithm is effective in deblocking and deringing low bit-rate H.264 video sequences.
Ehsan Nadernejad, Nino Burini, Søren Forchhammer
ICASSP3
2013 Distributed multi-hypothesis coding of depth maps using texture motion information and optical flow
abstract
Distributed Video Coding (DVC) is a video coding paradigm allowing a shift of complexity from the encoder to the decoder. Depth maps are images enabling the calculation of the distance of an object from the camera, which can be used in multiview coding in order to generate virtual views, but also in single view coding for motion detection or image segmentation. In this work, we address the problem of depth map video DVC encoding in a single-view scenario. We exploit the motion of the corresponding texture video which is highly correlated with the depth maps. In order to extract the motion information, a block-based and an optical flow-based methods are employed. Finally we fuse the proposed Side Informations using a multi-hypothesis DVC decoder, which allows us to exploit the strengths of all the proposed methods at the same time.
Matteo Salmistraro, Marco Zamarin, Lars Lau Rakêt, Søren Forchhammer
ICASSP4
2013 Texture side information generation for distributed coding of video-plus-depth
abstract
We consider distributed video coding in a monoview video-plus-depth scenario, aiming at coding textures jointly with their corresponding depth stream. Distributed Video Coding (DVC) is a video coding paradigm in which the complexity is shifted from the encoder to the decoder. The Side Information (SI) generation is an important element of the decoder, since the SI is the estimation of the to-be-decoded frame. Depth maps enable the calculation of the distance of an object from the camera. The motion between depth frames and their corresponding texture frames (luminance and chrominance components) is strongly correlated, so the additional depth information may be used to generate more accurate SI for the texture stream, increasing the efficiency of the system. In this paper we propose various methods for accurate texture SI generation, comparing them with other state-of-the-art solutions. The proposed system achieves gains on the reference decoder up to 1.49 dB.
Matteo Salmistraro, Lars Lau Rakêt, Marco Zamarin, Anna Ukhanova, Søren Forchhammer
ICIP5
2013 Game-theoretic rate-distortion-complexity optimization for HEVC
abstract
This paper presents an algorithm for rate-distortion-complexity optimization for the emerging High Efficiency Video Coding (HEVC) standard, whose high computational requirements urge the need for low-complexity optimization algorithms. Optimization approaches need to specify different complexity profiles in order to tailor the computational load to the different hardware and power-supply resources of devices. In this work, we focus on optimizing the quantization parameter and partition depth in HEVC via a game-theoretic approach. The proposed rate control strategy alone provides 0.2 dB improvement compared to the approach implemented in HEVC reference software, while rate-distortion-complexity optimization allows very accurate complexity control providing at the same time rate-distortion performance close to the optimal one.
Anna Ukhanova, Simone Milani, Søren Forchhammer
ICIP3
2013 OPtimal backlight scanning for 3D crosstalk reduction in LCD TV
abstract
This work presents a method to determine the optimal backlight scanning signals to minimize crosstalk for time-sequential stereoscopic 3D on LCD TV with active shutter glasses. The solution is obtained through optimization of the variables defined by a model of backlight scanning that considers important aspects like liquid crystal transitions and light diffusion, subject to constraints that ensure the rendition of a uniform backlight. Compared with basic backlight scanning, the proposed method can increase luminance at a given crosstalk level or reduce crosstalk at a given luminance level.
Nino Burini, Xiao Shu, Liangbao Jiao, Søren Forchhammer, Xiaolin Wu 0001
ICME4
2013 Edge-preserving intra depth coding based on context-coding and H.264/AVC
abstract
Depth map coding plays a crucial role in 3D Video communication systems based on the “Multi-view Video plus Depth” representation as view synthesis performance is strongly affected by the accuracy of depth information, especially at edges in the depth map image. In this paper an efficient algorithm for edge-preserving intra depth compression based on H.264/AVC is presented. The proposed method introduces a new Intra mode specifically targeted to depth macroblocks with arbitrarily shaped edges, which are typically not efficiently represented by DCT. Edge macroblocks are partitioned into two regions each approximated by a flat surface. Edge information is encoded by means of context-coding with an adaptive template. As a novel element, the proposed method allows exploiting the edge structure of previously encoded edge macroblocks during the context-coding step to further increase compression performance. Experiments show that the proposed Intra mode can improve view synthesis performance: average Bjøntegaard bit rate savings of 25% have been reported over a standard H.264/AVC Intra coder.
Marco Zamarin, Matteo Salmistraro, Søren Forchhammer, Antonio Ortega
ICME3
2013 Flicker reduction in LED-LCDs with local backlight
abstract
Local backlight dimming of LCD with LED backlight can reduce power consumption and improve quality of displayed images and videos. However, important variations of LED over time produce a visually annoying artifact called flickering. In this work, we propose a new algorithm to reduce flickering while maintaining video quality. The proposed algorithm uses an adaptive second order Infinite Impulse Response (IIR) in which coefficients are calculated from the local image features. Experimental results show that the proposed method can reduce flickering while simultaneously keeping similar video quality in terms of PSNR and MSE.
Ehsan Nadernejad, Claire Mantel, Nino Burini, Søren Forchhammer
MMSP4
2013 Multi-hypothesis distributed stereo video coding
abstract
Distributed Video Coding (DVC) is a video coding paradigm that exploits the source statistics at the decoder based on the availability of the Side Information (SI). Stereo sequences are constituted by two views to give the user an illusion of depth. In this paper, we present a DVC decoder for stereo sequences, exploiting an interpolated intra-view SI and two inter-view SIs. The quality of the SI has a major impact on the DVC Rate-Distortion (RD) performance. As the inter-view SIs individually present lower RD performance compared with the intra-view SI, we propose multi-hypothesis decoding for robust fusion and improved performance. Compared with a state-of-the-art single side information solution, the proposed DVC decoder improves the RD performance for all the chosen test sequences by up to 0.8 dB. The proposed multi-hypothesis decoder showed higher robustness compared with other fusion techniques.
Matteo Salmistraro, Marco Zamarin, Søren Forchhammer
MMSP3
2013 Adaptive mode decision with residual motion compensation for distributed video coding
abstract
Distributed video coding (DVC) is a coding paradigm that entails low complexity encoding by exploiting the source statistics at the decoder. To improve the DVC coding efficiency, this paper proposes a novel adaptive technique for mode decision to control and take advantage of skip mode and intra mode in DVC. The adaptive mode decision is not only based on quality of key frames but also the rate of Wyner-Ziv (WZ) frames. To improve noise distribution estimation for a more accurate mode decision, a residual motion compensation is proposed to estimate a current noise residue based on a previously decoded frame. The experimental results show that the proposed adaptive mode decision DVC significantly improves the rate distortion performance without increasing the encoding complexity. For a GOP size of 2 on the set of test sequences, the average bitrate saving of the proposed codec is 35.5% on WZ frames compared with the DISCOVER codec.
Huynh Van Luong, Søren Forchhammer, Jürgen Slowack, Jan De Cock, Rik Van de Walle
PCS2
2013 No-Reference Video Quality Assessment using MPEG analysis
abstract
We present a method for No-Reference (NR) Video Quality Assessment (VQA) for decoded video without access to the bitstream. This is achieved by extracting and pooling features from a NR image quality assessment method used frame by frame. We also present methods to identify the video coding and estimate the video coding parameters for MPEG-2 and H.264/AVC which can be used to improve the VQA. The analysis differs from most other video coding analysis methods since it is without access to the bitstream. The results show that our proposed method is competitive with other recent NR VQA methods for MPEG-2 and H.264/AVC.
Jacob Søgaard, Søren Forchhammer, Jari Korhonen
PCS2
2013 Modeling the color image and video quality on liquid crystal displays with backlight dimming
abstract
Objective image and video quality metrics focus mostly on the digital representation of the signal. However, the display characteristics are also essential for the overall Quality of Experience (QoE). In this paper, we use a model of a backlight dimming system for Liquid Crystal Display (LCD) and show how the modeled image can be used as an input to quality assessment algorithms. For quality assessment, we propose an image quality metric, based on Peak Signal-to-Noise Ratio (PSNR) computation in the CIE L*a*b* color space. The metric takes luminance reduction, color distortion and loss of uniformity in the resulting image in consideration. Subjective evaluations of images generated using different backlight dimming algorithms and clipping strategies show that the proposed metric estimates the perceived image quality more accurately than conventional PSNR.
Jari Korhonen, Claire Mantel, Nino Burini, Søren Forchhammer
VCIP4
2013 Enhancing perceived quality of compressed images and video with anisotropic diffusion and fuzzy filtering
Ehsan Nadernejad, Jari Korhonen, Søren Forchhammer, Nino Burini
Signal Process. Image Commun.3
2013 Optimal Local Dimming for LC Image Formation With Controllable Backlighting
abstract
Light emitting diode (LED)-backlit liquid crystal displays (LCDs) hold the promise of improving image quality while reducing the energy consumption with signal-dependent local dimming. However, most existing local dimming algorithms are mostly motivated by simple implementation, and they often lack concern for visual quality. To fully realize the potential of LED-backlit LCDs and reduce the artifacts that often occur in current systems, we propose a novel local dimming technique that can achieve the theoretical highest fidelity of intensity reproduction in either l(1) or l(2) metrics. Both the exact and fast approximate versions of the optimal local dimming algorithm are proposed. Simulation results demonstrate superior performances of the proposed algorithm in terms of visual quality and power consumption.
Xiao Shu, Xiaolin Wu 0001, Søren Forchhammer
IEEE Trans. Image Process.3
2012 Rate-Adaptive BCH Coding for Slepian-Wolf Coding of Highly Correlated Sources
abstract
This paper considers using BCH codes for distributed source coding using feedback. The focus is on coding using short block lengths for a binary source, X, having a high correlation between each symbol to be coded and a side information, Y, such that the marginal probability of each symbol, Xi in X, given Y is highly skewed. In the analysis, noiseless feedback and noiseless communication are assumed. A rate-adaptive BCH code is presented and applied to distributed source coding. Simulation results for a fixed error probability show that rate-adaptive BCH achieves better performance than LDPCA (Low-Density Parity-Check Accumulate) codes for high correlation between source symbols and the side information.
Søren Forchhammer, Matteo Salmistraro, Knud J. Larsen, Xin Huang 0004, Huynh Van Luong
DCC1
2012 Image dependent energy-constrained local backlight dimming
abstract
In this work, we consider and propose two extensions to an optimization-based image dependent backlight dimming algorithm. The first extension introduces error weighting based on human perception of luminance, aiming to improve the perceived image quality; the second extension adds an adjustable term for power consumption to the cost function, allowing flexible power management. Experimental results show that the proposed solution can achieve better results than other algorithms at several power consumption levels.
Nino Burini, Ehsan Nadernejad, Jari Korhonen, Søren Forchhammer, Xiaolin Wu 0001
ICIP4
2012 Objective assessment of the impact of frame rate on video quality
abstract
In this paper, we present a novel objective quality metric that takes the impact of frame rate into account. The proposed metric uses PSNR, frame rate and a content dependent parameter that can easily be obtained from spatial and temporal activity indices. The results have been validated on data from a subjective quality study, where the test subjects have been choosing the preferred path from the lowest quality to the best quality, at each step making a choice in favor of higher frame rate or lower distortion. A comparison with other relevant objective metrics shows that the proposed metric on average provides a more precise correlation with the subjective results.
Anna Ukhanova, Jari Korhonen, Søren Forchhammer
ICIP3
2012 Noise residual learning for noise modeling in distributed video coding
abstract
Distributed video coding (DVC) is a coding paradigm which exploits the source statistics at the decoder side to reduce the complexity at the encoder. The noise model is one of the inherently difficult challenges in DVC. This paper considers Transform Domain Wyner-Ziv (TDWZ) coding and proposes noise residual learning techniques that take residues from previously decoded frames into account to estimate the decoding residue more precisely. Moreover, the techniques calculate a number of candidate noise residual distributions within a frame to adaptively optimize the soft side information during decoding. A residual refinement step is also introduced to take advantage of correlation of DCT coefficients. Experimental results show that the proposed techniques robustly improve the coding efficiency of TDWZ DVC and for GOP=2 bit-rate savings up to 35% on WZ frames are achieved compared with DISCOVER.
Huynh Van Luong, Søren Forchhammer
PCS2
2012 Power consumption analysis of constant bit rate video transmission over 3G networks
Anna Ukhanova, Eugeniy Belyaev, Søren Forchhammer
Comput. Commun.4
2012 Cross-band noise model refinement for transform domain Wyner-Ziv video coding
Xin Huang 0004, Søren Forchhammer
Signal Process. Image Commun.2
2012 Side Information and Noise Learning for Distributed Video Coding Using Optical Flow and Clustering
abstract
Distributed video coding (DVC) is a coding paradigm that exploits the source statistics at the decoder side to reduce the complexity at the encoder. The coding efficiency of DVC critically depends on the quality of side information generation and accuracy of noise modeling. This paper considers transform domain Wyner-Ziv (TDWZ) coding and proposes using optical flow to improve side information generation and clustering to improve the noise modeling. The optical flow technique is exploited at the decoder side to compensate for weaknesses of block-based methods, when using motion-compensation to generate side information frames. Clustering is introduced to capture cross band correlation and increase local adaptivity in the noise modeling. This paper also proposes techniques to learn from previously decoded WZ frames. Different techniques are combined by calculating a number of candidate soft side information for low density parity check accumulate decoding. The proposed decoder side techniques for side information and noise learning (SING) are integrated in a TDWZ scheme. On test sequences, the proposed SING codec robustly improves the coding efficiency of TDWZ DVC. For WZ frames using a GOP size of 2, up to 4-dB improvement or an average (Bjøntegaard) bit-rate savings of 37% is achieved compared with DISCOVER.
Huynh Van Luong, Lars Lau Rakêt, Xin Huang 0004, Søren Forchhammer
IEEE Trans. Image Process.4
2011 Multiple LDPC decoding using bitplane correlation for Transform Domain Wyner-Ziv video coding
abstract
Distributed video coding (DVC) is an emerging video coding paradigm for systems which fully or partly exploit the source statistics at the decoder to reduce the computational burden at the encoder. This paper considers a Low Density Parity Check (LDPC) based Transform Domain Wyner-Ziv (TDWZ) video codec. To improve the LDPC coding performance in the context of TDWZ, this paper proposes a Wyner-Ziv video codec using bitplane correlation through multiple parallel LDPC decoding. The proposed scheme utilizes inter bitplane correlation to enhance the bitplane decoding performance. Experimental results show that the proposed scheme reduces the bit rate up to 3.9% and improves the rate-distortion (RD) performance of TDWZ.
Huynh Van Luong, Xin Huang 0004, Søren Forchhammer
ICASSP3
2011 Parallel iterative decoding of Transform Domain Wyner-Ziv video using cross bitplane correlation
abstract
In recent years, Transform Domain Wyner-Ziv (TDWZ) video coding has been proposed as an efficient Distributed Video Coding (DVC) solution, which fully or partly exploits the source statistics at the decoder to reduce the computational burden at the encoder. In this paper, a parallel iterative LDPC decoding scheme is proposed to improve the coding efficiency of TDWZ video codecs. The proposed parallel iterative LDPC decoding scheme is able to utilize cross bitplane correlation during decoding, by iteratively refining the soft-input, updating a modeled noise distribution and thereafter enhancing the bitplane decoding performance. Experimental results show that the proposed scheme reduces the bit rate of Wyner-Ziv frames up to 5.6% and improves the rate-distortion (RD) performance of TDWZ.
Huynh Van Luong, Xin Huang 0004, Søren Forchhammer
ICIP3
2011 Efficient depth map compression exploiting segmented color data
abstract
3D video representations usually associate to each view a depth map with the corresponding geometric information. Many compression schemes have been proposed for multi-view video and for depth data, but the exploitation of the correlation between the two representations to enhance compression performances is still an open research issue. This paper presents a novel compression scheme that exploits a segmentation of the color data to predict the shape of the different surfaces in the depth map. Then each segment is approximated with a parameterized plane. In case the approximation is sufficiently accurate for the target bit rate, the surface coefficients are compressed and transmitted. Otherwise, the region is coded using a standard H.264/AVC Intra coder. Experimental results show that the proposed scheme permits to outperformthe standardH.264/AVC Intra codec on depth data and can be effectively included into multi-view plus depth compression schemes.
Simone Milani, Pietro Zanuttigh, Marco Zamarin, Søren Forchhammer
ICME4
2011 Multi-hypothesis transform domain Wyner-Ziv video coding including optical flow
abstract
Transform Domain Wyner-Ziv (TDWZ) video coding is an efficient Distributed Video coding solution providing new features such as low complexity encoding, by mainly exploiting the source statistics at the decoder based on the availability of decoder side information. The accuracy of the decoder side information has a major impact on the performance of TDWZ. In this paper, a novel multi-hypothesis based TDWZ video coding is presented to exploit the redundancy between multiple side information and the source information. The decoder used optical flow for side information calculation. Compared with the best available single estimation mode TDWZ, the proposed multi-hypothesis based TDWZ achieves robustly better Rate-Distortion (RD) performance and the overall improvement is up to 0.6 dB at high bitrate and up to 2 dB compared with the DISCOVER TDWZ video codec.
Xin Huang 0004, Lars Lau Rakêt, Huynh Van Luong, Mads Nielsen, François Lauze, Søren Forchhammer
MMSP6
2011 Adaptive noise model for transform domain Wyner-Ziv video using clustering of DCT blocks
abstract
The noise model is one of the most important aspects influencing the coding performance of Distributed Video Coding. This paper proposes a novel noise model for Transform Domain Wyner-Ziv (TDWZ) video coding by using clustering of DCT blocks. The clustering algorithm takes advantage of the residual information of all frequency bands, iteratively classifies blocks into different categories and estimates the noise parameter in each category. The experimental results show that the coding performance of the proposed cluster level noise model is competitive with state-of-the-art coefficient level noise modelling. Furthermore, the proposed cluster level noise model is adaptively combined with a coefficient level noise model in this paper to robustly improve coding performance of TDWZ video codec up to 1.24 dB (by Bjontegaard metric) compared to the DISCOVER TDWZ video codec.
Huynh Van Luong, Xin Huang 0004, Søren Forchhammer
MMSP3
2011 No-reference analysis of decoded MPEG images for PSNR estimation and post-processing
Søren Forchhammer, Jakob Dahl Andersen
J. Vis. Commun. Image Represent.1
2011 Edge-based compression of cartoon-like images with homogeneous diffusion
Markus Mainberger, Andrés Bruhn, Joachim Weickert, Søren Forchhammer
Pattern Recognit.4
2010 Maximum Mutual Information Vector Quantization of Log-Likelihood Ratios for Memory Efficient HARQ Implementations
abstract
Modern mobile telecommunication systems, such as 3GPP LTE, make use of Hybrid Automatic Repeat reQuest (HARQ) for efficient and reliable communication between base stationsand mobile terminals. To this purpose, marginal posterior probabilities of the received bits are stored in the form of log-likelihood ratios (LLR) in order to combine information sent across different transmissions due to requests. To mitigate the effects of ever-increasing data rates that call for larger HARQ memory, vector quantization (VQ) is investigated as a technique for temporary compression of LLRs on the terminal. A capacity analysis leads to using maximum mutual information (MMI) as optimality criterion and in turn Kullback-Leibler (KL) divergence as distortion measure. Simulations run based on an LTE-like system have proven that VQ can be implemented in a computationally simple way at low rates of 2-3 bits per LLR value without compromising the system throughput.
Matteo Danieli, Søren Forchhammer, Jakob Dahl Andersen, Lars P. B. Christensen, Søren Skovgaard Christensen
DCC2
2010 Scalable-to-lossless transform domain distributed video coding
abstract
Distributed video coding (DVC) is a novel approach providing new features as low complexity encoding by mainly exploiting the source statistics at the decoder based on the availability of decoder side information. In this paper, scalable-to-lossless DVC is presented based on extending a lossy Transform Domain Wyner-Ziv (TDWZ) distributed video codec with feedback. The lossless coding is obtained by using a reversible integer DCT. Experimental results show that the performance of the proposed scalable-to-lossless TDWZ video codec can outperform alternatives based on the JPEG 2000 standard. The TDWZ codec provides frame by frame encoding. Comparing the lossless coding efficiency, the proposed scalable-to-lossless TDWZ video codec can save up to 5%-13% bits compared to JPEG LS and H.264 Intra frame lossless coding and do so as a scalable-to-lossless coding.
Xin Huang 0004, Anna Ukhanova, Anton Veselov, Søren Forchhammer, Marat Gilmutdinov
MMSP4
2010 Transform domain Wyner-Ziv video coding with refinement of noise residue and side information
abstract
Distributed Video Coding (DVC) is a video coding paradigm which mainly exploits the source statistics at the decoder based on the availability of side information at the decoder. This paper considers feedback channel based Transform Domain Wyner-Ziv (TDWZ) DVC. The coding efficiency of TDWZ video coding does not match that of conventional video coding yet, mainly due to the quality of side information and inaccurate noise estimation. In this context, a novel TDWZ video decoder with noise residue refinement (NRR) and side information refinement (SIR) is proposed. The proposed refinement schemes are successively updating the estimated noise residue for noise modeling and side information frame quality during decoding. Experimental results show that the proposed decoder can improve the Rate- Distortion (RD) performance of a state-of-the-art Wyner-Ziv video codec for the set of test sequences.
Xin Huang 0004, Søren Forchhammer
VCIP2
2009 Fast Compressed Domain Motion Detection in H.264 Video Streams for Video Surveillance Applications
abstract
This paper presents a novel approach to fast motion detection in H.264/MPEG-4 advanced video coding (AVC) compressed video streams for IP video surveillance systems. The goal is to develop algorithms which may be useful in a real-life industrial perspective by facilitating the processing of large numbers of video streams on a single server. The focus of the work is on using the information in coded video streams to reduce the computational complexity and memory requirements, which translates into reduced hardware requirements and costs. The devised algorithm detects and segments activity based on motion vectors embedded in the video stream without requiring a full decoding and reconstruction of video frames. To improve the robustness to noise, a confidence measure based on temporal and spatial clues is introduced to increase the probability of correct detection. The algorithm was tested on indoor surveillance H.264 sequences.
Krzysztof Szczerba, Søren Forchhammer, Jesper Støttrup-Andersen, Peder Tanderup Eybye
AVSS2
2009 Improved virtual channel noise model for transform domain Wyner-Ziv video coding
abstract
Distributed video coding (DVC) has been proposed as a new video coding paradigm to deal with lossy source coding using side information to exploit the statistics at the decoder to reduce computational demands at the encoder. A virtual channel noise model is utilized at the decoder to estimate the noise distribution between the side information frame and the original frame. This is one of the most important aspects influencing the coding performance of DVC. Noise models with different granularity have been proposed. In this paper, an improved noise model for transform domain Wyner-Ziv video coding is proposed, which utilizes cross-band correlation to estimate the Laplacian parameters more accurately. Experimental results show that the proposed noise model can improve the rate-distortion (RD) performance.
Xin Huang 0004, Søren Forchhammer
ICASSP2
2009 Postprocessing MPEG based on estimated quantization parameters
abstract
Postprocessing of MPEG(-2) video is widely used to attenuate the coding artifacts, especially deblocking but also deringing have been addressed. The focus has been on filters where the decoder has access to the code stream and e.g. utilizes information about the quantization parameter. We consider the case where the coded stream is not accessible, or from an architectural point of view not desirable to use, and instead estimate some of the MPEG stream parameters based on the decoded sequence. The I-frames are detected and the quantization parameters are estimated from the coded stream and used in the postprocessing. We focus on deringing and present a scheme which aims at suppressing ringing artifacts, while maintaining the sharpness of the texture. The goal is to improve the visual quality, so perceptual blur and ringing metrics are used in addition to PSNR evaluation. The performance of the new `pure' postprocessing compares favorable to a reference postprocessing filter which has access to the quantization parameters not only for I-frames but also on P and B-frames.
Søren Forchhammer
ICIP2
2009 Extending models for two-dimensional constraints
abstract
Random fields in two dimensions may be specified on 2 times 2 elements such that the probabilities of finite configurations and the entropy may be calculated explicitly. The Pickard random field is one example where probability of a new (non-boundary) element is conditioned on three previous elements. To extend the concept we consider extending such a field such that a vector or block of elements is conditioned on a larger set of previous elements. Given a stationary model defined on 2 times 2 elements, iterative scaling is used to define the extended model. The extended model may be used for models of two-dimensional constraints and as examples we apply it to the hard-square constraint and the no isolated bits (n.i.b) constraint. The iterative scaling can ensure that the entropy of the extension is optimized and that the entropy is increased compared to the initial model defined on 2 times 2 elements. Application to a simple stationary model with hidden states is also outlined. For the n.i.b constraint, the initial model is based on elements defined by blocks of (1 times 2) binary symbols.
Søren Forchhammer
ISIT1
2009 Distributed Video Coding with multiple side information
abstract
Distributed Video Coding (DVC) is a new video coding paradigm which mainly exploits the source statistics at the decoder based on the availability of some decoder side information. The quality of the side information has a major impact on the DVC rate-distortion (RD) performance in the same way the quality of the predictions had a major impact in predictive video coding. In this paper, a DVC solution exploiting multiple side information is proposed; the multiple side information is generated by frame interpolation and frame extrapolation targeting to improve the side information of a single estimation mode. Compared with the best available single side information solutions, the proposed DVC solution with multiple side information robustly improves the RD performance for the set of test sequences.
Xin Huang 0004, Catarina Brites, João Ascenso, Fernando Pereira 0001, Søren Forchhammer
PCS5
2009 MPEG2 video parameter and no reference PSNR estimation
abstract
MPEG coded video may be processed for quality assessment or postprocessed to reduce coding artifacts or transcoded. Utilizing information about the MPEG stream may be useful for these tasks. This paper deals with estimating MPEG parameter information from the decoded video stream without access to the MPEG stream. This may be used in systems and applications where the coded stream is not accessible. Detection of MPEG I-frames and DCT (discrete cosine transform) block size is presented. For the I-frames, the quantization parameters are estimated. Combining these with statistics of the reconstructed DCT coefficients, the PSNR is estimated from the decoded video without reference images. Tests on decoded fixed rate MPEG2 sequences demonstrate perfect detection rates and good performance of the PSNR estimation.
Søren Forchhammer
PCS2
2009 Block pickard models for two-dimensional constraints
abstract
In Pickard random fields (PRF), the probabilities of finite configurations and the entropy of the field can be calculated explicitly, but only very simple structures can be incorporated into such a field. Given two Markov chains describing a boundary, an algorithm is presented which determines whether a PRF consistent with the distribution on the boundary and a 2-D constraint exists. Iterative scaling is used as part of the algorithm, which also determines the conditional probabilities yielding the maximum entropy for the given boundary description if a solution exists. A PRF is defined for the domino tiling constraint represented by a quaternary alphabet. PRF models are also presented for higher order constraints, including the no isolated bits (n.i.b.) constraint, and a minimum distance 3 constraint by defining super symbols on blocks of binary symbols.
Søren Forchhammer, Jørn Justesen
IEEE Trans. Inf. Theory1
2008 Reduced complexity MPEG2 video post-processing for HD display
abstract
This paper presents MPEG(2) decoder post-processing for high definition (HD) flat panel displays. The focus is to design efficient post-processing to reduce blocking and ringing artifacts. Standard deblocking modules are improved to obtain a significant load reduction through a new DCT based control scheme. Standard deringing modules are enhanced through adaptive thresholding to improve the image quality. The schemes are implemented in a MPEG2 decoder for evaluation. The enhanced deblocking filter results in load reduction with an overall reduction in execution time of 41~46% over the basic implementation. The enhanced deringing combined with the deblocking achieves PSNR improvements on average of 0.5 dB over the basic deblocking and deringing on SDTV and HDTV test sequences. The deblocking and deringing models described in the paper are generic and applicable to a wide variety of common (8times8) DCT-block based real-time video schemes.
Kamran Virk, Søren Forchhammer
ICME3
2008 Improved side information generation for Distributed Video Coding
abstract
As a new coding paradigm, distributed video coding (DVC) deals with lossy source coding using side information to exploit the statistics at the decoder to reduce computational demands at the encoder. The performance of DVC highly depends on the quality of side information. With a better side information generation method, fewer bits will be requested from the encoder and more reliable decoded frames will be obtained. In this paper, a side information generation method is introduced to further improve the rate-distortion (RD) performance of transform domain distributed video coding. This algorithm consists of a variable block size based Y, U and V component motion estimation and an adaptive weighted overlapped block motion compensation (OBMC). The proposal is tested and compared with the results of an executable DVC codec released by DISCOVER group (DIStributed COding for Video sERvices). RD improvements on the set of test sequences are observed.
Xin Huang 0004, Søren Forchhammer
MMSP2
2007 A Multi-Frame Post-Processing Approach to Improved Decoding of H.264/AVC Video
abstract
Video compression techniques may yield visually annoying artifacts for limited bitrate coding. In order to improve video quality, a multi-frame based motion compensated filtering algorithm is reported based on combining multiple pictures to form a single super-resolution picture and decimation to the desired format. The algorithm is applied to H.264/AVC decoded sequences and the processing involves a quality estimation based on picture type and local quantization value. Compared with directly decoding, the peak signal to noise ratio (PSNR) of the sequence obtained by the proposed algorithm is improved, and annoying ringing artifacts are effectively suppressed.
Xin Huang 0004, Søren Forchhammer
ICIP (4)3
2007 Context Quantization by Minimum Adaptive Code Length
abstract
Context quantization is a technique to deal with the issue of context dilution in high-order conditional entropy coding. We investigate the problem of context quantizer design under the criterion of minimum adaptive code length. A property of such context quantizers is derived for binary symbols. A fast context quantizer design algorithm for conditioning binary symbols is presented and its complexity analyzed. It is conjectured that this algorithm is optimal. The context quantization is performed in what may be perceived as a probability simplex space rather than in the space of context instances.
Søren Forchhammer, Xiaolin Wu 0001
ISIT1
2007 Complexity control of fast motion estimation in H.264/MPEG-4 AVC with rate-distortion-complexity optimization
abstract
A complexity control algorithm for H.264 advanced video coding is proposed. The algorithm can control the complexity of integer inter motion estimation for a given target complexity. The Rate-Distortion-Complexity performance is improved by a complexity prediction model, simple analysis of the past statistics and a control scheme. The algorithm also works well for scene change condition. Test results for coding interlaced video (720 x576 PAL) are reported.
Mo Wu, Søren Forchhammer, Shankar Manuel Aghito
VCIP2
2007 Texture enhanced appearance models
Rasmus Larsen 0001, Mikkel B. Stegmann, Sune Darkner, Søren Forchhammer, Timothy F. Cootes, Bjarne K. Ersbøll
Comput. Vis. Image Underst.4
2007 Efficient Coding of Shape and Transparency for Video Objects
abstract
A novel scheme for coding gray-level alpha planes in object-based video is presented. Gray-level alpha planes convey the shape and the transparency information, which are required for smooth composition of video objects. The algorithm proposed is based on the segmentation of the alpha plane in three layers: binary shape layer, opaque layer, and intermediate layer. Thus, the latter two layers replace the single transparency layer of MPEG-4 Part 2. Different encoding schemes are specifically designed for each layer, utilizing cross-layer correlations to reduce the bit rate. First, the binary shape layer is processed by a novel video shape coder. In intra mode, the DSLSC binary image coder presented in [3] is used. This is extended here with an intermode utilizing temporal redundancies in shape image sequences. Then the opaque layer is compressed by a newly designed scheme which models the strong correlation with the binary shape layer by morphological erosion operations. Finally, three solutions are proposed for coding the intermediate layer. The knowledge of the two previously encoded layers is utilized in order to increase compression efficiency. Experimental results are reported demonstrating that the proposed techniques provide substantial bit rate savings coding shape and transparency when compared to the tools adopted in MPEG-4 Part 2.
Shankar Manuel Aghito, Søren Forchhammer
IEEE Trans. Image Process.2
2007 Entropy of Bit-Stuffing-Induced Measures for Two-Dimensional Checkerboard Constraints
abstract
A modified bit-stuffing scheme for two-dimensional (2-D) checkerboard constraints is introduced. The entropy of the scheme is determined based on a probability measure defined by the modified bit-stuffing. Entropy results of the scheme are given for 2-D constraints on a binary alphabet. The constraints considered are 2-D RLL(d,infin) for d=2,3 and 4 as well as for the constraint with a minimum 1-norm distance of 3 between 1s. For these results the entropy is within 1-2% of an upper bound on the capacity for the constraint. As a variation of the scheme, periodic merging arrays are also considered
Søren Forchhammer, Torben V. Laursen
IEEE Trans. Inf. Theory1
2006 A Model for the Two-Dimensional No Isolated Bits Constraint
abstract
A stationary model is presented for the two-dimensional (2-D) no isolated bits (n.i.b.) constraint over an extended alphabet defined by the elements within 1 by 2 blocks. This block-wise model is based on a set of sufficient conditions for a Pickard random field (PRF) over an m-ary alphabet. Iterative techniques are applied as part of determining the model parameters. Given two Markov chains describing a boundary, an algorithm is presented which determines whether a certain PRF consistent with the boundary exists. Iterative scaling is used as part of the algorithm, which also determines the conditional probabilities yielding the maximum entropy for the given boundary description if a solution exists. Optimizing over the parameters for a class of boundaries with certain symmetry properties, an entropy of 0.9156 is achieved for the n.i.b. constraint, providing a lower bound. An algorithm for iterative search for a PRF solution starting from a set of conditional probabilities is also presented
Søren Forchhammer, Torben V. Laursen
ISIT1
2006 Context-based coding of bilevel images enhanced by digital straight line analysis
abstract
A new efficient compression scheme for bilevel images containing locally straight edges is presented. This paper is especially focused on lossless (intra) coding of binary shapes for image and video objects, but other images with similar characteristics such as line drawings, layers of digital maps, or segmentation maps are also encoded efficiently. The algorithm is not targeted at document images with text, which can be coded efficiently with dictionary-based techniques as in JBIG2. The scheme is based on a local analysis of the digital straightness of the causal part of the object boundary, which is used in the context definition for arithmetic encoding. Tested on individual images of standard TV resolution binary shapes and the binary layers of a digital map, the proposed algorithm outperforms PWC, JBIG, JBIG2, and MPEG-4 CAE. On the binary shapes, the code lengths are reduced by 21%, 27%, 28%, and 41%, respectively. On the map layers, the reductions are 31%, 34%, 32%, and 64%, respectively. The algorithm is also more efficient on the test material than the state-of-the-art generic bilevel image coder free tree.
Shankar Manuel Aghito, Søren Forchhammer
IEEE Trans. Image Process.2
2004 Context Based Coding of Binary Shapes by Object Boundary Straightness Analysis
abstract
A new lossless compression scheme for bilevel images targeted at binary shapes of image and video objects is presented. The scheme is based on a local analysis of the digital straightness of the causal part of the object boundary, which is used in the context definition for arithmetic encoding. Tested on individual images of binary shapes and binary layers of digital maps the algorithm outperforms PWC, JBIG and MPEG-4 CAE. On the binary shapes the code lengths are reduced by 21%, 25%, and 42%, respectively. On the maps the reductions are 34%, 32%, and 59%, respectively. The algorithm is also more efficient than the state-of-the-art and more complex free tree coder for most of the binary shape and map test images.
Shankar Manuel Aghito, Søren Forchhammer
Data Compression Conference2
2004 Rate-distortion-complexity optimization of fast motion estimation in H.264/MPEG-4 AVC
abstract
This paper presents an operational method for optimizing integer motion estimation in real-time H.264 encoding with respect to the trade-off between rate-distortion and complexity. This three-parameter problem is converted into a more tractable two-parameter problem by a simple approximative elimination of either the rate or the distortion term by converting small differences in the term to be eliminated. This paper also presents an H.264-enhanced implementation of the fast EPZS motion estimation algorithm, which adds three early-stop criteria controlled by predefined thresholds allowing it to stop after 16 16 and 8 8 block types and for each additional reference frame. Results for interlaced SDTV material are presented with indications of the applicability of the operational method when applied to the presented motion estimation procedure. For example, a speed-up by a factor of 4 compared to a basic H.264-adapted EPZS implementation is achieved at only 1% increase in rate.
Jesper Støttrup-Andersen, Søren Forchhammer, Shankar Manuel Aghito
ICIP2
2004 Analysis of bit-stuffing codes and lower bounds on capacity for 2-D constrained arrays using quasistationary measures
abstract
A method for designing quasistationary probability measures for two-dimensional (2-D) constraints is presented. This measure is derived from a modified bit-stuff coding scheme and it gives the capacity of the coding scheme. This provides a constructive lower bound on the capacity of the 2-D constraint. The main examples are checkerboard codes with binary elements. The capacity for one instance of the modified bit-stuffing for the 2-D runlength-limited RLL(2,/spl infin/) constraint is calculated to be 0.4414 bits/symbol. For the constraint given by a minimum (1-norm) distance of 3 between 1s a code with capacity 0.3497 bits/symbol is given.
Søren Forchhammer
ISIT1
2004 Optimal context quantization in lossless compression of image data sequences
abstract
In image compression context-based entropy coding is commonly used. A critical issue to the performance of context-based image coding is how to resolve the conflict of a desire for large templates to model high-order statistic dependency of the pixels and the problem of context dilution due to insufficient sample statistics of a given input image. We consider the problem of finding the optimal quantizer Q that quantizes the K-dimensional causal context Ct = (Xt-t1,Xt-t2,...,X t-tK) of a source symbol Xt into one of a set of conditioning states. The optimality of context quantization is defined to be the minimum static or minimum adaptive code length of given a data set. For a binary source alphabet an optimal context quantizer can be computed exactly by a fast dynamic programming algorithm. Faster approximation solutions are also proposed. In case of m-ary source alphabet a random variable can be decomposed into a sequence of binary decisions, each of which is coded using optimal context quantization designed for the corresponding binary random variable. This optimized coding scheme is applied to digital maps and alpha-plane sequences. The proposed optimal context quantization technique can also be used to establish a lower bound on the achievable code length, and hence is a useful tool to evaluate the performance of existing heuristic context quantizers.
Søren Forchhammer, Xiaolin Wu 0001, Jakob Dahl Andersen
IEEE Trans. Image Process.1
2002 Progressive Coding of Palette Images and Digital Maps
abstract
A 2D version of PPM (Prediction by Partial Matching) coding is introduced simply by combining a 2D template with the standard PPM coding scheme. A simple scheme for resolution reduction is given and the 2D PPM scheme extended to resolution progressive coding by placing pixels in a lower resolution image layer. The resolution is increased by a factor of 2 in each step. The 2D PPM coding is applied to palette images and street maps. The sequential results are comparable to PWC. The PPM results are a little better for the palette images with few colors (up to 4-5 bpp) and a little worse for the images with more colors. For street maps the 2D PPM is slightly better. The PPM based resolution progressive coding provides a better result than coding the resolution layers as individual images. Compared to GIF the resolution progressive 2D PPM's coding efficiency is significantly better. An example of combined content-layer/spatial progressive coding is also given.
Søren Forchhammer, J. Martin Salinas
DCC1
2002 A unified approach to restoration, deinterlacing and resolution enhancement in decoding MPEG-2 video
abstract
The quality and spatial resolution of video can be improved by combining multiple pictures to form a single superresolution picture. We address the special problems associated with pictures of variable but somehow parameterized quality such as MPEG-decoded video. Our algorithm provides a unified approach to restoration, chrominance upsampling, deinterlacing, and resolution enhancement. A decoded MPEG-2 sequence for interlaced standard definition television (SDTV) in 4:2:0 is converted to: (1) improved quality interlaced SDTV in 4:2:0; (2) interlaced SDTV in 4:4:4; (3) progressive SDTV in 4:4:4; (4) interlaced high-definition TV (HDTV) in 4:2:0; (5) progressive HDTV in 4:2:0. These conversions also provide features such as freeze frame and zoom. The algorithm is mainly targeted at bit rates of 4-8 Mb/s. The algorithm is based on motion-compensated spatial upsampling from multiple images and decimation to the desired format. The processing involves an estimated quality of individual pixels based on MPEG image type and local quantization value. The mean-squared error (MSE) is reduced, compared to the directly decoded sequence, and annoying ringing artifacts, including mosquito noise, are effectively suppressed. The superresolution pictures obtained by the algorithm are of much higher visual quality and have lower MSE than superresolution pictures obtained by simple spatial interpolation.
Bo Martins, Søren Forchhammer
IEEE Trans. Circuits Syst. Video Technol.2
2002 Content layer progressive coding of digital maps
abstract
A new lossless context based method is presented for content progressive coding of limited bits/pixel images, such as maps, company logos, etc., common on the World Wide Web. Progressive encoding is achieved by encoding the image in content layers based on color level or other predefined information. Information from already coded layers are used when coding subsequent layers. This approach is combined with efficient template based context bilevel coding, context collapsing methods for multilevel images and arithmetic coding. Relative pixel patterns are used to collapse contexts. Expressions for calculating the resulting number of contexts are given. The new methods outperform existing schemes coding digital maps and in addition provide progressive coding. Compared to the state-of-the-art PWC coder, the compressed size is reduced to 50-70% on our layered map test images.
Søren Forchhammer, Ole Riis Jensen
IEEE Trans. Image Process.1
2001 Lossless Image Data Sequence Compression Using Optimal Context Quantization
abstract
Context based entropy coding often faces the conflict of a desire for large templates and the problem of context dilution. We consider the problem of finding the quantizer Q that quantizes the K-dimensional causal context C/sub i/=(X(i-t/sub 1/), X(i-t/sub 2/), ..., X(i-t/sub K/)) of a source symbol X/sub i/ into one of M conditioning states. A solution giving the minimum adaptive code length for a given data set is presented (when the cost of the context quantizer is neglected). The resulting context quantizers can be used for sequential coding of the sequence X/sub 0/, X/sub 1/, X/sub 2/, .... A coding scheme based on binary decomposition and context quantization for coding the binary decisions is presented and applied to digital maps and /spl alpha/-plane sequences. The optimal context quantization is also used to evaluate existing heuristic context quantizations.
Søren Forchhammer, Xiaolin Wu 0001, Jakob Dahl Andersen
Data Compression Conference1
2000 Content Layer Progressive Coding of Digital Maps
abstract
A new lossless context based method is presented for content progressive coding of limited bits/pixel images, such as maps, company logos, etc., common on the WWW. Progressive encoding is achieved by separating the image into content layers based on other predefined information. Information from already coded layers are used when coding subsequent layers. This approach is combined with efficient template based context bi-level coding, context collapsing methods for multi-level images and arithmetic coding. Relative pixel patterns are used to collapse contexts. The number of contexts are analyzed. The new methods outperform existing coding schemes coding digital maps and in addition provide progressive coding. Compared to the state-of-the-art PWC coder, the compressed size is reduced to 60-70% on our layered test images.
Søren Forchhammer, Ole Riis Jensen
Data Compression Conference1
2000 A Unified Approach to Restoration, Deinterlacing and Superresolution of MPEG-2 Decoded Video
abstract
The quality and the spatial resolution of video can be improved by combining multiple pictures to form a single superresolution picture. We address the special problems associated with pictures of variable but somehow parameterized quality such as MPEG-decoded video. Our algorithm provides a unified approach to restoration, chrominance upsampling, deinterlacing and superresolution as e.g. HDTV. The algorithm is mainly targeted at improving MPEG-2 decoding at high bit rates (4-8 Mbit/s). The mean squared error is reduced, compared to the directly decoded sequence, and annoying ringing artifacts including mosquito noise are effectively suppressed. The superresolution pictures obtained by the algorithm are of much higher visual quality and has lower mean squared error than superresolution pictures obtained by simple spatial interpolation.
Bo Martins, Søren Forchhammer
ICIP2
2000 Bounds on the capacity of constrained two-dimensional codes
abstract
Bounds on the capacity of constrained two-dimensional (2-D) codes are presented. The bounds of Calkin and Wilf (see SIAM J. Discr. Math., vol.11, no.1, p.54-60, 1998) apply to first-order symmetric constraints. The bounds are generalized in a weaker form to higher order and nonsymmetric constraints. Results are given for constraints specified by run-length limits or a minimum distance between pixels of a given value.
Søren Forchhammer, Jørn Justesen
IEEE Trans. Inf. Theory1
1999 Image Coding Using Markov Models with Hidden States
abstract
Summary form only given. Lossless image coding may be performed by applying arithmetic coding sequentially to probabilities conditioned on the past data. Therefore the model is very important. A new image model is applied to image coding. The model is based on a Markov process involving hidden states. An underlying Markov process called the slice process specifies D rows with the width of the image. Each new row of the image coincides with row N of an instance of the slice process. The N-1 previous rows are read from the causal part of the image and the last D-N rows are hidden. This gives a description of the current row conditioned on the N-1 previous rows. From the slice process we may decompose the description into a sequence of conditional probabilities, involving a combination of a forward and a backward pass. In effect the causal part of the last N rows of the image becomes the context. The forward pass obtained directly from the slice process starts from the left for each row with D-N hidden rows. The backward pass starting from the right additionally has the current row as hidden. The backward pass may be described as a completion of the forward pass. It plays the role of normalizing the possible completions of the forward pass for each pixel. The hidden states may effectively be represented in a trellis structure as in an HMM. For the slice process we use a state of D rows and V-1 columns, thus involving V columns in each transition. The new model was applied to a bi-level image (SO9 of the JBIG test set) in a two-part coding scheme.
Søren Forchhammer
Data Compression Conference1
1999 Virtual seminar room-modelling and experimentation in horizontal and vertical integration
abstract
The initial design considerations and research goals for an ATM network based virtual seminar room with five sites are presented. The basic observation behind the design of the virtual seminar room is, that besides the constant growth in available bandwidth for transmission in communication networks, many networks either already offer or are developing technologies to give quality of service (QoS) guarantees. This means that applications not only will support transmission of coded audio and video, but can be designed in a known and well-controlled network environment, which enables the applications to provide, in a broad sense, a high and reliable quality presented at the human computer interface.
Søren Forchhammer, Anders Fosgerau, Peter Søren Kirk Hansen, Steffen Duus Hansen, Ole Riis Jensen, Robin Sharp, John Aasted Sørensen
MMSP1
1999 Content progressive coding of limited bits/pixel images
abstract
We propose a new lossless context based method for content progressive coding of limited bits/pixel images, such as maps, company logos, etc., common on the WWW. Progressive encoding is done by separating the image into content layers based on color level or predefined information. By introducing skip-pixel coding the number of pixels encoded in each layer is reduced by using information from already coded layers. This approach is combined with efficient template based context bi-level coding, context collapsing methods for multi-level images and arithmetic coding. With our new method we are thus able to outperform existing coding schemes and in addition provide progressive coding. Compared to GIF we are able to reduce the compressed size with a factor of up to 3.
Ole Riis Jensen, Søren Forchhammer
MMSP2
1999 Adaptive partially hidden Markov models with application to bilevel image coding
abstract
Partially hidden Markov models (PHMMs) have previously been introduced. The transition and emission/output probabilities from hidden states, as known from the HMMs, are conditioned on the past. This way, the HMM may be applied to images introducing the dependencies of the second dimension by conditioning. In this paper, the PHMM is extended to multiple sequences with a multiple token version and adaptive versions of PHMM coding are presented. The different versions of the PHMM are applied to lossless bilevel image coding. To reduce and optimize the model cost and size, the contexts are organized in trees and effective quantization of the parameters is introduced. The new coding methods achieve results that are better than the JBIG standard on selected test images, although at the cost of increased complexity. By the minimum description length principle, the methods presented for optimizing the code length may apply as guidance for training (P)HMMs for, e.g., segmentation or recognition purposes. Thereby, the PHMM models provide a new approach to image modeling.
Søren Forchhammer, Tage S. Rasmussen
IEEE Trans. Image Process.1
1999 Lossless, near-lossless, and refinement coding of bilevel images
abstract
We present general and unified algorithms for lossy/lossless coding of bilevel images. The compression is realized by applying arithmetic coding to conditional probabilities. As in the current JBIG standard the conditioning may be specified by a template. For better compression, the more general free tree may be used. Loss may be introduced in a preprocess on the encoding side to increase compression. The primary algorithm is a rate-distortion controlled greedy flipping of pixels. Though being general, the algorithms are primarily aimed at material containing half-toned images as a supplement to the specialized soft pattern matching techniques that work better for text. Template based refinement coding is applied for lossy-to-lossless refinement. Introducing only a small amount of loss in half-toned test images, compression is increased by up to a factor of four compared with JBIG. Lossy, lossless, and refinement decoding speed and lossless encoding speed are less than a factor of two slower than JBIG. The (de)coding method is proposed as part of JBIG2, an emerging international standard for lossless/lossy compression of bilevel images.
Bo Martins, Søren Forchhammer
IEEE Trans. Image Process.2
1999 Correction to "lossless, near-lossless, and refinement coding of bilevel images"
Bo Martins, Søren Forchhammer
IEEE Trans. Image Process.2
1999 Entropy Bounds for Constrained Two-Dimensional Random Fields
abstract
The maximum entropy and thereby the capacity of two-dimensional (2-D) fields given by certain constraints on configurations is considered. Upper and lower bounds are derived. A new class of 2-D processes yielding good lower bounds is introduced. Asymptotically, the process achieves capacity for constraints with limited long-range effects. The processes are general and may also be applied to, e.g., data compression of digital images. Results are given for the binary hard square model, which is a 2-D run-length-limited model and some other 2-D models with simple constraints.
Søren Forchhammer, Jørn Justesen
IEEE Trans. Inf. Theory1
1998 Lossless Compression of Video Using Motion Compensation
abstract
Summary form only given. We investigate lossless coding of video using predictive coding and motion compensation. The new coding methods combine state-of-the-art lossless techniques as JPEG (context based prediction and bias cancellation, Golomb coding), with high resolution motion field estimation, 3D predictors, prediction using one or multiple (k) previous images, predictor dependent error modelling, and selection of motion field by code length. We treat the problem of precision of the motion field as one of choosing among a number of predictors. This way, we can incorporate 3D-predictors and intra-frame predictors as well. As proposed by Ribas-Corbera (see PhD thesis, University of Michigan, 1996), we use bi-linear interpolation in order to achieve sub-pixel precision of the motion field. Using more reference images is another way of achieving higher accuracy of the match. The motion information is coded with the same algorithm as is used for the data. For slow pan or slow zoom sequences, coding methods that use multiple previous images perform up to 20% better than motion compensation using a single previous image and up to 40% better than coding that does not utilize motion compensation.
Bo Martins, Søren Forchhammer
Data Compression Conference2
1998 The emerging JBIG2 standard
abstract
The Joint Bi-Level Image Experts Group (JBIG), an international study group affiliated with ISO/IEC and ITU-T, is in the process of drafting a new standard for lossy and lossless compression of bilevel images. The new standard, informally referred to as JBIG2, will support model-based coding for text and halftones to permit compression ratios up to three times those of existing standards for lossless compression. JBIG2 will also permit lossy preprocessing without specifying how it is to be done, In this case, compression ratios up to eight times those of existing standards may be obtained with imperceptible loss of quality. It is expected that JBIG2 will become an international standard by 2000.
Paul G. Howard, Faouzi Kossentini, Bo Martins, Søren Forchhammer, William Rucklidge
IEEE Trans. Circuits Syst. Video Technol.4
1998 Tree coding of bilevel images
abstract
Presently, sequential tree coders are the best general purpose bilevel image coders and the best coders of halftoned images. The current ISO standard, Joint Bilevel Image Experts Group (JBIG), is a good example. A sequential tree coder encodes the data by feeding estimates of conditional probabilities to an arithmetic coder. The conditional probabilities are estimated from co-occurrence statistics of past pixels, the statistics are stored in a tree. By organizing the code length calculations properly, a vast number of possible models (trees) reflecting different pixel orderings can be investigated within reasonable time prior to generating the code. A number of general-purpose coders are constructed according to this principle. Rissanen's one-pass algorithm, context, is presented in two modified versions. The baseline is proven to be a universal coder. The faster version, which is one order of magnitude slower than JBIG, obtains excellent and highly robust compression performance. A multipass free tree coding scheme produces superior compression results for all test images. A multipass free template coding scheme produces significantly better results than JBIG for difficult images such as halftones. By utilizing randomized subsampling in the template selection, the speed becomes acceptable for practical image coding.
Bo Martins, Søren Forchhammer
IEEE Trans. Image Process.2
1996 Bi-level Image Compression with Tree Coding
abstract
Presently, tree coders are the best bi-level image coders. The current ISO standard, JBIG, is a good example. By organising code length calculations properly a vast number of possible models (trees) can be investigated within reasonable time prior to generating code. Three general-purpose coders are constructed by this principle. A multi-pass free tree coding scheme produces superior compression results for all test images. A multi-pass fast free template coding scheme produces much better results than JBIG for difficult images, such as halftonings. Rissanen's algorithm 'Context' is presented in a new version that without sacrificing speed brings it close to the multi-pass coders in compression performance.
Bo Martins, Søren Forchhammer
Data Compression Conference2
1996 Partially hidden Markov models
abstract
Partially hidden Markov models (PHMM) are introduced. They differ from the ordinary HMMs in that both the transition probabilities of the hidden states and the output probabilities are conditioned on past observations. As an illustration they are applied to black and white image compression where the hidden variables may be interpreted as representing noncausal pixels.
Søren Forchhammer, Jorma Rissanen
IEEE Trans. Inf. Theory1
1995 Coding with Partially Hidden Markov Models
abstract
Partially hidden Markov models (PHMM) are introduced. They are a variation of the hidden Markov models (HMM) combining the power of explicit conditioning on past observations and the power of using hidden states. (P)HMM may be combined with arithmetic coding for lossless data compression. A general 2-part coding scheme for given model order but unknown parameters based on PHMM is presented. A forward-backward reestimation of parameters with a redefined backward variable is given for these models and used for estimating the unknown parameters. Proof of convergence of this reestimation is given. The PHMM structure and the conditions of the convergence proof allows for application of the PHMM to image coding. Relations between the PHMM and hidden Markov models (HMM) are treated. Results of coding bi-level images with the PHMM coding scheme is given. The results indicate that the PHMM can adapt to instationarities in the images.
Søren Forchhammer, Jorma Rissanen
Data Compression Conference1
1995 Filters involving derivatives with application to reconstruction from scanned halftone images
abstract
This paper presents a method for designing finite impulse response (FIR) filters for samples of a 2-D signal, e.g., an image, and its gradient. The filters, which are called blended filters, are decomposable in three filters, each separable in 1-D filters on subsets of the data set. Optimality in the minimum mean square error sense (MMSE) of blended filtering is shown for signals with separable autocorrelation function. Relations between correlation functions for signals and their gradients are derived. Blended filters may be composed from FIR Wiener filters using these relations. Simple blended filters are developed and applied to the problem of gray value image reconstruction from bilevel (scanned) clustered-dot halftone images, which is an application useful in the graphic arts. Reconstruction results are given, showing that reconstruction with higher resolution than the halftone grid is achievable with blended filters.
Søren Forchhammer, Kim S. Jensen
IEEE Trans. Image Process.1
1994 Data compression of scanned halftone images
abstract
A new method for coding scanned halftone images is proposed. It is information-lossy, but still preserving the image quality, compression rates of 16-35 have been achieved for a typical test image scanned on a high resolution scanner. The bi-level halftone images are filtered, in phase with the halftone grid, and converted to a gray level representation. A new digital description of (halftone) grids has been developed for this purpose. The gray level values are coded according to a scheme based on states derived from a segmentation of gray values. To enable real-time processing of high resolution scanner output, the coding has been parallelized and implemented on a transputer system. For comparison, the test image was coded using existing (lossless) methods giving compression rates of 2-7. The best of these, a combination of predictive and binary arithmetic coding was modified and optimized achieving a compression rate of 9.>
Søren Forchhammer, Kim S. Jensen
IEEE Trans. Commun.1
1989 Digital plane and grid point segments
Søren Forchhammer
Comput. Vis. Graph. Image Process.1
1988 Algorithms for coding scanned halftone pictures
abstract
A method for coding scanned documents containing halftone pictures, e.g. newspapers and magazines, for transmission purposes is proposed. The halftone screen is estimated and the grey value of each dot is found, thus giving a compact description. At the receiver the picture is rescreened. A novel data structure and related algorithms for handling the digital screen without restrictions on the screen parameters is presented. Data compression rates above 20 are obtained for the halftone pictures. The algorithms are suited for implementation with fast dedicated hardware. The rescreening can also be used as digital halftoning with arbitrary screens, using lookup tables.>
Søren Forchhammer, Morten Forchhammer
ICPR1
1988 Digital squares
abstract
Digital squares are defined and their geometric properties characterized. A linear time algorithm is presented that considers a convex digital region and determines whether or not it is a digital square. The algorithm also determines the range of the values of the parameter set of its preimages. The analysis involves transforming the boundary of a digital region into parameter space of slope and y-intercept.>
Søren Forchhammer, Chul E. Kim
ICPR1