Fatih Kamisli

dblp:93/6133 · DBLP profile ↗
← Back
21ranked-venue papers
15as first author
5since 2021 · last 2025
0000-0001-7112-1633ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 15 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 End-to-end learned block-based image compression with block-level masked convolutions and asymptotic closed-loop training
Fatih Kamisli
Multim. Tools Appl.1
2024 Variable-Rate Learned Image Compression with Multi-Objective Optimization and Quantization-Reconstruction Offsets
abstract
Achieving successful variable bitrate compression with computationally simple algorithms from a single end-to-end learned image or video compression model remains a challenge. Many approaches have been proposed, including conditional auto-encoders, channel-adaptive gains for the latent tensor or uniformly quantizing all elements of the latent tensor. This paper follows the traditional approach to vary a single quantization step size to perform uniform quantization of all latent tensor elements. However, three modifications are proposed to improve the variable rate compression performance. First, multi objective optimization is used for (post) training. Second, a quantization-reconstruction offset is introduced into the quantization operation. Third, variable rate quantization is used also for the hyper latent. All these modifications can be made on a pre-trained single-rate compression model by performing post training. The algorithms are implemented into three well-known image compression models and the achieved variable rate compression results indicate negligible or minimal compression performance loss compared to training multiple models. (Codes will be shared at https://github.com/InterDigitalInc/CompressAI)
Fatih Kamisli, Fabien Racapé, Hyomin Choi
DCC1
2024 A learned pixel-by-pixel lossless image compression method with 59K parameters and parallel decoding
Sinem Gümüs, Fatih Kamisli
Multim. Tools Appl.2
2023 Image compression with learned lifting-based DWT and learned tree-based entropy models
Ugur Berk Sahin, Fatih Kamisli
Multim. Syst.2
2023 Learned Lossless Image Compression Through Interpolation With Low Complexity
abstract
With the increasing popularity of deep learning in image processing, many learned lossless image compression methods have been proposed recently. One group of algorithms are based on scale-based auto-regressive models and can provide competitive compression performance while also allowing easily parallelized computations and short encoding/decoding times. However, they use large neural networks and have high computational requirements. This paper presents an interpolation based learned lossless image compression method which falls in the scale-based auto-regressive models group. The method achieves compression performance better than or on par with the recent scale-based auto-regressive models, yet requires more than 10x less neural network parameters (0.19M) and encoding/decoding computation complexity. These achievements are due to the contributions/findings in the overall system and neural network architecture design, such as sharing interpolator neural networks across different scales, using separate neural networks for different parameters of the probability distribution model and performing the processing in the YCoCg-R color space instead of the RGB color space.
Fatih Kamisli
IEEE Trans. Circuits Syst. Video Technol.1
2019 Single Image Noise Level Estimation Using Dark Channel Prior
abstract
Noise level is required as an input parameter in various image processing applications. In this work, we use the dark channel prior (DCP) to estimate the noise level of an image degraded by additive white Gaussian noise. We develop an approximate model of the probability density function of the dark channel of the noisy image. Using this model, the noise level is determined with the maximum likelihood estimation method from the dark channel intensity values of the noisy image. The results show that our method is faster than the state-of-the-art methods by about two orders of magnitude while providing slightly inferior estimation performance.
Aziz Berkay Yesilyurt, Aybüke Erol, Fatih Kamisli, A. Aydin Alatan
ICIP3
2019 Lossless Image and Intra-Frame Compression With Integer-to-Integer DST
abstract
Video coding standards are primarily designed for efficient lossy compression, but it is also desirable to support efficient lossless compression within video coding standards using small modifications to the lossy coding architecture. A simple approach is to skip transform and quantization, and simply entropy code the prediction residual. However, this approach is inefficient at compression. A more efficient and popular approach is to skip transform and quantization but also process the residual block in some modes with differential pulse code modulation (DPCM), along the horizontal or vertical direction, prior to entropy coding. This paper explores an alternative approach based on processing the residual block with integer-to-integer (i2i) transforms. I2i transforms can map integer pixels to integer transform coefficients without increasing the dynamic range and can be used for lossless compression. We focus on lossless intra coding and develop novel i2i approximations of the odd type-3 discrete sine transform (ODST-3). Experimental results with the high efficiency video coding (HEVC) reference software show that when the developed i2i approximations of the ODST-3 are used along the DPCM method of HEVC, an average 2.7% improvement of lossless intra frame compression efficiency is achieved over HEVC version 2, which uses only the DPCM method, without a significant increase in computational complexity.
Fatih Kamisli
IEEE Trans. Circuits Syst. Video Technol.1
2017 Statistical analysis and directional coding of layer-based HDR image coding residue
abstract
Existing methods for layer-based backward compatible high dynamic range (HDR) image and video coding mostly focus on the rate-distortion optimization of base layer while neglecting the encoding of the residue signal in the enhancement layer. Although some recent studies handle residue coding by designing function based fixed global mapping curves for 8-bit conversion and exploiting standard codecs on the resulting 8-bit images, they do not take the local characteristics of residue blocks into account. Inspired by the local anisotropic characteristics of the residue signal and directional methods for motion compensated low dynamic range (LDR) video coding, in this paper we first investigate whether HDR image coding residue exhibits also local anisotropic characteristics. Specifically, we verify directional structures in residue blocks by means of auto-covariance analysis for different bitrates, spatial activities and dynamic ranges as the main variables in HDR image coding. Then, we compare the rate distortion performances of directional coding methods with the baseline residue coding methods in the literature along with different combinations of 8-bit conversion methods. The experiments indicate that content dependent 8-bit conversions and directional coding significantly outperforms the existing function based 8-bit conversions and typical coding for residue coding.
Kutan Feyiz, Fatih Kamisli, Emin Zerman, Giuseppe Valenzise, Alper Koz, Frédéric Dufaux
MMSP2
2016 Lossless compression in HEVC with integer-to-integer transforms
abstract
Many approaches have been proposed to support lossless coding within video coding standards that are primarily designed for lossy coding. The simplest approach is to just skip transform and quantization and directly entropy code the prediction residual, which is used in HEVC version 1. However, this simple approach is inefficient for compression. More efficient approaches include processing the residual with DPCM prior to entropy coding. This paper explores an alternative approach based on processing the residual with integer-to-integer (i2i) transforms. I2i transforms map integers to integers, however, unlike the integer transforms used in HEVC for lossy coding, they do not increase the dynamic range at the output and can be used in lossless coding. Experiments with the HEVC reference software show competitive results.
Fatih Kamisli
MMSP1
2016 On lossless intra coding in HEVC with 3-tap filters
Saeed Ranjbar Alvar, Fatih Kamisli
Signal Process. Image Commun.2
2015 Multiview video compression with 1-D transforms
Burcu Karasoy, Fatih Kamisli
Signal Process. Image Commun.2
2015 Block-Based Spatial Prediction and Transforms Based on 2D Markov Processes for Image and Video Compression
abstract
Conventional intraframe coding is performed in two steps. First, a block of pixels are predicted by copying previously reconstructed neighbor pixels of the block along an angular direction inside the block. Then, the prediction residual block is transform coded with the well-known 2D discrete cosine transform (DCT). Recently, it has been shown that transforming the intraprediction residuals with the odd type-3 discrete sine transform along the prediction direction and the DCT along the perpendicular direction improves the compression performance. More recently, a recursive prediction approach has been proposed to improve intra prediction performance. Both of these recent approaches utilize Markov processes to develop improvements in either the transform or the prediction step but not in both. In this paper, both the intraprediction and the transform steps are obtained based on 2D Markov processes. The derived overall intraframe coding approaches can generalize the mentioned two approaches, provide improved coding gains and produce less blocking effects at low bitrates.
Fatih Kamisli
IEEE Trans. Image Process.1
2014 Recursive Prediction for Joint Spatial and Temporal Prediction in Video Coding
abstract
Video compression systems use prediction to reduce redundancies present in video sequences along the temporal and spatial dimensions. Standard video coding systems use either temporal or spatial prediction on a per block basis. If temporal prediction is used, spatial information is ignored. If spatial prediction is used, temporal information is ignored. This may be a computationally efficient approach, but it does not effectively combine temporal and spatial information. In this letter, we provide a framework where available temporal and spatial information can be combined effectively to perform joint spatial and temporal prediction in video coding. Experimental results obtained from one sample realization of this framework show its potential.
Fatih Kamisli
IEEE Signal Process. Lett.1
2013 Intra prediction based on Markov process modeling of images
abstract
In our previous work, we proposed a new approach to intra prediction, in which we model image pixels with a separable first-order Markov process. The used Markov process is separable and therefore the developed method was only applied to intra prediction with vertical, horizontal and DC modes. In this paper, we extend our previous work by developing intra prediction methods based on non-separable Markov models and apply them to intra prediction along any angular direction. Compared to general linear prediction approaches, in which each block pixel is predicted using a weighted sum of all neighbor pixels of the block, the proposed approach uses much fewer independent parameters and thus offers reduced memory or computation requirements, while achieving similar coding gains.
Fatih Kamisli
ICIP1
2013 Intra Prediction Based on Markov Process Modeling of Images
abstract
In recent video coding standards, intraprediction of a block of pixels is performed by copying neighbor pixels of the block along an angular direction inside the block. Each block pixel is predicted from only one or few directionally aligned neighbor pixels of the block. Although this is a computationally efficient approach, it ignores potentially useful correlation of other neighbor pixels of the block. To use this correlation, a general linear prediction approach is proposed, where each block pixel is predicted using a weighted sum of all neighbor pixels of the block. The disadvantage of this approach is the increased complexity because of the large number of weights. In this paper, we propose an alternative approach to intraprediction, where we model image pixels with a Markov process. The Markov process model accounts for the ignored correlation in standard intraprediction methods, but uses few neighbor pixels and enables a computationally efficient recursive prediction algorithm. Compared with the general linear prediction approach that has a large number of independent weights, the Markov process modeling approach uses a much smaller number of independent parameters and thus offers significantly reduced memory or computation requirements, while achieving similar coding gains with offline computed parameters.
Fatih Kamisli
IEEE Trans. Image Process.1
2012 Intra prediction based on statistical modeling of images
abstract
Intra prediction is an important part of intra-frame coding. A number of approaches have been proposed to improve intra prediction including a general linear prediction approach in which a weighted sum of all available neighbor pixels is used to predict each block pixel. An important part of this approach is the determination of the used weights. One method to determine the weights is to use the least-squares solution of an overdetermined linear system of weights. In this paper, we present an alternative approach where the weights are determined based on statistical modeling of image pixels. This approach results in an analytical expression for the weights and can achieve similar coding gains as methods based on least-squares solutions of overdetermined systems, while having several benefits such as reduced storage or computations.
Fatih Kamisli
VCIP1
2011 1-D Transforms for the Motion Compensation Residual
abstract
Transforms used in image coding are also commonly used to compress prediction residuals in video coding. Prediction residuals have different spatial characteristics from images, and it is useful to develop transforms that are adapted to prediction residuals. In this paper, we explore the differences between the characteristics of images and motion compensated prediction residuals by analyzing their local anisotropic characteristics and develop transforms adapted to the local anisotropic characteristics of these residuals. The analysis indicates that many regions of motion compensated prediction residuals have 1-D anisotropic characteristics and we propose to use 1-D directional transforms for these regions. We present experimental results with one example set of such transforms within the H.264/AVC codec and the results indicate that the proposed transforms can improve the compression efficiency of motion compensated prediction residuals over conventional transforms.
Fatih Kamisli, Jae S. Lim
IEEE Trans. Image Process.1
2010 Video compression with 1-D directional transforms in H.264/AVC
abstract
Typically the same transforms, such as the 2-D Discrete Cosine Transform (DCT), are used to compress both images in image compression and prediction residuals in video compression. However, these two signals have different spatial characteristics. In, we analyzed the difference between these two signals and proposed 1-D directional transforms for prediction residuals. In this paper, we provide further experimental results using these transforms in the H.264/AVC codec and present other related information which can provide insights in understanding the use of these transforms in video coding applications.
Fatih Kamisli, Jae S. Lim
ICASSP1
2009 Transforms for the motion compensation residual
abstract
The Discrete-Cosine-Transform (DCT) is the most widely used transform in image and video compression. Its use in image compression is often justified by the notion that it is the statistically optimal transform for first-order Markov signals, which have been used to model images. In standard video codecs, the motion-compensation residual (MC-residual) is also compressed with the DCT. The MC-residual may, however, possess different characteristics from an image. Hence, the question that arises is if other transforms can be developed that can perform better on the MC-residual than the DCT. Inspired by recent research on direction-adaptive image transforms, we provide an adaptive auto-covariance characterization for the MC-residual that shows some statistical differences between the MC-residual and the image. Based on this characterization, we propose a set of block transforms. Experimental results indicate that these transforms can improve the compression efficiency of the MC-residual.
Fatih Kamisli, Jae S. Lim
ICASSP1
2009 Directional wavelet transforms for prediction residuals in video coding
abstract
Various directional transforms have been developed recently to improve image compression. In video compression, however, prediction residuals of image intensities, such as the motion compensation residual or the resolution enhancement residual, are transformed. The applicability of the directional transforms on prediction residuals have not been carefully investigated. In this paper, we briefly discuss differing characteristics of prediction residuals and images, and propose directional transforms specifically designed for prediction residuals. We compare these transforms with the directional transforms proposed for images using prediction residuals. The results of the comparison indicate that our proposed directional transforms can provide better compression of prediction residuals than the directional transforms proposed for images.
Fatih Kamisli, Jae S. Lim
ICIP1
2007 Estimation of Fade and Dissolve Parameters for Weighted Prediction in H.264/AVC
abstract
Weighted prediction (WP) is one way of overcoming the limitations of simple motion compensation for scenes with gradual transitions such as fades or dissolves. In WP, the predictions for inter coded blocks are obtained from scaled versions of the reference frames. H.264/AVC is the first video coding standard that has incorporated WP tools. In this paper, we focus on the estimation of weights for WP in dissolves. Our findings indicate that the estimation approaches for dissolves should be different from the estimation approaches for fades. Specifically, estimating the weights jointly for the two reference frames of B-frames gives better performance for dissolves under most circumstances.
Fatih Kamisli, David M. Baylon
ICIP (5)1