Mounir Kaaniche

dblp:97/7871 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-1874-3243ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 ZHFD-Net: Zero-shot image dehazing via high-frequency enhancement and depth-guided optimization
Lishuang Qi, Mounir Kaaniche, Meng Zhao 0001, Yiqiao Wan
Signal Process. Image Commun.2
2026 Low-Bitrate Light Field Video Compression Through Key Sequences Encoding and Joint Reconstruction Network
abstract
Light field (LF) videos contain rich spatial, angular, and temporal information, resulting in immense data volumes and posing significant challenges for low-bitrate compression. Existing LF video compression methods focus on modifying the structure of traditional video codecs to encode all LF views, but they are insufficient to achieve low-bitrate compression of LF video. To address these limitations, we propose a low-bitrate LF video compression framework that exploits spatial-angular-temporal correlations through sparse coding and joint reconstruction. On the encoding side, we introduce a content-adaptive prediction structure for sparse key view sequences selection. This structure is adapted to LF video content, leveraging the most similar view as a reference to enhance prediction accuracy and significantly reduce bitrate. On the decoding side, we observe that pixels missing in the current view are often captured in adjacent angular and/or temporal views. As a result, we develop a spatial-angular-temporal based joint reconstruction network that integrates cues across the different domains. This approach supplements missing texture details near occlusion areas and reconstructs high-quality non-key views. Experimental results demonstrate the efficiency of our framework, achieving an average gain of about 60 % in terms of bitrate savings and 2 dB in terms of reconstruction quality compared to the state-of-the-art methods.
Xinpeng Huang, Chao Yang 0021, Mounir Kaaniche, Qiuwen Zhang, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Capture More, Synthesize Better: Video Frame Interpolation With Larger Receptive Field and Structural Priors
abstract
A large receptive field is crucial for the video frame interpolation (VFI) task. Existing video frame interpolation methods struggle with large motions due to their limited receptive fields. However, simply expanding the receptive field brings two challenges: a substantial computational burden and potential loss of texture details. In this paper, we first propose a novel spatial-temporal global window self-attention mechanism with an enlarged receptive field to enhance motion capture. Furthermore, to reduce the computational complexity introduced by the global window, we design a simple and effective separable fence window decomposition. Meanwhile, to better synthesize high-quality intermediate frames, we propose two complementary frame synthesis strategies. First, from the perspective of receptive field design, we introduce a progressive receptive field focusing module, enabling a smooth transition from global motion modeling to local detail preservation. Second, based on the VFI-specific property and the high structural similarity shared by the adjacent frames, we propose a structure-aware synthesis strategy, which incorporates structural priors to guide the generation of fine details. Subjective and objective experimental results demonstrate that our method effectively captures large motions while synthesizing texture details, outperforming state-of-the-art techniques on various datasets.
Baojun Zhou, Xinpeng Huang, Jieyu Chen, Mounir Kaaniche, Ping An 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Probabilistic-Based Learning for Joint Light Field Image Compression and Enhancement Under Low-Light Conditions
abstract
Light field (LF) imaging has attracted increasing research interest in challenging illumination conditions due to its ability to provide rich spatial and angular cues. However, such data present dual challenges: 1) the inherent multi-view structure introduces substantial data redundancy, creating high demands for efficient compression; 2) the insufficient illumination leads to severe quality degradation, which weakens inter-view consistency and visual perception. To address these coupled factors, we propose a Probabilistic-based learning for joint LF image compression and enhancement under low-light conditions (PrL-LFCE). The framework unifies structure-aware compression and feature enhancement mechanisms by introducing learnable probabilistic modeling into both feature coupling and latent distribution estimation to adaptively handle the uncertainty induced by illumination degradation and compression-related information loss. Specifically, we design a probability-based multi-directional feature coupling module that dynamically balances structural preservation and redundancy reduction across multiple directionally arranged sub-aperture images. Moreover, we introduce a swin-gated enhancement module that suppresses noise and highlights structurally salient regions in compression-aware feature representations through attention-guided gating. Extensive experiments show that PrL-LFCE consistently outperforms state-of-the-art methods, achieving at least 34.86% bitrate savings while maintaining excellent visual quality, demonstrating a strong joint compression and enhancement capability.
Deyang Liu, Jimin Wang, Mounir Kaaniche, Xiaofei Zhou 0003, Gangyi Jiang, Caifeng Shan
IEEE Trans. Image Process.4
2025 SR-MedTS: a Semantically-Robust Synergistic Framework for Multi-modal Medical Image Translation and Segmentation
abstract
Multi-modal medical image translation and segmentation are essential for achieving accurate diagnosis and treatment. However, existing methods often suffer from semantic shifts during modality translation, leading to issues such as vascular discontinuity and anatomical deformation, which further degrade downstream segmentation performance. To address these challenges, we propose a Semantic-Robust multimodal medical image Translation and Segmentation framework (SR-MedTS). In this respect, we design an end-to-end dualstream architecture composed of a generator, discriminator and semantic robustness module. Moreover, we propose a channelspatial attention mechanism embedded in the skip connections between the encoder and decoder of the segmentation generator network to enhance boundary recognition. Finally, a multiloss function is defined to optimize the overall architecture. Extensive experiments on a multi-modal abdominal dataset and a six-site prostate dataset demonstrate that SR-MedTS significantly improves cross-modal segmentation performance using only 10 % of annotated data, showing strong potential in lowresource medical imaging scenarios. Our code is available at https://github.com/susu337/SR-MedTS.
Mingbo Su, Yi Zhang 0111, Mounir Kaaniche, Meng Zhao 0001
BIBM3
2025 UNEM: UNrolled Generalized EM for Transductive Few-Shot Learning
abstract
Transductive few-shot learning has recently triggered wide attention in computer vision. Yet, current methods introduce key hyper-parameters, which control the prediction statistics of the test batches, such as the level of class balance, affecting performances significantly. Such hyper-parameters are empirically grid-searched over validation data, and their configurations may vary substantially with the target dataset and pre-training model, making such empirical searches both sub-optimal and computationally intractable. In this work, we advocate and introduce the unrolling paradigm, also referred to as "learning to optimize", in the context of few-shot learning, thereby learning efficiently and effectively a set of optimized hyperparameters. Specifically, we unroll a generalization of the ubiquitous Expectation-Maximization (EM) optimizer into a neural network architecture, mapping each of its iterates to a layer and learning a set of key hyper-parameters over validation data. Our unrolling approach covers various statistical feature distributions and pre-training paradigms, including recent foundational vision-language models and standard vision-only classifiers. We report comprehensive experiments, which cover a breadth of fine-grained downstream image classification tasks, showing significant gains brought by the proposed unrolled EM algorithm over iterative variants. The achieved improvements reach up to 10% and 7.5% on vision-only and vision-language benchmarks, respectively. The source code and learned parameters are available at https://github.com/ZhouLong0/UNEM-Transductive.
Fereshteh Shakeri, Aymen Sadraoui, Mounir Kaaniche, Jean-Christophe Pesquet, Ismail Ben Ayed
CVPR4
2025 Multi-task convolution neural network-based lifting scheme for image compression
Tassnim Dardouri, Mounir Kaaniche, Amel Benazza-Benyahia, Gabriel Dauphin
Pattern Recognit. Lett.2
2025 SU3plus: An Enhanced Swin-UNet for Synthesizing Vertically Integrated Liquid From Multiple Meteorological Satellite Data
abstract
Radar composite reflectivity is a crucial component of Earth observation data, playing a significant role in applications such as weather forecasting and climate disaster tracking. Due to the deployment challenges and limited coverage of meteorological radars, it is impossible to collect corresponding radar reflectivity in areas such as mountains and oceans. In such cases, using deep learning methods to reconstruct radar reflectivity from meteorological satellite data becomes an effective solution. However, Earth observation data is complex and exhibits strong long-range dependencies. Such data characteristics require the ability to model long-distance dependencies, making the Transformer more suitable for this specific data. With the ultimate goal of producing high reconstruction quality, we propose in this paper a novel architecture, designated by SU3plus, based on the Swin Transformer with a multi-scale feature fusion mechanism, while leveraging multiple channels of satellite data. Moreover, we define an appropriate loss function by combining the weighted versions of standard metrics, to take into account the data distribution imbalance and improve the reconstruction performance. Extensive experiments, carried out on the SEVIR storm dataset, confirm the effectiveness of the proposed approach compared to several state-of-the-art models.
Zhixuan Zhou, Xintong Zhao, Mounir Kaaniche, Ding Liu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Unrolled Projected Gradient Algorithm For Stain Separation In Digital Histopathological Images
abstract
This paper introduces a novel optimization approach for stain separation in digital histopathological images. Our stain separation cost function incorporates a smooth total variation regularization and is minimized by using a projected gradient algorithm. To enhance computational efficiency and enable supervised learning of the hyperparameters, we further unroll our algorithm into a neural network. The unrolled architecture is not only more efficient for solving the stain separation problem, but also allows to design a highly interpretable and flexible method. Experimental results demonstrate the effectiveness of the proposed unrolled projected gradient algorithm in achieving accurate and visually consistent stain separation.
Aymen Sadraoui, Astrid Laurent-Bellue, Mounir Kaaniche, Amel Benazza-Benyahia, Catherine Guettier, Jean-Christophe Pesquet
ICIP3
2024 DSNet: A dynamic squeeze network for real-time weld seam image segmentation
Fan Shi 0001, Mounir Kaaniche, Meng Zhao 0001, Yan Jing, Shengyong Chen
Eng. Appl. Artif. Intell.4
2024 Joint Learning of Fully Connected Network Models in Lifting Based Image Coders
abstract
The optimization of prediction and update operators plays a prominent role in lifting-based image coding schemes. In this paper, we focus on learning the prediction and update models involved in a recent Fully Connected Neural Network (FCNN)-based lifting structure. While a straightforward approach consists in separately learning the different FCNN models by optimizing appropriate loss functions, jointly learning those models is a more challenging problem. To address this problem, we first consider a statistical model-based entropy loss function that yields a good approximation to the coding rate. Then, we develop a multi-scale optimization technique to learn all the FCNN models simultaneously. For this purpose, two loss functions defined across the different resolution levels of the proposed representation are investigated. While the first function combines standard prediction and update loss functions, the second one aims to obtain a good approximation to the rate-distortion criterion. Experimental results carried out on two standard image datasets, show the benefits of the proposed approaches in the context of lossy and lossless compression.
Tassnim Dardouri, Mounir Kaaniche, Amel Benazza-Benyahia, Gabriel Dauphin, Jean-Christophe Pesquet
IEEE Trans. Image Process.2
2023 NNCD-IQA: A new neural networks based compressed database for image quality assessment
Zohaib Amjad Khan, Tassnim Dardouri, Mounir Kaaniche, Gabriel Dauphin
Multim. Tools Appl.3
2023 A novel multi-branch wavelet neural network for sparse representation based object classification
Tan-Sy Nguyen, Marie Luong, Mounir Kaaniche, Long H. Ngo, Azeddine Beghdadi
Pattern Recognit.3
2022 A New Video Quality Assessment Dataset for Video Surveillance Applications
abstract
In this paper, we propose a new comprehensive Video Surveillance Quality Assessment Dataset (VSQuAD) dedicated to Video Surveillance (VS) systems. In contrast to other public datasets, this one contains many more videos with distortions and diversified content from common video surveillance scenarios. These videos have been artificially degraded with various types of distortions (single distortion or multiple distortions simultaneously) at different severity levels. In order to improve the efficiency of the surveillance systems and the versatility of the video quality assessment dataset, night vision CCTV videos are also included. Furthermore, a comprehensive analysis of the content in terms of diversity and challenging problems is also presented in this study. The interest of such database is twofold. First, it will serve for benchmarking different video distortion detection and classification algorithms. Second, it will be useful for the design of learning models for various challenging VS problems such as identification and removal of the most common distortions. The complete dataset is made publicly available as part of a challenge session in this conference through the following link: https://www.l2ti.univ-paris13.fr/VSQuad/.
Azeddine Beghdadi, Muhammad Ali Qureshi, Borhen-Eddine Dakkar, Hammad Hassan Gillani, Zohaib Amjad Khan, Mounir Kaaniche, Mohib Ullah, Faouzi Alaya Cheikh
ICIP6
2022 Dynamic Neural Network for Lossy-to-Lossless Image Coding
abstract
Lifting-based wavelet transform has been extensively used for efficient compression of various types of visual data. Generally, the performance of such coding schemes strongly depends on the lifting operators used, namely the prediction and update filters. Unlike conventional schemes based on linear filters, we propose, in this paper, to learn these operators by exploiting neural networks. More precisely, a classical Fully Connected Neural Network (FCNN) architecture is firstly employed to perform the prediction and update. Then, we propose to improve this FCNN-based Lifting Scheme (LS) in order to better take into account the input image to be encoded. Thus, a novel dynamical FCNN model is developed, making the learning process adaptive to the input image contents for which two adaptive learning techniques are proposed. While the first one resorts to an iterative algorithm where the computation of two kinds of variables is performed in an alternating manner, the second learning method aims to learn the model parameters directly through a reformulation of the loss function. Experimental results carried out on various test images show the benefits of the proposed approaches in the context of lossy and lossless image compression.
Tassnim Dardouri, Mounir Kaaniche, Amel Benazza-Benyahia, Jean-Christophe Pesquet
IEEE Trans. Image Process.2
2021 A Neural Network Approach For Joint Optimization Of Predictors In Lifting-Based Image Coders
abstract
The objective of this paper is to investigate techniques for learning Fully Connected Network (FCN) models in a lifting based image coding scheme. More precisely, based on a 2D non separable lifting structure composed of three FCN-based prediction stages followed by an FCN-based update one, we first propose to resort to an $\ell_{p}$ loss function, with $p\in\{1,2\}$, to learn the three FCN prediction models. While the latter are separately learned in the first approach, a novel joint learning approach is then developed by minimizing a weighted $\ell_{p}$ loss function related to the global prediction error. Experimental results, carried out on the standard Challenge Learned Image Compression (CLIC) dataset, show the benefits of the proposed techniques in terms of rate-distortion performance.
Tassnim Dardouri, Mounir Kaaniche, Amel Benazza-Benyahia, Jean-Christophe Pesquet, Gabriel Dauphin
ICIP2
2020 Optimized Lifting Scheme Based on A Dynamical Fully Connected Network for Image Coding
abstract
Wavelet decompositions based on lifting schemes have been widely used in image coding. Generally, the efficiency of such compression methods strongly depends on the design of the lifting operators, namely the prediction and update filters. To improve their performance, we propose in this paper to optimize these filters by resorting to two learning strategies. In the first one, classical Fully Connected Networks (FCNs) are exploited to perform the prediction and update. In the second approach, we develop an adaptive learning method that takes into account the input image, yielding a dynamical model of FCN. Experimental results, carried out on the standard Challenge Learned Image Compression (CLIC) dataset, show the benefits that can be drawn from the proposed approaches compared to conventional ones.
Tassnim Dardouri, Mounir Kaaniche, Amel Benazza-Benyahia, Jean-Christophe Pesquet
ICIP2
2020 Residual Networks Based Distortion Classification and Ranking for Laparoscopic Image Quality Assessment
abstract
Laparoscopic images and videos are often affected by different types of distortion like noise, smoke, blur and nonuniform illumination. Automatic detection of these distortions, followed generally by application of appropriate image quality enhancement methods, is critical to avoid errors during surgery. In this context, a crucial step involves an objective assessment of the image quality, which is a two-fold problem requiring both the classification of the distortion type affecting the image and the estimation of the severity level of that distortion. Unlike existing image quality measures which focus mainly on estimating a quality score, we propose in this paper to formulate the image quality assessment task as a multi-label classification problem taking into account both the type as well as the severity level (or rank) of distortions. Here, this problem is then solved by resorting to a deep neural networks based approach. The obtained results on a laparoscopic image dataset show the efficiency of the proposed approach.
Zohaib Amjad Khan, Azeddine Beghdadi, Mounir Kaaniche, Faouzi Alaya Cheikh
ICIP3
2020 A Multi-Criteria Contrast Enhancement Evaluation Measure using Wavelet Decomposition
abstract
An effective contrast enhancement method should not only improve the perceptual quality of an image but should also avoid adding any artifacts or affecting naturalness of images. This makes Contrast Enhancement Evaluation (CEE) a challenging task in the sense that both the improvement in image quality and unwanted side-effects need to be checked for. Currently, there is no single CEE metric that works well for all kinds of enhancement criteria. In this paper, we propose a new Multi-Criteria CEE (MCCEE) measure which combines different metrics effectively to give a single quality score. In order to fully exploit the potential of these metrics, we have further proposed to apply them on the decomposed image using wavelet transform. This new metric has been tested on two natural image contrast enhancement databases as well as on medical Computed Tomography (CT) images. The results show a substantial improvement as compared to the existing evaluation metrics. The code for the metric is available at: https://github.com/zakopz/MCCEE-Contrast-Enhancement-Metric.
Zohaib Amjad Khan, Azeddine Beghdadi, Faouzi Alaya Cheikh, Mounir Kaaniche, Muhammad Ali Qureshi
MMSP4
2020 Convolution Autoencoder-Based Sparse Representation Wavelet for Image Classification
abstract
In this paper, we propose an effective Convolutional Autoencoder (AE) model for Sparse Representation (SR) in the Wavelet Domain for Classification (SRWC). The proposed approach involves an autoencoder with a sparse latent layer for learning sparse codes of wavelet features. The estimated sparse codes are used for assigning classes to test samples using a residual-based probabilistic criterion. Intensive experiments carried out on various datasets revealed that the proposed method yields better classification accuracy while exhibiting a significant reduction in the number of network parameters, compared to several recent deep learning-based methods.
Tan-Sy Nguyen, Long H. Ngo, Marie Luong, Mounir Kaaniche, Azeddine Beghdadi
MMSP4
2019 Depth-based color stereo images retrieval using joint multivariate statistical models
Emna Ghodhbani, Mounir Kaaniche, Amel Benazza-Benyahia
Signal Process. Image Commun.2
2019 Close Approximation of Kullback-Leibler Divergence for Sparse Source Retrieval
abstract
In this letter, we propose a fast and accurate approximation of the Kullback-Leibler divergence (KLD) between two Bernoulli-Generalized Gaussian (Ber-GG) distributions. Such a distribution has been found to be well suited for modeling sparse signals like wavelet-based representations. On the basis of high-bitrate approximations of the entropy of quantized Ber-GG sources, we provide a close approximation of the KLD without resorting to the conventional time-consuming Monte Carlo estimation approach. The developed approximation formula is then validated in the context of depth map and stereo image retrieval.
Emna Ghodhbani, Mounir Kaaniche, Amel Benazza-Benyahia
IEEE Signal Process. Lett.2
2019 Efficient Enhancement of Stereo Endoscopic Images Based on Joint Wavelet Decomposition and Binocular Combination
abstract
The success of minimally invasive interventions and the remarkable technological and medical progress have made endoscopic image enhancement a very active research field. Due to the intrinsic endoscopic domain characteristics and the surgical exercise, stereo endoscopic images may suffer from different degradations which affect its quality. Therefore, in order to provide the surgeons with a better visual feedback and improve the outcomes of possible subsequent processing steps, namely, a 3-D organ reconstruction/registration, it would be interesting to improve the stereo endoscopic image quality. To this end, we propose, in this paper, two joint enhancement methods which operate in the wavelet transform domain. More precisely, by resorting to a joint wavelet decomposition, the wavelet subbands of the right and left views are simultaneously processed to exploit the binocular vision properties. While the first proposed technique combines only the approximation subbands of both views, the second method combines all the wavelet subbands yielding an inter-view processing fully adapted to the local features of the stereo endoscopic images. Experimental results, carried out on various stereo endoscopic datasets, have demonstrated the efficiency of the proposed enhancement methods in terms of perceived visual image quality.
Bilel Sdiri, Mounir Kaaniche, Faouzi Alaya Cheikh, Azeddine Beghdadi, Ole Jakob Elle
IEEE Trans. Medical Imaging2
2018 Sparse optimization of non separable vector lifting scheme for stereo image coding
I. Bezzine, Mounir Kaaniche, Saadi Boudjit, Azeddine Beghdadi
J. Vis. Commun. Image Represent.2
2018 Efficient transform-based texture image retrieval techniques under quantization effects
Amani Chaker, Mounir Kaaniche, Amel Benazza-Benyahia, Marc Antonini
Multim. Tools Appl.2
2017 No-reference stereo image quality assessment based on joint wavelet decomposition and statistical models
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
Signal Process. Image Commun.2
2015 Block dependent dictionary based disparity compensation for stereo image coding
abstract
With the recent advances in stereoscopic display technologies, there is a growing demand for designing efficient stereo image compression techniques. For this reason, a great attention should be paid to the disparity/estimation process used to generate the residual image. In this paper, we propose to improve the disparity compensation process in a typical closed-loop-based stereo image coding scheme. A new formulation of this process, based on a block dependent dictionary, is developed. More specifically, the main idea aims to link together the disparities yielding similar compensations and assign a common disparity candidate to each subset of disparities. Experimental results have shown the interest of the proposed method in terms of bitrate saving and quality of reconstruction.
Gabriel Dauphin, Mounir Kaaniche, Anissa Zergaïnoh-Mokraoui
ICIP2
2015 Optimized lifting schemes based on ENO stencils for image approximation
abstract
In this paper, we propose to improve the classical lifting-based wavelet transforms by defining three classes of pixels which will be predicted differently. More specifically, the proposed idea is inspired by the Essentially Non-Oscillatory (ENO) transform and consists in shifting the stencil used for prediction in order to reduce the error near image singularities. Moreover, the different filters associated with these classes will be optimized in order to design a multiresolution representation well adapted to image characteristics. Our simulations show that the resulting multiscale representation leads to much lower amplitudes of the detail coefficients and improves the linear approximation properties.
Mounir Kaaniche, Basarab Matei, Sylvain Meignen
ICIP1
2015 Disparity based stereo image retrieval through univariate and bivariate models
Amani Chaker, Mounir Kaaniche, Amel Benazza-Benyahia
Signal Process. Image Commun.2
2015 Efficient Inter-View Bit Allocation Methods for Stereo Image Coding
abstract
In this paper, we present efficient bit allocation methods for stereo image coding purpose. Since the common idea behind most of the existing stereo compression schemes consists of encoding a reference and residual images as well as a disparity map, we mainly focus on the bit allocation issue between the reference and residual images. Generally, this problem is solved in an empirical manner by looking for the optimal rates leading to the minimum distortion value. Thanks to recent approximations of the entropy and distortion functions, we propose accurate and fast bit allocation schemes appropriate for the open-loop- and closed-loop-based stereo coding structures. Experimental results show the benefits which can be drawn from the proposed bit allocation methods.
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
IEEE Trans. Multim.2
2014 Accurate rate-distortion approximation for sparse Bernoulli-Generalized Gaussian models
abstract
The objective of this paper is to study rate-distortion properties of a quantized Bernoulli-Generalized Gaussian source. Such source model has been found to be well-adapted for signals having a sparse representation in a transformed domain. We provide here accurate approximations of the entropy and the distortion functions evaluated through a p-th order error measure. These theoretical results are then validated experimentally. Finally, the benefit that can be drawn from the proposed approximations in bit allocation problems is illustrated for a wavelet-based compression scheme.
Mounir Kaaniche, Aurélia Fraysse, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet
ICASSP1
2014 Exploiting disparity information for stereo image retrieval
abstract
The great interest of stereo images in several applications has led to the proliferation of huge and ever growing image databases. Therefore, there is an urgent demand for an effective Content Based Image Retrieval (CBIR) system devoted to stereo images. To meet such a demand, this paper proposes new wavelet-based retrieval approaches that exploit not only the visual contents of the Stereo Image (SI) pair but also its related disparity field. The first approach takes into account implicitly the disparity information by computing features from the disparity compensated left image and the right image. The second one aims at extracting relevant features directly from the left and right views, and the disparity map. Experimental results indicate that adding disparity information allows us to improve the retrieval performances of stereo images.
Amani Chaker, Mounir Kaaniche, Amel Benazza-Benyahia
ICIP2
2014 Rate distortion optimal bit allocation for stereo image coding
abstract
Many research works have been developed for stereo image compression purpose where most of them aim at encoding a reference image, a residual one and a disparity map. While the disparity field is often losslessly encoded, we are mainly interested in this paper in the bit allocation problem between the reference and residual images. Generally, the bit allocation is expressed as an optimization problem which involves the computation of the operational rate-distortion (RD) functions for all the wavelet subbands and for different quantization steps. However, this strategy is computationally intensive. To solve this problem, we consider the uniform scalar quantization of the wavelet subbands of both images modeled by a Generalized Gaussian distribution. Thanks to recent approximations of the entropy and distortion functions, we develop an optimal and fast bit allocation method. The obtained results confirm the efficiency of the proposed bit allocation method in the context of stereo image coding.
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
ICIP2
2014 A Bit Allocation Method for Sparse Source Coding
abstract
In this paper, we develop an efficient bit allocation strategy for subband-based image coding systems. More specifically, our objective is to design a new optimization algorithm based on a rate-distortion optimality criterion. To this end, we consider the uniform scalar quantization of a class of mixed distributed sources following a Bernoulli-generalized Gaussian distribution. This model appears to be particularly well-adapted for image data, which have a sparse representation in a wavelet basis. In this paper, we propose new approximations of the entropy and the distortion functions using piecewise affine and exponential forms, respectively. Because of these approximations, bit allocation is reformulated as a convex optimization problem. Solving the resulting problem allows us to derive the optimal quantization step for each subband. Experimental results show the benefits that can be drawn from the proposed bit allocation method in a typical transform-based coding application.
Mounir Kaaniche, Aurélia Fraysse, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet
IEEE Trans. Image Process.1
2013 An efficient retrieval strategy for wavelet-based quantized images
abstract
Recent research efforts have been devoted to the improvement of image retrieval systems when datasets are represented in a compressed form. In this context, new studies have shown that compression has a negative impact on the performances of the traditional retrieval systems. In this work, we are mainly interested in designing an efficient retrieval approach well adapted to wavelet-based compressed images. More precisely, we first propose to apply a compression scheme based on the Moment Preserving Quantization (MPQ). Then, the feature vectors will be defined in an appropriate way by focusing on the quantized subbands where some given statistical moments have been preserved. Experimental results indicate that the proposed approach outperforms the most recent one which involves the conventional uniform quantizer and constrains the query and the model images to have similar qualities during the retrieval step.
Amani Chaker, Mounir Kaaniche, Amel Benazza-Benyahia
ICASSP2
2012 Adaptive lifting schemes with a global ℓ1 minimization technique for image coding
abstract
Many existing works related to lossy-to-lossless image compression are based on the lifting concept. In this paper, we present a sparse optimization technique based on recent convex algorithms and applied to the prediction filters of a two-dimensional non separable lifting structure. The idea consists of designing these filters, at each resolution level, by minimizing the sum of the ℓ1-norm of the three detail subbands. Extending this optimization method in order to perform a global minimization over all resolution levels leads to a new optimization criterion taking into account linear dependencies between the generated coefficients. Simulations carried out on still images show the benefits which can be drawn from the proposed optimization techniques.
Mounir Kaaniche, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet, Amel Benazza-Benyahia
ICIP1
2012 A convex programming bit allocation method for sparse sources
abstract
The objective of this paper is to design an efficient bit allocation algorithm in the subband coding context based on an analytical approach. More precisely, we consider the uniform scalar quantization of subband coefficients modeled by a Generalized Gaussian distribution. This model appears to be particularly well-adapted for data having a sparse representation in the wavelet domain. Our main contribution is to reformulate the bit allocation problem as a convex programming one. For this purpose, we firstly define new convex approximations of the entropy and distortion functions. Then, we derive explicit expressions of the optimal quantization parameters. Finally, we illustrate the application of the proposed method to wavelet-based coding systems.
Mounir Kaaniche, Aurélia Fraysse, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet
PCS1
2011 Proximal splitting methods for depth estimation
abstract
Stereo matching is an active area of research in image processing. In a recent work, a convex programming approach was developed in order to generate a dense disparity field. In this paper, we address the same estimation problem and pro pose to solve it in a more general convex optimization frame work based on proximal methods. More precisely, unlike previous works where the criterion must satisfy some restrictive conditions in order to be able to numerically solve the minimization problem, this work offers a great flexibility in the choice of the involved criterion. The method is validated in a stereo image coding framework, and the results demonstrate the good performance of the proposed parallel proximal algorithm.
Mireille El Gheche, Jean-Christophe Pesquet, Joumana Farah, Mounir Kaaniche, Béatrice Pesquet-Popescu
ICASSP4
2011 Non-separable lifting scheme with adaptive update step for still and stereo image coding
Mounir Kaaniche, Amel Benazza-Benyahia, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet
Signal Process.1
2010 Two-dimensional non separable adaptive lifting scheme for still and stereo image coding
abstract
Many existing works related to lossy-to-lossless image compression are based on the lifting concept. However, it has been observed that the separable lifting scheme structure presents some limitations because of the separable processing performed along the image lines and columns. In this paper, we propose to use a 2D non separable lifting scheme decomposition that enables progressive reconstruction and exact decoding of images. More precisely, we focus on the optimization of all the involved decomposition operators. In this respect, we design the prediction filters by minimizing the variance of the detail signals. Concerning the update filters, we propose a new optimization criterion which aims at reducing the inherent aliasing artefacts. Simulations carried out on still and stereo images show the benefits which can be drawn from the proposed optimization of the lifting operators.
Mounir Kaaniche, Jean-Christophe Pesquet, Amel Benazza-Benyahia, Béatrice Pesquet-Popescu
ICASSP1
2009 Dense disparity map representations for stereo image coding
abstract
Research in stereo image coding has focused on the disparity estimation/compensation process to exploit the cross-view redundancies. Most of the reported methods use a classical block-based technique in order to estimate the disparity field. However, this estimation technique does not always provide an accurate disparity map, which may affect the disparity compensation step. In this paper, we propose to use an estimation method that produces a dense and smooth disparity map. Then, on the one hand, this map is segmented and efficiently coded by exploiting the high correlation between neighboring disparity values. On the other hand, we integrate the disparity information into a vector lifting scheme for stereo image coding. Experimental results indicate that the proposed coding scheme outperforms the conventional methods employing a block-based disparity estimation.
Mounir Kaaniche, Wided Miled, Béatrice Pesquet-Popescu, Amel Benazza-Benyahia, Jean-Christophe Pesquet
ICIP1
2009 Dense disparity estimation in multiview video coding
abstract
Multiview video coding is an emerging application where, in addition to classical temporal prediction, an efficient disparity prediction should be performed in order to achieve the best compression performance. A popular coder is the multiview video coding (MVC) extension of H.264/AVC, which uses a block-based disparity estimation (just like temporal prediction in H.264/AVC). In this paper, we propose to improve the MVC extension by using a dense estimation method that generates a smooth disparity map with ideally infinite precision. The obtained disparity is then segmented and efficiently encoded by using a rate-distortion optimization technique. Experimental results show that significant gains can be obtained compared to the block-based disparity estimation technique used in the MVC extension.
Ismaël Daribo, Mounir Kaaniche, Wided Miled, Marco Cagnazzo, Béatrice Pesquet-Popescu
MMSP2
2009 Vector Lifting Schemes for Stereo Image Coding
abstract
Many research efforts have been devoted to the improvement of stereo image coding techniques for storage or transmission. In this paper, we are mainly interested in lossy-to-lossless coding schemes for stereo images allowing progressive reconstruction. The most commonly used approaches for stereo compression are based on disparity compensation techniques. The basic principle involved in this technique first consists of estimating the disparity map. Then, one image is considered as a reference and the other is predicted in order to generate a residual image. In this paper, we propose a novel approach, based on vector lifting schemes (VLS), which offers the advantage of generating two compact multiresolution representations of the left and the right views. We present two versions of this new scheme. A theoretical analysis of the performance of the considered VLS is also conducted. Experimental results indicate a significant improvement using the proposed structures compared with conventional methods.
Mounir Kaaniche, Amel Benazza-Benyahia, Béatrice Pesquet-Popescu, Jean-Christophe Pesquet
IEEE Trans. Image Process.1