Faouzi Alaya Cheikh

dblp:c/FACheikh · DBLP profile ↗
← Back
48ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4823-5250ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Towards Scalable, Low-Latency Volumetric Streaming: A Hybrid Predictive Synchronization Framework
abstract
Volumetric video is emerging as a cornerstone for multi-user Extended Reality (XR) and metaverse applications. However, achieving synchronous playback across distributed clients remains challenging under heterogeneous network conditions (such as varying latency, jitter, packet loss, and asymmetric bandwidth). Existing synchronization approaches, such as state synchronization and fixed-frame buffering, struggle with scalability and often trade off latency for playback continuity. We present a predictive synchronization framework that combines lightweight time-series latency forecasting with adaptive buffering control. Our framework incorporates a hybrid predictor that adaptively selects the most suitable model (including EWMA, ARIMA, LSTM) for each client, escalating to heavier predictors only when residual errors exceed thresholds. This design balances accuracy with computational overhead, enabling lower latencies at larger scales. Our experimental results show that predictive synchronization reduces average inter-client skew by up to 76% compared to baselines, while also decreasing the average buffer size by 40% in a controlled multi-client testbed. The average per-frame latency remains below 20 ms and the 95th-percentile latency below 70 ms for up to 100 concurrent clients. These results demonstrate that prediction-aware, hybrid synchronization substantially improves quality of service while maintaining lightweight per-client overhead.
Robin Singh Sidhu, Hao Wang 0003, Faouzi Alaya Cheikh
ICC5
2025 PVD4RCV: A Photo-realistic Multi-Distortion Video Dataset for Benchmarking and Developing Robust Computer Vision Models
abstract
This work addresses a significant gap in existing image and video databases commonly used in computer vision applications by introducing a unique and comprehensive database named Photo-realistic Multi-Distortion Video Dataset for Benchmarking and Developing Robust Computer Vision Models (PVD4RCV). A key innovation of PVD4RCV lies in its incorporation of some relevant physical factors (e.g. depth information, interaction of light with scene contents) inherent to video signal acquisition in constrained and complex real-world environments, which are used to generate realistic distortions in video sequences (e.g. local motion blur, local defocus blur). PVD4RCV includes a diverse collection of videos featuring common distortions, real-world scenarios, and contextual variations. It includes both original and degraded video versions, along with detailed annotations to support the development of advanced learning models, particularly for tasks such as distortion classification and object detection. This resource aims to advance research and applications in computer vision by providing a robust foundation for model training and evaluation. The database is open source and available at the following link for the community: https://github.com/Aymanbegh/PVD4RCV
Ayman Beghdadi, Mohib Ullah, Azeddine Beghdadi, Borhen-Eddine Dakkar, Zohaib Amjad Khan, Faouzi Alaya Cheikh
VCIP6
2025 IoT-Driven Facial Expression Recognition for Personalized Healthcare in Industry 5.0
abstract
Facial emotion recognition (FER) plays a critical role in understanding human behavior, especially for individuals suffering from neurological disorders (NDs) like Parkinson’s disease (PD), Multiple Sclerosis (MS), and Stroke. Early and accurate detection of emotions is crucial for both the diagnosis of associated mood disorders and continuous monitoring. However, traditional methods often fall short in providing noninvasive, real-time solutions and lack the clinical expertise necessary to identify the specific emotion types associated with each ND category. In response, this research conducted under the ALAMEDA consortium presents an Internet of Things-based FER AI Toolkit designed to enhance early diagnosis and treatment for brain diseases. The toolkit is in line with the consortium’s clinical guidelines and provides a personalized, patient-focused solution that supports the goals of Industry 5.0 in healthcare. In line with Industry 5.0 principles, the FER AI Toolkit uses edge devices to collect real-time facial data while deep learning models running on cloud servers process this data. The recognized emotions are uploaded to the Semantic Knowledge Graph (SemKG) server. This allows healthcare professionals to make informed decisions based on real-time data. Additionally, the toolkit integrates seamlessly with key components of the ALAMEDA, including the Identity Authentication Manager (IAM) for secure access and the ALAMEDA Innovation Hub (AIH) for efficient resource management. By offering continuous and personalized healthcare insights, the FER AI Toolkit helps bridge the gap between diagnosis and patient well-being, ultimately advancing healthcare systems. Training materials and video demonstrations are available athttps://drive.google.com/drive/folders/1-iUz7FE2IrKHt5nMl3oGjsrCtk7Ps2bM?usp=sharingfor further learning.
Shehzad Ali, Ikhyun Lee, Faouzi Alaya Cheikh, Athena Cristina Ribigan, Ludovico Pedullà, Nikolaos Papagiannakis, Mohammad Hijji, Khan Muhammad 0001
IEEE Internet Things J.4
2025 Extrapolation Convolution for Data Prediction on a 2-D Grid: Bridging Spatial and Frequency Domains With Applications in Image Outpainting and Compressed Sensing
abstract
Extrapolation plays a critical role in machine/deep learning (ML/DL), enabling models to predict data points beyond their training constraints, particularly useful in scenarios deviating significantly from training conditions. This article addresses the limitations of current convolutional neural networks (CNNs) in extrapolation tasks within image restoration and compressed sensing (CS). While CNNs show potential in tasks such as image outpainting and CS, traditional convolutions are limited by their reliance on interpolation, failing to fully capture the dependencies needed for predicting values outside the known data. This work proposes an extrapolation convolution (EC) framework that models missing data prediction as an extrapolation problem using linear prediction within DL architectures. The approach is applied in two domains: first, image outpainting, where EC in encoder-decoder (EnDec) networks replaces conventional interpolation methods to reduce artifacts and enhance fine detail representation; second, Fourier-based CS-magnetic resonance imaging (CS-MRI), where it predicts high-frequency signal values from undersampled measurements in the frequency domain, improving reconstruction quality and preserving subtle structural details at high acceleration factors. Comparative experiments demonstrate that the proposed EC-DecNet and FDRN outperform traditional CNN-based models, achieving high-quality image reconstruction with finer details, as shown by improved peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and kernel inception distance (KID)/Frechet inception distance (FID) scores. Ablation studies and analysis highlight the effectiveness of larger kernel sizes and multilevel semi-supervised learning in FDRN for enhancing extrapolation accuracy in the frequency domain.
Vazim Ibrahim, Faouzi Alaya Cheikh, Vijayan K. Asari, Joseph Suresh Paul
IEEE Trans. Neural Networks Learn. Syst.2
2024 A Self-Supervised Diffusion Framework For Facial Emotion Recognition
abstract
In this paper, we introduced a novel Facial Emotion Recognition (FER) framework that utilizes a diffusion-based approach and an attention mechanism. The model is efficiently trained through self-supervised learning, leveraging labeled and unlabelled data. The proposed framework has been rigorously tested on the FER2013 and AffectNet datasets, achieving promising accuracies of $67.2 \%$ and $68.1 \%$, respectively. The quantitative results not only surpass the performance of existing state-of-the-art FER models but also demonstrate the synergistic effect of combining diffusion-based modeling with self-supervised learning and attention mechanisms within a solid architectural framework. Our approach sets a new benchmark in the field, offering a significant step forward in the accurate and efficient recognition of facial expressions.
Saif Hassan, Mohib Ullah, Ali Shariq Imran, Ghulam Mujtaba 0001, Muhammad Mudassar Yamin, Ehtesham Hashmi, Faouzi Alaya Cheikh, Azeddine Beghdadi
ICIP7
2024 CentralFormer: Centralized Spectral-Spatial Transformer for Hyperspectral Image Classification With Adaptive Relevance Estimation and Circular Pooling
Ningyang Li, Zhaohui Wang 0001, Faouzi Alaya Cheikh
IEEE Trans. Geosci. Remote. Sens.3
2023 Attention-Guided Self-supervised Framework for Facial Emotion Recognition
Saif Hassan, Mohib Ullah, Ali Shariq Imran, Faouzi Alaya Cheikh
PRICAI (3)4
2023 Deepview: Deep-Learning-Based Users Field of View Selection in 360° Videos for Industrial Environments
abstract
The industrial demands of immersive videos for virtual reality/augmented reality applications are crescendo, where the video stream provides a choice to the user viewing object of interest with the illusion of “being there.” However, in industry 4.0, streaming of such huge-sized video over the network consumes a tremendous amount of bandwidth, where the users are only interested in specific regions of the immersive videos. Furthermore, for delivering full excitement videos and minimizing the bandwidth consumption, the automatic selection of the user’s Region of Interest in a 360° video is very challenging because of subjectivity and difference in contentment. To tackle these challenges, we employ two efficient convolutional neural networks for salient object detection and memorability computation in a unified framework to find the most prominent portion of a 360° video. The proposed system is four-fold: 1) preprocessing; 2) intelligent visual interest predictor; 3) final viewport selection; and 4) virtual camera steerer. First, an input 360° video frame is split into three Field of Views (FoVs), each with a viewing angle of 120°. Next, each FoV is passed to the object detection and memorability prediction model for visual interestingness computation. Furthermore, the FoV is supplied as a viewport, containing the most salient and memorable objects. Finally, a virtual camera steerer is designed using enriched salient features from YOLO and LSTM that are forwarded to the dense optical flow to follow the salient object inside the immersive video. Performance evaluation of the proposed system over our own collected data from various Websites as well as on public data sets indicates the effectiveness for diverse categories of 360° videos and helps in the minimization of the bandwidth usage, making it suitable for industry 4.0 applications.
Khan Muhammad 0001, Khalid Mahmood 0003, Faouzi Alaya Cheikh, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque
IEEE Internet Things J.5
2022 A New Video Quality Assessment Dataset for Video Surveillance Applications
abstract
In this paper, we propose a new comprehensive Video Surveillance Quality Assessment Dataset (VSQuAD) dedicated to Video Surveillance (VS) systems. In contrast to other public datasets, this one contains many more videos with distortions and diversified content from common video surveillance scenarios. These videos have been artificially degraded with various types of distortions (single distortion or multiple distortions simultaneously) at different severity levels. In order to improve the efficiency of the surveillance systems and the versatility of the video quality assessment dataset, night vision CCTV videos are also included. Furthermore, a comprehensive analysis of the content in terms of diversity and challenging problems is also presented in this study. The interest of such database is twofold. First, it will serve for benchmarking different video distortion detection and classification algorithms. Second, it will be useful for the design of learning models for various challenging VS problems such as identification and removal of the most common distortions. The complete dataset is made publicly available as part of a challenge session in this conference through the following link: https://www.l2ti.univ-paris13.fr/VSQuad/.
Azeddine Beghdadi, Muhammad Ali Qureshi, Borhen-Eddine Dakkar, Hammad Hassan Gillani, Zohaib Amjad Khan, Mounir Kaaniche, Mohib Ullah, Faouzi Alaya Cheikh
ICIP8
2022 Deep learning for image-based liver analysis - A comprehensive review focusing on malignant lesions
abstract
Deep learning-based methods, in particular, convolutional neural networks and fully convolutional networks are now widely used in the medical image analysis domain. The scope of this review focuses on the analysis using deep learning of focal liver lesions, with a special interest in hepatocellular carcinoma and metastatic cancer; and structures like the parenchyma or the vascular system. Here, we address several neural network architectures used for analyzing the anatomical structures and lesions in the liver from various imaging modalities such as computed tomography, magnetic resonance imaging and ultrasound. Image analysis tasks like segmentation, object detection and classification for the liver, liver vessels and liver lesions are discussed. Based on the qualitative search, 91 papers were filtered out for the survey, including journal publications and conference proceedings. The papers reviewed in this work are grouped into eight categories based on the methodologies used. By comparing the evaluation metrics, hybrid models performed better for both the liver and the lesion segmentation tasks, ensemble classifiers performed better for the vessel segmentation tasks and combined approach performed better for both the lesion classification and detection tasks. The performance was measured based on the Dice score for the segmentation, and accuracy for the classification and detection tasks, which are the most commonly used metrics.
Shanmugapriya Survarachakan, Pravda Jith Ray Prasad, Rabia Naseem, Javier Pérez de Frutos, Rahul P. Kumar, Thomas Langø, Faouzi Alaya Cheikh, Ole Jakob Elle, Frank Lindseth
Artif. Intell. Medicine7
2022 Convolutional Neural Networks for Omnidirectional Image Quality Assessment: A Benchmark
abstract
In this paper, we conduct an extensive study on the use of pre-trained convolutional neural networks (CNNs) for omnidirectional image quality assessment (IQA). To cope with the lack of available IQA databases, transfer learning from seven pre-trained CNN models is investigated over retraining on standard 2D databases. In addition, we explore the influence of various image representations and training strategies on the model’s performance. A comparison of the use of projected versus radial content, and multichannel CNN versus patch-wise training is also covered. The experimental results on two publicly available databases are used to draw conclusions about which strategy best fits the visual quality prediction and at which computational cost. The analysis shows that retraining CNN models on 2D IQA databases improves the prediction accuracy. The latter and the required computational time are found to be significantly affected by the training strategy. Cross-database evaluations demonstrate that the nature and variety of the content impact the generalization ability of the models. Finally, we show that conclusions coming from other image processing communities may not hold for IQA. The provided discussion shall provide insights and recommendations when using pre-trained CNNs for omnidirectional IQA.
Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh
IEEE Trans. Circuits Syst. Video Technol.3
2021 3D-Resnet Fused Attention for Autism Spectrum Disorder Classification
Xiangjun Chen, Zhaohui Wang 0001, Faouzi Alaya Cheikh, Mohib Ullah
ICIG (2)3
2021 Perceptually-Weighted Cnn For 360-Degree Image Quality Assessment Using Visual Scan-Path And Jnd
abstract
Image quality assessment of immersive content and more specifically 360-degree one is still in its infancy. There are many challenges regarding sphere vs. projected representation, human visual system (HVS) properties in a 360-degree environment, etc. In this paper, we propose the use of CNNs to design a no reference model to predict visual quality of 360-degree images. Instead of feeding the CNN with ERPs, visually important viewports are extracted based on visual scan-path prediction and given to a multi-channel CNN using DenseNet-121. Moreover, information about visual fixations and just noticeable difference are used to account for the HVS properties and make the network closer to human judgment. The scan-path is also used to create multiple instances of the database so as to perform a robust generalization analysis and compensate for the lack of databases.
Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh
ICIP3
2021 Convolutional Neural Networks for Omnidirectional Image Quality Assessment: Pre-Trained or Re-Trained?
abstract
The use of convolutional neural networks (CNN) for image quality assessment (IQA) becomes many researcher’s focus. Various pre-trained models are fine-tuned and used for this task. In this paper, we conduct a benchmark study of seven state-of-the-art pre-trained models for IQA of omnidirectional images. To this end, we first train these models using an omnidirectional database and compare their performance with the pre-trained versions. Then, we compare the use of viewports versus equirectangular (ERP) images as inputs to the models. Finally, for the viewports-based models, we explore the impact of the input number of viewports on the models’ performance. Experimental results demonstrated the performance gain of the re-trained CNNs compared to their pre-trained versions. Also, the viewports-based approach outperformed the ERP-based one independently of the number of selected views.
Abderrezzaq Sendjasni, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh
ICIP3
2021 IR-SSL: Improved Regularization Based Semi-Supervised Learning For Land Cover Classification
abstract
Land cover classification has significant contributions in several applications including natural calamities estimation and response, observation of environmental changes, and urban planning to name a few. These types of applications demand the identification of different categories of land cover. Traditional land cover classification is significantly dependent on the availability of huge amount of labeled data. However, labeling satellite data is very time consuming and it often requires expert knowledge. To alleviate the dependency on labeled data, we propose a novel and improved regularization based deep semi-supervised learning (IR-SSL) method for land cover classification. Adaptation of deep semi-supervised learning approach in such a task gains reliability due to its robustness in feature learning. To consolidate the performance of our deep semi-supervised learning method, we combine it with a robust data augmentation technique. We perform experiments on a benchmark dataset. Considering limited labeled samples from the dataset, our method outperforms many state-of-the-art models.
Tawsin Uddin Ahmed, Mohib Ullah, Faouzi Alaya Cheikh
ICIP4
2020 Residual Networks Based Distortion Classification and Ranking for Laparoscopic Image Quality Assessment
abstract
Laparoscopic images and videos are often affected by different types of distortion like noise, smoke, blur and nonuniform illumination. Automatic detection of these distortions, followed generally by application of appropriate image quality enhancement methods, is critical to avoid errors during surgery. In this context, a crucial step involves an objective assessment of the image quality, which is a two-fold problem requiring both the classification of the distortion type affecting the image and the estimation of the severity level of that distortion. Unlike existing image quality measures which focus mainly on estimating a quality score, we propose in this paper to formulate the image quality assessment task as a multi-label classification problem taking into account both the type as well as the severity level (or rank) of distortions. Here, this problem is then solved by resorting to a deep neural networks based approach. The obtained results on a laparoscopic image dataset show the efficiency of the proposed approach.
Zohaib Amjad Khan, Azeddine Beghdadi, Mounir Kaaniche, Faouzi Alaya Cheikh
ICIP4
2020 A Bottom-Up Approach for Pig Skeleton Extraction Using RGB Data
Akif Quddus Khan, Mohib Ullah, Faouzi Alaya Cheikh
ICISP4
2020 A Multi-Criteria Contrast Enhancement Evaluation Measure using Wavelet Decomposition
abstract
An effective contrast enhancement method should not only improve the perceptual quality of an image but should also avoid adding any artifacts or affecting naturalness of images. This makes Contrast Enhancement Evaluation (CEE) a challenging task in the sense that both the improvement in image quality and unwanted side-effects need to be checked for. Currently, there is no single CEE metric that works well for all kinds of enhancement criteria. In this paper, we propose a new Multi-Criteria CEE (MCCEE) measure which combines different metrics effectively to give a single quality score. In order to fully exploit the potential of these metrics, we have further proposed to apply them on the decomposed image using wavelet transform. This new metric has been tested on two natural image contrast enhancement databases as well as on medical Computed Tomography (CT) images. The results show a substantial improvement as compared to the existing evaluation metrics. The code for the metric is available at: https://github.com/zakopz/MCCEE-Contrast-Enhancement-Metric.
Zohaib Amjad Khan, Azeddine Beghdadi, Faouzi Alaya Cheikh, Mounir Kaaniche, Muhammad Ali Qureshi
MMSP3
2019 Person Head Detection Based Deep Model for People Counting in Sports Videos
abstract
People counting in sports venues is emerging as a new domain in the field of video surveillance. People counting in these venues faces many key challenges, such as severe occlusions, few pixels per head, and significant variations in person's head sizes due to wide sport areas. We propose a deep model based method, which works as a head detector and takes into consideration the scale variations of heads in videos. Our method is based on the notion that head is the most visible part in the sports venues where large number of people are gathered. To cope with the problem of different scales, we generate scale aware head proposals based on scale map. Scale aware proposals are then fed to the Convolutional Neural Network (CNN) and it provides a response matrix containing the presence probabilities of people observed across scene scales. We then use non-maximal suppression to get the accurate head positions. For the performance evaluation, we carry out extensive experiments on two standard datasets and compare the results with state-of-the-art (SoA) methods. The results in terms of Average Precision (AvP), Average Recall (AvR), and Average F1-Score (AvF-Score) show that our method is better than SoA methods.
Sultan Daud Khan, Mohib Ullah, Nicola Conci, Faouzi Alaya Cheikh, Azeddine Beghdadi
AVSS5
2019 Just Noticeable Difference Model for Asymmetrically Distorted Stereoscopic Images
abstract
In this paper, we propose a saliency-weighted stereoscopic JND (SSJND) model constructed based on psychophysical experiments, accounting for binocular disparity and spatial masking effects of the human visual system (HVS). Specifically, a disparity-aware binocular JND model is first developed using psychophysical data, and then is employed to estimate the JND threshold for non-occluded pixel (NOP). In addition, to derive a reliable 3D-JND prediction, we determine the visibility threshold for occluded pixel (OP) by including a robust 2D-JND model. Finally, SSJND thresholds of one view are obtained by weighting the resulting JND for NOP and OP with their visual saliency. Based on subjective experiments, we demonstrate that the proposed model outperforms the other 3D-JND models in terms of perceptual quality at the same noise level.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne
ICASSP3
2019 Blind Stereopair Quality Assessment Using Statistics of Monocular and Binocular Image Structures
abstract
In this paper, we present a no-reference (NR) quality predictor for stereoscopic/3D images based on statistics aggregation of monocular and binocular local contrast features. In particular, for left and right views, we first extract statistical features of the image gradient magnitude (GM) and the Laplacian of Gaussian (LoG), describing the image local structures from different perspectives. The monocular statistical features are then combined to derive the binocular features based on a linear summation model using weightings based on LoG-response and image local-entropy, independently. These weights can effectively simulate the strength of the views dominance on binocular rivalry (BR) behavior of the human visual system. Subsequently, we further compute the GM features of the difference map between left and right views reflecting the distortion on disparity/depth information. Finally, the BR-inspired combined monocular and disparity-related binocular features associated with subjective quality scores are jointly used to construct a learned regression model relying on support vector machine regressor. Experimental results on three 3D-IQA benchmark databases demonstrate that our method achieves high quality prediction accuracy and competitive performance compared to state-of-the-art methods.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh
ICIP3
2019 An Image Based Prediction Model for Sleep Stage Identification
abstract
Human brain undergoes state changes during sleep which produces distinctive signal patterns when recorded by electroencephalography (EEG). Automatic identification of these stages is crucial to diagnosing and treating sleep related disorders. We propose an image processing based technique for automatic identification of sleep stages from EEG signals. We generate two dimensional image representations from the high dynamic range Fourier transform features of the one dimensional EEG signals. Using these representations, we learn a deep and dense convolutional neural network (CNN) model for prediction. The key advantage of the proposed method is its seamless use of the existing well studied and powerful deep CNN models designed for computer vision problems. Experiments on the popular Sleep-EDF database show that the proposed method significantly outperforms the compared methods for automatic sleep stage identification.
Saira Kanwal, Sultan Daud Khan, Mohib Ullah, Faouzi Alaya Cheikh
ICIP6
2019 Disam: Density Independent and Scale Aware Model for Crowd Counting and Localization
abstract
People counting in high density crowds is emerging as a new frontier in crowd video surveillance. Crowd counting in high density crowds encounters many challenges, such as severe occlusions, few pixels per head, and large variations in person's head sizes. In this paper, we propose a novel Density Independent and Scale Aware model (DISAM), which works as a head detector and takes into account the scale variations of heads in images. Our model is based on the intuition that head is the only visible part in high density crowds. In order to deal with different scales, unlike off-the-shelf Convolutional Neural Network (CNN) based object detectors which use general object proposals as inputs to CNN, we generate scale aware head proposals based on scale map. Scale aware proposals are then fed to the CNN and it renders a response matrix consisting of probabilities of heads. We then explore non-maximal suppression to get the accurate head positions. We conduct comprehensive experiments on two benchmark datasets and compare the performance with other state-of-theart methods. Our experiments show that the proposed DISAM outperforms the compared methods in both frame-level and pixel-level comparisons.
Sultan Daud Khan, Mohammad Uzair, Mohib Ullah, Rehanullah Khan, Faouzi Alaya Cheikh
ICIP6
2019 Efficient Enhancement of Stereo Endoscopic Images Based on Joint Wavelet Decomposition and Binocular Combination
abstract
The success of minimally invasive interventions and the remarkable technological and medical progress have made endoscopic image enhancement a very active research field. Due to the intrinsic endoscopic domain characteristics and the surgical exercise, stereo endoscopic images may suffer from different degradations which affect its quality. Therefore, in order to provide the surgeons with a better visual feedback and improve the outcomes of possible subsequent processing steps, namely, a 3-D organ reconstruction/registration, it would be interesting to improve the stereo endoscopic image quality. To this end, we propose, in this paper, two joint enhancement methods which operate in the wavelet transform domain. More precisely, by resorting to a joint wavelet decomposition, the wavelet subbands of the right and left views are simultaneously processed to exploit the binocular vision properties. While the first proposed technique combines only the approximation subbands of both views, the second method combines all the wavelet subbands yielding an inter-view processing fully adapted to the local features of the stereo endoscopic images. Experimental results, carried out on various stereo endoscopic datasets, have demonstrated the efficiency of the proposed enhancement methods in terms of perceived visual image quality.
Bilel Sdiri, Mounir Kaaniche, Faouzi Alaya Cheikh, Azeddine Beghdadi, Ole Jakob Elle
IEEE Trans. Medical Imaging3
2018 Deep Smoke Removal from Minimally Invasive Surgery Videos
abstract
During video-guided minimally invasive surgery, quality of frames may be degraded severely by cauterization-induced smoke and condensation of vapor. This degradation of quality creates discomfort for the operating surgeon, and causes serious problems for automatic follow-up processes such as registration, segmentation and tracking. This paper proposes a novel deep neural network based smoke removal solution that is able to enhance the quality of surgery video frames in real-time. It employs synthetically generated training dataset including smoke embedded and clean reference versions. Results calculated on the test set indicate that our network outperforms previous defogging methods in terms of quantitative and qualitative measures. While eliminating apparent smoke, it also successfully preserves the natural appearance of tissue surface. To the best of our knowledge, the presented method is the first deep neural network based approach for the surgical field smoke removal problem.
Sabri Bolkar, Faouzi Alaya Cheikh, Sule Yildirim Yayilgan
ICIP3
2018 No-Reference Quality Assessment of Stereoscopic Images Based on Binocular Combination of Local Features Statistics
abstract
No-reference (NR) stereoscopic 3D (S3D) image quality assessment (SIQA) is still challenging due to the poor understanding of how the human visual system (HVS) judges image quality based on binocular vision. In this paper, we propose an efficient opinion-aware NR Stereoscopic Quality predictor based on local contrast statistics combination (SQSC). Specifically, for left and right views, we first extract statistical features of the gradient magnitude (GM) and Laplacian of Gaussian (LoG) responses, describing the image local structures from different perspectives. The HVS is insensitive to low-order statistical redundancies that can be removed by LoG filtering. Hence, the monocular statistical features are then fused to derive the binocular features based on a linear combination model using LoG responses-based weightings. These weightings can efficiently simulate the binocular rivalry (BR) phenomenon. Finally, the binocular features and the subjective scores were jointly employed to construct a learned regression model obtained by the support vector regression (SVR) algorithm. Experimental results on three widely used 3D IQA databases demonstrate the high prediction performance of the proposed method when compared to recent well performing SIQA methods.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne
ICIP3
2018 Deep Feature Based End-to-End Transportation Network for Multi-Target Tracking
abstract
We propose an End-to-End Transportation Network (EETN) for multi-target tracking. In the EETN, we model the optimal set of trajectories through min-cost flow problem by exploring deep features to generate a graph. The transition cost among the nodes is found through statistical similarity metric. We consider dynamic programming to solve the optimization problem. For experimental evaluation, we compare our proposed EETN method with two state-of-the-art methods using four benchmark datasets. The quantitative analysis shows promising results of our EETN against state-of-the-art methods on precision/recall and F-score.
Mohib Ullah, Faouzi Alaya Cheikh
ICIP2
2017 Sparse Coded Handcrafted and Deep Features for Colon Capsule Video Summarization
abstract
Capsule endoscopy, which uses a wireless camera to take images of the digestive track, is emerging as an alternative to traditional wired colonoscopy. A single examination produces a sequence of approximately 50,000 frames. These sequences are manually reviewed, which is time consuming and typically takes about 45-90 minutes and requires the undivided concentration of the reviewer. In this paper, we propose a novel capsule video summarization framework using sparse coding and dictionary learning in feature space. Video frames are clustered into superframes based on power spectral density, and cluster representative frames are used for video summarization. Handcrafted and deep features that are extracted for representative frames are sparse coded using a learned dictionary. Sparse coded features are later used for training SVM classifier. The proposed method was compared with state-of-the-art methods based on sensitivity and specificity. The achieved results show that our proposed framework provides robust capsule video summarization without losing informative segments.
Ahmed Kedir Mohammed, Sule Yildirim Yayilgan, Marius Pedersen, Oistein Hovde, Faouzi Alaya Cheikh
CBMS5
2017 Stereoscopic image quality assessment based on the binocular properties of the human visual system
abstract
One of the most challenging issues in stereoscopic image quality assessment (IQA) is how to effectively model the binocular behaviors of the human visual system (HVS). The latter has a great impact on the perceptual stereoscopic 3D (S3D) quality. This paper presents a stereoscopic IQA metric based on the properties of the HVS. Instead of measuring the quality of the left and the right views separately, the proposed method predicts the quality of a cyclopean image to ensure that the overall S3D quality is as close as possible to the binocular vision. The cyclopean image is synthesized based on the local entropy of each view with the aim to simulate the phenomena of the binocular rivalry/suppression. A 2D IQA metric is employed to assess the quality of both the cyclopean image and the disparity map. Additionally, the quality of the cyclopean image is modulated according to the visual importance of each pixel defined by the just noticeable difference (JND). Finally, the 3D quality score is derived by combining the quality estimates of the cyclopean image and disparity map. Experimental results show that the proposed method outperforms many other state-of-the-art SIQA methods in terms of prediction accuracy and computational efficiency.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne
ICASSP3
2017 Full-reference stereoscopic image quality assessment accounting for binocular combination and disparity information
abstract
One of the most challenging issues in stereoscopic image quality assessment (SIQA) is how to effectively model the binocular behavior of the human visual system (HVS). The latter has a great impact on the perceptual 3D quality. In this paper, we propose a SIQA metric accounting for binocular combination properties and disparity information. Instead of computing the quality of the left and the right views separately, the proposed metric predicts the quality of a cyclopean image so as to have a good consistency with 3D human perception. The cyclopean image is synthesized based on the local entropy and the visual saliency of each view with the aim to simulate the phenomena of binocular fusion/rivalry. A 2D IQA metric is employed to assess the quality of both the cyclopean image and the disparity map. The obtained scores are used to derive the 3D quality score thanks to a pooling stage. Experimental results on three public 3D IQA databases show that the proposed method outperforms many other state-of-the-art SIQA methods, and achieves high prediction accuracy on these databases.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne
ICIP3
2017 A hierarchical feature model for multi-target tracking
abstract
We propose a novel Hierarchical Feature Model (HFM) for multi-target tracking. The traditional tracking algorithms use handcrafted features that cannot track targets accurately when the target model changes due to articulation, illumination intensity variation or perspective distortions. Our HFM explore deep features to model the appearance of targets. Then, we use an unsupervised dimensionality reduction for sparse representation of the feature vectors to cope with the time-critical nature of the tracking problem. Subsequently, a Bayesian filter is adopted as the tracker and a discrete combinatorial optimization is considered for target association. We compare our proposed HFM against 4 state-of-the-art trackers using 4 benchmark datasets. The experimental results show that our HFM outperforms all the state-of-the-art methods in terms of both Multi Object Tracking Accuracy (MOTA) and Multi Object Tracking Precision (MOTP).
Mohib Ullah, Ahmed Kedir Mohammed, Faouzi Alaya Cheikh, Zhaohui Wang 0001
ICIP3
2017 No-reference stereo image quality assessment based on joint wavelet decomposition and statistical models
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
Signal Process. Image Commun.4
2016 HoG based real-time multi-target tracking in Bayesian framework
abstract
Multi-target tracking is one of the most challenging tasks in computer vision. Several complex techniques have been proposed in the literature to tackle the problem. The main idea of such approaches is to find an optimal set of trajectories within a temporal window. The performance of such approaches are fairly good but their computational complexity is too high making them unpractical. In this paper, we propose a novel tracking-by-detection approach in a Bayesian filtering framework. The appearance of a target is modeled through HoG descriptor and the critical problem of target association is solved through combinatorial optimization. It is a simple yet very efficient approach and experimental results show that it achieves state-of-the-art performance in real time.
Mohib Ullah, Faouzi Alaya Cheikh, Ali Shariq Imran
AVSS2
2016 On the performance of 3D just noticeable difference models
abstract
The just noticeable difference (JND) notion reflects the maximum tolerable distortion. It has been extensively used for the optimization of 2D applications. For stereoscopic 3D (S3D) content, this notion is different since it relies on different mechanisms linked to our binocular vision. Unlike 2D, 3D-JND models appeared recently and the related literature is rather limited. These models can be used for the sake of compression and quality assessment improvement for S3D content. In this paper, we propose a deep and comparative study of the existing 3D-JND models. Additionally, in order to analyze their performance, the 3D-JND models have been integrated in recent metric dedicated to stereoscopic image quality assessment (SIQA). The results are reported on two widely used S3D image databases.
Yu Fan 0001, Mohamed-Chaker Larabi, Faouzi Alaya Cheikh, Christine Fernandez-Maloigne
ICIP3
2015 Automatic detection of colonoscopic anomalies using capsule endoscopy
abstract
Colon cancer and precancerous colon lesions are major health problems. The most frequent precancerous cancer lesions are polyps characterized by abnormal tissue growth in the human bowel. There are also anomalies such as inflammation and bleeding area. Colon capsule endoscopy (CCE) is a recent promising technology which enables to obtain videos of the inside of the intestine via an on board digital camera. The small capsule ingested by the patient is recuperated after it passes through the whole gastrointestinal track. Expert gastroenterologists analyze the video sequence visually in order to find out frames containing abnormalities related to cancer. This process is labor intensive and time consuming. In this paper we propose an algorithm that provides automatic video analysis in order to classify tissue regions in images into two categories: normal and abnormal. In order to achieve high correct classification rates, we first pre-process the image data by removing the noise in it and by normalizing the intensity values of the image pixels. Next, we use the well-known SIFT descriptor algorithm with Bag of Feature (BoF) approach for feature extraction. For training, Support Vector Machine (SVM) is trained on the extracted features using the training dataset. Finally a testing data set is used to assess the performance of the proposed algorithm. The early experimental results are very encouraging and show high correct classification rates, reaching up to 98.25% for images with polyps. The novelty of the proposed algorithm is the combination of using specific SIFT features with Bof and the vignetting correction to capture the local characteristics of the abnormalities in capsule video endoscopy images.
Limamou Gueye, Sule Yildirim Yayilgan, Faouzi Alaya Cheikh, Ilangko Balasingham
ICIP3
2015 Neural networks based visual attention model for surveillance videos
Fahad Fazal Elahi Guraya, Faouzi Alaya Cheikh
Neurocomputing2
2015 Efficient Inter-View Bit Allocation Methods for Stereo Image Coding
abstract
In this paper, we present efficient bit allocation methods for stereo image coding purpose. Since the common idea behind most of the existing stereo compression schemes consists of encoding a reference and residual images as well as a disparity map, we mainly focus on the bit allocation issue between the reference and residual images. Generally, this problem is solved in an empirical manner by looking for the optimal rates leading to the minimum distortion value. Thanks to recent approximations of the entropy and distortion functions, we propose accurate and fast bit allocation schemes appropriate for the open-loop- and closed-loop-based stereo coding structures. Experimental results show the benefits which can be drawn from the proposed bit allocation methods.
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
IEEE Trans. Multim.4
2014 Rate distortion optimal bit allocation for stereo image coding
abstract
Many research works have been developed for stereo image compression purpose where most of them aim at encoding a reference image, a residual one and a disparity map. While the disparity field is often losslessly encoded, we are mainly interested in this paper in the bit allocation problem between the reference and residual images. Generally, the bit allocation is expressed as an optimization problem which involves the computation of the operational rate-distortion (RD) functions for all the wavelet subbands and for different quantization steps. However, this strategy is computationally intensive. To solve this problem, we consider the uniform scalar quantization of the wavelet subbands of both images modeled by a Generalized Gaussian distribution. Thanks to recent approximations of the entropy and distortion functions, we develop an optimal and fast bit allocation method. The obtained results confirm the efficiency of the proposed bit allocation method in the context of stereo image coding.
Walid Hachicha, Mounir Kaaniche, Azeddine Beghdadi, Faouzi Alaya Cheikh
ICIP4
2014 Spatio-temporal analysis of eye fixations data in images
abstract
Computer vision algorithms such as image compression, image segmentation, context aware image resizing, image quality assessment and target detection, are designed with the aim of replicating our visual system. Human eye fixations recorded by using an eye tracker are typically used as a criterion for the optimisation of such algorithms. In this paper, we propose a method to analyze the spatio-temporal nature of fixations data for different observers. By studying the correlation matrix constructed based on the fixations data of different observers viewing the same image, it was found that 21 percent of the data can be accounted by one eigenvector. A visual inspection of this vector, shows that it represents the time sequence and locations of the objects in the image that observers deem as salient. The eigenvectors are used as a benchmark for evaluating the spatio-temporal performance of Itti's classic visual saliency model. Based on the results obtained from a comprehensive publicly available dataset, we show that the proposed method can be used as ground truth for evaluating the spatio-temporal performance of saliency models. Furthermore, it can be used to provide salient locations and their time sequence for a real-time image compression application with promising results.
Faouzi Alaya Cheikh, Jon Yngve Hardeberg
ICIP2
2013 Stereo image quality assessment using a binocular just noticeable difference model
abstract
This paper presents a novel full-reference Stereo Image Quality Assessment (SIQA) measure based on well understood characteristics of the human visual system (HVS), namely contrast sensitivity and frequency and directional selectivity. Additionally, the proposed metric takes into account the stereo interplay between the two views, where one view may affect our perception of the overall quality of the stereo image pair. Therefore, a Binocular Just Noticeable Difference (BJND) model is used to compute the distortion visibility threshold, and the binocular suppression theory is considered in the proposed metric. The scored 3D LIVE IQA database is used to evaluate the correlation of the proposed metric with the DMOS subjective score provided by the database. The obtained experimental results show that the proposed metric correlates much better with the DMOS score than the state-of-the-art metrics do.
Walid Hachicha, Azeddine Beghdadi, Faouzi Alaya Cheikh
ICIP3
2011 Blackboard content classification for lecture videos
abstract
In this paper, we propose a novel approach to understand the high level semantics of instructional video by identifying mid-level features from the lecture content. The lecture content in instructional videos can be divided into text, equations and figures. In unscripted lecture video, these visual contents can be useful visual cues to understand the high level semantics. For example, it could help us achieve efficient structuring and indexing of multimedia learning material. To understand the high level semantics from the content itself is however not a trivial task. To this end, we propose visual content classification system (VCCS) for multimedia lecture videos. We propose hybrid approach by combining support vector machine (SVM) and optical character recognition (OCR) to classify visual content into figures, text and equations. The initial results show overall classification accuracy above 85 percent.
Ali Shariq Imran, Faouzi Alaya Cheikh
ICIP2
2010 Multi-feature based visual saliency detection in surveillance video
abstract
The perception of video is different from that of image because of the motion information in video. Motion objects lead to the difference between two neighboring frames which is usually focused on. By far, most papers have contributed to image saliency but seldom to video saliency. Based on scene understanding, a new video saliency detection model with multi-features is proposed in this paper. First, background is extracted based on binary tree searching, then main features in the foreground is analyzed using a multi-scale perception model. The perception model integrates faces as a high level feature, as a supplement to other low-level features such as color, intensity and orientation. Motion saliency map is calculated using the statistic of the motion vector field. Finally, multi-feature conspicuities are merged with different weights. Compared with the gaze map from subjective experiments, the output of the multi-feature based video saliency detection model is close to gaze map.
Yubing Tong, Hubert Konik, Faouzi Alaya Cheikh, Fahad Fazal Elahi Guraya, Alain Trémeau
VCIP3
2004 Vector rational interpolation schemes for erroneous motion field estimation applied to MPEG-2 error concealment
abstract
A study on the use of vector rational interpolation for the estimation of erroneously received motion fields of MPEG-2 predictively coded frames is undertaken in this paper, aiming further at error concealment (EC). Various rational interpolation schemes have been investigated, some of which are applied to different interpolation directions. One scheme additionally uses the boundary matching error and another one attempts to locate the direction of minimal/maximal change in the local motion field neighborhood. Another one further adopts bilinear interpolation principles, whereas a last one additionally exploits available coding mode information. The methods present temporal EC methods for predictively coded frames or frames for which motion information pre-exists in the video bitstream. Their main advantages are their capability to adapt their behavior with respect to neighboring motion information, by switching from linear to nonlinear behavior, and their real-time implementation capabilities, enabling them for real-time decoding applications. They are easily embedded in the decoder model to achieve concealment along with decoding and avoid post-processing delays. Their performance proves to be satisfactory for packet error rates up to 2% and for video sequences with different content and motion characteristics and surpass that of other state-of-the-art temporal concealment methods that also attempt to estimate unavailable motion information and perform concealment afterwards.
Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas
IEEE Trans. Multim.2
2003 Relevance feedback for shape query refinement
abstract
In this paper we propose to incorporate a feedback loop, into the ordinal correlation framework and apply it to shape-based image retrieval. The user's feedback on the relevance of the retrieval results is used to tune the weights of the similarity measure. Statistics from the features of both relevant and irrelevant items are used to estimate the weights. Moreover, the information accumulated from previous retrieval iterations is used in the weights estimation. A simple measure of the discrimination power is proposed and used to show that the relevance feedback increases the capability of the ordinal correlation scheme to discriminate between relevant and irrelevant objects.
Faouzi Alaya Cheikh, Bogdan Cramariuc, Moncef Gabbouj
ICIP (1)1
2000 A Weighted Distance Approach to Relevance Feedback
abstract
Content-based image retrieval systems use low-level features like color and texture for image representation. Given these representations as feature vectors, similarity between images is measured by computing distances in the feature space. Unfortunately, these low-level features cannot always capture the high-level concept of similarity in human perception. Relevance feedback tries to improve the performance by allowing iterative retrievals where the feedback information from the user is incorporated into the database search. We present a weighted distance approach where the weights are the ratios of standard deviations of the feature values both for the whole database and also among the images selected as relevant by the user. The feedback is used for both independent and incremental updating of the weights and these weights are used to iteratively refine the effects of different features in the database search. Retrieval performance is evaluated using average precision and progress that are computed on a database of approximately 10,000 images and an average performance improvement of 19% is obtained after the first iteration.
Selim Aksoy, Robert M. Haralick, Faouzi Alaya Cheikh, Moncef Gabbouj
ICPR3
2000 Directional-rational approach for color image enhancement
abstract
In this paper, we present an unsharp masking-based approach for noise smoothing and edge enhancing in multichannel images. The proposed structure is similar to the conventional unsharp masking structure, however, the enhancement is allowed only in the direction of maximal change and the enhancement parameter is computed as a nonlinear function of the rate of change. The proposed scheme enhances the true details, limits the overshoot near sharp edges and attenuates noise in flat areas. Moreover the use of the control function eliminates the need for the subjective coefficient /spl lambda/ used in the conventional unsharp masking technique. Simulations results show that the processed image presents sharp edges which makes it more pleasant to the human eye. Moreover, the amount of noise in the image is clearly reduced.
Faouzi Alaya Cheikh, Moncef Gabbouj
ISCAS1
1999 Motion field estimation by vector rational interpolation for error concealment purposes
abstract
A study on the use of vector rational interpolation for the estimation of erroneously received motion fields of an MPEG-2 coded video bitstream has been performed. Four different motion vector interpolation schemes have been examined using motion information from available top and bottom adjacent blocks since left or right neighbours are usually lost. The presented interpolation schemes are capable of adapting their behaviour according to neighbouring motion information. Simulation results prove the satisfactory performance of the novel nonlinear interpolation schemes and the success of their application to the concealment of predictively coded frames. The motion vector rational interpolation concealment method proves to be a fast method, thus adequate for real-time applications.
Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas
ICASSP2
1996 Impulse noise removal in highly corrupted color images
abstract
We present a novel and efficient technique for the restoration of color images which are highly corrupted with impulse noise. This is a detection-estimation based approach in which outliers are first detected using a Teager-like operator followed by a locally adaptive threshold. Center pixels whose "energy" exceeds some threshold are replaced with the local marginal median. Simulation results show the superior performance of the proposed filtering algorithm compared to the renowned vector median (VM) and generalized vector directional filter (GVDF), which are commonly used for color image restoration. Monte Carlo simulations show the edge preservation and impulse noise attenuation capabilities of the proposed technique. The efficiency of the algorithm stems from its simple arithmetic operations compared with more demanding ones, e.g. computation of distances and angles in the case of VMF and GVDF, respectively.
Faouzi Alaya Cheikh, Ridha Hamila, Moncef Gabbouj, Jaakko Astola
ICIP (1)1