Marouane Tliba

dblp:279/3157 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-1696-7764ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Closing the Foveal Gap: Perceptually Grounded Scanpath Comparison with Disc IoU
abstract
Scanpath comparison metrics universally represent fixations as dimensionless (x, y) points, despite the human fovea subtending approximately 1° ∼ 2° of visual angle. This mismatch penalises spatially proximate fixations that process identical visual content. We propose FDISS (Foveal Disc IoU Scanpath Score), a metric that models each fixation as a foveal disc and computes similarity via bidirectional nearest-neighbour IoU matching, yielding interpretable Precision, Recall, and F1 scores. FDISS introduces a biologically grounded perceptual tolerance zone absent from all existing metrics and requires no free parameters at runtime.
Mohamed Amine Kerkouri, Marouane Tliba, Abdellah Zakaria Sellam, Cosimo Distante, Alessandro Bruno, Aladine Chetouani
ETRA2
2026 What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metrics
abstract
Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community.
Mohamed Amine Kerkouri, Marouane Tliba, Bin Wang 0068, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
ETRA2
2026 GazeVaLM: A Multi-Observer Eye-Tracking Benchmark for Evaluating Clinical Realism in AI-Generated X-Rays
abstract
We introduce GazeVaLM, a public eye-tracking dataset for studying clinical perception during chest radiograph authenticity assessment. The dataset comprises 960 gaze recordings from 16 expert radiologists interpreting 30 real and 30 synthetic chest X-rays (generated by diffusion based generative AI) under two conditions: diagnostic assessment and real-fake classification (Visual Turing test). For each image–observer pair, we provide raw gaze samples, fixation maps, scanpaths, saliency density maps, structured diagnostic labels, and authenticity judgments. We extend the protocol to 6 state-of-the-art multimodal LLMs, releasing their predicted diagnoses, authenticity labels, and confidence scores under matched conditions — enabling direct human–AI comparison at both decision and uncertainty levels. We further provide analyses of gaze agreement, inter-observer consistency, and benchmarking of radiologists versus LLMs in diagnostic accuracy and authenticity detection. GazeVaLM supports research in gaze modeling, clinical decision-making, human–AI comparison, generative image realism assessment, and uncertainty quantification. By jointly releasing visual attention data, clinical labels, and model predictions, we aim to facilitate reproducible research on how experts and AI systems perceive, interpret, and evaluate medical images. The dataset is available at https://huggingface.co/datasets/davidcwong/GazeVaLM.
David C. Wong 0005, Zeynep Isik, Bin Wang 0068, Marouane Tliba, Gorkem Durak, Elif Keles, Halil Ertugrul Aktas, Aladine Chetouani, Cagdas Topel, Nicolo Gennaro, Camila Lopes Vendrami, Tugce Agirlar Trabzonlu, Amir Ali Rahsepar, Laetitia Perronne, Matthew Antalek, Onural Ozturk, Gokcan Okur, Andrew C. Gordon, Ayis Pyrros, Frank H. Miller, Amir Borhani, Hatice Savas, Eric M. Hart, Elizabeth A. Krupinski, Ulas Bagci
ETRA4
2025 Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images
David C. Wong 0005, Bin Wang 0068, Gorkem Durak, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, A. Enis Çetin, Cagdas Topel, Nicolo Gennaro, Camila Lopes Vendrami, Tugce Agirlar Trabzonlu, Amir Ali Rahsepar, Laetitia Perronne, Matthew Antalek, Onural Ozturk, Gokcan Okur, Andrew C. Gordon, Ayis Pyrros, Frank H. Miller, Amir Borhani, Hatice Savas, Eric M. Hart, Elizabeth A. Krupinski, Ulas Bagci
ETRA4
2025 A Benchmark Dataset for Automated Diagnosis and Treatment Planning of Class III Malocclusion Using X-Rays and Profile Photos
abstract
In this paper, we introduce a benchmark dataset for the diagnosis and treatment planning of Class III malocclusion, a condition that requires precise evaluation to determine the necessity of surgical intervention. Our dataset comprises paired lateral cephalometric X-rays and profile photographs, each annotated with a treatment plan mainly indicating whether surgery is required or not. We assess state-of-the-art deep learning models on both imaging modalities to explore their potential for automating diagnosis and treatment decisions. Notably, the dataset facilitates research into non-radiographic diagnostic approaches, potentially enabling treatment planning based solely on profile photographs. We release this dataset as a resource to advance automated applications in orthodontics and maxillofacial surgery.
Omid Halimi Milani, Emadeldeen Hamdan, Marouane Tliba, Samim Taraji, Veerasathpurush Allareddy, Aladine Chetouani, Rachid Jennane, A. Enis Çetin, Mohammed H. Elnagar
ICIP3
2025 Gradient Attention Map Based Verification of Deep Convolutional Neural Networks with Application to X-ray Image Datasets
abstract
Deep learning models have great potential in medical imaging, including orthodontics and skeletal maturity assessment. However, using a model on data different from its training set can lead to unreliable predictions that may impact patient care. To address this, we introduce a Gradient Attention Map (GAM)-based framework that evaluates a model’s suitability for new data by examining its attention patterns. Using Grad-CAM, we generate attention maps and compare them with metrics such as IoU, Dice Similarity, SSIM, Cosine Similarity, Pearson Correlation, KL Divergence, and Wasserstein Distance. A Random Forest classifier then distinguishes between models that are well-suited and those that are misapplied. Experimental results show that our method effectively filters out unsuitable models, promoting safer and more reliable use of deep learning in medical imaging.
Omid Halimi Milani, Amanda Nikho, Lauren Mills, Marouane Tliba, A. Enis Çetin, Mohammed H. Elnagar
VTS4
2024 AVAtt : Art Visual Attention dataset for diverse painting styles
abstract
Preserving cultural heritage is paramount for societal and historical identity. Paintings, spanning ancient to modern eras, are pivotal subjects under constant scrutiny. As art reflects human creativity, studying human visual behavior toward paintings becomes increasingly vital. Thus, we introduce the AVAtt dataset, providing eye movement data for a diverse collection of painting styles across various ages and geographical origins. This dataset aims to facilitate the development and evaluation of computational saliency and scanpath prediction methods in the unique domain of painting (Dataset available at Github).
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
ETRA2
2024 Perceptual Evaluation of Masked AutoEncoder Emergent Properties Through Eye-Tracking-Based Policy
abstract
The advancement of image restoration, especially in reconstructing missing or damaged image areas, has benefited significantly from self-supervised learning techniques, notably through the recent Masked Auto Encoder (MAE) strategy. In this project, we leverage eye-tracking data to enhance image reconstruction quality, and more specifically, with fixation-based saliency combined with the MAE strategy. By examining the emergent properties of representation learning and drawing parallels to human perceptual observation, we focus on how eye-tracking data informs the selection of image patches for reconstruction, aligning computational methods with human visual perception. Our findings reveal the potential of integrating eye-tracking insights to improve the accuracy and perceptual relevance of self-supervised learning models in computer vision. This study thus underscores the synergy between computational image restoration methods and human perception, facilitated by eye-tracking technology, opening new directions and insights for both fields. Our experiments are available for reproducibility in this GitHub Repository.
Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Alessandro Bruno, Mohammed El Hassouni, Arzu Çöltekin
ETRA1
2024 Balancing Representation Abstractions and Local Details Preservation for 3d Point Cloud Quality Assessment
abstract
3D Point Clouds (PCs) have become a valuable tool for representing intricate 3D information. Assessing the quality of PCs remains a challenging task, especially when striving for optimal immersive experiences. This paper introduces a novel metric and training approach that leverages projection-based views to evaluate the quality of 3D content. Our approach addresses a critical issue related to the intrinsic bias of deep networks for image recognition towards building hierarchical representations including only the global semantic, at the expense of local details. This bias is a limiting factor in tasks like 3D point cloud quality assessment where instances of the same content with varying degrees and types of degradation can possess strikingly similar representations. We propose a novel point cloud quality metric using a dual supervised and unsupervised training strategy to balance semantic understanding and preservation of critical perceptual quality-relevant information. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics on two standard 3D PCs quality assessment benchmarks (3D PCQA).
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICASSP1
2024 Enhancing Immersive Experiences through 3D Point Cloud Analysis: A Novel Framework for Applying 2D Visual Saliency Models to 3D Point Clouds
abstract
In the new area of immersive multimedia environments, understanding and manipulating visual attention are crucial for enhancing user experience. This study introduces an innovative framework that extends traditional 2D saliency maps to the analysis of 3D point clouds, a step forward in adapting saliency prediction to more complex and immersive environments. Our framework centers on the orthographic projection of 3D point clouds onto 2D planes, enabling the application of established 2D saliency models to this novel context. We further delve into the evaluation of these models on a 3D point cloud eye-tracking dataset, exploring various projection settings and thresholding techniques to maintain the integrity of saliency information in the transition from 2D to 3D. This research not only bridges a gap in applying visual attention models to 3D data but also offers insights into the optimization of quality of experience in immersive multimedia systems.
Marouane Tliba, Xuemei Zhou, Irene Viola 0001, Pablo César, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
QoMEX1
2024 Special issue: Multimedia data analysis for smart city environment safety
Alessandro Bruno, Aladine Chetouani, Zoheir A. Sabeur, Marouane Tliba, Evangelos Maltezos, Miguel Gonzalez San Emeterio
Multim. Tools Appl.4
2023 Automatic diagnosis of knee osteoarthritis severity using Swin transformer
abstract
Knee osteoarthritis (KOA) is a widespread condition that can cause chronic pain and stiffness in the knee joint. Early detection and diagnosis are crucial for successful clinical intervention and management to prevent severe complications, such as loss of mobility. In this paper, we propose an automated approach that employs the Swin Transformer to predict the severity of KOA. Our model uses publicly available radiographic datasets with Kellgren and Lawrence scores to enable early detection and severity assessment. To improve the accuracy of our model, we employ a multi-prediction head architecture that utilizes multi-layer perceptron classifiers. Additionally, we introduce a novel training approach that reduces the data drift between multiple datasets to ensure the generalization ability of the model. The results of our experiments demonstrate the effectiveness and feasibility of our approach in predicting KOA severity accurately.
Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, Alessandro Bruno
CBMI4
2023 Detecting colour vision deficiencies via Webcam-based Eye-tracking: A case study
abstract
Webcam-based eye-tracking platforms have recently re-emerged due to improvements in machine learning-supported calibration processes and offer a scalable option for conducting eye movement studies. Although not yet comparable to the infrared-based ones regarding accuracy and frequency, some compelling performances have been observed, especially in those scenarios with medium-sized AOI (Areas of Interest) in images. In this study, we test the reliability of webcam-based eye-tracking on a specific task: Eye movement distribution analysis for CVD (Colour Vision Deficiency) detection. We introduce a new publicly available eye movement dataset based on a pilot study (n=12) on images with dominant red colour (previously shown to be difficult with dichromatic AOI to investigate CVD by comparing attention patterns obtained in webcam eye-tracking sessions). We hypothesized that webcam eye tracking without infrared support could detect differing attention patterns between CVD and non-CVD participants and observed statistically significant differences, allowing the retention of our hypothesis.
Alessandro Bruno, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Carlo Calogero Giunta, Arzu Çöltekin
ETRA2
2023 PCQA-Graphpoint: Efficient Deep-Based Graph Metric for Point Cloud Quality Assessment
abstract
Following the advent of immersive technologies and the increasing interest in representing interactive geometrical format, 3D Point Clouds (PC) have emerged as a promising solution and effective means to display 3D visual information. In addition to other challenges in immersive applications, objective and subjective quality assessments of compressed 3D content remain open problems and an area of research interest. Yet most of the efforts in the research area ignore the local geometrical structures between points representation. In this paper, we overcome this limitation by introducing a novel and efficient objective metric for Point Clouds Quality Assessment, by learning local intrinsic dependencies using Graph Neural Network (GNN). To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of our solution compared to state-of-the-art metrics.
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICASSP1
2023 An Inter-Observer Consistent Deep Adversarial Training for Visual Scanpath Prediction
abstract
The visual scanpath represents the fundamental concept upon which visual attention research is based. As a result, the ability to predict them has emerged as a crucial task in recent years. It is represented as a sequence of points through which the human gaze moves while exploring a scene. In this paper, we propose an inter-observer consistent adversarial training approach for scanpath prediction through a lightweight deep neural network. The proposed method employs a discriminative neural network as a dynamic loss that better models the natural stochastic phenomenon while maintaining consistency between the distributions related to the subjective nature of scanpaths traversed by different observers. The competitiveness of our approach against state-of-the-art methods is shown through a testing phase.
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
ICIP2
2022 A domain adaptive deep learning solution for scanpath prediction of paintings
abstract
Cultural heritage understanding and preservation is an important issue for society as it represents a fundamental aspect of its identity. Paintings represent a significant part of cultural heritage, and are the subject of study continuously. However, the way viewers perceive paintings is strictly related to the so-called HVS (Human Vision System) behaviour. This paper focuses on the eye-movement analysis of viewers during the visual experience of a certain number of paintings. In further details, we introduce a new approach to predicting human visual attention, which impacts several cognitive functions for humans, including the fundamental understanding of a scene, and then extend it to painting images. The proposed new architecture ingests images and returns scanpaths, a sequence of points featuring a high likelihood of catching viewers’ attention. We use an FCNN (Fully Convolutional Neural Network), in which we exploit a differentiable channel-wise selection and Soft-Argmax modules. We also incorporate learnable Gaussian distributions onto the network bottleneck to simulate visual attention process bias in natural scene images. Furthermore, to reduce the effect of shifts between different domains (i.e. natural images, painting), we urge the model to learn unsupervised general features from other domains using a gradient reversal classifier. The results obtained by our model outperform existing state-of-the-art ones in terms of accuracy and efficiency.
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno
CBMI2
2022 Representation Learning Optimization for 3D Point Cloud Quality Assessment Without Reference
abstract
Recent information and communication systems have employed 3D Point Cloud (PC) as an advanced geometrical representation modality for immersive applications. Like most multimedia data, PCs are often compressed for transmission and viewing purposes, which can impact the perceived quality. Developing robust and efficient objective quality metrics for PCs is still an open problem. In this paper, we propose an end-to-end deep approach for evaluating the perceptual effects of point cloud compression solutions without reference. Our approach focuses on leveraging the intrinsic point cloud characteristics to quantify the coding impairments from few distant randomly selected patches using supervised and unsupervised training strategies. To evaluate the performance of our method, two well-known datasets have been used. The results demonstrate the effectiveness and reliability of the proposed method compared to to state-of-the-art methods.
Marouane Tliba, Aladine Chetouani, Giuseppe Valenzise, Frédéric Dufaux
ICIP1
2022 Deep-Based Quality Assessment of Medical Images Through Domain Adaptation
abstract
Predicting the quality of multimedia content is often needed in different fields. In some applications, quality metrics are crucial with a high impact, and can affect decision making such as diagnosis from medical multimedia. In this paper, we focus on such applications by proposing an efficient and shallow model for predicting the quality of medical images without reference from a small amount of annotated data. Our model is based on convolution self-attention that aims to model complex representation from relevant local characteristics of images, which itself slide over the image to interpolate the global quality score. We also apply domain adaptation learning in unsupervised and semi-supervised manner. The proposed model is evaluated through a dataset composed of several images and their corresponding subjective scores. The obtained results showed the efficiency of the proposed method, but also, the relevance of the applying domain adaptation to generalize over different multimedia domains regarding the downstream task of perceptual quality prediction.1
Marouane Tliba, Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani
ICIP1
2021 Salypath: A Deep-Based Architecture For Visual Attention Prediction
abstract
Human vision is naturally more attracted by some regions within their field of view than others. This intrinsic selectivity mechanism, so-called visual attention, is influenced by both high- and low-level factors; such as the global environment (illumination, background texture, etc.), stimulus characteristics (color, intensity, orientation, etc.), and some prior visual information. Visual attention is useful for many computer vision applications such as image compression, recognition, and captioning. In this paper, we propose an end-to-end deep-based method, so-called SALYPATH (SALiencY and scanPATH), that efficiently predicts the scanpath of an image through features of a saliency model. The idea is predict the scanpath by exploiting the capacity of a deep-based model to predict the saliency. The proposed method was evaluated through 2 well-known datasets. The results obtained showed the relevance of the proposed framework comparing to state-of-the-art models.
Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Rachid Harba
ICIP2