VLDB 2026 Research / reviewers in the wild / expert
Mohamed Amine Kerkouri
dblp:296/3848
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-7479-6879ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closing the Foveal Gap: Perceptually Grounded Scanpath Comparison with Disc IoUabstractScanpath comparison metrics universally represent fixations as dimensionless (x, y) points, despite the human fovea subtending approximately 1° ∼ 2° of visual angle. This mismatch penalises spatially proximate fixations that process identical visual content. We propose FDISS (Foveal Disc IoU Scanpath Score), a metric that models each fixation as a foveal disc and computes similarity via bidirectional nearest-neighbour IoU matching, yielding interpretable Precision, Recall, and F1 scores. FDISS introduces a biologically grounded perceptual tolerance zone absent from all existing metrics and requires no free parameters at runtime. Mohamed Amine Kerkouri, Marouane Tliba, Abdellah Zakaria Sellam, Cosimo Distante, Alessandro Bruno, Aladine Chetouani |
ETRA | 1 |
| 2026 | What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metricsabstractScanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community. Mohamed Amine Kerkouri, Marouane Tliba, Bin Wang 0068, Aladine Chetouani, Ulas Bagci, Alessandro Bruno |
ETRA | 1 |
| 2025 | Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images
David C. Wong 0005, Bin Wang 0068, Gorkem Durak, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, A. Enis Çetin, Cagdas Topel, Nicolo Gennaro, Camila Lopes Vendrami, Tugce Agirlar Trabzonlu, Amir Ali Rahsepar, Laetitia Perronne, Matthew Antalek, Onural Ozturk, Gokcan Okur, Andrew C. Gordon, Ayis Pyrros, Frank H. Miller, Amir Borhani, Hatice Savas, Eric M. Hart, Elizabeth A. Krupinski, Ulas Bagci |
ETRA | 5 |
| 2024 | AVAtt : Art Visual Attention dataset for diverse painting stylesabstractPreserving cultural heritage is paramount for societal and historical identity. Paintings, spanning ancient to modern eras, are pivotal subjects under constant scrutiny. As art reflects human creativity, studying human visual behavior toward paintings becomes increasingly vital. Thus, we introduce the AVAtt dataset, providing eye movement data for a diverse collection of painting styles across various ages and geographical origins. This dataset aims to facilitate the development and evaluation of computational saliency and scanpath prediction methods in the unique domain of painting (Dataset available at Github). Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
ETRA | 1 |
| 2024 | Perceptual Evaluation of Masked AutoEncoder Emergent Properties Through Eye-Tracking-Based PolicyabstractThe advancement of image restoration, especially in reconstructing missing or damaged image areas, has benefited significantly from self-supervised learning techniques, notably through the recent Masked Auto Encoder (MAE) strategy. In this project, we leverage eye-tracking data to enhance image reconstruction quality, and more specifically, with fixation-based saliency combined with the MAE strategy. By examining the emergent properties of representation learning and drawing parallels to human perceptual observation, we focus on how eye-tracking data informs the selection of image patches for reconstruction, aligning computational methods with human visual perception. Our findings reveal the potential of integrating eye-tracking insights to improve the accuracy and perceptual relevance of self-supervised learning models in computer vision. This study thus underscores the synergy between computational image restoration methods and human perception, facilitated by eye-tracking technology, opening new directions and insights for both fields. Our experiments are available for reproducibility in this GitHub Repository. Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Alessandro Bruno, Mohammed El Hassouni, Arzu Çöltekin |
ETRA | 2 |
| 2023 | Automatic diagnosis of knee osteoarthritis severity using Swin transformerabstractKnee osteoarthritis (KOA) is a widespread condition that can cause chronic pain and stiffness in the knee joint. Early detection and diagnosis are crucial for successful clinical intervention and management to prevent severe complications, such as loss of mobility. In this paper, we propose an automated approach that employs the Swin Transformer to predict the severity of KOA. Our model uses publicly available radiographic datasets with Kellgren and Lawrence scores to enable early detection and severity assessment. To improve the accuracy of our model, we employ a multi-prediction head architecture that utilizes multi-layer perceptron classifiers. Additionally, we introduce a novel training approach that reduces the data drift between multiple datasets to ensure the generalization ability of the model. The results of our experiments demonstrate the effectiveness and feasibility of our approach in predicting KOA severity accurately. Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani, Marouane Tliba, Yassine Nasser, Rachid Jennane, Alessandro Bruno |
CBMI | 2 |
| 2023 | Detecting colour vision deficiencies via Webcam-based Eye-tracking: A case studyabstractWebcam-based eye-tracking platforms have recently re-emerged due to improvements in machine learning-supported calibration processes and offer a scalable option for conducting eye movement studies. Although not yet comparable to the infrared-based ones regarding accuracy and frequency, some compelling performances have been observed, especially in those scenarios with medium-sized AOI (Areas of Interest) in images. In this study, we test the reliability of webcam-based eye-tracking on a specific task: Eye movement distribution analysis for CVD (Colour Vision Deficiency) detection. We introduce a new publicly available eye movement dataset based on a pilot study (n=12) on images with dominant red colour (previously shown to be difficult with dichromatic AOI to investigate CVD by comparing attention patterns obtained in webcam eye-tracking sessions). We hypothesized that webcam eye tracking without infrared support could detect differing attention patterns between CVD and non-CVD participants and observed statistically significant differences, allowing the retention of our hypothesis. Alessandro Bruno, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Carlo Calogero Giunta, Arzu Çöltekin |
ETRA | 3 |
| 2023 | An Inter-Observer Consistent Deep Adversarial Training for Visual Scanpath PredictionabstractThe visual scanpath represents the fundamental concept upon which visual attention research is based. As a result, the ability to predict them has emerged as a crucial task in recent years. It is represented as a sequence of points through which the human gaze moves while exploring a scene. In this paper, we propose an inter-observer consistent adversarial training approach for scanpath prediction through a lightweight deep neural network. The proposed method employs a discriminative neural network as a dynamic loss that better models the natural stochastic phenomenon while maintaining consistency between the distributions related to the subjective nature of scanpaths traversed by different observers. The competitiveness of our approach against state-of-the-art methods is shown through a testing phase. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
ICIP | 1 |
| 2022 | A domain adaptive deep learning solution for scanpath prediction of paintingsabstractCultural heritage understanding and preservation is an important issue for society as it represents a fundamental aspect of its identity. Paintings represent a significant part of cultural heritage, and are the subject of study continuously. However, the way viewers perceive paintings is strictly related to the so-called HVS (Human Vision System) behaviour. This paper focuses on the eye-movement analysis of viewers during the visual experience of a certain number of paintings. In further details, we introduce a new approach to predicting human visual attention, which impacts several cognitive functions for humans, including the fundamental understanding of a scene, and then extend it to painting images. The proposed new architecture ingests images and returns scanpaths, a sequence of points featuring a high likelihood of catching viewers’ attention. We use an FCNN (Fully Convolutional Neural Network), in which we exploit a differentiable channel-wise selection and Soft-Argmax modules. We also incorporate learnable Gaussian distributions onto the network bottleneck to simulate visual attention process bias in natural scene images. Furthermore, to reduce the effect of shifts between different domains (i.e. natural images, painting), we urge the model to learn unsupervised general features from other domains using a gradient reversal classifier. The results obtained by our model outperform existing state-of-the-art ones in terms of accuracy and efficiency. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Alessandro Bruno |
CBMI | 1 |
| 2022 | Deep-Based Quality Assessment of Medical Images Through Domain AdaptationabstractPredicting the quality of multimedia content is often needed in different fields. In some applications, quality metrics are crucial with a high impact, and can affect decision making such as diagnosis from medical multimedia. In this paper, we focus on such applications by proposing an efficient and shallow model for predicting the quality of medical images without reference from a small amount of annotated data. Our model is based on convolution self-attention that aims to model complex representation from relevant local characteristics of images, which itself slide over the image to interpolate the global quality score. We also apply domain adaptation learning in unsupervised and semi-supervised manner. The proposed model is evaluated through a dataset composed of several images and their corresponding subjective scores. The obtained results showed the efficiency of the proposed method, but also, the relevance of the applying domain adaptation to generalize over different multimedia domains regarding the downstream task of perceptual quality prediction.1 Marouane Tliba, Aymen Sekhri, Mohamed Amine Kerkouri, Aladine Chetouani |
ICIP | 3 |
| 2021 | Salypath: A Deep-Based Architecture For Visual Attention PredictionabstractHuman vision is naturally more attracted by some regions within their field of view than others. This intrinsic selectivity mechanism, so-called visual attention, is influenced by both high- and low-level factors; such as the global environment (illumination, background texture, etc.), stimulus characteristics (color, intensity, orientation, etc.), and some prior visual information. Visual attention is useful for many computer vision applications such as image compression, recognition, and captioning. In this paper, we propose an end-to-end deep-based method, so-called SALYPATH (SALiencY and scanPATH), that efficiently predicts the scanpath of an image through features of a saliency model. The idea is predict the scanpath by exploiting the capacity of a deep-based model to predict the saliency. The proposed method was evaluated through 2 well-known datasets. The results obtained showed the relevance of the proposed framework comparing to state-of-the-art models. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Rachid Harba |
ICIP | 1 |