Seyed Ali Amirshahi

dblp:119/1545 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-0620-1535ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Few-Shot Supervised Contrastive Learning for Image/Video Distortion Classification
Riestiya Zain Fadillah, Seyed Ali Amirshahi, Marius Pedersen, Azeddine Beghdadi
ICPR (7)2
2026 Audio-driven visual attention for lightweight audio-visual quality assessment
abstract
Multimedia has become an integral component of the Internet and modern digital experiences, playing a vital role in our everyday lives. The majority of content that we consume daily combines both audio and video modalities. This emphasizes the demand for an effective audio-visual quality assessment. However, current state-of-the-art audio-visual quality assessment methods often require a high computational cost and execution time. Moreover, they ignore important factors, such as the signal’s context, that influence what users expect in terms of quality across different circumstances. In this study, we therefore introduce an approach that addresses these problems. Inspired by the complementary nature of audio and vision, our methods leverage audio cues to selectively amplify visually informative regions, allowing the model to capture the cross-modal contextual dependencies in a computationally efficient manner. Experimental results show that our proposed model significantly reduces the computational complexity by more than 80% of both parameter size and runtime while still providing a favorable performance compared to state-of-the-art methods. The source code of this work will be made publicly available.
Ha Thu Nguyen, Seyed Ali Amirshahi, Katrien De Moor, Mohamed-Chaker Larabi
Neurocomputing2
2026 Predicting image quality score distribution using deep neural networks
abstract
Blind objective image quality assessment methods typically predict a single Mean Opinion Score (MOS) to represent perceived image quality. However, MOS discards valuable information about observer variability, as different score distributions can produce the same mean value. Predicting the full quality score distribution provides a richer and more realistic representation of perceptual quality. In this work, we propose an Image Quality Distribution Network (IQDN) in five configurations: a custom convolutional neural network trained from scratch and four transfer learning variants based on ResNet50, VGG16, Xception, and DenseNet121 backbones. The models were trained to predict five-bin normalized quality score histograms on three datasets: KonIQ-10k, CID2013, and a newly introduced NAP540 dataset. Performance was evaluated using multiple point-wise error and distribution similarity metrics. Results show that the Image Quality Distribution Network with Xception backbone consistently performs better than existing state-of-the-art approaches in terms of Earth Mover’s Distance, capturing score distributions more effectively. Furthermore, a hybrid loss function combining Kullback–Leibler (KL) divergence and Huber loss improved performance across datasets. The predicted distributions also reconstruct MOS that closely correlate with ground-truth MOS. Finally, incorporating a quality-aware backbone demonstrates strong potential, particularly in cross-dataset testing scenarios. • Introduce the Image Quality Distribution Network (IQDN), a blind image quality assessment framework designed to predict full quality score distributions rather than single Mean Opinion Score (MOS) values. • Present a systematic evaluation of IQDN architectures on multiple image quality datasets, using a diverse set of distributional and point-wise error metrics. • Analyze the impact of different architectural configurations on the prediction performance, as well as cross-dataset generalization. • The proposed networks showed very good performance, both in predicting the distribution shape and in the Mean Opinion Score (MOS).
Nikola Plavac, Seyed Ali Amirshahi, Marius Pedersen, Sophie Triantaphillidou
J. Vis. Commun. Image Represent.2
2026 TransformAR: A light-weight transformer-based metric for Augmented Reality quality assessment
abstract
As Augmented Reality (AR) technology continues to gain traction in various sectors, ensuring a superior user experience has become an essential challenge for both academic researchers and industry professionals. However, the task of automatically predicting the quality of AR images remains difficult due to several inherent challenges, particularly the issue of visual confusion arising from the overlap of virtual and real-world elements. This paper introduces transformAR, a novel and efficient transformer-based framework designed to objectively assess the quality of AR images. The proposed model uses pre-trained vision transformers to capture content features from AR images, calculates distance vectors to measure the impact of distortions, and employs cross-attention-based decoders to effectively model the perceptual qualities of the AR images. Additionally, the training framework uses regularization techniques and label smoothing-like method to reduce the risk of overfitting. Through comprehensive experiments, we demonstrate that transformAR outperforms existing state-of-the-art approaches, offering a more reliable and scalable solution for AR image quality assessment.
Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi
Signal Process. Image Commun.3
2026 Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation
abstract
Augmented Reality (AR) is a major immersive media technology that enriches our perception of reality by overlaying digital content (the foreground) onto physical environments (the background). It has far-reaching applications, from entertainment and gaming to education, healthcare, and industrial training. Nevertheless, challenges such as visual confusion and classical distortions can result in user discomfort when using the technology. Evaluating AR quality of experience becomes essential to measure user satisfaction and engagement, facilitating the refinement necessary for creating immersive and robust experiences. Though the scarcity of data and the distinctive characteristics of AR technology render the development of effective quality assessment metrics challenging. This paper presents a deep learning-based objective metric designed specifically for assessing image quality for AR scenarios. The approach entails four key steps, (1) fine-tuning a self-supervised pre-trained vision transformer to extract prominent features from reference images and distilling this knowledge to improve representations of distorted images, (2) quantifying distortions by computing shift representations, (3) employing cross-attention-based decoders to capture perceptual quality features, and (4) integrating regularization techniques and label smoothing to address the overfitting problem. To validate the proposed approach, we conduct extensive experiments on the ARIQA dataset. The results showcase the superior performance of our proposed approach across all model variants, namely TransformAR, TransformAR-KD, and TransformAR-KD+ in comparison to existing state-of-the-art methods.
Aymen Sekhri, Seyed Ali Amirshahi, Mohamed-Chaker Larabi
IEEE Trans. Circuits Syst. Video Technol.2
2025 Uncertainty Quantification in Video Distortion Classification Under Dataset Shift
Riestiya Zain Fadillah, Seyed Ali Amirshahi, Marius Pedersen, Azeddine Beghdadi
ICANN (2)2
2025 Lightweight Image Quality Prediction Guided by Perceptual Ranking Feedback
abstract
Automatic Image Quality Assessment (IQA) remains a difficult challenge due to the complexity of mimicking the Human Visual System (HVS) and the limitations of traditional objective Image Quality Metrics (IQM). Existing learnable methods often involve high computational costs and fail to adequately capture the nuanced perceptual characteristics of the HVS, including the human ability to rank image quality and human sensitivity to differences in areas with high-frequency. In this study, we propose an effective approach that addresses these challenges by incorporating the characteristics of HVS and the perceptual classification into a lightweight IQM framework based on the transformer architecture. This allows our method to capture long-range dependencies effectively. Our approach leverages Objective Error Maps (OEMs) to enhance sensitivity to visual errors and employs a ranking module as an objective function, providing feedback on the perceptual quality at the feature level. Experimental results demonstrate that our approach not only achieves competitive performance compared to state-of-the-art IQMs but also significantly reduces computational complexity.
Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi
ICASSP3
2025 Energy Efficiency of Video Quality Assessment Metrics
abstract
Video Quality Assessment (VQA) metrics play a crucial role in modern video processing systems, yet their computational costs and environmental impact have received limited attention. This paper presents a comprehensive analysis of state-of-the-art VQA metrics across two standard datasets, including processing time, memory usage, and the ability to predict subjective scores. Results demonstrate that full-reference video metrics achieve superior accuracy compared to image-based metrics such as SSIM, but require up to more than 20 times more energy on average. No-reference metrics offer even higher accuracy but substantially larger energy usage. This difference in computational requirements has significant implications for energy consumption and environmental impact, particularly in large-scale deployments or high-resolution video processing scenarios.
Steven Le Moan, Ha Thu Nguyen, Seyed Ali Amirshahi
ICIP3
2025 ARaBIQA: A Novel Blind Image Quality Assessment Model for Augmented Reality
abstract
Ensuring the quality of Augmented Reality (AR) experiences is crucial for achieving user satisfaction in many applications such as navigation, education, and healthcare. However, automatic AR quality assessment is challenging due to limited data and the lack of a reference image notion in real-world scenarios. Hence, blind quality assessment appears to be the only plausible solution. Existing blind IQA metrics often struggle to capture perceptual features in AR content as effectively as they do in natural images. We propose ARaBIQA, the first blind image quality assessment (BIQA) method designed specifically for AR content. Using a self-supervised approach, ARaBIQA learns low-level AR-specific features, including distortions and visual confusion, and combines them with high-level content features through a joint fine-tuning strategy to produce robust quality predictions. The experimental results show that ARaBIQA outperforms existing blind IQA metrics, and ablation studies further validate its effectiveness.
Aymen Sekhri, Mohamed-Chaker Larabi, Seyed Ali Amirshahi
ICIP3
2024 Towards Light-Weight Transformer-Based Quality Assessment Metric for Augmented Reality
abstract
With the rise of Augmented Reality (AR) technology, which enhances the real world by overlaying computer-generated content, immersive experiences are being offered in education, entertainment, healthcare, … Assessing the quality of AR scenarios is crucial for understanding and improving user satisfaction and engagement. However, developing objective AR quality assessment methods is challenging due to the lack of data and the inherent complexity of technology, particularly in the presence of visual confusion. Existing convolution neural network-based approaches suffer from limited receptive fields and are not effective at capturing global information in visually confused AR scenarios. Additionally, to the best of our knowledge, exploring transformer capabilities for AR quality assessment is missing. Therefore, this study introduces transformAR, a lightweight transformer-based model for objective quality assessment in AR applications. This approach leverages pretrained vision transformer-based encoders to capture image content information, computes distance vectors to quantify distortions, and employs cross-attention-based decoders to model perceptual quality features. The model also integrates adapted regularization techniques and label smoothing to mitigate overfitting. Experimental results demonstrate the effectiveness of transformAR, outperforming the few existing state-of-the-art methods.
Aymen Sekhri, Seyed Ali Amirshahi, Mohamed-Chaker Larabi
MMSP2
2024 Quality of NeRF Changes with the Viewing Path an Observer Takes: A Subjective Quality Assessment of Real-time NeRF Model
abstract
Despite Neural Radiance Fields (NeRF) revolutionizing the 3D world with its free viewpoint synthesis, its quality evaluation still relies on static image datasets. While a limited number of studies have challenged this conventional evaluation process with video-based experiments, our approach introduces a well-structured subjective assessment emphasizing the importance of assessing NeRF scenes through multiple dynamic video sequences. A dataset "NeRF-4Scenes" is prepared that includes four scenes covering large areas from indoor and outdoor environments. The method generates three different camera trajectories within each scene and evaluates three NeRF models through pairwise comparison. Our comprehensive analysis has demonstrated the significance of evaluating NeRF scenes from dynamic viewpoints. The changing preferences among observers across varied paths highlight the impacts of evaluating NeRF scenes from multiple viewpoints. The dataset is available for download at https://www.ntnu.edu/colourlab/software.
Shaira Tabassum, Seyed Ali Amirshahi
QoMEX2