MohammadAli Hamidi

dblp:371/6070 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0002-4671-5274ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 A Hybrid Compression-Aware Ensemble Model for No-Reference Video Quality Assessment
MohammadAli Hamidi, Hadi Amirpour, Christian Timmerer, Luigi Atzori
QoMEX1
2026 QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu
QoMEX25
2025 MVAW-PCQA: A No-reference Point Cloud Quality Assessment via Multi-View Adaptive Weighting
abstract
Point cloud quality assessment (PCQA) is a critical research area focused on evaluating the perceptual Quality of Experience (QoE) of point clouds to enhance visual experiences of immersive multimedia applications for end users. To prevent the complex computations on 3D data applied by model-based methods, projection-based models have been developed to estimate the QoE by analysing 2D projection views of the point cloud. In this paper, we propose a novel projection-based No-Reference (NR) PCQA method, called Multi-View Adaptive Weighting Point Cloud Quality Assessment (MVAW-PCQA), to predict the QoE of distorted point clouds using six 2D projection views as the input of a convolutional neural network (CNN) architecture. First, multi-view involves independently extracting features from multiple projection views of a point cloud, guaranteeing view-specific features are learned without prematurely mixing spatial information, and preserving the unique contributions of each projection view to the final quality prediction. Then, an adaptive weighting fusion mechanism combines the features extracted from the different projection views by learning their relative importance. This design enables the model to focus on the most informative projections for predicting the point cloud quality. The experimental results demonstrate that our method outperforms state-of-the-art NR-PCQA methods on the SJTU-PCQA dataset in terms of root mean square error (RMSE) and correlation coefficients (Pearson, Spearman, and Kendall), while adopting a lightweight design with a reasonable number of parameters for the trained neural network.
MohammadAli Hamidi, Simone Porcu, Alessandro Floris, Luigi Atzori
QoMEX1
2025 VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and Results
abstract
This paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD.
Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde
VCIP5
2025 Learning-Based Objective Perceptual Quality Models Across Interactive and Immersive Media
abstract
Ensuring high-quality multimedia experiences at scale requires objective models of Quality of Experience (QoE) that approximate human perception reliably across diverse contents, devices, and networks. This doctoral thesis addresses the problem of defining and deploying such models by proposing a unified methodology structured around five principles: careful signal design, temporal modeling of event sequences, multi-view deep learning fusion under data constraints, perceptual alignment through thresholds such as JND/SUR, and efficiency for deployment in resource-limited settings. The methodology has been instantiated across multiple studies, ranging from real-time communication and adaptive streaming to point cloud and face image quality assessment. Across these use cases, the models show strong agreement with subjective judgments, robustness to content and device shifts, and efficiency compatible with online deployment. Collectively, these contributions establish a general recipe for perception-aware QoE modeling, bridging idea and operational practice.
MohammadAli Hamidi
VCIP1
2025 Perceptual JND Prediction for VMAF Using Content-Adaptive Dual-Path Attention
abstract
Just Noticeable Difference (JND) thresholds, particularly for quality metrics such as Video Multimethod Assessment Fusion (VMAF), are critical in streaming, helping identify when quality changes become perceptible and reducing redundant bitrate representations. The Satisfied User Ratio (SUR) complements JND by quantifying the percentage of users who do not perceive a difference, offering practical guidance for perceptually optimized streaming. This paper proposes a novel two-branch deep neural network (DNN) for predicting the 75% SUR for VMAF, the encoding level where 75% of viewers cannot perceive degradation. The framework combines handcrafted features (e.g., spatial and temporal indicators such as SI, TI, etc.) and deep learning-based (DL-based) representations extracted via a convolutional neural network (CNN) backbone. The DL-based branch employs a spatio-temporal attention mechanism and a Long Short-Term Memory (LSTM) to capture temporal dynamics, while the handcrafted branch encodes interpretable indicators through a fully connected layer. Both outputs are fused and passed through a lightweight Multilayer Perceptron (MLP) to predict 75% SUR. To improve robustness to noise and label uncertainty, the model is trained using the Smooth-L1 loss. Experiments on the VideoSet dataset show our method outperforms SOTA across all metrics, achieving a notably higher R2score (0.46 vs. 0.36), indicating improved prediction reliability and low computational complexity, making it suitable for real-time video streaming.
MohammadAli Hamidi, Hadi Amirpour, Christian Timmerer, Luigi Atzori
VCIP1
2024 Towards the Application of Multi-view Learning in Quality of Experience Collaborative Modelling
abstract
Multi-view (MV) learning is a machine learning technique for improving generalization efficiency by learning from different feature subsets derived from multiple sources. We believe this approach can help in Quality of Experience (QoE) modelling by integrating knowledge from different datasets generated by subjective tests conducted for the same or similar applications considering different QoE Influence Factors (IFs). To investigate this subject, in this paper, we present the experiments conducted starting from a complete dataset related to Web browsing sessions that has been artificially divided into two distinct subsets (views). The proposed MV learning approach implements a data fusion technique to integrate extracted features from different views into a unified feature space. To achieve a complete experiment on the entire problem space, all possible combinations of IFs (features) in two distinct partial views (PVs) are considered and trained in the MV approach; the full view (FV) approach, which utilizes the complete dataset, is also considered for performance comparison. Experimental results show the QoE estimation performance achieved by the MV (0.69) is comparable with that of the FV (0.72), although the 2 single views were used for training in the MV case. Moreover, the performance enhancement achieved by the MV compared with the PV is most noticeable when a lower number of features is used to train the models.
MohammadAli Hamidi, Simone Porcu, Alessandro Floris, Luigi Atzori
QoMEX1