Mohsen Azarmi

dblp:274/0708 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-0737-9204ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visualization and visual analytics · 70% Image and video coding · 30%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics › visual analytics
anomaly detection visualization
1.012026
OM4AnI: A Novel Overlap Measure for Anomaly Identification in Multi-Class Scatterplots · IEEE Trans. Vis. Comput. Graph. 2026
Visualization and visual analytics
scatterplot
1.012026
OM4AnI: A Novel Overlap Measure for Anomaly Identification in Multi-Class Scatterplots · IEEE Trans. Vis. Comput. Graph. 2026
Image and video coding › quality assessment
visual quality measure
1.012026
OM4AnI: A Novel Overlap Measure for Anomaly Identification in Multi-Class Scatterplots · IEEE Trans. Vis. Comput. Graph. 2026

Methods — techniques the papers use, named apart from their topics

pixel-level binning · 1.0
YearPublicationVenuePosition
2026 Identifying OM4AnI's Effectiveness in the Context of Explainable AI
abstract
Scatterplots are widely used in Explainable Artificial Intelligence (XAI) to investigate misclassifications and patterns across instances. However, a significant limitation of scatterplots is overplotting, especially when working with large datasets. Although several quality metrics have been proposed to measure the degree of overplotting, none have been demonstrated to be effective in the context of XAI. This paper aims to evaluate the effectiveness of a quality metric, called OM4AnI, in XAI scenarios. We begin by summarizing two visual patterns—cluster-based and regression-based patterns—that support three common XAI tasks: feature importance, feature dependency, and model accuracy. We also introduce how to select the parameters of OM4AnI based on these patterns. We construct two case studies to identify the effectiveness of OM4AnI using public datasets: Census Income dataset and MNIST dataset. OM4AnI is applied to both scenarios under various visual conditions (e.g., marker size and rendering order) to assess its effectiveness. The results demonstrate that OM4AnI serves as an effective quality metric for these two common XAI scenarios, paving the way for adapting other quality metrics to be scalable within XAI contexts.
Liqun Liu 0003, Leonid V. Bogachev, Mahdi Rezaei 0001, Nishant Ravikumar, Arjun Khara, Mohsen Azarmi, Roy A. Ruddle
PacificVis6
2026 OM4AnI: A Novel Overlap Measure for Anomaly Identification in Multi-Class Scatterplots
abstract
Scatterplots are widely used across various domains to identify anomalies in datasets, particularly in multi-class settings, such as detecting misclassified or mislabeled data. However, scatterplot effectiveness often declines with large datasets due to limited display resolution. This paper introduces a novel Visual Quality Measure (VQM) - OM4AnI (Overlap Measure for Anomaly Identification) - which quantifies the degree of overlap for identifying anomalies, helping users estimate how effectively anomalies can be observed in multi-class scatterplots. OM4AnI begins by computing anomaly index based on each data point's position relative to its class cluster. The scatterplot is then discretized into a matrix representation by binning the display space into cell-level (pixel-level) grids and computing the coverage for each pixel. It takes into account the anomaly index of data points covering these pixels and visual features (marker shapes, marker sizes, and rendering orders). Building on this foundation, we sum all the coverage information in each cell (pixel) of matrix representation to obtain the final quality score with respect to anomaly identification. We conducted an evaluation to analyze the efficiency, effectiveness, sensitivity of OM4AnI in comparison with six representative baseline methods that are based on different computation granularity levels: data level, marker level, and pixel level. The results show that OM4AnI outperforms baseline methods by exhibiting more monotonic trends against the ground truth and greater sensitivity to rendering order, unlike the baseline methods. It confirms that OM4AnI can inform users about how effectively their scatterplots support anomaly identification. Overall, OM4AnI shows strong potential as an evaluation metric and for optimizing scatterplots through automatic adjustment of visual parameters.
Liqun Liu 0003, Leonid V. Bogachev, Mahdi Rezaei 0001, Nishant Ravikumar, Arjun Khara, Mohsen Azarmi, Roy A. Ruddle
IEEE Trans. Vis. Comput. Graph.6
2025 Pedestrian Intention Prediction via Vision-Language Foundation Models
abstract
Prediction of pedestrian crossing intention is a critical function in autonomous vehicles. Conventional vision-based methods of crossing intention prediction often struggle with generalizability, context understanding, and causal reasoning. This study explores the potential of vision-language foundation models (VLFMs) for predicting pedestrian crossing intentions by integrating multimodal data through hierarchical prompt templates. The methodology incorporates contextual information, including visual frames, physical cues observations, and ego-vehicle dynamics, into systematically refined prompts to guide VLFMs effectively in intention prediction. Experiments were conducted on three common datasets—JAAD, PIE, and FU-PIP. Results demonstrate that incorporating vehicle speed, its variations over time, and time-conscious prompts significantly enhances the prediction accuracy up to 19.8%. Additionally, optimised prompts generated via an automatic prompt engineering framework yielded 12.5% further accuracy gains. These findings highlight the superior performance of VLFMs compared to conventional vision-based models, offering enhanced generalisation and contextual understanding for autonomous driving applications.
Mohsen Azarmi, Mahdi Rezaei 0001, He Wang 0002
IV1
2025 Driver-Net: Multi-Camera Fusion for Assessing Driver Take-Over Readiness in Automated Vehicles
abstract
Ensuring safe transition of control in automated vehicles requires an accurate and timely assessment of driver readiness. This paper introduces Driver-Net, a novel deep learning framework that fuses multi-camera inputs to estimate driver take-over readiness. Unlike conventional vision-based driver monitoring systems that focus on head pose or eye gaze, Driver-Net captures synchronised visual cues from the driver's head, hands, and body posture through a triple-camera setup. The model integrates spatio-temporal data using a dual-path architecture, comprising a Context Block and a Feature Block, followed by a cross-modal fusion strategy to enhance prediction accuracy. Evaluated on a diverse dataset collected from the University of Leeds Driving Simulator, the proposed method achieves an accuracy of up to 95.8% in driver readiness classification. This performance significantly enhances existing approaches and highlights the importance of multimodal and multi-view fusion. As a real-time, non-intrusive solution, Driver-Net contributes meaningfully to the development of safer and more reliable automated vehicles and aligns with new regulatory mandates and upcoming safety standards.
Mahdi Rezaei 0001, Mohsen Azarmi
IV2
2025 PIP-Net: Pedestrian Intention Prediction in the Wild
abstract
Accurate pedestrian intention prediction (PIP) by Autonomous Vehicles (AVs) is one of the current research challenges in this field. In this article, we introduce PIP-Net, a novel framework designed to predict pedestrian crossing intentions by AVs in real-world urban scenarios. We offer two variants of PIP-Net designed for different camera mounts and setups. Leveraging both kinematic data and spatial features from the driving scene, the proposed model employs a recurrent and temporal attention-based solution, outperforming state-of-the-art performance. To enhance the visual representation of road users and their proximity to the ego vehicle, we introduce a categorical depth feature map, combined with a local motion flow feature, providing rich insights into the scene dynamics. Additionally, we explore the impact of expanding the camera’s field of view, from one to three cameras surrounding the ego vehicle, leading to an enhancement in the model’s contextual perception. Depending on the traffic scenario and road environment, the model excels in predicting pedestrian crossing intentions up to 4 seconds in advance, which is a breakthrough in current research studies in pedestrian intention prediction. Finally, for the first time, we present the Urban-PIP dataset, a customised pedestrian intention prediction dataset, with multi-camera annotations in real-world automated driving scenarios.
Mohsen Azarmi, Mahdi Rezaei 0001, He Wang 0002
IEEE Trans. Intell. Transp. Syst.1
2024 AllWeather-Net: Unified Image Enhancement for Autonomous Driving Under Adverse Weather and Low-Light Conditions
Chenghao Qian, Mahdi Rezaei 0001, Saeed Anwar, Wenjing Li 0005, Tanveer Hussain 0001, Mohsen Azarmi, Wei Wang 0335
ICPR (30)6
2023 3D-Net: Monocular 3D object recognition for traffic monitoring
abstract
Machine Learning has played a major role in various applications including Autonomous Vehicles and Intelligent Transportation Systems. Utilizing a deep convolutional neural network, the article introduces a zero-calibration 3D Object recognition and tracking system for traffic monitoring. The model can accurately work on urban traffic cameras, regardless of their technical specification (i.e. resolution, lens, the field of view) and positioning (location, height, angle). For the first time, we introduce a novel satellite-ground inverse perspective mapping technique, which requires no camera calibrations and only needs the GPS position of the camera. This leads to an accurate environmental modeling solution that is capable of estimating road users’ 3D bonding boxes, speed, and trajectory using a monocular camera. We have also contributed to a hierarchical activity/traffic modeling solution using short- and long-term Spatio-temporal video analysis to understand the heatmap of the traffic flow, bottlenecks, and high-risk zones. The experiments are conducted on four datasets: MIO-TCD, UA-DETRAC, GRAM-RTM, and Leeds-Dataset including various use cases and traffic scenarios.
Mahdi Rezaei 0001, Mohsen Azarmi, Farzam Mohammad Pour Mir
Expert Syst. Appl.2