Amirhossein Zolfagharian

dblp:322/6284 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-2411-7938ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Software testing · 70% Program verification · 30%
Artificial intelligence
2 papers
Reinforcement learning · 70% Trustworthy machine learning · 17% Autonomous driving · 13%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
1.122025
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2025
A Search-Based Testing Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2023
Program verification › dynamic verification
runtime verification
0.912025
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2025
Software testing › deep learning testing
deep reinforcement learning agent testing
0.712023
A Search-Based Testing Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2023
Software testing
search-based software testing
0.712023
A Search-Based Testing Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2023
Software testing
test generation
0.712023
A Search-Based Testing Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2023
Robotics › Autonomous driving › safety validation
safety testing
0.212023
A Search-Based Testing Approach for Deep Reinforcement Learning Agents · IEEE Trans. Software Eng. 2023

Methods — techniques the papers use, named apart from their topics

machine learning · 3.1state abstraction · 1.7search-based testing · 1.3genetic algorithm · 1.3q-values · 0.9q-value · 0.9
YearPublicationVenuePosition
2025 SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
abstract
Deep Reinforcement Learning (DRL) has made significant advancements in various fields, such as autonomous driving, healthcare, and robotics, by enabling agents to learn optimal policies through interactions with their environments. However, the application of DRL in safety-critical domains presents challenges, particularly concerning the safety of the learned policies. DRL agents, which are focused on maximizing rewards, may select unsafe actions, leading to safety violations. Runtime safety monitoring is thus essential to ensure the safe operation of these agents, especially in unpredictable and dynamic environments. This paper introducesSMARLA, a black-box safety monitoring approach specifically designed for DRL agents.SMARLAutilizes machine learning to predict safety violations by observing the agent's behavior during execution. The approach is based on Q-values, which reflect the expected reward for taking actions in specific states.SMARLAemploys state abstraction to reduce the complexity of the state space, enhancing the predictive capabilities of the monitoring model. Such abstraction enables the early detection of unsafe states, allowing for the implementation of corrective and preventive measures before incidents occur. We quantitatively and qualitatively validatedSMARLAon three well-known case studies widely used in DRL research. Empirical results reveal thatSMARLAis accurate at predicting safety violations, with a low false positive rate, and can predict violations at an early stage, approximately halfway through the execution of the agent, before violations occur. We also discuss different decision criteria, based on confidence intervals of the predicted violation probabilities, to trigger safety mechanisms aiming at a trade-off between early detection and low false positive rates.
Amirhossein Zolfagharian, Manel Abdellatif, Lionel C. Briand, S. Ramesh 0002
IEEE Trans. Software Eng.1
2023 A Search-Based Testing Approach for Deep Reinforcement Learning Agents
abstract
Deep Reinforcement Learning (DRL) algorithms have been increasingly employed during the last decade to solve various decision-making problems such as autonomous driving, trading decisions, and robotics. However, these algorithms have faced great challenges when deployed in safety-critical environments since they often exhibit erroneous behaviors that can lead to potentially critical errors. One of the ways to assess the safety of DRL agents is to test them to detect possible faults leading to critical failures during their execution. This raises the question of how we can efficiently test DRL policies to ensure their correctness and adherence to safety requirements. Most existing works on testing DRL agents use adversarial attacks that perturb states or actions of the agent. However, such attacks often lead to unrealistic states of the environment. Furthermore, their main goal is to test the robustness of DRL agents rather than testing the compliance of the agents' policies with respect to requirements. Due to the huge state space of DRL environments, the high cost of test execution, and the black-box nature of DRL algorithms, exhaustive testing of DRL agents is impossible. In this paper, we propose a Search-based Testing Approach of Reinforcement Learning Agents (STARLA) to test the policy of a DRL agent by effectively searching for failing executions of the agent within a limited testing budget. We rely on machine learning models and a dedicated genetic algorithm to narrow the search toward faulty episodes (i.e., sequences of states and actions produced by the DRL agent). We apply STARLA on Deep-Q-Learning agents trained on two different RL problems widely used as benchmarks and show that STARLA significantly outperforms Random Testing by detecting more faults related to the agent's policy. We also investigate how to extract rules that characterize faulty episodes of the DRL agent using our search results. Such rules can be used to understand the conditions under which the agent fails and thus assess the risks of deploying it.
Amirhossein Zolfagharian, Manel Abdellatif, Lionel C. Briand, Mojtaba Bagherzadeh, S. Ramesh 0002
IEEE Trans. Software Eng.1