Alvine B. Belle

dblp:123/7745 · also Alvine Boaye Belle · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LLM-Based Safety Case Generation for Baidu Apollo: Are We there Yet?
abstract
Justifying the correct implementation of the non-functional requirements of mission-critical systems is crucial to prevent system failure. The latter could have severe consequences such as the death of people, and financial losses. Assurance cases (e.g., safety cases, security cases) can be used to prevent system failure. They are structured sets of arguments supported by evidence and aiming at demonstrating that a system's non-functional requirements have been correctly implemented. How-ever, although the availability of complete assurance cases is crucial to allow the research community to contribute to the system assurance field, it remains very challenging to access complete assurance cases due to several concerns such as confidentiality issues. Furthermore, assurance cases are usually very large documents. Still, their creation remains a manual, tedious, and error-prone process that heavily relies on domain expertise. Thus, exploring techniques to support their automatic instantiation becomes crucial. To fill these gaps, our experience paper first demonstrates the feasibility of an AMLAS-based design methodology on a case study aiming at manually creating a safety case for the ML-enabled trajectory prediction component of an open-source autonomous driving system i.e. Baidu Apollo. Our paper then reports our experience in using a Large Language Model (LLM) to automatically re-create the same safety case. The lessons we have drawn from this case study provide actionable insights that could benefit researchers and practitioners.
Oluwafemi Odu, Alvine B. Belle, Song Wang 0009
CAIN2
2025 Program Slicing in the Era of Large Language Models
abstract
Program slicing is a critical technique in software engineering, enabling developers to isolate relevant portions of code for tasks such as bug detection, code comprehension, and debugging. In this study, we investigate the application of large language models (LLMs) to both static and dynamic program slicing, with a focus on Java programs. We evaluate the performance of four state-of-the-art LLMs, i.e., GPT-4o, GPT-3.5 Turbo, Llama-2, and Gemma-7B, by leveraging advanced prompting techniques, including few-shot learning and chain-of-thought reasoning. Using a dataset of 100 Java programs derived from LeetCode problems, our experiments reveal that GPT-4o performs the best in both static and dynamic slicing across other LLMs, achieving an accuracy of 60.84% and 59.69%, respectively. Our results also show that the LLMs we experimented with are yet to achieve reasonable performance for either static slicing or dynamic slicing. Through a rigorous manual analysis, we developed a taxonomy of root causes and failure locations to explore the unsuccessful cases in more depth. We identified Complex Control Flow as the most frequent root cause of failures, with the majority of issues occurring in Variable Declarations and Assignments locations. To improve the performance of LLMs, we further examined a widely-used strategy for prompting guided by our taxonomy, i.e., prompt crafting, which involved refining the prompts to better guide the LLM through the slicing process. Our evaluation shows that prompt crafting can improve accuracy by 4%.
Kimya Khakzad Shahandashti, Mohammad Mahdi Mohajer, Alvine B. Belle, Song Wang 0009, Hadi Hemmati Lassonde
COMPSAC3
2025 SmartGSN: An Online Tool to Semi-automatically Manage Assurance Cases
Oluwafemi Odu, Daniel Méndez Beltrán, Emiliano Berrones Gutiérrez, Alvine B. Belle, Gerhard Yu, Melika Sherafat
SAFECOMP4
2025 Automatic instantiation of assurance cases from patterns using large language models
abstract
An assurance case is a structured set of arguments supported by evidence, demonstrating that a system’s nonfunctional requirements (e.g., safety, security, reliability) have been correctly implemented. Assurance case patterns serve as templates derived from previous successful assurance cases, aimed at facilitating the creation of new assurance cases. Despite using these patterns to generate assurance cases, their instantiation remains a largely manual and error-prone process that heavily relies on domain expertise. Thus, exploring techniques to support their automatic instantiation becomes crucial. This study aims to investigate the potential of Large Language Models (LLMs) in automating the generation of assurance cases that comply with specific patterns. Specifically, we formalize assurance case patterns using predicate-based rules and then utilize LLMs, i.e., GPT- 4o and GPT-4 Turbo, to automatically instantiate assurance cases from these formalized patterns. Our findings suggest that LLMs can generate assurance cases that comply with the given patterns. However, this study also highlights that LLMs may struggle with understanding some nuances related to pattern-specific relationships. While LLMs exhibit potential in the automatic generation of assurance cases, their capabilities still fall short compared to human experts. Therefore, a semi-automatic approach to instantiating assurance cases may be more practical at this time.
Oluwafemi Odu, Alvine B. Belle, Song Wang 0009, Segla Kpodjedo, Timothy Lethbridge, Hadi Hemmati
J. Syst. Softw.2
2024 Prompting GPT -4 to support automatic safety case generation
abstract
In the ever-evolving field of software engineering, the advent of large language models and conversational interfaces, exemplified by ChatGPT, represents a significant revolution. While their potential is evident in various domains, this paper expands upon our previous research, where we experimented with GPT –4, on its ability to create safety cases. A safety case is a structured argument supported by a body of evidence to demonstrate that a given system is safe to operate in a given environment. In this paper, we first determine GPT –4’s comprehension of the Goal Structuring Notation (GSN), a well-established notation for visually representing safety cases. Additionally, we conduct four distinct experiments using GPT –4 to evaluate its ability to generate safety cases within a specified system and application domain. To assess GPT –4’s performance in this context, we compare the results it produces with the ground-truth safety cases developed for an X-ray system, a machine learning-enabled component for tire noise recognition in a vehicle, and a lane management system from the automotive domain. This comparison enables us to gain valuable insights into the model’s generative capabilities. Our findings indicate that GPT –4 is able to generate moderately accurate and reasonable safety cases.
Mithila Sivakumar, Alvine B. Belle, Jinjun Shan, Kimya Khakzad Shahandashti
Expert Syst. Appl.2
2024 A PRISMA-driven systematic mapping study on system assurance weakeners
abstract
An assurance case is a structured hierarchy of claims aiming at demonstrating that a mission-critical system supports specific requirements (e.g., safety, security, privacy). The presence of assurance weakeners (i.e., assurance deficits, logical fallacies) in assurance cases reflects insufficient evidence, knowledge, or gaps in reasoning. These weakeners can undermine confidence in assurance arguments, potentially hindering the verification of mission-critical system capabilities which could result in catastrophic outcomes (e.g., loss of lives). Given the growing interest in employing assurance cases to ensure that systems are developed to meet their requirements, exploring the management of assurance weakeners becomes beneficial. As a stepping stone for future research on assurance weakeners, we aim to initiate the first comprehensive systematic mapping study on this subject. We followed the well-established PRISMA 2020 and SEGRESS guidelines to conduct our systematic mapping study. We searched for primary studies in five digital libraries and focused on the 2012–2023 publication year range. Our selection criteria focused on studies addressing assurance weakeners from a qualitative standpoint, resulting in the inclusion of 39 primary studies in our systematic review. Our systematic mapping study reports a taxonomy (map) that provides a uniform categorization of assurance weakeners and approaches proposed to manage them from a qualitative perspective. The taxonomy classifies weakeners in four categories: aleatory, epistemic, ontological, and argument uncertainty. Additionally, it classifies approaches supporting the management of weakeners in three main categories: representation, identification and mitigation approaches. Our study findings suggest that the SACM (Structured Assurance Case Metamodel) – a standard specified by the OMG (Object Management Group) – offers a comprehensive range of capabilities to capture structured arguments and reason about their potential assurance weakeners. Our findings also suggest novel assurance weakener management approaches should be proposed to better assure mission-critical systems.
Kimya Khakzad Shahandashti, Alvine B. Belle, Timothy Lethbridge, Oluwafemi Odu, Mithila Sivakumar
Inf. Softw. Technol.2
2023 Evidence-based decision-making: On the use of systematicity cases to check the compliance of reviews with reporting guidelines such as PRISMA 2020
abstract
Systematic reviews aim to provide high-quality evidence-based syntheses for efficacy under real-world conditions and allow understanding the correlations between exposures and outcomes. They are increasingly popular and have several stakeholders (e.g., healthcare providers, researchers, educators, students, journal editors, policy makers, managers) to whom they help make informed recommendations for practice or policy. Systematic reviews usually exhibit low methodological and reporting quality. To tackle this, reporting guidelines have been developed to support systematic reviews reporting and assessment. Following such guidelines is crucial to ensure that a review is transparent, complete, trustworthy, reproducible, and unbiased. However, systematic reviewers usually fail to adhere to existing reporting guidelines, which may significantly decrease the quality of the reviews they report and may result in systematic reviews that lack methodological rigor, yield low-credible findings and may mislead decision-makers. To assure that a review complies with reporting guidelines, we rely on assurance cases that are an emerging way of arguing and relaying various safety–critical systems’ requirements in an extensive manner, as well as checking the compliance of such systems with standards to support their certification. Since the nature of assurance cases makes them applicable to various domains and requirements/properties, we therefore propose a new type of assurance cases called systematicity cases. Systematicity cases focus on the systematicity property and allow arguing that a review is systematic i.e., that it sufficiently complies with the targeted reporting guideline. The most widespread reporting guidelines include PRISMA (Preferred Reporting Items for Systematic reviews and meta-Analyses). We measure the confidence in a systematicity case representing a review as a means to quantify the systematicity of that review i.e., the extent to which that review is systematic. We rely on rule-based Artificial Intelligence to create a knowledge-based system that automatically supports the inference mechanism that a given systematicity case embodies and that allows making a decision regarding the systematicity of a given review. An empirical evaluation performed on 25 reviews (self-identifying as systematic) showed that these reviews exhibit a suboptimal systematicity. More specifically, the systematicity of the analyzed reviews varies between 32.96% and 66.49% and its average is 54.42%. More efforts are therefore needed to report systematic reviews of higher quality. More experiments are also needed to further explore the factors hindering and/or assuring the systematicity of reviews. The main beneficiaries of our work are journal reviewers, journal editors, managers, policymakers, researchers, organizations developing reporting guidelines, peer reviewers, students, insurers, evidence users, as well as reporting guidelines developers.
Alvine B. Belle, Yixi Zhao
Expert Syst. Appl.1
2023 Bolstering the Persistence of Black Students in Undergraduate Computer Science Programs: A Systematic Mapping Study
abstract
Background: People who are racialized, gendered, or otherwise minoritized are underrepresented in computing professions in North America. This is reflected in undergraduate computer science (CS) programs, in which students from marginalized backgrounds continue to experience inequities that do not typically affect White cis-men. This is especially true for Black students in general, and Black women in particular, whose experience of systemic, anti-Black racism compromises their ability to persist and thrive in CS education contexts. Objectives: This systematic mapping study endeavours to (1) determine the quantity of existing non-deficit-based studies concerned with the persistence of Black students in undergraduate CS; (2) summarize the findings and recommendations in those studies; and (3) identify areas in which additional studies may be required. We aim to accomplish these objectives by way of two research questions: (RQ1) What factors are associated with Black students’ persistence in undergraduate CS programs?; and (RQ2) What recommendations have been made to further bolster Black students’ persistence in undergraduate CS education programs? Methods: This systematic mapping study was conducted in accordance with PRISMA 2020 and SEGRESS guidelines. Studies were identified by conducting keyword searches in seven databases. Inclusion and exclusion criteria were designed to capture studies illuminating persistence factors for Black students in undergraduate CS programs. To ensure the completeness of our search results, we engaged in snowballing and an expert-based search to identify additional studies of interest. Finally, data were collected from each study to address the research questions outlined above. Results: Using the methods outlined above, we identified 16 empirical studies, including qualitative, quantitative, and mixed-methods studies informed by a range of theoretical frameworks. Based on data collected from the primary studies in our sample, we identified 13 persistence factors across four categories: (I) social capital, networking, & support; (II) career & professional development; (III) pedagogical & programmatic interventions; and (IV) exposure & access. This data-collection process also yielded 26 recommendations across six stakeholder groups: (i) researchers; (ii) colleges and universities; (iii) the computing industry; (iv) K-12 systems and schools; (v) governments; and (vi) parents. Conclusion: This systematic mapping study resulted in the identification of numerous persistence factors for Black students in CS. Crucially, however, these persistence factors allow Black students to persist, but not thrive, in CS. Accordingly, we contend that more needs to be done to address the systemic inequities faced by Black people in general, and Black women in particular, in computing programs and professions. As evidenced by the relatively small number of primary studies captured by this systematic mapping study, there exists an urgent need for additional, asset-based empirical studies involving Black students in CS. In addition to foregrounding the intersectional experiences of Black women in CS, future studies should attend to the currently understudied experiences of Black men.
Alvine B. Belle, Callum Sutherland, Opeyemi Adesina, Segla Kpodjedo, Nathanael Ojong, Lisa Cole
ACM Trans. Comput. Educ.1
2022 A checklist-based approach to assess the systematicity of the abstracts of reviews self-identifying as systematic reviews
abstract
Systematic reviews are crucial for various stakeholders since they allow them to make evidence-based decisions without being overwhelmed by a large volume of research. Systematic reviews are increasingly popular in the software engineering field. The abstract is one of the most important systematic review’s components since it usually reflects the content of the review. It may be the only part of the review that most of the readers will read when needing to form an opinion on a given topic. Besides, the content of an abstract is usually the main information readers use to decide if they want to access the full content of the review or not. Since an abstract usually summarizes a review, readers may therefore mostly rely on that abstract to judge the quality of the review as well as its methodological rigor. However, abstracts are sometimes poorly written and may therefore give a misleading and even harmful picture of the reviews’ contents. To assess abstracts, we propose a measure that allows quantifying the systematicity of reviews’ abstracts i.e., the extent to which these abstracts exhibit good reporting quality. Experiments on 151 reviews published in the software engineering (SE) field showed that these reviews’ abstracts exhibit a suboptimal systematicity.
Alvine B. Belle, Yixi Zhao
APSEC1
2022 A new measure to assess the systematicity of the abstracts of reviews self-identifying as systematic reviews
abstract
Systematic reviews are crucial for various stakeholders since they allow them to make evidence-based decisions without being overwhelmed by a large volume of research. The abstract is one of the most important systematic review’s components since it usually reflects the content of the review. It may be the only part of the review that most of thereaders will read when needing to form an opinion on a given topic. Besides, the content of an abstract is usually the main information readers use to decide if they want to access the full content of the review or not. Since an abstract usually summarizes a review, readers may therefore mostly rely on that abstract to judge the quality of the review as well as its methodological rigor. However, abstracts are usually poorly written and may therefore give a misleading and even harmful picture of the reviews’ contents. To assess abstracts, we propose a measure that allows quantifying the systematicity of reviews’ abstracts i.e., the extent to which these abstracts exhibit good reporting quality. Experiments on 151 reviews published in the software engineering field showed that these reviews’ abstracts exhibit a suboptimal systematicity.
Alvine B. Belle, Yixi Zhao
APSEC1
2018 Improving formal analysis of state machines with particular emphasis on and-cross transitions
Opeyemi Adesina, Timothy Lethbridge, Stéphane S. Somé, Vahdat Abdelzad, Alvine B. Belle
Comput. Lang. Syst. Struct.5
2018 Design and implementation of distributed expert systems: On a control strategy to manage the execution flow of rule activation
Alvine B. Belle, Timothy Lethbridge, Miguel Garzón, Opeyemi Adesina
Expert Syst. Appl.1
2016 Combining lexical and structural information to reconstruct software layers
Alvine B. Belle, Ghizlane El-Boussaidi, Segla Kpodjedo
Inf. Softw. Technol.1
2015 The Layered Architecture Recovery as a Quadratic Assignment Problem
Alvine B. Belle, Ghizlane El-Boussaidi, Christian Desrosiers, Segla Kpodjedo, Hafedh Mili
ECSA1
2014 Recovering Software Layers from Object Oriented Systems
abstract
Recovering the architecture of existing software systems remains a challenge and an active research field in software engineering. In this paper, we propose an approach to recover the layered architecture of object oriented software systems. To do so, our approach first recovers clusters corresponding to the various responsibilities of the system; the challenge in this context is to find the appropriate level of granularity of these responsibilities. Then the recovered clusters are assigned to layers using an optimization algorithm that exploits the principles of the layering architectural style. The approach was validated on five Java open source systems.
Alvine B. Belle, Ghizlane El-Boussaidi, Hafedh Mili
ENASE1
2013 The Layered Architecture revisited: Is it an Optimization Problem?
Alvine B. Belle, Ghizlane El-Boussaidi, Christian Desrosiers, Hafedh Mili
SEKE1