VLDB 2026 Research / reviewers in the wild / expert
Saad Ezzini
dblp:216/8359
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-7657-4738ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to represent code changes
Xunzhu Tang, Haoye Tian, Weiguo Pian, Saad Ezzini, Abdoul Kader Kaboré, Andrew Habib, Kisub Kim, Jacques Klein, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 4 |
| 2026 | DarijaDB: Unlocking Text-to-SQL for Arabic DialectsabstractRecent advances in text-to-SQL models, which translate natural language questions (NLQs) into executable SQL queries, have made interacting with relational databases more accessible, even for those with limited technical ability. This task has seen significant improvement with the release of multiple English datasets and benchmarks such as WikiSQL, SPIDER, and BIRD, each covering different domains and levels of complexity. Non-English high-resource languages, such as Chinese, Russian, and Arabic, have also benefited from these advances, either through the translation of existing datasets or the creation of new ones. Dialect2SQL , a newly released text-to-SQL dataset, is dedicated to the Moroccan dialect (Darija), which is known for its complexity and distinctiveness compared to other Arabic dialects and Modern Standard Arabic. In this article, we conduct a comprehensive study on text-to-SQL for Darija by conducting several experiments mainly on the Dialect2SQL dataset using different approaches and configurations with two code-based large language models, StarCoder2 and Qwen-2.5-Coder. The experiments reveal the performance gap between models fine-tuned on English data and those fine-tuned on Darija. Additionally, the results illustrate the positive impact of incorporating multi-language datasets during training. In particular, the gap decreases from 10.1% to 6.7% in BLEU, and from 12.5% to 5.7% in TSED. Salmane Chafik, Saad Ezzini, Ismail Berrada |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2026 | Towards Automating Domain-Specific Data Generation for Text-to-SQL: A Comprehensive ApproachabstractAs software systems increasingly rely on natural language interfaces, ensuring the reliability of these systems is crucial. One critical component is the ability to accurately translate natural language queries into corresponding SQL queries, a field known as Text-to-SQL. However, the scarcity of high-quality, large-scale, and domain-specific Text-to-SQL datasets hinders the development of reliable and robust models. To tackle these challenges, we propose SelectCraft , a novel automatic generation approach designed to create realistic Text-to-SQL datasets tailored to specific domains. Our method leverages existing databases and their structures to generate complex text-SQL pairs that mirror real-world usage scenarios. As a proof of concept, we have successfully generated a substantial financial Text-to-SQL dataset, denominated as BanQies , encompassing over 1 million samples utilizing our proposed approach. Moreover, we introduce BanQL , a new large language model (LLM) based on StarCoder2 , a state-of-the-art code-based LLM, and fine-tuned on our newly created dataset. We evaluate BanQL performance against several state-of-the-art models, demonstrating significant enhancements in accuracy and generalizability, highlighting the advantages of incorporating domain-specific data in Text-to-SQL tasks. We firmly believe that our contributions have the potential to improve the overall reliability of Text-to-SQL software systems. CCS Concepts: • Software and its engineering; • Computing methodologies → Natural languageprocessing; • Information systems → Structured Query Language; Salmane Chafik, Saad Ezzini, Ismail Berrada |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2026 | Just-in-Time Detection of Silent Security PatchesabstractOpen source code is pervasive. In this setting, embedded vulnerabilities are spreading to downstream software at an alarming rate. Although such vulnerabilities are generally identified and addressed rapidly, inconsistent maintenance policies can cause security patches to go unnoticed. Indeed, security patches can be silent , i.e., they do not always come with comprehensive advisories such as CVEs. This lack of transparency leaves users oblivious to available security updates, providing ample opportunity for attackers to exploit unpatched vulnerabilities. Consequently, identifying silent security patches just in time when they are released is essential for preventing n-day attacks and for ensuring robust and secure maintenance practices. With llmda we propose to (1) leverage large language models (LLMs) to augment patch information with generated code change explanations, (2) design a representation learning approach that explores code-text alignment methodologies for feature combination, (3) implement a label-wise training with labeled instructions for guiding the embedding based on security relevance, and (4) rely on a probabilistic batch contrastive learning mechanism for building a high-precision identifier of security patches. We evaluate llmda on the PatchDB and SPI-DB literature datasets and show that our approach substantially improves over the state of the art, notably GraphSPD by 20% in terms of F-Measure on the SPI-DB benchmark. Xunzhu Tang, Kisub Kim, Saad Ezzini, Yewei Song, Haoye Tian, Jacques Klein, Tegawendé F. Bissyandé |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | The IoT Whisperer: A Framework for Intelligent IoT Service Composition Through LLMsabstractManually configuring Internet of Things (IoT) work-flows often involves intricate and time-consuming processes. This complexity can lead to an increase in development time and difficulty in managing and scaling IoT systems. The challenge lies in creating efficient and deployable workflows that seamlessly integrate diverse IoT devices and services. To address these challenges, we propose an approach that leverages IoT system orchestration, workflow development, and Large Language Models (LLMs) to convert high-level natural language system descriptions into executable workflows. Utilizing the Web of Things (WoT) framework, Node-RED and LLMs, we enabled the automatic generation of workflow components from natural language descriptions, enhancing the accessibility and usability of IoT system development for non-experts. We describe an assessment framework for the evaluation of abstract workflow descriptions and use this to analyze the implementation of our approach over systems within various domains, comparing the automatically generated workflows with those produced manu-ally. Our approach provides a powerful and versatile solution for IoT system development. This approach not only simplifies the process, but also provides a foundation for future innovations in the IoT landscape, paving the way for more adaptable and user-friendly IoT solutions. Ewan Warburton, Abdessalam Elhabbash, Saad Ezzini, Yehia El-khatib |
CLOUD | 3 |
| 2025 | CallNavi, A challenge and empirical study on LLM function calling and routingabstractpeer reviewed Yewei Song, Xunzhu Tang, Cedric Lothritz, Saad Ezzini, Jacques Klein, Tegawendé F. Bissyandé, Andrey Boytsov, Ulrick Ble, Anne Goujon |
EASE | 4 |
| 2025 | Correction to: App review driven collaborative bug finding
Xunzhu Tang, Haoye Tian, Pingfan Kong, Saad Ezzini, Kui Liu 0001, Xin Xia 0001, Jacques Klein, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 4 |
| 2024 | CodeAgent: Autonomous Communicative Agents for Code ReviewabstractXunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz, Bei Li, Saad Ezzini, Haoye Tian, Jacques Klein, Tegawendé F. Bissyandé. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xunzhu Tang, Kisub Kim, Yewei Song, Cedric Lothritz, Saad Ezzini, Haoye Tian, Jacques Klein, Tegawendé F. Bissyandé |
EMNLP | 6 |
| 2024 | App review driven collaborative bug findingabstractSoftware development teams generally welcome any effort to expose bugs in their code base. In this work, we build on the hypothesis that mobile apps from the same category (e.g., two web browser apps) may be affected by similar bugs in their evolution process. It is therefore possible to transfer the experience of one historical app to quickly find bugs in its new counterparts. This has been referred to as collaborative bug finding in the literature. Our novelty is that we guide the bug finding process by considering that existing bugs have been hinted within app reviews. Concretely, we design the BugRMSys approach to recommend bug reports for a target app by matching historical bug reports from apps in the same category with user app reviews of the target app. We experimentally show that this approach enables us to quickly expose and report dozens of bugs for targeted apps such as Brave (web browser app). BugRMSys 's implementation relies on DistilBERT to produce natural language text embeddings. Our pipeline considers similarities between bug reports and app reviews to identify relevant bugs. We then focus on the app review as well as potential reproduction steps in the historical bug report (from a same-category app) to reproduce the bugs. Overall, after applying BugRMSys to six popular apps, we were able to identify, reproduce and report 20 new bugs: among these, 9 reports have been already triaged, 6 were confirmed, and 4 have been fixed by official development teams. Xunzhu Tang, Haoye Tian, Pingfan Kong, Saad Ezzini, Kui Liu 0001, Xin Xia 0001, Jacques Klein, Tegawendé F. Bissyandé |
Empir. Softw. Eng. | 4 |
| 2023 | AI-based Question Answering Assistance for Analyzing Natural-language RequirementsabstractBy virtue of being prevalently written in natural language (NL), requirements are prone to various defects, e.g., inconsistency and incompleteness. As such, requirements are frequently subject to quality assurance processes. These processes, when carried out entirely manually, are tedious and may further overlook important quality issues due to time and budget pressures. In this paper, we propose QAssist - a question-answering (QA) approach that provides automated assistance to stakeholders, including requirements engineers, during the analysis of NL requirements. Posing a question and getting an instant answer is beneficial in various quality-assurance scenarios, e.g., incompleteness detection. Answering requirements-related questions automatically is challenging since the scope of the search for answers can go beyond the given requirements specification. To that end, QAssist provides support for mining external domain-knowledge resources. Our work is one of the first initiatives to bring together QA and external domain knowledge for addressing requirements engineering challenges. We evaluate QAssist on a dataset covering three application domains and containing a total of 387 question-answer pairs. We experiment with state-of-the-art QA methods, based primarily on recent large-scale language models. In our empirical study, QAssist localizes the answer to a question to three passages within the requirements specification and within the external domain-knowledge resource with an average recall of 90.1% and 96.5%, respectively. QAssist extracts the actual answer to the posed question with an average accuracy of 84.2%. Saad Ezzini, Sallam Abualhaija, Chetan Arora 0002, Mehrdad Sabetzadeh |
ICSE | 1 |
| 2022 | Automated Handling of Anaphoric Ambiguity in Requirements: A Multi-solution StudyabstractAmbiguity is a pervasive issue in natural-language requirements. A common source of ambiguity in requirements is when a pronoun is anaphoric. In requirements engineering, anaphoric ambiguity occurs when a pronoun can plausibly refer to different entities and thus be interpreted differently by different readers. In this paper, we develop an accurate and practical automated approach for handling anaphoric ambiguity in requirements, addressing both ambiguity detection and anaphora interpretation. In view of the multiple competing natural language processing (NLP) and machine learning (ML) technologies that one can utilize, we simultaneously pursue six alternative solutions, empirically assessing each using a collection of ≈1,350 industrial requirements. The alternative solution strategies that we consider are natural choices induced by the existing technologies; these choices frequently arise in other automation tasks involving natural-language requirements. A side-by-side empirical examination of these choices helps develop insights about the usefulness of different state-of-the-art NLP and ML technologies for addressing requirements engineering problems. For the ambiguity detection task, we observe that supervised ML outperforms both a large-scale language model, SpanBERT (a variant of BERT), as well as a solution assembled from off-the-shelf NLP coreference re-solvers. In contrast, for anaphora interpretation, SpanBERT yields the most accurate solution. In our evaluation, (1) the best solution for anaphoric ambiguity detection has an average precision of ≈60% and a recall of 100%, and (2) the best solution for anaphora interpretation (resolution) has an average success rate of ≈98%. Saad Ezzini, Sallam Abualhaija, Chetan Arora 0002, Mehrdad Sabetzadeh |
ICSE | 1 |
| 2022 | TAPHSIR: towards AnaPHoric ambiguity detection and ReSolution in requirementsabstractWe introduce TAPHSIR – a tool for anaphoric ambiguity detection and anaphora resolution in requirements. TAPHSIR facilities reviewing the use of pronouns in a requirements specification and revising those pronouns that can lead to misunderstandings during the development process. To this end, TAPHSIR detects the requirements which have potential anaphoric ambiguity and further attempts interpreting anaphora occurrences automatically. TAPHSIR employs a hybrid solution composed of an ambiguity detection solution based on machine learning and an anaphora resolution solution based on a variant of the BERT language model. Given a requirements specification, TAPHSIR decides for each pronoun occurrence in the specification whether the pronoun is ambiguous or unambiguous, and further provides an automatic interpretation for the pronoun. The output generated by TAPHSIR can be easily reviewed and validated by requirements engineers. TAPHSIR is publicly available on Zenodo (https://doi.org/10.5281/zenodo.5902117). Saad Ezzini, Sallam Abualhaija, Chetan Arora 0002, Mehrdad Sabetzadeh |
ESEC/SIGSOFT FSE | 1 |
| 2022 | WikiDoMiner: wikipedia domain-specific minerabstractWe introduce WikiDoMiner – a tool for automatically generating domain-specific corpora by crawling Wikipedia. WikiDoMiner helps requirements engineers create an external knowledge resource that is specific to the underlying domain of a given requirements specification (RS). Being able to build such a resource is important since domain-specific datasets are scarce. WikiDoMiner generates a corpus by first extracting a set of domain-specific keywords from a given RS, and then querying Wikipedia for these keywords. The output of WikiDoMiner is a set of Wikipedia articles relevant to the domain of the input RS. Mining Wikipedia for domain-specific knowledge can be beneficial for multiple requirements engineering tasks, e.g., ambiguity handling, requirements classification, and question answering. WikiDoMiner is publicly available on Zenodo under an open-source license (https: //doi.org/10.5281/zenodo.6672682) Saad Ezzini, Sallam Abualhaija, Mehrdad Sabetzadeh |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Using Domain-specific Corpora for Improved Handling of Ambiguity in RequirementsabstractAmbiguity in natural-language requirements is a pervasive issue that has been studied by the requirements engineering community for more than two decades. A fully manual approach for addressing ambiguity in requirements is tedious and time-consuming, and may further overlook unacknowledged ambiguity – the situation where different stakeholders perceive a requirement as unambiguous but, in reality, interpret the requirement differently. In this paper, we propose an automated approach that uses natural language processing for handling ambiguity in requirements. Our approach is based on the automatic generation of a domain-specific corpus from Wikipedia. Integrating domain knowledge, as we show in our evaluation, leads to a significant positive improvement in the accuracy of ambiguity detection and interpretation. We scope our work to coordination ambiguity (CA) and prepositional-phrase attachment ambiguity (PAA) because of the prevalence of these types of ambiguity in natural-language requirements [1]. We evaluate our approach on 20 industrial requirements documents. These documents collectively contain more than 5000 requirements from seven distinct application domains. Over this dataset, our approach detects CA and PAA with an average precision of 80% and an average recall of 89% (90% for cases of unacknowledged ambiguity). The automatic interpretations that our approach yields have an average accuracy of 85%. Compared to baselines that use generic corpora, our approach, which uses domain-specific corpora, has 33% better accuracy in ambiguity detection and 16% better accuracy in interpretation. Saad Ezzini, Sallam Abualhaija, Chetan Arora 0002, Mehrdad Sabetzadeh, Lionel C. Briand |
ICSE | 1 |