Mariam El Mezouar

dblp:195/6788 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-3317-7051ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Bots Get the Boot: Understanding Pull Request Rejections in the Era of AI Coders
abstract
Agent-authored pull requests (APRs) generated by autonomous coding agents are rejected more frequently than human-authored contributions. We study APR rejection using early and lifetime pull-request-level signals and additional annotations derived from the AIDev dataset. We augment AIDev with an APR Signal Set capturing code size, complexity, design indicators, social dynamics, and author experience, and observe that rejected APRs exhibit higher interaction volume, greater structural complexity, and increased code churn. We further label a uniformly sampled subset of 250 rejected APRs using rejection reasons adapted from an established taxonomy, finding that most rejections fall into a small number of recurring reasons, with only 10% remaining unclassified. Finally, we demonstrate how the augmented dataset can support automated analysis through a lightweight LLM-based classification example that assigns rejection reasons to rejected APRs.
Karla Gonzalez, Mariam El Mezouar
MSR2
2026 Is this build failure related to my patch? An empirical study of unrelated build failures in continuous integration
abstract
Abstract In a hectic Continuous Integration (CI) environment, where several builds are triggered concurrently, legitimate build failures (e.g., not caused by flaky tests) may not always be related to the current push. These unrelated build failures can burden developers as they devote hours to attest whether errors are truly associated with their present changes. In this paper, we extract 77,354 CI build failures from 7 open source projects to understand and identify unrelated build failures. We attempt to provide an indication for developers about whether a build failure is likely to be related to the current push or not. Our results reveal that developers likely invest a median of 4 hours to determine whether a build failure is (un)related to their pushes. We perform a document analysis on a sample of 371 unrelated build failures (based on the 95% confidence level and 5% confidence interval from 10,316 potentially unrelated failures) to understand why build failures are deemed as unrelated by developers. The themes generated from our document analysis reveal that unrelated tests failures represent 20% of the cases of why build failures are deemed unrelated by developers. To predict whether a build failure is unrelated to the current push, we extract 33 features from issue reports, issue comments, and from the commits pertaining to the triggering push. We build semi-supervised PU-learning models over seven Apache projects and achieve precision ranging from $$0.70 \pm 0.01$$ to $$0.88 \pm 0.02$$ , recall ranging from $$0.30 \pm 0.03$$ to $$1.00 \pm 0.00$$ , and F1-scores ranging from $$0.44 \pm 0.03$$ to $$0.91 \pm 0.00$$ , while the area under the ROC curve (AUC) spans $$0.63 \pm 0.02$$ to $$0.97 \pm 0.03$$ . Our analysis of feature importance reveals that (i) the time taken from a submitted patch to the build-triggering push (CI latency), (ii) build failures sharing similar error messages with recent failures, and (iii) the number of comments preceding the build failure, are all efficient indicators for identifying potential unrelated build failures. The semi-supervised approach proposed in this work can help developers identify build failures that are unrelated to their current push, providing actionable guidance such as re-running builds, inspecting infrastructure logs, or prioritizing code-level debugging based on prediction outcomes.
Yonghui Andie Huang, Daniel Alencar da Costa, Grant Dick, Mariam El Mezouar, Liwen Xiao
Empir. Softw. Eng.4
2026 Correction to: Is this build failure related to my patch? An empirical study of unrelated build failures in continuous integration
Yonghui Andie Huang, Daniel Alencar da Costa, Grant Dick, Mariam El Mezouar, Liwen Xiao
Empir. Softw. Eng.4
2026 Towards Refining Developer Questions Using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
abstract
In software engineering chatrooms, communication is often hindered by imprecise questions that cannot be answered. Recognizing key entities (e.g., programming languages and libraries) and user intent (e.g., learning or requesting a review) can be essential for improving question clarity and facilitating better exchange. However, existing research using natural language processing techniques often overlooks these softwarespecific nuances. In this paper, we introduceSoftwarE-specificNamed entity recognition,Intent detection, andResolution classification (SENIR), a labelling approach that leverages a Large Language Model to annotate entities, intents, and resolution status in developer chatroom conversations. To offer quantitative guidance for improving question clarity and resolvability, we build a resolution prediction model that leverages SENIR’s entity and intent labels along with additional predictive features. We evaluate SENIR on the DISCO dataset using a subset of annotated chatroom dialogues. SENIR achieves an 86% F-score for entity recognition, a 71% F-score for intent detection, and an 89% F-score for resolution status classification. Furthermore, our resolution prediction model, tested with various sampling strategies (random undersampling and oversampling with SMOTE) and evaluation methods (5-fold cross-validation, 10-fold cross-validation, and bootstrapping), demonstrates AUC values ranging from 0.7 to 0.8. Key factors influencing resolution include positive sentiment and entities such asProgramming LanguageandUser Variableacross multiple intents, while diagnostic entities (e.g.,Error Name) are more relevant in error-related questions. Moreover, resolution rates vary significantly by intent: questions aboutAPI UsageandAPI Changeachieve higher resolution rates, whereasDiscrepancyandReviewhave lower resolution rates. A Chi-Square analysis confirms the statistical significance of these differences.
Pouya Fathollahzadeh, Mariam El Mezouar, Hao Li 0094, Ying Zou 0001, Ahmed E. Hassan
IEEE Trans. Software Eng.2
2022 Exploring the Use of Chatrooms by Developers: An Empirical Study on Slack and Gitter
abstract
Communication is critical for the software development teams to maintain project awareness, facilitate project co-ordination and avoid misunderstandings. The features offered in the chatrooms, such as private messaging, group conversations, and code sharing help accommodate the communication needs of the software development teams. Therefore, chatrooms have been increasingly adopted among the developers. Since the last study on Slack performed by (Linet al.2016), the audience of Slack has more than doubled possibly leading to an evolution of the ways Slack is used; while another rich community formed around Gitter and remains unstudied. In this paper, we perform an investigative study using qualitative and quantitative techniques to gain insights on the use of popular modern chatrooms, specifically Slack and Gitter. Based on the survey responses from 163 developers, the interviews with 21 developers, and the chatroom data collected from 11 Slack and 770 Gitter rooms, we are able to uncover the reasons behind the use of Slack and Gitter, the perceived impact on the associated projects, and the quality determinants of the two chatrooms. We find that the developers seek knowledge from the chatrooms to obtain timely feedback from experts, and in return share their expertise to build the project community and their reputations. Furthermore, it is perceived by the Gitter developers that the chatrooms have an impact on prioritizing the new features and the bug fixes. In Slack, the most reported impact concerns an increased project awareness, in terms of a better tracking of the work progress. As reported on the developers’ survey, both Slack and Gitter chat services have a visible impact on mentoring developers, and sharing the best practices. In terms of quality determinants, a non-ephemeral history and a better history management (e.g., advanced search) could be keys for both chat services to reach their full potential.
Mariam El Mezouar, Daniel Alencar da Costa, Daniel M. Germán, Ying Zou 0001
IEEE Trans. Software Eng.1
2021 An Empirical Study of Developer Discussions in the Gitter Platform
abstract
Developer chatrooms (e.g., the Gitter platform) are gaining popularity as a communication channel among developers. In developer chatrooms, a developer ( asker ) posts questions and other developers ( respondents ) respond to the posted questions. The interaction between askers and respondents results in a discussion thread . Recent studies show that developers use chatrooms to inquire about issues, discuss development ideas, and help each other. However, prior work focuses mainly on analyzing individual messages of a chatroom without analyzing the discussion thread in a chatroom. Developer chatroom discussions are context-sensitive, entangled, and include multiple participants that make it hard to accurately identify threads. Therefore, prior work has limited capability to show the interactions among developers within a chatroom by analyzing only individual messages. In this article, we perform an in-depth analysis of the Gitter platform (i.e., developer chatrooms) by analyzing 6,605,248 messages of 709 chatrooms. To analyze the characteristics of the posted questions and the impact on the response behavior (e.g., whether the posted questions get responses), we propose an approach that identifies discussion threads in chatrooms with high precision (i.e., 0.81 F-score). Our results show that inactive members responded more often and unique questions take longer discussion time than simple questions. We also find that clear and concise questions are more likely to be responded to than poorly written questions. We further manually analyze a randomly selected sample of 384 threads to examine how respondents resolve the raised questions. We observe that more than 80% of the studied threads are resolved. Advanced-level/beginner-level questions along with the edited questions are the mostly resolved questions. Our results can help the project maintainers understand the nature of the discussion threads (e.g., the topic trends). Project maintainers can also benefit from our thread identification approach to spot the common repeated threads and use these threads as frequently asked questions (FAQs) to improve the documentation of their projects.
Osama Ehsan, Safwat Hassan, Mariam El Mezouar, Ying Zou 0001
ACM Trans. Softw. Eng. Methodol.3
2019 An empirical study on the teams structures in social coding using GitHub projects
Mariam El Mezouar, Feng Zhang 0001, Ying Zou 0001
Empir. Softw. Eng.1
2018 Are tweets useful in the bug fixing process? An empirical study on Firefox and Chrome
Mariam El Mezouar, Feng Zhang 0001, Ying Zou 0001
Empir. Softw. Eng.1