Farshad Kazemi

dblp:337/2915 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0001-4683-3484ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Interrogative Comments Posed by Review Comment Generators: An Empirical Study of Gerrit
abstract
Background: Review Comment Generators (RCGs) are models trained to automate code review tasks. Prior work shows that RCGs can generate review comments to initiate discussion threads; however, their ability to interact with author responses is unclear. This can be especially problematic if RCGs pose interrogative comments, i.e., comments that ask questions of other review participants. Aims: We set out to study the prevalence of RCG-generated interrogative code review comments, their similarity with the interrogative comments of humans, and the predictability of the generation of interrogative comments. Method: We study three task-specific RCGs and three RCGs based on Large Language Models (LLMs) on data from the Gerrit project using quantitative and qualitative methods. Results: We find that RCGs: (1) generate interrogative comments at a rate of$\mathbf{1 5. 6 \%} \boldsymbol{-} \mathbf{6 5. 2 6 \%}$; (2) differ from humans in generating such comments, which can stifle conversations if RCGs dissuade human reviewers from commenting deeply; and (3) produce interrogative comments with low predictability. Finally, we find that (4) the interrogative comments posed by LLMbased RCGs can differ even more substantially from human behaviour than those of task-specific RCGs. For example, the studied LLM-based RCGs pose rhetorical questions 3.16% of the time, whereas human-submitted interrogative comments pose rhetorical questions 8.74 % of the time. Conclusions: Our results suggest that neither task-specific nor LLM-based RCGs can replace human reviewers yet; however, we note opportunities for synergies. For example, RCGs tend to raise pertinent questions about exception handling of common APIs more frequently than human reviewers. Putting greater emphasis on technical comments generated by RCGs (rather than conversational ones, such as interrogative ones) will likely improve their perceived usefulness.
Farshad Kazemi, Maxime Lamothe, Shane McIntosh
ESEM1
2024 Reevaluating the Defect Proneness of Atoms of Confusion in Java Systems
abstract
Background:Code confusion concerns source code characteristics that make code harder for authors and reviewers to comprehend. Atoms of Confusions (AoCss) are a set of low-level programming idioms for C-like languages that have been proposed as a potential source of code confusion; previous studies have empirically evaluated the extent to which they (i) are confusing to developers and (ii) introduce risk to software products.
Guoshuai Shi, Farshad Kazemi, Michael W. Godfrey, Shane McIntosh
ESEM2
2024 Characterizing the Prevalence, Distribution, and Duration of Stale Reviewer Recommendations
abstract
The appropriate assignment of reviewers is a key factor in determining the value that organizations can derive from code review. While inappropriate reviewer recommendations can hinder the benefits of the code review process, identifying these assignments is challenging. Stale reviewers, i.e., those who no longer contribute to the project, are one type of reviewer recommendation that is certainly inappropriate. Understanding and minimizing this type of recommendation can thus enhance the benefits of the code review process. While recent work demonstrates the existence of stale reviewers, to the best of our knowledge, attempts have yet to be made to characterize and mitigate them. In this paper, we study the prevalence and potential effects. We then propose and assess a strategy to mitigate stale recommendations in existing code reviewer recommendation tools. By applying five code reviewer recommendation approaches (LearnRec, RetentionRec, cHRev, Sofia, and WLRRec) to three thriving open-source systems with 5,806 contributors, we observe that, on average, 12.59% of incorrect recommendations are stale due to developer turnover; however, fewer stale recommendations are made when the recency of contributions is considered by the recommendation objective function. We also investigate which reviewers appear in stale recommendations and observe that the top reviewers account for a considerable proportion of stale recommendations. For instance, in 15.31% of cases, the top-3 reviewers account for at least half of the stale recommendations. Finally, we study how long stale reviewers linger after the candidate leaves the project, observing that contributors who left the project 7.7 years ago are still suggested to review change sets. Based on our findings, we propose separating the reviewer contribution recency from the other factors that are used by the CRR objective function to filter out developers who have not contributed during a specified duration. By evaluating this strategy with different intervals, we assess the potential impact of this choice on the recommended reviewers. The proposed filter reduces the staleness of recommendations, i.e., the Staleness Reduction Ratio (SRR) improves between 21.44%–92.39%. Yet since the strategy may increase active reviewer workload, careful project-specific exploration of the impact of the cut-off setting is crucial.
Farshad Kazemi, Maxime Lamothe, Shane McIntosh
IEEE Trans. Software Eng.1
2022 Exploring the Notion of Risk in Code Reviewer Recommendation
abstract
Reviewing code changes allows stakeholders to improve the premise, content, and structure of changes prior to or after integration. However, assigning reviewing tasks to team members is challenging, particularly in large projects. Code reviewer recommendation has been proposed to assist with this challenge. Traditionally, the performance of reviewer recommenders has been derived based on historical data, where better solutions are those that recommend exactly which reviewers actually performed tasks in the past. More recent work expands the goals of recommenders to include mitigating turnover-based knowledge loss and avoiding overburdening the core development team. In this paper, we set out to explore how reviewer recommendation can incorporate the risk of defect proneness. To this end, we propose the Changeset Safety Ratio (CSR) – an evaluation measurement designed to capture the risk of defect proneness. Through an empirical study of three open source projects, we observe that: (1) existing approaches tend to improve one or two quantities of interest, such as core developers workload while degrading others (especially the CSR); (2) Risk Aware Recommender (RAR) – our proposed enhancement to multi-objective reviewer recommendation – achieves a 12.48% increase in expertise of review assignees and a 80% increase in CSR with respect to historical assignees, all while reducing the files at risk of knowledge loss by 19.39% and imposing a negligible 0.93% increase in workload for the core team; and (3) our dynamic method outperforms static and normalization-based tuning methods in adapting RAR to suit risk-averse and balanced risk usage scenarios to a significant degree (Conover's test, α < 0.05; small to large Kendall's W).
Farshad Kazemi, Maxime Lamothe, Shane McIntosh
ICSME1