VLDB 2026 Research / reviewers in the wild / expert
Ranim Khojah
dblp:331/4737
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-1090-3153ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emotional Strain and Frustration in LLM Interactions in Software EngineeringabstractLarge Language Models (LLMs) are increasingly integrated into various daily tasks in Software Engineering, such as coding and requirement elicitation. Despite their various capabilities and constant use, some interactions can lead to unexpected challenges (e.g. hallucinations or verbose answers) and, in turn, cause emotions that develop into frustration. Frustration can negatively impact engineers’ productivity and well-being if it escalates into stress and burnout. In this paper, we assess the impact of LLM interactions on software engineers’ emotional responses, specifically strains, and identify common causes of frustration when interacting with LLMs at work. Based on 62 survey responses from software engineers in industry and academia across various companies and universities, we found that a majority of our respondents experience frustrations or other related emotions regardless of the nature of their work. Additionally, our results showed that frustration mainly stemmed from issues with correctness and less critical issues, such as adaptability to context or specific format. While such issues may not cause frustration in general, artefacts that do not follow certain preferences, standards, or best practices can make the output unusable without extensive modification, causing frustration over time. In addition to the frustration triggers, our study offers guidelines to improve the software engineers’ experience, aiming to minimise long-term consequences on mental health. Cristina Martinez Montes, Ranim Khojah |
EASE | 2 |
| 2025 | The Impact of Prompt Programming on Function-Level Code GenerationabstractLarge Language Models (LLMs) are increasingly used by software engineers for code generation. However, limitations of LLMs such as irrelevant or incorrect code have highlighted the need for prompt programming (or prompt engineering) where engineers apply specific prompt techniques (e.g., chain-of-thought or input-output examples) to improve the generated code. While some prompt techniques have been studied, the impact of different techniques—and their interactions— on code generation is still not fully understood. In this study, we introduce CodePromptEval, a dataset of 7072 prompts designed to evaluate five prompt techniques (few-shot, persona, chain-of-thought, function signature, list of packages) and their effect on the correctness, similarity, and quality of complete functions generated by three LLMs (GPT-4o, Llama3, and Mistral). Our findings show that while certain prompt techniques significantly influence the generated code, combining multiple techniques does not necessarily improve the outcome. Additionally, we observed a trade-off between correctness and quality when using prompt techniques. Our dataset and replication package enable future research on improving LLM-generated code and evaluating new prompt techniques. Ranim Khojah, Francisco Gomes de Oliveira Neto, Mazen Mohamad, Philipp Leitner 0001 |
IEEE Trans. Software Eng. | 1 |
| 2023 | Evaluating the Trade-offs of Text-based Diversity in Test PrioritisationabstractDiversity-based techniques (DBT) have been cost-effective by prioritizing the most dissimilar test cases to detect faults at earlier stages of test execution. Diversity is measured on test specifications to convey how different test cases are from one another. However, there is little research on the trade-off of diversity measures based on different types of text-based specification (lexicographical or semantics). Particularly because the text content in test scripts vary widely from unit (e.g., code) to system-level (e.g., natural language). This paper compares and evaluates the cost-effectiveness in coverage and failures of different text-based diversity measures for different levels of tests. We perform an experiment on the test suites of 7 open source projects on the unit level, and 2 industry projects on the integration and system level. Our results show that test suites prioritised using semantic-based diversity measures causes a small improvement in requirements coverage, as opposed to lexical diversity that showed less coverage than random for system-level artefacts. In contrast, using lexical-based measures such as Jaccard or Levenshtein to prioritise code artefacts yield better failure coverage across all levels of tests. We summarise our findings into a list of recommendations for using semantic or lexical diversity on different levels of testing. Ranim Khojah, Chi Hong Chao, Francisco Gomes de Oliveira Neto |
AST | 1 |
| 2022 | Evaluating N-best Calibration of Natural Language Understanding for Dialogue SystemsabstractA Natural Language Understanding (NLU) component can be used in a dialogue system to perform intent classification, returning an N -best list of hypotheses with corresponding confidence estimates.We perform an in-depth evaluation of 5 NLUs, focusing on confidence estimation.We measure and visualize calibration for the 10 best hypotheses on model level and rank level, and also measure classification performance.The results indicate a trade-off between calibration and performance.In particular, Rasa (with Sklearn classifier) had the best calibration but the lowest performance scores, while Watson Assistant had the best performance but a poor calibration. Ranim Khojah, Alexander Berman, Staffan Larsson |
SIGDIAL | 1 |