VLDB 2026 Research / reviewers in the wild / expert
Mahir Arzoky
dblp:120/5542
· DBLP profile ↗
16ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An audit of machine learning experiments on software defect predictionabstractMachine learning algorithms are increasingly being proposed to solve the problem of predicting defect-prone software components. In this literature, computational experiments are the primary means of evaluating and comparing learners and the credibility of findings depends critically on their experimental design and reporting. This paper audits recent software defect prediction (SDP) experiments by assessing their experimental design, analysis and reporting practices against widely accepted norms from statistics, machine learning and empirical software engineering. Our aim is to characterise the current state of practice and evaluate the reproducibility of published findings. We undertook an audit of relevant studies published from the SCOPUS database (2019-2023) focusing on their experimental design and analysis choices e.g., the outcome variables such as F-measure and the type of out of sample (OOS) validation regime, e.g., cross-validation, plus the statistical analysis and inference mechanisms. In all, we evaluated nine different study issues. This was complemented by an assessment of reproducibility using the instrument proposed by González-Barahona and Robles. Our search located approximately 1,585 experiments in SDP (2019-2023), a substantial body of work. From this, we randomly sampled 101 ( $$ \approx 6.4\%$$ ) papers, 61 journal and 40 conference papers. Almost 50% are behind ‘paywalls’. We found considerable divergence in research practice. The number of datasets used ranged 1-365, the number of learners or learner variants evaluated from 1-34 and the number of performance metrics from 1 to 9. Approximately 45% of papers made use of formal statistical inference. We detected a total of 427 issues distributed across 101 papers (median=4) with only one paper being entirely issue-free. In terms of reproducibility, experiments ranged from near perfect to lacking almost all required information. We also found two examples of tortured phrases and potential “paper mill” activity. Approaches to designing and reporting computational experiments varied greatly, but almost half the studies provided insufficient information such that reproduction would be challenging. Overall, our audit suggests that as a research community, we have considerable scope for improvement. Fortunately, many improvements should be neither difficult nor costly to achieve. Giuseppe Destefanis, Leila Yousefi, Martin J. Shepperd, Allan Tucker, Stephen Swift, Steve Counsell, Mahir Arzoky |
Empir. Softw. Eng. | 7 |
| 2026 | ETM-F: an enriched topic modeling and filtration framework integrating ontologies and deep learning for biomedical trend analysis
Ahmad Altarawneh, Mahir Arzoky, Stephen Swift, Reem Qadan Al Fayez, Moh'd Belal Al-Zoubi, Bilal Sowan |
Knowl. Inf. Syst. | 2 |
| 2024 | Different Strokes for Different Folks: A Comparison of Developer and Tester Views on TestingabstractIn this paper, we re-analyse the data from a previous study of Straubinger et al., which asked 284 industrial IT staff about their views on testing. In that study, as well as developers, the dedicated role of tester was included in the data - both roles were treated as the same role. In this paper, we posit that the two roles (i.e., developer and tester) are so very different that we should analyse each role separately; testers will have unique insights into testing, so separating their views and experiences from developers is important. To this end, we analyse six of the same research questions as the original study, using separate developer and tester data. Results showed that for almost every question we re-visited, testers differed in their opinions from developers, whether on the type of testing they did, measures of code quality, effort to write tests and motivation for testing. Steve Counsell, Stephen Swift, Mahir Arzoky, Tracy Hall, Emily Winter 0001, Gareth Bennett, Thomas Shippey |
SEAA | 3 |
| 2022 | Estimating the Optimal Number of Clusters from Subsets of EnsemblesabstractThis research estimates the optimal number of clusters in a dataset using a novel ensemble technique - a preferred alternative to relying on the output of a single clustering. Combining clusterings from different algorithms can lead to a more stable and robust solution, often unattainable by any single clustering solution. Technically, we created subsets of ensembles as possible estimates; and evaluated them using a quality metric to obtain the best subset. We tested our method on publicly available datasets of varying types, sources and clustering difficulty to establish the accuracy and performance of our approach against eight standard methods. Our method outperforms all the techniques in the number of clusters estimated correctly. Due to the exhaustive nature of the initial algorithm, it is slow as the number of ensembles or the solution space increases; hence, we have provided an updated version based on the single-digit difference of Gray code that runs in linear time in terms of the subset size. Afees Adegoke Odebode, Allan Tucker, Mahir Arzoky, Stephen Swift |
DATA | 3 |
| 2021 | Opening the black box: Personalizing type 2 diabetes patients based on their latent phenotype and temporal associated complication rulesabstractAbstract It is widely considered that approximately 10% of the population suffers from type 2 diabetes. Unfortunately, the impact of this disease is underestimated. Patient's mortality often occurs due to complications caused by the disease and not the disease itself. Many techniques utilized in modeling diseases are often in the form of a “black box” where the internal workings and complexities are extremely difficult to understand, both from practitioners' and patients' perspective. In this work, we address this issue and present an informative model/pattern, known as a “latent phenotype,” with an aim to capture the complexities of the associated complications' over time. We further extend this idea by using a combination of temporal association rule mining and unsupervised learning in order to find explainable subgroups of patients with more personalized prediction. Our extensive findings show how uncovering the latent phenotype aids in distinguishing the disparities among subgroups of patients based on their complications patterns. We gain insight into how best to enhance the prediction performance and reduce bias in the models applied using uncertainty in the patients' data. Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
Comput. Intell. | 3 |
| 2020 | Using the Lexicon from Source Code to Determine Application DomainabstractContext: The vast majority of software engineering research is reported independently of the application domain: techniques and tools usage is reported without any domain context. As reported in previous research, this has not always been so: early in the computing era, the research focus was frequently application domain specific (for example, scientific and data processing). Andrea Capiluppi, Nemitari Ajienka, Nour Ali, Mahir Arzoky, Steve Counsell, Giuseppe Destefanis, Alina Dana Miron, Bhaveet Nagaria, Rumyana Neykova, Martin J. Shepperd, Stephen Swift, Allan Tucker |
EASE | 4 |
| 2020 | On Clones and Comments in Production and Test Classes: An Empirical Study
Steve Counsell, Steve Swift, Mahir Arzoky, Giuseppe Destefanis |
PROFES | 3 |
| 2020 | On the Link Between Refactoring Activity and Class Cohesion Through the Prism of Two Cohesion-Based MetricsabstractThe practice of refactoring has evolved over the past thirty years to become standard developer practice; for almost the same amount of time, proposals for measuring object-oriented cohesion have also been suggested. Yet, we still know very little about their inter-relationship empirically, despite the fact that classes exhibiting low cohesion would be strong candidates for refactoring. In this paper, we use a large set of refactorings to understand the characteristics of two cohesion metrics from a refactoring perspective. Firstly, through the well-known LCOM metric of Chidamber and Kemerer and, secondly, the C3 metric proposed more recently by Marcus et al. Our research question is motivated by the premise that different refactorings will be applied to classes with low cohesion compared with those applied to classes with high cohesion. We used three open-source systems as a basis of our analysis and on data from the lower and upper quartiles of metric data. Results showed that the set of refactoring types across both upper and lower quartiles was broadly the same, although very different in actual numbers. The `rename method' refactoring stood out from the rest, being applied over three times as often to classes with low cohesion than to classes with high cohesion. Steve Counsell, Giuseppe Destefanis, Steve Swift, Mahir Arzoky, Davide Taibi 0001 |
QRS | 4 |
| 2020 | Deep convolutional neural network designed for age assessment based on orthopantomography data
Seyed M. M. Kahaki, Md. Jan Nordin, Nazatul S. Ahmad, Mahir Arzoky, Waidah Ismail |
Neural Comput. Appl. | 4 |
| 2019 | An Empirical Study of the AGIS Visual Field Metric and Its Seasonal VariationsabstractThe severity of the glaucoma eye disease is usually measured by the Advanced Glaucoma Intervention Studies (AGIS) metric. The metric provides a value between zero and twenty inclusive, where the former represents no evidence of glaucoma and the latter the most advanced of measurements. In a previous study by Montolio et al., the season in which the test was undertaken was shown to affect the value of eye measurements; the lowest sensitivity was found in Summer and the highest sensitivity found in Winter and Spring. In this paper, we partially replicate that study with a different set of data from 2468 patients obtained from Moorfields Eye Hospital, London. We decomposed the data according to the four seasonal dates to determine if extra sensitivity meant that patients' results improved in Winter. Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis |
CBMS | 3 |
| 2019 | Opening the Black Box: Exploring Temporal Pattern of Type 2 Diabetes Complications in Patient Clustering Using Association Rules and Hidden Variable DiscoveryabstractThere is a great deal of debate over the importance of explanation in AI models inferred from health data. In particular, there is a balance that needs to be made between the accuracy of complex 'deep' models such as convolutional neural networks and the transparency of models that aim to model data in a more 'human' way such as expert systems. In this paper, we explore the use of temporal association rules to validate and uncover the meaning behind discrete hidden variables that have been inferred from clinical diabetes data. We use a recently published technique based upon the IC* (Induction Causation) algorithm that limits the number of hidden variables and places them within a network structure. Here, we take the hidden variables and compare their underlying discrete states to clusters that have been generated from temporal association rules. This allows us to characterise the hidden states based upon different sequences of complications. Results are very promising, with many hidden states aligning with the discovered clusters giving us a direct interpretation. Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
CBMS | 3 |
| 2019 | On the Relationship Between Coupling and Refactoring: An Empirical ViewpointabstractBackground: Refactoring has matured over the past twenty years to become part of a developer's toolkit. However, many fundamental research questions still remain largely unexplored. Aim: The goal of this paper is to investigate the highest and lowest quartile of refactoring-based data using two coupling metrics - the Coupling between Objects metric and the more recent Conceptual Coupling between Classes metric to answer this question. Can refactoring trends and patterns be identified based on the level of class coupling? Method: In this paper, we analyze over six thousand refactoring operations drawn from releases of three open-source systems to address one such question. Results: Results showed no meaningful difference in the types of refactoring applied across either lower or upper quartile of coupling for both metrics; refactorings usually associated with coupling removal were actually more numerous in the lower quartile in some cases. A lack of inheritance-related refactorings across all systems was also noted. Conclusions: The emerging message (and a perplexing one) is that developers seem to be largely indifferent to classes with high coupling when it comes to refactoring types - they treat classes with relatively low coupling in almost the same way. Steve Counsell, Mahir Arzoky, Giuseppe Destefanis, Davide Taibi 0001 |
ESEM | 2 |
| 2019 | The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi |
IDEAL (1) | 4 |
| 2018 | Opening the Black Box: Discovering and Explaining Hidden Variables in Type 2 Diabetic Patient Modelling
Leila Yousefi, Stephen Swift, Mahir Arzoky, Lucia Sacchi, Luca Chiovato, Allan Tucker |
BIBM | 3 |
| 2018 | Do Developers Really Worry About Refactoring Re-test? An Empirical Study of Open-Source Systems
Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis |
PROFES | 3 |
| 2014 | An Approach to Controlling the Runtime for Search Based Modularisation of Sequential Source Code Check-ins
Mahir Arzoky, Stephen Swift, Steve Counsell, James Cain 0002 |
IDA | 1 |