EDBT 2026 Demo / reviewers in the wild / expert
Jim Buckley
dblp:61/732
· DBLP profile ↗
46ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-6928-6746ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 42 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empirical pathways to developer experience: A facet-based synthesis of empirical designs and guidelinesabstractContext: Developer eXperience (Dev-X) focuses on ensuring a better experience for software developers, while still achieving development goals. Indeed, several existing studies suggest that improved Dev-X may actually result in improved development productivity. But a comprehensive understanding of how to empirically assess Dev-X remains limited. Aims and Method: This paper determines the empirical coverage of Dev-X facets to date by analyzing employed empirical designs and identifying quality issues or guidelines for assessing Dev-X. To this end, it presents a mapping study of 231 articles that employed empirical methods to explore the Dev-X landscape. Results: It identifies 16 facets of Dev-X —spanning emotions, values, and perceptions—and systematically maps the empirical methods, data types, and evaluation metrics used to assess these facets. The findings reveal gaps in facet coverage, highlight underutilized empirical designs, and stress the importance of quality assessment and methodological triangulation. This research offers a facet-wise synthesis of empirical designs and provides actionable guidelines to help researchers and practitioners empirically evaluate Dev-X. Conclusion: Based on SLR findings, this paper presents a Dev-X facet-based synthesis of empirical designs to guide selection of method, data, and metric combinations in future studies. By revealing design gaps (e.g., underexplored facets, overuse of surveys, weak corroboration) and highlighting dominant practices (e.g., survey-based ratings for emotions and values), our facet-wise schema offers an evidence-based guide for context-appropriate Dev-X evaluation strategies. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board . Goetz Botterweck, Qin Lai, Jim Buckley |
J. Syst. Softw. | 4 |
| 2025 | Industrial-Scale Neural Network Clone Detection with Disk-Based Similarity SearchabstractCode clones are similar code fragments that often arise from copy-and-paste programming. Neural networks can classify pairs of code fragments as clone/not-clone with high accuracy. However, finding clones in industrial-scale code needs a more scalable approach than pairwise comparison. We extend existing neural network-based clone detection schemes to handle codebases that far exceed available memory, using indexing and search methods for external storage such as disks and solid-state drives. We generate a high-dimensional vector embedding for each code fragment using a transformer-based neural network. We then find similar embeddings using efficient multidimensional nearest neighbor search algorithms on external storage to find similar embeddings without pairwise comparison. We identify specific problems with industrial-scale code bases, such as large sets of almost identical code fragments that interact poorly with k-nearest neighbour search algorithms, and provide an effective solution. We demonstrate that our disk-based clone search approach achieves similar clone detection accuracy as an equivalent in-memory technique. Using a solid-state drive as external storage, our approach is around 2 x slower than the in-memory approach for a problem size that can fit within memory. We further demonstrate that our approach can scale to over a billion lines of code, providing valuable insights into the trade-offs between indexing speed, query performance, and storage efficiency for industrial-scale code clone detection. Gul Aftab Ahmed, Muslim Chochlov, James Vincent Patten, Yuanhua Han, Guoxian Lu, Jim Buckley, David Gregg |
SANER | 7 |
| 2025 | A Framework and Taxonomy for Characterizing the Applicability of Software Architecture Recovery Approaches: A Tertiary-Mapping StudyabstractSummary Software architecture assists developers in addressing non‐functional requirements and in maintaining, debugging, and upgrading their software systems. Consequently, consistency between the designed architecture and the implemented software system itself is important; without this consistency the non‐functional requirements targeted may not be addressed and architectural documentation may mis‐direct maintenance efforts that target the associated code‐base. But often, when software is initially implemented or subsequently evolved, the designed architecture and software architecture become inconsistent, with the implemented structure degraded due to issues like developer time‐pressures, or ambiguous communication of the designed architecture. In such cases, Software Architecture Recovery (SAR) or consistency approaches can be applied to reconstruct the architecture of the software system and possibly to compare it to/re‐align it with the designed architecture. Many SAR approaches have been proposed in the research. However, choosing an appropriate architecture recovery approach for software systems is still an open issue. Consequently, this research aims to conduct a tertiary‐mapping study based on available secondary studies of architecture recovery approaches, to uncover important characteristics, towards the selection of appropriate SAR approaches. This research has aggregated 13 secondary studies and 10 primary studies beyond 2020 from 5 databases and, in doing so, identified 111 architecture recovery approaches. Based on these approaches, a taxonomy, containing nine main SAR‐selection categories is proposed and a framework (in the form of a supporting tool to help developers select an appropriate SAR approach) has been developed. Finally, this research identifies six potential open research gaps related to the underlying research that could be helpful for guiding research in the future. Abdul Qayum, Simon Colreavy-Donnelly, Muslim Chochlov, Jim Buckley, Dayi Lin, Ashish Rajendra Sai |
Softw. Pract. Exp. | 5 |
| 2024 | Nearest-neighbor, BERT-based, scalable clone detection: A practical approach for large-scale industrial code basesabstractAbstract Hidden code clones negatively impact software maintenance, but manually detecting them in large codebases is impractical. Additionally, automated approaches find detection of syntactically‐divergent clones very challenging. While recent deep neural networks (for example BERT‐based artificial neural networks) seem more effective in detecting such clones, their pairwise comparison of every code pair in the target system(s) is inefficient and scales poorly on large codebases. We present SSCD, a BERT‐based clone detection approach that targets high recall of Type 3 and Type 4 clones at a very large scale (in line with our industrial partner's requirements). It computes a representative embedding for each code fragment and finds similar fragments using a nearest neighbor search. Thus, SSCD avoids the pairwise‐comparison bottleneck of other neural network approaches, while also using a parallel, GPU‐accelerated search to tackle scalability. This article describes the approach, proposing and evaluating several refinements to improve Type 3/4 clone detection at scale. It provides a substantial empirical evaluation of the technique, including a speed/efficacy comparison of the approach against SourcererCC and Oreo, the only other neural‐network approach currently capable of scaling to hundreds of millions of LOC. It also includes a large in‐situ evaluation on our industrial collaborator's code base that assesses the original technique, the impact of the proposed refinements and illustrates the impact of incremental, active learning on its efficacy. We find that SSCD is significantly faster and more accurate than SourcererCC and Oreo. SAGA, a GPU‐accelerated traditional clone detection approach, is a little better than SSCD for T1/T2 clones, but substantially worse for T3/T4 clones. Thus, SSCD is both scalable to industrial code sizes, and comparatively more accurate than existing approaches for difficult T3/T4 clone searching. In‐situ evaluation on company datasets shows that SSCD outperforms the baseline approach (CCFinderX) for T3/T4 clones. Whitespace removal and active learning further improve SSCD effectiveness. Gul Aftab Ahmed, James Vincent Patten, Yuanhua Han, Guoxian Lu, David Gregg, Jim Buckley, Muslim Chochlov |
Softw. Pract. Exp. | 7 |
| 2023 | Decomposition of Monolith Applications Into Microservices Architectures: A Systematic ReviewabstractMicroservices architecture has gained significant traction, in part owing to its potential to deliver scalable, robust, agile, and failure-resilient software products. Consequently, many companies that use large and complex software systems are actively looking for automated solutions to decompose their monolith applications into microservices. This paper rigorously examines 35 research papers selected from well-known databases using a Systematic Literature Review (SLR) protocol and snowballing method, extracting data to answer the research questions, and presents the following four contributions. First, the Monolith to Microservices Decomposition Framework (M2MDF) which identifies the major phases and key elements of decomposition. Second, a detailed analysis of existing decomposition approaches, tools and methods. Third, we identify the metrics and datasets used to evaluate and validate monolith to microservice decomposition processes. Fourth, we propose areas for future research. Overall, the findings suggest that monolith decomposition into microservices remains at an early stage and there is an absence of methods for combining static, dynamic, and evolutionary data. Insufficient tool support is also in evidence. Furthermore, standardised metrics, datasets, and baselines have yet to be established. These findings can assist practitioners seeking to understand the various dimensions of monolith decomposition and the community's current capabilities in that endeavour. The findings are also of value to researchers looking to identify areas to further extend research in the monolith decomposition space. Yalemisew M. Abgaz, Andrew McCarren, Peter Elger, David Solan, Neil Lapuz, Marin Bivol, Glenn Jackson, Murat Yilmaz 0001, Jim Buckley, Paul M. Clarke |
IEEE Trans. Software Eng. | 9 |
| 2022 | Using a Nearest-Neighbour, BERT-Based Approach for Scalable Clone DetectionabstractCode clones can detrimentally impact software maintenance and manually detecting them in very large code-bases is impractical. Additionally, automated approaches find detection of Type 3 and Type 4 (inexact) clones very challenging. While the most recent artificial deep neural networks (for ex-ample BERT-based artificial neural networks) seem to be highly effective in detecting such clones, their pairwise comparison of every code pair in the target system(s) is inefficient and scales poorly on large codebases.We therefore introduce SSCD, a BERT-based clone detection approach that targets high recall of Type 3 and Type 4 clones at scale (in line with our industrial partner’s requirements). It does so by computing a representative embedding for each code fragment and finding similar fragments using a nearest neighbour search. SSCD thus avoids the pairwise-comparison bottleneck of other Neural Network approaches while also using parallel, GPU-accelerated search to tackle scalability.This paper details the approach and an empirical assessment towards configuring and evaluating that approach in industrial setting. The configuration analysis suggests that shorter input lengths and text-only based neural network models demonstrate better efficiency in SSCD, while only slightly decreasing effectiveness. The evaluation results suggest that SSCD is more effective than state-of-the-art approaches like SAGA and SourcererCC. It is also highly efficient: in its optimal setting, SSCD effectively locates clones in the entire 320 million LOC BigCloneBench (a standard clone detection benchmark) in just under three hours. Muslim Chochlov, Gul Aftab Ahmed, James Vincent Patten, Guoxian Lu, David Gregg, Jim Buckley |
ICSME | 7 |
| 2022 | The Effect of Feature Characteristics on the Performance of Feature Location TechniquesabstractFeature Location (FL)is a core software maintenance activity that aims to locate observable functionalities in the source code. Given its key role in software change, a vast array of Feature Location Techniques (FLTs) have been proposed but, as more and more FLTs are introduced, theselection of an appropriate FLTis an increasingly difficult problem. One consideration is thecharacteristics of the featuresbeing sought. For example, in the code associated with the feature, programmers may have named identifiers consistently, and with meaningful naming conventions, or not, and this may impact on the suitability of different FLTs. The suggestion that such characteristics matter has implicit support in the literature: An analysis of existing FLT empirical studies reveals that the system under study can often have a stronger impact on FLT performance than differing FLTs themselves. To understand this interaction between feature characteristics and FLTs better, this paper proposesa suite of feature-characteristic metricsthat are postulated to control FLTs’ performance, holistically across FLTs and impacting on individual FLTs to different degrees. To evaluate the suite, a controlled experiment is performed, using 878 features, to probe the relationship between the metrics and the performance of four FTL techniques: three commonly-used techniques and one state-of-the-art technique. The evaluation is performed using four commonly used evaluation measures and extended by employing 41 other established source-code metrics as extraneous variables. Results of the empirical evaluation suggest that the feature-metric suite presented impacts FLT performance holistically, and impacts different FLTs to different degrees. Thus, this paper moves towards the more standard selection of appropriate FLTs, with respect to the prominent feature characteristics in the software systems under study, and more rigorous consideration of the features selected to compare FLTs. Anthony Ventresque, Rainer Koschke, Andrea De Lucia, Jim Buckley |
IEEE Trans. Software Eng. | 5 |
| 2021 | Do Weibo Platform Experts Perform Better at Predicting Stock Market?
Ziyuan Ma, Conor Ryan, Jim Buckley, Muslim Chochlov |
EANN | 3 |
| 2021 | BoostNSift: A Query Boosting and Code Sifting Technique for Method Level Bug LocalizationabstractLocating bugs is an important, but effort-intensive and time-consuming task, when dealing with large-scale systems. To address this, Information Retrieval (IR) techniques are increasingly being used to suggest potential buggy source code locations, for given bug reports. While IR techniques are very scalable, in practice their effectiveness in accurately localizing bugs in a software system remains low. Results of empirical studies suggest that the effectiveness of bug localization techniques can be augmented by the configuration of queries used to locate buggy code. However, in most IR-based bug localization techniques, presented by researchers, the impact of the queries’ configurations is not fully considered. In a similar vein, techniques consider all code elements as equally suspicious of being buggy while localizing bugs, but this is not always the case either.In this paper, we present a new method-level, information-retrieval-based bug localization technique called "BoostNSift". BoostNSift exploits the important information in queries by ‘boost’ing that information, and then ‘sift’s the identified code elements, based on a novel technique that emphasizes the code elements’ specific relatedness to a bug report over its generic relatedness to all bug reports. To evaluate the performance of BoostNSift, we employed a state-of-the-art empirical design that has been commonly used for evaluating file level IR-based bug localization techniques: 6851 bugs are selected from commonly used Eclipse, AspectJ, SWT, and ZXing benchmarks and made openly available for method-level analyses. The performance of BoostNSift is compared with the openly-available state-of-the-art IR-based BugLocator, BLUiR, and BLIA techniques. Experiments show that BoostNSift improves on BLUiR by up to 324%, on BugLocator by up to 297%, and on BLIA up to 120%, in terms of Mean Reciprocal Rank (MRR). Similar improvements are observed in terms of Mean Average Precision (MAP) and Top-N evaluation measures. Jim Buckley, James Vincent Patten, Muslim Chochlov, Ashish Rajendra Sai |
SCAM | 2 |
| 2021 | Correction to: An empirical assessment of baseline feature location techniquesabstractA Correction to this paper has been published: 10.1007/s10664-020-09924-6 Andrew Le Gear, Christopher Exton, Jim Buckley |
Empir. Softw. Eng. | 4 |
| 2021 | Correction to: Characterizing the transfer of program comprehension in onboarding: an information-push perspectiveabstractA Correction to this paper has been published: 10.1007/s10664-020-09923-7 Rebecca Yates, Norah Power, Jim Buckley |
Empir. Softw. Eng. | 3 |
| 2021 | Taxonomy of centralization in public blockchain systems: A systematic literature reviewabstractBitcoin introduced delegation of control over a monetary system from a select few to all who participate in that system. This delegation is known as the decentralization of controlling power and is a powerful security mechanism for the ecosystem. After the introduction of Bitcoin, the field of cryptocurrency has seen widespread attention from industry and academia, so much so that the original novel contribution of Bitcoin, i.e., decentralization, may be overlooked, due to decentralizations’ assumed fundamental existence for the functioning of such crypto-assets. However, recent studies have observed a trend of increased centralization in cryptocurrencies such as Bitcoin and Ethereum. As this increased centralization has an impact the security of the blockchain, it is crucial that it is measured, towards adequate control. This research derives an initial taxonomy of centralization present in decentralized blockchains through rigorous synthesis using a systematic literature review. This is followed by iterative refinement through expert interviews. We systematically analyzed 89 research papers published between 2009 and 2019. Our study contributes to the existing body of knowledge by highlighting the multiple definitions and measurements of centralization in the literature. We identify different aspects of centralization and propose an encompassing taxonomy of centralization concerns. This taxonomy is based on empirically observable and measurable characteristics. It consists of 13 aspects of centralization, classified over six architectural layers: Governance, Network, Consensus, Incentive, Operational, and Application. We also discuss how the implications of centralization can vary depending on the aspects studied. We believe that this review and taxonomy provides a comprehensive overview of centralization in decentralized blockchains involving various conceptualizations and measures. Ashish Rajendra Sai, Jim Buckley, Brian Fitzgerald 0001, Andrew Le Gear |
Inf. Process. Manag. | 2 |
| 2020 | Inheritance software metrics on smart contractsabstractBlockchain systems have gained substantial traction recently, partly due to the potential of decentralized immutable mediation of economic activities. Ethereum is a prominent example that has the provision for executing stateful computing scripts known as Smart Contracts. These smart contracts resemble traditional programs, but with immutability being the core differentiating factor. Given their immutability and potential high monetary value, it becomes imperative to develop high-quality smart contracts. Software metrics have traditionally been an essential tool in determining programming quality. Given the similarity between smart contracts (written in Solidity for Ethereum) and object-oriented (OO) programming, OO metrics would appear applicable. In this paper, we empirically evaluate inheritance-based metrics as applied to smart contracts. We adopt this focus because, traditionally, inheritance has been linked to a more complex codebase which we posit is not the case with Solidity based smart contracts. In this work, we evaluate the hypothesis that, due to the differences in the context of smart contracts and OO programs, it may not be appropriate to use the same interpretation of inheritance based metrics for assessment. Ashish Rajendra Sai, Conor Holmes, Jim Buckley, Andrew Le Gear |
ICPC | 3 |
| 2020 | An empirical assessment of baseline feature location techniquesabstractFeature Location (FL) aims to locate observable functionalities in source code. Considering its key role in software maintenance, a vast array of automated and semi-automated Feature Location Techniques (FLTs) have been proposed. To compare FLTs, an open, standard set of non-subjective, reproducible "compare-to" FLT techniques (baseline techniques) should be used for evaluation. In order to relate the performance of FLTs compared against different baseline techniques, these compare-to techniques should be evaluated against each other. But evaluation across FLTs is confounded by empirical designs that incorporate different FL goals and evaluation criteria. This paper moves towards standardizing FLT comparability by assessing eight baseline techniques in an empirical design that addresses these confounding factors. These baseline techniques are assessed in twelve case studies to rank their performance. Results of the case studies suggest that different baseline techniques perform differently and that VSM-Lucene and LSI-Matlab performed better than other implementations. By presenting the relative performances of baseline techniques this paper facilitates empirical cross-comparison of existing and future FLTs. Finally, the results suggest that the performance of FLTs partially depends on system/benchmark characteristics, in addition to the FLTs themselves. Andrew Le Gear, Christopher Exton, Jim Buckley |
Empir. Softw. Eng. | 4 |
| 2020 | Characterizing the transfer of program comprehension in onboarding: an information-push perspective
Rebecca Yates, Norah Power, Jim Buckley |
Empir. Softw. Eng. | 3 |
| 2019 | Identifying Feature Clones: An Industrial Case StudyabstractDuring its software evolution, the original software system of our industrial partner was split into three variants. These have evolved over time, but retained a lot of common functionality. During strategical planning our industrial partner realized the need for consolidation of common code in a shared code base towards more efficient code maintenance and re-use. To support this agenda, a feature-clone identification approach was proposed, combining elements of feature location (to identify the relevant code in one system) and clone detection (to identify that common feature's code across systems) techniques. In this work, this approach is used (via our prototype tool CoRA) to locate three features that were identified by the industrial partner for re-factoring, and is evaluated. The methodology, involving a system expert, was designed to evaluate the discrete parts of the approach in isolation: textual and static analyses of feature location, and clone detection. It was found that the approach can effectively identify features and their clones. The hybrid textual/static feature location part is effective even for a relative system novice, showing results comparable to more optimal system expert's suggestions. Finally, more effective feature location increases the effectiveness of the clone detection part of the approach.11A preliminary version of this paper, explaining the motivation, approach and resultant tool was published in [1]. This paper extends that work with a discussion of the approach's in-vivo empirical evaluation. Muslim Chochlov, Michael English, Jim Buckley, Daniel Ilie, Maria Scanlon |
SANER | 3 |
| 2019 | Assessing the security implication of Bitcoin exchange rates
Ashish Rajendra Sai, Jim Buckley, Andrew Le Gear |
Comput. Secur. | 2 |
| 2019 | The State of Empirical Evaluation in Static Feature LocationabstractFeature location (FL) is the task of finding the source code that implements a specific, user-observable functionality in a software system. It plays a key role in many software maintenance tasks and a wide variety of Feature Location Techniques (FLTs), which rely on source code structure or textual analysis, have been proposed by researchers. As FLTs evolve and more novel FLTs are introduced, it is important to perform comparison studies to investigate “Which are the best FLTs?” However, an initial reading of the literature suggests that performing such comparisons would be an arduous process, based on the large number of techniques to be compared, the heterogeneous nature of the empirical designs, and the lack of transparency in the literature. This article presents a systematic review of 170 FLT articles, published between the years 2000 and 2015. Results of the systematic review indicate that 95% of the articles studied are directed towards novelty, in that they propose a novel FLT. Sixty-nine percent of these novel FLTs are evaluated through standard empirical methods but, of those, only 9% use baseline technique(s) in their evaluations to allow cross comparison with other techniques. The heterogeneity of empirical evaluation is also clearly apparent: altogether, over 60 different FLT evaluation metrics are used across the 170 articles, 272 subject systems have been used, and 235 different benchmarks employed. The review also identifies numerous user input formats as contributing to the heterogeneity. Analysis of the existing research also suggests that only 27% of the FLTs presented might be reproduced from the published material. These findings suggest that comparison across the existing body of FLT evaluations is very difficult. We conclude by providing guidelines for empirical evaluation of FLTs that may ultimately help to standardise empirical research in the field, cognisant of FLTs with different goals, leveraging common practices in existing empirical evaluations and allied with rationalisations. This is seen as a step towards standardising evaluation in the field, thus facilitating comparison across FLTs. Asanka Wasala, Christopher Exton, Jim Buckley |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2018 | [Engineering Paper] Identifying Feature Clones in a Suite of SystemsabstractAs part of a module re-unification project of an industrial partner's code, spanning one systems and two derivative systems, the feature-clone variants across these systems have to be extracted, to be later re-unified as singular code elements for re-use. To assist developers with this task, the CoRA (The Code Re-unification Application) tool was designed and implemented. An approach, and the subsequent design of the tool was derived from reflection on manual feature-location/clonedetection efforts on the company's systems, in the first phase of an action research cycle where the approach/implementation will be iteratively trialled, and subsequently refined, in-situ. A pilot study is discussed that leads to the proposed tool. The tool combines a hybrid (textual-static) feature location technique and a textual clone detection technique for featureclone identification. In this paper, the rationale behind the CoRA tool is presented, followed by a tool overview and its implementation details. Finally, an example use case shows how the tool is used to locate clones of a particular feature. Muslim Chochlov, Michael English, Jim Buckley, Daniel Ilie, Maria Scanlon |
SCAM | 3 |
| 2018 | Architecture consistency: State of the practice, challenges and requirementsabstractArchitecture Consistency (AC) aims to align implemented systems with their intended architectures. Several AC approaches and tools have been proposed and empirically evaluated, suggesting favourable results. In this paper, we empirically examine the state of practice with respect to Architecture Consistency, through interviews with nineteen experienced software engineers. Our goal is to identify 1) any practises that the companies these architects work for, currently undertake to achieve AC; 2) any barriers to undertaking explicit AC approaches in these companies; 3) software development situations where practitioners perceive AC approaches would be useful, and 4) AC tool needs, as perceived by practitioners. We also assess current commercial AC tool offerings in terms of these perceived needs. The study reveals that many practitioners apply informal AC approaches as there are barriers for adopting more formal and explicit approaches. These barriers are: 1) Difficulty in quantifying architectural inconsistency effects, and thus justifying the allocation of resources to fix them to senior management, 2) The near invisibility of architectural inconsistency to customers, 3) Practitioners’ reluctance towards fixing architectural inconsistencies, and 4) Practitioners perception that huge effort is required to map the system to the architecture when using more formal AC approaches and tools. Practitioners still believe that AC would be useful in supporting several of the software development activities such as auditing, evolution and ensuring quality attributes. After reviewing several commercial tools, we posit that AC tool vendors need to work on their ability to support analysis of systems made up of different technologies, that AC tools need to enhance their capabilities with respect to artefacts such as services and meta-data, and to focus more on non-maintainability architectural concerns. Nour Ali, Seán Baker, Ross O'Crowley, Sebastian Herold, Jim Buckley |
Empir. Softw. Eng. | 5 |
| 2018 | Erratum to: Architecture consistency: State of the practice, challenges and requirementsabstractThe article Architecture consistency: State of the practice, challenges and requirements, written by Nour Ali et al. was originally published electronically on the publisher’s internet portal on May 15, 2017 without open access. The original article has been corrected. Nour Ali, Seán Baker, Ross O'Crowley, Sebastian Herold, Jim Buckley |
Empir. Softw. Eng. | 5 |
| 2017 | A historical, textual analysis approach to feature location
Muslim Chochlov, Michael English, Jim Buckley |
Inf. Softw. Technol. | 3 |
| 2017 | An empirical study of data decomposition for software parallelization
Anne Meade, Deva Kumar Deeptimahanti, Jim Buckley, J. J. Collins |
J. Syst. Softw. | 3 |
| 2016 | Empirically Derived Recommendations for Personalised Text-Based Technical Support
Solomon Gizaw, Jim Buckley, Sarah Beecham |
SPICE | 2 |
| 2015 | Manually locating features in industrial source code: the search actions of software nomadsabstractExpert software engineers working on large systems often need to perform feature location when moving to work in unfamiliar areas. We hypothesise that leveraging the system-specific knowledge of these may help to improve semi-automated feature location techniques. In order to assess and understand how software nomads perform manual feature location searches, two expert professional software engineers were observed in-vivo following a think-aloud protocol while performing manual feature location on a large-scale heterogeneous system. The nomads' search actions were found to be around twice as effective as those reported in previous studies. This cannot be explained by sophisticated use of tools or complex queries. We conclude that system rules and conventions are frequently used by experts when constructing feature location search terms. Howell R. Jordan, Jacek Rosik, Sebastian Herold, Goetz Botterweck, Jim Buckley |
ICPC | 5 |
| 2015 | Characterising Users Through an Analysis of On-line Technical Support Forums
Solomon Gizaw, Jim Buckley, Sarah Beecham |
PROFES | 2 |
| 2015 | Using changeset descriptions as a data source to assist feature locationabstractFeature location attempts to assist developers in discovering functionality in source code. Many textual feature location techniques utilize information retrieval and rely on comments and identifiers of source code to describe software entities. An interesting alternative would be to employ the changeset descriptions of the code altered in that changeset as a data source to describe such software entities. To investigate this we implement a technique utilizing changeset descriptions and conduct an empirical study to observe this technique's overall performance. Moreover, we study how the granularity (i.e. file or method level of software entities) and changeset range inclusion (i.e. most recent or all historical changesets) affect such an approach. The results of a preliminary study with Rhino and Mylyn. Tasks systems suggest that the approach could lead to a potentially efficient feature location technique. They also suggest that it is advantageous in terms of the effort to configure the technique at method level granularity and that older changesets from older systems may reduce the effectiveness of the technique. Muslim Chochlov, Michael English, Jim Buckley |
SCAM | 3 |
| 2015 | Evaluating Pair-Programming for Non-Computer Science Major StudentsabstractThe ubiquitous nature of software has resulted in many non-computer-science (NCS) major students taking courses in computer programming. The appeal of learning computer programming for this cohort may be diminished, given that they do not usually have an initial desire to become expert programmers. This, coupled with the difficulty of learning computer programming, means that efforts to heighten their engagement with/learning of, programming skills are required. Clem O'Donnell, Jim Buckley, Abdulhussain E. Mahdi, John Nelson, Michael English |
SIGCSE | 2 |
| 2015 | Detection of violation causes in reflexion modelsabstractReflexion Modelling is a well-understood technique to detect architectural violations that occur during software architecture erosion. Resolving these violations can be difficult when erosion has reached a critical level and the causes of the violations are interwoven and difficult to understand. This article outlines a novel technique to automatically detect typical causes of violations in reflexion models, based on the definition and detection of typical symptoms for these causes. Preliminary results show that the proposed technique can support software architects' navigation through reflexion models of eroded systems to understand causes of violations and to systematically take actions against them. Sebastian Herold, Michael English, Jim Buckley, Steve Counsell, Mel Ó Cinnéide |
SANER | 3 |
| 2015 | Real-Time Reflexion Modelling in architecture reconciliation: A multi case study
Jim Buckley, Nour Ali, Michael English, Jacek Rosik, Sebastian Herold |
Inf. Softw. Technol. | 1 |
| 2015 | An empirically-based characterization and quantification of information seeking through mailing lists during Open Source developers' software evolution
Khaironi Yatim Sharif, Michael English, Nour Ali, Christopher Exton, J. J. Collins, Jim Buckley |
Inf. Softw. Technol. | 6 |
| 2013 | JITTAC: a just-in-time tool for architectural consistencyabstractArchitectural drift is a widely cited problem in software engineering, where the implementation of a software system diverges from the designed architecture over time causing architecture inconsistencies. Previous work suggests that this architectural drift is, in part, due to programmers' lack of architecture awareness as they develop code. JITTAC is a tool that uses a real-time Reflexion Modeling approach to inform programmers of the architectural consequences of their programming actions as, and often just before, they perform them. Thus, it provides developers with Just-In-Time architectural awareness towards promoting consistency between the as-designed architecture and the as-implemented system. JITTAC also allows programmers to give real-time feedback on introduced inconsistencies to the architect. This facilitates programmer-driven architectural change, when validated by the architect, and allows for more timely team-awareness of the actual architectural consistency of the system. Thus, it is anticipated that the tool will decrease architectural inconsistency over time and improve both developers' and architect's knowledge of their software's architecture. The JITTAC demo is available at: http://www.youtube.com/watch?v=BNqhp40PDD4. Jim Buckley, Sean Mooney, Jacek Rosik, Nour Ali |
ICSE | 1 |
| 2012 | Construct specific coupling measurement for C++ software
Michael English, Tony Cahill, Jim Buckley |
Comput. Lang. Syst. Struct. | 3 |
| 2011 | Assessing architectural drift in commercial software development: a case studyabstractAbstract Objectives: Software architecture is perceived as one of the most important artefacts created during a system's design. However, implementations often diverge from their intended architectures: a phenomenon called architectural drift. The objective of this research is to assess the occurrence of architectural drift in the context ofde novosoftware development, to characterize it, and to evaluate whether its detection leads to inconsistency removal.Method: Anin vivo, longitudinal case study was performed during the development of a commercial software system, where an approach based on Reflexion Modelling was employed to detect architectural drift. Observation and think‐aloud data, captured during the system's development, were assessed for the presence and types of architectural drift. When divergences were identified, the data were further analysed to see if identification led to the removal of these divergences.Results: The analysed system diverged from the intended architecture, during the initial implementation of the system. Surprisingly however, this work showed that Reflexion Modelling served to conceal some of the inconsistencies, a finding that directly contradicts the high regard that this technique enjoys as an architectural evaluation tool. Finally, the analysis illustrated that detection of inconsistencies was insufficient to prompt their removal, in the small, informal team context studied.Conclusions: Although the utility of the approach for detecting inconsistencies was demonstrated in most cases, it also served to hide several inconsistencies and did not act as a trigger for their removal. Hence additional efforts must be taken to lessen architectural drift and several improvements in this regard are suggested. Copyright © 2010 John Wiley & Sons, Ltd. Jacek Rosik, Andrew Le Gear, Jim Buckley, Muhammad Ali Babar 0001, Dave Connolly |
Softw. Pract. Exp. | 3 |
| 2010 | A replicated and refined empirical study of the use of friends in C++ software
Michael English, Jim Buckley, Tony Cahill |
J. Syst. Softw. | 2 |
| 2009 | An in-vivo study of the cognitive levels employed by programmers during software maintenanceabstractSeveral researchers have proposed Bloom's taxonomy as a framework within which to study the cognitive levels employed by programmers during software comprehension. But a review of empirical studies in this area illustrates that previous work has nearly exclusively focused on the lower cognitive levels of the taxonomy. However, the taxonomy was initially proposed as a dasiacumulative hierarchypsila, where less processing occurred at higher levels. This suggests that the focus of current software comprehension literature is appropriate. Given that there is mixed empirical evidence for this dasiacumulative hierarchypsila property, this work reports on the cognitive levels employed by 6 programmers, involved in in-vivo software maintenance and comprehension. It suggests that the cumulative hierarchy property is true of these contexts, thus adding legitimacy to the focus of the existing literature. However, it notes that processing at the higher cognitive levels does occur and is associated with specific maintenance sub-tasks. As this processing is effort and skill intensive, there is still a need for researchers to explore these higher cognitive levels. Tara Kelly, Jim Buckley |
ICPC | 2 |
| 2009 | Observation of Open Source programmers' information seekingabstractSeveral authors have proposed information seeking as an appropriate perspective for studying software evolution, and have characterized information seeking empirically in commercial software evolution settings. However, there is little research in the literature describing the information seeking behavior of open source programmers. This work describes a holistic schema for open-source (OS) programmers' information seeking, generated through open-coding of questions in OS developer mailing lists. It then reports on a study of the JDT mailing list, showing the types of information sought within this group and, for a subset of the data, it characterizes the responses obtained. Khaironi Yatim Sharif, Jim Buckley |
ICPC | 2 |
| 2009 | An empirical analysis of information retrieval based concept location techniques in software comprehension
Brendan Cleary, Christopher Exton, Jim Buckley, Michael English |
Empir. Softw. Eng. | 3 |
| 2008 | An industrial case study of architecture conformanceabstractA software designer often has little control over, or means of checking, whether his design is being adhered to, once the implementation begins. This 'architectural drift', where the original design of the system and the as-implemented design of the system diverge, can cause serious problems for evolution, maintenance and the comprehensibility of a system if it remains undocumented or uncorrected. The earlier such discrepancies can be identified, the better. Jacek Rosik, Andrew Le Gear, Jim Buckley, Muhammad Ali Babar 0001 |
ESEM | 3 |
| 2008 | Encapsulating targeted component abstractions using software Reflexion ModellingabstractAbstract Design abstractions such as components, modules, subsystems or packages are often not made explicit in the implementation of legacy systems. Indeed, often the abstractions that are made explicit turn out to be inappropriate for future evolution agendas. This can make the maintenance, evolution and refactoring of these systems difficult. In this publication, we carry out a fine‐grained evaluation of Reflexion Modelling as a technique for encapsulating user‐targeted components. This process is a prelude to component recovery, reuse and refactoring. The evaluation takes the form of twoin vivocase studies, where two professional software developers encapsulate components in a large, commercial software system. The studies demonstrate the validity of this approach and offer several best‐use guidelines. Specifically, they argue that users benefit from having a strong mental model of the system in advance of Reflexion Modelling, even if that model is flawed, and that users should expend effort exploring the expected relationships present in Reflexion Models. Copyright © 2008 John Wiley & Sons, Ltd. Jim Buckley, Andrew Le Gear, Christopher Exton, Ross Cadogan, Trevor Johnston, Bill Looby, Rainer Koschke |
J. Softw. Maintenance Res. Pract. | 1 |
| 2007 | Fine-Grained Software Metrics in PracticeabstractModularity is one of the key features of the Object- Oriented (00) paradigm. Low coupling and high cohesion help to achieve good modularity. Inheritance is one of the core concepts of the 00 paradigm which facilitates modularity. Previous research has shown that the use of the friend construct as a coupling mechanism in C+ + software is extensive. However, measures of the friend construct are scarse in comparison with measures of inheritance. In addition, these existing measures are coarse-grained, in spite of the widespread use of the friend mechanism. In this paper, a set of software metrics are proposed that measure the actual use of the friend construct, inheritance and other forms of coupling. These metrics are based on the interactions for which each coupling mechanism is necessary and sufficient. Previous work only considered the declaration of a relationship between classes. The software metrics introduced are empirically assessed using the LEDA software system. Our results indicate that the friend mechanism is used to a very limited extent to access hidden methods in classes. However, access to hidden attributes is more common. Michael English, Jim Buckley, Tony Cahill |
ESEM | 2 |
| 2006 | A Context-Aware Analysis Scheme for Bloom's TaxonomyabstractA large body of empirical work in the software comprehension area has focused on the cognitive processes that programmers undertake. However, as yet, little work exists on developing and assessing an encompassing framework within which one can compare comprehension studies with one another. Several authors have proposed that Bloom's taxonomy could provide such a framework and a lexical-analysis schema has been trialled to classify the data from empirical studies into this taxonomy. The schema is simple to apply but may result in ambiguity and reductionism. This paper proposes an alternative context-aware analysis schema. While such a schema undoubtedly consumes more effort, its value is illustrated by means of a pilot study, where its application is compared to that of the lexical-analysis schema Tara Kelly, Jim Buckley |
ICPC | 2 |
| 2005 | Empirically Studying Software Practitioners - Bridging the Gap between Theory and PracticeabstractIt is the view of many computer scientists that the standard of empirical software engineering research leaves scope for improvement. However, there is also an increasing awareness in the software engineering community that empirical studies are a vital aspect in the process of improving methods and tools, for software development and maintenance. This paper presents a review of the empirical work carried out to date in the area of program comprehension and illustrates that most of the evidence from these studies derives from lab-based experiments, thus implying a degree of artificial control. The paper argues that, in order to address the methodological shortfalls of the experimental paradigm, more qualitative methods need to be applied to accompany and support these quantitative studies, thus broadening the sources of data and increasing the 'body of evidence'. Michael P. O'Brien, Jim Buckley, Christopher Exton |
ICSM | 2 |
| 2005 | Reengineering towards components using "Reconn-exion"abstractContinuing to develop software from scratch will not be feasible indefinitely. Reusing existing software would seem to be a viable solution to this problem. The paradigm of component-based development (CBD) explicitly accounts for reuse in its process. Unfortunately the majority of existing software systems are not implemented using CBD, thus reusing portions of this software using CBD becomes difficult. Reengineering and maintenance research contains a plethora of software analysis and restructuring techniques that could be used to help us exploit legacy applications for reuse. This thesis focuses on two such techniques and combines variations of them for the purpose of component recovery: A feature location technique called Software Reconnaissance and a design recovery technique called Software Reflexion Modelling. Their combination is called "Component Reconn-exion." We describe the technique, highlight results and evaluation to date and finally discuss further work necessary to complete our contribution as a PhD. thesis. Andrew Le Gear, Jim Buckley |
ESEC/SIGSOFT FSE | 2 |
| 2005 | Towards a taxonomy of software changeabstractAbstract Previous taxonomies of software change have focused on the purpose of the change (i.e., the why) rather than the underlying mechanisms. This paper proposes a taxonomy of software change based on characterizing the mechanisms of change and the factors that influence these mechanisms. The ultimate goal of this taxonomy is to provide a framework that positions concrete tools, formalisms and methods within the domain of software evolution. Such a framework would considerably ease comparison between the various mechanisms of change. It would also allow practitioners to identify and evaluate the relevant tools, methods and formalisms for a particular change scenario. As an initial step towards this taxonomy, the paper presents a framework that can be used to characterize software change support tools and to identify the factors that impact on the use of these tools. The framework is evaluated by applying it to three different change support tools and by comparing these tools based on this analysis. Copyright © 2005 John Wiley & Sons, Ltd. Jim Buckley, Tom Mens, Matthias Zenger, Awais Rashid, Günter Kniesel-Wünsche |
J. Softw. Maintenance Res. Pract. | 1 |
| 2004 | Expectation-based, inference-based, and bottom-up software comprehensionabstractAbstract The software comprehension process has been conceptualized as being either ‘top‐down’ or ‘bottom‐up’ in nature. We formally distinguish between two comprehension processes that have previously been grouped together as ‘top‐down’. The first is ‘expectation‐based’ comprehension, where the programmer has pre‐generated expectations of the code's meaning. The second is ‘inference‐based’ comprehension, where the programmer derives meaning from clichéd implementations in the code. We identify the distinguishing features of the two variants, and use these characteristics as the basis for an empirical study. This study establishes the existence of the above‐mentioned processes, in conjunction with ‘bottom‐up’ comprehension. It also illustrates the relationship between these processes and programmers' application domain familiarity. Copyright © 2004 John Wiley & Sons, Ltd. Michael P. O'Brien, Jim Buckley, Teresa M. Shaft |
J. Softw. Maintenance Res. Pract. | 2 |