VLDB 2026 Research / reviewers in the wild / expert
Nicolas E. Gold
dblp:g/NicolasGold · also Nicolas Edwin Gold, Nicolas Gold
· DBLP profile ↗
39ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0002-2195-5995ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 34 · 10 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Movement Sonification of Familiar Music to Support the Agency of People with Chronic PainabstractFFAME (Filtering Familiar Audio for Movement Exploration) is a novel sonification framework aiming to facilitate movement in individuals with chronic back pain. Our personalised, music-based approach contrasts and extends prior work with predetermined tonal sonification. FFAME progressively filters selected music based on angles of the trunk. Through a qualitative analysis of reported experience of 15 participants with chronic pain and 5 physiotherapists, we identify how sonification parameters and musical characteristics affect movement and meaning-making. Music-based movement sonification proved impactful across multiple dimensions: (1) encouraging movement, (2) escaping pain-related rumination, (3) externalizing pain experiences, and (4) scaffolding physical activities. Drawing on enactivism and related philosophies, the study highlights how the semantic indeterminacy of music, combined with real-time movement sonification, created a rich, open-ended environment that supported user agency and exploration. Sonification for pain management can be creative and expressive, enabling people with pain to extend challenging movements and build movement confidence. Kyrill Potapov, Nicolas E. Gold, Temitayo A. Olugbade, Amanda C. de C. Williams, Christopher Dieter Overbeck, Danielle Lynch, Minna Orvokki Nygren, Nadia Bianchi-Berthouze |
CHI | 2 |
| 2025 | Causal program dependence analysisabstractDiscovering how program components affect one another plays a fundamental role in aiding engineers comprehend and maintain a software system. Despite the fact that the degree to which one program component depends upon another can vary in strength, traditional dependence analysis typically ignores such nuance. To account for this nuance in dependence-based analysis, we propose Causal Program Dependence Analysis (CPDA), a framework based on causal inference that captures the degree (or strength) of the dependence between program elements. For a given program, CPDA intervenes in the program execution to observe changes in value at selected points in the source code. It observes the association between program elements by constructing and executing modified versions of a program (requiring only light-weight parsing rather than sophisticated static analysis). CPDA applies causal inference to the observed changes to identify and estimate the strength of the dependence relations between program elements. We explore the advantages of CPDA's quantified dependence by presenting results for several applications. Our further qualitative evaluation demonstrates 1) that observing different levels of dependence facilitates grouping various functional aspects found in a program and 2) how focusing on the relative strength of the dependences for a particular program element provides a detailed context for that element. Furthermore, a case study that applies CPDA to debugging illustrates how it can improve engineer productivity. Seongmin Lee 0001, Dave W. Binkley, Robert Feldt, Nicolas E. Gold, Shin Yoo |
Sci. Comput. Program. | 4 |
| 2025 | A Programming Platform Design for Children to Create and Explore Hybrid Musical InstrumentsabstractABSTRACT Engaging young learners in coding activities is increasingly important to prepare them to participate in a contemporary computationally centred society. Engagement comes from finding appropriate motivation and the use of materials and methods that are suitable to the educational stage and learning opportunities. In this article, we present our experiences of designing and developing a Python‐based platform for children to use in our Music Making with Music and Making (MMMM) workshops, enabling school students to engage with hybrid digital music instrument design through programming synthesisers and undertaking LEGO construction. We document our understanding of the design and operation of the BrickPi3 and Raspberry Pi Build Hat systems and discuss our approach to using these as part of a simple software platform for schoolchildren to use. We present our design (and discuss alternatives considered) and the way it was successfully used by 61 11–13 year‐old children in workshop settings. We discuss the technical issues encountered and their resolution, and finally reflect on the positive feedback provided by teaching staff and potential for further development. The inherently integrated and inter‐disciplinary nature of our work aligns with and makes a positive contribution to the Science Technology Engineering Arts and Mathematics (STEAM) movement. Nicolas E. Gold, Evangelos Himonides, Ross Purves, Hazel Baxter |
Softw. Pract. Exp. | 1 |
| 2025 | The EmoPain@Home Dataset: Capturing Pain Level and Activity Recognition for People With Chronic Pain in Their HomesabstractChronic pain is a prevalent condition where fear of movement and pain interfere with everyday functioning. Yet, there is no open body movement dataset for people with chronic pain in everyday settings. Our EmoPain@Home dataset addresses this with capture from 18 people with and without chronic pain in their homes, while they performed their routine activities. The data includes labels for pain, worry, and movement confidence continuously recorded for activity instances for the people with chronic pain. We explored baseline two-level pain detection based on this dataset and obtained 0.62 mean F1 score. However, extension of the dataset led to deterioration in performance confirming high variability in pain expressions for real world settings. We investigated baseline activity recognition for this setting as a first step in exploring the use of the activity label as contextual information for improving pain level classification performance. We obtained mean F1 score of 0.43 for 9 activity types, highlighting its feasibility. Further exploration, however, showed that data from healthy people cannot be easily leveraged for improving performance because worry and low confidence alter activity strategies for people with chronic pain. Our dataset and findings lay critical groundwork for automatic assessment of pain experience and behaviour in the wild. Temitayo A. Olugbade, Raffaele Andrea Buono, Kyrill Potapov, Alex Bujorianu, Amanda C. de C. Williams, Santiago de Ossorno Garcia, Nicolas E. Gold, Catherine Holloway, Nadia Bianchi-Berthouze |
IEEE Trans. Affect. Comput. | 7 |
| 2024 | Movement Representation Learning for Pain Level ClassificationabstractSelf-supervised learning has shown value for uncovering informative movement features for human activity recognition. However, there has been minimal exploration of this approach for affect recognition where availability of large labelled datasets is particularly limited. In this paper, we propose a P-STEMR (Parallel Space-Time Encoding Movement Representation) architecture with the aim of addressing this gap and specifically leveraging the higher availability of human activity recognition datasets for pain-level classification. We evaluated and analyzed the architecture using three different datasets across four sets of experiments. We found statistically significant increase in average F1 score to 0.84 for pain level classification with two classes based on the architecture compared with the use of hand-crafted features. This suggests that it is capable of learning movement representations and transferring these from activity recognition based on data captured in lab settings to classification of pain levels with messier real-world data. We further found that the efficacy of transfer between datasets can be undermined by dissimilarities in population groups due to impairments that affect movement behaviour and in motion primitives (e.g. rotation versus flexion). Future work should investigate how the effect of these differences could be minimized so that data from healthy people can be more valuable for transfer learning. Temitayo A. Olugbade, Amanda C. de C. Williams, Nicolas E. Gold, Nadia Bianchi-Berthouze |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | EmoPain(at)Home: Dataset and Automatic Assessment within Functional Activity for Chronic Pain RehabilitationabstractWhile there is growing interest in developing tech-nology to support pain assessment, pain-related self-management, and healthcare personalisation, there are currently no datasets on nonverbal pain behaviour in the context of functional activities. To address this gap, we introduce the EmoPain(at)Home dataset which consists of motion capture data and self-reported pain, worry, and confidence intensities captured from people with chronic pain. The data were recorded during self-selected functional activities in the home, e.g. vacuuming. We include analysis of the dataset as well as baseline classification of pain levels with average F1 score of 0.61 for two classes. We additionally discuss inclusivity considerations for capture of datasets in naturalistic settings, based on lessons learnt within our study. Temitayo A. Olugbade, Raffaele Andrea Buono, Amanda C. de C. Williams, Santiago de Ossorno Garcia, Nicolas E. Gold, Catherine Holloway, Nadia Bianchi-Berthouze |
ACII | 5 |
| 2022 | Ethics in the mining of software repositoriesabstractAbstract Research in Mining Software Repositories (MSR) is research involving human subjects, as the repositories usually contain data about developers’ and users’ interactions with the repositories and with each other. The ethics issues raised by such research therefore need to be considered before beginning. This paper presents a discussion of ethics issues that can arise in MSR research, using the mining challenges from the years 2006 to 2021 as a case study to identify the kinds of data used. On the basis of contemporary research ethics frameworks we discuss ethics challenges that may be encountered in creating and using repositories and associated datasets. We also report some results from a small community survey of approaches to ethics in MSR research. In addition, we present four case studies illustrating typical ethics issues one encounters in projects and how ethics considerations can shape projects before they commence. Based on our experience, we present some guidelines and practices that can help in considering potential ethics issues and reducing risks. Nicolas E. Gold, Jens Krinke |
Empir. Softw. Eng. | 1 |
| 2021 | Observation-based approximate dependency modeling and its use for program slicing
Seongmin Lee 0001, Dave W. Binkley, Robert Feldt, Nicolas E. Gold, Shin Yoo |
J. Syst. Softw. | 4 |
| 2020 | Ethical Mining: A Case Study on MSR Mining ChallengesabstractResearch in Mining Software Repositories (MSR) is research involving human subjects, as the repositories usually contain data about developers' interactions with the repositories. Therefore, any research in the area needs to consider the ethics implications of the intended activity before starting. This paper presents a discussion of the ethics implications of MSR research, using the mining challenges from the years 2010 to 2019 as a case study to identify the kinds of data used. It highlights problems that one may encounter in creating such datasets, and discusses ethics challenges that may be encountered when using existing datasets, based on a contemporary research ethics framework. We suggest that the MSR community should increase awareness of ethics issues by openly discussing ethics considerations in published articles. Nicolas E. Gold, Jens Krinke |
MSR | 1 |
| 2020 | Evaluating lexical approximation of program dependence
Seongmin Lee 0001, Dave W. Binkley, Nicolas E. Gold, Syed S. Islam, Jens Krinke, Shin Yoo |
J. Syst. Softw. | 3 |
| 2019 | MOAD: Modeling Observation-Based Approximate DependencyabstractWhile dependency analysis is foundational to many applications of program analysis, the static nature of many existing techniques presents challenges such as limited scalability and inability to cope with multi-lingual systems. We present a novel dependency analysis technique that aims to approximate program dependency from a relatively small number of perturbed executions. Our technique, called MOAD (Modeling Observation-based Approximate Dependency), reformulates program dependency as the likelihood that one program element is dependent on another, instead of a more classical Boolean relationship. MOAD generates a set of program variants by deleting parts of the source code, and executes them while observing the impacts of the deletions on various program points. From these observations, MOAD infers a model of program dependency that captures the dependency relationship between the modification and observation points. While MOAD is a purely dynamic dependency analysis technique similar to Observation Based Slicing (ORBS), it does not require iterative deletions. Rather, MOAD makes a much smaller number of multiple, independent observations in parallel and infers dependency relationships for multiple program elements simultaneously, significantly reducing the cost of dynamic dependency analysis. We evaluate MOAD by instantiating program slices from the obtained probabilistic dependency model. Compared to ORBS, MOAD's model construction requires only 18.7% of the observations used by ORBS, while its slices are only 16% larger than the corresponding ORBS slice, on average. Seongmin Lee 0001, Dave W. Binkley, Robert Feldt, Nicolas E. Gold, Shin Yoo |
SCAM | 4 |
| 2019 | A comparison of tree- and line-oriented observational slicing
Dave W. Binkley, Nicolas E. Gold, Syed S. Islam, Jens Krinke, Shin Yoo |
Empir. Softw. Eng. | 2 |
| 2017 | Tree-Oriented vs. Line-Oriented Observation-Based SlicingabstractObservation-based slicing is a recently-introduced, language-independent slicing technique based on the dependencies observable from program behavior.The original algorithm processed traditional source code at the line-of-text level.A recent variation was developed to slice the tree-based XML representation of executable models.We ported the model slicer to source code using srcML to construct a tree-based representation of traditional source code.We present the results of a comparison of the two slicers using four experiments involving seventeen different programs, including classic benchmarks and larger production systems.The resulting slices had essentially the same size and quite often the same content.Where they differ, the use of tree structure traded an ability to remove unnecessary parts of a statement for the requirement of maintaining aspect of the code structure.Comparing the slicers finds that each has its advantages.For example, when the tree representation facilitates the deletion of large chunks of code, the tree slicer was over eight times faster.In contrast, when slicing C++ code it was over nine times slower because of the multitude of small trees created to support C++ syntax.Given the pros and cons of the two, the results suggest the value of their hybrid combination. Dave W. Binkley, Nicolas E. Gold, Syed S. Islam, Jens Krinke, Shin Yoo |
SCAM | 2 |
| 2017 | Generalized observational slicing for tree-represented modelling languagesabstractModel-driven software engineering raises the abstraction level making complex systems easier to understand than if written in textual code. Nevertheless, large complicated software systems can have large models, motivating the need for slicing techniques that reduce the size of a model. We present a generalization of observation-based slicing that allows the criterion to be defined using a variety of kinds of observable behavior and does not require any complex dependence analysis. We apply our implementation of generalized observational slicing for tree-structured representations to Simulink models. The resulting slice might be the subset of the original model responsible for an observed failure or simply the sub-model semantically related to a classic slicing criterion. Unlike its predecessors, the algorithm is also capable of slicing embedded Stateflow state machines. A study of nine real-world models drawn from four different application domains demonstrates the effectiveness of our approach at dramatically reducing Simulink model sizes for realistic observation scenarios: for 9 out of 20 cases, the resulting model has fewer than 25% of the original model's elements. Nicolas E. Gold, Dave W. Binkley, Mark Harman, Syed S. Islam, Jens Krinke, Shin Yoo |
ESEC/SIGSOFT FSE | 1 |
| 2016 | Musically Informed Sonification for Chronic Pain Rehabilitation: Facilitating Progress & Avoiding Over-DoingabstractIn self-directed chronic pain physical rehabilitation it is important that the individual can progress as physical capabilities and confidence grow. However, people with chronic pain often struggle to pass what they have identified as safe boundaries. At the same time, over-activity due to the desire to progress fast or function more normally, may lead to setbacks. We investigate how musically-informed movement sonification can be used as an implicit mechanism to both avoid overdoing and facilitate progress during stretching exercises. We sonify an end target-point in a stretch exercise, using a stable sound (i.e., where the sonification is musically resolved) to encourage movements ending and an unstable sound (i.e., musically unresolved) to encourage continuation. Results on healthy participants show that instability leads to progression further beyond the target-point while stability leads to a smoother stop beyond this point. We conclude discussing how these findings should generalize to the CP population. Joseph W. Newbold, Nadia Bianchi-Berthouze, Nicolas E. Gold, Ana Tajadura-Jiménez, Amanda C. de C. Williams |
CHI | 3 |
| 2015 | ORBS and the limits of static slicingabstractObservation-based slicing is a recently-introduced, language-independent slicing technique based on the dependencies observable from program behaviour. Due to the well-known limits of dynamic analysis, we may only compute an under-approximation of the true observation-based slice. However, because the observation-based slice captures all possible dependence that can be observed, even such approximations can yield insight into the limitations of static slicing. For example, a static slice, S, that is strictly smaller than the corresponding observation based slice is potentially unsafe. We present the results of three sets of experiments on 12 different programs, including benchmarks and larger programs, which investigate the relationship between static and observation-based slicing. We show that, in extreme cases, observation-based slices can find the true minimal static slice, where static techniques cannot. For more typical cases, our results illustrate the potential for observation-based slicing to highlight limitations in static slicers. Finally, we report on the sensitivity of observation-based slicing to test quality. Dave W. Binkley, Nicolas E. Gold, Mark Harman, Syed S. Islam, Jens Krinke, Shin Yoo |
SCAM | 2 |
| 2014 | ORBS: language-independent program slicingabstractCurrent slicing techniques cannot handle systems written in multiple programming languages. Observation-Based Slicing (ORBS) is a language-independent slicing technique capable of slicing multi-language systems, including systems which contain (third party) binary components. A potential slice obtained through repeated statement deletion is validated by observing the behaviour of the program: if the slice and original program behave the same under the slicing criterion, the deletion is accepted. The resulting slice is similar to a dynamic slice. We evaluate five variants of ORBS on ten programs of different sizes and languages showing that it is less expensive than similar existing techniques. We also evaluate it on bash and four other systems to demonstrate feasible large-scale operation in which a parallelised ORBS needs up to 82% less time when using four threads. The results show that an ORBS slicer is simple to construct, effective at slicing, and able to handle systems written in multiple languages without specialist analysis tools. Dave W. Binkley, Nicolas E. Gold, Mark Harman, Syed S. Islam, Jens Krinke, Shin Yoo |
SIGSOFT FSE | 2 |
| 2013 | Efficient Identification of Linchpin Vertices in Dependence ClustersabstractSeveral authors have found evidence of large dependence clusters in the source code of a diverse range of systems, domains, and programming languages. This raises the question of how we might efficiently locate the fragments of code that give rise to large dependence clusters. We introduce an algorithm for the identification of linchpin vertices, which hold together large dependence clusters, and prove correctness properties for the algorithm’s primary innovations. We also report the results of an empirical study concerning the reduction in analysis time that our algorithm yields over its predecessor using a collection of 38 programs containing almost half a million lines of code. Our empirical findings indicate improvements of almost two orders of magnitude, making it possible to process larger programs for which it would have previously been impractical. Dave W. Binkley, Nicolas E. Gold, Mark Harman, Syed S. Islam, Jens Krinke, Zheng Li 0002 |
ACM Trans. Program. Lang. Syst. | 2 |
| 2011 | Model projection: simplifying models in response to restricting the environmentabstractThis paper introduces Model Projection. Finite state models such as Extended Finite State Machines are being used in an ever increasing number of software engineering activities. Model projection facilitates model development by specializing models for a specific operating environment. A projection is useful in many design-level applications including specification reuse and property verification. Kelly Androutsopoulos, Dave W. Binkley, David Clark 0001, Nicolas E. Gold, Mark Harman, Kevin Lano, Zheng Li 0002 |
ICSE | 4 |
| 2011 | Knitting Music and Programming: Reflections on the Frontiers of Source Code AnalysisabstractSource Code Analysis and Manipulation (SCAM) underpins virtually every operational software system. Despite the impact and ubiquity of SCAM principles and techniques in software engineering, there are still frontiers to be explored. Looking "inward" to existing techniques, one finds frontiers of performance, efficiency, accuracy, and usability, looking "outward" one finds new languages, new problems, and thus new approaches. This paper presents a reflective framework for characterizing source languages and domains. It draws on current research projects in music program analysis, musical score processing, and machine knitting to identify new frontiers for SCAM. The paper also identifies opportunities for SCAM to inspire, and be inspired by, problems and techniques in other domains. Nicolas E. Gold |
SCAM | 1 |
| 2010 | Cloning and copying between GNOME projectsabstractThis paper presents an approach to automatically distinguish the copied clone from the original in a pair of clones. It matches the line-by-line version information of a clone to the pair's other clone. A case study on the GNOME Desktop Suite revealed a complex flow of reused code between the different subprojects. In particular, it showed that the majority of larger clones (with a minimal size of 28 lines or higher) exist between the subprojects and more than 60% of the clone pairs can be automatically separated into original and copy. Jens Krinke, Nicolas E. Gold, Yue Jia 0001, Dave W. Binkley |
MSR | 2 |
| 2009 | A theoretical and empirical study of EFSM dependenceabstractDependence analysis underpins many activities in software maintenance such as comprehension and impact analysis. As a result, dependence has been studied widely for programming languages, notably through work on program slicing. However, there is comparatively little work on dependence analysis at the model level and hitherto, no empirical studies. We introduce a slicing tool for Extended Finite State Machines (EFSMs) and use the tool to gather empirical results on several forms of dependence found in ten EFSMs, including well-known benchmarks in addition to real-world EFSM models. We investigate the statistical properties of dependence using statistical tests for correlation and formalize and prove four of the empirical findings arising from our empirical study. The paper thus provides the maintainer with both empirical data and foundational theoretical results concerning dependence in EFSM models. Kelly Androutsopoulos, Nicolas E. Gold, Mark Harman, Zheng Li 0002, Laurence Tratt |
ICSM | 2 |
| 2009 | Dependence clusters in source codeabstractA dependence cluster is a set of program statements, all of which are mutually inter-dependent. This article reports a large scale empirical study of dependence clusters in C program source code. The study reveals that large dependence clusters are surprisingly commonplace. Most of the 45 programs studied have clusters of dependence that consume more than 10% of the whole program. Some even have clusters consuming 80% or more. The widespread existence of clusters has implications for source code analyses such as program comprehension, software maintenance, software testing, reverse engineering, reuse, and parallelization. Mark Harman, Dave W. Binkley, Keith B. Gallagher, Nicolas E. Gold, Jens Krinke |
ACM Trans. Program. Lang. Syst. | 4 |
| 2008 | Evaluating Key Statements AnalysisabstractKey Statement Analysis extracts from a program, statements that form the core of the program’s computation. A good set of key statements is small but has a large impact. Key statements form a useful starting point for understanding and manipulating a program. An empirical investigation of three kinds of key statements is presented. The three are based on Bieman and Ott’s principal variables. To be effective, the key statements must have high impact and form a small, highly cohesive unit. Using a minor improvement of metrics for measuring impact and cohesion, key statements are shown to capture about 75% of the semantic effect of the function from which they are drawn. At the same time, they have cohesion about 20 percentage points higher than the corresponding function. A statistical analysis of the differences shows that key statements have higher average impact and higher average cohesion (p≪0.001). Dave W. Binkley, Nicolas E. Gold, Mark Harman, Zheng Li 0002, Kiarash Mahdavi |
SCAM | 2 |
| 2008 | Locating dependence structures using search-based slicing
Tao Jiang 0060, Nicolas E. Gold, Mark Harman, Zheng Li 0002 |
Inf. Softw. Technol. | 2 |
| 2008 | An empirical study of the relationship between the concepts expressed in source code and dependence
Dave W. Binkley, Nicolas E. Gold, Mark Harman, Zheng Li 0002, Kiarash Mahdavi |
J. Syst. Softw. | 2 |
| 2008 | Guest Editor's Introduction: 10th Conference on Software Maintenance and Reengineering
Giuseppe A. Di Lucca, Nicolas E. Gold, Giuseppe Visaggio |
J. Syst. Softw. | 2 |
| 2007 | An empirical study of static program slice sizeabstractThis article presents results from a study of all slices from 43 programs, ranging up to 136,000 lines of code in size. The study investigates the effect of five aspects that affect slice size. Three slicing algorithms are used to study two algorithmic aspects: calling-context treatment and slice granularity. The remaining three aspects affect the upstream dependencies considered by the slicer. These include collapsing structure fields, removal of dead code, and the influence of points-to analysis. The results show that for the most precise slicer, the average slice contains just under one-third of the program. Furthermore, ignoring calling context causes a 50% increase in slice size, and while (coarse-grained) function-level slices are 33% larger than corresponding statement-level slices, they may be useful predictors of the (finer-grained) statement-level slice size. Finally, upstream analyses have an order of magnitude less influence on slice size. Dave W. Binkley, Nicolas E. Gold, Mark Harman |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2006 | Allowing Overlapping Boundaries in Source Code using a Search Based Approach to Concept BindingabstractOne approach to supporting program comprehension involves binding concepts to source code. Previously proposed approaches to concept binding have enforced nonoverlapping boundaries. However, real-world programs may contain overlapping concepts. This paper presents techniques to allow boundary overlap in the binding of concepts to source code. In order to allow boundaries to overlap, the concept binding problem is reformulated as a search problem. It is shown that the search space of overlapping concept bindings is exponentially large, indicating the suitability of sampling-based search algorithms. Hill climbing and genetic algorithms are introduced for sampling the space. The paper reports on experiments that apply these algorithms to 21 COBOL II programs taken from the commercial financial services sector. The results show that the genetic algorithm produces significantly better solutions than both the hill climber and random search. Nicolas E. Gold, Mark Harman, Zheng Li 0002, Kiarash Mahdavi |
ICSM | 1 |
| 2006 | The Sound of Software: Using Sonification to Aid ComprehensionabstractProgram comprehension of unfamiliar software is a daunting task and existing comprehension environments, although helping significantly, do not fully alleviate the information overload involved. The visual medium has been well-explored in aiding software engineers understanding of source code and other artifacts concerned with maintaining existing software systems, but the use of non-visual representations, e.g. sound, has not gone far beyond simple noises to indicate error conditions or attract attention in a running program. This paper aims to explore the program comprehension problems that could usefully be addressed using sound. There are many dimensions to this problem and this paper addresses a new area open program comprehension research, defining the problem space and beginning to populate it with possible solutions. We expect the primary focus of this session to be on software comprehension and sound although many disciplines are likely to become involved in the research that flows from it Lewis Irwin Berman, Sebastian Danicic, Keith B. Gallagher, Nicolas E. Gold |
ICPC | 4 |
| 2005 | Unifying program slicing and concept assignment for higher-level executable source code extractionabstractAbstract Program slicing and concept assignment have both been proposed as source code extraction techniques. Unfortunately, each has a weakness that prevents wider application. For slicing, the extraction criterion is expressed at a very low level; constructing a slicing criterion requires detailed code knowledge which is often unavailable. The concept assignment extraction criterion is expressed at the domain level. However, unlike a slice, the extracted code is not executable as a separate subprogram in its own right. This paper introduces a unification of slicing and concept assignment which exploits their combined advantages, while overcoming these two individual weaknesses. Our ‘concept slices’ are executable programs extracted using high‐level criteria. The paper introduces four techniques that combine slicing and concept assignment and algorithms for each. These algorithms were implemented in two separate tools used to illustrate the application of the concept slicing algorithms in two very different case studies. The first is a commercially‐written COBOL module from a large financial organization, the second is an open source utility program written in C. Copyright © 2005 John Wiley & Sons, Ltd. Nicolas E. Gold, Mark Harman, Dave W. Binkley, Robert M. Hierons |
Softw. Pract. Exp. | 1 |
| 2005 | Spatial Complexity Metrics: An Investigation of UtilityabstractSoftware comprehension is one of the largest costs in the software lifecycle. In an attempt to control the cost of comprehension, various complexity metrics have been proposed to characterize the difficulty of understanding a program and, thus, allow accurate estimation of the cost of a change. Such metrics are not always evaluated. This paper evaluates a group of metrics recently proposed to assess the "spatial complexity" of a program (spatial complexity is informally defined as the distance a maintainer must move within source code to build a mental model of that code). The evaluation takes the form of a large-scale empirical study of evolving source code drawn from a commercial organization. The results of this investigation show that most of the spatial complexity metrics evaluated offer no substantially better information about program complexity than the number of lines of code. However, one metric shows more promise and is thus deemed to be a candidate for further use and investigation. Nicolas E. Gold, Andrew Mohan, Paul J. Layzell 0001 |
IEEE Trans. Software Eng. | 1 |
| 2004 | An Approach to Understanding Program Comprehensibility Using Spatial Complexity, Concept Assignment and Typographical StyleabstractThis paper has briefly presented an approach to identifying the comprehensibility of a program and initial results from its application. The results obtained so far, indicate that this approach is useful in modelling the comprehensibility of a program as it evolves. However further work is required to calibrate this approach to more accurately reflect comprehensibility and to identify at what point corrective action should be undertaken to maintain the quality of the program. Andrew Mohan, Nicolas E. Gold, Paul J. Layzell 0001 |
ICSM | 2 |
| 2003 | A Framework for Understanding Conceptual Changes in Evolving Source CodeabstractAs systems evolve, they become harder to understand because the implementation of concepts (e.g. business rules) becomes less coherent. To preserve source code comprehensibility, we need to be able to predict how this property will change. This would allow the construction of a tool to suggest what information should be added or clarified (e.g. in comments) to maintain the code's comprehensibility. We propose a framework to characterize types of concept change during evolution. It is derived from an empirical investigation of concept changes in evolving commercial COBOL II files. The framework describes transformations in the geometry and interpretation of regions of source code. We conclude by relating our observations to the types of maintenance performed and suggest how this work could be developed to provide methods for preserving code quality based on comprehensibility. Nicolas E. Gold, Andrew Mohan |
ICSM | 1 |
| 2003 | A Broker Architecture for Integrating Data Using a Web Services Environment
Keith H. Bennett, Nicolas E. Gold, Paul J. Layzell 0001, Fujun Zhu, Pearl Brereton, David Budgen, John A. Keane, Ioannis Kotsiopoulos, Mark Turner 0001, Jie Xu 0007, Orouba Almilaji, Jung-Ching Chen, Ali Owrak |
ICSOC | 2 |
| 2002 | From System Comprehension to Program ComprehensionabstractProgram and system comprehension are vital parts of the software maintenance process. We discuss the need for both perspectives and describe two methods that may be integrated to provide a smooth transition in understanding from the system level to the program level. Results from a qualitative survey of expert industrial software maintainers, their information needs and requirements when comprehending software are initially presented. We then review existing software tools which facilitate system level and program comprehension. Two successful methods from the fields of data mining and concept assignment are discussed, each addressing some of these requirements. We also describe how these methods can be coupled to produce a broader software comprehension method which partly satisfies all the requirements. Future directions including the closer integration of the techniques are also identified. Christos Tjortjis, Nicolas E. Gold, Paul J. Layzell 0001, Keith H. Bennett |
COMPSAC | 2 |
| 2001 | An Architectural Model for Service-Based Flexible SoftwareabstractThe urgent need to change software easily to meet evolving business requirements requires a radical shift in the development of software, with a more demand-centric view leading to software which will be delivered as a service, within the framework of an open marketplace. We describe a service architecture and its rationale, in which components may be bound instantly, just at the time they are needed and then the binding may, be disengaged. This allows highly flexible software services to be evolved in "internet time". The paper focuses on early results: some of the aims have been demonstrated and amplified through an experimental implementation based on e-Speak, an existing and available technology. It is concluded that technology such as e-Speak provides a useful infrastructure that rapidly enabled us to demonstrate the basic operation and viability of our approach. Keith H. Bennett, Jie Xu 0007, Malcolm Munro, Zhuang Hong, Paul J. Layzell 0001, Nicolas E. Gold, David Budgen, Pearl Brereton |
COMPSAC | 6 |
| 2001 | An Architectural Model for Service-Based Software with Ultra Rapid EvolutionabstractThere is an urgent industrial need for new approaches to software evolution that will lead to far faster implementation of software changes. For the past 40 years, the techniques, processes and methods of software development have been dominated by supply side issues, and as a result the software industry is oriented towards developers rather than users. Existing software maintenance processes are simply too slow to meet the needs of many businesses. To achieve the levels of functionality, flexibility and time to market of changes and updates required by users, a radical shift is required in the development of software, with a more demand-centric view leading to software which will be delivered as a service, within the framework of an open marketplace. Although there are some signs that this approach is being adopted by industry, it is in a very limited and restricted form. We summarise research that has resulted in a long term strategic view of software engineering innovation. Based on this foundation, we describe more recent work that has resulted in an innovative demand-led model for the future of software. We describe a service architecture in which components may be bound instantly, just at the time they are needed and then the binding may be disengaged. Such ultra late binding requires that many non-functional attributes of the software are capable of automatic negotiation and resolution. Some of these attributes have been demonstrated and amplified through a prototype implementation based on existing and available technology. Keith H. Bennett, Malcolm Munro, Nicolas E. Gold, Paul J. Layzell 0001, David Budgen, Pearl Brereton |
ICSM | 3 |
| 2001 | Hypothesis-Based Concept Assignment to Support Software MaintenanceabstractSoftware maintenance is typically the most expensive part of the software lifecycle, with program comprehension forming the most costly part of software maintenance. This paper outlines a method for assisting program comprehension by addressing the concept assignment problem. The method, termed Hypothesis-Based Concept Assignment, uses informal information contained within source code to reason plausibly about the concepts contained within the code. An extensive evaluation has shown that the method can accurately recognise concepts in a range of real-world programs. Nicolas E. Gold |
ICSM | 1 |