VLDB 2026 Research / reviewers in the wild / expert
Suzanne M. Embury
dblp:e/SMEmbury
· DBLP profile ↗
48ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0002-3711-0778ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 28 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 18 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 7Applied, interdisciplinary, general and emerging computing · 5Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Who Is Afraid of Test Smells? Assessing Technical Debt from Developer Actions
Zhongyan Chen, Suzanne M. Embury, Markel Vigo |
ICTSS | 2 |
| 2023 | The impact of unequal contributions in student software engineering team projects
Kamilla Kopec-Harding, Sukru Eraslan, Bowen Cai 0008, Suzanne M. Embury, Caroline Jay |
J. Syst. Softw. | 4 |
| 2022 | Placement of Workloads from Advanced RDBMS Architectures into Complex Cloud Infrastructure
Antony Higginson, Clive Bostock, Norman W. Paton, Suzanne M. Embury |
EDBT | 4 |
| 2022 | A unifying framework for the systematic analysis of Git workflows
Julio César Cortés Ríos, Suzanne M. Embury, Sukru Eraslan |
Inf. Softw. Technol. | 2 |
| 2020 | Database Workload Capacity Planning using Time Series Analysis and Machine LearningabstractWhen procuring or administering any I.T. system or a component of an I.T. system, it is crucial to understand the computational resources required to run the critical business functions that are governed by any Service Level Agreements. Predicting the resources needed for future consumption is like looking into the proverbial crystal ball. In this paper we look at the forecasting techniques in use today and evaluate if those techniques are applicable to the deeper layers of the technological stack such as clustered database instances, applications and groups of transactions that make up the database workload. The approach has been implemented to use supervised machine learning to identify traits such as reoccurring patterns, shocks and trends that the workloads exhibit and account for those traits in the forecast. An experimental evaluation shows that the approach we propose reduces the complexity of performing a forecast, and accurate predictions have been produced for complex workloads. Antony Higginson, Mihaela Dediu, Octavian Arsene, Norman W. Paton, Suzanne M. Embury |
SIGMOD Conference | 5 |
| 2020 | Characterising the Quality of Behaviour Driven Development SpecificationsabstractBehaviour Driven Development (BDD) is an agile testing technique that enables software requirements to be specified as example interactions with the system, using structured natural language. While (in theory) being readable by non-technical stakeholders, the examples can also be executed against the code base to identify behaviours that are not yet correctly implemented. Writing good BDD suites, however, is challenging. A typical suite can contain hundreds of individual scenarios, that must correctly specify the system as a whole as well as individually. Despite much discussion amongst practitioners and in the blogosphere, as yet no formal definition of what makes for a high quality BDD suite has been given. To shed light on this, we surveyed BDD practitioners, asking for their opinions on the quality criteria that are important for BDD suites. We proposed, and asked for opinions on, four quality principles, and gave practitioners the option to add more principles of their own. This paper reports on the results of the survey, and presents an approach to defining BDD suite quality. Leonard Peter Binamungu, Suzanne M. Embury, Nikolaos Konstantinou 0001 |
XP | 2 |
| 2020 | Integrating GitLab metrics into coursework consultation sessions in a software engineering course
Sukru Eraslan, Kamilla Kopec-Harding, Caroline Jay, Suzanne M. Embury, Robert Haines, Julio César Cortés Ríos, Peter Crowther |
J. Syst. Softw. | 4 |
| 2018 | SOURCERY: User Driven Multi-Criteria Source SelectionabstractData scientists are usually interested in a subset of sources with properties that are most aligned to intended data use. The SOURCERY system supports interactive multi-criteria user-driven source selection. SOURCERY allows a user to identify criteria they consider of importance and indicate their relative importance, and seeks a source selection result aligned to the user-supplied criteria preferences. The user is given an overview of the properties of the sources that are selected along with visual analyses contextualizing the result in relation to what is theoretically possible and what is possible given the set of available sources. The system also enables a user to interactively perform iterative fine-tuning to explore how changes to preferences may impact results. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Nurzety A. Azuan, Suzanne M. Embury |
CIKM | 8 |
| 2018 | Maintaining behaviour driven development specifications: Challenges and opportunitiesabstractIn Behaviour-Driven Development (BDD) the behaviour of a software system is specified as a set of example interactions with the system using a "Given-When-Then" structure. These examples are expressed in high level domain-specific terms, and are executable. They thus act both as a specification of requirements and as tests that can verify whether the current system implementation provides the desired behaviour or not. This approach has many advantages but also presents some problems. When the number of examples grows, BDD specifications can become costly to maintain and extend. Some teams find that parts of the system are effectively frozen due to the challenges of finding and modifying the examples associated with them. We surveyed 75 BDD practitioners from 26 countries to understand the extent of BDD use, its benefits and challenges, and specifically the challenges of maintaining BDD specifications in practice. We found that BDD is in active use amongst respondents, and that the use of domain specific terms, improving communication among stakeholders, the executable nature of BDD specifications, and facilitating comprehension of code intentions are the main benefits of BDD. The results also showed that BDD specifications suffer the same maintenance challenges found in automated test suites more generally. We map the survey results to the literature, and propose 10 research opportunities in this area. Leonard Peter Binamungu, Suzanne M. Embury, Nikolaos Konstantinou 0001 |
SANER | 2 |
| 2018 | User driven multi-criteria source selectionabstractSource selection is the problem of identifying a subset of available data sources that best meet a user’s needs. In this paper we propose a user-driven approach to source selection that seeks to identify sources that are most fit for purpose. The approach employs a decision support methodology to take account of a user’s context, to allow end users to tune their preferences by specifying the relative importance between different criteria, looking to find a trade-off solution aligned with his/her preferences. The approach is extensible to incorporate diverse criteria, not drawn from a fixed set, and solutions can use a subset of the data from each selected source, rather than require that sources are used in their entirety or not at all. The paper describes and motivates the approach, presenting a methodology for modelling a user’s context, and its collection of optimisation algorithms for exploring the space of solutions, and compares and evaluates the resulting algorithms using multiple real world data sets. The experiments show how source selection results are produced that are attuned to each user’s preferences, both with respect to overall weighted utility and through faithful representation of a user’s preferences within a result, while scaling to potentially thousands of sources. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Julio César Cortés Ríos, Nurzety A. Azuan, Suzanne M. Embury |
Inf. Sci. | 9 |
| 2018 | Data journeys: Identifying social and technical barriers to data movement in large, complex organisations
Iliada Eleftheriou, Suzanne M. Embury, Rebecca Moden, Peter Dobinson, Andy Brass |
J. Biomed. Informatics | 2 |
| 2017 | DBaaS Cloud Capacity Planning - Accounting for Dynamic RDBMS System that Employ Clustering and Standby Architectures
Antony Higginson, Norman W. Paton, Suzanne M. Embury, Clive Bostock |
EDBT | 3 |
| 2016 | Pay-as-you-go Data Integration: Experiences and Recurring Themes
Norman W. Paton, Khalid Belhajjame, Suzanne M. Embury, Alvaro A. A. Fernandes, Ruhaila Maskat |
SOFSEM | 3 |
| 2014 | Verification of Semantic Web Service Annotations Using Ontology-Based PartitioningabstractSemantic annotation of web services has been proposed as a solution to the problem of discovering services to fit a particular need and reusing them appropriately. While there exist tools that assist human users in the annotation task, e.g., Radiant and Meteor-S, no semantic annotation proposal considers the problem of verifying the accuracy of the resulting annotations. Early evidence from workflow compatibility checking suggests that the proportion of annotations that contain some form of inaccuracy is high, and yet no tools exist to help annotators to test the results of their work systematically before they are deployed for public use. In this paper, we adapt techniques from conventional software testing to the verification of semantic annotations for web service input and output parameters. We present an algorithm for the testing process and discuss ways in which manual effort from the annotator during testing can be reduced. We also present two adequacy criteria for specifying test cases used as input for the testing process. These criteria are based on structural coverage of the domain ontology used for annotation. The results of an evaluation exercise, based on a collection of annotations for bioinformatics web services, show that defects can be successfully detected by the technique. Khalid Belhajjame, Suzanne M. Embury, Norman W. Paton |
IEEE Trans. Serv. Comput. | 2 |
| 2013 | Reference Architectures to Measure Data Completeness across Integrated Databases
Nurul A. Emran, Suzanne M. Embury, Paolo Missier, Norashikin Ahmad |
ACIIDS (1) | 2 |
| 2013 | Measuring Data Completeness for Microbial Genomics Database
Nurul A. Emran, Suzanne M. Embury, Paolo Missier, Mohd Noor Mat Isa, Azah Kamilah Muda |
ACIIDS (1) | 2 |
| 2013 | Incrementally improving dataspaces based on user feedback
Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
Inf. Syst. | 3 |
| 2011 | User Feedback as a First Class Citizen in Information Integration Systems
Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Cornelia Hedeler, Suzanne M. Embury |
CIDR | 5 |
| 2011 | Pay-as-you-go mapping selection in dataspacesabstractThe vision of dataspaces proposes an alternative to classical data integration approaches with reduced up-front costs followed by incremental improvement on a pay-as-you-go basis. In this paper, we demonstrate DSToolkit, a system that allows users to provide feedback on results of queries posed over an integration schema. Such feedback is then used to annotate the mappings with their respective precision and recall. The system then allows a user to state the expected levels of precision (or recall) that the query results should exhibit and, in order to produce those results, the system selects those mappings that are predicted to meet the stated constraints. Cornelia Hedeler, Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Suzanne M. Embury, Lu Mao, Chenjuan Guo |
SIGMOD Conference | 5 |
| 2010 | Feedback-based annotation, selection and refinement of schema mappings for dataspacesabstractThe specification of schema mappings has proved to be time and resource consuming, and has been recognized as a critical bottleneck to the large scale deployment of data integration systems. In an attempt to address this issue, dataspaces have been proposed as a data management abstraction that aims to reduce the up-front cost required to setup a data integration system by gradually specifying schema mappings through interaction with end users in a pay-as-you-go fashion. As a step in this direction, we explore an approach for incrementally annotating schema mappings using feedback obtained from end users. In doing so, we do not expect users to examine mapping specifications; rather, they comment on results to queries evaluated using the mappings. Using annotations computed on the basis of user feedback, we present a method for selecting from the set of candidate mappings, those to be used for query evaluation considering user requirements in terms of precision and recall. In doing so, we cast mapping selection as an optimization problem. Mapping annotations may reveal that the quality of schema mappings is poor. We also show how feedback can be used to support the derivation of better quality mappings from existing mappings through refinement. An evolutionary algorithm is used to efficiently and effectively explore the large space of mappings that can be obtained through refinement. The results of evaluation exercises show the effectiveness of our solution for annotating, selecting and refining schema mappings. Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
EDBT | 3 |
| 2008 | Information quality in proteomicsabstractProteomics, the study of the protein complement of a biological system, is generating increasing quantities of data from rapidly developing technologies employed in a variety of different experimental workflows. Experimental processes, e.g. for comparative 2D gel studies or LC-MS/MS analyses of complex protein mixtures, involve a number of steps: from experimental design, through wet and dry lab operations, to publication of data in repositories and finally to data annotation and maintenance. The presence of inaccuracies throughout the processing pipeline, however, results in data that can be untrustworthy, thus offsetting the benefits of high-throughput technology. While researchers and practitioners are generally aware of some of the information quality issues associated with public proteomics data, there are few accepted criteria and guidelines for dealing with them. In this article, we highlight factors that impact on the quality of experimental data and review current approaches to information quality management in proteomics. Data quality issues are considered throughout the lifecycle of a proteomics experiment, from experiment design and technique selection, through data analysis, to archiving and sharing. David Stead, Norman W. Paton, Paolo Missier, Suzanne M. Embury, Cornelia Hedeler, Binling Jin, Alistair J. P. Brown, Alun D. Preece |
Briefings Bioinform. | 4 |
| 2008 | An ontology-based approach to handling information quality in e-ScienceabstractAbstract In this paper we outline a framework for managing information quality (IQ) in an e‐Science context. In contrast to previous approaches that take a very abstract view of IQ properties, we allow scientists to define the quality characteristics that are of importance to them in their particular domain. For example, ‘accuracy’ may be defined in terms of the conformance of experimental data to a particular standard. User‐scientists specify their IQ preferences against a formal ontology, so that the definitions are machine‐manipulable, allowing the environment to classify and organize domain‐specific quality characteristics within an overall quality management framework. As an illustration of our approach, we present an example Web service that computes IQ annotations for experiment datasets in transcriptomics. Copyright © 2007 John Wiley & Sons, Ltd. Alun D. Preece, Paolo Missier, Suzanne M. Embury, Binling Jin, Robert Mark Greenwood |
Concurr. Comput. Pract. Exp. | 3 |
| 2008 | Automatic annotation of Web services based on workflow definitionsabstractSemantic annotations of web services can support the effective and efficient discovery of services, and guide their composition into workflows. At present, however, the practical utility of such annotations is limited by the small number of service annotations available for general use. Manual annotation of services is a time consuming and thus expensive task, so some means are required by which services can be automatically (or semi-automatically) annotated. In this paper, we show how information can be inferred about the semantics of operation parameters based on their connections to other (annotated) operation parameters within tried-and-tested workflows. Because the data links in the workflows do not necessarily contain every possible connection of compatible parameters, we can infer only constraints on the semantics of parameters. We show that despite their imprecise nature these so-called loose annotations are still of value in supporting the manual annotation task, inspecting workflows and discovering services. We also show that derived annotations for already annotated parameters are useful. By comparing existing and newly derived annotations of operation parameters, we can support the detection of errors in existing annotations, the ontology used for annotation and in workflows. The derivation mechanism has been implemented, and its practical applicability for inferring new annotations has been established through an experimental evaluation. The usefulness of the derived annotations is also demonstrated. Khalid Belhajjame, Suzanne M. Embury, Norman W. Paton, Robert Stevens 0001, Carole A. Goble |
ACM Trans. Web | 2 |
| 2007 | Tool Support to Implementing Business Rules in Database ApplicationsabstractIn many cases, the programmer may require to encode business rules into the database applications. To do this, a large number of program elements may need to be examined by the programmer, to determine which have the capacity to violate a new rule and if so what minimal changes are required to prevent such violations. This process can be time-consuming, and even seasoned programmers can miss difficult and obscure cases in the mass of code. In this paper, we describe a static source code analysis technique to assist the programmer in enforcing business rules in a way that cuts down the amount of irrelevant code to be examined. Our technique derives all the possible ways in which a new business rule can be violated by the programs in the system being modified, and the specific program elements responsible. Liwen Lin, Suzanne M. Embury, Brian Warboys |
COMPSAC (1) | 2 |
| 2007 | Exploratory Design of Derivation Business Rules Using Query Rewriting
Roman Krenický, David Willmor, Suzanne M. Embury |
SEKE | 3 |
| 2007 | Managing information quality in e-science: the qurator workbenchabstractData-intensive e-science applications often rely on third-party data found in public repositories, whose quality is largely unknown. Although scientists are aware that this uncertainty may lead to incorrect scientific conclusions, in the absence of a quantitative characterization of data quality properties they find it difficult to formulate precise data acceptability criteria. We present an Information Quality management workbench, called Qurator, that supports data experts in the specification of personal quality models, and lets them derive effective criteria for data acceptability. The demo of our working prototype will illustrate our approach on a real e-science workflow for a bioinformatics application. Paolo Missier, Suzanne M. Embury, Robert Mark Greenwood, Alun D. Preece, Binling Jin |
SIGMOD Conference | 2 |
| 2006 | Towards the Management of Information Quality in ProteomicsabstractWe outline the application of a framework for managing information quality (IQ) in proteomics. The approach allows scientists to define the quality characteristics that are of importance in their particular domain, by extending a generic ontology of IQ concepts. Two quality indicators are defined for proteomic experiments: hit ratio and mass coverage. We describe how our framework allows experiments marked-up in a standard format (e.g. PEDRo) to be annotated with these computed indicators, and how the annotations can be viewed using a convenient plugin to the commonly-used Pedro data entry tool. Alun D. Preece, Binling Jin, Paolo Missier, Suzanne M. Embury, David Stead, Al Brown |
CBMS | 4 |
| 2006 | Managing Information Quality in e-Science Using Semantic Web Technology
Alun D. Preece, Binling Jin, Edoardo Pignotti, Paolo Missier, Suzanne M. Embury, David Stead, Al Brown |
ESWC | 5 |
| 2006 | An intensional approach to the specification of test cases for database applicationsabstractWhen testing database applications, in addition to creating in-memory fixtures it is also necessary to create an initial database state that is appropriate for each test case. Current approaches either require exact database states to be specified in advance, or else generate a single initial state (under guidance from the user) that is intended to be suitable for execution of all test cases. The first method allows large test suites to be executed in batch, but requires considerable programmer effort to create the test cases (and to maintain them). The second method requires less programmer effort, but increases the likelihood that test cases will fail in non-fault situations, due to unexpected changes to the content of the database. In this paper, we propose a new approach in which the database states required for testing are specified intensionally, as constrained queries, that can be used to prepare the database for testing automatically. This technique overcomes the limitations of the other approaches, and does not appear to impose significant performance overheads. David Willmor, Suzanne M. Embury |
ICSE | 2 |
| 2006 | Assessing Impacts of Changes to Business Rules through Data ExplorationabstractThe benefits of impact analysis in the maintenance and evolution of software systems are well known, and many forms of impact analysis, over different software life cycle objects, have been proposed. However, one form of impact from software change has yet to be explored by the research community: these are the impacts of changes to software on data. In particular, when the business rules enforced by a system change, it may be necessary to perform some cleanup or transformations on persistent data, in order to bring it into line with the new system functionality. Alternatively, the proposed new rule may itself have to be modified, if the data impacts are too costly to address. In this paper, we show how such impacts can arise, and propose an approach to assisting in their identification through user-driven exploration of hypothetical change-impact scenarios. Suzanne M. Embury, David Willmor, Lei Dang |
ICSEA | 1 |
| 2006 | Automatic Annotation of Web Services Based on Workflow Definitions
Khalid Belhajjame, Suzanne M. Embury, Norman W. Paton, Robert Stevens 0001, Carole A. Goble |
ISWC | 2 |
| 2006 | Quality Views: Capturing and Exploiting the User Perspective on Data Quality
Paolo Missier, Suzanne M. Embury, Robert Mark Greenwood, Alun D. Preece, Binling Jin |
VLDB | 2 |
| 2005 | Facilitating the Implementation and Evolution of Business RulesabstractMany software systems implement, amongst other things, a collection of business rules. However, the process of evolving the business rules associated with a system is both time consuming and error prone. In this paper, we propose a novel approach to facilitating business rule evolution through capturing information to assist with the evolution of rules at the point of implementation. We analyse the process of rule evolution, in order to determine the information that must be captured. Our approach allows programmers to implement rules by embedding them into application programs (giving the required performance and genericity), while still easing the problems of evolution. Liwen Lin, Suzanne M. Embury, Brian Warboys |
ICSM | 2 |
| 2005 | A Safe Regression Test Selection Technique for Database-Driven ApplicationsabstractRegression testing is a widely-used method for checking whether modifications to software systems have adversely affected the overall functionality. This is potentially an expensive process, since test suites can be large and time-consuming to execute. The overall costs can be reduced if tests that cannot possibly be affected by the modifications are ignored. Various techniques for selecting subsets of tests for re-execution have been proposed, as well as methods for proving that particular test selection criteria do not omit relevant tests. However, current selection techniques are focused on identifying the impact of modifications on program state. They assume that the only factor that can change the result of a test case is the set of input values given for it, while all other influences on the behavior of the program (such as external interrupts or hardware faults) will be constant for each re-execution of the test. This assumption is impractical in the case of an important class of software system, i.e. systems which make use of an external persistent state, such as a database management system, to share information between application invocations. If applied naively to such systems, existing regression test selection algorithms will omit certain test cases which could in fact be affected by the modifications to the code. In this paper, we show why this is the case, and propose a new definition of safety for regression test selection that takes into account the interactions of the program with a database state. We also present an algorithm and associated tool that safely performs test selection for database-driven applications, and (since efficiency is an important concern for test selection algorithms) we propose a variant that defines safety in terms of database state alone. This latter form of safety allows more efficient regression testing to be performed for applications in which program state is used only as a temporary holding space for data from the database. The claims of increased efficiency of both forms of safety are supported by the results of an empirical comparison with existing techniques. David Willmor, Suzanne M. Embury |
ICSM | 2 |
| 2004 | Program Slicing in the Presence of a Database StateabstractProgram slicing has long been recognised as a valuable technique for supporting the software maintenance process. However, many programs operate over some kind of external state, as well as the internal program state. Arguably, the most significant form of external state is that used to store data associated with the application, for example, in a database management system. We propose an approach to supporting slicing over both program and database state, which requires the introduction of two new forms of data dependency into the standard program dependency graph. Our method expands the usefulness of program slicing techniques to the considerable number of database application programs that are being maintained within industry and science today. David Willmor, Suzanne M. Embury, Jianhua Shao 0001 |
ICSM | 2 |
| 2004 | Algorithms for analysing related constraint business rules
Gaihua Fu, Jianhua Shao 0001, Suzanne M. Embury, W. Alex Gray |
Data Knowl. Eng. | 3 |
| 2003 | Analysing the Impact of Adding Integrity Constraints to Information Systems
Suzanne M. Embury, Jianhua Shao 0001 |
CAiSE | 1 |
| 2002 | Representing Constraint Business Rules Extracted from Legacy Systems
Gaihua Fu, Jianhua Shao 0001, Suzanne M. Embury, W. Alex Gray |
DEXA | 3 |
| 2001 | Querying Data-Intensive Programs for Data Design
Jianhua Shao 0001, Xingkun Liu, Gaihua Fu, Suzanne M. Embury, W. Alex Gray |
CAiSE | 4 |
| 2001 | Adapting integrity enforcement techniques for data reconciliation
Suzanne M. Embury, Sue M. Brandt, John S. Robinson, Iain Sutherland, Frank A. Bisby, W. Alex Gray, Andrew C. Jones, Richard J. White |
Inf. Syst. | 1 |
| 2000 | Advertising Database Capabilities for Information Sharing
Suzanne M. Embury, Jianhua Shao 0001, W. Alex Gray, Nigel Fishlock |
CAiSE | 1 |
| 2000 | Assisting the Integration of Taxonomic Data: The LITCHI ToolkitabstractWe demonstrate a prototype toolkit that uses constraints and constraint violation repair techniques to enable the automated detection and, where possible, the automated resolution of conflicts in taxonomic databases. Iain Sutherland, John S. Robinson, Sue M. Brandt, Andrew C. Jones, Suzanne M. Embury, W. Alex Gray, Richard J. White, Frank A. Bisby |
ICDE | 5 |
| 2000 | Techniques for Effective Integration, Maintenance and Evolution of Species DatabasesabstractThe LITCHI project is concerned with the integration and maintenance of databases of biological knowledge organised by species. We use constraints pertaining to good taxonomic practice in order to identify taxonomix conflicts in individual species databases and in databases formed by merging species databases from distinct sources. The LITCHI system can be used to resolve such conflicts incrementally. As the project has progressed, we have identified a number of distinctive features of the problem domain, and needs of the intended users, which have had a significant impact on the techniques and modes of operation that we found to be appropriate, especially in contrast with applications that handle rapidly-accumulating 'raw' data. It is upon these aspects of LITCHI that we concentrate in the present paper viewing LITCHI as an example of the more general problem of merging scientific data sets in which conflicts between the terminology used can occur. Andrew C. Jones, Iain Sutherland, Suzanne M. Embury, W. Alex Gray, Richard J. White, John S. Robinson, Frank A. Bisby, Sue M. Brandt |
SSDBM | 3 |
| 1999 | Conflict Detection for Integration of Taxonomic Data SourcesabstractOver recent years, international initiatives such as the 1993 UN Convention on Biological Diversity have highlighted the need for information about species diversity on a global scale. However, attempts to build global information systems by integrating smaller, independently created biodiversity databases have been hampered by differences in the sets of species names used. Some databases use different names to refer to the same species, while in other cases the same name can be applied to differing definitions of a species, or even entirely different species. The LITCHI project aims to assist biologists in the integration of databases by searching for conflicts within taxonomic checklists (i.e. lists of the species names used in a database and the relationships between them). In order to detect such conflicts, we have created a formal model of taxonomic practice, which describes (amongst other things) what it means for a checklist to be consistent and well-specified. This model has been used as the basis for a prototype tool that uses Prolog to search for naming conflicts within a relational database of checklists. We describe the background to our formal model and show how it has been used to implement the LITCHI system. Our prototype tool is already proving its worth by detecting conflicts and errors within real taxonomic checklists. Suzanne M. Embury, Andrew C. Jones, Iain Sutherland, W. Alex Gray, Richard J. White, John S. Robinson, Frank A. Bisby, Sue M. Brandt |
SSDBM | 1 |
| 1999 | LITCHI: Knowledge Integrity Testing for Taxonomic DatabasesabstractSummary form only given. The LITCHI project (Logic-based Integration of Taxonomic Conflicts in Heterogeneous Information Systems) aims to develop software to enable the automated detection and, where possible, resolution of conflicts in taxonomic checklists. A taxonomic checklist is a list of the names of species (and other taxa) used within a particular biological database. Since species names are typically used to gain access to data within biological databases, checklists provide a concise representation of the data values that can act as keys when querying such databases. More importantly, species names are also typically used as the join attribute when integrating several biological databases. However, naming of species is a subjective activity, and different scientific communities will have different ideas about the names that should be used for particular species. These conflicts of opinion arise as a result of the subjective nature of the classification process and geographical or historical differences in background knowledge. Some communities may use different names for the same species, while other groups of scientists may use the same name to refer to different species. Often, there is no one right naming scheme, but some consistent set of names must be used if biological databases are to be integrated. Therefore, there is a real need for a tool which will assist biologists in the integration of checklists, prior to the integration of species databases, so that these differences of opinion can be resolved. Iain Sutherland, Suzanne M. Embury, Andrew C. Jones, W. Alex Gray, Richard J. White, John S. Robinson, Frank A. Bisby, Sue M. Brandt |
SSDBM | 2 |
| 1999 | The Evolving Role of Constraints in the Functional Data Model
Peter M. D. Gray, Suzanne M. Embury, Kit-Ying Hui, Graham J. L. Kemp |
J. Intell. Inf. Syst. | 2 |
| 1997 | Distributing Semantic Constraints Between Heterogeneous DatabasesabstractIn recent years, research on distributing databases over networks has become increasingly important. In this paper, we concentrate on the issues of the interoperability of heterogeneous DBMSs and enforcing integrity across a multi-database made in this fashion. This has been done through a cooperative project between Aberdeen and Linko/spl uml/ping universities, with database modules distributed between the sites. In the process, we have shown the advantage of using DBMSs based on variants of the functional data model (FDM), which has made it remarkably straightforward to interoperate queries and schema definitions. Further, we have used the constraint transformation facilities of P/FDM (Prolog implementation of FDM) to compile global constraints into active rules installed locally on one or more AMOS (Active Mediators Object System) servers. We present the theory behind this, and the conditions for it to improve performance. Stefan Grufman, Fredrik Samson, Suzanne M. Embury, Peter M. D. Gray, Tore Risch |
ICDE | 3 |
| 1993 | On using Prolog to implement object-oriented databases
Norman W. Paton, Scott Leishman, Suzanne M. Embury, Peter M. D. Gray |
Inf. Softw. Technol. | 3 |