EDBT 2026 Demo / reviewers in the wild / expert
Norman W. Paton
dblp:p/NWPaton
· DBLP profile ↗
84ranked-venue papers in the field
6as first author
11since 2021 · last 2025
0000-0003-2008-6617ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 66 (6 first)Knowledge Engineering, Semantic Web & Information Systems · 8Information Retrieval & Web Search · 4Business Process & Enterprise Data · 3Data Mining & Knowledge Discovery · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gem: Gaussian Mixture Model Embeddings for Numerical Feature Distributions
Hafiz Tayyab Rauf, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
EDBT | 3 |
| 2025 | Taxonomy Inference for Tabular Data Using Large Language Models
Jiaoyan Chen 0001, Norman W. Paton |
ESWC (1) | 3 |
| 2025 | TableDC: Deep Clustering for Tabular DataabstractDeep clustering (DC), a fusion of deep representation learning and clustering, has recently demonstrated positive results in data science, particularly text processing and computer vision. However, joint optimization of feature learning and data distribution in the multi-dimensional space is domain-specific, so existing DC methods struggle to generalize to other application domains (such as data integration). In data management tasks, where high-density embeddings and overlapping clusters dominate, a data management-specific DC algorithm should be able to interact better with the data properties to support data integration tasks. This paper presents a deep clustering algorithm for tabular data (TableDC) that reflects the properties of data management applications that cluster tables (schema inference), rows (entity resolution) and columns (domain discovery). To address overlapping clusters, TableDC integrates Mahalanobis distance, which considers variance and correlation within the data, offering a similarity method suitable for tabular data in high-dimensional latent spaces. TableDC also shows higher tolerance to outliers through its heavy-tailed Cauchy distribution as the similarity kernel. The proposed similarity measure is particularly beneficial where the embeddings of raw data are densely packed and exhibit high degrees of overlap. Data integration tasks may also involve large numbers of clusters, which challenges the scalability of existing DC methods. TableDC learns data embeddings with a large number of clusters more efficiently than baseline DC methods, which scale in quadratic time. We evaluated TableDC with several existing DC, Standard Clustering (SC), and state-of-the-art bespoke methods over benchmark datasets. TableDC consistently outperforms existing DC, SC and bespoke methods. Hafiz Tayyab Rauf, André Freitas, Norman W. Paton |
Proc. ACM Manag. Data | 3 |
| 2025 | Front Matter
Sonia Bergamaschi, Sourav S. Bhowmick, Philippe Bonnet, Surajit Chaudhuri, Xiaoou Ding, Hakan Ferhatosmanoglu, Raul Castro Fernandez, Jana Giceva, Madelon Hulsebos, Alexandra Meliou, Nikos Ntarmos, Themis Palpanas, John Paparrizos, Norman W. Paton, Subhadeep Sarkar 0001, Giovanni Simonini, Nesime Tatbul, Jiuqi Wei, Jingren Zhou 0001 |
Proc. VLDB Endow. | 14 |
| 2024 | Dataset Discovery and Exploration: State-of-the-art, Challenges and Opportunities
Norman W. Paton |
EDBT | 1 |
| 2024 | Deep Clustering for Data Cleaning and Integration
Hafiz Tayyab Rauf, André Freitas, Norman W. Paton |
EDBT | 3 |
| 2022 | Voyager: Data Discovery and Integration for Onboarding in Data Science
Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, André Freitas |
EDBT | 2 |
| 2022 | Placement of Workloads from Advanced RDBMS Architectures into Complex Cloud Infrastructure
Antony Higginson, Clive Bostock, Norman W. Paton, Suzanne M. Embury |
EDBT | 3 |
| 2022 | Schema mapping generation in the wild
Lacramioara Mazilu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
Inf. Syst. | 2 |
| 2021 | Natural Language Inference over Tables: Enabling Explainable Data Exploration on Data Lakes
Mario Ramirez, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
ESWC | 3 |
| 2021 | Cost-effective Variational Active Entity ResolutionabstractAccurately identifying different representations of the same real-world entity is an integral part of data cleaning and many methods have been proposed to accomplish it. The challenges of this entity resolution task that demand so much research attention are often rooted in the task-specificity and user-dependence of the process. Adopting deep learning techniques has the potential to lessen these challenges. In this paper, we set out to devise an entity resolution method that builds on the robustness conferred by deep autoencoders to reduce human-involvement costs. Specifically, we reduce the cost of training deep entity resolution models by performing unsupervised representation learning. This unveils a transferability property of the resulting model that can further reduce the cost of applying the approach to new datasets by means of transfer learning. Finally, we reduce the cost of labeling training data through an active learning approach that builds on the properties conferred by the use of deep autoencoders. Empirical evaluation confirms the accomplishment of our cost-reduction desideratum, while achieving comparable effectiveness with state-of-the-art alternatives. Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, Stuart Davie, André Freitas |
ICDE | 2 |
| 2020 | Schema Mapping Generation in the Wild: A Demonstration with Open Government DataabstractSchema mapping generation identifies how data sets can be combined to create views that are relevant to an application. Where the data sets to be combined lack declared relationships, such as foreign keys, schema mapping generation can be considered to be in the wild. In this paper, we describe an approach to schema mapping generation in the context of open government data, in particular, the London Datastore. Mapping generation is informed by inferred profiling data about the data sets and their relationships, where the data sets are made available as csv files. We outline the mapping generation algorithm, and describe a demonstration of the approach, in which the user can: (i) specify the target to be populated by the generated mappings over a collection of sources from The London Datastore; (ii) browse the generated candidate mappings and the evidence that informed their creation; and (iii) steer the mapping generation process, to make use of preferred sources and dependable profiling results. Lacramioara Mazilu, Nikolaos Konstantinou 0001, Norman W. Paton, Alvaro A. A. Fernandes |
EDBT | 3 |
| 2020 | Dataset Discovery in Data LakesabstractData analytics stands to benefit from the increasing availability of datasets that are held without their conceptual relationships being explicitly known. When collected, these datasets form a data lake from which, by processes like data wrangling, specific target datasets can be constructed that enable value- adding analytics. Given the potential vastness of such data lakes, the issue arises of how to pull out of the lake those datasets that might contribute to wrangling out a given target. We refer to this as the problem of dataset discovery in data lakes and this paper contributes an effective and efficient solution to it. Our approach uses features of the values in a dataset to construct hash- based indexes that map those features into a uniform distance space. This makes it possible to define similarity distances between features and to take those distances as measurements of relatedness w.r.t. a target table. Given the latter (and exemplar tuples), our approach returns the most related tables in the lake. We provide a detailed description of the approach and report on empirical results for two forms of relatedness (unionability and joinability) comparing them with prior work, where pertinent, and showing significant improvements in all of precision, recall, target coverage, indexing and discovery times. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
ICDE | 3 |
| 2020 | Database Workload Capacity Planning using Time Series Analysis and Machine LearningabstractWhen procuring or administering any I.T. system or a component of an I.T. system, it is crucial to understand the computational resources required to run the critical business functions that are governed by any Service Level Agreements. Predicting the resources needed for future consumption is like looking into the proverbial crystal ball. In this paper we look at the forecasting techniques in use today and evaluate if those techniques are applicable to the deeper layers of the technological stack such as clustered database instances, applications and groups of transactions that make up the database workload. The approach has been implemented to use supervised machine learning to identify traits such as reoccurring patterns, shocks and trends that the workloads exhibit and account for those traits in the forecast. An experimental evaluation shows that the approach we propose reduces the complexity of performing a forecast, and accurate predictions have been produced for complex workloads. Antony Higginson, Mihaela Dediu, Octavian Arsene, Norman W. Paton, Suzanne M. Embury |
SIGMOD Conference | 4 |
| 2020 | Feedback driven improvement of data preparation pipelines
Nikolaos Konstantinou 0001, Norman W. Paton |
Inf. Syst. | 2 |
| 2019 | Feedback Driven Improvement of Data Preparation Pipelines
Nikolaos Konstantinou 0001, Norman W. Paton |
DOLAP | 2 |
| 2019 | Automating Data Preparation: Can We? Should We? Must We?
Norman W. Paton |
DOLAP | 1 |
| 2019 | SynthEdit: Format transformations by example using edit operationsabstractFormat transformation is one of the most labor intensive tasks of a data wrangling process. Recent advances in programming by example proposed synthesis algorithms that showed promising results on spreadsheet data. However, when employed on repositories consisting of multiple sources and large number of examples, such algorithms manifest scalability issues. This paper introduces a new transformation synthesis technique based on edit operations that enables efficient learning of transformation programs. Empirical results show comparable effectiveness and dramatic improvements in efficiency over the state-of-the art. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
EDBT | 3 |
| 2019 | Dynamap: Schema Mapping Generation in the WildabstractSchema mappings enable declarative and executable specification of transformations between different schematic representations of application concepts. Most work on mapping generation has assumed that the source and target schemas are well defined, e.g., with declared keys and foreign keys, and that the mapping generation processes exist to support the data engineer in the labour-intensive process of producing a high-quality integration. However, organizations increasingly have access to numerous independently produced data sets, e.g., in a data lake, with a requirement to produce rapid, best-effort integrations, without extensive manual effort. This paper introduces Dynamap, a mapping generation algorithm for such settings, where metadata about sources and the relationships between them is derived from automated data profiling, and where there may be many alternative ways of combining source tables. Our contributions include a dynamic programming algorithm for exploring the space of potential mappings, and techniques for propagating profiling data through mappings, so that the fitness of candidate mappings can be estimated. Experimental results show the effectiveness and scalability of the approach in a variety of synthetic and real-world scenarios. Lacramioara Mazilu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
SSDBM | 2 |
| 2018 | SOURCERY: User Driven Multi-Criteria Source SelectionabstractData scientists are usually interested in a subset of sources with properties that are most aligned to intended data use. The SOURCERY system supports interactive multi-criteria user-driven source selection. SOURCERY allows a user to identify criteria they consider of importance and indicate their relative importance, and seeks a source selection result aligned to the user-supplied criteria preferences. The user is given an overview of the properties of the sources that are selected along with visual analyses contextualizing the result in relation to what is theoretically possible and what is possible given the set of available sources. The system also enables a user to interactively perform iterative fine-tuning to explore how changes to preferences may impact results. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Nurzety A. Azuan, Suzanne M. Embury |
CIKM | 3 |
| 2018 | User driven multi-criteria source selectionabstractSource selection is the problem of identifying a subset of available data sources that best meet a user’s needs. In this paper we propose a user-driven approach to source selection that seeks to identify sources that are most fit for purpose. The approach employs a decision support methodology to take account of a user’s context, to allow end users to tune their preferences by specifying the relative importance between different criteria, looking to find a trade-off solution aligned with his/her preferences. The approach is extensible to incorporate diverse criteria, not drawn from a fixed set, and solutions can use a subset of the data from each selected source, rather than require that sources are used in their entirety or not at all. The paper describes and motivates the approach, presenting a methodology for modelling a user’s context, and its collection of optimisation algorithms for exploring the space of solutions, and compares and evaluates the resulting algorithms using multiple real world data sets. The experiments show how source selection results are produced that are attuned to each user’s preferences, both with respect to overall weighted utility and through faithful representation of a user’s preferences within a result, while scaling to potentially thousands of sources. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Julio César Cortés Ríos, Nurzety A. Azuan, Suzanne M. Embury |
Inf. Sci. | 3 |
| 2017 | Targeted Feedback Collection Applied to Multi-Criteria Source Selection
Julio César Cortés Ríos, Norman W. Paton, Alvaro A. A. Fernandes, Edward Abel, John A. Keane |
ADBIS | 2 |
| 2017 | Data context informed data wranglingabstractThe process of preparing potentially large and complex data sets for further analysis or manual examination is often called data wrangling. In classical warehousing environments, the steps in such a process have been carried out using Extract-Transform-Load platforms, with significant manual involvement in specifying, configuring or tuning many of them. Cost-effective data wrangling processes need to ensure that data wrangling steps benefit from automation wherever possible. In this paper, we define a methodology to fully automate an end-to-end data wrangling process incorporating data context, which associates portions of a target schema with potentially spurious extensional data of types that are commonly available. Instance-based evidence together with data profiling paves the way to inform automation in several steps within the wrangling process, specifically, matching, mapping validation, value format transformation, and data repair. The approach is evaluated with real estate data showing substantial improvements in the results of automated wrangling. Martin Koehler, Alex Teodor Bogatu, Cristina Civili, Nikolaos Konstantinou 0001, Edward Abel, Alvaro A. A. Fernandes, John A. Keane, Leonid Libkin, Norman W. Paton |
IEEE BigData | 9 |
| 2017 | DBaaS Cloud Capacity Planning - Accounting for Dynamic RDBMS System that Employ Clustering and Standby Architectures
Antony Higginson, Norman W. Paton, Suzanne M. Embury, Clive Bostock |
EDBT | 2 |
| 2017 | The VADA Architecture for Cost-Effective Data WranglingabstractData wrangling, the multi-faceted process by which the data required by an application is identified, extracted, cleaned and integrated, is often cumbersome and labor intensive. In this paper, we present an architecture that supports a complete data wrangling lifecycle, orchestrates components dynamically, builds on automation wherever possible, is informed by whatever data is available, refines automatically produced results in the light of feedback, takes into account the user's priorities, and supports data scientists with diverse skill sets. The architecture is demonstrated in practice for wrangling property sales and open government data. Nikolaos Konstantinou 0001, Martin Koehler, Edward Abel, Cristina Civili, Bernd Neumayr, Emanuel Sallinger, Alvaro A. A. Fernandes, Georg Gottlob, John A. Keane, Leonid Libkin, Norman W. Paton |
SIGMOD Conference | 11 |
| 2017 | Crowdsourcing for data management
Valter Crescenzi, Alvaro A. A. Fernandes, Paolo Merialdo, Norman W. Paton |
Knowl. Inf. Syst. | 4 |
| 2016 | Data Wrangling for Big Data: Challenges and OpportunitiesabstractData wrangling is the process by which the data required by an application is identified, extracted, cleaned and integrated, to yield adata set that is suitable for exploration and analysis. Although there are widely used Extract, Transform and Load (ETL) techniques and platforms, they often require manual work from technical and domain experts at different stages of the process. When confronted with the 4 V’s of big data (volume, velocity, variety and veracity),manual intervention may make ETL prohibitively expensive. This paper argues that providing cost-effective, highly-automated approaches to data wrangling involves significant research challenges,requiring fundamental changes to established areas such as data extraction,integration and cleaning, and to the ways in which these areas are brought together. Specifically, the paper discusses the importance of comprehensive support for context awareness within data wrangling, and the need for adaptive, pay-as-you-go solutions that automatically tune the wrangling process to the requirements and resources of the specific application. Tim Furche, Georg Gottlob, Leonid Libkin, Giorgio Orsi 0001, Norman W. Paton |
EDBT | 5 |
| 2016 | Structuring Linked Data Search Results Using Probabilistic Soft Logic
Duhai Alshukaili, Alvaro A. A. Fernandes, Norman W. Paton |
ISWC (1) | 3 |
| 2016 | Efficient Feedback Collection for Pay-as-you-go Source SelectionabstractTechnical developments, such as the web of data and web data extraction, combined with policy developments such as those relating to open government or open science, are leading to the availability of increasing numbers of data sources. Indeed, given these physical sources, it is then also possible to create further virtual sources that integrate, aggregate or summarise the data from the original sources. As a result, there is a plethora of data sources, from which a small subset may be able to provide the information required to support a task. The number and rate of change in the available sources is likely to make manual source selection and curation by experts impractical for many applications, leading to the need to pursue a pay-as-you-go approach, in which crowds or data consumers annotate results based on their correctness or suitability, with the resulting annotations used to inform, e.g., source selection algorithms. However, for pay-as-you-go feedback collection to be cost-effective, it may be necessary to select judiciously the data items on which feedback is to be obtained. This paper describes OLBP (Ordering and Labelling By Precision), a heuristics-based approach to the targeting of data items for feedback to support mapping and source selection tasks, where users express their preferences in terms of the trade-off between precision and recall. The proposed approach is then evaluated on two different scenarios, mapping selection with synthetic data, and source selection with real data produced by web data extraction. The results demonstrate a significant reduction in the amount of feedback required to reach user-provided objectives when using OLBP. Julio César Cortés Ríos, Norman W. Paton, Alvaro A. A. Fernandes, Khalid Belhajjame |
SSDBM | 2 |
| 2015 | Combining Syntactic and Semantic Evidence for Improving Matching over Linked Data Sources
Klitos Christodoulou, Alvaro A. A. Fernandes, Norman W. Paton |
WISE (1) | 3 |
| 2015 | Enabling community-driven information integration through clusteringabstractIt has become widely recognized that user feedback can play a fundamental role in facilitating information integration tasks, e.g., the construction of integration schema and the specification of schema mappings. While promising, existing proposals make the assumption that the users providing feedback expect the same results from the integration system. In practice, however, different users may anticipate different results, due, e.g., to their preferences or application of interest, in which case the feedback they provide may be conflicting, thereby deteriorating the quality of the services provided by the integration system. In this paper, we present clustering strategies for grouping information integration users into groups of users with similar expectations as to the results delivered by the integration system. As well as grouping information integration users, we show that clustering results can be used as inputs to a wide range of functionalities that are relevant in the context of crowd-driven information integration. Specifically, we show that clustering can be used to identify feedback of relevance to a given user by exploiting the feedback provided by other users in the same cluster. We report on evaluation exercises that assess the effectiveness of the clustering strategies we propose, and showcase the benefits community- and crowd-driven information integration can derive from clustering. Khalid Belhajjame, Norman W. Paton, Cornelia Hedeler, Alvaro A. A. Fernandes |
Distributed Parallel Databases | 2 |
| 2014 | SensorBench: benchmarking approaches to processing wireless sensor network dataabstractWireless sensor networks enable cost-effective data collection for tasks such as precision agriculture and environment monitoring. However, the resource-constrained nature of sensor nodes, which often have both limited computational capabilities and battery lifetimes, means that applications that use them must make judicious use of these resources. Research that seeks to support data intensive sensor applications has explored a range of approaches and developed many different techniques, including bespoke algorithms for specific analyses and generic sensor network query processors. However, all such proposals sit within a multi-dimensional design space, where it can be difficult to understand the implications of specific decisions and to identify optimal solutions. This paper presents a benchmark that seeks to support the systematic analysis and comparison of different techniques and platforms, enabling both development and user communities to make well informed choices. The contributions of the paper include: (i) the identification of key variables and performance metrics; (ii) the specification of experiments that explore how different types of task perform under different metrics for the controlled variables; and (iii) an application of the benchmark to investigate the behavior of several representative platforms and techniques. Ixent Galpin, Alan B. Stokes, George Valkanas, Alasdair J. G. Gray, Norman W. Paton, Alvaro A. A. Fernandes, Kai-Uwe Sattler, Dimitrios Gunopulos |
SSDBM | 5 |
| 2014 | Proactive adaptations in sensor network query processingabstractWireless sensor networks (WSN) are used by many applications for event and environmental monitoring. Due to the resource-limited nodes in WSNs, there has been much research into extending the functional lifetime of the network through energy-saving techniques. Sensor Network Query Processing (SNQP) is one such technique. SNQP uses information about a query and the WSN over which it is to be run, to generate an energy-efficient Query Execution Plan (QEP) that distributes processing in the form of QEP fragments to the nodes in the WSN. However, any QEP is likely to drain the batteries of the nodes unevenly, and, as a result, nodes used in a QEP may run out of energy when there are significant energy stocks still available in the WSN. An adaptive query processor could react to energy depletion, for example, by generating a revised plan that refrains from using the drained nodes. However, adapting only when a node has been depleted may provide few opportunities for the creation of effective new QEPs. In this paper, we introduce an approach that determines, at query compilation time, a sequence of QEPs with switch times for transitioning between successive plans with a view to extending the overall lifetime of the query. We describe how this approach has been implemented as an extension to an existing SNQP and present experimental results indicating that it can significantly increase QEP lifetimes. Alan B. Stokes, Norman W. Paton, Alvaro A. A. Fernandes |
SSDBM | 2 |
| 2013 | Incrementally improving dataspaces based on user feedback
Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
Inf. Syst. | 2 |
| 2013 | QoS-aware optimization of sensor network queries
Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton |
VLDB J. | 3 |
| 2011 | User Feedback as a First Class Citizen in Information Integration Systems
Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Cornelia Hedeler, Suzanne M. Embury |
CIDR | 2 |
| 2011 | A Semantically Enabled Service Architecture for Mashups over Streaming and Stored Data
Alasdair J. G. Gray, Raúl García-Castro, Kostis Kyzirakos, Manos Karpathiotakis, Jean-Paul Calbimonte, Kevin R. Page, Jason Sadler, Alex Frazer, Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton, Óscar Corcho, Manolis Koubarakis, David De Roure, Kirk Martinez, Asunción Gómez-Pérez |
ESWC (2) | 11 |
| 2011 | Deploying In-Network Data Analysis Techniques in Sensor NetworksabstractSensor Networks have received considerable attention recently, as they provide manifold benefits. Not only are they a means for data acquisition and monitoring of unexplored or inaccessible areas, they are also a low-cost alternative for sensing the environment, which greatly aids to better understand our surroundings. A major motivation in either occasion is to acknowledge endangering situations and take action(s) accordingly. To this end, we would like to enable data mining or analysis techniques on top or, even better, within such networks, due to the prohibitive cost of communication in this setting. In this work, we demonstrate running data mining algorithms on a set of sensors, which are of low-processing power. In addition to showcasing the execution of data analysis algorithms on resource-constrained hardware, our demo is intended to show how to take advantage of the properties of each algorithm to make better use of the sensors and their capabilities. We support the execution and monitoring of these algorithms with a graphical user interface (GUI). George Valkanas, Alexios Kotsifakos, Dimitrios Gunopulos, Ixent Galpin, Alasdair J. G. Gray, Alvaro A. A. Fernandes, Norman W. Paton |
Mobile Data Management (1) | 7 |
| 2011 | Pay-as-you-go mapping selection in dataspacesabstractThe vision of dataspaces proposes an alternative to classical data integration approaches with reduced up-front costs followed by incremental improvement on a pay-as-you-go basis. In this paper, we demonstrate DSToolkit, a system that allows users to provide feedback on results of queries posed over an integration schema. Such feedback is then used to annotate the mappings with their respective precision and recall. The system then allows a user to state the expected levels of precision (or recall) that the query results should exhibit and, in order to produce those results, the system selects those mappings that are predicted to meet the stated constraints. Cornelia Hedeler, Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Suzanne M. Embury, Lu Mao, Chenjuan Guo |
SIGMOD Conference | 3 |
| 2011 | SNEE: a query processor for wireless sensor networks
Ixent Galpin, Christian Y. A. Brenninkmeijer, Alasdair J. G. Gray, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
Distributed Parallel Databases | 6 |
| 2010 | Feedback-based annotation, selection and refinement of schema mappings for dataspacesabstractThe specification of schema mappings has proved to be time and resource consuming, and has been recognized as a critical bottleneck to the large scale deployment of data integration systems. In an attempt to address this issue, dataspaces have been proposed as a data management abstraction that aims to reduce the up-front cost required to setup a data integration system by gradually specifying schema mappings through interaction with end users in a pay-as-you-go fashion. As a step in this direction, we explore an approach for incrementally annotating schema mappings using feedback obtained from end users. In doing so, we do not expect users to examine mapping specifications; rather, they comment on results to queries evaluated using the mappings. Using annotations computed on the basis of user feedback, we present a method for selecting from the set of candidate mappings, those to be used for query evaluation considering user requirements in terms of precision and recall. In doing so, we cast mapping selection as an optimization problem. Mapping annotations may reveal that the quality of schema mappings is poor. We also show how feedback can be used to support the derivation of better quality mappings from existing mappings through refinement. An evolutionary algorithm is used to efficiently and effectively explore the large space of mappings that can be obtained through refinement. The results of evaluation exercises show the effectiveness of our solution for annotating, selecting and refining schema mappings. Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
EDBT | 2 |
| 2010 | Adaptive join processing in pipelined plansabstractIn adaptive query processing, the way in which a query is evaluated is changed in the light of feedback obtained from the environment during query evaluation. Such feedback may, for example, establish that misleading selectivity estimates were used when the query was compiled, leading to the optimizer choosing an inappropriate join order or unsuitable join algorithms. This paper describes how joins can be reordered, and the join algorithms used replaced, while they are being evaluated in pipelined plans. Where joins are reordered and/or replaced during their evaluation, the approach avoids duplicating work that has already been carried out, by resuming from where the previous plan left off. The approach has been evaluated empirically, and shown to be effective for improving query performance in the light of misleading selectivity estimates. Kwanchai Eurviriyanukul, Norman W. Paton, Alvaro A. A. Fernandes, Steven J. Lynden |
EDBT | 2 |
| 2010 | Fine-grained and efficient lineage querying of collection-based workflow provenanceabstractThe management and querying of workflow provenance data underpins a collection of activities, including the analysis of workflow results, and the debugging of workflows or services. Such activities require efficient evaluation of lineage queries over potentially complex and voluminous provenance logs. Näive implementations of lineage queries navigate provenance logs by joining tables that represent the flow of data between connected processors invoked from workflows. In this paper we provide an approach to provenance querying that: (i) avoids joins over provenance logs by using information about the workflow definition to inform the construction of queries that directly target relevant lineage results; (ii) provides fine grained provenance querying, even for workflows that create and consume collections; and (iii) scales effectively to address complex workflows, workflows with large intermediate data sets, and queries over multiple workflows. Paolo Missier, Norman W. Paton, Khalid Belhajjame |
EDBT | 2 |
| 2009 | Defining and Using Schematic Correspondences for Automatically Generating Schema Mappings
Lu Mao, Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes |
CAiSE | 3 |
| 2009 | Comprehensive Optimization of Declarative Sensor Network Queries
Ixent Galpin, Christian Y. A. Brenninkmeijer, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
SSDBM | 5 |
| 2009 | Adaptive workload allocation in query processing in autonomous heterogeneous environments
Anastasios Gounaris, Jim Smith 0001, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Paul Watson 0001 |
Distributed Parallel Databases | 3 |
| 2009 | Autonomic query parallelization using non-dedicated computers: an evaluation of adaptivity options
Norman W. Paton, Jorge Buenabad Chávez, Mengsong Chen, Vijayshankar Raman, Garret Swart, Inderpal Narang, Daniel M. Yellin, Alvaro A. A. Fernandes |
VLDB J. | 1 |
| 2008 | An Architecture for Query Optimization in Sensor NetworksabstractWe present a novel sensor network query processing architecture that (a) covers all the query optimization phases that are required to map a declarative query to executable code; and (b) does so for a more expressive query language than has heretofore been supported over sensor networks. The architecture is founded on the view that a sensor network truly is a distributed computing infrastructure, albeit a very constrained one. As such, we address the problem of how to develop a comprehensive optimizer for an expressive declarative continuous query language over acquisitional streams as one of finding extensions to a classical distributed query processing architecture that contend with the peculiarities of sensor networks as an environment for distributed computing. Ixent Galpin, Christian Y. A. Brenninkmeijer, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
ICDE | 5 |
| 2008 | A Comparative Evaluation of XML Difference Algorithms with Genomic Data
Cornelia Hedeler, Norman W. Paton |
SSDBM | 2 |
| 2008 | Automatic annotation of Web services based on workflow definitionsabstractSemantic annotations of web services can support the effective and efficient discovery of services, and guide their composition into workflows. At present, however, the practical utility of such annotations is limited by the small number of service annotations available for general use. Manual annotation of services is a time consuming and thus expensive task, so some means are required by which services can be automatically (or semi-automatically) annotated. In this paper, we show how information can be inferred about the semantics of operation parameters based on their connections to other (annotated) operation parameters within tried-and-tested workflows. Because the data links in the workflows do not necessarily contain every possible connection of compatible parameters, we can infer only constraints on the semantics of parameters. We show that despite their imprecise nature these so-called loose annotations are still of value in supporting the manual annotation task, inspecting workflows and discovering services. We also show that derived annotations for already annotated parameters are useful. By comparing existing and newly derived annotations of operation parameters, we can support the detection of errors in existing annotations, the ontology used for annotation and in workflows. The derivation mechanism has been implemented, and its practical applicability for inferring new annotations has been established through an experimental evaluation. The usefulness of the derived annotations is also demonstrated. Khalid Belhajjame, Suzanne M. Embury, Norman W. Paton, Robert Stevens 0001, Carole A. Goble |
ACM Trans. Web | 3 |
| 2006 | Practical Adaptation to Changing Resources in Grid Query ProcessingabstractGrid computational resources, as well as being heterogeneous, may also exhibit unpredictable, volatile behaviour. Therefore, query processing on the Grid needs to be adaptive in order to cope with evolving resource characteristics, such as machine load and availability. To address this challenge in a Grid environment, the non-adaptive OGSA-DQP1 system described in [1] has been enhanced with adaptive capabilities. Anastasios Gounaris, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Jim Smith 0001, Paul Watson 0001 |
ICDE | 2 |
| 2006 | Automatic Annotation of Web Services Based on Workflow Definitions
Khalid Belhajjame, Suzanne M. Embury, Norman W. Paton, Robert Stevens 0001, Carole A. Goble |
ISWC | 3 |
| 2006 | A novel approach to resource scheduling for parallel query processing on computational grids
Anastasios Gounaris, Rizos Sakellariou, Norman W. Paton, Alvaro A. A. Fernandes |
Distributed Parallel Databases | 3 |
| 2005 | Pedro Ontology Services: A Framework for Rapid Ontology Markup
Kevin L. Garwood, Phillip Lord, Helen E. Parkinson, Norman W. Paton, Carole A. Goble |
ESWC | 4 |
| 2005 | Spatio-Temporal Databases in Practice: Directly Supporting Previously Developed Land Data Using TripodabstractThis article presents a complete spatio-temporal object DBMS called Tripod, in the context of an application to support the management of previously developed land (PDL). In particular the article focuses on: the techniques that are available to realise the application data model; the facilities necessary to realise the operational semantics of updates, and the facilities that are available to identify interesting patterns of spatio-temporal change. Whilst there exist other proposals for data models and query languages for spatio-temporal object databases, Tripod is unique in that it provides: a complete implementation of an expressive spatio-temporal data model to allow modelling of (a)spatial time-varying data; the first example of an implementation of language bindings to support manipulation of stored data; and the first complete implementation of a spatio-temporal OQL that we know of. The article also focuses on the benefits arising from use of a spatio-temporal DBMS on an important category of application, namely land parcel management. Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Seung-Hyun Jeong 0003, Nassima Djafri, Keith T. Mason |
ICDE | 3 |
| 2004 | OGSA-DQP: A Service for Distributed Querying on the Grid
Mahmut Nedim Alpdemir, Arijit Mukherjee, Anastasios Gounaris, Norman W. Paton, Paul Watson 0001, Alvaro A. A. Fernandes, Desmond J. Fitzgerald |
EDBT | 4 |
| 2004 | A Generic Algorithmic Framework for Aggregation of Spatio-Temporal Data
Seung-Hyun Jeong 0003, Alvaro A. A. Fernandes, Norman W. Paton, Tony Griffiths |
SSDBM | 3 |
| 2004 | Self-monitoring query execution for adaptive query processing
Anastasios Gounaris, Norman W. Paton, Alvaro A. A. Fernandes, Rizos Sakellariou |
Data Knowl. Eng. | 2 |
| 2004 | The Tripod spatio-historical data model
Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Robert Barr |
Data Knowl. Eng. | 3 |
| 2004 | The Design, Implementation and Evaluation of an ODMG Compliant, Parallel Object Database Server
Jim Smith 0001, Sandra de F. Mendes Sampaio, Paul Watson 0001, Norman W. Paton |
Distributed Parallel Databases | 4 |
| 2003 | Grid Data Management Systems & Services
Arun Jagatheesan, Reagan W. Moore, Norman W. Paton, Paul Watson 0001 |
VLDB | 3 |
| 2003 | MOVIE: An incremental maintenance system for materialized object views
Muhammad Akhtar Ali, Alvaro A. A. Fernandes, Norman W. Paton |
Data Knowl. Eng. | 3 |
| 2003 | Estimating the quality of answers when querying over description logic ontologies
Martin Peim, Enrico Franconi, Norman W. Paton |
Data Knowl. Eng. | 3 |
| 2002 | Query Processing with Description Logic Ontologies Over Object-Wrapped DatabasesabstractThis paper presents an approach to answering queries over an ontology modelled using a description logic. The ontology acts as a global schema, providing a declarative description of the concepts of the domain, the instances of which are stored in (potentially many) object-wrapped sources. Queries are expressed using terms from the rich vocabulary of the ontology, and are translated into an equivalent calculus expression, which references only the objects available in the source databases. The query is then optimized on the basis of information from the ontology and the source databases. Distinctive features of the approach include: the use of the expressive ALCQI description logic, which supports both ontology definition and query expression; the adoption of a global-as-view approach to relating the ontology to the sources; and the use of the ontology to direct semantic optimization of queries phrased over specific sources. The approach is being developed in, and is illustrated using examples from, bioinformatics. Martin Peim, Enrico Franconi, Norman W. Paton, Carole A. Goble |
SSDBM | 3 |
| 2001 | An Experimental Performance Evaluation of Incremental Materialized View Maintenance in Object Databases
Muhammad Akhtar Ali, Norman W. Paton, Alvaro A. A. Fernandes |
DaWaK | 2 |
| 2001 | Tripod: A Comprehensive Model for Spatial and Aspatial Historical Objects
Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Keith T. Mason, Bo Huang 0001, Michael F. Worboys |
ER | 3 |
| 2001 | Information Management for Genome Level Bioinformatics
Norman W. Paton, Carole A. Goble |
VLDB | 1 |
| 2000 | Polar: An Architecture for a Parallel ODMG Compliant Object DatabaseabstractArticle Polar: an architecture for a parallel ODMG compliant object database Share on Authors: Jim Smith Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UK Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UKView Profile , Paul Watson Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UK Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UKView Profile , Sandra de F. Mendes Sampaio Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UK Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UKView Profile , Norman Paton Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UK Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UKView Profile Authors Info & Claims CIKM '00: Proceedings of the ninth international conference on Information and knowledge managementNovember 2000 Pages 352–359https://doi.org/10.1145/354756.354840Online:06 November 2000Publication History 12citation401DownloadsMetricsTotal Citations12Total Downloads401Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jim Smith 0001, Paul Watson 0001, Sandra de F. Mendes Sampaio, Norman W. Paton |
CIKM | 4 |
| 2000 | Incremental Maintenance of Materialized OQL ViewsabstractThe importance of materialized views has grown signi cantly with the advent of data warehousing and OLAP technology.This increases the relevance of solutions to the problem of incrementally maintaining materialized views.So far, most w orkon this problem has been con ned to relational settings.Proposals that apply to object databases have either used non-standard models or fallen short of providing a comprehensive framework.This paper contributes a solution to the incremental view maintenance problem for a large class of views expressed in OQL, the query language of the ODMG standard for object databases.The solution applies to immediate update propagation, and works for any update operation on views de ned over a substantial subset of ODMG types.The approach presen ted has been fully implemented and preliminary performance results are reported. Muhammad Akhtar Ali, Alvaro A. A. Fernandes, Norman W. Paton |
DOLAP | 3 |
| 2000 | User Interface Modelling with UML
Paulo Pinheiro 0001, Norman W. Paton |
EJC | 2 |
| 2000 | Query processing in DOQL: A deductive database language for the ODMG model
Pedro R. Falcone Sampaio, Norman W. Paton |
Data Knowl. Eng. | 2 |
| 1999 | Database Challenges for Genome Information in the Post Sequencing Phase
Fouzia Moussouni-Marzolf, Norman W. Paton, Andy Hayes, Steve Oliver, Carole A. Goble, Andy Brass |
DEXA | 2 |
| 1999 | Query Processing in the TAMBIS Bioinformatics Source Integration SystemabstractConducting bioinformatic analyses involves biologists in expressing requests over a range of highly heterogeneous information sources and software tools. Such activities are laborious, and require detailed knowledge of the data structures and call interfaces of the different sources. The TAMBIS (Transparent Access to Multiple Bioinformatics Information Sources) project seeks to make the diversity in data structures, call interfaces and locations of bioinformatics sources transparent to users. In TAMBIS, queries are expressed in terms of an ontology implemented using a description logic, and queries over the ontology are rewritten to a middleware level for execution over the diverse sources. The paper describes query processing in TAMBIS, focusing in particular on the way source-independent concepts in the ontology are related to source-dependent middleware calls, and describing how the planner identifies efficient ways of evaluating user queries. Norman W. Paton, Robert Stevens 0001, Patricia G. Baker, Carole A. Goble, Sean Bechhofer, Andy Brass |
SSDBM | 1 |
| 1999 | TAMBIS Online: A Bioinformatics Source Integration ToolabstractConducting bioinformatic analyses involves biologists in expressing requests over a range of heterogeneous information sources. The TAMBIS (Transparent Access to Multiple Bioinformatics Information Sources) project seeks to make the diversity in data structures, call interfaces and locations of bioinformatics sources transparent to users. TAMBIS is available at. Robert Stevens 0001, Norman W. Paton, Patricia G. Baker, Gary Ng, Carole A. Goble, Sean Bechhofer, Andy Brass |
SSDBM | 2 |
| 1999 | Active Rule Analysis and Optimisation in the Rock & Roll Deductive Object-Oriented Database
Andrew Dinn, Norman W. Paton, M. Howard Williams |
Inf. Syst. | 2 |
| 1998 | Formalizing and Validating Behavioral Models Through the Event Calculus
Oscar Díaz 0001, Norman W. Paton, Jon Iturrioz |
Inf. Syst. | 2 |
| 1997 | Stimuli and Business Policies as Modelling Constructs: Their Definition and Validation Through the Event Calculus
Oscar Díaz 0001, Norman W. Paton |
CAiSE | 2 |
| 1997 | ROCK & ROLL: A Deductive Object-Oriented Database with Active and Spatial ExtensionsabstractROCK & ROLL is a deductive object oriented database system that supports two languages, one imperative and the other deductive, both derived from the same object oriented data model. As the languages share a common type system, they can be integrated without manifesting impedance mismatches, and thus programmers can conveniently exploit both deductive and imperative features in a single application. The components of ROCK & ROLL are as follows: data model OM-OM supports a range of conventional modelling constructs, such as sets, sequences, aggregations and (both single and multiple) inheritance; deductive language ROLL-ROLL is a conventional first order deductive database language, which differs from Datalog with negation in being strictly typed (through type inference), having a structured clause base that associates rules with classes, and in that the extensional database is that of OM, rather than the relational model; and imperative language ROCK-ROCK is a conventional imperative object oriented programming language, with facilities for creating and manipulating OM objects, iteration, I/O, etc. The use of two languages has allowed us to keep the logic language ROLL simple, as facilities such as updates are not handled within ROLL, but rather in the closely integrated imperative language ROCK. The basic ROCK & ROLL system provides comprehensive modelling and programming facilities, but recent work has extended it with both active rules and spatial data types. Andrew Dinn, M. Howard Williams, Norman W. Paton |
ICDE | 3 |
| 1995 | Design and implementation of ROCK & ROLL: a deductive object-oriented database system
Maria L. Barja, Alvaro A. A. Fernandes, Norman W. Paton, M. Howard Williams, Andrew Dinn, Alia I. Abdelmoty |
Inf. Syst. | 3 |
| 1994 | Geographic Data Handling in a Deductive Object-Oriented Database
Alia I. Abdelmoty, Norman W. Paton, M. Howard Williams, Alvaro A. A. Fernandes, Maria L. Barja, Andrew Dinn |
DEXA | 2 |
| 1994 | An Effective Deductive Object-Oriented Database Through Language Integration
Maria L. Barja, Norman W. Paton, Alvaro A. A. Fernandes, M. Howard Williams, Andrew Dinn |
VLDB | 2 |
| 1993 | Combining Active Rules and Metaclasses for Enhanced Extensibility in Object-Oriented Systems
Norman W. Paton, Oscar Díaz 0001, Maria L. Barja |
Data Knowl. Eng. | 1 |
| 1991 | Rule Management in Object Oriented Databases: A Uniform Approach
Oscar Díaz 0001, Norman W. Paton, Peter M. D. Gray |
VLDB | 2 |
| 1988 | A Prolog Interface to a Functional Data Model Database
Peter M. D. Gray, David S. Moffat, Norman W. Paton |
EDBT | 3 |