VLDB 2026 Research / reviewers in the wild / expert
Alvaro A. A. Fernandes
dblp:f/AAAFernandes · also Alvaro Adolfo Antunes Fernandes
· DBLP profile ↗
69ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-6100-7199ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 50 · 1 since 2021Artificial intelligence and machine learning · 10Systems, architecture and hardware · 7Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Data integration and cleaning · 40% Distributed and cloud data management · 20% Information retrieval · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 64% Distributed systems · 36% | |
| Computer networks
2 papers |
Internet of things and sensor networks · 100% |
Topics — the 14 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed and cloud data management
data lake |
0.4 | 1 | 2020 | Dataset Discovery in Data Lakes · ICDE 2020 |
Data integration and cleaning › data discovery
dataset discovery |
0.4 | 1 | 2020 | Dataset Discovery in Data Lakes · ICDE 2020 |
Information retrieval
similarity search |
0.4 | 1 | 2020 | Dataset Discovery in Data Lakes · ICDE 2020 |
Data integration and cleaning
data wrangling |
0.4 | 2 | 2020 | The VADA Architecture for Cost-Effective Data Wrangling · SIGMOD Conference 2017 Dataset Discovery in Data Lakes · ICDE 2020 |
Query processing and optimization › query optimization › distributed query optimization
sensor network query optimization |
0.2 | 2 | 2013 | QoS-aware optimization of sensor network queries · VLDB J. 2013 An Architecture for Query Optimization in Sensor Networks · ICDE 2008 |
Query processing and optimization
query optimization |
0.1 | 1 | 2008 | An Architecture for Query Optimization in Sensor Networks · ICDE 2008 |
Internet of things and sensor networks
wireless sensor network |
0.1 | 1 | 2008 | An Architecture for Query Optimization in Sensor Networks · ICDE 2008 |
Query processing and optimization
adaptive query processing |
0.1 | 1 | 2006 | Practical Adaptation to Changing Resources in Grid Query Processing · ICDE 2006 |
Distributed and cloud data management › large-scale data management
grid data management |
0.1 | 1 | 2006 | Practical Adaptation to Changing Resources in Grid Query Processing · ICDE 2006 |
Spatial and temporal data management › spatial data model
spatio-temporal data model |
0.1 | 1 | 2005 | Spatio-Temporal Databases in Practice: Directly Supporting Previously Developed Land Data Using Tripod · ICDE 2005 |
Spatial and temporal data management
spatio-temporal query language |
0.1 | 1 | 2005 | Spatio-Temporal Databases in Practice: Directly Supporting Previously Developed Land Data Using Tripod · ICDE 2005 |
Internet of things and sensor networks
sensor network query processing |
0.0 | 1 | 2013 | QoS-aware optimization of sensor network queries · VLDB J. 2013 |
Distributed systems › distributed database
distributed query processing |
0.0 | 1 | 2008 | An Architecture for Query Optimization in Sensor Networks · ICDE 2008 |
Database theory › deductive database
deductive object-oriented database |
0.0 | 1 | 1994 | An Effective Deductive Object-Oriented Database Through Language Integration · VLDB 1994 |
Methods — techniques the papers use, named apart from their topics
hashing · 0.4feature-based indexing · 0.4distance-based similarity · 0.4continuous query language · 0.2query feedback · 0.1precision and recall prediction · 0.1adaptivity evaluation · 0.1OGSA-DQP · 0.1language integration · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Schema mapping generation in the wild
Lacramioara Mazilu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
Inf. Syst. | 3 |
| 2021 | Incorporating Data Context to Cost-Effectively Automate End-to-End Data WranglingabstractThe process of preparing potentially large and complex data sets for further analysis or manual examination is often called data wrangling. In classical warehousing environments, the steps in such a process are carried out using Extract-Transform-Load platforms, with significant manual involvement in specifying, configuring or tuning many of them. In typical big data applications, we need to ensure that all wrangling steps, including web extraction, selection, integration and cleaning, benefit from automation wherever possible. Towards this goal, in the paper we: (i) introduce a notion of data context, which associates portions of a target schema with extensional data of types that are commonly available; (ii) define a scalable methodology to bootstrap an end-to-end data wrangling process based on data profiling; (iii) describe how data context is used to inform automation in several steps within wrangling, specifically, matching, value format transformation, data repair, and mapping generation and selection to optimise the accuracy, consistency and relevance of the result; and (iv) we evaluate the approach with real estate data and financial data, showing substantial improvements in the results of automated wrangling. Martin Koehler, Edward Abel, Alex Teodor Bogatu, Cristina Civili, Lacramioara Mazilu, Nikolaos Konstantinou 0001, Alvaro A. A. Fernandes, John A. Keane, Leonid Libkin, Norman W. Paton |
IEEE Trans. Big Data | 7 |
| 2020 | Schema Mapping Generation in the Wild: A Demonstration with Open Government DataabstractSchema mapping generation identifies how data sets can be combined to create views that are relevant to an application. Where the data sets to be combined lack declared relationships, such as foreign keys, schema mapping generation can be considered to be in the wild. In this paper, we describe an approach to schema mapping generation in the context of open government data, in particular, the London Datastore. Mapping generation is informed by inferred profiling data about the data sets and their relationships, where the data sets are made available as csv files. We outline the mapping generation algorithm, and describe a demonstration of the approach, in which the user can: (i) specify the target to be populated by the generated mappings over a collection of sources from The London Datastore; (ii) browse the generated candidate mappings and the evidence that informed their creation; and (iii) steer the mapping generation process, to make use of preferred sources and dependable profiling results. Lacramioara Mazilu, Nikolaos Konstantinou 0001, Norman W. Paton, Alvaro A. A. Fernandes |
EDBT | 4 |
| 2020 | Dataset Discovery in Data LakesabstractData analytics stands to benefit from the increasing availability of datasets that are held without their conceptual relationships being explicitly known. When collected, these datasets form a data lake from which, by processes like data wrangling, specific target datasets can be constructed that enable value- adding analytics. Given the potential vastness of such data lakes, the issue arises of how to pull out of the lake those datasets that might contribute to wrangling out a given target. We refer to this as the problem of dataset discovery in data lakes and this paper contributes an effective and efficient solution to it. Our approach uses features of the values in a dataset to construct hash- based indexes that map those features into a uniform distance space. This makes it possible to define similarity distances between features and to take those distances as measurements of relatedness w.r.t. a target table. Given the latter (and exemplar tuples), our approach returns the most related tables in the lake. We provide a detailed description of the approach and report on empirical results for two forms of relatedness (unionability and joinability) comparing them with prior work, where pertinent, and showing significant improvements in all of precision, recall, target coverage, indexing and discovery times. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
ICDE | 2 |
| 2020 | Targeted evidence collection for uncertain supplier selection
Edward Abel, Julio César Cortés Ríos, Norman W. Paton, John A. Keane, Alvaro A. A. Fernandes |
Expert Syst. Appl. | 5 |
| 2019 | SynthEdit: Format transformations by example using edit operationsabstractFormat transformation is one of the most labor intensive tasks of a data wrangling process. Recent advances in programming by example proposed synthesis algorithms that showed promising results on spreadsheet data. However, when employed on repositories consisting of multiple sources and large number of examples, such algorithms manifest scalability issues. This paper introduces a new transformation synthesis technique based on edit operations that enables efficient learning of transformation programs. Empirical results show comparable effectiveness and dramatic improvements in efficiency over the state-of-the art. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
EDBT | 2 |
| 2019 | Dynamap: Schema Mapping Generation in the WildabstractSchema mappings enable declarative and executable specification of transformations between different schematic representations of application concepts. Most work on mapping generation has assumed that the source and target schemas are well defined, e.g., with declared keys and foreign keys, and that the mapping generation processes exist to support the data engineer in the labour-intensive process of producing a high-quality integration. However, organizations increasingly have access to numerous independently produced data sets, e.g., in a data lake, with a requirement to produce rapid, best-effort integrations, without extensive manual effort. This paper introduces Dynamap, a mapping generation algorithm for such settings, where metadata about sources and the relationships between them is derived from automated data profiling, and where there may be many alternative ways of combining source tables. Our contributions include a dynamic programming algorithm for exploring the space of potential mappings, and techniques for propagating profiling data through mappings, so that the fitness of candidate mappings can be estimated. Experimental results show the effectiveness and scalability of the approach in a variety of synthetic and real-world scenarios. Lacramioara Mazilu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
SSDBM | 3 |
| 2019 | Towards Automatic Data Format Transformations: Data Wrangling at ScaleabstractAbstract Data wrangling is the process whereby data are cleaned and integrated for analysis. Data wrangling, even with tool support, is typically a labour intensive process. One aspect of data wrangling involves carrying out format transformations on attribute values, for example so that names or phone numbers are represented consistently. Recent research has developed techniques for synthesizing format transformation programs from examples of the source and target representations. This is valuable, but still requires a user to provide suitable examples, something that may be challenging in applications in which there are huge datasets or numerous data sources. In this paper, we investigate the automatic discovery of examples that can be used to synthesize format transformation programs. In particular, we propose two approaches to identifying candidate data examples and validating the transformations that are synthesized from them. The approaches are evaluated empirically using datasets from open government data. Alex Teodor Bogatu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
Comput. J. | 3 |
| 2018 | SOURCERY: User Driven Multi-Criteria Source SelectionabstractData scientists are usually interested in a subset of sources with properties that are most aligned to intended data use. The SOURCERY system supports interactive multi-criteria user-driven source selection. SOURCERY allows a user to identify criteria they consider of importance and indicate their relative importance, and seeks a source selection result aligned to the user-supplied criteria preferences. The user is given an overview of the properties of the sources that are selected along with visual analyses contextualizing the result in relation to what is theoretically possible and what is possible given the set of available sources. The system also enables a user to interactively perform iterative fine-tuning to explore how changes to preferences may impact results. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Nurzety A. Azuan, Suzanne M. Embury |
CIKM | 4 |
| 2018 | User driven multi-criteria source selectionabstractSource selection is the problem of identifying a subset of available data sources that best meet a user’s needs. In this paper we propose a user-driven approach to source selection that seeks to identify sources that are most fit for purpose. The approach employs a decision support methodology to take account of a user’s context, to allow end users to tune their preferences by specifying the relative importance between different criteria, looking to find a trade-off solution aligned with his/her preferences. The approach is extensible to incorporate diverse criteria, not drawn from a fixed set, and solutions can use a subset of the data from each selected source, rather than require that sources are used in their entirety or not at all. The paper describes and motivates the approach, presenting a methodology for modelling a user’s context, and its collection of optimisation algorithms for exploring the space of solutions, and compares and evaluates the resulting algorithms using multiple real world data sets. The experiments show how source selection results are produced that are attuned to each user’s preferences, both with respect to overall weighted utility and through faithful representation of a user’s preferences within a result, while scaling to potentially thousands of sources. Edward Abel, John A. Keane, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler, Nikolaos Konstantinou 0001, Julio César Cortés Ríos, Nurzety A. Azuan, Suzanne M. Embury |
Inf. Sci. | 4 |
| 2017 | Targeted Feedback Collection Applied to Multi-Criteria Source Selection
Julio César Cortés Ríos, Norman W. Paton, Alvaro A. A. Fernandes, Edward Abel, John A. Keane |
ADBIS | 3 |
| 2017 | Data context informed data wranglingabstractThe process of preparing potentially large and complex data sets for further analysis or manual examination is often called data wrangling. In classical warehousing environments, the steps in such a process have been carried out using Extract-Transform-Load platforms, with significant manual involvement in specifying, configuring or tuning many of them. Cost-effective data wrangling processes need to ensure that data wrangling steps benefit from automation wherever possible. In this paper, we define a methodology to fully automate an end-to-end data wrangling process incorporating data context, which associates portions of a target schema with potentially spurious extensional data of types that are commonly available. Instance-based evidence together with data profiling paves the way to inform automation in several steps within the wrangling process, specifically, matching, mapping validation, value format transformation, and data repair. The approach is evaluated with real estate data showing substantial improvements in the results of automated wrangling. Martin Koehler, Alex Teodor Bogatu, Cristina Civili, Nikolaos Konstantinou 0001, Edward Abel, Alvaro A. A. Fernandes, John A. Keane, Leonid Libkin, Norman W. Paton |
IEEE BigData | 6 |
| 2017 | Quantifying integration quality using feedback on mapping resultsabstractTraditional data integration delivers high integration quality but requires significant upfront effort because of the need for expensive experts to be involved. The pay-as-you-go approach to data integration aims to reduce this effort by relying on a bootstrap phase where algorithms replace experts in identifying or validating source-to-target semantic correspondences and executable mappings. Since the results of this phase are expected to be of lower quality, a continuous improvement phase is then launched where user feedback is collected and assimilated in order to improve the integration. It is crucial, therefore, to quantify integration quality. This paper presents a solution to this problem using feedback on mapping results as evidence. We contribute a methodology for quantifying integration quality while taking into account the inherent uncertainty of user feedback. The approach is evaluated in synthetic and real-world integration scenarios and shown to accurately and cost-effectively quantify their quality as a conditional probability. Fernando R. S. Serrano, Alvaro A. A. Fernandes, Klitos Christodoulou |
iiWAS | 2 |
| 2017 | The VADA Architecture for Cost-Effective Data WranglingabstractData wrangling, the multi-faceted process by which the data required by an application is identified, extracted, cleaned and integrated, is often cumbersome and labor intensive. In this paper, we present an architecture that supports a complete data wrangling lifecycle, orchestrates components dynamically, builds on automation wherever possible, is informed by whatever data is available, refines automatically produced results in the light of feedback, takes into account the user's priorities, and supports data scientists with diverse skill sets. The architecture is demonstrated in practice for wrangling property sales and open government data. Nikolaos Konstantinou 0001, Martin Koehler, Edward Abel, Cristina Civili, Bernd Neumayr, Emanuel Sallinger, Alvaro A. A. Fernandes, Georg Gottlob, John A. Keane, Leonid Libkin, Norman W. Paton |
SIGMOD Conference | 7 |
| 2017 | Crowdsourcing for data management
Valter Crescenzi, Alvaro A. A. Fernandes, Paolo Merialdo, Norman W. Paton |
Knowl. Inf. Syst. | 2 |
| 2016 | Structuring Linked Data Search Results Using Probabilistic Soft Logic
Duhai Alshukaili, Alvaro A. A. Fernandes, Norman W. Paton |
ISWC (1) | 2 |
| 2016 | Pay-as-you-go Data Integration: Experiences and Recurring Themes
Norman W. Paton, Khalid Belhajjame, Suzanne M. Embury, Alvaro A. A. Fernandes, Ruhaila Maskat |
SOFSEM | 4 |
| 2016 | Efficient Feedback Collection for Pay-as-you-go Source SelectionabstractTechnical developments, such as the web of data and web data extraction, combined with policy developments such as those relating to open government or open science, are leading to the availability of increasing numbers of data sources. Indeed, given these physical sources, it is then also possible to create further virtual sources that integrate, aggregate or summarise the data from the original sources. As a result, there is a plethora of data sources, from which a small subset may be able to provide the information required to support a task. The number and rate of change in the available sources is likely to make manual source selection and curation by experts impractical for many applications, leading to the need to pursue a pay-as-you-go approach, in which crowds or data consumers annotate results based on their correctness or suitability, with the resulting annotations used to inform, e.g., source selection algorithms. However, for pay-as-you-go feedback collection to be cost-effective, it may be necessary to select judiciously the data items on which feedback is to be obtained. This paper describes OLBP (Ordering and Labelling By Precision), a heuristics-based approach to the targeting of data items for feedback to support mapping and source selection tasks, where users express their preferences in terms of the trade-off between precision and recall. The proposed approach is then evaluated on two different scenarios, mapping selection with synthetic data, and source selection with real data produced by web data extraction. The results demonstrate a significant reduction in the amount of feedback required to reach user-provided objectives when using OLBP. Julio César Cortés Ríos, Norman W. Paton, Alvaro A. A. Fernandes, Khalid Belhajjame |
SSDBM | 3 |
| 2015 | Combining Syntactic and Semantic Evidence for Improving Matching over Linked Data Sources
Klitos Christodoulou, Alvaro A. A. Fernandes, Norman W. Paton |
WISE (1) | 2 |
| 2015 | Enabling community-driven information integration through clusteringabstractIt has become widely recognized that user feedback can play a fundamental role in facilitating information integration tasks, e.g., the construction of integration schema and the specification of schema mappings. While promising, existing proposals make the assumption that the users providing feedback expect the same results from the integration system. In practice, however, different users may anticipate different results, due, e.g., to their preferences or application of interest, in which case the feedback they provide may be conflicting, thereby deteriorating the quality of the services provided by the integration system. In this paper, we present clustering strategies for grouping information integration users into groups of users with similar expectations as to the results delivered by the integration system. As well as grouping information integration users, we show that clustering results can be used as inputs to a wide range of functionalities that are relevant in the context of crowd-driven information integration. Specifically, we show that clustering can be used to identify feedback of relevance to a given user by exploiting the feedback provided by other users in the same cluster. We report on evaluation exercises that assess the effectiveness of the clustering strategies we propose, and showcase the benefits community- and crowd-driven information integration can derive from clustering. Khalid Belhajjame, Norman W. Paton, Cornelia Hedeler, Alvaro A. A. Fernandes |
Distributed Parallel Databases | 4 |
| 2014 | SensorBench: benchmarking approaches to processing wireless sensor network dataabstractWireless sensor networks enable cost-effective data collection for tasks such as precision agriculture and environment monitoring. However, the resource-constrained nature of sensor nodes, which often have both limited computational capabilities and battery lifetimes, means that applications that use them must make judicious use of these resources. Research that seeks to support data intensive sensor applications has explored a range of approaches and developed many different techniques, including bespoke algorithms for specific analyses and generic sensor network query processors. However, all such proposals sit within a multi-dimensional design space, where it can be difficult to understand the implications of specific decisions and to identify optimal solutions. This paper presents a benchmark that seeks to support the systematic analysis and comparison of different techniques and platforms, enabling both development and user communities to make well informed choices. The contributions of the paper include: (i) the identification of key variables and performance metrics; (ii) the specification of experiments that explore how different types of task perform under different metrics for the controlled variables; and (iii) an application of the benchmark to investigate the behavior of several representative platforms and techniques. Ixent Galpin, Alan B. Stokes, George Valkanas, Alasdair J. G. Gray, Norman W. Paton, Alvaro A. A. Fernandes, Kai-Uwe Sattler, Dimitrios Gunopulos |
SSDBM | 6 |
| 2014 | Proactive adaptations in sensor network query processingabstractWireless sensor networks (WSN) are used by many applications for event and environmental monitoring. Due to the resource-limited nodes in WSNs, there has been much research into extending the functional lifetime of the network through energy-saving techniques. Sensor Network Query Processing (SNQP) is one such technique. SNQP uses information about a query and the WSN over which it is to be run, to generate an energy-efficient Query Execution Plan (QEP) that distributes processing in the form of QEP fragments to the nodes in the WSN. However, any QEP is likely to drain the batteries of the nodes unevenly, and, as a result, nodes used in a QEP may run out of energy when there are significant energy stocks still available in the WSN. An adaptive query processor could react to energy depletion, for example, by generating a revised plan that refrains from using the drained nodes. However, adapting only when a node has been depleted may provide few opportunities for the creation of effective new QEPs. In this paper, we introduce an approach that determines, at query compilation time, a sequence of QEPs with switch times for transitioning between successive plans with a view to extending the overall lifetime of the query. We describe how this approach has been implemented as an extension to an existing SNQP and present experimental results indicating that it can significantly increase QEP lifetimes. Alan B. Stokes, Norman W. Paton, Alvaro A. A. Fernandes |
SSDBM | 3 |
| 2013 | Efficiently implementable algebra for distributed in-network spatial analysisabstractExisting sensor network query processors (SNQPs) have demonstrated that in-network processing is an effective and efficient means of interacting with wireless sensor networks (WSNs) for data collection tasks. Inspired by these findings, this article investigates the question as to whether spatial analysis over WSNs can be built upon established distributed query processing techniques, but, here, emphasis is on the spatial aspects of sensed data, which are not adequately addressed in the existing SNQPs. By spatial analysis, we mean the ability to detect topological relationships between spatially referenced entities (e.g. whether mist intersects a vineyard or is disjoint from it) and to derive representations grounded on such relationships (e.g. the geometrical extent of that part of a vineyard that is covered by mist). To support the efficient representation, querying and manipulation of spatial data, we use an algebraic approach. We revisit a previously proposed centralized spatial algebra comprising a set of spatial data types and a comprehensive collection of operations. We have redefined and re-conceptualized the algebra for distributed evaluation and shown that it can be efficiently implemented for in-network execution. This article provides rigorous, formal definitions of the spatial data types, points, lines and regions, together with spatial-valued and topological operations over them. The article shows how the algebra can be used to characterize complex and expressive topological relationships between spatial entities and spatial phenomena that, due to their dynamic, evolving nature, cannot be represented a priori. Farhana Jabeen, Alvaro A. A. Fernandes |
Int. J. Geogr. Inf. Sci. | 2 |
| 2013 | Incrementally improving dataspaces based on user feedback
Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
Inf. Syst. | 4 |
| 2013 | QoS-aware optimization of sensor network queries
Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton |
VLDB J. | 2 |
| 2012 | Utility-driven adaptive query workload execution
Norman W. Paton, Marcelo A. T. Aragão, Alvaro A. A. Fernandes |
Future Gener. Comput. Syst. | 3 |
| 2012 | An algorithmic strategy for in-network distributed spatial analysis in wireless sensor networks
Farhana Jabeen, Alvaro A. A. Fernandes |
J. Parallel Distributed Comput. | 2 |
| 2011 | User Feedback as a First Class Citizen in Information Integration Systems
Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Cornelia Hedeler, Suzanne M. Embury |
CIDR | 3 |
| 2011 | A Semantically Enabled Service Architecture for Mashups over Streaming and Stored Data
Alasdair J. G. Gray, Raúl García-Castro, Kostis Kyzirakos, Manos Karpathiotakis, Jean-Paul Calbimonte, Kevin R. Page, Jason Sadler, Alex Frazer, Ixent Galpin, Alvaro A. A. Fernandes, Norman W. Paton, Óscar Corcho, Manolis Koubarakis, David De Roure, Kirk Martinez, Asunción Gómez-Pérez |
ESWC (2) | 10 |
| 2011 | Deploying In-Network Data Analysis Techniques in Sensor NetworksabstractSensor Networks have received considerable attention recently, as they provide manifold benefits. Not only are they a means for data acquisition and monitoring of unexplored or inaccessible areas, they are also a low-cost alternative for sensing the environment, which greatly aids to better understand our surroundings. A major motivation in either occasion is to acknowledge endangering situations and take action(s) accordingly. To this end, we would like to enable data mining or analysis techniques on top or, even better, within such networks, due to the prohibitive cost of communication in this setting. In this work, we demonstrate running data mining algorithms on a set of sensors, which are of low-processing power. In addition to showcasing the execution of data analysis algorithms on resource-constrained hardware, our demo is intended to show how to take advantage of the properties of each algorithm to make better use of the sensors and their capabilities. We support the execution and monitoring of these algorithms with a graphical user interface (GUI). George Valkanas, Alexios Kotsifakos, Dimitrios Gunopulos, Ixent Galpin, Alasdair J. G. Gray, Alvaro A. A. Fernandes, Norman W. Paton |
Mobile Data Management (1) | 6 |
| 2011 | Pay-as-you-go mapping selection in dataspacesabstractThe vision of dataspaces proposes an alternative to classical data integration approaches with reduced up-front costs followed by incremental improvement on a pay-as-you-go basis. In this paper, we demonstrate DSToolkit, a system that allows users to provide feedback on results of queries posed over an integration schema. Such feedback is then used to annotate the mappings with their respective precision and recall. The system then allows a user to state the expected levels of precision (or recall) that the query results should exhibit and, in order to produce those results, the system selects those mappings that are predicted to meet the stated constraints. Cornelia Hedeler, Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes, Suzanne M. Embury, Lu Mao, Chenjuan Guo |
SIGMOD Conference | 4 |
| 2011 | Utility functions for adaptively executing concurrent workflowsabstractAbstract Workflows are widely used in applications that require coordinated use of computational resources. Workflow definition languages typically abstract over some aspects of the way in which a workflow is to be executed, such as the level of parallelism to be used or the physical resources to be deployed. As a result, a workflow management system has the responsibility of establishing how best to map tasks within a workflow to the available resources. As workflows are typically run over shared resources, and thus face unpredictable and changing resource capabilities, there may be benefit to be derived from adapting the task‐to‐resource mapping while a workflow is executing. This paper describes the use of utility functions to express the relative merits of alternative mappings; in essence, a utility function can be used to give a score to a candidate mapping, and the exploration of alternative mappings can be cast as an optimization problem. In this approach, changing the utility function allows adaptations to be carried out with a view to meeting different objectives. The contributions of this paper include: (i) a description of how adaptive workflow execution can be expressed as an optimization problem where the objective of the adaptation is to maximize a utility function; (ii) a description of how the approach has been applied to support adaptive workflow execution in execution environments consisting of multiple resources, such as grids or clouds, in which adaptations are coordinated across multiple workflows; and (iii) an experimental evaluation of the approach with utility measures based on response time and profit using the Pegasus workflow system. Copyright © 2010 John Wiley & Sons, Ltd. Kevin Lee 0006, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes |
Concurr. Comput. Pract. Exp. | 4 |
| 2011 | SNEE: a query processor for wireless sensor networks
Ixent Galpin, Christian Y. A. Brenninkmeijer, Alasdair J. G. Gray, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
Distributed Parallel Databases | 5 |
| 2010 | Feedback-based annotation, selection and refinement of schema mappings for dataspacesabstractThe specification of schema mappings has proved to be time and resource consuming, and has been recognized as a critical bottleneck to the large scale deployment of data integration systems. In an attempt to address this issue, dataspaces have been proposed as a data management abstraction that aims to reduce the up-front cost required to setup a data integration system by gradually specifying schema mappings through interaction with end users in a pay-as-you-go fashion. As a step in this direction, we explore an approach for incrementally annotating schema mappings using feedback obtained from end users. In doing so, we do not expect users to examine mapping specifications; rather, they comment on results to queries evaluated using the mappings. Using annotations computed on the basis of user feedback, we present a method for selecting from the set of candidate mappings, those to be used for query evaluation considering user requirements in terms of precision and recall. In doing so, we cast mapping selection as an optimization problem. Mapping annotations may reveal that the quality of schema mappings is poor. We also show how feedback can be used to support the derivation of better quality mappings from existing mappings through refinement. An evolutionary algorithm is used to efficiently and effectively explore the large space of mappings that can be obtained through refinement. The results of evaluation exercises show the effectiveness of our solution for annotating, selecting and refining schema mappings. Khalid Belhajjame, Norman W. Paton, Suzanne M. Embury, Alvaro A. A. Fernandes, Cornelia Hedeler |
EDBT | 4 |
| 2010 | Adaptive join processing in pipelined plansabstractIn adaptive query processing, the way in which a query is evaluated is changed in the light of feedback obtained from the environment during query evaluation. Such feedback may, for example, establish that misleading selectivity estimates were used when the query was compiled, leading to the optimizer choosing an inappropriate join order or unsuitable join algorithms. This paper describes how joins can be reordered, and the join algorithms used replaced, while they are being evaluated in pipelined plans. Where joins are reordered and/or replaced during their evaluation, the approach avoids duplicating work that has already been carried out, by resuming from where the previous plan left off. The approach has been evaluated empirically, and shown to be effective for improving query performance in the light of misleading selectivity estimates. Kwanchai Eurviriyanukul, Norman W. Paton, Alvaro A. A. Fernandes, Steven J. Lynden |
EDBT | 3 |
| 2010 | Distributed Spatial Analysis in Wireless Sensor NetworksabstractEnvironmental monitoring is an important application area for wireless sensor networks (WSNs). An important problem for environmental WSNs is the characterization of the dynamic behaviour of transient physical phenomena over space. In the case of mote-level WSNs, a solution that is computed inside the WSN is essential for energy efficiency. In this context, the main contributions of this paper to the literature on in network processing in WSNs are threefold. The paper further develops an algebraic framework with which one can express and evaluate complex topological relationships over geometrical representations of permanent features (e.g., buildings, or geographical features such as lakes and rivers) and of transient phenomena (e.g., areas of mist over a cultivated field). The paper then describes distributed implementations of spatial-algebraic operations over the regions represented by that framework, thereby enabling identification of topological relationships between regions. Finally, the paper presents experimental evidence that the techniques described lead to efficient runtime behaviour. Taken together, these contributions constitute a further step towards enabling the high-level specification of expressive spatial analyses for efficient execution inside a WSN. Farhana Jabeen, Alvaro A. A. Fernandes |
ICPADS | 2 |
| 2009 | Defining and Using Schematic Correspondences for Automatically Generating Schema Mappings
Lu Mao, Khalid Belhajjame, Norman W. Paton, Alvaro A. A. Fernandes |
CAiSE | 4 |
| 2009 | Utility Driven Adaptive Work?ow ExecutionabstractWorkflows are widely used in applications that require coordinated use of computational resources. Workflow definition languages typically abstract over some aspects of the way in which a workflow is to be executed, such as the level of parallelism to be used or the physical resources to be deployed. As a result, a workflow management system has responsibility for establishing how best to map tasks within a workflow to the available resources. As workflows are typically run over shared resources, and thus face unpredictable and changing resource capabilties, there may be benefit to be derived from adapting the task-to-resource mapping while a workflow is executing. This paper describes the use of utility functions to express the relative merits of alternative mappings; in essence, a utility function can be used to give a score to a candidate mapping, and the exploration of alternative mappings can be cast as an optimization problem. In this approach, changing the utility function allows adaptations to be carried out with a view to meeting different objectives. The contributions of this paper include: (i) a description of how adaptive workflow execution can be expressed as an optimization problem where the objective of the adaptation is to maximize some property expressed as a utility function; (ii) a description of how the approach has been applied to support adaptive workflow execution in grids; and (iii) an experimental evaluation of the resulting approach for alternative utility measures based on response time and profit. Kevin Lee 0006, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes |
CCGRID | 4 |
| 2009 | Time-completeness trade-offs in record linkage using adaptive query processingabstractApplications that involve data integration among multiple sources often require a preliminary step of data reconciliation in order to ensure that tuples match correctly across the sources. In dynamic settings such as data mashups, however, traditional offline data reconciliation techniques that require prior availability of the data may not be applicable. The alternative, performing similarity joins at query time, is computationally expensive, while ignoring the mismatch problem altogether leads to an incomplete integration. In this paper we make the assumption that, in some dynamic integration scenarios, users may agree to trade the completeness of a join result in return for a faster computation. We explore the consequences of this assumption by proposing a novel, hybrid join algorithm that involves a combination of exact and approximate join operators, managed using adaptive query processing techniques. The algorithm is optimistic: it can switch between physical join operators multiple times throughout query processing, but it only resorts to approximate join operators when there is statistical evidence that result completeness is compromised. Our experiments show that sensible savings in join execution time can be achieved in practice, at the expense of a modest reduction in result completeness. Roald Lengu, Paolo Missier, Alvaro A. A. Fernandes, Giovanna Guerrini, Marco Mesiti |
EDBT | 3 |
| 2009 | Comprehensive Optimization of Declarative Sensor Network Queries
Ixent Galpin, Christian Y. A. Brenninkmeijer, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
SSDBM | 4 |
| 2009 | Adaptive workflow processing and execution in PegasusabstractAbstract Workflows are widely used in applications that require coordinated use of computational resources. Workflow definition languages typically abstract over some aspects of the way in which a workflow is to be executed, such as the level of parallelism to be used or the physical resources to be deployed. As a result, a workflow management system has the responsibility of establishing how best to execute a workflow given the available resources. The Pegasus workflow management system compiles abstract workflows into concrete execution plans, and has been widely used in large‐scale e‐Science applications. This paper describes an extension to Pegasus whereby resource allocation decisions are revised during workflow evaluation, in the light of feedback on the performance of jobs at runtime. The contributions of this paper include: (i) a description of how adaptive processing has been retrofitted to an existing workflow management system; (ii) a scheduling algorithm that allocates resources based on runtime performance; and (iii) an experimental evaluation of the resulting infrastructure using grid middleware over clusters. Copyright © 2009 John Wiley & Sons, Ltd. Kevin Lee 0006, Norman W. Paton, Rizos Sakellariou, Ewa Deelman, Alvaro A. A. Fernandes, Gaurang Mehta |
Concurr. Comput. Pract. Exp. | 5 |
| 2009 | Adaptive workload allocation in query processing in autonomous heterogeneous environments
Anastasios Gounaris, Jim Smith 0001, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Paul Watson 0001 |
Distributed Parallel Databases | 5 |
| 2009 | The design and implementation of OGSA-DQP: A service-based distributed query processor
Steven J. Lynden, Arijit Mukherjee, Alastair C. Hume, Alvaro A. A. Fernandes, Norman W. Paton, Rizos Sakellariou, Paul Watson 0001 |
Future Gener. Comput. Syst. | 4 |
| 2009 | Autonomic query parallelization using non-dedicated computers: an evaluation of adaptivity options
Norman W. Paton, Jorge Buenabad Chávez, Mengsong Chen, Vijayshankar Raman, Garret Swart, Inderpal Narang, Daniel M. Yellin, Alvaro A. A. Fernandes |
VLDB J. | 8 |
| 2008 | An Architecture for Query Optimization in Sensor NetworksabstractWe present a novel sensor network query processing architecture that (a) covers all the query optimization phases that are required to map a declarative query to executable code; and (b) does so for a more expressive query language than has heretofore been supported over sensor networks. The architecture is founded on the view that a sensor network truly is a distributed computing infrastructure, albeit a very constrained one. As such, we address the problem of how to develop a comprehensive optimizer for an expressive declarative continuous query language over acquisitional streams as one of finding extensions to a classical distributed query processing architecture that contend with the peculiarities of sensor networks as an environment for distributed computing. Ixent Galpin, Christian Y. A. Brenninkmeijer, Farhana Jabeen, Alvaro A. A. Fernandes, Norman W. Paton |
ICDE | 4 |
| 2006 | Practical Adaptation to Changing Resources in Grid Query ProcessingabstractGrid computational resources, as well as being heterogeneous, may also exhibit unpredictable, volatile behaviour. Therefore, query processing on the Grid needs to be adaptive in order to cope with evolving resource characteristics, such as machine load and availability. To address this challenge in a Grid environment, the non-adaptive OGSA-DQP1 system described in [1] has been enhanced with adaptive capabilities. Anastasios Gounaris, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Jim Smith 0001, Paul Watson 0001 |
ICDE | 4 |
| 2006 | A novel approach to resource scheduling for parallel query processing on computational grids
Anastasios Gounaris, Rizos Sakellariou, Norman W. Paton, Alvaro A. A. Fernandes |
Distributed Parallel Databases | 4 |
| 2005 | Spatio-Temporal Databases in Practice: Directly Supporting Previously Developed Land Data Using TripodabstractThis article presents a complete spatio-temporal object DBMS called Tripod, in the context of an application to support the management of previously developed land (PDL). In particular the article focuses on: the techniques that are available to realise the application data model; the facilities necessary to realise the operational semantics of updates, and the facilities that are available to identify interesting patterns of spatio-temporal change. Whilst there exist other proposals for data models and query languages for spatio-temporal object databases, Tripod is unique in that it provides: a complete implementation of an expressive spatio-temporal data model to allow modelling of (a)spatial time-varying data; the first example of an implementation of language bindings to support manipulation of stored data; and the first complete implementation of a spatio-temporal OQL that we know of. The article also focuses on the benefits arising from use of a spatio-temporal DBMS on an important category of application, namely land parcel management. Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Seung-Hyun Jeong 0003, Nassima Djafri, Keith T. Mason |
ICDE | 2 |
| 2005 | A Language-Based Approach for Comprehensively Supporting the In Silico Experimental ProcessabstractOne of the challenges for bioinformaticians is to approximate, in silico, tried and tested research methods used in vitro. One of the problems standing in their way is the lack of a concrete framework for designing and expressing in silico experiments that aim at being isomorphic to in vitro experiments. This paper introduces such a framework in the form of a specification language called ISXL. ISXL projects to biologists a model of in silico experiments that approximates the research method they are most familiar with, as follows. An ISXL-specified experiment (1) conforms to a conceptual model that explicitly captures the basic constituents of experiments in the empirical sciences; (2) may be defined in relation to explicit hypothesis formulation and validation rather than simply taking the form of an evidence gathering process as in alternative approaches; (3) may be long-lived and evolve over time, in the sense that there is built-in support for denoting past versions of specifications, past results, past hypotheses, past validation criteria; (4) may denote other experiments and their constituent parts, thereby reflecting the interrelatedness of scientific processes. Features (1)-(4) above are made possible by endowing ISXL with certain characteristics of a persistent workflow environment. This allows ISXL experiments to be rich in metadata without imposing too great a burden on the biologist. The metadata in turn open the way for ISXL experiments to be capable of introspection and reflection. This paper focuses on describing of ISXL conceptually and syntactically, and indicates how ISXL experiments are given a formal semantics. Ane Tröger, Alvaro A. A. Fernandes |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2004 | An Abstract Algebra for Knowledge Discovery in Databases
Luciano Gerber, Alvaro A. A. Fernandes |
ADBIS | 2 |
| 2004 | A Language for Comprehensively Supporting the In Vitro Experimental Process In SilicoabstractOne of the challenges for bioinformaticians is to approximate, in silico, tried and tested research methods used in vitro. One of the problems standing in their way is the lack of a concrete framework for designing and expressing in silico experiments that aims at being isomorphic to in vitro experiments. This paper introduces such a framework in the form of a specification language called ISXL. ISXL projects to biologists a model of in silico experiments that approximates the research method they are most familiar with, as follows. An ISXL-specified experiment (1) conforms to a conceptual model that explicitly captures the basic constituents of experiments in the empirical sciences; (2) may be defined in relation to explicit hypothesis formulation and validation rather than simply taking the form of an evidence gathering process as in alternative approaches; (3) may be long-lived and evolve over lime, in the sense that there is built-in support for denoting past versions of specifications, past results, past hypotheses, past validation criteria; (4) may denote other experiments and their constituent parts, thereby reflecting the interrelatedness of scientific processes. Features (1)-(4) above are made possible by endowing ISXL with certain characteristics of a persistent workflow environment This allows ISXL experiments to be rich in metadata without imposing too great a burden on the biologist. The metadata in turn open the way for ISXL experiments to be capable of introspection and reflection. This paper focuses on describing of ISXL conceptually and syntactically, and indicates how ISXL experiments are given a formal semantics. Ane Tröger, Alvaro A. A. Fernandes |
BIBE | 2 |
| 2004 | Seamlessly Supporting Combined Knowledge Discovery and Query Answering: A Case Study
Marcelo A. T. Aragão, Alvaro A. A. Fernandes |
Discovery Science | 2 |
| 2004 | OGSA-DQP: A Service for Distributed Querying on the Grid
Mahmut Nedim Alpdemir, Arijit Mukherjee, Anastasios Gounaris, Norman W. Paton, Paul Watson 0001, Alvaro A. A. Fernandes, Desmond J. Fitzgerald |
EDBT | 6 |
| 2004 | Logic-Based Integration of Query Answering and Knowledge Discovery
Marcelo A. T. Aragão, Alvaro A. A. Fernandes |
FQAS | 2 |
| 2004 | A Generic Algorithmic Framework for Aggregation of Spatio-Temporal Data
Seung-Hyun Jeong 0003, Alvaro A. A. Fernandes, Norman W. Paton, Tony Griffiths |
SSDBM | 2 |
| 2004 | Self-monitoring query execution for adaptive query processing
Anastasios Gounaris, Norman W. Paton, Alvaro A. A. Fernandes, Rizos Sakellariou |
Data Knowl. Eng. | 3 |
| 2004 | The Tripod spatio-historical data model
Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Robert Barr |
Data Knowl. Eng. | 2 |
| 2003 | Service-Based Distributed Querying on the Grid
Mahmut Nedim Alpdemir, Arijit Mukherjee, Norman W. Paton, Paul Watson 0001, Alvaro A. A. Fernandes, Anastasios Gounaris, Jim Smith 0001 |
ICSOC | 5 |
| 2003 | MOVIE: An incremental maintenance system for materialized object views
Muhammad Akhtar Ali, Alvaro A. A. Fernandes, Norman W. Paton |
Data Knowl. Eng. | 2 |
| 2001 | An Experimental Performance Evaluation of Incremental Materialized View Maintenance in Object Databases
Muhammad Akhtar Ali, Norman W. Paton, Alvaro A. A. Fernandes |
DaWaK | 3 |
| 2001 | Tripod: A Comprehensive Model for Spatial and Aspatial Historical Objects
Tony Griffiths, Alvaro A. A. Fernandes, Norman W. Paton, Keith T. Mason, Bo Huang 0001, Michael F. Worboys |
ER | 2 |
| 2001 | A Query Calculus for Spatio-Temporal Object DatabasesabstractThe development of any comprehensive proposal for spatio-temporal databases involves significant extensions to many aspects of a non-spatio-temporal architecture. One aspect that has received less attention than most is the development of a query calculus that can be used to provide a semantics for spatio-temporal queries and underpin an effective query optimization and evaluation framework. We show how a query calculus for spatio-temporal object databases that builds upon the monoid calculus proposed by Fegaras and Maier (2000) for ODMG-compliant database systems can be developed. The paper shows how an extension of the ODMG type system with spatial and temporal types can be accommodated into the monoid approach. It uses several queries over historical (possibly spatial) data to illustrate how, by mapping them into monoid comprehensions, the way is open for the application of a logical optimizer based on the normalization algorithm proposed by Fegaras and Maier. Tony Griffiths, Alvaro A. A. Fernandes, Nassima Djafri, Norman W. Paton |
TIME | 2 |
| 2000 | Incremental Maintenance of Materialized OQL ViewsabstractThe importance of materialized views has grown signi cantly with the advent of data warehousing and OLAP technology.This increases the relevance of solutions to the problem of incrementally maintaining materialized views.So far, most w orkon this problem has been con ned to relational settings.Proposals that apply to object databases have either used non-standard models or fallen short of providing a comprehensive framework.This paper contributes a solution to the incremental view maintenance problem for a large class of views expressed in OQL, the query language of the ODMG standard for object databases.The solution applies to immediate update propagation, and works for any update operation on views de ned over a substantial subset of ODMG types.The approach presen ted has been fully implemented and preliminary performance results are reported. Muhammad Akhtar Ali, Alvaro A. A. Fernandes, Norman W. Paton |
DOLAP | 2 |
| 1999 | Extending a deductive object-oriented database system with spatial data handling facilities
Alvaro A. A. Fernandes, Andrew Dinn, Norman W. Paton, M. Howard Williams, Olive Liew |
Inf. Softw. Technol. | 1 |
| 1997 | The formalisation of ROCK & ROLL: A deductive object-oriented database system
Alvaro A. A. Fernandes, Maria L. Barja, Norman W. Paton, M. Howard Williams |
Inf. Softw. Technol. | 1 |
| 1995 | Design and implementation of ROCK & ROLL: a deductive object-oriented database system
Maria L. Barja, Alvaro A. A. Fernandes, Norman W. Paton, M. Howard Williams, Andrew Dinn, Alia I. Abdelmoty |
Inf. Syst. | 2 |
| 1994 | Geographic Data Handling in a Deductive Object-Oriented Database
Alia I. Abdelmoty, Norman W. Paton, M. Howard Williams, Alvaro A. A. Fernandes, Maria L. Barja, Andrew Dinn |
DEXA | 4 |
| 1994 | An Effective Deductive Object-Oriented Database Through Language Integration
Maria L. Barja, Norman W. Paton, Alvaro A. A. Fernandes, M. Howard Williams, Andrew Dinn |
VLDB | 3 |
| 1992 | Approaches to deductive object-oriented databases
Alvaro A. A. Fernandes, Norman W. Paton, M. Howard Williams, Andrew Bowles |
Inf. Softw. Technol. | 1 |