Giansalvatore Mecca

dblp:m/GMecca · DBLP profile ↗
← Back
49ranked-venue papers
16as first author
1since 2021 · last 2024
0000-0002-1189-1481ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 43 · 15 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorTheory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
25 papers
Data integration and cleaning · 83% Database theory · 10% Data models and query languages · 4%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data preprocessing
data cleaning
0.832020
Cleaning data with Llunatic · VLDB J. 2020
That's All Folks! LLUNATIC Goes Open Source · Proc. VLDB Endow. 2014
The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013
Data integration and cleaning
schema mapping
0.762014
Mapping and cleaning · ICDE 2014
++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange · Proc. VLDB Endow. 2011
Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010
Data integration and cleaning › data preprocessing › data cleaning
data repair
0.632016
BART in Action: Error Generation and Empirical Evaluations of Data-Cleaning Systems · SIGMOD Conference 2016
Mapping and cleaning · ICDE 2014
The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013
Data integration and cleaning
data exchange
0.442014
++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange · Proc. VLDB Endow. 2011
Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010
Core schema mappings · SIGMOD Conference 2009
Database theory › data dependencies
chase procedure
0.312017
Benchmarking the Chase · PODS 2017
Data integration and cleaning › data preprocessing › data cleaning
interactive data cleaning
0.212016
Interactive and Deterministic Data Cleaning · SIGMOD Conference 2016
Data integration and cleaning
data transformation
0.212014
IQ-METER - An evaluation tool for data-transformation systems · ICDE 2014
Data integration and cleaning › data preprocessing › data cleaning › data repair
constraint-based data repair
0.212013
The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013
Database theory › data dependencies
equality-generating dependency
0.112010
Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010
Data integration and cleaning › data extraction
web data extraction
0.132004
An Automatic Data Grabber for Large Web Sites · VLDB 2004
RoadRunner: Towards Automatic Data Extraction from Large Web Sites · VLDB 2001
ULIXES: Building Relational Views over the Web · ICDE 1997
Data integration and cleaning › data exchange
core computation
0.112009
Core schema mappings · SIGMOD Conference 2009
Data integration and cleaning › schema mapping
schema mapping discovery
0.112009
Concise and Expressive Mappings with +Spicy · Proc. VLDB Endow. 2009
Database theory › query answering
query answering under dependencies
0.112017
Benchmarking the Chase · PODS 2017
Natural language and speech › Information extraction and text analysis
web information extraction
0.122004
Automatic information extraction from large websites · J. ACM 2004
RoadRunner: automatic data extraction from data-intensive web sites · SIGMOD Conference 2002
Mathematical optimization › combinatorial optimization
greedy algorithm
0.112015
Messing Up with BART: Error Generation for Evaluating Data-Cleaning Algorithms · Proc. VLDB Endow. 2015
Data integration and cleaning
data quality
0.112014
IQ-METER - An evaluation tool for data-transformation systems · ICDE 2014
Database theory
integrity constraints
0.012013
The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013
Natural language and speech › Information extraction and text analysis › web information extraction
wrapper induction
0.012004
Automatic information extraction from large websites · J. ACM 2004
Automata and formal languages
grammatical inference
0.012004
Automatic information extraction from large websites · J. ACM 2004
Data models and query languages
datalog
0.022001
Query Languages for Sequence Databases: Termination and Complexity · IEEE Trans. Knowl. Data Eng. 2001
Sequences, Datalog and Transducers · PODS 1995
Data integration and cleaning
web data management
0.021998
The Araneus Web-Base Management System · SIGMOD Conference 1998
To Weave the Web · VLDB 1997
Query processing and optimization
view maintenance
0.012002
Efficient Queries over Web Views · IEEE Trans. Knowl. Data Eng. 2002
Database theory › dependency theory
functional dependency
0.012010
Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010
Data integration and cleaning
data extraction
0.012001
RoadRunner: Towards Automatic Data Extraction from Large Web Sites · VLDB 2001
Data models and query languages › query language
sequence query language
0.012001
Query Languages for Sequence Databases: Termination and Complexity · IEEE Trans. Knowl. Data Eng. 2001
Query processing and optimization
query rewriting
0.012009
Concise and Expressive Mappings with +Spicy · Proc. VLDB Endow. 2009
Requirements engineering and software design
model-driven engineering
0.012000
Homer: a Model-Based CASE Tool for Data-Intensive Web Sites · SIGMOD Conference 2000
Data integration and cleaning
schema matching
0.012008
The Spicy system: towards a notion of mapping quality · SIGMOD Conference 2008
Data models and query languages
semistructured data
0.011998
The Araneus Web-Base Management System · SIGMOD Conference 1998
Data integration and cleaning
web data integration
0.011997
To Weave the Web · VLDB 1997

Methods — techniques the papers use, named apart from their topics

benchmarking · 0.8experimental evaluation · 0.6greedy algorithm · 0.4multi-hop search · 0.2SQL update queries · 0.2DFS · 0.2BFS · 0.2scalable implementation · 0.2chase-based algorithm · 0.2chase algorithm · 0.2prefix mark-up languages · 0.1grammar inference · 0.1model-theoretic semantics · 0.0logic programming · 0.0fixpoint semantics · 0.0
YearPublicationVenuePosition
2024 Similarity Measures For Incomplete Database Instances
abstract
International audience
Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti, Donatello Santoro, Enzo Veltri
EDBT2
2020 Cleaning data with Llunatic
Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro
VLDB J.2
2019 INDIANA: An interactive system for assisting database exploration
Antonio Giuzio, Giansalvatore Mecca, Elisa Quintarelli, Manuel Roveri, Donatello Santoro, Letizia Tanca
Inf. Syst.2
2017 Benchmarking the Chase
abstract
The chase is a family of algorithms used in a number of data management tasks, such as data exchange, answering queries under dependencies, query reformulation with constraints, and data cleaning. It is well established as a theoretical tool for understanding these tasks, and in addition a number of prototype systems have been developed. While individual chase-based systems and particular optimizations of the chase have been experimentally evaluated in the past, we provide the first comprehensive and publicly available benchmark---test infrastructure and a set of test scenarios---for evaluating chase implementations across a wide range of assumptions about the dependencies and the data. We used our benchmark to compare chase-based systems on data exchange and query answering tasks with one another, as well as with systems that can solve similar tasks developed in closely related communities. Our evaluation provided us with a number of new insights concerning the factors that impact the performance of chase implementations.
Michael Benedikt, George Konstantinidis 0001, Giansalvatore Mecca, Boris Motik, Paolo Papotti, Donatello Santoro, Efthymia Tsamoura
PODS3
2016 GROM: a General Rewriter of Semantic Mappings
abstract
We present GROM, a tool conceived to handle high-level schema mappings between semantic descriptions of a source and a target database. GROM rewrites mappings between the virtual, view-based semantic schemas, in terms of mappings between the two physical databases, and then executes them. The system serves the purpose of teaching two main lessons. First, designing mappings among higher-level descriptions is often simpler than working with the original schemas. Second, as soon as the view-definition language becomes more expressive, to handle, for example, negation, the mapping problem becomes extremely challenging from the technical viewpoint, so that one needs to find a proper trade-off between expressiveness and scalability.
Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente
EDBT1
2016 Interactive and Deterministic Data Cleaning
abstract
We present Falcon, an interactive, deterministic, and declarative data cleaning system, which uses SQL update queries as the language to repair data. Falcon does not rely on the existence of a set of pre-defined data quality rules. On the contrary, it encourages users to explore the data, identify possible problems, and make updates to fix them. Bootstrapped by one user update, Falcon guesses a set of possible sql update queries that can be used to repair the data. The main technical challenge addressed in this paper consists in finding a set of sql update queries that is minimal in size and at the same time fixes the largest number of errors in the data. We formalize this problem as a search in a lattice-shaped space. To guarantee that the chosen updates are semantically correct, Falcon navigates the lattice by interacting with users to gradually validate the set of sql update queries. Besides using traditional one-hop based traverse algorithms (e.g., BFS or DFS), we describe novel multi-hop search algorithms such that Falcon can dive over the lattice and conduct the search efficiently. Our novel search strategy is coupled with a number of optimization techniques to further prune the search space and efficiently maintain the lattice. We have conducted extensive experiments using both real-world and synthetic datasets to show that Falcon can effectively communicate with users in data repairing.
Enzo Veltri, Donatello Santoro, Guoliang Li 0001, Giansalvatore Mecca, Paolo Papotti, Nan Tang 0001
SIGMOD Conference5
2016 BART in Action: Error Generation and Empirical Evaluations of Data-Cleaning Systems
abstract
Repairing erroneous or conflicting data that violate a set of constraints is an important problem in data management. Many automatic or semi-automatic data-repairing algorithms have been proposed in the last few years, each with its own strengths and weaknesses. Bart is an open-source error-generation system conceived to support thorough experimental evaluations of these data-repairing systems. The demo is centered around three main lessons. To start, we discuss how generating errors in data is a complex problem, with several facets. We introduce the important notions of detectability and repairability of an error, that stand at the core of Bart. Then, we show how, by changing the features of errors, it is possible to influence quite significantly the performance of the tools. Finally, we concretely put to work five data-repairing algorithms on dirty data of various kinds generated using Bart, and discuss their performance.
Donatello Santoro, Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti
SIGMOD Conference4
2015 Ontology-based mappings
Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente
Data Knowl. Eng.1
2015 Messing Up with BART: Error Generation for Evaluating Data-Cleaning Algorithms
abstract
We study the problem of introducing errors into clean databases for the purpose of benchmarking data-cleaning algorithms. Our goal is to provide users with the highest possible level of control over the error-generation process, and at the same time develop solutions that scale to large databases. We show in the paper that the error-generation problem is surprisingly challenging, and in fact, NP-complete. To provide a scalable solution, we develop a correct and efficient greedy algorithm that sacrifices completeness, but succeeds under very reasonable assumptions. To scale to millions of tuples, the algorithm relies on several non-trivial optimizations, including a new symmetry property of data quality constraints. The trade-off between control and scalability is the main technical contribution of the paper.
Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti, Donatello Santoro
Proc. VLDB Endow.3
2014 Mapping and cleaning
abstract
We address the challenging and open problem of bringing together two crucial activities in data integration and data quality, i.e., transforming data using schema mappings, and fixing conflicts and inconsistencies using data repairing. This problem is made complex by several factors. First, schema mappings and data repairing have traditionally been considered as separate activities, and research has progressed in a largely independent way in the two fields. Second, the elegant formalizations and the algorithms that have been proposed for both tasks have had mixed fortune in scaling to large databases. In the paper, we introduce a very general notion of a mapping and cleaning scenario that incorporates a wide variety of features, like, for example, user interventions. We develop a new semantics for these scenarios that represents a conservative extension of previous semantics for schema mappings and data repairing. Based on the semantics, we introduce a chase-based algorithm to compute solutions. Appropriate care is devoted to developing a scalable implementation of the chase algorithm. To the best of our knowledge, this is the first general and scalable proposal in this direction.
Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro
ICDE2
2014 IQ-METER - An evaluation tool for data-transformation systems
abstract
We call a data-transformation system any system that maps, translates and exchanges data across different representations. Nowadays, data architects are faced with a large variety of transformation tasks, and there is huge number of different approaches and systems that were conceived to solve them. As a consequence, it is very important to be able to evaluate such alternative solutions, in order to pick up the right ones for the problem at hand. To do this, we introduce IQ-Meter, the first comprehensive tool for the evaluation of data-transformation systems. IQ-Meter can be used to benchmark, test, and even learn the best usage of data-transformation tools. It builds on a number of novel algorithms to measure the quality of outputs and the human effort required by a given system, and ultimately measures “how much intelligence” the system brings to the solution of a data-translation task.
Giansalvatore Mecca, Paolo Papotti, Donatello Santoro
ICDE1
2014 That's All Folks! LLUNATIC Goes Open Source
abstract
It is widely recognized that whenever different data sources need to be integrated into a single target database errors and inconsistencies may arise, so that there is a strong need to apply data-cleaning techniques to repair the data. Despite this need, database research has so far investigated mappings and data repairing essentially in isolation. Unfortunately, schema-mappings and data quality rules interact with each other, so that applying existing algorithms in a pipelined way -- i.e., first exchange then data, then repair the result -- does not lead to solutions even in simple settings. We present the Llunatic mapping and cleaning system, the first comprehensive proposal to handle schema mappings and data repairing in a uniform way. Llunatic is based on the intuition that transforming and cleaning data are different facets of the same problem, unified by their declarative nature. This holistic approach allows us to incorporate unique features into the system, such as configurable user interaction and a tunable trade-off between efficiency and quality of the solutions.
Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro
Proc. VLDB Endow.2
2013 Semantic-Based Mappings
Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente
ER1
2013 The LLUNATIC Data-Cleaning Framework
abstract
Data-cleaning (or data-repairing) is considered a crucial problem in many database-related tasks. It consists in making a database consistent with respect to a set of given constraints. In recent years, repairing methods have been proposed for several classes of constraints. However, these methods rely on ad hoc decisions and tend to hard-code the strategy to repair conflicting values. As a consequence, there is currently no general algorithm to solve database repairing problems that involve different kinds of constraints and different strategies to select preferred values. In this paper we develop a uniform framework to solve this problem. We propose a new semantics for repairs, and a chase-based algorithm to compute minimal solutions. We implemented the framework in a DBMS-based prototype, and we report experimental results that confirm its good scalability and superior quality in computing repairs.
Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro
Proc. VLDB Endow.2
2012 What is the IQ of your data transformation system?
abstract
Mapping and translating data across different representations is a crucial problem in information systems. Many formalisms and tools are currently used for this purpose, to the point that developers typically face a difficult question: "what is the right tool for my translation task?" In this paper, we introduce several techniques that contribute to answer this question. Among these, a fairly general definition of a data transformation system, a new and very efficient similarity measure to evaluate the outputs produced by such a system, and a metric to estimate user efforts. Based on these techniques, we are able to compare a wide range of systems on many translation tasks, to gain interesting insights about their effectiveness, and, ultimately, about their "intelligence".
Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Donatello Santoro
CIKM1
2012 Core schema mappings: Scalable core computations in data exchange
Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich
Inf. Syst.1
2011 ++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange
Bruno Marnette, Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Donatello Santoro
Proc. VLDB Endow.2
2010 Scalable Data Exchange with Functional Dependencies
abstract
The recent literature has provided a solid theoretical foundation for the use of schema mappings in data-exchange applications. Following this formalization, new algorithms have been developed to generate optimal solutions for mapping scenarios in a highly scalable way, by relying on SQL. However, these algorithms suffer from a serious drawback: they are not able to handle key constraints and functional dependencies on the target, i.e., equality generating dependencies (egds). While egds play a crucial role in the generation of optimal solutions, handling them with first-order languages is a difficult problem. In fact, we start from a negative result: it is not always possible to compute solutions for scenarios with egds using an SQL script. Then, we identify many practical cases in which this is possible, and develop a best-effort algorithm to do this. Experimental results show that our algorithm produces solutions of better quality with respect to those produced by previous algorithms, and scales nicely to large databases.
Bruno Marnette, Giansalvatore Mecca, Paolo Papotti
Proc. VLDB Endow.2
2009 Core schema mappings
abstract
Research has investigated mappings among data sources under two perspectives. On one side, there are studies of practical tools for schema mapping generation; these focus on algorithms to generate mappings based on visual specifications provided by users. On the other side, we have theoretical researches about data exchange. These study how to generate a solution - i.e., a target instance - given a set of mappings usually specified as tuple generating dependencies. However, despite the fact that the notion of a core of a data exchange solution has been formally identified as an optimal solution, there are yet no mapping systems that support core computations. In this paper we introduce several new algorithms that contribute to bridge the gap between the practice of mapping generation and the theory of data exchange. We show how, given a mapping scenario, it is possible to generate an executable script that computes core solutions for the corresponding data exchange problem. The algorithms have been implemented and tested using common runtime engines to show that they guarantee very good performances, orders of magnitudes better than those of known algorithms that compute the core as a post-processing step.
Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich
SIGMOD Conference1
2009 Concise and Expressive Mappings with +Spicy
abstract
We introduce the +Spicy mapping system. The system is based on a number of novel algorithms that contribute to increase the quality and expressiveness of mappings. +Spicy integrates the computation of core solutions in the mapping generation process in a highly efficient way, based on a natural rewriting of the given mappings. This allows for an efficient implementation of core computations using common runtime languages like SQL or XQuery and guarantees very good performances, orders of magnitude better than those of previous algorithms. The rewriting algorithm can be applied both to mappings generated by the system, or to pre-defined mappings provided as part of the input. To do this, the system was enriched with a set of expressive primitives, so that +Spicy is the first mapping system that brings together a sophisticate and expressive mapping generation algorithm with an efficient strategy to compute core solutions.
Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Marcello Buoncristiano
Proc. VLDB Endow.1
2008 Schema mapping verification: the spicy way
abstract
Schema mapping algorithms rely on value correspondences - i.e., correspondences among semantically related attributes - to produce complex transformations among data sources. These correspondences are either manually specified or suggested by separate modules called schema matchers. The quality of mappings produced by a mapping generation tool strongly depends on the quality of the input correspondences. In this paper, we introduce the Spicy system, a novel approach to the problem of verifying the quality of mappings. Spicy is based on a three-layer architecture, in which a schema matching module is used to provide input to a mapping generation module. Then, a third module, the mapping verification module, is used to check candidate mappings and choose the ones that represent better transformations of the source into the target. At the core of the system stands a new technique for comparing the structure and actual content of trees, called structural analysis. Experimental results show that, by carefully designing the comparison algorithm, it is possible to achieve both good scalability and high precision in mapping selection.
Angela Bonifati, Giansalvatore Mecca, Alessandro Pappalardo, Salvatore Raunich, Gianvito Summa
EDBT2
2008 The Spicy system: towards a notion of mapping quality
abstract
We introduce the Spicy system, a novel approach to the problem of automatically selecting the best mappings among two data sources. Known schema mapping algorithms rely on value correspondences -- i.e. correspondences among semantically related attributes -- to produce complex transformations among data sources. Spicy brings together schema matching and mapping generation tools to further automate this process. A key observation, here, is that the quality of the mappings is strongly influenced by the quality of the input correspondences. To address this problem, Spicy adopts a three-layer architecture, in which a schema matching module is used to provide input to a mapping generation module. Then, a third module, the mapping verification module, is used to check candidate mappings and choose the ones that represent better transformations of the source into the target. At the core of the system stands a new technique for comparing the structure and actual content of trees, called structural analysis. Experimental results show that our mapping discovery algorithm achieves both good scalability and high precision in mapping selection.
Angela Bonifati, Giansalvatore Mecca, Alessandro Pappalardo, Salvatore Raunich, Gianvito Summa
SIGMOD Conference2
2007 Noodles: A Clustering Engine for the Web
Giansalvatore Mecca, Salvatore Raunich, Alessandro Pappalardo, Donatello Santoro
ICWE1
2007 A new algorithm for clustering search results
Giansalvatore Mecca, Salvatore Raunich, Alessandro Pappalardo
Data Knowl. Eng.1
2004 An Automatic Data Grabber for Large Web Sites
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo, Paolo Missier
VLDB2
2004 Automatic information extraction from large websites
abstract
Information extraction from websites is nowadays a relevant problem, usually performed by software modules called wrappers. A key requirement is that the wrapper generation process should be automated to the largest extent, in order to allow for large-scale extraction tasks even in presence of changes in the underlying sites. So far, however, only semi-automatic proposals have appeared in the literature.We present a novel approach to information extraction from websites, which reconciles recent proposals for supervised wrapper induction with the more traditional field of grammar inference. Grammar inference provides a promising theoretical framework for the study of unsupervised---that is, fully automatic---wrapper generation algorithms. However, due to some unrealistic assumptions on the input, these algorithms are not practically applicable to Web information extraction tasks.The main contributions of the article stand in the definition of a class of regular languages, called the prefix mark-up languages, that abstract the structures usually found in HTML pages, and in the definition of a polynomial-time unsupervised learning algorithm for this class. The article shows that, differently from other known classes, prefix mark-up languages and the associated algorithm can be practically used for information extraction purposes.A system based on the techniques described in the article has been implemented in a working prototype. We present some experimental results on known Websites, and discuss opportunities and limitations of the proposed approach.
Valter Crescenzi, Giansalvatore Mecca
J. ACM2
2003 Automatic annotation of data extracted from large Web sites
Luigi Arlotta, Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo
WebDB3
2003 Design and development of data-intensive web sites: The araneus approach
abstract
Data-intensive Web sites are large sites based on a back-end database, with a fairly complex hypertext structure. The paper develops two main contributions: (a) a specific design methodology for data-intensive Web sites, composed of a set of steps and design transformations that lead from a conceptual specification of the domain of interest to the actual implementation of the site; (b) a tool called H omer , conceived to support the site design and implementation process, by allowing the designer to move through the various steps of the methodology, and to automate the generation of the code needed to implement the actual site.Our approach to site design is based on a clear separation between several design activities, namely database design, hypertext design, and presentation design. All these activities are carried on by using high-level models, all subsumed by an extension of the nested relational model; the mappings between the models can be nicely expressed using an extended relational algebra for nested structures. Based on the design artifacts produced during the design process, and on their representation in the algebraic framework, H omer is able to generate all the code needed for the actual generation of the site, in a completely automatic way.
Paolo Merialdo, Paolo Atzeni, Giansalvatore Mecca
ACM Trans. Internet Techn.3
2002 RoadRunner: automatic data extraction from data-intensive web sites
abstract
No abstract available.
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo
SIGMOD Conference2
2002 Efficient Queries over Web Views
abstract
Large Web sites are becoming repositories of structured information that can benefit from being viewed and queried as relational databases. However, querying these views efficiently requires new techniques. Data usually resides at a remote site and is organized as a set of related HTML documents, with network access being a primary cost factor in query evaluation. This cost can be reduced by exploiting the redundancy often found in site design. We use a simple data model, a subset of the Araneus data model, to describe the structure of a Web site. We augment the model with link and inclusion constraints that capture the redundancies in the site. We map relational views of a site to a navigational algebra and show how to use the constraints to rewrite algebraic expressions, reducing the number of network accesses. We show that similar techniques can be used to maintain materialized views over sets of HTML pages.
Giansalvatore Mecca, Alberto O. Mendelzon, Paolo Merialdo
IEEE Trans. Knowl. Data Eng.1
2001 RoadRunner: Towards Automatic Data Extraction from Large Web Sites
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo
VLDB2
2001 Query Languages for Sequence Databases: Termination and Complexity
abstract
This paper develops a query language for sequence databases, such as genome databases and text databases. Unlike relational data, queries over sequential data can easily produce infinite answer sets since the universe of sequences is infinite, even for a finite alphabet. The challenge is to develop query languages that are both highly expressive and finite. This paper develops such a language as a subset of a logic for string databases called Sequence Datalog. The main idea is to use safe recursion to control and limit unsafe recursion. The main results are the definition of a finite form of recursion, called domain-bounded recursion, and a characterization of its complexity and expressive power. Although finite, the resulting class of programs is highly expressive since its data complexity is complete for the elementary functions.
Giansalvatore Mecca, Anthony J. Bonner
IEEE Trans. Knowl. Data Eng.1
2001 Data-Intensive Web Sites: Design and Maintenance
Paolo Atzeni, Paolo Merialdo, Giansalvatore Mecca
World Wide Web3
2000 Homer: a Model-Based CASE Tool for Data-Intensive Web Sites
abstract
We present HOMER, a CASE tool for building and maintaining complex, data-intensive Web sites. In HOMER the processes of creation and maintenance of a Web site are completely based on the adoption of suitable models, to describe the various aspects of the site (content navigation structure, presentation). The development of a site does not require any code writing activity: based on the results of the design process, the system automatically creates programs to implement the site, statically and/or dynamically, as needed; also, the system does not depend on any specific tool or language: it has a modular architecture, which integrates external servers for specific tasks; finally, the system sup- ports site administrators for several maintenance activities, which can involve changes over the site at different levels.
Paolo Merialdo, Paolo Atzeni, Marco Magnante, Giansalvatore Mecca, Marco Pecorone
SIGMOD Conference4
2000 Querying Sequence Databases with Transducers
Anthony J. Bonner, Giansalvatore Mecca
Acta Informatica2
2000 Database Cooperation: Classification and Middleware Tools
abstract
We propose new criteria for the classification of systems for database cooperation, based on the nature of the component databases. In fact, the traditional criteria - heterogeneity, distribution, and autonomy - are often constraints for the design process, rather than design parameters. In this case, other features of the component databases should be addressed. We claim that a more useful classification can be based on three new criteria: (a) degree of transparency, (b) complexity of operations, and (c) level of liveliness of data. This leads us to distinguish three main categories of systems: (i) multidatabases, (ii) data warehouses, and (iii) local information systems with external data. For each of these categories, we discuss implementations based on tools offered by currently available technology
Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca
J. Database Manag.3
1999 In Search of the Lost Schema
Stéphane Grumbach, Giansalvatore Mecca
ICDT2
1999 Cut and Paste
Giansalvatore Mecca, Paolo Atzeni
J. Comput. Syst. Sci.1
1998 Design and Maintenance of Data-Intensive Web Sites
Paolo Atzeni, Giansalvatore Mecca, Paolo Merialdo
EDBT2
1998 Efficient Queries over Web Views
Giansalvatore Mecca, Alberto O. Mendelzon, Paolo Merialdo
EDBT1
1998 The Araneus Web-Base Management System
abstract
The paper describes the ARANEUS Wel-Base Management System, a system developed at University Roma Tre, which represents a proposal towards the definition of a new kind of data-repository, designed to manage Web data in the database style. We call a Web-Base a collection of data of heterogeneous nature, and more specifically: (i) highly structured data, such as the ones typically stored in relational or object-oriented database systems; (ii) semistructured data, in the Web style. We can simplify by saying that it incorporates both databases and Web sites. A Web-Base Management System (WBMS) is a system for managing such Web-bases
Giansalvatore Mecca, Paolo Atzeni, Alessandro Masci 0001, Paolo Merialdo, Giuseppe Sindoni
SIGMOD Conference1
1998 Grammars Have Exceptions
Valter Crescenzi, Giansalvatore Mecca
Inf. Syst.2
1998 Sequences, Datalog, and Transducers
Anthony J. Bonner, Giansalvatore Mecca
J. Comput. Syst. Sci.2
1997 ULIXES: Building Relational Views over the Web
abstract
The authors consider structured Web sites, those sites in which structures are so tight and regular that one can assimilate the site, from the logical viewpoint, to a conventional database. They have argued that, with respect to structured Web servers, it is possible to apply ideas from traditional database techniques, specifically with respect to design, query, and update. They focus on the querying process, which consists of associating a scheme with a server and then using this scheme to pose queries in a high level query language. To describe the scheme, they use a specific data model, called the ARANEUS Data Model (ADM). They call ADM a page oriented model, in the sense that the main construct of the model is that of a page scheme, used to describe the structure of sets of homogeneous pages in the server. ADM schemes are then offered to the user, who can query them using the ULIXES language, whose expressions produce relations as results. These are essentially relational views over Web data and can therefore be queried using any relational query language. It should be noted that the approach inherited some ideas from other proposals for query languages for the Web. However, these approaches are mainly based on a loose notion of structure, and tend to view the Web as a huge collection of unstructured objects, organized as a graph. In contrast, the approach explicitly considers structure, both in the information source (the Web) and in the derived information (the relational views).
Paolo Atzeni, Alessandro Masci 0001, Giansalvatore Mecca, Paolo Merialdo, Elena Tabet
ICDE3
1997 Cut & Paste
abstract
The paper develops EDITOR, a language for manipulating semi-structured documents, such as the ones typically available on the Web. EDITOR programs allow to search and restructure a document. They are based on two simple ideas, taken from text editors: Search” instructions are used to select regions of interest in a document, and “cut & paste” to restructure them. We study the expressive power and the complexity of these programs. We show that they are computationally complete, in the sense that any computable document restructuring can be expressed in EDITOR. We also study the complexity of a safe subclass of programs, showing that it captures exactly the class of polynomial-time restructurings. The language has been implemented in Java, and is used in the ARANEUS project to build database views over Web sites.
Paolo Atzeni, Giansalvatore Mecca
PODS2
1997 To Weave the Web
Paolo Atzeni, Giansalvatore Mecca, Paolo Merialdo
VLDB2
1997 ISALOG(¬): A Deductive Language with Negation for Complex-Object Databases with Hierarchies
Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca
Data Knowl. Eng.3
1995 Sequences, Datalog and Transducers
abstract
This paper develops a query language for sequence databases, such as genome databases and text databases. The language, called Sequence Dataiog, extends classical Datalog with interpreted function symbols for manipulating sequences. It has both a clear operational and declarative semantics, based on a new notion called the extended active domain of a database. The extended domain contains all the sequences in the database and all their subsequences. This idea leads to a clear distinction between safe and unsafe recursion over sequences: safe recursion stays inside the extended active domain, while unsafe recursion does not. By carefully limiting the amount of unsafe recursion, the paper develops a safe and expressive subset of Sequence Datalog. As part of the development, a new type of transducer is introduced, called a generalizect sequence transducer. Unsafe recursion is allowed only within these generalized transducers. Generalized transducers extend ordinary transducers by allowing them to invoke other transducers as “subroutines.” Generalized transducers can be implemented in Sequence Datalog in a straightforward way. Moreover, their introduction into the language leads to simple conditions that guarantee safety and finiteness. This paper develops two such conditions, The first condition expresses exactly the class of PTIME sequence functions; and the second expresses exactly the class of elementary sequence functions.
Giansalvatore Mecca, Anthony J. Bonner
PODS1
1993 IsaLog: A declarative language for complex objects with hierarchies
abstract
The IsaLog model and language are presented. The model has complex objects with classes, relations, and is a hierarchies. The language is strongly types and declarative. The main issue is the definition of the semantics of the language, given in three different ways that are shown to be equivalent: a model-theoretic semantics, a reduction to logic programming with function symbols, and a fixpoint semantics. Each of the semantics presents new aspects with respect to existing proposals because of the interaction of oid-invention with general is a hierarchies. The solutions are based on the explicit Skolem functors, which provide a powerful tool for manipulating object-identifiers.>
Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca
ICDE3