VLDB 2026 Research / reviewers in the wild / expert
Giansalvatore Mecca
dblp:m/GMecca
· DBLP profile ↗
49ranked-venue papers
16as first author
1since 2021 · last 2024
0000-0002-1189-1481ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 43 · 15 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorTheory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
25 papers |
Data integration and cleaning · 83% Database theory · 10% Data models and query languages · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning › data preprocessing
data cleaning |
0.8 | 3 | 2020 | Cleaning data with Llunatic · VLDB J. 2020 That's All Folks! LLUNATIC Goes Open Source · Proc. VLDB Endow. 2014 The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013 |
Data integration and cleaning
schema mapping |
0.7 | 6 | 2014 | Mapping and cleaning · ICDE 2014 ++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange · Proc. VLDB Endow. 2011 Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010 |
Data integration and cleaning › data preprocessing › data cleaning
data repair |
0.6 | 3 | 2016 | BART in Action: Error Generation and Empirical Evaluations of Data-Cleaning Systems · SIGMOD Conference 2016 Mapping and cleaning · ICDE 2014 The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013 |
Data integration and cleaning
data exchange |
0.4 | 4 | 2014 | ++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange · Proc. VLDB Endow. 2011 Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010 Core schema mappings · SIGMOD Conference 2009 |
Database theory › data dependencies
chase procedure |
0.3 | 1 | 2017 | Benchmarking the Chase · PODS 2017 |
Data integration and cleaning › data preprocessing › data cleaning
interactive data cleaning |
0.2 | 1 | 2016 | Interactive and Deterministic Data Cleaning · SIGMOD Conference 2016 |
Data integration and cleaning
data transformation |
0.2 | 1 | 2014 | IQ-METER - An evaluation tool for data-transformation systems · ICDE 2014 |
Data integration and cleaning › data preprocessing › data cleaning › data repair
constraint-based data repair |
0.2 | 1 | 2013 | The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013 |
Database theory › data dependencies
equality-generating dependency |
0.1 | 1 | 2010 | Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010 |
Data integration and cleaning › data extraction
web data extraction |
0.1 | 3 | 2004 | An Automatic Data Grabber for Large Web Sites · VLDB 2004 RoadRunner: Towards Automatic Data Extraction from Large Web Sites · VLDB 2001 ULIXES: Building Relational Views over the Web · ICDE 1997 |
Data integration and cleaning › data exchange
core computation |
0.1 | 1 | 2009 | Core schema mappings · SIGMOD Conference 2009 |
Data integration and cleaning › schema mapping
schema mapping discovery |
0.1 | 1 | 2009 | Concise and Expressive Mappings with +Spicy · Proc. VLDB Endow. 2009 |
Database theory › query answering
query answering under dependencies |
0.1 | 1 | 2017 | Benchmarking the Chase · PODS 2017 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.1 | 2 | 2004 | Automatic information extraction from large websites · J. ACM 2004 RoadRunner: automatic data extraction from data-intensive web sites · SIGMOD Conference 2002 |
Mathematical optimization › combinatorial optimization
greedy algorithm |
0.1 | 1 | 2015 | Messing Up with BART: Error Generation for Evaluating Data-Cleaning Algorithms · Proc. VLDB Endow. 2015 |
Data integration and cleaning
data quality |
0.1 | 1 | 2014 | IQ-METER - An evaluation tool for data-transformation systems · ICDE 2014 |
Database theory
integrity constraints |
0.0 | 1 | 2013 | The LLUNATIC Data-Cleaning Framework · Proc. VLDB Endow. 2013 |
Natural language and speech › Information extraction and text analysis › web information extraction
wrapper induction |
0.0 | 1 | 2004 | Automatic information extraction from large websites · J. ACM 2004 |
Automata and formal languages
grammatical inference |
0.0 | 1 | 2004 | Automatic information extraction from large websites · J. ACM 2004 |
Data models and query languages
datalog |
0.0 | 2 | 2001 | Query Languages for Sequence Databases: Termination and Complexity · IEEE Trans. Knowl. Data Eng. 2001 Sequences, Datalog and Transducers · PODS 1995 |
Data integration and cleaning
web data management |
0.0 | 2 | 1998 | The Araneus Web-Base Management System · SIGMOD Conference 1998 To Weave the Web · VLDB 1997 |
Query processing and optimization
view maintenance |
0.0 | 1 | 2002 | Efficient Queries over Web Views · IEEE Trans. Knowl. Data Eng. 2002 |
Database theory › dependency theory
functional dependency |
0.0 | 1 | 2010 | Scalable Data Exchange with Functional Dependencies · Proc. VLDB Endow. 2010 |
Data integration and cleaning
data extraction |
0.0 | 1 | 2001 | RoadRunner: Towards Automatic Data Extraction from Large Web Sites · VLDB 2001 |
Data models and query languages › query language
sequence query language |
0.0 | 1 | 2001 | Query Languages for Sequence Databases: Termination and Complexity · IEEE Trans. Knowl. Data Eng. 2001 |
Query processing and optimization
query rewriting |
0.0 | 1 | 2009 | Concise and Expressive Mappings with +Spicy · Proc. VLDB Endow. 2009 |
Requirements engineering and software design
model-driven engineering |
0.0 | 1 | 2000 | Homer: a Model-Based CASE Tool for Data-Intensive Web Sites · SIGMOD Conference 2000 |
Data integration and cleaning
schema matching |
0.0 | 1 | 2008 | The Spicy system: towards a notion of mapping quality · SIGMOD Conference 2008 |
Data models and query languages
semistructured data |
0.0 | 1 | 1998 | The Araneus Web-Base Management System · SIGMOD Conference 1998 |
Data integration and cleaning
web data integration |
0.0 | 1 | 1997 | To Weave the Web · VLDB 1997 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 0.8experimental evaluation · 0.6greedy algorithm · 0.4multi-hop search · 0.2SQL update queries · 0.2DFS · 0.2BFS · 0.2scalable implementation · 0.2chase-based algorithm · 0.2chase algorithm · 0.2prefix mark-up languages · 0.1grammar inference · 0.1model-theoretic semantics · 0.0logic programming · 0.0fixpoint semantics · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Similarity Measures For Incomplete Database InstancesabstractInternational audience Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti, Donatello Santoro, Enzo Veltri |
EDBT | 2 |
| 2020 | Cleaning data with Llunatic
Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro |
VLDB J. | 2 |
| 2019 | INDIANA: An interactive system for assisting database exploration
Antonio Giuzio, Giansalvatore Mecca, Elisa Quintarelli, Manuel Roveri, Donatello Santoro, Letizia Tanca |
Inf. Syst. | 2 |
| 2017 | Benchmarking the ChaseabstractThe chase is a family of algorithms used in a number of data management tasks, such as data exchange, answering queries under dependencies, query reformulation with constraints, and data cleaning. It is well established as a theoretical tool for understanding these tasks, and in addition a number of prototype systems have been developed. While individual chase-based systems and particular optimizations of the chase have been experimentally evaluated in the past, we provide the first comprehensive and publicly available benchmark---test infrastructure and a set of test scenarios---for evaluating chase implementations across a wide range of assumptions about the dependencies and the data. We used our benchmark to compare chase-based systems on data exchange and query answering tasks with one another, as well as with systems that can solve similar tasks developed in closely related communities. Our evaluation provided us with a number of new insights concerning the factors that impact the performance of chase implementations. Michael Benedikt, George Konstantinidis 0001, Giansalvatore Mecca, Boris Motik, Paolo Papotti, Donatello Santoro, Efthymia Tsamoura |
PODS | 3 |
| 2016 | GROM: a General Rewriter of Semantic MappingsabstractWe present GROM, a tool conceived to handle high-level schema mappings between semantic descriptions of a source and a target database. GROM rewrites mappings between the virtual, view-based semantic schemas, in terms of mappings between the two physical databases, and then executes them. The system serves the purpose of teaching two main lessons. First, designing mappings among higher-level descriptions is often simpler than working with the original schemas. Second, as soon as the view-definition language becomes more expressive, to handle, for example, negation, the mapping problem becomes extremely challenging from the technical viewpoint, so that one needs to find a proper trade-off between expressiveness and scalability. Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente |
EDBT | 1 |
| 2016 | Interactive and Deterministic Data CleaningabstractWe present Falcon, an interactive, deterministic, and declarative data cleaning system, which uses SQL update queries as the language to repair data. Falcon does not rely on the existence of a set of pre-defined data quality rules. On the contrary, it encourages users to explore the data, identify possible problems, and make updates to fix them. Bootstrapped by one user update, Falcon guesses a set of possible sql update queries that can be used to repair the data. The main technical challenge addressed in this paper consists in finding a set of sql update queries that is minimal in size and at the same time fixes the largest number of errors in the data. We formalize this problem as a search in a lattice-shaped space. To guarantee that the chosen updates are semantically correct, Falcon navigates the lattice by interacting with users to gradually validate the set of sql update queries. Besides using traditional one-hop based traverse algorithms (e.g., BFS or DFS), we describe novel multi-hop search algorithms such that Falcon can dive over the lattice and conduct the search efficiently. Our novel search strategy is coupled with a number of optimization techniques to further prune the search space and efficiently maintain the lattice. We have conducted extensive experiments using both real-world and synthetic datasets to show that Falcon can effectively communicate with users in data repairing. Enzo Veltri, Donatello Santoro, Guoliang Li 0001, Giansalvatore Mecca, Paolo Papotti, Nan Tang 0001 |
SIGMOD Conference | 5 |
| 2016 | BART in Action: Error Generation and Empirical Evaluations of Data-Cleaning SystemsabstractRepairing erroneous or conflicting data that violate a set of constraints is an important problem in data management. Many automatic or semi-automatic data-repairing algorithms have been proposed in the last few years, each with its own strengths and weaknesses. Bart is an open-source error-generation system conceived to support thorough experimental evaluations of these data-repairing systems. The demo is centered around three main lessons. To start, we discuss how generating errors in data is a complex problem, with several facets. We introduce the important notions of detectability and repairability of an error, that stand at the core of Bart. Then, we show how, by changing the features of errors, it is possible to influence quite significantly the performance of the tools. Finally, we concretely put to work five data-repairing algorithms on dirty data of various kinds generated using Bart, and discuss their performance. Donatello Santoro, Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti |
SIGMOD Conference | 4 |
| 2015 | Ontology-based mappings
Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente |
Data Knowl. Eng. | 1 |
| 2015 | Messing Up with BART: Error Generation for Evaluating Data-Cleaning AlgorithmsabstractWe study the problem of introducing errors into clean databases for the purpose of benchmarking data-cleaning algorithms. Our goal is to provide users with the highest possible level of control over the error-generation process, and at the same time develop solutions that scale to large databases. We show in the paper that the error-generation problem is surprisingly challenging, and in fact, NP-complete. To provide a scalable solution, we develop a correct and efficient greedy algorithm that sacrifices completeness, but succeeds under very reasonable assumptions. To scale to millions of tuples, the algorithm relies on several non-trivial optimizations, including a new symmetry property of data quality constraints. The trade-off between control and scalability is the main technical contribution of the paper. Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti, Donatello Santoro |
Proc. VLDB Endow. | 3 |
| 2014 | Mapping and cleaningabstractWe address the challenging and open problem of bringing together two crucial activities in data integration and data quality, i.e., transforming data using schema mappings, and fixing conflicts and inconsistencies using data repairing. This problem is made complex by several factors. First, schema mappings and data repairing have traditionally been considered as separate activities, and research has progressed in a largely independent way in the two fields. Second, the elegant formalizations and the algorithms that have been proposed for both tasks have had mixed fortune in scaling to large databases. In the paper, we introduce a very general notion of a mapping and cleaning scenario that incorporates a wide variety of features, like, for example, user interventions. We develop a new semantics for these scenarios that represents a conservative extension of previous semantics for schema mappings and data repairing. Based on the semantics, we introduce a chase-based algorithm to compute solutions. Appropriate care is devoted to developing a scalable implementation of the chase algorithm. To the best of our knowledge, this is the first general and scalable proposal in this direction. Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro |
ICDE | 2 |
| 2014 | IQ-METER - An evaluation tool for data-transformation systemsabstractWe call a data-transformation system any system that maps, translates and exchanges data across different representations. Nowadays, data architects are faced with a large variety of transformation tasks, and there is huge number of different approaches and systems that were conceived to solve them. As a consequence, it is very important to be able to evaluate such alternative solutions, in order to pick up the right ones for the problem at hand. To do this, we introduce IQ-Meter, the first comprehensive tool for the evaluation of data-transformation systems. IQ-Meter can be used to benchmark, test, and even learn the best usage of data-transformation tools. It builds on a number of novel algorithms to measure the quality of outputs and the human effort required by a given system, and ultimately measures “how much intelligence” the system brings to the solution of a data-translation task. Giansalvatore Mecca, Paolo Papotti, Donatello Santoro |
ICDE | 1 |
| 2014 | That's All Folks! LLUNATIC Goes Open SourceabstractIt is widely recognized that whenever different data sources need to be integrated into a single target database errors and inconsistencies may arise, so that there is a strong need to apply data-cleaning techniques to repair the data. Despite this need, database research has so far investigated mappings and data repairing essentially in isolation. Unfortunately, schema-mappings and data quality rules interact with each other, so that applying existing algorithms in a pipelined way -- i.e., first exchange then data, then repair the result -- does not lead to solutions even in simple settings. We present the Llunatic mapping and cleaning system, the first comprehensive proposal to handle schema mappings and data repairing in a uniform way. Llunatic is based on the intuition that transforming and cleaning data are different facets of the same problem, unified by their declarative nature. This holistic approach allows us to incorporate unique features into the system, such as configurable user interaction and a tunable trade-off between efficiency and quality of the solutions. Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro |
Proc. VLDB Endow. | 2 |
| 2013 | Semantic-Based Mappings
Giansalvatore Mecca, Guillem Rull, Donatello Santoro, Ernest Teniente |
ER | 1 |
| 2013 | The LLUNATIC Data-Cleaning FrameworkabstractData-cleaning (or data-repairing) is considered a crucial problem in many database-related tasks. It consists in making a database consistent with respect to a set of given constraints. In recent years, repairing methods have been proposed for several classes of constraints. However, these methods rely on ad hoc decisions and tend to hard-code the strategy to repair conflicting values. As a consequence, there is currently no general algorithm to solve database repairing problems that involve different kinds of constraints and different strategies to select preferred values. In this paper we develop a uniform framework to solve this problem. We propose a new semantics for repairs, and a chase-based algorithm to compute minimal solutions. We implemented the framework in a DBMS-based prototype, and we report experimental results that confirm its good scalability and superior quality in computing repairs. Floris Geerts, Giansalvatore Mecca, Paolo Papotti, Donatello Santoro |
Proc. VLDB Endow. | 2 |
| 2012 | What is the IQ of your data transformation system?abstractMapping and translating data across different representations is a crucial problem in information systems. Many formalisms and tools are currently used for this purpose, to the point that developers typically face a difficult question: "what is the right tool for my translation task?" In this paper, we introduce several techniques that contribute to answer this question. Among these, a fairly general definition of a data transformation system, a new and very efficient similarity measure to evaluate the outputs produced by such a system, and a metric to estimate user efforts. Based on these techniques, we are able to compare a wide range of systems on many translation tasks, to gain interesting insights about their effectiveness, and, ultimately, about their "intelligence". Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Donatello Santoro |
CIKM | 1 |
| 2012 | Core schema mappings: Scalable core computations in data exchange
Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich |
Inf. Syst. | 1 |
| 2011 | ++Spicy: an OpenSource Tool for Second-Generation Schema Mapping and Data Exchange
Bruno Marnette, Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Donatello Santoro |
Proc. VLDB Endow. | 2 |
| 2010 | Scalable Data Exchange with Functional DependenciesabstractThe recent literature has provided a solid theoretical foundation for the use of schema mappings in data-exchange applications. Following this formalization, new algorithms have been developed to generate optimal solutions for mapping scenarios in a highly scalable way, by relying on SQL. However, these algorithms suffer from a serious drawback: they are not able to handle key constraints and functional dependencies on the target, i.e., equality generating dependencies (egds). While egds play a crucial role in the generation of optimal solutions, handling them with first-order languages is a difficult problem. In fact, we start from a negative result: it is not always possible to compute solutions for scenarios with egds using an SQL script. Then, we identify many practical cases in which this is possible, and develop a best-effort algorithm to do this. Experimental results show that our algorithm produces solutions of better quality with respect to those produced by previous algorithms, and scales nicely to large databases. Bruno Marnette, Giansalvatore Mecca, Paolo Papotti |
Proc. VLDB Endow. | 2 |
| 2009 | Core schema mappingsabstractResearch has investigated mappings among data sources under two perspectives. On one side, there are studies of practical tools for schema mapping generation; these focus on algorithms to generate mappings based on visual specifications provided by users. On the other side, we have theoretical researches about data exchange. These study how to generate a solution - i.e., a target instance - given a set of mappings usually specified as tuple generating dependencies. However, despite the fact that the notion of a core of a data exchange solution has been formally identified as an optimal solution, there are yet no mapping systems that support core computations. In this paper we introduce several new algorithms that contribute to bridge the gap between the practice of mapping generation and the theory of data exchange. We show how, given a mapping scenario, it is possible to generate an executable script that computes core solutions for the corresponding data exchange problem. The algorithms have been implemented and tested using common runtime engines to show that they guarantee very good performances, orders of magnitudes better than those of known algorithms that compute the core as a post-processing step. Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich |
SIGMOD Conference | 1 |
| 2009 | Concise and Expressive Mappings with +SpicyabstractWe introduce the +Spicy mapping system. The system is based on a number of novel algorithms that contribute to increase the quality and expressiveness of mappings. +Spicy integrates the computation of core solutions in the mapping generation process in a highly efficient way, based on a natural rewriting of the given mappings. This allows for an efficient implementation of core computations using common runtime languages like SQL or XQuery and guarantees very good performances, orders of magnitude better than those of previous algorithms. The rewriting algorithm can be applied both to mappings generated by the system, or to pre-defined mappings provided as part of the input. To do this, the system was enriched with a set of expressive primitives, so that +Spicy is the first mapping system that brings together a sophisticate and expressive mapping generation algorithm with an efficient strategy to compute core solutions. Giansalvatore Mecca, Paolo Papotti, Salvatore Raunich, Marcello Buoncristiano |
Proc. VLDB Endow. | 1 |
| 2008 | Schema mapping verification: the spicy wayabstractSchema mapping algorithms rely on value correspondences - i.e., correspondences among semantically related attributes - to produce complex transformations among data sources. These correspondences are either manually specified or suggested by separate modules called schema matchers. The quality of mappings produced by a mapping generation tool strongly depends on the quality of the input correspondences. In this paper, we introduce the Spicy system, a novel approach to the problem of verifying the quality of mappings. Spicy is based on a three-layer architecture, in which a schema matching module is used to provide input to a mapping generation module. Then, a third module, the mapping verification module, is used to check candidate mappings and choose the ones that represent better transformations of the source into the target. At the core of the system stands a new technique for comparing the structure and actual content of trees, called structural analysis. Experimental results show that, by carefully designing the comparison algorithm, it is possible to achieve both good scalability and high precision in mapping selection. Angela Bonifati, Giansalvatore Mecca, Alessandro Pappalardo, Salvatore Raunich, Gianvito Summa |
EDBT | 2 |
| 2008 | The Spicy system: towards a notion of mapping qualityabstractWe introduce the Spicy system, a novel approach to the problem of automatically selecting the best mappings among two data sources. Known schema mapping algorithms rely on value correspondences -- i.e. correspondences among semantically related attributes -- to produce complex transformations among data sources. Spicy brings together schema matching and mapping generation tools to further automate this process. A key observation, here, is that the quality of the mappings is strongly influenced by the quality of the input correspondences. To address this problem, Spicy adopts a three-layer architecture, in which a schema matching module is used to provide input to a mapping generation module. Then, a third module, the mapping verification module, is used to check candidate mappings and choose the ones that represent better transformations of the source into the target. At the core of the system stands a new technique for comparing the structure and actual content of trees, called structural analysis. Experimental results show that our mapping discovery algorithm achieves both good scalability and high precision in mapping selection. Angela Bonifati, Giansalvatore Mecca, Alessandro Pappalardo, Salvatore Raunich, Gianvito Summa |
SIGMOD Conference | 2 |
| 2007 | Noodles: A Clustering Engine for the Web
Giansalvatore Mecca, Salvatore Raunich, Alessandro Pappalardo, Donatello Santoro |
ICWE | 1 |
| 2007 | A new algorithm for clustering search results
Giansalvatore Mecca, Salvatore Raunich, Alessandro Pappalardo |
Data Knowl. Eng. | 1 |
| 2004 | An Automatic Data Grabber for Large Web Sites
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo, Paolo Missier |
VLDB | 2 |
| 2004 | Automatic information extraction from large websitesabstractInformation extraction from websites is nowadays a relevant problem, usually performed by software modules called wrappers. A key requirement is that the wrapper generation process should be automated to the largest extent, in order to allow for large-scale extraction tasks even in presence of changes in the underlying sites. So far, however, only semi-automatic proposals have appeared in the literature.We present a novel approach to information extraction from websites, which reconciles recent proposals for supervised wrapper induction with the more traditional field of grammar inference. Grammar inference provides a promising theoretical framework for the study of unsupervised---that is, fully automatic---wrapper generation algorithms. However, due to some unrealistic assumptions on the input, these algorithms are not practically applicable to Web information extraction tasks.The main contributions of the article stand in the definition of a class of regular languages, called the prefix mark-up languages, that abstract the structures usually found in HTML pages, and in the definition of a polynomial-time unsupervised learning algorithm for this class. The article shows that, differently from other known classes, prefix mark-up languages and the associated algorithm can be practically used for information extraction purposes.A system based on the techniques described in the article has been implemented in a working prototype. We present some experimental results on known Websites, and discuss opportunities and limitations of the proposed approach. Valter Crescenzi, Giansalvatore Mecca |
J. ACM | 2 |
| 2003 | Automatic annotation of data extracted from large Web sites
Luigi Arlotta, Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo |
WebDB | 3 |
| 2003 | Design and development of data-intensive web sites: The araneus approachabstractData-intensive Web sites are large sites based on a back-end database, with a fairly complex hypertext structure. The paper develops two main contributions: (a) a specific design methodology for data-intensive Web sites, composed of a set of steps and design transformations that lead from a conceptual specification of the domain of interest to the actual implementation of the site; (b) a tool called H omer , conceived to support the site design and implementation process, by allowing the designer to move through the various steps of the methodology, and to automate the generation of the code needed to implement the actual site.Our approach to site design is based on a clear separation between several design activities, namely database design, hypertext design, and presentation design. All these activities are carried on by using high-level models, all subsumed by an extension of the nested relational model; the mappings between the models can be nicely expressed using an extended relational algebra for nested structures. Based on the design artifacts produced during the design process, and on their representation in the algebraic framework, H omer is able to generate all the code needed for the actual generation of the site, in a completely automatic way. Paolo Merialdo, Paolo Atzeni, Giansalvatore Mecca |
ACM Trans. Internet Techn. | 3 |
| 2002 | RoadRunner: automatic data extraction from data-intensive web sitesabstractNo abstract available. Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo |
SIGMOD Conference | 2 |
| 2002 | Efficient Queries over Web ViewsabstractLarge Web sites are becoming repositories of structured information that can benefit from being viewed and queried as relational databases. However, querying these views efficiently requires new techniques. Data usually resides at a remote site and is organized as a set of related HTML documents, with network access being a primary cost factor in query evaluation. This cost can be reduced by exploiting the redundancy often found in site design. We use a simple data model, a subset of the Araneus data model, to describe the structure of a Web site. We augment the model with link and inclusion constraints that capture the redundancies in the site. We map relational views of a site to a navigational algebra and show how to use the constraints to rewrite algebraic expressions, reducing the number of network accesses. We show that similar techniques can be used to maintain materialized views over sets of HTML pages. Giansalvatore Mecca, Alberto O. Mendelzon, Paolo Merialdo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2001 | RoadRunner: Towards Automatic Data Extraction from Large Web Sites
Valter Crescenzi, Giansalvatore Mecca, Paolo Merialdo |
VLDB | 2 |
| 2001 | Query Languages for Sequence Databases: Termination and ComplexityabstractThis paper develops a query language for sequence databases, such as genome databases and text databases. Unlike relational data, queries over sequential data can easily produce infinite answer sets since the universe of sequences is infinite, even for a finite alphabet. The challenge is to develop query languages that are both highly expressive and finite. This paper develops such a language as a subset of a logic for string databases called Sequence Datalog. The main idea is to use safe recursion to control and limit unsafe recursion. The main results are the definition of a finite form of recursion, called domain-bounded recursion, and a characterization of its complexity and expressive power. Although finite, the resulting class of programs is highly expressive since its data complexity is complete for the elementary functions. Giansalvatore Mecca, Anthony J. Bonner |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2001 | Data-Intensive Web Sites: Design and Maintenance
Paolo Atzeni, Paolo Merialdo, Giansalvatore Mecca |
World Wide Web | 3 |
| 2000 | Homer: a Model-Based CASE Tool for Data-Intensive Web SitesabstractWe present HOMER, a CASE tool for building and maintaining complex, data-intensive Web sites. In HOMER the processes of creation and maintenance of a Web site are completely based on the adoption of suitable models, to describe the various aspects of the site (content navigation structure, presentation). The development of a site does not require any code writing activity: based on the results of the design process, the system automatically creates programs to implement the site, statically and/or dynamically, as needed; also, the system does not depend on any specific tool or language: it has a modular architecture, which integrates external servers for specific tasks; finally, the system sup- ports site administrators for several maintenance activities, which can involve changes over the site at different levels. Paolo Merialdo, Paolo Atzeni, Marco Magnante, Giansalvatore Mecca, Marco Pecorone |
SIGMOD Conference | 4 |
| 2000 | Querying Sequence Databases with Transducers
Anthony J. Bonner, Giansalvatore Mecca |
Acta Informatica | 2 |
| 2000 | Database Cooperation: Classification and Middleware ToolsabstractWe propose new criteria for the classification of systems for database cooperation, based on the nature of the component databases. In fact, the traditional criteria - heterogeneity, distribution, and autonomy - are often constraints for the design process, rather than design parameters. In this case, other features of the component databases should be addressed. We claim that a more useful classification can be based on three new criteria: (a) degree of transparency, (b) complexity of operations, and (c) level of liveliness of data. This leads us to distinguish three main categories of systems: (i) multidatabases, (ii) data warehouses, and (iii) local information systems with external data. For each of these categories, we discuss implementations based on tools offered by currently available technology Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca |
J. Database Manag. | 3 |
| 1999 | In Search of the Lost Schema
Stéphane Grumbach, Giansalvatore Mecca |
ICDT | 2 |
| 1999 | Cut and Paste
Giansalvatore Mecca, Paolo Atzeni |
J. Comput. Syst. Sci. | 1 |
| 1998 | Design and Maintenance of Data-Intensive Web Sites
Paolo Atzeni, Giansalvatore Mecca, Paolo Merialdo |
EDBT | 2 |
| 1998 | Efficient Queries over Web Views
Giansalvatore Mecca, Alberto O. Mendelzon, Paolo Merialdo |
EDBT | 1 |
| 1998 | The Araneus Web-Base Management SystemabstractThe paper describes the ARANEUS Wel-Base Management System, a system developed at University Roma Tre, which represents a proposal towards the definition of a new kind of data-repository, designed to manage Web data in the database style. We call a Web-Base a collection of data of heterogeneous nature, and more specifically: (i) highly structured data, such as the ones typically stored in relational or object-oriented database systems; (ii) semistructured data, in the Web style. We can simplify by saying that it incorporates both databases and Web sites. A Web-Base Management System (WBMS) is a system for managing such Web-bases Giansalvatore Mecca, Paolo Atzeni, Alessandro Masci 0001, Paolo Merialdo, Giuseppe Sindoni |
SIGMOD Conference | 1 |
| 1998 | Grammars Have Exceptions
Valter Crescenzi, Giansalvatore Mecca |
Inf. Syst. | 2 |
| 1998 | Sequences, Datalog, and Transducers
Anthony J. Bonner, Giansalvatore Mecca |
J. Comput. Syst. Sci. | 2 |
| 1997 | ULIXES: Building Relational Views over the WebabstractThe authors consider structured Web sites, those sites in which structures are so tight and regular that one can assimilate the site, from the logical viewpoint, to a conventional database. They have argued that, with respect to structured Web servers, it is possible to apply ideas from traditional database techniques, specifically with respect to design, query, and update. They focus on the querying process, which consists of associating a scheme with a server and then using this scheme to pose queries in a high level query language. To describe the scheme, they use a specific data model, called the ARANEUS Data Model (ADM). They call ADM a page oriented model, in the sense that the main construct of the model is that of a page scheme, used to describe the structure of sets of homogeneous pages in the server. ADM schemes are then offered to the user, who can query them using the ULIXES language, whose expressions produce relations as results. These are essentially relational views over Web data and can therefore be queried using any relational query language. It should be noted that the approach inherited some ideas from other proposals for query languages for the Web. However, these approaches are mainly based on a loose notion of structure, and tend to view the Web as a huge collection of unstructured objects, organized as a graph. In contrast, the approach explicitly considers structure, both in the information source (the Web) and in the derived information (the relational views). Paolo Atzeni, Alessandro Masci 0001, Giansalvatore Mecca, Paolo Merialdo, Elena Tabet |
ICDE | 3 |
| 1997 | Cut & PasteabstractThe paper develops EDITOR, a language for manipulating semi-structured documents, such as the ones typically available on the Web. EDITOR programs allow to search and restructure a document. They are based on two simple ideas, taken from text editors: Search” instructions are used to select regions of interest in a document, and “cut & paste” to restructure them. We study the expressive power and the complexity of these programs. We show that they are computationally complete, in the sense that any computable document restructuring can be expressed in EDITOR. We also study the complexity of a safe subclass of programs, showing that it captures exactly the class of polynomial-time restructurings. The language has been implemented in Java, and is used in the ARANEUS project to build database views over Web sites. Paolo Atzeni, Giansalvatore Mecca |
PODS | 2 |
| 1997 | To Weave the Web
Paolo Atzeni, Giansalvatore Mecca, Paolo Merialdo |
VLDB | 2 |
| 1997 | ISALOG(¬): A Deductive Language with Negation for Complex-Object Databases with Hierarchies
Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca |
Data Knowl. Eng. | 3 |
| 1995 | Sequences, Datalog and TransducersabstractThis paper develops a query language for sequence databases, such as genome databases and text databases. The language, called Sequence Dataiog, extends classical Datalog with interpreted function symbols for manipulating sequences. It has both a clear operational and declarative semantics, based on a new notion called the extended active domain of a database. The extended domain contains all the sequences in the database and all their subsequences. This idea leads to a clear distinction between safe and unsafe recursion over sequences: safe recursion stays inside the extended active domain, while unsafe recursion does not. By carefully limiting the amount of unsafe recursion, the paper develops a safe and expressive subset of Sequence Datalog. As part of the development, a new type of transducer is introduced, called a generalizect sequence transducer. Unsafe recursion is allowed only within these generalized transducers. Generalized transducers extend ordinary transducers by allowing them to invoke other transducers as “subroutines.” Generalized transducers can be implemented in Sequence Datalog in a straightforward way. Moreover, their introduction into the language leads to simple conditions that guarantee safety and finiteness. This paper develops two such conditions, The first condition expresses exactly the class of PTIME sequence functions; and the second expresses exactly the class of elementary sequence functions. Giansalvatore Mecca, Anthony J. Bonner |
PODS | 1 |
| 1993 | IsaLog: A declarative language for complex objects with hierarchiesabstractThe IsaLog model and language are presented. The model has complex objects with classes, relations, and is a hierarchies. The language is strongly types and declarative. The main issue is the definition of the semantics of the language, given in three different ways that are shown to be equivalent: a model-theoretic semantics, a reduction to logic programming with function symbols, and a fixpoint semantics. Each of the semantics presents new aspects with respect to existing proposals because of the interaction of oid-invention with general is a hierarchies. The solutions are based on the explicit Skolem functors, which provide a powerful tool for manipulating object-identifiers.> Paolo Atzeni, Luca Cabibbo, Giansalvatore Mecca |
ICDE | 3 |