Sonia Bergamaschi

dblp:b/SBergamaschi · DBLP profile ↗
← Back
65ranked-venue papers in the field
24as first author
16since 2021 · last 2026
0000-0001-8087-6587ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 42 (14 first)Knowledge Engineering, Semantic Web & Information Systems · 7 (3 first)Business Process & Enterprise Data · 5 (3 first)Information Retrieval & Web Search · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 4 (1 first)Data Mining & Knowledge Discovery · 2 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Text-to-SQL with Large Language Models: Challenges Revisited and New Dimensions
Luca Sala, Giovanni Sullutrone, Sonia Bergamaschi
DATA (1)3
2025 Evaluation of Dataframe Libraries for Data Preparation on a Single Machine
Angelo Mozzillo, Luca Zecchini, Luca Gagliardelli, Adeel Aslam, Sonia Bergamaschi, Giovanni Simonini
EDBT5
2025 Front Matter
Sonia Bergamaschi, Sourav S. Bhowmick, Philippe Bonnet, Surajit Chaudhuri, Xiaoou Ding, Hakan Ferhatosmanoglu, Raul Castro Fernandez, Jana Giceva, Madelon Hulsebos, Alexandra Meliou, Nikos Ntarmos, Themis Palpanas, John Paparrizos, Norman W. Paton, Subhadeep Sarkar 0001, Giovanni Simonini, Nesime Tatbul, Jiuqi Wei, Jingren Zhou 0001
Proc. VLDB Endow.1
2024 PRECEDE: Climate and Energy Forecasts to Support Energy Communities with Deep Learning Models
abstract
Energy optimization is crucial for environmental sustainability, as it reduces resource consumption, minimizes greenhouse gas emissions, and promotes the use of renewable energy. Efficient energy use helps combat climate change and preserves natural ecosystems for future generations. In this paper, a system to support the distribution of photovoltaic energy for Emilia Romagna Energy Communities is proposed. The system will manage and integrate large amounts of data and offer innovative services based on them for calculating climate and energy forecasts. To enable more reliable production estimates and efficient energy storage and distribution, the system will use a platform for managing and integrating data from Regional Climate Models. It will incorporate Machine Learning and Deep Learning models for accurate climate forecasts and optimize energy flows by considering consumption profiles, production forecasts, and storage characteristics. The application background, the proposed methodology, and the current challenges related to the domain will be discussed, with a particular focus on data sources and management operations.
Francesco Dattola, Pasquale Iaquinta, Miriam Iusi, Deborah Federico, Raffaele Greco, Marco Talerico, Valentina Coscarella, Luca Legato, Ivana Pellegrino, Sonia Bergamaschi, Mirko Orsini, Riccardo Martoglia, Andrea Livaldi, Abeer Jelali, Simone Sbreglia, Tommaso Ruga, Ester Zumpano, Luciano Caroprese, Camilla Lops, Sergio Montelpare, Mariano Pierantozzi, Maira Aracne
IEEE Big Data10
2024 Sensitive Topics Retrieval in Digital Libraries: A Case Study of ḥadīṯ collections
Giovanni Sullutrone, Riccardo Amerigo Vigliermo, Luca Sala, Sonia Bergamaschi
TPDL (2)4
2024 Stream-aware indexing for distributed inequality join processing
Adeel Aslam, Giovanni Simonini, Luca Gagliardelli, Luca Zecchini, Sonia Bergamaschi
Inf. Syst.5
2024 GSM: A generalized approach to Supervised Meta-blocking for scalable entity resolution
abstract
Entity Resolution (ER) constitutes a core data integration task that relies on Blocking in order to tame its quadratic time complexity. Schema-agnostic blocking achieves very high recall, requires no domain knowledge and applies to data of any structuredness and schema heterogeneity. This comes at the cost of many irrelevant candidate pairs (i.e., comparisons), which can be significantly reduced through Meta-blocking techniques, i.e., techniques that leverage the co-occurrence patterns of entities inside the blocks: first, a weighting scheme assigns a score to every pair of candidate entities in proportion to the likelihood that they are matching and then, a pruning algorithm discards the pairs with the lowest scores. Supervised Meta-blocking goes beyond this approach by combining multiple scores per comparison into a feature vector that is fed to a binary classifier. By using probabilistic classifiers, Generalized Supervised Meta-blocking associates every pair of candidates with a score that can be used: (i) by any pruning algorithm for retaining the set of candidate comparisons; and (ii) by state-of-the-art progressive ER methods to identify the most promising candidates as early as possible (when time is a critical component for the downstream applications that consume the data). For higher effectiveness, new weighting schemes are examined as features. Through an extensive experimental analysis, we identify the best pruning algorithms, their optimal sets of features as well as the minimum possible size of the training set. The resulting approaches achieve excellent performance across several established benchmark datasets.
Luca Gagliardelli, George Papadakis 0001, Giovanni Simonini, Sonia Bergamaschi, Themis Palpanas
Inf. Syst.4
2024 Determining the Largest Overlap between Tables
abstract
Both on the Web and in data lakes, it is possible to detect much redundant data in the form of largely overlapping pairs of tables. In many cases, this overlap is not accidental and provides significant information about the relatedness of the tables. Unfortunately, efficiently quantifying the overlap between two tables is not trivial. In particular, detecting their largest overlap, i.e., their largest common subtable, is a computationally challenging problem. As the information overlap may not occur in contiguous portions of the tables, only the ability to permute columns and rows can reveal it. The detection of the largest overlap can help us in relevant tasks such as the discovery of multiple coexisting versions of the same table, which can present differences in the completeness and correctness of the conveyed information. Automatically detecting these highly similar, matching tables would allow us to guarantee their consistency through data cleaning or change propagation, but also to eliminate redundancy to free up storage space or to save additional work for the editors. We present the first formal definition of this problem, and with it Sloth, our solution to efficiently detect the largest overlap between two tables. We experimentally demonstrate on real-world datasets its efficacy in solving this task, analyzing its performance and showing its impact on multiple use cases.
Luca Zecchini, Tobias Bleifuß, Giovanni Simonini, Sonia Bergamaschi, Felix Naumann
Proc. ACM Manag. Data4
2023 The REThinkWASTE data integration and analytics platform for intelligent waste management
abstract
The use of big data has grown rapidly in recent years, finding its way into various fields of use, from medicine to industry, from traffic flow optimisation to environmental protection. In the field of waste management, the European REthinkWASTE 1 project provided the opportunity for research and testing of new methods for intelligent waste management, leading to a better understanding of the problems in this area and solutions to solve them. This paper describes the results obtained in the REThinkWASTE project by exploiting MOMIS (Mediator EnvirOnment for Multiple Information Sources) [4], the open-source data integration system developed by UniMoRe and DataRiver. First, the architecture of the REThinkWASTE data integration and analysis platform and the choices made during the problem analysis phase are described. Next, the technologies used for Key Performance Indicator extraction are described and compared with the alternatives evaluated. The aim of the paper is to provide a valid example of the effectiveness of Big Data management and analysis technologies in a real-world scenario.
Andrea Livaldi, Sonia Bergamaschi, Mirko Orsini, Luca Magnotta, Riccardo Venturi, Stefano Gabri
BDCAT2
2023 A Big Data Platform for the Management of Local Energy Communities Data
abstract
In this paper we present a big data platform designed to collect and analyze energy data of Local Energy Communities, with the goal to improve the conscious use of energy by the users. The platform, originally commissioned by ENEA, is designed to acquire and manage different kinds of data (e.g., energy consumption and production, weather data, etc.) coming from multiple sources in many formats. In this work, we present the designed architecture, and several dataflows which show the real-use cases that highlight the main strengths offered by our platform.
Sonia Bergamaschi, Luca Gagliardelli
IEEE Big Data1
2023 [Vision Paper] Privacy-Preserving Data Integration
abstract
The digital transformation of different processes and the resulting availability of vast amounts of data describing people and their behaviors offer significant promise to advance multiple research areas and enhance both the public and private sectors. Exploiting the full potential of this vision requires a unified representation of different autonomous data sources to facilitate detailed data analysis capacity. Collecting and processing sensitive data about individuals leads to consideration of privacy requirements and confidentiality concerns. This vision paper provides a concise overview of the research field concerning Privacy-Preserving Data Integration (PPDI), the associated challenges, opportunities, and unexplored aspects, with the primary aim of designing a novel and comprehensive PPDI framework based on a Trusted Third-Party microservices architecture.
Lisa Trigiante, Domenico Beneventano, Sonia Bergamaschi
IEEE Big Data3
2023 HKS: Efficient Data Partitioning for Stateful Streaming
Adeel Aslam, Giovanni Simonini, Luca Gagliardelli, Angelo Mozzillo, Sonia Bergamaschi
DaWaK5
2023 BrewER: Entity Resolution On-Demand
abstract
The task of entity resolution (ER) aims to detect multiple records describing the same real-world entity in datasets and to consolidate them into a single consistent record. ER plays a fundamental role in guaranteeing good data quality, e.g., as input for data science pipelines. Yet, the traditional approach to ER requires cleaning the entire data before being able to run consistent queries on it; hence, users struggle to tackle common scenarios with limited time or resources (e.g., when the data changes frequently or the user is only interested in a portion of the dataset for the task). We previously introduced BrewER, a framework to evaluate SQL SP queries on dirty data while progressively returning results as if they were issued on cleaned data, according to a priority defined by the user. In this demonstration, we show how BrewER can be exploited to ease the burden of ER, allowing data scientists to save a significant amount of resources for their tasks.
Luca Zecchini, Giovanni Simonini, Sonia Bergamaschi, Felix Naumann
Proc. VLDB Endow.3
2022 Generalized Supervised Meta-blocking
abstract
Entity Resolution is a core data integration task that relies on Blocking to scale to large datasets. Schema-agnostic blocking achieves very high recall, requires no domain knowledge and applies to data of any structuredness and schema heterogeneity. This comes at the cost of many irrelevant candidate pairs (i.e., comparisons), which can be significantly reduced by Meta-blocking techniques that leverage the entity co-occurrence patterns inside blocks: first, pairs of candidate entities are weighted in proportion to their matching likelihood, and then, pruning discards the pairs with the lowest scores. Supervised Meta-blocking goes beyond this approach by combining multiple scores per comparison into a feature vector that is fed to a binary classifier. By using probabilistic classifiers, Generalized Supervised Meta-blocking associates every pair of candidates with a score that can be used by any pruning algorithm. For higher effectiveness, new weighting schemes are examined as features. Through extensive experiments, we identify the best pruning algorithms, their optimal sets of features, as well as the minimum possible size of the training set.
Luca Gagliardelli, George Papadakis 0001, Giovanni Simonini, Sonia Bergamaschi, Themis Palpanas
Proc. VLDB Endow.4
2022 Entity Resolution On-Demand
abstract
Entity Resolution (ER) aims to identify and merge records that refer to the same real-world entity. ER is typically employed as an expensive cleaning step on the entire data before consuming it. Yet, determining which entities are useful once cleaned depends solely on the user's application, which may need only a fraction of them. For instance, when dealing with Web data, we would like to be able to filter the entities of interest gathered from multiple sources without cleaning the entire, continuously-growing data. Similarly, when querying data lakes, we want to transform data on-demand and return the results in a timely manner---a fundamental requirement of ELT ( Extract-Load-Transform ) pipelines. We propose BrewER , a framework to evaluate SQL SP queries on dirty data while progressively returning results as if they were issued on cleaned data. BrewER tries to focus the cleaning effort on one entity at a time, following an ORDER BY predicate. Thus, it inherently supports top-k and stop-and-resume execution. For a wide range of applications, a significant amount of resources can be saved. We exhaustively evaluate and show the efficacy of BrewER on four real-world datasets.
Giovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix Naumann
Proc. VLDB Endow.3
2021 Reproducible experiments on Three-Dimensional Entity Resolution with JedAI
Georgios M. Mandilaras, George Papadakis 0001, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis, Alicia Lara-Clares, Antonio Fariña
Inf. Syst.7
2020 RulER: Scaling Up Record-level Matching Rules
Luca Gagliardelli, Giovanni Simonini, Sonia Bergamaschi
EDBT3
2020 Three-dimensional Entity Resolution with JedAI
George Papadakis 0001, Georgios M. Mandilaras, Luca Gagliardelli, Giovanni Simonini, Emmanouil Thanos, George Giannakopoulos, Sonia Bergamaschi, Themis Palpanas, Manolis Koubarakis
Inf. Syst.7
2019 SparkER: Scaling Entity Resolution in Spark
abstract
We present SparkER, an ER tool that can scale practitioners' favorite ER algorithms.SparkER has been devised to take full advantage of parallel and distributed computation as well (running on top of Apache Spark).The first SparkER version was focused on the blocking step and implements both schema-agnostic and Blast meta-blocking approaches (i.e. the state-of-the-art ones); a GUI for SparkER, to let non-expert users to use it in an unsupervised mode, was developed.The new version of SparkER to be shown in this demo, extends significantly the tool.Entity matching and Entity Clustering modules have been added.Moreover, in addition to the completely unsupervised mode of the first version, a supervised mode has been added.The user can be assisted in supervising the entire process and in injecting his knowledge in order to achieve the best result.During the demonstration, attendees will be shown how SparkER can significantly help in devising and debugging ER algorithms.
Luca Gagliardelli, Giovanni Simonini, Domenico Beneventano, Sonia Bergamaschi
EDBT4
2019 Computing inter-document similarity with Context Semantic Analysis
Fabio Benedetti, Domenico Beneventano, Sonia Bergamaschi, Giovanni Simonini
Inf. Syst.3
2019 Scaling entity resolution: A loosely schema-aware approach
Giovanni Simonini, Luca Gagliardelli, Sonia Bergamaschi, H. V. Jagadish
Inf. Syst.3
2019 Schema-Agnostic Progressive Entity Resolution
abstract
Entity Resolution (ER) is the task of finding entity profiles that correspond to the same real-world entity. Progressive ER aims to efficiently resolve large datasets when limited time and/or computational resources are available. In practice, its goal is to provide the best possible partial solution by approximating the optimal comparison order of the entity profiles. So far, Progressive ER has only been examined in the context of structured (relational) data sources, as the existing methods rely on schema knowledge to save unnecessary comparisons: they restrict their search space to similar entities with the help of schema-based blocking keys (i.e., signatures that represent the entity profiles). As a result, these solutions are not applicable in Big Data integration applications, which involve large and heterogeneous datasets, such as relational and RDF databases, JSON files, Web corpus etc. To cover this gap, we propose a family of schema-agnostic Progressive ER methods, which do not require schema information, thus applying to heterogeneous data sources of any schema variety. First, we introduce two naïve schema-agnostic methods, showing that straightforward solutions exhibit a poor performance that does not scale well to large volumes of data. Then, we propose four different advanced methods. Through an extensive experimental evaluation over 7 real-world, established datasets, we show that all the advanced methods outperform to a significant extent both the naïve and the state-of-the-art schema-based ones. We also investigate the relative performance of the advanced methods, providing guidelines on the method selection.
Giovanni Simonini, George Papadakis 0001, Themis Palpanas, Sonia Bergamaschi
IEEE Trans. Knowl. Data Eng.4
2018 Schema-Agnostic Progressive Entity Resolution
abstract
Entity Resolution (ER) is the task of finding entity profiles that correspond to the same real-world entity. Progressive ER aims to efficiently resolve large datasets when limited time and/or computational resources are available. In practice, its goal is to provide the best possible partial solution by approximating the optimal comparison order of the entity profiles. So far, Progressive ER has only been examined in the context of structured (relational) data sources, as the existing methods rely on schema knowledge to save unnecessary comparisons: they restrict their search space to similar entities with the help of schema-based blocking keys (i.e., signatures that represent the entity profiles). As a result, these solutions are not applicable in Big Data integration applications, which involve large and heterogeneous datasets, such as relational and RDF databases, JSON files, Web corpus etc. To cover this gap, we propose a family of schema-agnostic Progressive ER methods, which do not require schema information, thus applying to heterogeneous data sources of any schema variety. First, we introduce a naïve schema-agnostic method, showing that the straightforward solution exhibits a poor performance that does not scale well to large volumes of data. Then, we propose three different advanced methods. Through an extensive experimental evaluation over 7 real-world, established datasets, we show that all the advanced methods outperform to a significant extent both the naïve and the state-of-the-art schema-based ones. We also investigate the relative performance of the advanced methods, providing guidelines on the method selection.
Giovanni Simonini, George Papadakis 0001, Themis Palpanas, Sonia Bergamaschi
ICDE4
2016 Context Semantic Analysis: A Knowledge-Based Technique for Computing Inter-document Similarity
Fabio Benedetti, Domenico Beneventano, Sonia Bergamaschi
SISAP3
2016 Combining user and database perspective for solving keyword queries over relational databases
Sonia Bergamaschi, Francesco Guerra 0001, Matteo Interlandi, Raquel Trillo Lado, Yannis Velegrakis
Inf. Syst.1
2016 BLAST: a Loosely Schema-aware Meta-blocking Approach for Entity Resolution
abstract
Identifying records that refer to the same entity is a fundamental step for data integration. Since it is prohibitively expensive to compare every pair of records, blocking techniques are typically employed to reduce the complexity of this task. These techniques partition records into blocks and limit the comparison to records co-occurring in a block. Generally, to deal with highly heterogeneous and noisy data (e.g. semi-structured data of the Web), these techniques rely on redundancy to reduce the chance of missing matches. Meta-blocking is the task of restructuring blocks generated by redundancy-based blocking techniques, removing superfluous comparisons. Existing meta-blocking approaches rely exclusively on schema-agnostic features. In this paper, we demonstrate how "loose" schema information (i.e., statistics collected directly from the data) can be exploited to enhance the quality of the blocks in a holistic loosely schema-aware (meta-)blocking approach that can be used to speed up your favorite Entity Resolution algorithm. We call it B last (Blocking with Loosely-Aware Schema Techniques). We show how B last can automatically extract this loose information by adopting a LSH-based step for efficiently scaling to large datasets. We experimentally demonstrate, on real-world datasets, how B last outperforms the state-of-the-art unsupervised meta-blocking approaches, and, in many cases, also the supervised one.
Giovanni Simonini, Sonia Bergamaschi, H. V. Jagadish
Proc. VLDB Endow.2
2015 Open Data for Improving Youth Policies
abstract
The Open Data \textit{philosophy} is based on the idea that certain data should be made ​​available to all citizens, in an open form, without any copyright restrictions, patents or other mechanisms of control. Various government have started to publish open data, first of all USA and UK in 2009, and in 2015, the Open Data Barometer project (www.opendatabarometer.org) states that on 77 diverse states across the world, over 55 percent have developed some form of Open Government Data initiative. We claim Public Administrations, that are the main producers and one of the consumers of Open Data, might effectively extract important information by integrating its own data with open data sources.This paper reports the activities carried on during a one-year research project on Open Data for Youth Policies. The project was mainly devoted to explore the youth situation in the municipalities and provinces of the Emilia Romagna region (Italy), in particular, to examine data on population, education and work.The project goals were: to identify interesting data sources both from the open data community and from the private repositories of local governments of Emilia Romagna region related to the Youth Policies; to integrate them and, to show up the result of the integration by means of a useful navigator tool; in the end, to publish new information on the web as Linked Open Data. This paper also reports the main issues encountered that may seriously affect the entire process of consumption, integration till the publication of open data.
Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po
KEOD2
2015 Driving Innovation in Youth Policies with Open Data
Domenico Beneventano, Sonia Bergamaschi, Luca Gagliardelli, Laura Po
IC3K2
2015 Visual Querying LOD sources with LODeX
abstract
The Linked Open Data (LOD) Cloud has more than tripled its sources in just three years (from 295 sources in 2011 to 1014 in 2014). While the LOD data are being produced at a increasing rate, LOD tools lack in producing an high level representation of datasets and in supporting users in the exploration and querying of a source. To overcome the above problems and significantly increase the number of consumers of LOD data, we devised a new method and a tool, called LODeX, that promotes the understanding, navigation and querying of LOD sources both for experts and for beginners. It also provides a standardized and homogeneous summary of LOD sources and supports user in the creation of visual queries on previously unknown datasets.
Fabio Benedetti, Sonia Bergamaschi, Laura Po
K-CAP2
2015 Exploiting semantics for filtering and searching knowledge in a software development context
Sonia Bergamaschi, Riccardo Martoglia, Serena Sorrentino
Knowl. Inf. Syst.1
2014 Big Data Integration - State of the Art & Challenges
Sonia Bergamaschi
KEOD1
2013 QUEST: A Keyword Search System for Relational Data based on Semantic and Machine Learning Techniques
abstract
We showcase QUEST (QUEry generator for STructured sources), a search engine for relational databases that combines semantic and machine learning techniques for transforming keyword queries into meaningful SQL queries. The search engine relies on two approaches: the forward, providing mappings of keywords into database terms (names of tables and attributes, and domains of attributes), and the backward, computing the paths joining the data structures identified in the forward step. The results provided by the two approaches are combined within a probabilistic framework based on the Dempster-Shafer Theory. We demonstrate QUEST capabilities, and we show how, thanks to the flexibility obtained by the probabilistic combination of different techniques, QUEST is able to compute high quality results even with few training data and/or with hidden data sources such as those found in the Deep Web.
Sonia Bergamaschi, Francesco Guerra 0001, Matteo Interlandi, Raquel Trillo Lado, Yannis Velegrakis
Proc. VLDB Endow.1
2012 A Supervised Method for Lexical Annotation of Schema Labels Based on Wikipedia
Serena Sorrentino, Sonia Bergamaschi, Elena Parmiggiani
ER2
2012 A Meta-language for MDX Queries in eLog Business Solution
abstract
The adoption of business intelligence technology in industries is growing rapidly. Business managers are not satisfied with ad hoc and static reports and they ask for more flexible and easy to use data analysis tools. Recently, application interfaces that expand the range of operations available to the user, hiding the underlying complexity, have been developed. The paper presents eLog, a business intelligence solution designed and developed in collaboration between the database group of the University of Modena and Reggio Emilia and eBilling, an Italian SME supplier of solutions for the design, production and automation of documentary processes for top Italian companies. eLog enables business managers to define OLAP reports by means of a web interface and to customize analysis indicators adopting a simple meta-language. The framework translates the user's reports into MDX queries and is able to automatically select the data cube suitable for each query. Over 140 medium and large companies have exploited the technological services of eBilling S.p.A. to manage their documents flows. In particular, eLog services have been used by the major media and telecommunications Italian companies and their foreign annex, such as Sky, Media set, H3G, Tim Brazil etc. The largest customer can provide up to 30 millions mail pieces within 6 months (about 200 GB of data in the relational DBMS). In a period of 18 months, eLog could reach 150 millions mail pieces (1 TB of data) to handle.
Sonia Bergamaschi, Matteo Interlandi, Mario Longo, Laura Po, Maurizio Vincini
ICDE1
2011 The list Viterbi training algorithm and its application to keyword search over databases
abstract
Hidden Markov Models (HMMs) are today employed in a variety of applications, ranging from speech recognition to bioinformatics. In this paper, we present the List Viterbi training algorithm, a version of the Expectation-Maximization (EM) algorithm based on the List Viterbi algorithm instead of the commonly used forward-backward algorithm. We developed the batch and online versions of the algorithm, and we also describe an interesting application in the context of keyword search over databases, where we exploit a HMM for matching keywords into database terms. In our experiments we tested the online version of the training algorithm in a semi-supervised setting that allows us to take into account the feedbacks provided by the users.
Silvia Rota, Sonia Bergamaschi, Francesco Guerra 0001
CIKM2
2011 A Hidden Markov Model Approach to Keyword-Based Search over Relational Databases
Sonia Bergamaschi, Francesco Guerra 0001, Silvia Rota, Yannis Velegrakis
ER1
2011 NORMS: An automatic tool to perform schema label normalization
abstract
Schema matching is the problem of finding relationships among concepts across heterogeneous data sources (heterogeneous in format and structure). Schema matching systems usually exploit lexical and semantic information provided by lexical databases/thesauri to discover intra/inter semantic relationships among schema elements. However, most of them obtain poor performance on real world scenarios due to the significant presence of “non-dictionary words”. Non-dictionary words include compound nouns, abbreviations and acronyms. In this paper, we present NORMS (NORMalizer of Schemata), a tool performing schema label normalization to increase the number of comparable labels extracted from schemata.
Serena Sorrentino, Sonia Bergamaschi, Maciej Gawinecki
ICDE2
2011 Keyword search over relational databases: a metadata approach
abstract
Keyword queries offer a convenient alternative to traditional SQL in querying relational databases with large, often unknown, schemas and instances. The challenge in answering such queries is to discover their intended semantics, construct the SQL queries that describe them and used them to retrieve the respective tuples. Existing approaches typically rely on indices built a-priori on the database content. This seriously limits their applicability if a-priori access to the database content is not possible. Examples include the on-line databases accessed through web interface, or the sources in information integration systems that operate behind wrappers with specific query capabilities. Furthermore, existing literature has not studied to its full extend the inter-dependencies across the ways the different keywords are mapped into the database values and schema elements. In this work, we describe a novel technique for translating keyword queries into SQL based on the Munkres (a.k.a. Hungarian) algorithm. Our approach not only tackles the above two limitations, but it offers significant improvements in the identification of the semantically meaningful SQL queries that describe the intended keyword query semantics. We provide details of the technique implementation and an extensive experimental evaluation.
Sonia Bergamaschi, Elton Domnori, Francesco Guerra 0001, Raquel Trillo Lado, Yannis Velegrakis
SIGMOD Conference1
2011 A semantic approach to ETL technologies
Sonia Bergamaschi, Francesco Guerra 0001, Mirko Orsini, Claudio Sartori 0001, Maurizio Vincini
Data Knowl. Eng.1
2011 Editorial
Carlo Batini, Domenico Beneventano, Sonia Bergamaschi, Tiziana Catarci
Inf. Syst.3
2011 Using semantic techniques to access web data
Raquel Trillo Lado, Laura Po, Sergio Ilarri, Sonia Bergamaschi, Eduardo Mena
Inf. Syst.4
2010 Automatic Lexical Annotation Applied to the SCARLET Ontology Matcher
Laura Po, Sonia Bergamaschi
ACIIDS (2)2
2010 Schema label normalization for improving schema matching
Serena Sorrentino, Sonia Bergamaschi, Maciej Gawinecki, Laura Po
Data Knowl. Eng.2
2010 Keymantic: Semantic Keyword-based Searching in Data Integration Systems
abstract
We propose the demonstration of Keymantic , a system for keyword-based searching in relational databases that does not require a-priori knowledge of instances held in a database. It finds numerous applications in situations where traditional keyword-based searching techniques are inapplicable due to the unavailability of the database contents for the construction of the required indexes.
Sonia Bergamaschi, Elton Domnori, Francesco Guerra 0001, Mirko Orsini, Raquel Trillo Lado, Yannis Velegrakis
Proc. VLDB Endow.1
2009 Schema Normalization for Improving Schema Matching
Serena Sorrentino, Sonia Bergamaschi, Maciej Gawinecki, Laura Po
ER2
2009 Toward a Unified View of Data and Services
Sonia Bergamaschi, Andrea Maurino
WISE1
2007 An Incremental Method for the Lexical Annotation of Domain Ontologies
abstract
In this article, we present MELIS (Meaning Elicitation and Lexical Integration System), a method and a software tool for enabling an incremental process of automatic annotation of local schemas (e.g. relational database schemas, directory trees) with lexical information. The distinguishing and original feature of MELIS is the incremental process: the higher the number of schemas which are processed, the more background/ domain knowledge is cumulated in the system (a portion of domain ontology is learned at every step), the better the performance of the systems on annotating new schemas. MELIS has been tested as a component of the MOMIS-Ontology Builder, a framework able to create a domain ontology representing a set of selected data sources, described with a standard W3C language wherein concepts and attributes are annotated according to the lexical reference database. We describe the MELIS component within the MOMIS-Ontology Builder framework and provide some experimental results of MELIS as a standalone tool and as a component integrated in MOMIS.
Sonia Bergamaschi, Paolo Bouquet, Daniel Giacomuzzi, Francesco Guerra 0001, Laura Po, Maurizio Vincini
Int. J. Semantic Web Inf. Syst.1
2004 TUCUXI: The InTelligent Hunter Agent for Concept Understanding and LeXical ChaIning
abstract
In this paper we present TUCUXI, an intelligent hunter agent that replaces traditional keywords-based queries on the Web with a user-provided domain ontology, where meanings to be searched are not ambiguous. TUCUXI judges the relevance of the retrieved pages by matching the domain ontology against a simplified, but semantically rich, document representation (Map of Meanings). The Map of Meanings extraction involves the Lexical Chaining technique, from the Natural Language Processing (NLP) research field.
Roberta Benassi, Sonia Bergamaschi, Maurizio Vincini
Web Intelligence2
2003 WINK: A Web-Based System for Collaborative Project Management in Virtual Enterprises
abstract
The increasing of globalization and flexibility required to the companies has generated, in the last decade, new issues, related to the managing of large scale projects within geographically distributed networks and to the cooperation of enterprises. ICT support systems are required to allow enterprises to share information, guarantee data-consistency and establish synchronized and collaborative processes. In this paper, we present a collaborative project management system that integrates data coming from aerospace industries with two main goals: avoiding inconsistencies generated by updates at the sources' level and minimizing data replications. The proposed system is composed of a collaborative project management component supported by a Web interface, a multiagent data integration component, which supports information sharing and querying, and SOAP enabled Web-services which ensure the whole interoperability of the software components. The system was developed by the University of Modena and Reggio Emilia, Gruppo Formula S.p.A. and Alenia Spazio S.p.A. within the EU WINK Project (Web-linked Integration of Network based Knowledge - IST-2000-28221).
Sonia Bergamaschi, Gionata Gelati, Francesco Guerra 0001, Maurizio Vincini
WISE1
2003 Description logics for semantic query optimization in object-oriented database systems
abstract
Semantic query optimization uses semantic knowledge (i.e., integrity constraints) to transform a query into an equivalent one that may be answered more efficiently. This article proposes a general method for semantic query optimization in the framework of Object-Oriented Database Systems. The method is effective for a large class of queries, including conjunctive recursive queries expressed with regular path expressions and is based on three ingredients. The first is a Description Logic, ODL RE , providing a type system capable of expressing: class descriptions, queries, views, integrity constraint rules and inference techniques, such as incoherence detection and subsumption computation. The second is a semantic expansion function for queries, which incorporates restrictions logically implied by the query and the schema (classes + rules) in one query. The third is an optimal rewriting method of a query with respect to the schema classes that rewrites a query into an equivalent one, by determining more specialized classes to be accessed and by reducing the number of factors. We implemented the method in a tool providing an ODMG-compliant interface that allows a full interaction with OQL queries, wrapping underlying Description Logic representation and techniques to the user.
Domenico Beneventano, Sonia Bergamaschi, Claudio Sartori 0001
ACM Trans. Database Syst.2
2002 A Data Integration Framework for e-Commerce Product Classification
Sonia Bergamaschi, Francesco Guerra 0001, Maurizio Vincini
ISWC1
2002 Momis: Exploiting Agents to Support Information Integration
abstract
Information overloading introduced by the large amount of data that is spread over the Internet must be faced in an appropriate way. The dynamism and the uncertainty of the Internet, along with the heterogeneity of the sources of information are the two main challenges for today's technologies related to information management. In the area of information integration, this paper proposes an approach based on mobile software agents integrated in the MOMIS (Mediator envirOnment for Multiple Information Sources) infrastructure, which enables semi-automatic information integration to deal with the integration and query of multiple, heterogeneous information sources (relational, object, XML and semi-structured sources). The exploitation of mobile agents in MOMIS can significantly increase the flexibility of the system. In fact, their characteristics of autonomy and adaptability well suit the distributed and open environments, such as the Internet. The aim of this paper is to show the advantages of the introduction in the MOMIS infrastructure of intelligent and mobile software agents for the autonomous management and coordination of integration and query processing over heterogeneous data sources.
Giacomo Cabri, Francesco Guerra 0001, Maurizio Vincini, Sonia Bergamaschi, Letizia Leonardi, Franco Zambonelli
Int. J. Cooperative Inf. Syst.4
2001 Semantic integration of heterogeneous information sources
Sonia Bergamaschi, Silvana Castano, Maurizio Vincini, Domenico Beneventano
Data Knowl. Eng.1
2000 Information Integration: The MOMIS Project Demonstration
Domenico Beneventano, Sonia Bergamaschi, Silvana Castano, Alberto Corni, R. Guidetti, G. Malvezzi, Michele Melchiori, Maurizio Vincini
VLDB2
1998 Chrono: A Conceptual Design Framework for Temporal Entities
Sonia Bergamaschi, Claudio Sartori 0001
ER1
1998 Consistency Checking in Complex Object Database Schemata with Integrity Constraints
abstract
Integrity constraints are rules that should guarantee the integrity of a database. Provided an adequate mechanism to express them is available, the following question arises: is there any way to populate a database which satisfies the constraints supplied by a database designer? That is, does the database schema, including constraints, admit at least a nonempty model? This work answers the above question in a complex object database environment, providing a theoretical framework, including the following ingredients: (1) two alternative formalisms, able to express a relevant set of state integrity constraints with a declarative style; (2) two specialized reasoners, based on the tableaux calculus, able to check the consistency of complex objects database schemata expressed with the two formalisms. The proposed formalisms share a common kernel, which supports complex objects and object identifiers, and which allow the expression of acyclic descriptions of: classes, nested relations and views, built up by means of the recursive use of record, quantified set, and object type constructors and by the intersection, union, and complement operators. Furthermore, the kernel formalism allows the declarative formulation of typing constraints and integrity rules. In order to improve the expressiveness and maintain the decidability of the reasoning activities, we extend the kernel formalism into two alternative directions. The first formalism, OLCP, introduces the capability of expressing path relations. Because cyclic schemas are extremely useful, we introduce a second formalism, OLCD, with the capability of expressing cyclic descriptions but disallowing the expression of path relations. In fact, we show that the reasoning activity in OLCDP (i.e., OLCP with cycles) is undecidable.
Domenico Beneventano, Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001
IEEE Trans. Knowl. Data Eng.2
1997 ODB-QOPTIMIZER: A Tool for Semantic Query Optimization in OODB
abstract
ODB-QOPTIMIZER is a ODMG 93 compliant tool for the schema validation and semantic query optimization. The approach is based on two fundamental ingredients. The first one is the OCDL description logics (DLs) proposed as a common formalism to express class descriptions, a relevant set of integrity constraint rules (IC rules) and queries. The second one are DLs inference techniques, exploited to evaluate the logical implications expressed by IC rules and thus to produce the semantic expansion of a given query.
Sonia Bergamaschi, Domenico Beneventano, Claudio Sartori 0001, Maurizio Vincini
ICDE1
1997 Incoherence and Subsumption for Recursive Views and Queries in Object-Oriented Data Models
Domenico Beneventano, Sonia Bergamaschi
Data Knowl. Eng.2
1994 The E/S Knowledge Representation System
Sonia Bergamaschi, Stefano Lodi, Claudio Sartori 0001
Data Knowl. Eng.1
1992 Subsumption for Complex Object Data Models
Domenico Beneventano, Sonia Bergamaschi
ICDT2
1992 On Taxonomic Reasoning in Conceptual Design
abstract
Taxonomic reasoning is a typical task performed by many AI knowledge representation systems. In this paper, the effectiveness of taxonomic reasoning techniques as an active support to knowledge acquisition and conceptual schema design is shown. The idea developed is that by extending conceptual models with defined concepts and giving them rigorous logic semantics, it is possible to infer isa relationships between concepts on the basis of their descriptions. From a theoretical point of view, this approach makes it possible to give a formal definition for consistency and minimality of a conceptual schema. From a pragmatic point of view it is possible to develop an active environment that allows automatic classification of a new concept in the right position of a given taxonomy, ensuring the consistency and minimality of a conceptual schema. A formalism that includes the data semantics of models giving prominence to type constructors (E/R, TAXIS, GALILEO) and algorithms for taxonomic inferences are presented: their soundness, completeness, and tractability properties are proved. Finally, an extended formalism and taxonomic inference algorithms for models giving prominence to attributes (FDM, IFO) are given.
Sonia Bergamaschi, Claudio Sartori 0001
ACM Trans. Database Syst.1
1988 Entity-Situation: A Model for the Knowledge Representation Module of a KBMS
Sonia Bergamaschi, Flavio Bonfatti, Claudio Sartori 0001
EDBT1
1988 On Taxonomic Reasoning in E/R Environment
Sonia Bergamaschi, Lorenzo Cavedoni, Claudio Sartori 0001, Paolo Tiberio
ER1
1988 Relational data base design for the intensional aspects of a knowledge base
Sonia Bergamaschi, Flavio Bonfatti, Lorenza Cavazza, Claudio Sartori 0001, Paolo Tiberio
Inf. Syst.1
1986 Choice of the optimal number of blocks for data access by an index
Sonia Bergamaschi, Maria Rita Scalas
Inf. Syst.1