Paolo Manghi

dblp:28/2137 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0001-7291-3210ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 5Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2023 A Graph Neural Network Approach for Evaluating Correctness of Groups of Duplicates
abstract
Abstract Unlabeled entity deduplication is a relevant task already studied in the recent literature. Most methods can be traced back to the following workflow: entity blocking phase, in-block pairwise comparisons between entities to draw similarity relations, closure of the resulting meshes to create groups of duplicate entities, and merging group entities to remove disambiguation. Such methods are effective but still not good enough whenever a very low false positive rate is required. In this paper, we present an approach for evaluating the correctness of “groups of duplicates”, which can be used to measure the group’s accuracy hence its likelihood of false-positiveness. Our novel approach is based on a Graph Neural Network that exploits and combines the concept of Graph Attention and Long Short Term Memory (LSTM). The accuracy of the proposed approach is verified in the context of Author Name Disambiguation applied to a curated dataset obtained as a subset of the OpenAIRE Graph that includes PubMed publications with at least one ORCID identifier.
Michele De Bonis, Filippo Minutella, Fabrizio Falchi, Paolo Manghi
TPDL4
2023 Tracing Data Footprints: Formal and Informal Data Citations in the Scientific Literature
Ornella Irrera, Andrea Mannocci, Paolo Manghi, Gianmaria Silvello
TPDL3
2022 "Knock Knock! Who's There?" A Study on Scholarly Repositories' Availability
Andrea Mannocci, Miriam Baglioni, Paolo Manghi
TPDL3
2020 Context-Driven Discoverability of Research Data
Miriam Baglioni, Paolo Manghi, Andrea Mannocci
TPDL2
2019 The OpenAIRE Research Community Dashboard: On Blending Scientific Workflows and Scientific Publishing
Miriam Baglioni, Alessia Bardi, Argiro Kokogiannaki, Paolo Manghi, Katerina Iatropoulou, Pedro Príncipe, André Vieira, Lars Holm Nielsen, Harry Dimitropoulos, Ioannis Foufoulas, Natalia Manola, Claudio Atzori, Sandro La Bruzzo, Emma Lazzeri, Michele Artini, Michele De Bonis, Andrea Dell'Amico
TPDL4
2018 GDup: De-Duplication of Scholarly Communication Big Graphs
abstract
Today, several online services offer functionalities to access information from big scholarly communication graphs, which interlink entities such as publications, authors, datasets, organizations, etc. Such graphs are often populated over time as aggregations of multiple sources and therefore suffer from entity duplication problems. Although deduplication of graphs is a known and actual problem, solutions tend to be dedicated and address a few of the underlying challenges. In this paper, we propose the GDup system, an integrated, scalable, general-purpose system for entity deduplication over big information graphs. GDup supports practitioners with the functionalities needed to realize a fully-fledged entity deduplication workflow over a generic input graph, inclusive of Ground Truth support, end-user feedback, and strategies for identifying and merging duplicates to obtain an output disambiguated graph. GDup is today one of the core components of the OpenAIRE infrastructure production system, monitoring Open Science trends on behalf of the European Commission.
Claudio Atzori, Paolo Manghi, Alessia Bardi
BDCAT2
2018 De-duplicating the OpenAIRE Scholarly Communication Big Graph
abstract
The OpenAIRE infrastructure populates a scholarly communication big graph interlinking metadata objects of publications, datasets, software, organizations, funders, and projects. In order to de-duplicate this graph, OpenAIRE has developed GDup, an integrated, scalable, general-purpose system for entity deduplication over big information graphs. GDup offers functionalities to realize a fully-fledged entity deduplication workflow over a generic input graph, inclusive of Ground Truth support, end-user feedback, and strategies for identifying and merging duplicates to obtain an output disambiguated graph.
Claudio Atzori, Paolo Manghi, Alessia Bardi
eScience2
2018 Developing sustainable Open Science solutions in the frame of EU funded research: the OpenUP case
abstract
Open Access and Open Scholarship have revolutionized the way scholarly artefacts are evaluated and published, while the introduction of new technologies and media in scientific workflows has changed the "how" and to "whom" science is communicated, and how stakeholders interact with the scientific community and the broader public. The EU funded project OpenUP is connecting people, information and tools and provides a knowledge hub and a validated framework for the review, assessment and dissemination aspects of the research lifecycle, under the prism of a gender-sensitive Open Science.
Eleni Toli, Electra Sifacaki, Natalia Manola, Yannis E. Ioannidis, Anthony Ross 0001, Edit Görögh, Michela Vignoli, Vilte Banelyte, Paolo Manghi, Saskia Woutersen-Windhouwer
OpenSym9
2016 DataQ: A Data Flow Quality Monitoring System for Aggregative Data Infrastructures
Andrea Mannocci, Paolo Manghi
TPDL2
2015 Data journals: A survey
abstract
Data occupy a key role in our information society. However, although the amount of published data continues to grow and terms such as data deluge and big data today characterize numerous (research) initiatives, much work is still needed in the direction of publishing data in order to make them effectively discoverable, available, and reusable by others. Several barriers hinder data publishing, from lack of attribution and rewards, vague citation practices, and quality issues to a rather general lack of a data‐sharing culture. Lately, data journals have overcome some of these barriers. In this study of more than 100 currently existing data journals, we describe the approaches they promote for data set description, availability, citation, quality, and open access. We close by identifying ways to expand and strengthen the data journals approach as a means to promote data set access and exploitation.
Leonardo Candela, Donatella Castelli, Paolo Manghi, Alice Tani
J. Assoc. Inf. Sci. Technol.3
2013 Data Searchery - Preliminary Analysis of Data Sources Interlinking
Paolo Manghi, Andrea Mannocci
TPDL1
2010 Connecting the local and the online in information management
abstract
With the popularity of social media sites, digital content is increasingly stored and managed online. At the same time, the desktop and local storage continues to provide a personal environment in which users perform their daily tasks. Thus, to accomplish their tasks, users need to continuously switch between local and remote resources and applications, often carrying the burden of coordinating and synchronizing these in a consistent way. In this demonstration, we describe a system, called ScholarLynk, that bridges the local and online worlds and allows users to manage both local and online resources in a uniform way and in collaboration with others.
Gabriella Kazai, Natasa Milic-Frayling, Tim Haughton, Natalia Manola, Katerina Iatropoulou, Antonis Lempesis, Paolo Manghi, Marko Mikulicic
CIKM7
2007 Scalable Query Dissemination in XPeer
abstract
This paper presents XPeer, a data sharing system for massively distributed XML data. XPeer allows users to publish and query heterogeneous information without any significant administration efforts. XPeer tries to dispatch any given query to all and only the potentially relevant peers, exploiting a superpeer network to this aim.
Giovanni Conforti, Giorgio Ghelli, Paolo Manghi, Carlo Sartiani
IDEAS3
2006 Static analysis for path correctness of XML queries
abstract
A part of a query that will never contribute data to the query answer should be regarded as an error. This principle has been recently accepted into mainstream XML query languages, but was still waiting for a complete treatment. We provide here a precise definition for this class of errors, and define a type system that is sound and complete, in its search for such errors, for a core language, under mild restrictions on the use of recursion in type definitions. In the process, we describe a dichotomy among existential and universal type systems, which is essential to understand some specific features of our type system.
Dario Colazzo, Giorgio Ghelli, Paolo Manghi, Carlo Sartiani
J. Funct. Program.3
2004 Types for path correctness of XML queries
abstract
If a subexpression in a query will never contribute data to the query answer, this should be regarded as an error. This principle has been recently accepted into mainstream XML query languages, but was still waiting for a complete treatment. We provide here a precise definition for this class of errors, and define a type system that is sound and complete, in its search for such errors, for a core language, under mild restrictions on the use of recursion in type definitions. In the process, we describe a dichotomy among existential and universal type systems, which is useful to understand some unusual features of our type system.
Dario Colazzo, Giorgio Ghelli, Paolo Manghi, Carlo Sartiani
ICFP3
2002 Types for Correctness of Queries over Semistructured Data
Dario Colazzo, Giorgio Ghelli, Paolo Manghi, Carlo Sartiani
WebDB3
2002 The Query Language TQL
Giovanni Conforti, Giorgio Ghelli, Antonio Albano, Dario Colazzo, Paolo Manghi, Carlo Sartiani
WebDB5
2002 An Approach to High-Level Language Bindings to XML
Fabio Simeoni, Paolo Manghi, David Lievens, Richard Connor 0001, Steve Neely
Inf. Softw. Technol.2
2002 A typed text retrieval query language for XML documents
abstract
Abstract XML is nowadays considered the standard meta‐language for document markup and data representation. XML is widely employed in Web‐related applications as well as in database applications, and there is also a growing interest for it by the literary community to develop tools for supporting document‐oriented retrieval operations. The purpose of this article is to show the basic new requirements of this kind of applications and to present the main features of a typed query language, called Tequyla‐TX, designed to support them.
Dario Colazzo, Carlo Sartiani, Antonio Albano, Paolo Manghi, Giorgio Ghelli, Luca Lini, Michele Paoli
J. Assoc. Inf. Sci. Technol.4
1998 On the Unification of Persistent Programming and the World Wide Web
Richard Connor 0001, Keith Sibson, Paolo Manghi
WebDB3