Silvio Peroni

dblp:48/1099 · DBLP profile ↗
← Back
28ranked-venue papers in the field
11as first author
6since 2021 · last 2024
0000-0003-0530-4305ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 15 (9 first)Information Retrieval & Web Search · 11 (2 first)Data Mining & Knowledge Discovery · 1Business Process & Enterprise Data · 1
YearPublicationVenuePosition
2024 Developing Application Profiles for Enhancing Data and Workflows in Cultural Heritage Digitisation Processes
abstract
Abstract As a result of the proliferation of 3D digitisation in the context of cultural heritage projects, digital assets and digitisation processes – being considered as proper research objects – must prioritise adherence to FAIR principles. Existing standards and ontologies, such as CIDOC-CRM, play a crucial role in this regard, but they are often over-engineered for the need of a particular application context, thus making their understanding and adoption difficult. Application profiles of a given standard – defined as sets of ontological entities drawn from one or more semantic artefacts for a particular context or application – are usually proposed as tools for promoting interoperability and reuse while being tied entirely to the particular application context they refer to. In this paper, we present an adaptation and application of an ontology development methodology, i.e. SAMOD, to guide the creation of robust, semantically sound application profiles of large standard models. Using an existing pilot study we have developed in a project dedicated to leveraging virtual technologies to preserve and valorise cultural heritage, we introduce an application profile named CHAD-AP, that we have developed following our customised version of SAMOD. We reflect on the use of SAMOD and similar ontology development methodologies for this purpose, highlighting its strengths and current limitations, future developments, and possible adoption in other similar projects.
Sebastian Barzaghi, Ivan Heibi, Arianna Moretti, Silvio Peroni
ISWC (3)4
2022 Structured References from PDF Articles: Assessing the Tools for Bibliographic Reference Extraction and Parsing
Alessia Cioffi, Silvio Peroni
TPDL2
2022 Enabling Portability and Reusability of Open Science Infrastructures
Giuseppe Grieco, Ivan Heibi, Arcangelo Massari, Arianna Moretti, Silvio Peroni
TPDL5
2022 The Way We Cite: Common Metadata Used Across Disciplines for Defining Bibliographic References
Erika Alves dos Santos, Silvio Peroni, Marcos Luiz Mucheroni
TPDL2
2022 A Programming Interface for Creating Data According to the SPAR Ontologies and the OpenCitations Data Model
Simone Persiani, Marilena Daquino, Silvio Peroni
ESWC3
2021 BiblioDAP'21: The 1st Workshop on Bibliographic Data Analysis and Processing
abstract
Automatic processing of bibliographic data becomes very important in digital libraries, data science and machine learning due to its importance in keeping pace with the significant increase of published papers every year from one side and to the inherent challenges from the other side. This processing has several aspects including but not limited to I) Automatic extraction of references from PDF documents, II) Building an accurate citation graph, III) Author name disambiguation, etc. Bibliographic data is heterogeneous by nature and occurs in both structured (e.g. citation graph) and unstructured (e.g. publications) formats. Therefore, it requires data science and machine learning techniques to be processed and analysed. Here we introduce BiblioDAP'21: The 1st Workshop on Bibliographic Data Analysis and Processing.
Zeyd Boukhers, Philipp Mayr 0001, Silvio Peroni
KDD3
2020 The OpenCitations Data Model
abstract
A variety of schemas and ontologies are currently used for the machine-readable description of bibliographic entities and citations. This diversity, and the reuse of the same ontology terms with different nuances, generates inconsistencies in data. Adoption of a single data model would facilitate data integration tasks regardless of the data supplier or context application. In this paper we present the OpenCitations Data Model (OCDM), a generic data model for describing bibliographic entities and citations, developed using Semantic Web technologies. We also evaluate the effective reusability of OCDM according to ontology evaluation practices, mention existing users of OCDM, and discuss the use and impact of OCDM in the wider open science community.
Marilena Daquino, Silvio Peroni, David M. Shotton, Giovanni Colavizza, Behnam Ghavimi, Anne Lauscher, Philipp Mayr 0001, Matteo Romanello, Philipp Zumstein
ISWC (2)2
2018 The SPAR Ontologies
abstract
Abstract Over the past eight years, we have been involved in the development of a set of complementary and orthogonal ontologies that can be used for the description of the main areas of the scholarly publishing domain, known as the SPAR (Semantic Publishing and Referencing) Ontologies. In this paper, we introduce this suite of ontologies, discuss the basic principles we have followed for their development, and describe their uptake and usage within the academic, institutional and publishing communities.
Silvio Peroni, David M. Shotton
ISWC (2)1
2017 The RASH JavaScript Editor (RAJE): A Wordprocessor for Writing Web-first Scholarly Articles
abstract
The most used format for submitting and publishing papers in the academic domain is the Portable Document Format (PDF), since its possibility of being rendered in the same way independently from the device used for visualising it. However, the PDF format has some important issues as well, among which the lack of interactivity and the low degree of accessibility. In order to address these issues, recently some journals, conferences, and workshops have started to accept also HTML as Web-first submission/publication format. However, most of the people are not able to produce a well-formed HTML5 article from scratch, and they would, thus, need an appropriate interface, e.g. a word processor, for creating such HTML-compliant scholarly article. To provide a solution to the aforementioned issue, in this paper we introduce the RASH JavaScript Editor (a.k.a. RAJE), which is a multi platform word processor for writing scholarly articles in HTML natively. RAJE allows authors to write research papers by means of a user-friendly interface hiding the complexities of HTML5. We also discuss the outcomes of a user study where we asked some researchers to write a scientific paper using RAJE.
Gianmarco Spinaci, Silvio Peroni, Angelo Di Iorio, Francesco Poggi, Fabio Vitali
DocEng2
2017 UNDO: The United Nations System Document Ontology
abstract
Akoma Ntoso is an OASIS Committee Specification Draft standard for the electronic representations of parliamentary, normative and judicial documents in XML. Recently, it has been officially adopted by the United Nations (UN) as the main electronic format for making UN documents machine-processable. However, Akoma Ntoso does not force nor define any formal ontology for allowing the description of real-world objects, concepts and relations mentioned in documents. In order to address this gap, in this paper we introduce the United Nations System Document Ontology (UNDO), i.e. an OWL 2 DL ontology developed and adopted by the United Nations that aims at providing a framework for the formal description of all these entities.
Silvio Peroni, Monica Palmirani, Fabio Vitali
ISWC (2)1
2017 One Year of the OpenCitations Corpus - Releasing RDF-Based Scholarly Citation Data into the Public Domain
Silvio Peroni, David M. Shotton, Fabio Vitali
ISWC (2)1
2017 Interfacing fast-fashion design industries with Semantic Web technologies: The case of Imperial Fashion
Silvio Peroni, Fabio Vitali
J. Web Semant.1
2016 The Role of Ontology Design Patterns in Linked Data Projects
Valentina Presutti, Giorgia Lodi, Andrea Giovanni Nuzzolese, Aldo Gangemi, Silvio Peroni, Luigi Asprino
ER5
2016 FOOD: FOod in Open Data
abstract
This paper describes the outcome of an e-government project named FOOD, FOod in Open Data, which was carried out in the context of a collaboration between the Institute of Cognitive Sciences and Technologies of the Italian National Research Council, the Italian Ministry of Agriculture (MIPAAF) and the Italian Digital Agency (AgID). In particular, we implemented several ontologies for describing protected names of products (wine, pasta, fish, oil, etc.). In addition, we present the process carried out for producing and publishing a LOD dataset containing data extracted from existing Italian policy documents on such products and compliant with the aforementioned ontologies.
Silvio Peroni, Giorgia Lodi, Luigi Asprino, Aldo Gangemi, Valentina Presutti
ISWC (2)1
2015 Exploring Scholarly Papers Through Citations
abstract
Bibliographies are fundamental components of academic papers and both the scientific research and its evaluation are fundamentally organized around the correct examination and classification of scientific bibliographies. Currently, most digital libraries publish bibliographic information about their content for free, and many include the citations (outgoing and in some cases even incoming) to the papers they manage. Unfortunately no sophistication is spent for these lists: monolithic pieces of text where it is even difficult to tell automatically the authors, the title and publication details, and where users are provided with no mechanisms to filter and access full context of each citation. For instance, there is no way to know in which sentence a work was cited (the citation context) and why (the citation function).
Angelo Di Iorio, Raffaele Giannella, Francesco Poggi, Silvio Peroni, Fabio Vitali
DocEng4
2014 Evaluating Citation Functions in CiTO: Cognitive Issues
Paolo Ciancarini, Angelo Di Iorio, Andrea Giovanni Nuzzolese, Silvio Peroni, Fabio Vitali
ESWC4
2014 Dealing with structural patterns of XML documents
abstract
Evaluating collections of XML documents without paying attention to the schema they were written in may give interesting insights into the expected characteristics of a markup language, as well as any regularity that may span vocabularies and languages, and that are more fundamental and frequent than plain content models. In this paper we explore the idea of structural patterns in XML vocabularies, by examining the characteristics of elements as they are used, rather than as they are defined. We introduce from the ground up a formal theory of 8 plus 3 structural patterns for XML elements, and verify their identifiability in a number of different XML vocabularies. The results allowed the creation of visualization and content extraction tools that are completely independent of the schema and without any previous knowledge of the semantics and organization of the XML vocabulary of the documents.
Angelo Di Iorio, Silvio Peroni, Francesco Poggi, Fabio Vitali
J. Assoc. Inf. Sci. Technol.2
2013 Recognising document components in XML-based academic articles
abstract
Recognising textual structures (paragraphs, sections, etc.) provides abstract and more general mechanisms for describing documents independent of the particular semantics of specific markup schemas, tools and presentation stylesheets. In this paper we propose an algorithm that allows us to identify the structural role of each element in a set of homogeneous scientific articles stored as XML files.
Angelo Di Iorio, Silvio Peroni, Francesco Poggi, Fabio Vitali, David M. Shotton
ACM Symposium on Document Engineering2
2013 Tools for the Automatic Generation of Ontology Documentation: A Task-Based Evaluation
abstract
Ontologies are knowledge constructs essential for creation of the Web of Data. Good documentation is required to permit people to understand ontologies and thus employ them correctly, but this is costly to create by tradition authorship methods, and is thus inefficient to create in this way until an ontology has matured into a stable structure. The authors describe three tools, LODE, Parrot and the OWLDoc-based Ontology Browser, that can be used automatically to create documentation from a well-formed OWL ontology at any stage of its development. They contrast their properties and then report on the authors’ evaluation of their effectiveness and usability, determined by two task-based user testing sessions.
Silvio Peroni, David M. Shotton, Fabio Vitali
Int. J. Semantic Web Inf. Syst.1
2012 A first approach to the automatic recognition of structural patterns in XML documents
abstract
XML is among the preferred formats for storing the structure of documents such as scientific articles, manuals, documentation, literary works, etc. Sometimes publishers adopt established and well-known vocabularies such as DocBook and TEI, other times they create partially or entirely new ones that better deal with the particular requirements of their documents. The (explicit and implicit) requirements of use in these vocabularies often follow well-established patterns, creating meta-structures (the block, the container, the inline element, etc.) that persist across vocabularies and authors and that describe a truer and more general conceptualization of the documents' building blocks. Addressing such meta-structures not only gives a better insight of what documents really are composed of, but provides abstract and more general mechanisms to work on documents regardless of the availability of specific schemas, tools and presentation stylesheets. In this paper we introduce a schemaindependent theory based on eleven structural patterns. We provide a definition of such patterns and how they synthesize characteristics emerging from real markup documents. Additionally, we propose an algorithm that allows us to identify the pattern of each element in a set of homogeneous markup documents.
Angelo Di Iorio, Silvio Peroni, Francesco Poggi, Fabio Vitali
ACM Symposium on Document Engineering2
2012 Faceted documents: describing document characteristics using semantic lenses
abstract
The semantic enhancement of a traditional scientific paper is not a straightforward operation, since it involves many different aspects or facets. In this paper we propose eight different semantic lenses through which these facets may be viewed, and describe and exemplify the ontologies by which these lenses may be implemented.
Silvio Peroni, David M. Shotton, Fabio Vitali
ACM Symposium on Document Engineering1
2012 The Live OWL Documentation Environment: A Tool for the Automatic Generation of Ontology Documentation
Silvio Peroni, David M. Shotton, Fabio Vitali
EKAW1
2012 Latest Developments to LODE
Silvio Peroni, David M. Shotton, Fabio Vitali
EKAW1
2012 FaBiO and CiTO: Ontologies for describing bibliographic resources and citations
Silvio Peroni, David M. Shotton
J. Web Semant.1
2011 A Novel Approach to Visualizing and Navigating Ontologies
Enrico Motta, Paul Mulholland, Silvio Peroni, Mathieu d'Aquin, José Manuél Gómez-Pérez, Victor Mendez, Fouad Zablith
ISWC (1)3
2011 A Semantic Web approach to everyday overlapping markup
abstract
Overlapping structures in XML are not symptoms of a misunderstanding of the intrinsic characteristics of a text document nor evidence of extreme scholarly requirements far beyond those needed by the most common XML-based applications. On the contrary, overlaps have started to appear in a large number of incredibly popular applications hidden under the guise of syntactical tricks to the basic hierarchy of the XML data format. Unfortunately, syntactical tricks have the drawback that the affected structures require complicated workarounds to support even the simplest query or usage. In this article, we present Extremely Annotational Resource Description Framework (RDF) Markup (EARMARK), an approach to overlapping markup that simplifies and streamlines the management of multiple hierarchies on the same content, and provides an approach to sophisticated queries and usages over such structures without the need of ad-hoc applications, simply by using Semantic Web tools and languages. We compare how relevant tasks (e.g., the identification of the contribution of an author in a word processor document) are of some substantial complexity when using the original data format and become more or less trivial when using EARMARK. We finally evaluate positively the memory and disk requirements of EARMARK documents in comparison to Open Office and Microsoft Word XML-based formats.
Angelo Di Iorio, Silvio Peroni, Fabio Vitali
J. Assoc. Inf. Sci. Technol.2
2010 Handling Markup Overlaps Using OWL
Angelo Di Iorio, Silvio Peroni, Fabio Vitali
EKAW2
2009 Annotations with EARMARK for arbitrary, overlapping and out-of order markup
abstract
In this paper we propose a novel approach to markup, called Extreme Annotational RDF Markup (EARMARK), using RDF and OWL to annotate features in text content that cannot be mapped with usual markup languages. EARMARK provides a unifying framework to handle tree-based XML features as well as more complex markup for non-XML scenarios such as overlapping elements, repeated and non-contiguous ranges and structured attributes. EARMARK includes and expands the principles of XML markup, RDFa inline annotations and existing approaches to overlapping markup such as LMNL and TexMecs. EARMARK documents can also be linearized into plain XML by choosing any of a number of strategies to express a tree-based subset of the annotations as an XML structure and fitting in the remaining annotations through a number of "tricks", markup expedients for hierarchical linearization of non-hierarchical features. EARMARK provides a solid platform for providing vocabulary-independent declarative support to advanced document features such as transclusion, overlapping and out-of-order annotations within a conceptually insensitive environment such as XML, and does so by exploiting recent semantic web concepts and languages.
Silvio Peroni, Fabio Vitali
ACM Symposium on Document Engineering1