Stefano Ferilli

dblp:46/4447 · DBLP profile ↗
← Back
26ranked-venue papers in the field
14as first author
6since 2021 · last 2024
0000-0003-1118-0601ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 10 (4 first)Other / Interdisciplinary · 6 (3 first)Data Mining & Knowledge Discovery · 5 (3 first)Information Retrieval & Web Search · 5 (4 first)
YearPublicationVenuePosition
2024 Mining Literary Trends: A Tool for Digital Library Analysis
Eleonora Bernasconi, Stefano Ferilli
TPDL (1)2
2024 An Holistic Approach to Diagnostic and Therapeutic Care Pathways Management
abstract
The European Commission EU4Health program (2021-2027) is launched after the severe health crisis caused by COVID-19 to support member states in long-term health challenges to build more resilient health systems aimed at reducing inequalities in access to healthcare. In Italy, the PNRR program has among its goals the enhancement of the Diagnostic Therapeutic Care Pathways, particularly their complete informatisation to reduce the gap currently present at the regional and, in some cases, at the hospital level. This paper describes a possible AI framework as a starting point for a potential solution to this goal. The proposed solution involves the use of GraphDB for information persistence and evolved process management methods for the implementation of Care Pathways.
Domenico Redavid, Stefano Ferilli
KEOD2
2023 Holistic Graph-Based Document Representation and Management for Open Science
Stefano Ferilli, Davide Di Pierro 0001, Domenico Redavid
TPDL1
2023 Graph Databases for Diachronic Language Data Modelling
Barbara McGillivray, Pierluigi Cassotti, Davide Di Pierro 0001, Paola Marongiu, Anas Fahad Khan, Stefano Ferilli, Pierpaolo Basile
LDK6
2023 The World Literature Knowledge Graph
Marco Stranisci, Eleonora Bernasconi, Viviana Patti, Stefano Ferilli, Miguel Ceriani, Rossana Damiano
ISWC4
2022 Holistic Graph-Based Representation and AI for Digital Library Management
Stefano Ferilli
TPDL1
2019 Ensembles of density estimators for positive-unlabeled learning
Teresa M. A. Basile, Nicola Di Mauro, Floriana Esposito, Stefano Ferilli, Antonio Vergari
J. Intell. Inf. Syst.4
2019 Activity prediction in process mining using the WoMan framework
Stefano Ferilli, Sergio Angelastro
J. Intell. Inf. Syst.1
2018 Extending expressivity and flexibility of abductive logic programming
Stefano Ferilli
J. Intell. Inf. Syst.1
2018 A multi-strategy approach to structural analogy making
Fabio Leuzzi, Stefano Ferilli
J. Intell. Inf. Syst.2
2016 A sentence structure-based approach to unsupervised author identification
Stefano Ferilli
J. Intell. Inf. Syst.1
2016 Predicate invention-based specialization in Inductive Logic Programming
Stefano Ferilli
J. Intell. Inf. Syst.1
2015 Sentiment analysis as a text categorization task: A study on feature and algorithm selection for Italian language
abstract
The availability on the Internet of huge amounts of blog posts, messages and comments allows to study the attitude of people on various topics. Sentiment Analysis, Opinion Mining and Emotion Analysis denote the area of research in Computer Science aimed at studying, analyzing and classifying text documents based on the underlying opinions expressed by their authors on various topics. While this is a tough task, because it is related to psychological aspects that are not always immediately evident in the lexical and syntactical aspects of the sentences, its importance may be paramount for several applications such as market analysis, political polls, etc. Fundamental pre-processing techniques for this task come from the area of Natural Language Processing, which may pose additional problems when the language of interest is different than English, and thus less (or less reliable) resources are available to extract the needed data from the text. This paper studies the performance of Sentiment Analysis, seen as a Text Categorization task, depending on the use of different classifiers and different features. While the approach is general, we focus on texts in Italian. The outcomes suggest which experimental settings can be most profitably used in this landscape, and show that significantly good results can be obtained.
Stefano Ferilli, Berardina De Carolis, Floriana Esposito, Domenico Redavid
DSAA1
2015 Logic-Based Incremental Process Mining
Stefano Ferilli, Domenico Redavid, Floriana Esposito
ECML/PKDD (3)1
2014 Abstract argumentation for reading order detection
abstract
Detecting the reading order among the layout components of a document's page is fundamental to ensure effectiveness or even applicability of subsequent content extraction steps. While in single-column documents the reading flow can be straightforwardly determined, in more complex documents the task may become very hard. This paper proposes an automatic strategy for identifying the correct reading order of a document page's components based on abstract argumentation. The technique is unsupervised, and works on any kind of document based only on general assumptions about how humans behave when reading documents. Experimental results show that it is effective in more complex cases, and requires less background knowledge, than previous solutions that have been proposed in the literature.
Stefano Ferilli, Domenico Grieco, Domenico Redavid, Floriana Esposito
ACM Symposium on Document Engineering1
2014 Guidelines and Tool for Meaningful OWL-S Services Annotations
abstract
The current tools to create OWL-S annotations have been designed starting from the knowledge engineer’s point of view. Unfortunately, the formalisms underlying Semantic Web languages are often incomprehensible to the developers of Web services. To bridge this gap, it is desirable that developers are provided with suitable tools that do not necessarily require knowledge of these languages in order to create annotations on Web services. With reference to some characteristics of the involved technologies, this work addresses these issues, proposing guidelines that can improve the annotation activity of Web service developers. Following these guidelines, we also designed a tool that allows a Web service developer to annotate Web services without requiring him to have a deep knowledge of Semantic Web languages. A prototype of such a tool is presented and discussed in this paper.
Domenico Redavid, Stefano Ferilli, Berardina De Carolis, Floriana Esposito
KEOD2
2013 Hi-Fi HTML rendering of multi-format documents in DoMinUS
abstract
Digital Libraries collect, organize and provide to end users large quantities of selected documents. While these documents come in a variety of formats, it is desirable that they are delivered to final users in a uniform way. Web formats are a suitable choice for this purpose. While Web documents are very flexible as to layout presentation, that is determined at runtime by the interpreter, documents coming from a library should preserve their original layout when displayed to final users. Using raster images would not allow the user to access the actual content of the document's components (text and images). This paper presents a technique to render in an HTML file the original layout of a document, preserving the peculiarity of its components (text, images, formulas, tables, algorithms). It builds on the DoMInUS framework, that can process documents in several source formats.
Stefano Ferilli, Floriana Esposito, Domenico Redavid
ACM Symposium on Document Engineering1
2013 Finding Critical Cells in Web Tables with SRL: Trying to Uncover the Devil's Tease
abstract
Tables are extremely important components of documents, because they bear very informative content in a compact and structured way. Being able to understand a table's internal organization would allow to extract and reuse the data they contain. This can be reduced to recognizing critical cells only. Since purely algorithmic approaches are unable to deal with the many different table layouts designed to represent particular kinds of information and/or particular perspectives on them, Machine Learning may represent an effective solution. On one hand, the spatial organization of tables puts a strong emphasis on the relationships among cells, on the other, the extreme variability in style, size, and aims of tables requires flexible approaches. This paper proposes the exploitation of a Statistical Relational Learning approach, that is able to model the complex spatial relationships involved in a table structure, by mixing the power of a relational representation formalism with the flexibility of a statistical learning tool. Experiments on a real-world dataset are reported both for single cell classification and for overall table structure recognition, whose results prove the validity of the proposed approach.
Nicola Di Mauro, Floriana Esposito, Stefano Ferilli
ICDAR3
2011 A Contour-Based Progressive Technique for Shape Recognition
abstract
Information Retrieval in large digital document repositories is at the same time a hard and crucial task. While the primary type of information available in documents is usually text, images play a very important role because they pictorially describe concepts that are dealt with in the document. Unfortunately, the semantic gap separating such a visual content from the underlying meaning is very wide. Additionally image processing techniques are usually very demanding in computational resources. Hence, only recently the area of Content-Based Image Retrieval has gained more attention. In this paper we describe a new technique to identify known objects in a picture based on a comparison of the shapes to known models. The comparison works by progressive approximations to save computational resources, and relies on novel algorithmic and representational solutions to improve preliminary shape extraction.
Stefano Ferilli, Teresa M. A. Basile, Floriana Esposito, Marenglen Biba
ICDAR1
2010 A histogram-based technique for automatic threshold assessment in a run length smoothing-based algorithm
abstract
Document layout analysis is crucial in the automatic document processing workflow, because its outcome affects all subsequent processing steps. A first problem concerns the possibility of dealing not only with documents having easy layout, but with so-called non-Manhattan layout documents as well. Another problem is that most available techniques can be applied to scanned document, due to the emphasis in previous decades being put on legacy documents digitization. Conversely, nowadays most documents come directly in digital format, and thus new techniques must be developed. A famous approach proposed in the literature for layout analysis was the RLSA, suitable to scanned black&white images and based the application of Run Length Smoothing and the AND logical operator. A recent variant thereof is based on the application of the OR operator, for which reason has been called RLSO. It exploits a bottom-up approach that proved able to handle even non-Manhattan layouts, on both scanned and natively digital documents. Like RLSA, it is based on the definition of thresholds for the smoothing operator, but the different approach requires different criteria than those that work in RLSA to define proper values. Since this is a hard and unnatural task for an (even expert) user, this paper proposes a technique to automatically define such thresholds for each single document, based on the distribution of spacing therein. Application on selected samples of documents, that aimed at covering a significant landscape of real cases, revealed that the approach is satisfactory for documents characterized by the use of a uniform text font size. It can provide a useful basis also for handling more complex cases.
Stefano Ferilli, Teresa M. A. Basile, Floriana Esposito
Document Analysis Systems1
2009 A Distance-Based Technique for Non-Manhattan Layout Analysis
abstract
Layout analysis is a fundamental step in automatic document processing. Many different techniques have been proposed to perform this task. Some follow a top-down approach: they start by identifying the high level components of the page structure and then recursively split them until basic blocks are found. On the other hand, bottom-up approaches start with the smallest elements (e.g., the pixels in case of digitized document) and then recursively merge them into higher level components. A first limitation of such methods is that most of them are designed to deal only with digitized documents and hence are not applicable to native digital documents which are nowadays pervasive. Furthermore, top-down and most of bottom-up methods are able to process Manhattan layout documents only. In this work, we propose a general bottom-up strategy to tackle the layout analysis of (possibly) non-Manhattan documents, and two specializations of it to handle both bitmap and PS/PDF sources. It was successfully embedded and tested in the DOMINUS document management system.
Stefano Ferilli, Marenglen Biba, Floriana Esposito, Teresa M. A. Basile
ICDAR1
2007 Incremental Learning of First Order Logic Theories for the Automatic Annotations of Web Documents
abstract
Organizing large repositories spread throughout the most diverse Web sites rises the problem of effective storage and efficient retrieval of documents. This can be obtained by selectively extracting from them the significant textual information, contained in peculiar layout components, that in turn depend on the identification of the correct document class. The continuous flow of new and different documents in a weakly structured environment like the Web calls for in- crementality, as the ability to continuously update or revise a faulty knowledge previously acquired, while the need to express structural relations among layout components suggest the exploitation of a powerful and symbolic representation language. This paper proposes the application of incremental first-order logic learning techniques in the document layout preprocessing steps, supported by good results obtained in experiments on a real dataset.
Floriana Esposito, Stefano Ferilli, Nicola Di Mauro, Teresa M. A. Basile
ICDAR2
2007 Inference of abduction theories for handling incompleteness in first-order learning
Floriana Esposito, Stefano Ferilli, Teresa M. A. Basile, Nicola Di Mauro
Knowl. Inf. Syst.2
2005 Learning User Profiles from Text in e-Commerce
Marco de Gemmis, Pasquale Lops, Stefano Ferilli, Nicola Di Mauro, Teresa M. A. Basile, Giovanni Semeraro
ADMA3
2005 On the LearnAbility of Abstraction Theories from Observations for Relational Learning
Stefano Ferilli, Teresa M. A. Basile, Nicola Di Mauro, Floriana Esposito
ECML1
2005 Intelligent Document Processing
abstract
Digital repositories raise the need for an effective and efficient retrieval of the stored material. In this paper, we propose the intensive application of intelligent techniques to the steps of document layout analysis, document image classification and understanding on digital documents. Specifically, the complex interrelation existing among layout components, that are fundamental to assign them the proper semantic role, suggest the exploitation of first-order representations in some learning steps. Results obtained in a prototypical system for scientific conference management prove that the proposed approach can be beneficial both for the layout recognition and for the selection of interesting components of the document, from which extracting the text for categorizing the document according to its topic.
Floriana Esposito, Stefano Ferilli, Teresa M. A. Basile, Nicola Di Mauro
ICDAR2