EDBT 2026 Demo / reviewers in the wild / expert
Francesco Osborne
dblp:38/10105
· DBLP profile ↗
31ranked-venue papers in the field
10as first author
11since 2021 · last 2026
0000-0001-6557-3131ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 24 (10 first)Information Retrieval & Web Search · 5Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large language models for scholarly ontology generation: An extensive analysis in the engineering fieldabstractOntologies of research topics are crucial for structuring scientific knowledge, enabling scientists to navigate vast amounts of research, and forming the backbone of intelligent systems such as search engines and recommendation systems. However, manual creation of these ontologies is expensive, slow, and often results in outdated and overly general representations. As a solution, researchers have been investigating ways to automate or semi-automate the process of generating these ontologies. One of the key challenges in this domain is accurately assessing the semantic relationships between pairs of research topics. This paper presents an analysis of the capabilities of large language models (LLMs) in identifying such relationships, with a specific focus on the field of engineering. To this end, we introduce a novel benchmark based on the IEEE Thesaurus for evaluating the task of identifying three types of semantic relations between pairs of topics: broader , narrower , and same-as . Our study evaluates the performance of seventeen LLMs, which differ in scale, accessibility (open vs. proprietary), and model type (full vs. quantised), while also assessing four zero-shot reasoning strategies. Several models with varying architectures and sizes have achieved excellent results on this task, including Mixtral-8 × 7B, Dolphin-Mistral-7B, and Claude 3 Sonnet, with F1-scores of 0.847, 0.920, and 0.967, respectively. Furthermore, our findings demonstrate that smaller, quantised models, when optimised through prompt engineering, can achieve strong performance while requiring very limited computational resources. Tanay Aggarwal, Angelo A. Salatino, Francesco Osborne, Enrico Motta |
Inf. Process. Manag. | 3 |
| 2026 | Evaluating the effectiveness of fine-tuning in financial NLP: The case of Social Trading Action DetectionabstractFinancial Natural Language Processing crucially leverages social media for market insights. However, most existing methods for this purpose rely on simple sentiment analysis models, which fail to capture the concrete trading intentions expressed in these discussions. While Large Language Models (LLMs) offer a promising alternative to simplistic sentiment analysis, the actual benefits of fine-tuning across different model families remain unclear in noisy, domain-specific contexts like online forums. To address this gap, we present a comprehensive assessment of the advantages and limitations of fine-tuning for Social Trading Action Detection (STAD), a novel task that aims to classify online posts into actionable categories, namely buy, sell, or other. In addition, we introduce FinReddit-2K, a manually annotated dataset consisting of 2123 Reddit posts, designed to serve as a benchmark for this task. Our experimental analysis goes beyond standard performance metrics and identifies both the types of errors that fine-tuning can successfully mitigate and those that it may inadvertently introduce. Through a systematic evaluation of 57 models, comparing 14 traditional models with 23 zero-shot LLMs and 20 fine-tuned variants, our results show that fine-tuning yields an average F1-score improvement of +15.1%. The best-performing model, a fine-tuned Mistral-7B, achieves an F1-score of 86.0%, although our analysis reveals that fine-tuning fails to produce meaningful performance gains in several scenarios. Simone D'Amico, Andrea Maurino, Francesco Osborne, Giancarlo Sperlì |
Inf. Process. Manag. | 3 |
| 2025 | Knowledge graph validation by integrating LLMs and human-in-the-loopabstractEnsuring the quality of knowledge graphs (KGs) is crucial for the success of the intelligent applications they support. Recent advances in large language models (LLMs) have demonstrated human-level performance across various tasks, raising the question of their potential for KG validation. In this work, we explore the role of LLMs in human-centric KG validation workflows, examining different collaboration strategies between LLMs and domain experts. We propose and evaluate nine distinct approaches, ranging from fully automated validation to hybrid methods that combine expert oversight with AI assistance. These workflows are tested within a real-world KG construction pipeline used to generate the Computer Science Knowledge Graph (CS-KG), a large-scale resource designed to support scientometric tasks such as trend forecasting and hypothesis generation. CS-KG comprises 41 million statements represented as 350 million triples within the Computer Science domain. Our findings show that integrating LLMs into the CS-KG verification process enhances precision by 12%, improving alignment with expert-level validation. However, this comes at the cost of recall, resulting in a 5% decrease in the overall F1 score. In contrast, a hybrid approach which involves both human-in-the-loop and LLM modules, yields the best overall results, improving F1 score by 5% with minimal human involvement. • LLMs demonstrate weak performance as standalone knowledge graph validators. • LLMs in combination with other automated validation methods reach human-level quality. • Human-LLM collaboration balances trade-offs between precision and recall. • HiL involvement for conflicts between automated validators reduces manual effort. Stefani Tsaneva, Danilo Dessì, Francesco Osborne, Marta Sabou |
Inf. Process. Manag. | 3 |
| 2024 | Capturing the Viewpoint Dynamics in the News Domain
Enrico Motta, Francesco Osborne, Martino M. L. Pulici, Angelo A. Salatino, Iman Naja |
EKAW | 2 |
| 2024 | Large Language Models for Scientific Question Answering: An Extensive Analysis of the SciQA Benchmark
Jens Lehmann 0001, Antonello Meloni, Enrico Motta, Francesco Osborne, Diego Reforgiato Recupero, Angelo A. Salatino, Sahar Vahdati |
ESWC (1) | 4 |
| 2024 | Workshop on Deep Learning and Large Language Models for Knowledge Graphs (DL4KG)abstractThe use of Knowledge Graphs (KGs) which constitute large networks of real-world entities and their interrelationships, has grown rapidly. A substantial body of research has emerged, exploring the integration of deep learning (DL) and large language models (LLMs) with KGs. This workshop aims to bring together leading researchers in the field to discuss and foster collaborations on the intersection of KG and DL/LLMs. Mehwish Alam, Davide Buscaldi, Michael Cochez, Genet Asefa Gesese, Francesco Osborne, Diego Reforgiato Recupero |
KDD | 5 |
| 2024 | Citation prediction by leveraging transformers and natural language processing heuristicsabstractIn scientific papers, it is common practice to cite other articles to substantiate claims, provide evidence for factual assertions, reference limitations, and research gaps, and fulfill various other purposes. When authors include a citation in a given sentence, there are two considerations they need to take into account: (i) where in the sentence to place the citation and (ii) which citation to choose to support the underlying claim. In this paper, we focus on the first task as it allows multiple potential approaches that rely on the researcher’s individual style and the specific norms and conventions of the relevant scientific community. We propose two automatic methodologies that leverage transformers architecture for either solving a Mask-Filling problem or a Named Entity Recognition problem. On top of the results of the proposed methodologies, we apply ad-hoc Natural Language Processing heuristics to further improve their outcome. We also introduce s2orc-9K, an open dataset for fine-tuning models on this task. A formal evaluation demonstrates that the generative approach significantly outperforms five alternative methods when fine-tuned on the novel dataset. Furthermore, this model’s results show no statistically significant deviation from the outputs of three senior researchers. Davide Buscaldi, Danilo Dessì, Enrico Motta, Marco Murgia, Francesco Osborne, Diego Reforgiato Recupero |
Inf. Process. Manag. | 5 |
| 2023 | Ontology-Based Generation of Data Platform AssetsabstractThe design and management of modern big data platforms are extremely complex. It requires carefully integrating multiple storage and computational platforms as well as implementing approaches to protect and audit data access. Therefore, onboarding new data and implementing new data transformation processes is typically time-consuming and expensive. In many cases, enterprises construct their data platforms without a clear distinction between logical and technical concerns. Consequently, these platforms lack sufficient abstraction and are closely tied to particular technologies, making the adaptation to technological evolution very costly. This paper illustrates a novel approach to designing data platform models based on a formal ontology that structures various domain components into an accessible knowledge graph. We also describe the preliminary version of AGILE-DM, a novel ontology that we built for this purpose. Our solution is flexible, technologically agnostic, and more adaptable to changes and technical advancements. Vincenzo De Leo, Gianni Fenu, David Greco, Nicolo Bidotti, Paolo Platter, Enrico Motta, Andrea Giovanni Nuzzolese, Francesco Osborne, Diego Reforgiato Recupero |
IEEE Big Data | 8 |
| 2023 | AIDA-Bot 2.0: Enhancing Conversational Agents with Knowledge Graphs for Analysing the Research Landscape
Antonello Meloni, Simone Angioni, Angelo A. Salatino, Francesco Osborne, Aliaksandr Birukou, Diego Reforgiato Recupero, Enrico Motta |
ISWC | 4 |
| 2022 | Leveraging Knowledge Graph Technologies to Assess Journals and Conferences at Springer Nature
Simone Angioni, Angelo A. Salatino, Francesco Osborne, Aliaksandr Birukou, Diego Reforgiato Recupero, Enrico Motta |
ISWC | 3 |
| 2022 | CS-KG: A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta |
ISWC | 2 |
| 2020 | ResearchFlow: Understanding the Knowledge Flow Between Academia and Industry
Angelo A. Salatino, Francesco Osborne, Enrico Motta |
EKAW | 2 |
| 2020 | AI-KG: An Automatically Generated Knowledge Graph of Artificial IntelligenceabstractScientific knowledge has been traditionally disseminated and preserved through research articles published in journals, conference proceedings, and online archives. However, this article-centric paradigm has been often criticized for not allowing to automatically process, categorize, and reason on this knowledge. An alternative vision is to generate a semantically rich and interlinked description of the content of research publications. In this paper, we present the Artificial Intelligence Knowledge Graph (AI-KG), a large-scale automatically generated knowledge graph that describes 820K research entities. AI-KG includes about 14M RDF triples and 1.2M reified statements extracted from 333K research publications in the field of AI, and describes 5 types of entities (tasks, methods, metrics, materials, others) linked by 27 relations. AI-KG has been designed to support a variety of intelligent services for analyzing and making sense of research dynamics, supporting researchers in their daily job, and helping to inform decision-making in funding bodies and research policymakers. AI-KG has been generated by applying an automatic pipeline that extracts entities and relationships using three tools: DyGIE++, Stanford CoreNLP, and the CSO Classifier. It then integrates and filters the resulting triples using a combination of deep learning and semantic technologies in order to produce a high-quality knowledge graph. This pipeline was evaluated on a manually crafted gold standard, yielding competitive results. AI-KG is available under CC BY 4.0 and can be downloaded as a dump or queried via a SPARQL endpoint. Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta, Harald Sack |
ISWC (2) | 2 |
| 2019 | The CSO Classifier: Ontology-Driven Detection of Research Topics in Scholarly Articles
Angelo A. Salatino, Francesco Osborne, Thiviyan Thanapalasingam, Enrico Motta |
TPDL | 2 |
| 2019 | Improving Editorial Workflow and Metadata Quality at Springer Nature
Angelo A. Salatino, Francesco Osborne, Aliaksandr Birukou, Enrico Motta |
ISWC (2) | 2 |
| 2018 | Pragmatic Ontology Evolution: Reconciling User Requirements and Application Performance
Francesco Osborne, Enrico Motta |
ISWC (1) | 1 |
| 2018 | The Computer Science Ontology: A Large-Scale Taxonomy of Research AreasabstractOntologies of research areas are important tools for characterising, exploring, and analysing the research landscape. Some fields of research are comprehensively described by large-scale taxonomies, e.g., MeSH in Biology and PhySH in Physics. Conversely, current Computer Science taxonomies are coarse-grained and tend to evolve slowly. For instance, the ACM classification scheme contains only about 2K research topics and the last version dates back to 2012. In this paper, we introduce the Computer Science Ontology (CSO), a large-scale, automatically generated ontology of research areas, which includes about 26K topics and 226K semantic relationships. It was created by applying the Klink-2 algorithm on a very large dataset of 16M scientific articles. CSO presents two main advantages over the alternatives: (i) it includes a very large number of topics that do not appear in other classifications, and (ii) it can be updated automatically by running Klink-2 on recent corpora of publications. CSO powers several tools adopted by the editorial team at Springer Nature and has been used to enable a variety of solutions, such as classifying research publications, detecting research communities, and predicting research trends. To facilitate the uptake of CSO we have developed the CSO Portal, a web application that enables users to download, explore, and provide granular feedback on CSO at different levels. Users can use the portal to rate topics and relationships, suggest missing relationships, and visualise sections of the ontology. The portal will support the publication of and access to regular new releases of CSO, with the aim of providing a comprehensive resource to the various communities engaged with scholarly data. Angelo A. Salatino, Thiviyan Thanapalasingam, Andrea Mannocci, Francesco Osborne, Enrico Motta |
ISWC (2) | 4 |
| 2018 | Ontology-Based Recommendation of Editorial Products
Thiviyan Thanapalasingam, Francesco Osborne, Aliaksandr Birukou, Enrico Motta |
ISWC (2) | 2 |
| 2017 | Forecasting the Spreading of Technologies in Research CommunitiesabstractTechnologies such as algorithms, applications and formats are an important part of the knowledge produced and reused in the research process. Typically, a technology is expected to originate in the context of a research area and then spread and contribute to several other fields. For example, Semantic Web technologies have been successfully adopted by a variety of fields, e.g., Information Retrieval, Human Computer Interaction, Biology, and many others. Unfortunately, the spreading of technologies across research areas may be a slow and inefficient process, since it is easy for researchers to be unaware of potentially relevant solutions produced by other research communities. In this paper, we hypothesise that it is possible to learn typical technology propagation patterns from historical data and to exploit this knowledge i) to anticipate where a technology may be adopted next and ii) to alert relevant stakeholders about emerging and relevant technologies in other fields. To do so, we propose the Technology-Topic Framework, a novel approach which uses a semantically enhanced technology-topic model to forecast the propagation of technologies to research areas. A formal evaluation of the approach on a set of technologies in the Semantic Web and Artificial Intelligence areas has produced excellent results, confirming the validity of our solution. Francesco Osborne, Andrea Mannocci, Enrico Motta |
K-CAP | 1 |
| 2016 | Ontology Forecasting in Scientific Literature: Semantic Concepts Prediction Based on Innovation-Adoption Priors
Amparo Elizabeth Cano, Francesco Osborne, Angelo A. Salatino |
EKAW | 2 |
| 2016 | TechMiner: Extracting Technologies from Academic Publications
Francesco Osborne, Hélène de Ribaupierre, Enrico Motta |
EKAW | 1 |
| 2016 | Automatic Classification of Springer Nature Proceedings with Smart Topic Miner
Francesco Osborne, Angelo A. Salatino, Aliaksandr Birukou, Enrico Motta |
ISWC (2) | 1 |
| 2015 | Klink-2: Integrating Multiple Web Sources to Generate Semantic Topic Networks
Francesco Osborne, Enrico Motta |
ISWC (1) | 1 |
| 2015 | Property-based Semantic Similarity and Relatedness for Improving Recommendation Accuracy and DiversityabstractThe authors introduce new measures of semantic similarity and relatedness for ontological concepts, based on the properties associated to them. They consider two concepts similar if, for some properties they have in common, they also have the same values assigned to these properties. On the other hand, the authors consider two concepts related if they have the same values assigned to different properties. These measures are used in the propagation of user interest values in ontology-based user models to other similar or related concepts in the domain. The authors tested their algorithm in event recommendation domain and in recipe domain and showed that property-based propagation based on similarity outperforms the standard edge-based propagation. Adding relatedness as a criterion for propagation improves diversity without sacrificing accuracy. In addition, assigning a certain relevance to each property improves the accuracy of recommendation. Finally, the property-based spreading activation is effective for cross-domain recommendation. Silvia Likavec, Francesco Osborne, Federica Cena |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2014 | Inferring Semantic Relations by User Feedback
Francesco Osborne, Enrico Motta |
EKAW | 1 |
| 2014 | A Hybrid Semantic Approach to Building Dynamic Maps of Research Communities
Francesco Osborne, Giuseppe Scavo, Enrico Motta |
EKAW | 1 |
| 2014 | Identifying Diachronic Topic-Based Research Communities by Clustering Shared Research Trajectories
Francesco Osborne, Giuseppe Scavo, Enrico Motta |
ESWC | 1 |
| 2014 | User data discovery and aggregation: The CS-UDD algorithm
Francesca Carmagnola, Francesco Osborne, Ilaria Torre 0001 |
Inf. Sci. | 2 |
| 2013 | Exploring Scholarly Data with Rexplore
Francesco Osborne, Enrico Motta, Paul Mulholland |
ISWC (1) | 1 |
| 2013 | Anisotropic propagation of user interests in ontology-based user models
Federica Cena, Silvia Likavec, Francesco Osborne |
Inf. Sci. | 3 |
| 2012 | Mining Semantic Relations between Research Areas
Francesco Osborne, Enrico Motta |
ISWC (1) | 1 |