VLDB 2026 Research / reviewers in the wild / expert
Danilo Dessì
dblp:204/9864
· DBLP profile ↗
14ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-3843-3285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Research Knowledge Graphs: The Shifting Paradigm of Scholarly Information Representation
Matthäus Zloch, Danilo Dessì, Jennifer D'Souza 0001, Leyla Jael Castro, Benjamin Zapilko, Saurav Karmakar, Brigitte Mathiak, Markus Stocker, Wolfgang Otto 0002, Sören Auer, Stefan Dietze |
ESWC (2) | 2 |
| 2025 | TeleScope A Longitudinal Dataset for Investigating Online Discourse and Information Interaction on TelegramabstractTelegram is a globally popular instant messaging platform known for its strong emphasis on security, privacy, and unique social networking features. It has recently emerged as the host for various cross-domain analysis and research works, such as social media influence, propaganda studies, and extremism. This paper introduces TeleScope, an extensive dataset suite that, to our knowledge, is the largest of its kind. It comprises metadata for about 500K Telegram channels and downloaded message metadata for about 71K public channels, accounting for around 120M crawled messages. We also release channel connections and user interaction data built using Telegram’s message-forwarding feature to study multiple use cases, such as information spread and message-forwarding patterns. In addition, we provide data enrichments, such as language detection, active message posting periods for each channel, and Telegram entities extracted from messages, that enable online discourse analysis beyond what is possible with the original data alone. The dataset is designed for diverse applications, independent of specific research objectives, and sufficiently versatile to facilitate the replication of social media studies comparable to those conducted on platforms like X (formerly Twitter). Susmita Gangopadhyay, Danilo Dessì, Dimitar Dimitrov 0002, Stefan Dietze |
ICWSM | 2 |
| 2025 | Knowledge graph validation by integrating LLMs and human-in-the-loopabstractEnsuring the quality of knowledge graphs (KGs) is crucial for the success of the intelligent applications they support. Recent advances in large language models (LLMs) have demonstrated human-level performance across various tasks, raising the question of their potential for KG validation. In this work, we explore the role of LLMs in human-centric KG validation workflows, examining different collaboration strategies between LLMs and domain experts. We propose and evaluate nine distinct approaches, ranging from fully automated validation to hybrid methods that combine expert oversight with AI assistance. These workflows are tested within a real-world KG construction pipeline used to generate the Computer Science Knowledge Graph (CS-KG), a large-scale resource designed to support scientometric tasks such as trend forecasting and hypothesis generation. CS-KG comprises 41 million statements represented as 350 million triples within the Computer Science domain. Our findings show that integrating LLMs into the CS-KG verification process enhances precision by 12%, improving alignment with expert-level validation. However, this comes at the cost of recall, resulting in a 5% decrease in the overall F1 score. In contrast, a hybrid approach which involves both human-in-the-loop and LLM modules, yields the best overall results, improving F1 score by 5% with minimal human involvement. • LLMs demonstrate weak performance as standalone knowledge graph validators. • LLMs in combination with other automated validation methods reach human-level quality. • Human-LLM collaboration balances trade-offs between precision and recall. • HiL involvement for conflicts between automated validators reduces manual effort. Stefani Tsaneva, Danilo Dessì, Francesco Osborne, Marta Sabou |
Inf. Process. Manag. | 2 |
| 2025 | Research hypothesis generation over scientific knowledge graphsabstractGenerating research hypotheses is a crucial step in scientific investigation that involves the creation of precise, verifiable, and logically valid statements that can be empirically examined. Therefore, many efforts have been made to automate or assist this process through the use of various Artificial Intelligence solutions. However, most existing methods are tailored to very specific domains, particularly within the biomedical field. There have been recent attempts to formalize hypothesis generation as a link prediction task over knowledge graphs. This solution is potentially domain-independent and applicable across diverse disciplines. Nevertheless, current approaches for link prediction, which typically rely on embedding models or path-based methods, have shown limited success in accurately predicting new hypotheses. To address these limitations, this paper introduces ResearchLink, an innovative and domain-independent methodology for hypothesis generation over knowledge graphs. ResearchLink combines path-based features and knowledge graph embeddings with text embeddings, capturing the semantic context of entities within a given corpus, and integrates additional information from bibliometric databases to improve research collaboration predictions. To conduct a rigorous evaluation of ResearchLink, we constructed CSKG-600, a new dataset for hypothesis generation, consisting of 600 statements that were manually labelled by domain experts. ResearchLink achieved outstanding performance (78.7% P@20), significantly outperforming alternative approaches such as TransH (71.8%), TransD (71.7%), and RotatE (70.7%). Agustín Borrego, Danilo Dessì, Daniel Ayala Hernández, Inma Hernández, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, David Ruiz 0001, Enrico Motta |
Knowl. Based Syst. | 2 |
| 2024 | Citation prediction by leveraging transformers and natural language processing heuristicsabstractIn scientific papers, it is common practice to cite other articles to substantiate claims, provide evidence for factual assertions, reference limitations, and research gaps, and fulfill various other purposes. When authors include a citation in a given sentence, there are two considerations they need to take into account: (i) where in the sentence to place the citation and (ii) which citation to choose to support the underlying claim. In this paper, we focus on the first task as it allows multiple potential approaches that rely on the researcher’s individual style and the specific norms and conventions of the relevant scientific community. We propose two automatic methodologies that leverage transformers architecture for either solving a Mask-Filling problem or a Named Entity Recognition problem. On top of the results of the proposed methodologies, we apply ad-hoc Natural Language Processing heuristics to further improve their outcome. We also introduce s2orc-9K, an open dataset for fine-tuning models on this task. A formal evaluation demonstrates that the generative approach significantly outperforms five alternative methods when fine-tuned on the novel dataset. Furthermore, this model’s results show no statistically significant deviation from the outputs of three senior researchers. Davide Buscaldi, Danilo Dessì, Enrico Motta, Marco Murgia, Francesco Osborne, Diego Reforgiato Recupero |
Inf. Process. Manag. | 2 |
| 2022 | CS-KG: A Large-Scale Knowledge Graph of Research Entities and Claims in Computer Science
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta |
ISWC | 1 |
| 2022 | Guest Editorial of the FGCS Special Issue on Advances in Intelligent Systems for Online Education
Geoffray Bonnin, Danilo Dessì, Gianni Fenu, Martin Hlosta, Mirko Marras, Harald Sack |
Future Gener. Comput. Syst. | 2 |
| 2022 | LexTex: a framework to generate lexicons using WordNet word senses in domain specific categories
Danilo Dessì, Diego Reforgiato Recupero |
J. Intell. Inf. Syst. | 1 |
| 2022 | SCICERO: A deep learning and NLP approach for generating scientific knowledge graphs in the computer science domain
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta |
Knowl. Based Syst. | 1 |
| 2021 | L2D 2021: First International Workshop on Enabling Data-Driven Decisions from Learning on the WebabstractBy offering courses and resources, learning platforms on the Web have been attracting lots of participants, and the interactions with these systems have generated a vast amount of learning-related data. Their collection, processing and analysis have promoted a significant growth of learning analytics and have opened up new opportunities for supporting and assessing educational experiences. To provide all the stakeholders involved in the educational process with a timely guidance, being able to understand student's behavior and enable models which provide data-driven decisions pertaining to the learning domain is a primary property of online platforms, aiming at maximizing learning outcomes. In this workshop, we focus on collecting new contributions in this emerging area and on providing a common ground for researchers and practitioners (Website: https://mirkomarras.github.io/l2d-wsdm2021). Danilo Dessì, Tanja Käser, Mirko Marras, Elvira Popescu, Harald Sack |
WSDM | 1 |
| 2021 | Generating knowledge graphs by employing Natural Language Processing and Machine Learning techniques within the scholarly domain
Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta |
Future Gener. Comput. Syst. | 1 |
| 2020 | AI-KG: An Automatically Generated Knowledge Graph of Artificial IntelligenceabstractScientific knowledge has been traditionally disseminated and preserved through research articles published in journals, conference proceedings, and online archives. However, this article-centric paradigm has been often criticized for not allowing to automatically process, categorize, and reason on this knowledge. An alternative vision is to generate a semantically rich and interlinked description of the content of research publications. In this paper, we present the Artificial Intelligence Knowledge Graph (AI-KG), a large-scale automatically generated knowledge graph that describes 820K research entities. AI-KG includes about 14M RDF triples and 1.2M reified statements extracted from 333K research publications in the field of AI, and describes 5 types of entities (tasks, methods, metrics, materials, others) linked by 27 relations. AI-KG has been designed to support a variety of intelligent services for analyzing and making sense of research dynamics, supporting researchers in their daily job, and helping to inform decision-making in funding bodies and research policymakers. AI-KG has been generated by applying an automatic pipeline that extracts entities and relationships using three tools: DyGIE++, Stanford CoreNLP, and the CSO Classifier. It then integrates and filters the resulting triples using a combination of deep learning and semantic technologies in order to produce a high-quality knowledge graph. This pipeline was evaluated on a manually crafted gold standard, yielding competitive results. AI-KG is available under CC BY 4.0 and can be downloaded as a dump or queried via a SPARQL endpoint. Danilo Dessì, Francesco Osborne, Diego Reforgiato Recupero, Davide Buscaldi, Enrico Motta, Harald Sack |
ISWC (2) | 1 |
| 2018 | COCO: Semantic-Enriched Collection of Online Courses at Scale with Experimental Use Cases
Danilo Dessì, Gianni Fenu, Mirko Marras, Diego Reforgiato Recupero |
WorldCIST (2) | 1 |
| 2018 | SuperNoder: a tool to discover over-represented modular structures in networksabstractBACKGROUND: Networks whose nodes have labels can seem complex. Fortunately, many have substructures that occur often ("motifs"). A societal example of a motif might be a household. Replacing such motifs by named supernodes reduces the complexity of the network and can bring out insightful features. Doing so repeatedly may give hints about higher level structures of the network. We call this recursive process Recursive Supernode Extraction. RESULTS: This paper describes algorithms and a tool to discover disjoint (i.e. non-overlapping) motifs in a network, replacing those motifs by new nodes, and then recursing. We show applications in food-web and protein-protein interaction (PPI) networks where our methods reduce the complexity of the network and yield insights. CONCLUSIONS: SuperNoder is a web-based and standalone tool which enables the simplification of big graphs based on the reduction of high frequency motifs. It applies various strategies for identifying disjoint motifs with the goal of enhancing the understandability of networks. Danilo Dessì, Jacopo Cirrone, Diego Reforgiato Recupero, Dennis E. Shasha |
BMC Bioinform. | 1 |