Slavko Zitnik

dblp:120/1960 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-3452-1106ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 MOOC on Linguistic Linked Data
Jorge Gracia, Slavko Zitnik, Maxim Ionov, Christian Chiarcos, Dagmar Gromann, Francesco Mambrini, Marco Passarotti, Armando Stellato, John P. McCrae, Gilles Sérasset, Andon Tchechmedjiev, Sara Carvalho, Penny Labropoulou, Rute Costa
ESWC (2)2
2024 MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations
abstract
Understanding the relation between the meanings of words is an important part of comprehending natural language. Prior work has either focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs), with some exceptions. Given the rarity of highly multilingual benchmarks, it is unclear to what extent PLMs capture relational knowledge and are able to transfer it across languages. To start addressing this question, we propose MultiLexBATS, a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages, such as Bambara, Lithuanian, and Albanian. As experiment on cross-lingual transfer of relational knowledge, we test the PLMs’ ability to (1) capture analogies across languages, and (2) predict translation targets. We find considerable differences across relation types and languages with a clear preference for hypernymy and antonymy as well as romance languages.
Dagmar Gromann, Hugo Gonçalo Oliveira, Lucia Pitarch, Elena Apostol, Jordi Bernad, Eliot Bytyci, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabík, Jorge Gracia, Letizia Granata, Anas Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia di Buono, Ana Ostroski Anic, Sigita Rackeviciene, Ricardo Rodrigues 0001, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stankovic, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
LREC/COLING30
2024 SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene
abstract
This paper introduces the upgrade of a training corpus for linguistic annotation of modern standard Slovene. The enhancement spans both the size of the corpus and the depth of annotation layers. The revised SUK 1.0 corpus, building on its predecessor ssj500k 2.3, has doubled in size, containing over a million tokens. This expansion integrates three preexisting open-access datasets, all of which have undergone automatic tagging and meticulous manual review across multiple annotation layers, each represented in varying proportions. These layers span tokenization, segmentation, lemmatization, MULTEXT-East morphology, Universal Dependencies, JOS-SYN syntax, semantic role labeling, named entity recognition, and the newly incorporated coreferences. The paper illustrates the annotation processes for each layer while also presenting the results of the new CLASSLA-Stanza annotation tool, trained on the SUK corpus data. As one of the fundamental language resources of modern Slovene, the SUK corpus calls for constant development, as outlined in the concluding section.
Spela Arhar Holdt, Jaka Cibej, Kaja Dobrovoljc, Tomaz Erjavec, Polona Gantar, Simon Krek, Tina Munda, Nejc Robida, Luka Tercon, Slavko Zitnik
LREC/COLING10
2024 Multimodal learning for temporal relation extraction in clinical texts
abstract
OBJECTIVES: This study focuses on refining temporal relation extraction within medical documents by introducing an innovative bimodal architecture. The overarching goal is to enhance our understanding of narrative processes in the medical domain, particularly through the analysis of extensive reports and notes concerning patient experiences. MATERIALS AND METHODS: Our approach involves the development of a bimodal architecture that seamlessly integrates information from both text documents and knowledge graphs. This integration serves to infuse common knowledge about events into the temporal relation extraction process. Rigorous testing was conducted on diverse clinical datasets, emulating real-world scenarios where the extraction of temporal relationships is paramount. RESULTS: The performance of our proposed bimodal architecture was thoroughly evaluated across multiple clinical datasets. Comparative analyses demonstrated its superiority over existing methods reliant solely on textual information for temporal relation extraction. Notably, the model showcased its effectiveness even in scenarios where not provided with additional information. DISCUSSION: The amalgamation of textual data and knowledge graph information in our bimodal architecture signifies a notable advancement in the field of temporal relation extraction. This approach addresses the critical need for a more profound understanding of narrative processes in medical contexts. CONCLUSION: In conclusion, our study introduces a pioneering bimodal architecture that harnesses the synergy of text and knowledge graph data, exhibiting superior performance in temporal relation extraction from medical documents. This advancement holds significant promise for improving the comprehension of patients' healthcare journeys and enhancing the overall effectiveness of extracting temporal relationships in complex medical narratives.
Timotej Knez, Slavko Zitnik
J. Am. Medical Informatics Assoc.2
2023 Word in context task for the Slovene language
Timotej Knez, Slavko Zitnik
LDK2
2023 Online-Notes System: Real-Time Speech Recognition and Translation of Lectures
Tjasa Jelovsek, Marko Bajec, Iztok Lebar Bajec, Kaja Gantar, Slavko Zitnik
RCIS5
2023 Temporal Relation Extraction from Clinical Texts Using Knowledge Graphs
Timotej Knez, Slavko Zitnik
RCIS2
2022 ANGLEr: A Next-Generation Natural Language Exploratory Framework
Timotej Knez, Marko Bajec, Slavko Zitnik
RCIS3
2022 Target-level sentiment analysis for news articles
Slavko Zitnik, Neli Blagus, Marko Bajec
Knowl. Based Syst.1
2020 Punctuation Restoration System for Slovene Language
Marko Bajec, Marko Jankovic, Slavko Zitnik, Iztok Lebar Bajec
RCIS3
2018 Social media comparison and analysis: The best data source for research?
abstract
In the past decade, social media has become an important part of our everyday life. The employment of different social media changes the way we communicate, collaborate, gather information and consequently perceive the world around us. Thus, researchers from different fields exploit the social media to provide deeper insight into human behaviour. Each social media possesses its own privacy politics and access to publicly available data. In this paper, we present a generic framework along with the tools to analyse different social media. The analysis shows basic usage statistics, reach and engagement differences, language, sentiment and gender identification of each social network data. We compare data from Twitter, Facebook, Tumblr, Google+ and YouTube. The results reveal specifics of each social media, which to some extent also depend on the data available and the selected seed keywords. We uncover that popularity of selected topics in social media is proportional to the number of hits on Google, celebrities and politicians are the most talked topics and that behaviour of users across social media is different. For example, Twitter users prefer to post more, while Facebook and Youtube users prefer to comment. The majority of all social media posts are in English, larger number of them are negative and often written by male users. The results of the proposed framework should serve as a tool to identify the appropriate source of data for the representative analysis of social media.
Neli Blagus, Slavko Zitnik
RCIS2
2018 Process models of interrelated speech intentions from online health-related conversations
Elena V. Epure, Dario Compagno, Camille Salinesi, Rébecca Deneckère, Marko Bajec, Slavko Zitnik
Artif. Intell. Medicine6
2017 LogMap+: Relational data enrichment and linked data resources matching
abstract
Relational database to ontology mapping and ontology matching techniques are mostly addressed separately, even though it is known that the real power of semantic data lies in data interconnection. The latter is especially important when designing a new ontology, which often includes at least some of the concepts that already exist in the linked open data cloud. Thus, in this paper we describe a new end-to-end tool LogMap+ for transformation of relational data into an ontology and matching it against a pre-existent semantic source. Apart from offering the efficient web-based application, the main contributions are the improvements of the domain specific LogMap system. We evaluate our general tool against OAEI 2014 challenge datasets and achieve comparable results to the top performing algorithms and also outperform the domain specific LogMap tool.
Slavko Zitnik, Marko Bajec, Dejan Lavbic
RCIS1
2015 Iterative joint extraction of entities, relationships and coreferences from text sources
abstract
Machine understanding of textual documents has been challenging since the early computer era. Since the information extraction research field emerged it has inferred multiple natural language processing tasks, such as named entities recognition, relationships extraction and coreference resolution. Even though for the purpose of the end-to-end information extraction all of the three tasks are crucial, existing work has been focusing merely on one specific task at the time or at best on their connection in a pipeline. In this paper we introduce a novel iterative and joint information extraction system that interconnects all the three tasks together using iterative feature functions which use the advantage of the intermediate extractions. Furthermore, we introduce a special transformation of data into skip-mention sequences to enable the extraction of relations and coreferences using fast first-order graphical models. Additionally, the system uses an ontology as its knowledge source, as a list of inferred extraction rules, and as a data schema of extracted results. Experimental results show that the accuracy of extractions improves after each iteration. In particular, our model obtained a 15% error reduction on named entity recognition over individual models.
Slavko Zitnik, Marko Bajec
RCIS1
2015 Sieve-based relation extraction of gene regulatory networks from biological literature
abstract
BACKGROUND: Relation extraction is an essential procedure in literature mining. It focuses on extracting semantic relations between parts of text, called mentions. Biomedical literature includes an enormous amount of textual descriptions of biological entities, their interactions and results of related experiments. To extract them in an explicit, computer readable format, these relations were at first extracted manually from databases. Manual curation was later replaced with automatic or semi-automatic tools with natural language processing capabilities. The current challenge is the development of information extraction procedures that can directly infer more complex relational structures, such as gene regulatory networks. RESULTS: We develop a computational approach for extraction of gene regulatory networks from textual data. Our method is designed as a sieve-based system and uses linear-chain conditional random fields and rules for relation extraction. With this method we successfully extracted the sporulation gene regulation network in the bacterium Bacillus subtilis for the information extraction challenge at the BioNLP 2013 conference. To enable extraction of distant relations using first-order models, we transform the data into skip-mention sequences. We infer multiple models, each of which is able to extract different relationship types. Following the shared task, we conducted additional analysis using different system settings that resulted in reducing the reconstruction error of bacterial sporulation network from 0.73 to 0.68, measured as the slot error rate between the predicted and the reference network. We observe that all relation extraction sieves contribute to the predictive performance of the proposed approach. Also, features constructed by considering mention words and their prefixes and suffixes are the most important features for higher accuracy of extraction. Analysis of distances between different mention types in the text shows that our choice of transforming data into skip-mention sequences is appropriate for detecting relations between distant mentions. CONCLUSIONS: Linear-chain conditional random fields, along with appropriate data transformations, can be efficiently used to extract relations. The sieve-based architecture simplifies the system as new sieves can be easily added or removed and each sieve can utilize the results of previous ones. Furthermore, sieves with conditional random fields can be trained on arbitrary text data and hence are applicable to broad range of relation extraction tasks and data domains.
Slavko Zitnik, Marinka Zitnik, Blaz Zupan, Marko Bajec
BMC Bioinform.1