VLDB 2026 Research / reviewers in the wild / expert
Ranka Stankovic
dblp:49/8114
· DBLP profile ↗
22ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-5123-6273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Development of Serbian QA Datasets through Prompt-Based Generation and Human ValidationabstractLLMs capable of answering questions, fulfilling diverse user requests, and functioning as chatbots rely heavily on extensive datasets. However, for the Serbian language, there is a significant lack of high-quality datasets structured in a question-and-answer (QA) format. To address this, we extracted a portion of the SQuAD-sr dataset, which, to the best of our knowledge, is the largest QA dataset in Serbian and contains over 87k samples. While this dataset is an incredibly valuable resource, it was translated using an adapted Translate-Align-Retrieve method and contains errors and terminological inaccuracies. In this work, we systematically reviewed and corrected more than 7k samples from the SQuAD-sr dataset, significantly improving the dataset’s reliability and quality. We call this modified subset of the SQuAD-sr dataset, the SQuAD-sr-md dataset. The corrections that were made are crucial for training accurate and robust QA models in Serbian, ensuring that AI systems can leverage the full potential of this dataset. We also introduce an additional QA dataset generated from encyclopedia articles, Wikipedia pages, and scientific paper abstracts using LLMs, which contains 74k samples. We name this dataset the SerbianQA-Gen. Jovana Radenovic, Olivera Kitanovic, Ranka Stankovic, Mihailo Skoric |
LREC | 3 |
| 2026 | PARSEME 2.0 Multilingual Corpus of Multiword ExpressionsabstractInternational audience Agata Savary, Manon Scholivet, Carlos Ramisch, Takuya Nakamura, Eric Bilinski, Sara Stymne, Voula Giouli, Stella Markantonatou, Vasile Florian Pais, Maria Mitrofan, Louis Estève, Bruno Guillaume, Verginica Barbu Mititelu, Jaka Cibej, Roberto Díaz Hernández, Victoria Fendel, Polona Gantar, Olha Kanishcheva, Cvetana Krstev, Chaya Liebeskind, Irina Lobzhanidze, Aleksandra M. Markovic, Gunta Nespore-Berzkalne, Adriana S. Pagano, Mehrnoush Shamsfard, Ranka Stankovic, Vahideh Tajalli, Carole Tiberius, Aakanksha Padhye |
LREC | 26 |
| 2026 | Integrating TEI, NER/NEL, Textometry, and Linked Data for a Semantically Enriched Interview CorpusabstractThis paper presents a pipeline that converts unstructured interview transcripts into a semantically enriched, queryable knowledge resource. The texts from the Digitalne Ikone 20+ interview collection were first encoded in TEI XML (Text Encoding Initiative), marking interview boundaries, paragraph breaks, speaker turns with identifiers, dates, and topics. This structural encoding underpins downstream NLP and enables structured querying (e.g., by speaker). We then applied Named Entity Recognition to identify persons, places, organizations, and events, and embedded the results directly in TEI. In the third stage, Named Entity Linking mapped entity mentions to canonical Wikidata identifiers via context-aware disambiguation; missing entries were added to Wikidata when necessary. The resulting TEI+NER/NEL corpus, serialized as linked data, follows the NIF (NLP Interchange Framework). The pipeline also supports retrieval-augmented summarization that retrieves evidence passages and prompts LLMs (implemented with DSPy) to produce faithful interview summaries. We discuss design choices (TXM for textometry with JeRTeh resources; TESLA models for NER/NEL), report qualitative gains in interpretability through semantic links, and outline future work on domain-adapted NER/NEL, graph-based completion, and more expressive RAG architectures. The approach is replicable for other oral-history or media corpora and advances practical, evidence-grounded access to cultural archives and beyond. Ranka Stankovic, Tamara Vucenovic, Biljana Rujevic, Milica Ikonic Nesic, Mihailo Skoric |
LREC | 1 |
| 2024 | Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrACabstractOntoLex, the dominant community standard for machine-readable lexical resources in the context of RDF, Linked Data and Semantic Web technologies, is currently extended with a designated module for Frequency, Attestations and Corpus-based Information (OntoLex-FrAC). We propose a novel component for OntoLex-FrAC, addressing the incorporation of corpus queries for (a) linking dictionaries with corpus engines, (b) enabling RDF-based web services to exchange corpus queries and responses data dynamically, and (c) using conventional query languages to formalize the internal structure of collocations, word sketches, and colligations. The primary field of application of the query extension is in digital lexicography and corpus linguistics, and we present a proof-of-principle implementation in backend components of a novel platform designed to support digital lexicography for the Serbian language. Christian Chiarcos, Ranka Stankovic, Maxim Ionov, Gilles Sérasset |
LREC/COLING | 2 |
| 2024 | MultiLexBATS: Multilingual Dataset of Lexical Semantic RelationsabstractUnderstanding the relation between the meanings of words is an important part of comprehending natural language. Prior work has either focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs), with some exceptions. Given the rarity of highly multilingual benchmarks, it is unclear to what extent PLMs capture relational knowledge and are able to transfer it across languages. To start addressing this question, we propose MultiLexBATS, a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages, such as Bambara, Lithuanian, and Albanian. As experiment on cross-lingual transfer of relational knowledge, we test the PLMs’ ability to (1) capture analogies across languages, and (2) predict translation targets. We find considerable differences across relation types and languages with a clear preference for hypernymy and antonymy as well as romance languages. Dagmar Gromann, Hugo Gonçalo Oliveira, Lucia Pitarch, Elena Apostol, Jordi Bernad, Eliot Bytyci, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabík, Jorge Gracia, Letizia Granata, Anas Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia di Buono, Ana Ostroski Anic, Sigita Rackeviciene, Ricardo Rodrigues 0001, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stankovic, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova |
LREC/COLING | 27 |
| 2024 | Topic Modeling of the SrpELTeC Corpus: A Comparison of NMF, LDA, and BERTopicabstractTopic modeling is an effective way to gain insight into large amounts of data.Some of the most widely used topic models are Latent Dirichlet allocation (LDA) and Nonnegative Matrix Factorization (NMF).However, new ways to mine topics have emerged with the rise of self-attention models and pretrained language models.BERTopic represents the current stateof-the-art when it comes to modeling topics.In this paper, we compared LDA, NMF, and BERTopic performance on literary texts in the Serbian language, both quantitatively by measuring Topic Coherency (TC) and Topic Diversity (TD), and by conducting a qualitative evaluation of the obtained topics.Additionally, for BERTopic, we compared multilingual sentence transformer embeddings with the Jerteh-355 monolingual embeddings for Serbian.NMF yielded the best Topic Coherency results, while BERTopic with Jerteh-355 embeddings gave the best Topic Diveristy.The monolingual Serbian Jerteh-355 embeddings also outperformed sentence transformer embeddings in both TC and TD. Teodora Mihajlov, Milica Ikonic Nesic, Ranka Stankovic, Olivera Kitanovic |
FedCSIS | 3 |
| 2024 | SrpCNNeL: Serbian Model for Named Entity LinkingabstractThis paper presents the development of a Named Entity Linking (NEL) model to the Wikidata knowledge base for the Serbian language, named SrpCNNeL.The model was trained to recognize and link seven different named entity types (persons, locations, organizations, professions, events, demonyms, and works of art) on a dataset containing sentences from novels, legal documents, as well as sentences generated from the Wikidata knowledge base and the Leximirka lexical database.The resulting model demonstrated robust performance, achieving an F1 score of 0.8 on the test set.Considering that the dataset contains the highest number of locations linked to the knowledge base, an evaluation was conducted on an independent dataset and compared to the baseline Spacy Entity Linker for locations only. Milica Ikonic Nesic, Sasa Petalinkar, Ranka Stankovic, Milos Utvic, Olivera Kitanovic |
FedCSIS | 3 |
| 2023 | Football terminology: compilation and transformation into OntoLex-Lemon resource
Jelena Lazarevic, Ranka Stankovic, Mihailo Skoric, Biljana Rujevic |
LDK | 2 |
| 2023 | Towards ELTeC-LLOD: European Literary Text Collection Linguistic Linked Open Data
Ranka Stankovic, Christian Chiarcos, Milos Utvic, Olivera Kitanovic |
LDK | 1 |
| 2022 | Distant Reading in Digital Humanities: Case Study on the Serbian Part of the ELTeC CollectionabstractIn this paper we present the Serbian part of the ELTeC multilingual corpus of novels written in the time period 1840-1920. The corpus is being built in order to test various distant reading methods and tools with the aim of re-thinking the European literary history. We present the various steps that led to the production of the Serbian sub-collection: the novel selection and retrieval, text preparation, structural annotation, POS-tagging, lemmatization and named entity recognition. The Serbian sub-collection was published on different platforms in order to make it freely available to various users. Several use examples show that this sub-collection is usefull for both close and distant reading approaches. Ranka Stankovic, Cvetana Krstev, Branislava Sandrih, Dusko Vitas, Mihailo Skoric, Milica Ikonic Nesic |
LREC | 1 |
| 2021 | A Twitter Corpus and Lexicon for Abusive Speech Detection in SerbianabstractAbusive speech in social media, including profanities, derogatory and hate speech, has reached the level of a pandemic. A system that would be able to detect such texts could help in making the Internet and social media a better and more respectful virtual space. Research and commercial application in this area were so far focused mainly on the English language. This paper presents the work on building AbCoSER, the first corpus of abusive speech in Serbian. The corpus consists of 6,436 manually annotated tweets, out of which 1,416 were labelled as tweets using some kind of abusive speech. Those 1,416 tweets were further sub-classified, for instance to those using vulgar, hate speech, derogatory language, etc. In this paper, we explain the process of data acquisition, annotation, and corpus construction. We also discuss the results of an initial analysis of the annotation quality. Finally, we present an abusive speech lexicon structure and its enrichment with abusive triggers extracted from the AbCoSER dataset. Danka Jokic, Ranka Stankovic, Cvetana Krstev, Branislava Sandrih |
LDK | 2 |
| 2020 | A Multilingual Evaluation Dataset for Monolingual Word Sense AlignmentabstractAligning senses across resources and languages is a challenging task with beneficial applications in the field of natural language processing and electronic lexicography. In this paper, we describe our efforts in manually aligning monolingual dictionaries. The alignment is carried out at sense-level for various resources in 15 languages. Moreover, senses are annotated with possible semantic relationships such as broadness, narrowness, relatedness, and equivalence. In comparison to previous datasets for this task, this dataset covers a wide range of languages and resources and focuses on the more challenging task of linking general-purpose language. We believe that our data will pave the way for further advances in alignment and evaluation of word senses by creating new solutions, particularly those notoriously requiring data such as neural networks. Our resources are publicly available at https://github.com/elexis-eu/MWSA. Sina Ahmadi, John P. McCrae, Sanni Nimb, Anas Fahad Khan, Monica Monachini, Bolette S. Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José-Luis Sancho-Gómez, Rafael-J. Ureña-Ruiz, Jordi Porta-Zamorano, Kiril Ivanov Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stankovic, Andrej Perdih, Dejan Gabrovsek |
LREC | 40 |
| 2020 | Machine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for SerbianabstractThe training of new tagger models for Serbian is primarily motivated by the enhancement of the existing tagset with the grammatical category of a gender. The harmonization of resources that were manually annotated within different projects over a long period of time was an important task, enabled by the development of tools that support partial automation. The supporting tools take into account different taggers and tagsets. This paper focuses on TreeTagger and spaCy taggers, and the annotation schema alignment between Serbian morphological dictionaries, MULTEXT-East and Universal Part-of-Speech tagset. The trained models will be used to publish the new version of the Corpus of Contemporary Serbian as well as the Serbian literary corpus. The performance of developed taggers were compared and the impact of training set size was investigated, which resulted in around 98% PoS-tagging precision per token for both new models. The sr_basic annotated dataset will also be published. Ranka Stankovic, Branislava Sandrih, Cvetana Krstev, Milos Utvic, Mihailo Skoric |
LREC | 1 |
| 2020 | Two approaches to compilation of bilingual multi-word terminology lists from lexical resourcesabstractAbstract In this paper, we present two approaches and the implemented system for bilingual terminology extraction that rely on an aligned bilingual domain corpus, a terminology extractor for a target language, and a tool for chunk alignment. The two approaches differ in the way terminology for the source language is obtained: the first relies on an existing domain terminology lexicon, while the second one uses a term extraction tool. For both approaches, four experiments were performed with two parameters being varied. In the experiments presented in this paper, the source language was English, and the target language Serbian, and a selected domain was Library and Information Science, for which an aligned corpus exists, as well as a bilingual terminological dictionary. For term extraction, we used the FlexiTerm tool for the source language and a shallow parser for the target language, while for word alignment we used GIZA++. The evaluation results show that for the first approach the F1 score varies from 29.43% to 51.15%, while for the second it varies from 61.03% to 71.03%. On the basis of the evaluation results, we developed a binary classifier that decides whether a candidate pair, composed of aligned source and target terms, is valid. We trained and evaluated different classifiers on a list of manually labeled candidate pairs obtained after the implementation of our extraction system. The best results in a fivefold cross-validation setting were achieved with the Radial Basis Function Support Vector Machine classifier, giving a F1 score of 82.09% and accuracy of 78.49%. Branislava Sandrih, Cvetana Krstev, Ranka Stankovic |
Nat. Lang. Eng. | 3 |
| 2018 | Using English Baits to Catch Serbian Multi-Word Terminology
Cvetana Krstev, Branislava Sandrih, Ranka Stankovic, Miljana Mladenovic |
LREC | 3 |
| 2016 | Rule-based Automatic Multi-word Term Extraction and Lemmatization
Ranka Stankovic, Cvetana Krstev, Ivan Obradovic, Biljana Lazic, Aleksandra Trtovac |
LREC | 1 |
| 2012 | A tool for enhanced search of multilingual digital libraries of e-journals
Ranka Stankovic, Cvetana Krstev, Ivan Obradovic, Aleksandra Trtovac, Milos Utvic |
LREC | 1 |
| 2010 | A Description of Morphological Features of Serbian: a Revision using Feature System Declaration
Cvetana Krstev, Ranka Stankovic, Dusko Vitas |
LREC | 2 |
| 2010 | GIS Application Improvement with Multilingual Lexical and Terminological Resources
Ranka Stankovic, Ivan Obradovic, Olivera Kitanovic |
LREC | 1 |
| 2008 | The Usage of Various Lexical Resources and Tools to Improve the Performance of Web Search Engines
Cvetana Krstev, Ranka Stankovic, Dusko Vitas, Ivan Obradovic |
LREC | 2 |
| 2006 | WS4LR: A Workstation for Lexical Resources
Cvetana Krstev, Ranka Stankovic, Dusko Vitas, Ivan Obradovic |
LREC | 2 |
| 2004 | Combining Heterogeneous Lexical Resources
Cvetana Krstev, Dusko Vitas, Ranka Stankovic, Ivan Obradovic, Gordana Pavlovic-Lazetic |
LREC | 3 |