EDBT 2026 Demo / reviewers in the wild / expert
Nikola Ljubesic
dblp:40/8154
· DBLP profile ↗
41ranked-venue papers
14as first author
18since 2021 · last 2026
0000-0001-7169-9152ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 14 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and SerbianabstractParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic fashion from the ParlaMint transcripts and their corresponding metadata, which were aligned to the speech recordings of each corresponding parliament. In this release of the dataset, each of the corpora is significantly enriched with various automatic annotation layers. The textual modality of all four corpora has been enriched with linguistic annotations and sentiment predictions. Similar to that, their spoken modality has been automatically enriched with occurrences of filled pauses, the most frequent disfluency in typical speech. Two out of the four languages have been additionally enriched with detailed word- and grapheme-level alignments, and the automatic annotation of the position of primary stress in multisyllabic words. With these enrichments, the usefulness of the underlying corpora has been drastically increased for downstream research across multiple disciplines, which we showcase through an analysis of acoustic correlates of sentiment. All the corpora are made available for download in JSONL and TextGrid formats, as well as for search through a concordancer. Nikola Ljubesic, Peter Rupnik, Ivan Porupski, Taja Kuzman Pungersek |
LREC | 1 |
| 2026 | The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
Taja Kuzman Pungersek, Peter Rupnik, Vit Suchomel, Nikola Ljubesic |
LREC | 4 |
| 2026 | ROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian
Kaja Dobrovoljc, Darinka Verdonik, Jaka Cibej, Peter Rupnik, Nikola Ljubesic |
LREC | 5 |
| 2025 | Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 12 |
| 2025 | Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
Nikola Ljubesic, Ivan Porupski, Peter Rupnik |
INTERSPEECH | 1 |
| 2024 | A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper CollectionabstractPreparing historical newspaper collections is a complicated endeavour, consisting of multiple steps that have to be carefully adapted to the specific content in question, including imaging, layout prediction, optical character recognition, and linguistic annotation. To address the high costs associated with the process, we present a lightweight approach to producing high-quality corpora and apply it to a massive collection of Slovenian historical newspapers from the 18th, 19th and 20th century resulting in a billion-word giga-corpus. We start with noisy OCR-ed data produced by different technologies in varying periods by the National and University Library of Slovenia. To address the inherent variability in the quality of textual data, a challenge commonly encountered in digital libraries globally, we perform a targeted post-digitisation correction procedure, coupled with a robust curation mechanism for noisy texts via language model inference. Subsequently, we subject the corrected and filtered output to comprehensive linguistic annotation, enriching the corpus with part-of-speech tags, lemmas, and named entity labels. Finally, we perform an analysis through topic modeling at the noun lemma level, along with a frequency analysis of the named entities, to confirm the viability of our corpus preparation method. Filip Dobranic, Bojan Evkoski, Nikola Ljubesic |
LREC/COLING | 3 |
| 2024 | CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre AnnotationabstractThis paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language space. The collection of these corpora comprises a total of 13 billion tokens of texts from 26 million documents. The comparability of the corpora is ensured by a comparable crawling setup and the usage of identical crawling and post-processing technology. All the corpora were linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline, and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier, which further enhances comparability at the level of linguistic annotation and metadata enrichment. The genre-focused analysis of the resulting corpora shows a rather consistent distribution of genres throughout the seven corpora, with variations in the most prominent genre categories being well-explained by the economic strength of each language community. A comparison of the distribution of genre categories across the corpora indicates that web corpora from less developed countries primarily consist of news articles. Conversely, web corpora from economically more developed countries exhibit a smaller proportion of news content, with a greater presence of promotional and opinionated texts. Nikola Ljubesic, Taja Kuzman |
LREC/COLING | 1 |
| 2024 | The ParlaSent Multilingual Training Dataset for Sentiment Identification in Parliamentary ProceedingsabstractThe paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. The paper additionally introduces the first domain-specific multilingual transformer language model for political science applications, which was additionally pre-trained on 1.72 billion words from parliamentary proceedings of 27 European parliaments. We present experiments demonstrating how the additional pre-training on parliamentary data can significantly improve the model downstream performance, in our case, sentiment identification in parliamentary proceedings. We further show that our multilingual model performs very well on languages not seen during fine-tuning, and that additional fine-tuning data from other languages significantly improves the target parliament’s results. The paper makes an important contribution to multiple disciplines inside the social sciences, and bridges them with computer science and computational linguistics. Lastly, the resulting fine-tuned language model sets up a more robust approach to sentiment analysis of political texts across languages, which allows scholars to study political sentiment from a comparative perspective using standardized tools and techniques. Michal Mochtak, Peter Rupnik, Nikola Ljubesic |
LREC/COLING | 3 |
| 2024 | Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 LanguagesabstractLarge, curated, web-crawled corpora play a vital role in training language models (LMs). They form the lion’s share of the training data in virtually all recent LMs, such as the well-known GPT, LLaMA and XLM-RoBERTa models. However, despite this importance, relatively little attention has been given to the quality of these corpora. In this paper, we compare four of the currently most relevant large, web-crawled corpora (CC100, MaCoCu, mC4 and OSCAR) across eleven lower-resourced European languages. Our approach is two-fold: first, we perform an intrinsic evaluation by performing a human evaluation of the quality of samples taken from different corpora; then, we assess the practical impact of the qualitative differences by training specific LMs on each of the corpora and evaluating their performance on downstream tasks. We find that there are clear differences in quality of the corpora, with MaCoCu and OSCAR obtaining the best results. However, during the extrinsic evaluation, we actually find that the CC100 corpus achieves the highest scores. We conclude that, in our experiments, the quality of the web-crawled corpora does not seem to play a significant role when training LMs. Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubesic, Miquel Esplà-Gomis, Gema Ramírez-Sánchez, Antonio Toral |
LREC/COLING | 4 |
| 2024 | Gos 2: A New Reference Corpus of Spoken SlovenianabstractThis paper introduces a new version of the Gos reference corpus of spoken Slovenian, which was recently extended to more than double the original size (300 hours, 2.4 million words) by adding speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public academic speech, and the Artur speech recognition database. We describe this process by first presenting the criteria guiding the balanced selection of the newly added data and the challenges encountered when merging language resources with divergent designs, followed by the presentation of other major enhancements of the new Gos corpus, such as improvements in lemmatization and morphosyntactic annotation, word-level speech alignment, a new XML schema and the development of a specialized online concordancer. Darinka Verdonik, Kaja Dobrovoljc, Tomaz Erjavec, Nikola Ljubesic |
LREC/COLING | 4 |
| 2024 | Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 10 |
| 2024 | Universal NER: A Gold-Standard Multilingual Named Entity Recognition BenchmarkabstractStephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa, Hila Gonen, Joseph Marvin Imperial, Börje F. Karlsson, Peiqin Lin, Nikola Ljubešić, LJ Miranda, Barbara Plank, Arij Riabi, Yuval Pinter. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Stephen Mayhew 0002, Terra Blevins, Shuheng Liu 0002, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson 0001, Peiqin Lin, Nikola Ljubesic, Lester James V. Miranda, Barbara Plank, Arij Riabi, Yuval Pinter |
NAACL-HLT | 9 |
| 2024 | Can cross-domain term extraction benefit from cross-lingual transfer and nested term labeling?abstractAbstract Automatic term extraction (ATE) is a natural language processing task that eases the effort of manually identifying terms from domain-specific corpora by providing a list of candidate terms. In this paper, we treat ATE as a sequence-labeling task and explore the efficacy of XLMR in evaluating cross-lingual and multilingual learning against monolingual learning in the cross-domain ATE context. Additionally, we introduce NOBI, a novel annotation mechanism enabling the labeling of single-word nested terms. Our experiments are conducted on the ACTER corpus, encompassing four domains and three languages (English, French, and Dutch), as well as the RSDO5 Slovenian corpus, encompassing four additional domains. Results indicate that cross-lingual and multilingual models outperform monolingual settings, showcasing improved F1-scores for all languages within the ACTER dataset. When incorporating an additional Slovenian corpus into the training set, the multilingual model exhibits superior performance compared to state-of-the-art approaches in specific scenarios. Moreover, the newly introduced NOBI labeling mechanism enhances the classifier’s capacity to extract short nested terms significantly, leading to substantial improvements in Recall for the ACTER dataset and consequentially boosting the overall F1-score performance. Tran Thi Hong Hanh, Matej Martinc, Andraz Repar, Nikola Ljubesic, Antoine Doucet, Senja Pollak |
Mach. Learn. | 4 |
| 2024 | Geographic Adaptation of Pretrained Language ModelsabstractAbstract While pretrained language models (PLMs) have been shown to possess a plethora of linguistic knowledge, the existing body of research has largely neglected extralinguistic knowledge, which is generally difficult to obtain by pretraining on text alone. Here, we contribute to closing this gap by examining geolinguistic knowledge, i.e., knowledge about geographic variation in language. We introduce geoadaptation, an intermediate training step that couples language modeling with geolocation prediction in a multi-task learning setup. We geoadapt four PLMs, covering language groups from three geographic areas, and evaluate them on five different tasks: fine-tuned (i.e., supervised) geolocation prediction, zero-shot (i.e., unsupervised) geolocation prediction, fine-tuned language identification, zero-shot language identification, and zero-shot prediction of dialect features. Geoadaptation is very successful at injecting geolinguistic knowledge into the PLMs: The geoadapted PLMs consistently outperform PLMs adapted using only language modeling (by especially wide margins on zero-shot prediction tasks), and we obtain new state-of-the-art results on two benchmarks for geolocation prediction and language identification. Furthermore, we show that the effectiveness of geoadaptation stems from its ability to geographically retrofit the representation space of the PLMs. Valentin Hofmann, Goran Glavas, Nikola Ljubesic, Janet B. Pierrehumbert, Hinrich Schütze |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languagesabstractWe present the most relevant results of the project MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages in its second year. To date, parallel and monolingual corpora have been produced for seven low-resourced European languages by crawling large amounts of textual data from selected top-level domains of the Internet; both human and automatic evaluation show its usefulness. In addition, several large language models pretrained on MaCoCu data have been published, as well as the code used to collect and curate the data. Marta Bañón, Malina Chichirau, Miquel Esplà-Gomis, Mikel L. Forcada, Aarón Galiano Jiménez, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Jaume Zaragoza-Bernabeu |
EAMT | 7 |
| 2023 | Quantifying the impact of context on the quality of manual hate speech annotationabstractAbstract The quality of annotations in manually annotated hate speech datasets is crucial for automatic hate speech detection. This contribution focuses on the positive effects of manually annotating online comments for hate speech within the context in which the comments occur. We quantify the impact of context availability by meticulously designing an experiment: Two annotation rounds are performed, one in-context and one out-of-context, on the same English YouTube data (more than 10,000 comments), by using the same annotation schema and platform, the same highly trained annotators, and quantifying annotation quality through inter-annotator agreement. Our results show that the presence of context has a significant positive impact on the quality of the manual annotations. This positive impact is more noticeable among replies than among comments, although the former is harder to consistently annotate overall. Previous research reporting that out-of-context annotations favour assigning non-hate-speech labels is also corroborated, showing further that this tendency is especially present among comments inciting violence, a highly relevant category for hate speech research and society overall. We believe that this work will improve future annotation campaigns even beyond hate speech and motivate further research on the highly relevant questions of data annotation methodology in natural language processing, especially in the light of the current expansion of its scope of application. Nikola Ljubesic, Igor Mozetic, Petra Kralj Novak |
Nat. Lang. Eng. | 1 |
| 2022 | MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languagesabstractWe introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual and parallel corpora for under-resourced European languages. The approach followed consists of crawling large amounts of textual data from carefully selected top-level domains of the Internet, and then applying a curation and enrichment pipeline. In addition to corpora, the project will release successive versions of the free/open-source web crawling and curation software used. Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero, Taja Kuzman, Nikola Ljubesic, Rik van Noord, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Peter Rupnik, Vit Suchomel, Antonio Toral, Tobias van der Werff, Jaume Zaragoza |
EAMT | 6 |
| 2022 | The GINCO Training Dataset for Web Genre Identification of Documents Out in the WildabstractThis paper presents a new training dataset for automatic genre identification GINCO, which is based on 1,125 crawled Slovenian web documents that consist of 650,000 words. Each document was manually annotated for genre with a new annotation schema that builds upon existing schemata, having primarily clarity of labels and inter-annotator agreement in mind. The dataset consists of various challenges related to web-based data, such as machine translated content, encoding errors, multiple contents presented in one document etc., enabling evaluation of classifiers in realistic conditions. The initial machine learning experiments on the dataset show that (1) pre-Transformer models are drastically less able to model the phenomena, with macro F1 metrics ranging around 0.22, while Transformer-based models achieve scores of around 0.58, and (2) multilingual Transformer models work as well on the task as the monolingual models that were previously proven to be superior to multilingual models on standard NLP tasks. Taja Kuzman, Peter Rupnik, Nikola Ljubesic |
LREC | 3 |
| 2020 | CoSimLex: A Resource for Evaluating Graded Word Similarity in ContextabstractState of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far. Carlos Santos Armendariz, Matthew Purver, Matej Ulcar, Senja Pollak, Nikola Ljubesic, Mark Granroth-Wilding |
LREC | 5 |
| 2020 | Gigafida 2.0: The Reference Corpus of Written Standard SloveneabstractWe describe a new version of the Gigafida reference corpus of Slovene. In addition to updating the corpus with new material and annotating it with better tools, the focus of the upgrade was also on its transformation from a general reference corpus, which contains all language variants including non-standard language, to the corpus of standard (written) Slovene. This decision could be implemented as new corpora dedicated specifically to non-standard language emerged recently. In the new version, the whole Gigafida corpus was deduplicated for the first time, which facilitates automatic extraction of data for the purposes of compilation of new lexicographic resources such as the collocations dictionary and the thesaurus of Slovene. Simon Krek, Spela Arhar Holdt, Tomaz Erjavec, Jaka Cibej, Andraz Repar, Polona Gantar, Nikola Ljubesic, Iztok Kosem, Kaja Dobrovoljc |
LREC | 7 |
| 2019 | How to tag non-standard language: Normalisation versus domain adaptation for Slovene historical and user-generated textsabstractAbstract Part-of-speech (PoS) tagging of non-standard language with models developed for standard language is known to suffer from a significant decrease in accuracy. Two methods are typically used to improve it: word normalisation, which decreases the out-of-vocabulary rate of the PoS tagger, and domain adaptation where the tagger is made aware of the non-standard language variation, either through supervision via non-standard data being added to the tagger’s training set, or via distributional information calculated from raw texts. This paper investigates the two approaches, normalisation and domain adaptation, on carefully constructed data sets encompassing historical and user-generated Slovene texts, in particular focusing on the amount of labour necessary to produce the manually annotated data sets for each approach and comparing the resulting PoS accuracy. We give quantitative as well as qualitative analyses of the tagger performance in various settings, showing that on our data set closed and open class words exhibit significantly different behaviours, and that even small inconsistencies in the PoS tags in the data have an impact on the accuracy. We also show that to improve tagging accuracy, it is best to concentrate on obtaining manually annotated normalisation training data for short annotation campaigns, while manually producing in-domain training sets for PoS tagging is better when a more substantial annotation campaign can be undertaken. Finally, unsupervised adaptation via Brown clustering is similarly useful regardless of the size of the training data available, but improvements tend to be bigger when adaptation is performed via in-domain tagging data. Katja Zupan, Nikola Ljubesic, Tomaz Erjavec |
Nat. Lang. Eng. | 2 |
| 2016 | TweetGeo - A Tool for Collecting, Processing and Analysing Geo-encoded Linguistic DataabstractIn this paper we present a newly developed tool that enables researchers interested in spatial variation of language to define a geographic perimeter of interest, collect data from the Twitter streaming API published in that perimeter, filter the obtained data by language and country, define and extract variables of interest and analyse the extracted variables by one spatial statistic and two spatial visualisations. We showcase the tool on the area and a selection of languages spoken in former Yugoslavia. By defining the perimeter, languages and a series of linguistic variables of interest we demonstrate the data collection, processing and analysis capabilities of the tool. Nikola Ljubesic, Tanja Samardzic, Curdin Derungs |
COLING | 1 |
| 2016 | Collaborative Development of a Rule-Based Machine Translator between Croatian and Serbian
Filip Klubicka, Gema Ramírez-Sánchez, Nikola Ljubesic |
EAMT | 3 |
| 2016 | Dealing with Data Sparseness in SMT with Factured Models and Morphological Expansion: a Case Study on Croatian
Víctor M. Sánchez-Cartagena, Nikola Ljubesic, Filip Klubicka |
EAMT | 2 |
| 2016 | Corpus vs. Lexicon Supervision in Morphosyntactic Tagging: the Case of Slovene
Nikola Ljubesic, Tomaz Erjavec |
LREC | 1 |
| 2016 | Corpus-Based Diacritic Restoration for South Slavic Languages
Nikola Ljubesic, Tomaz Erjavec, Darja Fiser |
LREC | 1 |
| 2016 | Producing Monolingual and Parallel Web Corpora at the Same Time - SpiderLing and Bitextor's Love Affair
Nikola Ljubesic, Miquel Esplà-Gomis, Antonio Toral, Sergio Ortiz-Rojas, Filip Klubicka |
LREC | 1 |
| 2016 | New Inflectional Lexicons and Training Corpora for Improved Morphosyntactic Annotation of Croatian and Serbian
Nikola Ljubesic, Filip Klubicka, Zeljko Agic, Ivo-Pavao Jazbec |
LREC | 1 |
| 2016 | Croatian Error-Annotated Corpus of Non-Professional Written Language
Vanja Stefanec, Nikola Ljubesic, Jelena Kuvac Kraljevic |
LREC | 2 |
| 2015 | Abu-MaTran: Automatic building of Machine Translation
Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Raphaël Rubino, Miquel Esplà-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic |
EAMT | 11 |
| 2014 | Standardizing Tweets with Character-Level Machine Translation
Nikola Ljubesic, Tomaz Erjavec, Darja Fiser |
CICLing (2) | 1 |
| 2014 | The SETimes.HR Linguistically Annotated Corpus of Croatian
Zeljko Agic, Nikola Ljubesic |
LREC | 2 |
| 2014 | Comparing two acquisition systems for automatically building an English-Croatian parallel corpus from multilingual websites
Miquel Esplà-Gomis, Filip Klubicka, Nikola Ljubesic, Sergio Ortiz-Rojas, Vassilis Papavassiliou, Prokopis Prokopidis |
LREC | 3 |
| 2014 | TweetCaT: a tool for building Twitter corpora of smaller languages
Nikola Ljubesic, Darja Fiser, Tomaz Erjavec |
LREC | 1 |
| 2014 | caWaC - A web corpus of Catalan and its application to language modeling and machine translation
Nikola Ljubesic, Antonio Toral |
LREC | 1 |
| 2014 | Quality Estimation for Synthetic Parallel Data Generation
Raphaël Rubino, Antonio Toral, Nikola Ljubesic, Gema Ramírez-Sánchez |
LREC | 3 |
| 2012 | Efficient Discrimination Between Closely Related Languages
Jörg Tiedemann, Nikola Ljubesic |
COLING | 2 |
| 2012 | Addressing polysemy in bilingual lexicon extraction from comparable corpora
Darja Fiser, Nikola Ljubesic, Ozren Kubelka |
LREC | 2 |
| 2010 | Towards Sentiment Analysis of Financial Texts in Croatian
Zeljko Agic, Nikola Ljubesic, Marko Tadic |
LREC | 2 |
| 2010 | Building a Gold Standard for Event Detection in Croatian
Nikola Ljubesic, Tomislava Lauc, Damir Boras |
LREC | 1 |
| 2008 | Generating a Morphological Lexicon of Organization Entity Names
Nikola Ljubesic, Tomislava Lauc, Damir Boras |
LREC | 1 |