EDBT 2026 Demo / reviewers in the wild / expert
Senja Pollak
dblp:75/8154
· DBLP profile ↗
42ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0002-4380-0863ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 1 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Social Bias in Slovenia: The EEC-SL Dataset
Jaya Caporusso, Damar Hoogland, Boshko Koloski, Matthew Purver, Senja Pollak, Spela Vintar |
LREC | 5 |
| 2026 | Slovene Morphological and Word Formation Segmentation: A Novel Dataset and Evaluation
Marko Pranjic, Boris Kern, Ines Vorsic, Senja Pollak |
LREC | 4 |
| 2026 | Mono- and cross-lingual evaluation of representation language models on less-resourced languagesabstractThe current dominance of large language models in natural language processing is based on their contextual awareness. For text classification, text representation models, such as ELMo, BERT, and BERT derivatives, are typically fine-tuned for a specific problem. Most existing work focuses on English; in contrast, we present a large-scale multilingual empirical comparison of several monolingual and multilingual ELMo and BERT models using 14 classification tasks in nine languages. The results show, that the choice of best model largely depends on the task and language used, especially in a cross-lingual setting. In monolingual settings, monolingual BERT models tend to perform the best among BERT models. Among ELMo models, the ones trained on large corpora dominate. Cross-lingual knowledge transfer is feasible on most tasks already in a zero-shot setting without losing much performance. Matej Ulcar, Ales Zagar, Carlos Santos Armendariz, Andraz Repar, Senja Pollak, Matthew Purver, Marko Robnik-Sikonja |
Comput. Speech Lang. | 5 |
| 2026 | FuDoBa: Fusing Document and Knowledge Graph Based Representations with Bayesian OptimisationabstractAbstract Building on the success of large language models (LLMs), LLM-based representations have dominated the document representation landscape, achieving strong performance on document embedding benchmarks. However, high-dimensional, computationally expensive LLM embeddings can be too generic or inefficient for domain-specific and resource-scarce applications. To address these limitations, we introduce FuDoBa—a Bayesian optimisation-based representation learning method that integrates LLM embeddings with domain-specific structured knowledge, sourced both locally and from external repositories such as WikiData. This fusion produces low-dimensional, task-relevant representations while reducing training complexity and yielding interpretable early-fusion weights for improved classification performance. We demonstrate the effectiveness of our approach on six datasets across two domains, showing that when paired with robust AutoML-based classifiers, our method performs on par with, or surpasses, proprietary LLM-only embedding baselines, while offering modality-wise interpretability and a smaller dimensional footprint. Boshko Koloski, Senja Pollak, Roberto Navigli, Blaz Skrlj |
Mach. Learn. | 2 |
| 2024 | A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News MediaabstractDehumanisation involves the perception and/or treatment of a social group’s members as less than human. This phenomenon is rarely addressed with computational linguistic techniques. We adapt a recently proposed approach for English, making it easier to transfer to other languages and to evaluate, introducing a new sentiment resource, the use of zero-shot cross-lingual valence and arousal detection, and a new method for statistical significance testing. We then apply it to study attitudes to migration expressed in Slovene newspapers, to examine changes in the Slovene discourse on migration between the 2015-16 migration crisis following the war in Syria and the 2022-23 period following the war in Ukraine. We find that while this discourse became more negative and more intense over time, it is less dehumanising when specifically addressing Ukrainian migrants compared to others. Jaya Caporusso, Damar Hoogland, Mojca Brglez, Boshko Koloski, Matthew Purver, Senja Pollak |
LREC/COLING | 6 |
| 2024 | Denoising Labeled Data for Comment Moderation Using Active LearningabstractNoisily labeled textual data is ample on internet platforms that allow user-created content. Training models, such as offensive language detection models for comment moderation, on such data may prove difficult as the noise in the labels prevents the model to converge. In this work, we propose to use active learning methods for the purposes of denoising training data for model training. The goal is to sample examples the most informative examples with noisy labels with active learning and send them to the oracle for reannotation thus reducing the overall cost of reannotation. In this setting we tested three existing active learning methods, namely DBAL, Variance of Gradients (VoG) and BADGE. The proposed approach to data denoising is tested on the problem of offensive language detection. We observe that active learning can be effectively used for the purposes of data denoising, however care should be taken when choosing the algorithm for this purpose. Andraz Pelicon, Mladen Karan, Ravi Shekhar, Matthew Purver, Senja Pollak |
LREC/COLING | 5 |
| 2024 | LLMSegm: Surface-level Morphological Segmentation Using Large Language ModelabstractMorphological word segmentation splits a given word into its morphemes (roots and affixes), the smallest meaning-bearing units of language. We introduce a novel approach, called LLMSegm, to surface-level morphological segmentation leveraging large language models (LLMs). The proposed approach is applicable in low-data settings as well as for low-resourced languages. We show how to transform the surface-level morphological segmentation task to a binary classification problem and train LLMs to solve it efficiently. For input, we leverage the information from the default LLM subword tokenisation, and a custom morphological segmentation using novel encoding. The evaluation of LLMSegm across seven morphologically diverse languages demonstrates substantial gains in minimally-supervised settings as well as for low-resourced languages, compared to several existing competitive approaches. In terms of F1-scores and accuracy, we achieve improved results compared to the competing methods in six out of seven datasets. Keywords: morphological segmentation, surface-level segmentation, large language models, low-resource settings Marko Pranjic, Marko Robnik-Sikonja, Senja Pollak |
LREC/COLING | 3 |
| 2024 | AutoML-Guided Fusion of Entity and LLM-Based Representations for Document Classification
Boshko Koloski, Senja Pollak, Roberto Navigli, Blaz Skrlj |
DS (1) | 2 |
| 2024 | AHAM: Adapt, Help, Ask, Model Harvesting LLMs for Literature Mining
Boshko Koloski, Nada Lavrac, Bojan Cestnik, Senja Pollak, Blaz Skrlj, Andrej Kastrin |
IDA (1) | 4 |
| 2024 | Can cross-domain term extraction benefit from cross-lingual transfer and nested term labeling?abstractAbstract Automatic term extraction (ATE) is a natural language processing task that eases the effort of manually identifying terms from domain-specific corpora by providing a list of candidate terms. In this paper, we treat ATE as a sequence-labeling task and explore the efficacy of XLMR in evaluating cross-lingual and multilingual learning against monolingual learning in the cross-domain ATE context. Additionally, we introduce NOBI, a novel annotation mechanism enabling the labeling of single-word nested terms. Our experiments are conducted on the ACTER corpus, encompassing four domains and three languages (English, French, and Dutch), as well as the RSDO5 Slovenian corpus, encompassing four additional domains. Results indicate that cross-lingual and multilingual models outperform monolingual settings, showcasing improved F1-scores for all languages within the ACTER dataset. When incorporating an additional Slovenian corpus into the training set, the multilingual model exhibits superior performance compared to state-of-the-art approaches in specific scenarios. Moreover, the newly introduced NOBI labeling mechanism enhances the classifier’s capacity to extract short nested terms significantly, leading to substantial improvements in Recall for the ACTER dataset and consequentially boosting the overall F1-score performance. Tran Thi Hong Hanh, Matej Martinc, Andraz Repar, Nikola Ljubesic, Antoine Doucet, Senja Pollak |
Mach. Learn. | 6 |
| 2022 | Can Cross-Domain Term Extraction Benefit from Cross-lingual Transfer?
Tran Thi Hong Hanh, Matej Martinc, Antoine Doucet, Senja Pollak |
DS | 4 |
| 2022 | Retrieval-Efficiency Trade-Off of Unsupervised Keyword Extraction
Blaz Skrlj, Boshko Koloski, Senja Pollak |
DS | 3 |
| 2022 | EMBEDDIA project: Cross-Lingual Embeddings for Less- Represented Languages in European News MediaabstractEMBEDDIA project developed a range of resources and methods for less-resourced EU languages, focusing on applications for media industry, including keyword extraction, comment moderation and article generation. Senja Pollak, Andraz Pelicon |
EAMT | 1 |
| 2022 | Evaluation of Curriculum Learning Algorithms using Computational Creativity Inspired Metrics
Benjamin Fele, Jan Babic, Senja Pollak, Martin Znidarsic |
ICCC | 3 |
| 2022 | Out of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?abstractKeyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into ones that require training - supervised and ones that do not - unsupervised. In this study, we are interested in settings, where for a language under investigation, no training data is available. More specifically, we explore whether pretrained multilingual language models can be employed for zero-shot cross-lingual keyword extraction on low-resource languages with limited or no available labeled training data and whether they outperform state-of-the-art unsupervised keyword extractors. The comparison is conducted on six news article datasets covering two high-resource languages, English and Russian, and four low-resource languages, Croatian, Estonian, Latvian, and Slovenian. We find that the pretrained models fine-tuned on a multilingual corpus covering languages that do not appear in the test set (i.e. in a zero-shot setting), consistently outscore unsupervised models in all six languages. Boshko Koloski, Senja Pollak, Blaz Skrlj, Matej Martinc |
LREC | 2 |
| 2022 | Embeddings models for Buddhist SanskritabstractThe paper presents novel resources and experiments for Buddhist Sanskrit, broadly defined here including all the varieties of Sanskrit in which Buddhist texts have been transmitted. We release a novel corpus of Buddhist texts, a novel corpus of general Sanskrit and word similarity and word analogy datasets for intrinsic evaluation of Buddhist Sanskrit embeddings models. We compare the performance of word2vec and fastText static embeddings models, with default and optimized parameter settings, as well as contextual models BERT and GPT-2, with different training regimes (including a transfer learning approach using the general Sanskrit corpus) and different embeddings construction regimes (given the encoder layers). The results show that for semantic similarity the fastText embeddings yield the best results, while for word analogy tasks BERT embeddings work the best. We also show that for contextual models the optimal layer combination for embedding construction is task dependant, and that pretraining the contextual embeddings models on a reference corpus of general Sanskrit is beneficial, which is a promising finding for future development of embeddings for less-resourced languages and domains. Ligeia Lugli, Matej Martinc, Andraz Pelicon, Senja Pollak |
LREC | 4 |
| 2022 | Extracting and Analysing Metaphors in Migration Media Discourse: towards a Metaphor Annotation SchemeabstractThe study of metaphors in media discourse is an increasingly researched topic as media are an important shaper of social reality and metaphors are an indicator of how we think about certain issues through references to other things. We present a neural transfer learning method for detecting metaphorical sentences in Slovene and evaluate its performance on a gold standard corpus of metaphors (classification accuracy of 0.725), as well as on a sample of a domain specific corpus of migrations (precision of 0.40 for extracting domain metaphors and 0.74 if evaluated only on a set of migration related sentences). Based on empirical results and findings of our analysis, we propose a novel metaphor annotation scheme containing linguistic level, conceptual level, and stance information. The new scheme can be used for future metaphor annotations of other socially relevant topics. Ana Zwitter Vitez, Mojca Brglez, Marko Robnik-Sikonja, Tadej Skvorc, Andreja Vezovnik, Senja Pollak |
LREC | 6 |
| 2022 | Knowledge graph informed fake news classification via heterogeneous representation ensemblesabstractIncreasing amounts of freely available data both in textual and relational form offers exploration of richer document representations, potentially improving the model performance and robustness. An emerging problem in the modern era is fake news detection—many easily available pieces of information are not necessarily factually correct, and can lead to wrong conclusions or are used for manipulation. In this work we explore how different document representations, ranging from simple symbolic bag-of-words, to contextual, neural language model-based ones can be used for efficient fake news identification. One of the key contributions is a set of novel document representation learning methods based solely on knowledge graphs, i.e., extensive collections of (grounded) subject-predicate-object triplets. We demonstrate that knowledge graph-based representations already achieve competitive performance to conventionally accepted representation learners. Furthermore, when combined with existing, contextual representations, knowledge graph-based document representations can achieve state-of-the-art performance. To our knowledge this is the first larger-scale evaluation of how knowledge graph-based representations can be systematically incorporated into the process of fake news classification. Boshko Koloski, Timen Stepisnik Perdih, Marko Robnik-Sikonja, Senja Pollak, Blaz Skrlj |
Neurocomputing | 4 |
| 2022 | TNT-KID: Transformer-based neural tagger for keyword identificationabstractAbstract With growing amounts of available textual data, development of algorithms capable of automatic analysis, categorization, and summarization of these data has become a necessity. In this research, we present a novel algorithm for keyword identification, that is, an extraction of one or multiword phrases representing key aspects of a given document, called Transformer-Based Neural Tagger for Keyword IDentification (TNT-KID). By adapting the transformer architecture for a specific task at hand and leveraging language model pretraining on a domain-specific corpus, the model is capable of overcoming deficiencies of both supervised and unsupervised state-of-the-art approaches to keyword extraction by offering competitive and robust performance on a variety of different datasets while requiring only a fraction of manually labeled data required by the best-performing systems. This study also offers thorough error analysis with valuable insights into the inner workings of the model and an ablation study measuring the influence of specific components of the keyword identification workflow on the overall performance. Matej Martinc, Blaz Skrlj, Senja Pollak |
Nat. Lang. Eng. | 3 |
| 2021 | Prioritization of COVID-19-Related Literature via Unsupervised Keyphrase Extraction and Document Representation Learning
Blaz Skrlj, Marko Jukic, Nika Erzen, Senja Pollak, Nada Lavrac |
DS | 4 |
| 2021 | Scientific Question Generation: Pattern-Based and Graph-Based RoboCHAIR Methods
Senja Pollak, Vid Podpecan, Janez Kranjc, Borut Lesjak, Nada Lavrac |
ICCC | 1 |
| 2021 | Supervised and Unsupervised Neural Approaches to Text ReadabilityabstractAbstract We present a set of novel neural supervised and unsupervised approaches for determining the readability of documents. In the unsupervised setting, we leverage neural language models, whereas in the supervised setting, three different neural classification architectures are tested. We show that the proposed neural unsupervised approach is robust, transferable across languages, and allows adaptation to a specific readability task and data set. By systematic comparison of several neural architectures on a number of benchmark and new labeled readability data sets in two languages, this study also offers a comprehensive analysis of different neural approaches to readability classification. We expose their strengths and weaknesses, compare their performance to current state-of-the-art classification approaches to readability, which in most cases still rely on extensive feature engineering, and propose possibilities for improvements. Matej Martinc, Senja Pollak, Marko Robnik-Sikonja |
Comput. Linguistics | 2 |
| 2021 | Emotion recognition in low-resource settings: An evaluation of automatic feature selection methods
Fasih Haider, Senja Pollak, Pierre Albert, Saturnino Luz |
Comput. Speech Lang. | 2 |
| 2021 | tax2vec: Constructing Interpretable Features from Taxonomies for Short Text ClassificationabstractThe use of background knowledge is largely unexploited in text classification tasks. This paper explores word taxonomies as means for constructing new semantic features, which may improve the performance and robustness of the learned classifiers. We propose tax2vec, a parallel algorithm for constructing taxonomy-based features, and demonstrate its use on six short text classification problems: prediction of gender, personality type, age, news topics, drug side effects and drug effectiveness. The constructed semantic features, in combination with fast linear classifiers, tested against strong baselines such as hierarchical attention neural networks, achieves comparable classification results on short text documents. The algorithm's performance is also tested in a few-shot learning setting, indicating that the inclusion of semantic features can improve the performance in data-scarce situations. The tax2vec capability to extract corpus-specific semantic keywords is also demonstrated. Finally, we investigate the semantic space of potential features, where we observe a similarity with the well known Zipf's law. Blaz Skrlj, Matej Martinc, Jan Kralj, Nada Lavrac, Senja Pollak |
Comput. Speech Lang. | 5 |
| 2021 | autoBOT: evolving neuro-symbolic representations for explainable low resource text classificationabstractLearning from texts has been widely adopted throughout industry and science. While state-of-the-art neural language models have shown very promising results for text classification, they are expensive to (pre-)train, require large amounts of data and tuning of hundreds of millions or more parameters. This paper explores how automatically evolved text representations can serve as a basis for explainable, low-resource branch of models with competitive performance that are subject to automated hyperparameter tuning. We present autoBOT (automatic Bags-Of-Tokens), an autoML approach suitable for low resource learning scenarios, where both the hardware and the amount of data required for training are limited. The proposed approach consists of an evolutionary algorithm that jointly optimizes various sparse representations of a given text (including word, subword, POS tag, keyword-based, knowledge graph-based and relational features) and two types of document embeddings (non-sparse representations). The key idea of autoBOT is that, instead of evolving at the learner level, evolution is conducted at the representation level. The proposed method offers competitive classification performance on fourteen real-world classification tasks when compared against a competitive autoML approach that evolves ensemble models, as well as state-of-the-art neural language models such as BERT and RoBERTa. Moreover, the approach is explainable, as the importance of the parts of the input space is part of the final solution yielded by the proposed optimization procedure, offering potential for meta-transfer learning. Blaz Skrlj, Matej Martinc, Nada Lavrac, Senja Pollak |
Mach. Learn. | 4 |
| 2020 | COVID-19 Therapy Target Discovery with Context-Aware Literature Mining
Matej Martinc, Blaz Skrlj, Sergej Pirkmajer, Nada Lavrac, Bojan Cestnik, Martin Marzidovsek, Senja Pollak |
DS | 7 |
| 2020 | Bisociative Literature-Based Discovery: Lessons Learned and New Prospects
Nada Lavrac, Matej Martinc, Senja Pollak, Bojan Cestnik |
ICCC | 3 |
| 2020 | A Study on Reproducibility in Computational Creativity Research
Martin Marzidovsek, Senja Pollak |
ICCC | 2 |
| 2020 | Tackling the ADReSS Challenge: A Multimodal Approach to the Automated Recognition of Alzheimer's Dementia
Matej Martinc, Senja Pollak |
INTERSPEECH | 2 |
| 2020 | CoSimLex: A Resource for Evaluating Graded Word Similarity in ContextabstractState of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far. Carlos Santos Armendariz, Matthew Purver, Matej Ulcar, Senja Pollak, Nikola Ljubesic, Mark Granroth-Wilding |
LREC | 4 |
| 2020 | Leveraging Contextual Embeddings for Detecting Diachronic Semantic ShiftabstractWe propose a new method that leverages contextual embeddings for the task of diachronic semantic shift detection by generating time specific word representations from BERT embeddings. The results of our experiments in the domain specific LiverpoolFC corpus suggest that the proposed method has performance comparable to the current state-of-the-art without requiring any time consuming domain adaptation on large corpora. The results on the newly created Brexit news corpus suggest that the method can be successfully used for the detection of a short-term yearly semantic shift. And lastly, the model also shows promising results in a multilingual settings, where the task was to detect differences and similarities between diachronic semantic shifts in different languages. Matej Martinc, Petra Kralj Novak, Senja Pollak |
LREC | 3 |
| 2019 | Combining n-grams and deep convolutional features for language variety classificationabstractAbstract This paper presents a novel neural architecture capable of outperforming state-of-the-art systems on the task of language variety classification. The architecture is a hybrid that combines character-based convolutional neural network (CNN) features with weighted bag-of-n-grams (BON) features and is therefore capable of leveraging both character-level and document/corpus-level information. We tested the system on the Discriminating between Similar Languages (DSL) language variety benchmark data set from the VarDial 2017 DSL shared task, which contains data from six different language groups, as well as on two smaller data sets (the Arabic Dialect Identification (ADI) Corpus and the German Dialect Identification (GDI) Corpus, from the VarDial 2016 ADI and VarDial 2018 GDI shared tasks, respectively). We managed to outperform the winning system in the DSL shared task by a margin of about 0.4 percentage points and the winning system in the ADI shared task by a margin of about 0.2 percentage points in terms of weighted F1 score without conducting any language group-specific parameter tweaking. An ablation study suggests that weighted BON features contribute more to the overall performance of the system than the CNN-based features, which partially explains the uncompetitiveness of deep learning approaches in the past VarDial DSL shared tasks. Finally, we have implemented our system in a workflow, available in the ClowdFlows platform, in order to make it easily available also to the non-programming members of the research community. Matej Martinc, Senja Pollak |
Nat. Lang. Eng. | 2 |
| 2018 | Simplified Hybrid Approach for Detection of Semantic Orientations in Economic Texts
Jan Stihec, Martin Znidarsic, Senja Pollak |
ECIR | 3 |
| 2018 | Conceptualising Computational Creativity: Towards automated historiography of a research field
Geraint A. Wiggins, Nada Lavrac, Vid Podpecan, Senja Pollak |
ICCC | 4 |
| 2018 | BISLON: BISociative SLOgaN generation based on stylistic literary devices
Martin Znidarsic, Matej Martinc, Andraz Repar, Senja Pollak |
ICCC | 4 |
| 2018 | SAAMEAT: Active Feature Transformation and Selection Methods for the Recognition of User Eating ConditionsabstractAutomatic recognition of eating conditions of humans could be a useful technology in health monitoring. The audio-visual information can be used in automating this process, and feature engineering approaches can reduce the dimensionality of audio-visual information. The reduced dimensionality of data (particularly feature subset selection) can assist in designing a system for eating conditions recognition with lower power, cost, memory and computation resources than a system which is designed using full dimensions of data. This paper presents Active Feature Transformation (AFT) and Active Feature Selection (AFS) methods, and applies them to all three tasks of the ICMI 2018 EAT Challenge for recognition of user eating conditions using audio and visual features. The AFT method is used for the transformation of the Mel-frequency Cepstral Coefficient and ComParE features for the classification task, while the AFS method helps in selecting a feature subset. Transformation by Principal Component Analysis (PCA) is also used for comparison. We find feature subsets of audio features using the AFS method (422 for Food Type, 104 for Likability and 68 for Difficulty out of 988 features) which provide better results than the full feature set. Our results show that AFS outperforms PCA and AFT in terms of accuracy for the recognition of user eating conditions using audio features. The AFT of visual features (facial landmarks) provides less accurate results than the AFS and AFT sets of audio features. However, the weighted score fusion of all the feature set improves the results. Fasih Haider, Senja Pollak, Eleni Zarogianni, Saturnino Luz |
ICMI | 2 |
| 2018 | Reusable workflows for gender prediction
Matej Martinc, Senja Pollak |
LREC | 2 |
| 2016 | Optimality Principles in Computational Approaches to Conceptual Blending: Do We Need Them (at) All?
Pedro Martins 0003, Senja Pollak, Tanja Urbancic, Amílcar Cardoso |
ICCC | 2 |
| 2016 | Computational Creativity Conceptualisation Grounded on ICCC Papers
Senja Pollak, Biljana Mileva-Boshkoska, Dragana Miljkovic, Geraint A. Wiggins, Nada Lavrac |
ICCC | 1 |
| 2015 | The Good, the Bad, and the AHA! Blends
Pedro Martins 0003, Tanja Urbancic, Senja Pollak, Nada Lavrac, Amílcar Cardoso |
ICCC | 3 |
| 2012 | Irregularity Detection in Categorized Document Corpora
Borut Sluban, Senja Pollak, Roel Coesemans, Nada Lavrac |
LREC | 2 |
| 2010 | Learning to Mine Definitions from Slovene Structured and Unstructured Knowledge-Rich Resources
Darja Fiser, Senja Pollak, Spela Vintar |
LREC | 2 |