VLDB 2026 Research / reviewers in the wild / expert
Boshko Koloski
dblp:277/0961
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-7330-0579ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Social Bias in Slovenia: The EEC-SL Dataset
Jaya Caporusso, Damar Hoogland, Boshko Koloski, Matthew Purver, Senja Pollak, Spela Vintar |
LREC | 3 |
| 2026 | Fully- and semi-supervised hierarchical multi-label image classification with graph learning
Marjan Stoimchev, Boshko Koloski, Jurica Levatic, Dragi Kocev, Saso Dzeroski |
Inf. Sci. | 2 |
| 2026 | FuDoBa: Fusing Document and Knowledge Graph Based Representations with Bayesian OptimisationabstractAbstract Building on the success of large language models (LLMs), LLM-based representations have dominated the document representation landscape, achieving strong performance on document embedding benchmarks. However, high-dimensional, computationally expensive LLM embeddings can be too generic or inefficient for domain-specific and resource-scarce applications. To address these limitations, we introduce FuDoBa—a Bayesian optimisation-based representation learning method that integrates LLM embeddings with domain-specific structured knowledge, sourced both locally and from external repositories such as WikiData. This fusion produces low-dimensional, task-relevant representations while reducing training complexity and yielding interpretable early-fusion weights for improved classification performance. We demonstrate the effectiveness of our approach on six datasets across two domains, showing that when paired with robust AutoML-based classifiers, our method performs on par with, or surpasses, proprietary LLM-only embedding baselines, while offering modality-wise interpretability and a smaller dimensional footprint. Boshko Koloski, Senja Pollak, Roberto Navigli, Blaz Skrlj |
Mach. Learn. | 1 |
| 2025 | Make Literature-Based Discovery Great Again Through Reproducible Pipelines
Bojan Cestnik, Andrej Kastrin, Boshko Koloski, Nada Lavrac |
IDA | 3 |
| 2025 | HorNets: learning from discrete and continuous signals with routing neural networksabstractAbstract Construction of neural network architectures suitable for learning from both continuous and discrete tabular data is challenging, as contemporary high-dimensional tabular data sets are often characterized by a relatively small set of instances and the request for efficient learning. We propose HorNets (Horn Networks), a neural network architecture with state-of-the-art performance on synthetic and real-life data sets from scarce-data tabular domains. HorNets are based on a clipped polynomial-like activation function, extended by a custom discrete-continuous routing mechanism that decides which part of the neural network to optimize based on the input’s cardinality. By explicitly modeling parts of the feature combination space or combining whole space in a linear attention-like manner, HorNets dynamically decide which mode of operation is the most suitable for a given piece of data with no explicit supervision. This architecture is one of the few approaches that reliably retrieves logical clauses (including noisy XNOR) and achieves state-of-the-art classification performance on 14 real-life biomedical high-dimensional data sets. HorNets are made freely available under a permissive license alongside a synthetic generator of categorical benchmarks. Boshko Koloski, Nada Lavrac, Blaz Skrlj |
Mach. Learn. | 1 |
| 2024 | A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News MediaabstractDehumanisation involves the perception and/or treatment of a social group’s members as less than human. This phenomenon is rarely addressed with computational linguistic techniques. We adapt a recently proposed approach for English, making it easier to transfer to other languages and to evaluate, introducing a new sentiment resource, the use of zero-shot cross-lingual valence and arousal detection, and a new method for statistical significance testing. We then apply it to study attitudes to migration expressed in Slovene newspapers, to examine changes in the Slovene discourse on migration between the 2015-16 migration crisis following the war in Syria and the 2022-23 period following the war in Ukraine. We find that while this discourse became more negative and more intense over time, it is less dehumanising when specifically addressing Ukrainian migrants compared to others. Jaya Caporusso, Damar Hoogland, Mojca Brglez, Boshko Koloski, Matthew Purver, Senja Pollak |
LREC/COLING | 4 |
| 2024 | AutoML-Guided Fusion of Entity and LLM-Based Representations for Document Classification
Boshko Koloski, Senja Pollak, Roberto Navigli, Blaz Skrlj |
DS (1) | 1 |
| 2024 | AHAM: Adapt, Help, Ask, Model Harvesting LLMs for Literature Mining
Boshko Koloski, Nada Lavrac, Bojan Cestnik, Senja Pollak, Blaz Skrlj, Andrej Kastrin |
IDA (1) | 1 |
| 2022 | Retrieval-Efficiency Trade-Off of Unsupervised Keyword Extraction
Blaz Skrlj, Boshko Koloski, Senja Pollak |
DS | 2 |
| 2022 | Out of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?abstractKeyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into ones that require training - supervised and ones that do not - unsupervised. In this study, we are interested in settings, where for a language under investigation, no training data is available. More specifically, we explore whether pretrained multilingual language models can be employed for zero-shot cross-lingual keyword extraction on low-resource languages with limited or no available labeled training data and whether they outperform state-of-the-art unsupervised keyword extractors. The comparison is conducted on six news article datasets covering two high-resource languages, English and Russian, and four low-resource languages, Croatian, Estonian, Latvian, and Slovenian. We find that the pretrained models fine-tuned on a multilingual corpus covering languages that do not appear in the test set (i.e. in a zero-shot setting), consistently outscore unsupervised models in all six languages. Boshko Koloski, Senja Pollak, Blaz Skrlj, Matej Martinc |
LREC | 1 |
| 2022 | Knowledge graph informed fake news classification via heterogeneous representation ensemblesabstractIncreasing amounts of freely available data both in textual and relational form offers exploration of richer document representations, potentially improving the model performance and robustness. An emerging problem in the modern era is fake news detection—many easily available pieces of information are not necessarily factually correct, and can lead to wrong conclusions or are used for manipulation. In this work we explore how different document representations, ranging from simple symbolic bag-of-words, to contextual, neural language model-based ones can be used for efficient fake news identification. One of the key contributions is a set of novel document representation learning methods based solely on knowledge graphs, i.e., extensive collections of (grounded) subject-predicate-object triplets. We demonstrate that knowledge graph-based representations already achieve competitive performance to conventionally accepted representation learners. Furthermore, when combined with existing, contextual representations, knowledge graph-based document representations can achieve state-of-the-art performance. To our knowledge this is the first larger-scale evaluation of how knowledge graph-based representations can be systematically incorporated into the process of fake news classification. Boshko Koloski, Timen Stepisnik Perdih, Marko Robnik-Sikonja, Senja Pollak, Blaz Skrlj |
Neurocomputing | 1 |