Marko Grobelnik

dblp:36/6757 · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-7373-5591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 MRM3: Machine Readable ML Model Metadata
Andrej Cop, Blaz Bertalanic, Marko Grobelnik, Carolina Fortuna
MobiSys3
2025 Improving stochastic models by smart denoising and latent representation optimization
abstract
This paper introduces an innovative deep learning-based optimization method specifically designed for data derived from stochastic processes . Addressing the prevalent issue of rapid overfitting in real-world scenarios with limited historical data , our approach focuses on denoising optimization. The method effectively balances the simultaneous optimization of latent data representation and target variables, leading to enhanced model performance. We rigorously test our approach using five diverse real-world datasets. Our study is structured into three parts: an ablation study to validate the individual components of our method, a statistical analysis using the Wilcoxon rank-sum test to confirm the superiority of our method against five research hypotheses, and a detailed exploration of parameter visualization and fine-tuning. The comprehensive evaluation demonstrates that our method not only outperforms existing techniques but also significantly contributes to the advancement of deep learning models for stochastic processes. The findings underscore the potential of our method as a robust solution to the challenges in modeling stochastic processes with deep learning , offering new avenues for efficient and accurate predictions .
Jakob Jelencic, M. Besher Massri, Ljupco Todorovski, Marko Grobelnik, Dunja Mladenic
Inf. Sci.4
2025 News dissemination: a semantic approach to barrier classification
abstract
Abstract The dissemination of information worldwide is significantly facilitated by the news media, with many events having global relevance across various regions. However, certain news events receive limited coverage restricted to specific geographic areas, due to the barriers that hinder the spread of information. These barriers can be attributed to political, geographical, economic, cultural, or linguistic factors. In this research, we propose an approach for classifying these barriers by extracting semantic information from news articles using Wikipedia-concepts. Our methodology involves the collection of news articles, each annotated to indicate the specific barrier types, leveraging metadata from news publishers. Subsequently, we employ Wikipedia-concepts, in conjunction with the content of the news articles, as features to determine the barriers to news dissemination. Our approach is then compared with traditional text classification techniques, deep learning methods, and transformer-based models. We have performed experiments on news articles from ten categories of topics including health, sports, business, etc. The findings indicate that 1) Utilizing semantic knowledge yields distinct concepts across the ten categories, thereby enhancing the effectiveness and speed of the classification model. 2) The proposed approach, incorporating Wikipedia-concepts-based semantic knowledge, leads to improved performance in barrier classification when compared to using solely the body text of news articles. Specifically, there is an increase in the average F1-scores for four out of five barriers, with the economic barrier rising from 0.65 to 0.68, the linguistic barrier from 0.71 to 0.72, the political barrier from 0.68 to 0.70, and the geographical barrier from 0.63 to 0.68.
Abdul Sittar, Dunja Mladenic, Marko Grobelnik
J. Intell. Inf. Syst.3
2023 Seeking information about assistive technology: Exploring current practices, challenges, and the need for smarter systems
abstract
Ninety percent of the 1.2 billion people who need assistive technology (AT) do not have access. Information seeking practices directly impact the ability of AT producers, procurers, and providers (AT professionals) to match a user's needs with appropriate AT, yet the AT marketplace is interdisciplinary and fragmented, complicating information seeking. We explored common limitations experienced by AT professionals when searching information to develop solutions for a diversity of users with multi-faceted needs. Through Template Analysis of 22 expert interviews, we find current search engines do not yield the necessary information, or appropriately tailor search results, impacting individuals’ awareness of products and subsequently their availability and the overall effectiveness of AT provision. We present value-based design implications to improve functionality of future AT-information seeking platforms, through incorporating smarter systems to support decision-making and need-matching whilst ensuring ethical standards for disability fairness remain.
Jamie Danemayer, Catherine Holloway, Youngjun Cho, Nadia Bianchi-Berthouze, Aneesha Singh, William Bhot, Ollie Dixon, Marko Grobelnik, John Shawe-Taylor
Int. J. Hum. Comput. Stud.8
2023 A commonsense-infused language-agnostic learning framework for enhancing prediction of political bias in multilingual news headlines
abstract
Predicting the political bias of news headlines is a challenging task that becomes even more challenging in a multilingual setting with low-resource languages. To deal with this, we propose to utilise Inferential Commonsense Knowledge via a Translate-Retrieve-Translate strategy to introduce a learning framework. To begin with, we use the translate-retrieve-translate strategy to acquire inferential knowledge in the target language. We then employ an attention mechanism to emphasise important inferences. We finally integrate the attended inferences into a multilingual, pre-trained language model for the task of bias prediction. To evaluate the effectiveness of our framework, we present a dataset of over 62.6K multilingual news headlines annotated with their respective political biases in five low-resource European languages. We evaluate several state-of-the-art multilingual pre-trained language models since their performance tends to vary across languages (low or high resource). Evaluation results demonstrate that our proposed framework is effective regardless of the models employed. Overall, the best-performing model trained with only headlines shows 0.90 accuracy and F1 and a 0.83 Jaccard score. With attended knowledge in our framework, the same model shows an increase in 2.2% accuracy and F1 and a 3.6% Jaccard score. Extending our experiments to individual languages reveals that the models we analyse for Slovenian perform significantly worse than other languages in our dataset. To investigate this, we assess the effect of translation quality on prediction performance. It indicates that the disparity in performance is most likely due to poor translation quality. We release our dataset and scripts at https://github.com/Swati17293/KG-Multi-Bias for future research. Our framework has the potential to benefit journalists, social scientists, news producers, and consumers.
Adrian Mladenic Grobelnik, Dunja Mladenic, Marko Grobelnik
Knowl. Based Syst.4
2022 Analysis of information cascading and propagation barriers across distinctive news events
abstract
News reporting, on events that occur in our society, can have different styles and structures, as well as different dynamics of news spreading over time. News publishers have the potential to spread their news and reach out to a large number of readers worldwide. In this paper we would like to understand how well they are doing it and which kind of obstacles the news may encounter when spreading. The news to be spread wider cross multiple barriers such as linguistic (the most evident one, as they get published in other natural languages), economic, geographical, political, time zone, and cultural barriers. Observing potential differences between spreading of news on different events published by multiple publishers can bring insights into what may influence the differences in the spreading patterns. There are multiple reasons, possibly many hidden, influencing the speed and geographical spread of news. This paper studies information cascading and propagation barriers, applying the proposed methodology on three distinctive kinds of events: Global Warming, earthquakes, and FIFA World Cup. Our findings suggest that 1) the scope of a specific event significantly effects the news spreading across languages, 2) geographical size of a news publisher's country is directly proportional to the number of publishers and articles reporting on the same information, 3) countries with shorter time-zone differences and similar cultures tend to propagate news between each other, 4) news related to Global Warming comes across economic barriers more smoothly than news related to FIFA World Cup and earthquakes and 5) events which may in some way involve political benefits are mostly published by those publishers which are not politically neutral.
Abdul Sittar, Dunja Mladenic, Marko Grobelnik
J. Intell. Inf. Syst.3
2022 Why is a document relevant? Understanding the relevance scores in cross-lingual document retrieval
abstract
Modern cross-lingual document retrieval models are capable of finding documents relevant to the query. However, they do not have the capabilities for explaining why the document is relevant. This paper proposes a novel learning-to-rank model named LM-EMD that uses the multilingual BERT language model and Earth Mover’s Distance (EMD) to measure the document’s relevancy to the input query and provide interpretable insights into why a document is relevant. The model uses the query and document token’s contextual embeddings generated with multilingual BERT to measure their distances in the embedding space, which are then used by EMD to calculate the document’s relevance score and identify which document tokens contribute the most to its relevancy. We evaluate the model on five language pairs of varying degrees of similarity and analyze its performance. We find that the model (1) performs similar as the best performing comparing model on high-resource languages, (2) is less effective on low-resource languages, and (3) provides insight into why a document is relevant to the query.
Erik Novak, Luka Bizjak, Dunja Mladenic, Marko Grobelnik
Knowl. Based Syst.4
2021 WIT: Workshop on deriving Insights from user-generated Text
abstract
User-Generated text is a rich source of user insights and experiences that can be very helpful in many different daily life situations, such as when deciding what product to buy, what hotel to stay, what company to apply for a job, what region to buy a house, etc. This kind of text also plays a very relevant role in current research efforts in academic research groups, technology companies, as well as big publishers, telecommunications players, recruiting and job-market focuses organizations, etc. The goal of this new workshop is to bring together researchers interested in the application of novel techniques in AI/ML/NLP and Knowledge Discovery to address challenges around harnessing text-heavy user-generated data that is available to organizations and over the Web. The workshop program contains invited speakers, contributed talks, poster sessions and a discussion panel
Estevam Hruschka, Tom M. Mitchell, Marko Grobelnik, Behzad Golshan
KDD3
2021 NewsMeSH: A new classifier designed to annotate health news with MeSH headings
João Pita Costa, Luis Rei, Luka Stopar, Flavio Fuart, Marko Grobelnik, Dunja Mladenic, Inna Novalija, Anthony Staines, Jarmo Pääkkönen, Jenni Konttila, Joseba Bidaurrazaga, Oihana Belar, Christine Henderson, Gorka Epelde, Monica Arrue, Paul Carlin, Jonathan G. Wallace
Artif. Intell. Medicine5
2019 StreamStory: Exploring Multivariate Time Series on Multiple Scales
abstract
This paper presents an approach for the interactive visualization, exploration and interpretation of large multivariate time series. Interesting patterns in such datasets usually appear as periodic or recurrent behavior often caused by the interaction between variables. To identify such patterns, we summarize the data as conceptual states, modeling temporal dynamics as transitions between the states. This representation can visualize large datasets with potentially billions of examples. We extend the representation to multiple spatial granularities allowing the user to find patterns on multiple scales. The result is an interactive web-based tool called StreamStory. StreamStory couples the abstraction with several tools that map the abstractions back to domain-specific concepts using techniques from statistics and machine learning. It is aimed at users who are not experts in data analytics, minimizing the number of parameters to configure out-of-the-box. We use three real-world datasets to demonstrate how StreamStory can be used to perform three main visual analytics tasks: identify the main states of a complex system and map them back to data-specific concepts, find high-level and long-term periodic behavior and traverse the scales to identify which scales exhibit interesting phenomena. We find and interpret several known, as well as previously unknown patterns in these datasets.
Luka Stopar, Primoz Skraba, Marko Grobelnik, Dunja Mladenic
IEEE Trans. Vis. Comput. Graph.3
2018 Real-Time Data-Intensive Telematics Functionalities at the Extreme Edge of the Network: Experience with the PrEstoCloud Project
abstract
In recent years, use of different sensors connected to vehicles is dramatically increasing in order to enhance transportation efficiency. The current Big Data technologies are predominantly used to store large amount of telematics data especially in the cloud, and they are only able to perform simple querying for the purpose of reporting. While all the data is stored in the cloud-centric datacenters, these telematics systems are not capable of exploiting other functionalities offered by advanced real-time analytics such as run-time anomaly detection. In this paper, we propose an advanced telematics system orchestrated upon an edge computing framework in the context of the PrEstoCloud (Proactive Cloud Resources Management at the Edge for Efficient Real-Time Big Data Processing) project. This telematics system is a real-time data-intensive application running at the extreme edge of the network for drivers' behavior profiling and triggering run-time alerts. Such functionalities may be useful in order to notify stakeholders for example drivers and logistic centers on situations where a possible accident may occur or attention is required.
Salman Taherizadeh, Blaz Novak, Marija Komatar, Marko Grobelnik
COMPSAC (2)4
2017 News Across Languages - Cross-Lingual Document Similarity and Event Tracking (Extended Abstract)
abstract
In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream. Within a recently developed system Event Registry we examine two aspects of this problem: how to compare articles in different languages and how to link collections of articles in different languages which refer to the same event. Building on previous work, we show there are methods which scale well and can compute a meaningful similarity between articles from languages with little or no direct overlap in the training data.Using this capability, we then propose an approach to link clusters of articles across languages which represent the same event.
Jan Rupnik, Andrej Muhic, Gregor Leban, Blaz Fortuna, Marko Grobelnik
IJCAI5
2017 Curious Cat-Mobile, Context-Aware Conversational Crowdsourcing Knowledge Acquisition
abstract
Scaled acquisition of high-quality structured knowledge has been a longstanding goal of Artificial Intelligence research. Recent advances in crowdsourcing, the sheer number of Internet and mobile users, and the commercial availability of supporting platforms offer new tools for knowledge acquisition. This article applies context-aware knowledge acquisition that simultaneously satisfies users’ immediate information needs while extending its own knowledge using crowdsourcing. The focus is on knowledge acquisition on a mobile device, which makes the approach practical and scalable; in this context, we propose and implement a new KA approach that exploits an existing knowledge base to drive the KA process, communicate with the right people, and check for consistency of the user-provided answers. We tested the viability of the approach in experiments using our platform with real users around the world, and an existing large source of common-sense background knowledge. These experiments show that the approach is promising: the knowledge is estimated to be true and useful for users 95% of the time. Using context to proactively drive knowledge acquisition increased engagement and effectiveness (the number of new assertions/day/user increased for 175%). Using pre-existing and newly acquired knowledge also proved beneficial.
Luka Bradesko, Michael Witbrock, Janez Starc, Zala Herga, Marko Grobelnik, Dunja Mladenic
ACM Trans. Inf. Syst.5
2016 News Across Languages - Cross-Lingual Document Similarity and Event Tracking
abstract
In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream. Within a recently developed system Event Registry we examine two aspects of this problem: how to compare articles in different languages and how to link collections of articles in different languages which refer to the same event. Taking a multilingual stream and clusters of articles from each language, we compare different cross-lingual document similarity measures based on Wikipedia. This allows us to compute the similarity of any two articles regardless of language. Building on previous work, we show there are methods which scale well and can compute a meaningful similarity between articles from languages with little or no direct overlap in the training data. Using this capability, we then propose an approach to link clusters of articles across languages which represent the same event. We provide an extensive evaluation of the system as a whole, as well as an evaluation of the quality and robustness of the similarity measure and the linking algorithm.
Jan Rupnik, Andrej Muhic, Gregor Leban, Primoz Skraba, Blaz Fortuna, Marko Grobelnik
J. Artif. Intell. Res.6
2015 Editorial
Dunja Mladenic, Estevam Hruschka, Marko Grobelnik
J. Web Semant.3
2014 The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite
LREC12
2011 Search and mining entity-relationship data
abstract
This paper summarizes the details of the first international workshop on search and mining entity-relationship data. This workshop will bridge between IR, DB, and KM researchers to seek novel solutions for search and data mining of rich entity-relationship data and their applications in various domains. We first provide an overview about the workshop. We then briefly discuss the workshop program.
Haggai Roitman, Ralf Schenkel, Marko Grobelnik
CIKM3
2010 AnswerArt - Contextualized Question Answering
Lorand Dali, Delia Rusu, Blaz Fortuna, Dunja Mladenic, Marko Grobelnik
ECML/PKDD (3)5
2010 Video search: are algorithms all we need?
abstract
This panel will debate various approaches to improving video search, and explore how professional cataloguing, crowd sourced metadata, and improvements in search algorithms will evolve over the next ten years. Panelists will explore the needs of large scale video archives, and compare these against the current capabilities of video search.
Jeff Ubois, Jamie Davidson, Marko Grobelnik, Paul Over, Hans Westerhof
WWW3
2009 Demo: HistoryViz - Visualizing Events and Relations Extracted from Wikipedia
Ruben Sipos, Abhijit Bhole, Blaz Fortuna, Marko Grobelnik, Dunja Mladenic
ESWC4
2009 Document Visualization Based on Semantic Graphs
abstract
In this paper, we present a document visualization technique for data analysis based on the semantic representation of text in the form of a directed graph, referred to as semantic graph. It is derived using natural language processing as follows. Firstly subject– verb – object triplets are automatically extracted from the Penn Treebank parse tree obtained for each sentence in the document. Secondly, the triplets are further enhanced by linking them to their corresponding co-referenced named entity, by resolving pronominal anaphors as well as attaching the associated WordNet synset. Starting from the document's semantic graph and the list of extracted triplets we automatically generate the document summary, for which we also derive the semantic representation.
Delia Rusu, Blaz Fortuna, Dunja Mladenic, Marko Grobelnik, Ruben Sipos
IV4
2009 Guest editors' introduction: special issue of selected papers from ECML PKDD 2009
Alek Kolcz, Dunja Mladenic, Wray L. Buntine, Marko Grobelnik, John Shawe-Taylor
Data Min. Knowl. Discov.4
2009 Guest editors' introduction: Special Issue from ECML PKDD 2009
Alek Kolcz, Dunja Mladenic, Wray L. Buntine, Marko Grobelnik, John Shawe-Taylor
Mach. Learn.4
2008 Monitoring Network Evolution using MDL
abstract
Given publication titles and authors, what can we say about the evolution of scientific topics and communities over time? Which communities shrunk, which emerged, and which split, over time? And, when in time were the turning points? We propose TimeFall, which can automatically answer these questions given a social network/graph that evolves over time. The main novelty of the proposed approach is that it needs no user-defined parameters, relying instead on the principle of minimum description length (MDL), to extract the communities, and to find good cut-points in time when communities change abruptly: a cut-point is good, if it leads to shorter data description. We illustrate our algorithm on synthetic and large real datasets, and we show that the results of the TimeFall agree with human intuition.
Jure Ferlez, Christos Faloutsos, Jure Leskovec, Dunja Mladenic, Marko Grobelnik
ICDE5
2008 Cross-lingual search over 22 european languages
abstract
In this paper we present a system for cross-lingual information retrieval, which can handle tens of languages and millions of documents. Functioning of the system is demonstrated on corpus of European Legislation (22 languages, more than 400,000 documents per language). The system uses an interactive web-interface, which can take advantage of a predefined thesaurus allowing the user to dynamically re-rank the retrieval results based on the mapping onto a predefined thesaurus.
Blaz Fortuna, Jan Rupnik, Bostjan Pajntar, Marko Grobelnik, Dunja Mladenic
SIGIR4
2006 Background knowledge for ontology construction
abstract
In this paper we describe a solution for incorporating background knowledge into the OntoGen system for semi-automatic ontology construction. This makes it easier for different users to construct different and more personalized ontologies for the same domain. To achieve this we introduce a word weighting schema to be used in the document representation. The weighting schema is learned based on the background knowledge provided by user. It is than used by OntoGen's machine learning and text mining algorithms.
Blaz Fortuna, Marko Grobelnik, Dunja Mladenic
WWW2
2005 Impact of Linguistic Analysis on the Semantic Graph Coverage and Learning of Document Extracts
Jure Leskovec, Natasa Milic-Frayling, Marko Grobelnik
AAAI3
2004 Feature selection using linear classifier weights: interaction with classification models
abstract
This paper explores feature scoring and selection based on weights from linear classification models. It investigates how these methods combine with various learning models. Our comparative analysis includes three learning algorithms: Naïve Bayes, Perceptron, and Support Vector Machines (SVM) in combination with three feature weighting methods: Odds Ratio, Information Gain, and weights from linear models, the linear SVM and Perceptron. Experiments show that feature selection using weights from linear SVMs yields better classification performance than other feature weighting methods when combined with the three explored learning algorithms. The results support the conjecture that it is the sophistication of the feature weighting method rather than its apparent compatibility with the learning algorithm that improves classification performance.
Dunja Mladenic, Janez Brank, Marko Grobelnik, Natasa Milic-Frayling
SIGIR3
2003 Feature selection on hierarchy of web documents
Dunja Mladenic, Marko Grobelnik
Decis. Support Syst.2
2000 Text mining (workshop session - title only)
Marko Grobelnik, Dunja Mladenic, Natasa Milic-Frayling
KDD1
1999 Feature Selection for Unbalanced Class Distribution and Naive Bayes
Dunja Mladenic, Marko Grobelnik
ICML2
1994 Using Machine Learning Techniques to Interpret Results from Discrete Event Simulation
Dunja Mladenic, Ivan Bratko, Ray J. Paul, Marko Grobelnik
ECML4
1992 Stochastic Search in Inductive Logic Programming
Matevz Kovacic, Nada Lavrac, Marko Grobelnik, Darko Zupanic, Dunja Mladenic
ECAI3