Marko Grobelnik

dblp:36/6757 · DBLP profile ↗
← Back
17ranked-venue papers in the field
1as first author
4since 2021 · last 2025
0000-0001-7373-5591ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6Data Mining & Knowledge Discovery · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 5Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 Improving stochastic models by smart denoising and latent representation optimization
abstract
This paper introduces an innovative deep learning-based optimization method specifically designed for data derived from stochastic processes . Addressing the prevalent issue of rapid overfitting in real-world scenarios with limited historical data , our approach focuses on denoising optimization. The method effectively balances the simultaneous optimization of latent data representation and target variables, leading to enhanced model performance. We rigorously test our approach using five diverse real-world datasets. Our study is structured into three parts: an ablation study to validate the individual components of our method, a statistical analysis using the Wilcoxon rank-sum test to confirm the superiority of our method against five research hypotheses, and a detailed exploration of parameter visualization and fine-tuning. The comprehensive evaluation demonstrates that our method not only outperforms existing techniques but also significantly contributes to the advancement of deep learning models for stochastic processes. The findings underscore the potential of our method as a robust solution to the challenges in modeling stochastic processes with deep learning , offering new avenues for efficient and accurate predictions .
Jakob Jelencic, M. Besher Massri, Ljupco Todorovski, Marko Grobelnik, Dunja Mladenic
Inf. Sci.4
2025 News dissemination: a semantic approach to barrier classification
abstract
Abstract The dissemination of information worldwide is significantly facilitated by the news media, with many events having global relevance across various regions. However, certain news events receive limited coverage restricted to specific geographic areas, due to the barriers that hinder the spread of information. These barriers can be attributed to political, geographical, economic, cultural, or linguistic factors. In this research, we propose an approach for classifying these barriers by extracting semantic information from news articles using Wikipedia-concepts. Our methodology involves the collection of news articles, each annotated to indicate the specific barrier types, leveraging metadata from news publishers. Subsequently, we employ Wikipedia-concepts, in conjunction with the content of the news articles, as features to determine the barriers to news dissemination. Our approach is then compared with traditional text classification techniques, deep learning methods, and transformer-based models. We have performed experiments on news articles from ten categories of topics including health, sports, business, etc. The findings indicate that 1) Utilizing semantic knowledge yields distinct concepts across the ten categories, thereby enhancing the effectiveness and speed of the classification model. 2) The proposed approach, incorporating Wikipedia-concepts-based semantic knowledge, leads to improved performance in barrier classification when compared to using solely the body text of news articles. Specifically, there is an increase in the average F1-scores for four out of five barriers, with the economic barrier rising from 0.65 to 0.68, the linguistic barrier from 0.71 to 0.72, the political barrier from 0.68 to 0.70, and the geographical barrier from 0.63 to 0.68.
Abdul Sittar, Dunja Mladenic, Marko Grobelnik
J. Intell. Inf. Syst.3
2022 Analysis of information cascading and propagation barriers across distinctive news events
abstract
News reporting, on events that occur in our society, can have different styles and structures, as well as different dynamics of news spreading over time. News publishers have the potential to spread their news and reach out to a large number of readers worldwide. In this paper we would like to understand how well they are doing it and which kind of obstacles the news may encounter when spreading. The news to be spread wider cross multiple barriers such as linguistic (the most evident one, as they get published in other natural languages), economic, geographical, political, time zone, and cultural barriers. Observing potential differences between spreading of news on different events published by multiple publishers can bring insights into what may influence the differences in the spreading patterns. There are multiple reasons, possibly many hidden, influencing the speed and geographical spread of news. This paper studies information cascading and propagation barriers, applying the proposed methodology on three distinctive kinds of events: Global Warming, earthquakes, and FIFA World Cup. Our findings suggest that 1) the scope of a specific event significantly effects the news spreading across languages, 2) geographical size of a news publisher's country is directly proportional to the number of publishers and articles reporting on the same information, 3) countries with shorter time-zone differences and similar cultures tend to propagate news between each other, 4) news related to Global Warming comes across economic barriers more smoothly than news related to FIFA World Cup and earthquakes and 5) events which may in some way involve political benefits are mostly published by those publishers which are not politically neutral.
Abdul Sittar, Dunja Mladenic, Marko Grobelnik
J. Intell. Inf. Syst.3
2021 WIT: Workshop on deriving Insights from user-generated Text
abstract
User-Generated text is a rich source of user insights and experiences that can be very helpful in many different daily life situations, such as when deciding what product to buy, what hotel to stay, what company to apply for a job, what region to buy a house, etc. This kind of text also plays a very relevant role in current research efforts in academic research groups, technology companies, as well as big publishers, telecommunications players, recruiting and job-market focuses organizations, etc. The goal of this new workshop is to bring together researchers interested in the application of novel techniques in AI/ML/NLP and Knowledge Discovery to address challenges around harnessing text-heavy user-generated data that is available to organizations and over the Web. The workshop program contains invited speakers, contributed talks, poster sessions and a discussion panel
Estevam Hruschka, Tom M. Mitchell, Marko Grobelnik, Behzad Golshan
KDD3
2017 Curious Cat-Mobile, Context-Aware Conversational Crowdsourcing Knowledge Acquisition
abstract
Scaled acquisition of high-quality structured knowledge has been a longstanding goal of Artificial Intelligence research. Recent advances in crowdsourcing, the sheer number of Internet and mobile users, and the commercial availability of supporting platforms offer new tools for knowledge acquisition. This article applies context-aware knowledge acquisition that simultaneously satisfies users’ immediate information needs while extending its own knowledge using crowdsourcing. The focus is on knowledge acquisition on a mobile device, which makes the approach practical and scalable; in this context, we propose and implement a new KA approach that exploits an existing knowledge base to drive the KA process, communicate with the right people, and check for consistency of the user-provided answers. We tested the viability of the approach in experiments using our platform with real users around the world, and an existing large source of common-sense background knowledge. These experiments show that the approach is promising: the knowledge is estimated to be true and useful for users 95% of the time. Using context to proactively drive knowledge acquisition increased engagement and effectiveness (the number of new assertions/day/user increased for 175%). Using pre-existing and newly acquired knowledge also proved beneficial.
Luka Bradesko, Michael Witbrock, Janez Starc, Zala Herga, Marko Grobelnik, Dunja Mladenic
ACM Trans. Inf. Syst.5
2015 Editorial
Dunja Mladenic, Estevam Hruschka, Marko Grobelnik
J. Web Semant.3
2011 Search and mining entity-relationship data
abstract
This paper summarizes the details of the first international workshop on search and mining entity-relationship data. This workshop will bridge between IR, DB, and KM researchers to seek novel solutions for search and data mining of rich entity-relationship data and their applications in various domains. We first provide an overview about the workshop. We then briefly discuss the workshop program.
Haggai Roitman, Ralf Schenkel, Marko Grobelnik
CIKM3
2010 AnswerArt - Contextualized Question Answering
Lorand Dali, Delia Rusu, Blaz Fortuna, Dunja Mladenic, Marko Grobelnik
ECML/PKDD (3)5
2010 Video search: are algorithms all we need?
abstract
This panel will debate various approaches to improving video search, and explore how professional cataloguing, crowd sourced metadata, and improvements in search algorithms will evolve over the next ten years. Panelists will explore the needs of large scale video archives, and compare these against the current capabilities of video search.
Jeff Ubois, Jamie Davidson, Marko Grobelnik, Paul Over, Hans Westerhof
WWW3
2009 Demo: HistoryViz - Visualizing Events and Relations Extracted from Wikipedia
Ruben Sipos, Abhijit Bhole, Blaz Fortuna, Marko Grobelnik, Dunja Mladenic
ESWC4
2009 Guest editors' introduction: special issue of selected papers from ECML PKDD 2009
Alek Kolcz, Dunja Mladenic, Wray L. Buntine, Marko Grobelnik, John Shawe-Taylor
Data Min. Knowl. Discov.4
2008 Monitoring Network Evolution using MDL
abstract
Given publication titles and authors, what can we say about the evolution of scientific topics and communities over time? Which communities shrunk, which emerged, and which split, over time? And, when in time were the turning points? We propose TimeFall, which can automatically answer these questions given a social network/graph that evolves over time. The main novelty of the proposed approach is that it needs no user-defined parameters, relying instead on the principle of minimum description length (MDL), to extract the communities, and to find good cut-points in time when communities change abruptly: a cut-point is good, if it leads to shorter data description. We illustrate our algorithm on synthetic and large real datasets, and we show that the results of the TimeFall agree with human intuition.
Jure Ferlez, Christos Faloutsos, Jure Leskovec, Dunja Mladenic, Marko Grobelnik
ICDE5
2008 Cross-lingual search over 22 european languages
abstract
In this paper we present a system for cross-lingual information retrieval, which can handle tens of languages and millions of documents. Functioning of the system is demonstrated on corpus of European Legislation (22 languages, more than 400,000 documents per language). The system uses an interactive web-interface, which can take advantage of a predefined thesaurus allowing the user to dynamically re-rank the retrieval results based on the mapping onto a predefined thesaurus.
Blaz Fortuna, Jan Rupnik, Bostjan Pajntar, Marko Grobelnik, Dunja Mladenic
SIGIR4
2006 Background knowledge for ontology construction
abstract
In this paper we describe a solution for incorporating background knowledge into the OntoGen system for semi-automatic ontology construction. This makes it easier for different users to construct different and more personalized ontologies for the same domain. To achieve this we introduce a word weighting schema to be used in the document representation. The weighting schema is learned based on the background knowledge provided by user. It is than used by OntoGen's machine learning and text mining algorithms.
Blaz Fortuna, Marko Grobelnik, Dunja Mladenic
WWW2
2004 Feature selection using linear classifier weights: interaction with classification models
abstract
This paper explores feature scoring and selection based on weights from linear classification models. It investigates how these methods combine with various learning models. Our comparative analysis includes three learning algorithms: Naïve Bayes, Perceptron, and Support Vector Machines (SVM) in combination with three feature weighting methods: Odds Ratio, Information Gain, and weights from linear models, the linear SVM and Perceptron. Experiments show that feature selection using weights from linear SVMs yields better classification performance than other feature weighting methods when combined with the three explored learning algorithms. The results support the conjecture that it is the sophistication of the feature weighting method rather than its apparent compatibility with the learning algorithm that improves classification performance.
Dunja Mladenic, Janez Brank, Marko Grobelnik, Natasa Milic-Frayling
SIGIR3
2000 Text mining (workshop session - title only)
Marko Grobelnik, Dunja Mladenic, Natasa Milic-Frayling
KDD1
1994 Using Machine Learning Techniques to Interpret Results from Discrete Event Simulation
Dunja Mladenic, Ivan Bratko, Ray J. Paul, Marko Grobelnik
ECML4