Jakub Simko

dblp:09/8578 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-0239-4237ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
abstract
Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models.However, a comparison of various generation strategies for low-resource language settings is lacking.While various prompting strategies have been proposed-such as demonstrations, labelbased summaries, and self-revision-their comparative effectiveness remains unclear, especially for low-resource languages.In this paper, we systematically evaluate the performance of these generation strategies and their combinations across 11 typologically diverse languages, including several extremely low-resource ones.Using three NLP tasks and four open-source LLMs, we assess downstream model performance on generated versus gold-standard data.Our results show that strategic combinations of generation methods-particularly targetlanguage demonstrations with LLM-based revisions-yield strong performance, narrowing the gap with real data to as little as 5% in some settings.We also find that smart prompting techniques can reduce the advantage of larger LLMs, highlighting efficient generation strategies for synthetic data generation in lowresource scenarios with smaller models.
Tatiana Anikina, Ján Cegin, Jakub Simko, Simon Ostermann 0002
EMNLP3
2025 LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
abstract
Jan Cegin, Jakub Simko, Peter Brusilovsky. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ján Cegin, Jakub Simko, Peter Brusilovsky
NAACL (Long Papers)2
2024 Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
abstract
Jan Cegin, Branislav Pecher, Jakub Simko, Ivan Srba, Maria Bielikova, Peter Brusilovsky. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ján Cegin, Branislav Pecher, Jakub Simko, Ivan Srba, Mária Bieliková, Peter Brusilovsky
ACL (1)3
2024 Multilinguality in the VIGILANT project
abstract
VIGILANT (Vital IntelliGence to Investigate ILlegAl DisiNformaTion) is a three-year Horizon Europe project that will equip European Law Enforcement Agencies (LEAs) with advanced disinformation detection and analysis tools to investigate and prevent criminal activities linked to disinformation. These include disinformation instigating violence towards minorities, promoting false medical cures, and increasing tensions between groups causing civil unrest and violent acts. VIGILANT’s four LEAs require support for English, Spanish, Catalan, Greek, Estonian, Romanian and Russian. Therefore, multilinguality is a major challenge and we present the current status of our tools and our plans to improve their performance.
Brendan Spillane, Carolina Scarton, Róbert Móro, Petar Ivanov, Andrey Tagarev, Jakub Simko, Ibrahim Abu Farha, Gary Munnelly, Filip Uhlárik, Freddy Heppell
EAMT (2)6
2023 ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness
abstract
The emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing?Traditionally, crowdsourcing has been used for acquiring solutions to a wide variety of human-intelligence tasks, including ones involving text generation, modification or evaluation.For some of these tasks, models like ChatGPT can potentially substitute human workers.In this study, we investigate whether this is the case for the task of paraphrase generation for intent classification.We apply data collection methodology of an existing crowdsourcing study (similar scale, prompts and seed data) using ChatGPT and Falcon-40B.We show that ChatGPT-created paraphrases are more diverse and lead to at least as robust models.
Ján Cegin, Jakub Simko, Peter Brusilovsky
EMNLP2
2023 MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
abstract
Dominik Macko, Robert Moro, Adaku Uchendu, Jason Lucas, Michiharu Yamashita, Matúš Pikuliak, Ivan Srba, Thai Le, Dongwon Lee, Jakub Simko, Maria Bielikova. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Dominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas, Michiharu Yamashita, Matús Pikuliak, Ivan Srba, Thai Le, Dongwon Lee 0001, Jakub Simko, Mária Bieliková
EMNLP10
2023 Multilingual Previously Fact-Checked Claim Retrieval
abstract
Matúš Pikuliak, Ivan Srba, Robert Moro, Timo Hromadka, Timotej Smoleň, Martin Melišek, Ivan Vykopal, Jakub Simko, Juraj Podroužek, Maria Bielikova. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Matús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka, Timotej Smolen, Martin Melisek, Ivan Vykopal, Jakub Simko, Juraj Podrouzek, Mária Bieliková
EMNLP8
2023 Auditing YouTube's Recommendation Algorithm for Misinformation Filter Bubbles
abstract
In this article, we present results of an auditing study performed over YouTube aimed at investigating how fast a user can get into a misinformation filter bubble, but also what it takes to “burst the bubble,” i.e., revert the bubble enclosure. We employ a sock puppet audit methodology, in which pre-programmed agents (acting as YouTube users) delve into misinformation filter bubbles by watching misinformation-promoting content. Then they try to burst the bubbles and reach more balanced recommendations by watching misinformation-debunking content. We record search results, home page results, and recommendations for the watched videos. Overall, we recorded 17,405 unique videos, out of which we manually annotated 2,914 for the presence of misinformation. The labeled data was used to train a machine learning model classifying videos into three classes (promoting, debunking, neutral) with the accuracy of 0.82. We use the trained model to classify the remaining videos that would not be feasible to annotate manually. Using both the manually and automatically annotated data, we observe the misinformation bubble dynamics for a range of audited topics. Our key finding is that even though filter bubbles do not appear in some situations, when they do, it is possible to burst them by watching misinformation-debunking content (albeit it manifests differently from topic to topic). We also observe a sudden decrease of misinformation filter bubble effect when misinformation-debunking videos are watched after misinformation-promoting videos, suggesting a strong contextuality of recommendations. Finally, when comparing our results with a previous similar study, we do not observe significant improvements in the overall quantity of recommended misinformation content.
Ivan Srba, Róbert Móro, Matús Tomlein, Branislav Pecher, Jakub Simko, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Adrian Gavornik, Mária Bieliková
Trans. Recomm. Syst.5
2022 Black-box Audit of YouTube's Video Recommendation: Investigation of Misinformation Filter Bubble Dynamics (Extended Abstract)
abstract
In this paper, we describe a black-box sockpuppeting audit which we carried out to investigate the creation and bursting dynamics of misinformation filter bubbles on YouTube. Pre-programmed agents acting as YouTube users stimulated YouTube's recommender systems: they first watched a series of misinformation promoting videos (bubble creation) and then a series of misinformation debunking videos (bubble bursting). Meanwhile, agents logged videos recommended to them by YouTube. After manually annotating these recommendations, we were able to quantify the portion of misinformative videos among them. The results confirm the creation of filter bubbles (albeit not in all situations) and show that these bubbles can be bursted by watching credible content. Drawing a direct comparison with a previous study, we do not see improvements in overall quantities of misinformation recommended.
Matús Tomlein, Branislav Pecher, Jakub Simko, Ivan Srba, Róbert Móro, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Mária Bieliková
IJCAI3
2022 Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims
abstract
False information has a significant negative influence on individuals as well as on the whole society. Especially in the current COVID-19 era, we witness an unprecedented growth of medical misinformation. To help tackle this problem with machine learning approaches, we are publishing a feature-rich dataset of approx. 317k medical news articles/blogs and 3.5k fact-checked claims. It also contains 573 manually and more than 51k automatically labelled mappings between claims and articles. Mappings consist of claim presence, i.e., whether a claim is contained in a given article, and article stance towards the claim. We provide several baselines for these two tasks and evaluate them on the manually labelled part of the dataset. The dataset enables a number of additional tasks related to medical misinformation, such as misinformation characterisation studies or studies of misinformation diffusion between sources.
Ivan Srba, Branislav Pecher, Matús Tomlein, Róbert Móro, Elena Stefancova, Jakub Simko, Mária Bieliková
SIGIR6
2021 An Audit of Misinformation Filter Bubbles on YouTube: Bubble Bursting and Recent Behavior Changes
abstract
The negative effects of misinformation filter bubbles in adaptive systems have been known to researchers for some time. Several studies investigated, most prominently on YouTube, how fast a user can get into a misinformation filter bubble simply by selecting “wrong choices” from the items offered. Yet, no studies so far have investigated what it takes to “burst the bubble”, i.e., revert the bubble enclosure. We present a study in which pre-programmed agents (acting as YouTube users) delve into misinformation filter bubbles by watching misinformation promoting content (for various topics). Then, by watching misinformation debunking content, the agents try to burst the bubbles and reach more balanced recommendation mixes. We recorded the search results and recommendations, which the agents encountered, and analyzed them for the presence of misinformation. Our key finding is that bursting of a filter bubble is possible, albeit it manifests differently from topic to topic. Moreover, we observe that filter bubbles do not truly appear in some situations. We also draw a direct comparison with a previous study. Sadly, we did not find much improvements in misinformation occurrences, despite recent pledges by YouTube.
Matús Tomlein, Branislav Pecher, Jakub Simko, Ivan Srba, Róbert Móro, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Mária Bieliková
RecSys3
2019 Web-Navigation Skill Assessment Through Eye-Tracking Data
Patrik Hlavac, Jakub Simko, Mária Bieliková
ADBIS2
2019 Impact of English Reading Comprehension Abilities on Processing Magazine Style Narrative Visualizations and Implications for Personalization
abstract
In this paper, we present research to uncover how the level of reading comprehension abilities impacts how users process textual documents in English with embedded visualizations (i.e., Magazine Style Narrative Visualizations or MSNVs). We analyze performance and gaze data of users processing MSNVs from two user studies, one run in Canada and one in a non-English speaking European country. Our findings provide important insights toward developing automatic, real-time support to MSNV processing personalized according to users' English reading comprehension abilities.
Dereck Toker, Róbert Móro, Jakub Simko, Mária Bieliková, Cristina Conati
UMAP3
2019 Screen recording segmentation to scenes for eye-tracking analysis
Jakub Simko, Jakub Vrba
Multim. Tools Appl.1
2013 Classsourcing: Crowd-Based Validation of Question-Answer Learning Objects
Jakub Simko, Marián Simko, Mária Bieliková, Jakub Sevcech, Roman Burger
ICCCI1
2013 Human computation: Image metadata acquisition based on a single-player annotation game
Jakub Simko, Michal Tvarozek, Mária Bieliková
Int. J. Hum. Comput. Stud.1
2011 Semantics Discovery via Human Computation Games
abstract
The effective acquisition of (semantic) metadata is crucial for many present day applications. Games with a purpose address this issue by transforming computational problems into computer games. The authors present a novel approach to metadata acquisition via Little Search Game (LSG) – a competitive web search game, whose purpose is the creation of a term relationship network. From a player perspective, the goal is to reduce the number of search results returned for a given search term by adding negative search terms to a query. The authors describe specific aspects of the game’s design, including player motivation and anti-cheating issues. The authors have performed a series of experiments with Little Search Game, acquired real-world player input, gathered qualitative feedback from the players, constructed and evaluated term relationship network from the game logs and examined the types of created relationships.
Jakub Simko, Michal Tvarozek, Mária Bieliková
Int. J. Semantic Web Inf. Syst.1