VLDB 2026 Research / reviewers in the wild / expert
Ivan Srba
dblp:06/9076
· DBLP profile ↗
10ranked-venue papers in the field
6as first author
6since 2021 · last 2026
0000-0003-3511-5337ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7 (4 first)Data Mining & Knowledge Discovery · 3 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language ModelsabstractIn the age of social media and generative AI, the ability to automatically assess the credibility of online content has become increasingly critical, complementing traditional approaches to false information detection. Credibility assessment relies on aggregating diverse credibility signals—small units of information, such as content subjectivity, bias or a presence of persuasion techniques—into a final credibility label/score. However, current research in automatic credibility assessment and credibility signals detection remains highly fragmented, with many signals studied in isolation and lacking integration. Notably, there is a scarcity of approaches that detect and aggregate multiple credibility signals simultaneously. These challenges are further exacerbated by the absence of a comprehensive and up-to-date overview of research works that connects these research efforts under a common framework and identifies shared trends, challenges and open problems. In this survey, we address this gap by presenting a systematic and comprehensive literature review of 175 research papers, focusing on textual credibility signals within the field of Natural Language Processing (NLP), which undergoes a rapid transformation due to advancements in Large Language Models (LLMs). While positioning the NLP research into the broader multidisciplinary landscape, we examine both automatic credibility assessment methods as well as the detection of nine categories of credibility signals. We provide an in-depth analysis of three key categories: (1) factuality, subjectivity and bias, (2) persuasion techniques and logical fallacies and (3) check-worthy and fact-checked claims. In addition to summarising existing methods, datasets and tools, we outline future research direction and emerging opportunities, with particular attention to evolving challenges posed by generative AI. Ivan Srba, Olesya Razuvayevskaya, João Augusto Leite, Róbert Móro, Ipek Baris Schlicht, Sara Tonelli, Francisco Moreno García, Santiago Barrio Lottmann, Denis Teyssou, Valentin Porcellini, Carolina Scarton, Kalina Bontcheva, Mária Bieliková |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Task Prompt Vectors: Effective Initialization Through Multi-task Soft Prompt Transfer
Róbert Belanec, Simon Ostermann 0002, Ivan Srba, Mária Bieliková |
ECML/PKDD (8) | 3 |
| 2025 | Revisiting Algorithmic Audits of TikTok: Poor Reproducibility and Short-term Validity of FindingsabstractSocial media platforms are constantly shifting towards algorithmically curated content based on implicit or explicit user feedback. Regulators, as well as researchers, are calling for systematic social media algorithmic audits as this shift leads to enclosing users in filter bubbles and leading them to more problematic content. An important aspect of such audits is the reproducibility and generalisability of their findings, as it allows to draw verifiable conclusions and audit potential changes in algorithms over time. In this work, we study the reproducibility of the existing sockpuppeting audits of TikTok recommender systems, and the generalizability of their findings. In our efforts to reproduce the previous works, we find multiple challenges stemming from social media platform changes and content evolution, but also the research works themselves. These drawbacks limit the audit reproducibility and require an extensive effort altogether with inevitable adjustments to the auditing methodology. Our experiments also reveal that these one-shot audit findings often hold only in the short term, implying that the reproducibility and generalizability of the audits heavily depend on the methodological choices and the state of algorithms and content on the platform. This highlights the importance of reproducible audits that allow us to determine how the situation changes in time. Matej Mosnar, Adam Skurla, Branislav Pecher, Matus Tibensky, Jan Jakubcik, Adrian Bindas, Peter Sakalik, Ivan Srba |
SIGIR | 8 |
| 2023 | Auditing YouTube's Recommendation Algorithm for Misinformation Filter BubblesabstractIn this article, we present results of an auditing study performed over YouTube aimed at investigating how fast a user can get into a misinformation filter bubble, but also what it takes to “burst the bubble,” i.e., revert the bubble enclosure. We employ a sock puppet audit methodology, in which pre-programmed agents (acting as YouTube users) delve into misinformation filter bubbles by watching misinformation-promoting content. Then they try to burst the bubbles and reach more balanced recommendations by watching misinformation-debunking content. We record search results, home page results, and recommendations for the watched videos. Overall, we recorded 17,405 unique videos, out of which we manually annotated 2,914 for the presence of misinformation. The labeled data was used to train a machine learning model classifying videos into three classes (promoting, debunking, neutral) with the accuracy of 0.82. We use the trained model to classify the remaining videos that would not be feasible to annotate manually. Using both the manually and automatically annotated data, we observe the misinformation bubble dynamics for a range of audited topics. Our key finding is that even though filter bubbles do not appear in some situations, when they do, it is possible to burst them by watching misinformation-debunking content (albeit it manifests differently from topic to topic). We also observe a sudden decrease of misinformation filter bubble effect when misinformation-debunking videos are watched after misinformation-promoting videos, suggesting a strong contextuality of recommendations. Finally, when comparing our results with a previous similar study, we do not observe significant improvements in the overall quantity of recommended misinformation content. Ivan Srba, Róbert Móro, Matús Tomlein, Branislav Pecher, Jakub Simko, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Adrian Gavornik, Mária Bieliková |
Trans. Recomm. Syst. | 1 |
| 2022 | Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked ClaimsabstractFalse information has a significant negative influence on individuals as well as on the whole society. Especially in the current COVID-19 era, we witness an unprecedented growth of medical misinformation. To help tackle this problem with machine learning approaches, we are publishing a feature-rich dataset of approx. 317k medical news articles/blogs and 3.5k fact-checked claims. It also contains 573 manually and more than 51k automatically labelled mappings between claims and articles. Mappings consist of claim presence, i.e., whether a claim is contained in a given article, and article stance towards the claim. We provide several baselines for these two tasks and evaluate them on the manually labelled part of the dataset. The dataset enables a number of additional tasks related to medical misinformation, such as misinformation characterisation studies or studies of misinformation diffusion between sources. Ivan Srba, Branislav Pecher, Matús Tomlein, Róbert Móro, Elena Stefancova, Jakub Simko, Mária Bieliková |
SIGIR | 1 |
| 2021 | An Audit of Misinformation Filter Bubbles on YouTube: Bubble Bursting and Recent Behavior ChangesabstractThe negative effects of misinformation filter bubbles in adaptive systems have been known to researchers for some time. Several studies investigated, most prominently on YouTube, how fast a user can get into a misinformation filter bubble simply by selecting “wrong choices” from the items offered. Yet, no studies so far have investigated what it takes to “burst the bubble”, i.e., revert the bubble enclosure. We present a study in which pre-programmed agents (acting as YouTube users) delve into misinformation filter bubbles by watching misinformation promoting content (for various topics). Then, by watching misinformation debunking content, the agents try to burst the bubbles and reach more balanced recommendation mixes. We recorded the search results and recommendations, which the agents encountered, and analyzed them for the presence of misinformation. Our key finding is that bursting of a filter bubble is possible, albeit it manifests differently from topic to topic. Moreover, we observe that filter bubbles do not truly appear in some situations. We also draw a direct comparison with a previous study. Sadly, we did not find much improvements in misinformation occurrences, despite recent pledges by YouTube. Matús Tomlein, Branislav Pecher, Jakub Simko, Ivan Srba, Róbert Móro, Elena Stefancova, Michal Kompan, Andrea Hrckova, Juraj Podrouzek, Mária Bieliková |
RecSys | 4 |
| 2017 | Educational Question Routing in Online Student CommunitiesabstractStudents' performance in Massive Open Online Courses (MOOCs) is enhanced by high quality discussion forums or recently emerging educational Community Question Answering (CQA) systems. Nevertheless, only a small number of students answer questions asked by their peers. This results in instructor overload, and many unanswered questions. To increase students' participation, we present an approach for recommendation of new questions to students who are likely to provide answers. Existing approaches to such question routing proposed for non-educational CQA systems tend to rely on a few experts, what is not applicable in educational domain where it is important to involve all kinds of students. In tackling this novel educational question routing problem, our method (1) goes beyond previous question-answering data as it incorporates additional non-QA data from the course (to improve prediction accuracy and to involve more of the student community) and (2) applies constraints on users' workload (to prevent user overloading). We use an ensemble classifier for predicting students' willingness to answer a question, as well as students' expertise for answering. We conducted an online evaluation of the proposed method using an A/B experiment in our CQA system deployed in edX MOOC. The proposed method outperformed a baseline method (non-educational question routing enhanced with workload restriction) by improving recommendation accuracy, keeping more community members active, and increasing an average number of their contributions. Jakub Macina, Ivan Srba, Joseph Jay Williams, Mária Bieliková |
RecSys | 2 |
| 2016 | Design of CQA Systems for Flexible and Scalable Deployment and Evaluation
Ivan Srba, Mária Bieliková |
ICWE | 1 |
| 2016 | A Comprehensive Survey and Classification of Approaches for Community Question AnsweringabstractCommunity question-answering (CQA) systems, such as Yahoo! Answers or Stack Overflow, belong to a prominent group of successful and popular Web 2.0 applications, which are used every day by millions of users to find an answer on complex, subjective, or context-dependent questions. In order to obtain answers effectively, CQA systems should optimally harness collective intelligence of the whole online community, which will be impossible without appropriate collaboration support provided by information technologies. Therefore, CQA became an interesting and promising subject of research in computer science and now we can gather the results of 10 years of research. Nevertheless, in spite of the increasing number of publications emerging each year, so far the research on CQA systems has missed a comprehensive state-of-the-art survey. We attempt to fill this gap by a review of 265 articles published between 2005 and 2014, which were selected from major conferences and journals. According to this evaluation, at first we propose a framework that defines descriptive attributes of CQA approaches. Second, we introduce a classification of all approaches with respect to problems they are aimed to solve. The classification is consequently employed in a review of a significant number of representative approaches, which are described by means of attributes from the descriptive framework. As a part of the survey, we also depict the current trends as well as highlight the areas that require further attention from the research community. Ivan Srba, Mária Bieliková |
ACM Trans. Web | 1 |
| 2015 | Utilizing Non-QA Data to Improve Questions Routing for Users with Low QA Activity in CQAabstractCommunity Question Answering (CQA) systems, such as Yahoo! Answers and Stack Overflow, represent a well-known example of collective intelligence. The existing CQA systems, despite their overall successfulness and popularity, fail to answer a significant amount of questions in required time. One option for scaffolding collaboration in CQA systems is a recommendation of new questions to users who are suitable candidates for providing correct answers (so called question routing). Various methods have been proposed so far to find appropriate answerers, but almost all approaches heavily depend on previous users' activities in a particular CQA system (i.e. QA-data). In our work, we attempt to involve a whole community including users with no or minimal previous activity (e.g. newcomers or lurkers). We proposed a question routing method which analyses users' non-QA data from a CQA system itself as well as from external services and platforms, such as blogs, microblogs or social networking sites, in order to estimate users' interests and expertise early and more precisely. Consequently, we can recommend new questions to a wider part of a community as well as more accurately. Evaluation on a dataset from Stack Exchange platform showed that considering non-QA data leads not only to better recognition of users with low activity as suitable answerers, but also to higher overall precision of the recommendations. It implies that non-QA data can supplement QA data during expertise estimation in question routing and thus also improve a success rate of a questions answering process. Ivan Srba, Marek Grznar, Mária Bieliková |
ASONAM | 1 |