VLDB 2026 Research / reviewers in the wild / expert
Pablo Aragón
dblp:95/9656
· DBLP profile ↗
11ranked-venue papers in the field
5as first author
5since 2021 · last 2026
0000-0002-6017-4577ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 9 (4 first)Data Mining & Knowledge Discovery · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual Reference Need Assessment System for WikipediaabstractWikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content needs to be verifiable, meaning that readers can check that claims are backed by references to reliable sources. This depends on manual verification by editors, an effective but labor-intensive process, especially given the high volume of daily edits. To address this challenge, we introduce a multilingual machine learning system to assist editors in identifying claims requiring citations. Our approach is tested in 10 language editions of Wikipedia, outperforming existing benchmarks for reference need assessment. We not only consider machine learning evaluation metrics but also system requirements, allowing us to explore the trade-offs between model accuracy and computational efficiency under real-world infrastructure constraints. We deploy our system in production and release data and code to support further research. Aitolkyn Baigutanova, Francisco Navas, Pablo Aragón, Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Miriam Redi, Diego Sáez-Trumper |
WWW | 3 |
| 2026 | Language-Agnostic Modeling of Source Reliability on WikipediaabstractOver the last few years, verifying the credibility of information sources has become a fundamental need to combat disinformation. Here, we present a language-agnostic model designed to assess the reliability of web domains as sources in references across multiple language editions of Wikipedia. Utilizing editing activity data, the model evaluates domain reliability within different articles of varying controversiality, such as Climate Change, COVID-19, History, Media, and Biology topics. Crafting features that express domain usage across articles, the model effectively predicts domain reliability, achieving an F1 Macro score of approximately 0.80 for English and other high-resource languages. For mid-resource languages, we achieve 0.65, while the performance of low-resource languages varies. In all cases, the time the domain remains present in the articles (which we dub as permanence ) is one of the most predictive features. We highlight the challenge of maintaining consistent model performance across languages of varying resource levels and demonstrate that adapting models from higher-resource languages can improve performance. We believe these findings can assist Wikipedia editors in their ongoing efforts to verify citations and may offer useful insights for other user-generated content communities. Jacopo D'Ignazi, Andreas Kaltenbrunner, Yelena Mejova, Michele Tizzani, Kyriaki Kalimeri, Mariano G. Beiró, Pablo Aragón |
ACM Trans. Web | 7 |
| 2025 | Characterizing Knowledge Manipulation in a Russian Wikipedia ForkabstractWikipedia is powered by MediaWiki, a free and open-source software that is also the infrastructure for many other wiki-based online encyclopedias. These include the recently launched website Ruwiki, which has copied and modified the original Russian Wikipedia content to conform to Russian law. To identify practices and narratives that could be associated with different forms of knowledge manipulation, this article presents an in-depth analysis of this Russian Wikipedia fork. We propose a methodology to characterize the main changes with respect to the original version. The foundation of this study is a comprehensive comparative analysis of more than 1.9M articles from Russian Wikipedia and its fork. Using meta-information and geographical, temporal, categorical, and textual features, we explore the changes made by Ruwiki editors. Furthermore, we present a classification of the main topics of knowledge manipulation in this fork, including a numerical estimation of their scope. This research not only sheds light on significant changes within Ruwiki, but also provides a methodology that could be applied to analyze other Wikipedia forks and similar collaborative projects. Mykola Trokhymovych, Oleksandr Kosovan, Nathan Forrester, Pablo Aragón, Diego Sáez-Trumper, Ricardo Baeza-Yates |
ICWSM | 4 |
| 2024 | Language-Agnostic Modeling of Wikipedia Articles for Content Quality Assessment across LanguagesabstractWikipedia is the largest web repository of free knowledge. Volunteer editors devote time and effort to creating and expanding articles in more than 300 language editions. As content quality varies from article to article, editors also spend substantial time rating articles with specific criteria. However, keeping these assessments complete and up-to-date is largely impossible given the ever-changing nature of Wikipedia. To overcome this limitation, we propose a novel computational framework for modeling the quality of Wikipedia articles. State-of-the-art approaches to model Wikipedia article quality have leveraged machine learning techniques with language-specific features. In contrast, our framework is based on language-agnostic structural features extracted from the articles, a set of universal weights, and a language version-specific normalization criterion. Therefore, we ensure that all language editions of Wikipedia can benefit from our framework, even those that do not have their own quality assessment scheme. Using this framework, we have built datasets with the feature values and quality scores of all revisions of all articles in the existing language versions of Wikipedia. We provide a descriptive analysis of these resources and a benchmark of our framework. In addition, we discuss possible downstream tasks to be addressed with these datasets, which are released for public use. Paramita Das, Isaac L. Johnson, Diego Sáez-Trumper, Pablo Aragón |
ICWSM | 4 |
| 2023 | A Comparative Study of Reference Reliability in Multiple Language Editions of WikipediaabstractInformation presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language editions. We quantify the cross-lingual patterns of the perennial sources list, a collection of reliability labels for web domains identified and collaboratively agreed upon by Wikipedia editors. We discover that some sources (or web domains) deemed untrustworthy in one language (i.e., English) continue to appear in articles in other languages. This trend is especially evident with sources tailored for smaller communities. Furthermore, non-authoritative sources found in the English version of a page tend to persist in other language versions of that page. We finally present a case study on the Chinese, Russian, and Swedish Wikipedias to demonstrate a discrepancy in reference reliability across cultures. Our finding highlights future challenges in coordinating global knowledge on source reliability. Aitolkyn Baigutanova, Diego Sáez-Trumper, Miriam Redi, Meeyoung Cha, Pablo Aragón |
CIKM | 5 |
| 2018 | Interactive Discovery System for Direct DemocracyabstractDecide Madrid is the civic technology of Madrid City Council which allows users to create and support online petitions. Despite the initial success, the platform is encountering problems with the growth of petition signing because petitions are far from the minimum number of supporting votes they must gather. Previous analyses have suggested that this problem is produced by the interface: a paginated list of petitions which applies a non-optimal ranking algorithm. For this reason, we present an interactive system for the discovery of topics and petitions. This approach leads us to reflect on the usefulness of data visualization techniques to address relevant societal challenges. Pablo Aragón, Yago Bermejo, Vicenç Gómez, Andreas Kaltenbrunner |
ASONAM | 1 |
| 2018 | Big Data-Driven Platform for Cross-Media MonitoringabstractThe abundance of online media content requires highly scalable architectures to allow cross-media monitoring. This paper presents an innovative big data-as-a-service platform for analysing large complex networks in order to enhance cross-media monitoring. In contrast to the existing media monitoring systems, the platform equips marketers with several distinctive features. First, while most of the systems perform quantitative exploratory analysis of social media, our platform applies graph analytics in order to reveal social interaction types, hidden patterns in the cross-media network and the information diffusion over time. Second, our platform integrates and implements distributed versions of graph analytics algorithms (Louvain, HITS and others) that can scale to a large volume of data. Third, the creation of cross-media graphs is triggered by user-defined queries that can be easily specified by marketers. Thus, end-users can build and analyse different graphs according to specific goals of the study. Finally, the platform allows reducing Hadoop cluster usage costs due to executing the graph mining algorithms on demand triggered by user-defined queries. Instead of running costly streaming processes that continuously listen for new queries, we implemented Spark-as-a-service approach via Apache Livy REST interface. Liana Napalkova, Pablo Aragón, Juan Carlos Castro Robles |
DSAA | 2 |
| 2018 | Online Petitioning Through Data Exploration and What We Found There: A Dataset of Petitions from Avaaz.org
Pablo Aragón, Diego Sáez-Trumper, Miriam Redi, Scott A. Hale, Vicenç Gómez, Andreas Kaltenbrunner |
ICWSM | 1 |
| 2017 | To Thread or Not to Thread: The Impact of Conversation Threading on Online Discussion
Pablo Aragón, Vicenç Gómez, Andreas Kaltenbrunner |
ICWSM | 1 |
| 2016 | Visualization Tool for Collective Awareness in a Platform of Citizen Proposals
Pablo Aragón, Vicenç Gómez, Andreas Kaltenbrunner |
ICWSM | 1 |
| 2016 | When a Movement Becomes a Party: Computational Assessment of New Forms of Political Organization in Social Media
Pablo Aragón, Yana Volkovich, David Laniado, Andreas Kaltenbrunner |
ICWSM | 1 |