VLDB 2026 Research / reviewers in the wild / expert
Miriam Redi
dblp:85/9997
· DBLP profile ↗
22ranked-venue papers in the field
7as first author
7since 2021 · last 2026
0000-0002-0581-0251ORCID · reported
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 21 (7 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual Reference Need Assessment System for WikipediaabstractWikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In Wikipedia, content needs to be verifiable, meaning that readers can check that claims are backed by references to reliable sources. This depends on manual verification by editors, an effective but labor-intensive process, especially given the high volume of daily edits. To address this challenge, we introduce a multilingual machine learning system to assist editors in identifying claims requiring citations. Our approach is tested in 10 language editions of Wikipedia, outperforming existing benchmarks for reference need assessment. We not only consider machine learning evaluation metrics but also system requirements, allowing us to explore the trade-offs between model accuracy and computational efficiency under real-world infrastructure constraints. We deploy our system in production and release data and code to support further research. Aitolkyn Baigutanova, Francisco Navas, Pablo Aragón, Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Miriam Redi, Diego Sáez-Trumper |
WWW | 7 |
| 2023 | A Comparative Study of Reference Reliability in Multiple Language Editions of WikipediaabstractInformation presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to assess the reliability of references in multiple language editions. We quantify the cross-lingual patterns of the perennial sources list, a collection of reliability labels for web domains identified and collaboratively agreed upon by Wikipedia editors. We discover that some sources (or web domains) deemed untrustworthy in one language (i.e., English) continue to appear in articles in other languages. This trend is especially evident with sources tailored for smaller communities. Furthermore, non-authoritative sources found in the English version of a page tend to persist in other language versions of that page. We finally present a case study on the Chinese, Russian, and Swedish Wikipedias to demonstrate a discrepancy in reference reliability across cultures. Our finding highlights future challenges in coordinating global knowledge on source reliability. Aitolkyn Baigutanova, Diego Sáez-Trumper, Miriam Redi, Meeyoung Cha, Pablo Aragón |
CIKM | 3 |
| 2023 | AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content CreationabstractThis paper presents the AToMiC (Authoring Tools for Multi media Content) dataset, designed to advance research in image/text cross-modal retrieval. While vision--language pretrained transformers have led to significant improvements in retrieval effectiveness, existing research has relied on image-caption datasets that feature only simplistic image--text relationships and underspecified user models of retrieval tasks. To address the gap between these oversimplified settings and real-world applications for multimedia content creation, we introduce a new approach for building retrieval test collections. We leverage hierarchical structures and diverse domains of texts, styles, and types of images, as well as large-scale image--document associations embedded in Wikipedia. We formulate two tasks based on a realistic user model and validate our dataset through retrieval experiments using baseline models. AToMiC offers a testbed for scalable, diverse, and reproducible multimedia retrieval research. Finally, our dataset provides the basis for a dedicated track at the 2023 Text Retrieval Conference (TREC), and is publicly available at https://github.com/TREC-AToMiC/AToMiC. Jheng-Hong Yang, Carlos Eduardo Rosar Kós Lassance, Rafael S. Rezende, Krishna Srinivasan, Miriam Redi, Stéphane Clinchant, Jimmy Lin |
SIGIR | 5 |
| 2023 | Longitudinal Assessment of Reference Quality on WikipediaabstractWikipedia plays a crucial role in the integrity of the Web. This work analyzes the reliability of this global encyclopedia through the lens of its references. We operationalize the notion of reference quality by defining reference need (RN), i.e., the percentage of sentences missing a citation, and reference risk (RR), i.e., the proportion of non-authoritative references. We release Citation Detective, a tool for automatically calculating the RN score, and discover that the RN score has dropped by 20 percent point in the last decade, with more than half of verifiable statements now accompanying references. The RR score has remained below 1% over the years as a result of the efforts of the community to eliminate unreliable references. We propose pairing novice and experienced editors on the same Wikipedia article as a strategy to enhance reference quality. Our quasi-experiment indicates that such a co-editing experience can result in a lasting advantage in identifying unreliable sources in future edits. As Wikipedia is frequently used as the ground truth for numerous Web applications, our findings and suggestions on its reliability can have a far-reaching impact. We discuss the possibility of other Web services adopting Wiki-style user collaboration to eliminate unreliable content. Aitolkyn Baigutanova, Jaehyeon Myung, Diego Sáez-Trumper, Ai-Jou Chou, Miriam Redi, Changwook Jung, Meeyoung Cha |
WWW | 5 |
| 2022 | Visual Gender Biases in Wikipedia: A Systematic Evaluation across the Ten Most Spoken Languages
Pablo Beytía, Pushkal Agarwal, Miriam Redi, Vivek K. Singh 0001 |
ICWSM | 3 |
| 2021 | Wiki-Reliability: A Large Scale Dataset for Content Reliability on WikipediaabstractWikipedia is the largest online encyclopedia, used by algorithms and web users as a central hub of reliable information on the web.The quality and reliability of Wikipedia content is maintained by a community of volunteer editors. Machine learning and information retrieval algorithms could help scale up editors' manual efforts around Wikipedia content reliability. However, there is a lack of large-scale data to support the development of such research. To fill this gap, in this paper, we propose Wiki-Reliability, the first dataset of English Wikipedia articles annotated with a wide set of content reliability issues. To build this dataset, we rely on Wikipedia "templates". Templates are tags used by expert Wikipedia editors to indicate content issues, such as the presence of "non-neutral point of view" or "contradictory articles", and serve as a strong signal for detecting reliability issues in a revision. We select the 10 most popular reliability-related templates on Wikipedia, and propose an effective method to label almost 1M samples of Wikipedia article revisions as positive or negative with respect to each template. Each positive/negative example in the dataset comes with the full article text and 20 features from the revision's metadata. We provide an overview of the possible downstream tasks enabled by such data, and show that Wiki-Reliability can be used to train large-scale models for content reliability prediction. We release all data and code for public use. KayYen Wong, Miriam Redi, Diego Sáez-Trumper |
SIGIR | 2 |
| 2021 | On the Value of Wikipedia as a Gateway to the WebabstractBy linking to external websites, Wikipedia can act as a gateway to the Web. To date, however, little is known about the amount of traffic generated by Wikipedia’s external links. We fill this gap in a detailed analysis of usage logs gathered from Wikipedia users’ client devices. Our analysis proceeds in three steps: First, we quantify the level of engagement with external links, finding that, in one month, English Wikipedia generated 43M clicks to external websites, in roughly even parts via links in infoboxes, cited references, and article bodies. Official links listed in infoboxes have by far the highest click-through rate (CTR), 2.47% on average. In particular, official links associated with articles about businesses, educational institutions, and websites have the highest CTR, whereas official links associated with articles about geographical content, television, and music have the lowest CTR. Second, we investigate patterns of engagement with external links, finding that Wikipedia frequently serves as a stepping stone between search engines and third-party websites, effectively fulfilling information needs that search engines do not meet. Third, we quantify the hypothetical economic value of the clicks received by external websites from English Wikipedia, by estimating that the respective website owners would need to pay a total of $7–13 million per month to obtain the same volume of traffic via sponsored search. Overall, these findings shed light on Wikipedia’s role not only as an important source of information, but also as a high-traffic gateway to the broader Web ecosystem. Tiziano Piccardi, Miriam Redi, Giovanni Colavizza, Robert West 0001 |
WWW | 2 |
| 2020 | Quantifying Engagement with Citations on WikipediaabstractWikipedia is one of the most visited sites on the Web and a common source of information for many users. As an encyclopedia, Wikipedia was not conceived as a source of original information, but as a gateway to secondary sources: according to Wikipedia’s guidelines, facts must be backed up by reliable sources that reflect the full spectrum of views on the topic. Although citations lie at the heart of Wikipedia, little is known about how users interact with them. To close this gap, we built client-side instrumentation for logging all interactions with links leading from English Wikipedia articles to cited references during one month, and conducted the first analysis of readers’ interactions with citations. We find that overall engagement with citations is low: about one in 300 page views results in a reference click (0.29% overall; 0.56% on desktop; 0.13% on mobile). Matched observational studies of the factors associated with reference clicking reveal that clicks occur more frequently on shorter pages and on pages of lower quality, suggesting that references are consulted more commonly when Wikipedia itself does not contain the information sought by the user. Moreover, we observe that recent content, open access sources, and references about life events (births, deaths, marriages, etc.) are particularly popular. Taken together, our findings deepen our understanding of Wikipedia’s role in a global information economy where reliability is ever less certain, and source attribution ever more vital. Tiziano Piccardi, Miriam Redi, Giovanni Colavizza, Robert West 0001 |
WWW | 2 |
| 2019 | Citation Needed: A Taxonomy and Algorithmic Assessment of Wikipedia's VerifiabilityabstractWikipedia is playing an increasingly central role on the web, and the policies its contributors follow when sourcing and fact-checking content affect million of readers. Among these core guiding principles, verifiability policies have a particularly important role. Verifiability requires that information included in a Wikipedia article be corroborated against reliable secondary sources. Because of the manual labor needed to curate Wikipedia at scale, however, its contents do not always evenly comply with these policies. Citations (i.e. reference to external sources) may not conform to verifiability requirements or may be missing altogether, potentially weakening the reliability of specific topic areas of the free encyclopedia. In this paper, we aim to provide an empirical characterization of the reasons why and how Wikipedia cites external sources to comply with its own verifiability guidelines. First, we construct a taxonomy of reasons why inline citations are required, by collecting labeled data from editors of multiple Wikipedia language editions. We then crowdsource a large-scale dataset of Wikipedia sentences annotated with categories derived from this taxonomy. Finally, we design algorithmic models to determine if a statement requires a citation, and to predict the citation reason . We evaluate the accuracy of such models across different classes of Wikipedia articles of varying quality, and on external datasets of claims annotated for fact-checking purposes. Miriam Redi, Besnik Fetahu, Jonathan T. Morgan, Dario Taraborelli |
WWW | 1 |
| 2018 | Online Petitioning Through Data Exploration and What We Found There: A Dataset of Petitions from Avaaz.org
Pablo Aragón, Diego Sáez-Trumper, Miriam Redi, Scott A. Hale, Vicenç Gómez, Andreas Kaltenbrunner |
ICWSM | 3 |
| 2017 | Bridging the Aesthetic Gap: The Wild Beauty of Web ImageryabstractTo provide good results, image search engines need to rank not just the most relevant images, but also the highest quality images. To surface beautiful pictures, existing computational aesthetic models are trained with datasets from photo contest websites, dominated by professional photos. Such models fail completely in real web scenarios, where images are extremely diverse in terms of quality and type (e.g. drawings, clip-art, etc). This work aims at bridging and understanding this "aesthetic gap". We collect a dataset of around 100K web images with `quality' and `type' (photo vs non-photo) annotations. We design a set of visual features to describe image pictorial characteristics, and deeply analyse the peculiar beauty of web images as opposed to appealing professional images. Finally, we build a set of computational aesthetic frameworks based on deep learning and hand-crafted features that take into account the diverse quality of web images, and show that they significantly outperform traditional computational aesthetics methods on our dataset. Miriam Redi, Frank Z. Liu, Neil O'Hare |
ICMR | 1 |
| 2017 | Beautiful and Damned. Combined Effect of Content Quality and Social Ties on User EngagementabstractUser participation in online communities is driven by the intertwinement of the social network structure with the crowd-generated content that flows along its links. These aspects are rarely explored jointly and at scale. By looking at how users generate and access pictures of varying beauty on Flickr, we investigate how the production of quality impacts the dynamics of online social systems. We develop a deep learning computer vision model to score images according to their aesthetic value and we validate its output through crowdsourcing. By applying it to over 15 B Flickr photos, we study for the first time how image beauty is distributed over a large-scale social system. Beautiful images are evenly distributed in the network, although only a small core of people get social recognition for them. To study the impact of exposure to quality on user engagement, we set up matching experiments aimed at detecting causality from observational data. Exposure to beauty is double-edged: following people who produce high-quality content increases one's probability of uploading better photos; however, an excessive imbalance between the quality generated by a user and the user's neighbors leads to a decline in engagement. Our analysis has practical implications for improving link recommender systems. Luca Maria Aiello, Rossano Schifanella, Miriam Redi, Stacey Svetlichnaya, Frank Z. Liu, Simon Osindero |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from VideosabstractThumbnails play such an important role in online videos. As the most representative snapshot, they capture the essence of a video and provide the first impression to the viewers; ultimately, a great thumbnail makes a video more attractive to click and watch. We present an automatic thumbnail selection system that exploits two important characteristics commonly associated with meaningful and attractive thumbnails: high relevance to video content and superior visual aesthetic quality. Our system selects attractive thumbnails by analyzing various visual quality and aesthetic metrics of video frames, and performs a clustering analysis to determine the relevance to video content, thus making the resulting thumbnails more representative of the video. On the task of predicting thumbnails chosen by professional video editors, we demonstrate the effectiveness of our system against six baseline methods, using a real-world dataset of 1,118 videos collected from Yahoo Screen. In addition, we study what makes a frame a good thumbnail by analyzing the statistical relationship between thumbnail frames and non-thumbnail frames in terms of various image quality features. Our study suggests that the selection of a good thumbnail is highly correlated with objective visual quality metrics, such as the frame texture and sharpness, implying the possibility of building an automatic thumbnail selection system based on visual aesthetics. Yale Song, Miriam Redi, Jordi Vallmitjana, Alejandro Jaimes |
CIKM | 2 |
| 2016 | Complura: Exploring and Leveraging a Large-scale Multilingual Visual Sentiment OntologyabstractWhat would someone from another culture think of this photograph I just took? Would they think my picture of this "wilted flower" was also sentimentally positive or would they perceive it negatively instead? Or what if I wanted to find other photographs that are semantically related to my image as well as sentimentally sensitive, but from other cultures? In fact, this cultural and sentimental relevancy are features that we would expect of any recommender system and query expansion engine, respectively. Motivated by this, we present an online demonstration of a system called Complura. Our system implements three major functions: an interactive multilingual ontology browser, a cross-lingual image-based sentiment analyzer, and a culturally-coherent, sentiment-aware image query expansion engine. We ground our system on a multilingual visual sentiment ontology, containing over 10k sentiment-polarized visual concepts over 12 languages and over 7.3M images. Hongyi Liu 0004, Brendan Jou, Tao Chen 0015, Mercan Topkara, Nikolaos Pappas 0002, Miriam Redi, Shih-Fu Chang |
ICMR | 6 |
| 2016 | Multilingual Visual Sentiment Concept MatchingabstractThe impact of culture in visual emotion perception has recently captured the attention of multimedia research. In this study, we provide powerful computational linguistics tools to explore, retrieve and browse a dataset of 16K multilingual affective visual concepts and 7.3M Flickr images. First, we design an effective crowdsourcing experiment to collect human judgements of sentiment connected to the visual concepts. We then use word embeddings to represent these concepts in a low dimensional vector space, allowing us to expand the meaning around concepts, and thus enabling insight about commonalities and differences among different languages. We compare a variety of concept representations through a novel evaluation task based on the notion of visual semantic relatedness. Based on these representations, we design clustering schemes to group multilingual visual concepts, and evaluate them with novel metrics based on the crowdsourced sentiment annotations as well as visual semantic relatedness. The proposed clustering framework enables us to analyze the full multilingual dataset in-depth and also show an application on a facial data subset, exploring cultural insights of portrait-related affective visual concepts. Nikolaos Pappas 0002, Miriam Redi, Mercan Topkara, Brendan Jou, Hongyi Liu 0004, Tao Chen 0015, Shih-Fu Chang |
ICMR | 2 |
| 2016 | Predicting Pre-click Quality for Native AdvertisementsabstractNative advertising is a specific form of online advertising where ads replicate the look-and-feel of their serving platform. In such context, providing a good user experience with the served ads is crucial to ensure long-term user engagement. In this work, we explore the notion of ad quality, namely the effectiveness of advertising from a user experience perspective. We design a learning framework to predict the pre-click quality of native ads. More specifically, we look at detecting offensive native ads, showing that, to quantify ad quality, ad offensive user feedback rates are more reliable than the commonly used click-through rate metrics. We then conduct a crowd-sourcing study to identify which criteria drive user preferences in native advertising. We translate these criteria into a set of ad quality features that we extract from the ad text, image and advertiser, and then use them to train a model able to identify offensive ads. We show that our model is very effective in detecting offensive ads, and provide in-depth insights on how different features affect ad quality. Finally, we deploy a preliminary version of such model and show its effectiveness in the reduction of the offensive ad feedback rate. Ke Zhou 0003, Miriam Redi, Andrew Haines, Mounia Lalmas-Roelleke |
WWW | 2 |
| 2015 | Like Partying? Your Face Says It All. Predicting the Ambiance of Places with Profile Pictures
Miriam Redi, Daniele Quercia, Lindsay T. Graham, Samuel D. Gosling |
ICWSM | 1 |
| 2015 | An Image Is Worth More than a Thousand Favorites: Surfacing the Hidden Beauty of Flickr Pictures
Rossano Schifanella, Miriam Redi, Luca Maria Aiello |
ICWSM | 2 |
| 2013 | Semantic indexing and computational aesthetics: interactions, bridgesand boundariesabstractSemantic Indexing and Computational Aesthetics are two closely related fields. For some aspects they are similar, complementary for others, and sometimes completely disjoint. Semantic Indexing is about automatically identifying content in natural images, namely recognizing objects and scenes. Computational Aesthetics provides a set of techniques to automatically assign a beauty degree to a given image. In our work, we enrich both types of visual analysis by exploring the synergy of those two fields. We investigate the role of Semantic Indexing techniques for Computational Aesthetics Frameworks, and, vice versa, the importance of Aesthetic features for Semantic Indexing prediction. We show the benefits and the limits of this synergy, and propose some improvements in this direction. Miriam Redi |
ICMR | 1 |
| 2013 | Direct modeling of image keypoints distribution through copula-based image signaturesabstractLocal Image Descriptors (LID) aggregation models such as Bag of Words and Fisher Vectors represent an image based on the distribution of its LIDs given a global model, e.g. a visual codebook or a Gaussian Mixture. Miriam Redi, Bernard Mérialdo |
ICMR | 1 |
| 2012 | Exploring two spaces with one feature: kernelized multidimensional modeling of visual alphabetsabstractMarginal Alphabets (MEDA) were proposed as an alternative to Bag of Words (BoW) for image representation. They aggregate sets of locally extracted descriptors (LEDs) by using visual alphabets based on the marginal approximation of the LED components. Compared to the exponential complexity of the BoW codebooks, the MEDA model is very efficient because each dimension of the LED is quantized independently. However, MEDA lacks of considering the relations between the LED components, loosing precious information for image representation. Miriam Redi, Bernard Mérialdo |
ICMR | 1 |
| 2011 | Saliency moments for image categorizationabstractIn this paper we present Saliency Moments, a new, holistic descriptor for image recognition inspired by two biological vision principles: the gist perception and the selective visual attention. While traditional image features extract either local or global discriminative properties from the visual content, we use a hybrid approach that exploits some coarsely localized information, i.e. the salient regions shape and contours, to build a global, low-dimensional image signature. Results show that this new type of image description outperforms the traditional global features on scene and object categorization, for a variety of challenging datasets. Moreover, we show that, when combined with other existing descriptors (SIFT, Color Moments, Wavelet Feature and Edge Histogram), the saliency-based features provide complementary information, improving the precision of a retrieval system we build for the TRECVID 2010. Miriam Redi, Bernard Mérialdo |
ICMR | 1 |