VLDB 2026 Research / reviewers in the wild / expert
Jahna Otterbacher
dblp:40/3725
· DBLP profile ↗
24ranked-venue papers in the field
10as first author
7since 2021 · last 2025
0000-0002-7655-7118ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 21 (7 first)Data Mining & Knowledge Discovery · 3 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KeepA(n)I: Social Stereotypes in and Social Norms for Computer VisionabstractThe KeepA(n)I platform facilitates the auditing of computer vision systems that tag images, which aid visual communication on the Web and social media, from content moderation to the development of new apps and tools. In particular, KeepA(n)I enables a broad set of stakeholders to scrutinize a process of interest that embeds an image tagger for issues of social stereotyping, while also examining the social norms that humans apply to the observed AI behaviors. KeepA(n)I’s approach, and its use of the power of the crowd, can aid the stakeholders in receiving responses to both descriptive and normative questions (i.e., which stereotyping behaviors are observed and if they are considered problematic by a given “crowd” for an intended context). We provide an overview of the platform, its key features, and a discussion via a use case on the diverse set of stakeholders that can benefit from it. Evgenia Christoforou, Nicolas C. Nicolaou, Efstathios Stavrakis, Jahna Otterbacher |
ICWSM | 4 |
| 2024 | Subjectivity, Polarity and the Aspect of Time in the Evolution of Crowd-Sourced Biographies
Constantinos Romantzis, Alexandros Karakasidis 0001, Evangelos Mathioudis, Ioannis Katakis 0001, Pantelis Agathangelou, Jahna Otterbacher |
ICWE | 6 |
| 2024 | Generative AI in Crowdwork for Web and Social Media Research: A Survey of Workers at Three PlatformsabstractCrowdsourcing plays an important role in Web and social media research, from data annotation, to online experiments and user surveys. With the emergence of Generative AI (GenAI), researchers are considering how models and tools such as GPT might replace crowdwork. Many have already evaluated GPT on annotation tasks. However, it is less clear how GenAI might impact other types of tasks, or to what extent crowdworkers have already incorporated it into their work processes. Thus, we asked crowdworkers directly regarding their use of GenAI, via a survey at two points in time, across three commercial platforms. We found evidence that workers' self-reported use of GenAI did not change over time, but rather, was strongly correlated to the platform in which they operate, with MTurk workers using GenAI much more often than those operating at Clickworker and Prolific. As most respondents reported that survey completion is their "usual type of task", we discuss the implication of the use of GenAI in user surveys, via specific examples of ICWSM research. Evgenia Christoforou, Gianluca Demartini, Jahna Otterbacher |
ICWSM | 3 |
| 2023 | Just Another Day on Twitter: A Complete 24 Hours of Twitter DataabstractAt the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change. Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter |
ICWSM | 12 |
| 2022 | How Does the Crowd Impact the Model? A Tool for Raising Awareness of Social Bias in Crowdsourced Training DataabstractIt is increasingly easy for interested parties to play a role in the development of predictive algorithms, with a range of available tools and platforms for building datasets, as well as for training and evaluating machine learning (ML) models. For this reason, it is essential to create awareness among practitioners on the ethical challenges, such as the presence of social bias in training data. We present RECANT (Raising Awareness of Social Bias in Crowdsourced Training Data), a tool that allows users to explore the behaviors of four biometric models -- predicting the gender and race, as well as the perceived attractiveness and trustworthiness, of the person depicted in an input image. These models have been trained on a crowdsourced dataset of passport-style people images, where crowd annotators described attributes of the images, and reported their own demographic characteristics. With RECANT, users can explore the correct and wrong predictions made by each model, when using different subsets of the data in training, based on annotator attributes. We present its features, along with sample exercises, as a hands-on tool for raising awareness of potential pitfalls in data practices surrounding ML. Periklis Perikleous, Andreas Kafkalias, Zenonas Theodosiou, Pinar Barlas, Evgenia Christoforou, Jahna Otterbacher, Gianluca Demartini, Andreas Lanitis |
CIKM | 6 |
| 2022 | Shifting Our Awareness, Taking Back Tags: Temporal Changes in Computer Vision Services' Social Behaviors
Pinar Barlas, Maximilian Krahn, Styliani Kleanthous, Kyriakos Kyriakou, Jahna Otterbacher |
ICWSM | 5 |
| 2021 | Do you see what I see? Images of the COVID-19 pandemic through the lens of GoogleabstractDuring times of crisis, information access is crucial. Given the opaque processes behind modern search engines, it is important to understand the extent to which the "picture" of the Covid-19 pandemic accessed by users differs. We explore variations in what users "see" concerning the pandemic through Google image search, using a two-step approach. First, we crowdsource a search task to users in four regions of Europe, asking them to help us create a photo documentary of Covid-19 by providing image search queries. Analysing the queries, we find five common themes describing information needs. Next, we study three sources of variation - users' information needs, their geo-locations and query languages - and analyse their influences on the similarity of results. We find that users see the pandemic differently depending on where they live, as evidenced by the 46% similarity across results. When users expressed a given query in different languages, there was no overlap for most of the results. Our analysis suggests that localisation plays a major role in the (dis)similarity of results, and provides evidence of the diverse "picture" of the pandemic seen through Google. Monica Lestari Paramita, Kalia Orphanou, Evgenia Christoforou, Jahna Otterbacher, Frank Hopfgartner |
Inf. Process. Manag. | 4 |
| 2019 | How Do We Talk about Other People? Group (Un)Fairness in Natural Language Image DescriptionsabstractCrowdsourcing plays a key role in developing algorithms for image recognition or captioning. Major datasets, such as MS COCO or Flickr30K, have been built by eliciting natural language descriptions of images from workers. Yet such elicitation tasks are susceptible to human biases, including stereotyping people depicted in images. Given the growing concerns surrounding discrimination in algorithms, as well as in the data used to train them, it is necessary to take a critical look at this practice. We conduct experiments at Figure Eight using a controlled set of people images. Men and women of various races are positioned in the same manner, wearing a grey t-shirt. We prompt workers for 10 descriptive labels, and consider them using the human-centric approach, which assumes reporting bias. We find that “what’s worth saying” about these uniform images often differs as a function of the gender and race of the depicted person, violating the notion of group fairness. Although this diversity in natural language people descriptions is expected and often beneficial, it could result in automated disparate impact if not managed properly. Jahna Otterbacher, Pinar Barlas, Styliani Kleanthous, Kyriakos Kyriakou |
HCOMP | 1 |
| 2019 | Social B(eye)as: Human and Machine Descriptions of People Images
Pinar Barlas, Kyriakos Kyriakou, Styliani Kleanthous, Jahna Otterbacher |
ICWSM | 4 |
| 2019 | Fairness in Proprietary Image Tagging Algorithms: A Cross-Platform Audit on People Images
Kyriakos Kyriakou, Pinar Barlas, Styliani Kleanthous, Jahna Otterbacher |
ICWSM | 4 |
| 2018 | Social Cues, Social Biases: Stereotypes in Annotations on People ImagesabstractHuman computation is often subject to systematic biases. We consider the case of linguistic biases and their consequences for the words that crowdworkers use to describe people images in an annotation task. Social psychologists explain that when describing oth- ers, the subconscious perpetuation of stereotypes is in- evitable, as we describe stereotype-congruent people and/or in-group members more abstractly than others. In an MTurk experiment we show evidence of these bi- ases, which are exacerbated when an image’s “popular tags” are displayed, a common feature used to provide social information to workers. Underscoring recent calls for a deeper examination of the role of training data quality in algorithmic biases, results suggest that it is rather easy to sway human judgment. Jahna Otterbacher |
HCOMP | 1 |
| 2018 | Investigating User Perception of Gender Bias in Image Search: The Role of SexismabstractThere is growing evidence that search engines produce results that are socially biased, reinforcing a view of the world that aligns with prevalent social stereotypes. One means to promote greater transparency of search algorithms - which are typically complex and proprietary - is to raise user awareness of biased result sets. However, to date, little is known concerning how users perceive bias in search results, and the degree to which their perceptions differ and/or might be predicted based on user attributes. One particular area of search that has recently gained attention, and forms the focus of this study, is image retrieval and gender bias. We conduct a controlled experiment via crowdsourcing using participants recruited from three countries to measure the extent to which workers perceive a given image results set to be subjective or objective. Demographic information about the workers, along with measures of sexism, are gathered and analysed to investigate whether (gender) biases in the image search results can be detected. Amongst other findings, the results confirm that sexist people are less likely to detect and report gender biases in image search results. Jahna Otterbacher, Alessandro Checco, Gianluca Demartini, Paul D. Clough |
SIGIR | 1 |
| 2017 | Headlines Matter: Using Headlines to Predict the Popularity of News Articles on Twitter and Facebook
Alicja Piotrkowicz, Vania Dimitrova, Jahna Otterbacher, Katja Markert |
ICWSM | 3 |
| 2015 | Linguistic Bias in Collaboratively Produced Biographies: Crowdsourcing Social Stereotypes?
Jahna Otterbacher |
ICWSM | 1 |
| 2014 | Write Like I Write: Herding in the Language of Online Reviews
Loizos Michael, Jahna Otterbacher |
ICWSM | 2 |
| 2013 | Gender, writing and ranking in review forums: a case study of the IMDb
Jahna Otterbacher |
Knowl. Inf. Syst. | 1 |
| 2011 | Evolutionary timeline summarization: a balanced optimization framework via iterative substitutionabstractClassic news summarization plays an important role with the exponential document growth on the Web. Many approaches are proposed to generate summaries but seldom simultaneously consider evolutionary characteristics of news plus to traditional summary elements. Therefore, we present a novel framework for the web mining problem named Evolutionary Timeline Summarization (ETS). Given the massive collection of time-stamped web documents related to a general news query, ETS aims to return the evolution trajectory along the timeline, consisting of individual but correlated summaries of each date, emphasizing relevance, coverage, coherence and cross-date diversity. ETS greatly facilitates fast news browsing and knowledge comprehension and hence is a necessity. We formally formulate the task as an optimization problem via iterative substitution from a set of sentences to a subset of sentences that satisfies the above requirements, balancing coherence/diversity measurement and local/global summary quality. The optimized substitution is iteratively conducted by incorporating several constraints until convergence. We develop experimental systems to evaluate on 6 instinctively different datasets which amount to 10251 documents. Performance comparisons between different system-generated timelines and manually created ones by human editors demonstrate the effectiveness of our proposed framework in terms of ROUGE metrics. Rui Yan 0001, Xiaojun Wan 0001, Jahna Otterbacher, Liang Kong 0001, Xiaoming Li 0001, Yan Zhang 0004 |
SIGIR | 3 |
| 2010 | Inferring gender of movie reviewers: exploiting writing style, content and metadataabstractDespite differences in the way that men and women experience goods and communicate their perspectives, online review communities typically do not provide participants' gender. We propose to infer author gender, given a set of reviews of a particular item, and experiment on reviews posted at the Internet Movie Database (IMDb). Using logistic regression, we explore the contribution of three types of information: 1) style, 2) content, and 3) metadata (e.g. review age, social feedback). Our results concur with previous research, in that there are salient differences in writing style and content between reviews authored by men versus women. However, in comparison to literary or scientific texts, to which classification tasks are often applied, reviews are brief and occur within the context of an ongoing discourse. Therefore, to compensative for the brevity of reviews, content and stylistic features can be augmented with metadata. We find in particular that the perceived utility of a review is an important correlate of gender. The model incorporating all features has a classification accuracy of 73.7% and is not as sensitive to review length as are those based only on stylistic or content features. Jahna Otterbacher |
CIKM | 1 |
| 2009 | Biased LexRank: Passage retrieval using random walks with question-based priors
Jahna Otterbacher, Günes Erkan, Dragomir R. Radev |
Inf. Process. Manag. | 1 |
| 2008 | Hierarchical summarization for delivering information to mobile devices
Jahna Otterbacher, Dragomir R. Radev, Omer Kareem |
Inf. Process. Manag. | 1 |
| 2006 | Fact-focused novelty detection: a feasibility studyabstractMethods for detecting sentences in an input document set, which are both relevant and novel with respect to an information need, would be of direct benefit to many systems, such as extractive text summarizers. However, satisfactory levels of agreement between judges performing this task manually have yet to demonstrated, leaving researchers to conclude that the task is too subjective. In previous experiments, judges were asked to first identify sentences that are relevant to a general topic, and then to eliminate sentences from the list that do not contain new information. Currently, a new task is proposed, in which annotators perform the same procedure, but within the context of a specific, factual information need. In the experiment, satisfactory levels of agreement between independent annotators were achieved on the first step of identifying sentences containing relevant information relevant. However, the results indicate that judges do not agree on which sentences contain novel information. Jahna Otterbacher, Dragomir R. Radev |
SIGIR | 1 |
| 2006 | News to go: hierarchical text summarization for mobile devicesabstractWe present an evaluation of a novel hierarchical text summarization method that allows users to view summaries of Web documents from small, mobile devices. Unlike previous approaches, ours does not require the documents to be in HTML since it infers a hierarchical structure automatically. Currently, the method is used to summarize news articles sent to a Web mail account in plain text format. Subjects used a Web-enabled mobile phone emulator to access the account's inbox and view the summarized news articles. They then used the summaries to complete several information-seeking tasks, which involved answering factual questions about the stories. In comparing the hierarchical text summary setting to that in which subjects were given the full text articles, there was no significant difference in task accuracy or the time taken to complete the task. However, in the hierarchical summarization setting, the number of bytes transferred per user request is less than half that of the full text case. Finally, in comparing the new method to three other summarization methods, subjects achieved significantly better accuracy on the tasks when using hierarchical summaries. Jahna Otterbacher, Dragomir R. Radev, Omer Kareem |
SIGIR | 1 |
| 2005 | Hierarchical text summarization for WAP-enabled mobile devicesabstractWe present WAP MEAD, a WAP-enabled text summarization system. It incorporates a state-of-the art text summarizer enhanced to produce hierarchical summaries that are appropriate for various types of mobile devices, including cellular phones. Dragomir R. Radev, Omer Kareem, Jahna Otterbacher |
SIGIR | 3 |
| 2003 | Learning cross-document structural relationships using boostingabstractMulti-document discoure analysis has emerged with the potential of improving various information retrieval applications. Based on the newly proposed Cross-document Structure Theory (CST), this paper describes an empirical study that uses boosting to classify CST relationships between sentence pairs extracted from topically related documents. We show that the binary classifier for determining existence of structural relationships significantly outperforms the baseline. We also achieve promising results on the multi-class case in which the full taxonomy of relationships are considered. Jahna Otterbacher, Dragomir R. Radev |
CIKM | 2 |