Daniel Kluver

dblp:118/7181 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2810-7243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Why They Come And Go: A Case Study of Productive Flyby Users and Their Rating Integrity Challenge in Movie Recommenders
abstract
We present a case study of productive flyby users (PFB users) on a recommendation website.These users exhibit counterintuitive behavior: they input a large amount of data during their first visit but never return.This phenomenon can have both positive and negative impacts on the system.On the positive side, their high productivity contributes a substantial amount of data.On the negative side, they may input inappropriate ratings that violate the assumptions of recommendation algorithms, potentially undermining system performance.To better understand the nature and causes of this behavior, we investigated their motivations, expectations, reasons for leaving, and the potential risks associated with their ratings using a mixed-methods approach.Specifically, we conducted interviews with 11 users, surveyed 41 users, and analyzed the impact of 1,000 PFB users on the performance of recommendation algorithms for regular users.Our findings revealed diverse motivations among PFB users.Some engaged with the system merely to pass the time, while others had unrealistic expectations of the recommender system.Regarding rating quality, 27% of surveyed users admitted to rating movies they had not seen, citing reasons such as browsing too quickly or attempting to manipulate the algorithm.Notably, users who reported leaving because they were "just killing time and forgot about the website" were the most likely to rate unseen movies.Overall, PFB users significantly influence recommendation algorithms and their performance for regular users.While some subgroups negatively affect prediction accuracy, others provide
Ruixuan Sun, Ruoyan Kong, Ashlee Milton, Daniel Kluver, Ian Paterson, Joseph A. Konstan
CHIIR4
2024 The MovieLens Beliefs Dataset: Collecting Pre-Choice Data for Online Recommender Systems
abstract
An increasingly important aspect of designing recommender systems involves considering how recommendations will influence consumer choices. This paper addresses this issue by introducing a method for collecting user beliefs about un-experienced goods – a critical predictor of choice behavior. We implemented this method on the MovieLens platform, resulting in a rich dataset that combines user ratings, beliefs, and observed recommendations. We document challenges to such data collection, including selection bias in response and limited coverage of the product space. This unique resource empowers researchers to delve deeper into user behavior and analyze user choices absent recommendations, measure the effectiveness of recommendations, and prototype algorithms that leverage user belief data, ultimately leading to more impactful recommender systems. The dataset can be found at https://grouplens.org/datasets/movielens/ml_belief_2024/.
Guy Aridor, Duarte Gonçalves, Ruoyan Kong, Daniel Kluver, Joseph A. Konstan
RecSys4
2023 The Economics of Recommender Systems: Evidence from a Field Experiment on MovieLens
abstract
We conduct a 6 month field experiment on a movie-recommendation platform to identify if and how recommendation systems affect consumption. We use within-consumer randomization at the good level and elicit beliefs about unconsumed goods to disentangle exposure from informational effects. We have three experimental groups: (a) control, (b) exposed, and (c) recommended + exposed goods where only goods in (c) are recommended and we elicit beliefs about goods in (b) and (c). Comparing across these treatment arms we find recommendations increase consumption beyond its role in exposing goods to consumers. We provide support for an informational mechanism: recommendations affect consumers' beliefs, which in turn explain consumption. Recommendations reduce uncertainty about goods consumers are most uncertain about and induce information acquisition. Finally, we find evidence for spatial correlation in beliefs.
Guy Aridor, Duarte Gonçalves, Daniel Kluver, Ruoyan Kong, Joseph A. Konstan
EC3
2021 Effective Strategies for Crowd-Powered Cognitive Reappraisal Systems: A Field Deployment of the Flip*Doubt Web Application for Mental Health
abstract
Online technologies offer great promise to expand models of delivery for therapeutic interventions to help users cope with increasingly common mental illnesses like anxiety and depression. For example, "cognitive reappraisal" is a skill that involves changing one's perspective on negative thoughts in order to improve one's emotional state. In this work, we present Flip*Doubt, a novel crowd-powered web application that provides users with cognitive reappraisals ("reframes") of negative thoughts. A one-month field deployment of Flip*Doubt with 13 graduate students yielded a data set of negative thoughts paired with positive reframes, as well as rich interview data about how participants interacted with the system. Through this deployment, our work contributes: (1) an in-depth qualitative understanding of how participants used a crowd-powered cognitive reappraisal system in the wild; and (2) detailed codebooks that capture informative context about negative input thoughts and reframes. Our results surface data-derived hypotheses that may help to explain what types of reframes are helpful for users, while also providing guidance to future researchers and developers interested in building collaborative systems for mental health. In our discussion, we outline implications for systems research to leverage peer training and support, as well as opportunities to integrate AI/ML-based algorithms to support the cognitive reappraisal task. (Note: This paper includes potentially triggering mentions of mental health issues and suicide.)
C. Estelle Smith, William Lane 0003, Hannah Miller Hillberg, Daniel Kluver, Loren G. Terveen, Svetlana Yarosh
Proc. ACM Hum. Comput. Interact.4
2021 Exploring author gender in book rating and recommendation
Michael D. Ekstrand, Daniel Kluver
User Model. User Adapt. Interact.2
2018 Exploring author gender in book rating and recommendation
abstract
Collaborative filtering algorithms find useful patterns in rating and consumption data and exploit these patterns to guide users to good items. Many of the patterns in rating datasets reflect important real-world differences between the various users and items in the data; other patterns may be irrelevant or possibly undesirable for social or ethical reasons, particularly if they reflect undesired discrimination, such as gender or ethnic discrimination in publishing. In this work, we examine the response of collaborative filtering recommender algorithms to the distribution of their input data with respect to a dimension of social concern, namely content creator gender. Using publicly-available book ratings data, we measure the distribution of the genders of the authors of books in user rating profiles and recommendation lists produced from this data. We find that common collaborative filtering algorithms differ in the gender distribution of their recommendation lists, and in the relationship of that output distribution to user profile distribution.
Michael D. Ekstrand, Mucun Tian, Mohammed R. Imran Kazi, Hoda Mehrpouyan, Daniel Kluver
RecSys5
2018 What I See is What You Don't Get: The Effects of (Not) Seeing Emoji Rendering Differences across Platforms
abstract
Emoji are popular in digital communication, but they are rendered differently on different viewing platforms (e.g., iOS, Android). It is unknown how many people are aware that emoji have multiple renderings, or whether they would change their emoji-bearing messages if they could see how these messages render on recipients' devices. We developed software to expose the multi-rendering nature of emoji and explored whether this increased visibility would affect how people communicate with emoji. Through a survey of 710 Twitter users who recently posted an emoji-bearing tweet, we found that at least 25% of respondents were unaware that the emoji they posted could appear differently to their followers. Additionally, after being shown how one of their tweets rendered across platforms, 20% of respondents reported that they would have edited or not sent the tweet. These statistics reflect millions of potentially regretful tweets shared per day because people cannot see emoji rendering differences across platforms. Our results motivate the development of tools that increase the visibility of emoji rendering differences across platforms, and we contribute our cross-platform emoji rendering software to facilitate this effort.
Hannah Miller Hillberg, Zachary Levonian, Daniel Kluver, Loren G. Terveen, Brent J. Hecht
Proc. ACM Hum. Comput. Interact.3
2017 Understanding Emoji Ambiguity in Context: The Role of Text in Emoji-Related Miscommunication
Hannah Miller Hillberg, Daniel Kluver, Jacob Thebault-Spieker, Loren G. Terveen, Brent J. Hecht
ICWSM2
2017 Simulation Experiments on (the Absence of) Ratings Bias in Reputation Systems
abstract
As the gig economy continues to grow and freelance work moves online, five-star reputation systems are becoming more and more common. At the same time, there are increasing accounts of race and gender bias in evaluations of gig workers, with negative impacts for those workers. We report on a series of four Mechanical Turk-based studies in which participants who rated simulated gig work did not show race- or gender bias, while manipulation checks showed they reliably distinguished between low- and high-quality work. Given prior research, this was a striking result. To explore further, we used a Bayesian approach to verify absence of ratings bias (as opposed to merely not detecting bias). This Bayesian test let us identify an upper- bound: if any bias did exist in our studies, it was below an average of 0.2 stars on a five-star scale. We discuss possible interpretations of our results and outline future work to better understand the results.
Jacob Thebault-Spieker, Daniel Kluver, Maximilian A. Klein, Aaron Halfaker, Brent J. Hecht, Loren G. Terveen, Joseph A. Konstan
Proc. ACM Hum. Comput. Interact.2
2015 Letting Users Choose Recommender Algorithms: An Experimental Study
abstract
Recommender systems are not one-size-fits-all; different algorithms and data sources have different strengths, making them a better or worse fit for different users and use cases. As one way of taking advantage of the relative merits of different algorithms, we gave users the ability to change the algorithm providing their movie recommendations and studied how they make use of this power. We conducted our study with the launch of a new version of the MovieLens movie recommender that supports multiple recommender algorithms and allows users to choose the algorithm they want to provide their recommendations. We examine log data from user interactions with this new feature to under-stand whether and how users switch among recommender algorithms, and select a final algorithm to use. We also look at the properties of the algorithms as they were experienced by users and examine their relationships to user behavior. We found that a substantial portion of our user base (25%) used the recommender-switching feature. The majority of users who used the control only switched algorithms a few times, trying a few out and settling down on an algorithm that they would leave alone. The largest number of users prefer a matrix factorization algorithm, followed closely by item-item collaborative filtering; users selected both of these algorithms much more often than they chose a non-personalized mean recommender. The algorithms did produce measurably different recommender lists for the users in the study, but these differences were not directly predictive of user choice.
Michael D. Ekstrand, Daniel Kluver, F. Maxwell Harper, Joseph A. Konstan
RecSys2
2015 User Session Identification Based on Strong Regularities in Inter-activity Time
abstract
Session identification is a common strategy used to develop metrics for web analytics and perform behavioral analyses of user-facing systems. Past work has argued that session identification strategies based on an inactivity threshold is inherently arbitrary or has advocated that thresholds be set at about 30 minutes. In this work, we demonstrate a strong regularity in the temporal rhythms of user initiated events across several different domains of online activity (incl. video gaming, search, page views and volunteer contributions). We describe a methodology for identifying clusters of user activity and argue that the regularity with which these activity clusters appear implies a good rule-of-thumb inactivity threshold of about 1 hour. We conclude with implications that these temporal rhythms may have for system design based on our observations and theories of goal-directed human activity.
Aaron Halfaker, Oliver Keyes, Daniel Kluver, Jacob Thebault-Spieker, Tien T. Nguyen, Kenneth Shores, Anuradha Uduwage, Morten Warncke-Wang
WWW3
2014 Evaluating recommender behavior for new users
abstract
The new user experience is one of the important problems in recommender systems. Past work on recommending for new users has focused on the process of gathering information from the user. Our work focuses on how different algorithms behave for new users. We describe a methodology that we use to compare representatives of three common families of algorithms along eleven different metrics. We find that for the first few ratings a baseline algorithm performs better than three common collaborative filtering algorithms. Once we have a few ratings, we find that Funk's SVD algorithm has the best overall performance. We also find that ItemItem, a very commonly deployed algorithm, performs very poorly for new users. Our results can inform the design of interfaces and algorithms for new users.
Daniel Kluver, Joseph A. Konstan
RecSys1
2013 Rating support interfaces to improve user experience and recommender accuracy
abstract
One of the challenges for recommender systems is that users struggle to accurately map their internal preferences to external measures of quality such as ratings. We study two methods for supporting the mapping process: (i) reminding the user of characteristics of items by providing personalized tags and (ii) relating rating decisions to prior rating decisions using exemplars. In our study, we introduce interfaces that provide these methods of support. We also present a set of methodologies to evaluate the efficacy of the new interfaces via a user experiment. Our results suggest that presenting exemplars during the rating process helps users rate more consistently, and increases the quality of the data.
Tien T. Nguyen, Daniel Kluver, Ting-Yu Wang, Pik-Mai Hui, Michael D. Ekstrand, Martijn C. Willemsen, John Riedl
RecSys2
2012 How many bits per rating?
abstract
Most recommender systems assume user ratings accurately represent user preferences. However, prior research shows that user ratings are imperfect and noisy. Moreover, this noise limits the measurable predictive power of any recommender system. We propose an information theoretic framework for quantifying the preference information contained in ratings and predictions. We computationally explore the properties of our model and apply our framework to estimate the efficiency of different rating scales for real world datasets. We then estimate how the amount of information predictions give to users is related to the scale ratings are collected on. Our findings suggest a tradeoff in rating scale granularity: while previous research indicates that coarse scales (such as thumbs up / thumbs down) take less time, we find that ratings with these scales provide less predictive value to users. We introduce a new measure, preference bits per second, to quantitatively reconcile this tradeoff.
Daniel Kluver, Tien T. Nguyen, Michael D. Ekstrand, Shilad Sen, John Riedl
RecSys1