EDBT 2026 Demo / reviewers in the wild / expert
Michael D. Ekstrand
dblp:69/7949
· DBLP profile ↗
46ranked-venue papers in the field
22as first author
23since 2021 · last 2026
0000-0003-2467-0108ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 44 (22 first)Database Systems & Data Management · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Justice, Emancipation, Democracy, and Information Access (JEDI): The SIGIR Workshop on Resisting Corporate and Authoritarian Capture of Information Access Platforms
Bhaskar Mitra 0001, Dana McKay, Michael D. Ekstrand, Sanne Vrijenhoek, Maria Murray |
SIGIR | 3 |
| 2026 | Recall, Robustness, and Lexicographic EvaluationabstractAlthough originally developed to evaluate sets of items, recall is often used to evaluate rankings of items, including those produced by recommender, retrieval, and other machine learning systems. The application of recall without a formal evaluative motivation has led to criticism of recall as a vague or inappropriate measure. In light of this debate, we reflect on the measurement of recall in rankings from a formal perspective. Our analysis is composed of three tenets: recall, robustness, and lexicographic evaluation. First, we formally define “recall orientation” as the sensitivity of a metric to a user interested in finding every relevant item. Second, we analyze recall orientation from the perspective of robustness with respect to possible content consumers and providers, connecting recall to recent conversations about fair ranking. Finally, we extend this conceptual and theoretical treatment of recall by developing a practical preference-based evaluation method based on lexicographic comparison. Through extensive empirical analysis across multiple recommendation and retrieval tasks, we establish that our new evaluation method, lexirecall, has convergent validity (i.e., it is correlated with existing recall metrics) and exhibits substantially higher sensitivity in terms of discriminative power and stability in the presence of missing labels. Our conceptual, theoretical, and empirical analysis substantially deepens our understanding of recall and motivates its adoption through connections to robustness and fairness. Fernando Diaz 0001, Michael D. Ekstrand, Bhaskar Mitra 0001 |
Trans. Recomm. Syst. | 2 |
| 2025 | Fairness in Information Access Conceptual Foundations and New Directions
Michael D. Ekstrand |
ECIR (5) | 1 |
| 2025 | FAccTRec 2025: The 8th Workshop on Responsible RecommendationabstractThe 8th Workshop on Responsible Recommendation (FAccTRec 2025) was held in conjunction with the 19th ACM Conference on Recommender Systems in September, 2025 at Prague, Czech Republic, in a hybrid format.This workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns.It served to advance research and discussion of these topics in the recommender systems space, and incubate ideas for future development and refinement.For 2025, we highlight (1) the increasing importance of pre-trained models in recommendation; and (2) shifting regulatory, organizational, and political landscapes. Michael D. Ekstrand, Toshihiro Kamishima, Amifa Raj, Karlijn Dinnissen |
RecSys | 1 |
| 2024 | Not Just Algorithms: Strategically Addressing Consumer Impacts in Information Retrieval
Michael D. Ekstrand, Lex Beattie, Maria Soledad Pera, Henriette Cramer |
ECIR (4) | 1 |
| 2024 | Multiple Testing for IR and Recommendation System Experiments
Ngozi Ihemelandu, Michael D. Ekstrand |
ECIR (3) | 2 |
| 2024 | Towards Optimizing Ranking in Grid-Layout for Provider-Side Fairness
Amifa Raj, Michael D. Ekstrand |
ECIR (5) | 2 |
| 2024 | Conducting Recommender Systems User Studies Using POPROXabstractThe Platform for OPen Recommendation and Online eXperimentation (POPROX) is a new resource to allow RecSys researchers to conduct online user research without having to develop all of the necessary infrastructure and recruit users. Our first domain is personalized news recommendations – POPROX 1.0 provides a daily newsletter (with content from the Associated Press) to users who have already consented to participate in research, along with interfaces and protocols to support researchers in conducting studies that assign subsets of users to various experimental algorithms and/or interfaces. Robin D. Burke, Joseph A. Konstan, Michael D. Ekstrand |
RecSys | 3 |
| 2024 | FAccTRec 2024: The 7th Workshop on Responsible RecommendationabstractThe 7th Workshop on Responsible Recommendation (FAccTRec 2024) was held in conjunction with the 18th ACM Conference on Recommender Systems on October, 2024 at Bari, Italy, in a hybrid format. This workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. It served to advance research and discussion of these topics in the recommender systems space, and incubate ideas for future development and refinement. For 2024, the workshop highlights i) the possible tensions between the factors related to social responsibility, and ii) the challenges as a result of AI-related regulations in the European Union. Michael D. Ekstrand, Toshihiro Kamishima, Amifa Raj, Karlijn Dinnissen |
RecSys | 1 |
| 2024 | AltRecSys: A Workshop on Alternative, Unexpected, and Critical Ideas in RecommendationabstractThe AltRecsys workshop, held in conjunction with the 18th edition of the ACM Conference on Recommender Systems (RecSys) in Bari, Italy, provides a platform for highlighting “alternative” work in recommender systems. Modeled after alt.chi and the CRAFT sessions at the FAccT conference, AltRecSys offers a space to discuss interesting, preliminary, offbeat, unexpected, and critical ideas in recommender systems that do not (yet) fit well into the kinds of publications and formats for the main conference or traditional workshops. This workshop is not a venue to showcase research advances. Instead, it is envisioned as a forum where researchers, (industry) practitioners, and other associated stakeholders can exchange ideas and together identify areas of study and new questions to expand the discussions and research agendas of the RecSys community in future years. The call for contributions and the workshop sessions are centered around the question “what are the vital questions, needs, or opportunities that the RecSys community is currently overlooking?” Michael D. Ekstrand, Maria Soledad Pera, Alan Said |
RecSys | 1 |
| 2024 | It's Not You, It's Me: The Impact of Choice Models and Ranking Strategies on Gender Imbalance in Music RecommendationabstractAs recommender systems are prone to various biases, mitigation approaches are needed to ensure that recommendations are fair to various stakeholders. One particular concern in music recommendation is artist gender fairness. Recent work has shown that the gender imbalance in the sector translates to the output of music recommender systems, creating a feedback loop that can reinforce gender biases over time. Andres Ferraro, Michael D. Ekstrand, Christine Bauer 0001 |
RecSys | 2 |
| 2024 | Distributionally-Informed Recommender System EvaluationabstractCurrent practice for evaluating recommender systems typically focuses on point estimates of user-oriented effectiveness metrics or business metrics, sometimes combined with additional metrics for considerations such as diversity and novelty. In this article, we argue for the need for researchers and practitioners to attend more closely to various distributions that arise from a recommender system (or other information access system) and the sources of uncertainty that lead to these distributions. One immediate implication of our argument is that both researchers and practitioners must report and examine more thoroughly the distribution of utility between and within different stakeholder groups. However, distributions of various forms arise in many more aspects of the recommender systems experimental process, and distributional thinking has substantial ramifications for how we design, evaluate, and present recommender systems evaluation and research results. Leveraging and emphasizing distributions in the evaluation of recommender systems is a necessary step to ensure that the systems provide appropriate and equitably distributed benefit to the people they affect. Michael D. Ekstrand, Ben Carterette, Fernando Diaz 0001 |
Trans. Recomm. Syst. | 1 |
| 2024 | Building Human Values into Recommender Systems: An Interdisciplinary SynthesisabstractRecommender systems are the algorithms which select, filter, and personalize content across many of the world's largest platforms and apps. As such, their positive and negative effects on individuals and on societies have been extensively theorized and studied. Our overarching question is how to ensure that recommender systems enact the values of the individuals and societies that they serve. Addressing this question in a principled fashion requires technical knowledge of recommender design and operation, and also critically depends on insights from diverse fields including social science, ethics, economics, psychology, policy, and law. This article is a multidisciplinary effort to synthesize theory and practice from different perspectives, with the goal of providing a shared language, articulating current design approaches, and identifying open problems. We collect a set of values that seem most relevant to recommender systems operating across different domains, and then examine them from the perspectives of current industry practice, measurement, product design, and policy approaches. Important open problems include multi-stakeholder processes for defining values and resolving trade-offs, better values-driven measurements, recommender controls that people use, non-behavioral algorithmic feedback, optimization for long-term outcomes, causal inference of recommender effects, academic-industry research collaborations, and interdisciplinary policy-making. Jonathan Stray, Alon Y. Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Chloé Bakalar, Lex Beattie, Michael D. Ekstrand, Claire Leibowicz, Connie Moon Sehat, Sara Johansen, Lianne Kerlin, David Vickrey, Spandana Singh, Sanne Vrijenhoek, Amy X. Zhang, McKane Andrus, Natali Helberger, Polina Proutskova, Tanushree Mitra, Nina Vasan |
Trans. Recomm. Syst. | 9 |
| 2023 | Much Ado About Gender: Current Practices and Future Recommendations for Appropriate Gender-Aware Information AccessabstractInformation access research (and development) sometimes makes use of gender, whether to report on the demographics of participants in a user study, as inputs to personalized results or recommendations, or to make systems gender-fair, amongst other purposes. This work makes a variety of assumptions about gender, however, that are not necessarily aligned with current understandings of what gender is, how it should be encoded, and how a gender variable should be ethically used. In this work, we present a systematic review of papers on information retrieval and recommender systems that mention gender in order to document how gender is currently being used in this field. We find that most papers mentioning gender do not use an explicit gender variable, but most of those that do either focus on contextualizing results of model performance, personalizing a system based on assumptions of user gender, or auditing a model’s behavior for fairness or other privacy-related issues. Moreover, most of the papers we review rely on a binary notion of gender, even if they acknowledge that gender cannot be split into two categories. We connect these findings with scholarship on gender theory and recent work on gender in human-computer interaction and natural language processing. We conclude by making recommendations for ethical and well-grounded use of gender in building and researching information access systems. Christine Pinney, Amifa Raj, Alex Hanna, Michael D. Ekstrand |
CHIIR | 4 |
| 2023 | FAccTRec 2023: The 6th Workshop on Responsible RecommendationabstractThe 6th Workshop on Responsible Recommendation (FAccTRec 2023) was held in conjunction with the 17th ACM Conference on Recommender Systems on September, 2023 at Singapore, in a hybrid format. This workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. It served to advance research and discussion of these topics in the recommender systems space, and incubate ideas for future development and refinement. Michael D. Ekstrand, Jean Garcia-Gathright, Nasim Sonboli, Amifa Raj, Karlijn Dinnissen |
RecSys | 1 |
| 2023 | Introducing LensKit-Auto, an Experimental Automated Recommender System (AutoRecSys) ToolkitabstractLensKit is one of the first and most popular Recommender System libraries. While LensKit offers a wide variety of features, it does not include any optimization strategies or guidelines on how to select and tune LensKit algorithms. LensKit developers have to manually include third-party libraries into their experimental setup or implement optimization strategies by hand to optimize hyperparameters. We found that 63.6% (21 out of 33) of papers using LensKit algorithms for their experiments did not select algorithms or tune hyperparameters. Non-optimized models represent poor baselines and produce less meaningful research results. This demo introduces LensKit-Auto. LensKit-Auto automates the entire Recommender System pipeline and enables LensKit developers to automatically select, optimize, and ensemble LensKit algorithms. Tobias Vente, Michael D. Ekstrand, Jöran Beel |
RecSys | 2 |
| 2023 | Inference at Scale: Significance Testing for Large Search and Recommendation ExperimentsabstractA number of information retrieval studies have been done to assess which statistical techniques are appropriate for comparing systems. However, these studies are focused on TREC-style experiments, which typically have fewer than 100 topics. There is no similar line of work for large search and recommendation experiments; such studies typically have thousands of topics or users and much sparser relevance judgements, so it is not clear if recommendations for analyzing traditional TREC experiments apply to these settings. In this paper, we empirically study the behavior of significance tests with large search and recommendation evaluation data. Our results show that the Wilcoxon and Sign tests show significantly higher Type-1 error rates for large sample sizes than the bootstrap, randomization and t-tests, which were more consistent with the expected error rate. While the statistical tests displayed differences in their power for smaller sample sizes, they showed no difference in their power for large sample sizes. We recommend the sign and Wilcoxon tests should not be used to analyze large scale evaluation results. Our result demonstrate that with Top-N recommendation and large search evaluation data, most tests would have a 100% chance of finding statistically significant results. Therefore, the effect size should be used to determine practical or scientific significance. Ngozi Ihemelandu, Michael D. Ekstrand |
SIGIR | 2 |
| 2023 | Patterns of Gender-Specializing Query ReformulationabstractUsers of search systems often reformulate their queries by adding query terms to reflect their evolving information need or to more precisely express their information need when the system fails to surface relevant content. Analyzing these query reformulations can inform us about both system and user behavior. In this work, we study a special category of query reformulations that involve specifying demographic group attributes, such as gender, as part of the reformulated query (e.g., ''olympic 2021 soccer results'' -> ''olympic 2021 women's soccer results"). There are many ways a query, the search results, and a demographic attribute such as gender may relate, leading us to hypothesize different causes for these reformulation patterns, such as under-representation on the original result page or based on the linguistic theory of markedness. This paper reports on an observational study of gender-specializing query reformulations---their contexts and effects---as a lens on the relationship between system results and gender, based on large-scale search log data from Bing. We find that these reformulations sometimes correct for and other times reinforce gender representation on the original result page, but typically yield better access to the ultimately-selected results. The prevalence of these reformulations---and which gender they skew towards---differ by topical context. However, we do not find evidence that either group under-representation or markedness alone adequately explains these reformulations. We hope that future research will use such reformulations as a probe for deeper investigation into gender (and other demographic) representation on the search result page. Amifa Raj, Bhaskar Mitra 0001, Nick Craswell, Michael D. Ekstrand |
SIGIR | 4 |
| 2022 | Measuring Fairness in Ranked Results: An Analytical and Empirical ComparisonabstractInformation access systems, such as search and recommender systems, often use ranked lists to present results believed to be relevant to the user's information need. Evaluating these lists for their fairness along with other traditional metrics provides a more complete understanding of an information access system's behavior beyond accuracy or utility constructs. To measure the (un)fairness of rankings, particularly with respect to the protected group(s) of producers or providers, several metrics have been proposed in the last several years. However, an empirical and comparative analyses of these metrics showing the applicability to specific scenario or real data, conceptual similarities, and differences is still lacking. Amifa Raj, Michael D. Ekstrand |
SIGIR | 2 |
| 2021 | SimuRec: Workshop on Synthetic Data and Simulation Methods for Recommender Systems ResearchabstractThere is significant interest lately in using synthetic data and simulation infrastructures for various types of recommender systems research. However, there are not currently any clear best practices around how best to apply these methods. We proposed a workshop to bring together researchers and practitioners interested in simulating recommender systems and their data to discuss the state of the art of such research and the pressing open methodological questions. The workshop resulted in a report authored by the participants that documents currently-known best practices on which the group has consensus and lays out an agenda for further research over the next 3–5 years to fill in places where we currently lack the information needed to make methodological recommendations. Michael D. Ekstrand, Allison Chaney, Pablo Castells, Robin D. Burke, David Rohde, Manel Slokom |
RecSys | 1 |
| 2021 | FAccTRec 2021: The 4th Workshop on Responsible RecommendationabstractThe Fourth Workshop on Responsible Recommendation (FAccTRec 2021) was held in conjunction with the 15th ACM Conference on Recommender Systems on September, 2021 at Amsterdam, Netherlands, in a hybrid format. This workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. It served to advance research and discussion of these topics in the recommender systems space, and incubate ideas for future development and refinement. Michael D. Ekstrand, Pierre-Nicolas Schwab, Toshihiro Kamishima, Nasim Sonboli |
RecSys | 1 |
| 2021 | Baby Shark to Barracuda: Analyzing Children's Music Listening BehaviorabstractMusic is an important part of childhood development, with online music listening platforms being a significant channel by which children consume music. Children’s offline music listening behavior has been heavily researched, yet relatively few studies explore how their behavior manifests online. In this paper, we use data from LastFM 1 Billion and the Spotify API to explore online music listening behavior of children, ages 6–17, using education levels as lenses for our analysis. Understanding the music listening behavior of children can be used to inform the future design of recommender systems. Lawrence Spear, Ashlee Milton, Garrett Allen, Amifa Raj, Michael D. Ekstrand, Maria Soledad Pera |
RecSys | 6 |
| 2021 | Estimation of Fair Ranking Metrics with Incomplete JudgmentsabstractThere is increasing attention to evaluating the fairness of search system ranking decisions. These metrics often consider the membership of items to particular groups, often identified using protected attributes such as gender or ethnicity. To date, these metrics typically assume the availability and completeness of protected attribute labels of items. However, the protected attributes of individuals are rarely present, limiting the application of fair ranking metrics in large scale systems. In order to address this problem, we propose a sampling strategy and estimation technique for four fair ranking metrics. We formulate a robust and unbiased estimator which can operate even with very limited number of labeled items. We evaluate our approach using both simulated and real world data. Our experimental results demonstrate that our method can estimate this family of fair ranking metrics and provides a robust, reliable alternative to exhaustive or random data annotation. Ömer Kirnap, Fernando Diaz 0001, Asia J. Biega, Michael D. Ekstrand, Ben Carterette, Emine Yilmaz |
WWW | 4 |
| 2020 | Estimating Error and Bias in Offline Evaluation ResultsabstractOffline evaluations of recommender systems attempt to estimate users' satisfaction with recommendations using static data from prior user interactions. These evaluations provide researchers and developers with first approximations of the likely performance of a new system and help weed out bad ideas before presenting them to users. However, offline evaluation cannot accurately assess novel, relevant recommendations, because the most novel items were previously unknown to the user, so they are missing from the historical data and cannot be judged as relevant. We present a simulation study to estimate the error that such missing data causes in commonly-used evaluation metrics in order to assess its prevalence and impact. We find that missing data in the rating or observation process causes the evaluation protocol to systematically mis-estimate metric values, and in some cases erroneously determine that a popularity-based recommender outperforms even a perfect personalized recommender. Substantial breakthroughs in recommendation quality, therefore, will be difficult to assess with existing offline techniques. Mucun Tian, Michael D. Ekstrand |
CHIIR | 2 |
| 2020 | Evaluating Stochastic Rankings with Expected ExposureabstractWe introduce the concept of expected exposure as the average attention ranked items receive from users over repeated samples of the same query. Furthermore, we advocate for the adoption of the principle of equal expected exposure: given a fixed information need, no item should receive more or less expected exposure than any other item of the same relevance grade. We argue that this principle is desirable for many retrieval objectives and scenarios, including topical diversity and fair ranking. %Leveraging user models from existing retrieval metrics, we propose a general evaluation methodology based on expected exposure and draw connections to related metrics in information retrieval evaluation. Importantly, this methodology relaxes classic information retrieval assumptions, allowing a system, in response to a query, to produce a distribution over rankings instead of a single fixed ranking. We study the behavior of the expected exposure metric and stochastic rankers across a variety of information access conditions, including ad hoc retrieval and recommendation. %We believe that measuring and optimizing expected exposure metrics using randomization opens a new area for retrieval algorithm development and progress. Fernando Diaz 0001, Bhaskar Mitra 0001, Michael D. Ekstrand, Asia J. Biega, Ben Carterette |
CIKM | 3 |
| 2020 | LensKit for Python: Next-Generation Software for Recommender Systems ExperimentsabstractLensKit is an open-source toolkit for building, researching, and learning about recommender systems. First released in 2010 as a Java framework, it has supported diverse published research, small-scale production deployments, and education in both MOOC and traditional classroom settings. In this paper, I present the next generation of the LensKit project, re-envisioning the original tool's objectives as flexible Python package for supporting recommender systems research and development. LensKit for Python (LKPY) enables researchers and students to build robust, flexible, and reproducible experiments that make use of the large and growing PyData and Scientific Python ecosystem, including scikit-learn, and TensorFlow. To that end, it provides classical collaborative filtering implementations, recommender system evaluation metrics, data preparation routines, and tools for efficiently batch running recommendation algorithms, all usable in any combination with each other or with other Python software. Michael D. Ekstrand |
CIKM | 1 |
| 2020 | 3rd FAccTRec Workshop: Responsible RecommendationabstractThe third Workshop on Responsible Recommendation (FAccTRec 2020) was held in conjunction with the 14th ACM Conference on Recommender Systems on September 26th, 2020 as a virtual event with the conference home base in Brazil. This full-day workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. It served to advance research and discussion of these topics in the recommender systems space, and incubate ideas for future development and refinement. Michael D. Ekstrand, Pierre-Nicolas Schwab, Jean Garcia-Gathright, Toshihiro Kamishima, Nasim Sonboli |
RecSys | 1 |
| 2019 | Fairness and discrimination in recommendation and retrievalabstractFairness and related concerns have become of increasing importance in a variety of AI and machine learning contexts. They are also highly relevant to recommender systems and related problems such as information retrieval, as evidenced by the growing literature in RecSys, FAT*, SIGIR, and special sessions such as the FATREC and FACTS-IR workshops and the Fairness track at TREC 2019; however, translating algorithmic fairness constructs from classification, scoring, and even many ranking settings into recommendation and other information access scenarios is not a straightforward task. This tutorial will help orient RecSys researchers to algorithmic fairness, understand how concepts do and do not translate from other settings, and provide an introduction to the growing literature on this topic. Michael D. Ekstrand, Robin D. Burke, Fernando Diaz 0001 |
RecSys | 1 |
| 2019 | StoryTime: eliciting preferences from children for book recommendationsabstractWe present StoryTime, a book recommender for children. Our web-based recommender is co-designed with children and uses images to elicit their preferences. By building on existing solutions related to both visual interfaces and book recommendation strategies for children, StoryTime can generate suggestions without historical data or adult guidance. We discuss the benefits of StoryTime as a starting point for further research exploring the cold start problem, incorporating historical data, and needs related to children as a complex audience to enhance the recommendation process. Ashlee Milton, Adam Keener, Joshua Ames, Michael D. Ekstrand, Maria Soledad Pera |
RecSys | 5 |
| 2019 | Fairness and Discrimination in Retrieval and RecommendationabstractFairness and related concerns have become of increasing importance in a variety of AI and machine learning contexts. They are also highly relevant to information retrieval and related problems such as recommendation, as evidenced by the growing literature in SIGIR, FAT*, RecSys, and special sessions such as the FATREC workshop and the Fairness track at TREC 2019; however, translating algorithmic fairness constructs from classification, scoring, and even many ranking settings into information retrieval and recommendation scenarios is not a straightforward task. This tutorial will help to orient IR researchers to algorithmic fairness, understand how concepts do and do not translate from other settings, and provide an introduction to the growing literature on this topic. Michael D. Ekstrand, Robin D. Burke, Fernando Diaz 0001 |
SIGIR | 1 |
| 2019 | Workshop on Fairness, Accountability, Confidentiality, Transparency, and Safety in Information Retrieval (FACTS-IR)abstractThis workshop explores challenges in responsible information retrieval system development and deployment. The focus is on determining actionable research agendas on five key dimensions of responsible information retrieval: fairness, accountability, confidentiality, transparency, and safety. Rather than just a mini-conference, this workshop is an event during which participants are expected to work. The workshop brings together a diverse set of researchers and practitioners interested in contributing to the development of a technical research agenda for responsible information retrieval. Alexandra Olteanu, Jean Garcia-Gathright, Maarten de Rijke, Michael D. Ekstrand |
SIGIR | 4 |
| 2018 | Exploring author gender in book rating and recommendationabstractCollaborative filtering algorithms find useful patterns in rating and consumption data and exploit these patterns to guide users to good items. Many of the patterns in rating datasets reflect important real-world differences between the various users and items in the data; other patterns may be irrelevant or possibly undesirable for social or ethical reasons, particularly if they reflect undesired discrimination, such as gender or ethnic discrimination in publishing. In this work, we examine the response of collaborative filtering recommender algorithms to the distribution of their input data with respect to a dimension of social concern, namely content creator gender. Using publicly-available book ratings data, we measure the distribution of the genders of the authors of books in user rating profiles and recommendation lists produced from this data. We find that common collaborative filtering algorithms differ in the gender distribution of their recommendation lists, and in the relationship of that output distribution to user profile distribution. Michael D. Ekstrand, Mucun Tian, Mohammed R. Imran Kazi, Hoda Mehrpouyan, Daniel Kluver |
RecSys | 1 |
| 2018 | 2nd FATREC workshop: responsible recommendationabstractThe second Workshop on Responsible Recommendation (FATREC 2018) was held in conjunction with the 12th ACM Conference on Recommender Systems on October 6th, 2018 in Vancouver, Canada. This full-day workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. Toshihiro Kamishima, Pierre-Nicolas Schwab, Michael D. Ekstrand |
RecSys | 3 |
| 2017 | FATREC Workshop on Responsible RecommendationabstractThe first Workshop on Responsible Recommendation (FATREC) was held in conjunction with the 11th ACM Conference on Recommender Systems in August, 2017 in Como, Italy. This full-day workshop brought together researchers and practitioners to discuss several topics under the banner of social responsibility in recommender systems: fairness, accountability, transparency, privacy, and other ethical and social concerns. Michael D. Ekstrand, Amit Sharma 0007 |
RecSys | 1 |
| 2016 | Behaviorism is Not Enough: Better Recommendations through Listening to UsersabstractBehaviorism is the currently-dominant paradigm for building and evaluating recommender systems. Both the operation and the evaluation of recommender system applications are most often driven by analyzing the behavior of users. In this paper, we argue that listening to what users say about the items and recommendations they like, the control they wish to exert on the output, and the ways in which they perceive the system and not just observing what they do will enable important developments in the future of recommender systems. We provide both philosophical and pragmatic motivations for this idea, describe the various points in the recommendation and evaluation processes where explicit user input may be considered, and discuss benefits that may result from considered incorporation of user preferences at each of these points. In particular, we envision recommender applications that aim to support users' better selves: helping them live the life that they desire to lead. For example, recommender-assisted behavior change requires algorithms to predict not what users choose or do now, inferable from behavioral data, but what they should choose or do in the future to become healthier, fitter, more sustainable, or culturally aware. We hope that our work will spur useful discussion and many new ideas for recommenders that empower their users. Michael D. Ekstrand, Martijn C. Willemsen |
RecSys | 1 |
| 2015 | Letting Users Choose Recommender Algorithms: An Experimental StudyabstractRecommender systems are not one-size-fits-all; different algorithms and data sources have different strengths, making them a better or worse fit for different users and use cases. As one way of taking advantage of the relative merits of different algorithms, we gave users the ability to change the algorithm providing their movie recommendations and studied how they make use of this power. We conducted our study with the launch of a new version of the MovieLens movie recommender that supports multiple recommender algorithms and allows users to choose the algorithm they want to provide their recommendations. We examine log data from user interactions with this new feature to under-stand whether and how users switch among recommender algorithms, and select a final algorithm to use. We also look at the properties of the algorithms as they were experienced by users and examine their relationships to user behavior. We found that a substantial portion of our user base (25%) used the recommender-switching feature. The majority of users who used the control only switched algorithms a few times, trying a few out and settling down on an algorithm that they would leave alone. The largest number of users prefer a matrix factorization algorithm, followed closely by item-item collaborative filtering; users selected both of these algorithms much more often than they chose a non-personalized mean recommender. The algorithms did produce measurably different recommender lists for the users in the study, but these differences were not directly predictive of user choice. Michael D. Ekstrand, Daniel Kluver, F. Maxwell Harper, Joseph A. Konstan |
RecSys | 1 |
| 2014 | User perception of differences in recommender algorithmsabstractRecent developments in user evaluation of recommender systems have brought forth powerful new tools for understanding what makes recommendations effective and useful. We apply these methods to understand how users evaluate recommendation lists for the purpose of selecting an algorithm for finding movies. This paper reports on an experiment in which we asked users to compare lists produced by three common collaborative filtering algorithms on the dimensions of novelty, diversity, accuracy, satisfaction, and degree of personalization, and to select a recommender that they would like to use in the future. We find that satisfaction is negatively dependent on novelty and positively dependent on diversity in this setting, and that satisfaction predicts the user's final selection. We also compare users' subjective perceptions of recommendation properties with objective measures of those same characteristics. To our knowledge, this is the first study that applies modern survey design and analysis techniques to a within-subjects, direct comparison study of recommender algorithms. Michael D. Ekstrand, F. Maxwell Harper, Martijn C. Willemsen, Joseph A. Konstan |
RecSys | 1 |
| 2013 | Rating support interfaces to improve user experience and recommender accuracyabstractOne of the challenges for recommender systems is that users struggle to accurately map their internal preferences to external measures of quality such as ratings. We study two methods for supporting the mapping process: (i) reminding the user of characteristics of items by providing personalized tags and (ii) relating rating decisions to prior rating decisions using exemplars. In our study, we introduce interfaces that provide these methods of support. We also present a set of methodologies to evaluate the efficacy of the new interfaces via a user experiment. Our results suggest that presenting exemplars during the rating process helps users rate more consistently, and increases the quality of the data. Tien T. Nguyen, Daniel Kluver, Ting-Yu Wang, Pik-Mai Hui, Michael D. Ekstrand, Martijn C. Willemsen, John Riedl |
RecSys | 5 |
| 2012 | RecStore: an extensible and adaptive framework for online recommender queries inside the database engineabstractMost recommendation methods (e.g., collaborative filtering) consist of (1) a computationally intense offline phase that computes a recommender model based on users' opinions of items, and (2) an online phase consisting of SQL-based queries that use the model (generated offline) to derive user preferences and provide recommendations for interesting items. Current application usage trends require a completely online recommender process, meaning the recommender model must update in real time as new opinions enter the system. To tackle this problem, we propose RecStore, a DBMS storage engine module capable of efficient online model maintenance. Externally, models managed by RecStore behave as relational tables, thus existing SQL-based recommendation queries remain unchanged while gaining online model support. RecStore maintains internal statistics and data structures aimed at providing efficient incremental updates to the recommender model, while employing an adaptive strategy for internal maintenance and load shedding to realize a balance between efficiency in updates or query processing based on system workloads. RecStore is also extensible, supporting a declarative syntax for defining recommender models. The efficacy of RecStore is demonstrated by providing the implementation details of three state-of-the-art collaborative filtering models. We provide an extensive experimental evaluation of a prototype of RecStore, built inside the storage engine of PostgreSQL, using a real-life recommender system workload. Justin J. Levandoski, Mohamed Sarwat, Mohamed F. Mokbel, Michael D. Ekstrand |
EDBT | 4 |
| 2012 | When recommenders fail: predicting recommender failure for algorithm selection and combinationabstractHybrid recommender systems --- systems using multiple algorithms together to improve recommendation quality --- have been well-known for many years and have shown good performance in recent demonstrations such as the NetFlix Prize. Modern hybridization techniques, such as feature-weighted linear stacking, take advantage of the hypothesis that the relative performance of recommenders varies by circumstance and attempt to optimize each item score to maximize the strengths of the component recommenders. Less attention, however, has been paid to understanding what these strengths and failure modes are. Understanding what causes particular recommenders to fail will facilitate better selection of the component recommenders for future hybrid systems and a better understanding of how individual recommender personalities can be harnessed to improve the recommender user experience. We present an analysis of the predictions made by several well-known recommender algorithms on the MovieLens 10M data set, showing that for many cases in which one algorithm fails, there is another that will correctly predict the rating. Michael D. Ekstrand, John Riedl |
RecSys | 1 |
| 2012 | How many bits per rating?abstractMost recommender systems assume user ratings accurately represent user preferences. However, prior research shows that user ratings are imperfect and noisy. Moreover, this noise limits the measurable predictive power of any recommender system. We propose an information theoretic framework for quantifying the preference information contained in ratings and predictions. We computationally explore the properties of our model and apply our framework to estimate the efficiency of different rating scales for real world datasets. We then estimate how the amount of information predictions give to users is related to the scale ratings are collected on. Our findings suggest a tradeoff in rating scale granularity: while previous research indicates that coarse scales (such as thumbs up / thumbs down) take less time, we find that ratings with these scales provide less predictive value to users. We introduce a new measure, preference bits per second, to quantitatively reconcile this tradeoff. Daniel Kluver, Tien T. Nguyen, Michael D. Ekstrand, Shilad Sen, John Riedl |
RecSys | 3 |
| 2011 | Rethinking the recommender research ecosystem: reproducibility, openness, and LensKitabstractRecommender systems research is being slowed by the difficulty of replicating and comparing research results. Published research uses various experimental methodologies and metrics that are difficult to compare. It also often fails to sufficiently document the details of proposed algorithms or the evaluations employed. Researchers waste time reimplementing well-known algorithms, and the new implementations may miss key details from the original algorithm or its subsequent refinements. When proposing new algorithms, researchers should compare them against finely-tuned implementations of the leading prior algorithms using state-of-the-art evaluation methodologies. With few exceptions, published algorithmic improvements in our field should be accompanied by working code in a standard framework, including test harnesses to reproduce the described results. To that end, we present the design and freely distributable source code of LensKit, a flexible platform for reproducible recommender systems research. LensKit provides carefully tuned implementations of the leading collaborative filtering algorithms, APIs for common recommender system use cases, and an evaluation framework for performing reproducible offline evaluations of algorithms. We demonstrate the utility of LensKit by replicating and extending a set of prior comparative studies of recommender algorithms --- showing limitations in some of the original results --- and by investigating a question recently raised by a leader in the recommender systems community on problems with error-based prediction evaluation. Michael D. Ekstrand, Michael Ludwig, Joseph A. Konstan, John Riedl |
RecSys | 1 |
| 2011 | LensKit: a modular recommender frameworkabstractLensKit is a new recommender systems toolkit aiming to be a platform for recommender research and education. It provides a common API for recommender systems, modular implementations of several collaborative filtering algorithms, and an evaluation framework for consistent, reproducible offline evaluation of recommender algorithms. In this demo, we will showcase the ease with which LensKit allows recommenders to be configured and evaluated. Michael D. Ekstrand, Michael Ludwig, Jack Kolb, John Riedl |
RecSys | 1 |
| 2011 | UCERSTI 2: second workshop on user-centric evaluation of recommender systems and their interfacesabstractNo abstract available. Martijn C. Willemsen, Dirk G. F. M. Bollen, Michael D. Ekstrand |
RecSys | 3 |
| 2011 | RecBench: Benchmarks for Evaluating Performance of Recommender System Architectures
Justin J. Levandoski, Michael D. Ekstrand, Michael Ludwig, Ahmed Eldawy, Mohamed F. Mokbel, John Riedl |
Proc. VLDB Endow. | 2 |
| 2010 | Automatically building research reading listsabstractAll new researchers face the daunting task of familiarizing themselves with the existing body of research literature in their respective fields. Recommender algorithms could aid in preparing these lists, but most current algorithms do not understand how to rate the importance of a paper within the literature, which might limit their effectiveness in this domain. We explore several methods for augmenting existing collaborative and content-based filtering algorithms with measures of the influence of a paper within the web of citations. We measure influence using well-known algorithms, such as HITS and PageRank, for measuring a node's importance in a graph. Among these augmentation methods is a novel method for using importance scores to influence collaborative filtering. We present a task-centered evaluation, including both an offline analysis and a user study, of the performance of the algorithms. Results from these studies indicate that collaborative filtering outperforms content-based approaches for generating introductory reading lists. Michael D. Ekstrand, Praveen Kannan, James A. Stemper, John T. Butler, Joseph A. Konstan, John Riedl |
RecSys | 1 |