Kim Falk

dblp:274/7552 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-3573-9257ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Diversity in News Recommendations Increases Click-Through Rates: Insights from an Online Experiment and User Study
abstract
Diversity is a widely studied beyond-accuracy aspect of recommender systems, particularly in the news domain. Extensive research has explored its theoretical foundations and proposed algorithmic strategies to promote it, with most evaluations conducted through offline experiments. This work presents the results of deploying and evaluating diversification methods in a large-scale production news recommender system. Motivated by the goal of upholding editorial values, we compare three diversification methods: Interleaving and two implementations of Intra-List Diversification (ILD), relying on Term Frequency-Inverse Document Frequency (TF-IDF) and Bidirectional Encoder Representations from Transformers (BERT) embeddings, respectively. Across a two-week online experiment (A/B test) and a follow-up user study on a large-scale production news platform, ILD with BERT embeddings improved diversity as measured by a reduction in Intra-List Similarity (ILS) and increased Click-Through Rates (CTRs), while also improving users’ perceived relevance.
Robin Verachtert, Kim Falk, Christine Bauer 0001
UMAP2
2025 Multi-Armed Bandits in the Wild
Kim Falk
RecSys1
2024 A Framework and Toolkit for Testing the Correctness of Recommendation Algorithms
abstract
Evaluating recommender systems adequately and thoroughly is an important task. Significant efforts are dedicated to proposing metrics, methods, and protocols for doing so. However, there has been little discussion in the recommender systems’ literature on the topic of testing. In this work, we adopt and adapt concepts from the software testing domain, e.g., code coverage, metamorphic testing, or property-based testing, to help researchers to detect and correct faults in recommendation algorithms. We propose a test suite that can be used to validate the correctness of a recommendation algorithm, and thus identify and correct issues that can affect the performance and behavior of these algorithms. Our test suite contains both black box and white box tests at every level of abstraction, i.e., system, integration, and unit. To facilitate adoption, we release RecPack Tests , an open-source Python package containing template test implementations. We use it to test four popular Python packages for recommender systems: RecPack , PyLensKit , Surprise , and Cornac . Despite the high test coverage of each of these packages, we find that we are still able to uncover undocumented functional requirements and even some bugs. This validates our thesis that testing the correctness of recommendation algorithms can complement traditional methods for evaluating recommendation algorithms.
Lien Michiels, Robin Verachtert, Andres Ferraro, Kim Falk, Bart Goethals
Trans. Recomm. Syst.4
2023 Recommenders In the wild - Practical Evaluation Methods
abstract
The gap between training a recommender model and actually having a recommender system in production is a topic often neglected. A recommender system is far more than a model which produces good metrics in an offline evaluation. Specifically, the evaluation of various recommendation engines in production is often very different from offline evaluations on a laptop. This tutorial will go through many practical steps and focus on the development, evaluation and, in particular, metrics and A/B tests.
Kim Falk, Morten Arngren
RecSys1
2022 Optimizing product recommendations for millions of merchants
abstract
At Shopify, we serve product recommendations to customers across millions of merchants’ online stores. It is a challenge to provide optimized recommendations to all of these independent merchants; one model might lead to an overall improvement in our metrics on aggregate, but significantly degrade recommendations for some stores. To ensure we provide high quality recommendations to all merchant segments, we develop several models that work best in different situations as determined in offline evaluation. Learning which strategy best works for a given segment also allows us to start off new stores with good recommendations, without necessarily needing to rely on an individual store amassing large amounts of traffic. In production, the system will start out with the best strategy for a given merchant, and then adjust to the current environment using multi-armed bandits. Collectively, this methodology allows us to optimize the types of recommendations served on each store.
Kim Falk, Chen Karako
RecSys1