VLDB 2026 Research / reviewers in the wild / expert
Faisal Shehzad
dblp:256/1185
· DBLP profile ↗
5ranked-venue papers in the field
5as first author
5since 2021 · last 2025
0000-0001-6239-8198ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | "We Share Our Code Online": Why This Is Not Enough to Ensure Reproducibility and Progress in Recommender Systems ResearchabstractIssues with reproducibility have been identified as a major factor hampering progress in recommender systems research. In response, researchers increasingly share the code of their models. However, the provision of only the code of the proposed model is usually not sufficient to ensure reproducibility. In many works, the central claim is that a new model is advancing the state of the art. Thus, it is crucial that the entire experiment is reproducible, including the configuration and the results of the considered baselines. With this work, our goal is to gauge the level of reproducibility in algorithms research in recommender systems. We systematically analyzed the reproducibility level of 65 papers published at a top-ranked conference during the last three years. Our results are sobering. While the model code is shared in about two thirds of the papers, the code of the baselines is provided only in eight cases. The hyperparameters of the baselines are reported even less frequently, and how these were exactly determined is not explained in any paper. As a result, it is commonly not only impossible to reproduce the full result tables reported in the papers, it is also unclear if the claimed improvements over the state of the art were actually achieved. Overall, we conclude that the research community has not reached the required level of reproducibility yet. We therefore call for more rigorous reproducibility standards to ensure progress in this field. Faisal Shehzad, Timo Breuer 0002, Maria Maistro, Dietmar Jannach |
RecSys | 1 |
| 2025 | Revisiting the Performance of Graph Neural Networks for Session-based Recommendation
Faisal Shehzad, Dietmar Jannach |
RecSys | 1 |
| 2025 | A Worrying Reproducibility Study of Intent-Aware Recommendation ModelsabstractLately, we have observed a growing interest in intent-aware recommender systems (IARS). The promise of such systems is that they are capable of generating better recommendations by predicting and considering the underlying motivations and short-term goals of consumers. From a technical perspective, various sophisticated neural models were recently proposed in this emerging and promising area. In the broader context of complex neural recommendation models, a growing number of research works unfortunately indicates that (i) reproducing such works is often difficult and (ii) that the true benefits of such models may be limited in reality, e.g., because the reported improvements were obtained through comparisons with untuned or weak baselines. In this work, we investigate if recent research in IARS is similarly affected by such problems. Specifically, we tried to reproduce five contemporary IARS models that were published in top-level outlets, and we benchmarked them against a number of traditional non-neural recommendation models. In two of the cases, running the provided code with the optimal hyperparameters reported in the paper did not yield the results reported in the paper. Worryingly, we find that all examined IARS approaches are consistently outperformed by at least one traditional model. These findings point to sustained methodological issues and to a pressing need for more rigorous scholarly practices. Faisal Shehzad, Maurizio Ferrari Dacrema, Dietmar Jannach |
SIGIR | 1 |
| 2024 | Performance Comparison of Session-Based Recommendation Algorithms Based on GNNs
Faisal Shehzad, Dietmar Jannach |
ECIR (4) | 1 |
| 2023 | Everyone's a Winner! On Hyperparameter Tuning of Recommendation ModelsabstractThe performance of a recommender system algorithm in terms of common offline accuracy measures often strongly depends on the chosen hyperparameters. Therefore, when comparing algorithms in offline experiments, we can obtain reliable insights regarding the effectiveness of a newly proposed algorithm only if we compare it to a number of state-of-the-art baselines that are carefully tuned for each of the considered datasets. While this fundamental principle of any area of applied machine learning is undisputed, we find that the tuning process for the baselines in the current literature is barely documented in much of today’s published research. Ultimately, in case the baselines are actually not carefully tuned, progress may remain unclear. In this paper, we exemplify through a computational experiment involving seven recent deep learning models how every method in such an unsound comparison can be reported to be outperforming the state-of-the-art. Finally, we iterate appropriate research practices to avoid unreliable algorithm comparisons in the future. Faisal Shehzad, Dietmar Jannach |
RecSys | 1 |