VLDB 2026 Research / reviewers in the wild / expert
Saikishore Kalloori
dblp:185/2855
· DBLP profile ↗
13ranked-venue papers in the field
4as first author
10since 2021 · last 2024
0000-0002-9669-0602ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 11 (4 first)Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving German News Clustering with Contrastive LearningabstractAutomatic news articles clustering is one of the most important tasks for news publishers. Traditional unsupervised models exploit generic text representation (e.g., BERT) and typically do not consider the relationships between each paragraph in news articles. Such depth learning from news articles is important for clustering full-length articles. Recently contrastive learning (CL) has shown to be a popular method for representation learning that uses positive and negative data pairs generated using data augmentation techniques to improve the representation in the latent space. In this work, we propose text augmentation methods and use contrastive learning to cluster daily growing full-length German news articles. Our experiments on four German news article datasets (one labeled and three unlabeled datasets) demonstrate that contrastive learning and our text augmentation methods significantly improve the representation of news articles compared to generic pre-trained text representation and have high performance for clustering tasks. Piriyakorn Piriyatamwong, Saikishore Kalloori, Fabio Zünd |
CIKM | 2 |
| 2024 | Unified Argument Retrieval System from German News Articles Using Large Language ModelsabstractThe rapid growth in the number of news articles published daily can create challenges for users to explore specific topics and gather different perspectives around the topics to make neutral and unbiased conclusions. The system's ability to intelligently cluster news articles from multiple sources and retrieve concise (pro/con) relevant arguments is necessary for users' well-informed decision-making. In this paper, we introduce our unified argument retrieval system that uses our clustering model to cluster news articles and subsequently extracts the core arguments from news articles using the argument prediction model. We conducted a user study to understand the system's usability and users' satisfaction with the quality of clusters and arguments extracted. Piriyakorn Piriyatamwong, Saikishore Kalloori, Fabio Zünd |
CIKM | 2 |
| 2024 | RecSys Challenge 2024: Balancing Accuracy and Editorial Values in News RecommendationsabstractThe RecSys Challenge 2024 aims to advance news recommendation by addressing both the technical and normative challenges inherent in designing effective and responsible recommender systems for news publishing. This paper describes the challenge, including its objectives, problem setting, and the dataset provided by the Danish news publishers Ekstra Bladet and JP/Politikens Media Group (“Ekstra Bladet”). The challenge explores the unique aspects of news recommendation, such as modeling user preferences based on behavior, accounting for the influence of the news agenda on user interests, and managing the rapid decay of news items. Additionally, the challenge embraces normative complexities, investigating the effects of recommender systems on news flow and their alignment with editorial values. We summarize the challenge setup, dataset characteristics, and evaluation metrics. Finally, we announce the winners and highlight their contributions. The dataset is available at: https://recsys.eb.dk. Johannes Kruse 0002, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Claudio Pomo, Abhishek Srivastava 0004, Anshuk Uppal, Michael Riis Andersen, Jes Frellsen |
RecSys | 3 |
| 2023 | RecSys Challenge 2023: Deep Funnel Optimization with a Focus on User PrivacyabstractThe RecSys 2023 Challenge involved a conversion prediction task in the online advertising space. The dataset was provided by ShareChat (Mohalla Tech Pvt Ltd). The challenge data represents a sample of ad impressions served to the users over a period of 22 days and the task is for a given ad impression, to predict a conversion (install an app) will happen or not. The challenge ran for 3 months with a public dashboard. There were 519 teams registered and 231 teams made at least one submission. The task setting represents an important research area of modeling ad recommendations under user privacy. We identify interesting themes in feature engineering, addressing sparsity and calibrating across multi-step predictions. Rahul Agrawal, Sarang Brahme, Sourav Maitra, Saikishore Kalloori, Abhishek Srivastava 0004, Yong Liu 0020, Athirai Aravazhi Irissappane |
RecSys | 4 |
| 2023 | A Retrieval System for Images and Videos based on Aesthetic Assessment of VisualsabstractAttractive images or videos are the visual backbones of journalism and social media to gain the user's attention. From trailers to teaser images to image galleries, appealing visuals have only grown in importance over the years. However, selecting eye-catching shots from a video or the perfect image from large image collections is a challenging and time-consuming task. We present our tool that can assess image and video content from an aesthetic standpoint. We discovered that it is possible to perform such an assessment by combining expert knowledge with data-driven information. We combine the relevant aesthetic features and machine learning algorithms into an aesthetics retrieval system, which enables users to sort uploaded visuals based on an aesthetic score and interact with additional photographic, cinematic, and person-specific features. Daniel Vera Nieto, Saikishore Kalloori, Fabio Zünd, Clara Fernandez-Labrador, Marc Willhaus, Severin Klingler, Markus Gross 0001 |
SIGIR | 2 |
| 2022 | RecSys Challenge 2022: Fashion Purchase PredictionabstractThe RecSys 2022 Challenge was a session-based recommendation task in the fashion domain. The dataset was supplied by Dressipi. Given session data consisting of views and purchases, as well as content data representing the fashion characteristics of the items, the task was to predict which item was purchased at the end of the session. The challenge ran for 3 months with a public leaderboard and final result on a separate hidden test set. There were over 300 teams that submitted a solution to the leaderboard and about 50 that submitted a solution for the final test set. The winning team achieved a MRR score of 0.216 which means that the correct target item was on average ranked 5th in the list of predictions. We identify some interesting common themes among the solutions in this paper and the winning approaches are presented in the workshop. Nick Landia, Frederick Cheung, Donna North, Saikishore Kalloori, Abhishek Srivastava 0004, Bruce Ferwerda |
RecSys | 4 |
| 2022 | Personalized Information Retrieval for Touristic Attractions in Augmented RealityabstractThe rapid advances and increasing accessibility of augmented reality (AR) in recent years opened up many new possibilities to incorporate AR into our daily lives. A very interesting area for AR is tourism where one can enhance attractions with virtual elements and provide tourists with additional information about the places they are visiting. In this paper, we present our prototype, an AR application that augments various points of interest (POIs) by showing images and facts about each POI. We also developed a simple recommender system that ensures the facts are selected based on user preferences, thus creating a unique and personalized experience for each user. Furthermore, we also conducted a live user study to assess the usability of our prototype and the usefulness of our personalization system. Felix Yang, Saikishore Kalloori, Ribin Chalumattu, Markus Gross 0001 |
WSDM | 2 |
| 2021 | RecSys 2021 Challenge Workshop: Fairness-aware engagement prediction at scale on Twitter's Home TimelineabstractThe workshop features presentations of accepted contributions to the RecSys Challenge 2021, organized by Politecnico di Bari, ETH Zürich, Jönköping University, and the data set is provided by Twitter. The challenge focuses on a real-world task of tweet engagement prediction in a dynamic environment. For 2021, the challenge considers four different engagement types: Likes, Retweet, Quote, and replies. This year’s challenge brings the problem even closer to Twitter’s real recommender systems by introducing latency constraints. We also increases the data size to encourage novel methods. Also, the data density is increased in terms of the graph where users are considered to be nodes and interactions as edges. The goal is twofold: to predict the probability of different engagement types of a target user for a set of Tweets based on heterogeneous input data while providing fair recommendations. In fact, multi-goal optimization considering accuracy and fairness is particularly challenging. However, we believed that the recommendation community was nowadays mature enough to face the challenge of providing accurate and, at the same time, fair recommendations. To this end, Twitter has released a public dataset of close to 1 billion data points, > 40 million each day over 28 days. Week 1 − 3 will be used for training and week 4 for evaluation and testing. Each datapoint contains the tweet along with engagement features, user features, and tweet features. A peculiarity of this challenge is related to keeping the dataset updated with the platform: if a user deletes a Tweet, or their data from Twitter, the dataset is promptly updated. Moreover, each change in the dataset implied new evaluations of all submissions and the update of the leaderboard metrics. The challenge was well received with 578 registered users, and 386 submissions. Vito Walter Anelli, Saikishore Kalloori, Bruce Ferwerda, Luca Belli, Alykhan Tejani, Frank Portman, Alexandre Lung-Yut-Fong, Benjamin Paul Chamberlain, Yuanpu Xie, Jonathan J. Hunt, Michael M. Bronstein, Wenzhe Shi |
RecSys | 2 |
| 2021 | Horizontal Cross-Silo Federated Recommender SystemsabstractRecommender systems (RSs) completely rely on the knowledge of training information to generate recommendations. However, due to privacy, ownership, and protection of users’ information, such training information is not easily accessible or shared with an RS. Moreover, with recent regulations in privacy laws (e.g, GDPR), collecting user preferences and perform centralized training may not be feasible. Federated Learning (FL) is a form of machine learning technique where the goal is to learn a high-quality recommendation model without never directly accessing raw training data. In this work, we specifically focus on situations where multiple stakeholders (referred to as corporate companies like e-commerce business partners, hospitals, banks, news media publishers) participate in federated learning to build a shared recommendation model. We performed offline experiments by simulating a real federated learning setup and investigated the benefits federated learning brings to stakeholders in terms of ranking compared to an RS model trained without participating in federated learning. Our experimental results reveal that stakeholders can significantly benefit from federated learning to generate accurate recommendations. Moreover, we also study the use and benefits of federated learning in situations when there are not enough preferences available for users. Saikishore Kalloori, Severin Klingler |
RecSys | 1 |
| 2021 | A Practical Federated Learning Framework for Small Number of StakeholdersabstractFederated Learning (FL) allows to collaboratively build machine learning models between different entities without the need for sharing or gathering the data. In FL, typically there is a global server and a set of clients (stakeholders) to build shared machine learning models. In contrast to distributed machine learning, the controller of the training process (here the global server) never sees the data of the stakeholders participating in FL. Every stakeholder owns his own data and doesn't share it. During the training and learning process, only the model updates (e.g. gradients) are shared. To our best of knowledge, we did not find a publicly available practical federated learning framework for stakeholders. We have built a framework that enables FL for a small number of stakeholders. In the paper, we describe the framework architecture, communication protocol, and algorithms. Our framework is open-sourced and it is easy to set up for stakeholders and ensures that no private information is leaked during the training process. Christian Schneebeli, Saikishore Kalloori, Severin Klingler |
WSDM | 2 |
| 2019 | Item Recommendation by Combining Relative and Absolute Feedback DataabstractUser preferences in the form of absolute feedback, s.a., ratings, are widely exploited in Recommender Systems (RSs). Recent research has explored the usage of preferences expressed with pairwise comparisons, which signal relative feedback. It has been shown that pairwise comparisons can be effectively combined with ratings, but, it is important to fine tune the technique that leverages both types of feedback. Previous approaches train a single model by converting ratings into pairwise comparisons, and then use only that type of data. However, we claim that these two types of preferences reveal different information about users interests and should be exploited differently. Hence, in this work, we develop a ranking technique that separately exploits absolute and relative preferences in a hybrid model. In particular, we propose a joint loss function which is computed on both absolute and relative preferences of users. Our proposed ranking model uses pairwise comparisons data to predict the user's preference order between pairs of items and uses ratings to push high rated (relevant) items to the top of the ranking. Experimental results on three different data sets demonstrate that the proposed technique outperforms competitive baseline algorithms on popular ranking-oriented evaluation metrics. Saikishore Kalloori, Tianyu Li 0007, Francesco Ricci 0001 |
SIGIR | 1 |
| 2018 | Eliciting pairwise preferences in recommender systemsabstractPreference data in the form of ratings or likes for items are widely used in many Recommender Systems. However, previous research has shown that even item comparisons, which generate pairwise preference data, can be used to model user preferences. Moreover, pairwise preferences can be effectively combined with ratings to compute recommendations. In such hybrid approaches, the Recommender System requires to elicit both types of preference data from the user. In this work, we aim at identifying how and when to elicit pairwise preferences, i.e., when this form of user preference data is more meaningful for the user to express and more beneficial for the system. We conducted an online A/B test and compared a rating-only based system variant with another variant that allows the user to enter both types of preferences. Our results demonstrate that pairwise preferences are valuable and useful, especially when the user is focusing on a specific type of items. By incorporating pairwise preferences, the system can generate better recommendations than a state of the art rating-only based solution. Additionally, our results indicate that there seems to be a dependency between the user's personality, the perceived system usability and the satisfaction for the preference elicitation procedure, which varies if only ratings or a combination of ratings and pairwise preferences are elicited. Saikishore Kalloori, Francesco Ricci 0001, Rosella Gennari |
RecSys | 1 |
| 2016 | Pairwise Preferences Based Matrix Factorization and Nearest Neighbor Recommendation TechniquesabstractMany recommendation techniques rely on the knowledge of preferences data in the form of ratings for items. In this paper, we focus on pairwise preferences as an alternative way for acquiring user preferences and building recommendations. In our scenario, users provide pairwise preference scores for a set of item pairs, indicating how much one item in each pair is preferred to the other. We propose a matrix factorization (MF) and a nearest neighbor (NN) prediction techniques for pairwise preference scores. Our MF solution maps users and items pairs to a joint latent features vector space, while the proposed NN algorithm leverages specific user-to-user similarity functions well suited for comparing users preferences of that type. We compare our approaches to state of the art solutions and show that our solutions produce more accurate pairwise preferences and ranking predictions. Saikishore Kalloori, Francesco Ricci 0001, Marko Tkalcic |
RecSys | 1 |