VLDB 2026 Research / reviewers in the wild / expert
Leandro Balby Marinho
dblp:59/4973 · also Leandro Balby
· DBLP profile ↗
23ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-7599-372XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmark Data Contamination in Underrepresented Languages: A Comprehensive Analysis Using Brazilian Data
Iriedson Souto Maior de Moraes Vilar, David Candeia Maia, João Brunet, Fábio Morais 0001, Leandro Balby Marinho |
LREC | 5 |
| 2026 | User Perceptions of Personalized and Generic Explanations in LLM-Driven Recommender SystemsabstractThe adoption of Recommender Systems (RSs) in various domains has become increasingly popular, but concerns have been raised about their lack of transparency and interpretability. While significant advancements have been made in creating explainable RSs, there is still a shortage of automated approaches that can deliver meaningful and contextual human-centered explanations. Numerous studies have evaluated explanations based on human-generated recommendations and explanations to address this gap. However, such approaches do not scale for real-world systems. Building on recent research that exploits Large Language Models (LLMs) for RSs, we propose leveraging the conversational capabilities of ChatGPT to provide users with personalized, human-like, and meaningful explanations for recommended items. Our article presents a user study with 94 participants that measures users’ perceptions of ChatGPT-generated explanations when acting as a recommender system. Regarding recommendations, we assess whether users prefer ChatGPT over random (but popular) recommendations. Concerning explanations, we assess users’ perceptions of personalization, effectiveness, and persuasiveness. We also break down the explanations in its constituent arguments and investigate the differences in argument types between generic and user-specific explanations. Our results show that participants rated ChatGPT’s recommendations significantly higher than random ones ( \(\beta=-.53\) , \({\textrm{p}} < .001\) ). Surprisingly, user-based explanations that explicitly referenced participants’ preferences were not perceived as more personalized or persuasive than generic ones, except when the recommended movie was unfamiliar, where user-based explanations became more effective ( \(\beta=0.35\) , \({\textrm{p}} < .05\) ). Overall, our findings highlight both the promise and the limitations of using LLMs for generating personalized explanations, suggesting that personalization is most impactful when users lack prior knowledge of the item. Itallo Silva, Leandro Balby Marinho, Alan Said, Martijn C. Willemsen |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2025 | Leveraging Query Terms for Efficient Legal Document Recommendation
André Rolim, Leandro Balby Marinho, Edleno Silva de Moura, Marcos Aurélio Domingues, Ricardo S. Oliveira |
ECIR (3) | 2 |
| 2024 | Leveraging ChatGPT for Automated Human-centered Explanations in Recommender SystemsabstractThe adoption of recommender systems (RSs) in various domains has become increasingly popular, but concerns have been raised about their lack of transparency and interpretability. While significant advancements have been made in creating explainable RSs, there is still a shortage of automated approaches that can deliver meaningful and contextual human-centered explanations. Numerous researchers have evaluated explanations based on human-generated recommendations and explanations to address this gap. However, such approaches do not scale for real-world systems. Building on recent research that exploits Large Language Models (LLMs) for RSs, we propose leveraging the conversational capabilities of ChatGPT to provide users with personalized, human-like, and meaningful explanations for recommended items. Our paper presents one of the first user studies that measure users’ perceptions of ChatGPT-generated explanations while acting as an RS. Regarding recommendations, we assess whether users prefer ChatGPT over random (but popular) recommendations. Concerning explanations, we assess users’ perceptions of personalization, effectiveness, and persuasiveness. Our findings reveal that users tend to prefer ChatGPT-generated recommendations over popular ones. Additionally, personalized rather than generic explanations prove to be more effective when the recommended item is unfamiliar. Itallo Silva, Leandro Balby Marinho, Alan Said, Martijn C. Willemsen |
IUI | 2 |
| 2024 | A Tool for Explainable Pension Fund Recommendations using Large Language ModelsabstractIn this demo, we present a prototype tool designed to help financial advisors recommend private pension funds to investors based on their preferences, offering personalized investment suggestions. The tool leverages Large Language Models (LLMs), which enhance explainability by providing clear and understandable rationales for recommendations and effectively handles both sequential and cold-start scenarios. We outline the design, implementation, and results of a user-based evaluation using real-world data. The evaluation shows a high recommendation acceptance rate among financial advisors, highlighting the tool’s potential to improve decision-making in financial advisory services. Eduardo Alves da Silva, Leandro Balby Marinho, Edleno Silva de Moura, Altigran S. da Silva |
RecSys | 2 |
| 2023 | Leveraging Large Language Models for Goal-driven Interactive RecommendationsabstractWe present a proof of concept application for interactive recommendations and explanations leveraging the capabilities of Large Language Models (LLMs). The application creates a highly interactive user-driven setting for recommendations giving users the possibility to explicitly tailor recommendations to their needs. Using the possibilities brought by LLMs, the application further generates convincing explanations of recommendations, aligned with the explicitly stated goals of the users. The web application continuously improves by incorporating user feedback and updating recommendations and explanations as needed. Alan Said, Martijn C. Willemsen, Leandro Balby Marinho, Itallo Silva |
HAI | 3 |
| 2023 | Evaluating Pre-training Strategies for Collaborative FilteringabstractPre-training is essential for effective representation learning models, especially in natural language processing and computer vision-related tasks. The core idea is to learn representations, usually through unsupervised or self-supervised approaches on large and generic source datasets, and use those pre-trained representations (aka embeddings) as initial parameter values during training on the target dataset. Seminal works in this area show that pre-training can act as a regularization mechanism placing the model parameters in regions of the optimization landscape closer to better local minima than random parameter initialization. However, no systematic studies evaluate the effectiveness of pre-training strategies on model-based collaborative filtering. This paper conducts a broad set of experiments to evaluate different pre-training strategies for collaborative filtering using Matrix Factorization (MF) as the base model. We show that such models equipped with pre-training in a transfer learning setting can vastly improve the prediction quality compared to the standard random parameter initialization baseline, reaching state-of-the-art results in standard recommender systems benchmarks. We also present alternatives for the out-of-vocabulary item problem (i.e., items present in target but not in source datasets) and show that pre-training in the context of MF acts as a regularizer, explaining the improvement in model generalization. Júlio B. G. Costa, Leandro Balby Marinho, Rodrygo L. T. Santos, Denis Parra |
UMAP | 2 |
| 2023 | Towards automatic Privacy-Preserving Record Linkage: A Transfer Learning based classification step
Thiago Pereira da Nóbrega, Carlos Eduardo S. Pires, Dimas C. Nascimento, Leandro Balby Marinho |
Data Knowl. Eng. | 4 |
| 2022 | Similarity-Based Explanations meet Matrix Factorization via Structure-Preserving EmbeddingsabstractEmbeddings are core components of modern model-based Collaborative Filtering (CF) methods, such as Matrix Factorization (MF) and Deep Learning variations. In essence, embeddings are mappings of the original sparse representation of categorical features (e.g., user and items) to dense low-dimensional representations. A well-known limitation of such methods is that the learned embeddings are opaque and hard to explain to the users. On the other hand, a key feature of simpler KNN-based CF models (aka user/item-based CF) is that they naturally yield similarity-based explanations, i.e., similar users/items as evidence to support model recommendations. Unlike related works that try to attribute explicit meaning (via metadata) to the learned embeddings, in this paper, we propose to equip the learned embeddings of MF with meaningful similarity-based explanations. First, we show that the learned user/item embeddings of MF do not preserve the distances between users (or items) in the original rating matrix. Next, we propose a novel approach that initializes Stochastic Gradient Descent (SGD) with user/item embeddings that preserve the structural properties of the original input data. We conduct a broad set of experiments and show that our method enables explanations, very similar to the ones provided by KNN-based approaches, without harming the prediction performance. Moreover, we show that fine-tuning the structure-preserving embeddings may unlock better local minima in the optimization space, leading simple vanilla MF to reach competitive performances with the best-known models for the rating prediction task. Leandro Balby Marinho, Júlio Barreto Guedes da Costa, Denis Parra, Rodrygo L. T. Santos |
IUI | 1 |
| 2021 | Assessing Media Bias in Cross-Linguistic and Cross-National Populations
Allan Sales da Costa Melo, Albin Zehe, Leandro Balby Marinho, Adriano Veloso, Andreas Hotho, Janna Omeliyanenko |
ICWSM | 3 |
| 2020 | Computing with Subjectivity LexiconsabstractIn this paper, we introduce a new set of lexicons for expressing subjectivity in text documents written in Brazilian Portuguese. Besides the non-English idiom, in contrast to other subjectivity lexicons available, these lexicons represent different subjectivity dimensions (other than sentiment) and are more compact in number of terms. This last feature was designed intentionally to leverage the power of word embedding techniques, i.e., with the words mapped to an embedding space and the appropriate distance measures, we can easily capture semantically related words to the ones in the lexicons. Thus, we do not need to build comprehensive vocabularies and can focus on the most representative words for each lexicon dimension. We showcase the use of these lexicons in three highly non-trivial tasks: (1) Automated Essay Scoring in the Presence of Biased Ratings, (2) Subjectivity Bias in Brazilian Presidential Elections and (3) Fake News Classification Based on Text Subjectivity. All these tasks involve text documents written in Portuguese. Caio Libânio Melo Jerônimo, Cláudio Elízio Calazans Campelo, Leandro Balby Marinho, Allan Sales da Costa Melo, Adriano Veloso, Roberta Viola |
LREC | 3 |
| 2019 | Fake News Classification Based on Subjective LanguageabstractWhile many works investigate spread patterns of fake news in social networks, we focus on the textual content. Instead of relying on syntactic representations of documents (aka Bag of Words) as many works do, we seek more robust representations that may better differentiate fake from legitimate news. We propose to consider the subjectivity of news under the assumption that the subjectivity levels of legitimate and fake news are significantly different. For computing the subjectivity level of news, we rely on a set subjectivity lexicons built by Brazilian linguists. We then build subjectivity feature vectors for each news article by calculating the Word Mover's Distance (WMD) between the news and these lexicons considering the embedding the news words lie in, in order to classify the documents. The results demonstrate that our method is more robust than classical text classification approaches, especially in scenarios where training and test domains are different. Caio Libânio Melo Jerônimo, Leandro Balby Marinho, Cláudio Elízio Calazans Campelo, Adriano Veloso, Allan Sales da Costa Melo |
iiWAS | 2 |
| 2016 | A Content-Based Approach for Recommending UML Sequence DiagramsabstractSoftware engineers usually have to face a large space of choices during the development process, including libraries/APIs, frameworks and UML models, which undermines their ability in finding the ones that best fit their needs.Recommender Systems appear as a solution to this problem since they have been applied successfully in other domains that suffer from similar issues.In this paper we propose to recommend UML Sequence Diagrams, a popular software artifact in many development processes, as an attempt to mitigate this problem.Our approach consists of: (i) a suitable representation of the users' information needs and sequence diagrams' content; and (ii) two content-based recommendation algorithms to recommend sequence diagrams that match the users' preferences.We performed a study with computer science subjects, where we generated recommendations with (ii) and measured the users' satisfaction upon these recommendations.Our preliminary results show that both algorithms are able to provide accurate recommendations. Thaciana G. O. Cerqueira, Franklin Ramalho, Leandro Balby Marinho |
SEKE | 3 |
| 2015 | Context-Aware Event Recommendation in Event-based Social NetworksabstractThe Web has grown into one of the most important channels to communicate social events nowadays. However, the sheer volume of events available in event-based social networks (EBSNs) often undermines the users' ability to choose the events that best fit their interests. Recommender systems appear as a natural solution for this problem, but differently from classic recommendation scenarios (e.g. movies, books), the event recommendation problem is intrinsically cold-start. Indeed, events published in EBSNs are typically short-lived and, by definition, are always in the future, having little or no trace of historical attendance. To overcome this limitation, we propose to exploit several contextual signals available from EBSNs. In particular, besides content-based signals based on the events' description and collaborative signals derived from users' RSVPs, we exploit social signals based on group memberships, location signals based on the users' geographical preferences, and temporal signals derived from the users' time preferences. Moreover, we combine the proposed signals for learning to rank events for personalized recommendation. Thorough experiments using a large crawl of Meetup.com demonstrate the effectiveness of our proposed contextual learning approach in contrast to state-of-the-art event recommenders from the literature. Augusto Q. de Macedo, Leandro Balby Marinho, Rodrygo L. T. Santos |
RecSys | 2 |
| 2015 | Are Real-World Place Recommender Algorithms Useful in Virtual World Environments?abstractLarge scale virtual worlds such as massive multiplayer online games or 3D worlds gained tremendous popularity over the past few years. With the large and ever increasing amount of content available, virtual world users face the information overload problem. To tackle this issue, game-designers usually deploy recommendation services with the aim of making the virtual world a more joyful environment to be connected at. In this context, we present in this paper the results of a project that aims at understanding the mobility patterns of virtual world users in order to derive place recommenders for helping them to explore content more efficiently. Our study focus on the virtual world SecondLife, one of the largest and most prominent in recent years. Since SecondLife is comparable to real-world Location-based Social Networks (LBSNs), i.e., users can both check-in and share visited virtual places, a natural approach is to assume that place recommenders that are known to work well on real-world LBSNs will also work well on SecondLife. We have put this assumption to the test and found out that (i) while collaborative filtering algorithms have compatible performances in both environments, (ii) existing place recommenders based on geographic metadata are not useful in SecondLife. Leandro Balby Marinho, Christoph Trattner, Denis Parra |
RecSys | 1 |
| 2015 | SPS'15: 2015 International Workshop on Social Personalization & SearchabstractNo abstract available. Christoph Trattner, Denis Parra, Peter Brusilovsky, Leandro Balby Marinho |
SIGIR | 4 |
| 2014 | DYSCS: A platform to build geographically and semantically enhanced social content sites
Luciana Cavalcante de Menezes, Cláudio de Souza Baptista, Ana Gabrielle Ramos Falcão, Maxwell Guimarães de Oliveira, Leandro Balby Marinho |
J. Syst. Softw. | 5 |
| 2010 | Semi-supervised Tag Recommendation - Using Untagged Resources to Mitigate Cold-Start Problems
Christine Preisach, Leandro Balby Marinho, Lars Schmidt-Thieme |
PAKDD (1) | 2 |
| 2009 | Learning optimal ranking with tensor factorization for tag recommendationabstractTag recommendation is the task of predicting a personalized list of tags for a user given an item. This is important for many websites with tagging capabilities like last.fm or delicious. In this paper, we propose a method for tag recommendation based on tensor factorization (TF). In contrast to other TF methods like higher order singular value decomposition (HOSVD), our method RTF ('ranking with tensor factorization') directly optimizes the factorization model for the best personalized ranking. RTF handles missing values and learns from pairwise ranking constraints. Our optimization criterion for TF is motivated by a detailed analysis of the problem and of interpretation schemes for the observed data in tagging systems. In all, RTF directly optimizes for the actual problem using a correct interpretation of the data. We provide a gradient descent algorithm to solve our optimization problem. We also provide an improved learning and prediction method with runtime complexity analysis for RTF. The prediction runtime of RTF is independent of the number of observations and only depends on the factorization dimensions. Besides the theoretical analysis, we empirically show that our method outperforms other state-of-the-art tag recommendation methods like FolkRank, PageRank and HOSVD both in quality and prediction runtime. Steffen Rendle, Leandro Balby Marinho, Alexandros Nanopoulos, Lars Schmidt-Thieme |
KDD | 2 |
| 2008 | Folksonomy-Based Collabulary Learning
Leandro Balby Marinho, Krisztián Búza, Lars Schmidt-Thieme |
ISWC | 1 |
| 2007 | Tag Recommendations in Folksonomies
Robert Jäschke, Leandro Balby Marinho, Andreas Hotho, Lars Schmidt-Thieme, Gerd Stumme |
PKDD | 2 |
| 2007 | A domain model of Web recommender systems based on usage mining and collaborative filtering
Rosario Girardi, Leandro Balby Marinho |
Requir. Eng. | 2 |
| 2005 | A system of agent-based software patterns for user modeling based on usage miningabstractIn adaptive hypermedia systems, a user can select explicitly an adaptation effect or he/she can leave the system execute some of these functions. An important component of an adaptive system is the ability to model the users of the system according to their goals and preferences. Web usage mining aims at discover interesting patterns of use by analyzing Web usage data. This information can be used to capture implicitly user models and used them for the adaptation of systems. User modeling and system adaptability can be approached through the agent paradigm. This article summarizes a system of architectural and detailed design patterns describing known agent-based solutions to recurrent problems of user modeling based on usage mining along with the description of a general purpose problem-solving architectural pattern used by some of the first ones. Patterns are derived from recurrent designs of specific agent-based applications. The proposed patterns are being developed in the context of a Multi-Agent Domain Engineering research project, which approaches software complexity and productivity through the construction of techniques and tools promoting software reuse in Multi-Agent Domain Engineering. Rosario Girardi, Leandro Balby Marinho, Ismênia Ribeiro de Oliveira |
Interact. Comput. | 2 |