EDBT 2026 Demo / reviewers in the wild / expert
Hugo Manuel Proença
dblp:241/5165
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-7315-5925ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Converted Data is All You Need for Causal Optimization of e-Commerce PromotionsabstractPromotional campaigns are essential drivers of customer engagement and revenue in e-commerce. Maintaining these campaigns within budget constraints requires targeted allocation, traditionally achieved through causal uplift models that rely on vast datasets of user interactions, including non-converted sessions, which introduce challenges such as noisy data, attribution complexity and imbalanced outcomes. We propose a novel approach using converted-only data, which reduces training data size, simplifies attribution, improves efficiency, and mitigates the impact of non-converted interactions. We present a generalized framework for budget constrained promotion allocation with converted-only data and validate it through a benchmarking study and multiple large-scale deployments at Booking.com, positively impacting the experience of millions of customers worldwide. Our results demonstrate that the proposed method is competitive with standard modeling approaches and, in some cases, significantly outperforms them. Dmitri Goldenberg, Hugo Manuel Proença, Amit Livne, Felipe Moraes, Javier Albert, Bracha Shapira |
CIKM | 2 |
| 2025 | Discovering multiple antibiotic resistance phenotypes using diverse top-k subgroup list discoveryabstractAntibiotic resistance is one of the major global threats to human health and occurs when antibiotics lose their ability to combat bacterial infections. In this problem, a clinical decision support system could use phenotypes in order to alert clinicians of the emergence of patterns of antibiotic resistance in patients. Patient phenotyping is the task of finding a set of patient characteristics related to a specific medical problem such as the one described in this work. However, a single explanation of a medical phenomenon might be useless in the eyes of a clinical expert and be discarded. The discovery of multiple patient phenotypes for the same medical phenomenon would be useful in such cases. Therefore, in this work, we define the problem of mining diverse top-k phenotypes and propose the EDSLM algorithm, which is based on the Subgroup Discovery technique, the subgroup list model, and the Minimum Description Length principle. Our proposal provides clinicians with a method with which to obtain multiple and diverse phenotypes of a set of patients. We show a real use case of phenotyping in antimicrobial resistance using the well-known MIMIC-III dataset. Antonio Lopez-Martinez-Carrasco, Hugo Manuel Proença, Jose M. Juarez, Matthijs van Leeuwen, Manuel Campos |
Artif. Intell. Medicine | 2 |
| 2023 | Novel Approach for Phenotyping Based on Diverse Top-K Subgroup Lists
Antonio Lopez-Martinez-Carrasco, Hugo Manuel Proença, Jose M. Juarez, Matthijs van Leeuwen, Manuel Campos |
AIME | 2 |
| 2023 | Uplift Modeling: From Causal Inference to PersonalizationabstractUplift modeling is a collection of machine learning techniques for estimating causal effects of a treatment at the individual or subgroup levels. Over the last years, causality and uplift modeling have become key trends in personalization at online e-commerce platforms, enabling the selection of the best treatment for each user in order to maximize the target business metric. Uplift modeling can be particularly useful for personalized promotional campaigns, where the potential benefit caused by a promotion needs to be weighed against the potential costs. In this tutorial we will cover basic concepts of causality and introduce the audience to state-of-the-art techniques in uplift modeling. We will discuss the advantages and the limitations of different approaches and dive into the unique setup of constrained uplift modeling. Finally, we will present real-life applications and discuss challenges in implementing these models in production. Felipe Moraes, Hugo Manuel Proença, Anastasiia Kornilova, Javier Albert, Dmitri Goldenberg |
CIKM | 2 |
| 2023 | Discovering Diverse Top-K Characteristic Lists
Antonio Lopez-Martinez-Carrasco, Hugo Manuel Proença, Jose M. Juarez, Matthijs van Leeuwen, Manuel Campos |
IDA | 2 |
| 2022 | Robust subgroup discoveryabstractAbstract We introduce the problem ofrobust subgroup discovery, i.e., finding a set of interpretable descriptions of subsets that 1) stand out with respect to one or more target attributes, 2) are statistically robust, and 3) non-redundant. Many attempts have been made to mine eitherlocallyrobust subgroups or to tackle the pattern explosion, but we are the first to address both challenges at the same time from aglobalmodelling perspective. First, we formulate the broad model class of subgroup lists, i.e., ordered sets of subgroups, for univariate and multivariate targets that can consist of nominal or numeric variables, including traditional top-1 subgroup discovery in its definition. This novel model class allows us to formalise the problem of optimal robust subgroup discovery using the Minimum Description Length (MDL) principle, where we resort to optimal Normalised Maximum Likelihood and Bayesian encodings for nominal and numeric targets, respectively. Second, finding optimal subgroup lists is NP-hard. Therefore, we propose SSD++, a greedy heuristic that finds good subgroup lists and guarantees that the most significant subgroup found according to the MDL criterion is added in each iteration. In fact, the greedy gain is shown to be equivalent to a Bayesian one-sample proportion, multinomial, or t-test between the subgroup and dataset marginal target distributions plus a multiple hypothesis testing penalty. Furthermore, we empirically show on 54 datasets that SSD++ outperforms previous subgroup discovery methods in terms of quality, generalisation on unseen data, and subgroup list size. Hugo Manuel Proença, Peter Grünwald, Thomas Bäck, Matthijs van Leeuwen |
Data Min. Knowl. Discov. | 1 |
| 2020 | Discovering Outstanding Subgroup Lists for Numeric Targets Using MDL
Hugo Manuel Proença, Peter Grünwald, Thomas Bäck, Matthijs van Leeuwen |
ECML/PKDD (1) | 1 |
| 2020 | Interpretable multiclass classification by MDL-based rule lists
Hugo Manuel Proença, Matthijs van Leeuwen |
Inf. Sci. | 1 |
| 2016 | Optimizing probabilistic fuzzy systems for classification using metaheuristicsabstractTwo new methods for the optimization of probabilistic fuzzy classifiers are proposed. Probabilistic fuzzy systems are specially attractive due to their explicit and simultaneous modelling of two kinds of uncertainty, namely vagueness in linguistic terms (fuzziness) and probabilistic uncertainty. The current method uses the maximization of the likelihood with the stochastic gradient descent, which not only converges to local minima but also does not guarantee the minimization of the misclassification error. The proposed methods address this specific problem by incorporating global search techniques. The first algorithm proposed is a genetic algorithm with simple crossover and mutation operations. The other is a first generation memetic algorithm which combines the genetic algorithm with the stochastic gradient descent. A total of five benchmarks were used to compare the three algorithms. The results show that the proposed methods have an average relative improvement of 2% and 6% for the accuracy with the genetic and memetic algorithms, respectively. Hugo Manuel Proença, Susana M. Vieira, Uzay Kaymak, Rui Jorge Almeida, João Miguel da Costa Sousa |
FUZZ-IEEE | 1 |