VLDB 2026 Research / reviewers in the wild / expert
Zahra Donyavi
dblp:274/2643
· DBLP profile ↗
5ranked-venue papers
4as first author
4since 2021 · last 2025
0009-0001-8631-5884ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Match: A Maximum-Likelihood Approach for Classification under Label ShiftabstractMachine learning models often suffer from performance degradation when dealing with class distributions that differ from the training distribution, a scenario commonly referred to as label shift. Addressing this challenge, this paper introduces Match, a novel adjustment approach that maximizes the likelihood of predicted probabilities under class prevalence constraints. Unlike existing methods such as retraining with instance re-weighting and the Bayes update rule, Match ensures that the adjusted class distribution aligns precisely with the prevalence estimates from quantifiers. By formulating the adjustment process as a binary integer linear optimization problem, Match benefits from efficient mixed-integer solvers. Extensive experiments demonstrate that Match outperforms the state-of-the-art in classifier adjustment with statistical significance, particularly in handling scenarios with imbalanced distributions. Zahra Donyavi, Feiyu Li, Yunrui Zhang, Diego Furtado Silva, Gustavo Batista |
KDD (2) | 1 |
| 2024 | MC-SQ and MC-MQ: Ensembles for Multi-Class QuantificationabstractQuantification research proposes methods to estimate the class distribution in an independent sample. Quantification methods find applications in areas that rely on estimated aggregated quantities, such as epidemiology, sentiment analysis, political research, and ecological surveillance. For instance, epidemiologists are often concerned with the dynamics of the number of disease cases across space and time. Thus, while classification predicts individual subjects, quantifiers are the methods that directly estimate the number of cases. Although quantification is a thriving area of research, with numerous approaches proposed in the last decade, most focus has been on binary-class quantifiers. One common approach for multi-class quantification is the one-versus-all (OVA) approach, but empirical evidence suggests its performance is suboptimal. This paper's first contribution is to elucidate why OVA quantifiers struggle to perform well in multi-class settings due to a distribution shift. To circumvent this problem, our second proposal is two new multi-class quantifiers based on ensemble learning that significantly improve performance for binary and multi-class settings. Our comprehensive experimental setup with 37 state-of-the-art (single and ensemble) quantifiers shows that our ensembles are the best-performing quantifiers and rank first in a recent quantification competition. Zahra Donyavi, Adriane Beatriz de Souza Serapião, Gustavo Batista |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Ensembles of Classifiers and Quantifiers with Data Fusion for Quantification Learning
Adriane Beatriz de Souza Serapião, Zahra Donyavi, Gustavo Batista |
DS | 2 |
| 2023 | MC-SQ: A Highly Accurate Ensemble for Multi-class QuantificationabstractQuantification research proposes methods to estimate the class distribution in an independent sample. Many areas, such as epidemiology, sentiment analysis, political research and ecological surveillance, rely on quantification methods to estimate aggregated quantities. For instance, epidemiologists are often concerned with the dynamics of the number of disease cases across space and time. Thus, while classification predicts individual subjects, quantification is the class of methods that directly estimate the number of cases. Quantification is a thriving research area, and the community has proposed several approaches in the last decade. Nevertheless, most quantification research has focused on binary-class quantifiers, expecting these approaches to extend to multi-class using the one-versus-all (OVA) approach. However, there is enough empirical evidence indicating the performance of OVA multi-class quantifiers is subpar. This paper has two main contributions. First, we demonstrate why OVA quantifiers are doomed to underperform in multi- class settings due to a distribution shift they cannot handle. Second, we propose a new class of quantifiers based on ensemble learning that boosts the performance of the base quantifiers in the binary and, more importantly, multi-class settings. In one of the most comprehensive experimental setups ever attempted in quantification research, we show that our ensembles are the best-performing quantifiers compared with 33 state-of-the-art (single and ensemble) quantifiers and rank first in a recent quantification competition. Zahra Donyavi, Adriane Serapio, Gustavo Batista |
SDM | 1 |
| 2020 | Diverse training dataset generation based on a multi-objective optimization for semi-Supervised classification
Zahra Donyavi, Shahrokh Asadi |
Pattern Recognit. | 1 |