EDBT 2026 Demo / reviewers in the wild / expert
Anastasia Makarova
dblp:244/2207
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Optimization for machine learning · 41% Language models and text generation · 19% Reinforcement learning · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 60% Medical and health informatics · 40% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
2.4 | 4 | 2024 | Adversarial Causal Bayesian Optimization · ICLR 2024 Model-based Causal Bayesian Optimization · ICLR 2023 Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
causal bayesian optimization |
1.4 | 2 | 2024 | Adversarial Causal Bayesian Optimization · ICLR 2024 Model-based Causal Bayesian Optimization · ICLR 2023 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward model training |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Learning theory
online learning |
0.8 | 1 | 2024 | Adversarial Causal Bayesian Optimization · ICLR 2024 |
Machine learning › Reinforcement learning
regret minimization |
0.8 | 1 | 2024 | Adversarial Causal Bayesian Optimization · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling |
0.5 | 1 | 2021 | Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
low-rank tensor decomposition |
0.5 | 1 | 2021 | Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
risk-averse bayesian optimization |
0.5 | 1 | 2021 | Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
tensor decomposition |
0.5 | 1 | 2021 | Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021 |
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition |
0.5 | 1 | 2021 | Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Smart cities and intelligent transportation
shared mobility |
0.2 | 1 | 2024 | Adversarial Causal Bayesian Optimization · ICLR 2024 |
Medical and health informatics › medical imaging
medical image analysis |
0.1 | 1 | 2021 | Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.1 | 1 | 2020 | Mixed-Variable Bayesian Optimization · IJCAI 2020 |
Methods — techniques the papers use, named apart from their topics
submodular optimization · 1.5multiplicative weights · 1.5counterfactual reasoning · 1.5neural network encoder · 1.0cross-approximation · 1.0data augmentation · 0.9causal framework · 0.9structural causal model · 0.7tensor-train decomposition · 0.5tensor train decomposition · 0.5acquisition function · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 12 |
| 2024 | Adversarial Causal Bayesian OptimizationabstractIn Causal Bayesian Optimization (CBO), an agent intervenes on a structural causal model with known graph but unknown mechanisms to maximize a downstream reward variable. In this paper, we consider the generalization where other agents or external events also intervene on the system, which is key for enabling adaptiveness to non-stationarities such as weather changes, market forces, or adversaries. We formalize this generalization of CBO as Adversarial Causal Bayesian Optimization (ACBO) and introduce the first algorithm for ACBO with bounded regret: Causal Bayesian Optimization with Multiplicative Weights (CBO-MW). Our approach combines a classical online learning strategy with causal modeling of the rewards. To achieve this, it computes optimistic counterfactual reward estimates by propagating uncertainty through the causal graph. We derive regret bounds for CBO-MW that naturally depend on graph-related quantities. We further propose a scalable implementation for the case of combinatorial interventions and submodular rewards. Empirically, CBO-MW outperforms non-causal and non-adversarial Bayesian optimization methods on synthetic environments and environments based on real-word data. Our experiments include a realistic demonstration of how CBO-MW can be used to learn users' demand patterns in a shared mobility system and reposition vehicles in strategic areas. Scott Sussex, Pier Giuseppe Sessa, Anastasia Makarova, Andreas Krause 0001 |
ICLR | 3 |
| 2023 | Model-based Causal Bayesian Optimization
Scott Sussex, Anastasia Makarova, Andreas Krause 0001 |
ICLR | 2 |
| 2021 | Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-ApproximationabstractWe propose an end-to-end trainable framework that processes large-scale visual data tensors by looking at a fraction of their entries only. Our method combines a neural network encoder with a tensor train decomposition to learn a low-rank latent encoding, coupled with cross-approximation (CA) to learn the representation through a subset of the original samples. CA is an adaptive sampling algorithm that is native to tensor decompositions and avoids working with the full high-resolution data explicitly. Instead, it actively selects local representative samples that we fetch out-of-core and on demand. The required number of samples grows only logarithmically with the size of the input. Our implicit representation of the tensor in the network enables processing large grids that could not be otherwise tractable in their uncompressed form. The proposed approach is particularly useful for large-scale multidimensional grid data (e.g., 3D tomography), and for tasks that require context over a large receptive field (e.g., predicting the medical condition of entire organs). The code is available at https://github.com/aelphy/c-pic. Mikhail Usvyatsov, Anastasia Makarova, Rafael Ballester-Ripoll, Maksim Rakhuba, Andreas Krause 0001, Konrad Schindler |
ICCV | 2 |
| 2021 | Risk-averse Heteroscedastic Bayesian OptimizationabstractMany black-box optimization tasks arising in high-stakes applications require risk-averse decisions. The standard Bayesian optimization (BO) paradigm, however, optimizes the expected value only. We generalize BO to trade mean and input-dependent variance of the objective, both of which we assume to be unknown a priori. In particular, we propose a novel risk-averse heteroscedastic Bayesian optimization algorithm (RAHBO) that aims to identify a solution with high return and low noise variance, while learning the noise distribution on the fly. To this end, we model both expectation and variance as (unknown) RKHS functions, and propose a novel risk-aware acquisition function. We bound the regret for our approach and provide a robust rule to report the final decision point for applications where only a single solution must be identified. We demonstrate the effectiveness of RAHBO on synthetic benchmark functions and hyperparameter tuning tasks. Anastasia Makarova, Ilnura Usmanova, Ilija Bogunovic, Andreas Krause 0001 |
NeurIPS | 1 |
| 2020 | Mixed-Variable Bayesian OptimizationabstractThe optimization of expensive to evaluate, black-box, mixed-variable functions, i.e. functions that have continuous and discrete inputs, is a difficult and yet pervasive problem in science and engineering. In Bayesian optimization (BO), special cases of this problem that consider fully continuous or fully discrete domains have been widely studied. However, few methods exist for mixed-variable domains and none of them can handle discrete constraints that arise in many real-world applications. In this paper, we introduce MiVaBo, a novel BO algorithm for the efficient optimization of mixed-variable functions combining a linear surrogate model based on expressive feature representations with Thompson sampling. We propose an effective method to optimize its acquisition function, a challenging problem for mixed-variable domains, making MiVaBo the first BO method that can handle complex constraints over the discrete variables. Moreover, we provide the first convergence analysis of a mixed-variable BO algorithm. Finally, we show that MiVaBo is significantly more sample efficient than state-of-the-art mixed-variable BO algorithms on several hyperparameter tuning tasks, including the tuning of deep generative models. Erik A. Daxberger, Anastasia Makarova, Matteo Turchetta, Andreas Krause 0001 |
IJCAI | 2 |