Anastasia Makarova

dblp:244/2207 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Optimization for machine learning · 41% Language models and text generation · 19% Reinforcement learning · 15%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 60% Medical and health informatics · 40%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
2.442024
Adversarial Causal Bayesian Optimization · ICLR 2024
Model-based Causal Bayesian Optimization · ICLR 2023
Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
causal bayesian optimization
1.422024
Adversarial Causal Bayesian Optimization · ICLR 2024
Model-based Causal Bayesian Optimization · ICLR 2023
Natural language and speech › Language models and text generation
alignment
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Natural language and speech › Language models and text generation › alignment
reward hacking
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Machine learning › Reinforcement learning › reward learning
reward model training
0.912025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Machine learning › Learning theory
online learning
0.812024
Adversarial Causal Bayesian Optimization · ICLR 2024
Machine learning › Reinforcement learning
regret minimization
0.812024
Adversarial Causal Bayesian Optimization · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › probabilistic regression
heteroscedastic modeling
0.512021
Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
low-rank tensor decomposition
0.512021
Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
risk-averse bayesian optimization
0.512021
Risk-averse Heteroscedastic Bayesian Optimization · NeurIPS 2021
Machine learning › Representation and self-supervised learning
tensor decomposition
0.512021
Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition
0.512021
Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.312025
RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025
Smart cities and intelligent transportation
shared mobility
0.212024
Adversarial Causal Bayesian Optimization · ICLR 2024
Medical and health informatics › medical imaging
medical image analysis
0.112021
Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation · ICCV 2021
Machine learning › Optimization for machine learning
hyperparameter optimization
0.112020
Mixed-Variable Bayesian Optimization · IJCAI 2020

Methods — techniques the papers use, named apart from their topics

submodular optimization · 1.5multiplicative weights · 1.5counterfactual reasoning · 1.5neural network encoder · 1.0cross-approximation · 1.0data augmentation · 0.9causal framework · 0.9structural causal model · 0.7tensor-train decomposition · 0.5tensor train decomposition · 0.5acquisition function · 0.5
YearPublicationVenuePosition
2025 RRM: Robust Reward Model Training Mitigates Reward Hacking
abstract
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%.
Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh
ICLR12
2024 Adversarial Causal Bayesian Optimization
abstract
In Causal Bayesian Optimization (CBO), an agent intervenes on a structural causal model with known graph but unknown mechanisms to maximize a downstream reward variable. In this paper, we consider the generalization where other agents or external events also intervene on the system, which is key for enabling adaptiveness to non-stationarities such as weather changes, market forces, or adversaries. We formalize this generalization of CBO as Adversarial Causal Bayesian Optimization (ACBO) and introduce the first algorithm for ACBO with bounded regret: Causal Bayesian Optimization with Multiplicative Weights (CBO-MW). Our approach combines a classical online learning strategy with causal modeling of the rewards. To achieve this, it computes optimistic counterfactual reward estimates by propagating uncertainty through the causal graph. We derive regret bounds for CBO-MW that naturally depend on graph-related quantities. We further propose a scalable implementation for the case of combinatorial interventions and submodular rewards. Empirically, CBO-MW outperforms non-causal and non-adversarial Bayesian optimization methods on synthetic environments and environments based on real-word data. Our experiments include a realistic demonstration of how CBO-MW can be used to learn users' demand patterns in a shared mobility system and reposition vehicles in strategic areas.
Scott Sussex, Pier Giuseppe Sessa, Anastasia Makarova, Andreas Krause 0001
ICLR3
2023 Model-based Causal Bayesian Optimization
Scott Sussex, Anastasia Makarova, Andreas Krause 0001
ICLR2
2021 Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation
abstract
We propose an end-to-end trainable framework that processes large-scale visual data tensors by looking at a fraction of their entries only. Our method combines a neural network encoder with a tensor train decomposition to learn a low-rank latent encoding, coupled with cross-approximation (CA) to learn the representation through a subset of the original samples. CA is an adaptive sampling algorithm that is native to tensor decompositions and avoids working with the full high-resolution data explicitly. Instead, it actively selects local representative samples that we fetch out-of-core and on demand. The required number of samples grows only logarithmically with the size of the input. Our implicit representation of the tensor in the network enables processing large grids that could not be otherwise tractable in their uncompressed form. The proposed approach is particularly useful for large-scale multidimensional grid data (e.g., 3D tomography), and for tasks that require context over a large receptive field (e.g., predicting the medical condition of entire organs). The code is available at https://github.com/aelphy/c-pic.
Mikhail Usvyatsov, Anastasia Makarova, Rafael Ballester-Ripoll, Maksim Rakhuba, Andreas Krause 0001, Konrad Schindler
ICCV2
2021 Risk-averse Heteroscedastic Bayesian Optimization
abstract
Many black-box optimization tasks arising in high-stakes applications require risk-averse decisions. The standard Bayesian optimization (BO) paradigm, however, optimizes the expected value only. We generalize BO to trade mean and input-dependent variance of the objective, both of which we assume to be unknown a priori. In particular, we propose a novel risk-averse heteroscedastic Bayesian optimization algorithm (RAHBO) that aims to identify a solution with high return and low noise variance, while learning the noise distribution on the fly. To this end, we model both expectation and variance as (unknown) RKHS functions, and propose a novel risk-aware acquisition function. We bound the regret for our approach and provide a robust rule to report the final decision point for applications where only a single solution must be identified. We demonstrate the effectiveness of RAHBO on synthetic benchmark functions and hyperparameter tuning tasks.
Anastasia Makarova, Ilnura Usmanova, Ilija Bogunovic, Andreas Krause 0001
NeurIPS1
2020 Mixed-Variable Bayesian Optimization
abstract
The optimization of expensive to evaluate, black-box, mixed-variable functions, i.e. functions that have continuous and discrete inputs, is a difficult and yet pervasive problem in science and engineering. In Bayesian optimization (BO), special cases of this problem that consider fully continuous or fully discrete domains have been widely studied. However, few methods exist for mixed-variable domains and none of them can handle discrete constraints that arise in many real-world applications. In this paper, we introduce MiVaBo, a novel BO algorithm for the efficient optimization of mixed-variable functions combining a linear surrogate model based on expressive feature representations with Thompson sampling. We propose an effective method to optimize its acquisition function, a challenging problem for mixed-variable domains, making MiVaBo the first BO method that can handle complex constraints over the discrete variables. Moreover, we provide the first convergence analysis of a mixed-variable BO algorithm. Finally, we show that MiVaBo is significantly more sample efficient than state-of-the-art mixed-variable BO algorithms on several hyperparameter tuning tasks, including the tuning of deep generative models.
Erik A. Daxberger, Anastasia Makarova, Matteo Turchetta, Andreas Krause 0001
IJCAI2