Weronika Ormaniec

dblp:322/8772 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 36% Probabilistic and Bayesian machine learning · 33% Optimization for machine learning · 21%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.912025
Standardizing Structural Causal Models · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.912025
Standardizing Structural Causal Models · ICLR 2025
Machine learning › Optimization for machine learning
hessian analysis
0.912025
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis · ICLR 2025
Machine learning › Deep learning architectures and training
loss landscape
0.912025
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural causal model
0.912025
Standardizing Structural Causal Models · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis · ICLR 2025
Machine learning › Deep learning architectures and training › transformer
transformer optimization
0.912025
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis · ICLR 2025
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.812024
Transition Constrained Bayesian Optimization via Markov Decision Processes · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
constraint-based planning
0.812024
Transition Constrained Bayesian Optimization via Markov Decision Processes · NeurIPS 2024
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.312025
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis · ICLR 2025
Data integration and cleaning › data generation
synthetic data generation
0.312025
Standardizing Structural Causal Models · ICLR 2025

Methods — techniques the papers use, named apart from their topics

standardization · 1.7identifiability analysis · 1.7matrix calculus · 0.9hessian analysis · 0.9reinforcement learning · 0.8markov decision process · 0.8linearization · 0.8
YearPublicationVenuePosition
2025 What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
abstract
The Transformer architecture has inarguably revolutionized deep learning, overtaking classical architectures like multi-layer perceptions (MLPs) and convolutional neural networks (CNNs). At its core, the attention block differs in form and functionality from most other architectural components in deep learning—to the extent that, in comparison to MLPs/CNNs, Transformers are more often accompanied by adaptive optimizers, layer normalization, learning rate warmup, etc. The root causes behind these outward manifestations and the precise mechanisms that govern them remain poorly understood. In this work, we bridge this gap by providing a fundamental understanding of what distinguishes the Transformer from the other architectures—grounded in a theoretical comparison of the (loss) Hessian. Concretely, for a single self-attention layer, (a) we first entirely derive the Transformer’s Hessian and express it in matrix derivatives; (b) we then characterize it in terms of data, weight, and attention moment dependencies; and (c) while doing so further highlight the important structural differences to the Hessian of classical networks. Our results suggest that various common architectural and optimization choices in Transformers can be traced back to their highly non-linear dependencies on the data and weight matrices, which vary heterogeneously across parameters. Ultimately, our findings provide a deeper understanding of the Transformer’s unique optimization landscape and the challenges it poses.
Weronika Ormaniec, Felix Dangel, Sidak Pal Singh
ICLR1
2025 Standardizing Structural Causal Models
abstract
Synthetic datasets generated by structural causal models (SCMs) are commonly used for benchmarking causal structure learning algorithms. However, the variances and pairwise correlations in SCM data tend to increase along the causal ordering. Several popular algorithms exploit these artifacts, possibly leading to conclusions that do not generalize to real-world settings. Existing metrics like $\operatorname{Var}$-sortability and $\operatorname{R^2}$-sortability quantify these patterns, but they do not provide tools to remedy them. To address this, we propose internally-standardized structural causal models (iSCMs), a modification of SCMs that introduces a standardization operation at each variable during the generative process. By construction, iSCMs are not $\operatorname{Var}$-sortable. We also find empirical evidence that they are mostly not $\operatorname{R^2}$-sortable for commonly-used graph families. Moreover, contrary to the post-hoc standardization of data generated by standard SCMs, we prove that linear iSCMs are less identifiable from prior knowledge on the weights and do not collapse to deterministic relationships in large systems, which may make iSCMs a useful model in causal inference beyond the benchmarking problem studied here. Our code is publicly available at: https://github.com/werkaaa/iscm.
Weronika Ormaniec, Scott Sussex, Lars Lorch, Bernhard Schölkopf, Andreas Krause 0001
ICLR1
2024 Transition Constrained Bayesian Optimization via Markov Decision Processes
abstract
Bayesian optimization is a methodology to optimize black-box functions. Traditionally, it focuses on the setting where you can arbitrarily query the search space. However, many real-life problems do not offer this flexibility; in particular, the search space of the next query may depend on previous ones. Example challenges arise in the physical sciences in the form of local movement constraints, required monotonicity in certain variables, and transitions influencing the accuracy of measurements. Altogether, such *transition constraints* necessitate a form of planning. This work extends classical Bayesian optimization via the framework of Markov Decision Processes. We iteratively solve a tractable linearization of our utility function using reinforcement learning to obtain a policy that plans ahead for the entire horizon. This is a parallel to the optimization of an *acquisition function in policy space*. The resulting policy is potentially history-dependent and non-Markovian. We showcase applications in chemical reactor optimization, informative path planning, machine calibration, and other synthetic examples.
Jose Pablo Folch, Calvin Tsay, Robert M. Lee, Behrang Shafei, Weronika Ormaniec, Andreas Krause 0001, Mark van der Wilk, Ruth Misener, Mojmír Mutný
NeurIPS5