Yaniv Gurwicz

dblp:83/4274 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Probabilistic and Bayesian machine learning · 71% Deep learning architectures and training · 15% Trustworthy machine learning · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 44% Information retrieval · 44% Recommender systems · 13%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
3.052025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Causal Interpretation of Self-Attention in Pre-Trained Transformers · NeurIPS 2023
From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent Confounders · ICML 2023
Environmental and earth informatics › climate modeling
climate model emulation
1.522025
Causal Climate Emulation with Bayesian Filtering · NeurIPS 2025
ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning · NeurIPS 2023
Environmental and earth informatics
climate modeling
1.522025
Causal Climate Emulation with Bayesian Filtering · NeurIPS 2025
ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders
1.222023
From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent Confounders · ICML 2023
Iterative Causal Discovery in the Possible Presence of Latent Confounders and Selection Bias · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
causal world model
0.912025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention
0.912025
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning
0.722018
Constructing Deep Neural Networks by Bayesian Network Structure Learning · NeurIPS 2018
Bayesian Structure Learning by Recursive Bootstrap · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
constraint-based causal discovery
0.712023
From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent Confounders · ICML 2023
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.712023
Causal Interpretation of Self-Attention in Pre-Trained Transformers · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
time series causal discovery
0.712023
From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent Confounders · ICML 2023
Information retrieval › evaluation
benchmark dataset
0.712023
ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning · NeurIPS 2023
Data mining
dataset construction
0.712023
ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning · NeurIPS 2023
Machine learning › Trustworthy machine learning › dataset bias
selection bias
0.512021
Iterative Causal Discovery in the Possible Presence of Latent Confounders and Selection Bias · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412019
Modeling Uncertainty by Learning a Hierarchy of Deep Neural Connections · NeurIPS 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412019
Modeling Uncertainty by Learning a Hierarchy of Deep Neural Connections · NeurIPS 2019
Machine learning › Representation and self-supervised learning
causal representation learning
0.312025
Causal Climate Emulation with Bayesian Filtering · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment classification
0.212023
Causal Interpretation of Self-Attention in Pre-Trained Transformers · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.112019
Modeling Uncertainty by Learning a Hierarchy of Deep Neural Connections · NeurIPS 2019
Computer vision › Image recognition and object detection
image classification
0.112018
Constructing Deep Neural Networks by Bayesian Network Structure Learning · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

constraint-based causal discovery · 1.8causal representation learning · 1.7bayesian filtering · 1.7structural equation model · 1.3partial correlation · 1.3machine learning emulation · 1.3attention mechanism analysis · 0.9structural vector autoregressive process · 0.7statistical test · 0.7conditional independence test · 0.5bayesian neural network · 0.4
YearPublicationVenuePosition
2025 A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
abstract
Are generative pre-trained transformer (GPT) models, trained only to predict the next token, implicitly learning a world model from which sequences are generated one token at a time? We address this question by deriving a causal interpretation of the attention mechanism in GPT and presenting a causal world model that arises from this interpretation. Furthermore, we propose that GPT models, at inference time, can be utilized for zero-shot causal structure learning for input sequences, and introduce a corresponding confidence score. Empirical tests were conducted in controlled environments using the setups of the Othello and Chess strategy games. A GPT, pre-trained on real-world games played with the intention of winning, was tested on out-of-distribution synthetic data consisting of sequences of random legal moves. We find that the GPT model is likely to generate legal next moves for out-of-distribution sequences for which a causal structure is encoded in the attention mechanism with high confidence. In cases where it generates illegal moves, it also fails to capture a causal structure.
Raanan Y. Rohekar, Yaniv Gurwicz, Sungduk Yu, Estelle Aflalo, Vasudev Lal
ICML2
2025 Causal Climate Emulation with Bayesian Filtering
abstract
Traditional models of climate change use complex systems of coupled equations to simulate physical processes across the Earth system. These simulations are highly computationally expensive, limiting our predictions of climate change and analyses of its causes and effects. Machine learning has the potential to quickly emulate data from climate models, but current approaches are not able to incorporate physically-based causal relationships. Here, we develop an interpretable climate model emulator based on causal representation learning. We derive a novel approach including a Bayesian filter for stable long-term autoregressive emulation. We demonstrate that our emulator learns accurate climate dynamics, and we show the importance of each one of its components on a realistic synthetic dataset and data from two widely deployed climate models.
Sebastian Hickman, Ilija Trajkovic, Julia Kaltenborn, Francis Pelletier, Alex Archibald, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard
NeurIPS6
2023 From Temporal to Contemporaneous Iterative Causal Discovery in the Presence of Latent Confounders
abstract
We present a constraint-based algorithm for learning causal structures from observational time-series data, in the presence of latent confounders. We assume a discrete-time, stationary structural vector autoregressive process, with both temporal and contemporaneous causal relations. One may ask if temporal and contemporaneous relations should be treated differently. The presented algorithm gradually refines a causal graph by learning long-term temporal relations before short-term ones, where contemporaneous relations are learned last. This ordering of causal relations to be learnt leads to a reduction in the required number of statistical tests. We validate this reduction empirically and demonstrate that it leads to higher accuracy for synthetic data and more plausible causal graphs for real-world data compared to state-of-the-art algorithms.
Raanan Y. Rohekar, Shami Nisimov, Yaniv Gurwicz, Gal Novik
ICML3
2023 ClimateSet: A Large-Scale Climate Model Dataset for Machine Learning
abstract
Climate models have been key for assessing the impact of climate change and simulating future climate scenarios. The machine learning (ML) community has taken an increased interest in supporting climate scientists’ efforts on various tasks such as climate model emulation, downscaling, and prediction tasks. Many of those tasks have been addressed on datasets created with single climate models. However, both the climate science and ML communities have suggested that to address those tasks at scale, we need large, consistent, and ML-ready climate model datasets. Here, we introduce ClimateSet, a dataset containing the inputs and outputs of 36 climate models from the Input4MIPs and CMIP6 archives. In addition, we provide a modular dataset pipeline for retrieving and preprocessing additional climate models and scenarios. We showcase the potential of our dataset by using it as a benchmark for ML-based climate model emulation. We gain new insights about the performance and generalization capabilities of the different ML models by analyzing their performance across different climate models. Furthermore, the dataset can be used to train an ML emulator on several climate models instead of just one. Such a “super emulator” can quickly project new climate change scenarios, complementing existing scenarios already provided to policymakers. We believe ClimateSet will create the basis needed for the ML community to tackle climate-related tasks at scale.
Julia Kaltenborn, Charlotte E. E. Lange, Venkatesh Ramesh, Philippe Brouillard, Yaniv Gurwicz, Chandni Nagda, Jakob Runge, Peer Nowack, David Rolnick
NeurIPS5
2023 Causal Interpretation of Self-Attention in Pre-Trained Transformers
abstract
We propose a causal interpretation of self-attention in the Transformer neural network architecture. We interpret self-attention as a mechanism that estimates a structural equation model for a given input sequence of symbols (tokens). The structural equation model can be interpreted, in turn, as a causal structure over the input symbols under the specific context of the input sequence. Importantly, this interpretation remains valid in the presence of latent confounders. Following this interpretation, we estimate conditional independence relations between input symbols by calculating partial correlations between their corresponding representations in the deepest attention layer. This enables learning the causal structure over an input sequence using existing constraint-based algorithms. In this sense, existing pre-trained Transformers can be utilized for zero-shot causal-discovery. We demonstrate this method by providing causal explanations for the outcomes of Transformers in two tasks: sentiment classification (NLP) and recommendation.
Raanan Y. Rohekar, Yaniv Gurwicz, Shami Nisimov
NeurIPS2
2021 Iterative Causal Discovery in the Possible Presence of Latent Confounders and Selection Bias
abstract
We present a sound and complete algorithm, called iterative causal discovery (ICD), for recovering causal graphs in the presence of latent confounders and selection bias. ICD relies on the causal Markov and faithfulness assumptions and recovers the equivalence class of the underlying causal graph. It starts with a complete graph, and consists of a single iterative stage that gradually refines this graph by identifying conditional independence (CI) between connected nodes. Independence and causal relations entailed after any iteration are correct, rendering ICD anytime. Essentially, we tie the size of the CI conditioning set to its distance on the graph from the tested nodes, and increase this value in the successive iteration. Thus, each iteration refines a graph that was recovered by previous iterations having smaller conditioning sets---a higher statistical power---which contributes to stability. We demonstrate empirically that ICD requires significantly fewer CI tests and learns more accurate causal graphs compared to FCI, FCI+, and RFCI algorithms.
Raanan Y. Rohekar, Shami Nisimov, Yaniv Gurwicz, Gal Novik
NeurIPS3
2019 Modeling Uncertainty by Learning a Hierarchy of Deep Neural Connections
abstract
Modeling uncertainty in deep neural networks, despite recent important advances, is still an open problem. Bayesian neural networks are a powerful solution, where the prior over network weights is a design choice, often a normal distribution or other distribution encouraging sparsity. However, this prior is agnostic to the generative process of the input data, which might lead to unwarranted generalization for out-of-distribution tested data. We suggest the presence of a confounder for the relation between the input data and the discriminative function given the target label. We propose an approach for modeling this confounder by sharing neural connectivity patterns between the generative and discriminative networks. This approach leads to a new deep architecture, where networks are sampled from the posterior of local causal structures, and coupled into a compact hierarchy. We demonstrate that sampling networks from this hierarchy, proportionally to their posterior, is efficient and enables estimating various types of uncertainties. Empirical evaluations of our method demonstrate significant improvement compared to state-of-the-art calibration and out-of-distribution detection methods.
Raanan Y. Rohekar, Yaniv Gurwicz, Shami Nisimov, Gal Novik
NeurIPS2
2018 Bayesian Structure Learning by Recursive Bootstrap
abstract
We address the problem of Bayesian structure learning for domains with hundreds of variables by employing non-parametric bootstrap, recursively. We propose a method that covers both model averaging and model selection in the same framework. The proposed method deals with the main weakness of constraint-based learning---sensitivity to errors in the independence tests---by a novel way of combining bootstrap with constraint-based learning. Essentially, we provide an algorithm for learning a tree, in which each node represents a scored CPDAG for a subset of variables and the level of the node corresponds to the maximal order of conditional independencies that are encoded in the graph. As higher order independencies are tested in deeper recursive calls, they benefit from more bootstrap samples, and therefore are more resistant to the curse-of-dimensionality. Moreover, the re-use of stable low order independencies allows greater computational efficiency. We also provide an algorithm for sampling CPDAGs efficiently from their posterior given the learned tree. That is, not from the full posterior, but from a reduced space of CPDAGs encoded in the learned tree. We empirically demonstrate that the proposed algorithm scales well to hundreds of variables, and learns better MAP models and more reliable causal relationships between variables, than other state-of-the-art-methods.
Raanan Y. Rohekar, Yaniv Gurwicz, Shami Nisimov, Guy Koren, Gal Novik
NeurIPS2
2018 Constructing Deep Neural Networks by Bayesian Network Structure Learning
abstract
We introduce a principled approach for unsupervised structure learning of deep neural networks. We propose a new interpretation for depth and inter-layer connectivity where conditional independencies in the input distribution are encoded hierarchically in the network structure. Thus, the depth of the network is determined inherently. The proposed method casts the problem of neural network structure learning as a problem of Bayesian network structure learning. Then, instead of directly learning the discriminative structure, it learns a generative graph, constructs its stochastic inverse, and then constructs a discriminative graph. We prove that conditional-dependency relations among the latent variables in the generative graph are preserved in the class-conditional discriminative graph. We demonstrate on image classification benchmarks that the deepest layers (convolutional and dense) of common networks can be replaced by significantly smaller learned structures, while maintaining classification accuracy---state-of-the-art on tested benchmarks. Our structure learning algorithm requires a small computational cost and runs efficiently on a standard desktop CPU.
Raanan Y. Rohekar, Shami Nisimov, Yaniv Gurwicz, Guy Koren, Gal Novik
NeurIPS3
2011 Multiclass object classification for real-time video surveillance systems
Yaniv Gurwicz, Raanan Y. Rohekar, Boaz Lachover
Pattern Recognit. Lett.1
2005 Bayesian network classification using spline-approximated kernel density estimation
Yaniv Gurwicz, Boaz Lerner
Pattern Recognit. Lett.1