EDBT 2026 Demo / reviewers in the wild / expert
Çagatay Yildiz
dblp:202/7085 · also Cagatay Yildiz
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Representation and self-supervised learning · 22% Probabilistic and Bayesian machine learning · 18% Deep learning architectures and training · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 28 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
1.5 | 3 | 2023 | Modulated Neural ODEs · NeurIPS 2023 Latent Neural ODEs with Sparse Bayesian Multiple Shooting · ICLR 2023 Continuous-time Model-based Reinforcement Learning · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.9 | 2 | 2022 | Learning interacting dynamical systems with latent Gaussian process ODEs · NeurIPS 2022 Learning unknown ODE models with Gaussian processes · ICML 2018 |
Machine learning › Representation and self-supervised learning › causal representation learning › identifiability
identifiable representation learning |
0.9 | 1 | 2025 | Identifying latent state transitions in non-linear dynamical systems · ICLR 2025 |
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA |
0.9 | 1 | 2025 | Identifying latent state transitions in non-linear dynamical systems · ICLR 2025 |
Robotics › Motion planning and robot control › system identification
nonlinear system identification |
0.9 | 1 | 2025 | Identifying latent state transitions in non-linear dynamical systems · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
continual pre-training |
0.8 | 1 | 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model
large language model adaptation |
0.8 | 1 | 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? · EMNLP 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.7 | 1 | 2023 | Latent Neural ODEs with Sparse Bayesian Multiple Shooting · ICLR 2023 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.7 | 1 | 2023 | Modulated Neural ODEs · NeurIPS 2023 |
Machine learning › Graph learning
interacting dynamical systems |
0.6 | 1 | 2022 | Learning interacting dynamical systems with latent Gaussian process ODEs · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.5 | 2 | 2021 | ODE2VAE: Deep generative second order ODEs with Bayesian neural networks · NeurIPS 2019 Continuous-time Model-based Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
actor-critic methods |
0.5 | 1 | 2021 | Continuous-time Model-based Reinforcement Learning · ICML 2021 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.5 | 1 | 2021 | Continuous-time Model-based Reinforcement Learning · ICML 2021 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2019 | ODE2VAE: Deep generative second order ODEs with Bayesian neural networks · NeurIPS 2019 |
Machine learning › Optimization for machine learning › parallel optimization
asynchronous parallel optimization |
0.3 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Optimization for machine learning
second-order optimization |
0.3 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.3 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? · EMNLP 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.2 | 1 | 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? · EMNLP 2024 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.2 | 1 | 2023 | Modulated Neural ODEs · NeurIPS 2023 |
Machine learning › Graph learning
graph neural network |
0.2 | 1 | 2022 | Learning interacting dynamical systems with latent Gaussian process ODEs · NeurIPS 2022 |
Machine learning › Graph learning › graph neural network
message passing |
0.2 | 1 | 2022 | Learning interacting dynamical systems with latent Gaussian process ODEs · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
latent dynamics model |
0.1 | 1 | 2019 | ODE2VAE: Deep generative second order ODEs with Bayesian neural networks · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.1 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
stochastic gradient MCMC |
0.1 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Bioinformatics and computational biology › systems biology
dynamical system modeling |
0.1 | 1 | 2018 | Learning unknown ODE models with Gaussian processes · ICML 2018 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
asynchronous parallel algorithms |
0.1 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 0.9nonlinear ICA · 0.9token-level analysis · 0.8perplexity analysis · 0.8time-invariant modulator variables · 0.7sparse bayesian learning · 0.7neural ODE · 0.7multiple shooting · 0.7variational sparse gaussian process inference · 0.6ODE · 0.6non-parametric modeling · 0.3gaussian process · 0.3ergodic convergence analysis · 0.3asynchronous parallelization · 0.3L-BFGS · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying latent state transitions in non-linear dynamical systemsabstractThis work aims to recover the underlying states and their time evolution in a latent dynamical system from high-dimensional sensory measurements. Previous works on identifiable representation learning in dynamical systems focused on identifying the latent states, often with linear transition approximations. As such, they cannot identify nonlinear transition dynamics, and hence fail to reliably predict complex future behavior. Inspired by the advances in nonlinear ICA, we propose a state-space modeling framework in which we can identify not just the latent states but also the unknown transition function that maps the past states to the present. Our identifiability theory relies on two key assumptions: (i) sufficient variability in the latent noise, and (ii) the bijectivity of the augmented transition function. Drawing from this theory, we introduce a practical algorithm based on variational auto-encoders. We empirically demonstrate that it improves generalization and interpretability of target dynamical systems by (i) recovering latent state dynamics with high accuracy, (ii) correspondingly achieving high future prediction accuracy, and (iii) adapting fast to new environments. Additionally, for complex real-world dynamics, (iv) it produces state-of the-art future prediction results for long horizons, highlighting its usefulness for practical scenarios. Caglar Hizli, Çagatay Yildiz, Matthias Bethge, S. T. John, Pekka Marttinen |
ICLR | 2 |
| 2024 | Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?abstractIn the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions.Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora curated from the Internet with minimal human intervention, and (iii) trained in an online fashion.These stark contrasts prevent researchers from transferring lessons learned on model generalization and adaptation in deep learning contexts to LLMs.To this end, our short paper introduces empirical observations that aim to shed light on further training of already pretrained language models.Specifically, we demonstrate that training a model on a text domain could degrade its perplexity on the test portion of the same domain.We observe with our subsequent analysis that the performance degradation is positively correlated with the similarity between the additional and the original pretraining dataset of the LLM.Our further token-level perplexity observations reveals that the perplexity degradation is due to a handful of tokens that are not informative about the domain.We hope these findings will guide us in determining when to adapt a model vs when to rely on its foundational capabilities. Firat Öncel, Matthias Bethge, Beyza Ermis, Mirco Ravanelli, Cem Subakan, Çagatay Yildiz |
EMNLP | 6 |
| 2023 | Latent Neural ODEs with Sparse Bayesian Multiple Shooting
Valerii Iakovlev, Çagatay Yildiz, Markus Heinonen, Harri Lähdesmäki |
ICLR | 2 |
| 2023 | Modulated Neural ODEsabstractNeural ordinary differential equations (NODEs) have been proven useful for learning non-linear dynamics of arbitrary trajectories. However, current NODE methods capture variations across trajectories only via the initial state value or by auto-regressive encoder updates. In this work, we introduce Modulated Neural ODEs (MoNODEs), a novel framework that sets apart dynamics states from underlying static factors of variation and improves the existing NODE methods. In particular, we introduce *time-invariant modulator variables* that are learned from the data. We incorporate our proposed framework into four existing NODE variants. We test MoNODE on oscillating systems, videos and human walking trajectories, where each trajectory has trajectory-specific modulation. Our framework consistently improves the existing model ability to generalize to new dynamic parameterizations and to perform far-horizon forecasting. In addition, we verify that the proposed modulator variables are informative of the true unknown factors of variation as measured by $R^2$ scores. Ilze Amanda Auzina, Çagatay Yildiz, Sara Magliacane, Matthias Bethge, Efstratios Gavves |
NeurIPS | 2 |
| 2022 | Learning interacting dynamical systems with latent Gaussian process ODEsabstractWe study uncertainty-aware modeling of continuous-time dynamics of interacting objects. We introduce a new model that decomposes independent dynamics of single objects accurately from their interactions. By employing latent Gaussian process ordinary differential equations, our model infers both independent dynamics and their interactions with reliable uncertainty estimates. In our formulation, each object is represented as a graph node and interactions are modeled by accumulating the messages coming from neighboring objects. We show that efficient inference of such a complex network of variables is possible with modern variational sparse Gaussian process inference techniques. We empirically demonstrate that our model improves the reliability of long-term predictions over neural network based alternatives and it successfully handles missing dynamic or static information. Furthermore, we observe that only our model can successfully encapsulate independent dynamics and interaction information in distinct functions and show the benefit from this disentanglement in extrapolation scenarios. Çagatay Yildiz, Melih Kandemir, Barbara Rakitsch |
NeurIPS | 1 |
| 2022 | Variational multiple shooting for Bayesian ODEs with Gaussian processesabstractRecent machine learning advances have proposed black-box estimation of \textit{unknown continuous-time system dynamics} directly from data. However, earlier works are based on approximative solutions or point estimates. We propose a novel Bayesian nonparametric model that uses Gaussian processes to infer posteriors of unknown ODE systems directly from data. We derive sparse variational inference with decoupled functional sampling to represent vector field posteriors. We also introduce a probabilistic shooting augmentation to enable efficient inference from arbitrarily long trajectories. The method demonstrates the benefit of computing vector field posteriors, with predictive uncertainty scores outperforming alternative methods on multiple ODE learning tasks. Pashupati Hegde, Çagatay Yildiz, Harri Lähdesmäki, Samuel Kaski, Markus Heinonen |
UAI | 2 |
| 2021 | Continuous-time Model-based Reinforcement LearningabstractModel-based reinforcement learning (MBRL) approaches rely on discrete-time state transition models whereas physical systems and the vast majority of control tasks operate in continuous-time. To avoid time-discretization approximation of the underlying process, we propose a continuous-time MBRL framework based on a novel actor-critic method. Our approach also infers the unknown state evolution differentials with Bayesian neural ordinary differential equations (ODE) to account for epistemic uncertainty. We implement and test our method on a new ODE-RL suite that explicitly solves continuous-time control systems. Our experiments illustrate that the model is robust against irregular and noisy data, and can solve classic control problems in a sample-efficient manner. Çagatay Yildiz, Markus Heinonen, Harri Lähdesmäki |
ICML | 1 |
| 2019 | ODE2VAE: Deep generative second order ODEs with Bayesian neural networksabstractWe present Ordinary Differential Equation Variational Auto-Encoder (ODE2VAE), a latent second order ODE model for high-dimensional sequential data. Leveraging the advances in deep generative models, ODE2VAE can simultaneously learn the embedding of high dimensional trajectories and infer arbitrarily complex continuous-time latent dynamics. Our model explicitly decomposes the latent space into momentum and position components and solves a second order ODE system, which is in contrast to recurrent neural network (RNN) based time series models and recently proposed black-box ODE techniques. In order to account for uncertainty, we propose probabilistic latent ODE dynamics parameterized by deep Bayesian neural networks. We demonstrate our approach on motion capture, image rotation, and bouncing balls datasets. We achieve state-of-the-art performance in long term motion prediction and imputation tasks. Çagatay Yildiz, Markus Heinonen, Harri Lähdesmäki |
NeurIPS | 1 |
| 2018 | Learning unknown ODE models with Gaussian processesabstractIn conventional ODE modelling coefficients of an equation driving the system state forward in time are estimated. However, for many complex systems it is practically impossible to determine the equations or interactions governing the underlying dynamics. In these settings, parametric ODE model cannot be formulated. Here, we overcome this issue by introducing a novel paradigm of nonparametric ODE modelling that can learn the underlying dynamics of arbitrary continuous-time systems without prior knowledge. We propose to learn non-linear, unknown differential functions from state observations using Gaussian process vector fields within the exact ODE formalism. We demonstrate the model’s capabilities to infer dynamics from sparse data and to simulate the system forward into future. Markus Heinonen, Çagatay Yildiz, Henrik Mannerström, Jukka Intosalmi, Harri Lähdesmäki |
ICML | 2 |
| 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex OptimizationabstractRecent studies have illustrated that stochastic gradient Markov Chain Monte Carlo techniques have a strong potential in non-convex optimization, where local and global convergence guarantees can be shown under certain conditions. By building up on this recent theory, in this study, we develop an asynchronous-parallel stochastic L-BFGS algorithm for non-convex optimization. The proposed algorithm is suitable for both distributed and shared-memory settings. We provide formal theoretical analysis and show that the proposed method achieves an ergodic convergence rate of ${\cal O}(1/\sqrt{N})$ ($N$ being the total number of iterations) and it can achieve a linear speedup under certain conditions. We perform several experiments on both synthetic and real datasets. The results support our theory and show that the proposed algorithm provides a significant speedup over the recently proposed synchronous distributed L-BFGS algorithm. Umut Simsekli, Çagatay Yildiz, Thanh Huy Nguyen 0001, A. Taylan Cemgil, Gaël Richard |
ICML | 2 |