EDBT 2026 Demo / reviewers in the wild / expert
Baturay Saglam
dblp:302/3987
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-8324-5980ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
actor-critic methods |
0.8 | 1 | 2024 | Actor Prioritized Experience Replay (Abstract Reprint) · AAAI 2024 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.8 | 1 | 2024 | Actor Prioritized Experience Replay (Abstract Reprint) · AAAI 2024 |
Machine learning › Reinforcement learning › actor-critic methods
off-policy actor-critic |
0.8 | 1 | 2024 | Actor Prioritized Experience Replay (Abstract Reprint) · AAAI 2024 |
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
prioritized experience replay |
0.8 | 1 | 2024 | Actor Prioritized Experience Replay (Abstract Reprint) · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
temporal difference learning · 0.8prioritized experience replay · 0.8policy gradient · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Risk-Averse Constrained Reinforcement Learning with Optimized Certainty EquivalentsabstractConstrained optimization provides a common framework for dealing with conflicting objectives in reinforcement learning (RL). In most of these settings, the objectives (and constraints) are expressed though the expected accumulated reward. However, this formulation neglects risky or even possibly catastrophic events at the tails of the reward distribution, and is often insufficient for high-stakes applications in which the risk involved in outliers is critical. In this work, we propose a framework for risk-aware constrained RL, which exhibits per-stage robustness properties jointly in reward values and time using optimized certainty equivalents (OCEs). Our framework ensures an exact equivalent to the original constrained problem within a parameterized strong Lagrangian duality framework under appropriate constraint qualifications, and yields a simple algorithmic recipe which can be wrapped around standard RL solvers, such as PPO. Lastly, we establish the convergence of the proposed algorithm and verify the risk-aware properties of our approach through several numerical experiments. Jane Lee, Baturay Saglam, Spyridon Pougkakiotis, Amin Karbasi, Dionysios S. Kalogerias |
NeurIPS | 2 |
| 2024 | Actor Prioritized Experience Replay (Abstract Reprint)abstractA widely-studied deep reinforcement learning (RL) technique known as Prioritized Experience Replay (PER) allows agents to learn from transitions sampled with non-uniform probability proportional to their temporal-difference (TD) error. Although it has been shown that PER is one of the most crucial components for the overall performance of deep RL methods in discrete action domains, many empirical studies indicate that it considerably underperforms off-policy actor-critic algorithms. We theoretically show that actor networks cannot be effectively trained with transitions that have large TD errors. As a result, the approximate policy gradient computed under the Q-network diverges from the actual gradient computed under the optimal Q-function. Motivated by this, we introduce a novel experience replay sampling framework for actor-critic methods, which also regards issues with stability and recent findings behind the poor empirical performance of PER. The introduced algorithm suggests a new branch of improvements to PER and schedules effective and efficient training for both actor and critic networks. An extensive set of experiments verifies our theoretical findings, showing that our method outperforms competing approaches and achieves state-of-the-art results over the standard off-policy actor-critic algorithms. Baturay Saglam, Furkan B. Mutlu, Dogan Can Çiçek, Suleyman Serdar Kozat |
AAAI | 1 |
| 2024 | Parameter-Free Reduction of the Estimation Bias in Deep Reinforcement Learning for Deterministic Policy GradientsabstractAbstract Approximation of the value functions in value-based deep reinforcement learning induces overestimation bias, resulting in suboptimal policies. We show that when the reinforcement signals received by the agents have a high variance, deep actor-critic approaches that overcome the overestimation bias lead to a substantial underestimation bias. We first address the detrimental issues in the existing approaches that aim to overcome such underestimation error. Then, through extensive statistical analysis, we introduce a novel, parameter-free Deep Q-learning variant to reduce this underestimation bias in deterministic policy gradients. By sampling the weights of a linear combination of two approximate critics from a highly shrunk estimation bias interval, our Q-value update rule is not affected by the variance of the rewards received by the agents throughout learning. We test the performance of the introduced improvement on a set of MuJoCo and Box2D continuous control tasks and demonstrate that it outperforms the existing approaches and improves the baseline actor-critic algorithm in most of the environments tested. Baturay Saglam, Furkan B. Mutlu, Dogan Can Çiçek, Suleyman Serdar Kozat |
Neural Process. Lett. | 1 |
| 2023 | Actor Prioritized Experience ReplayabstractA widely-studied deep reinforcement learning (RL) technique known as Prioritized Experience Replay (PER) allows agents to learn from transitions sampled with non-uniform probability proportional to their temporal-difference (TD) error. Although it has been shown that PER is one of the most crucial components for the overall performance of deep RL methods in discrete action domains, many empirical studies indicate that it considerably underperforms off-policy actor-critic algorithms. We theoretically show that actor networks cannot be effectively trained with transitions that have large TD errors. As a result, the approximate policy gradient computed under the Q-network diverges from the actual gradient computed under the optimal Q-function. Motivated by this, we introduce a novel experience replay sampling framework for actor-critic methods, which also regards issues with stability and recent findings behind the poor empirical performance of PER. The introduced algorithm suggests a new branch of improvements to PER and schedules effective and efficient training for both actor and critic networks. An extensive set of experiments verifies our theoretical findings, showing that our method outperforms competing approaches and achieves state-of-the-art results over the standard off-policy actor-critic algorithms. Baturay Saglam, Furkan B. Mutlu, Dogan Can Çiçek, Suleyman Serdar Kozat |
J. Artif. Intell. Res. | 1 |
| 2023 | Deep intrinsically motivated exploration in continuous control
Baturay Saglam, Suleyman Serdar Kozat |
Mach. Learn. | 1 |
| 2021 | AWD3: Dynamic Reduction of the Estimation BiasabstractValue-based deep Reinforcement Learning (RL) algorithms suffer from the estimation bias primarily caused by function approximation and temporal difference (TD) learning. This problem induces faulty state-action value estimates and therefore harms the performance and robustness of the learning algorithms. Although several techniques were proposed to tackle, learning algorithms still suffer from this bias. Here, we introduce a technique that eliminates the estimation bias in off-policy continuous control algorithms using the experience replay mechanism. We adaptively learn the weighting hyper-parameter beta in the Weighted Twin Delayed Deep Deterministic Policy Gradient algorithm. Our method is named Adaptive-WD3 (AWD3). We show through continuous control environments of OpenAI gym that our algorithm matches or outperforms the state-of-the-art off-policy policy gradient learning algorithms. Dogan Can Çiçek, Enes Duran, Baturay Saglam, Kagan Kaya, Furkan B. Mutlu, Suleyman Serdar Kozat |
ICTAI | 3 |
| 2021 | Off-Policy Correction for Deep Deterministic Policy Gradient Algorithms via Batch Prioritized Experience ReplayabstractThe experience replay mechanism allows agents to use the experiences multiple times. In prior works, the sampling probability of the transitions was adjusted according to their importance. Reassigning sampling probabilities for every transition in the replay buffer after each iteration is highly inefficient. Therefore, experience replay prioritization algorithms recalculate the significance of a transition when the corresponding transition is sampled to gain computational efficiency. However, the importance level of the transitions changes dynamically as the policy and the value function of the agent are updated. In addition, experience replay stores the transitions are generated by the previous policies of the agent that may significantly deviate from the most recent policy of the agent. Higher deviation from the most recent policy of the agent leads to more off-policy updates, which is detrimental for the agent. In this paper, we develop a novel algorithm, Batch Prioritizing Experience Replay via KL Divergence (KLPER), which prioritizes batch of transitions rather than directly prioritizing each transition. Moreover, to reduce the off-policyness of the updates, our algorithm selects one batch among a certain number of batches and forces the agent to learn through the batch that is most likely generated by the most recent policy of the agent. We combine our algorithm with Deep Deterministic Policy Gradient and Twin Delayed Deep Deterministic Policy Gradient and evaluate it on various continuous control tasks. KLPER provides promising improvements for deep deterministic continuous control algorithms in terms of sample efficiency, final performance, and stability of the policy during the training. Dogan Can Çiçek, Enes Duran, Baturay Saglam, Furkan B. Mutlu, Suleyman Serdar Kozat |
ICTAI | 3 |
| 2021 | Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic MethodsabstractIn value-based deep reinforcement learning methods, approximation of value functions induces overestimation bias and leads to suboptimal policies. We show that in deep actor-critic methods that aim to overcome the overestimation bias, if the reinforcement signals received by the agent have a high variance, a significant underestimation bias arises. To minimize the underestimation, we introduce a parameter-free, novel deep Q-learning variant. Our Q-value update rule combines the notions behind Clipped Double Q-learning and Maxmin Q-learning by computing the critic objective through the nested combination of maximum and minimum operators to bound the approximate value estimates. We evaluate our modification on the suite of several OpenAI Gym continuous control tasks, improving the state-of-the-art in every environment tested. Baturay Saglam, Enes Duran, Dogan Can Çiçek, Furkan B. Mutlu, Suleyman Serdar Kozat |
ICTAI | 1 |