Ali Devran Kara

dblp:217/2136 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 81% Planning, search and constraint satisfaction · 10% Motion planning and robot control · 10%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
markov decision process
1.222023
Q-Learning for MDPs with General Spaces: Convergence and Near Optimality via Quantization under Weak Continuity · J. Mach. Learn. Res. 2023
Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › function approximation
linear function approximation
0.912025
Learning with Linear Function Approximations in Mean-Field Control · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning
mean field control
0.912025
Learning with Linear Function Approximations in Mean-Field Control · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Learning with Linear Function Approximations in Mean-Field Control · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.712023
Q-Learning for MDPs with General Spaces: Convergence and Near Optimality via Quantization under Weak Continuity · J. Mach. Learn. Res. 2023
Robotics › Motion planning and robot control › motion planning › motion planning under uncertainty
belief space planning
0.612022
Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes · J. Mach. Learn. Res. 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process
0.612022
Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
mean field games
0.312025
Learning with Linear Function Approximations in Mean-Field Control · J. Mach. Learn. Res. 2025

Methods — techniques the papers use, named apart from their topics

linear function approximation · 0.9independent learning · 0.9coordinated learning · 0.9weak continuity · 0.7quantization · 0.7POMDP reduction · 0.7finite window approximation · 0.6filter stability · 0.6belief discretization · 0.6
YearPublicationVenuePosition
2025 Learning with Linear Function Approximations in Mean-Field Control
abstract
The paper focuses on mean-field type multi-agent control problems with finite state and action spaces where the dynamics and cost structures are symmetric and homogeneous, and are affected by the distribution of the agents. A standard solution method for these problems is to consider the infinite population limit as an approximation and use symmetric solutions of the limit problem to achieve near optimality. The control policies, and in particular the dynamics, depend on the population distribution in the finite population setting, or the marginal distribution of the state variable of a representative agent for the infinite population setting. Hence, learning and planning for these control problems generally require estimating the reaction of the system to all possible state distributions of the agents. To overcome this issue, we consider linear function approximation for the control problem and provide coordinated and independent learning methods. We rigorously establish error upper bounds for the performance of learned solutions. The performance gap stems from (i) the mismatch due to estimating the true model with a linear one, and (ii) using the infinite population solution in the finite population problem as an approximate control. The provided upper bounds quantify the impact of these error sources on the overall performance.
Erhan Bayraktar, Ali Devran Kara
J. Mach. Learn. Res.2
2023 Q-Learning for MDPs with General Spaces: Convergence and Near Optimality via Quantization under Weak Continuity
abstract
Reinforcement learning algorithms often require finiteness of state and action spaces in Markov decision processes (MDPs) (also called controlled Markov chains) and various efforts have been made in the literature towards the applicability of such algorithms for continuous state and action spaces. In this paper, we show that under very mild regularity conditions (in particular, involving only weak continuity of the transition kernel of an MDP), Q-learning for standard Borel MDPs via quantization of states and actions (called Quantized Q-Learning) converges to a limit, and furthermore this limit satisfies an optimality equation which leads to near optimality with either explicit performance bounds or which are guaranteed to be asymptotically optimal. Our approach builds on (i) viewing quantization as a measurement kernel and thus a quantized MDP as a partially observed Markov decision process (POMDP), (ii) utilizing near optimality and convergence results of Q-learning for POMDPs, and (iii) finally, near-optimality of finite state model approximations for MDPs with weakly continuous kernels which we show to correspond to the fixed point of the constructed POMDP. Thus, our paper presents a very general convergence and approximation result for the applicability of Q-learning for continuous MDPs.
Ali Devran Kara, Naci Saldi, Serdar Yüksel
J. Mach. Learn. Res.1
2022 Near Optimality of Finite Memory Feedback Policies in Partially Observed Markov Decision Processes
abstract
In the theory of Partially Observed Markov Decision Processes (POMDPs), existence of optimal policies have in general been established via converting the original partially observed stochastic control problem to a fully observed one on the belief space, leading to a belief-MDP. However, computing an optimal policy for this fully observed model, and so for the original POMDP, using classical dynamic or linear programming methods is challenging even if the original system has finite state and action spaces, since the state space of the fully observed belief-MDP model is always uncountable. Furthermore, there exist very few rigorous value function approximation and optimal policy approximation results, as regularity conditions needed often require a tedious study involving the spaces of probability measures leading to properties such as Feller continuity. In this paper, we study a planning problem for POMDPs where the system dynamics and measurement channel model are assumed to be known. We construct an approximate belief model by discretizing the belief space using only finite window information variables. We then find optimal policies for the approximate model and we rigorously establish near optimality of the constructed finite window control policies in POMDPs under mild non-linear filter stability conditions and the assumption that the measurement and action sets are finite (and the state space is real vector valued). We also establish a rate of convergence result which relates the finite window memory size and the approximation error bound, where the rate of convergence is exponential under explicit and testable exponential filter stability conditions. While there exist many experimental results and few rigorous asymptotic convergence results, an explicit rate of convergence result is new in the literature, to our knowledge.
Ali Devran Kara, Serdar Yüksel
J. Mach. Learn. Res.1