Donghwan Lee 0002

dblp:19/173-2 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-4962-8478ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 85% Motion planning and robot control · 8% Optimization for machine learning · 7%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
temporal difference learning
1.932025
A Primal-dual Perspective for Distributed TD-learning · IJCAI 2025
Backstepping Temporal Difference Learning · ICLR 2023
Target-Based Temporal-Difference Learning · ICML 2019
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
1.222024
Regularized Q-Learning · NeurIPS 2024
A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms · NeurIPS 2020
Machine learning › Reinforcement learning › function approximation
linear function approximation
0.812024
Regularized Q-Learning · NeurIPS 2024
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning convergence
0.812024
Regularized Q-Learning · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning with function approximation
0.812024
Regularized Q-Learning · NeurIPS 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
Regularized Q-Learning · NeurIPS 2024
Robotics › Motion planning and robot control › robot control › nonlinear control
backstepping control
0.712023
Backstepping Temporal Difference Learning · ICLR 2023
Machine learning › Optimization for machine learning
convergence analysis
0.522020
A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms · NeurIPS 2020
Target-Based Temporal-Difference Learning · ICML 2019
Mathematical optimization
dynamical systems
0.412020
A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms · NeurIPS 2020
Machine learning › Reinforcement learning
target network
0.412019
Target-Based Temporal-Difference Learning · ICML 2019
Mathematical optimization
distributed optimization
0.312025
A Primal-dual Perspective for Distributed TD-learning · IJCAI 2025
Machine learning › Reinforcement learning
function approximation
0.112020
A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

primal-dual ODE dynamics · 1.7distributed optimization · 1.7asymptotic stability · 0.9affine switching systems · 0.9ODE analysis · 0.9switching system models · 0.8regularization · 0.8temporal difference learning · 0.7backstepping · 0.7asymptotic convergence analysis · 0.4
YearPublicationVenuePosition
2026 Finite-time analysis of simultaneous double Q-learning
Hyunjun Na, Donghwan Lee 0002
Neurocomputing2
2025 A Primal-dual Perspective for Distributed TD-learning
abstract
The goal of this paper is to investigate distributed temporal difference (TD) learning for a networked multi-agent Markov decision process. The proposed approach is based on distributed optimization algorithms, which can be interpreted as primal-dual ordinary differential equation (ODE) dynamics subject to null-space constraints. Based on the exponential convergence behavior of the primal-dual ODE dynamics subject to null-space constraints, we examine the behavior of the final iterate in various distributed TD-learning scenarios, considering both constant and diminishing step-sizes and incorporating both i.i.d. and Markovian observation models. Unlike existing methods, the proposed algorithm does not require the assumption that the underlying communication network structure is characterized by a doubly stochastic matrix.
Han-Dong Lim, Donghwan Lee 0002
IJCAI2
2024 Regularized Q-Learning
abstract
Q-learning is widely used algorithm in reinforcement learning (RL) community. Under the lookup table setting, its convergence is well established. However, its behavior is known to be unstable with the linear function approximation case. This paper develops a new Q-learning algorithm, called RegQ, that converges when linear function approximation is used. We prove that simply adding an appropriate regularization term ensures convergence of the algorithm. Its stability is established using a recent analysis tool based on switching system models. Moreover, we experimentally show that RegQ converges in environments where Q-learning with linear function approximation has known to diverge. An error bound on the solution where the algorithm converges is also given.
Han-Dong Lim, Donghwan Lee 0002
NeurIPS2
2024 Harnessing membership function dynamics for stability analysis of T-S fuzzy systems
Donghwan Lee 0002, Do Wan Kim
Inf. Sci.1
2024 Relaxed Conditions for Parameterized Linear Matrix Inequality in the Form of Double Fuzzy Summation
abstract
The aim of this study is to investigate less conservative conditions for a parameterized linear matrix inequality (PLMI) expressed in the form of a double convex sum. This type of PLMI frequently appears in Takagi–Sugeno (T–S) fuzzy control system analysis and design problems. In this letter, we derive new less conservative linear matrix inequalities (LMIs) for the PLMI by employing the proposed sum relaxation method based on Young's inequality. The derived LMIs are proven to be less conservative than the existing conditions related to this topic in the literature. The proposed technique is applicable to various stability analysis and control design problems for T–S fuzzy systems, which are formulated as solving the PLMIs in the form of a double convex sum. Furthermore, examples are provided to illustrate the reduced conservatism of the derived LMIs.
Do Wan Kim, Donghwan Lee 0002
IEEE Trans. Fuzzy Syst.2
2023 Backstepping Temporal Difference Learning
Han-Dong Lim, Donghwan Lee 0002
ICLR2
2020 A Unified Switching System Perspective and Convergence Analysis of Q-Learning Algorithms
abstract
This paper develops a novel and unified framework to analyze the convergence of a large family of Q-learning algorithms from the switching system perspective. We show that the nonlinear ODE models associated with Q-learning and many of its variants can be naturally formulated as affine switching systems. Building on their asymptotic stability, we obtain a number of interesting results: (i) we provide a simple ODE analysis for the convergence of asynchronous Q-learning under relatively weak assumptions; (ii) we establish the first convergence analysis of the averaging Q-learning algorithm; and (iii) we derive a new sufficient condition for the convergence of Q-learning with linear function approximation.
Donghwan Lee 0002, Niao He
NeurIPS1
2019 Target-Based Temporal-Difference Learning
abstract
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side. In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms that maintain two separate learning parameters {–} the target variable and online variable. We propose three members in the family, the averaging TD, double TD, and periodic TD, where the target variable is updated through an averaging, symmetric, or periodic fashion, respectively, mirroring those techniques used in deep Q-learning practice. We establish asymptotic convergence analyses for both averaging TD and double TD and a finite sample analysis for periodic TD. In addition, we provide some simulation results showing potentially superior convergence of these target-based TD algorithms compared to the standard TD-learning. While this work focuses on linear function approximation and policy evaluation setting, we consider this as a meaningful step towards the theoretical understanding of deep Q-learning variants with target networks.
Donghwan Lee 0002, Niao He
ICML1
2017 Local Model Predictive Control for T-S Fuzzy Systems
abstract
In this paper, a new linear matrix inequality-based model predictive control (MPC) problem is studied for discrete-time nonlinear systems described as Takagi-Sugeno fuzzy systems. A recent local stability approach is applied to improve the performance of the proposed MPC scheme. At each time k , an optimal state-feedback gain that minimizes an objective function is obtained by solving a semidefinite programming problem. The local stability analysis, the estimation of the domain of attraction, and feasibility of the proposed MPC are proved. Examples are given to demonstrate the advantages of the suggested MPC over existing approaches.
Donghwan Lee 0002, Jianghai Hu
IEEE Trans. Cybern.1
2016 Local state-feedback stabilization of nonlinear fuzzy systems with measurement errors in the state
abstract
This paper deals with the local state-feedback control design problems for continuous-time Takagi-Sugeno (T-S) fuzzy systems with the measurement errors in the state. The state-feedback controller is designed in such a way that the closed-loop system is locally asymptotically stable. The stability is guaranteed to be robust against the measurement errors in the state. In addition, an inner estimation of the domain of attraction (DA) is obtained as a sublevel set of the obtained quadratic Lyapunov function. The design procedure is formulated as optimizations subject to linear matrix inequalities (LMIs).
Donghwan Lee 0002, Jianghai Hu
FUZZ-IEEE1
2016 Local stabilization of discrete-time T-S fuzzy systems with magnitude- and energy-bounded disturbances
Donghwan Lee 0002, Jianghai Hu
Inf. Sci.1